Here is a quick question for you. How long does it take an attacker to move from initial compromise to full lateral movement in your network?
In 2019, it was about 84 minutes. By 2023, it had dropped to just 62 minutes. Today, we are seeing breakout times of less than an hour! These figures come from CrowdStrike’s annual Global Threat Reports, which track adversary speed year over year.
Minus 60 minutes till game over.
Now here is the real question. How long does it take your Cyber Threat Intelligence (CTI) team to identify, analyze, and respond to a sophisticated threat? If you are like most organizations, you measure it in hours or even days.
You are not losing because adversaries are smarter. You are losing because they are faster. They are faster because they have figured out something that many CTI teams are still missing!
This guide will teach you how to leverage data engineering for CTI. You will learn what modern data-driven threat intelligence looks like, how to build the right infrastructure, and practical steps to get started without becoming a data scientist. Let’s jump in!
The Problem: Why Traditional CTI Cannot Keep Up
Let me paint a picture of what most CTI teams look like today. You have a team of smart analysts. People who understand Tactics, Techniques, and Procedures (TTPs).
- They can read through threat reports.
- They know the MITRE ATT&CK framework inside and out.
- They are good at what they do.
But here is what their day actually looks like. Every single day, they are being hit with a tsunami of data. Endpoint logs. Network flow data. Authentication records. Cloud telemetry. It is coming from everywhere, and it never stops.
What are they doing with all this data?
They are manually pivoting through Security Information and Event Management (SIEM) queries. They are copying and pasting IP addresses from threat intel feeds. They are trying to correlate three different data sources using spreadsheets. They are reading through hundreds of pages of threat reports, trying to find that one paragraph that actually matters to their business.
Does any of this sound familiar?

Modern adversaries are using automated ransomware frameworks. They are deploying polymorphic malware that changes its signature every time it runs. They are using Domain Generation Algorithms (DGAs) to create thousands of new command-and-control domains on the fly. Increasingly, they are using AI to write better phishing emails, better exploit code, and adapt faster than ever before.
You cannot fight automation with manual processes. You cannot fight machine-speed attacks with human-speed analysis.
What It Takes: Building a Modern Data-Driven CTI Capability
So let’s talk about what it actually takes to build a modern, data-driven CTI capability.
- Faster Detection
Move from hours of analysis to millisecond-level alerting on emerging threats. - Better Pattern Recognition
Identify sophisticated attacks that slip through traditional signature-based defenses. - Scalable Intelligence
Process terabytes of security telemetry without burning out your team. - Proactive Hunting
Shift from reactive blocking to proactive threat hunting using behavioral analysis.
The foundation is data engineering. It may sound boring, but it is absolutely critical: getting your data house in order.
The Speed Problem: Why You Need Real-Time Streaming Data Pipelines
Think about it this way. Before you can predict where an attacker is going. Before you can train any machine learning model. Before you can do any kind of advanced analysis. You need to get your data house in order.
For most organizations, that house is currently on fire!
Traditional security setups use what is called batch processing. Your logs are collected and processed every hour, every half hour, or even every 15 minutes. But remember those breakout times. Less than 60-minutes. If you are processing logs in hourly batches, you have already lost. The attacker has moved laterally, established persistence, and is halfway to achieving their objective before your first alert even fires.
This is why modern CTI teams (and security teams in general) are moving to streaming architectures. Tools like Apache Kafka act as the central nervous system for your security data. Instead of collecting logs and waiting, you are processing events as they happen. Millisecond latency.
Plus, that same event stream can feed your real-time alerting, your long-term data lake for historical analysis, and your machine learning models for pattern detection. All simultaneously.
Apache Kafka is not the only option here. Apache Flink provides millisecond-level latency for stateful detection logic, while Apache Spark Structured Streaming offers a “micro-batch” architecture with slightly higher latency but a unified API for both batch and streaming workloads. The choice depends on your specific latency requirements and existing infrastructure.
Action: Identify one critical log source in your environment (such as authentication logs or DNS queries) and document its current ingestion latency. How long does it take for an event to be searchable in your SIEM after it occurs? Write that number down. That is your baseline to improve.
The Relationship Problem: Why You Need Graph Databases
But speed is only half the battle. The other half is understanding relationships.
Here is an example. An IP address connects to your network. Your standard question might be: “Is this IP malicious?” But that is the wrong question.
The right questions are:
- What other IPs are connected to the same infrastructure?
- Which domains resolve to similar IP ranges?
- Do they share SSL certificates?
- What malware families have historically used this infrastructure?
- What threat actors prefer this hosting provider?
- What are the timing patterns related to this infrastructure?
The key is no longer looking at an isolated indicator. Instead, look at the entire ecosystem of relationships.
If you want to see how graph-based thinking applies to threat profiling, check out my guide on the Diamond Model. It provides a structured way to map relationships between adversaries, capabilities, infrastructure, and victims.
Unfortunately, this is where traditional relational databases often fall apart. They are designed for structured queries like “give me all records where X equals Y.” But threat intelligence is not structured. It is a web of connections. You need a database that thinks in graphs, not tables.
Graph databases like Neo4j let you ask questions like: “Show me all domains that share infrastructure with known malicious IPs, that were registered in the last 72 hours, by the same registrant used in previous APT campaigns.” Companies like Intuit have used Neo4j to map over 500,000 network endpoints and their interdependencies, enabling their security team to visualize complex relationships in milliseconds rather than hours.

Try writing an SQL query for that. Good luck!
Action: Take one IOC from a recent investigation and manually map its relationships on paper or a whiteboard. What domains does it connect to? What other IPs share that infrastructure? What threat actors have used similar setups? If this takes you more than an hour, you have just identified why you need graph-based tooling.
Machine Learning for Threat Intelligence: Turning Data into Detection
Okay, so now you have got your data pipeline humming. Events are streaming in real time. You have a graph database that understands relationships. Your infrastructure can actually handle the scale of modern security telemetry.
Now comes the intelligence part. Turning raw data into actionable intelligence through data analysis for CTI.

Finding Needles in Haystacks: Anomaly Detection
Let’s start with a classic CTI problem. How do you find the sophisticated, low-and-slow attacks that do not trigger any of your signature-based rules?
This is where unsupervised learning can shine. Techniques like Isolation Forest and Autoencoders do not need you to tell them what “bad” looks like. Instead, they learn what “normal” looks like and then flag anything that deviates from that baseline.
Imagine you have a user who normally logs in from London between 9 AM and 5 PM every day. They access five specific servers during their daily work and download about 5 megabytes of data each day.
Then one Tuesday, that account logs in from Singapore at 3 AM, accesses 30 different servers, and downloads 10 gigabytes of data.
Every individual action might be technically allowed by your security controls. But the pattern… that is not normal. An unsupervised model would flag it immediately. Not because it matches a known attack signature, but because it violates the behavioral baseline established for this user.
This is how you catch the stuff that slips through traditional defenses.
Key Machine Learning Algorithms for Threat Intelligence
Isolation Forest
Works by randomly partitioning datasets. Anomalies are isolated faster (shorter path lengths) than normal points. Excellent for high-dimensional network data. Learn more in this IEEE research on network anomaly detection.
Autoencoders
Neural networks are trained to compress and reconstruct input. When fed anomalous data they have not seen before, they produce high reconstruction error, which serves as the anomaly score.
LSTM Networks
A type of Recurrent Neural Network designed for sequence data. Treats domain names as character sequences to identify DGA-generated domains.
XGBoost
Gradient boosting on decision trees. High accuracy on tabular data for malware classification tasks.
Detecting Domain Generation Algorithms
Here is another one for you. What do you do about Domain Generation Algorithms? How do you detect them?
Sophisticated malware uses these algorithms to generate thousands of seemingly random domain names every day. The malware tries to connect to these domains, and the attacker only needs to register one or two of them to maintain command and control.
Classic indicator sharing does not work here. By the time you have added one malicious domain to your block list, the malware has already moved to the next 500.
But machine learning models, particularly recurrent neural networks, can spot DGA domains by analyzing the character patterns. You see, legitimate domains tend to have pronounceable structures and meaningful words, whereas DGA domains look like something has just smashed the keyboard to pieces.
A trained model can identify DGA domains in real time with high accuracy, even if it has never seen that specific domain before. Research implementations have achieved accuracy rates exceeding 98%, though real-world performance varies depending on the specific DGA families encountered and how well the model is tuned to your environment.
You are no longer blocking known bad. You are blocking entire classes of malicious behavior.
This approach aligns perfectly with the Pyramid of Pain. Instead of chasing fleeting indicators like IP addresses and domain names at the bottom of the pyramid, machine learning allows you to detect TTPs and behavioral patterns near the top, which are much harder for adversaries to change
The LLM Factor: Both Friend and Foe
Finally, we need to talk about the elephant in the room. Generative AI and Large Language Models (LLMs).
Here is where it gets interesting. Because LLMs are both a defensive and adversarial threat.
On the defensive side, LLMs are already being used to:
- Automatically summarize 100-page threat reports and turn them into actionable intelligence
- Convert natural language questions into complex database queries
- Generate detection rules from plain English descriptions
- Correlate information across dozens of threat intelligence sources
That is hours of analyst work compressed into seconds.
But here is the flip side. Adversaries have access to the same technology. We are seeing threat actors use LLMs to craft more convincing phishing emails, write better exploit code, and even conduct social engineering at scale. Google’s Threat Intelligence Group has documented this shift in its AI Threat Tracker reports.
The implication is clear. If you are not using AI to defend, you are at a structural disadvantage against adversaries who are using it to attack.
Action: Pick one repetitive task your team does weekly, such as summarizing a threat report or writing a SIEM query. Spend 30 minutes testing whether an LLM can assist with it. Document the time saved and the quality of output. This gives you data to justify further investment.
Getting Started: Practical Steps for Your CTI Team
Right now, you might be thinking, “Adam, this all sounds great, but I am not a data scientist, and I do not know how to code in Python or build streaming pipelines. My team barely has time to keep up with daily alerts, let alone implement machine learning models.”
I get it. However, you do not need to become a data scientist to benefit from data science. What you need is data science thinking. You need to understand what is possible, where the bottlenecks are, and how to ask the right questions.
Which of the challenges I have described resonates most with your team? Keep that specific pain point in mind as we walk through the solution.
Step One: Start with Your Pain Points
Do not try to boil the ocean. Start with the specific problems that are burning the most analyst hours in your team.
- Are you spending hours manually correlating IOCs across platforms?
- Are you constantly fighting false positives in your alerts?
- Are you missing sophisticated attacks because they do not match any of your known signatures?
- Are you unable to keep up with the volume of threat intelligence reports?
Just pick one. Just one. Then ask: could this be automated or augmented with the right data infrastructure or machine learning model?
Prove value on that one single use case. Show your leadership team that implementing streaming ingestion reduced detection time from two hours to just three minutes. Show them that the anomaly detection model caught three intrusions that signature-based rules might have missed.
Then, once you have that concrete piece of evidence, scale from there.
Action: Create a simple “Time Audit” spreadsheet. For one week, have each analyst log how they spend their time in 30-minute blocks. At the end of the week, identify the top three tasks consuming the most hours. Those are your automation candidates.
Step Two: Build and Buy Strategically
You do not have to build everything from scratch. There are commercial platforms now that provide streaming ingestion, graph databases, and pre-trained machine learning models out of the box.
Microsoft Sentinel, Splunk, and Elastic all have built-in machine learning capabilities. You can start using their anomaly detection without even writing a single line of code.
But, and this is important, you need to understand the fundamentals so that you can evaluate these tools effectively, configure them properly, and interpret their results intelligently.
A vendor’s “AI-powered threat detection” means nothing if you do not understand the model they are using, the data it was trained on, and its false-positive rate. Before buying any “AI-powered” security tool, ask three questions:
- What type of model is it using (supervised, unsupervised, LLM-based)?
- What data was it trained on, and how similar is it to my environment?
- What is its documented false positive rate? Can you tune this?
If the vendor cannot answer these clearly, proceed with caution.
Action: Review the ML or AI features in one security tool you already own. Find the documentation on how it works. Can you answer the three questions above? If not, schedule a call with your vendor to get clarity before your next renewal.
Step Three: Invest in Hybrid Skills
The most valuable people in CTI moving forward will be those who can bridge two worlds.
First, they have a deep understanding of adversary tradecraft and intelligence operations.
Second, they have practical data analysis skills: querying databases, writing basic Python scripts, and understanding statistical concepts.
You do not need a PhD in machine learning. But learning the basics of SQL, Python, and some data analysis skills… that is achievable, and it will transform your effectiveness as a threat intelligence analyst.
There are resources specifically designed for security practitioners looking to improve their data analysis skills. Jupyter notebooks for threat hunting, Python libraries for working with threat intelligence formats, and open source frameworks for building detections.
The tools exist. The knowledge is accessible. What is required is the commitment to evolve alongside the threat landscape.
The table below maps the key skill categories to specific competencies and where to start learning them. Use it as a roadmap for your professional development or when building training plans for your team.
| Skill Category | Key Skills | Where to Start |
|---|---|---|
| Data Querying | SQL, KQL, Splunk SPL | Your SIEM’s official documentation and training. The KC7 project. |
| Scripting | Python basics, API interaction | Python threat hunting tools series or boot.dev for the fundamentals. |
| Statistical Analysis | Baselining, Z-scores, anomaly concepts | Khan Academy statistics or DataCamp. |
| ML Literacy | Understanding model types, evaluating outputs | Vendor documentation for tools you own, Udemy, Coursera, or Codecademy. |
Action: Commit to learning one new skill from the table above this quarter. Block 30 minutes per week on your calendar for deliberate practice. Consistency beats intensity.
The era of the spreadsheet-wielding analyst is ending. The era of the data-driven CTI professional has begun!
Frequently Asked Questions
What is Data Engineering for CTI?
Data engineering for CTI involves building the infrastructure to collect, process, and store security telemetry at scale. This includes streaming data pipelines (such as Apache Kafka and Flink), data lakes for long-term storage, and standardized formats (STIX and TAXII) for interoperability. Good data engineering ensures your threat intelligence team can access the right data, at the right speed, in the right format to detect and respond to threats.
Do I Need to Know How to Code to Use Data Analysis for CTI?
No, you do not need to be a programmer to benefit from data analysis for CTI. Many commercial platforms, such as Microsoft Sentinel, Splunk, and Elastic, offer built-in machine learning capabilities that require no coding. However, learning basic skills like SQL querying and Python scripting will significantly expand what you can accomplish. The goal is to develop “data science thinking” so you can ask the right questions, even if someone else builds the technical solution.
What Is the Difference Between Batch Processing and Streaming in CTI?
Batch processing collects and processes security logs at fixed intervals (e.g., every hour or 15 minutes), introducing delays between when an event occurs and when it is analyzed. Streaming processing analyzes events in near real time as they happen, with millisecond-level latency. Given that modern attacker breakout times are less than 60-minutes, streaming architectures are essential for detecting threats before they achieve their objectives.
How Do Graph Databases Help With Threat Intelligence?
Graph databases store data as nodes (entities) and relationships (edges), making them ideal for modeling the interconnected nature of threat infrastructure.
Unlike traditional relational databases that require complex queries to traverse connections, graph databases can quickly answer questions like “What domains share infrastructure with this malicious IP and were registered by the same entity used in previous campaigns?” This enables analysts to move from collecting isolated indicators to understanding entire adversarial ecosystems.
What Machine Learning Techniques Are Most Useful for CTI?
The most useful machine learning techniques for CTI include Isolation Forests and Autoencoders for anomaly detection (finding unknown threats based on behavioral deviation), LSTM neural networks for detecting Domain Generation Algorithms, and classification models like XGBoost for malware categorization.
Additionally, Explainable AI techniques like SHAP help analysts understand why a model flagged something as suspicious, which is essential for building trust in automated detection.




