As a cyber threat intelligence (CTI) analyst, do you ever feel like you’re trying to drink from a firehose? In-depth research papers, breaking news articles, cryptic social media chatter, and dozens of raw, unvetted threat feeds. The key to overcoming this challenge lies in understanding the two fundamental forms of data we deal with every day: structured and unstructured threat intelligence.
This guide will demystify these concepts and provide clarity in the chaos. We’ll explore the narrative power of unstructured intelligence, the automated speed of structured threat intelligence, and why you need both to build a complete and effective defense.
We’ll even provide a simple cheat sheet to help them distinguish between the two and demonstrate how they work together to transform a developing story into a precise, machine-readable alert that your security tools can act on instantly. Let’s dive in!
Want to listen on the go? Check out this article in podcast form!
Types of Cyber Threat Intelligence
Before we start exploring data formats, it’s helpful to understand the different types of threat intelligence.

Just like in traditional intelligence, CTI operates on multiple levels to serve different audiences and purposes. Each level addresses different questions and supports various functions within an organization, ranging from the boardroom to the security operations center (SOC).
For a deeper dive, you can check out our quick guide to Cyber Threat Intelligence.
Strategic Intelligence
This is the 30,000-foot view, tailored for executive leadership (CISOs, CIOs, and even the Board of Directors). It focuses on broad risks, trends, and the overall threat landscape to inform high-level business strategy.
Strategic intelligence answers questions like, “Which threat actor groups are most likely to target our industry in the next year?” or “What is the potential business impact of the rise in AI-powered phishing attacks?”
The goal is to provide foresight that guides major decisions about budget, resource allocation, and risk tolerance. It’s typically delivered in the form of briefings, white papers, and long-term reports.
Tactical Intelligence
This level provides insight into specific, impending attacks. It’s about the “who,” “why,” and “how” of an adversary’s campaign, helping defenders understand their intent and capabilities. It focuses on the specific Tactics, Techniques, and Procedures (TTPs) that adversaries use as well as their infrastructure.
Tactical intelligence provides details on how threat actors operate, including the specific malware they deploy, the vulnerabilities they exploit, and the command-and-control protocols they utilize.
For example, a tactical report might explain how a ransomware group uses a particular PowerShell command for lateral movement. This allows defenders to create highly specific detection rules (like YARA or Sigma), tune their security tools, and proactively threat hunt for adversary behaviors within their network, using frameworks like the MITRE ATT&CK.
Operational Intelligence
This is for the hands-on defenders—the security analysts and engineers on the front lines who defend against active or imminent incidents and campaigns.
This kind of intelligence is crucial for network defenders involved in the day-to-day operations of a business. They require atomic indicators (e.g., IP addresses, domains, and hashes), malware signatures, and vulnerability details to prioritize their detection and response efforts.
Operational intelligence is typically delivered in technical, short-form reports, emails, or messages. However, it can also be automated and delivered through threat feeds for the streamlined ingestion and distribution of data to security tools.
Some CTI analysts will define Indicators of Compromise (IOCs), such as malicious IPs, hashes, domains, or URLs, as “technical intelligence.” This intelligence is characterized by high volume and volatility, meaning it’s primarily used for automated security systems. Firewalls, EDRs, and SIEMs utilize these IOCs for quick detection and blocking. Most analysts group this under operational intelligence.
These intelligence levels are built from two data formats: structured and unstructured. Let’s explore these two formats.
What is Unstructured Threat Intelligence? The Story of the Attack
Imagine you’re an analyst starting your day. You come across a detailed 15-page PDF report from a major security vendor that details a new ransomware family.
The report doesn’t just list indicators; it tells a complete story. It describes the suspected motivations of the threat actor, perhaps linking them to a specific geopolitical region and their economic goals. It walks you through their entire attack kill chain, from the initial spear-phishing email to the final encryption of network shares. It even includes screenshots of the C2 panel and quotes from the group’s dark web leak site where they taunt their victims.
This is unstructured threat intelligence.
It’s the rich, human-readable narrative that provides the “who, what, where, when, and why” behind an attack. It doesn’t follow a rigid, predefined format because stories are complex and messy. This intelligence is designed for human consumption and interpretation, relying on an analyst’s expertise to connect the dots.
Unstructured Threat Intelligence Examples
The sources are vast and varied for unstructured threat intelligence. They can include:
- Long-form technical blog posts from independent researchers.
- Urgent security advisories from government bodies like CISA.
- A flurry of tweets from an expert reverse-engineering a new malware sample in real-time.
- Chatter on clandestine dark web forums.
- Even a presentation from a security conference.
The Good of Unstructured Threat Intelligence
This is where the true “intelligence” in CTI lies. Unstructured data is a “gold mine” of critical context. It answers the crucial questions: Why are we being targeted? How does this adversary operate? What can we do about it?
It’s the difference between knowing a domain is bad and knowing it’s part of the C2 infrastructure for a specific APT group known to target your industry.
The Challenge with Unstructured Threat Intelligence
Unstructured threat intelligence is not inherently machine-readable. A firewall or SIEM cannot simply “read” a blog post and understand the nuances. Manually extracting key indicators—the IPs, domains, hashes, and TTPs—from these sources is incredibly tedious and prone to error.
An analyst might spend hours sifting through reports, copying and pasting IOCs, and attempting to validate them. This manual process not only consumes an analyst’s valuable time but also introduces a critical delay between the discovery of a threat and the ability to defend against it.
What is Structured Threat Intelligence? The Actionable Data
Now, let’s continue with that 15-page PDF. After reading the story, the analyst’s job is to distill that narrative into pure, actionable data. They meticulously extract every concrete indicator mentioned in the report, transforming the story into a set of facts:
- IP Address:
198.51.100.10(Identified as a C2 Server) - File Hash (SHA-256):
e3b0c44298fc1c149afbf4c8...(Identified as the malware payload) - Domain Name:
maliciousexample.com(Identified as the phishing source)
When you organize this data into a predefined, consistent format that a computer can understand, you’ve created structured threat intelligence.
Structured intelligence is organized for machines (machine-readable). Its primary purpose is to be ingested, parsed, and acted upon by security tools with maximum speed and precision.
The gold standard for this is STIX (Structured Threat Information Expression), a data format designed specifically to make threat intelligence machine-readable. STIX doesn’t just list indicators; it provides a framework to describe them as objects (like an ‘Indicator’ or ‘Malware’ object) and define the relationships between them.
Structured Threat Intelligence Examples
Structured threat intelligence can be a simple list of IOCs or a JSON file detailing interconnections between atomic indicators, malware families, and threat actors.
For example, a STIX JSON file can contain an ‘Indicator’ object with a pattern describing a SHA256 hash (e.g. [file:hashes.'SHA-256' = 'e3b0c442...']). This can then link to a ‘Malware’ object named NewRansomware, which in turn can be linked to a ‘Threat Actor’ object named EvilCorp.
Other examples include simple CSV files of IOCs for quick ingestion, or entries in a relational database used by a Threat Intelligence Platform (TIP).
The Good of Structured Threat Intelligence
The key benefit of structured threat intelligence is the ability to automate at scale. Security tools can ingest structured data and act without human intervention:
- A SIEM can take a STIX pattern and automatically generate a query to search historical logs for a file hash across thousands of endpoints.
- A firewall can parse a feed of malicious IPs and update its blocklist in seconds.
This enables organizations to scale their defenses far beyond what a human team could manage manually, responding to threats in near real-time and drastically reducing the window of opportunity for an attacker.
The Challenge with Structured Threat Intelligence
Structured threat intelligence often lacks the rich, narrative context that its unstructured counterpart provides. A list of IP addresses doesn’t tell you why they are malicious, who is behind them, or what their ultimate goal is.
Without this context, security teams are forced into a reactive posture, playing a frustrating game of “whack-a-mole.” They block an IP, and the attacker, whose motivations and campaign strategy are unknown, simply switches to a new one.
This is why structured data is most powerful when it’s derived from, and can be traced back to, its unstructured origins.
Structured vs. Unstructured Threat Intelligence: Head-to-Head Comparison
To help illustrate the key differences between structured and unstructured threat intelligence, here is a simple cheat sheet you can reference. Use it to decide which format to use when delivering strategic, tactical, and operational CTI.
| Feature | Unstructured Threat Intelligence | Structured Threat Intelligence |
|---|---|---|
| Format | Free-form, narrative-driven, and without a strict schema (e.g., PDFs, blogs, news articles, text). | Follows a predefined, consistent data model with a rigid schema (e.g., STIX/TAXII, JSON, XML, CSV). |
| Audience | Primarily human analysts who use experience and intuition to interpret nuance and intent. | Primarily machines (SIEM, Firewall, TIP) that require a predictable format for automated parsing and action. |
| Key Value | Provides rich, strategic context—the “why” and “how”—that informs proactive defense and planning. | Delivers operational speed, scalability, and automation, enabling defense against threats at machine speed. |
| Pros | Often the most timely source of new threats; provides deep, detailed understanding of adversary TTPs. | Fully machine-readable, enabling near-instantaneous response and drastically reducing MTTD/MTTR. |
| Cons | Difficult and slow to process automatically; often contains “noise” and requires significant effort to validate. | Lacks narrative context, which can lead to simplistic actions (e.g., blocking a shared IP) without understanding the “why.” |
| Example | A detailed analysis of a FIN7 campaign on a vendor blog, with attack chain diagrams and adversary quotes. | A STIX 2.1 JSON object defining the relationship between the ‘FIN7’ threat actor and the C2 domains it uses. |
Don’t let this head-to-head comparison fool you! It’s not a question of structured versus unstructured threat intelligence; it’s about combining them both to paint a comprehensive picture.
Why You Need Both for a Complete Picture
The most effective CTI programs don’t choose to deliver either structured or unstructured threat intelligence; they build a seamless operational workflow that transforms human insight into machine-speed action.
This process is the core engine of a modern CTI team – they take unstructured data, turn it into structured threat intelligence, and then deliver unstructured intelligence when necessary to inform key decision makers.

Think of it as a sophisticated intelligence assembly line, constantly running within the Threat Intelligence Lifecycle:
- Collect (Unstructured): It begins with data collection. An analyst, monitoring OSINT sources, discovers a detailed blog post from a security researcher about a new phishing campaign. This isn’t a formal feed; it’s a piece of prose, rich with screenshots, analysis, and speculation about the attacker’s motives and targets, which happen to align with the business’s industry.
- Process (Analyze): Here, human expertise is paramount. The analyst doesn’t simply copy and paste the indicators. They read the report to gain a thorough understanding of the adversary’s playbook. They map the described behaviors to the MITRE ATT&CK framework, identifying specific TTPs like “T1566.001 – Phishing: Spearphishing Attachment” and “T1059.001 – Command and Scripting Interpreter: PowerShell.” This crucial step adds a layer of behavioral context that a simple IOC list would miss.
- Produce (Structured): Now, the transformation happens. The analyst uses a TIP (e.g., OpenCTI, MISP, etc.) to create a rich, structured intelligence object. They don’t just create a list of IOCs; they build a threat graph. The malicious file hash is encapsulated in a STIX ‘Indicator’ object, which is then explicitly linked via a ‘Relationship’ object to a ‘Malware’ object named NewPhisher. This, in turn, is linked to the identified TTPs and a ‘Threat Actor’ profile. This creates a multi-dimensional, machine-readable product that retains the context of the original report.
- Disseminate (Automation): This is where the magic happens. The structured STIX object is pushed from the TIP to the entire security ecosystem:
- The firewall instantly ingests the malicious domain and IP addresses, blocking any future communication.
- The EDR platform receives the file hash and immediately initiates a hunt across all endpoints for any sign of the malware, past or present.
- The SIEM creates a new high-priority correlation rule, ensuring that if any IOC from this specific campaign is seen anywhere on the network, it generates an immediate, context-rich alert for the SOC team.
This hybrid approach creates a virtuous cycle. It leverages the irreplaceable cognitive skills of human analysts to find the narrative and understand the “why,” then translates that understanding into a structured format. The added structure allows machines to execute the “what” (blocking, alerting, hunting) at a scale and speed no human team could ever hope to match.
It also enables you to transform structured data into unstructured intelligence without losing any context.
For instance, if you are reporting strategic intelligence to an executive, they don’t want a list of indicators or a raw JSON file. They want a nicely formatted, organized, and succinct report that is human-readable. To produce that report, you can use structured data in your TIP. This data still retains its original context (thanks to STIX); it just requires you to analyze it, extract strategic insights, and deliver these in a report.
It’s important to note that many TIPs provide you with the ability to ingest STIX/TAXII feeds, without you having to collect and turn unstructured intelligence into structured data yourself. This automated ingestion of structured threat intelligence is why TIPs are so valuable; they serve as the factory floor for this assembly line, accelerating the CTI lifecycle and freeing up analysts to focus on the next emerging threat.
Tools for Bridging the Gap: Unstructured to Structured Threat Intelligence
The process of manually reading a blog post and copying and pasting every IP address, domain, and file hash into a spreadsheet is a significant bottleneck for any CTI team. It’s slow, prone to error, and simply doesn’t scale.
To truly operationalize the intelligence lifecycle, analysts need tools that can accelerate the transformation of unstructured text into structured, actionable data. Fortunately, the open-source community has produced some excellent utilities to help bridge this gap.
- Obstracts: A tool for transforming blog posts into organized threat intelligence. It extracts indicators (IOCs and TTPs) from any blog post you provide, helping to automate the extraction of information and convert unstructured data into machine-readable threat intelligence.
- Stixify: A tool similar to Obstracts, but instead of blog posts, it takes any uploaded unstructured data (e.g., PDFs, Word docs, PowerPoints, emails, Slack messages, etc.) and generates a STIX object of this data. This makes it easy to integrate with other security tools and saves you from creating your own STIX bundles.
Together, these tools form a powerful, open-source workflow. An analyst can use obstracts to quickly rip all the technical observables from any blog post or stixify to create a context-rich, machine-readable STIX package from any unstructured data they discover.
By leveraging tools like these, CTI teams can dramatically increase their speed and efficiency, freeing up valuable analyst time to focus on what they do best: understanding the adversary.
Conclusion
CTI teams often face an overwhelming amount of data. The key to staying on top is mastering the flow between unstructured and structured threat intelligence.
- Unstructured threat intelligence provides rich narrative context—the “why” behind an adversary’s actions. Ignoring it means flying blind, reacting to isolated incidents without understanding the broader campaign, leading to a perpetual game of whack-a-mole.
- Structured threat intelligence is machine-readable data that powers automated defenses—the “what,” or specific IOCs that your tools can block instantly. Without it, brilliant analyst insights remain trapped, failing to translate into real-time protection, resulting in slow and manual responses.
A mature CTI program transforms human-readable narratives into actionable, structured intelligence. This enables proactive defense, allowing you to translate new TTPs (unstructured) into custom SIEM detection rules (structured) to hunt adversary behavior.
Mastering this relationship elevates the CTI analyst, harnessing machine speed without sacrificing human wisdom, defending networks at scale, anticipating threats, and turning data into actionable defense.
Frequently Asked Questions
What is the Purpose of STIX?
STIX (Structured Threat Information Expression) standardizes cyber threat intelligence, enabling seamless sharing and understanding. Its machine-readable format utilizes defined objects, such as ‘Indicator’, ‘Malware’, and ‘Threat Actor’, which are linked by ‘Relationship’ objects. This structure allows tools to understand complex threat data, facilitating intelligent automation, vendor interoperability, and collaborative defense.
What are the Main Types of Threat Intelligence?
While sometimes grouped differently, threat intelligence is often discussed across three main levels, each serving a different purpose and audience:
- Strategic: This is high-level intelligence for executives. It focuses on the “big picture” of the threat landscape, including which adversary groups are likely to target the organization’s industry and their motivations. It answers the question, “What are the major cyber risks to our business goals?”
- Tactical: This focuses on the Tactics, Techniques, and Procedures (TTPs) and infrastructure of adversaries. It’s about the “who, why, and how” of a threat, providing insight into an adversary’s capabilities.
- Operational: This intelligence provides details on specific, active, or imminent attack campaigns. It focuses on granular indicators (IOCs), such as IP addresses, file hashes, and domain names. It is highly time-sensitive.
What is Unstructured Data in Cyber Security?
Unstructured data in cyber security includes blogs, news, and social media, and is a “gold mine” for threat intelligence. It provides crucial attack context (the “how” and “why”) missing from raw indicators. For instance, a blog can reveal the intent behind a phishing campaign, unlike a simple list of malicious domains. The challenge lies in its automated processing, which requires either human or advanced AI/NLP analysis to extract meaningful signals.



