I have been thinking a lot about this problem in cyber threat intelligence where experts are faced with describing and mapping the kill chain an attacker followed, this is usually done with mapping actions within the report to MITRE ATT&CK techniques.
At the first layer of this onion, it seems simple to me. You have a CTI report. You want the model or human to tell you which techniques show up. Maybe PowerShell. Maybe registry persistence. Maybe command and control. Feed the report to an LLM and ask for the ATT&CK IDs.
Not so simple when you dig in
Not really when you dig into it deeper. It's very interesting as even a good portion of cyber security experts will disagree on the correct ID for them.
A single sentence can imply more than one technique. A whole report can contain a bunch of techniques. Some techniques show up constantly in training data, while others barely appear at all (the current benchmarks for this domain are pretty limited and are on the smaller size).
On 2026-04-12, activity attributed to actor CIRCE-ORBIT was observed targeting a Windows environment. Initial access was followed by execution of a PowerShell stager registered as a scheduled task that re-launched at every user login. Post-exploitation, the implant used rundll32.exe to sideload a payload from disk and beaconed to a C2 endpoint over HTTPS in fifteen minute intervals.
Hundreds of yes/no questions per sentence
So the model is not just answering one question. It is answering hundreds of tiny yes or no questions at once in some form.
Most of the time the answer is no. That creates a weird learning problem for this issue, we are overwhelmed by the negatives depending on the benchmark. The model sees an ocean of negative examples and only a few drops of positives for techniques.
- T1059.001PowerShell
- T1059.003Windows Command Shell
- T1053.005Scheduled Task
- T1547.001Registry Run Keys
- T1105Ingress Tool Transfer
- T1027Obfuscated Files
- T1055Process Injection
- T1218.011Rundll32
- T1071.001C2 · Web Protocols
- T1140Deobfuscate Files
The benchmarks are frozen
We can even see it with the industry standard benchmark of TRAM2 and AnnoCTR, one of which is made by MITRE themselves, these benchmarks don't cover nearly enough techniques, and for AnnoCTR we still need to figure out the upper limit of the benchmark itself.
This issue persists to evaluation time as well, we can see that sentence predictions get rolled up to the report level. That means false positives can pile up across a report, but missed evidence can also kill recall. Sounds fun, right?
The 72% ceiling
Most of the industry has either been throwing larger models at the problem or larger cybersecurity training data, the industry standard right now scores a 72% across TRAM2, this is pretty low. This leads me to think two things here; 1. It's possible the benchmark is part of the problem too, and honestly, it kind of is. 2. We should be looking at how we fine-tune the encoder rather than brute-forcing this benchmark.
LLMs are very good at a lot of things nowadays, but CTI extraction is not just a vibes based summarization task. It is a sparse multi label extraction problem with a weird label space, limited expert data, and a metric that punishes the model in ways that are not always obvious.
Bigger models don't deploy
A bigger model can help, but it also brings deployment problems. A lot of defender environments are not going to casually run an 8B model for every report. They may not have the GPU budget. They may not want sensitive reports leaving their environment, etc.
The bounty that doesn't exist
What's strange to me is nobody's really built the “Bugcrowd for CTI attribution” thing yet. Vulnerability disclosure figured this out years ago: you can't automate finding every zero-day, so you build a marketplace where people get paid to fill the gap. Attribution has basically the same shape of problem, it's sparse, people tend to disagree a lot etc.
Right now what exists is either some vendor's internal tagging pipeline that is essentially a blackbox in some cases, or upwards or 40k a month lol,
Or you run into benchmarks like TRAM2 or AnnoCTR that got built once and then just sits there while ATT&CK keeps evolving around it. Neither one actually closes the loop. Nobody's paying analysts to hash out the hard sentences and turning that into something useful.
Disagreement is data
This is kind of the main theme I come back to, the disagreement itself is data. When two solid analysts read the same sentence and land on different technique IDs, that's not just noise you average away. That's telling you something, either the taxonomy is fuzzy right there, or the report itself didn't give enough to go on.
A bounty-style setup could actually use that in a way, multiple people weigh in on a report, you pay for agreement and you pay for disagreement that's actually informative, and over time you end up with a dataset that keeps growing instead of one that was annotated once and frozen.
Who this helps
For a defender perspective, we could hopefully take the load off of them, we constantly hear of alert fatigue within blue teams and how to resolve it. Instead of just trusting a model that's right maybe 72% of the time on a good benchmark, or a black box, you could get a second opinion from people who've read a thousand of these reports before you decide what to do about it. For the field, it's a way to actually grow the label space instead of everyone fighting over the same small set of examples forever, or relying on these larger companies to refine these benchmarks.
Where this gets hard
Kind of ironic in a way though with what I mentioned above, the reports worth attributing are usually the ones a company least wants leaving the building, which is the exact same reason nobody wants to run an 8B model on them either. And incentives are tricky in a different way than bug bounties: “is this a real RCE” has a pretty clean yes/no answer, but “is this T1059.003 or T1059.001” can stay genuinely arguable even between two senior people. You'd need to pay for good judgment, not just volume, or you'd get the same gaming problems bounty programs always run into, or what we see now, the war on AI slop reports.
Where that leaves us
This doesn't fix the modeling side of things I've been ranting about above, better encoders, better negative sampling, better eval design still all matter. But it'd actually go after the data problem underneath it, which is honestly the bigger bottleneck in the world of machine learning.