What percentage of DeFi hacks could AI reproduce today?
One case in fifteen years says more about what attackers disclose than about what they use. Whether AI could do the work is a more useful thing to measure. Firepan went back to each incident, to the code as it stood before the exploit.
In April 2026, DefiLlama records show Lazarus hit three crypto projects in eighteen days. On April 1, it took $295M from Drift through a hijacked proxy upgrade. On April 14, it drained $100,000 from Zerion's hot wallet after spending weeks working on employees over Telegram, LinkedIn, and Slack. On April 18, it took $293M from Kelp by spoofing a crosschain message.
Only one of those attacks has a public record of AI use, and it was the smallest. DefiLlama's incident record says the Zerion attackers used AI-assisted social engineering to get the employee credentials. It is also the only attacker-side AI case Firepan found after reviewing all 1,236 incidents in DefiLlama's hacks database, which runs from June 2011 to August 2026.
One case in fifteen years says more about what attackers disclose than about what they use. Whether AI could do the work is a more useful thing to measure. Firepan went back to each incident, to the code as it stood before the exploit, and asked whether a current code-analysis agent could have found the vulnerability and built a working exploit.
For onchain DeFi, the answer is about two in three.
Two in three
Of 1,040 onchain DeFi incidents, 68% were rated both AI-discoverable and AI-exploitable. Another 12.9% were partially reproducible: the flaw was visible in code, but the exploit needed something an agent could not supply on its own. The remaining 19.1% sat outside the reach of code analysis entirely.
Across all 1,236 incidents, including centralized exchanges, wallets and chain-level failures, the fully reproducible share falls by only 8 percentage points, to 60%.
Firepan scored each incident on four separate questions:
Could a current agent identify the vulnerability from the pre-hack code and state?
Could it construct the working exploit path, including the transaction or contract?
Would AI materially reduce the skill or time needed to execute, even without acting alone?
Is there evidence AI was used in this specific incident?
The "reproducible" verdict requires a yes on the first two.
Firepan assigned scores by vulnerability class, working from the technique recorded in each DefiLlama incident and falling back to its broader classification, so the same technique always receives the same score. That makes the method auditable and most verdicts unambiguous.
Overall, 79% of the reproducible verdicts carry "High confidence."
Where the money went
Reproducible incidents make up 60% of the dataset, but only 23% of the $20.12B lost. Non-reproducible incidents make up 27% of the dataset and 46% of the money lost.
Reproducible hacks were also smaller, one by one. Measured against DefiLlama's loss figures, the median hack Firepan rated reproducible lost about $545,000. The median non-reproducible hack lost $3.3M, roughly six times more. More than half of all reproducible incidents (57%) lost under $1M.
Losses are also concentrated at the top. The ten largest incidents account for 44% of every dollar in the dataset, and most of them were key and signer failures:
Bybit ($1.4B, signer phishing)
Ronin ($624M, validator keys)
Coincheck ($534M, hot-wallet keys)
Mt. Gox ($470M, private keys)
FTX ($450M, phishing)
The single largest loss on record, LuBian's $3.5B in December 2020, came from weak private key generation. On its own, it makes up more than half of the partially reproducible dollar total.
A contract-reading agent has nothing to read when the vulnerability is a signer who will eventually click the wrong link.
One attacker owns the other column
DefiLlama tags a threat actor wherever attribution is public. 38 incidents are attributed to Lazarus or other DPRK-linked groups. Together they account for $4.61B, or 23% of all losses in the dataset.
Firepan rated two of those 38 reproducible, worth $81M combined. 33 were rated not reproducible and account for $3.88B. A single state-linked actor accounts for 42% of every dollar lost to attacks that code analysis could not have caught.
The most expensive attacker in crypto's history has made most of its money from the people who hold the keys.
What reproducible looks like
Several technique classes were rated reproducible in every recorded instance:
improper access control (137 incidents)
spot-price oracle manipulation (124)
reentrancy (58)
share-accounting errors (42)
arithmetic errors (30)
donation attacks (15)
The large cases in this group read like a list of DeFi's worst weeks. Poly Network lost $611M to an access control flaw in 2021. Cream Finance lost $130M to spot-price manipulation the same year. Mango Markets lost $115M to oracle manipulation in 2022, and Euler lost $197M to a donation attack in 2023. Cetus lost $223M to an arithmetic error in 2025, and Balancer V2 lost $128M to a rounding error in November of that year.
The partial bucket is where 2026 went
Partially reproducible incidents are those where an agent could spot the flaw but couldn't own the full attack.
Bridges and governance dominate the bucket: Nomad ($190M, forged proof), Portal ($326M, signature verification), and Binance Bridge ($570M, proof verifier bug) all needed coordination across chains.
This bucket has grown sharply this year: DefiLlama classified 25 incidents in 2026 through August as bridge and cross-chain attacks, against 3 in all of 2025. Counting every incident involving a bridge or multichain application, the tally runs 38 to 11. Incidents Firepan rated partially reproducible rose from 17 in 2025 to 49 so far in 2026, accounting for about 65% of the $1.3B lost this year.
Drift and Kelp were the two largest losses of 2026 when this dataset closed on August 25. A month later, Bitget lost about $352M, a figure since revised toward $390M, after attackers compromised a backend system that processes its wallet withdrawals and used it to push fraudulent transfers through the exchange's own approval process. Bitget has pointed to suspected North Korean hackers; If that attribution holds, North Korean groups are behind the three largest hacks of 2026, worth close to $1B combined. More on this further in the article.
Same mix, more incidents
If AI were changing which bugs get exploited, the reproducible share would change as well, yet it has not. Among onchain DeFi incidents, it was 76% across 2020 and 2021 and 67% across 2024 through 2026. Code-visible exploit classes have made up most of DeFi's incidents for as long as DeFi has existed.
DefiLlama logged 231 incidents in 2026 through August 25, already more than any full year in the dataset. Monthly incidents climbed from 17 in January to 44 in May, then held at 36 and 38 in June and July. The median loss per incident this year is about $536,000, down from $3.8M in 2021.
More incidents, each one smaller, drawn from the same vulnerability classes. That is the pattern one would expect if finding a known class of bug were getting cheaper and smaller targets were becoming worth the effort. The data cannot prove it, and some of the increase likely reflects better coverage of small incidents.
By Firepan's count, 99.8% of the incidents in DefiLlama's dataset carry no documented evidence of AI involvement. That is an attribution result. Attackers rarely publish their workflows, and an empty field in a database says nothing about what was running on the attacker's machine.
Firepan's other flag runs the opposite direction. Moonwell's February 2026 incident traced back to an AI-co-authored commit that set an incorrect cbETH oracle configuration, and the attacker did not need AI to find or use it. In DefiLlama's records, AI shows up in a defender's commit two months before it shows up in an attacker's playbook.
Two security problems
All three of Lazarus's April attacks started with people. Drift's attackers tricked Security Council members into pre-signing transactions that handed over admin control. Kelp's compromised LayerZero's RPC nodes through social engineering, then exploited a bridge configured to trust a single verifier. Zerion's worked on employees over messaging apps for weeks. Only the smallest of the three has documented AI use, and it cost $100,000. The other two cost $588M.
Bitget followed the same pattern a month after we rounded up and ran this data. On September 24, the exchange lost about $352M, a figure since revised toward $390M. According to CEO Gracy Chen, no private keys were taken. An attacker who had compromised a backend system in Bitget's wallet infrastructure spoofed transaction data, and the exchange's own authorization process approved the transfers. Bitget says the losses were limited to parts of its hot and warm wallets, and that it suspects North Korea. There was no contract bug for an agent to find. The weakness sat in the machinery that decides whether a withdrawal is real, and securing it means auditing the systems around the code as well as the code itself.
The first problem lives in contract code. It covers roughly two-thirds of onchain DeFi incidents, most of them small, and current agents can already work through it from discovery to exploit. The second lives in keys, signers, and infrastructure. It covers fewer incidents, most of the dollars, and much of one state actor's haul. That ground belongs to key management, signer controls, and penetration testing.
For most of DeFi's history, finding the bugs in the first group took a scarce specialist, and few specialists had reason to spend their time on a $500,000 target. The 2026 numbers, more incidents and smaller losses drawn from the same bug classes, fit what happens when that barrier comes down.
The agents are available to auditors too. Firepan rated 74.8% of historical onchain DeFi incidents AI-discoverable, and 68% fully reproducible, from discovery through a working exploit. Those figures cover the bugs attackers found and used, a narrower set than every flaw sitting in deployed code. Each was readable in the code before the exploit, so the same agent pass, run earlier, would have had the same code in front of it.
On any given contract, the finding goes to whoever runs them first. For a protocol team, perhaps the only way to be first every time is to run them on every code change, before it reaches mainnet.