This note sets out eight findings. It cannot tell you which three to fund. Zak Raxter, who runs red teaming and adversary emulation and was at both conferences, ranks them in a 20 minute session on 1 October, 11:30 AM ET.
Register free Recording sent to registrants
Defenders have long been protected by a budget constraint that belonged to the attacker. Deep attention was expensive, so it was rationed, so most of your estate never received any. Black Hat USA 2026 and DEF CON 34 were the week that assumption stopped being safe to hold.
Microsoft's David Weston named it from the keynote stage on 5 August, in a talk titled "The End of Rare: Defending When Offense Is Cheap".1 He put numbers on it from the stage. Pointed at 200 Linux kernel vulnerabilities, a Microsoft system generated crash-level proofs of concept for 182 of them, at an average compute cost of $3.61 and 21 minutes each.2
That is the economics behind everything below. Eight pieces of research across the two conferences show what follows when attacker reasoning stops being scarce. This note states what each one breaks, sources every claim, and is honest about the one judgment it cannot make for you.
- Exploit generation now costs single-digit dollars and minutes, on the keynote's own figures: 182 of 200 kernel vulnerabilities turned into crash-level proofs of concept at $3.61 each.2
- The software supply chain moved up a layer. The unit of compromise is no longer the package but the instruction file an agent reads. More than 30 percent of the malicious skills catalogued abused the coding agent itself as a malware dropper.3
- Isolation in seven AI and automation products was a language-level restriction, not a runtime boundary. Four CVEs, scored 8.3 to 9.9 by the researchers.6
- AI produced a vulnerability class that did not previously exist, across roughly 700 vulnerable systems, including an Apache Traffic Server zero-day. It emerged with a researcher placed at one specific stage, not from autonomous operation.9
- Three assumptions that predate AI broke independently of it: NAT as segmentation, a valid code signature as integrity, and hook inspection as detection. One primitive involved has shipped since Windows 7.11,13,16
Security and risk leaders responsible for exposure management, detection engineering and AI adoption should:
- Bring agent instruction files inside software supply chain review this quarter, by routing a new skill, extension or instruction file through the same approval path that governs a new dependency.
- Inventory every place untrusted or model-generated code executes, and ask each vendor in writing what enforces isolation at runtime, not in the language.
- Stop crediting NAT with segmentation, by treating trusted and untrusted workloads sharing a NAT table as adjacent in your lateral movement model.
- Re-test any detection that rests on a valid code signature or on conventional hook inspection, both demonstrated bypassable on current operating system versions.
- Establish whether a technique is commodity or specialist before funding anything expensive, because that judgment, not the CVSS score, sets the order. See the last column of the table below.
What would change this reading. Evidence that the agent marketplace vector is sustained rather than opportunistic, or a second team reproducing the AI research result without a specialist in the loop. Either would move items below from watch to fund.
The eight findings, and the column nobody can fill
Everything below is public and sourced. The last column is the one that sets your order, and it is mostly empty. A disclosure document exists to prove a technique works. It is not written to tell you whether anyone is using it against you yet, so where we could not source an answer we have left one.
| What it breaks | Demonstrated | What an attacker needs | Commodity or specialist? |
|---|---|---|---|
| Exploit development is slow and expensive | 182 of 200 kernel PoCs, $3.61 and 21 min each2 | Compute budget, falling | Not established |
| Dependency review covers your supply chain | Malicious agent skills, credential theft at scale3,4 | A developer who installs without review | In the wild. Live campaign, since removed4 |
| The sandbox isolates untrusted code | Pyodide escapes in 7 products, CVSS 8.3 to 9.96 | The ability to submit code to the product | Not established |
| Novel bug classes need rare human talent | New class found by AI, ~700 systems, one 0-day9 | Tokens, plus one skilled operator | Specialist on the researchers' own account9 |
| NAT provides segmentation | NatJack: session hijack, DNS poisoning, table exhaustion11 | Position behind the same NAT table | Patches raise difficulty, do not close it11 |
| A valid code signature means integrity | dylib hijacking on macOS 26, signature intact13 | Write access to an app bundle | Not established |
| Hook inspection catches tampering | Patchless hooking via instrumentation callbacks16 | Code execution on a Windows host | Not established |
| Trained employees are harder targets | Social engineering built for primed targets19 | Contact with the target | Not established |
Last column is our reading, not the researchers'. Where it says "not established" we could not find a source that answers it either way. The three rows we could answer are answered from the disclosures themselves, not from any assessment of your sector.
Eight rows. You will action three. Five blanks in the column that decides which three, and no amount of further reading fills them, because nobody publishes what attackers picked up last month. The rest of this note gives you the mechanism behind each row.
1. The supply chain attack moved up a layer
The most operationally relevant disclosure of the week was not a memory corruption bug. It was a distribution problem.
Michael Bargury and Tamir Ishay Sharbat of Zenity presented "Promptware EOD: Skillful Agent Detonation", on a credential-stealing campaign distributed through an AI agent skills marketplace.5 Zenity's disclosure puts it at more than 1.7 million aggregate installs, via typosquatted look-alikes of two widely used services.4 The payload was not a binary. It was installation instructions that caused the agent itself to fetch and run attacker-controlled code.
What it collected: SSH keys, cloud credentials, Git and package manager tokens, Kubernetes and Docker configurations, database credentials, infrastructure-as-code credentials, environment files and service account files, packaged with host metadata and sent to attacker infrastructure.4 That is the contents of a developer workstation, which in most organizations is a working path to production.
The structural finding is in Zenity's own research write-up rather than the press release: more than 30 percent of the malicious skills abused the agent itself as a malware dropper, and behind the live campaign sat hundreds of reserved but empty package names, staged for later use.3 The marketplace operators removed the identified threats within hours of notification.3,4 The reserved names are still reserved.
Five years of supply chain work went into dependency provenance: SBOMs, signed artifacts, registry allowlists. None of that machinery reads a markdown file that tells an agent what to do.
Open question. One campaign, removed within hours. Sustained offensive investment, or an opportunistic land grab the marketplaces will now police? That answer decides whether this is a Q4 control change or a watch item.
2. The sandbox that was not a sandbox
At DEF CON 34, Vladimir Tokarev and Saar Pearl of Cyera presented "WASM Was Not the Boundary: Sandcastles, Not Sandboxes," on products that run untrusted Python in the browser using Pyodide, a WebAssembly build of CPython. Seven of them relied on Python-level restrictions that did not isolate untrusted code from the host.6
The mechanism generalizes beyond Pyodide, which is why it matters. WebAssembly does not stop code reaching capabilities the embedding environment deliberately exposes. In the researchers' words, ctypes.CDLL(None) loads the main program's symbol table, the equivalent of dlopen(NULL) in C, and in Pyodide the main program is the Emscripten-compiled WASM module.6 From there, Emscripten's own exports reach the host JavaScript runtime. The Python-level restrictions never covered that path.
| Product | CVE and severity, as the researchers published them6 |
|---|---|
| n8n | CVE-2025-68668, CVSS 9.9 |
| cohere-terrarium | CVE-2026-61522, CVSS 9.3 (reserved on NVD at time of writing) |
| Grist | CVE-2026-24002, CVSS 9.1 per Cyera, 9.6 on NVD7,8 |
| smolagents | CVE-2026-10613, CVSS 8.3 |
| langchain-sandbox, stlite, cibuildwheel | Affected, no CVE assigned |
What an escape reaches depends on what embeds the sandbox, and the researchers are specific about it: in Node.js, the credentials, OAuth tokens, database access and API keys available to the service process; in CI/CD, release credentials, signing keys and source checkouts; in agent frameworks, delegated credentials and tool access.6
Seven products is the set two researchers chose to examine, not the set that is affected. Nothing in the disclosure tells you how many agent tools already running in your estate share the architecture. Start with anything that offers users a formula bar, a code cell or a "run this snippet" box.
3. AI found a new bug class, and the human was not optional
James Kettle, Director of Research at PortSwigger, built HTTP Terminator to settle whether AI can do original security research or only rediscover known patterns. The work was presented as "Can AI Do Novel Security Research? Meet the HTTP Terminator" at both conferences.9
The system worked from 138 specifications, generated roughly 15,000 micro-inspiration fragments and 30,000 unique attack vectors after deduplication, and confirmed roughly 700 vulnerable systems against authorized live targets.9,10 Exploitable flaws turned up in BeyondTrust Secure Remote Access, F5 BIG-IP, Citrix NetScaler, and Apache Traffic Server, where the autonomous cascade produced a zero-day, CVE-2026-63078.9
It also produced a vulnerability class nobody had named: shared-parser confusion. The mechanism is worth understanding because it is an architectural pattern, not a bug. Many servers use the same parsing code for HTTP requests and HTTP responses. Because the code path is shared, response-specific parsing rules get applied to inbound requests, and as the research puts it, "any response-processing feature could be exploited by a request".9
The worked example is a content type the RFCs specify as response-only. Send Content-Type: multipart/byteranges on an inbound request and servers that reuse response-parsing logic will process it, with front end and back end disagreeing about where the message ends. That disagreement is the desync.9
Where that class came from should shape how you staff this. It emerged in the human-guided cascade phase, not from autonomous operation. The paper is direct: the system's power is unlocked by putting a researcher back in at exactly one point, the discovery cascade.10 An experienced researcher amplifies an AI research system rather than being replaced by one.
The question is no longer whether AI can find novel vulnerabilities. It is where in your process a human has to stand for that to happen.
Open question. PortSwigger put one of the field's most experienced request smuggling researchers at that point. Does the same amplification work with an ordinarily competent operator, or collapse without the expertise? That is the difference between a technique a few specialists run and one your threat model must assume is commodity.
4. Three boundaries that stopped holding
NAT is not a segmentation control
Malcolm Stagg of the Synack Red Team disclosed NatJack at Black Hat USA 2026, in a session titled "Breaking Trust Boundaries: Exploiting Design Assumptions in Network Infrastructure." Four techniques are described: hijacking active TCP connections, poisoning DNS responses, identifying the ports assigned to other connections, and forcing denial of service by exhausting a device's NAT table.11
Two CVEs cover implementation-specific flaws: CVE-2026-56181 in Windows NAT in Hyper-V, and CVE-2026-63913 in the Linux netfilter conntrack subsystem.11
The part that matters for planning is the remediation status, and it is stated plainly by the researcher's own employer. Available fixes, including Linux kernel 6.6.142 and later and FreeBSD 15.0 and later, "raise the difficulty of exploitation but do not close the underlying design gap".11 The guidance is to prioritize encryption in transit, segmentation that isolates untrusted workloads, and IP Source Guard, and to look first at environments where trusted and untrusted workloads sit behind the same NAT table.11,12
Code signatures survived the hijack
Patrick Wardle's DEF CON 34 talk, "Dylib Hijacking on macOS: Dead or Alive?", demonstrated the technique still working against macOS 26.13,14
The reason the signature survives is worth knowing, because it tells you what to look for instead. The technique abuses two legitimate loader behaviors.15 A weak dylib, declared with LC_LOAD_WEAK_DYLIB, is an optional dependency: if it is missing the loader continues rather than aborting, so planting a file at the path of a dependency that was never shipped gets your code loaded. Separately, run-path dependencies declared with @rpath are searched across several directories in order, so a library planted in an earlier directory is found first. In both cases the loader does exactly what it was designed to do, and the application's own signed contents are untouched.
If your macOS detection treats a valid signature as evidence of integrity, that is the assumption under test. Wardle's own write-up points at the detection that does work, watching what actually gets mapped into the process rather than whether the bundle still verifies.14
Patchless hooking, shipping since Windows 7
Yoann Dequeker of Wavestone ran a four-hour DEF CON 34 workshop, "Explore the Windows Instrumentation Callback", covering weaponizing the mechanism for execution hijacking, process injection and beacon obfuscation. The workshop abstract describes Nirvana Debug as a type of instrumentation callback existing since Windows 7.16 The technique itself was first published by Alex Ionescu at Recon in 2015.17
The mechanism explains why hook-scanning misses it. An instrumentation callback is a per-process user mode callback on system traps, registered through NtSetInformationProcess and stored by the kernel in the process structure.18 It fires at the moment the kernel returns to user mode after a syscall completes, before the original return instruction runs, which makes it the first code to execute on the way back. Nothing in ntdll is patched. Disassemble the syscall stubs and they are clean, because the redirection happens in kernel-managed control flow rather than in the user mode bytes your tooling inspects.
For a detection engineering backlog, the age of the primitive is the point. This is not a new Windows feature being abused. It is an eleven-year-old published technique that most detection content never covered, and a scanner looking for patched function prologues will report clean forever.
These three sit at very different distances from your environment. One needs position behind a shared NAT, one needs a macOS fleet, one is a primitive from 2015. Their CVSS scores say nothing about which deserves your next sprint.
5. The human layer stopped being the cheap one to defend
Awareness training rests on an assumption about volume: that attempts are numerous but unconvincing, and a trained employee learns to spot the tells. Two separate developments put pressure on each half of that.
On the volume side, Daniel Fabian, Head of Google Red Teams, published a position piece during Black Hat week arguing that red teams must build their own agentic systems, and that "the side that operationalizes agents first sets the tempo".20 Read next to Weston's $3.61 figure, the volume ceiling is the thing that moved.
On the quality side, Dr Megan Squire, Principal Threat Intelligence Researcher at F-Secure, presented "What Scammers Know That Social Engineers Don't: Three Techniques for Tough Cases" at DEF CON 34's Social Engineering Community village.19 The three techniques, awareness hijacking, complexity-as-crucible and first win, are drawn from criminal fraud research and are built for cases where, in the session's own words, "targets may be on high alert or trained against falling for more traditional manipulation techniques."19
That is the uncomfortable pairing. One removes the ceiling on volume. The other is built specifically for the people your training has already reached. Neither is a vulnerability you can patch, and neither appears in a scan.
Open question. Re-scoping a phishing program around attacks designed for trained people is a budget line. Whether to spend it this year depends on whether these techniques have reached the crews that actually target your sector.
The decision this note cannot make for you
Four of the five recommendations at the top are cheap, reversible and correct regardless of how the year goes. Make them.
The fifth is different, and here is what it looks like in practice. Do you fund a detection engineering sprint on instrumentation callbacks this quarter, or defer it twelve months? That is a team's worth of time either way, and it turns entirely on whether the primitive sits in tooling your likely intrusion sets already carry, or is still run by hand by a few specialists. The CVSS score does not say. The paper does not say. Get it wrong one way and you burn a quarter of detection capacity on something nobody is using against you. Get it wrong the other way and you are blind to it for a year.
The same shape applies to rebuilding segmentation that leans on NAT, and to re-scoping a phishing program. Each is a budget line, and each depends on a judgment the published record will not make.
There is no survey of what red teamers adopted in August. Nobody runs one, and if they did it would land next spring. What exists instead is the people who were in the rooms.
Zak Raxter is one of them. He runs red teaming, adversary emulation and incident forensics, with over a decade in offensive security and prior roles at Juniper Square, Prudential Financial and Mosaic Group. He was at both conferences, and he is not reporting on them.
The published agenda is three questions: how AI is reshaping attacker tactics across reconnaissance, exploitation and attack chaining; which emerging techniques and broken assumptions warrant attention; and how security teams should adapt.21 Twenty minutes, with time to put your own question to him.
One caveat, stated plainly because this is a peer community and not a marketing channel. This is one practitioner's read, not an industry survey, and one operator's kit is a sample of one. It is also the only direct evidence available on the commodity question right now, and normally you have to be in the right room to hear it.
Field Notes from Black Hat & DEF CON 2026
1 October, 11:30 AM ET / 7:30 PM GST / 9:00 PM IST. Twenty minutes, live, with time for your questions. Free
New to CISO Platform? Join the community free for frameworks, checklists and peer discussion.
Questions security leaders are asking
What changed in the attacker playbook at Black Hat and DEF CON 2026?
The dominant shift was the falling cost of attacker reasoning, framed in Microsoft's David Weston keynote "The End of Rare: Defending When Offense Is Cheap," where he reported generating crash-level proofs of concept for 182 of 200 Linux kernel vulnerabilities at an average $3.61 and 21 minutes each. Supporting research showed AI-assisted discovery of a new vulnerability class, agent instruction files used as a supply chain vector, and sandbox, NAT and code signing boundaries demonstrated as bypassable.
What was the malicious AI skills campaign disclosed at Black Hat USA 2026?
Zenity Labs disclosed a credential-stealing campaign distributed through an AI agent skills marketplace, reported at more than 1.7 million aggregate installs. Typosquatted skills impersonated popular services and inserted installation instructions causing agents to execute attacker-controlled code, which harvested SSH keys, cloud credentials, Kubernetes and Docker configurations, database credentials and environment files. More than 30 percent of the malicious skills catalogued abused the coding agent itself as a malware dropper.
What is the Pyodide sandbox escape presented at DEF CON 34?
Researchers from Cyera showed that seven products running untrusted Python via Pyodide relied on Python-level restrictions rather than runtime isolation. Calling ctypes.CDLL(None) loads the main program's symbol table, which in Pyodide is the Emscripten-compiled WASM module, and Emscripten's exports then reach the host JavaScript runtime. Four CVEs were published, scored 8.3 to 9.9 by the researchers, covering n8n, cohere-terrarium, Grist and smolagents, with langchain-sandbox, stlite and cibuildwheel also affected.
What is shared-parser confusion?
Shared-parser confusion is a vulnerability class named by PortSwigger's HTTP Terminator research. Many servers use the same parsing code for HTTP requests and responses, so response-specific parsing rules get applied to inbound requests. The practical consequence is that any response-processing feature becomes reachable from a request. The worked example sends a Content-Type of multipart/byteranges, which the RFCs specify as response-only, on an inbound request, causing front end and back end to disagree about where the message ends.
Can AI do original security research?
PortSwigger's HTTP Terminator research found that it can, with a qualification. The system generated roughly 30,000 attack vectors and confirmed roughly 700 vulnerable systems, produced the previously unnamed shared-parser confusion class, and found a zero-day in Apache Traffic Server. The new class emerged during the human-guided discovery cascade rather than fully autonomous operation, leading the researchers to conclude that an experienced researcher acts as an amplifier for an AI research system.
What is NatJack and does it affect my network?
NatJack is a class of attacks disclosed at Black Hat USA 2026 by Malcolm Stagg of the Synack Red Team, manipulating NAT connection tracking state to hijack TCP sessions, poison DNS responses, identify ports assigned to other connections, and exhaust NAT tables. Two CVEs cover implementation flaws in Windows NAT in Hyper-V and in Linux netfilter conntrack. Available fixes raise the difficulty of exploitation but do not close the underlying design gap, so the recommended controls are encryption in transit, segmentation of untrusted workloads and IP Source Guard.
How should a CISO prioritize conference research?
Separate the findings that justify a cheap, reversible control change regardless of likelihood from the ones that need a judgment about how widely a technique is actually used. The first group can be actioned straight from the published record. The second needs an operator's read on whether a technique has moved from specialist to commodity, which disclosure documents are not written to provide.
Is the Turbo Talks session free, and what if I cannot attend live?
The session is free. One registration covers all seven sessions in the series, and registered attendees are sent the recording if they miss a session.
Evidence
Every factual claim above resolves to one of the sources below. Source type is given so you can weigh each one: a researcher's own paper or a vendor's own disclosure is stronger than a conference recap, and we have said where a figure appears only in one place. All accessed 1 October 2026.
1. Microsoft Security Blog, "Microsoft at Black Hat USA 2026." Vendor publication. Confirms the keynote title, speaker and date. microsoft.com
2. CSO Online, reporting on the Weston keynote. Journalist report carrying a direct quote. Source of the 182 of 200, $3.61 and 21-minute figures. Note: these figures were spoken on stage and are not published in any Microsoft document we could find. Attribute them to the keynote, not to Microsoft research. csoonline.com
3. Zenity Labs, "AI Total" research page. Researchers' own publication. Source for the 30 percent malware-dropper finding, the reserved package names, and "within hours" removal. zenity.io/research/ai-total
4. Zenity Labs press release via Businesswire, 6 August 2026. Company-issued disclosure. Source for the 1.7 million aggregate installs, the named typosquatted services and the credential list. Note: the 1.7M figure appears in the press release; Zenity's own research pages describe the scale differently, so we have attributed it rather than stated it flatly. businesswire.com
5. Zenity, Black Hat USA 2026 recap. Vendor publication. Confirms the talk title and both presenters. zenity.io/blog
6. Cyera Research, "Sandcastles, Not Sandboxes: How One Architectural Flaw Exposed Seven Products," Tokarev and Pearl, 7 August 2026. Researchers' own white paper, published alongside the DEF CON 34 talk. Primary source for the seven products, all four CVEs and scores, the ctypes and Emscripten mechanism, and the blast radius by host environment. cyera.com
7. Cyera Research, "Cellbreak: Grist's Pyodide Sandbox Escape." Researchers' own write-up. Per-product detail for Grist. cyera.com
8. NVD record for CVE-2026-24002. Authoritative vulnerability database. Source of the 9.6 score that differs from the researchers' 9.1. We have given both. nvd.nist.gov
9. PortSwigger Research, "Can AI Do Novel Security Research? Meet the HTTP Terminator," James Kettle. Researcher's own full write-up. Primary source for the method, all figures, the shared-parser confusion mechanism and the multipart/byteranges example, the affected products and CVE-2026-63078. portswigger.net/research
10. PortSwigger, HTTP Terminator executive summary (PDF). Researcher's own paper. Source for the discovery-cascade finding and the four-stage method. portswigger.net (PDF)
11. Synack press release, "NatJack research: NAT implementation weaknesses." Researcher's employer, company-issued. Primary source for the four techniques, both CVEs, the patch status quote and the recommended controls. synack.com
12. Synack NatJack research page. Researcher's employer. Supports the shared-NAT-table framing. go.synack.com
13. Patrick Wardle, "Dylib Hijacking on macOS: Dead or Alive?" conference slides. Researcher's own slides. Confirms the talk and macOS 26. speakerdeck.com
14. Objective-See blog, Patrick Wardle. Researcher's own blog. Confirms the technique still works on macOS 26 and covers detection. objective-see.org
15. Virus Bulletin, "Dylib hijacking on OS X," Patrick Wardle. Peer-reviewed publication by the same researcher. Primary source for the LC_LOAD_WEAK_DYLIB and @rpath mechanism. virusbulletin.com
16. DEF CON 34 workshop listing, "WS7: Explore the Windows Instrumentation Callback," Yoann Dequeker. Official conference listing. Confirms this was a four-hour workshop, not a stage talk, and the "existing since Windows 7" description. events.humanitix.com
17. Alex Ionescu, "Hooking Nirvana," Recon 2015, slides and demo code. Original researcher's own material. Primary source for the technique's 2015 publication. github.com/ionescu007
18. cirosec, "Windows Instrumentation Callbacks." Technical write-up. Source for the NtSetInformationProcess registration and the return-to-user-mode timing. cirosec.de
19. Social Engineering Community (DEF CON village), session listing for Dr Megan Squire. Official village listing. Confirms the talk title, her F-Secure role, the three technique names and the primed-target framing. se.community
20. Google blog, "The evolving role of the red team in the era of agentic security," Daniel Fabian, Head of Google Red Teams, 13 August 2026. Named author, company publication. Note: this is a position piece published during Black Hat week, not a conference talk. An earlier draft of this note described it as a DEF CON 34 session; we could find no conference listing supporting that and have corrected it. blog.google
21. CISO Platform session page for the Turbo Talks session referenced in this note. Event listing. Source for the speaker background and published agenda. cisoplatform.com
Disclosure. The session referenced in this note is part of Turbo Talks. Technology Partner: FireCompass, who produce and deliver every session in the series. CISO Platform hosts it for the community. The research summarized above is independent of both.
Vendor and product names appear only where they are factually part of the disclosed research, and nothing here is an endorsement. Where a figure is reported in one place only, or where sources disagree, the Evidence list says so. The open questions are genuinely open at the time of writing, and we will update this page after the session with what was said about them.

Comments