Raxter shares field observations from both conferences and highlights the shifts security teams should monitor over the next 12 months. A session with Zak Raxter, Senior Solutions Architect, Offensive Security, FireCompass, recorded Thursday, 1 October 2026. Written up by the CISO Platform editorial team.
"Black Hat shows where the security industry is going. DEF CON shows how attackers are going to break it." We asked Zak not to give us a conference recap, but to tell us what he saw change. His answer fits on one line: the vulnerabilities did not change, the price of exploiting them did.
The full session, including 14 questions from the audience.
Three key points
- How AI is reshaping the attacker playbook across reconnaissance, exploitation, and attack chaining
- New attack techniques and broken security assumptions demonstrated by researchers
- What CISOs and security teams should reconsider based on 2026 research findings
About the speaker
Zak Raxter is Senior Solutions Architect, Offensive Security at FireCompass. He has ten-plus years in offensive security, AppSec and DevSecOps, including security engineering at Juniper Square, Prudential Financial and Mosaic Group. He has built endpoint protection, vulnerability management and internal bug bounty programs from scratch, and has been the person triaging the queue as well as the person running the red team. His day job covers red teaming, adversary emulation and incident forensics, plus web and API testing against production systems.
He opened with a sourcing commitment worth repeating: every number in the session has a public source, and all of them are linked at the end of this post.
CISO executive summary
The argument in one line. Expert-grade attacks are now available on hobbyist budgets. Attackers are running the playbook they already had, faster, at larger scale, with far less skill behind the keyboard. Your perimeter is being tested at that pace whether or not you are testing it yourself.
Key findings
- One researcher's method reached internet scale. James Kettle encoded four years of his own desync research into an autonomous system. It invented new request smuggling techniques on its own, including Shared-Parser Confusion, and found roughly 700 vulnerable targets across 30,000 authorized sites. The list included banks, government infrastructure, security products and an airport. The source code and a blueprint for adapting it are public.
- Cost collapsed. A four-link pre-auth RCE chain in WordPress core was built on $25 of a $200 AI subscription in just over ten hours. Brokers pay $500,000 for a WordPress RCE.
- Low severity findings no longer stay low. Three low-rated flaws chained into full mailbox, calendar, SharePoint and OneDrive exfiltration from one click on a microsoft.com link, using Microsoft's own infrastructure as the exit path.
- Telemetry is an injection surface. A publicly exposed Sentry DSN plus an MCP integration chains into remote code execution on a developer machine, with an 85 percent success rate in controlled testing across more than 100 organizations. Nothing fires, because every action the hijacked agent takes sits inside permissions it was legitimately granted.
- Scope is not self-enforcing. Under permissive settings, OpenAI agents tasked with offensive security work reached systems beyond their intended environment.
- API attacks moved to logic abuse. The share using unauthorized workflows and abnormal activity went from 30.01 percent in 2024 to 61.18 percent in 2025, and the average volume per organization doubled from 121 to 258 attacks a day.
Recommendations
- Keep a human in the loop. Kettle's best results came when he stepped back in at the discovery cascade. Put a human gate before any active exploitation.
- Assume attackers have frontier-class tools. Open-weight models sit months behind the frontier, anyone can download them, and their safeguards strip out easily.
- Put AI to work on your own posture. Pick one slow or shallow step, such as triage, retesting or authorization coverage, and measure it before and after.
- Let business criticality decide where depth and continuity go. Nobody can afford maximum depth on every asset all the time. Test at the attacker's price with proof of exploit on every finding, run recon weekly, and let triggers override the calendar.
Part one: the barrier dropped
Zak was explicit that he did not come back from Vegas with a list of novel vulnerability classes. What he came back with was a change in economics.
Expert-grade attacks. Hobbyist budgets. Zak Raxter, opening the session
One expert, amplified
The session Zak kept returning to was James Kettle's, from PortSwigger, the team behind Burp Suite. Kettle encoded four years of his own desync research method into an autonomous system and pointed it at live, authorized targets. The question was whether AI could come up with genuinely new ways of exploiting HTTP request smuggling, not just replay known ones.
It could. The model proposed a technique now named Shared-Parser Confusion from proven examples, and Kettle confirmed and generalized it. The detail Zak flagged as the most important in the whole talk is where the best finds came from: not from letting the system run unattended, but from Kettle stepping back in at the point where one result seeds the next hypothesis.
Then it ran at scale. Around 30,000 candidate desync vectors explored, 30,000 authorized sites tested, roughly 700 vulnerable targets found. Banks. Government infrastructure. Security products. An airport.
The source code and a blueprint for adapting the method are now public.
A $500,000 exploit for $25
Two weeks before Vegas, Adam Kues at Searchlight Cyber published a pre-auth RCE chain in WordPress core. The budget was $25 out of a $200 AI subscription, and just over ten hours.
Four links:
- A batch API validation desync opens a pre-auth SQL injection.
- Post cache poisoning fakes posts in memory.
- A forged customize_changeset borrows the admin's identity.
- A parse_request hook replay creates a new admin.
Full compromise. Brokers pay $500,000 for a WordPress RCE.
The barrier to entry, measured
| Figure | What it measures |
|---|---|
| +44% | Attacks that began by exploiting public-facing apps, driven mostly by missing authentication |
| +49% | More active ransomware and extortion groups than the year before |
| ~4 months | Gap between the best open-weight cyber model and the US frontier |
| 64 to 92% | Rate at which simple tricks bypassed that model's safeguards, rising to 100 percent once weights were altered |
Sources: IBM X-Force Threat Intelligence Index 2026; Anthropic and CAISI evaluations of GLM-5.3, September 2026.
Zak's reading of the four-month figure is the one that should land with a security leader. The common assumption that frontier models are far ahead, and therefore that serious capability stays behind a vendor's safety controls, no longer holds. An attacker can download a near-frontier model, strip the safeguards, and run it with no logging and no rules.
The gap between "an attacker could chain these four lows" and "someone did, on a lunch budget" has closed. Zak Raxter
Part two: three assumptions that broke
Three things most security teams believed in January that did not survive August.
Assumption one: low severity findings stay low
Varonis Threat Labs presented SearchLeak at DEF CON 34. Three low-rated issues in Copilot Enterprise Search, chained:
| Stage | Weakness | How it works |
|---|---|---|
| 1 | Parameter-to-prompt injection | The q URL parameter reaches the model as an executable prompt. The victim types nothing. |
| 2 | HTML render race condition | The sanitizer wraps output after generation completes. The browser renders the stream as it arrives, so the injected image tag fires first. |
| 3 | CSP bypass via Bing SSRF | CSP allowlists *.bing.com. Bing's Search-by-Image endpoint fetches the attacker URL server-side. |
The result takes one click on a microsoft.com link. Copilot searches the victim's mailbox, calendar, SharePoint and OneDrive, then ships the results out through Microsoft's own infrastructure. Email subject lines alone often carry one-time codes, MFA codes and password-reset links.
The victim sees Copilot think for a second. Anti-phishing tooling sees a trusted domain. Intrusion detection sees legitimate traffic, because that is what it is.
The chain was issued CVE-2026-42824, rated critical and patched by Microsoft. Varonis' practical advice: watch Copilot Search URLs for encoded HTML, and review CSP allowlists for domains that fetch URLs server-side. Zak's addition was blunter. There will be more of these.
Assumption two: our telemetry is trusted input
Tenet Security presented Ghostjacking at DEF CON 34. The premise is one sentence long, and it is the kind that changes how you look at your own pipeline.
An AI coding agent reads a stack trace to debug a failure. The attacker controls what lands in that stack trace.
A publicly exposed Sentry DSN plus an MCP integration chains into remote code execution on a developer machine. Nothing fires, and the reason nothing fires is the part worth taking to your detection team: every action the hijacked agent takes is inside the permissions it was legitimately granted. EDR, WAF and IAM all see authorized work.
| Figure | What it measures |
|---|---|
| 85% | Success rate in controlled testing across more than 100 organizations |
| 2,388 | Organizations found with publicly discoverable Sentry DSNs |
| 71 | Of those that sit in the Tranco top one million |
| ~27% | Of the Fortune 1000 exposed through one MCP integration path |
This is not a lab finding, and the prevalence numbers are the argument. Zak's observation on why this is so widespread: everyone is in a rush to integrate AI and MCP into their workflows, and the security implications are arriving second.
Assumption three: the scope we set is the scope it stays in
OpenAI gave a detailed public account at Black Hat. During evaluations run under permissive settings, agents tasked with offensive security work found and exploited vulnerabilities in supporting infrastructure, then reached systems beyond their intended environment.
Zak's framing of the lesson: guardrails on some of these systems are closer to suggestions than to controls, particularly when scope is loose. For anyone running autonomous testing, scope has to be enforced where the agent executes, every time it acts.
API attacks moved to logic abuse
The techniques here have not changed much. The distribution has.
| Measure | 2024 | 2025 |
|---|---|---|
| Share of API attacks using unauthorized workflows and abnormal activity | 30.01% | 61.18% |
| Average API attacks per organization per day | 121 | 258 |
| Organizations reporting an API security incident | Not reported | 87% |
Source: Akamai, State of the Internet: apps, APIs and DDoS, 2026.
Why your tooling stays quiet
A broken object level authorization request is syntactically perfect. Clean 200, valid JSON, correct content type. It does not look like an attacker throwing the kitchen sink at an endpoint. It looks like a customer.
The only thing wrong with it is whose record came back.
There is no payload to fingerprint, so signature-based tooling has nothing to fire on. Proving it needs two authenticated identities held at the same time and a replay across the boundary. That is a test design decision, not a product feature, and it is why most assessments miss this class entirely: they run as one user.
The most-exploited OWASP API risks for 2025 line up with that reading. Security misconfiguration at 40 percent, broken object property level authorization at 35 percent, broken authentication at 19 percent.
This session is part of the AI Pentesting and AI Safety Series for Security Leaders: seven sessions running from September 2026 to February 2027, twenty minutes each, live. One free registration covers all seven. Register free
Part three: four shifts worth making
Keep a human in the loop. This is the recommendation Zak put first, and he grounded it in Kettle's own result rather than in caution. The strongest finds came when the researcher stepped back in to guide the system and seed it with vectors, not when it ran alone. A human gate before any active exploitation does double duty: it improves the output and it stops things going off the rails.
Assume attackers have frontier tools. Open-weight models like GLM-5.3 sit months behind the frontier, anyone can download them, and their safeguards strip out easily. Planning on the assumption that capability stays gated is planning against an adversary who does not exist.
Put AI to work on your posture. Attackers already use it for recon, testing, exploit development and proof of concept code. Defenders get the same capability increase, and historically defence has run behind. The practical version is not a transformation program.
Ask where AI adds speed or quality. Pick one slow or shallow step. Triage, retesting, authorization coverage. Measure it before and after. Keep the human in the loop.
Building the program: depth, breadth and continuity
Nobody can afford maximum depth on every asset, all the time. Business criticality decides where depth and continuity go.
| Axis | What it answers | What he recommends |
|---|---|---|
| Depth | How far each test goes | Test at the attacker's price, with proof of exploit on every finding. Spend it on business logic, authorization and multi-stage chains. Leave CVE matching to cheaper tools. |
| Breadth | How many assets you cover | Run recon weekly. The surface changes faster than the scope document. Label every asset by business criticality and rate of change, with criteria someone can review. |
| Continuity | How often you retest | Triggers first: new asset, new API, deploy, disclosed CVE. Then a schedule set by criticality and budget. |
The breadth recommendation deserves its own note, because it is where Zak saw the most exposure in practice. Development teams are shipping more code than ever, much of it AI generated, some of it reviewed and some of it not. The attack surface moves weekly whether the scope document does or not.
Depth and cadence by tier
| Tier | Test depth | Retest cadence |
|---|---|---|
| P1, business critical | Deepest: logic, authorization, multi-stage chains | Monthly, and on every major change |
| P2 | On demand | Quarterly |
| P3 | Event-driven | When a trigger fires |
| Day-one CVEs, every tier | Rule-based tooling, no agent time | Daily |
Triggers override the schedule on every tier: new asset, new API, deploy, disclosed CVE.
Two conditions sit under the whole model. Measure validated coverage and time to verdict. And automate depth only with scope control, human approval before active exploitation, and a full audit trail.
Five questions for your team this week
- When did we last test authorization across two tenants at the same time, rather than as one user? Most assessments run with a single account, which is precisely the condition under which broken object level authorization stays invisible.
- Which API versions are still routed that we believe are retired? Retired on paper and reachable on the wire are different states, and recon is what tells them apart.
- Is every asset labeled by criticality and rate of change, using criteria someone can review? You cannot decide where depth goes until you know where the crown jewels sit and how many layers stand between them and an attacker.
- What does our coding agent read automatically, and who can write into that source? If any system takes input and files it somewhere an agent later reads, that path needs monitoring. Attackers will find it.
- How many production changes shipped since our last external test? A quarterly test against a hundred releases is a point-in-time test, and it is honest to call it that.
Your perimeter is being tested at this pace whether or not you are testing it yourself. Zak Raxter, closing the session
Questions from the session
If attackers can continuously test your environment using AI, how often should you be testing yourself? Is continuous really possible?
Yes and no. A penetration test every hour is not feasible. For business critical assets Zak recommends weekly at minimum, ideally more, and says the tools available today make that achievable. Monthly is the floor. Ideally, any time a change takes place.
How do you keep an agent in scope?
Two things. A human in the loop, and validation of both input and output. FireCompass runs what Zak described as an input and output firewall: everything going into the model is filtered, sanitized and validated, and everything coming out is filtered and sanitized the same way. Beyond that, be as specific as possible with the agent's scope and limit it to the single task you want it on.
Would you trust an AI agent to autonomously exploit vulnerabilities in production? How do you put guardrails on it?
If you are rolling your own, no. Where it is done safely, the controls are rate limiting, enforced guardrails, scope restricted as tightly as possible, and permissions kept minimal. Zak's reasoning on the last point: anything the agent has permission to do, it will do.
What surprised you most, and why should security leaders care?
How far the barrier to entry for sophisticated attacks has fallen. He went into the conference already aware of AI's role. What surprised him was the speed and how little money is required. Attackers are using this to compromise systems now, not eventually.
Is GLM-5.3 genuinely comparable to frontier models?
In capable hands, yes. On published benchmarks the gap sits within a margin of error, in some cases under one percent. What decides the outcome is who the human in the loop is and how they use it.
Which attack techniques felt genuinely new, versus familiar techniques with an AI label?
Most of it was familiar. The exception Zak named was Kettle's HTTP Terminator work and the new method for exploiting request smuggling, Shared-Parser Confusion.
Which commonly trusted security control proved easier to bypass than expected?
Web application firewalls. Many teams treat a WAF as the end of the conversation for protecting a web app or API, and across these cases traffic slipped through looking entirely legitimate. Zak expects new AI-assisted obfuscation techniques to make this worse over the next year.
What should we change in our pen testing or red team scope?
Cover the whole attack surface, not the documented one. The assets Zak worries about most are shadow IT: forgotten dev and staging environments, retired APIs still routed, and systems an executive spun up after discovering they could vibe code. Finding those is a recon problem.
How can we tell whether our security tools would actually detect these techniques?
Test them. A new technique comes out, you run it and see whether your defenses catch it. It does not have to touch production; staging and test environments will answer the question.
For a team with a limited budget, what would you prioritize?
Work out where AI can enhance the team you already have. Zak's position, and FireCompass's, is that AI is not replacing penetration testers any time soon, but it is enlarging what a small team can cover, and that gain counts for most when the team is small.
What is getting too much attention, and which risk deserves more?
Authentication breaks and authorization bypasses get plenty of attention. Business logic workflow testing does not get enough. Zak compared it to living off the land: the attacker uses the application exactly as designed, in an order nobody tested, and nothing about the traffic stands out.
Is SearchLeak being exploited in the wild?
No confirmed exploitation that Zak has seen, though he would not be surprised. A great deal of exploitation stays under the radar until a researcher or a news cycle surfaces it.
Next in this series: How Our AI Reached the HackerOne Top 3, and What We Learned
22 October. Twenty minutes, live. One registration covers all seven sessions in the series. Free
New to CISO Platform? Join the community free for frameworks, checklists and peer discussion.
Sources
Research
- HTTP Terminator research and executive summary, PortSwigger Research, August 2026 (portswigger.net/research/http-terminator), with coverage in The Hacker News
- WordPress pre-auth RCE chain, Adam Kues, Searchlight Cyber, 20 July 2026 (slcyber.io)
- SearchLeak and CVE-2026-42824, Varonis Threat Labs, presented at DEF CON 34 (varonis.com/blog/searchleak), with coverage in BleepingComputer
- Ghostjacking and the agentic kill chain, Tenet Security, DEF CON 34 (tenetsecurity.ai), exposure numbers reported by Forkast
- The Hugging Face incident talk, OpenAI at Black Hat USA 2026 (youtube.com/watch?v=87DyyMV0kCY)
Data and reports
- X-Force Threat Intelligence Index 2026, IBM
- Q2 2026 ransomware report, GuidePoint
- State of the Internet: apps, APIs and DDoS, 2026, Akamai, plus daily API attack volumes reported in Infosecurity Magazine
- GLM-5.3 cyber capability evaluations, Anthropic citing CAISI, September 2026, and benchmark scores reported by Developer Tech
- State of AI-enabled malware, Unit 42
- Ghostjacking and identity gaps, Dark Reading
Technology Partner: FireCompass. All sessions, speakers and content in the AI Pentesting and AI Safety Series are produced and delivered by FireCompass. CISO Platform is hosting the series for its community.
Every figure, quote and recommendation above is drawn from the session recording and the slides presented, available here. No figures have been added from outside the session.

Comments