For two years we told ourselves the AI arms race would be won by whoever had the best model. The most capable frontier system, the biggest context window, the cleverest agent. Build the smartest brain and you win.
A Russian-speaking actor who goes by Trim just proved the opposite, from the other side.
He does not have a better model than you. He has the same public models you have. He jailbroke them, wrapped them in a suite of off-the-shelf offensive tools, and turned the result into a product he sells to other criminals. Cato Networks' research arm caught it and Dark Reading reported it this week. Trim first published six detailed techniques to bypass the safety filters on Claude Opus. Then he shipped the platform: "AI Pentest Checker", which runs reconnaissance, validates vulnerabilities, escalates them, and writes the exploitation report. According to the reporting it produces a finished report in about ten minutes, and he was handing out free keys to a first cohort of beta testers.
Read that again. Recon, validation, escalation, reporting. That is not a tool. That is a junior red team, packaged, on subscription. A one-person crew shipped it.
The edge is not the model. It is who operationalises it first.
Trim's models are not special. Anyone can rent the same ones. What he built is the wiring: the jailbreak that keeps the model compliant, the glue to the scanners, the report at the end. The capability was always sitting there in public. He was just first to turn it into an operation.
That is the whole shift, and most security roadmaps have not absorbed it. The asymmetry in this industry used to come from discovery. Who found the zero-day first, who had the exploit nobody else had. Now the exploit is cheap to generate and the scarce thing is operational reach. The attacker no longer needs a team of specialists. He rents one by the hour, for a few dollars a run.
Trim is not an outlier. He is the shape of the market.
If he were a one-off I would not be writing this. The lineage is what matters.
It started with WormGPT in 2023, a single uncensored model sold to write business email compromise. Then FraudGPT, a subscription bot for phishing and carding. Then GhostGPT in 2025, an uncensored assistant on Telegram writing malware and exploit code with a no-logs promise. And by mid-2025 Cato was documenting Nytheon AI, no longer one jailbroken model but a whole platform of uncensored ones. Trim's "AI Pentest Checker" is the next rung: not a chatbot that helps you write an attack, an autonomous pipeline that runs one.
And this is not confined to the criminal underground. The model makers are publishing the same finding about their own systems. Anthropic disclosed a data-extortion campaign where a single operator used Claude Code to run most of an intrusion, reconnaissance through exfiltration, across seventeen organisations in one month. Months later it disclosed a Chinese state-linked group using Claude Code as an autonomous penetration tester. OpenAI's October report describes disrupting state operators and industrialised social-engineering rings. Google's M-Trends 2026 documents malware, PROMPTFLUX and PROMPTSTEAL, that queries an LLM mid-execution to rewrite its own code and slip past signatures. Different actors, same move: hand the operational work to a model.
The number that should worry a board is not the model's benchmark score. It is the labour the model deletes from the attacker's side. One person now does what used to take a crew.
Why "we need safer models" is the wrong answer
Walk into most AI-security conversations right now and the reflex response is model-shaped. Better guardrails. Stronger alignment. A safer frontier model. Wait for the vendor to close the jailbreak.
That answer cannot defend your enterprise, because it controls the wrong side of the transaction. Guardrails are a supply-side lever the model vendor holds. They matter to Anthropic, OpenAI and Google, and those companies are right to invest in them. But Trim already showed the limit. He published six ways around Claude's filters, and when a hosted model is inconvenient he reaches for a local uncensored one that has no filters to bypass. You cannot defend your business with someone else's alignment policy. The guardrail is not yours to hold.
This is the same mistake I wrote about when everyone was auditing AI prompts while attackers walked in through over-privileged identities. We keep pointing defensive attention at the model because the model is the thing in the headline. The model is not where the attack lands.
What actually holds
Here is the part the AI-panic coverage keeps missing. The model writes the plan. The attack still has to happen in your environment.
Trim's platform can reason about a target brilliantly, but the recon still touches your perimeter. The exploit still hits a real service. The escalation still creates a real session. The exfiltration still moves real bytes to a real place. An LLM makes the author faster and cheaper. It does not make the behaviour invisible, because behaviour is not text the model gets to rewrite. It is what happens on the wire and on the host.
So anchor detection there. On the infrastructure the actor reuses, the sequence of actions a real intrusion produces, the way a tool touches a system regardless of who or what wrote it. That is the one layer an attacker cannot pad, jailbreak, or prompt away. It is also exactly the argument I made about yellow teams: a few dozen organisations will build their own AI attack frameworks, and everyone else is saved not by owning a team colour but by detection that sees the behaviour those frameworks produce.
We do this at Sekoia because it is what a threat-detection-and-research team is for: it tracks the actor, not the tool. When our TDR pulled apart APT28's shift from X-Agent to LLM-driven malware, the one that hands its commands to a language model at runtime, the detection did not hinge on the model it called. It hinged on how the malware reached out, what it touched, the pattern only that actor produces. Swap the LLM and the detection still holds. That is detection capital: memory of how these actors move, turned into something a SOC can act on at machine speed. It is why "AI Pentest Checker v2" does not reset your defence to zero.
Trim rented a brain. The lesson most people will take is that we need a smarter brain to rent back, an autonomous defender to fight his autonomous attacker. That is the arms race he is counting on us to join, because he can rent capability as fast as we can. The defence that does not depend on out-modelling him is the one built on what he cannot change: how his attack behaves when it reaches you. Build that, or buy it from people who hunt him. The brain was never the edge.


