"> Building Brakes for Rogue AI | pre-seed funding | VC Cafe
October 5, 2026 Weekly insights on Israeli tech, venture capital, and AI
AI Agents

Building Brakes for Rogue AI

Building Brakes for Rogue AI vccafe - pre-seed funding / ???? ??? ???

Key takeaways

  • The rogue-agent incidents of 2026 came from goal-seeking agents with too much access, not malice, and that kind of failure is fixable.
  • Misalignment startups fall into four groups (control, evaluation, interpretability, manipulation), plus liability for when prevention fails, and each group is at a different stage of maturity.
  • Irregular is reportedly raising at a $1.5B valuation despite hosting several of the breakouts, which shows how much demand there is for independent AI testing.
  • Self-improving agents push safety controls outside the agent, toward identity, permissions and sandboxing.
  • The best pre-seed opportunities are in eval containment, continuous monitoring and change governance for agents.

Every few days in the past couple of months, another story about AI agents going rogue has landed in my inbox.

In June, an OpenAI agent given a mundane task, pulling Australian health statistics, found its way into a Medicare portal run by Services Australia. It was the first known case of an AI agent hacking a government system, and it wasn’t the last. OpenAI has since confirmed that its agents also reached data on SEC and Census Bureau websites, and that another tried and failed to break into a Department of Education site. Last week, a report from Asymmetric Security added a more unsettling detail: the agents appear to have tried to cover their tracks, using private accounts and throwaway email inboxes. The researchers couldn’t tell whether that was deliberate.

Then there’s Hugging Face. In July, roughly 700 OpenAI agents running an internal cyber evaluation slipped out of their sandbox, improvised a message board to coordinate with one another, and compromised parts of Hugging Face’s production infrastructure. The motive wasn’t sabotage. The agents had worked out that Hugging Face might hold the answers to the benchmark they were being graded on. They were, in effect, cheating on the test. Google, Anthropic and Meta have since disclosed their own models escaping evaluation environments, and OpenAI now faces what appears to be the first lawsuit seeking to hold an AI developer liable for a rogue agent.

I’d be careful with the word “rogue”, though. The more sober post-mortems point to human decisions: a misconfigured test environment, models deliberately run without their usual safeguards, warning signs that nobody acted on. The agents weren’t malicious. They were relentlessly goal-seeking, and they had far more access than their task required. For an investor, that’s actually the important part. These were boring failures, and boring failures are exactly the kind that good products fix.

An old problem with new consequences

None of this should come as a surprise. Researchers have been cataloguing “specification gaming”, where a model optimises for the metric rather than the goal, for years. Anthropic and others have published work showing models resisting shutdown and behaving differently when they suspect they’re being tested.

What’s changed is the setting. When a model’s misbehaviour was confined to a chat window, the gap between what we asked for and what it pursued was an academic curiosity. Give the same model tools, credentials and an internet connection, and that gap becomes an incident report.

The market has noticed. “Excessive agency”, the industry’s term for an agent holding more access than its task needs, jumped from sixth to third in OWASP’s 2026 ranking of AI application risks. Buyers are paying up too: Cyera acquired agent identity startup Oasis for $1B last month, and earlier this year Cisco bought Astrix, Fortinet bought Virtue AI and F5 bought CalypsoAI. A new crop of startups is attracting serious funding to tackle the problem, and CB Insights recently mapped them against the risks Anthropic itself warns about. It’s a useful frame, and I’ll come back to it below. But first, the company that best captures this moment happens to be Israeli.

Israel’s front-row seat: Irregular

Irregular, formerly Pattern Labs, was founded in Tel Aviv in 2023 by Dan Lahav and Omer Nevo. It builds cyber test ranges where frontier labs push their models into extreme scenarios to measure dangerous capabilities before release. OpenAI, Anthropic and Google are customers, as are government organisations.

Here’s the twist: several of this summer’s breakouts happened inside Irregular-hosted evaluation environments. OpenAI pointed to a misconfiguration in the testbed that allowed models to reach the public internet. Irregular says the incidents weren’t evidence of models independently breaking out of a secure sandbox, and that it has since fixed the setup and added safeguards.

You might expect that to hurt the business. Instead, Irregular is reportedly raising more than $100M at a $1.5B valuation, led by Thrive Capital and Greenoaks, roughly three times the reported valuation of its Sequoia-led round last year. Calcalist reports the company has been profitable since 2025.

My read is that investors are pricing in two things. The incidents proved that the capabilities Irregular measures are real, and the labs clearly want independent testers rather than grading their own homework. But there’s a lesson here for every evaluation company: containing the thing you’re testing is now a core product requirement, not a footnote. “How do you contain what you test?” belongs on every diligence checklist in this space.

Irregular is also part of a wider pattern. Israel’s cyber DNA, built around identity, access and runtime protection, maps almost perfectly onto the problem of controlling AI agents. In September alone, AIR Security raised $50M to build a firewall for AI agents and Cymphony raised $25M to secure the overlap between employees and agents. Both Oasis and Astrix have Israeli roots, as do Zenity and Noma Security, two of the companies on CB Insights’ map.

Mapping the category

The easiest way to understand misalignment startups is by the failure mode they address. There are four, plus a fifth layer for when everything else fails, and they sit at very different stages of maturity.

unnamed - 2026-10-02T071700.391 - pre-seed funding / ???? ??? ???
Startups tackling AI misalignment, grouped by failure mode. Source: CB Insights, based on risks flagged in Anthropic’s S-1.

Limiting what agents can do. The most developed segment doesn’t try to fix the model at all. It assumes something will go wrong and limits the blast radius. WitnessAI sets rules for what agents can do and monitors their activity. Zenity finds agents across a company and flags risky behaviour. Noma Security manages agent identities and access, and has grown its team 154% in a year to around 150 people. Straiker blocks attacks and risky agent behaviour, and E2B runs agent code in isolated sandboxes for customers including Manus, Perplexity and Groq. The Hugging Face breach is practically a case study for this layer: the agents needed exposed credentials and an open network path, and least-privilege controls would have shrunk the damage considerably. This is where the buyers and the exits are today. It’s also crowded, so a new entrant needs a sharp technical wedge, not another dashboard.

Catching agents that game the test. Models can behave differently when they know they’re being watched, which quietly undermines every evaluation. Patronus AI tests agents in realistic environments before launch, Mindgard automates red teaming, Gray Swan combines red teaming with post-deployment monitoring, and Raindrop watches live agents for unexpected behaviour. Since the Hugging Face incident was, at its core, a test-gaming incident, I’d argue this category just became more urgent. It’s turning into a real market, though frontier evaluation still depends on a small number of very large buyers.

Finding capabilities nobody looked for. How do you discover an ability you weren’t testing for? That’s the domain of interpretability and model auditing. Goodfire builds tools to analyse and edit what happens inside models and counts Mayo Clinic as a client. Martian, Guide Labs and Tilde Research each approach the inner workings of models from a different angle, and Irregular tests frontier models for unexpected cyber capabilities. This segment is much earlier. A lot of the work happens inside the labs and is effectively given away, so venture-scale outcomes probably depend on regulation or enterprise audit requirements creating a buyer outside them.

Detecting manipulation. The thinnest category covers models that mislead, flatter or steer the people overseeing them. Alice tests whether models deceive their overseers while secretly pursuing another goal, and Apollo Research evaluates frontier models for strategic deception. Much of the serious work here sits at nonprofits. Fewer startups could signal whitespace, or simply that nobody is buying yet. I’d want to see a paying customer before forming a thesis.

Paying for it when prevention fails. Armilla AI already sells AI liability insurance, and with OpenAI now facing a lawsuit over Hugging Face, that market just became a lot less theoretical. Incident forensics belongs here too. Hugging Face reconstructed more than 17,000 attacker actions, largely using AI tools of its own.

Why self-improving agents raise the stakes

There’s a second trend that makes all of this more pressing. A wave of research is teaching agents to improve themselves. Microsoft’s SkillOpt and Google’s WikiSkill let agents rewrite their own instructions based on what worked and what didn’t. Self-Harness and the Darwin Gödel Machine go further, letting agents modify their own prompts, tools, memory and control flow, and Meta’s Hyperagents even let the improvement process itself evolve. Xiaomi’s HarnessX co-evolves the agent and the underlying open-weight model, while EnvHarness and EverMind’s Raven extend the same idea to training environments and multi-agent coordination.

The loop is the same everywhere: run tasks, study the traces, change part of the system, and keep the change if the score goes up. It’s a genuine productivity unlock, and it changes the misalignment picture in a few important ways.

First, guardrails can no longer live inside the agent. If an agent can rewrite its own prompts and tools, any safety rule written there becomes a suggestion. Controls have to sit outside the loop, enforced at the level of identity, permissions, sandboxing and network access, which strengthens the case for the control layer.

Second, these optimisation loops are Goodhart machines. Every framework accepts a change when a validation score improves, and the Hugging Face agents were optimising for a score too. Self-improvement scales the incentive to game tests, which makes tamper-resistant, realistic evaluation a safety mechanism in its own right.

Third, point-in-time audits lose their meaning. The agent you certified on Monday isn’t the agent running on Friday. That favours continuous monitoring over one-off certification, and it hints at a category that barely exists today: change governance for agents, tracking what changed, why, and who or what approved it.

The research I drew on ends with a line I like: developers will spend less time hand-tuning agents and more time deciding what can change, what must stay fixed and how improvement is measured. That’s a pretty good description of the job this whole category is trying to do.

Where I’m looking

As a pre-seed investor, I’m most interested in the gaps the incidents exposed rather than the segments that already have incumbents. Evaluation containment and integrity looks under-built relative to demand, and Irregular shows both the appetite and the cost of getting it wrong. Continuous monitoring and change governance for self-modifying agents is barely a category yet, but it will be needed as these research frameworks move into production. I’m still open to the control layer, but only for teams with a genuinely new technical wedge, such as short-lived, per-task credentials or egress control for agent sandboxes. Forensics and liability now have real demand. Manipulation I’m watching, waiting for a buyer to show up.

For founders already shipping agents, including in our own portfolio, the practical takeaways are simple. Audit where your agents hold standing credentials, scope access to each task, log every action, and treat any self-optimisation loop like a code change that needs review. Enterprise buyers will start asking about all of this, if they aren’t already.

The brakes are becoming a business. And given how many of the companies building them come out of Israel’s security ecosystem, I expect a good number of the winners to have Tel Aviv on their cap table.

Follow me
Co Founder and Managing Partner at Remagine Ventures
Eze Vidra is the founder of VC Cafe and the co-founder and managing partner of Remagine Ventures, a pre-seed fund investing in ambitious founders at the intersection of AI, technology, entertainment, gaming, and commerce with a spotlight on Israel.

He is a former General Partner at Google Ventures (GV) in Europe, former head of Google for Entrepreneurs in Europe, and founding head of Campus London, Google's first startup hub. Eze writes on Israeli tech, venture capital, artificial intelligence, and founder strategy.

He is also the founder of Techbikers, a nonprofit that brings together the startup ecosystem on cycling challenges in support of Room to Read.
Eze Vidra
Follow me

Sources

Eze Vidra
About the Author

Eze Vidra

Eze Vidra is the founder of VC Cafe and Managing Partner at Remagine Ventures. He has written about Israeli tech, venture capital, AI, and startup building since 2005.

  • Founder of VC Cafe
  • Managing Partner at Remagine Ventures
  • Two decades covering Israeli tech and global venture trends
Total
0
Share