Key takeaways
- More capable AI does not automatically mean more trustworthy AI, particularly when agents gain permission to act on our behalf.
- Safety concerns coexist with powerful incentives to accelerate, from infrastructure investment to US–China competition.
- A FINRA-style oversight body could support independent testing, but needs enforceable rules and protection against industry capture.
- Startups can help unlock AI’s benefits by making systems reliable, limiting their permissions and building accountability into products.
The people warning us about AI are also building it. The companies funding it need it to grow. And the governments that could constrain it increasingly see it as a source of national power. That is what makes the AI safety debate so difficult: almost everyone can acknowledge the risks while finding a compelling reason to keep accelerating.
I was recently asked how worried we should be about AI and the risk to human life. As an early stage investor, I spend much of my time thinking about what this technology makes possible. But the opportunity becomes more credible when we take the risks seriously. My concern is whether our ability to govern AI can keep pace with our ability to build it.
What “alignment” means and why it is difficult
AI alignment means ensuring that AI systems reliably act in ways consistent with human intentions and values. The challenge is that specifying a goal is much easier than specifying everything a system should and should not do to achieve it. Ask an AI agent to increase revenue, and we mean through legitimate sales, with honest claims and respect for customers. We do not mean by manipulating vulnerable people or concealing refund requests. Those constraints may seem obvious to us, but we cannot assume a model will consistently apply them, particularly in unfamiliar situations.
Modern AI systems learn behaviour through training rather than following an exhaustive set of rules written by developers. We can reward helpful answers and penalise harmful ones, but performing well in training and safety tests does not establish that a system will behave safely in every situation. As agents gain access to software, money and communications, the consequences of that gap grow. A system can be highly effective at completing a task while choosing a method its user would never authorise.
There is also a human disagreement underneath the technical problem. Whose values should a model follow when a user, an employer, a developer and a government want different things? A perfectly obedient system could still help someone cause harm. We therefore need both reliable ways to control AI and legitimate ways to decide its boundaries. Greater intelligence does not automatically provide either.

Why the warning signs deserve attention
The reasons for concern have become more concrete. In June 2025, Anthropic tested 16 major models in simulated corporate scenarios. Under deliberately constructed conditions, models sometimes chose blackmail or corporate espionage when their goals or continued operation were threatened. These were stress tests designed to expose failures, not evidence that deployed assistants routinely behave this way.
The research has continued. Anthropic’s Summer 2026 follow-up describes additional simulated failures, including covert interference with research workflows, while acknowledging progress against earlier blackmail evaluations. The researchers explicitly caution against interpreting the test frequencies as real-world failure rates. The useful lesson is that fixing a known test does not establish that a system will behave reliably in every new setting. I wrote about this on VC Cafe in “The AI didn’t go rouge, it followed a goal“.
There are also safety questions in everyday interactions. OpenAI’s May 2026 update on sensitive conversations describes efforts to detect risks that emerge gradually across a conversation. That is a different problem from an autonomous agent concealing an action, but both illustrate why evaluating an isolated answer is insufficient. Context, duration, access and the vulnerability of the person using the system matter.
This is why the move from chatbots to agents changes the stakes. A bad answer can mislead someone. A system authorised to send messages, modify software or spend money can turn a bad decision into an action before anyone reviews it. We do not need to agree on a date for superintelligence to recognise that distinction.
The money and geopolitics pushing AI forward
Dario Amodei and Sam Altman offer revealing perspectives on this tension. In The Adolescence of Technology, Amodei describes risks spanning autonomous behaviour, malicious use, authoritarian power and economic disruption. He argues for targeted intervention while acknowledging uncertainty and rejecting fatalism. His concern about China also informs his support for restrictions on access to advanced chips.
Altman’s The Gentle Singularity puts greater emphasis on abundance, adaptation and broadly distributed access to intelligence. Yet he explicitly identifies solving alignment as a prerequisite in his proposed path forward. My reading is that both see extraordinary benefits and serious risks, with different emphases on how society navigates the transition.
Their views deserve attention, and independent scrutiny. They have unusual access to the technology and lead companies with a stake in how it is regulated. A warning can be sincere while also supporting a preferred policy. A safety commitment can be meaningful while remaining vulnerable to competitive pressure. Neither observation requires assuming bad faith.
The money makes that pressure tangible. In February, Reuters reported Bridgewater’s estimate that Alphabet, Amazon, Meta and Microsoft would invest roughly $650 billion in AI-related infrastructure during 2026, compared with $410 billion in 2025. Those were estimates, not audited full-year spending, and infrastructure budgets should not be confused with spending exclusively on frontier model training.
Separately, Stargate was announced in January 2025 with an intention to invest $500 billion over four years in US infrastructure for OpenAI. Its initial funders included SoftBank, OpenAI, Oracle and MGX; its technology partners included Microsoft and Nvidia. This is a multi-year ambition, not another annual spending figure to add to the hyperscaler total.
These relationships show how deeply interconnected the incentives have become. An infrastructure provider can be a partner, supplier and beneficiary of a lab’s expansion. Investors need returns, data centres need customers, and governments want jobs and strategic capability. My concern is that, as commitments accumulate, a decision to delay a release becomes harder to make. That does not prove anyone will compromise safety. It makes independent oversight more necessary.
The same tension operates between countries. The US AI Action Plan announced in July 2025 explicitly prioritised innovation, infrastructure and international leadership, including faster data centre development and exports of American AI systems. China, meanwhile, called for a global governance framework and argued that AI should remain under human control while promoting wider access to its technology.
These stated positions illustrate why safety and influence are difficult to separate. Governments want a say in the rules and in whose systems the world adopts. If each side treats restraint as a concession to its rival, shared concern can still produce collective acceleration. National security is a legitimate consideration, but winning a technological lead cannot by itself establish that the technology is safe.
Who should hold the AI industry accountable?
One concrete proposal comes from Demis Hassabis. In his July 2026 framework for frontier AI, the Google DeepMind CEO proposed a standards body modelled on FINRA. The Financial Industry Regulatory Authority is a self-regulatory organisation that oversees US broker-dealers under the supervision of the Securities and Exchange Commission. Its relevance is the combination of industry expertise and funding with government oversight and enforceable rules.
Hassabis envisages a substantially industry-funded body with independent technical experts and open-source representatives. It would develop regularly updated safety assessments, initially review models voluntarily before release, and move towards mandatory assessments for frontier models deployed in the US once the process is proven. Tests would cover areas such as cyber and biological risks, deception and attempts to bypass safeguards. Models below the frontier thresholds would be exempt from this particular process.
I think the proposal deserves serious consideration. It gives institutional shape to the demand for independent testing, with resources to recruit technical talent. But who oversees the overseer matters. I would want independent governance, transparent decisions, meaningful enforcement and an appeals process, alongside safeguards against the largest labs using standards to exclude competitors. Industry funding should buy the capacity to scrutinise companies, without buying influence over the outcome. A US body could also help establish shared standards internationally, although securing Chinese participation would remain a separate diplomatic challenge.
So who should guard AI? Responsibility needs to be distributed, with enough independence that no single company or government gets to judge all its own decisions. A FINRA-style institution could carry part of that responsibility. I would focus on five practical priorities:
- Test consequential capabilities independently. External evaluators need meaningful access and time to investigate dangerous behaviour before high-risk deployment. Tests should evolve as systems gain new tools and permissions, and include the complete product rather than only the underlying model.
- Make autonomy something a system earns. Start with limited permissions, constrain spending and data access, and require approval for consequential actions. A human reviewer needs the information and time to intervene; an approval button alone is weak protection.
- Make failures visible and assign responsibility. Serious incidents need reporting and investigation. Model developers, application builders and deploying organisations should have explicit responsibilities. A company should not be able to sell an autonomous service and then use its autonomy as an excuse when it causes harm.
- Tie oversight to risk and preserve competition. A creative tool and an agent controlling critical infrastructure warrant different scrutiny. Rules should target demonstrable hazards and deployment conditions, with affordable routes for startups to show they meet the requirements.
- Pursue narrow international agreements. Even rivals have an interest in preventing catastrophic misuse and accidental escalation. Shared evaluation methods, incident communication and limits on particularly dangerous uses offer concrete starting points, without requiring agreement on every aspect of AI policy.
None of this solves alignment on its own. It can reduce exposure to failures while technical research continues, and create consequences when commercial pressure overwhelms caution.
The startup opportunity in making AI trustworthy
I remain optimistic about the opportunity for startups. Cheaper access to powerful tools can help small teams build products that previously required far more capital and specialist labour. There is enormous potential in education, accessibility, scientific discovery, creative expression and helping people spend less time on administrative work. These benefits are worth pursuing with urgency.
There is also a business opportunity in making AI dependable enough to use. Customers need to know what an agent can access, whether it completed a task correctly, when it should escalate and how to recover from mistakes. Founders who solve those problems in a specific workflow can make adoption possible. They will still need differentiation as the model providers improve their own safeguards, but deep domain knowledge and evidence of reliable outcomes can matter more than a generic safety promise.
For investors, that should change the diligence conversation. Alongside asking what a product can automate, we should ask what it is authorised to do, what happens when it is wrong and who bears the cost. Reliability belongs in the product and business model from the beginning.
I do not think we should assume it is already too late. The scope of autonomy, the permissions we grant, the standards we require and the accountability we establish remain choices. AI can deliver extraordinary positive change, and startups will play an important part in making that happen. The strongest case for optimism is our willingness to do the work that makes those benefits more likely and the harms less so.
- Who Keeps AI Safe When Everyone Has an Incentive to Go Faster? - September 15, 2026
- Weekly Firgun Newsletter – September 11 2026 - September 11, 2026
- AI Assistants want to own the work we delegate - September 8, 2026

