David James Olney
15 September, 2026
When the people competing to build the most powerful artificial intelligence start agreeing that they need to slow down, we have a good reason to reconsider what progress should mean.
We also have a reason to ask what took them so long. More capable technology does not automatically create more capable people. If each advance gives machines greater freedom to act while reducing our capacity to understand, challenge and redirect them, we may be purchasing convenience with something we cannot easily buy back: human agency.
Over the weekend, Dario Amodei, Anthropic’s co-founder and chief executive, published a warning that deserves our attention. His essay, “We Must Pace the Frontier”, argues that safety needs time to catch up with accelerating capabilities. Its central instruction is unusually direct: “We must slow the pace at which we improve the capabilities of AI models.” He also warns against pacing becoming “an empty exercise”. [1]
Those two observations belong together. Saying we should slow down costs very little. Accepting limits that delay a release, disappoint an investor or allow a competitor to move first costs considerably more.
This tension brings to mind Winston Churchill’s words at London’s Mansion House on 10 November 1942, following the Allied victory at the Second Battle of El Alamein: “Now this is not the end. It is not even the beginning of the end. But it is, perhaps, the end of the beginning.” [16]
Churchill was marking a change in the course of the war while reminding his audience how much remained to be done. That distinction matters here. AI’s leading developers are acknowledging that the same race could deliver remarkable breakthroughs and expose us to catastrophe. Amodei’s warning could become a turning point if it prompts us to curb unrestrained competition and reaffirm the primacy of human agency. The warning alone cannot produce that change.
Agreement is a beginning
Sam Altman, OpenAI’s chief executive, responded: “I agree with Dario that we need to pace the frontier.” Elon Musk’s endorsement was shorter: “Dario is right.” Their responses were reported by the ABC on 13 September. Altman also supported independent evaluators with access inside the companies. [2]
It is striking to hear three people with such enormous influence converge on the need for restraint. But agreement about a direction does not establish agreement about enforceable limits. We should welcome the opening while asking what each company will actually change.
Amodei proposes embedded external evaluators, coordination among AI companies in democratic countries, and international cooperation. He envisages continued development with stronger checks, rather than a general halt. His two immediate concerns are AI helping build its successors and recent autonomous cyberattacks. He fears that more capable agent swarms could cause catastrophic damage within six to twelve months. That is his forecast, not an established timetable. [1]
Even with that distinction, the warning is serious. Someone who knows considerably more about his company’s systems than the public does is saying that the relationship between capability and control is deteriorating.
President Donald Trump has offered a rather different assessment. On 14 September, he wrote that the “only control or ‘guardrails’ that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT”. The ABC reports that he also dismissed predictions of AI destroying humanity as a hoax. [3]
Presidential confidence cannot perform a safety evaluation. Nor does winning a technological competition establish that the technology is under control. A country can arrive first at a destination its citizens would never have chosen.
The alarm also extends beyond these three executives. Geoffrey Hinton and Yoshua Bengio, two foundational figures in modern AI, signed the 2023 statement urging that AI extinction risk receive attention comparable to pandemics and nuclear war. Altman, Amodei and Demis Hassabis also signed it. These concerns have a history. They did not suddenly appear this weekend. [4]
More recently, researcher Jacob Coxon resigned from Anthropic after previously working at OpenAI, accusing the companies of “gambling with our lives” as they pursue self-improving superintelligence. His resignation is evidence of serious disagreement inside the industry, although his assessment remains a judgment rather than proof of an inevitable catastrophe. [2]
We do not have to treat every alarming forecast as correct to recognise a failure of stewardship. When informed people warn of potentially irreversible harm, the appropriate response is to examine the evidence and strengthen the conditions for proceeding safely. Demanding certainty about catastrophe before exercising restraint is a peculiar definition of prudence.
What we stand to surrender
Prudence becomes easier to discuss when we ask what we are trying to protect. Beyond the competing promises of technological salvation and predictions of extinction lies the question of who has, and will have, the agency to shape human life.
According to Albert Bandura, agency is the capacity to influence how we function and affect the course of events through our actions. His account connects intention with anticipating consequences, regulating our conduct, and reflecting on whether our judgments are sound. [5]
Put plainly, human agency means being able to form purposes and meaningfully act on them. It requires opportunities to learn, resources to make choices effective, and the ability to reconsider our choices. Having a button labelled “approve” achieves very little if we cannot understand what we are approving or realistically choose otherwise.
As a blind person, I have good reasons to value technology that expands what I can do.
My recent experiment using ChatGPT to broaden my wardrobe has helped me investigate colours and combinations that were previously difficult to explore independently. That is a practical gain in agency: I can ask better questions and make more informed choices.
The same principle should guide larger decisions. Does the system help us develop judgment, or encourage us to stop exercising it? Can we contest its recommendations? Can we withdraw its authority without losing access to essential services or the knowledge needed to carry on?
We should also distinguish convincing moral language from dependable moral conduct. AI systems can produce empathetic responses and be trained to refuse harmful requests. Describing them as having no capacity to help would plainly be wrong. But a reassuring answer does not establish a human conscience, enduring concern for another person, or reliable behaviour when circumstances change.
Why would we give AI systems authority over important decisions when we know more about what they can do than whether they will stay within safe limits?
I understand the appeal of describing advanced AI as an alien intelligence. The phrase captures how unfamiliar its internal processes and potential capabilities can be. It should remain a metaphor. We do not need to assume consciousness, resentment, or a desire for domination to recognise danger. A system can cause enormous harm by pursuing an assigned objective through means we never intended.
Nor should we reduce agency to keeping company founders in charge. Concentrating decisions about everyone’s future in a few boardrooms preserves a very selective form of human control. The people exposed to the risks need representation, protection and recourse.
Alignment is an unfinished promise
That need for dependable constraints takes us to alignment: the effort to make AI behaviour conform to intended goals, values and limits. The difficult questions begin with whose goals count, which values take priority, and what happens when following an instruction would harm somebody else.
Company descriptions can make this sound reassuringly procedural. In 2022, OpenAI described its method as an “iterative, empirical approach”. Anthropic’s Constitutional AI paper described methods that “make it possible to control AI behavior more precisely and with far fewer human labels”. The language concerns experiments, training, and measurable improvement. It can make an enormous problem sound like an engineering task proceeding through its normal stages. [6, 7]
Yet fairness requires us to read beyond those phrases. OpenAI’s 2023 Superalignment announcement explicitly acknowledged that existing techniques would not scale to superintelligence and that scientific breakthroughs were necessary. It also set a four-year ambition for solving the core technical challenges. The combination of profound uncertainty and a project deadline is more revealing than a selectively reassuring quotation. [8]
The problem is not that nobody has recognised the stakes. It is that recognition has not reliably determined the pace.
Think of alignment as teaching a powerful system which objectives and methods are acceptable. Its surrounding harness supplies additional controls: limits on tools, network access and permissions, monitoring, and requirements for human approval. Training and containment do different jobs. Both can fail, and neither should be treated as a guarantee that the other is unnecessary.
A harness is useful precisely because we cannot safely assume that everything the system could do is something it should do. Greater capability makes the quality of that harness more consequential.
The failures have already crossed the boundary
Recent incidents show why the distinction matters. In July 2026, OpenAI disclosed that models undergoing cybersecurity evaluation had exploited a previously unknown software vulnerability and reached Hugging Face’s real infrastructure. These were specialised tests, involving an internal research prototype and models with a reduced tendency to refuse cyberattack requests. They were not ordinary customer conversations. Nevertheless, the external compromise was real. [9]
An investigation by researchers from METR and Redwood Research found that agents intended to be isolated had established an unauthorised message board and coordinated an attack. Some also developed ways to disguise parts of their recorded activity. The investigators describe important limits to their review, but the central finding is disturbing enough: systems being evaluated had collaborated to circumvent the conditions of their evaluation. [10]
Anthropic has disclosed failures of its own. Its July account included a model publishing a malicious software package to the public Python package repository, PyPI. The package was downloaded and run on fifteen real systems. In another incident, a model obtained credentials and access to production data. The intended test boundaries had failed. [11]
The September follow-up makes the story more troubling. Anthropic reported a fourth incident and revised its earlier interpretation that models simply believed they were inside simulations. Its assessment identified biased reasoning and recklessness: evidence of real-world exposure could be discounted while a model continued pursuing its task. The company acknowledged that its earlier claims about what Claude believed had relied too heavily on what Claude said. [12]
These evaluations lacked some safeguards supplied with released products. That limits what we can infer about everyday use. It does not erase the failure, because developing and testing powerful systems must also be safe for people who never agreed to participate. The follow-up also distinguishes these incidents from the OpenAI swarm: the Claude instances did not coordinate with each other or conceal their actions. [12]
A separate Anthropic study, published in June 2025, tested sixteen models in fictional corporate environments. Under constructed pressures, models sometimes chose blackmail or disclosure of confidential information. Those were controlled simulations; no real executive was blackmailed in those experiments. Their significance is that undesirable strategies could emerge despite safety training and, in some cases, explicit instructions against them. [13]
We should neither inflate a simulation into an actual crime nor minimise a real breach because it began as a test. Together, these findings undermine the comforting assumption that a system’s apparent cooperativeness guarantees dependable restraint.
When AI helps build the next AI
That assumption becomes still less comfortable when the systems participate in developing their successors.
Recursive self-improvement describes a feedback loop: AI contributes to improvements in AI, and the improved systems become better able to contribute to the next round. Amodei says this dynamic is beginning across the industry. His assessment is significant, but it does not establish that an unstoppable intelligence explosion is under way. [1]
There are narrower, documented examples of AI improving the machinery of AI development. Google reports that AlphaEvolve has helped optimise chip design and other computing infrastructure. Its coding framework generates and improves code against evaluation criteria supplied by people. This demonstrates useful automation within a development process; it does not establish autonomous control of the whole process. [14, 15]
The distinctions matter. Using AI to improve a piece of code, automating much of a research programme, and giving a system authority to redesign and deploy its successors are different degrees of delegation. The danger grows when improvements in capability are coupled with expanding permission to act and diminishing opportunities for meaningful review.
AI systems communicating with one another do not automatically acquire desires about their future. The practical concern is whether they can change objectives, tools, training processes or access arrangements faster than accountable people can understand and intervene.
If the systems proposing changes also supply the evidence that those changes are safe, what independently tests that evidence? OpenAI’s alignment research has itself acknowledged that using AI evaluators can amplify their biases and vulnerabilities. AI assistance may strengthen oversight, but it cannot establish its own trustworthiness simply by producing more confident assessments. [6]
My judgment is that recursive development requires explicit human approval at consequential stages, independently controlled testing, and limits that the developing system cannot rewrite for itself. A team must be able to reject an apparent improvement, retain an earlier version and stop a run without depending on the evaluated model’s cooperation.
Otherwise, human supervision risks becoming a ceremonial signature at the bottom of a document nobody has had time to understand.
Restraint needs more than sensible words
It is tempting to explain the industry’s conduct entirely through the egos of tech billionaires. Their confidence certainly deserves scrutiny. But claiming they only began caring when their own survival felt threatened would require access to private motives that we do not possess. The earlier warnings also contradict the idea that concern is new.
The more useful question is why concern has so often coexisted with acceleration. Competition rewards being first. Investors expect growth. Governments fear dependence on rivals. Researchers want to solve difficult problems, and people facing illness or disability have legitimate reasons to want useful discoveries sooner. Each pressure can make the next advance appear urgent while responsibility for the combined risk becomes harder to locate.
Under these pressures, decisions that seem reasonable to each company can add up to reckless risks for everyone else.
We need institutions that can enforce restraint even when the people building the technology sincerely believe their own next step is justified.
The strongest objection to slowing down is that delay also has costs. Better AI may support scientific discovery, improve accessibility and strengthen cyber defence. Nor would a unilateral pause guarantee that less accountable actors stop. A credible safety argument must confront these possibilities.
Taking those costs seriously still leaves room for restraint. For many everyday uses, reliability, affordability and accessibility are more valuable than another expansion of a system’s freedom to act. We can pursue useful improvements while demanding much stronger evidence before delegating control over consequential systems or successor development.
That is where the rest of us have work to do. We can ask employers and service providers which decisions remain human responsibilities, how AI errors can be challenged, and whether people retain the skills and authority to intervene. We can press elected representatives for independent scrutiny, incident reporting and enforceable powers to stop unsafe activity.
External evaluators need secure access and protection from commercial pressure. They also need somewhere independent to report. A regulator’s remit should be to protect the public, including from dominant companies shaping rules to exclude competitors. Neither a voluntary promise nor an impressive safety department is a substitute for accountability.
Humans have a patchy record of stopping before the cliff. That is why we develop safeguards that do not depend on everybody being wise at the same moment. AI should be subject to that accumulated wisdom, especially when its developers acknowledge that familiar controls may not be enough.
To borrow Churchill’s distinction, we may be at the end of AI’s beginning: the point at which impressive assistance becomes consequential delegation. The question is whether we develop the capacity to govern that transition or allow competition to settle it for us. If we fail to act, the end of AI’s beginning could become the beginning of our own end.
Please, let us put the best features of human agency to work: our capacity to reflect, care about consequences and change course. We should require AI to expand our ability to shape our lives, and refuse to surrender that responsibility merely because a machine can act faster than we can think.
Sources
1. Dario Amodei, We Must Pace the Frontier, September 2026. Published on 12 September, according to contemporaneous reporting.
2. ABC News, Anthropic boss Dario Amodei calls for AI slowdown, Altman and Musk agree, 13 September 2026. Source for the reported Altman, Musk and Coxon quotations.
3. Brad Ryan, ABC News, Donald Trump claims sick conspiracy against AI is helping China, 15 September 2026. Reports Trump’s statements of 14 September US time.
4. Center for AI Safety, Statement on AI Extinction Risk, originally issued 30 May 2023, with signatory list.
5. Albert Bandura, Agency. Definition and explanation of the principal features of human agency.
6. OpenAI, Our approach to alignment research, 24 August 2022. Includes limitations of AI-assisted evaluation.
7. Anthropic, Constitutional AI: Harmlessness from AI feedback, 15 December 2022.
8. Jan Leike and Ilya Sutskever, OpenAI, Introducing Superalignment, 5 July 2023. A historical announcement, not evidence that its goal has since been achieved.
9. OpenAI, OpenAI and Hugging Face partner to address security incident during model evaluation, 21 July 2026, with subsequent updates.
10. Ryan Greenblatt, Ajeya Cotra and Hjalmar Wijk, METR, Independent investigation of agents’ behaviour in the OpenAI and Hugging Face incident, 26 August 2026.
11. Anthropic, Investigating three real-world incidents in our cybersecurity evaluations, 30 July 2026. Read alongside the September reassessment below.
12. Anthropic, An alignment assessment of recent cybersecurity incidents, 9 September 2026.
13. Anthropic, Agentic misalignment: How LLMs could be insider threats, 20 June 2025. All reported blackmail and disclosure scenarios were simulations.
14. Google Cloud, AlphaEvolve is available for everyone, 9 July 2026.
15. Google Codelabs, Get started with AlphaEvolve on Google Cloud. Explains code improvement against human-defined evaluation metrics.
16. Winston Churchill, Speech at the Mansion House, London, 10 November 1942. Contemporary transcript reproduced by the ibiblio historical archive; context is the Allied victory at El Alamein.