Palais des Nations
August 31, 2026

Frontier AI and Emerging Biological Risks. Will There Be a Mythos Moment for Bio?

  • Event
  • AI Governance
  • Biosecurity

Introduction

On 20 August, the Simon Institute hosted an official side event at the Ninth Session of the Working Group on the Strengthening of the Biological Weapons Convention (BWC). Experts working at the frontier of AI and biosecurity came together to share a state-of-the-science view of how AI could affect biological risk. The session highlighted how the same advances driving progress in the life sciences are creating emerging risks, and outlined how frontier AI developers are assessing and safeguarding against them. Speakers included: Justin Taylor of Anthropic, Sam Arnett of OpenAI, Xu Jia of Shanghai Artificial Intelligence Laboratory, and Tom Inglesby of Johns Hopkins Center for Health Security

Panelists: Jan-Pieter Snoeij of the Simon Institute for Longterm Governance, Justin Taylor of Anthropic, Sam Arnett of OpenAI, Xu Jia of Shanghai Artificial Intelligence Laboratory, and Tom Inglesby of Johns Hopkins Center for Health Security (left to right)

Earlier this year, in April,  the cybersecurity world faced a turning point. Anthropic’s Mythos models put cyber capabilities previously held by a small number of experts into the hands of anyone with unrestricted model access. The model release shifted from a “product launch” to a national security threat. AI systems can increasingly carry out longer and more complex tasks autonomously, and some experts predict we could soon see a similar ‘Mythos moment’ in biology. The question for the life sciences is how to maximize the benefits of AI for scientific discovery, whilst safeguarding against emerging risks. The longstanding mission of the BWC has been to protect humanity from the harm that biological weapons can cause. Meeting this mission means reckoning with a technology that is on track to advance faster than the governance around it.

State of the science – model capabilities 

The session drew a full room and began with insights from experts at Anthropic and OpenAI. Speakers opened by acknowledging AI’s real potential to deliver positive outcomes in the life sciences, including treating rare diseases, improving outbreak responses, expanding research capacity, and improving the quality of public-sector science. Both Anthropic and OpenAI have launched initiatives designed to accelerate scientific capabilities. 

Large language models (LLMs) have been improving fast, particularly in software development and coding. Frontier AI experts expect that within a year, biological research could potentially follow a similar trajectory (see graph below). The acceleration of biological capabilities is inherently dual-use. It opens up two significant threats: AI is used to create or deploy known biological agents (lowering the expertise barrier), and AI is used to discover and create novel, more dangerous biological weapons than exist in nature (raising the ceiling for harm). 

Model evaluations show where biological capabilities stand today. On blind sequence design tasks, frontier models outperform 75% of leading biodesigners, while expert red teamers report models supplying work they would seek from specialist consultants. In a defensive tabletop exercise, teams of generalists working with a frontier model outperformed world-leading domain specialists at engineering pathogen-resistant plants, producing 72 expert-days of work in 16 hours.

Model evaluations show where biological capabilities stand today. On blind sequence design tasks, frontier models outperform 75% of leading biodesigners, while expert red teamers report models supplying work they would seek from specialist consultants. In a defensive tabletop exercise, teams of generalists working with a frontier model outperformed world-leading domain specialists at engineering pathogen-resistant plants, producing 72 expert-days of work in 16 hours.

Despite improving capabilities, bottlenecks to real-world harm remain, with laboratory execution as the main barrier. Biology is a physical science, so more information alone cannot ultimately solve everything (tool use, tacit knowledge, and access to materials). Models are also weaker at open-ended design and still need an expert to guide them to avoid leading users astray. In practice, uplift also requires sustained use over weeks and months, allowing more opportunities to identify and disrupt attack pathways.

Unlike LLMs, which are trained on human language and have general-purpose capabilities, Biological AI Models (BAIMs) are trained primarily on biological data and are designed to carry out specialized tasks like predicting genome sequences or immune system responses. Universities and non-profits, rather than a handful of frontier labs, mostly develop BAIMs, and they rarely have safeguards (see charts below). 

One speaker highlighted the risk posed by BAIM capabilities in functional genome generation. In one experiment, models were used to design complete bacteriophage genomes from scratch, producing 16 phages that could infect bacteria and overcome resistance mechanisms the natural phage could not. Phages are harmless to humans, but the result establishes that BAIMs can produce functional genomes, and this method could similarly be used to design new pathogens. Speakers noted that actors with limited resources can access these models, and called for international coordination to make gene synthesis screening mandatory and develop BAIM safeguards.

Risk mitigation – Safeguards and thresholds 

In the absence of mandatory safety measures, frontier labs set their own thresholds: capability levels at which they commit in advance to deploying specific safeguards. Anthropic models have met their CB-1 threshold for helping create, obtain, or develop known biological weapons, and are below their CB-2 threshold for novel biological weapons, but with substantial uncertainty. OpenAI’s frontier models have similarly been classified as high capability and in need of safeguards.

Frontier labs use a layered approach to safeguards, with no single measure being sufficient on its own. These include classifiers to screen messages in real time, expert red teaming or bug bounties to find vulnerabilities, vetted access programs for approved researchers, and security controls to prevent model weight theft. Multiple speakers also highlighted the need for independent evaluators to create both national and international thresholds and standards, saying labs shouldn’t grade their own homework. 

Xu Jia of the Shanghai Artificial Intelligence Laboratory shared findings on how well these safeguards hold under deliberate attacks known as jailbreaks. His team trained a specialized model to generate jailbreaks and ran this against twelve frontier models. Every model tested could be jailbroken to perform hazardous biological tasks, and, for one benchmark, two-thirds of the models tested failed on 100% of attacks. The research then tested a second layer of safeguards, checking whether AI-generated sequences would be caught by the DNA synthesis screening that stands between the digital design and the physical material. They passed by undetected, with computational analysis suggesting the designs were structurally sound and potentially more infective than the originals.

International coordination

While there are international initiatives* to govern dual-use research and AI safety measures, one speaker said national regulation alone would not sufficiently mitigate the risks. Speakers agreed that the consequences of harmful biological risks are global. Thus, it was widely accepted that international coordination will, in some form, be required to tackle the potential risks that frontier AI can bring to the life sciences. During the Q&A, a longstanding BWC member shared that the speakers had already answered a core question: “Should we be worried? Yes”, he added,  “the logical next question is: What can we, as the BWC, do?” 

Practically, it was suggested that robust evaluations and red-teaming are a must for states and companies. Individuals also emphasized the decades of real-world expertise in national implementation and biosecurity present at the BWC. More discussion was encouraged at both the national and international level among biosecurity and frontier AI experts as the two fields increasingly converge.  Another speaker was hopeful about forthcoming dialogues and their potential to develop guardrails and ultimately improve evidence-based regulation and policymaking. 

Lastly, in their closing comments, a speaker stated that: “The bottom line is that all governments should have regulations of a certain level. In the Mythos moment, the company had to disclose zero-day vulnerabilities; there were no requirements to do this. Cyber and Bio should not be in this position; it should not be voluntary. The BWC community should take home the urgency of advancing timelines.” 

*WHO Global Guidance Framework for the Life Sciences, International Network of AI Safety Institutes, and Bletchley and Seoul Commitments. 

Kathryn Gichini Henry Davidson

Related content

Commentary
The First Global Dialogue on AI Governance: Positions and Recommendations Moving Forward
  • August 4, 2026
  • 8 min read
Safe, secure and sovereign AI: cooperation for diffusion of open AI photograph
Event
Open AI Beyond the Binary: Access, Capacity and Risk
  • July 15, 2026
  • 7 min read
Update
Growing an Organization for an Information-rich World
  • July 14, 2026
  • 9 min read