We simulate the futures we still have time to steer.
Lida Safety Research is a technical AI safety lab working on white-box interpretability, AI security, and world-scale simulation for governance — and on getting those findings in front of the people who set policy.
Fig. 1 · Sky chart, RA 0h–24hThe sky turns, predictably. The futures of AI don’t — yet.
24SPAR Fall 2026 fellows, with 3 mentors
3research projects running this term
2first-place hackathon wins, Apart and Redwood
18hof time zones between our furthest collaborators
Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.
Signed by 600+ AI researchers and executives
Focus areas
Three lines of work, one question
What does a model do when no one is watching, and what would it take to notice in time? We attack that from the weights, from the attack surface, and from the world the systems will be deployed into.
White-box interpretability
Building model organisms of scheming and sandbagging, then reading the weights for the structure those behaviours share. If deception has a signature, we want a detector that does not depend on catching a lie in the output.
AI security
Fragmentation attacks split a prohibited task into pieces that each look harmless. We study how those attacks compose across calls and models, and what a defence that reassembles intent has to look like.
Simulation for governance
Policy proposals are argued about far more often than they are tested. We build agentic simulations of countries, labs and markets so a proposal can be run before it is written into law.
LidaSim
Wargaming
A world-scale wargame for catastrophic AI scenarios
LidaSim puts agentic AI and human players into the same global-scale simulation — states, labs, supply chains and the incentives that connect them — so a scenario can be played out rather than only described.
Step 1
Set the scenario
A capability jump, a leaked model, a race between labs — the starting conditions we want to stress.
Step 2
Populate the world
AI agents play countries, companies and markets; human players take seats alongside them.
Step 3
Play it out, many times
Run the same world under different policies and starting conditions instead of arguing one story.
Step 4
Measure what a policy buys
Compare outcomes with and without the intervention to see how much risk it actually removes.
It is where our forecasting work becomes something a policymaker can sit down and lose at — and the same engine drives our Simulating AI Policies fellowship project this term.
Selected results
1stApart Research AI Governance hackathon — LidaSim—
1stRedwood Research Alignment Faking hackathonOct 2025
4thApart Research def/acc hackathon, 1,000+ participants—
Fellowship
Fall 2026
What we are running now
Three fellowship projects are active this term, run with the SPAR Fall 2026 cohort — 24 fellows and 3 mentors across 18 hours of time zones. Each project is meant to end in something usable: a detector, a prediction, or a simulation others can run.
Interpretability
The Signature of Scheming
We create model organisms for several distinct types of scheming, then apply modern interpretability techniques to ask whether they share a common structural pattern — one signature a monitor could look for instead of a taxonomy it has to enumerate.
9 fellows
Multi-agent
On Schelling Coordination
Can models collude without communicating? We test coordination across model families and fine-tunes, and use interpretability to predict, in advance, which pairs are likely to find the same focal point.
6 fellows
Governance
Simulating AI Policies
An agentic simulation of the world — countries, companies, and the pressures between them — used to evaluate how much a given policy actually reduces AI risk, and where it backfires.
9 fellows
Publications and code
Steering models away from sandbagging by grafting a reference activation into the forward pass.
Twenty-four fellows and three mentors, split across the three projects above. Doctoral and master’s candidates and postdocs at UCLA, Northwestern, NYU, Deakin and IDEAS NCBR; graduates of MIT, UC Berkeley, Minerva and UW–Madison; and working engineers from Amex, Instacart, Netflix, Luminance, Applied Intuition and Bain — spread over eighteen hours of time zones, from the Pacific coast to Australia.
A small nonprofit group of researchers and engineers, deliberately spread across countries and languages — because AI safety work that only happens in English, in two cities, is not finished work.
David Williams-King
Co-founder · Executive Director
AI safety researcher and research manager focused on x-risk reduction. He works at the AI safety fellowship ERA and runs AI security events like FAST in Singapore, and enjoys mentoring people new to the field.
Linh Le
Co-founder · Research Director
Independent AI safety researcher affiliated with the Oxford AI Governance Initiative; previously Mila and the University of Technology Sydney. Dreams of running simulations of AI policies to avoid catastrophic outcomes for the world.
Researchers, engineers and fellows
Hong Kiat Tan
Researcher · SPAR mentor
Finishing a PhD in Mathematics at UCLA. Your average mech interp enjoyer — developing and applying mechanistic interpretability to safety and alignment.
Melanie Bui
Researcher
AI safety and governance researcher, focused on coordination failures and governance solutions for middle powers navigating the AI transition.
Zachary Schlosser
Researcher
Researcher, educator and convener on AI policy. Previously led the AI program for a US policy accelerator; now builds the field around formal methods for AGI uncontainability.
Cedric Lam
Researcher
AI safety researcher from NYU with a Trust & Safety background. Works on cybersecurity and AI governance for creative spaces, especially AI-generated content.
My (Chiffon) Nguyen
Researcher
Works on keeping AI safe and empowering — reducing loss of control and disempowerment, multi-agent systems, multilingual AI and AI for good.
Arthur Collé
Agent backend engineer
Distributed systems engineer. Spent the last three years building production multi-agent systems and doing autonomous agent research. Based in Washington, DC.
Astha Mehta
SPAR fellow
Lead ML engineer at Bain & Company, working on AI safety and security.
Niruthiha Selvanayagam
SPAR fellow
PhD student at École de technologie supérieure (ÉTS Montréal), working on AI security.
Get involved
Three ways to work with us
One inbox, read by both directors. Tell us which of these you have in mind and we will route it.
Collaborate on research
Interpretability, AI security or simulation — if your question overlaps with ours, we would rather work on it together than in parallel.