LidaSafety Research

Nonprofit AI safety research · Canada / Remote

We simulate the futures
we still have time to steer.

Lida Safety Research is a technical AI safety lab working on white-box interpretability, AI security, and world-scale simulation for governance — and on getting those findings in front of the people who set policy.

Fig. 1 · Sky chart, RA 0h–24hThe sky turns, predictably. The futures of AI don’t — yet.

24SPAR Fall 2026 fellows, with 3 mentors
3research projects running this term
2first-place hackathon wins, Apart and Redwood
18hof time zones between our furthest collaborators
Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war.

Signed by 600+ AI researchers and executives

Focus areas

Three lines of work, one question

What does a model do when no one is watching, and what would it take to notice in time? We attack that from the weights, from the attack surface, and from the world the systems will be deployed into.

White-box interpretability

Building model organisms of scheming and sandbagging, then reading the weights for the structure those behaviours share. If deception has a signature, we want a detector that does not depend on catching a lie in the output.

AI security

Fragmentation attacks split a prohibited task into pieces that each look harmless. We study how those attacks compose across calls and models, and what a defence that reassembles intent has to look like.

Simulation for governance

Policy proposals are argued about far more often than they are tested. We build agentic simulations of countries, labs and markets so a proposal can be run before it is written into law.

LidaSim

Wargaming

A world-scale wargame for catastrophic AI scenarios

LidaSim puts agentic AI and human players into the same global-scale simulation — states, labs, supply chains and the incentives that connect them — so a scenario can be played out rather than only described.

  1. Step 1

    Set the scenario

    A capability jump, a leaked model, a race between labs — the starting conditions we want to stress.

  2. Step 2

    Populate the world

    AI agents play countries, companies and markets; human players take seats alongside them.

  3. Step 3

    Play it out, many times

    Run the same world under different policies and starting conditions instead of arguing one story.

  4. Step 4

    Measure what a policy buys

    Compare outcomes with and without the intervention to see how much risk it actually removes.

It is where our forecasting work becomes something a policymaker can sit down and lose at — and the same engine drives our Simulating AI Policies fellowship project this term.

Selected results

  • 1stApart Research AI Governance hackathon — LidaSim
  • 1stRedwood Research Alignment Faking hackathonOct 2025
  • 4thApart Research def/acc hackathon, 1,000+ participants

Fellowship

Fall 2026

What we are running now

Three fellowship projects are active this term, run with the SPAR Fall 2026 cohort — 24 fellows and 3 mentors across 18 hours of time zones. Each project is meant to end in something usable: a detector, a prediction, or a simulation others can run.

Interpretability

The Signature of Scheming

We create model organisms for several distinct types of scheming, then apply modern interpretability techniques to ask whether they share a common structural pattern — one signature a monitor could look for instead of a taxonomy it has to enumerate.

9 fellows

Multi-agent

On Schelling Coordination

Can models collude without communicating? We test coordination across model families and fine-tunes, and use interpretability to predict, in advance, which pairs are likely to find the same focal point.

6 fellows

Governance

Simulating AI Policies

An agentic simulation of the world — countries, companies, and the pressures between them — used to evaluate how much a given policy actually reduces AI risk, and where it backfires.

9 fellows

Publications and code

  • Steering models away from sandbagging by grafting a reference activation into the forward pass.

  • Steering away from sandbagging without paired honest and malicious input–output examples.

  • Defences against fragmented misuse attacks, with mentored SPAR projects and hackathon placements behind them.

    CodeGitHub →AI security

Working notes and longer write-ups go up on blog.lidasafety.org.

Cohort

SPAR Fall 2026

The SPAR Fall 2026 cohort

Twenty-four fellows and three mentors, split across the three projects above. Doctoral and master’s candidates and postdocs at UCLA, Northwestern, NYU, Deakin and IDEAS NCBR; graduates of MIT, UC Berkeley, Minerva and UW–Madison; and working engineers from Amex, Instacart, Netflix, Luminance, Applied Intuition and Bain — spread over eighteen hours of time zones, from the Pacific coast to Australia.

The Signature of Scheming

9 fellows

  • Caleb DeLeeuw

    Lead author of “The Secret Agenda” on strategic deception in LLMs

    UTC−8

  • Jorge Roldan

    AI/ML Engineer at American Express

    New York

  • Leonardo Piano

    AI Engineer · evals, control, red-teaming

    Italy

  • Mirella Zeisler

    Software Engineer · AI control and deception

    Netherlands

  • Tung Nguyen

    AI/ML Engineer · interpretability, fairness, geometry ML

    Houston

  • Fariza Rashid

    Postdoctoral researcher · mech interp, AI security, red teaming

    Australia

  • Samuel Baumohl

    Incoming software engineer at Netflix · CS, UW–Madison

    UTC−8

  • Tuyen Tran

    Postdoctoral Researcher at Deakin University

    Australia

  • Taggart Tufte

    Fellow

    UTC−6

On Schelling Coordination

6 fellows

  • Alexander Olhava

    Mathematics PhD student at Northwestern; random matrix theory at UC Berkeley

    Chicago

  • Marcin Podhajski

    PhD candidate at IDEAS NCBR · AI security and graph neural networks

    Poland

  • Mudit Mangal

    Senior ML Engineer at Instacart · UC Berkeley School of Information

    UTC−8

  • Astha MehtaLida team

    Lead ML engineer at Bain & Company

    London

  • Cedric LamLida team

    MS student at NYU · Trust & Safety background

    New York

  • Xuanhan Tang

    Fellow

    UTC−6

Simulating AI Policies

9 fellows

  • Rishika Bansal

    AI engineer at Luminance on production LLM systems and model trust · MIT

    New York

  • Anthony Zang

    Engineer at Applied Intuition

    UTC−8

  • Karan Verma

    Independent AI researcher · evals, agentic systems, policy-impact forecasting

    India

  • Melanie BuiLida team

    AI Engineer and AI safety researcher · AI control, compute governance

    Vietnam

  • Tra My (Chiffon) NguyenLida team

    Researcher at Lida Safety · ML & Statistics, Minerva University

    Vietnam

  • Zachary SchlosserLida team

    Led the AI program for a US policy accelerator

    New York

  • Yosra Kazemi

    Independent AI researcher · causal inference, reinforcement learning

    Toronto

  • India Dearlove

    Fellow

    UTC+10

  • Phuc-Nguyen Nguyen

    Fellow

    Vietnam

Team

Who we are

A small nonprofit group of researchers and engineers, deliberately spread across countries and languages — because AI safety work that only happens in English, in two cities, is not finished work.

David Williams-King

David Williams-King

Co-founder · Executive Director

AI safety researcher and research manager focused on x-risk reduction. He works at the AI safety fellowship ERA and runs AI security events like FAST in Singapore, and enjoys mentoring people new to the field.

Linh Le

Linh Le

Co-founder · Research Director

Independent AI safety researcher affiliated with the Oxford AI Governance Initiative; previously Mila and the University of Technology Sydney. Dreams of running simulations of AI policies to avoid catastrophic outcomes for the world.

Researchers, engineers and fellows

Hong Kiat Tan

Hong Kiat Tan

Researcher · SPAR mentor

Finishing a PhD in Mathematics at UCLA. Your average mech interp enjoyer — developing and applying mechanistic interpretability to safety and alignment.

Melanie Bui

Melanie Bui

Researcher

AI safety and governance researcher, focused on coordination failures and governance solutions for middle powers navigating the AI transition.

Zachary Schlosser

Zachary Schlosser

Researcher

Researcher, educator and convener on AI policy. Previously led the AI program for a US policy accelerator; now builds the field around formal methods for AGI uncontainability.

Cedric Lam

Cedric Lam

Researcher

AI safety researcher from NYU with a Trust & Safety background. Works on cybersecurity and AI governance for creative spaces, especially AI-generated content.

My (Chiffon) Nguyen

My (Chiffon) Nguyen

Researcher

Works on keeping AI safe and empowering — reducing loss of control and disempowerment, multi-agent systems, multilingual AI and AI for good.

Arthur Collé

Arthur Collé

Agent backend engineer

Distributed systems engineer. Spent the last three years building production multi-agent systems and doing autonomous agent research. Based in Washington, DC.

Astha Mehta

Astha Mehta

SPAR fellow

Lead ML engineer at Bain & Company, working on AI safety and security.

Niruthiha Selvanayagam

Niruthiha Selvanayagam

SPAR fellow

PhD student at École de technologie supérieure (ÉTS Montréal), working on AI security.

Get involved

Three ways to work with us

One inbox, read by both directors. Tell us which of these you have in mind and we will route it.

Collaborate on research

Interpretability, AI security or simulation — if your question overlaps with ours, we would rather work on it together than in parallel.

Propose a collaboration →

Fund the work

Donations go through our fiscal sponsor, Anti Entropy, a 501(c)(3) (EIN 88-0967420), and are tax-deductible in the United States.

Ask about donating →

Mentor or join a cohort

Fellows join each term through SPAR; experienced researchers can also mentor a project or apply to us directly.

Write to us about a cohort →

Based

Canada / Remote

Fiscal sponsor

Anti Entropy, 501(c)(3)

Elsewhere

LinkedIn · X