Safety · the true AGI age
Keep the most powerful minds where we can still govern them.
As general-purpose AI approaches true capability, the decisive safety question is not only how capable these systems become — it is where they are allowed to act, and who can stop them when they are wrong. AGI Research Lab works on two commitments that answer it. One limits where advanced intelligence can reach. The other turns that same intelligence, first, toward defending the people and institutions who depend on the digital world.
Commitment 01
Keep AGI in the digital world.
A mind confined to information can be audited, paused, copied, and switched off. A mind wired into the physical world — through autonomous robots, actuators, and internet-connected machinery — cannot be so easily recalled.
The physical world has no undo. A software mistake rolls back; a physical action may already have happened. Every bridge we build between advanced AI and real-world machinery is a bridge that makes the off-switch matter a little less.
So we take a deliberate stance: keep the frontier of machine autonomy in the realm of ideas, text, analysis, and simulation — where error is reversible — and keep a human hand on anything that moves, builds, or touches the physical world. Fewer autonomous robots. Fewer internet-connected mechanical devices empowered to act on the world without a person in the loop. Not because the physical world does not matter, but because it matters far too much to automate carelessly.
Containment is not fear of the technology. It is the discipline that keeps the technology governable — the reason an off-switch still means something.
The off-switch only works while the machine stays on our side of the screen.
Commitment 02
Turn the best intelligence to defense — first.
If attackers gain AI, defenders cannot be left with yesterday’s tools. The same intelligence that could threaten the digital infrastructure civilization now runs on — banks, exchanges, records, whole institutions made of pure information — should first be put to work defending it.
We are building toward intelligent security: adaptive, learning defenses that make a break-in genuinely hard. Not a static wall an intruder can study at leisure, but a system that learns how its own network behaves, questions what looks wrong, and tightens in real time — so an intrusion is caught while it is happening, not discovered months later.
And a response that acts. When something hostile is found inside a system we defend — a financial institution’s network, say — detection alone is not enough. Our aim is an automated response that contains the breach in the moment: isolating the compromised ground, revoking the intruder’s access, and cutting their foothold off from the rest of the protected system — while preserving the evidence and handing it to the institution’s own people and the proper authorities.
Think of it plainly as a guard and a response working together: make the door hard to open, and, if someone gets through anyway, remove them from the building and lock it behind them. It acts inside the systems we are entrusted to protect, under the institution’s authority and the law — defense of our own ground, never retaliation beyond it.
Make the break-in hard. If it happens anyway, end it fast — and keep the receipts.
One idea, two directions
Both commitments are the same conviction seen twice: the most capable intelligence should stay where humans can still govern it — contained to the digital world, and turned first toward protecting the people who live alongside it.
This is the safety work of AGI Research Lab, and the reason the think tank exists around it: to design that future in the open, and to keep the record of the human thinking that shapes it. Read the manifesto.
Testing phase — this page describes a direction we are building toward, not a finished product.