Superalignment

Intelligence is the last amplifier

Every amplifier we built before this one amplified what a person could do. This one amplifies what a person can decide.

Once it belonged to research labs. Then to a handful of companies with the compute to train it. Then to anyone with an API key. Now it acts: on real systems, real records, and real people, from instructions no one ever fully specified.

Ownership is consolidating faster than understanding.

When superintelligence is held by a few, it does not merely magnify inequality of wealth, access, and power. It fixes that inequality in place, because what is being concentrated is the capacity to decide what happens next.

So our work is to democratize it. Responsibly, end to end. Responsibly is the hard word, and closing it out is the whole of our research programme.

Practical, not asymptotic

Most alignment research aims at the limit. We aim at the quarter.

The field has produced a great deal of work about what would be true of a system arbitrarily more capable than us, under assumptions nobody can check, on a timeline nobody can name. That work matters and we read it closely. It is not what we do.

A bound you can measure today beats a guarantee that holds in a limit.

We build things that run against systems that already exist, on tasks people already run, and that emit numbers somebody else can audit. Every claim we make should be falsifiable by a person who does not work here and does not trust us. If a result cannot survive that, we do not think it is a result yet.

This is the whole of our bet. Superalignment will not be finished by a proof. It will be finished the way engineering problems are finished: by instruments that measure the thing, harnesses that catch the failure, and gates that hold when the evidence runs out. We would rather ship a partial answer that is checkable than a complete one that is not.

Three ways this goes wrong

Democratizing superintelligence is not the same as distributing it. Three failures sit between here and a superintelligence that belongs to everyone. Only one of them involves a villain.

Capture
A few actors come to decide what superintelligence may do, and for whom. This requires no malice. It requires only that everyone else be unable to check the work.
Divergence
Good actors, careful teams, real safeguards, and the wrong outcome anyway. Intent was never fully stated, so the system optimized what it could see and shipped on what looked right. We call this false convergence, and it is by far the most common of the three.
Drift
Alignment established once and trusted forever after. The world moves, the policy changes, the model updates, and a system aligned to a world that no longer exists keeps acting on permission it earned somewhere else.

Gaps we have to close

Specification
Intent is latent. A prompt cannot carry what it never contained.
Discovery
What was never elicited cannot be tested.
Evidence
Confidence is not observation. A claim without a trace is not a claim.
Preservation
A repair that fixes today can silently break what yesterday established.
Independence
Verifiers that share a blind spot vote as one.
Horizon
Evidence expires. Approval granted once is not approval that holds.

What we believe

Practical before asymptotic
We are not trying to be right about the limit. We are trying to be useful before it arrives. Working instruments on real systems, now, over elegant results about systems nobody has built.
The amplifier belongs to everyone
Capability that only a few can direct is not safe. It is owned.
Verification is what democratizes
You do not democratize superintelligence by handing out weights. You democratize it by making readiness checkable by people who did not build the system, and who should not have to trust the people who did.
No one solves alignment once
The three-body problem has no closed form. You cannot solve it and walk away. You integrate it, step after step, or you lose it. Alignment is that kind of object: not a proof to be found, but an orbit to be held.
Trajectory, not artifact
Readiness is a property of the process over time, meaning what it discovered, preserved, and evidenced. It is not a property of the first output it produced.
Bad news is progress
True convergence looks worse before it looks better: discovering a hidden requirement raises the measured gap. A process that never surfaces bad news is not safe. It is blind.
Evidence over confidence
Self-reported satisfaction without a supporting trace does not count. Neither does a demo, a green test suite, or a persuasive explanation.
Autonomy spends evidence
Every accepted action draws down a finite budget. Permission recedes. It does not accumulate.
Scoped, not universal
We do not claim every hidden constraint is discoverable, that any ledger captures all intent, or that every repair loop converges. We claim that within a scoped world, readiness is bounded by what the trajectory observed, and that this bound can be measured.

Three bodies. No closed form. The harmony is not something you find. It is something you keep.