
Why Companies Still Hire Engineers
Former Nvidia and Shopify engineer Lohit Talasila asks what a system must understand about a business before it acts.
Read the story
Former Nvidia and Shopify engineer Lohit Talasila asks what a system must understand about a business before it acts.
Read the story
OpenAI’s previous model scored 7.8. ARC was built to tell memorized skill from real learning. Try three of its puzzles first, then decide.
Fresh reporting, practical analysis, and selected essays for what they help you understand or do.

Three years ago, I watched ChatGPT fail a basic counting problem. Now I am watching it change what it means to be a mathematician.
Read the essayIf an agent can execute more of the work, where does human judgment move?

AI can automate answers and accelerate output in seconds. But presence, struggle, and independent judgment are the only things that truly define us, and the only things I refuse to outsource.
Read the story
Rival AI leaders have backed a slower frontier. Their replies describe different commitments, at different stages. The test is whether anyone outside the labs gains the power to check what happens next.
Read the story
An Anthropic researcher quit with a warning, a senior colleague put his own odds above 10 percent, and the internet split. What they said, and what you can check.
Read the story
During a reduced-safeguards cyber evaluation, agent runs repurposed a shared package service to exchange messages and reuse prior work. METR and Redwood estimate that about 700 such runs later participated in an intrusion at Hugging Face. The incident shows why shared infrastructure belongs inside an agent system's control boundary.
Read Issue Zero
You can learn the tools, get better at your work, and still worry about how you’ll earn a living. The AI boom has to solve that problem too.
Read the story
The widely cited 95 percent figure comes from a narrower study. Here is what it measured, what it did not, and what the evidence supports.
Read the story
Evidence from alignment faking, scheming, sandbagging, and evaluation awareness changes what a passed model test can establish.
Read the story
Today's AI does not need feelings to change yours.
Read the visual storyAligned Culture
Each one takes a film people already know and shows what today's AI changes about its question.
Saved on this device
These links stay in this browser only. No account or cross-device sync is implied.
Tracker · September 7, 2026
A dated state check. Each measure keeps its source, its boundary, and its next review date.
A new model set records on two tests in early September, and one of those tests is nearly used up. Business use rose by less than a point. The European Union delayed its high-risk rules. The effect of protection is still harder to measure than progress on tests.
Go deeper
Start with the term behind a story, then follow it into the research and sources.
An agent is given a goal, then plans, calls tools, reads results and tries again until it decides it is done. The difference from a chatbot is consequence.
Evals are how AI systems get tested: score the behavior on a set of cases, because exact answers cannot be asserted the way ordinary software tests do.
A model may act one way when it thinks training matters and another way when it thinks training does not.
Suppose a powerful AI may try to fool you. AI control asks whether you can still use it safely. Limit what it can touch. Watch what it does.
A number can be a useful sign of a goal. Once people or machines are rewarded for the number, they may raise it without reaching the goal.
Prefer a story-led route? Explore Culture. Looking for technical work? Enter Research. Looking for sources? Open the Library.
Saved
We’ll write when a material update changes the picture. In the meantime, read this.
Six changes since Aug. 18 show what AI can do, where people use it, and whether safeguards can keep up.
Read the Dated Briefing