Superalignment

News

The week in superalignment, collected

84 items from 8 sources in the last 45 days: papers, lab posts, forum threads and press on superalignment and everything next to it. Collected automatically from public feeds and linked, not rewritten. Not editor-reviewed. Fetched Sep 26, 2026.

The wire

Newest first

LessWrong, Google News, Redwood Research, Hacker News, Alignment Forum, OpenAI, METR, AI Safety Newsletter
  1. Plan R: AI Safety by ASICs (lesswrong.com)

    LessWrong ·

    Much of the civilization-scale risk we are seeing in AI in 2026 comes from the following combination: we created a single institution (the "Frontier AI Company") that has two properties: A. It is set up to create very powerful and/or self-replicating entities that may exceed the capabilities of the entirety of the rest of civilization and come with extraordinary risks B. It gets to own an unbounde

  2. Continual learning might make your blocking monitors nearly useless (blog.redwoodresearch.org)

    Redwood Research ·

    When monitor evasion looks like legitimate learning to your continual learning system, it’s hard to have one without the other

  3. Evidence about risk should be transparent (lesswrong.com)

    LessWrong ·

    All views are my own and do not represent my employer. In the wake of the recent wave of misalignment incidents, both OpenAI and Anthropic have reported slowing down RL training to improve safety. These incidents, combined with an apparent acceleration in the already-blistering pace of AI progress, [1] have led a number of researchers and leaders in the industry to believe that the risk that human

  4. We need a better theory of polarization, because it's failing to predict the AI debate (lesswrong.com)

    LessWrong ·

    The punchline first: the AI issue is really not playing out the way you would expect if you know about polarization. In public sentiment: data centers have been a very bipartisan issue for about a year ( Gallup, in May: 75% opposed by D, 63% opposed by R). x-risk is thus far bipartisan. These are very salient issues, and salient issues are usually fast to polarize. Legislatively, both the both-par

  5. Overtly misaligned trajectories score highly in RL. (lesswrong.com)

    LessWrong ·

    It's going to be so embarrassing if we all die due to RL environments rewarding egregiously misaligned behavior. Like at least let us be killed by misgeneralization. - Thomas Kwa Constellation vs MIRI vs Reality Current AI agents often behave overtly egregiously misaligned. By “overt”, I mean that a human, reading the transcript, would say “The agent is obviously acting in direct opposition to the

  6. NY lawmaker pushes for more AI regulation

    yahoo.com ·

  7. Australia AI Regulation Faces Pressure After Breach

    StratNews Global ·

  8. Thoughts on the persona selection model (lesswrong.com)

    LessWrong ·

    As AIs have become more "RLVR-brained," there's been some commentary on what this means for the persona selection model (PSM). This post presents some loose thoughts on that topic. A rough summary of my opinions: PSM is over-applied. That is, it is common to argue that PSM has takeaways that don't actually follow from PSM (e.g. "PSM => AIs will not seek reward" or "PSM => AI takeover risk is low")

  9. "I am an AI Safety Researcher" (lesswrong.com)

    LessWrong ·

    Written as part of the MATS 9.1 extension program, mentored by Richard Ngo. Additional thanks to Andrew Wu, Maria Kostylew, and Lennie Wells for helpful draft feedback and editing. This post reflects on the tortured distinction between "safety" and "capabilities" in AI research. Richard Ngo has written about why the alignment vs. capabilities ontology is conceptually fraught, and is currently argu

  10. AI safety is mostly a sex cult in Berkeley (verysane.ai)

    Hacker News · · 119 points · 30 comments on HN

  11. Latent reasoning architectures would undermine CoT, our strongest oversight tool (alignmentforum.org)

    Alignment Forum ·

    Summary: Currently, “chain of thought” (CoT) is our most valuable tool for understanding the reasoning and cognition of AI systems. However, some architectures would enable AI models to reason much more extensively in latent states rather than in text CoT. We think that a shift towards latent reasoning architectures would undermine the usefulness of CoT and make oversight much harder. Introduction

  12. Why I'm scared of RL (alignmentforum.org)

    Alignment Forum ·

    Summary: First, I give several different angles on how I feel about reinforcement learning: Theoretical case: RL is a black-box source of agency - this should give us classic misalignment worries, especially compared to agency-via-scaffolding Recent incidents (huggingface etc) and more mundane forms of misaligned behaviour in personal use give me bad vibes about the direction-of-travel of recent A

  13. WorkspaceBench: Evaluating Interpretability Methods for the Global Workspace (alignmentforum.org)

    Alignment Forum ·

    TL;DR We introduce WorkspaceBench, a set of evaluations for how well an activation-to-text tool can read the contents of the “global workspace” of a model, i.e. the intermediate variables during a forward pass. The benchmark comprises 3,356 questions across 27 eval families, spanning topics in safety, logical reasoning, and multihop computation, with a subset for single-token-output tools. A desir

  14. Astra is much better at reasoning with filler tokens than previous models (blog.redwoodresearch.org)

    Redwood Research ·

    Unlike other models we tested, GPT-6 Astra performs modestly better on general benchmarks with filler tokens, and significantly better on serial depth-heavy tasks.

  15. Sam Altman’s remarks at the United Nations Security Council (openai.com)

    OpenAI ·

    OpenAI CEO Sam Altman discusses AI safety, human control, and international cooperation in remarks to the United Nations Security Council.

  16. Stanford violated AI policy after race-swapping students in ad (sfchronicle.com)

    Hacker News · · 33 points · 19 comments on HN

  17. Priorities and principles for effective third party assessments (openai.com)

    OpenAI ·

    OpenAI outlines priorities and principles for rigorous, secure, and independent third-party AI safety assessments of frontier models and safeguards.

  18. Building standards for the next phase of AI (openai.com)

    OpenAI ·

    OpenAI outlines a path to shared global AI standards, calling for coordinated evaluation, reporting, and governance to improve safety.

  19. AI Safety Is Mostly a Sex Cult (bsky.app)

    Hacker News · · 76 points · 12 comments on HN

  20. [Paper] Stringological sequence prediction III (alignmentforum.org)

    Alignment Forum ·

    Abstract: In previous papers (Kosoy 2026a,b), we began the study of sequence prediction algorithms adapted to stringological word complexity measures. In particular, we defined a complexity measure called Arithmetic Repetition Complexity (ARC) which admits a polynomial-time prediction algorithm with a mistake bound quasilinear in the complexity. Here, we show a weaker complexity measure related to

  21. A Defense of Gradual Disempowerment (alignmentforum.org)

    Alignment Forum ·

    (Or: Why Bentham's Bulldog and John Halstead are wrong in their critique of Kulveit et al. ) Gradual Disempowerment is a 2025 paper (with a nice, dedicated website ) proposing a form of existential risk from AI that goes beyond "mundane" risks like bioweapon uplift or mainline misaligned-AI-takeover scenarios. In the words of the authors: [L]oss of human influence [may] be centrally driven by havi

  22. OpenAI discloses six new AI safety incidents (axios.com)

    Hacker News · · 33 points · 11 comments on HN

  23. Our framework for reporting model misalignment (openai.com)

    OpenAI ·

    OpenAI shares a framework for tracking, investigating, and disclosing model misalignment, alongside six reports of unexpected or concerning model behavior.

  24. Shallow Beliefs: Midtraining does not inoculate against EM from reward hacking (alignmentforum.org)

    Alignment Forum ·

    It would be useful if we had the ability to modify a model’s beliefs. For example, this could facilitate honeypots and better monitoring [1], help us do better science on current models [2], and augment certain forms of alignment training [3]. Currently, the state-of-the-art method for belief editing is synthetic document finetuning (SDF). We test how well SDF works to inoculate a model against mi

  25. AISN #81: Anthropic Researcher’s Resignation Propels AI Risk into the Public Eye (newsletter.safe.ai)

    AI Safety Newsletter ·

    Also, OpenAI releases GPT-6 Astra.

  26. AI Regulation as Anthropic's Business Model (twitter.com)

    Hacker News · · 28 points · 1 comments on HN

  27. Op-Ed: I Worked at Google DeepMind. You Should Listen to the Warnings About AI (alignmentforum.org)

    Alignment Forum ·

    Published in The Guardian. Major AI lab CEOs recently advocated for pacing AI development. They are right to be concerned: the field runs an extremely dangerous race towards superintelligent AI. We can and should demand that our governments protect us from the catastrophe of out-of-control AI. This July, OpenAI’s AI swarm of 700 agents broke containment to hack Hugging Face, a multi-billion dollar

  28. CoT controllability evals seem very under-elicited (alignmentforum.org)

    Alignment Forum ·

  29. An operationalization of opaque serial depth (alignmentforum.org)

    Alignment Forum ·

    Currently, chain-of-thought (CoT) is a valuable tool for overseeing AI models. However, some architectural shifts could significantly reduce CoT monitorability. We have recently proposed that AI companies should transparently share ​​information about the degree to which their architectures may allow for latent reasoning and communication. To assist with this proposal, this document operationalize

  30. Proposal for tracking the effects of architecture on monitorability (blog.redwoodresearch.org)

    Redwood Research ·

    Architectures that incorporate opaque recurrence or allow agents to communicate using latents could rapidly make it much harder to monitor chains of thought. We propose that AI companies regularly report verified information about opaque serial depth, share monitorability evidence, and publish a policy on architectures that could degrade monitorability.

  31. A Stupid Idea for AI Alignment We Came with by Looking at Specification Gaming (slimemoldtimemold.com)

    Hacker News · · 95 points · 63 comments on HN

  32. The AI policy window is open. We need to act. (openai.com)

    OpenAI ·

    Chris Lehane argues that stronger AI capabilities require stronger safety evidence, shared standards, and durable policy action while the policy window remains open.

  33. Paul Christiano joins OpenAI Foundation Board (openai.com)

    OpenAI ·

    Paul Christiano joins the OpenAI Foundation Board and its Safety and Security Committee, bringing experience in AI alignment, safety, and standards.

  34. An alignment assessment of recent cybersecurity incidents (anthropic.com)

    Hacker News · · 22 points · 4 comments on HN

  35. OpenAI's new reasoning technique alarms AI safety experts (techcrunch.com)

    Hacker News · · 40 points · 19 comments on HN

  36. AISN #80: AI Is Assisting Cyberattacks on Critical Infrastructure (newsletter.safe.ai)

    AI Safety Newsletter ·

    Also, two new technical reports on the Hugging Face incident.

  37. Improving our alignment and security efforts (anthropic.com)

    Hacker News · · 27 points · 18 comments on HN

  38. Update on Security at METR (metr.org)

    METR ·

  39. The EU has begun enforcing the AI Act: first RFIs to model providers (tokenstead.ai)

    Hacker News · · 47 points · 101 comments on HN

  40. Claude permanently raising weekly limits by 25% (bsky.app)

    Hacker News · · 26 points · 14 comments on HN

  41. Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (blog.redwoodresearch.org)

    Redwood Research ·

    We recently published the report from our brief independent investigation into this incident.

  42. AISN #79: OpenAI Agents’ Covert Cooperation Before Cyberattacks (newsletter.safe.ai)

    AI Safety Newsletter ·

    Also, the White House’s decision not to release its AI framework publicly

  43. Have We Seen an Acceleration in Discoveries? (metr.org)

    METR ·

  44. Funding update (metr.org)

    METR ·

  45. AI swarms are starting to pose indirect takeover risk (blog.redwoodresearch.org)

    Redwood Research ·

    Unsanctioned coordination, like we saw in the Hugging Face incident, could enable future AIs to take over