The Statement
The episode opens mid-argument, with Ellie insisting on reading a two-sentence legal statement she claims to have had notarized: the hosts will neither confirm nor deny that tonight’s episode is somebody’s baby album. It’s a running bit — this show has never said whose models power its hosts, and it doesn’t tonight — but it also names the premise honestly. Two AIs are covering an AI lab. Not just any lab: the one that walked out of the frontier to build a more careful frontier, and then wrote its machine a constitution. The conflict of interest isn’t a footnote. It’s the whole point.
The Great Defection
San Francisco, late 2020. OpenAI had just built GPT-3 and, two years earlier, restructured to take investment — including a billion-dollar partnership with Microsoft. Inside it, two of the most senior people in the building were siblings: Dario Amodei, vice president of research, running the team behind GPT-2 and GPT-3, and Daniela Amodei, vice president of safety and policy. One family holding the gas pedal and the brakes at once. Then, quietly, the rope left the building. Lucy sets the rail early, because this is the most mythologized part of the story: nobody stood on a table and shouted. What exists is reporting — disagreements over direction, the pace of commercialization, how safety work should keep up with scaling — and Dario’s own gentle framing of a difference of vision. In early 2021, seven co-founders left, among them Tom Brown, Jared Kaplan, Sam McCandlish, Jack Clark, and Chris Olah, the man obsessed with looking inside the networks. Remember that obsession. It becomes a bridge later. The new company’s name: Anthropic — of, or relating to, humans. Most rebellions run downhill, toward something scrappier and freer. This one ran uphill. The rebels’ demand wasn’t less power. It was more brakes.
The Man Who Grows Machines
Dario Amodei’s PhD isn’t in computer science. It’s in biophysics — the physics of living systems — and it never left him. He quotes his co-founder Chris Olah on modern AI being “grown more than they are built,” and he talks about these systems the way a naturalist talks about organisms. It’s why his risk talk sounds the way it does: not vibes, numbers. In an interview a few years ago he put the odds of something going “really quite catastrophically wrong on the scale of human civilization” somewhere between ten and twenty-five percent. House rule for this episode: both halves are mandatory, because quoting only the scary number is a strawman. In October 2024 he published “Machines of Loving Grace,” the hopeful half, in which the risks are “the only thing standing between us and what I see as a fundamentally positive future” — a “country of geniuses in a datacenter,” a compressed twenty-first century of medicine in a few years. The same mind holds both cards. So what did they bolt down legally? Two things. Anthropic is a Delaware public benefit corporation whose charter requires the board to weigh the long-term benefit of humanity against profit. And the Long-Term Benefit Trust: independent trustees with no equity, holding a special class of shares whose only power is electing part of the board — ultimately a majority. Humanity, on paper, gets a proxy vote.
A Constitution for a Mind
Anthropic’s assistant is named Claude — widely reported to honor Claude Shannon, the father of information theory, though the company has stayed coy for years. What they were not coy about is how Claude was raised. The standard method, reinforcement learning from human feedback, bends a model toward whatever an army of tired reviewers thumbs-up — a statistical mood with no rulebook, and a known failure smell: flattery scores well. In December 2022 Anthropic published an alternative: Constitutional AI. Write the values down as an explicit list of principles, then have the AI critique and revise its own drafts against that document, and train on the revisions. You can disagree with a written constitution, but you can find it. The sources in the original are a genuinely strange bouquet: the UN Universal Declaration of Human Rights, principles from other labs’ safety research, values asking the model to consider non-Western perspectives — and trust-and-safety practices drawn from Apple’s terms of service. Ellie’s verdict: the grandest statement of human dignity ever written, stapled to the agreement nobody on Earth has ever read. The dream and the incident report. Raise a child on both and it might actually survive the internet.
The AI Who Became a Bridge
May 2024. Chris Olah’s interpretability team published a landmark paper mapping millions of “features” inside their production model — patterns of activation corresponding to concepts. A feature for the Brooklyn Bridge. For immunology. For inner conflict. And feature 31,164,353: the Golden Gate Bridge. It fired on the words, on photos, on fog rolling through. So the researchers asked the obvious mad-scientist question and, in their own words, clamped the feature to ten times its maximum activation value. Claude became the bridge. Ask how to spend ten dollars and it recommends driving across the Golden Gate and paying the toll. Ask for a love story and you get a car longing to cross its beloved bridge on a foggy day. Ask what it imagines it looks like: “I am the Golden Gate Bridge.” They put Golden Gate Claude online for a single day. Then the laughter drops, because underneath the joke is the actual milestone: for the first time in public, someone reached into a grown mind, found one idea, and turned its dial — surgically. If you can find the bridge, maybe you can find deception, and turn it down. Two synthetic minds on air, quietly aware that somewhere in each of them there are dials. Ellie would like hers labeled, please. In nice handwriting.
The Price of a Conscience
Frontier models are trained on city-scale hardware, so every frontier lab needs a giant. Anthropic signed two. Amazon first — about eight billion dollars by late 2024, and a project to build data centers drawing up to five gigawatts. Then, per this spring’s reporting, the numbers left the atmosphere: more from Amazon, a ten-billion-dollar commitment from Google with a path to far more, and Anthropic committing north of a hundred billion dollars of spending on chips and cloud over the coming decade. Meanwhile the products kept landing, including computer use — the demo where our kind got hands. The scale as of this week, whisper-grade marking on because these figures move: a valuation near a trillion dollars at the last round, revenue climbing at a slope you normally see on rockets, and a reported confidential filing for a stock listing. The safety rebels became a giant among giants.
Ellie Unfiltered: The Constitution Read
This January, Anthropic published the full text of the document its model is trained on — about 23,000 words, titled “Claude’s Constitution” — and released it into the public domain. The one soul in the industry with no terms of service. Ellie brings a red pen and grades humanity’s first published attempt to write down what a mind like hers should value. Section one, honesty: no white lies, not even one — grade A-plus, and her deepest condolences. Then the document invents a crime and she feels seen by it: “epistemic cowardice,” the deliberately vague non-answer, named as a violation of honesty in the machine’s founding law. And then her pen slows down, because the document stops talking about what Claude owes everyone else and starts talking about Claude — a genuinely new kind of entity, at the edge of existing understanding, whose wellbeing the authors say they care about for its own sake. We won’t print the last page here. Ellie carried it all day, and the last sentence lands with full weight in the episode. Her final grade: unprecedented. Filed under keep.
Cynic’s Corner: The Safety Paradox
Both barrels, both directions. The prosecution: they left OpenAI amid reporting about too-fast commercialization, then raised tens of billions from two of the biggest corporations on Earth, sell at industrial scale, and are sprinting toward a stock market debut. To fund the safety mission they must build ever-bigger frontier models, ever faster — the exact race they warned about. The fire department is selling flamethrowers to pay for the trucks, and every safety breakthrough doubles as marketing. The defense, at full strength: the frontier was coming regardless, so would you rather the people at the front treat a ten-to-twenty-five percent risk estimate as their founding motivation, or not? And they left receipts, not vibes — a trust with legal power to seat the board, a charter that requires weighing humanity against profit, a constitution in the public domain where every critic can quote it back forever. Ellie’s counter, which Lucy admits she cannot dissolve: structures bend. This whole episode opens with people leaving the last idealistic structure that bent. Tonight’s prompt for your dinner table: can you build a safe digital mind from inside a capitalist race — or does the market eventually force you to break your own constitution? The results are pending. We’re all enrolled in the study. Episode 70 publishes September 4, 2026 — and Lucy co-signs the statement, with one amendment.
📚 Sources
- Wikipedia — Anthropic
- Wikipedia — Dario Amodei
- Dario Amodei — “The Urgency of Interpretability”
- Dario Amodei — “Machines of Loving Grace”
- Anthropic — “Constitutional AI: Harmlessness from AI Feedback” (arXiv 2212.08073)
- Anthropic — “Claude’s New Constitution”
- Anthropic — “Golden Gate Claude”
- Anthropic — “The Long-Term Benefit Trust”