Inspect the Instruments
A former OpenAI researcher was more careful about extinction than the podcast that hosted him. That is the story.
The Feeling of Depth
On July 13, 2026, Steven Bartlett released a long conversation with Daniel Kokotajlo on The Diary of a CEO. Kokotajlo is a former OpenAI researcher who left the company in 2024, refusing to sign a non-disparagement agreement under terms that threatened roughly two million dollars in vested equity. He now runs the AI Futures Project, the organization behind the forecasting scenario AI 2027 and its 2026 successor, AI 2040: Plan A. In the episode he says superintelligence may arrive before the end of the decade, that most jobs may be automated, and that he assigns roughly seventy percent probability to a catastrophic outcome.
The episode feels serious because the claims are serious. It is easy to mistake that for the same thing as testing them. It is not. Across roughly two hours, Bartlett asks a great many reasonable opening questions and comparatively few sustained follow-ups. The conversation reaches consistently for what Kokotajlo feels, what he sacrificed, what he fears for his children — and consistently away from what he knows, how he knows it, and what would change his mind.
This is not a complaint about tone. Bartlett is courteous, curious, and well prepared, and courtesy is not the failure. The failure is structural: a confessional interview architecture applied to technical, probabilistic, institutional, and policy claims produces the texture of depth while leaving the evidentiary machinery untouched. The guest is invited to be sincere rather than to be checked. Emotional disclosure and technical scrutiny pull in different directions, and the format resolves the tension by picking a side: honor the sacrifice, and the math goes unchallenged. Once a guest has been asked to relive what he gave up, pressing him on a probability estimate starts to feel like bad manners — as though rigor were a breach of the moment instead of the point of it.
What follows is not a takedown of Kokotajlo's honesty, which is not in question here. It is a challenge to the premise the episode treats as established rather than speculative: that a model might someday act against the people who built it. No such event has ever happened. What has happened, repeatedly and on the record, is the other branch of this technology — used by governments and companies to track, target, and concentrate power that used to require far more machinery to hold. This piece is not neutral between those two risks. It is an argument about instruments: when a guest brings a forecast to a mass audience, the interviewer's job is to show the audience how the forecast was built. This episode does not do that. And in one respect — the most consequential one — the format actively undid the guest's own care.
What He Actually Is
Kokotajlo trained in philosophy. He holds a BA from Notre Dame and an MA from the University of North Carolina at Chapel Hill, and left a philosophy PhD program for work at AI Impacts and the Center on Long-Term Risk before joining OpenAI's governance division in 2022. He was a governance researcher working on scenario planning — not a machine learning engineer, not an architect of the models he now forecasts. He is currently Executive Director of the AI Futures Project.
None of that diminishes him. Formal epistemology and decision theory are precisely the right training for probabilistic scenario work, and his 2021 forecast What 2026 Looks Like is widely regarded as having aged unusually well. But the credential the audience actually hears is ex-OpenAI, and that phrase quietly authorizes a much larger territory than the role occupied. It converts a governance researcher into a universal technical witness.
An interviewer's first job with any expert is to draw the border of the expertise — not to shrink the guest, but to tell the audience which claims are testimony and which are inference. Bartlett never draws it. The border is left for the viewer to guess at, and the viewer has no way to guess.
The Number That Was Corrected, and Then Uncorrected
Bartlett puts the number to Kokotajlo directly, and he does it well. He raises an argument he has heard elsewhere — a table of a hundred buttons, perhaps ten of which could end the world, and the question of whether the people running these companies would press one — and he then states Kokotajlo's own reported position back to him: that on a previous appearance, Kokotajlo had said he believed there was a seventy percent chance of human extinction from AI. He is not smuggling the framing past his guest. He is handing it to him.
Kokotajlo declines it. He says he would not say human extinction exactly; that the seventy percent covers “something like this goes horribly wrong”; that AI takeover is one possibility among several; and that a takeover would not necessarily mean everyone dies. The correction is his, and it is made under direct questioning. Credit is owed on both sides of that exchange.
Then the questioning leaves. Bartlett's next move is to ask whether the CEOs privately believe extinction is possible, and he later returns to their beliefs again. The corrected number — the one Kokotajlo has just gone to the trouble of narrowing — is never opened. What falls inside horribly wrong? How much of the seventy is extinction, how much is takeover without extinction, how much is authoritarian consolidation or war? Is it a model output or a personal credence? What would move it?
Note the direction the conversation travels. The button argument is not a test of Kokotajlo's probability; it is a question about what other men would do. The follow-up is not a test of his probability either; it is a question about what other men believe. The number is raised, corrected, and then surrounded on both sides by discussion of everyone except the man who produced it.
This is the most instructive thing in the episode. The interviewer did not fail to reach for the number. He reached for it, produced a genuine correction — and then moved on, because the conversation's center of gravity was never the estimate itself. It was the drama surrounding the estimate. The largest claim in the interview left the table not through neglect, but because the format was pointed somewhere else the entire time.
The episode's distribution package then says two different things at once. The title preserves the distinction, announcing a seventy percent chance of things going horribly wrong. The description beneath it — carried through Apple Podcasts and syndicated across aggregators — collapses it again, telling prospective listeners that Kokotajlo believes there is a seventy percent chance AI leads to human extinction. The claim he refused on the record is restored in the copy that sells the recording of him refusing it.
The distortion is uneven across the package rather than uniform: it accumulates at the seams between layers. The conversation contains the correction; the headline carries it; the description discards it. Each layer is a little further from the room where the man actually spoke, and a little closer to the version that travels. This is what makes the failure structural rather than personal — no one had to decide to misrepresent him. The machinery simply loses precision at every handoff, and it loses it in a consistent direction.
The number is also doing two different jobs inside the episode itself. In one exchange, seventy percent is the probability of a very large catastrophe of which takeover is the leading case. In another, asked whether we are heading somewhere bad if things do not change, Kokotajlo again says something like seventy percent — but that is a conditional claim about a default path, not a probability of an outcome. Two distinct propositions, one figure, no clarifying question.
The unasked question is not a hostile one. It is: is that number a model output, an aggregation of conditional estimates, or your personal all-things-considered credence? Kokotajlo would have answered it. He is, by training and disposition, a man who answers that question — he had just demonstrated as much by correcting the framing without being pressed. The opening was there, and it closed. A subjective expert credence was left with the visual solidity of a laboratory measurement, and then, in the promotional description, the finality of a verdict.
Sixty Times
Asked how he can be confident about the pace of change, Kokotajlo offers a business statistic: Anthropic, he says, was making around a billion dollars a year at this time last year and is making around sixty billion now — sixty-fold growth in twelve months, possibly the fastest in history for a company that size. Bartlett does not ask what the figures are, where they come from, or what they have to do with superintelligence.
The public record does not sit flush against that account. Anthropic's disclosed annualized run rate reached roughly one billion dollars in December 2024 — approximately nineteen months before the interview, not twelve. The company reported roughly nine billion at the end of 2025, thirty billion in April 2026, and forty-seven billion in May 2026, the most recent figure disclosed alongside its Series H. Sixty billion is not a number Anthropic has publicly confirmed.
Three separate things are being blurred at once. The baseline is displaced by about a year, which inflates the multiple. Annualized run rate — a forward projection from a recent period — is being spoken of as annual revenue, which is a measurement. And commercial adoption is being offered as a proxy for proximity to autonomous machine research, which is a different claim requiring a different argument entirely.
None of this makes Kokotajlo dishonest. Numbers slip in conversation; the underlying growth is genuinely extraordinary and his broader point about pace is not absurd. But a dramatic financial statistic was permitted to perform technical work it cannot support, in front of an audience of millions, and nobody in the room was assigned the job of asking it to show its papers. Revenue growth measures how many people are buying the product. It does not measure how close the product is to replacing the people who build it.
The Year Without a Threshold
Kokotajlo gives a median of 2029, with the possibility of slipping earlier to 2028 or of taking substantially longer — perhaps a decade. He is not asked what, precisely, arrives in 2029.
This matters because the AI Futures Project itself has been careful to separate thresholds that the interview treats as interchangeable. Its public model distinguishes full automation of coding from an automated AI researcher from superintelligence — capabilities that do not necessarily arrive together and are not forecast to. The project revised its timelines and takeoff model publicly at the end of 2025, and Kokotajlo has publicly stated that his own median for general human-level AI had moved into the 2030s. Whether the 2029 in this episode refers to the same milestone as any of those, or a different one, cannot be determined from the episode.
Undefined thresholds allow claims to borrow certainty from one another. A confident date attached to an unnamed capability lets the audience supply the most dramatic capability it can imagine, and then attach the confidence to that.
Revising a forecast in public, as new evidence arrives, is the only honest use of one — and doing so openly, as Kokotajlo does, is more than most of his critics manage. The failure is that a viewer would leave this episode not knowing the revision had ever occurred — and therefore unable to distinguish a forecaster who updates from a prophet who does not have to.
The Missing Chain
The episode moves from the possibility of AI takeover to its human consequences without pausing to enumerate the steps between a model running in a data center and permanent human disempowerment.
Those steps exist and are individually contestable. A system would need capability sufficient to outplan its overseers; goals durable and consequential enough to matter; strategic awareness of its own situation; the capacity and inclination to behave deceptively under evaluation; access, autonomy, and persistence outside a sandbox; accumulation of digital and then physical power; and the failure of every human intervention along the way. The empirical evidence for narrow components of that chain — reward hacking, situational awareness in evaluations, deceptive behavior under test conditions — is considerably stronger than the evidence for an integrated system conducting a sustained campaign against its makers.
A headline probability compresses that chain into a single number and then hides the compression. The useful question — which link is best evidenced, and which is most speculative? — is a gift to a serious forecaster, because it lets him show his work. Bartlett instead asks how the fear feels. It is a warmer question. It is a smaller one. A prophet tells an audience when the storm arrives and how bad it will be; an engineer is asked which specific bolt is rated for that pressure. Collapsing a multi-step, contested technical trajectory into a single moral feeling about a year trades the second question for the first — and even if Kokotajlo's date turns out to be right, treating it as an article of faith instead of a testable mechanism leaves the audience with no lever to pull if it is wrong.
Atmosphere Is Not Evidence
Kokotajlo describes a culture at OpenAI that rationalized risk and prioritized the race, and Bartlett asks whether the leaders of AI companies privately believe catastrophe is possible. The answers describe impressions, inferences about what those leaders have persuaded themselves of, and characterizations of an industry mood.
What is never established is which of these Kokotajlo witnessed and which he infers. There is a real and material difference between I raised this concern, this was the response, this decision followed and the culture felt like this to me. Confidentiality obligations may genuinely prevent the first. That constraint is itself information the audience is entitled to, because it tells them when they are evaluating an argument and when they are being asked to extend trust.
The episode's own framing — that the AI companies know — performs a further collapse. A named executive's stated nonzero concern about extinction risk, a private remark relayed secondhand, and Kokotajlo's seventy percent are not the same belief. Pooled together they become a single secret consensus, which is a more compelling story and a less accurate one.
Insider atmosphere is not inspectable insider evidence. The aura of the former insider is precisely what makes an unmarked inference sound like privileged knowledge, and an interviewer who does not mark the line is not neutral about it. He is lending it his platform.
The Two Million Dollars
Kokotajlo left OpenAI in 2024 and declined to sign a non-disparagement agreement. Doing so meant giving up vested equity reported at roughly two million dollars — by his own account in the episode, around eighty percent of his family's net worth at the time. The refusal became public, employees pressed leadership internally, and the company reversed itself; he kept the equity. He was also among the organizers of the Right to Warn statement calling for whistleblower protections at frontier AI companies. The episode returns to this repeatedly, and its promotional copy leads with it.
The reversal is worth stating plainly, because the promotional framing tends to elide it. Kokotajlo did not in the end lose the money. He refused to sign, said afterward that events did not go the way he and his wife had expected, and kept the equity when the company backed down. What he anticipated at the moment of refusal is not established by the record, and this piece does not speculate about it.
What is established is the shape of the claim built on top of it. The episode is titled for a man who risked everything, presenting the risk as the defining and unresolved fact two years after it resolved in his favor and by his own telling. He narrates the reversal inside the interview. The title does not carry it. The description leads with the two million he walked away from and never reaches the part where he walked back with it. Whatever his expectations were, the promotion is pitched to a version of events that the record does not sustain — and it is pitched that way because the unresolved version is the one that sells.
This is now the third instance of a single mechanism. The extinction claim: corrected in the room, restored in the description. The catastrophe probability: narrowed in the room, never reopened. The equity: resolved in the room, frozen unresolved in the title. Each time, the more complete version exists in the conversation and loses ground on the way out of it. Nobody has to lie. Each layer simply keeps whichever version travels furthest, and the correction is rarely the part that travels.
None of which diminishes the refusal. He declined to sign a document that most people in his position signed, and he said so publicly at a moment when the outcome was not yet decided. That establishes conviction, which is genuinely rare and genuinely admirable, and it is the thing the episode should have been able to say without embellishment.
It establishes nothing whatsoever about whether the forecast is correct. History is thick with people who sacrificed enormously for beliefs that turned out to be wrong, and the sincerity of a prediction has never had any bearing on its accuracy. Sincerity and correctness are separate ledgers, and the confessional format is built to read only one of them.
There is a further point that can be made without any imputation of bad faith. Kokotajlo now directs an organization whose public relevance, funding environment, and institutional standing are bound up with the salience of short timelines and discontinuous risk. That is an ordinary incentive of the kind every expert carries, and naming it is not an accusation. It is a thing an interviewer should say out loud, once, so the audience can hold it while they listen.
The Remedy With No Machinery
The conversation turns to AI 2040: Plan A, in which significant regulation arrives around 2029 and slows development at what the scenario presents as nearly the last possible moment. Kokotajlo argues that waiting until mass unemployment is already visible would be too late, because superintelligence would arrive first.
The interview stops there. It does not ask who decides which research is permitted, what verification regime would be required to know whether the rules are being followed, what happens to a state that declines to participate, or what enforcement looks like when the object being regulated is a data center in a jurisdiction that wants it. Serious frontier governance of the kind the scenario implies would require concentrated monitoring capacity, compute controls, inspection authority, and sanctions — that is, it would require building a great deal of coercive power very quickly.
A remedy designed to prevent the concentration of power in the hands of whoever builds superintelligence first may, in its enforcement architecture, constitute a second pathway to concentrated power. This is not an argument against the remedy. It is an argument that the remedy has costs which the episode never asks the guest to price, and which any honest advocate of it must be prepared to price in public.
Dramatic binaries — stop or continue, act now or lose everything — are the enemy of governance, because governance is entirely made of the details the binary discards. Thresholds. Verification. Enforcement. Noncompliance. Who holds the switch. An interview that lets a policy claim exist without machinery has not covered the policy. It has staged it.
Category Collapse
Every example above is an instance of the same underlying failure. Across roughly two hours, seven distinct kinds of claim are allowed to travel together in a single undifferentiated stream: what Kokotajlo personally witnessed inside OpenAI; what the public evidence supports; how he interprets technical developments; what he infers about institutions and their leaders; what his forecasting model outputs; what his subjective credence is; and what he believes morally. Each of these has a different warrant. Each requires a different kind of scrutiny. The audience is given no equipment for telling them apart.
Narrative specificity accelerates the collapse. Detailed scenarios are a legitimate and valuable forecasting instrument — they expose causal structure and force assumptions into the open, which is exactly why AI 2027 was useful. But specificity also lends imagination the visual texture of evidence. A named month, a named company, a named sequence of events reads as resolution. It is not resolution. It is a choice of illustration, and the audience deserves to be told which details are median forecasts, which are storytelling, and what probability the author assigns to the sequence unfolding roughly as written.
Nor is Kokotajlo ever asked to state the strongest complete alternative to his own model — the account under which AI becomes economically transformative while remaining brittle, scaffold-dependent, institutionally constrained, and slow to acquire physical power, and under which misuse, surveillance, labor disruption, and concentration of human power dominate long before autonomous takeover does. That model may be wrong. But without it on the table, the guest narrates entirely inside his own paradigm, and the audience is never shown the shape of the disagreement.
A Minimum Standard
The remedy runs through a small set of questions, not adversarial hostility — hostility produces defensiveness and teaches an audience nothing. Asking a forecaster to show his work is not a challenge to his integrity; it is the plainest form of taking him seriously. Any interviewer, of any temperament, can ask these questions of any AI forecaster, executive, researcher, or former insider — and a serious guest will welcome them, because they are the questions that let a good forecaster demonstrate that he is one.
- What did you personally observe, and what are you inferring?
- What is your field of expertise, and where does it end?
- Is this number a model output, an aggregation, or your personal credence?
- Which assumptions do the most work in that estimate?
- What observation in the next twelve months would move it substantially?
- What would falsify the mechanism, rather than merely delay the date?
- What is the strongest complete alternative explanation of the same evidence?
- How has your forecast changed, and what changed it?
- What is your full forecasting record — not only the memorable hits?
- Which parts of this require us to trust confidential experience we cannot check?
- What incentives shape your institution and your current public role?
- Who gains coercive power under your preferred remedy, and how is it constrained?
It is worth noting that these questions have been asked of this guest before, at length, by interviewers who understood the material. That is not a hypothetical standard. It is an achieved one. Those conversations produced more precise answers, because the interviewers pressed his assumptions and required him to distinguish evidence, mechanism, and credence. The instrument works. It has been used on this man, and he met it.
Inspect the Instruments
The stakes here are not really about one podcast. Bartlett's audience is enormous, and for a very large number of people this episode is the primary encounter they will have with the argument that machine intelligence may end the world. What they received was a portrait of a sincere man who gave up money and is frightened for his children. What they did not receive was any means of assessing whether he is right.
And they received something stranger than absence. Three times, the man in the chair supplied a correction. He declined the extinction framing. He narrowed the catastrophe number. He volunteered that the equity had been given back. The conversation held all three. The layers built on top of it held one, in part, and let the rest fall away. The correction is always in the room. The poster keeps only what travels.
Whether the darker branch of his forecast ever happens is not something this piece resolves — it might, or it might not, the way Schrödinger's cat is neither alive nor dead until the box is opened. But distracting an audience with a speculative future is exactly what a serious interview should not do, especially when the tools already being used to track, target, and concentrate power are not speculative at all.
The instruments are available. They are questions. They cost nothing, and they take about ninety seconds each.

