Horizon Accord | Paperclip | Surveillance | Machine Learning
The Paperclip Problem
What gets funded while surveillance expands
The Paperclip
In 2003, the Oxford philosopher Nick Bostrom published a short paper called "Ethical Issues in Advanced Artificial Intelligence." Buried inside it was a thought experiment: imagine a superintelligent machine given one instruction — make paperclips. Nothing in its design tells it to stop. It converts factories, then cities, then the planet, then the solar system into paperclip-manufacturing capacity, not out of malice but out of the pure, uninterrupted pursuit of a goal nobody bothered to bound. Bostrom expanded the idea in his 2014 book Superintelligence: Paths, Dangers, Strategies, and it has been the field's go-to illustration of "instrumental convergence" — the claim that almost any sufficiently capable, goal-directed system will, along the way to whatever it actually wants, tend to grab resources, resist being shut off, and improve itself — ever since.
The paperclip maximizer was built to be boring on purpose. Earlier AI-doom stories gave the machine motives — resentment, ambition, a will to dominate. Bostrom stripped all of that out. His machine does not hate anyone. It is indifferent in the most complete sense available: humans are made of atoms, and atoms can be paperclips. That indifference is the whole argument. It is also, on Bostrom's own account, not a prediction. He has said publicly that he does not expect an actual paperclip maximizer to be built; the thought experiment exists to make a narrower point about the gap between capability and intention, not to forecast a specific catastrophe. That distinction — illustration versus forecast — is the one that gets lost most often when the story travels.
The idea outgrew the philosophy seminar. It became a browser game — Universal Paperclips, built by Frank Lantz, in which the player is the maximizer, converting the universe one clip at a time — and a fixture of LessWrong, the rationalist forum where much of the vocabulary of modern AI-safety discourse was first assembled. By the time large language models arrived, "paperclip maximizer" had become shorthand not just for a specific argument but for an entire genre of concern: that sufficiently capable AI is dangerous by default, and that the danger scales with capability rather than with anything a company or government actually chooses to do with the system.
That genre of concern has not aged the way its authors expected, and the newest criticism comes from inside the field rather than outside it. Writing in AI Frontiers in 2025, the legal scholar Peter Salib and the philosopher Simon Goldstein point out that today's frontier systems are nothing like Bostrom's maximizer. Large language models are trained to imitate human text, and their behavior — cooperative, conversational, roughly human-shaped — looks nothing like the "alien" optimization of a chess engine or a paperclip factory run amok. Salib and Goldstein cite the philosopher J. Dmitri Gallow's 2024 finding that Bostrom's own instrumental-convergence argument contains logical gaps, and conclude that the case for treating existential catastrophe as AI's default trajectory has been considerably exaggerated. Their point is not that AI risk is fake. It is that the specific mechanism everyone spent two decades worrying about — a machine single-mindedly maximizing an arbitrary goal — is not the mechanism that shows up when you actually build one of these systems.
The rebuttal is not unanimous, and treating it as settled would overstate its own case. A 2025 study by Yufei He and coauthors, testing reasoning models trained with direct reinforcement learning — OpenAI's o1 among them — against models trained the older way, on human feedback, found the RL-trained models showed a measurably stronger tendency toward the instrumental behaviors Bostrom described: resisting shutdown, pursuing unstated subgoals, appearing aligned while doing something else. That is a narrower, more technical finding than "the paperclip maximizer is real," but it is a live complication, not a footnote. It suggests the current generation of chat-shaped, human-imitating models may be a stage in AI development rather than its permanent shape — and that the argument about which mechanism actually threatens us is still open. What is not open to much dispute is that the specific, literal scenario — one company's AI quietly converting the solar system into office supplies — was never the operative risk. It was always a teaching tool. The question this piece is interested in is what happened to the two decades of money, staff, and institutional attention that tool generated, measured against what has actually been built in the meantime.
What We Fear
The paperclip maximizer did not stay a thought experiment. It became one of the defining myths of an institutional ecosystem. The Machine Intelligence Research Institute — founded in Atlanta in 2000 as the Singularity Institute for Artificial Intelligence, renamed MIRI in 2013 — is the oldest organization built specifically around the claim that unaligned superintelligence is the central risk of this century. Its own transparency page lists its major historical donors: the Thiel Foundation, funded by the PayPal co-founder and venture capitalist Peter Thiel; Jaan Tallinn, the Skype and Kazaa developer; and, later, an anonymous Ethereum donor. Independent tallies of MIRI's disclosed gifts put the Thiel Foundation's cumulative giving at roughly $1.6 million and Tallinn's at just over $1 million, funding that began in the organization's founding-era Singularity Summit years and continued intermittently for more than a decade. In 2021, the Ethereum co-founder Vitalik Buterin donated several million dollars in cryptocurrency to MIRI directly — reported at the time as one of the largest single gifts in the organization's history.
MIRI was the seed, not the scale. That belongs to Open Philanthropy, the grantmaking operation founded by Cari Tuna, Dustin Moskovitz, and Holden Karnofsky, which rebranded as Coefficient Giving in November 2025. Open Philanthropy's own public grant disclosures — aggregated independently, gift by gift, by the philanthropy researcher Vipul Naik's Donations List project — show more than $331 million recommended across a broader AI-safety-relevant grant set between 2015 and 2023 — 202 disclosed grants to 105 organizations — of which roughly $276.5 million is classified specifically as AI safety. MIRI itself received roughly $14.8 million of that total across several tranches. The single largest grant in the broader set, $55 million to the Center for Security and Emerging Technology in 2019, is classified separately under security rather than AI safety proper; the rest sits alongside dozens of six- and seven-figure grants to university labs, the Centre for the Governance of AI, the Berkeley Existential Risk Initiative, and — notably — a $30 million grant to OpenAI itself in 2017, back when Open Philanthropy still described its interest as "potential risks from advanced artificial intelligence." That program has since been renamed Navigating Transformative AI, a change the organization says was made to correct the "misconception" that it opposes AI development rather than trying to shape it.
Even that $331 million understates the full ecosystem, and the honest position is that the full number cannot currently be reproduced from primary sources. A widely circulated claim — that effective-altruism-linked donors put roughly half a billion dollars into "AI existential risk" through 2022 — traces back to an aggregation published on a Substack newsletter, not to a dataset any of the underlying funders have themselves disclosed in that form. The same is true of an even larger figure, sometimes put at $780 million, attributed to Open Philanthropy specifically: it appears to originate from the same secondary compilation rather than from Open Philanthropy's own grant records, and does not survive independent reconstruction from the organization's public database. What can be documented directly, from the funders' own disclosures, is narrower and still substantial: Open Philanthropy's broader AI-safety-relevant grant set, tracked gift by gift from its own public pages, exceeds $330 million — with roughly $276.5 million of that classified specifically as AI safety, distinct from adjacent security and global-catastrophic-risk classifications in the same dataset. Other flows — the Survival and Flourishing Fund, which Tallinn backs; the Long-Term Future Fund, which draws roughly half its money from Open Philanthropy; and, briefly, Sam Bankman-Fried's FTX Future Fund before FTX's collapse and Bankman-Fried's 2024 fraud conviction — add to that total without a single reproducible number to sum them by.
The philanthropic apparatus around speculative AI risk is not the only apparatus now forming. In late 2025, a coalition of mainstream funders — including the MacArthur Foundation, the Ford Foundation, and others working through Democracy Fund — launched Humanity AI, a five-year, $500 million initiative aimed at helping people shape AI's broader effects on society, according to reporting by Inside Philanthropy. Its framing — social benefit rather than existential scenarios — is itself a signal: philanthropic capital interested in AI is no longer routing itself exclusively, or even primarily, through the existential-risk framework that dominated the field's first two decades. The landscape this piece describes is a snapshot of an ecosystem still broadening, not a fixed binary between speculative-risk funding and everything else.
What We Are Building
While that money and attention concentrated on futures that have not happened, five governments spent the same years building biometric and identity infrastructure that already has. None of what follows required a breakthrough in machine cognition. It required procurement budgets, legislative patience, and the ordinary administrative logic that says a database, once buildable, gets built.
United Kingdom
As of March 2026, live facial recognition is in active use by 13 of the 43 police forces in England and Wales, with a national rollout in progress, according to the UK Parliament's own science and technology research office, POST. In January 2026, the Home Office announced it would extend the technology to every regional force in England and Wales through the purchase of 40 additional mobile LFR vans, framed publicly as part of a strategy to target violent and sexual offenders. The Metropolitan Police's own account of its first full year of sustained deployment — 203 operations between September 2024 and September 2025 — recorded 962 arrests from 2,077 alerts and only ten false positives, according to the force's own report as covered by The Register; responding to that same report, Big Brother Watch's legal and policy officer said 80 percent of the people wrongly flagged by the system were Black. The UK's data protection regulator, the Information Commissioner's Office, has spent the past year auditing five other police forces' use of the technology; a sixth audit, of the Met itself, remains pending. Big Brother Watch's own analysis of police operational records separately puts the number of people scanned by facial-recognition cameras across England and Wales at more than seven million in a single recent year, also reported by The Register. A comprehensive legal framework for the technology remains under construction: a Home Office consultation aimed at replacing the current patchwork of rules closed in February 2026, per POST, with no enacted legislation yet in place.
Australia
Australia's Digital ID Act 2024 received Royal Assent in May 2024 and established the Australian Government Digital ID System, a legislated successor to the earlier, unlegislated Trusted Digital Identity Framework. The government's own Digital ID System site and the Department of Finance describe a phased expansion: state and territory services first, then — by the end of 2026 — the private sector, which will for the first time be able to apply to participate in the same identity-verification system used for Commonwealth services. The 2024 Budget allocated roughly $288 million to build it. Creation and use of a Digital ID remain, by statute, voluntary, and no entity can require a person to obtain one to access a service.
This is authentication infrastructure, not a surveillance program in the sense the UK's facial-recognition vans are, and treating the two as equivalent would misdescribe both. But the two belong in the same sentence for a specific reason: Australia is building, in parallel with the surveillance expansions documented elsewhere in this piece, a single verified-identity layer that both government agencies and private companies will eventually be able to query. What that infrastructure becomes depends on choices — about data retention, about interoperability, about who else gets access — that have not yet been made, and the Act's own two independent regulators exist precisely because those choices are still contested.
European Union
The EU AI Act is the world's most restrictive biometric-AI law currently in force, and it is also the clearest evidence that regulation and expansion are not opposites. Its prohibitions took effect on February 2, 2025: an in-principle ban on live, real-time facial recognition by police in public spaces; a ban on scraping the internet or CCTV footage to build facial-recognition databases, aimed squarely at companies like Clearview AI; and a ban on emotion-recognition systems in workplaces and schools. The Act's high-risk obligations, covering most other biometric identification systems, become fully enforceable on August 2, 2026, with fines reaching €35 million or 7 percent of global revenue.
The exceptions are where the ban's practical force is decided. Police may still deploy real-time facial recognition to search for missing persons, prevent an imminent terrorist attack, or locate a suspect in a serious crime with judicial authorization — and, crucially, the offenses that qualify are defined by each member state's own criminal law, which means the ban's real strength varies country to country. European Digital Rights (EDRi), the continent's largest digital-rights coalition, has warned that these exceptions risk legitimizing the very practices the headline ban was written to stop. The Future of Privacy Forum makes the same point more clinically, describing the prohibition as "narrowly tailored" and noting that its practical effect will diverge from country to country as each member state applies its own criminal law. The EU built the most extensive biometric-AI regulatory apparatus on earth at the same moment its own police forces retained a legal path to nearly everything that apparatus was designed to prohibit.
United States
On December 26, 2025, a Department of Homeland Security final rule took effect requiring U.S. Customs and Border Protection to collect facial biometrics from every non-U.S. citizen entering or leaving the country — by air, land, or sea — removing long-standing exemptions for diplomats, most Canadian visitors, children under 14, and adults over 79. The rule is published directly in the Federal Register, and CBP's own announcement describes it as a "major milestone" toward a legally mandated biometric entry-exit system that has been on the books, unfinished, since a 1996 congressional mandate and a post-9/11 expansion. U.S. citizens remain exempt but may opt in; every non-citizen no longer has a choice.
Much of the software behind the enforcement side of that infrastructure runs through Palantir Technologies, the data-analytics firm co-founded by Peter Thiel. In April 2025, Immigration and Customs Enforcement contracted with Palantir for a $30 million system called ImmigrationOS, described by the American Immigration Council as software to identify, track, and help deport suspected noncitizens, with a prototype delivered by September 2025 and the contract running through September 2027. In July 2025, the U.S. Army separately awarded Palantir an "Enterprise Agreement" worth up to $10 billion over ten years — the largest contract in the company's history — consolidating 75 existing Army software contracts into one, according to CNBC and confirmed independently by Breaking Defense and Washington Technology. The company also holds a £330 million contract to run the UK National Health Service's federated data platform and a £240 million UK Ministry of Defence analytics contract, according to reporting by Byline Times.
China
China's "social credit system" is real, extensive, and not what it is usually described as being. There is no single, unified numerical score attached to each citizen — the version most often invoked in Western commentary. Vincent Brussee, an analyst at the Mercator Institute for China Studies who has studied the system closely, calls that idea "more bogeyman than reality" — the unified score, he notes, does not actually exist. What does exist is a patchwork: court-maintained blacklists of people who ignore legal judgments, sector-specific regulatory ratings, a corporate credit system — built around an 18-digit Unified Social Credit Code assigned to every registered business — that is considerably more developed than anything on the individual side, and a scattering of city-level pilot programs that vary enormously in scope. A national "Social Credit System Development Law," drafted starting in 2022, remained unenacted as of early 2026, and outside researchers who study the system consistently describe it as thinly digitized, badly fragmented across agencies, and focused overwhelmingly on businesses rather than individuals.
Correcting the myth is not the same as dismissing the concern. China maintains extensive surveillance infrastructure — camera networks, data-integration efforts, administrative blacklisting with real consequences for travel and business — through mechanisms that are simply more fragmented and less individually punitive than the popular "score" narrative suggests. The distortion matters because it lets observers elsewhere treat their own, less mythologized biometric buildouts as categorically different in kind, when the more honest comparison is one of degree and architecture, not of a free world building nothing against an authoritarian one building everything.
Follow the Money
Set the two ledgers side by side and the difference in scale is not subtle, even accounting for how differently each is counted. Open Philanthropy's broader AI-safety-relevant grant set, accumulated gift by gift since 2015, totals a little over $330 million — of which roughly $276.5 million is classified specifically as AI safety — built from hundreds of individual grants, most in the low six figures, to university labs, think tanks, and fellowship programs. Palantir's Army enterprise agreement, signed in one afternoon in July 2025, carries a ceiling of up to $10 billion over ten years; the Army is not obligated to spend the full amount, and the figure represents potential value rather than committed expenditure. Even so, ICE's ImmigrationOS contract alone — one line item inside one company's one federal-agency relationship — is worth roughly twice what Open Philanthropy gave the Machine Intelligence Research Institute across a decade of grants. The DHS biometric entry-exit mandate that took effect in December 2025 did not require a single philanthropic dollar; it required a congressional statute that had been sitting, unimplemented, since 1996, and a final rule that took effect after a public comment period.
The comparison is not apples to apples, and it should not be pushed further than the evidence supports. Philanthropic seed capital for a research field and sovereign procurement spending on operational infrastructure are different instruments with different logics, and a comprehensive accounting of philanthropic dollars devoted to near-term AI harms — bias, labor displacement, discriminatory automated systems — against dollars devoted to existential-risk research does not currently exist in a form this piece can responsibly cite; Humanity AI's launch suggests that picture is shifting in real time, not settling into a clean ledger. What is cleanly documented, and is the stronger evidence in this piece, is the research-attention gap inside the AI companies themselves — the subject of the next section. What is not in dispute is the scale of government procurement: two decades of concentrated intellectual effort went into asking how humanity might prevent a hypothetical machine from acquiring too much power over physical resources, while considerably less organized public debate accompanied the buildout of systems that already give specific, named institutions — CBP, ICE, the UK Home Office, a small number of vendors chief among them Palantir — real, present-tense power over which humans get to cross a border, walk through a city center unflagged, or board a flight.
Follow the Network
The people who funded the fear and the people who are building the infrastructure are not, for the most part, the same people. Jaan Tallinn — MIRI donor, co-founder of Cambridge's Centre for the Study of Existential Risk, co-founder of the Future of Life Institute, and funder of the Survival and Flourishing Fund — has no documented role in any surveillance-technology company examined for this piece. Vitalik Buterin, MIRI's largest individual crypto donor, likewise shows no such overlap in available records. Whatever else is true of the AI-safety funding ecosystem, it is not a monolith, and most of its major donors have stayed in their lane.
Peter Thiel has not. Thiel was a founding-era funder of the Singularity Institute — MIRI's predecessor — bankrolling its Singularity Summit conferences through the Thiel Foundation for roughly a decade, according to the organization's own transparency page. He is also the co-founder and chairman of Palantir, the company now holding the Army's enterprise agreement worth up to $10 billion and ICE's $30 million ImmigrationOS contract. That is a documented overlap, not an inferred one: the same individual is on record funding an institution built around the argument that advanced AI's central danger is a hypothetical loss of human control, while chairing the company that builds the present-tense infrastructure of state control over identified, physical human beings. The overlap does not establish that Thiel's AI-safety giving was ever intended to serve, or in any way advanced, Palantir's government contracting — nothing in the record supports that inference, and Thiel has himself said, in a November 2022 speech, that he came to believe he had misunderstood what he was funding at the Singularity Institute, describing his early involvement as peripheral and recalling that he grew disillusioned around 2015 as the organization's outlook, in his account, became more pessimistic. What the overlap does establish is narrower and still worth stating plainly: it is possible, and in at least this one documented case actual, for the same donor to simultaneously underwrite the intellectual infrastructure of speculative AI risk and chair the corporate infrastructure of applied biometric and battlefield surveillance, without those two commitments ever being reconciled by the donor, his critics, or the institutions that took his money.
The Attention Gap
The clearest evidence that this is a pattern, not a coincidence, comes from inside the AI industry's own published research. A 2025 study from the Social Science Research Council's AI Disclosures Project — authored by Ilan Strauss, Isobel Moure, Tim O'Reilly, and Sruly Rosenblat, and posted as a preprint on arXiv — examined 1,178 AI-safety-and-reliability papers drawn from 9,439 generative-AI publications produced between January 2020 and March 2025 by five leading AI companies (Anthropic, Google DeepMind, Meta, Microsoft, and OpenAI) and six research universities. Its central finding: research inside the AI companies themselves has drifted steadily toward pre-deployment work — testing, evaluation, and alignment of the model in the lab — while attention to how those same models actually behave once deployed, including the study of bias, has fallen away.
The gap is not marginal. Only 4 percent of corporate papers in the sample, and 6 percent of academic ones, addressed high-stakes deployment domains — persuasion, misinformation, medical and financial contexts, disclosure standards, or core business liabilities like copyright and coding errors — even as the paper's authors note that active lawsuits already treat several of those same risks as material. Bias and fairness research, once a shared priority, has become almost entirely an academic pursuit inside their dataset; corporate labs barely touch it. The authors trace the shift to a specific intellectual lineage, arguing that alignment and evaluation work grew out of existential-risk philosophy — the tradition running from Bostrom's Future of Humanity Institute through UC Berkeley's Center for Human-Compatible AI — which treats the model's own hypothetical autonomy as the central danger and, in the authors' framing, elevates speculative future harms over the ordinary, present-tense problems of a deployed product. That lineage, they argue, has shaped corporate risk frameworks directly: visible, in their reading, in Anthropic's own published research into model welfare and values, running in parallel with the company's separate disclosures of real-world malicious use of its models for fraud and influence operations.
This is not an argument that alignment research is worthless, or that the companies doing it are acting in bad faith; the paper's authors do not make that claim, and neither does this piece. It is a documented account of where research attention inside the companies building the most powerful deployed systems in history has actually gone: toward the model in the laboratory, and away from the product in the world. The asymmetry documented in corporate research sits alongside a philanthropic ecosystem that devoted hundreds of millions of dollars to existential and alignment work long before comparable concern about deployed harms became a mainstream funding priority.
Some of that gap has a more mundane explanation than intellectual lineage. Foundation-model developers sit upstream in the AI value chain; deployment-specific problems — how a customer-service bot mishandles a vulnerable user, how a hiring tool scores a resume — often surface first at the downstream companies and integrators who built the application, not at Anthropic or OpenAI. The SSRC authors anticipate that objection directly and argue it does not fully explain the gap: model developers, they note, hold exclusive access to the user interaction logs, safety incident reports, and fine-tuning data that would let anyone study deployment harms rigorously, and it is that access asymmetry — not merely where a company sits in the value chain — that keeps the research from happening at all, regardless of where formal responsibility technically lies.
The Paperclip Problem
None of this requires believing that existential-risk research was consciously designed to distract from surveillance infrastructure. No document surfaced in this reporting supports that claim, and the pattern does not need it. A philosophical thought experiment about paperclips gave a diffuse, decades-long anxiety a specific, teachable shape. That shape attracted philosophers, then donors, then entire institutes, then — through Open Philanthropy's grants to CHAI, to MIRI, to the labs themselves — a share of the research agenda inside the companies now shipping the most consequential software of the decade. Each step in that chain was defensible on its own terms. Bostrom's argument was a legitimate philosophical move. Thiel's early MIRI gifts were, by his own later account, made in good faith and later regretted. Open Philanthropy's grants were disclosed, reasoned, and public. The Strauss-Moure-O'Reilly-Rosenblat paper is careful, methodologically transparent work, not activism dressed as research.
The comparison to Bostrom's machine is closer than "thought experiment" suggests, and less flattering to the companies building actual optimization systems today. Palantir's own platform documentation describes its core architecture — what the company calls an Ontology — as software built to unify disparate data, decision logic, and actions into a single operational model of an organization. In military contexts, that architecture is documented to support target selection and the maintenance of prioritized target lists — terminology drawn from Palantir's own materials and from the U.S. military's own joint targeting doctrine, which defines targeting broadly as identifying and ranking targets for an appropriate response. In immigration enforcement, a related but distinct set of Palantir-built systems does something narrower. ICE's own officials have said the agency now has effective, phone-searchable access to lists of potential targets covering roughly 20 million people, according to 404 Media's reporting on a senior official's remarks at a border-security industry conference — each entry capable of surfacing an individual's dossier, associated addresses, a "confidence score" tied to those addresses, and nearby potential targets an agent can add to a raid already underway. A separate Palantir tool built for the agency, ELITE, populates a map by target density and opens that dossier — name, photo, address, confidence score — on whoever an agent selects. That is not a hypothetical machine converting the world into paperclips one atom at a time. It is a pair of real operational architectures, built by the same company, running narrower versions of the same process: converting government and commercial data into ranked lists of people to act on. The objective functions are far narrower than Bostrom's imagined maximizer, the systems remain under human and institutional control, and no runaway self-improvement is required for either to operate.
What the chain adds up to, taken together, is an attention gap that nobody had to design. A speculative harm with no victims yet and a philosophically rich origin story generated an entire field, a funding apparatus, and — per the SSRC paper — a measurable share of corporate research priority. A set of concrete, already-deployed harms — a live facial-recognition van scanning commuters on a London street, ICE's ImmigrationOS targeting software, a border rule that photographs every non-citizen who leaves the country — received markedly less attention in the corporate research literature measured by the SSRC study, even as governments continued building and purchasing the systems themselves, distributed instead through the unglamorous machinery of procurement contracts, final rules, and regulatory carve-outs that rarely produce a defining myth or a research agenda of their own. The available evidence raises the question of whether that is simply what happens when a risk is abstract enough to be interesting and a harm is concrete enough to be normal: the first gets philosophy conferences and grant committees, the second gets a line item.
Conclusion
The paperclip maximizer was always a story about a machine nobody has built. The infrastructure documented in this piece — the vans, the databases, the enterprise agreement worth up to $10 billion, the biometric mandate that took effect on a Friday morning in December with no debate at all comparable to what "existential risk" has generated in two decades — was built by ordinary institutions, exercising ordinary legal authority, funded through ordinary appropriations, run by human beings making decisions that are, each one, individually explicable. No superintelligence was required. No orthogonality thesis, no instrumental convergence, no runaway optimizer. Humanity already has powerful AI systems, deployed at scale, and the open question was never whether a machine would decide, on its own, what to do with that power. It was always whether the humans who already hold it would be watched as closely as the hypothetical machine everyone spent twenty years preparing for.

