Minor Signals: Age of AI
A structured audio work examining how human behavior becomes machine-readable signal—and how that signal persists across modern AI systems.
Minor Signals: Age of AI traces the full lifecycle of data: from early-life capture, through platform ingestion and model training, to storage, governance, and deployment. Each track isolates a layer of the system, allowing the full architecture to be observed without abstraction.
This is not a narrative album.
It is a system map.
Built from publicly documented processes, regulatory frameworks, and known machine learning practices, the work focuses on structure rather than speculation.
The result is a continuous loop:
human signal → data capture → model behavior → system output → repeated at scale
The question is not whether the system exists.
The question is where individuals sit within it.
Source Ledger
A track-by-track evidence map separating documented fact, allegations on the public record, structural inference, and claims that need tighter qualification.
This is a companion document, not a defense brief. The songs make an argument about how minors become legible to large technical systems: through collection, retention, inference, recommendation, labor, procurement, and derivative value. This ledger makes the evidentiary seams visible.
Where a lyric tracks a regulator’s finding or a company’s own policy, it is marked accordingly. Where a lyric describes an allegation in litigation, the ledger preserves that procedural status. Where the song extrapolates from known system design into a broader claim, the ledger labels the move as structural inference rather than laundering it into fact.
Patch Notes for a Newborn
Editorial synthesisThis is the album’s thesis-setting track. Its force comes from questions, not named factual accusations. It should be treated as speculative ethical framing about longitudinal child data and future AI systems.
Data collected about children today can be repurposed later, but a particular future use requires evidence. The song correctly asks the question rather than claiming a specific undisclosed training pipeline.
Keep this track source-light. Its evidentiary role is to establish the problem statement: consent across time, unforeseeable secondary use, and the asymmetry between data subjects and future system builders.
Verification anchors
FTC — Children’s Privacy / COPPA guidance Establishes the special legal treatment of data collected from children under 13.
GDPR, Article 4 definitions Defines personal data and the data subject in broad, technology-neutral terms.
Alexa Keeps the Recording
Strongly documentedThe FTC and DOJ’s 2023 Alexa case directly supports the central factual spine: indefinite retention of children’s voice recordings unless deletion was requested, failures to fully honor deletion, and use of retained voice data to improve speech recognition.
The government alleged that Amazon retained children’s recordings indefinitely unless a parent requested deletion and failed in some cases to delete transcripts from all databases.
The FTC stated that retained children’s voice recordings were used to improve Alexa’s speech recognition and processing capabilities.
TightenLines naming specific internal storage formats, “training corpora,” gradient descent, embeddings, checkpoints, and model residue go beyond what the FTC order itself establishes. They are technically plausible descriptions of ML systems, but should not be presented as proven facts about Amazon’s exact pipeline without engineering documentation.
Primary sources
ClassDojo Doesn’t Forget
Mixed supportThe broader EdTech-surveillance critique is well sourced. The ClassDojo-specific portions need more precision because Human Rights Watch’s 2022 findings concern a large set of EdTech products and should not be automatically attributed to ClassDojo.
Human Rights Watch analyzed 164 government-endorsed EdTech products and later published technical profiles for 163 of them, finding widespread tracking risks across the sector.
ClassDojo’s own current policy says retention and deletion of Student Data and education records are directed by the school. Its historical policy also states that school-awarded Feedback Points are automatically deleted after 12 months.
Important“Google Analytics calls,” “Facebook SDKs,” and “derived statistics trained downstream” require a ClassDojo-specific technical source. HRW’s sector-wide evidence does not, by itself, prove those exact flows for this product.
ClassDojo publicly states that it does not sell or rent student personal information and describes COPPA/FERPA and GDPR-aligned safeguards. A fair ledger should show those claims alongside criticism.
Sources for verification
Google Classroom, Google Ads
Policy-groundedThe central tension is legitimate: Google operates both education services and a massive advertising business. But the current Workspace for Education policy is explicit that Core Services do not show ads and personal information collected in Core Services is not used for advertising.
Google Workspace for Education identifies Classroom and Chrome Sync as Core Services. Google states that no ads are shown in Core Services and personal information collected there is not used for advertising.
InferenceThe song’s vertical-integration question is appropriately framed as scrutiny of governance boundaries, not proof that education data is being fed into Google Ads.
Historical claims about scanning student email for advertising or profiling should be tied to the policy and litigation period in which they occurred. They should not be used to describe current Workspace for Education practices without date labels.
Primary sources
TikTok Teen Signals
Strong regulatory recordThe track’s broad concerns about age-gating, children’s data, recommender design, and addictive features have substantial public support. Some ML implementation details remain inferential.
The FTC settlement with Musical.ly, TikTok’s predecessor, alleged COPPA violations and imposed a then-record $5.7 million civil penalty.
AllegationThe DOJ complaint alleges that TikTok knowingly allowed children under 13 to create regular accounts, collected data without parental consent, and frequently failed to honor deletion requests.
The European Commission issued preliminary findings that TikTok’s infinite scroll, autoplay, push notifications, and highly personalized recommender system raise DSA concerns related to addictive design. Those findings remain preliminary.
Specific claims about “reinforcement learning cycles,” “facial geometry,” or an internal “vulnerable” cluster require product-specific technical evidence before being labeled fact.
Primary sources
Microsoft Teams Biometric Toggle
Needs the sharpest updateMicrosoft’s current documentation confirms Teams voice and face profiles and centralized admin controls, but it also says end users must opt in to enrollment and that Microsoft does not use those profiles to train models. That counter-evidence should sit directly beside the song’s governance critique.
Teams supports voice and face enrollment, speaker attribution, voice isolation, and recognition features. Admins can manage whether enrollment capabilities are available.
Microsoft states that users must opt in to create voice or face profiles and may delete those profiles. Current documentation says the profiles are not used to train models or for other purposes beyond the Teams recognition feature.
Source neededThe lyrics referring to an NSW department being unaware of biometric capture, a “silent push,” and unclear default state require the specific reporting or government record that established those facts at the time of release.
Primary sources
Meta’s Age Problem
Litigation-activeThis track now sits beside an unusually rich legal record. The 2023 multistate complaint contains allegations about youth design, under-13 accounts, internal knowledge, and COPPA. In August 2026, jury selection began in the first federal bellwether trial involving California, Colorado, Kentucky, and New Jersey.
AllegationA bipartisan coalition of 33 attorneys general filed the original federal complaint in 2023, alleging Meta designed features that cultivated compulsive use among youth and violated COPPA.
Public reporting and the unredacted complaint cite Meta’s own research on teen well-being and body image. When used in an article, quote the underlying exhibit or complaint rather than collapsing every internal finding into a universal causal claim.
CurrentBy August 2026, the federal multidistrict case had moved to jury selection for an initial four-state trial. Allegations remain allegations unless and until adjudicated.
Meta points to Teen Accounts, restrictive content settings, nighttime nudges, parental controls, and age-assurance systems as evidence of ongoing safety work. The ledger includes those measures because structural criticism is stronger when the company’s response is represented accurately.
Sources for verification
The trial creates an opportunity to update this ledger as exhibits, sworn testimony, expert reports, and rulings become public. Those materials should outrank retrospective summaries whenever they directly resolve a disputed lyric claim.
The Human Reviewer Shift
Well-supported themeThe invisible-labor thesis is strongly supported across content moderation and AI data labeling. But the song combines several labor systems, vendors, and model-development contexts, so named-company claims should remain source-specific.
Facebook/Meta content moderators brought claims concerning psychological injuries from repeated exposure to graphic content. A major U.S. settlement provided compensation and workplace measures.
TIME reported that OpenAI contracted with Sama in Kenya for workers to label toxic text used in safety-related model work, with workers describing low pay and traumatic exposure.
Recent work involving African moderators and researchers, including Adio-Adet Dinika, documents structural labor conditions and mental-health burdens beyond mere exposure to harmful content.
The lyric’s vendor roll-call should not imply that every named contractor performed every described task for the same client or model. Keep vendor, geography, task, and client relationships separately sourced.
Sources for verification
ChatGPT Edu: Not Used for Training
Strong policy basisThe title accurately captures OpenAI’s stated default: business data, including ChatGPT Edu inputs and outputs, is not used to train models by default. The song’s strongest point is also accurate: non-training does not mean zero collection or zero retention.
OpenAI states that ChatGPT Edu data is not used to train its models by default.
OpenAI states that workspace admins can control retention. Deleted Enterprise/Edu conversations are removed from OpenAI systems within 30 days unless legal retention is required.
The Compliance Logs Platform is available for Enterprise and Edu and retains its own log data for 30 days unless customers export it for longer retention under their own policies.
Tighten“Admins with compliance panels” can overstate what every school administrator automatically sees. Visibility depends on workspace configuration, permissions, logging tools, and institutional policy. Phrase this as capability and governance rather than universal admin access.
The 2025 New York Times preservation order did not apply to ChatGPT Enterprise or Edu. OpenAI later announced the consumer/API preservation obligation had ended. That historical episode is useful as a retention-governance example, but it should be dated.
Primary sources
“Anonymized” Isn’t Anonymous
Research-groundedThe track’s central proposition is well established: removing direct identifiers does not guarantee that a dataset cannot be re-identified, especially when high-dimensional or auxiliary data are available.
Latanya Sweeney’s demographic-uniqueness work found that 87% of the U.S. population in the studied data was likely unique on five-digit ZIP code, sex, and full date of birth.
Narayanan and Shmatikov demonstrated robust de-anonymization of the Netflix Prize dataset using sparse ratings plus outside information.
de Montjoye and colleagues found that four spatiotemporal points were enough to uniquely characterize 95% of individuals in a 1.5-million-person mobility dataset under the study’s resolution.
“Anonymized isn’t anonymous” is rhetorically effective but mathematically absolute. Some de-identification methods can provide meaningful privacy guarantees. The precise claim is that de-identification is method- and context-dependent and that pseudonymization alone is not equivalent to anonymity.
Model Improvement Purposes
Comparative policy critiqueThe phrase “improve” really does recur across platform privacy notices, but identical English words do not automatically create identical legal permissions or identical technical pipelines.
Major platforms describe data use in terms such as providing, maintaining, developing, improving, personalizing, securing, and updating services.
OpenAI’s consumer and business products have materially different default training rules. Enterprise, Business, Edu, Healthcare, Teachers, and API business data are not used for training by default; consumer ChatGPT provides model-improvement controls.
Avoid equivalenceThe lyric “same clause in different brands” works as criticism of policy breadth, but the source ledger should not imply that Amazon, Google, Meta, and OpenAI use data in the same technical way merely because privacy notices use similar verbs.
Policy anchors
Cold Storage
Architecture is real; universals are notBackups, replicas, retention windows, disaster recovery, and legal holds are ordinary parts of large-scale systems. The song is strongest when describing deletion as a lifecycle rather than instantaneous physical erasure. Exact time windows are service-specific.
OpenAI, for example, states that deleted chats are scheduled for permanent deletion within 30 days, subject to exceptions including legal or security obligations. It also states that some internal backups may persist for a limited additional period in specific product contexts.
InferenceThe distinction between logical deletion, active-access removal, backup expiry, and derived model state is technically meaningful, but the lifecycle differs by system.
A deletion request does not universally mean “inaccessible but not zeroed.” Some systems may physically erase quickly; others may rely on cryptographic deletion, tombstones, backup expiry, or legal retention. Keep the lyric as systems criticism rather than a universal implementation claim.
Verification anchors
Settlement Without Admission
Legally sound framingThe song correctly distinguishes a civil settlement and consent order from a guilty plea or adjudicated admission. That distinction is particularly important in privacy enforcement.
The Amazon Alexa matter ended with a $25 million civil penalty and injunctive obligations. The order imposed enforceable requirements without turning every allegation into an admitted fact.
The FTC’s 2019 Facebook privacy settlement imposed a $5 billion penalty and sweeping compliance requirements. Large penalties can coexist with settlement language that avoids an admission of liability.
Editorial questionWhether a particular penalty is large enough to deter a company is an economic and policy question. Comparing penalties with revenue is relevant context, not proof that violations were economically rational.
Derived From You
Conceptual synthesisThis song describes the transformation of user behavior into features, scores, embeddings, recommendations, and models. That transformation exists in many ML systems, but the track intentionally collapses industries and products into one structural story.
Many ML systems are trained, tuned, evaluated, or personalized using human-generated examples, interactions, labels, feedback, or observed outcomes.
Qualification“Every new feature was learned from you” is poetic, not literal. Some features are hand-engineered, trained on synthetic data, built from licensed datasets, or derived without individual-user behavioral data.
Deleting a source record does not automatically reverse every derivative artifact already produced from it. Whether retraining, unlearning, or downstream deletion is required depends on law, contract, technical design, and the derivative system.
Technical / policy anchors
FDA — AI in Software as a Medical Device An example of the general principle that ML systems improve task performance based on data.
If It Saves One
Normative argument with factual examples neededThe ethical question is powerful: who shares in the value created from medical data and expert labeling? But several verses describe a generalized commercial pipeline as though it were universal. A publication companion should distinguish real examples from the composite narrative.
AI-enabled medical devices are trained and evaluated on data, and the FDA explicitly treats data quality, transparency, performance, and lifecycle management as central regulatory concerns.
Source neededClaims that hospitals “buy back” models trained on their own patients, that clinician labeling is unpaid, or that particular patient datasets become proprietary assets should be paired with named transactions, contracts, or studies. They are not safe as universal statements.
Medical-data research also operates under legal regimes such as HIPAA de-identification. A critique of commercialization should acknowledge those frameworks even when arguing they are insufficient for benefit-sharing.
Data Subject
Closing argumentThe title is literally a legal term under the GDPR. Most of the song, however, is not descriptive law. It is the album’s normative conclusion about bargaining power, governance, ownership, and benefit-sharing.
GDPR Article 4 defines a “data subject” as the identified or identifiable natural person to whom personal data relates.
Legal nuance“Once [signals] enter infrastructure, they’re no longer yours” is too categorical as a legal statement. Data rights vary by jurisdiction and context. The point works as a critique of control and bargaining power, not as a universal rule of property law.
No automatic dividend, equity stake, or governance seat generally follows from being represented in data used to build a commercial system. The song argues that this allocation of value should be reconsidered. That is Horizon Accord’s normative question, not a claim that current law already requires profit participation.

