# Extended corpus: coverage and quality

Snapshot: September 26, 2026. Experimental automated analysis, not a validated research index.

45 materials, 20 participants, 7,367 speech units and 112,355 words. The model processed all 7,367 units; 18 remain unresolved. All 20 participants have material from at least two months, but dates are uneven and annual coverage is incomplete.

## Extraction

Participant utterances were separated using speaker labels or speech boundaries. Other speakers’ questions, marked translations, readings of fiction, film dialogue inserts, editorial notes and repeated sentences by the same person were removed. Extraction beginnings and endings were checked technically; this was not independent listening to every recording.

A unit is a sentence or part of a turn at a paragraph boundary, unlike the shorter semantic passages in the original pilot. Counts are not directly comparable. Mixed temporal references may receive `other`; one-word turns are excluded. Completeness refers to selected transcripts, not all of a person’s speech during the year.

Neal Mohan’s January 21 letter is written material marked `authored_publication_date`. Newsom’s January 8 prepared remarks are marked `prepared_remarks_date`; correspondence with delivered speech is unverified. `publication_date_proxy` is not a confirmed recording date. Participants may quote other people within their own turns; the model is instructed to analyze the participant’s assertion, but compliance is not guaranteed.

## Preliminary model evaluation

Model: mlx-community/Qwen3.5-4B-4bit, revision 0e7ffd5c629ef7719d4cbc04069232580bfa9d9c. Temperature 0; prompt and hash are retained in analyze_panel.py and annotation records. Agreement with prior assistant labels: 103/130 (79.23%), with 5 unresolved answers counted as disagreements.

This compares two annotations on a convenient short-excerpt sample, not demonstrated accuracy on long materials. No independent human gold standard exists. Some disagreements need more context; expected labels were not changed after the check.

| Assistant reference | Model responses |
|---|---|
| past | past: 36, other: 1, present: 4, unresolved: 1 |
| present | future: 7, present: 46, other: 3, past: 1, unresolved: 2 |
| future | future: 16, present: 4, unresolved: 2 |
| other | other: 5, present: 2 |

| Participant | Sources | Dates | Months | Words | Annotated |
|---|---:|---:|---|---:|---:|
| Wagner Moura | 2 | 2 | 2026-02, 2026-08 | 3952 | 219 |
| Ethan Hawke | 2 | 2 | 2026-01, 2026-02 | 1193 | 74 |
| Tayari Jones | 2 | 2 | 2026-02, 2026-08 | 1805 | 117 |
| Shannon Minter | 2 | 2 | 2026-03, 2026-04 | 1709 | 71 |
| Scottie Scheffler | 3 | 3 | 2026-03, 2026-05, 2026-08 | 7821 | 476 |
| Sundar Pichai | 3 | 3 | 2026-02, 2026-06, 2026-07 | 4985 | 303 |
| Marco Rubio | 2 | 2 | 2026-01, 2026-02 | 5019 | 250 |
| Donald Trump | 2 | 2 | 2026-02, 2026-03 | 15166 | 1247 |
| Mark Carney | 2 | 2 | 2026-01, 2026-04 | 5382 | 379 |
| Zohran Mamdani | 3 | 3 | 2026-04, 2026-07, 2026-09 | 4707 | 251 |
| Benjamin Netanyahu | 2 | 2 | 2026-05, 2026-09 | 14482 | 1203 |
| Steve Witkoff | 2 | 2 | 2026-01, 2026-03 | 2365 | 161 |
| Gavin Newsom | 2 | 2 | 2026-01, 2026-06 | 5879 | 348 |
| Gwynne Shotwell | 3 | 2 | 2026-02, 2026-06 | 5974 | 385 |
| Mark Kelly | 3 | 3 | 2026-05, 2026-06, 2026-08 | 2912 | 178 |
| Jacob Frey | 2 | 2 | 2026-01, 2026-02 | 3084 | 187 |
| Lando Norris | 2 | 2 | 2026-07, 2026-08 | 5183 | 322 |
| Dario Amodei | 2 | 2 | 2026-02, 2026-09 | 6805 | 369 |
| John Furner | 2 | 2 | 2026-02, 2026-05 | 6927 | 391 |
| Neal Mohan | 2 | 2 | 2026-01, 2026-05 | 7005 | 418 |

## Limits and next steps

Do not call this a complete 2026 archive, a representative TIME100 panel or a predictive correlation model. Equal weighting applies only to observed people in each window. Daily correlation thresholds of 5 people and 100 units are interface safeguards, not statistical guarantees.

The annotation review page hides model labels and saves answers in the browser for JSON export; it does not automatically modify the index. Next steps are an extraction audit, two independent human annotators on a randomly selected stratified sample of the extended corpus, and broader monthly coverage for each participant.
