# TAF Tense Index — pilot methodology

Version 0.2 · September 26, 2026. Period: January 1–September 26, 2026, inclusive.

## Two separate corpora

The default panel contains 20 participants and automated analysis of selected transcripts and authored materials. The original short-excerpt pilot remains a separate mode. Both modes offer the full TIME100 list of 100 people; inclusion in the list does not mean that speech has been processed.

See panel-report-en.md for coverage and model evaluation. Neal Mohan’s January 21 letter is written material, not verified spoken language. Gavin Newsom’s January 8 prepared remarks have not been checked against audio. These types are explicitly marked in the data and cards. Other materials are transcripts; recording may precede broadcast. Extraction rules have been technically checked, but an independent audit of every utterance remains incomplete.

## Cohort and original pilot

TIME100 2026 comprises 100 people across 95 profiles. Five paired profiles were split into individual participants to retain equal personal weights. TIME100 is not an ordinal ranking. Excerpts currently cover 50 people; the corpus is not representative. The list was published in April, while collection extends retrospectively to January. It cannot support claims of real-time prediction.

The pilot uses participants’ own words: direct interview quotations, speeches, prepared remarks, authored letters and statements. Journalistic paraphrase and translations are excluded. Prepared text does not establish that the same words were spoken on stage; source type is retained. Speech dates are preferred; publication dates are explicitly marked in `date_basis` when used instead.

The original pilot contains short reviewed excerpts rather than continuous transcripts. Assistant selection creates source and excerpt selection bias. `reviewed` means assistant review, not independent human validation.

## Semantic annotation

A unit is a passage with one dominant temporal reference relative to the moment of speech:

- `past`: events that have already happened;
- `present`: current states, actions, evaluations or general statements;
- `future`: expected events, promises, intentions or goals;
- `other`: ambiguous, conditional or temporally unspecified statements.

Grammar alone does not determine the class. A clear future intention is `future`, even if expressed modally. Mixed statements should be split by meaning. The interface displays an English explanation for each pilot decision; original research annotations are retained in the data files.

The pilot was annotated by the assistant. The separate panel uses local Qwen3.5-4B-4bit via MLX, with `machine_annotated` labels kept distinct from `reviewed`. Agreement on 130 control excerpts is 103/130. There is no independent human accuracy assessment.

## Equal weights and windows

Within each window, class shares are computed per participant, then averaged across participants with material in that window. More excerpts do not increase a participant’s weight. Missing participants are not treated as zero; panel composition and effective weights change over time.

The four shares sum to 100%. The main chart shows three temporal categories; `other` remains in the denominator and appears in the composition panel. Composition refers to the entire selected period.

Windows of 1, 7 and 30 calendar days end at each chart date and are clipped to the selected period. A window pools excerpts and recalculates equal weights; it is not an average of daily percentages. Empty windows have no value. A rolling window may have a value on a day with no new speech if recent material remains within it. Missing values are not interpolated. Repeated quotations from the same person are removed after case and punctuation normalization; paraphrased repetition requires separate review.

## External series and correlations

Public FRED series: S&P 500 (`SP500`), expected volatility (`VIXCLS`) and daily US economic policy uncertainty (`USEPUINDXD`). Each series has its own latest available date. Weekends and gaps are not filled.

The comparison chart standardizes each displayed series within the selected period: z = (x − mean) / standard deviation. It compares patterns, not physical units. The speech comparison line uses the selected window.

Correlations always use unsmoothed daily profiles on matching dates. Levels compares speech shares with external index levels. Changes compares speech-share differences in percentage points with percentage changes in the external series between consecutive matching dates. Intervals may vary and are not necessarily one-calendar-day returns.

Pearson correlation requires at least 20 pairs and nonzero variation in both series. Each cell shows n. This is an interface threshold, not a guarantee of statistical reliability. No significance testing, multiple-comparison adjustment, lag analysis or causal inference is provided. Panel composition, excerpt selection, autocorrelation and irregular intervals may distort the result.

## Local operation and admission

The interface, calculations and stored data work locally without external requests. The server binds to 127.0.0.1. Refreshing external series and collecting new sources require internet access; source links open their original websites.

A discovered link does not enter calculations automatically: the panel needs a source manifest, speaker extraction rule and model annotation. `unresolved` units are excluded from the denominator. Automated-panel correlations additionally require at least 5 observed people and 100 speech units per day. Blank correlation cells are expected when coverage is insufficient.

`publication_date_proxy` means the recording date is unknown and the interview publication date was used. These cards are marked; this adds uncertainty to daily comparisons. New pipeline additions require verified original English speech, speaker attribution and an event or broadcast date within the study period.
