# NYC mayoral ecological inference

This analysis applies the current `2020-Vote-Calc` atlas model to the eight
official 2021 Democratic mayoral primary rounds, three official 2025 Democratic
mayoral primary rounds, and the 2025 mayoral general election. Each stage has
baseline, no-poll, and double-poll-strength outputs. Stages without a matched
demographic polling target share the same numerical fit across these settings.

## Certified votes and rounds

Original [2021 BOE cast-vote records](https://www.vote.nyc/sites/default/files/pdf/election_results/2021/20210622Primary%20Election/cvr/PE2021_CVR_Final.zip)
and [2025 BOE cast-vote records](https://www.vote.nyc/sites/default/files/pdf/election_results/2025/20250624Primary%20Election/rcv/2025_Primary_CVR_2025-07-17.zip)
are reconstructed with candidate identities from each archive's lookup workbook.
Unmarked ranks are skipped; a rank with an overvote stops further preferences;
duplicate candidates are ignored. Initial-invalid/blank ballots are excluded
from the valid RCV universe. Every candidate and inactive count in every round
is checked against the [certified 2021 report](https://vote.nyc/sites/default/files/pdf/election_results/2021/20210622Primary%20Election/rcv/024306_1.html)
and [certified 2025 report](https://vote.nyc/sites/default/files/pdf/election_results/2025/20250624Primary%20Election/rcv/026916_1.html).
No precinct or citywide candidate calibration is required.

The official reports use simultaneous eliminations. In 2025 the first round
counts all candidates, the second redistributes write-ins, and the third
redistributes all eliminated candidates to Mamdani/Cuomo or inactive. These are
the official three rounds; CUNY's nine-stage display and one-candidate-at-a-time
tabulations are different stage definitions. The existing eight-map CUNY 2021
series also groups stages differently and does not match all certified totals.
Its original audit is retained in `../ei_source_audit.json`; it is not used in
these fits.

The [2025 general-election precinct CSV](https://vote.nyc/sites/default/files/pdf/election_results/2025/20251104General%20Election/00000100000Citywide%20Mayor%20Citywide%20EDLevel.csv)
is read once per candidate/party line; fusion lines for the same candidate are
summed. Public Counter, Absentee, Affidavit, and Emergency are ballot-mode
counts, not additional candidate votes. Named candidates remain separate;
Scattered is retained as pooled write-ins. Totals reconcile to the
[certified recap](https://vote.nyc/sites/default/files/pdf/election_results/2025/20251104General%20Election/00000100000Citywide%20Mayor%20Citywide%20Recap.pdf),
including 2,194,204 valid mayoral votes. General-election nonparticipation
includes the recap's 24,443 unrecorded mayoral votes.

## Geography and demographic inputs

Election-day precinct boundaries are DCP releases 21B for the June 2021
primary, 25B for the June 2025 primary, and 25D for the November 2025 general.
The public download URLs and SHA-256 hashes are in `source_manifest.json`.
Precinct keys are two-digit Assembly District plus three-digit Election
District. All positive-vote precincts have native geometry. Retired/combined
general-election EDs with zero candidate counts are omitted.

The existing atlas demographic caches supply eight disjoint groups:
White college, White non-college, Black, Hispanic, Asian, Native American,
Pacific Islander, and multiracial. Race categories are non-Hispanic except
Hispanic, which includes any race. Correct Census CVAP identities are Black
line 5 and Asian line 4. White education is the atlas's non-Hispanic White
age-25+ tract education proxy, not a direct cross-tabulation of CVAP education.
The 2021 election uses the 2017–2021 ACS CVAP estimate; 2025 uses 2020–2024 ACS
CVAP, the existing completed pre-election cache, on canonical 2020 block groups.
Neither is an election-day census. Cache hashes are recorded per contest.

The atlas's reusable `allocation_cache.build` intersects precincts with block
groups in EPSG:5070 and weights overlap by uniform within-block-group CVAP
density. Column-normalized weights conserve every precinct's candidate count.
Every candidate and inactive total is independently checked after allocation.
There are no nearest-populated-block-group allocation fallbacks in these jobs.
Exact constraints apply to the resulting allocated block-group counts; public
data do not observe votes by block group or demographic.

The neighborhood graph uses 20 nearest **polygon representative points** in
EPSG:5070, excluding self, restricted to NYC. Inverse distance has the atlas's
250-meter floor and is row-normalized. This is a documented point-location
choice differing from the presidential atlas's Census internal points. Areas
outside NYC have no observations for these races and are not inserted as zeros.

For each contest, a common feasible CVAP margin is used across its rounds.
Where allocated valid RCV/contest ballots exceed ACS CVAP, local CVAP rises to
that ballot count while retaining the demographic proportions. The audits show
2 adjusted block groups in 2021 (+47.07 CVAP), 1 in the 2025 primary (+136.48),
and 26 in the 2025 general (+2,553.82). No zero-CVAP composition imputations
were required. These adjustments are model repairs, not observed population.
Inactive ballots remain a separate category from nonparticipation.

## Model and polling

Unknown tables have exact adjusted demographic row margins and allocated
candidate/inactive/nonparticipation column margins. The objective is the
atlas's population-weighted geographic neighbor prediction error plus soft
demographic polling penalties and relative entropy against the local independent
table, with epsilon 0.02. Groups under 1% locally use the observed citywide
category-share baseline in neighbor prediction, as in the atlas. They retain
their demographic margins in the fitted table.

`candidate_eot.py` extends sparse `RegionalObjective` moments to arbitrary
candidate columns. Each poll explicitly lists disjoint candidate bins; its
denominator includes only those bins. Inactive/nonparticipation never enter
candidate preference targets. The unchanged atlas `eot.fit` and log-domain
Sinkhorn oracle perform feasible updates, full-objective line searches and
primal-dual-corrected gap certification. Local equality-preserving Newton
directions accelerate the solve; their acceptance still uses the complete
coupled objective. The gap threshold is 2e-7. Tests check gradients, candidate
bin denominators, zero columns, convergence and exact margins.

Matched polling targets:

| Election/stage | Survey | Targets |
| --- | --- | --- |
| 2021 primary round 1 | [Marist, June 3–9, 2021](https://maristpoll.marist.edu/wp-content/uploads/2021/06/20210612_WNBC_Telemundo-47_POLITICO_Marist-Poll_NYC-NOS-and-Tables_RCV_20210611256.pdf), page 3, n=876 | White, Black, Latino first choices, conditioned on named candidates and excluding undecided |
| 2021 primary round 7 | Same Marist survey, page 5 | White, Black, Latino late-round Adams/Garcia/Wiley preferences |
| 2025 primary round 1 | [Marist, June 9–12, 2025](https://maristpoll.marist.edu/wp-content/uploads/2025/06/Marist-Poll_June-NYC-NOS-and-Tables_202506162130-2.pdf), page 5, n=1,350 | White, Black, Latino first choices, conditioned on named candidates and excluding undecided |
| 2025 general | [CNN/SSRS voter poll](https://www.cnn.com/election/2025/exit-polls/new-york-city/general/mayor/0), [public numerical feed](https://politics.api.cnn.io/results/exit-poll/2025-MG-NYC.json), n=4,744 | White college/non-college, Black, Latino, Asian, pooled other races; conditioned on Mamdani/Cuomo/Sliwa |

Rounded percentages are normalized to one. Reliability is
`0.5 * min(1, approximate_subgroup_n / 400)`; preelection targets receive an
additional factor 0.5. The subgroup sample sizes are approximated by total
sample times published group share and are not measured design-effective sample
sizes. The pooled White general-election target is replaced by its disjoint
education subgroups; overlapping rows from the same survey are not double-counted.
The general's other-race poll constrains Native/Pacific/multiracial jointly.
Primary Asian/Native/Pacific/multiracial preference targets are unavailable.

Only matched stages receive polling. First-choice targets are not transferred
to later rounds; a survey's different candidate elimination sequence does not
create an interchangeable demographic target. The 2025 late-three poll table
does not constrain its official final-two stage. Baseline uses strength 1;
no-poll uses 0 and strong-poll 2. Preelection targets may reflect preferences
that changed before voting. Sensitivity differences are not confidence intervals.

## Interpretation and output

Stages are fitted independently. Changes in subgroup candidate shares can
reflect inferred active-ballot composition and the stage-specific objective.
They are **not demographic transfer counts** and do not guarantee persistent
individual voter identities or demographic turnout across rounds. CVAP includes
citizens outside the Democratic registered electorate, so primary
participation is valid RCV ballots divided by adjusted CVAP, not registered
Democratic turnout. Positive entropy produces a unique regularized optimum,
not identification of individual demographic ballots or calibrated uncertainty.

`demographic_summary.csv` contains every candidate/category for every
demographic, official stage, polling setting, citywide and five boroughs.
`poll_sensitivity.csv` records citywide candidate-share differences in percentage
points. Each fit directory has labeled `estimates.npz`, canonical float64
`counts.npy`, convergence history for independently optimized fits, poll target
definitions and fingerprints. Shared geometry/CVAP/graph and raw-to-adjusted
population audits are in `data/processed/ei/<contest>/`. The offline explorer is
`index.html`; all baseline stages have eight-demographic PNG maps in `maps/`.

Maps shade each demographic's estimated leading candidate, with darker colors
for a larger active-vote share. Groups below 1% of local CVAP, 25 citizens, or
10 estimated active votes are masked. Tiny Native/Pacific subgroup estimates
should be interpreted especially cautiously. Final `validation.json` reports
optimizer certificates and independently recomputed saved-table margin errors.

## Reproduction

Keep the sibling `2020-Vote-Calc` checkout and its existing demographic caches
and full 2020 NY block-group geometry. Python dependencies are NumPy, pandas,
SciPy, GeoPandas, Shapely, Matplotlib, requests, BeautifulSoup, python-calamine,
and pypdf. From the repository root:

```powershell
python scripts/fetch_nyc_ei_sources.py
python scripts/nyc_sources.py --year 2021
python scripts/nyc_sources.py --year 2025
python scripts/prepare_nyc_ei.py
python -m unittest discover -s scripts -p test_candidate_eot.py -v
python scripts/run_nyc_ei.py
python scripts/audit_nyc_ei.py
python scripts/build_ei_report.py
python scripts/build_ei_maps.py
```

Existing source and input caches are reused. For changed raw sources or
preparation assumptions, regenerate the affected prepared input directory in
a fresh location rather than interpreting an old cache as a new analysis.
Fit fingerprints bind numerical margins, graph, metadata, polls, solver code,
epsilon and poll strength. Model output is not published to aweakprior.com by
these scripts.
