# Overlap, original-source cross-checks and selection

1. **Exact retrospective cells:** Iceland 2009 boys/girls are read from B34/C34 and N34/O34 of II.B1.7.28/.33/.38. Reading background means are E30/F30, H30/I30 and N30/O30 of II.B1.9.10. K30/L30 remain `c`. Composition is II.B1.9.9 row 30. No data are transcribed from rounded chart/PDF labels.
2. **Independent 2009 publication:** annex I tables I.2.3/I.3.3/I.3.6 and annex II T.II.4.1 are compared against the retrospective workbook. Iceland differences are below 0.0002 points/percentage points and reproduce the displayed PDF rounding. `data/source_revision_checks.csv` records both exact values and hashes, including other countries and suppression changes. These files are not overwritten to force identity.
3. **2015/2018 overlap:** `data/country_overlap_checks.csv` compares each acquired historical mean and SE with the existing frozen release. All numeric differences for these cycles are below 0.0001; observed extremes are approximately ±0.000005. Published precision accounts for that level of difference. This validates the new extraction's row/column/group mapping and its overlap with the established series.
4. **2012 overlap:** the already selected Iceland rows match their original frozen workbooks within floating-point precision. The new OECD-29 2012 point uses the previously unselected country rows of that same frozen workbook, not a different source vintage or changing OECD aggregate.
5. **2022 revisions:** the older `wh9d4z` Iceland gender cells differ from the newer selected cells by approximately −0.013256 to −0.054789 points. Their SEs also differ. This known vintage difference is preserved in `data/overlap_checks.csv`. The archive establishes different vintages, but does not establish the detailed cause of each revision. **All current 2022 values are kept**; old observations are not silently substituted.
6. **Fixed aggregates:** 48 reconstructed current means and SEs reconcile within 0.0001. Candidate aggregates require every member and are marked as our country-level aggregation. Missing member cells block the whole point.
7. **Selection identity:** every old candidate input row carries `candidate_origin=existing selected observation, unchanged`. Tests compare its estimate, SE, year and grouping against the copied selected CSV. Only the documented earlier points are added; no interpolation or source-vintage replacement is performed.

For exact endpoint changes use `data/changed_annotations.csv`; for girls-minus-boys gaps use `data/gender_gaps.csv`. Their numeric precision is retained, while proposed prose rounds only the final descriptive differences.
