# Quality report - h1b-lca-fy2010-fy2025-2026-10-09

This is the short measured summary. The full reconciliation, the identity audit and the independent review are in `RELEASE_CERTIFICATION_REPORT.md`, `IDENTITY_AUDIT_ERAS.md` and `qa/`.

| Item | Result |
|---|---|
| Case rows / unique applications | 8,918,577 / 8,238,223 (377,296 repeated year-to-date snapshot rows kept with lineage) |
| Source traceability | 8,918,577 of 8,918,577 rows tie to the source workbook cells (case number, status, decision date); an independent sample of 23,867 rows: 0 mismatches |
| Row key uniqueness | `(source_key, source_row)`: 0 duplicates |
| Worksites | 9,375,219 rows, 0 without a case row |
| Panel | 1,916,599 rows; 11 measures re-derived from the case files for every cell, 0 mismatches; 27 measure columns checked by an independent reviewer after patch 1, 0 mismatches on the definitions documented in METHODOLOGY |
| Reconciliation to DOL published totals | exact on all four statuses for 21 of 29 sources with an official target; 8 sources differ (listed in the certification report); worker positions not reproduced |
| Missingness | per column in `DATA_DICTIONARY.csv` (`null_share`) |
| Excluded fields and free text | no excluded column in any table; 4,608 identifier-like cells in free-text columns redacted (patch 2); 84,759 hash-derived ids replaced by opaque ids; post-redaction scan found no remaining e-mail, tax-id-format, phone or 9+ digit match outside documented exceptions (`qa/PRIVACY_VERIFICATION.json`) |
| Employer identity | provisional (R5-v1); 120 audited groups across five eras; no false merge seen among 40 spelling merges; 17 of 40 sampled splits cannot be decided from filing fields |
| Known exceptions | FY2010 assumptions; estimated FY2010-FY2015 certified/withdrawn split; provisional certified/withdrawn rule; FY2025 Q1 difference of 1,733 unresolved; unstandardized SOC codes; PW level absent FY2010-FY2014 and FY2016 |
| Not independently reviewed | SOC vintage assignment, employer-id quality, wage annualization correctness |
| Formats | Parquet only; no CSV files in this edition |
| Review status | builder QA plus one AI reviewer; no independent human review; no legal opinion on rights or license |


---

# Release Certification Report - H-1B LCA dataset FY2010-FY2025, Edition 2026-10-09

Prepared by the Persistfolio build agent. Reviewer independence: the checks below were written and run by the same agent that built the pipeline; where possible the check is a separate re-implementation that shares no pipeline code (marked INDEPENDENT). No external auditor and no counsel has reviewed this report.

## 1. Verdict

**Engineering: CERTIFIED WITH DOCUMENTED EXCEPTIONS.** Every structural and arithmetic check passes. DOL reconciliation is exact on all four statuses for 21 of the 29 sources that have an official target; the other eight (FY2021 Q1-Q3, FY2023 Q1-Q2, FY2025 Q1, FY2014, FY2015) have documented differences, and the cause of the six certified-count differences is unproven. FY2010 and the FY2010-FY2015 certified/withdrawn split are assumption-based and labeled as such. Employer identity is provisional and measurably imperfect.
**Commercial release:** published by the founder on the basis of the verification below and of DOL's public records. This is not a legal clearance: no lawyer has reviewed the rights position or the license, and no independent human audit has been done. The license is `LICENSE.txt` (the placeholder inside the ZIP is superseded).

## 2. What was certified
34 DOL files (SHA-256 in `SOURCE_INVENTORY.csv`), FY2010-FY2025, every fiscal year complete. Released: 8,918,577 H-1B case rows (71 columns), 9,375,219 worksite rows, 8,238,223 unique applications, 1,916,599 panel rows (45 columns); Parquet 809 MB (a CSV.gz rendering was built during development but is not shipped in this edition).

## 3. Reproducible QA results
| # | Check | Method | Result |
|---|---|---|---|
| 1 | Source integrity | SHA-256 recorded for all 34 files at receipt; registry hash = manifest hash | pass |
| 2 | Every source row traceable | release vs source workbook cells, key (source_key, source_row): case number, status, decision date compared on all rows (`release_vs_source.py`, INDEPENDENT of the pipeline SQL) | 8,918,577 rows, 0 discrepancies, row counts equal per file |
| 3 | Row roles | pure-Python re-implementation from raw parquet (`verify_roles_multi.py`, INDEPENDENT) | 0 mismatches over 8,918,577 rows; 8,215,630 original decisions; 377,296 repeated snapshots; 8,238,223 unique case numbers |
| 4 | Role logic fixtures | `test_quarter_dedup.py` (synthetic cases incl. repeated snapshots) | 24 of 24 |
| 5 | Panel rebuilt from released case files | pandas only, released Parquet only (`release_tieout.py`, INDEPENDENT) | 1,916,599 of 1,916,599 cells; 11 measures; 0 mismatches (builder-written). Second reviewer: 27 measure columns; found the 4 `pw_level` columns wrong (fixed in patch 1), the rest 0 mismatches |
| 6 | Same, free sample | same script on the sample files | 4,074 of 4,074 cells, 0 mismatches |
| 7 | Applications table | one row per case number; original-loaded count equals original-decision rows; snapshot sum equals snapshot rows | pass |
| 8 | Worksites | every worksite row has a case row | 0 orphans over 9,375,219 |
| 9 | Row key unique | (source_key, source_row) | 0 duplicates |
| 10 | Excluded fields absent | column-name scan of all released tables, dictionary completeness | 0 hits; every column documented |
| 11 | Reconciliation to DOL | `reconcile_official.py`, raw layer, all visa classes; `reference/DOL_RECONCILIATION.csv` | see section 4 |
| 12 | Panel vs DOL fiscal-year totals | `qa/FY_PANEL_VS_DOL_PUBLISHED.csv` | H-1B records are 97.5% to 98.5% of DOL all-visa processed in every comparable year (section 5) |
| 13 | Identity audit | 120 groups, 5 eras, 3 strata (`IDENTITY_AUDIT_ERAS.md`) | no false merge seen; 11 of 40 splits likely false; see section 7 |
| 14 | Cumulative-file check | section 6 | pass after a defect was found and fixed |
All scripts are in `reproduction/pilot_pipeline/` (branch `experiment/sku6-h1b-feasibility`); seeds are fixed (audit 20261011, sample 20261012). The database build takes about 30 minutes on 2 cores and 8 GB.

## 4. Reconciliation to DOL published statistics (exceptions preserved)
Compared per source: processed, certified, denied, withdrawn (all visa classes; DOL's published figure from the latest report that states it).
- **Exact on all four:** 21 sources - FY2017, FY2018, FY2019, FY2020 Q1-Q4, FY2021 Q4, FY2022 Q1-Q4, FY2023 Q3-Q4, FY2024 Q1-Q4, FY2025 Q2-Q4.
- **Certified (and processed) higher than DOL; denied and withdrawn exact:**

| Source | Release minus DOL | % of DOL certified |
|---|---|---|
| FY2021 Q1 | +1,799 | 2.42% |
| FY2021 Q2 (cumulative target) | +2,754 | 1.37% |
| FY2021 Q3 (cumulative target) | +3,300 | 0.86% |
| FY2023 Q1 | +1,567 | 1.79% |
| FY2023 Q2 (cumulative target) | +2,299 | 1.09% |
| FY2025 Q1 | +1,733 | 1.77% |
Cause **unresolved**. For the three Q1 files the difference is within 0 to 3 of the count of that quarter's certified cases that reappear as certified-withdrawn in later files; Q1 files with no such cases tie exactly. That is a strong association, not a demonstrated mechanism. DOL restated FY2025 Q1 between its own reports (106,278 in the FY2025 Q3 report, 105,681 in the FY2025 Q4 report; the file has 107,414). Detail: `FY2025_Q1_VARIANCE_INVESTIGATION.md`.
- **FY2014 and FY2015:** processed +1 and -6 against DOL; the as-filed status columns differ because the files have no original certification date. With the submit-date proxy the certified and withdrawn totals land within 0.06% (certified) and 0.9% (withdrawn) of DOL (`qa/CW_PROXY_CHECK_FY2014_FY2015.csv`). Aggregate agreement only.
- **No official target:** FY2010, FY2011, FY2012, FY2013, FY2016 (internal consistency only).
- **Worker positions:** do not reproduce (+2.4% to +3.5% on four checked files). Unsupported.

## 5. Panel fiscal-year totals vs DOL (`qa/FY_PANEL_VS_DOL_PUBLISHED.csv`)
The panel's `fy_status_records_dol_processed_basis` summed over a fiscal year, H-1B only, divided by DOL's all-visa processed total: FY2014 97.9%, FY2015 97.9%, FY2017 97.7%, FY2018 97.7%, FY2019 97.7%, FY2020 97.7%, FY2021 98.5%, FY2022 97.6%, FY2023 97.9%, FY2024 97.5%, FY2025 97.9%. A stable ratio in that range is expected (H-1B is most of the all-visa total, the rest is E-3 and H-1B1). It would not hold if cumulative-file rows were double counted (see section 6).

## 6. Cumulative quarterly files (FY2021 Q2, FY2021 Q3, FY2023 Q2)
**Question:** are repeated year-to-date snapshots counted as independent applications or status events?
**Applications:** no. Unique applications are keyed on case number; 8,238,223 case numbers, one application each, in every version of the build.
**Status events:** before this check, yes. Rows flagged "an earlier row of this case exists" were all labeled `LATER_STATUS_EVENT`. Measured against the case's latest earlier-file row (same status, decision date and original certification date):

| File | Rows with an earlier-file row | Identical copy | Status changed |
|---|---|---|---|
| FY2021 Q2 | 80,020 | 77,006 (96.2%) | 3,014 |
| FY2021 Q3 | 209,382 | 205,756 (98.3%) | 3,626 |
| FY2023 Q2 | 100,782 | 94,531 (93.8%) | 6,251 |
| every quarter-only and annual file (25 files) | about 4,000 to 30,000 each | 0 (3 in FY2012) | all |
Identical copies are year-to-date repeats, not events. Changed-status rows are real events and stay `LATER_STATUS_EVENT`.
**Fix:** a separate role `REPEATED_SNAPSHOT` (377,296 rows). Rows are kept with full lineage (`source_key`, `source_row`, evidence `IDENTICAL_TO_EARLIER_FILE_ROW`) and excluded from the `fy_*` panel counts and the application status history.
**Effect:** FY2021 panel status records 803,733 -> 520,971 (152% -> 98.5% of DOL), FY2023 626,887 -> 532,356 (115% -> 97.9%). Original decisions, unique applications and every DOL reconciliation on the raw files are unchanged. Independent role re-implementation agrees on all 377,296 rows.

## 7. Employer identity (R5-v1)
See `IDENTITY_AUDIT_ERAS.md` (120 groups, five eras, three strata, small samples, one reviewer). No false merge was seen among 40 spelling merges (11 plausible, 29 supported). Of 40 sampled R5 splits, 12 look correct, 11 look like one company split across offices, 17 cannot be decided from filing fields. Of 40 unmerged name-variant groups, 23 are the same company on address evidence; all share a `candidate_group_id`. Unknown and not claimed: any population false-merge rate, any accuracy rate, corporate-family resolution.

## 8. Defects found and fixed during certification
1. Cumulative repeats labeled as status events (section 6). Fixed, re-verified.
2. Free-sample and tie-out dedupe key dropped legitimate rows when one case appears in two files of one fiscal year. Fixed in the earlier full run; both tie-outs now pass.
3. SOC fallback order for SOC 2000 files; header drift inside a fiscal year; five new layout families. Fixed during onboarding.
4. A data-quality finding, not fixed by design: DOL's FY2011/FY2012 files hold SOC codes that Excel turned into dates; kept as filed.

## 9. Material limitations
Full list in `LIMITATIONS.md`. The ones that change how a buyer should use the data: FY2010 assumptions; provisional certified/withdrawn rule and estimated FY2010-FY2015 split (no error bound claimed); eight sources with the differences above; unsupported worker positions; provisional employer ids; weaker SOC readability in FY2011-FY2013; PW level absent FY2010-FY2014 and FY2016; DOL restates its own figures.

## 10. Open review items
1. No legal opinion has been obtained on the rights questions or on the license. Employer names that look like natural persons are kept as in DOL's public file; the correction and removal process is in `LICENSE.txt` section 7.
2. No independent human second reviewer has re-run `release_tieout.py` and `release_vs_source.py` from the stored source files.

## 11. Independent review and patch 1 (2026-10-09, after the first build)
A separate reviewer agent (a different AI instance that did not read the pipeline SQL or the builder's QA scripts) re-checked the package from the released Parquet files and the 34 source workbooks. It is still AI review by the same vendor, not a human audit, and an independent human second review (section 10) has not been done. Result: **PASS WITH EXCEPTIONS.**
- Integrity: 59 of 59 listed files match; all 34 source hashes match.
- Traceability: 23,867 sampled rows from all 16 fiscal years, 0 mismatches on case number, decision date, wage and employer name beyond whitespace. Some values are normalized, not literally as filed (status upper-cased, 'Certified - Withdrawn' written CERTIFIED-WITHDRAWN from FY2020, names whitespace-trimmed, Excel formula wrappers stripped from SOC codes).
- Unique applications: 8,238,223 case numbers, 377,296 repeated snapshots, 8,215,630 original decisions all reproduced, including the FY2021 and FY2023 cumulative files.
- Panel: 4 `pw_level` columns were 0 for FY2020-FY2025 (the matcher looked for 'Level I' but those files say 'I'). **Fixed in patch 1**: the corrected panel differs from the first one only in those four columns; all other cells are identical.
- Excluded fields: no excluded column exists in any table, but filer-typed free text contained FEIN-like, e-mail and phone strings (LIMITATIONS 12). **Fixed in patch 2** (section 12). Whether the redaction is sufficient has not been reviewed by a privacy professional.
- Documentation corrections (patch 1): SOC date-corruption counts, `case_status` normalization, `source_row` definition, `wage_annual_from_calc` for AMBIGUOUS rows, percentile rounding, manifest wording, the 14-day proxy threshold, and scope of the 'independent' claims.
- Not covered by the reviewer: DOL reconciliation, SOC vintage assignment, employer-id quality, wage annualization correctness, worksites beyond the leak scan.
Referenced repository documents (`RIGHTS_ASSESSMENT_RELEASE_FIELDS.md`, `FY2025_Q1_VARIANCE_INVESTIGATION.md`, `reproduction/pilot_pipeline/`) live in the project repository and are not shipped in this package.

## 12. Patch 2: privacy redaction (2026-10-09, before any sale)
Trigger: the independent review (section 11) found FEIN-format, e-mail and phone strings in free-text fields; the follow-up found that short hashes used as ids (`evidence_component_id`, split-fragment `employer_id`) were computed from phone and address keys of excluded fields.
Action: 4,608 cells redacted; 84,759 ids replaced one-to-one by opaque ids; every Parquet table rewritten with the same rows in the same order; free sample rebuilt from the same rules. Details: METHODOLOGY (last section), LIMITATIONS 12 and 15, `qa/PRIVACY_REDACTION_LEDGER.csv`.
Verification results (all must pass): I1 PASS; I2 PASS; I3 PASS; I4 PASS; I5 PASS; I6 PASS; P1 PASS; P2 PASS; P3 PASS. Full output: `qa/PRIVACY_VERIFICATION.json`. Panel rebuilt from the redacted case files: 1,916,599 of 1,916,599 cells re-derived, 0 mismatches (`qa/release_tieout_patch2.json`).
Limits: pattern-based detection finds only the listed patterns; personal names, unusual formats and numbers shorter than 9 digits are not detected. This is AI-built verification, not a human privacy audit or a legal opinion.
