# Limitations (read before use)

1. **FY2010 is built on assumptions.** The file has no visa-class column: every row is assumed H-1B (E-3 and H-1B1 rows may be mixed in). It has no offered-wage unit column: the unit is inferred from the prevailing-wage unit. Its certified/withdrawn split is a proxy flagging 5.3% of certified-withdrawn rows as later events, against 59% in FY2013, and there is no official split to check it against.
2. **Certified vs withdrawn is a rule, not a DOL field.** DOL does not document how it splits CERTIFIED-WITHDRAWN into its published certified and withdrawn counts. The rule used (FY2016 on) is provisional. For FY2010 to FY2015 the split is an estimate from submit dates. FY2014 and FY2015 agree with DOL's totals within 0.9% in aggregate; FY2010 to FY2013 have no official target. The number of rows the estimate may misclassify is not known and no error bound is claimed.
3. **Six quarters do not match DOL's published certified counts.** FY2021 Q1 (+1,799), FY2021 Q2 (+2,754 on the cumulative target), FY2021 Q3 (+3,300), FY2023 Q1 (+1,567), FY2023 Q2 (+2,299), FY2025 Q1 (+1,733). All are in the certified column only; denied and withdrawn are exact. For the three Q1 files the difference is within 0 to 3 of the number of that quarter's certified cases that appear later as certified-withdrawn, and quarters with no such cases tie exactly, so a restatement-type mechanism is likely, but it is not proven. DOL also restated FY2025 Q1 between its own reports. The difference is unresolved. FY2014 processed differs by +1 and FY2015 by -6.
4. **No official target** was available for FY2010 to FY2013 and FY2016, so those years are checked only for internal consistency and plausibility.
5. **Worker positions do not reproduce.** Sums of `total_worker_positions` exceed DOL's published position totals by 2% to 4%. Positions are as filed; use them as filed, not as DOL totals.
6. **Employer identity is provisional.** `employer_id` is a conservative analytical key. Some companies appear under several ids (R5 splits across offices, legal-suffix variants), and distinct ids are not distinct companies. Name collisions between unrelated companies are split only when evidence disconnects them. No corporate families, subsidiaries, renames or acquisitions are resolved. No false-merge rate is claimed because it was not measurable.
7. **SOC coverage varies.** About 11% of FY2012 and 7% of FY2013 rows, 3% of FY2011 and 1% of FY2010 rows have no vintage-resolved code, mostly the legacy code 15-1799. DOL's FY2011 and FY2012 files contain SOC codes that Excel converted to dates (measured in the independent review: FY2011 4,411 rows, FY2012 16,864, FY2010 128, FY2022 242 across Q1-Q4); they are unreadable and kept as filed. The earlier figures 4,487, 16,943 and 1,983 were the counts of codes the pipeline could not parse, which is a larger group. About 1,923 FY2022 Q2 codes have a readable form such as 15-1132.00.00 that the pipeline did not standardize (about 2,545 such codes in total end up `INVALID_OR_MISSING` in the panel); this is a known gap, kept as is in this edition. Split mappings to SOC 2018 are not resolved to a single code. Vintages mix inside FY2022 Q4 to FY2023 Q4.
8. **Prevailing wage level** is absent in FY2010 to FY2014 and FY2016 files, so level counts are 0 there. (Patch 1 fixed a defect in the first build that left the panel level counts at 0 for FY2020-FY2025 as well; the filed values were always present in the case files.)
9. **Cumulative quarterly files.** FY2021 Q2, Q3 and FY2023 Q2 repeat earlier rows. Those rows (377,293, plus 3 identical duplicates in FY2012) are labeled `REPEATED_SNAPSHOT` and excluded from status counts but present in the case table. If you sum rows without the role filter you will double count.
10. **Open status.** An application whose last loaded status is CERTIFIED may be withdrawn later (`final_status_confidence`). The last file ends at FY2025. 22,593 case numbers have no original decision in the loaded files.
11. **Flags are per filing.** `willful_violator` and `h1b_dependent` are as filed on that filing and are not attributes of a company or of an `employer_id`.
12. **Excluded and redacted on purpose.** The FEIN, contact, attorney and agent, employer street address, phone and law-firm columns are not released. Filers also type identifiers into free-text fields. In patch 2 these were redacted before release: e-mail addresses; tax-id-format numbers NN-NNNNNNN (this includes DUNS numbers typed in that format); formatted phone numbers; runs of 9 or more digits in name-like fields; digit-only values of 9 or more digits in the wage fields; and values of 10 or more digits in postal fields. Replacement tokens are `[REDACTED-EMAIL]`, `[REDACTED-PHONE]`, `[REDACTED-ID-NUMBER]` and `[REDACTED-NUMERIC-ID]`; the rest of the string is kept. 4,608 cells were changed (cases 4,384, worksites 82, panel 142). `qa/PRIVACY_REDACTION_LEDGER.csv` lists every changed cell by `source_key` and `source_row` (plus `slot`, or the panel key) with column and type, and contains no original values, so the original row can be looked up in the DOL file. **Not removed:** web addresses, numbers shorter than 9 digits, 9-digit ZIP+4 postal codes, wage amounts with cents, and personal names that filers used as employer names. The redaction is pattern-based; it is not a guarantee that no personal data remains and it is not a legal anonymisation opinion. No privacy professional or lawyer has reviewed it.
13. **Source file defects** are preserved, not repaired (blank padding rows, truncated names, odd status values: 14 FY2013 rows "PENDING QUALITY AND COMPLIANCE REVIEW - UNASSIGNED", 2 FY2014 rows "REJECTED").
14. **Not an official or complete record.** DOL may have restated, corrected or withdrawn records after publication. This is a snapshot of the files listed in `SOURCE_INVENTORY.csv`. No updates are promised; a later year is a separate edition.
15. **Some identifiers are opaque.** In the first build, `evidence_component_id` and the `employer_id` of split fragments were short hashes computed from the filing's raw name together with its phone number and street-address key (both excluded fields), and the ids of 46 employer name keys that were themselves tax-id-like, phone-like or e-mail-like strings were hashes of those strings. Short hashes of such inputs can be guessed back, so patch 2 replaced 84,759 ids (2,969 employer ids, 81,783 evidence component ids, 7 candidate group ids) with opaque values. The replacement is one-to-one in every table (verified), so joins, counts and the panel are unchanged. These ids can no longer be recomputed from the published columns. `employer_key_l1` of the affected employers shows redaction tokens, so two different employers can show the same key text, but never the same id.
16. **No independent review.** Verification is the builder's QA plus one independent AI reviewer (same vendor). No human has audited the data, the privacy redaction or the identity audit, and no lawyer has reviewed the rights position or the license. Treat the redaction, the license and the identity work as not checked by outside experts.
17. **Parquet only.** This edition contains Parquet files and no CSV files. Parquet types (dates, integers) are as documented in `DATA_DICTIONARY.csv`.
18. **Reconciliation exceptions.** Eight of the 29 sources with an official DOL target do not tie exactly (items 2 and 3 above); worker positions are not reproduced (item 5); the FY2025 Q1 difference of +1,733 is unresolved.
19. **Employer names are public-record names, some of them personal.** Names are reproduced from DOL's public files. Some employers are sole proprietors or professional practices whose name is, or contains, an individual's name. A heuristic screen found about 8,500 distinct names (about 0.4% of case rows) that may look like personal names; the screen has many false positives and unknown recall, and the redaction does not detect personal names. Public availability does not remove every legal obligation that may apply to your use. Correction and removal requests: `LICENSE.txt` section 7.
