# Methodology

## Source
IRS Tax Exempt Organization Search bulk e-file data: per-year index files (`index_2022.csv` ... `index_2026.csv`) and monthly zip batches of return XML under `https://apps.irs.gov/pub/epostcard/990/xml/`. IRS Exempt Organizations Business Master File extracts (`eo_ca.csv` for identifying California foundations; national `eo1.csv` to `eo4.csv`, last modified 2026-09-07, as independent evidence that a recipient is an organization). Source SHA-256 values for the index files and BMF are recorded in `metadata/release_metadata.json`; every filing row carries its XML file's SHA-256 and zip member CRC-32.

## Population
1. All index rows of type 990PF whose EIN is in the California BMF extract and whose period ends between 2022-01 and 2025-12 (41,001 index rows).
2. Each XML was opened by range-reading its zip batch. Some index rows point at the wrong monthly batch (the file is in a sibling batch); these were recovered by object-ID lookup across all batches. 1 object ID could not be found in any batch and is listed as an exclusion.
3. Final inclusion is decided by the filing itself: filer address state = CA and tax year (`TaxYr`) in 2022-2024. 518 filings whose own address is not California were excluded; 7,113 outside the tax-year range were excluded.
Result: 33,369 filings parsed in scope, 33,046 operative after supersession (10,504 / 11,185 / 11,357 for TY2022 / 2023 / 2024), from 12,083 foundations.

## Extraction
Grants come from the Part XV supplementary-information groups: `GrantOrContributionPdDurYrGrp` (paid, line 3a) and `GrantOrContriApprvForFutGrp` (approved for future payment, line 3b). Each filing is parsed twice, by an XML-tree parser and by an independent text-pattern parser that does not share code with it, and the two are compared row by row; differences found: 0. Row sums are reconciled with the foundation's own reported Part XV total.

## Duplicate and amended filings
311 foundation-period groups had two or three filings (e.g. an original and an amended return). We tested the groups before choosing a rule: in 231 groups the grant content of the filings was identical, in 22 the same number of lines had different values and in 58 the lines differed. The rule adopted is: within a foundation and period end, the filing with the latest IRS submission timestamp is operative and the others are marked `SUPERSEDED` (with `superseded_by_object_id`). In 75 groups the latest-timestamp filing is not the one with the highest object ID, so the choice mattered; the timestamp is the filer's actual submission time. Ties are broken by index year, then object ID. Superseded filings are kept in the filings file with their own totals but contribute no rows to the grant files. A foundation that changed its fiscal year can have two filings with the same tax year but different period ends; both are operative.

## Recipient classes
Every grant line is classified before release; `ORG_CONFIRMED` and `ORG_PROBABLE` are released, `INDIVIDUAL`, `AMBIGUOUS`, `ORG_UNVERIFIED` and `AGGREGATE` (a placeholder or bucket such as VARIOUS, OTHER CONTRIBUTIONS or SEE ATTACHED SCHEDULE instead of a named recipient) are withheld. `ORG_CONFIRMED` needs corroboration: the foundation's status text corroborated across unrelated foundations, or an exact normalized name plus state match in the IRS Exempt Organizations Business Master File (national extract, retrieved 2026-10-11; used only as corroboration because it is not exhaustive: foreign organizations, governments, churches and recent organizations can be absent). `ORG_PROBABLE` needs several indicators of an organization without such corroboration. A person-looking name whose only organization evidence is the filer's status text is released as `ORG_PROBABLE` unless it carries a personal-name signal (a name token that filers themselves labelled as individuals elsewhere) and no institution word or BMF match; those are withheld as `ORG_UNVERIFIED`. The classifier uses the recipient's status text as filed (for example PC, 501(c)(3), I), the XML name field, organization words in the name, purpose wording (scholarship, tuition, hardship), whether the same name appears as an organization in other foundations' filings, whether the funder mostly makes individual grants, name text that itself says the grant is to an individual (for example a personal name followed by a parenthetical "grant to individual"), a person's name routed care-of an organization, a Schedule K-1 payee line naming a person, a named crowd-funding campaign, and a name lexicon learned from lines the filers themselves labelled as individuals. The reasons are in `recipient_class_reason`. Do not read `RecipientPersonNm` as meaning an individual; the XML puts many organizations in that field.

## Fields we derived
`grant_id`, `recipient_class`, `recipient_class_reason`, `row_flags`, the ZIP5 columns (street addresses are never released), and the supersession columns. Everything else is as filed. Amounts are whole dollars as filed; negative, zero and missing amounts are flagged, not corrected.

## Reproducibility
Two independent full builds, one with ten and one with seven parallel shards, produced byte-identical merged tables and an identical population list (SHA-256 recorded in `QUALITY_REPORT.md`). All output tables are sorted and contain no timestamps.
