0002 — Registry as catalog, not gate; eras as rows, not fields¶
Date: 2026-09-09 Status: Decided
This ADR records the registry design decision. Examples of additional era rows describe how to extend the catalog; they do not imply that those rows ship in the current registry. See Datasets for the implemented rows.
Context¶
Dataset facts were spread across datasets.py, inventory.py and fetch.py,
with monthly stated twice and nothing cross-checking them. The SIA/APAC
family (7 datasets) was registered in all three but reachable from neither
the public API nor the CLI. Unifying the facts into one row was the obvious
fix, and it raised two design risks worth recording.
Decision 1 — the registry is a catalog, not a gate¶
A single-source registry can become a lock: "the eleven keys in this dict are the only things you may ingest." We reject that reading.
- The pipeline takes
Datasetvalues:import_scope(dataset: Dataset, ...). REGISTRYis a dict of pre-built values — convenience, not authority.- An uncurated dataset is ingested by constructing a
Dataset(...)and passing it through the same code path. There is no second, untyped mode. - The row holds identity and location only. Behaviour goes in an importer module (the existing hybrid pattern), never a flag on the row.
Decision 2 — era variants are separate rows¶
Live listing of ftp.datasus.gov.br shows era directories for three of four
families (SIM/CID9 vs CID10, SIASUS/199407_200712 vs 200801_,
SIHSUS/199201_200712 vs 200801_). A single ftp_dir: str cannot reach
them.
We considered an eras: tuple[Era, ...] field and rejected it: one row implies
one dictionary, but CID9 and CID10 have different columns. The row would claim
a schema it does not have.
Instead, each supported era gets its own row with its own dictionary (for
example, sim_obitos and sim_obitos_cid9). A cross-era union is a deliberate act by the user, which is
correct: it surfaces a real schema break instead of hiding it.
Note (later): the CID-9 era is now registered as its own row, sim_obitos_cid9.
Consequences¶
ftp_dir: strsurvives;sincebecomescoverage: (first, last | None).cadence(how DATASUS publishes) andpartition_by(how we store) are separate fields. Deriving one from the other would let a storage change silently alter filename generation.dictionaryis a field so an ad-hocDatasetcan point at a YAML outside the package. Uncurated does not mean schemaless: without a YAML the import fails fast rather than falling back to all-strings.- Adding an era is the same operation as adding a dataset: one row, one YAML.
- The row's schema becomes semi-public (users may construct it). Changes are additive and keyword-only.
- Tier 3 tests (spec §6) validate each row's
ftp_dirandcoverageagainst the live server on a schedule, because centralizing facts centralizes the blast radius of a wrong one.
Amendment 2026-09-12 — release directories share a row; names are readable¶
DATASUS publishes preliminary files beside the final ones under the same
names (SIM, SINASC, SINAN, e-SUS Notifica; verified live 2026-09-11/12) and
the layout does not change at that boundary. Decision 2 splits rows on
schema, so a publication status is not a new row: the row declares
prelim_dir, discovery reads both directories, every lake row carries
_source_release, and outdated() names the scopes whose release moved.
Eras with different columns in different directories still split.
Row names are <sistema>_<conteúdo> in full Portuguese words matching the
DATASUS file-type description. ALIASES and get_config() were removed;
there is no compatibility surface.
Amendment 2026-09-21 — outdated() compares files, not only the release¶
outdated() now compares each active publication's recorded files with what
the server lists today (path, and size/mtime when recorded), not only the
release, so a moved directory or a same-name republish is reported too.