All randomness flows from explicit seeds. Datasets: a master seed spawns
per-record independent streams (numpy.random.SeedSequence.spawn), so the
full dataset is bit-reproducible AND any single record can be regenerated
alone (tests/test_generator.py verifies both).
Every simulated record's provenance holds: generator version, HITRAN data
version, per-record seed/spawn key, instrument config id, and every sampled
noise parameter value.
Baseline trainings pin seeds and hyperparameters; hyperparams.json is
written at train time including validation curves.
Physics correctness is enforced by dual independent implementations
(Faddeeva vs quadrature; time-domain lock-in vs Fourier coefficients) that
must agree to 0.1% / 1% on 1000 random points — see Quality gates.
v0.2 introduces the spektran generate CLI as the primary generation
interface; scripts/generate_dataset.py remains available as an alternative.
v0.5 adds spektran train --baseline <name> for one-command baseline
training with automatic data generation. All commands support --json
output for AI agent integration (see AGENTS.md).
v0.6 extends full reproducibility to CRDS, FTIR, and DOAS modalities.
Each new generator uses the same SeedSequence.spawn per-record pattern
as TDLAS and NDIR. CRDS and FTIR use HITRAN demo line lists; DOAS uses
synthetic cross sections (by design — UV/Vis cross sections have different
structure from IR line-by-line parameters). All three modalities share the
same instrument-sampling and noise-chain infrastructure, and their
provenance records are schema-validated identically.