Categorical FE evidence and release limits

Common categorical time effects are an experimental FE extension in the 3.3.1 release. The release decision accepts the combined evidence below for experimental use, with distribution through the private GitHub repository for authorised users under the unchanged proprietary licence. Experimental status does not establish statistical qualification or grant a public-use licence.

Version 3.3.2 retains this experimental feature and its evidence limits. The results on this page belong to the original 3.3.1 development and release candidates; the namespace rename is not a new statistical qualification. Use the original runtime for their fitted artefacts. See the 3.3.2 upgrade guide.

Supported contract

The extension requires a balanced panel of fitted units, a Normal likelihood, shared proper priors for structural parameters and one common categorical time-effect column. The identified unit/category nuisance coefficients use the flat-prior error-contrast construction. The transformed media/control design must pass the declared estimability checks.

Level forecasts require known categories and all fitted units. Reusing a category assumes its fitted effect still applies. Media-only predictions, manual scenarios and historical incrementality retain their existing model-implied contracts. Categorical CRE, named RE, categorical pointwise LOO/WAIC, calibration and fixed-budget optimisation remain unavailable.

See Fixed-effects Estimator for the exact likelihood, scaling, persistence and operation boundaries, and Choose an Estimator for the unit-only presets.

Development evidence before E1

This is a historical development snapshot dated 13 September 2026, before the prospective E1 screen. Later release-specific evidence must be reported separately; it must not replace the original results below.

EvidenceCompleted resultLimit
Software and numerical contractsGeometry, likelihood, estimability, persistence and operation-gate checks were completed for the implementation candidatePassing tests do not establish repeated-data recovery or interval calibration
Bounded gridOne synthetic dataset in each of ten cells met the monitored sampling diagnosticsOne dataset per cell does not estimate a cell-specific success rate
Initial linear posterior references136 of 140 comparisons met their fixed MCSE-plus-quadrature thresholds; four did notThese comparisons share posterior draws and are not independent trials
Supplemental linear target checksDensity and gradient comparisons passed at 45 fixed points across five linear cellsChecks at selected points do not prove agreement everywhere
Retained-draw investigationMCSE reproduction, batch/CDF sensitivity and conditional-moment checks were completed without new posterior fitsThese checks investigate discrepancies on the existing data
Sampler-seed follow-upSix fits across three existing datasets, with two new sampler seeds per dataset, met the monitored diagnostics and all 168 reference comparisonsThe four original mismatches remain recorded; this is not a fresh-data qualification study

The original mismatches concerned L1’s sigma mean, L2’s first-channel amplitude variance and 97th percentile, and L3’s second-channel amplitude 97th percentile. The follow-up supports finite Monte Carlo variability as an explanation. It does not prove universal calibration or the absence of all implementation defects. Counts from the original and follow-up checks must remain separate.

The stochastic examples cover linear media, geometric adstock followed by logistic saturation, delayed adstock followed by logistic saturation, and logistic saturation followed by geometric adstock. Weak within-category variation is included and can produce wide effect intervals. These examples do not statistically validate every transformation exposed by the library.

E1 outcome: original gate failed

The fixed campaign completed on 13 September 2026 in 484.96 seconds (8 minutes 5 seconds), within the 30-minute total cap. Its three jobs took 221.79, 113.25 and 149.85 seconds, each below ten minutes. All ten fresh-data cases completed without retries, replacement datasets or timeouts.

E1 failed its release gate. The delayed-adstock case N3 recorded five sampling divergences; the required count was zero. All monitored R-hat and bulk/tail effective sample sizes met their thresholds in every case. These other diagnostics do not cancel the divergence failure.

CellResultDivergencesMaximum R-hatMinimum bulk ESSMinimum tail ESS
L1Passed01.002723942.22485.7
N1Passed01.001651888.01411.7
L2Passed01.004004586.92274.6
N2Passed01.002681823.9963.5
L3Passed01.001573885.72771.1
N3Failed: divergences51.002161651.31412.7
N4Passed01.002361821.91446.4
L1_weakPassed01.000213438.12541.9
N1_weakPassed01.004691596.01394.2
L1_unassociatedPassed01.003143887.22455.3

All 45 independent linear density/gradient points and all 140 linear posterior comparisons passed their fixed tolerances. The 54 focused software tests and installed-wheel verification also passed. These are distinct evidence types; software and reference checks do not override the failed sampling gate. The retained records include all parameter/estimand errors, intervals, widths, prediction results, warnings and posterior files. A single dataset in each cell cannot estimate that cell’s repeated-data coverage or success rate.

E1 sampled the frozen candidate labelled 3.4.0rc1. While it was running, the release target was corrected to 3.3.1. That initial correction changed version and disclosures only. The final candidate also includes the later delayed-adstock numerical fix described below. E1 was not restarted or relabelled. Its source identity is distinct from the corrected numerical implementation and final distributions. Version 3.4.0rc1 was not published; 3.3.1 is the authorised release version.

The original development mismatches remain in the historical section above. This fresh-data result neither erases those mismatches nor qualifies Q1. The subsequent release decision accepts the combined evidence described below. The separately declared N3-S1 check is not a retry within E1 and does not alter its outcome.

Numerical correction and N3-S1 follow-up

Normalised delayed-adstock weights could all underflow to zero for some positive-alpha inputs allowed by the existing prior. The correction subtracts the minimum squared lag distance before exponentiation, cancelling a common factor in the normalised ratio. The mathematical kernel, priors, public signature, convolution and unnormalised branch are unchanged. Alpha-zero cases with an undefined normaliser remain undefined.

The implementation passed 654 deterministic tests, with 18 sampling-related cases deselected. These include independent high-precision value/gradient references, the demonstrated underflow cases, endpoint behaviour, batching, convolution modes, half-life paths and categorical FE contracts. Those checks verify the correction; they do not establish the cause of E1’s divergences. The failing leapfrog states were not retained in E1, and its saved draws do not demonstrate the underflow mechanism.

The separately declared N3-S1 check completed on 13 September 2026 in 63.15 seconds, within its three-minute total cap. It used the corrected implementation, the original N3 dataset and the same data/sampler seeds, with four chains, 1,000 tuning steps and 1,000 retained draws per chain. It completed once, with no retry or timeout.

N3-S1 diagnosticResult
Divergences0
Maximum R-hat1.00324
Minimum bulk ESS1961.62
Minimum tail ESS1701.74
Non-finite maximum energy errors0
Maximum tree-depth hits0

All required retained values were finite. Four Python warnings about forking a multithreaded process were retained; the run completed without a deadlock. All ten retained output hashes and 648 source hashes were verified, and the nine reported estimand means and intervals were independently reproduced from retained draws.

N3-S1 passed its bounded checks. It supports successful sampling for this corrected implementation on this one dataset. It does not prove that the fix caused the disappearance of divergences, establish recovery on new datasets or convert E1 into a passed campaign. The original four development reference mismatches and E1’s five divergences remain part of the evidence.

Evidence applicable to release 3.3.1

Release 3.3.1 uses the same statistical source and dependency contract as N3-S1. Subsequent changes concern release disclosures and the package verifier’s retention regression tests. Relative to E1, the statistical source change is confined to the normalised delayed-adstock transform and its docstring. The linear, geometric and saturation-first E1 cells use unchanged statistical paths; their results remain evidence for those paths, not fresh executions of the final candidate. N3-S1 supplies the separate delayed-adstock check on the corrected source. The complete ten-cell grid was not rerun after the correction. Exact package and documentation verification is retained separately from sampling evidence.

Release decision

On 13 September 2026, the release owner accepted E1’s original failure, the deterministic numerical correction and the separately declared N3-S1 pass as the basis for experimental categorical FE in 3.3.1. This combination does not satisfy E1’s original all-cells-pass rule. The decision explicitly accepts that departure; the failed campaign, original development mismatches and evidence limits remain recorded.

Distribution is through the existing private GitHub repository for authorised users under the unchanged proprietary licence. No public-use licence or generally supported categorical FE status is granted. Release approval does not complete Q1 or establish repeated-data calibration or causal identification.

Original predeclared experimental release requirement

The original E1 protocol scheduled one new dataset in each of the same ten cells, with fixed seeds, four chains and unchanged primary chain lengths. Its total execution cap is 30 elapsed minutes, divided into jobs of at most ten minutes. The cap includes setup, checks, sampling, reference calculations, reporting and artefact retention. There are no diagnostic retries or replacement datasets.

Every required example, diagnostic, numerical comparison and retained artefact must pass the predeclared engineering gates. A timeout, missing result or failed mandatory check leaves release on hold. Successful execution permits experimental release review; it does not itself authorise publication.

The release review must report all ten planned cells, completed and failed attempts, diagnostics, numerical comparisons, realised errors, intervals, widths and any budget stop. Different cells must not be pooled into a replication count or coverage rate. The completed release-specific report must identify the exact source and distribution hashes checked.

Claims that remain unqualified

Q1, the separate repeated-data statistical qualification protocol, remains unexecuted. E1 cannot establish repeated-data bias, interval calibration, general robustness or a high probability of successful inference on new datasets. Its small number of examples has limited precision.

Categorical effects remove the configured common category components. They do not by themselves establish causal identification or eliminate every time-varying confounder. Media contributions and spend interventions remain model-implied quantities. Use an explicit identification argument and appropriate substantive evidence for causal decisions.