Tests of Accuracy with Random Points (Lemos et al. 2023). For each trial a true parameter is drawn from the prior, data are simulated, and posterior samples are drawn conditioned on those data. Given a random reference point, the fraction of posterior samples closer to the reference than the truth is the credibility level of the smallest distance-based credible region that contains the truth. For a calibrated posterior these fractions are uniform, so the expected coverage probability (ECP) at credibility level alpha equals alpha.
Arguments
- fit
An
nsbi_npefit fromnpe(), annsbi_nlefit fromnle(), or annsbi_nrefit fromnre(). With an NLE or NRE fit every trial is a separate MCMC run, so start with a smalln_tarpand raise it once the cost is known.- simulator
The simulator used for inference; called once per trial (see nsbi_simulator).
- prior
The prior to draw the true parameters from, and the reference points when
references = "prior"(defaults tofit$prior). As insbc(), the coverage claim is about the prior the fit was trained on, so the default is what tests this fit; another prior tests how the fit does on parameters it was not calibrated against. It must cover the same parameters as the fit.- n_tarp
Number of TARP trials (fresh (theta, x) pairs). At least 2: the true draws are standardized by their own spread before distances are computed, and a single draw leaves that spread undefined.
- n_posterior_samples
Posterior draws per trial.
- references
How to draw reference points:
"uniform"(default, uniform over the hyper-rectangle spanned by the true parameter draws, as in the paper) or"prior"(draws from the prior).- sim_args
Named list of extra arguments passed to every simulator call; see nsbi_simulator.
- seed
Optional seed.
- ...
Passed to
posterior(), which is how the MCMC controls (n_chains,warmup,thin,sampler) reach an NLE or NRE fit.
Value
An object of class nsbi_tarp with the per-trial coverage values
and the ECP curve. Plot it with plot_tarp().
Details
Unlike sbc(), which ranks each parameter marginally, TARP is a joint
test: it can detect posteriors whose marginals are calibrated but whose
correlation structure is wrong. Distances are computed after z-scoring each
parameter (using the spread of the true draws), so parameters on different
scales contribute comparably.
A trial whose simulation returns non-finite output is dropped, which lowers
the effective n_tarp. A trial whose posterior comes back with fewer than
n_posterior_samples draws is an error, for the same reason as in sbc():
a trial scored on a different number of draws is not comparable to the rest.
References
Lemos, Coogan, Hezaveh & Perreault-Levasseur (2023), "Sampling-based accuracy testing of posterior estimators for general inference", ICML. doi:10.48550/arXiv.2302.03026
