RASS: Risk-Audited Budget Selection for Compact NeRF Benchmark Subsets

NeurIPS 2026 · Evaluations & Datasets Track · Poster

73×fewer scenes: 3,521 → 48 (RASS-48)
0.182Wilson 95% lower bound at 48 scenes, target 0.08
22 / 22constraints RASS-96 passes in the four-method audit
64 / 140DL3DV scenes certified for four methods

TL;DR

Evaluate a new NeRF method on 48 scenes instead of 3,521, and know how often that shortcut holds.

RASS does not hand-pick one small subset. It audits the procedure that draws subsets: every candidate must match the full benchmark's PSNR, SSIM, and LPIPS means and distributions within declared tolerances, and RASS recommends the smallest size whose 95% Wilson lower bound on the pass rate meets a target. On Nutrition5k, RASS-48 is the screening subset (73× fewer scenes) and RASS-96 is the recommended reporting subset, which also passes joint audits over three and four methods. The protocol transfers to DL3DV.

Abstract

Full-dataset NeRF evaluation can require training and evaluating thousands of scenes, making exhaustive comparison expensive even when the benchmark is fixed. We introduce RASS, a risk-audited protocol for constructing compact NeRF benchmark subsets with explicit fidelity criteria. Given full-dataset per-scene logs, RASS repeatedly samples candidate subsets, checks whether their global metric means and score distributions match the full benchmark within user-defined tolerances, and recommends the smallest evaluated subset size whose Wilson lower confidence bound exceeds a target pass probability. We instantiate RASS on 3,522 Nutrition5k-derived food scenes using Zip-NeRF logs for the formal audit and PSNR, SSIM, LPIPS, and Kolmogorov–Smirnov (KS) distributional tests as fidelity criteria. The released 48-scene balanced operating point, RASS-48, is the lowest-cost evaluated screening subset that passes our declared global-fidelity audit and is intended for low-cost early ablations and rapid screening. A stricter target pass probability, pmin = 0.20, selects RASS-96, which we recommend as the compact reporting subset: it also passes joint fidelity audits over three and four view-synthesis methods. We also evaluate alternative subset generators, including uniform random sampling, farthest-first selection, and facility-location diversity sampling, and find that the RASS audit can validate different generators under the same fidelity contract. On DL3DV, a re-declared contract certifies 64 of 140 scenes for four methods under proportional allocation. We release the RASS-48 and RASS-96 scenes, configurations, descriptors, regime labels, seeds, audit scripts, and permitted logs to enable reproducible, compact NeRF evaluation.

How RASS works

The four steps of RASS: describe, generate, test, audit
RASS audits a subset-generation procedure, not one lucky subset. Each scene gets a 57-D descriptor, and K-Means splits the benchmark into six regimes. For every budget, RASS draws 400 balanced candidates (equal scenes per regime) and tests each one against six constraints: the PSNR, SSIM, and LPIPS mean gaps and the KS distance of each metric's distribution. The recommended budget is the smallest one whose 95% Wilson lower confidence bound on the pass rate reaches the target pmin. One passing draw is then exported by a fixed rule and audited on its own.
The 400 candidate draws at 48 scenes
At 48 scenes, 88 of the 400 candidate draws pass all six constraints. Seven more match every mean but shift a distribution's tail, and the KS guardrail rejects them. The pass rate is 0.22 and its Wilson lower bound 0.182, above pmin = 0.08. A yield of at least 0.08 means a passing subset needs at most 12.5 cheap log checks on average.

Results

Cost-reliability frontier with the RASS-48 and RASS-96 operating points
The frontier trades scenes for reliability. RASS-48 is the smallest evaluated budget that meets pmin = 0.08, and RASS-96 is the first that meets the stricter pmin = 0.20.
Operating pointScenesPassing drawsWilson LCBUse
RASS-484888 / 4000.182Screening and ablations (pmin = 0.08)
RASS-9696113 / 4000.241Compact reporting (pmin = 0.20)
Full benchmark3,521––Leaderboards and per-regime claims
Share of each tolerance that RASS-48 uses
The exported RASS-48 subset, eight scenes from each of the six regimes, passes the global audit on its own. Its largest use of any tolerance is 80%, for the LPIPS KS distance.

Beyond one method and one dataset

The same audit runs as a joint event over several methods: a candidate passes only if every method meets all of its constraints at once. Equal allocation per regime is biased when regime sizes differ, and sampling each regime in proportion to its size removes that bias.

PopulationMethods in the joint eventSmallest certified sizeRASS-48RASS-96
Nutrition5k, 3,473 scenesZip-NeRF, Feature-Splatting, Instant-NGP7214 / 1616 / 16
Nutrition5k, 3,473 scenes+ nerfacto9619 / 2222 / 22
Nutrition5k, 3,473 scenes+ BioNeRF12023 / 2826 / 28
DL3DV, 140 scenesnerfacto32––
DL3DV, 140 scenes+ splatfacto (3DGS)40––
DL3DV, 140 scenes+ TensoRF64––
DL3DV, 140 scenes+ Instant-NGP64––

Smallest certified size: Wilson LCB ≥ 0.08 with proportional allocation (equal allocation does not reach the target up to 120 scenes on Nutrition5k or up to 80 on DL3DV). Subset columns: constraints passed. Adding the pairwise method orderings to the three-method event changes no pass count. On DL3DV, Instant-NGP was trained for 16k steps, so its scores are not comparable across methods.

Using RASS

Workflow for a new method
  1. Train and evaluate only on RASS-48 (screening) or RASS-96 (reporting).
  2. Compare with the released per-scene logs on the same scenes, where paired differences remove additive scene effects.
  3. Report the subset version and the thresholds.
  4. Train on the full benchmark only for leaderboard placement or to contribute logs.
  5. Avoid comparing subset results with full-benchmark results.

The audit runs on a CPU in seconds from the released logs. This reproduces the paper's frontier, the RASS-48 and RASS-96 audits, and regenerates RASS-96 exactly:

git clone https://github.com/GCVCG/RASS
cd RASS/rass_kaggle_artifact_anonymous
pip install numpy pandas scipy pyyaml
python scripts/reproduce_tables.py
Scope. RASS certifies global fidelity to the full benchmark's means and distributions, for the declared methods and tolerances. It does not certify regime-level fidelity: with per-regime constraints, no budget up to 360 scenes passes. RASS-96 passes the three- and four-method audits but not the five-method one, and the multi-method and DL3DV budgets hold under proportional allocation, whose DL3DV frontier was added after the audit contract was fixed. Neither subset replaces full-benchmark leaderboards.

Slides

RASS slides, title slide

Released artifact

  • RASS-48 and RASS-96 scene lists, the 3,521-scene audit population, and the cross-method scene list.
  • 57-D scene descriptors, their definitions, and the six regime labels.
  • Per-scene Zip-NeRF, Feature-Splatting, Instant-NGP, nerfacto, and BioNeRF logs on the Nutrition5k-derived scenes, and DL3DV per-scene metrics for four methods.
  • Thresholds, seeds, versioned event configurations, and the audit scripts.
  • Croissant metadata with Responsible AI fields, a dataset card, and a validator.

Code: github.com/GCVCG/RASS. Data: Kaggle. Raw Nutrition5k imagery and complete NeRF outputs are not redistributed; DL3DV-derived files are released under CC BY-NC 4.0.

Acknowledgments

This work was partially supported by TRIVISION (PID2025-173459NB-C21), PULSAR (PID2025-173459NB-C22) and PID2022-141566NB-I00, financed by MICIU/AEI/10.13039/501100011033 and by FEDER/UE; AIA2025-163919-C51 (AEI, MICIU); CNS2022-135480 (MICIU/AEI, FEDER, NextGenerationEU/PRTR); ICREA Academia'2022 (Generalitat de Catalunya); the FPI doctoral fellowship; and RES resources at BSC MareNostrum5 (IM-2025-3-0008).

Universitat de Barcelona Universitat Pompeu Fabra Barcelona Supercomputing Center

Citation

@inproceedings{almughrabi2026rass, title = {{RASS}: Risk-Audited Budget Selection for Compact {NeRF} Benchmark Subsets}, author = {AlMughrabi, Ahmad and Serret, Flavi{\`a} and Marques, Ricardo and Radeva, Petia}, booktitle = {Advances in Neural Information Processing Systems (NeurIPS), Evaluations and Datasets Track}, year = {2026} }