How often is a code check right? Data published with the study at https://www.vallit.net/research/how-often-is-a-code-check-right summary.json Every measured result in the study, computed by scripts/research-export.mjs in the vallit-website repository from the benchmark's raw result files (their SHA-256 is listed under source.files). For each tier of the field runs, "sample" says whether the whole tier was judged ("all") or a sample, and a sample lists how it was spread over the codebases ("byRepository"). Both rejected samples were spread unevenly, so they carry no interval ("precisionCI": null). A set's verifier cost covers its result file; set E's second reading needed two runs, and "allRuns" gives what the result log records for both. lab-cases.csv All 560 test cases of holdouts A to E, one row per case and scoring: the frozen detector that scored it (and its SHA-256), static analysis alone or with the model verifier, the expected problems, what was detected, and the true positives, false positives and false negatives. Lists inside a cell are separated by "; ". A problem is written check:rule:file. score.mjs Recounts the lab table from lab-cases.csv: node score.mjs lab-cases.csv rules.md The labelling rules for the two checks, written before any test case existed. label-corrections.md Every label that changed after its set was first scored, with the rule sentence that decides it. The first scores in lab-cases.csv are the recorded ones, before these corrections. field-repositories.csv The 30 open-source repositories of the three field runs, each at the commit that was analysed, with its size and how long the static analysis took. field-judgements.csv Every finding that was judged in the field runs: the run, the tier the checks gave it (reported, likely, possible, or a sample of the rejected ones), the rule, and the verdict (real or false alarm). The repository, file and line are left out on purpose: where a real problem sits is for its maintainers to hear first. Please cite as Vallit (2026). How often is a code check right? Measuring on 560 test cases and 30 real codebases. https://www.vallit.net/research/how-often-is-a-code-check-right