Detection

How often is a code check right?

A small scale model of a classical temple on a stone plinth by a still sea; far behind it the real temple of the same design, vast and half lost in fog.

We measured two of our code checks on 560 test cases, then on 30 open-source codebases. On the test cases, 84% to 92% of what they reported was real. The first time they ran on real code, 4 of 131 findings were real (3.1%).

Many apps built with AI tools hold other people’s data and take other people’s money. The mistakes that hurt them most are rarely a missing header. They are mistakes about who may do what: a record anyone can change, a price the browser gets to choose, a payment the app believes because someone said so.

Broken access control is the first entry of the OWASP Top 10 of 2021. Of the applications behind that list, 94% were tested for some form of it, and it had more occurrences in the data than any other category. In 2025, an insufficient row-level security policy in apps generated by one AI builder was recorded as a vulnerability of its own. The builder disputes it: each customer, it says, is responsible for their application’s data.

Vallit has checks that read an app’s code to find these mistakes. Before a check tells someone their app is broken, we want to know two things: how often what it reports is real, and how often it misses a real problem. A check that raises false alarms teaches people to ignore it, and then they ignore the one that matters too. A check that misses gives false comfort. So we measured both, first on test cases written without access to the detector’s code, then on real code.

On five test sets, each scored once by a version of the detector that had not seen it, 84% to 92% of the reported problems were real. On the first ten real codebases, 4 of 131 (3.1%) were. With a second reading by two language models, 9 of the 25 findings reported as Likely on ten other codebases were real (36%). On ten more, nothing was reported as Likely, and none of the 5 Possible findings or the 20 rejected candidates we sampled was real.

These are small samples from open-source projects rather than apps built with AI tools, and no outside security expert has reviewed the labels yet. The rest of this study describes how we measured and what each result can and cannot show. The chart below puts the test sets and the first two runs on real code on one scale.

Share of reported findings that were real

Mostly real on test cases, far less often on real code. Each test set with analysis alone, and the first two runs on real code; on the second, only findings reported as Likely count. Lines show 95% intervals.
The numbers
WhereChecksReal / reportedPrecision95% interval
Test set Av1, analysis alone29 / 3387.9%72.7% to 95.2%
Test set Bv2, analysis alone54 / 6090.0%79.9% to 95.3%
Test set Cv3, analysis alone61 / 6692.4%83.5% to 96.7%
Test set Dv4, analysis alone50 / 5590.9%80.4% to 96.1%
Test set Ev5, analysis alone59 / 7084.3%74.0% to 91.0%
First ten codebasesv3, analysis alone4 / 1313.1%1.2% to 7.6%
Next ten codebasesv4, reported as Likely9 / 2536.0%20.2% to 55.5%

What we measured

The study measures two code checks. The first, Who can reach which data, follows each way into an app that it knows to each database call. It knows route handlers, API routes, Server Actions, tRPC procedures, Express and Hono routes and Supabase Edge Functions. The second, Where money can leak, follows every checkout and payment webhook it finds and every place that grants paid access. Between them they ask ten questions, each a rule with a written definition of when it is a problem and when it is not (the rules).

Who can reach which data asks four:

  • Can anyone change or delete data without signing in?
  • Can users give themselves more rights?
  • Can signed-in users reach other customers’ records?
  • Does the request get to say which user it is?

Where money can leak asks six:

  • Can anyone tell your app a payment went through?
  • Can customers choose what they pay?
  • Can paid features be switched on without paying?
  • Does a repeated payment event credit twice?
  • Do customers who cancel keep paid access?
  • Will real payment events fail their verification?

Both checks start with a static analysis that reads the code without running it. It follows each request through helpers and data layers in other files and records what was established on the way: a signed-in user, a role, an owner, a verified signature. Then it proposes candidates, without a language model.

Then comes a second reading. Each candidate goes to Claude Haiku 4.5 with the code path the analysis traced and the guards on it. The model may ask for up to four further files or declarations by name, twice. It is asked to name the line that decides. If Haiku rejects the candidate, it is dropped. Otherwise Claude Sonnet 5 reads it again. When Sonnet confirms, the finding is reported as Likely. When Haiku confirmed and Sonnet rejects, or the deciding code is out of reach, it is reported as Possible. When Haiku was unsure and Sonnet rejects, it is dropped. If Sonnet cannot be reached, Haiku’s verdict stands; if neither model can be reached, the finding is kept and marked Possible rather than dropped. The models never receive the whole repository at once.

How we measured

The ten rules were written down before a single test case existed. They describe what the application does, not how any detector works. Each test set was then written without access to the detector’s code. A case is a small repository of one to nine files that reads like real application code. Of the 560 cases, 308 contain one or more real problems and 252 are near misses: the same shape, built safely. When a label later contradicted the rules, it was corrected in a published log, with the sentence of the rule that decides it. The first score of that set stayed as it was recorded.

Before a set was opened, the detector was committed and its source files hashed. The set was scored once, and that score is the one we report. Only set E’s second reading needed two runs: a rate limit cut the first short, and the verdicts that had failed were read again. Set A was scored by version 1, B by version 2, C by 3, D by 4 and E by 5. Each set then became a regression test for the versions after it. Every test-set result names the SHA-256 of the detector bytes it measured; the later runs on real code name the detector’s commit.

A report counts as right only when the check, the rule and the file all match a label. Precision is the share of reported problems that are real. Recall is the share of real problems that were reported. Every interval is a 95% Wilson score interval; with sets this small, the intervals are wide.

Versions 3, 4 and 5 each ran once on ten open-source repositories they had never seen. Each finding was judged against the rules from the finding and the repository’s code alone, without the detector’s code or the models’ verdicts. The sample included candidates the models had rejected, so that real problems thrown away would show. After the last run, three repositories were also audited in part for problems the checks might have missed.

On test cases

On the five sets, 84% to 92% of what the analysis reported was real, and it found 64% to 84% of the real problems. Among the safe near misses it raised an alarm on 1 to 5 cases per set.

Every set was new to the detector that scored it, and on every one that detector missed real problems. Recall rose from 64% on A to 84% on C, fell to 68% on D and came back to 80% on E. The number moves with the set as much as with the detector, which is why we report all five and not the best one.

Some kinds of problem stay hard. Over the five first scores, the checks found 93% of payment webhooks that skip the signature. They found only 53% of subscriptions that keep access after a cancellation, and 58% of requests that name their own user. The numbers under the chart below list every rule.

Precision and recall on each test set

Precision stayed high; recall moved with every new set. Analysis alone, before the second reading: each set scored once, by a version of the detector that had not seen it. Lines show 95% intervals.
The numbers
SetDetectorCasesPrecision (95% interval)Recall (95% interval)Cases with their problem foundSafe cases flagged
Av18087.9% (72.7% to 95.2%)64.4% (49.8% to 76.8%)29 of 441 of 36
Bv212090.0% (79.9% to 95.3%)75.0% (63.9% to 83.6%)52 of 662 of 54
Cv312092.4% (83.5% to 96.7%)83.6% (73.4% to 90.3%)55 of 662 of 54
Dv412090.9% (80.4% to 96.1%)67.6% (56.3% to 77.1%)49 of 662 of 54
Ev512084.3% (74.0% to 91.0%)79.7% (69.2% to 87.3%)54 of 665 of 54
QuestionFoundRecall (95% interval)False alarms
Can anyone change or delete data without signing in?36 of 4383.7% (70.0% to 91.9%)7
Can users give themselves more rights?28 of 3971.8% (56.2% to 83.5%)1
Can signed-in users reach other customers’ records?42 of 5182.4% (69.7% to 90.4%)12
Does the request get to say which user it is?28 of 4858.3% (44.3% to 71.2%)2
Can anyone tell your app a payment went through?25 of 2792.6% (76.6% to 97.9%)1
Can customers choose what they pay?24 of 2982.8% (65.5% to 92.4%)3
Can paid features be switched on without paying?26 of 3672.2% (56.0% to 84.2%)0
Does a repeated payment event credit twice?15 of 2560.0% (40.7% to 76.6%)3
Do customers who cancel keep paid access?10 of 1952.6% (31.7% to 72.7%)2
Will real payment events fail their verification?19 of 2190.5% (71.1% to 97.3%)0

Scoring the same sets again

After a set is scored, its misses and false alarms are studied and many are fixed. A few labels are corrected under the written rules, and the set joins the regression tests. Scored again on September 23, 2026, sets A and B reach 100% precision and recall, C 97.3% on both, and D 94.5% and 93.2%. That is what a detector that has learned from its test looks like. Only a first score says anything about code the detector has not seen, which is why we wrote a new set for every version instead of improving against one. The chart below shows each set the first time and later.

Each test set, first score and later

Most of the gap is the detector learning from the sets. Each set the first time and later as a regression test; some of the gap is labels corrected after the first score. E, the newest set, still scores what it scored the first time.
The numbers
SetPrecision, firstPrecision, laterRecall, firstRecall, later
A87.9%100.0%64.4%100.0%
B90.0%100.0%75.0%100.0%
C92.4%97.3%83.6%97.3%
D90.9%94.5%67.6%93.2%
E84.3%84.3%79.7%79.7%

On real code

Then version 3, the most precise of the static versions on its own test set (92.4%), ran once on ten open-source codebases. They were Dokploy, Dub, Langfuse, Linkwarden, LobeHub, Rallly, Next SaaS Stripe Starter, Midday, Unkey and Vercel’s Chatbot. It reported 131 problems; 4 of 131 (3.1%) were real.

61 of the 127 false alarms were writes the analysis believed anyone could make. Real applications enforce sign-in and permissions in their own wrappers, ability checks and admin clients, and an analysis that recognises the common ways to do it does not recognise theirs. The test cases were realistic, but they were small, and they did not contain enough of that variety.

We suspect this is not specific to our checks. A check built and tested on small written examples tends to score better on new examples of that kind than on real applications.

A second reading

The analysis is fast and explainable, but it reads permissions the way a pattern does. So we added a reader that follows the code path the way a reviewer would. On the two test sets where it ran, it removed 2 of 5 false alarms on D and 4 of 11 false alarms on E and dropped no real problem. A reader can only take findings away, so recall stays where the analysis left it.

Version 4 with the second reading ran once on ten more codebases: Cap, Zero, Typebot, Cal.com, Civitai, Formbricks, Chatbot UI, Saasfly, Unsend and Vercel Platforms. The analysis proposed 205 candidates. The models confirmed 25, left 41 as Possible and rejected 139. Of the 25 reported as Likely, 9 were real (36%). Of the 41 reported as Possible, 2 were real. Of 31 rejected candidates we judged, none was real. The chart below shows the judged findings in each tier.

Judged findings on real code, by tier

Likely holds most of the real problems. Judged findings from the second and third runs on real code. The rejected candidates were sampled unevenly across codebases, so they carry no interval.
The numbers
RunTierIn the tierReal / judgedShare real95% interval
Second run on real code, version 4Likely259 / 2536.0%20.2% to 55.5%
Second run on real code, version 4Possible412 / 414.9%1.3% to 16.1%
Second run on real code, version 4Rejected1390 / 310.0%n/a
Third run on real code, version 5Likely00 / 0n/an/a
Third run on real code, version 5Possible50 / 50.0%0.0% to 43.4%
Third run on real code, version 5Rejected670 / 200.0%n/a

The 11 real problems were of three kinds:

  • 5 places where a signed-in user could reach records of another organisation
  • 4 checkouts that take the price id from the request without checking it against a list
  • 2 Server Actions that write to the database without checking who called

We name the repositories we measured, not where the problems are. That is for their maintainers to hear first.

Most false Likely findings came from three patterns:

  • writes that only record a status the server chose, behind a view policy
  • random invite or reservation ids used as keys
  • a server token compared inside a nested wrapper

Version 5 was changed to handle all three; we have not measured that change on those codebases again, because they are no longer unseen. Version 5 ran once on ten further codebases: Puter, OpenPanel, Ghost, Bigcapital, Karakeep, Firecrawl, Papra, Stack Auth, next-forge and Webstudio. It proposed 72 candidates; the models confirmed none, left 5 as Possible and rejected 67. All 5 Possible findings and 20 of the 67 rejected ones were judged, and all 25 were false alarms.

Three of those codebases were then audited for problems of the ten kinds that the checks might have missed. Each audit covered part of the code, traced every payment path and recorded what it left out. They read entry points or route files: 160 in Karakeep, 107 in Papra and 60 in Stack Auth. None found a problem of the ten kinds in what it covered. So on the last run, nothing false was reported as Likely and nothing real was missed where we looked. With no real problem found in what was read, this cannot tell us how many a check would find.

Both checks now run with the second reading on every connected repository. What the models reject never reaches the report, and the report shows the rest without a label for how sure we are.

What we don’t know yet

The 30 codebases are open-source projects, from large products such as Cal.com and Ghost to starters and templates. They were not chosen as apps built with AI tools. Apps made with Lovable, Bolt, v0, Cursor or Replit, the ones Vallit is for, can only be measured when their owners ask. The test cases and the judgements on real code were made against the written rules without access to the detector’s code. No outside security expert has reviewed them yet. Borderline calls are recorded with their reasons.

The precision of Likely rests on 25 judged findings from one run, with a 95% interval from 20% to 55%. The next run reported none as Likely, and one more run would move it. The rejected candidates were sampled unevenly: on the second run, 10 of the 99 rejected on Formbricks were judged, and none was real. With so few judged, the 95% interval there reaches 28%, so we cannot say how many real problems the models threw away. Recall on real code rests on partial audits of three repositories, in which no problem of the ten kinds was found. NestJS decorator controllers, Remix and SvelteKit are not discovered as entry points; Bigcapital, a NestJS server, contributed 0 entry points to the last run. In the product, an inexpensive scoring model rates every candidate before the language models read it. It only records the rating and drops none, so its absence from these runs changes no result.

Two newer checks, Who your app believes and Who can spend your AI budget, have been measured on test sets only. Their final set was scored on September 23, 2026 with the second reading. One rule missed the bar written down for it: 2 of 4 cases of someone posing as another user were found. No real repository has been run for them yet. The next study we want to run is on apps built with AI tools, with their owners’ permission.

We publish the data behind this study. Every measured result on this page is computed from the benchmark’s raw result files, whose SHA-256 it lists. Beside it are the 560 test cases as scored, every judged finding on real code without where it sits, and the 30 repositories at the commits we analysed. So are the labelling rules and every later correction to a label. A script that recounts precision and recall on each test set comes with them, and a readme that says what each file holds. The test cases’ code and the detector are not public. If you research this problem and want to check our work more closely, or you maintain one of these repositories, write to info@vallit.net.

Citation

@techreport{vallit2026codechecks,
  title       = {How often is a code check right? Measuring on 560 test cases and 30 real codebases},
  author      = {{Vallit}},
  institution = {Vallit},
  year        = {2026},
  month       = sep,
  url         = {https://www.vallit.net/research/how-often-is-a-code-check-right}
}