How accurate is Voco?

Last verified 2026-08-09

Every calorie tracker claims accuracy. Almost none publish how they measured it, or how often they were wrong.

This is our scorecard. It is re-run and re-published on every release. If the numbers get worse, they get published anyway.

Which way we miss · mean signed error +2.11%

We miss slightly high, never quietly under. A tracker that quietly runs low stalls your cut for weeks before you notice. Ours leans 2% the safe way.

Signed calorie error vs the spec band of −3% to +5%
200 cases · 2026-08-09
How big the miss is · calorie error 9.67%

The average miss on the 159 cases where a published source says what the answer is. On a 700-calorie meal, about 68 calories.

159 hard-answer cases · 2026-08-09
False badges ✓ Verified 0 / 295

295 trap cases built to bait a wrong ✓, run three times. Zero false badges, every run.

295 traps × 3 runs · 2026-08-09
How often ✓ appears ~30%

of logged items currently earn the checkmark. Everything else is honestly graded High or Estimated, on the item.

411 items · 28 users · 2026-08-01

Every number on this page is the average of three full runs of the whole test, back to back, on the exact parser that ships, release 5.2. No single run failed a bar. A single run is a screenshot, not a measurement.

How to read this

Where we land, food by food

The test: 200 written food logs, phrased the way people talk, scored against sources we cite per case. USDA, current manufacturer labels, and chains’ own published nutrition. No consumer aggregators.

Calorie error by category · 200 cases · 2026-08-09
scored against a published number scored against a range, see below

Chain food is our worst number: 15.08%. Order a burrito bowl without weighing it and expect the estimate to be off by roughly a sixth. It is also the category the verified match layer exists to intercept: when a named chain item matches a curated record, the estimate is replaced by the published numbers.

The gap between weighed and unweighed is the argument for owning a food scale. Saying “about five ounces” instead of weighing costs roughly eight points of accuracy.

One more defence against the quiet undercount: the cooking oil you did not mention still gets counted 87.9% of the time (same 200 cases, 2026-08-09). That misses about one in eight, and we say so.

The receipts · bars set in the spec before the parser was chosen · 2026-08-09
MetricResultBar
Mean signed error+2.11%−3% to +5%
Calorie error, all 200 cases8.04%≤ 12%
Calorie error, weighed food5.05%≤ 8%
Protein error8.50%≤ 12%
Macros add up to the calories96.50%≥ 90%

The all-200 figure, 8.04%, is lower than the 9.67% headline only because 41 of the 200 cases are range-scored. We publish both and lead with the higher one.

The 1.84% we will not brag about

Vague inputs like “a plate of stir fry” score 1.84% (30 cases, 2026-08-09), the best number on this page, and it is not a precision result. Nobody publishes what was on your plate, so those cases are scored against a range built from our own written portion assumptions, and the app runs on the same assumptions, so most answers land inside the band by construction. It measures consistency, not truth. Compare us on 9.67%, not on 1.84%.

When we stamp ✓ Verified

✓ means the item’s numbers came from a curated record, a manufacturer label, USDA, or the chain’s own published nutrition, and replaced the estimate outright. You also have to have stated a measure. Most items do not earn it, by design.

The traps are the near-misses that have actually bitten us: a sports drink versus its zero-sugar twin, a bar versus its mini, a food that is not in the database at all but looks like one that is. Zero false badges across 295 traps, in each of three runs. With 295 traps, the strongest honest claim is a true false-badge rate below about 1.01%, at 95% confidence. We will not print a bare “100%”.

Behind the badge: 1,687 curated records, and a matcher that declines to guess about 85% of the time on purpose. A missed badge costs a moment. A wrong one costs the reason to trust all the others.

A sealed set of 56 further traps has never been run. It runs exactly once, at App Store launch, and gets published whatever it says. That is the only way to know we have not been tuning against our own answer key.

Reserved · badge precision on real user traffic not yet measured

Every ✓ the system would issue is being logged against the live result. When that read-out exists, the number goes here, whatever it says. The ~30% coverage figure above is how often the badge appears, never a substitute for whether it was right.

What this page does not measure

Version history

One row per release, each from a fresh run. No row is ever deleted.

2026-05-30
Parser locked after a six-way comparison. 8.84% · +2.30%
2026-07-25
Per-item confidence added; deterministic settings pinned. The measured +4.45% broke the original ±3% bar, so the bar was widened. 8.4% · +4.45%
2026-07-31
Test set rebuilt from source; roughly 40% of the old answers were wrong and were corrected. The ruler changed, not the app. 7.7% · +1.66%
2026-08-01
A revised parser tested and not adopted. 8.9%, unshipped
2026-08-02
Shipping parser re-measured. 7.32% · +1.59%
2026-08-09
This page. Release 5.2: example foods our parsing instructions shared with the test set were removed, and everything re-measured. The trap set grew from 29 to 295 cases. 9.67% hard-answer · 8.04% all · +2.11%

Want the long version? Every case carries its source ID and retrieval date, and we intend to open-source the harness and the answer key so anyone can re-run this. Until then, ask us anything: hello@voco.fit.