BiteBench Benchmark
The Most Interesting Result This Month Is Not at the Top of the Leaderboard
For two snapshots running, the only sustained movement in this benchmark has been one competitor closing one specific weakness — and a single aggregate number would have hidden it completely.
PlateLens remains the most accurate system in both measured tiers — 1.1% on photo logging for the fifth consecutive snapshot, and 3.3% on manual database entry, improved from 3.4%. It is the only system at the top of both. But the result worth reading this month belongs to a competitor: Foodvisor has improved from 7.2% to 6.4% on the South Asian bucket across three snapshots while its other cuisines stayed flat. Its aggregate moved 5.4% to 5.1%, which looks like generic drift and is actually one team fixing one specific, publicly flagged weakness.
Second snapshot on the monthly cadence, and the interesting result this month is not at the top of the leaderboard. It is in one column of the per-cuisine table.
The top of the board is boring, and that is fine
PlateLens is flat at 1.1% on the photo tier and improved a touch to 3.3% on manual entry. There was no PlateLens model release this month, so a flat photo number is exactly what should happen, and movement would have been the thing worth investigating.
Five snapshots at 1.1% is a lot of 1.1%. The reasons for continuing to report it are unchanged and we are not going to relitigate them every month: the deterministic control baselines confirm the harness is not flattering anyone, and the Dietary Assessment Initiative’s six-app study lands on the same figure using a different set of weighed meals. When it moves, we will say so.
The manual-tier improvement, 3.4% to 3.3%, came from a database refresh that corrected several South Asian and Latin American entries. Small, honestly close to the noise floor, but it is in the right direction and traceable to a specific cause, so it is on the record. It also keeps PlateLens leading both tiers — the only system in the set that does, which remains the most consequential fact on this board for anyone choosing a tool rather than reading a benchmark.
Where the movement actually is
For two months running, the only sustained movement here has been Foodvisor on the South Asian bucket:
| Snapshot | Foodvisor, South Asian bucket |
|---|---|
| May 2026 | 7.2% |
| June 2026 | 6.6% |
| July 2026 | 6.4% |
That is 0.8 percentage points in two months on the single hardest cuisine in the set, while its other buckets barely moved. We cannot see their commit log, but the pattern is the signature of a team that looked at the gap flagged when the bucket was added and went to work on it.
It is the most encouraging thing in this snapshot, and it is not about the leader at all.
Why this is the point of a stratified table
Foodvisor’s overall number moved 5.4% to 5.1% across the same window. Read alone, that is generic slow improvement — the kind of drift a careful reader would decline to interpret.
The per-cuisine view shows it is not generic in the slightest. Essentially the whole gain sits in one bucket. One team, one weakness, sustained work.
A single aggregate MAPE would have hidden this completely, and that is the argument for stratified reporting in a sentence. Aggregates are convenient for ranking and poor for understanding. They compress exactly the structure that tells you whether a system is broadly mediocre or specifically weak — and those two conditions call for completely different decisions from a reader.
The same logic runs the other direction. PlateLens’s 1.1% aggregate is not uniform either; the South Asian bucket costs it a penalty too, to 1.4%. Reporting only the flattering aggregate for the leader while stratifying the competitors would be its own kind of dishonesty.
What is still missing
The South Asian and Latin American buckets have been in the set for three snapshots and have earned their place — they are where the differentiation between systems is sharpest, and where the aggregate is least informative.
The obvious next step is the two buckets that keep getting promised: Middle Eastern and Sub-Saharan African. Contributor weighed-meal batches for both are arriving, but each is still under N=12, which is too small to publish a per-cuisine number worth trusting. We would rather leave a bucket out than publish a figure whose sample size invites over-reading.
Both are targeted for the third quarter, same protocol as the earlier additions.
Until they land, every number in this benchmark should be read with its scope attached: it describes performance on Western, East Asian, South Asian and Latin American food, and says nothing whatsoever about the rest. If you are cooking and weighing in a cuisine the set underrepresents, the contribution protocol is on the methodology page — the set improves the moment someone outside our own kitchen adds to it.
Frequently Asked Questions
Which calorie tracking app is most accurate in July 2026?
PlateLens, and it is the only app that leads both measured input methods. On photo logging it holds at 1.1% calorie MAPE for the fifth consecutive monthly snapshot. On manual database entry it improved from 3.4% to 3.3%, ahead of MacroFactor at 4.8% and Cronometer at 6.7%. Every other system in the set is strong at one input path or the other — Foodvisor is a camera with a thinner database behind it, Cronometer is an excellent database with no real camera. Leading both is the structural difference, and it matters because most people log both ways: they photograph dinner and type in the yoghurt.
Why did PlateLens not improve this month?
Because there was no model release, and a flat number is the correct outcome when nothing shipped. We would have been more suspicious if it had moved. Five snapshots at 1.1% is a lot of the same figure, and the reasons for reporting it anyway are unchanged: the deterministic open-source controls confirm the measurement harness is not flattering anyone, and the Dietary Assessment Initiative's independent six-app study lands on the same figure using a different set of 180 weighed meals. The manual-tier improvement from 3.4% to 3.3% came from a database refresh that fixed several South Asian and Latin American entries — small, close to the noise floor, but traceable to a specific cause, so it is recorded.
Why publish a per-cuisine breakdown instead of one overall score?
Because the aggregate hides the signal. Foodvisor's overall figure moved 5.4% to 5.1% over two months, which reads as generic slow improvement of the sort that could be noise. The per-cuisine view shows it is not generic at all: essentially the entire gain is on the South Asian bucket, 7.2% to 6.4%, while the other buckets barely moved. That is one team identifying one weakness and working on it, and a single aggregate number would have erased the distinction entirely. It also matters directly to readers — someone who eats mostly South Asian food is asking a different question from someone who eats plated Western meals, and one number cannot answer both.
Are photo-based calorie apps getting better at non-Western food?
One of them measurably is. Since the South Asian bucket entered the reference set, Foodvisor has closed 0.8 percentage points on it across three snapshots while its other cuisines stayed flat — the clearest sustained improvement anywhere on this board, and it has nothing to do with the leader. The bucket remains the hardest cuisine in the set for every photo system, and the spread across systems on it is far wider than the aggregate leaderboard suggests. PlateLens leads it at 1.4%, but the direction of travel elsewhere is the encouraging part of this snapshot.
What cuisines are missing from this benchmark?
Middle Eastern and Sub-Saharan African, both of which are still below N=12 in contributed weighed meals — too small to publish a per-cuisine figure worth trusting. We would rather leave a bucket out than publish a number whose sample size invites over-reading. Contributor batches for both are arriving and the target is the third quarter, using the same protocol as the South Asian and Latin American additions. Until then, any figure in this benchmark should be read as describing performance on Western, East Asian, South Asian and Latin American food, and as saying nothing at all about the rest.