Original research
How accurate are AI calorie counting apps? We tested 30 meals.
Every AI calorie tracker shows you a number. None of them show you an error bar. We build one of these apps, so we wanted to know how wide that bar actually is — not from a vendor’s marketing page, but from our own run.
So we tested it. Thirty meal photographs, three vision models, two passes each: 180 scans, all of them through the exact prompt and schema Biteline uses in production. The short answer is that photo estimation is imprecise at every price point. On the same plate, the three models’ calorie estimates were a median of 49.2 per cent apart. The most expensive model was not the steadiest. And a model can be badly wrong about what is on the plate while producing numbers that pass every arithmetic check you could write against them.
That is an awkward thing for a calorie-tracking app to publish. It is also the most useful thing we can tell you before you decide how much weight to put on any number a camera gives you — ours included.
What we found
- 30 openly-licensed meal photos, three Claude vision models, two passes each — 180 scans in total.
- Across the 29 photos containing food, the highest and lowest model estimate for the same plate differed by a median of 49.2%.
- The widest gap was a bowl of ramen: 250 kcal, 631 kcal and 739 kcal for one photograph.
- Asked twice, the same model moved its own answer by 8.1% (Sonnet 5), 10.0% (Opus 4.8) and 21.8% (Haiku 4.5) on average.
- All three models passed an Atwater arithmetic check — including the one that read tomato soup as peanut soup and nearly tripled its calories.
- This measures agreement and repeatability, not accuracy: none of these meals was weighed, so treat the numbers as a floor on the uncertainty.
Contents▾
How we tested it
The test had to be reproducible and it had to be honest about what it could and could not establish, so the design is deliberately plain.
- 30 photographs of meals from Wikimedia Commons, all openly licensed. They span what people actually photograph: pizza, a cheeseburger and fries, a Caesar salad, spaghetti bolognese, a sushi platter, curry and rice, steak frites, a full English breakfast, a club sandwich, tomato soup, tacos, ramen, fried rice, pancakes, an omelette, a roast dinner, paella, falafel, a burrito, dumplings, a stir fry, lasagne, fish and chips, a doner kebab, porridge, a yoghurt bowl, a smoothie bowl, couscous, shakshuka — plus one photograph of an empty banquet table as a negative control.
- Three models, served through Amazon Bedrock in the EU: Claude Haiku 4.5, Claude Sonnet 5 and Claude Opus 4.8. They sit at three clearly different price points, which is the axis every app in this category is really choosing along.
- Two passes each. Every model saw every photo twice, in independent runs. Three models × thirty photos × two passes is 180 scans.
- The production rubric. The prompt, the scoring guidance and the JSON schema were imported directly from Biteline’s own analysis code rather than rewritten for the benchmark. What the test measured is what the app would have shown you.
- Photo only. No dish names, no hints, no context. In the app you can type “lamb tagine with apricots” before you scan; here nothing was supplied, because the point was to measure the hard case.
One thing this design does not give us is ground truth. Nobody put these plates on a kitchen scale — they are photographs on the internet. That limit shapes everything below, and we come back to it in full.
On the same plate, the models were about half apart
The first thing to look at is not whether any model is right. It is whether three capable models, given the same photograph, arrive anywhere near each other. If they do not, at least one of them is wrong — and no amount of interface polish fixes that.
They did not. Taking the highest and lowest first-pass estimate for each photo and expressing the gap as a share of the lowest, the median spread across the 29 photos containing food was 49.2 per cent. On a bowl of ramen the three answers were 250, 631 and 739 kcal. On a bowl of tomato soup they were 315, 911 and 324.
| Dish | Haiku 4.5 | Sonnet 5 | Opus 4.8 | Spread |
|---|---|---|---|---|
| Ramen | 631 | 250 | 739 | 195.6% |
| Tomato soup | 911 | 315 | 324 | 189.2% |
| Fish and chips | 532 | 948 | 1199 | 125.4% |
| Sushi platter | 795 | 1036 | 1765 | 122.0% |
| Shakshuka | 892 | 418 | 479 | 113.4% |
| Falafel plate | 588 | 677 | 1213 | 106.3% |
| Porridge | 363 | 180 | 193 | 101.7% |
| Omelette | 790 | 514 | 410 | 92.7% |
| Caesar salad | 896 | 605 | 494 | 81.4% |
| Fried rice | 436 | 668 | 762 | 74.8% |
| Smoothie bowl | 651 | 1061 | 1131 | 73.7% |
| Paella | 1203 | 885 | 720 | 67.1% |
| Dumplings | 516 | 640 | 823 | 59.5% |
| Club sandwich | 1103 | 1726 | 1210 | 56.5% |
| Doner kebab | 1045 | 1490 | 1559 | 49.2% |
| Margherita pizza | 861 | 995 | 1265 | 46.9% |
| Cheeseburger and fries | 1125 | 1610 | 1632 | 45.1% |
| Roast dinner | 815 | 1042 | 736 | 41.6% |
| Steak frites | 1220 | 1160 | 1599 | 37.8% |
| Spaghetti bolognese | 712 | 948 | 972 | 36.5% |
| Burrito | 1303 | 1001 | 1057 | 30.2% |
| Chicken curry and rice | 815 | 1055 | 1012 | 29.4% |
| Vegetable stir fry | 300 | 388 | 359 | 29.3% |
| Tacos | 864 | 946 | 1098 | 27.1% |
| Couscous tagine | 695 | 628 | 753 | 19.9% |
| Lasagne | 720 | 640 | 765 | 19.5% |
| Yoghurt and granola | 681 | 607 | 692 | 14.0% |
| Full English breakfast | 1307 | 1380 | 1376 | 5.6% |
| Pancakes with syrup | 909 | 863 | 864 | 5.3% |
The pattern in this sample is worth naming, carefully, as an observation and not a law. The two narrowest results were a stack of pancakes (5.3 per cent) and a full English breakfast (5.6 per cent) — dishes made of countable, standard-sized units sitting in plain view. The widest were bowls: ramen, soup, shakshuka, porridge. In a bowl the portion is hidden below the rim and the fat is dissolved into the liquid, and there is simply less visual information to work from. If you eat a lot of soups, stews and curries, expect your log to be noisier than someone who eats plated food.
The same model, the same photo, twice
Disagreement between models is only half the picture. The other half is whether a single model agrees with itself. That is the noise floor: if a model swings by 20 per cent when nothing at all has changed, then a 20 per cent gap to another model tells you nothing.
| Model | Repeat difference | Widest single change | Valid responses |
|---|---|---|---|
| Claude Haiku 4.5 | 21.8% | 185% (tomato soup) | 58 / 60 |
| Claude Sonnet 5 | 8.1% | 34% (club sandwich) | 60 / 60 |
| Claude Opus 4.8 | 10.0% | 54% (tomato soup) | 60 / 60 |
Two things stand out. The first is that paying more did not buy repeatability: Claude Opus 4.8, the most capable and by far the most expensive model in the test, moved its own answer by 10.0 per cent on average, slightly more than Claude Sonnet 5’s 8.1 per cent. The second is that the cheapest model was roughly two and a half times noisier than either, and it was the only one that ever failed to return a valid response at all — twice, both in its second pass.
Put plainly: an app that tells you 250 kcal today and 300 kcal tomorrow for the same bowl of ramen is not lying to you twice. It is telling you, badly, that it does not know.
A number can be perfectly consistent and still be wrong
The obvious way to defend against a bad estimate is a validation rule. The cheapest one available is the Atwater cross-check: protein × 4 plus carbs × 4 plus fat × 9 should reconcile with the stated calories. It is arithmetic, not opinion, so it needs no reference meal — and it catches output that is internally incoherent.
All three models passed it. Mean reconciliation error was 2.9 per cent for Haiku 4.5, 1.3 per cent for Sonnet 5 and 1.4 per cent for Opus 4.8. Across 180 scans exactly one response was more than 10 per cent out.
That is the uncomfortable finding, not a reassuring one. Consider the bowl of tomato soup. Sonnet 5 returned “Creamy tomato soup”, 315 kcal and 5.6 g of protein — the identical answer on both passes. Opus 4.8 returned “Creamy tomato soup”, 324 kcal and 5.8 g. Haiku 4.5 returned “Peanut soup”: 911 kcal, 37 g of protein and 69 g of fat.
Check the arithmetic on that wrong answer
37 g protein × 4, plus 39 g carbs × 4, plus 69 g fat × 9, is 925 kcal against a stated 911 — an error of 1.5 per cent. The numbers reconcile. They are the internally consistent macros of a real bowl of peanut soup. The photograph was of tomato soup.
No runtime validation catches this. The model did not produce nonsense; it produced a coherent description of a different meal. On its second pass the same model looked at the same photograph, called it tomato soup, and returned 320 kcal. The first answer was 185 per cent above the second; put the other way round, the second was 65 per cent below the first. Same photograph, same model, same rubric.
Two results in the whole run were physically implausible, and both came from the same model on the same photograph: 109.5 g of protein in a 795 kcal sushi platter on the first pass, and 76.4 g in 601 kcal on the second — 55 and 51 per cent of the meal’s calories in protein. Those two a rule could have caught. The peanut soup it could not.
This is the single strongest argument for an app that lets you edit everything. If the tracker had shown you 911 kcal for a bowl of soup, the only thing standing between that and your daily total is you noticing.
The one thing all three models got right
The thirtieth photograph was a negative control: an empty banquet table, laid but with no food on it. All three models, on both passes, returned zero calories and an empty ingredient list. They titled it things like “Empty banquet table setting” and “No meal present”.
That matters more than it sounds. The failure mode of a vision model in this task is imprecision, not invention. None of them hallucinated a meal onto a bare table to have something to report. When these models are wrong, they are wrong about how much — and sometimes about which dish — but they are not manufacturing food that is not there.
What a scan costs, and how long it takes
Cost and latency are the two things in this test that are facts rather than judgements, and they explain something readers often find odd: why AI food scanners cap their free tiers.
| Model | Mean cost per scan | Median latency | p90 latency | Tokens in / out |
|---|---|---|---|---|
| Claude Haiku 4.5 | $0.0058 | 3,131 ms | 3,995 ms | 3,554 / 450 |
| Claude Sonnet 5 | $0.0195 | 5,417 ms | 6,975 ms | 4,156 / 467 |
| Claude Opus 4.8 | $0.0330 | 5,790 ms | 8,364 ms | 4,092 / 501 |
A photograph is an expensive thing to send to a model: around 4,000 input tokens each time. The cheapest model here is 5.7 times cheaper per scan than the most expensive and about 1.8 times faster at the median — and it is also the one that called tomato soup peanut soup. Cost is the easy axis to optimise. It is not the one that decides whether the number in your diary is any good.
Every scan you take costs the app that runs it real money, which is why Biteline’s free tier is capped at 3 AI photo scans a day while manual logging stays unlimited. Any app offering unlimited free photo scans is either subsidising them, using the cheapest model available, or both.
What this test does not show
This is the section most benchmark write-ups leave out, and it is the one that decides whether the rest is worth citing.
- There is no ground truth here. These are photographs, not weighed meals. We can say that three models put a bowl of ramen between 250 and 739 kcal. We cannot say what was actually in it. This measures agreement and repeatability — not accuracy.
- Agreement is a floor, not a ceiling. Where the models disagree, at least one is wrong. Where they agree, they could still all be wrong in the same direction, and this design would never see it.
- Thirty photos shows the shape of the problem, not a ranking. It is enough to establish that the uncertainty is large. It is nowhere near enough to declare one model better than another on any particular dish.
- These photographs are easier than yours. Wikimedia Commons food photography is well lit, well framed and shot from a flattering angle. A phone photo taken in a dim restaurant is a harder input, not an easier one.
- Two passes measures short-run variance only. It tells you nothing about how a model drifts over months.
- Models change. These figures were measured in August 2026, on those three model versions. They describe those models on that day, and nothing else.
If anyone — us included — quotes a single accuracy percentage for photo-based calorie estimation, ask what it was measured against. A percentage without a weighed reference meal behind it is a number someone chose.
What to do if you track with a camera
None of this makes photo tracking useless. It makes it a particular kind of tool, with a particular way of being used well.
- Read the number as a range. If your app says 620 kcal, the honest reading is “roughly 500 to 800, probably”. Decide things on that basis and you will not be misled by it.
- Fix the portion before anything else. Portion is the biggest single lever on the result, and it is the one you can actually judge — you are standing in front of the plate. Hand-size heuristics and reference objects are more reliable than they sound, and better than accepting a default.
- Scan the barcode when there is a label. A manufacturer’s figures are not an estimate. Photo recognition is for the plates that have no label.
- Be more sceptical of bowls than of plates. Soups, stews, curries and anything under a sauce hide the very things the estimate depends on. So do restaurant and takeaway meals, where the oil in the pan is invisible and usually generous.
- Judge the tracker on the trend, not the meal. A log that is biased in the same direction every day still tells you whether you are eating more than you think and whether that is moving your weight. Give it three to four weeks and watch the scale, not any individual entry.
- Distrust photo-perfect claims. Nothing we measured supports them, at any price point, for any model.
What we did with these results
We ran the benchmark to make a decision, and the decision was Claude Sonnet 5. Against the most capable model available it differed by 18.1 per cent on calories, where the cheapest model differed by 40.3 per cent; on protein specifically that gap was 17.5 per cent against 77.2 per cent. It was the steadiest on repeats, it returned valid, complete output on all 60 of its runs, and it costs about 40 per cent less per scan than the top model. The cheapest option was not a saving. It was a different product.
The rest of what we did was not about model choice at all:
- Every number in the result sheet is editable, and the portion slider sits at the top of it, because portion is where the error concentrates and rescaling it rescales every macro live.
- The app calls an estimate an estimate, in the interface, at the moment you are looking at one.
- Barcode scanning and manual entry are free on every plan, so the accurate route is never the paid one.
- We do not publish an accuracy percentage, because we have not run a weighed-food validation study. When we do, we will publish the method with it.
You can read the rest of what Biteline does and how it works on the home page.
Method, in full
- Photos. 30 openly-licensed meal photographs from Wikimedia Commons, used at their published resolution. One is an empty banquet table, included as a negative control.
- Models. Claude Haiku 4.5, Claude Sonnet 5 and Claude Opus 4.8, served through Amazon Bedrock in the
eu-west-3region. - Prompt and schema. Imported directly from Biteline’s production analysis provider, unmodified.
- Passes. Two per model per photo, run independently. 180 scans in total.
- Spread. (highest − lowest) ÷ lowest, across the three models’ first-pass calorie estimates for one photo.
- Repeat difference. Mean absolute percentage difference between pass one and pass two for the same model and photo. The negative control is excluded from every percentage, because its calorie count is zero.
- Atwater check. protein × 4 + carbs × 4 + fat × 9, compared against the model’s own stated calories for the same meal.
- Failures. Claude Haiku 4.5 returned schema-invalid output on two photographs in its second pass; those two are excluded from its repeat figure. Claude Sonnet 5 and Claude Opus 4.8 returned valid, complete output on all 60 of their runs each.
- Cost and latency. Measured per request at published rates, not estimated from a price list.
- Date. August 2026.
If you are writing about this and want the underlying per-scan results, email contact@getbiteline.com and we will send them.
Keep reading
- How to count calories without weighing your foodA scale answers one of the four questions a meal asks. Learn what your own hand holds, spend your precision on the oil, and eyeball the broccoli.
- How to log restaurant and takeaway mealsYou cannot see the oil, weigh the rice or read the recipe. Break the plate into parts, compute the drinks exactly, and round up.
- How many calories should I eat to lose weight?Maintenance minus a deficit, worked through with real numbers — including the case where the safety floor overrides the pace you asked for.
- Protein, carbs and fat: what split actually matters for weight loss?Only one of the three numbers is worth setting on purpose, and it is protein. The rules, the guard rails and two worked examples you can check.
Track a meal in about ten seconds
Biteline photographs the plate, names the dish and breaks it down into calories, protein, carbs and fat — then lets you correct every one of those numbers. 3 AI photo scans a day are free and manual logging always is.
Get Biteline