Viewpoint Past Its Used-By Date Software does not spoil. It does not smell, or separate, or change...
Accuracy & Oversight
Measuring the Signal
The Paper Trail Behind Our Accuracy — what happened when we logged every correction, rejection, approval and challenge to Prodeen’s output across a cross-section of food and beverage companies for 90 days: what the log showed, and what we fixed.
| Research Window | 15 May – 14 Aug 2026 · 90 days |
| Outputs Produced | 3,499 |
| Food & Beverage Companies | 16 |
| Verdicts Logged | 397 |
When someone on your team gets something wrong, you already know how to handle it: someone catches it, it gets logged, it gets fixed, and you can point to the record if anyone ever asks. That's not a special process. It's just how you run a team you're accountable for.
Most of what gets sold to teams like yours doesn't hold itself to that standard, and it fails in two different directions.
One is the legacy database: decades of entries, every jurisdiction you can name, a pitch built entirely on size. What it won't tell you is how often what it hands you still matches the rule as it stands today, not as it stood when the entry was last touched. Nothing published, nothing to check. A bigger haystack doesn't make the needle easier to find. Volume was never proof of accuracy.
The other is AI with nothing checking it. Fast, confident, unaudited. Whatever it gets wrong reaches you looking exactly like what it gets right. No flag. Nothing catching it before it's already your problem.
Prodeen sets the standard the other two skip. Across a cross-section of food and beverage companies, every session is on the record: every correction, rejection, approval, or challenge to what the AI produced gets logged, classified, and counted, on live work, not a sample.
That's the check the legacy databases don't publish and the unchecked AI tools don't have at all. It also has to keep running on a schedule, not just once, because the regulations underneath it don't hold still either.
Two questions worth asking whoever you use today:
-
How often does a person actually check what it produces, and what happens when they disagree with it?
That's the human-in-the-loop test. -
Can you show me the number, and the method behind it, not just the claim?
That's the published-transparency test. If you're not sure how your current provider would answer either one, that's worth noticing.
The short version
-
A cross-section of food and beverage companies used Prodeen for live work over 90 days, and every correction, rejection, approval, or question logged against what the AI gave them was recorded: 397 pieces of direct feedback.
-
The regulatory reading itself held up in 95% of that feedback. The other 19 went to a person, got fixed, and became a permanent test we now run before shipping anything new.
-
Most of what got sent back wasn't wrong information at all. It was the wrong format, a missing source link, or the tool creating a new document instead of updating the one already open.
-
The five most common causes of rework are already identified, and most of the fixes are live this week.
-
We publish this every quarter, on the same terms every time, not just when the numbers look good.
Why we bother measuring this
A wrong answer in food regulatory work is not a minor inconvenience. It's a mislabeled product, a filing built on the wrong market's limits, a recall. Nobody in your seat wants someone's word that a tool is trustworthy, whether that word comes from an AI sales deck or a database that's been around so long nobody thinks to question it anymore. You want evidence, and you want it to keep showing up, quarter after quarter, not just in the pitch.
That's also why every output runs past a person, not just a model. Every conversation gets reviewed. Every time the log shows a reviewer pushing back on something, we record it, classify it, and publish what it adds up to.
What the log caught this quarter
A cross-section of food and beverage companies used Prodeen for live work over 90 days: real dossiers, real compliance questions, real market-access calls. Every session is logged, and we tracked every time the log showed a reaction to something the AI produced, classifying each one.
That happened 397 times. That's not a coverage gap: most usage doesn't prompt a reaction one way or the other, the same way most satisfied customers don't leave a review. The 397 is how many times the log recorded a reaction, not how many outputs we chose to check, and it's the number every rate in this report is built on.
In 378 of those cases, 95%, the regulatory reading itself held up. The other 19 needed an actual correction to the finding, and each one went to a person for individual review before anything moved forward: exactly what a review layer exists to catch.
| 95% | 19 |
| Of the 397 times the log recorded direct feedback, the regulatory reading itself held up without correction. | Cases where it didn't, and the finding was corrected. Every one went to a person before anything moved forward. |
That number, 19, matters more than any other figure in this report, because it's the only one that means the underlying regulation was actually wrong rather than just delivered in a way someone wanted to change.
Each one was fixed in the moment, not left open in a queue: the log shows the right answer recorded before anyone moved on. Every one of those 19 also becomes a permanent test we run before shipping any change to the product, so the same miss can't happen the same way twice, and that check runs continuously, not once a quarter.
WHERE EVERY ONE OF THE 3,499 OUTPUTS WENT

UPDATED — Check two (29 May – 27 Aug 2026): where every one of the 3,581 outputs went

INSIDE THE 397 THAT GOT DIRECT FEEDBACK

UPDATED — Check two (29 May – 27 Aug 2026): inside the 464 that got direct feedback

Of the 25 rejections, 17 were a filing mix-up, not a wrong answer. Of the 254 corrections, seven in ten were formatting or a citation fix, not a challenge to the finding.
None of this waits for a quarterly report to get acted on. What the log catches feeds directly into what ships next, which is what the next section walks through.
Where the mistakes cluster, and what we're doing about it
Almost nine out of every ten corrections and rejections this quarter trace back to five repeat causes, not a scattered list of one-offs, ranked by how often they came up.
Building a new document instead of updating the one you already had (17 of 25 rejections): the single biggest cause of a full rejection. Fix shipping this week: the AI defaults to updating an existing document whenever you reference one, and says so before it acts.
Citing a regulation by description instead of a working link (59 cases): the single most repeated complaint in the log. Every citation now has to be a real, clickable link, or it gets flagged for you to check, not buried mid-document.
Leaving a batch job half done without saying so (62 cases): covering things like updating a list of markets or substances and quietly stopping partway through. From now on, the AI lists what it's covering up front and reports what's left if it can't finish.
Boilerplate nobody asked for (58 cases): auto-added "limitations" sections, summaries that run too long, data that should be a table arriving as a paragraph instead. These defaults go away this week unless you ask for them.
Treating "only use this source" as a suggestion (fewer cases, highest cost): a named jurisdiction or source not respected as a hard rule, the failure that costs the most trust when it happens. A named source or market is now a hard limit, and the AI says so if it can't answer within it.
None of these five is the AI getting the science or the regulation wrong. They're the AI not doing the job the way it was asked, which, as any quality team knows, is a cheaper problem to fix than a real knowledge gap. Three longer-term changes make these fixes permanent rather than just prompted: a check that blocks publishing if a citation isn't a real link, a screen that shows which document will be updated before the AI runs so a routing mistake gets caught before it happens, and visible progress on any multi-item job so nothing finishes silently half done.
We're up to roughly 19 regulatory test cases and about 250 more covering the formatting and citation issues above: the start of a test library built entirely from real mistakes real regulatory teams caught, not a generic benchmark.
Signal, not noise
It would be easy to sell this on volume: how many regulations we track, how many jurisdictions we watch, how many alerts go out each week. That's not the number that matters to you, and it isn't the one we open with.
You don't need every regulatory update in every market. You need the ones that touch what you actually make, sell, and file: your categories, your additives, your markets, delivered as something you can act on without doing the filtering yourself.
A feed that covers everything and leaves you to sort out what's relevant isn't thorough. It's more work with your name on it. Scoping every signal to your actual portfolio, instead of surfacing everything and calling it coverage, is what makes the accuracy numbers above worth having in the first place: a shorter list of things you're expected to act on has to earn your trust in a way a firehose never does.
The standard we're setting
We're publishing this methodology, this quarter's numbers, and the fix list above, and we'll keep doing it every quarter, on the same terms every time. That's the actual commitment here, not a fixed accuracy percentage. Anyone can assert a number like that. Few will show you the log behind it, and fewer still will commit to updating it in public on a fixed schedule regardless of what it shows. That's the standard we're setting, and we intend to keep raising it: tighter sampling, clearer rules, and a wider net, every quarter.
Where your data goes
Accuracy is the subject of this report, but it's not the only question worth asking before proprietary formulation or filing data goes into any tool. Prodeen is ISO 27001 certified: the standard for how sensitive information gets managed and protected. That's a separate certification from the numbers above, and worth confirming independently rather than taking on trust, the same as everything else in this report.
See what it finds in your own portfolio
Run Prodeen against your actual categories, additives, and markets. Test it on your portfolio →
Appendix
A note on methodology, and what it doesn’t cover yet
This report counts what happened when the log recorded a reaction to Prodeen's output: corrected it, approved it, rejected it, or challenged it. It doesn't grade the output nobody reacted to, and it doesn't claim to. Most of what Prodeen produced in the 90-day window was used without a recorded reaction either way, so the figures above describe what happened in the cases the log shows close engagement, not a claim about everything Prodeen ever produced.
For reference, counting every output Prodeen produced in the window, including the ones nobody flagged, the same correction figure works out to about 1 in 200 (0.54%). We don't lead with that number in this report because it includes a lot of output nobody actually checked, which makes it a weaker signal than the 397 figure above.
This is us grading our own work, not an independent audit. What makes it worth something anyway is that the classification rules were fixed before the quarter started and don't move to flatter the result.
Methodology note: figures in this paper come from a systematic review of the log of recorded working sessions between Prodeen and reviewers, drawn from a cross-section of food and beverage companies, 15 May to 14 August 2026. Company and individual identities are withheld throughout; every number here is reported across companies and total review volume, never by account or by name.

