HomePage
All grantsPage
R&D GrantGrant
SchaalklaarGrant
Ecologiepremie+Grant
Horizon EuropeGrant
Innovation DeductionGrant
Strategic Transformation SupportGrant
Payroll Tax ReliefGrant
Why an expert?Page
CompanyPage
ProjectsPage
NewsPage
JobsPage
Contact usPage
Grant scanTool
Privacy PolicyPage
HomeWhitepapersThe yardstick keeps moving

The yardstick keeps moving

Whitepaper

Why nobody knows exactly what an AI system can do, and why that became a business risk this year. Of 445 benchmark papers examined, 16% report a margin of uncertainty beside the number, and the European legislator postponed the rules for high-risk AI to December 2027 because nobody had written down in time how to measure conformity.

View whitepaperDownload PDF

Book a free grant scan →

← Back to whitepapers
/ Frequently asked questions

Frequently asked questions

Why is it so hard to measure what an AI system can actually do?+

Because the yardstick itself moves. The best-documented measure shows that the length of tasks a model can handle has roughly doubled every three months since 2024 — but when the organisation behind that measure replaced its task set and test environment in January 2026, the scores of the same models shifted 57% down and 55% up. What you measure depends heavily on how you measure.

What does it mean that AI benchmarks are unstable?+

That a number on a leaderboard is not a fixed property of the model, but shifts with the chosen test set. Of 445 benchmark papers examined, only 16% report any uncertainty margin on the published figure. A score without an error margin looks exact, yet says little about what the model does in your context.

Why has measuring AI performance become a business risk?+

Because every figure ultimately drives a decision: an investment, a hire, a contract clause or a tender. If that figure is not stable, the decision rests on sand. The only randomised field trial with real developers even found a 19% slowdown where a time saving was expected — the measured effect diverged from the assumed one.

What does the EU AI Act change for high-risk AI conformity, and why was it delayed?+

The high-risk obligations require a system to demonstrably meet measurable criteria. The European legislator hit the same wall as the industry — nobody had defined in time how to measure that conformity — and postponed the high-risk AI rules to December 2027.

How can a company reliably assess an AI system's capabilities before buying?+

By not relying on a single leaderboard figure, but testing on your own tasks, in your own environment, with an explicit uncertainty margin. Ask for the test conditions behind every number, and treat a score without an error margin with suspicion. The question is not whether AI works, but whether you can establish that it works for your case.