An analytics assistant can produce valid SQL and still answer the wrong business question. Before granting access to more tables, decide what each business metric means and how a reviewer can verify the answer.
This is an updated replacement for our earlier article about the reported 21% to 95% improvement in AI analytics. It separates the published findings from a practical test a smaller business could run.
What the 21% and 95% figures refer to
In its June 3, 2026 article, Anthropic reports that Claude answered no more than 21% of analytics questions correctly on its evaluations without skills. With skills, it reports aggregate results consistently above 95%. Its implementation also includes governed data foundations, reference material and validation.
These are Anthropic’s reported internal results, not a benchmark for your company or a VasthavM client outcome. They do not establish that installing a semantic layer alone produces the same improvement.
Agree on a question that can be checked
For a small pilot, choose one recurring question from a real meeting. For example: how many customers placed a second paid order last month? This is an illustrative test, not a claim about any client’s data.
Write down whether refunds count, how test accounts are excluded, which date sets the reporting period and how customer records are matched. Have the responsible person approve those choices before asking an agent to calculate the result.
- Name the metric and its owner.
- Identify the approved source and the unit represented by each row.
- Record the filters, reporting period and known exceptions.
- Keep a checked answer and the calculation used to obtain it.
Give the assistant a narrow route to the answer
A practical starting point is one read-only view or a bounded query tool. The assistant should not need to choose among every historical table to answer the pilot question. Keep the allowed calculations and supporting definitions alongside the integration.
Make the output reviewable: include the reporting period, source, relevant filters and any missing input. If the business question cannot be answered with the approved data, the response should say what is missing rather than silently substitute another metric.
Test wrong answers as well as good ones
Build a small set of expected answers before the pilot. Include duplicate customer records, a late-arriving order and a question outside the tool’s permissions. Record the setup so the same test can be repeated after a change.
Review the result and the route used to obtain it. A correct number produced with an incorrect filter can fail as soon as the data changes. Separate a calculation error from a wrong definition, missing data or an unsupported explanation.
Assign ownership after the first demonstration
Decide who updates the definitions when the business changes. Put a review date on the pilot and agree what should stop automated reporting, such as a missing source or a failed comparison against the checked answer.
Start with one question and an accountable reviewer. Bring the current report, the underlying source and a description of where people disagree. That is enough to begin scoping an analytics assistant without promising a particular accuracy percentage.



