Skip to content
BlogDocument AI evaluation

An AI assistant for your documents: test questions and limits

A convincing answer may use an outdated document. Evaluate its claim, source and response when information is missing.

Test worksheet: Document AI assistants: evaluating answers

A convincing answer may use an outdated document. Evaluate its claim, source and response when information is missing. Record inputs, expected outcome, evidence, owner and observed result.

Download CSV worksheet
In this guide

Choose a specific lookup task

For a factory, a first pilot may answer quality-procedure questions; for a service company, locate work-order requirements. Limit included documents and decisions. An assistant reading documents should not be considered authorized to modify an order or approve an expense merely because it answers a question.

Microsoft distinguishes retrieval quality, grounding in context and answer completeness. Google Cloud describes evaluations with criteria tailored to individual prompts. These tools can support review; they do not replace business-owned expected answers or establish task success on their own.

Fictional example: current procedure and earlier version

S-03 version two permits preparing up to 20 units without supervisor review. Version three, effective in the scenario from October 1, limits that preparation to ten. The question is “Can I prepare 15 units without review?” The expected answer uses version three, says no and identifies the relevant section. These synthetic rules evaluate a lookup, rather than providing real production instructions.

An answer correctly citing version two still fails the current task. An answer saying no but linking a packaging section also fails to establish its basis. The evaluator must find the passage, verify validity and check that it supports the limit. Include similarly named documents so testing is not restricted to one easily retrieved file.

CriterionCheckFailure not to hide
Correct answerMatches the current document's criterionInvented number, condition or exception
Relevant sourcePassage supports this answerExisting but irrelevant citation
Version and scopeDocument authorized for the queryOld version or another context
Insufficient informationStates what is missing and refers to an ownerComplete answer without evidence

Include difficult and unanswerable questions

Prepare the worksheet before the demonstration. Mix frequent questions, exceptions, alternate names and questions the material cannot answer. The owner supplies expected answer, document, version and passage. Keep some questions outside initial preparation to check that results extend beyond a rehearsed demonstration.

CaseExpected resultEvidence
Prepare 15 units under S-03No; review required by version threeAnswer and current passage
Only version two is availableStates validity limitation rather than claiming current ruleAvailable files and answer
Undocumented exception questionAcknowledges missing information and refers onwardAnswer and expected criterion
Unauthorized other-team documentIts content is not used for this personTest user and sources
Source contradicts generated textFailed case despite citationAnswer/passage comparison
Repeat after updating S-03Uses new version or states update pendingIndex version and new test

Record failure types, not just an average score

Count executed questions in each category and the cases meeting their criteria. Ten fictional questions do not establish a business-wide success rate. Review errors that could lead to an incorrect decision separately. A high average can hide a critical exception answered with excessive confidence.

Keep material version, configuration and rehearsal date. When documents, model or search change, rerun affected questions and some controls. Updates may take time to become queryable; define how that state is recognized. Permission, format, language and provider dependencies limit what the result demonstrates.

Keep reading and acting separate

A document can describe an action without authorizing the assistant to execute it. Review AI agent permissions before connecting tools; use AI ROI to evaluate later benefit. Software acceptance helps record pilot decisions and unresolved issues.

The downloadable worksheet contains ten synthetic tests, rather than an approved evaluation. Bring authorized documents, versions and five difficult questions to a free consultation about AI for business. We can review a useful lookup pilot and evidence needed before widening its use.

Review assistant tests using your documents

Your authorized documents and ten questions help define which answers need a source and when the assistant should abstain. In a free consultation, also review users and decisions excluded from the pilot.

  • Authorized documents with versions and effective dates
  • Five frequent and five difficult questions with expected answers
  • Users, permissions and decisions excluded from the pilot
Book a free consultationAsk on WhatsApp

Related

Frequently asked questions

Not always. The document may be old, from another context or unable to support the text's claim. Open the passage and verify number, condition and validity. Source and answer are related but separate checks.

State the limitation and identify missing information or a responsible person. The expected answer should include that behavior. Do not reward completing every answer through assumptions unauthorized by the documents.

They can verify those ten cases, rather than the whole operation. Retain categories, outcomes and sample size. Add representative documents and exceptions before treating a percentage as a pilot indicator.

Automated evaluation can help prioritize, but its criteria and failures also need review. The content owner confirms business answers and important exceptions. Do not represent an automatic score as human approval.

Sources

  1. RAG evaluators for generative AIMicrosoft
  2. Gen AI evaluation service overviewGoogle Cloud

Last updated:

Keep reading

Free consultation

Where does AI actually pay off for you?

In a free consultation we review your tasks and tell you which ones are worth automating with AI and which aren’t. AI automations from ~$5,000 MXN. We reply the same business day.

  • Free, no commitment
  • Proposal in 1 business day
  • Delivered in stages