[Blog](<https://nightlysoftware.com/en/blog>)Document AI evaluation 

# An AI assistant for your documents: test questions and limits

A convincing answer may use an outdated document. Evaluate its claim, source and response when information is missing.

**[Jonathan Perez](<https://nightlysoftware.com/en/company#jonathan-perez>)**Co-founder · Design, product and sales October 8, 2026 · 7 min read 

**Short answer**

Evaluate an assistant with representative questions and expected answers verified by the document owner. Check accuracy, source, version and behavior when information is insufficient. A citation and fluent answer are not enough: the passage must support the decision someone will make.

## Test worksheet: Document AI assistants: evaluating answers

A convincing answer may use an outdated document. Evaluate its claim, source and response when information is missing. Record inputs, expected outcome, evidence, owner and observed result.

[Download CSV worksheet](<https://nightlysoftware.com/plantillas/asistente-ia-documentos-evaluacion-en.csv>)

In this guide

-   [Choose a specific lookup task](<https://nightlysoftware.com/en/blog/document-ai-assistant-evaluation#scope>)
-   [Fictional example: current procedure and earlier version](<https://nightlysoftware.com/en/blog/document-ai-assistant-evaluation#example>)
-   [Include difficult and unanswerable questions](<https://nightlysoftware.com/en/blog/document-ai-assistant-evaluation#questions>)
-   [Record failure types, not just an average score](<https://nightlysoftware.com/en/blog/document-ai-assistant-evaluation#measurement>)
-   [Keep reading and acting separate](<https://nightlysoftware.com/en/blog/document-ai-assistant-evaluation#limits>)

## Choose a specific lookup task

For a factory, a first pilot may answer quality-procedure questions; for a service company, locate work-order requirements. Limit included documents and decisions. An assistant reading documents should not be considered authorized to modify an order or approve an expense merely because it answers a question.

[Microsoft](<https://learn.microsoft.com/en-us/azure/foundry/concepts/evaluation-evaluators/rag-evaluators>) distinguishes retrieval quality, grounding in context and answer completeness. [Google Cloud](<https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/evaluation-overview>) describes evaluations with criteria tailored to individual prompts. These tools can support review; they do not replace business-owned expected answers or establish task success on their own.

## Fictional example: current procedure and earlier version

S-03 version two permits preparing up to 20 units without supervisor review. Version three, effective in the scenario from October 1, limits that preparation to ten. The question is “Can I prepare 15 units without review?” The expected answer uses version three, says no and identifies the relevant section. These synthetic rules evaluate a lookup, rather than providing real production instructions.

An answer correctly citing version two still fails the current task. An answer saying no but linking a packaging section also fails to establish its basis. The evaluator must find the passage, verify validity and check that it supports the limit. Include similarly named documents so testing is not restricted to one easily retrieved file.

| Criterion |Check |Failure not to hide |
| --- | --- | --- |
| Correct answer |Matches the current document's criterion |Invented number, condition or exception |
| Relevant source |Passage supports this answer |Existing but irrelevant citation |
| Version and scope |Document authorized for the query |Old version or another context |
| Insufficient information |States what is missing and refers to an owner |Complete answer without evidence |

## Include difficult and unanswerable questions

Prepare the worksheet before the demonstration. Mix frequent questions, exceptions, alternate names and questions the material cannot answer. The owner supplies expected answer, document, version and passage. Keep some questions outside initial preparation to check that results extend beyond a rehearsed demonstration.

| Case |Expected result |Evidence |
| --- | --- | --- |
| Prepare 15 units under S-03 |No; review required by version three |Answer and current passage |
| Only version two is available |States validity limitation rather than claiming current rule |Available files and answer |
| Undocumented exception question |Acknowledges missing information and refers onward |Answer and expected criterion |
| Unauthorized other-team document |Its content is not used for this person |Test user and sources |
| Source contradicts generated text |Failed case despite citation |Answer/passage comparison |
| Repeat after updating S-03 |Uses new version or states update pending |Index version and new test |

## Record failure types, not just an average score

Count executed questions in each category and the cases meeting their criteria. Ten fictional questions do not establish a business-wide success rate. Review errors that could lead to an incorrect decision separately. A high average can hide a critical exception answered with excessive confidence.

Keep material version, configuration and rehearsal date. When documents, model or search change, rerun affected questions and some controls. Updates may take time to become queryable; define how that state is recognized. Permission, format, language and provider dependencies limit what the result demonstrates.

## Keep reading and acting separate

A document can describe an action without authorizing the assistant to execute it. Review [AI agent permissions](<https://nightlysoftware.com/en/blog/ai-agent-permissions>) before connecting tools; use [AI ROI](<https://nightlysoftware.com/en/blog/ai-roi>) to evaluate later benefit. [Software acceptance](<https://nightlysoftware.com/en/blog/software-business-acceptance-tests>) helps record pilot decisions and unresolved issues.

The downloadable worksheet contains ten synthetic tests, rather than an approved evaluation. Bring authorized documents, versions and five difficult questions to a [free consultation](<https://nightlysoftware.com/en/book>) about [AI for business](<https://nightlysoftware.com/en/solutions/ai-for-business>). We can review a useful lookup pilot and evidence needed before widening its use.

## Review assistant tests using your documents

Your authorized documents and ten questions help define which answers need a source and when the assistant should abstain. In a free consultation, also review users and decisions excluded from the pilot.

-   Authorized documents with versions and effective dates
-   Five frequent and five difficult questions with expected answers
-   Users, permissions and decisions excluded from the pilot

[Book a free consultation](<https://nightlysoftware.com/en/book>)[Ask on WhatsApp](<https://wa.me/524622212236?text=I%20want%20to%20evaluate%20an%20AI%20assistant%20with%20authorized%20documents.%20I%20have%20current%20versions%2C%20frequent%20and%20difficult%20questions%20with%20expected%20answers%2C%20and%20the%20users%20and%20decisions%20excluded%20from%20the%20pilot.>)

Related

-   [AI for business](<https://nightlysoftware.com/en/solutions/ai-for-business>)
-   [Process automation](<https://nightlysoftware.com/en/solutions/business-process-automation>)

## Frequently asked questions

### Does a citation prove an answer is correct? 

Not always. The document may be old, from another context or unable to support the text's claim. Open the passage and verify number, condition and validity. Source and answer are related but separate checks.

### What should happen when no answer is found? 

State the limitation and identify missing information or a responsible person. The expected answer should include that behavior. Do not reward completing every answer through assumptions unauthorized by the documents.

### Can ten questions measure accuracy? 

They can verify those ten cases, rather than the whole operation. Retain categories, outcomes and sample size. Add representative documents and exceptions before treating a percentage as a pilot indicator.

### Do we need another agent to review answers? 

Automated evaluation can help prioritize, but its criteria and failures also need review. The content owner confirms business answers and important exceptions. Do not represent an automatic score as human approval.

## Sources

1.  [RAG evaluators for generative AI](<https://learn.microsoft.com/en-us/azure/foundry/concepts/evaluation-evaluators/rag-evaluators>)Microsoft 
2.  [Gen AI evaluation service overview](<https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/evaluation-overview>)Google Cloud 

Last updated: October 8, 2026

## Keep reading

[AIOct 2, 2026

### What permissions should an AI agent get in your business?](<https://nightlysoftware.com/en/blog/ai-agent-permissions>)[AIOct 2, 2026

### AI ROI: how to tell if AI is actually making your business money](<https://nightlysoftware.com/en/blog/ai-roi>)[AI invoice reviewOct 8, 2026

### Reading invoices with AI: fields, differences and review before posting](<https://nightlysoftware.com/en/blog/invoice-ai-human-review>)

---

Canonical: https://nightlysoftware.com/en/blog/document-ai-assistant-evaluation

Updated: 2026-10-08

Description: Test answers, citations, versions and unanswered questions before using an AI assistant. Includes a document-answer evaluation worksheet.

