Consider an illustrative case: a customer-service team tests an AI assistant on twenty common questions. It drafts clear replies, recognises basic requests and reduces the time needed to prepare a response. The pilot is declared successful. In the first week of live use, a customer asks about an account affected by a recent policy change; the assistant retrieves an outdated document, gives a plausible answer and the agent sends it without noticing. The error does not come from a dramatic model failure. It comes from a deployment decision made without testing the conditions under which the service actually operates.

A demonstration answers whether a model can do something. It does not answer whether an organisation should rely on it in a particular workflow. AI evaluation needs to function as a release gate: a documented decision that the system has met context-specific evidence thresholds, has clear limits and can be monitored, paused and improved after it goes live. That is more demanding than a pilot scorecard, but it is also more useful. It moves attention from impressive examples to the quality of a real operational decision.

For African institutions, the gap between the two can be wide. Systems may draw on fragmented records, operate across multilingual teams, serve customers on mobile and assisted channels, and encounter changing regulations or uneven connectivity. A model selected on generic benchmarks, or tested only by a central innovation team, may perform well in a controlled setting while creating quiet failure modes for the people required to deliver and supervise the service.

A pilot is evidence, not a verdict

Pilots are valuable. They reveal whether staff find a tool usable and whether a data connection is feasible. Their limitation is that they are often designed to demonstrate success: clean examples, stable systems and a narrow prompt set. Production is less accommodating.

NIST’s AI Risk Management Framework: Generative AI Profile, published in July 2024, is a voluntary companion resource to the AI RMF. It addresses trustworthiness across the AI lifecycle; it neither prescribes one approval process nor acts as a compliance certificate.

The deployment claim should be specific. “The assistant answers questions well” is not an operational claim. “The assistant can draft a response to standard delivery-status enquiries, using an approved current source, while a trained agent remains responsible for sending it” is one. The latter can be tested. It also tells an organisation what it must not assume: that the system can resolve an exception, interpret a policy change or act without review.

Test the decision, not the model’s personality

A model can sound helpful while failing the task that matters. Evaluation should begin with the decision or action the organisation is considering: prioritising a case, classifying a document, drafting a response, extracting a field, flagging a risk or proposing the next step. The test set should contain the records, ambiguity and consequences that define that decision.

The NIST AI RMF Playbook offers voluntary suggested actions across the framework’s Govern, Map, Measure and Manage functions. Its Measure guidance includes connecting measurement to the deployment context and using input from domain experts and relevant users. That is a useful corrective to evaluating only through aggregate model scores. A customer-service leader, clinical reviewer, procurement officer or field supervisor may recognise an unsafe answer that a generic text-similarity metric will miss.

Good cases include the inconvenient ones. Test standard requests, but also incomplete records, conflicting sources, misspellings, code-switching, policy changes, requests outside the authorised scope, attempts to elicit sensitive information and situations in which the correct response is to escalate. If a system will work with a knowledge base, include documents that are obsolete, duplicated or difficult to interpret. If it will call a tool, test unavailable services and mismatched identifiers.

The aim is not to catch a model out. It is to determine the boundary at which the system can provide useful support and the boundary at which a person or another workflow must take over. A test suite that contains only questions with known easy answers does not locate that boundary. Keep a held-out set separate from examples used to tune prompts or retrieval; otherwise a rising score may reflect familiarity with the test rather than dependable performance.

Context has to include the last mile

Local deployment conditions change the evaluation problem. A service may be delivered by staff using shared devices, through a call centre with brief interaction windows, in a mixture of English and local languages, or alongside records kept by partner organisations. A field team may have to resume work after a connectivity interruption. An enterprise may rely on a vendor platform whose documents, permissions and model version can change without warning.

An evaluation designed exclusively from headquarters can overlook these conditions. It may confirm that an English prompt produces a sensible answer while failing to test whether a user can understand it, whether a supervisor can correct it, or whether the response remains safe when an approved source is unavailable. This is not a claim that every model must be built locally. It is a requirement that the institution evaluate the model in the context for which it will be accountable.

Language is part of system behaviour. Where users and staff work in more than one language, testing needs to include the actual language patterns and domain terms they use, not only a translated interface. The relevant measure may be whether a customer understands an eligibility explanation, whether a health worker can identify uncertainty, or whether a field officer can detect a mistranslated instruction. The model’s apparent fluency in English is not evidence of those outcomes.

A release gate needs more than an accuracy number

No single metric decides whether an AI system is ready. Accuracy matters for extraction and classification tasks, but it may be insufficient for an assistant whose serious failures concern privacy, unsupported advice, unequal treatment or a failure to escalate. The appropriate measures follow the task and risk. Record task success, retrieval failures, response latency, cost per completed task and reviewer effort; examine results by the languages and channels used rather than hiding weak performance inside one average.

OpenAI’s evaluation guidance recommends defining the task, building test data and analysing results rather than relying on intuition. It is product documentation, not an independent assurance standard, but it captures an operationally sound practice: compare candidate systems and changes against representative cases with criteria that match the intended outcome.

For a release gate, leaders need a concise evidence record rather than a technical appendix nobody will use. It should state the intended use and prohibited uses; the authoritative data or knowledge sources; the test cases and who reviewed them; the pass criteria; the errors observed; the required human oversight; the monitoring signals; and the pause or rollback route. This turns a claim of “pilot success” into an accountable decision that can be revisited when the system, policy or workflow changes.

Thresholds should be consequential. A document-extraction tool may be permitted to pre-fill fields only after its scores have been calibrated against reviewed examples, with uncertain cases routed to a person. A model’s own statement of confidence is not sufficient evidence of reliability. A public-information assistant may be allowed to draft only from approved current materials and must provide an assisted channel for unresolved cases. A system that proposes credit, eligibility or enforcement action may require a much stronger level of validation and a meaningful human decision. The threshold is not a universal percentage; it is the level of evidence appropriate to the harm of getting the task wrong.

Monitoring begins where the pilot ends

A completed test run does not freeze performance. Knowledge sources change, a supplier updates a model, a workflow is redesigned, seasonal demand shifts and users learn new ways of asking for help. A model can therefore regress even when no one has deliberately changed the application.

Production feedback must become evidence for the next evaluation cycle. The NIST Playbook’s Measure guidance recommends documenting pre- and post-deployment performance and regularly assessing whether metrics and controls remain effective. These are voluntary suggestions, not regulatory requirements.

Monitor the failures that matter. Track unsupported answers, overrides, escalations, source-retrieval failures, complaint outcomes, time spent correcting outputs and performance across relevant channels or language groups. A higher completion rate can conceal a worsening error pattern if staff have learned to repair the system quietly. Conversely, a rise in escalations may signal that staff understand the system’s limits and are using the control route correctly.

Make evaluation a shared operational practice

Evaluation is sometimes handed to a small technical team and presented to leadership as a benchmark chart. That arrangement cannot capture the full question. Domain specialists understand the decision; operations teams understand service failure; security and privacy teams understand control requirements; front-line staff understand the last mile; and leadership must determine the risk the organisation is willing to accept.

Discover the decision boundary. Identify the task, user, affected party, source records, authorised action and credible failure modes. Write the point at which the system must stop or ask for human intervention before testing model options.

Design representative evidence. Build and maintain a small, governed set of real or safely anonymised cases. Include ordinary work, known exceptions and cases that should be refused or escalated. Ask domain reviewers to assess whether the outputs are accurate, understandable and appropriate for the service.

Release with controls. Set the pass criteria, named approver, user guidance, audit evidence and rollback mechanism. Limit authority at launch; expand it only when production evidence shows that the system is dependable for the next level of action.

Xelius supports organisations through applied research, AI evaluation, workflow design, data architecture and platform implementation. A useful starting point is the AI use case already being called a successful pilot, then defining the real decision it will influence and the evidence required before it enters a live workflow.

The mature question is whether the system can be stopped

Operational value will not be decided by who can produce the most convincing demonstration. It will be decided by who can say, with evidence, what a system is authorised to do, where it fails, who can intervene and when to reassess it.

A release gate does not slow responsible adoption for its own sake. It creates the conditions in which an institution can deploy useful AI without converting uncertainty into an untraceable service decision. That is how a pilot becomes a capability that can be governed, improved and trusted.

Tags