Originally drafted 31 July 2026. First published 7 October 2026; revised for publication.

Consider an illustrative ministry seeking to reduce the time required to answer public enquiries. A supplier demonstrates an AI assistant that can search policy documents, draft replies and route complex cases. The demonstration is persuasive. The procurement brief asks for accuracy, security and a delivery date. It says little about which answers a frontline worker may rely on, what records the assistant may access, how a citizen can challenge a wrong response, or what happens when the supplier updates the model after launch.

Those omissions are not minor contract details. They decide whether the ministry has bought a useful support tool or quietly transferred part of a public decision to a system it cannot explain, supervise or correct.

The African Union’s Continental Artificial Intelligence Strategy, published in July 2024, provides an Africa-centred direction for responsible AI development and use. Turning that direction into public value will depend on practical choices in service design, procurement and operation.

An AI purchase must secure a bounded capability for a defined public service, not an opaque promise that a vendor will make government smarter. That distinction is particularly important where records are fragmented, specialist capacity is scarce and the public service must still work across shared devices, multiple languages and uneven connectivity.

The procurement brief is already a governance decision

Governance is often discussed after a tool has been selected: as a policy, ethics review or training programme added during implementation. Yet the most consequential choices are commonly made earlier. Procurement determines what problem is being solved, the data a supplier can access, the evidence a bidder must provide, who carries operational responsibility and which changes require approval.

A vague requirement such as “an AI platform for citizen services” invites a vague response. It can bundle together retrieval, content generation, profiling, workflow routing and decision support without distinguishing their risks. It can also make later accountability difficult because no one has specified the boundary between a staff member’s judgement and the system’s output.

A better brief begins with the service consequence. Is the objective to help staff find current guidance, triage incomplete applications, translate routine information, or identify cases requiring human follow-up? Each needs different records, measures, safeguards and escalation paths.

Clearer scope can reduce the risk of buying the wrong capability. The OECD’s 2025 review of AI in government identifies access to high-quality data as a common challenge, and uncertain costs and outdated legacy systems as further implementation challenges. A model licence cannot solve them alone.

Buy a service outcome, not a model claim

A supplier may advertise a foundation model, an accuracy figure or a polished conversational interface. None tells a public buyer whether the system improves an actual service under local conditions.

Accuracy has to be defined against the job. A tool that correctly retrieves a document may still produce an unsuitable response if the document is outdated, the question is ambiguous or the user needs an assisted channel. A triage system may perform well in a test set yet create harm if it deprioritises cases with incomplete records, a condition that can be common in services involving informal workers, rural residents or people whose identity documentation is inconsistent.

Public buyers should ask bidders to demonstrate performance on representative, lawfully obtained and appropriately governed material. They should specify what the tool must not do, as well as what it must do. If the tool may draft but not decide, the interface, workflow and performance measures must preserve that line.

Useful evaluation evidence includes how the system handles uncertainty, unsupported questions and conflicting documents; which languages and channels were actually tested; how a frontline worker corrects an output; and whether the supplier can show the provenance of retrieved material. A generic benchmark score is not a substitute for evidence from the service environment.

This is also a budget discipline. A cheaper licence can become expensive when an agency must retrofit integration, data cleaning, human review, training and incident handling. Procurement should price the operating model, not only the tool.

Suppliers must provide evidence, not just assurances

Public institutions cannot responsibly accept “responsible AI” as a marketing claim. They need information that allows them to assess a system’s capabilities, limitations and dependencies before deployment, and to re-evaluate it when material changes occur.

The UK’s Algorithmic Transparency Recording Standard guidance offers a useful reference point, not an African legal template. Mandatory organisational scope covers UK government departments and certain arm’s-length bodies providing public/frontline services or direct public interaction. Tool scope and exemptions also apply; publication is not mandatory for every public body or every algorithm. The mandatory scope and exemptions policy should be read alongside the guidance. Its underlying lesson travels well: transparency is strongest when it documents purpose, ownership, data, human role and the broader process around a tool.

A public buyer can require an AI supplier to disclose, at a proportionate level:

  • the intended purpose, known limitations and prohibited uses;
  • the models, external services and material subcontractors on which the service depends;
  • how agency data is stored, isolated, retained, deleted and used for training or improvement;
  • testing methods, performance evidence and conditions under which performance may degrade;
  • logging, monitoring and incident-notification arrangements;
  • the process for material model, dataset, feature or hosting changes; and
  • exit, export and transition provisions if the service is discontinued or the supplier changes.

These are not demands to expose source code or security-sensitive details. They are the evidence needed to manage a public service responsibly. Where intellectual property limits disclosure, a supplier should still explain the operational boundary, governance controls and assurance route sufficiently for a defensible decision.

Human oversight has to be designed into the workflow

“Human in the loop” can become an empty phrase when staff are expected to review large volumes quickly, lack authority to override the tool, or cannot see why an output was generated. A person clicking approve does not create meaningful oversight.

A stronger approach defines the role before the system is procured. Assistive use supports drafting, search or translation while the staff member remains responsible for the response. Advisory use supplies a recommendation that a qualified official must independently assess. Decision-influencing use affects an entitlement, enforcement or prioritisation process and demands greater safeguards, transparency and review. The final category should not be hidden inside a generic software purchase.

The design must include intervention. Staff need a clear way to correct an output, escalate a concern and report a recurring failure. Citizens and businesses need a route to understand a consequential decision and seek review. Managers need evidence of overrides, errors, complaints and service impact, not merely a dashboard showing conversations completed.

This is where local operating realities matter. If a tool is introduced into a service delivered partly through call centres, community agents or regional offices, oversight cannot exist only in a central digital team. Training, permissions and feedback routes must reach the people who actually meet service users. If connectivity is intermittent, the process needs an alternative rather than treating the absence of a live model response as staff failure.

Data access should be the narrowest useful access

An AI system’s data architecture is a service decision. Giving a supplier broad access to historical records may seem convenient, but it can create risks that are unrelated to the immediate use case: sensitive data exposure, unclear retention, secondary use and difficulty deleting information later.

Start with the question the system must answer. A knowledge assistant for internal staff may only need curated, approved documents. A casework aid may require specific structured fields and current policy rules, with strict access controls. A public-facing information tool may need no personal data at all.

The principle is simple: use the minimum data necessary, for the minimum time necessary, under controls that the institution can audit. It is compatible with innovation and more likely to produce a system that can be trusted. It also reduces a common form of vendor lock-in: the gradual accumulation of public data in a proprietary environment that is difficult to inspect or leave.

Data quality deserves equal attention. AI can make inconsistent records appear coherent, which is useful until a polished answer masks the fact that the underlying record is wrong. Public organisations need ownership of source-data correction, clear provenance for knowledge content and a process for withdrawing outdated guidance. A model cannot compensate for an institution that has not decided which record is authoritative.

Pilot contracts should earn the right to scale

Pilots are valuable when they test a clear hypothesis. They become wasteful when they create a demonstration without evidence about operating cost, user impact, accuracy under real conditions or accountability at scale.

A responsible pilot contract should set a bounded purpose, duration, user group and data scope. It should name the service owner, the supplier obligations and the independent assurance or review required. It should define measures that matter to the service: resolution time, error rate, escalation rate, staff workload, user comprehension, exclusion and complaints. It should also state the conditions under which the tool will be paused, changed or not scaled.

The OECD’s finding that many government AI efforts remain early-stage is a reason to learn carefully. Governments can move quickly where services are well defined and risks are low, while reserving stronger controls for uses that influence rights, access or enforcement.

Build capability alongside the contract. A supplier can deliver software; it cannot substitute for public ownership. Procurement teams, legal advisers, data stewards, managers and frontline staff need enough shared understanding to set requirements, challenge evidence and manage change.

The public buyer remains responsible for the public service

AI will enter public services through many routes: direct procurement, cloud platforms, system integrators and tools already used by staff. The procurement moment is still the best opportunity to make the public interest explicit before technical and commercial dependencies harden.

The goal is not to make every agency build its own model or reject external suppliers. It is to retain control of the service purpose, the data boundary, the human decision, the audit trail and the exit route. That is how an institution can benefit from specialist technology without outsourcing its responsibilities to it.

These are operating-model recommendations, not a statement of identical legal duties across jurisdictions. Contract terms and service safeguards need review against applicable procurement, data-protection and sector law.

Government leaders planning AI-enabled services can begin with one high-volume, bounded workflow where staff already understand the failure modes. Xelius supports public institutions through service and workflow research, data architecture, responsible AI design and implementation support.

The most credible public-sector AI procurement makes a real service more dependable while leaving responsibility visible, contestable and public.