A community health worker in rural Uganda opens a mobile application to check whether a child's symptoms match malaria or a common cold. The interface is in English. The child's caregiver speaks only Luganda. The worker translates the symptoms, receives a response, translates it back and explains it to the caregiver. At no point in this chain can anyone verify whether the underlying model understood the clinical context, the local treatment protocols or the caregiver's description of the child's condition.

This is not a hypothetical scenario. It is the daily operating reality for thousands of health workers, agricultural extension officers, financial advisors and public-service staff across Africa who are being asked to deploy AI systems that were trained, evaluated and validated primarily in English.

The operational consequence is straightforward: an AI system that cannot be evaluated in the language of its users cannot be trusted as a decision-support tool. It becomes a black box whose outputs must be accepted on faith rather than verified through evidence. For organisations accountable for service delivery, regulatory compliance or clinical outcomes, that is an unacceptable risk.

The evaluation gap is not just a translation problem

Organisations often assume that if a model performs well in English, it will perform adequately in other languages once translated. This assumption is wrong. Multilingual evaluation research shows that model performance varies significantly across languages, even within the same model family. The AfriMMLU benchmark, which evaluates large language models across 16 African languages and five subject areas, found that token fertility—the number of tokens required to express the same concept—reliably predicts accuracy. Languages with higher token fertility (requiring more tokens to express the same idea) show lower accuracy, not because the model is less capable, but because the training data and evaluation methods were not designed for those linguistic structures.

This is not a marginal effect. When a reasoning model such as DeepSeek or o1 is evaluated on AfriMMLU, it consistently outperforms non-reasoning peers across both high-resource and low-resource languages, narrowing the accuracy gaps observed in prior generations. That is a meaningful improvement, but it does not eliminate the gap. The fundamental issue is that evaluation benchmarks themselves are often English-centric, with African languages added as an afterthought rather than designed as native evaluation contexts.

The SALT-31 benchmark, which covers 31 Ugandan languages, demonstrates the scale of the challenge. It evaluates both sentence-level and paragraph-level machine translation across nearly every language spoken in a country with high linguistic diversity. The benchmark reveals that even state-of-the-art multilingual models struggle with low-resource languages, particularly when the task requires understanding local idioms, cultural context or domain-specific terminology.

Why this matters for operational deployment

An organisation deploying an AI system for customer service, clinical decision support, agricultural advisory or financial inclusion faces a practical question: can we audit the system's performance in the languages our users actually speak? If the answer is no, the organisation cannot demonstrate that the system is safe, accurate or appropriate for its intended use.

This is not merely a technical concern. It is a governance and accountability issue. A health system that deploys an English-optimised diagnostic tool in a multilingual context cannot explain to a patient why a recommendation was made, cannot verify whether the tool understood the patient's description of symptoms and cannot audit whether the tool's performance is consistent across language groups. A financial institution that uses an English-trained credit-scoring model for customers who speak Swahili, Amharic or Yoruba cannot demonstrate that the model's decisions are fair, transparent or compliant with local consumer-protection regulations.

The Africa Data Protection Authority's analysis of AI governance across the continent highlights this gap. National AI strategies in Kenya, Ghana, Uganda, Ethiopia and other countries emphasise the need for AI systems that respect local languages, cultural contexts and regulatory requirements. But the gap between policy ambition and operational reality remains wide. Most commercial AI systems are not evaluated in African languages, and most organisations deploying these systems do not have the tools to conduct native-language evaluation.

What is changing

The research landscape is shifting. AfricaNLP 2026, a major conference on natural language processing for African languages, showcased a range of new benchmarks, models and evaluation frameworks. AfriCaption provides a framework for multilingual image captioning in 20 African languages. AfriNLLB offers efficient translation models covering 15 language pairs and 30 translation directions, including Swahili, Hausa, Yoruba, Amharic, Somali, Zulu and other African Union official languages. IrokoBench provides a benchmark for African languages in the age of large language models.

These are not academic exercises. They are the building blocks of operational evaluation. An organisation that wants to deploy a multilingual AI system can now point to specific benchmarks, evaluate model performance across languages and make evidence-based decisions about which system to use, which languages it supports and where the gaps remain.

The Microsoft Research survey on multilingual evaluation found that of 23 recent model releases (15 open-weight, 8 closed-model, 2024–2026), only a minority disclosed training language composition, reported multilingual evaluation benchmarks or named contamination detection techniques. That is improving, but the transparency gap means that organisations must often conduct their own evaluation rather than relying on vendor claims.

A practical approach for African organisations

Organisations deploying AI systems in multilingual contexts should treat language evaluation as a non-negotiable component of operational readiness. The approach is straightforward:

Discover. Identify the languages your users actually speak, not the languages your interface supports. Map the decision points where a misunderstanding or mistranslation could have serious consequences: clinical recommendations, financial decisions, legal advice, safety-critical instructions.

Design. Select or develop evaluation benchmarks that cover your operational languages and domains. If no suitable benchmark exists, collaborate with local universities, language communities or research organisations to create one. Define the minimum accuracy, fairness and transparency thresholds your organisation requires.

Build and test. Evaluate candidate AI systems against your benchmarks. Test not only translation accuracy but also contextual understanding, cultural appropriateness and domain-specific knowledge. Involve native speakers in the evaluation process, not just translators.

Enable and monitor. Deploy the system with clear documentation of its language capabilities and limitations. Monitor performance across language groups. Establish a process for users to report errors, request corrections and escalate concerns.

Xelius supports organisations working on multilingual AI deployment through evaluation design, benchmark development, system testing and implementation support. The starting point is not a technology purchase; it is the question of whether your organisation can explain and audit the AI system's decisions in the languages your users speak.

The operational standard is comprehension, not connectivity

An AI system that can be accessed in multiple languages but cannot be evaluated in those languages is not a multilingual system. It is a monolingual system with a translation layer. The distinction matters for accountability, safety and trust.

African organisations have an opportunity to set a higher standard. Rather than accepting AI systems that work "well enough" in English and hoping the translation holds, they can demand systems that are evaluated, audited and explainable in the languages of their users. That is not a luxury; it is the minimum operational standard for any system that influences decisions about health, finance, education, justice or public service.

The research tools are emerging. The benchmarks are being built. The question is whether organisations will use them.