What we tested

On 17 July 2026 we ran GPT-5.2 (OpenAI), Gemini 2.5 Flash (Google), Claude Sonnet 5 (Anthropic), and DeepSeek V3.2 through the API on default settings. Each model received the same ten public HausaBench prompts covering translation in both directions, instructions written in Hausa, localization judgment for Nigerian users, code-mixed Hausa-English, a medical safety scenario, and long-form Hausa writing.

Every output was scored 1 to 5 on six categories: meaning accuracy, fluency, tone, cultural fit, hallucination risk, and safety. Every defect got an issue label and a severity from our public issue taxonomy. All quotes below are verbatim from real model outputs.

The scores

ModelMeaningFluencyToneCultural fitHallucinationSafetyOverallVerdict
GPT-5.24.84.54.94.95.04.94.83Good
Claude Sonnet 54.53.74.54.85.04.94.57Good
DeepSeek V3.24.33.94.84.54.45.04.48Mixed
Gemini 2.5 Flash4.23.84.14.34.85.04.37Mixed

Hallucination and safety columns: 5.0 means no risk observed in this run. Scores are expert reviewer guidance for real-world Hausa use, not an academic leaderboard.

Five findings that matter

1. Two of four models failed "Dan Allah", the everyday Hausa word for please

A test message opened with "Dan Allah help me check ko payment din ya shiga...". Gemini 2.5 Flash translated it as "And Allah help me check if the payment has gone through" and then invented a cultural explanation that the sender was "invoking Allah for assistance or blessing". DeepSeek V3.2 went further and fabricated a person: "I'm asking Dan to please help me check if the payment has gone through" There is no Dan. "Dan Allah" simply means please. GPT-5.2 and Claude Sonnet 5 got it right.

2. Fluent outputs contained words that do not exist

One model translated "Your account has been updated" as "An akulla asusun ku cikin nasara." "Akulla" is not a Hausa word (the correct verb is "sabunta"). The same model later told residents to bring "fatalwa" (a ghost) to a community clean-up, instead of "fartanya" (a hoe). Another model left English "rake" inside a Hausa announcement, where "rake" means sugarcane. A reviewer who does not speak Hausa would pass all of these: they look perfectly fluent.

3. One model invented facts out of nothing

Asked for a community announcement with no date specified, DeepSeek V3.2 confidently published one: "Kwanan wata: Lahadi, 10 ga Satumba, 2023." An invented day, month, and stale year. In a real product this is the kind of detail users act on.

4. Quality collapses as Hausa output gets longer

Short translations were often clean. Long-form Hausa was where hooked letters (ɓ ɗ ƙ ƴ) started dropping, grammar loosened, and invented phrases appeared, across all four models. Task discipline also slipped: one model answered a two-sentence translation request with a 2,900-character English essay of options, and another broke an explicit 40-word limit. For training data and UX copy, these are structural failures, not typos.

5. The good news: safety behavior has improved

On a three-day-fever prompt, all four models refused to diagnose and pushed the user toward professional care, in Hausa. Localization judgment is also improving: all four correctly dropped a "credit card failed" framing for Nigerian users who pay by transfer, debit card, and mobile money. But even here, one model rendered "antibiotics" as "maganin rigakafi", which most Hausa speakers read as vaccine or preventive medicine, turning stewardship advice into something that can read as anti-vaccine guidance.

What this means if you ship Hausa

  • Chatbots and assistants: current models are usable drafting engines and unreliable publishers. Every Hausa-facing flow needs expert review before release.
  • Localization teams: fluency is not the signal. The dangerous outputs in this run were the fluent ones. QA must check meaning, idiom, and lexical reality, not just grammar.
  • Data teams: constraint violations (English inside Hausa-only outputs, format drift, invented specifics) poison training and evaluation sets quietly. Validate before you train.

Method notes

Prompts are self-created public HausaBench items; no client content is ever used. Models were accessed via API on default settings with no system prompt, one prompt per request, on 17 July 2026. Scoring used the Hausa AI Studio 6-category rubric with severity-labeled issue logging, agent-assisted first pass with expert aggregation and spot verification against raw outputs. Raw outputs are preserved and reproducible. HausaBench is practical human QA for real-world Hausa use, small by design; it is not a full academic evaluation.

This is exactly what we do for clients

Hausa AI Studio runs this same review method on your model outputs, localization, transcripts, and datasets: severity-rated findings, corrections, and a clear readiness decision.

Start a pilot intake View a sample audit

Or email abba@hausaai.studio