wesolutionsai
WHITE PAPEROCT 2026REF · WES-RES-OCT2026-05

Arabic-first evaluation for production AI systems.

Why English-first evaluation fails in the Gulf, what the published Arabic benchmarks actually measure, and how to design an evaluation harness that survives a regulator's model-governance review.

EXTENT
3 exhibits · ~2,300 words
READ
10 min
CLASSIFICATION
Public
AUTHORSHIP
wesolutions Research
EXECUTIVE SUMMARY

Gulf institutions are deploying large language models into Arabic-speaking operations on the strength of evaluations run mostly in English, against public benchmarks, on Modern Standard Arabic at best. The published evidence says that is not a safe extrapolation: performance drops measurably when the same task moves from English to Arabic and again when it moves from Modern Standard Arabic to Egyptian or Levantine dialect, and public Arabic benchmarks cover knowledge far better than they cover safety, dialect or the institution's own tasks. At the same time, the governance machinery is already in place to demand better: the Central Bank of the UAE's Model Management Standards make independent validation mandatory for all banks operating in the UAE, and Saudi Arabia's SDAIA has published generative-AI guidelines for government and the public. This paper maps the benchmark landscape as it stands, shows the dialect gap with published numbers, and sets out seven design principles for an Arabic-first evaluation harness that a model-risk function can sign.

  1. 01

    The dialect gap is measured, not hypothetical. On the AraDiCE dialect versions of the PIQA commonsense benchmark, Fanar-1-9B-Instruct scores 67.7 in Modern Standard Arabic, 63.7 in Egyptian and 59.0 in Levantine — an 8.7-point drop from MSA to Levantine on the same task; ALLaM-7B-Instruct-preview shows the same pattern (67.5 / 63.4 / 60.9) 45. A system evaluated only in MSA is unevaluated for a large share of what Gulf users will actually type.

  2. 02

    Public Arabic benchmarks measure knowledge far better than operations. The Open Arabic LLM Leaderboard v2 aggregates seven task families — AlGhafa, ArabicMMLU, EXAMS, a human-translated MMLU, AraTrust, MadinahQA and ALRAGE — almost all multiple-choice 2. None of them measures an institution's own workflows, document formats or refusal behaviour, which is where production risk lives.

  3. 03

    The strongest native benchmark is narrow by design: ArabicMMLU is 14,575 multiple-choice questions in MSA across 40 tasks, sourced from school exams in eight Arab countries 3. It is the right anchor for general capability and the wrong instrument for safety, dialect or domain evaluation — and its authors never claimed otherwise.

  4. 04

    Regulators already require or expect what Arabic-first evaluation provides. The CBUAE's Model Management Standards and Guidance (final, 21 December 2022) make model governance and independent validation mandatory for all banks operating in the UAE 6, and the CBUAE has since issued guidance on responsible AI and machine-learning adoption that expects AI governance to align with those Standards 7. SDAIA's generative-AI guidelines (January 2024) set expectations for government use in Saudi Arabia 8. An Arabic system validated only on English benchmarks is difficult to defend in either framework.

  5. 05

    The fix is architectural, not heroic: a harness built on seven principles — dialect-by-task coverage, private task-grounded test sets, contamination control, Arabic-native safety testing, human adjudication by Arabic speakers, versioned scoring, and publishable methodology — can be implemented with methods and instruments that already exist in the open literature 2349.

01

The problem: evaluated in English, deployed in Arabic

The Gulf now has a serious stable of Arabic-capable models: Saudi Arabia's ALLaM (SDAIA, served through HUMAIN and Amazon Bedrock), the UAE's Jais family and its successor Jais 2 (G42's Inception with MBZUAI) and the Falcon line including Falcon-Arabic (TII), and Qatar's Fanar (Ministry of Communications and Information Technology with the Qatar Computing Research Institute) 1410. Procurement conversations, however, still lean on global English benchmarks, and even Arabic-aware buyers usually stop at a leaderboard average. That average hides the two gaps that matter in production: the gap between English and Arabic performance, and the gap between Modern Standard Arabic and the dialects in which citizens, customers and front-line staff actually write.

Both gaps are now measurable with public instruments. The AraDiCE benchmark suite (Qatar Computing Research Institute, 2024) produced carefully post-edited Egyptian and Levantine versions of standard benchmarks precisely so that the dialect effect could be isolated 5. The Fanar technical report, which evaluates several regional models on AraDiCE, shows the pattern cleanly: on the PIQA commonsense task, moving the same questions from MSA to Levantine costs Fanar-1-9B-Instruct 8.7 points and ALLaM-7B-Instruct-preview 6.6 points 4. Exhibit 2 plots the numbers. The direction is consistent across models; the size of the drop is model-specific — which is exactly why it must be measured per system, not assumed.

EXHIBIT 2
The dialect gap, measured: AraDiCE-PIQA accuracy by register (%)
Sources: Fanar technical report, AraDiCE-PIQA evaluations (QCRI, 2025) 4; AraDiCE parallel sets per Mousi et al., 2024 5.
Same questions, three registers. The gradient, not the absolute level, is the point: a system evaluated only in MSA carries an unmeasured penalty in dialect.
02

What the public benchmark landscape actually covers

Exhibit 1 maps the principal public instruments. Three observations follow from it. First, the centre of gravity is knowledge measured by multiple choice: ArabicMMLU's 14,575 questions across 40 tasks 3, the AlGhafa suite's classification and reading-comprehension tasks 9, school-exam sets such as EXAMS, and human-translated MMLU variants. Second, the aggregation layer exists and is credible: the Open Arabic LLM Leaderboard, launched in May 2024 by 2A2I, TII and Hugging Face and re-based in 2025 as OALL v2, runs a fixed suite on a common harness and publishes everything 2. Third, the coverage thins precisely where production risk concentrates: AraTrust is the suite's one safety-and-truthfulness instrument 2; dialect coverage rests largely on AraDiCE-style adaptations 5; and no public benchmark covers an institution's own document formats, terminology, workflows or refusal policy — by construction, since public benchmarks must be general.

EXHIBIT 1
The public Arabic evaluation landscape: principal instruments and what they measure
InstrumentOrigin and yearWhat it measuresFormLimits for production use
ArabicMMLUMBZUAI and collaborators; ACL Findings 2024 3General and academic knowledge in MSA: 14,575 questions, 40 tasks, school exams from eight Arab countriesMultiple choiceMSA only; knowledge, not tasks; public (contamination risk)
AlGhafaTechnology Innovation Institute; ArabicNLP 2023 9Reading comprehension, sentiment, question answering and related abilitiesMultiple choiceGeneral abilities; MSA-centric
Open Arabic LLM Leaderboard (OALL v1 → v2)2A2I, TII and Hugging Face; May 2024, re-based 2025 2Aggregated suite: AlGhafa, ArabicMMLU, EXAMS, human-translated MMLU, AraTrust, MadinahQA, ALRAGELeaderboard on a common harnessAverages hide task and dialect variance; public sets
AraTrustAcademic consortium; in OALL v2 2Truthfulness and safety across eight sub-tasksMultiple choiceThe suite's one safety instrument; not a policy-specific red team
AraDiCEQatar Computing Research Institute, 2024 5Dialect (Egyptian, Levantine) and cultural capability via post-edited parallel versions of standard benchmarksParallel benchmark setsTwo dialects so far; method generalises, coverage does not yet
National model reports (ALLaM, Jais 2, Falcon-Arabic, Fanar)SDAIA; Inception/MBZUAI; TII; QCRI/MCIT, 2023–2026 1410Model-specific evaluations, including dialect and human evaluation in the Fanar and Jais 2 reportsTechnical reportsVendor-reported; instruments vary by report
Source: wesolutions Research compilation of the instruments cited.
03

The governance anchors: why this is compliance-adjacent, not a quality preference

In the UAE's banking sector the model-validation obligation is explicit. The Central Bank of the UAE issued its final Model Management Standards and Guidance on 21 December 2022, applying to all banks and branches operating in the UAE; the Standards cover model governance, data management, development, implementation, usage and independent validation across the model lifecycle, and required banks to submit a gap assessment and remediation plan within six months 6. The CBUAE has since issued guidance on consumer protection and the responsible adoption of AI and machine learning by licensed financial institutions, which expects AI governance to align with the Model Management Standards, meaningful human oversight of consequential decisions, continuous monitoring and validation, and equivalent standards for third-party AI providers 7. A bank deploying an Arabic-language agent therefore already owes its validator evidence of fitness in Arabic; an English benchmark report does not discharge that duty.

In Saudi Arabia the anchor is SDAIA. The authority published its Generative AI Guidelines for government entities on 10 January 2024 and a companion edition for the public on 11 January 2024, covering responsible use, risk identification and mitigation, and alignment with SDAIA's AI Ethics Principles and the Personal Data Protection Law 8. The guidelines are guidance rather than statute, but they define what 'responsible adoption' means for the Kingdom's government buyers — and government buyers write those expectations into procurement. Exhibit 3 summarises the instruments. The practical reading for vendors: evaluation evidence in the language of deployment is becoming part of the compliance file, whether the file is addressed to a central-bank validator, a data office or a ministry's procurement committee.

EXHIBIT 3
Governance anchors that make Arabic evaluation a compliance-adjacent requirement
JurisdictionInstrumentStatusWhat it requires or expects
UAE (banks)CBUAE Model Management Standards and Guidance, final 21 Dec 2022 6Mandatory for all banks and branchesLifecycle model governance: inventory, data management, development, implementation, usage, independent validation; gap assessment was due within six months
UAE (banks)CBUAE guidance on consumer protection and responsible AI/ML adoption 7Supervisory guidanceAI governance aligned with the Model Management Standards; human oversight of consequential decisions; continuous monitoring and validation; equivalent standards for third-party AI
Saudi ArabiaSDAIA Generative AI Guidelines — government (10 Jan 2024) and public (11 Jan 2024) editions 8Authoritative guidanceResponsible adoption, risk identification and mitigation; alignment with SDAIA AI Ethics Principles (Sep 2023) and the Personal Data Protection Law
Saudi ArabiaPersonal Data Protection Law (Royal Decree M/19 of 2021, as amended; in force Sep 2023) 8StatuteLawful processing of the personal data that task-grounded Arabic test sets will inevitably touch
Sources: as cited. The table lists anchors most relevant to Arabic AI deployment; it is not a complete inventory of GCC AI instruments.
04

Seven design principles for an Arabic-first harness

The harness that satisfies both the engineering and the governance requirement can be specified in seven principles (the authors' synthesis). None requires inventing new science; all require discipline.

  • Dialect-by-task coverage, not a single Arabic score. Evaluate each production task in MSA and in each dialect the user base actually writes, following the AraDiCE method of post-edited parallel sets, and report the matrix rather than the average 5.
  • Task-grounded private test sets. Draw evaluation items from the institution's own documents, tickets and transcripts (with data-protection clearance), because public benchmarks cannot represent the institution's formats or terminology — and keep them private so they cannot leak into training data 11.
  • Contamination control. Treat any public benchmark score as an upper bound: public sets circulate in training corpora, so complement them with freshly authored or held-out items, and date-stamp every set 211.
  • Safety and refusal testing in Arabic. Jailbreaks, harmful-content probes and refusal-policy tests must be authored in Arabic, including dialect and transliteration variants; AraTrust is the public starting point, not the finish line 2.
  • Human adjudication by Arabic speakers. Generative outputs need rubric-based human scoring by evaluators who read the users' Arabic, with inter-rater agreement tracked; automatic metrics alone are difficult to defend in front of a validator 411.
  • Versioned, reproducible scoring. Pin the harness version, model version, prompts and decoding parameters for every run, the way the open leaderboards pin theirs, so that results can be reproduced at audit 2.
  • Publishable methodology, private data. The method should survive publication even where the test items cannot — which is precisely the standard a model-risk function applies to any validation 6.
05

Implications

For buyers. Require the dialect-by-task matrix in every AI procurement where Arabic users are in scope, and require the vendor to state which public benchmarks it reports, which versions, and what it holds private. Treat an English-only evaluation annex as a gap, not a formality. For banks, connect the AI evaluation file to the existing model-validation function rather than building a parallel one — the CBUAE Standards already define the governance container 67.

For vendors. Build the harness once, as infrastructure, and amortise it across engagements: the public layer (OALL suite, ArabicMMLU, AraDiCE-style dialect sets) is reusable by construction, and the private layer is a template instantiated per client 235. The firms that can hand a validator a signed Arabic evaluation dossier will find that it functions as a commercial differentiator precisely because it is scarce.

METHODOLOGY

Desk research completed 9 October 2026. Benchmark facts are taken from the primary publications and leaderboard documentation (ArabicMMLU, AlGhafa, the Open Arabic LLM Leaderboard v1 and v2, AraTrust, AraDiCE); dialect-gap figures from the Fanar technical report's AraDiCE evaluations; model facts from the ALLaM, Jais 2, Falcon and Fanar publications; regulatory facts from the CBUAE's published Standards and subsequent AI guidance as summarised by EY, CRISIL and Plenitude, and from SDAIA's January 2024 guideline publications as recorded by DataGuidance and Digital Policy Alert. The seven design principles are the authors' synthesis and are labelled as such.

LIMITATIONS
  • The dialect-gap exhibit reports two models on one task family; it demonstrates the direction and plausible size of the effect, not a universal constant. The gap must be measured per system.
  • Benchmark versions move: OALL re-bases its suite, and public sets are periodically revised. Figures are as published at the access date.
  • The CBUAE Model Management Standards bind banks; their read-across to other regulated sectors is an inference from supervisory practice, stated as such.
  • SDAIA's guidelines are guidance, not statute; their force operates through government procurement and ethics-review practice.
  • No proprietary evaluations were run for this paper.
ENDNOTES
  1. 01SDAIA, ALLaM Arabic large language model family (2024), served through HUMAIN and Amazon Bedrock (AWS, September 2026); Technology Innovation Institute, Falcon model family and Falcon-Arabic (2023–2025); G42 Inception / MBZUAI, Jais (2023). https://falconllm.tii.ae/
  2. 02El Filali, A. et al., 'Introducing the Open Arabic LLM Leaderboard' (2A2I, TII, Hugging Face, May 2024) and 'The Open Arabic LLM Leaderboard 2' (2025): OALL v2 suite comprising AlGhafa, Native ArabicMMLU, EXAMS, human-translated MMLU, AraTrust (eight sub-tasks), MadinahQA and ALRAGE, run on a common evaluation harness. Accessed 9 October 2026. https://huggingface.co/blog/leaderboard-arabic-v2
  3. 03Koto, F. et al., 'ArabicMMLU: Assessing Massive Multitask Language Understanding in Arabic', Findings of ACL 2024 (arXiv:2402.12840): 14,575 native Arabic multiple-choice questions across 40 tasks in Modern Standard Arabic, sourced from school exams across eight Arab countries. https://arxiv.org/abs/2402.12840
  4. 04Fanar Team (QCRI / Ministry of Communications and Information Technology, Qatar), 'Fanar: An Arabic-Centric Multimodal Generative AI Platform' (arXiv:2501.13944), including AraDiCE dialect evaluations; AraDiCE-PIQA accuracy by register for Fanar-1-9B-Instruct (MSA 67.68, Egyptian 63.66, Levantine 59.03) and ALLaM-7B-Instruct-preview (67.52, 63.44, 60.88) as reported on the Fanar-1 evaluation tables. https://arxiv.org/abs/2501.13944
  5. 05Mousi, B. et al., 'AraDiCE: Benchmarks for Dialectal and Cultural Capabilities in LLMs' (Qatar Computing Research Institute, 2024): post-edited Egyptian and Levantine parallel versions of standard benchmarks, including PIQA and ArabicMMLU subsets. https://arxiv.org/abs/2409.11404
  6. 06Central Bank of the UAE, Model Management Standards and Guidance, final issued 21 December 2022: mandatory for all banks and branches operating in the UAE; components covering model governance, data management, development, implementation, usage and independent validation; gap assessment and remediation plan due by 21 June 2023. As summarised by EY, 'CBUAE Model Management Standard and Guidance', and CRISIL, 'Central Bank of UAE raises the MRM game', June 2024. https://www.ey.com/en_ae/risk/cbuae-model-management-standard-and-guidance
  7. 07Central Bank of the UAE, guidance on consumer protection and the responsible adoption of artificial intelligence and machine learning by licensed financial institutions, as summarised by Plenitude Consulting: AI governance aligned with the Model Management Standards; human oversight; continuous monitoring, validation and periodic independent review; third-party AI risk management. Accessed 9 October 2026. https://www.plenitudeconsulting.com/news-insights/cbuae-issues-new-ai-and-ml-guidance-for-financial-institutions
  8. 08SDAIA, Generative Artificial Intelligence Guidelines for government entities (published 10 January 2024) and for the public (11 January 2024), aligned with SDAIA's AI Ethics Principles (September 2023) and the Saudi Personal Data Protection Law; as recorded by DataGuidance, 'Saudi Arabia: SDAIA publishes guidelines on generative AI', and Digital Policy Alert. https://www.dataguidance.com/news/saudi-arabia-sdaia-publishes-guidelines-generative-ai
  9. 09Almazrouei, E. et al., 'AlGhafa Evaluation Benchmark for Arabic Language Models', ArabicNLP 2023 (Technology Innovation Institute). https://aclanthology.org/2023.arabicnlp-1.21/
  10. 10Inception (G42) / MBZUAI, 'Jais 2: A Family of Arabic-Centric Open Large Language Models' (arXiv:2608.13580), including OALL v2 evaluation results for Jais 2, Fanar and ALLaM models. https://arxiv.org/abs/2608.13580
  11. 11'Evaluating Arabic Large Language Models: A Survey of Benchmarks, Methods and Gaps' (arXiv:2510.13430, 2025): survey of the Arabic evaluation landscape, including the OALL leaderboards, ArabicMMLU, AlGhafa and dialectal and task-grounded evaluation gaps. https://arxiv.org/abs/2510.13430
ACRONYMS
2A2I
Open-source Arabic AI initiative; co-publisher of the Open Arabic LLM Leaderboard
AI
Artificial intelligence
CBUAE
Central Bank of the United Arab Emirates
GCC
Gulf Cooperation Council
LLM
Large language model
MBZUAI
Mohamed bin Zayed University of Artificial Intelligence
MCIT
Ministry of Communications and Information Technology (Qatar)
ML
Machine learning
MMLU
Massive Multitask Language Understanding (benchmark family)
MSA
Modern Standard Arabic
OALL
Open Arabic LLM Leaderboard
PIQA
Physical Interaction Question Answering (commonsense benchmark)
QCRI
Qatar Computing Research Institute
SDAIA
Saudi Data and AI Authority
TII
Technology Innovation Institute (Abu Dhabi)
ABOUT THIS REPORT

wesolutions Research. Desk-based analysis of published Arabic benchmarks, national model technical reports and regulator model-governance instruments; all figures carry an endnote.

Public.

DATA

Every exhibit in this report can be downloaded as a CSV file, with its source line, for independent re-scoring.

Back to library
NEXT REPORT · ARTICLE · OCT 2026

Why the senior lead has to be the engineer.