Arabic OCR

Benchmark published · vendor-conducted, methodology open

We didn’t claim Arabic OCR leadership. We measured it.

5,724 documents across four benchmark tracks. Against 16 leading OCR and AI systems, YaiGlobal ranked #1 of 17 on both accuracy metrics that matter.

0.125% CER 0.250% WER #1 of 17 systems evaluated

The proof

Arabic depth you can measure — not just a claim.

3,760 documents. 12 benchmark datasets. 16 competing systems, including Google, OpenAI and Microsoft. YaiGlobal ranked #1 overall on both accuracy metrics.

  • 0.250%Average Word Error Rate — best of 17 systems
  • 0.125%Average Character Error Rate — best of 17 systems
  • 3,760Documents evaluated across all datasets
  • 12Datasets — ranked #1 overall on both metrics
Average Word Error Rate by system Lower is better · 3,760 documents

Average Word Error Rate, lower is better. Seven of the 17 evaluated systems are shown — YaiGlobal with the closest commercial and document-parsing competitors. YaiGlobal also ranked #1 on Character Error Rate and chrF.

Four KITAB-Bench tracks, not just OCR accuracy.

The same evaluation measures table structure, end-to-end PDF conversion and page layout across 5,724 processed items.

Image-to-text OCR

Avg. CER0.125%

Avg. WER0.250%

chrF85.23

Rank 1 of 17 systems on all three

Table extraction

TEDS97.99%

Table types evaluated13

Weakest type — merged cells93.75%

Rank 1 — 7.44 points ahead of LlamaParse

PDF-to-Markdown

MARS overall79.97%

Text — chrF79.81

Tables — TEDS80.12

Rank 1 on MARS and TEDS — 3.54 points ahead of LlamaParse

Layout detection

Precision0.916

F10.801

Recall0.711

Rank 1 on precision and F1

mAP and recall trail DETR (Docling)

See it happen

From a scanned page, out of the dark, into structured knowledge.

A single illustrative page, watched through YaiGlobal’s pipeline — read, detect regions, parse and construct the layout, in sequence.

Plays while in view · pick a stage to jump

Scanned document

Heading1

حفظ التراث المخطوط العربي

Text2

تشهد المكتبات الرقمية اليوم تحولاً جذرياً في طرق حفظ الوثائق العربية ومعالجتها، إذ تتيح تقنيات التعرف الضوئي الحديثة استخلاص النصوص بدقة عالية.

Table3
القسمالصفحاتالحالة
الفصل الأول١–٤٨مكتمل
الفصل الثاني٤٩–١١٢قيد المراجعة
الملاحق١١٣–١٣٠مكتمل

Structured output

Illustrative page composed for this demo, not a customer document. Labels reflect structure detected — not per-field accuracy scores.

Built for Arabic

Arabic isn’t an afterthought at YaiGlobal.

A native Arabic-speaking team, working from two offices, built the engine that leads this benchmark — Arabic first, everything else extending from it.

  • Arabic-native intelligence

    Optimized and tested on the hardest Arabic documents — connected letterforms, calligraphy, degraded historical print.

  • Visual grounding

    Every extracted element traces back to the exact region of the page it came from.

  • Structural intelligence

    Headings, tables and reading order preserved — not just a stream of recognized characters.

  • Two offices, one product

    Santa Clara, USA and Ariana, Tunisia — engineering and native-language depth in the same team.

Deploy on Your Terms.

Flexibility without Compromising Security.

Fully Managed Cloud

Instant SaaS Integration

Access our high-performance web platform with zero infrastructure overhead.

  • Instant access to online OCR platform and REST API.
  • High-availability, multi-zone infrastructure.
  • Fully managed updates, maintenance, and scaling.

On-Premises

Your Infrastructure, Complete Control

Tailored for enterprises where strict data sovereignty, privacy, or regulatory compliance require total infrastructure control.

  • 100% on your local hardware or private cloud.
  • Total ownership of your deployment.
  • Zero accuracy loss with the freedom of local infrastructure control.

Custom solutions

Built for organizations with serious archives.

Libraries, universities, government, publishers, research institutions and enterprises with large Arabic document collections.

  • Digital library creation

    From scanned collection to structured, searchable library.

  • Search & discovery

    Citation linking, full-text search and retrieval-augmented Q&A over your collection.

  • Historical & institutional archives

    Manuscripts, degraded print and institutional repositories.

  • Custom integration

    API and workflow-specific tuning for existing systems.

Have Arabic documents?
Let’s benchmark them.

Send a sample. We’ll run it through the same pipeline that scored #1 of 17 in this report.