Image-to-text OCR
Avg. CER0.125%
Avg. WER0.250%
chrF85.23
Rank 1 of 17 systems on all three
Benchmark published · vendor-conducted, methodology open
5,724 documents across four benchmark tracks. Against 16 leading OCR and AI systems, YaiGlobal ranked #1 of 17 on both accuracy metrics that matter.
0.125% CER 0.250% WER #1 of 17 systems evaluated
The proof
3,760 documents. 12 benchmark datasets. 16 competing systems, including Google, OpenAI and Microsoft. YaiGlobal ranked #1 overall on both accuracy metrics.
Average Word Error Rate, lower is better. Seven of the 17 evaluated systems are shown — YaiGlobal with the closest commercial and document-parsing competitors. YaiGlobal also ranked #1 on Character Error Rate and chrF.
The same evaluation measures table structure, end-to-end PDF conversion and page layout across 5,724 processed items.
Image-to-text OCR
Avg. CER0.125%
Avg. WER0.250%
chrF85.23
Rank 1 of 17 systems on all three
Table extraction
TEDS97.99%
Table types evaluated13
Weakest type — merged cells93.75%
Rank 1 — 7.44 points ahead of LlamaParse
PDF-to-Markdown
MARS overall79.97%
Text — chrF79.81
Tables — TEDS80.12
Rank 1 on MARS and TEDS — 3.54 points ahead of LlamaParse
Layout detection
Precision0.916
F10.801
Recall0.711
Rank 1 on precision and F1
mAP and recall trail DETR (Docling)
See it happen
A single illustrative page, watched through YaiGlobal’s pipeline — read, detect regions, parse and construct the layout, in sequence.
Scanned document
تشهد المكتبات الرقمية اليوم تحولاً جذرياً في طرق حفظ الوثائق العربية ومعالجتها، إذ تتيح تقنيات التعرف الضوئي الحديثة استخلاص النصوص بدقة عالية.
| القسم | الصفحات | الحالة |
|---|---|---|
| الفصل الأول | ١–٤٨ | مكتمل |
| الفصل الثاني | ٤٩–١١٢ | قيد المراجعة |
| الملاحق | ١١٣–١٣٠ | مكتمل |
Structured output
Illustrative page composed for this demo, not a customer document. Labels reflect structure detected — not per-field accuracy scores.
Built for Arabic
A native Arabic-speaking team, working from two offices, built the engine that leads this benchmark — Arabic first, everything else extending from it.
Optimized and tested on the hardest Arabic documents — connected letterforms, calligraphy, degraded historical print.
Every extracted element traces back to the exact region of the page it came from.
Headings, tables and reading order preserved — not just a stream of recognized characters.
Santa Clara, USA and Ariana, Tunisia — engineering and native-language depth in the same team.
Deploy on Your Terms.
Fully Managed Cloud
Access our high-performance web platform with zero infrastructure overhead.
On-Premises
Tailored for enterprises where strict data sovereignty, privacy, or regulatory compliance require total infrastructure control.
Custom solutions
Libraries, universities, government, publishers, research institutions and enterprises with large Arabic document collections.
From scanned collection to structured, searchable library.
Citation linking, full-text search and retrieval-augmented Q&A over your collection.
Manuscripts, degraded print and institutional repositories.
API and workflow-specific tuning for existing systems.
Send a sample. We’ll run it through the same pipeline that scored #1 of 17 in this report.