Getting Data Out of Arabic Documents
Invoices, delivery notes and government letters arrive as photographs. Turning them into fields your system can use is a solved problem — if you scope it properly.
Valeur X Team
Technology & Industry

The finance team receives a supplier invoice as a photograph in a WhatsApp message. Someone types nine fields into the ERP. Multiply that by four hundred a month and you have a full-time job nobody applied for.
Extraction, not reading
The goal is not for the machine to understand the document. It is to return the same nine fields in the same shape every time, with a confidence score and a picture of where each one was found. Anything uncertain goes to a person who corrects it in seconds instead of typing it from scratch.
- Supplier invoices, and the VAT number printed on them
- Delivery notes matched against a purchase order
- Bank letters, customs papers and certificates of origin
- Contracts, where only the dates, parties and values are needed
Arabic OCR is the hard part
Connected script, diacritics that may or may not be present, Arabic and Latin on the same line, a handwritten note in the margin and a stamp across the total. Quality here is decided by scanner settings and pre-processing far more than by the choice of model.
Aim for ninety per cent of fields extracted and every uncertain one flagged. Chasing a hundred per cent unattended is how these projects run over.
Put it where the work already happens
The extracted data has to land in the system the person already uses, with the original document attached beside it. If they have to open a second tool to check the result, they will go back to typing.
ZATCA e-invoicing means more of this arrives structured every year. But the supplier who still sends a photograph will be with us for a long time yet, and the finance team should not be the one paying for it.

