Translation that scales · Lesson 5

Arabic is a craft, not a checkbox.

Direction mirrors the whole page, punctuation has Arabic forms, two numeral systems compete, and two letters cause half the spelling complaints. The details every Arabic reader sees instantly, and how to get them right at scale.

Every lesson so far applies to any language pair. This one is for the pair that anchors institutional translation in this region: Arabic and English. Arabic is written right to left, its letters connect and change shape by position, its punctuation and numerals have their own forms, and its orthography has traps that turn a correct translation into a careless-looking document. If you publish in Arabic, these details are your quality bar, because every native reader sees them instantly.

Direction is a document property, not a text property

The obvious fact: Arabic runs right to left. The less obvious consequence: the entire document mirrors. Page one's visual anchor is the top right. Sidebars sit on the right. Table columns run right to left, so the "first" column is the rightmost. In presentations, a photo placed on the left of an Arabic slide should generally sit on the right of its English twin. Reading gravity flips, and a layout that ignores the flip feels subtly backwards to native readers even when every word is correct: emphasis lands in the wrong place, the eye enters the page at the exit.

This is why lesson two called direction a document-level property. Translating AR to EN means re-basing the whole layout, shapes and images included, not just re-aligning paragraphs.

Punctuation has an Arabic form

Arabic has its own punctuation marks: the question mark ؟ (mirrored, opening leftward), the comma ، (raised and reversed), and the semicolon ؛. Using Latin ? and , inside Arabic text is the typographic equivalent of writing English with upside-down question marks: readable, and wrong. It is also one of the most common machine-output defects, because models trained mostly on web text absorb sloppy punctuation from it. The good news is that this error is perfectly mechanical, which makes it a QA check rather than a judgment call.

Numerals: two systems, one rule

Arabic text can carry two numeral systems: Arabic-Indic digits (٠١٢٣٤٥٦٧٨٩) and the Latin digits (0123456789) that English also uses. Which is correct? Convention varies by country and institution: Gulf publications often keep Latin digits in Arabic text, especially in financial material, while other traditions and more formal contexts prefer Arabic-Indic. Both are legitimate. What is not legitimate is mixing them arbitrarily, ٢٠٢٥ in the heading and 2025 in the table below it. Pick a convention per document type, write it into your style guide, and enforce it. One more trap: even with Arabic-Indic digits, multi-digit numbers read left to right (٢٠٢٥ is two-zero-two-five), which regularly confuses layout software handling mixed-direction lines.

The numeral decision belongs in governance, not in each translator's head. It is a one-line style-guide entry that prevents a thousand inconsistencies.

Kashida and diacritics: the craft details

Kashida (tatweel) is the elongation stroke that stretches the connection between Arabic letters, historically used by typesetters for justification and emphasis. In digital documents it is mostly a hazard: typists insert it decoratively, and it then breaks searching, matching, and glossary lookups, because "شركة" and "شركـــة" are different strings to software. Institutional style guides generally say: no kashida in body text, let the font and justification engine do their work.

Diacritics (tashkeel, the short-vowel marks) follow a similar rule. Fully vocalized text belongs in the Quran, poetry, and children's books. Modern institutional prose leaves them out, adding a single diacritic only where a word would otherwise be genuinely ambiguous. Sprinkling them randomly, or stripping one that was disambiguating, is another mark of carelessness a reviewer catches quickly.

Two letters that cause half the spelling complaints

Arabic orthography has two famous traps. Taa marbuta (ة), the feminine ending, looks like haa (ه) with two dots, and dropping the dots changes the word: mudira (مديرة, a female director) becomes mudirah-spelled-wrong, and in some pairs the meaning shifts entirely. Hamza (ء) is harder: the glottal stop sits on different "seats" depending on surrounding vowels, on alif (أ / إ), waw (ؤ), yaa (ئ), or alone (ء), and the rules fill pages. Native writers get hamza seats wrong constantly; machines trained on native writing inherit the errors. Both are checkable: Arabic-aware QA can flag suspect taa marbuta and hamza patterns mechanically, turning a proofreader's pet peeve into a lint rule.

Mixed directions: the acronym in the middle of the sentence

Real institutional Arabic is full of embedded Latin text: an English acronym, a product name, a URL, a legal citation. Each flips direction mid-line, and the Unicode bidirectional algorithm decides how the pieces visually order themselves. Mostly it guesses right; the classic failures are punctuation jumping to the wrong end of a line that ends with Latin text, parentheses appearing reversed, and adjacent Latin fragments swapping order. There is no prose fix, only correct handling in the layout engine, and it is a major reason "the translation is fine but the document looks broken" happens with naive tools.

One line, two directions

An Arabic sentence mentioning "KPI" and "2025" contains three directional runs: RTL Arabic, an LTR acronym, LTR digits, then RTL Arabic again. The words are trivial to translate. Rendering the line so it reads correctly is pure typography, and it is where naive pipelines visibly fail.

Fonts: the last mile

Arabic script connects, so a font is not just letterforms but a system of positional shapes and ligatures. Not every font family ships an Arabic companion, and substituting a mismatched fallback gives you the familiar defect of Arabic text that looks pasted into an English document: wrong weight, wrong height, broken rhythm against neighboring Latin text. Document translation has to carry font intent across the language boundary, choosing an Arabic face that matches the original's weight and role, which is part of the layout preservation work from lesson two.

Where this leaves you

None of these details is difficult individually. What makes Arabic document quality hard is that there are a dozen of them, they are invisible to non-readers of Arabic, and every one of them is instantly visible to your actual audience. The mature answer is the one this whole course has been building: conventions written into a style guide and glossary (lesson three), applied by the machine, verified by Arabic-aware QA checks, and judged by a human reviewer (lesson four), inside documents whose structure and direction survived translation (lesson two). This is the flagship territory for TranslateX: native Arabic and RTL handling, from punctuation and typography nuance to hamza and taa marbuta QA checks, with layouts mirrored when direction flips. That closes the course. The machine does the labor, your conventions do the consistency, and humans do the accountability: translation that scales.

Mirroring

The page flips with the language.

The Arabic original anchors top right and reads leftward; its English twin anchors top left. Same chart, same table, same hierarchy, opposite reading gravity. Words translated into an unmirrored skeleton feel backwards to every native reader.

  • Reading gravity: the eye enters where the direction starts
  • Tables reverse column order; sidebars and images swap sides
  • Presentation shapes mirror position when direction flips
Annual Report 2025.pdfArabic original · RTLAR

التقرير السنوي ٢٠٢٥

Annual Report 2025 (EN).pdfEnglish copy · LTR, mirroredEN

Annual Report 2025

The details

Small marks, loud signals.

Arabic punctuation, numeral consistency, taa marbuta, hamza seats, and mixed-direction lines: each is a small mechanical detail, and each one tells an Arabic reader whether the document was produced with care. Mechanical details deserve mechanical checks.

DetailThe trap
؟ ، ؛Latin ? , ; left in Arabic textPunctuation
١٢٣ vs 123Numeral system inconsistent within one documentNumerals
ة vs هTaa marbuta typed as haa changes the wordSpelling
أ إ آ اWrong hamza seat, or hamza dropped entirelyHamza
AR + EN in one lineMixed-direction text reorders around acronymsBidi

Five details that mark a document as carelessly produced to any Arabic reader, and that automated Arabic-aware QA can flag mechanically.

Frequently asked questions

What does it mean that a layout mirrors in RTL?

The whole visual structure flips, not just the text alignment: the page anchors top right, sidebars and images swap sides, table columns run right to left, and presentation shapes sit in mirrored positions. A translated document that keeps the old skeleton feels backwards to native readers even when every word is correct.

Should Arabic documents use Arabic-Indic or Latin numerals?

Both are legitimate; convention varies by country, institution, and document type. Gulf financial material often keeps Latin digits, while more formal traditions prefer Arabic-Indic (٠١٢٣). The real rule is consistency: pick one convention per document type, record it in the style guide, and never mix systems within a document.

What are the taa marbuta and hamza problems?

Taa marbuta (ة) loses its two dots and becomes haa (ه), silently changing the word. Hamza sits on different seats (أ إ ؤ ئ ء) depending on rules subtle enough that native writers get them wrong, and models trained on native writing inherit the mistakes. Both are mechanical enough for Arabic-aware QA checks to flag automatically.

Why do lines mixing Arabic and English sometimes render in the wrong order?

A mixed line contains multiple directional runs (RTL Arabic, LTR acronyms, LTR digits), and the Unicode bidirectional algorithm decides their visual order. It usually guesses right, but punctuation at run boundaries, parentheses, and adjacent Latin fragments are classic failure points, which only correct handling in the layout engine fixes.

Is kashida wrong in Arabic text?

Not historically: it is a legitimate typesetting device for justification and emphasis. In digital body text it mostly causes harm, breaking search, matching, and glossary lookups because elongated and plain spellings are different strings. Most institutional style guides ban it from body text and let the justification engine do the work.

Translate the document. Keep the design.

Right-click a file, work inside Office, or press F6 on anything on screen. TranslateX returns the same document in the other language: fonts, tables, and layout intact, terminology on brand.

Arabic, English & more · Layout preserved · On-premises available