PDF documents
A PDF has no paragraphs, only glyphs placed at coordinates on a page. The platform reconstructs a document from that geometry, lets you translate it like any other file, and delivers a PDF whose pages look like the original with the translated text in place. Because the layout is fixed, translation length matters more here than in any other format, and this page explains what to watch for.
Which PDFs work best
Section titled “Which PDFs work best”For the best result, upload a text-based PDF: one exported from a word processor, layout application or browser, where you can select and copy the text in a PDF viewer. Its text, fonts and positions are read directly, so the reconstruction is exact and nothing depends on recognition.
Scans and images
Section titled “Scans and images”A scan or photo of paper has no text layer, and neither does an image file. OCR support for these is the next step for the platform: the page is reconstructed into positioned text frames just like a text-based PDF, and translated the same way. Until it ships, the upload step asks for a file with selectable text; run the scan through an OCR tool first, or upload the original document the scan was made from. Even with OCR, a text-based original gives a cleaner result than a scan of it, so use the original when you have it.
A PDF that mixes text pages and scanned pages is accepted; the text pages are translated and the scanned pages come through without their content.
Password-protected PDFs cannot be opened. Remove the password and upload again.
How the layout is reconstructed
Section titled “How the layout is reconstructed”The pages are analysed to recover what the author meant:
- Reading order and columns — text is ordered top-to-bottom and left-to-right, with multi-column layouts read column by column.
- Paragraphs, headings and lists — inferred from font size, position and bullet glyphs.
- Tables — detected from the grid lines and cell alignment, and translated cell by cell.
- Running headers, footers and tables of contents — recognised and kept in place.
- Images and vector graphics — kept exactly where they are. Text inside images is not translated.
A tagged PDF (one saved with accessibility structure, which Word, InDesign and most browsers can produce) gives far better results: its headings, lists and tables are taken from the tags rather than guessed. If you have the choice, export a tagged PDF.
Each reconstructed paragraph becomes a row in the editor. In the preview, a PDF renders as fixed pages with every text frame at its true position and size, so you see the page as it will be delivered.
When the translation is longer
Section titled “When the translation is longer”Every text frame on a PDF page has a fixed box. A translation that is longer than the source has to go somewhere:
- A frame first grows downward and into the empty space around it, using the same rules the delivery rebuild applies, so most modest length differences are absorbed.
- Text that still does not fit is cut off in the delivered file.
The editor tells you before that happens: rows in length-sensitive frames
show a character budget (for example 87/112) rendered in the document’s
own font, turning amber as you approach the limit and red past it, with a
tooltip naming what would overflow. See
when space is limited.
Sweep the preview before delivery for red budgets, and shorten those
translations.
Where a document’s font is not available at rebuild, a metrically close substitute is used, which can shift line breaks slightly. The character budget and the preview both use the same substitution, so what you see is what ships.
Delivery
Section titled “Delivery”The delivered file is a PDF, never a Word document. It is created from the Delivery tab like any other format; see Translated files & XLIFF. Right-to-left targets are laid out right-to-left, with per-element direction controls in the preview.
Related
Section titled “Related”- Supported file formats
- Shortcuts & preview — the page preview
- Style & formatting — overflow budgets