Skip to content

Convert a Bangla PDF to Word

A Bengali PDF — a circular, a form, a book chapter, a scan — into a .docx you can edit, with the conjuncts and vowel signs in the right order and OCR set to Bangla for scanned pages. Nothing is stored.

  • Never stored
  • No queue, no waiting
  • No signup, no watermark

How it works

1

Add your PDF

A text PDF or a scan. The tool looks at the first page and switches OCR on for a scan; the reader is already set to Bengali.

2

Pick the mode

Editable text gives a normal Word document, page for page. Exact look pins every line where it was over a picture of the page, for printing rather than editing.

3

Convert and download

A .docx in a Unicode Bangla font, with headings, paragraphs, images and links.

What is different about a Bangla PDF

The reordering is the whole point. A PDF text layer holds glyphs in drawing order, and a Bangla shaping engine draws the pre-base vowel signs (ি, ে, ৈ) before the consonant and splits ো and ৌ into two glyphs around it; the text that comes out of a plain extractor is therefore "স্ব␃াভািবক" where the page says "স্বাভাবিক". Each line is put back into logical order — vowel sign after its cluster, reph as র্ before the cluster it sits on, two-part vowels as one code point — before anything else happens. OCR output needs none of this, because the engine already emits logical order.

Headings, paragraphs, links, images and page breaks are worked out exactly as on the general PDF to Word page: words sharing a baseline are a line, a widening gap starts a paragraph, a short line at 1.6 times the body size is a heading, and a link annotation becomes a Word hyperlink. Exact look renders the page artwork at 144 DPI and places the real text over it; Editable text rebuilds clean paragraphs in reading order.

Bengali digits are kept as Bengali digits, and a Latin word on the page — an email address, a reference number — stays Latin. For a scan, the Bengali pack is the one that is loaded, so a page of Bangla with an English letterhead reads the Bangla correctly and may read the letterhead as Bangla-shaped nonsense; add English in the language list if the English matters.

When you want something else

A PDF made from a Word file is a reconstruction waiting to happen; ask for the .docx if you can. And a Bijoy .docx is better converted with Bijoy to Unicode than printed to PDF and converted back — the keystrokes convert without a single error, the formatting survives, and nothing is re-read from pixels.

For an archive of scanned Bangla books, a desktop OCR tool with the Bengali model reads a folder overnight, and OCRmyPDF adds a searchable layer to each PDF without changing how it looks — the right shape of tool for hundreds of files, where this page is for one.

Frequently asked questions

Is my PDF stored?

No. Nothing is stored; the file is never kept. For a scan the reader and the Bengali pack come from this site once and are kept for next time — the document itself goes nowhere.

Why does Bangla from a PDF usually come out misspelt?

Because a PDF stores the glyphs in the order they are drawn, which for Bangla is visual order: the i-kar is drawn before the consonant it follows, and the two halves of ো sit either side of it. Most converters copy that order out as it is, so বাংলাদেশ comes out with its vowel signs in front of the wrong letters and every word is misspelt in a way a spell-checker cannot fix. This converter puts the signs back where Unicode wants them — the same reordering every Bangla PDF needs, including ones this site made.

Does it handle a PDF made from a Bijoy document?

If the PDF has a text layer in SutonnyMJ, the text comes out as Bijoy keystrokes, because that is what is in the file; run the result through Bijoy to Unicode and the formatting survives. If the PDF is a scan, OCR reads the picture and gives Unicode straight away.

How good is the OCR on Bangla scans?

Clean 300 dpi scans of printed Bangla read well; faded photocopies and phone photos lose conjuncts. The language is fixed to Bengali on this page, so a Latin letterhead or a reference number cannot push the reader into English for the whole page. Pages scanned askew are straightened first.

What Word format do I get?

A standard .docx, set in a Unicode Bangla font, that opens in Word, Google Docs and LibreOffice. The text is real Unicode: searchable, and typeable with Avro or Ridmik.

What about tables and forms?

Ruled forms in a scan come out as real Word tables in Editable text mode. Tables in a text PDF come out as positioned text, like every PDF-to-Word converter — for a spreadsheet, PDF to Excel is the better tool.

Good to know: Fonts are substituted: the .docx uses a Unicode Bangla font, not the PDF's own, so the look is close rather than identical. Text PDFs set in a Bijoy font convert as Bijoy keystrokes and need the Bijoy to Unicode converter afterwards. OCR ignores photos and logos, leaves handwriting as a picture, and takes a few seconds per page, since your own device does the reading.

Put this tool on your website

Free for any blog, class page or help article. Paste one snippet and your visitors can use it right on your page.