Hindi Converter

Unicode ⇄ Kruti Dev 010 · Full Formatting

Free Forever · No Signup · Files Deleted Instantly

Hindi PDF & Scan → Word with the tables rebuilt

Upload a scanned PDF, a phone photo, or an image of a Hindi document and get back an editable Unicode Word file — with bordered tables detected and rebuilt as real Word tables, not pasted as a picture.

Drop your file here or click to browse

.docx or .xlsx files — up to 10 at once — max 15 MB per file

(Ratio: KD 15pt = English 10pt)

Your file is converted first — review it side-by-side, then download

The problem with most PDF-to-Word tools and Hindi

Two different things get called "a Hindi PDF", and they fail in two different ways. A PDF exported from a Word file typed in Kruti Dev, Kokila or Utsaah often carries a text layer that looks fine on screen but is garbage underneath — copy from it and you get letters in the wrong order, floating matras and stray halants. Generic converters trust that layer and hand you the garbage. A scanned PDF has no text layer at all, and a converter that only reads text layers returns an empty document or an image pasted into a page.

This tool checks. Every page's text layer is scored against the spelling rules of Devanagari — a word cannot begin with a vowel sign, cannot carry two vowel signs in a row, cannot end in a bare halant. A healthy page breaks those rules in under 2% of its words; a broken export breaks them in 22% or more. Pages that fail are re-read with OCR instead of being trusted, and a page that has no text layer goes to OCR automatically.

What happens to a bad photo

Real uploads are not clean scans. They are photographs taken at an angle, in a room with one window, of a photocopy of a photocopy. Before any reading happens the image is binarised (plain Otsu thresholding, or adaptive thresholding when the background brightness varies enough to indicate a shadow), auto-rotated to the correct orientation, deskewed, and checked for which script it is actually in. Measured on real government documents, that preprocessing chain lifted word recall from 93.9% to 95.8% on a clean scan and from 79% to 88% on a difficult Hindi table page.

Tables

Table detection in Devanagari has a specific trap: every Hindi word carries a shirorekha, the long horizontal bar across the top, and line-detection algorithms happily mistake a row of those bars for a table rule. A seven-row table comes back as eighteen rows. Detected row lines are therefore checked for whether they reach across the table's full width like a real rule, or stop at word boundaries like a headline. Bordered tables come back as real, editable Word tables.

Honest limits

This is not magic and the number that matters is measured, not claimed. Handwriting, heavily underlined Devanagari headings, and borderless tables are the weak spots — a borderless table falls back to plain text lines rather than being invented. Files go through a first-in-first-out queue so a busy moment does not crash the server; you are told how many files are ahead of you.

स्कैन की हुई पीडीएफ और फोटो से वर्ड फाइल — हिन्दी में

आपके पास सिर्फ कागज़ की स्कैन कॉपी है, या मोबाइल से खींची हुई फोटो — और चाहिए एडिट होने वाली वर्ड फाइल? यही काम यह टूल करता है, और बॉर्डर वाली टेबल को असली वर्ड टेबल बनाकर देता है, फोटो चिपकाकर नहीं।

कई पीडीएफ ऐसी होती हैं जो देखने में ठीक लगती हैं पर उनका अंदरूनी टेक्स्ट खराब होता है (कृति देव, कोकिला या उत्साह से बनी पीडीएफ में यह आम है) — कॉपी करते ही मात्राएँ इधर-उधर हो जाती हैं। यह टूल हर पेज के टेक्स्ट को देवनागरी के नियमों पर जाँचता है और खराब मिलने पर उसे भरोसा करने के बजाय OCR से दोबारा पढ़ता है।

टेढ़ी-मेढ़ी, धुँधली या छाया वाली फोटो के लिए पढ़ने से पहले इमेज को साफ किया जाता है — काला-सफेद किया जाता है, अपने आप सीधा घुमाया जाता है, तिरछापन ठीक होता है, और लिपि पहचानी जाती है। असली सरकारी दस्तावेज़ों पर नापने पर इस सुधार से शुद्धता 93.9% से 95.8% (साफ स्कैन) और 79% से 88% (मुश्किल हिन्दी टेबल) तक बढ़ी।

साफ बात: हाथ की लिखावट, नीचे लाइन खिंची हुई हिन्दी हेडिंग, और बिना बॉर्डर वाली टेबल अभी भी कमज़ोर जगहें हैं। बिना बॉर्डर की टेबल को अंदाज़े से बनाने के बजाय सादी लाइनों में दिया जाता है।

फाइलें कतार (फर्स्ट इन, फर्स्ट आउट) में लगती हैं और आपको बताया जाता है कि आपसे पहले कितनी फाइलें हैं — ताकि सर्वर पर एक साथ ज़्यादा बोझ न पड़े।

Frequently Asked Questions

Can it read a scanned PDF that has no text in it at all?

Yes. A page with no usable text layer is rendered as an image and read with OCR automatically. You do not have to tell it which kind of PDF you have.

Will tables in my scan come out as real Word tables?

Bordered tables are detected and rebuilt as editable Word tables. Borderless tables fall back to plain text lines rather than being guessed at.

How accurate is the Hindi OCR?

Measured on real government documents, word recall is around 95-96% on a clean scan and around 88% on a difficult table-heavy page. Handwriting and heavily underlined headings are noticeably worse.

My photo is crooked and badly lit. Is that a problem?

Usually not. The image is auto-rotated, deskewed and thresholded before reading, and adaptive thresholding kicks in when the lighting is uneven, which is the normal case for phone photos.

How many pages can I upload?

Up to 15 MB and up to 10 pages per file. Files are queued first-in-first-out and you are shown how many are ahead of yours.