0.00 GB / 1.00 GB monthly quota
0.00 GB / 1.00 GB additional quota
0 / 5 daily conversions
/month
Email with pasword reset link sent.
Enter your email address and we'll send you a link to reset your password.
PDF (OCR) is a plain PDF with one addition: a text layer. A scan or a photograph of a page is, to a computer, a picture — searching it finds nothing, and no text can be selected from it. Optical character recognition reads the shapes in that picture and works out which letters they are.
The result keeps the original image exactly as it was, and writes the recognised words underneath it, invisibly, in the position each word occupies on the page. Nothing looks different when you open the file. But Ctrl+F now finds words, text can be selected and copied, and search engines, document managers and screen readers can read the contents.
This is the format to pick when the appearance of the original matters — a signed contract, an old book page, a receipt — and you also need the words to be findable. If you only want the words and do not care what the page looked like, TXT (OCR) gives you the text alone in a much smaller file.
Recognition is very good on clean, straight, printed pages and gets worse as the image gets worse. A blurred photograph taken at an angle, a faint carbon copy or handwriting will produce mistakes, and the text layer is only ever as accurate as the reading was. The picture itself is untouched either way, so nothing is lost.