PDF guide
Text PDF or scanned PDF? The difference matters.
Both may look identical on screen, but they require different extraction methods and carry different accuracy risks.
Text-based PDF
A text PDF contains selectable characters. The converter can usually read dates, descriptions and amounts directly, making processing faster and more reliable. Try selecting a transaction description in your PDF viewer: if the words highlight individually, it is probably text-based.
Scanned or image-only PDF
A scanned statement contains page images rather than usable characters. The converter renders affected pages and runs optical character recognition locally. OCR estimates the text from pixels and can confuse similar characters, punctuation, minus signs and decimal separators.
Improve a scan before conversion
- Use a straight, complete scan at 300 DPI where possible.
- Avoid shadows, fingers, folds, glare and cropped margins.
- Keep pages upright and use consistent orientation.
- Prefer the original bank download over a photograph or screenshot.
- Check every low-confidence amount against the original.
Hybrid statements
Some PDFs contain selectable text on certain pages and images on others. The converter assesses pages individually and may use direct extraction and OCR in the same conversion.