Two PDFs can look identical on screen — same layout, same text, same appearance — while being fundamentally different underneath. One might be a genuine digital document, built from real text and vector graphics. The other might be a scan: a photograph of a printed page, saved as a PDF, with no real text inside it at all. This single distinction explains why the same task can behave completely differently depending on which kind of file you are actually working with, and it is worth being able to check quickly rather than guessing.
The fastest test: try to select text
Open the PDF in any standard viewer and try to click and drag to select a line of text, the same way you would select text on a website. If individual words and letters highlight the way normal text does, the document is digital — it contains real, selectable text. If nothing highlights, or the whole page behaves like one single image no matter where you click, the document is scanned; what looks like text is actually part of a picture of the page.
Why this distinction matters so much in practice
- Compression: a scanned PDF is essentially a stack of photographs and compresses dramatically; a digital PDF is mostly small text data and often has little to compress in the first place.
- Converting to Word or Excel: a digital PDF converts cleanly, since the real text and structure are already there to work from; a scanned PDF needs its text recognised first, which can introduce the occasional misread character.
- Copy-pasting: works instantly on digital text; does nothing at all on a scan, since there is no text there to select.
- Searching within the document: works on digital PDFs; does not work on a scan unless it has been through a separate text-recognition process.
A second test: check the file size relative to page count
If selecting text is inconclusive for some reason, file size is a useful secondary clue — a digital, text-based PDF is typically very small per page (often well under 100 KB per page for plain text), while a scanned page, even compressed, tends to run noticeably larger, since it fundamentally holds more raw visual data. A ten-page document at 200 KB total is almost certainly digital; the same ten pages at 8 MB total is almost certainly a stack of scans.
A less common middle case: a scanned PDF with recognised text added
Some scanned documents have been through a text-recognition process at some point, which adds an invisible layer of selectable text behind the scanned image — you can select and copy text from a document like this, but the visible page you are looking at is still fundamentally an image, not real digital text and layout. This kind of file behaves like a digital PDF for text selection and searching, but still behaves like a scan for compression, since the underlying visual content is still an image that needs re-encoding, not vector text that barely needs compressing at all.
Why a mixed document can confuse this check
Some PDFs combine both kinds of content — several pages of real digital text, with one or two pages that are scanned images pasted in (a signed page, a stamped certificate, a photographed attachment). Checking only the first page of a document like this can give a misleading impression of the whole file; it is worth checking a few different pages, particularly any that look visually different from the rest, before concluding the whole document is one type or the other.
What this means for planning your next step
Once you know which kind of PDF you are working with, the right next step becomes much clearer. For a digital PDF that needs to be smaller, compression may have limited room to work with, and reducing page count is often more effective; see our guide on what to do when a PDF will not compress. For a scanned PDF, compression usually has plenty of room to work, covered in our guide to reducing the size of a scanned PDF.
Checking before you commit to a workflow
This quick check is worth doing before starting any significant task on an unfamiliar PDF — before promising someone a quick copy-paste of a quote, before assuming a compression pass will dramatically shrink a file, before expecting a conversion to come out perfectly clean. Thirty seconds spent confirming which kind of document you actually have avoids a good deal of confusion about why a tool is not behaving the way you expected partway through a task.
Why this is worth checking even for a document you created yourself
It is easy to assume you already know which kind of PDF you have because you know how it was created, but this can be misleading — a document you typed and exported might have had a scanned signature page inserted into it later, quietly turning part of it into scanned content without the rest of the process changing. Checking directly, rather than relying on memory of how the document originated, avoids being caught out by exactly this kind of mixed document.
One more useful signal to check
Zoom in significantly on the document — a digital PDF's text stays perfectly crisp at any zoom level, while a scanned page will visibly show its underlying pixel grain once you zoom in far enough, which is a reliable secondary confirmation alongside the text-selection test.
Frequently asked questions
Can a PDF be partly digital and partly scanned?
Yes, this is common enough to be worth specifically checking for in any document you have not looked at closely before, particularly one assembled from several different sources.
Does a scanned PDF always look lower quality than a digital one?
Not necessarily at normal viewing size, especially with a good-quality scan; the difference is more about what is happening underneath than how it looks on screen at a glance.
If I can select text, does that guarantee the document is fully digital?
Mostly yes, though as noted, a scanned document with a recognised text layer added can also allow text selection while still being fundamentally an image underneath for other purposes like compression.
Is there a way to convert a scanned PDF into a genuinely digital one?
Converting it to Word with PDF to Word and then exporting that back to PDF produces a document with real, recognised text, though it will not be pixel-identical to the scan's original appearance.
Does this distinction matter for printing?
Not really — both print the same way; the distinction mainly matters for compression, conversion, searching and copy-pasting rather than for print output.
Why would someone create a scanned PDF instead of a digital one on purpose?
Usually because the source document only exists on paper — a signed contract, an official stamped document — and scanning is the only way to get it into a PDF at all.
Is checking this free, and does it require any special tool?
No special tool is needed — any standard PDF viewer lets you try selecting text, which is the quickest and most reliable check.
Does the same check work on a phone, not just a desktop viewer?
Yes, most mobile PDF viewers support text selection the same way, so the same quick test works regardless of the device you happen to be using.