Scanning & Digitization

Compress a PDF That Mixes Text and Scanned Pages

Compress a PDF That Mixes Text and Scanned Pages

A great many real documents are not purely digital or purely scanned — they are a typed cover letter with a scanned ID attached, an application form followed by a photographed supporting document, a contract with one scanned signature page. This kind of mixed document compresses a little differently than either a fully digital or fully scanned file, since the pages behave differently from each other within the same compression pass.

Why a mixed document is not quite the same problem as either pure case

A fully digital document is usually already small and needs little compression. A fully scanned document is uniformly image-heavy and compresses predictably across every page. A mixed document sits in between: most of the file size usually comes from the small number of scanned pages, while the digital pages barely contribute, which means the overall file's behaviour under compression is really being driven by just the scanned portion, even if that portion is a minority of the total pages.

Steps

  1. Open the PDF Compressor and upload the full mixed document as one file.
  2. Choose a compression level based on the scanned pages' needs, since they are what actually determines the file's size — a document that is mostly scanned images needs a stronger setting than one with just one or two scanned pages among many digital ones.
  3. Check the result, paying particular attention to the scanned pages specifically, since the digital pages will look identical at any compression level.

Why the text pages are essentially unaffected regardless of setting

Compression works by re-encoding embedded raster images; genuinely digital text has no image data for the compressor to act on, so a strong compression setting applied to the whole file has no visible effect on those pages at all — they remain exactly as sharp and selectable as before. This means you can generally choose the compression level based purely on what the scanned pages need, without worrying about degrading the digital portions of the document, since there is nothing there for the compression to degrade.

When just one or two pages are scanned among many digital ones

If a 20-page mostly digital document has just one scanned signature page attached, that single scanned page is likely responsible for a large share of the total file size on its own. In this situation, a strong compression setting applied to the whole document costs you nothing on the 19 digital pages and meaningfully shrinks the one scanned page, which is usually an easy, low-risk choice — there is little reason to hold back with a lighter setting purely out of caution for pages that will not be affected either way.

When the scanned portion is the larger part of the document

The calculation shifts if the scanned pages make up most of the document — a report with a two-page digital summary followed by fifteen pages of scanned supporting evidence, for instance. Here, treat the compression level decision the way you would for a fully scanned document, since the scanned pages are doing almost all of the work in determining the final size, and the small digital portion is along for the ride either way.

Checking each type of page separately after compressing

Once compressed, it is worth reviewing the digital pages and the scanned pages as two separate checks rather than one general skim — confirm the digital text is still sharp and selectable (it should be, at any setting), and separately confirm the scanned pages are still clearly legible at whatever level you chose. This catches the specific failure mode of a mixed document: assuming the whole file looks fine because the digital majority of pages looks fine, while missing that the one scanned page has become too soft to read clearly.

An alternative approach: compressing the scanned portion before combining

If you are building the mixed document yourself — adding a scanned attachment to an existing digital file — it is sometimes cleaner to compress the scanned portion on its own first, before merging it with the digital pages, rather than compressing everything together afterwards. This lets you fine-tune the scanned page's compression level in isolation, confirm it looks right, and then merge it into the final document with confidence, rather than compressing the whole combined file and hoping the balance comes out right.

What this means for hitting a specific overall size target

If the combined document needs to hit a specific size limit, remember that the scanned pages are almost certainly where the reduction needs to come from — there is little to gain by worrying about the digital pages, since they are already contributing very little to the total. Focus your attention and any troubleshooting specifically on the scanned portion if the combined file is still too large after a first compression pass.

A practical example that shows the pattern clearly

Take a 12-page rental application: a two-page digital form filled in on a computer, followed by a scanned ID and a scanned proof-of-address letter, then a scanned bank statement covering eight pages. The two-page digital form might weigh 40 KB in total, essentially irrelevant to the final size. The ten scanned pages, at a typical scan resolution, might weigh several megabytes on their own. If this document needs to fit under a 2 MB portal limit, the entire compression conversation is really about those ten scanned pages — the digital form could be compressed at any strength with zero visible change, so there is no trade-off to weigh there at all.

Frequently asked questions

Will compression ever make my digital text pages look worse?

No, at any compression level, since there is no image data on a genuinely digital text page for compression to act on.

How do I know which pages in my document are scanned versus digital?

Try selecting text directly on each page in a PDF viewer — if you can select and highlight it, that page is digital; if the whole page behaves as one image, it is scanned. See our guide on telling if a PDF is scanned or digital for more detail.

Should I split the document and compress the scanned pages separately?

You can, and it gives you finer control, but compressing the whole document together at a level chosen for the scanned pages usually achieves the same practical result with less effort.

What if only a small part of one page is a scanned image, not the whole page?

The compressor will still re-encode that embedded image regardless of how much of the page it occupies; the surrounding genuine text on the same page remains unaffected.

Does the order of scanned versus digital pages in the document matter for compression?

No, compression is applied per page based on that page's actual content, regardless of where it sits in the overall document.

Is there a risk in choosing too strong a setting for a mostly-digital document?

Very little risk to the digital pages themselves; the only thing to watch is whether the few scanned pages become too soft at a very strong setting, which is worth checking after compressing.

Is the tool free and private?

Yes. No sign-up, no watermark, and files are removed from the server automatically about an hour after processing.

Related guides

← Back to all posts

Keep reading