Data Compression

intermediate25 min

Learning objectives

  • Distinguish between lossless and lossy compression
  • Recommend suitable compression techniques
  • Justify compression choices for different media

Learn

AQA 4.5.7 — Data compression

Retrieval: the previous two lessons calculated genuinely large file sizes for images and audio. Compression is how those files are made smaller for storage and transmission, without necessarily storing every original bit.

Key vocabulary

  • Compression — reducing a file's size by re-encoding its data more efficiently.
  • Lossless compression — shrinks a file in a way that can be perfectly reversed; decompressing recreates the exact original data, bit for bit.
  • Lossy compression — shrinks a file by permanently discarding some information judged least noticeable, achieving much greater size reduction at the cost of not being perfectly reversible.

Understand — two fundamentally different strategies

Lossless compression works by finding and eliminating genuine redundancy — patterns that repeat, or information that's predictable from context — and re-encoding them more compactly, in a way that can always be undone exactly. Lossy compression works differently: it deliberately throws away information the algorithm has judged the human eye or ear is unlikely to notice missing, which lets it compress far more aggressively, but that discarded detail can never be recovered.

See it — a simple illustration of lossless compression

A very simple lossless technique, run-length encoding, replaces a run of repeated values with a (value, count) pair instead of writing the value out every time:

Original:   AAAAAAAABBBCCCCCCCCCCDD
Compressed: (A,8)(B,3)(C,10)(D,2)

23 original characters become 4 short pairs — and critically, the original 23-character string can be reconstructed exactly, with nothing lost, by simply expanding each pair back out.

Worked comparison — why the strategies suit different data

A plain text document compresses very well losslessly, because language has huge amounts of redundancy (repeated words, predictable letter patterns) — and text absolutely must decompress back to the exact original wording; losing even one character could change its meaning. A photograph compresses far better lossily, because natural images contain fine colour/brightness detail the human eye barely perceives anyway — discarding some of it can shrink the file dramatically while looking almost identical, which lossless compression alone could never achieve on the same image.

Common mistake

Assuming "lossy" always means "noticeably worse quality." Lossy compression exists on a spectrum, controlled by a quality setting — a lightly-compressed photo can be visually indistinguishable from the original while still being meaningfully smaller; only aggressive lossy compression produces obviously visible or audible degradation. The trade-off is a dial, not a single fixed choice.

Recommend and justify — matching compression to media

MediaRecommended approachJustification
A legal contract (text document)LosslessEvery character must be preserved exactly; the file is small regardless.
A holiday photo for social mediaLossyLarge file, human eye tolerates some detail loss, big size saving matters more than pixel-perfect accuracy.
A medical X-ray imageLosslessLosing subtle detail could hide a genuine diagnosis-relevant feature — accuracy matters more than file size here.
Streamed musicLossyLarge files, most listeners cannot perceive the discarded detail at a well-chosen quality setting, and smaller files stream more reliably.

Apply it — justify a compression choice yourself

A hospital needs to store thousands of X-ray images long-term, where storage cost is a real concern, but any loss of diagnostic detail is unacceptable. A separate department needs to email preview thumbnails of the same X-rays for quick internal reference only. Recommend and justify a compression approach for each of these two use cases, even though they involve the same underlying images.

Challenge

Identify one real application you use regularly that relies on lossy compression (e.g. a streaming video/music service, a messaging app's image sharing) and one that relies on lossless compression (e.g. a ZIP archive, a software installer). For each, explain in a sentence why that choice is appropriate given what the data is used for.

Looking ahead: the final lesson of this sequence (Data Storage Calculations) ties together every file-size formula from this sequence — images, sound, and the effect compression has on all of them — into calculations spanning the full range of storage units.

Test yourself

Check your understanding with exam-style questions.

Go to Exam Practice
Log in to track this lesson on your progress dashboard.
Log in