Sound Representation
intermediate30 minLearning objectives
- Explain how sound is digitised
- Calculate audio storage requirements
- Evaluate the impact of sampling rate and bit depth on sound quality
Learn
AQA 4.5.6 — Sound representation
Retrieval: the previous lesson stored a continuous picture as a grid of discrete pixel samples. Sound presents the same fundamental challenge in a different form: real sound is a smooth, continuous wave, but a computer can only store discrete numbers — so how does digitisation actually work?
Key vocabulary
- Analogue signal — the original, continuously-varying sound wave (air pressure changing smoothly over time).
- Sampling — measuring the wave's amplitude (height) at regular, discrete moments in time.
- Sampling rate — how many samples are taken per second, measured in Hertz (Hz); e.g. 44,100 Hz means 44,100 samples every second.
- Bit depth (or resolution) — how many bits are used to store each individual sample's amplitude value.
Understand — turning a continuous wave into discrete numbers
A microphone produces a continuously varying electrical signal. To store it digitally, the signal is measured — sampled — at fixed, regular time intervals, and each measurement is rounded to the nearest value that fits in the available bit depth. The stored file is really just a long list of these individual sample values, played back in rapid sequence to recreate the sound.
See it — sampling a wave
Imagine a smooth wave being measured at 8 evenly-spaced instants (a very low sampling rate, chosen here just to make each individual sample visible):
Amplitude
^
3 | ●
2 | ● ●
1 | ●
0 |● ●
-1 | ●
-2 | ●
+---------------------------------> time
s0 s1 s2 s3 s4 s5 s6 s7
Each ● is one sample: a single amplitude value, recorded at one instant, and nothing else about the wave between those instants is stored at all. A higher sampling rate would place far more ● marks closer together, capturing the wave's shape more faithfully.
Sampling rate and bit depth — two separate quality dimensions
Sampling rate controls how often amplitude is measured — a low sampling rate can miss rapid changes in the wave, especially high-pitched sounds, producing a less accurate recreation. Bit depth controls how precisely each individual measurement is stored — a low bit depth rounds each sample to a coarser value, adding audible "noise" or distortion even if the timing is captured perfectly. These are independent: you can have frequent-but-imprecise samples, or infrequent-but-precise ones, and real audio quality depends on both together.
Calculate it — the file size formula
File size (in bits) = sampling rate × bit depth × duration (seconds) × number of channels.
Worked example 1 (mono): 10 seconds of audio, sampled at 44,100 Hz, 16-bit depth, 1 channel (mono): 44,100 × 16 × 10 × 1 = 7,056,000 bits, which is 882,000 bytes (about 861 KB).
Worked example 2 (stereo — double the channels): the exact same 10 seconds, sampling rate and bit depth, but stereo (2 channels): 44,100 × 16 × 10 × 2 = 14,112,000 bits — precisely double, because stereo stores two independent streams (left ear, right ear) rather than one.
Worked example 3 (lower quality, shorter clip): 5 seconds at a lower 22,050 Hz sampling rate, 8-bit depth, mono: 22,050 × 8 × 5 × 1 = 882,000 bits, which is 110,250 bytes — dramatically smaller than Example 1, from halving the sampling rate, halving the bit depth, and halving the duration all at once.
Check your understanding — calculate it yourself
Calculate the file size, in bytes, for 8 seconds of mono audio at 16,000 Hz sampling rate and 8-bit depth. (16,000 × 8 × 8 × 1 = 1,024,000 bits = 128,000 bytes.)
Evaluate — sampling rate, bit depth and quality trade-offs
A phone voice call can use a low sampling rate and bit depth, since human speech doesn't need high fidelity to remain understandable, keeping bandwidth low. A music streaming service uses a much higher sampling rate and bit depth, since listeners can perceive the extra detail, at the direct cost of a larger file (and, for streaming, more bandwidth used per second). Neither choice is "wrong" — each is matched to what the specific application actually needs.
Challenge
A podcast episode is 20 minutes long, recorded in mono at 44,100 Hz, 16-bit. Calculate its file size in MB (1 MB = 1,000,000 bytes). Then calculate how much larger the file would be if it had been recorded in stereo instead, and explain in a sentence why podcasts (mostly speech) are often recorded in mono deliberately, unlike music.
Looking ahead: the next lesson (Data Compression) directly addresses the very large file sizes these calculations have just revealed — asking how much of that data can be reduced, and by what methods, without unacceptably damaging quality.