Introduction to Big Data
advanced40 minLearning objectives
- Explain the characteristics of Big Data using the four Vs
- Explain why traditional relational databases become unsuitable at sufficiently large scale
- Identify realistic examples of Big Data in different industries
Learn
AQA 4.11.1 — Introduction to Big Data
Retrieval: the previous five lessons built genuine competence with a small, well-behaved relational database. This lesson asks what changes when the data involved is no longer small — and explicitly retrieves Year 12 Sequence 5's algorithm-performance reasoning to explain precisely why scale is the actual problem, not merely "a lot of data."
Key vocabulary — the four Vs
- Volume — the sheer quantity of data: terabytes, petabytes, or more, far beyond what fits in memory or even on a single machine's storage.
- Velocity — the speed data is generated and must be processed, sometimes requiring near-instant handling (a live stream of sensor readings) rather than periodic batches.
- Variety — data arriving in many different forms at once: structured (like this course's neat tables), semi-structured (JSON, log files), and unstructured (images, video, free text).
- Veracity — the trustworthiness and accuracy of the data itself; large-scale data sources are often noisy, incomplete, or contradictory in ways a small, carefully-curated dataset usually isn't.
Understand — why volume alone breaks a traditional relational approach
This course's sample database has 5 customers and 7 orders — every query so far runs instantly, even with a full table scan. Year 12 Sequence 5 established that an algorithm's cost genuinely depends on the size of its input: an O(n) linear search over 7 rows is trivial, but the same O(n) operation over several billion rows is not, even on powerful hardware — and many operations a single-machine relational database relies on (joining, sorting, maintaining indexes) get proportionally more expensive as volume grows, sometimes far worse than linearly.
See it — a genuine scale comparison
| This course's sample database | A real large-scale system | |
|---|---|---|
| Orders table size | 7 rows | Billions of rows, growing every second |
| Fits in memory? | Trivially | Often does not fit on a single machine at all |
A single SELECT ... WHERE | Instant | May need to be distributed across many machines to complete in reasonable time |
Understand — velocity and variety, concretely
A social media platform must ingest and process posts, likes, and messages continuously, often needing near-real-time responses (velocity) — a relational database designed for periodic, structured batch updates struggles to keep pace. A hospital system storing structured patient records alongside unstructured MRI scan images and free-text clinical notes (variety) cannot represent everything cleanly as rows and columns the way this course's schema does.
Real-life examples across industries
- Social media — billions of posts, likes and interactions generated continuously worldwide.
- IoT sensor networks — thousands of connected devices (smart meters, industrial sensors) each streaming readings every few seconds.
- Genomics — a single human genome sequence is itself several gigabytes; research studies combine thousands of them.
- Financial trading — markets generating millions of price-change events per second, requiring near-instant processing.
Reason about veracity — a genuinely distinct problem from the other three
Even if volume, velocity and variety were all solved, large-scale data sources are often unreliable: sensor readings drift or fail, user-submitted data contains typos and duplicates, and different data sources may genuinely disagree. Veracity is not a storage problem at all — it's a trust problem, requiring separate techniques (cleaning, validation, cross-checking) that this course's small, hand-curated sample data never needed.
Common mistake
Treating "Big Data" as simply meaning "a large database," and assuming a sufficiently powerful single server eventually solves it. The four Vs show scale alone is only one dimension — velocity, variety and veracity are genuinely different kinds of challenge that raw processing power or storage capacity don't directly solve.
Check your understanding
A weather-monitoring network has 50,000 sensors worldwide, each reporting a reading every 10 seconds, in a mixture of numeric readings and free-text fault reports, with some sensors known to occasionally send corrupted data. Identify which of the four Vs is most clearly illustrated by each aspect of this scenario. (4 marks)
(50,000 sensors each reporting every 10 seconds = Velocity (continuous, high-frequency data). The sheer accumulated total over time = Volume. A mixture of numeric readings and free-text fault reports = Variety (structured and unstructured data together). Sensors occasionally sending corrupted data = Veracity (the data cannot be fully trusted at face value).)
Challenge
For an online retailer processing millions of transactions daily across a global customer base, identify one genuine example of each of the four Vs specific to that business (not simply repeating this lesson's own examples).
Looking ahead: the next lesson covers the genuine opportunities, ethical issues and technical approaches organisations use once they've recognised they're dealing with genuine Big Data.