27. MongoDB and the Document Model¶
Lecture 26 named MongoDB as the leading example of a document store; now we open it up properly. MongoDB is the most widely used NoSQL database in the world, and it is the one you are overwhelmingly likely to meet in a real job, so this lecture and the next work through it in genuine depth: how it is structured, what a "document" actually is on disk, and how to read, write, update, and delete data with its real query syntax. Keep the relational habits from Units 3–5 close at hand — the most useful thing you can do in this lecture is notice, concretely, where MongoDB agrees with everything you already know and where it deliberately does something else.
In This Lecture¶
- MongoDB's architecture: the
mongodserver process, databases, collections, and documents - The document-oriented data model, and how it compares to the relational row/table model
- BSON — why MongoDB stores Binary JSON instead of plain text JSON
- MongoDB's core data types, including
ObjectId, dates, arrays, and embedded documents - Modeling relationships by embedding — and how that contrasts with relational normalization
- Real CRUD syntax:
insertOne/insertMany,find,updateOne/updateMany,deleteOne/deleteMany - Query operators (
$gt,$in,$and) and projections - Realistic use cases for a document database
Introduction to MongoDB¶
MongoDB is an open-source, document-oriented NoSQL database, first released in 2009 and now maintained by MongoDB, Inc. Instead of rows in tables, MongoDB stores data as documents — flexible, nested, JSON-like structures — grouped into collections. It was built from the ground up around the horizontal-scaling and schema-flexibility priorities from Lecture 26, while still offering a rich query language, indexing (the same core idea as Lecture 25, adapted to documents), and — since version 4.0 — multi-document ACID transactions.
MongoDB Architecture¶
A running MongoDB deployment centers on mongod, the core server process that stores
data and answers queries. Clients — the mongosh shell, or a driver in your application's
language — connect to mongod (directly, or in a cluster, via a router; Lecture 28 covers
that) and issue commands against a four-level hierarchy:
MongoDB's storage hierarchy
mongod server process — one running MongoDB instance
Database — e.g. university
Collection — e.g. students, courses — analogous to a relational table
Document — e.g. one student's record — analogous to a relational row
Documents and Collections¶
A collection is MongoDB's rough analog of a relational table — but unlike a table, it enforces no fixed schema: two documents in the same collection can have entirely different fields. A document is the analog of a row, written as a set of field/value pairs — but unlike a row, a document's values can themselves be arrays or nested documents, not just flat scalars.
// A single document in the "students" collection
{
_id: ObjectId("66f1a2b3c4d5e6f7a8b9c0d1"),
regNo: "FA21-BCS-045",
name: "Ayesha Khan",
gpa: 3.72,
department: "Computer Science",
isActive: true,
enrollmentDate: ISODate("2021-09-01")
}
Relational model
TableFixed columns, every row has the same shape
Document model
CollectionNo fixed columns, documents can vary in shape
BSON Data Format¶
MongoDB does not store documents as plain-text JSON. It stores them as BSON ("Binary JSON"), a binary-encoded serialization format designed for exactly two things JSON itself can't do well:
- Richer data types — plain JSON only has strings, numbers, booleans, null, arrays, and
objects. BSON adds types JSON has no native concept of, including dates, binary data, and a
dedicated 12-byte
ObjectIdtype — all of which you'll see below. - Efficient traversal — every BSON field is stored with a length prefix, so the database can skip directly past a field it doesn't need instead of parsing character-by-character the way a text-based JSON parser must.
You never write BSON by hand — you write ordinary-looking JSON-like syntax in mongosh or a
driver, and MongoDB encodes it to BSON internally, decoding it back to a readable
representation whenever you view it.
MongoDB Data Types¶
| Type | Example | Notes |
|---|---|---|
ObjectId |
ObjectId("66f1a2b3c4d5e6f7a8b9c0d1") |
A 12-byte unique identifier, auto-generated as _id if you don't supply one |
String |
"Ayesha Khan" |
UTF-8 text |
Int32 / Int64 |
42, 9223372036854775807 |
Whole numbers, sized like a relational INT/BIGINT |
Double |
3.72 |
Default type for a decimal literal |
Boolean |
true, false |
|
Date |
ISODate("2021-09-01") / new Date() |
Stored internally as milliseconds since the Unix epoch |
Array |
["CSC270", "CSC211"] |
An ordered list — values need not even share a type |
Embedded Document |
{ street: "6 Lawrence St", city: "Islamabad" } |
A document nested as the value of a field |
Null |
null |
Represents a deliberately absent value |
Embedded Documents and Arrays¶
This is the single biggest mental shift coming from the relational model. Instead of splitting related data into separate tables joined by a foreign key, a document database can embed related data directly inside the parent document — trading normalization for the convenience of retrieving everything about one entity in a single read, with no join at all.
One-to-Few: Embedding a List of Addresses¶
A student who has lived at, at most, a handful of addresses is the classic "one-to-few" case — the related data is small, bounded, and almost always read together with its parent:
db.students.insertOne({
regNo: "FA21-BCS-045",
name: "Ayesha Khan",
addresses: [
{ type: "home", city: "Islamabad", street: "6 Lawrence St" },
{ type: "hostel", city: "Islamabad", street: "COMSATS Hostel Block C" }
]
});
One-to-Many: Embedding Enrolled Courses¶
Now consider a student's enrolled courses — the relationship Lecture 16 would model with a
Student table, a Course table, and an Enrollment junction table connected by foreign
keys. MongoDB can instead embed the enrollment details directly inside the student document:
db.students.insertOne({
regNo: "FA21-BCS-045",
name: "Ayesha Khan",
gpa: 3.72,
enrolledCourses: [
{ courseCode: "CSC270", title: "Database Systems", creditHours: 3, semester: "Fall 2025" },
{ courseCode: "CSC211", title: "Algorithms and Data Structures", creditHours: 3, semester: "Fall 2025" }
]
});
Relational approach (Lecture 16)
Student tableregNo, name, gpa
Enrollment tableregNo (FK), courseCode (FK), semester
Course tablecourseCode, title, creditHours
Reading a student's courses needs a two-table join
Document approach (this lecture)
One student documentenrolledCourses embedded directly as an array — no join needed to read it
Embedding is not the only option in MongoDB, and it is not always the right one. When
related data is large, unbounded, or needs to be queried independently of its "parent" — many
students sharing the same course, say — MongoDB documents can instead reference each other
by storing an ObjectId, much closer to a relational foreign key:
// Referencing instead of embedding -- closer to the relational foreign-key style
db.students.insertOne({
regNo: "FA21-BCS-045",
name: "Ayesha Khan",
enrolledCourseIds: [ ObjectId("66f2b1c0..."), ObjectId("66f2b1c1...") ]
});
Embed for data read together; reference for data queried independently
Compare this directly against Lecture 16 — Mapping EER Models to Relational Schemas: the relational model normalizes first and joins at query time; MongoDB asks you to design around your actual read pattern first — embed what you almost always read together, reference what you need to query, update, or share independently.
MongoDB CRUD Operations¶
MongoDB's four CRUD verbs map directly onto SQL's INSERT, SELECT, UPDATE, and DELETE —
the syntax simply speaks documents instead of rows.
Create¶
db.students.insertOne({
regNo: "FA21-BCS-046",
name: "Bilal Ahmed",
gpa: 3.15,
department: "Computer Science",
isActive: true
});
db.students.insertMany([
{ regNo: "FA21-BCS-047", name: "Sara Malik", gpa: 3.90, department: "Software Engineering", isActive: true },
{ regNo: "FA21-BCS-048", name: "Hamza Tariq", gpa: 2.85, department: "Computer Science", isActive: false }
]);
Read¶
// Every document in the collection
db.students.find();
// Documents matching an exact value
db.students.find({ department: "Computer Science" });
// Just the first match
db.students.findOne({ regNo: "FA21-BCS-045" });
Update¶
db.students.updateOne(
{ regNo: "FA21-BCS-045" },
{ $set: { gpa: 3.80 } }
);
db.students.updateMany(
{ isActive: false },
{ $set: { status: "alumni" } }
);
Delete¶
updateOne/deleteOne without a unique filter is a silent trap
updateOne and deleteOne act on the first document MongoDB happens to match — if your
filter isn't actually unique, you may not modify the document you meant to. Filter on a
genuinely unique field (like regNo, or _id) whenever you intend to touch exactly one
document, exactly the same discipline a relational WHERE clause on a primary key gives you.
MongoDB Querying¶
Beyond exact-match filters, MongoDB provides query operators — always written with a
leading $ — for comparisons and logical combinations:
// $gt: greater than
db.students.find({ gpa: { $gt: 3.5 } });
// $in: value is one of a given set
db.students.find({ department: { $in: ["Computer Science", "Software Engineering"] } });
// $and: every condition must hold
db.students.find({
$and: [
{ gpa: { $gte: 3.0 } },
{ isActive: true }
]
});
A second argument to find() is a projection — which fields to return, mirroring a
relational SELECT column_list instead of SELECT *:
// Return only name and gpa (and the always-included _id, suppressed here with 0)
db.students.find(
{ gpa: { $gt: 3.5 } },
{ name: 1, gpa: 1, _id: 0 }
);
| Operator | Meaning | Relational equivalent |
|---|---|---|
$eq |
Equal to | = |
$gt / $gte |
Greater than / or equal | > / >= |
$lt / $lte |
Less than / or equal | < / <= |
$in |
Value is in a given list | IN (...) |
$and / $or |
Combine conditions | AND / OR |
MongoDB Use Cases¶
Document databases like MongoDB fit particularly well where records naturally vary in shape or are usually read as one self-contained unit: content-management systems, product catalogs with wildly different attributes per category, user profiles, mobile and single-page application backends, and any system whose schema is still evolving quickly during active product development.
Key Takeaways¶
- MongoDB stores documents — flexible, JSON-like structures — grouped into collections,
served by the
mongodprocess; roughly: collection ≈ table, document ≈ row, but neither enforces a fixed shape. - MongoDB stores data as BSON, not plain text JSON, for richer data types (
ObjectId,Date, distinct integer/double types) and faster binary traversal. - Relationships can be modeled by embedding related data directly in the parent document
(best for data almost always read together) or by referencing another document's
ObjectId(closer to a relational foreign key, better for independently-queried data) — a genuinely different design question than the normalize-first relational approach from Lecture 16. - The CRUD verbs —
insertOne/insertMany,find/findOne,updateOne/updateMany,deleteOne/deleteMany— map directly onto SQL'sINSERT,SELECT,UPDATE,DELETE. - Query operators like
$gt,$in, and$and, plus projections, give MongoDB's query language most of the expressive power of a SQLWHEREandSELECTclause, in document form.
A single mongod on one machine is exactly the "one server" picture from Lecture 26's CAP
theorem discussion — fine until you need the availability or the capacity a single server
cannot give you. The next lecture shows how MongoDB actually achieves both. Continue to
Lecture 28 — MongoDB Sharding and Replication.