24. Paging: Method and Hardware Support¶
Lecture 23 ended on segmentation's unresolved problem: letting memory pieces vary in size, to match how a programmer thinks about a program, is exactly what produces external fragmentation — scattered holes, individually too small, with no single block big enough for the next request. Paging sidesteps that problem with a deceptively simple move: stop letting pieces vary in size at all. Every piece of memory a process can be given is the exact same fixed size, which means any free piece can satisfy any request. This lecture works through paging's basic method, the hardware — the page table, the frame table, and the Translation Look-aside Buffer — that makes it fast enough to use on every single memory access, and the protection it provides almost for free.
In This Lecture¶
- Paging's basic method — fixed-size frames and pages — and exactly why it eliminates external fragmentation entirely
- The page table and a complete, numeric address-translation walkthrough, bit by bit
- The system-wide frame table
- The Translation Look-aside Buffer (TLB) and a worked Effective Access Time (EAT) calculation
- Protection via the valid-invalid bit and read/write/execute permission bits
- A brief look at hashed and inverted page tables for large address spaces
Paging: Basic Method¶
Paging divides physical memory into fixed-size blocks called frames, and divides a process's logical memory into blocks of the exact same fixed size, called pages. Any free frame, anywhere in physical memory, can hold any page of any process — there is no requirement that a process's pages sit next to each other in physical memory at all.
Paging — logical pages scattered freely across physical frames
Process logical memoryPage 0, Page 1, Page 2, Page 3 — contiguous in the process's own view
Physical memoryFrame 5, Frame 2, Frame 9, Frame 1 — scattered wherever free frames happen to be
Because any free frame fits any page, external fragmentation disappears completely — there is never a situation where "enough total free memory exists but no single piece is big enough," since every piece is the same fixed size and a request for n pages only ever needs n free frames, wherever they happen to be. The cost paid in exchange is a (much smaller) return of internal fragmentation: a process's memory requirement is essentially never an exact multiple of the page size, so its last page is typically only partially used, and whatever's left over inside that one final frame is wasted — but bounded by at most one page's worth, per process, rather than the unpredictable waste segmentation and dynamic partitioning could produce.
The Page Table¶
Each process gets its own page table, a per-process structure mapping page number → frame number. A logical address generated by the CPU is split into exactly two fields:
- The page number — used as an index into this process's page table, yielding the frame number that page currently lives in.
- The offset — the position within that page/frame, which (crucially) is identical whether you're looking at the logical address or the final physical address, since a page and its frame are the same fixed size.
Address Translation: A Worked Example¶
Take a page size of 4 KB = 4096 bytes = 2¹², so the offset field is 12 bits. Suppose
this process's logical address space is 64 KB = 2¹⁶ bytes — a 16-bit logical address —
which leaves 16 − 12 = 4 bits for the page number, i.e. up to 16 pages per process
(numbered 0–15).
Splitting a 16-bit logical address: 4-bit page number + 12-bit offset
Page number — 4 bits 0001 = decimal 1
Offset — 12 bits 0011 1000 1000 = decimal 904
Take the logical address 5000 (decimal). In 16-bit binary:
Split it: the top 4 bits, 0001, give page number = 1. The bottom 12 bits,
0011 1000 1000, give offset = 904. Check the arithmetic directly: 1 × 4096 + 904 =
4096 + 904 = 5000 ✓ — matches the original logical address exactly.
Now look up page 1 in this process's page table:
| Page number | Frame number |
|---|---|
| 0 | 5 |
| 1 | 2 |
| 2 | 9 |
| 3 | 1 |
Page 1 maps to frame 2. The physical address is formed by concatenating the frame number with the same offset computed above:
In binary, with frame numbers also occupying 4 bits in this example system: frame 2 =
0010, offset unchanged at 0011 1000 1000, giving physical address 0010 0011 1000 1000 =
9096 decimal — 2 × 4096 + 904 = 9096 ✓, consistent both ways.
Translation summary
Logical address 5000= page 1, offset 904
Page table lookupPage 1 → Frame 2
Physical address 9096= frame 2 × 4096 + offset 904
The offset never changes — only the page/frame field does
This is the entire trick that makes paging's translation so fast: the low-order bits (the offset) pass straight through, untouched. Only the high-order bits (the page number) need an actual table lookup, to be replaced by the frame number. Everything below the page-size boundary is already in the right place.
The Frame Table¶
While each process has its own private page table, the operating system also maintains exactly one system-wide frame table — tracking, for every physical frame in the machine, whether it is currently free or allocated, and if allocated, which process and which page of that process currently occupies it. The frame table is what the OS consults whenever any process needs a new frame (to find a free one) and whenever a frame needs to be reclaimed (to know who to notify that it's being taken away).
| Frame | Status | Owner (process, page) |
|---|---|---|
| 0 | Free | — |
| 1 | Allocated | P7, page 3 |
| 2 | Allocated | P4, page 1 |
| 3 | Free | — |
| 4 | Allocated | P7, page 0 |
Hardware Support: The Translation Look-aside Buffer (TLB)¶
A naive implementation of paging doubles every memory access: one access to read the page table entry, and a second to actually fetch the data — and the page table itself usually lives in main memory, making even that first lookup no faster than an ordinary memory reference. The Translation Look-aside Buffer (TLB) fixes this with a small, very fast piece of associative-memory hardware that caches the most recently used page-number → frame-number translations, checked before the full page table on every single memory access.
- TLB hit — the needed translation is already cached; the frame number comes back almost immediately, and the full page table is never even touched.
- TLB miss — the translation isn't cached; the CPU must walk the real page table in memory to find it (one extra memory access), and then cache that result in the TLB before finally accessing the actual data.
Effective Access Time (EAT)¶
Because a hit and a miss cost different amounts of time, and most memory accesses are a random mix of both, we describe overall average cost with the Effective Access Time:
The 2 × memory time on a miss accounts for the extra memory access needed to walk the
real page table, in addition to the memory access that actually fetches the data once the
frame number is known.
Worked example. Suppose TLB access takes 10 ns, a main-memory access takes 100 ns, and the TLB hit ratio is 80% (0.8), so the miss ratio is 20% (0.2):
EAT = 0.8 × (10 + 100) + 0.2 × (10 + 2 × 100)
= 0.8 × 110 + 0.2 × (10 + 200)
= 88 + 0.2 × 210
= 88 + 42
= 130 ns
The Effective Access Time is 130 ns — noticeably more than a single 100 ns memory access on its own (reflecting the real overhead paging adds), but far closer to that single-access cost than the 210 ns a miss alone would require, precisely because 80% of the time the TLB hit avoids the extra page-table walk altogether.
Hit ratio is the whole game
Every term in this formula is fixed by the hardware except the hit ratio, which depends on program behavior (how often nearby memory references reuse the same pages — their locality). A higher hit ratio drives EAT closer to a single memory access; a lower one drives it closer to double that. This is exactly why real programs' performance is so sensitive to how well their memory-access patterns exhibit locality.
Protection in Paging¶
Each entry in a page table carries more than just a frame number — it carries the bits that make paging safe to use at all.
- Valid-invalid bit. Marks whether the page is actually part of the process's legitimate logical address space right now. A process's logical address space is nearly always smaller than what its page number field could theoretically address (a 4-bit page number can address 16 pages even if the process only actually uses 6) — every unused entry is marked invalid. Accessing a page marked invalid traps to the operating system immediately, which is exactly the hardware-enforced mechanism that catches a process trying to touch memory outside its own address space.
- Read/write/execute permission bits. Independently of validity, each entry can restrict what kind of access is allowed — a code page marked read/execute but not write (stopping a process from overwriting its own instructions, deliberately or by a bug), a data page marked read/write but not execute (stopping injected data from ever being run as code).
| Page | Frame | Valid | Read | Write | Execute |
|---|---|---|---|---|---|
| 0 (code) | 5 | ✓ | ✓ | ✗ | ✓ |
| 1 (heap) | 2 | ✓ | ✓ | ✓ | ✗ |
| 2 (unused) | — | ✗ | — | — | — |
Advanced: Hashed and Inverted Page Tables¶
A straightforward, linear, one-entry-per-page page table works well while a process's address space is modest — but on a modern 64-bit system, a process's potential address space is enormous, and a linear page table covering all of it, even mostly empty, would itself demand an impractical amount of memory just to exist.
- Hashed page tables. Instead of indexing directly by page number into a (potentially huge) linear array, the page number is hashed into a much smaller table. Each hashed entry holds the page number itself (to resolve collisions, since different pages can hash to the same bucket) and its frame number, chained to the next entry on a collision. This handles large, sparse address spaces efficiently — an address space is sparse when the pages actually in use are spread thinly across a vast potential range, exactly the common case for large modern programs.
- Inverted page tables. The more radical restructuring: instead of one table per process with one entry per page, keep exactly one table, system-wide, with one entry per physical frame. Each entry records which process and which page currently occupies that frame. This saves an enormous amount of memory on a 64-bit system — the table's size is bounded by the number of physical frames that actually exist, not by the size of every process's virtual address space combined — but it comes at the cost of a slower lookup: finding the entry for a given (process, page) pair now means searching (typically via a hash, same idea as above) through the whole frame-indexed table, rather than a single direct index.
Same trade-off, different scale
Both techniques trade some lookup speed for a dramatic reduction in page-table memory — the right choice whenever a process's address space is far larger, and far sparser, than its actual memory usage, which describes almost every modern 64-bit process.
Key Takeaways¶
- Paging divides physical memory into fixed-size frames and logical memory into same-size pages — any free frame fits any page, which eliminates external fragmentation entirely, at the cost of (bounded) internal fragmentation in a process's last page.
- A page table maps page number → frame number; translating a logical address means splitting it into a page-number field and an offset field, looking up the frame number, and concatenating frame number + (unchanged) offset into the physical address.
- The frame table is one system-wide structure tracking every physical frame's allocation status and current owner.
- The TLB caches recent translations and is checked before the full page table; EAT blends the hit cost and the (more expensive, extra-memory-access) miss cost, weighted by the hit ratio — in the worked example, an 80% hit ratio with 10 ns TLB and 100 ns memory access gave an EAT of 130 ns.
- Protection comes from the valid-invalid bit (catches access to pages outside the process's real address space) and read/write/execute permission bits per entry.
- Hashed page tables handle large, sparse address spaces efficiently by hashing the page number; inverted page tables keep one entry per physical frame, system-wide, trading slower lookups for dramatically less memory on large 64-bit systems.
Paging, as presented here, assumes every one of a process's pages is loaded into a frame before the process runs. The next lecture relaxes that assumption entirely — loading pages only as they are actually needed — and asks what happens when physical memory is smaller than every process's combined demand for it. Continue to Lecture 25 — Virtual Memory and Demand Paging.