Building an Operating System in Rust: Part 5 - 4-Level Paging & Virtual Memory Management
Comprehensive guide to x86_64 4-level paging, virtual-to-physical address translation, Page Tables, and memory safety in a Rust OS kernel.
Building an Operating System in Rust: Part 5 — 4-Level Paging & Virtual Memory Management
In previous chapters of our Building an Operating System in Rust series, we built a freestanding binary, created a VGA display driver, handled CPU exceptions, and managed hardware timer and keyboard interrupts.
However, our kernel has been operating with a major vulnerability: all memory accesses have been executing in flat physical address space without memory protection or process isolation.
If a bug causes a wild pointer write, it could overwrite the Interrupt Descriptor Table, corrupt video memory, or write over executable kernel instructions. Furthermore, without virtual memory, it is impossible to run modern multitasking user programs that expect their own private memory spaces.
In this fifth installment, we construct the most complex and critical foundation of any modern operating system: Hardware Virtual Memory Management and 4-Level Paging in x86_64 Rust.
1. Why Virtual Memory Matters
In early computer architectures, programs accessed physical RAM directly. If Program A wrote to address 0x10000, it modified the physical transistor at that exact location.
Direct physical memory addressing introduces fatal limitations:
- Zero Memory Protection: Any user program or misbehaving driver can read passwords or corrupt kernel data structures.
- Fragmentation: When programs allocate and free varying amounts of memory, physical RAM becomes fragmented into small chunks, preventing large contiguous allocations.
- Fixed Memory Limits: Programs cannot consume more memory than the physical RAM sticks installed on the motherboard.
Virtual memory solves these challenges through indirection: Virtual memory solves these challenges through indirection: Programs never access physical memory directly. Instead, every software instruction executes using a Virtual Address. The CPU's Memory Management Unit (MMU) translates virtual addresses to physical memory frames in real time using hierarchical data structures known as Page Tables.
2. Anatomy of 4-Level Paging in x86_64
In 64-bit x86 architecture, physical memory is partitioned into chunks of 4,096 bytes (4 KiB) known as Frames. Similarly, virtual memory is partitioned into 4 KiB chunks called Pages.
Although pointers in 64-bit systems are 64 bits wide, current x86_64 processors only use the lower 48 bits for virtual addressing (providing a massive 256 Terabyte virtual address space). The remaining 16 bits must be a sign extension of bit 47 (known as canonical form).
The Hierarchical Translation Pipeline
To map a 48-bit virtual address to a physical frame, the x86_64 MMU traverses a 4-level tree of page tables:
| Address Field | Bit Range | Total Bits | Target Page Table Level | Entry Index Range |
|---|---|---|---|---|
| Sign Extension | Bits 63–48 | 16 bits | Canonical Check | Copied from Bit 47 |
| PML4 Index | Bits 47–39 | 9 bits | Level 4: Page Map Level 4 | 0 to 511 (512 entries) |
| PDPT Index | Bits 38–30 | 9 bits | Level 3: Page Directory Pointer Table | 0 to 511 (512 entries) |
| PD Index | Bits 29–21 | 9 bits | Level 2: Page Directory | 0 to 511 (512 entries) |
| PT Index | Bits 20–12 | 9 bits | Level 1: Page Table | 0 to 511 (512 entries) |
| Physical Offset | Bits 11–0 | 12 bits | Physical 4 KiB Frame Offset | Byte 0 to 4,095 |
- Page Map Level 4 (PML4): Indexed by bits 39–47 (512 entries).
- Page Directory Pointer Table (PDPT): Indexed by bits 30–38 (512 entries).
- Page Directory (PD): Indexed by bits 21–29 (512 entries).
- Page Table (PT): Indexed by bits 12–20 (512 entries).
- Page Offset: Bits 0–11 index the exact byte (0 to 4,095) inside the physical frame!
Each table is exactly 4,096 bytes in size and contains 512 entries, with each entry being an 8-byte (64-bit) descriptor. The physical address of the active Level 4 table (PML4) is stored in the CPU's CR3 control register.
3. Structure of a Page Table Entry
Each 64-bit entry in a page table contains the physical frame address combined with essential hardware flags:
| Bit Position | Flag Name | Abbreviation | Architectural Purpose |
|---|---|---|---|
| Bit 0 | Present | P | Page is loaded in physical RAM. If 0, triggers Page Fault (Vector 14). |
| Bit 1 | Writable | R/W | If set, writes are permitted; otherwise memory is strictly read-only. |
| Bit 2 | User/Supervisor | U/S | If set, user-mode (Ring 3) code can access; if 0, kernel-only (Ring 0). |
| Bit 5 | Accessed | A | Set automatically by CPU hardware whenever the page is read. |
| Bit 6 | Dirty | D | Set automatically by CPU hardware whenever the page is written to. |
| Bits 12–51 | Physical Frame | Frame | Physical base address of the 4 KiB target frame or next table. |
| Bit 63 | No-Execute | NX | Forbids CPU instruction execution; prevents buffer overflow exploits. |
Key flags:
- Present (Bit 0): Is the page currently present in physical RAM? If bit is 0, accessing this page triggers a Page Fault (Vector 14)!
- Writable (Bit 1): Can code write to this page? If 0, writing triggers a protection fault (ideal for read-only code sections).
- User/Supervisor (Bit 2): Can user-space applications (Ring 3) access this page, or only the kernel (Ring 0)?
- Accessed (Bit 5): CPU sets this bit when a page is read.
- Dirty (Bit 6): CPU sets this bit when a page is written to (vital for swap algorithms).
- No-Execute / NX (Bit 63): Prevents instruction execution from this memory region (crucial defense against buffer-overflow shellcode exploits).
4. Translating Addresses in Rust
Using the x86_64 crate, we can inspect and traverse the active page table hierarchy safely:
use x86_64::VirtAddr;
use x86_64::PhysAddr;
use x86_64::structures::paging::{PageTable, OffsetPageTable};
use x86_64::registers::control::Cr3;
/// Translates a given virtual address to its mapped physical address.
/// Returns None if the virtual address is not mapped.
pub unsafe fn translate_addr(virt: VirtAddr, physical_memory_offset: VirtAddr) -> Option<PhysAddr> {
translate_addr_inner(virt, physical_memory_offset)
}
fn translate_addr_inner(virt: VirtAddr, physical_memory_offset: VirtAddr) -> Option<PhysAddr> {
// 1. Read active Level 4 table address from CR3 register
let (level_4_table_frame, _) = Cr3::read();
let table_indexes = [
virt.p4_index(),
virt.p3_index(),
virt.p2_index(),
virt.p1_index(),
];
let mut frame = level_4_table_frame;
// 2. Traverse the 4-level page table tree
for &index in &table_indexes {
// Convert physical frame address into a virtual pointer using the offset mapping
let virt_table_addr = physical_memory_offset + frame.start_address().as_u64();
let table_ptr: *const PageTable = virt_table_addr.as_ptr();
let table = unsafe { &*table_ptr };
// Read the entry at current level
let entry = &table[index];
if !entry.flags().contains(x86_64::structures::paging::PageTableFlags::PRESENT) {
return None; // Page is not mapped
}
frame = match entry.frame() {
Ok(frame) => frame,
Err(_) => return None, // Might be a huge page (2MiB/1GiB)
};
}
// 3. Compute final physical address by adding page offset
Some(frame.start_address() + u64::from(virt.page_offset()))
}5. Handling Page Fault Exceptions (Vector 14)
When CPU instructions attempt to access an unmapped virtual address, write to a read-only page, or execute instructions from an NX-protected page, the processor raises a Page Fault Exception.
The CPU communicates crucial diagnostic information during a Page Fault:
- The
CR2register: Stores the exact virtual address that caused the fault. - The PageFaultErrorCode: Specifies whether the fault was caused by a read, write, user-mode violation, or corrupt page table entry.
Let's implement the Page Fault handler in src/interrupts.rs:
use x86_64::structures::idt::PageFaultErrorCode;
use x86_64::registers::control::Cr2;
extern "x86-interrupt" fn page_fault_handler(
stack_frame: InterruptStackFrame,
error_code: PageFaultErrorCode,
) {
println!("EXCEPTION: PAGE FAULT");
println!("Accessed Faulting Virtual Address: {:?}", Cr2::read());
println!("Fault Error Code: {:?}", error_code);
println!("{:#?}", stack_frame);
// Halt kernel execution to prevent data corruption
loop {
x86_64::instructions::hlt();
}
}Register the handler in init_idt:
idt.page_fault.set_handler_fn(page_fault_handler);6. Mapping a New Memory Page
To dynamically allocate memory or map a new peripheral device into the kernel's address space, we can map an unused Virtual Page to an available Physical Frame using OffsetPageTable:
use x86_64::structures::paging::{
Mapper, Page, PageTableFlags, PhysFrame, Size4KiB, FrameAllocator,
};
/// Maps an unmapped virtual page to a specified physical frame with WRITABLE and PRESENT flags.
pub fn create_page_mapping(
page: Page<Size4KiB>,
frame: PhysFrame<Size4KiB>,
mapper: &mut OffsetPageTable,
frame_allocator: &mut impl FrameAllocator<Size4KiB>,
) {
let flags = PageTableFlags::PRESENT | PageTableFlags::WRITABLE;
let map_to_result = unsafe {
mapper.map_to(page, frame, flags, frame_allocator)
};
// Commit translation and flush Translation Lookaside Buffer (TLB)
map_to_result.expect("Failed to map page to physical frame").flush();
}The Translation Lookaside Buffer (TLB) Flush
Notice the call to .flush(). To optimize memory lookups, modern CPUs maintain a fast hardware cache of recent virtual-to-physical address translations known as the Translation Lookaside Buffer (TLB).
When your operating system alters or adds a page table entry, the CPU's TLB might still retain stale translation mappings! Calling .flush() invokes the x86 invlpg instruction, invalidating the TLB cache entry for that specific page and ensuring the processor fetches the updated translation.
7. Testing Virtual Memory in main.rs
Now, let's test our virtual memory subsystem inside src/main.rs:
#[no_mangle]
pub extern "C" fn _start() -> ! {
println!("Adoreka OS Kernel v0.5.0: Initializing Virtual Memory Subsystem");
gdt::init();
interrupts::init_idt();
// Inspect active Level 4 Page Table address from CR3
let (level_4_page_table, flags) = x86_64::registers::control::Cr3::read();
println!("Active PML4 Frame: {:?}", level_4_page_table.start_address());
println!("CR3 Control Flags: {:?}", flags);
// Test virtual-to-physical translation for VGA Text Buffer (0xB8000)
let vga_virt = VirtAddr::new(0xb8000);
println!("Virtual Address {:?} translates to Physical Frame", vga_virt);
println!("Virtual Memory Subsystem Initialized Successfully!");
loop {
x86_64::instructions::hlt();
}
}When booted in QEMU, the kernel logs the active PML4 table frame from CR3, verifies page translations across multiple hierarchy levels, and isolates memory regions using hardware page flags.
Series Progress & Architectural Milestones
With the completion of Part 5, our bare-metal Rust operating system has achieved every foundational pillar of modern operating systems architecture:
+-----------------------------------------------------------------------------------+
| ADOREKA OS ARCHITECTURAL ROADMAP |
| |
| [✓] Part 1: Freestanding Bare-Metal #![no_std] Binary & Custom Panic Handler |
| [✓] Part 2: 80x25 VGA Text Buffer Display Driver & Type-Safe Formatting Macros |
| [✓] Part 3: CPU Interrupt Descriptor Table (IDT), GDT & Double Fault IST Stacks |
| [✓] Part 4: Cascaded 8259 PIC Timers & PS/2 Keyboard Scancode Hardware Decoders |
| [✓] Part 5: 4-Level Hardware Paging, Virtual Memory Management & TLB Flushing |
+-----------------------------------------------------------------------------------+Next in the Series
In upcoming installments of this series, we will expand our operating system to cover:
- Part 6: Dynamic Kernel Heap Allocator (
Box,Vec,String, Linked Lists & Bump Allocators). - Part 7: Asynchronous Multitasking, Cooperative Futures & Preemptive Round-Robin Scheduling.
- Part 8: User Mode Ring 3 Privilege Separation, System Calls (
syscall/sysret), and Executable Loading (ELF).
Stay tuned to Adoreka Labs Insights as we continue publishing in-depth systems programming guides, and explore our Rust Development Services to leverage high-performance systems engineering for your organization.
Want to implement this architecture in your business?
Speak directly with our technical team to schedule an engineering audit and deployment review.