Author: Eric Kim ([email protected])

Reference kernel version: 7.2-rc kernels(The mainline ones)

The goal of this document is to describe how common paths for memory accessing works in a kernel. It describes each data structure’s general members and few general case of its use. Additionally, it describes general execution path of page fault and other memory faults/allocation helpers. Mostly, the code referenced is about memory allocation.

Note that this is merely a note I took while reading linux kernel mm source, so some info may be inaccurate. Please cross-reference with sources listed below:

Also, This is mostly based on x86_64, sizes and design may vary across architectures.

Memory model for dummies

This section explains memory abstraction easy as possible, and go through advanced features. If you know nothing about memory(but know at least python or C), after reading this, you would somewhat understand the rest of the document.

If one simplify the computer far enough, computers become only the CPU and memory. In essence, cpu requests address to memory, and memory returns data in that address. CPU would deal with datas in that memory. The abstraction makes computers look simple like this on all programming languages, high and low-level, and even assembly. Request a data in an address, and it would fetch without explicit fetch implementation. After virtual memory is enabled, all addresses are regarded as virtual, and cpu will go through page tables to find the data. But actually, no. modern CPUs have cache that holds recently used data, and CPU would not go to memory unless the requested page is not in cache(cache miss).

For each task, page tables exist as a map to virtual address space of that task. page tables are what is used to translate virtual address to physical address. The page tables are, just like any other data, in memory. Compared to cache, accessing page table in memory would be much slower, and this is a problem because when virtual memory is enabled, it would mean having to access relatively slow memory at every access. TLB comes to play to solve this. TLB stores recent virtual-to-physical address translations of current task. If the translation is not on TLB(TLB miss), it would walk through the page table to find the target page. When changing tasks, TLB is flushed as it holds previous address translations.

More techy stuff for one who read the previous

physical memory is divided into nodes in NUMA computers(with most computers, it is single node, but enabled nonetheless), and inside the nodes, it is divided into zones. Most memory space are pages from ZONE_NORMAL. Each zone’s pages are managed by buddy allocator, which allocates sizes of physical memory by power of two size. It manages 4KB, 8KB, 16KB size blocks up to 4MB(on x86_64), and each has list of available blocks. If none available, it allocates from bigger blocks by splitting it. Allocated pages could be used for user space and kernel.

Pages are managed with pte or in case of hugepages pmd, with multiple levels of directories, with top directory being page global directory. VMA(Virtual Memory Area) defines a certain range of memory addresses that are mapped to physical memory, and multiple VMAs are formed to create the full memory mappings of a process. Each process has different memory mappings to implement virtual memory. Kernel obtains memory objects to use from slub, which will be explained later. User requested memory is reserved on page table after adding in valid memory area(VMA) to task’s kernel object and returning the address of that area,and are actually allocated with page faults hinting accessing valid pages(in current task context), by getting a page, making the address valid by adding the pte entry to page table, and returning accessed data from user.

mm_struct(collection of VMA and other descriptors) is just a descriptor for kernel’s virtual address space management, and page tables can be accessed with pgd(page global directory) object in mm_struct. pgd has multiple hierarchies of entries with pte at the end, in the order pgd→pud→pmd→pte. PTE accounts for 4KiB page, and pmd, pud, pgd, each stores 512 entries of each child entries. So pmd can hold 2MiB, pud 1GiB, and pgd 512GiB.

The philosophy of linux mm is lazy allocating(In fact most modern computing practices are done lazily). The actual memory is allocated when it is used; not when malloc() or mmap(), when you access the pointer of that address. Each process can map large portions of virtual memory address, even exceeding physical memory capacity, as it would not allocate actual pages right away. Page fault exception path is how kernel deals with allocation. If the physical address on pte is invalid, the CPU starts executing page fault exception path, which would figure out the nature of the page fault using the address accessed and such, and handle it the proper way.

Physical Memory

Page(struct page)