Comparing cache simulation
During my internship in 2025, I tried to compare QEMU and gem5, two famous computer architecture simulators, in order to understand the origin of the differences they produce in their statistics during cache simulation.
Why would we want to simulate cache ?
There are a lot of reasons someone would want to simulate the cache of a system :
- If you have ever attended a computer architecture lesson, you’re very likely to have heard about the cache optimization that allows an important speedup for some algorithms. To do so, you want to have statistics about your cache usage during execution (such has number of hits and miss).You could do so using the performance registers of your CPU while executing the benchmark. However, this doesn’t allow to change the cache size, associavity, structure, … and some CPUs may not have these registers. Thus, it is pretty common to simulate the cache accesses during execution.
- In the academic world, RISC-V has gained a lot of interest these last few years, because of its business model. Thus, a lot of studies (such as this internship) use RISC-V as the architecture for the experiment, but it is not easy to run RISC-V on real hardware, because of the rarity of RISC-V processors and systems, which requires simulation of the RISC-V architecture, which means that we can’t use our host performance registers. Another very important reason to use cache simulators in academic world is also to run experiment on the memory hierarchy, and compare the impacts of the changes made for them. For instance, my internship tutor, Julie Dumas, and some other people in the laboratory I was in are interested in cache coherency protocols, and all of their experiment requires them to write new protocol and check the performance results.
If you run the same architecture, there are a lot of solution, such as Cachegrind, but for other architecture, you will require complete architecture simulators that we will see next :
Our protagonists
There are multiple different computer architecture simulator online; among them, two got my interest during this internship, for different reasons. Let’s try to understand their pros and their cons.
gem5
gem5 is a simulator that is very popular in computer architecture academy, because it allows :
- great fidelity in the simulation : gem5 is trying to reproduce the most it can the real way a processor work (without simulating the RTL), by using
SimObjectswhich can represent cache, different CPU, memories, clocks, … - simple modification and usage of the framework : all thos
SimObjectsare written in C++, and then are bound together within a pyhon script (namedconfig) to create a specific system.
I’m mentioned two times previously SimObject: this is a superclass of all simulated object, which can be clocked (through ClockedObject), and can have ports which allows communication between the SimObject where one object is a Requestor and the other is the Responder :
The communication is based on an ACKnowledge system: when the requestor is requesting anything from the responder, three things can happen
- Either the responder can answer, and the requestor is also available: the responder ACK the request, and then gives the answer to the requestor, which ACK too the answer.
- Maybe the responder is not already available, and when the requestor will ask something, it will answer that it is not yet available. The requestor will then have to reschedule a “try again” answer a little bit later.
- The same thing can happen the other way : maybe when the responder has its answer, the requestor is not anymore available. The requestor will then have to reschedule a “try again” answer too, in order for the responder to send again its answer.
This is detailed more precisely within gem5 tutorial.
QEMU
QEMU on the other side is mainly focused on two goals : having a simulation that works, and its speed. To do so, one of its main feature is called Tiny Code Generation (TCG): it is a Dynamic Binary Translation (DBT) framework.
TCG uses an Intermediate Representation (IR) to allow executing any simulated Instruction Set Architecture (ISA) on any other host ISA, and it uses cache of previously translated Translation Blocks (sequence on instruction that finishes with jump/branches), to lower the overhead cost of translating too frequently :
On Figure 2, we can also notice that TCG has implemented a plugin system : it allows C code to be compiled into dynamic library, which will be loaded when QEMU starts, and can register hooks to different part of QEMU internals, such as :
- start/stop/reset of QEMU
- Translation of blocks : depending of the translated instruction, we can hook different callbacks
- Executing an instruction : when an instruction is executed, we can intercept it, and intercept its memory access
QEMU has a few integrated plugins, such as limiter of Instruction-per-Second, git bisect like execution, and the one which will gain our interest : cache simulation.
This previous cache simulation plugin is great, but is has a very important slowdown due to the very frequent callback which are pretty slow with the cache simulation. Marie Badaroux proposed in 2023 another way to simulate cache, which compute cache interaction during the translation of the blocks, instead of computing them on each executed instruction.
Simulation environment
Both of these simulators can run in quite different ways :
- Bare Metal : we can compile our benchmarks to run directly on the hardware without any abstraction. We must include in the binary a
crt0to initialize hardware, memory and registers. - Full System simulation : we must compile a linux kernel with some basic tools within the disk image (we chose busybox), and add our benchmarks compiled for linux which will run on it too.
- Syscall Emulation : instead of simulating the whole OS to run the linux benchmark, we can intercept the syscall the benchmark does to interface it with the kernel that is running on the host, in order to improve performances
On these simulators, we chose to run two different benchmarks :
- Embench-iot is a small, minimal requirement benchmark, made for embedded platforms. Thus, it doesn’t require any stdlib, and when all benchmark are linked together, it is in text and in data/rodata/bss, with strong spacial locality
- PARSEC is a benchmark made especially for shared memory system (which is usefull for manycore systems). It is built for linux, and each benchmark is in text and in data/rodata/bss.
Already found origin of differences
This article is not yet finished. Coming soon :p