A patch that compiles is not a patch that works. This is a statement that most kernel developers would find obvious, yet the majority of automated code generation tools treat compilation success as their primary — and often only — measure of correctness. In the context of kernel driver migration, this gap between "compiles" and "correct" is not a minor inconvenience. It is the difference between a driver that loads and a driver that corrupts memory under load.
When KDRIFT generates a migration patch, that patch must survive four independent stages of validation before we consider it successful. Each stage catches a distinct class of bugs that the previous stages cannot detect. The attrition rate at each stage tells us something important about the nature of kernel code correctness.
This post walks through each stage, explains what it catches, and shows why skipping any one of them means shipping patches that pass the easiest test and hoping for the best.
The first stage is the one everyone knows. Run make with the appropriate defconfig and check the exit code. Compilation catches the obvious problems: syntax errors, missing symbols, type mismatches between function arguments and parameters, undeclared identifiers, and incompatible pointer types.
For a migration patch, compilation failure usually means the patch generator got the API signature wrong. Maybe it used the old function name, passed the wrong number of arguments, or forgot that a struct field was renamed. These are the low-hanging fruit of migration errors, and any competent code generation system should handle them.
But here is the critical point: compilation verifies syntax, not semantics. The compiler confirms that your code is well-formed according to the C language specification. It does not confirm that your code does what you intend. It does not check that you are holding the right lock when accessing shared data. It does not verify that your DMA buffer lifetime matches the hardware's access pattern. It does not know that the function you're calling now has different error return semantics than it did in the previous kernel version.
A patch that compiles cleanly is a patch that has cleared the lowest bar. In our data, 100% of patches that reach Stage 2 compiled successfully — that is the entry condition. What happens next is where things get interesting.
Sparse is a semantic checker for the Linux kernel, originally written by Linus Torvalds. Smatch is a static analysis tool built on top of Sparse that performs additional data flow analysis. Together, they catch a class of bugs that the compiler cannot see: lock ordering violations, null pointer dereferences, type confusion between kernel and user space pointers, and incorrect bitwise operations on endian-sensitive data.
The canonical example is __user annotation violations. The kernel uses type annotations to distinguish between pointers to user-space memory and pointers to kernel-space memory. The compiler treats both as plain pointers — they're the same type at the C level. But mixing them is a security vulnerability. Sparse catches this.
Consider a migration where an ioctl handler is updated. The old API passed a kernel-space buffer; the new API passes a __user pointer that must be copied with copy_from_user(). A naive migration might update the function signature but forget to add the copy:
This compiles. The types are compatible at the C level. But Sparse flags it immediately: dereference of noderef expression. The correct patch would use copy_from_user() or get_user() to safely transfer the data. Without Sparse, this patch ships with a kernel information leak and a potential arbitrary write.
Smatch goes further. It tracks data flow across function boundaries to find null pointer dereferences after failed allocations, integer overflow in size calculations passed to kmalloc, and cases where error paths don't release resources. For migration patches, Smatch is particularly valuable when API changes alter error handling semantics — the kind of change where the return value type stays the same but its meaning shifts.
In our benchmark, 11% of patches that compile cleanly are flagged by Sparse or Smatch. These are patches that would pass any test based purely on compilation. They would be deployed, and they would contain latent bugs ranging from information leaks to memory corruption.
Static analysis reasons about code without executing it. Stage 3 does the opposite: it builds a complete kernel with the patched driver and boots it in QEMU. This catches a class of bugs that no static tool can detect — driver probe failures, resource conflicts, init ordering bugs, and device tree binding mismatches.
The setup is straightforward but carefully controlled. For each migration case in DRIFTBENCH, we maintain a minimal defconfig that enables the target driver and its dependencies. The kernel is compiled, loaded into a QEMU virtual machine with the appropriate machine type, and booted with an automated timeout. We capture the full dmesg output and parse it for errors.
What we're looking for at this stage is not subtle. We're looking for hard failures: the driver fails to probe, the kernel panics during init, a resource allocation fails because the new API requires a different registration order, or the module fails to load because a symbol dependency changed.
These failures are common in migrations that involve subsystem registration changes. For example, when the platform driver registration API changes how it handles deferred probing, a patch that updates the function call but doesn't adjust the probe return handling will compile fine, pass Sparse, and then fail to load the driver in a running kernel.
The QEMU stage also catches build system issues that don't manifest as compilation errors. A Kconfig dependency that changed between kernel versions can cause a driver to be silently excluded from the build. The compilation succeeds (the driver simply isn't compiled), but the QEMU boot reveals that the module is missing.
Stage 3 eliminates an additional 11% of patches — patches that compiled, passed static analysis, and still contained bugs that only manifest at runtime.
The final stage subjects the booted kernel to synthetic workload with dynamic analysis enabled. We enable CONFIG_PROVE_LOCKING (lockdep) and CONFIG_KASAN (Kernel Address Sanitizer) and then exercise the driver through its primary code paths.
Lockdep is the kernel's lock dependency validator. It constructs a graph of all lock acquisitions at runtime and checks for potential deadlocks — not just actual deadlocks, but lock ordering violations that could deadlock under different scheduling conditions. A migration patch that changes lock acquisition order (because the new API requires calling a function that internally takes a lock) may never deadlock in testing but will deadlock in production under memory pressure.
KASAN detects memory corruption at runtime: use-after-free, out-of-bounds access, double-free, and memory leaks. These bugs are particularly common in migrations that change object lifetime semantics. If the old API allocated and freed a buffer internally, but the new API requires the caller to manage the buffer, a patch that doesn't add the corresponding kfree() will leak memory on every call. Under sustained load, the system runs out of memory.
These are the bugs that only manifest under specific conditions — certain code paths, certain timing, certain workloads. A driver that boots cleanly and handles a single ioctl call correctly may corrupt memory on the thousandth call, or deadlock when two threads access it simultaneously. Without dynamic analysis under load, these bugs ship to production.
Stage 4 catches an additional 22% of patches that survived all previous stages. These are patches that compile, pass static analysis, boot in a VM, and still contain correctness bugs.
Across DRIFTBENCH, here is the attrition at each stage when evaluating migration patches from frontier LLMs:
Read that from right to left: 44% of patches that compile correctly contain bugs that are only caught by subsequent validation stages. Nearly half. If your migration tool only checks compilation, you are shipping patches with a coin-flip chance of containing latent bugs.
The distribution across stages is also informative. Stage 2 (static analysis) and Stage 3 (boot testing) each catch roughly the same proportion of bugs — about 11% of the total each. Stage 4 (dynamic analysis) catches twice as many as either of them alone. This suggests that the hardest bugs to detect are behavioral, not structural. They are about what the code does at runtime, not how it looks to a static tool.
This is consistent with our understanding of kernel API evolution. The changes that are hardest to migrate are not the ones that rename functions or change argument types. Those are easy — the compiler catches them immediately. The hard changes are the ones that alter behavioral contracts: error handling semantics, locking requirements, object lifetime expectations. These are the changes that pass compilation and static analysis but fail under load.
Staged validation is not optional. It is not a nice-to-have quality improvement. It is the minimum bar for generating kernel patches that you can trust. Without it, you are selecting for patches that pass the easiest test — compilation — and ignoring the 44% of cases where that test is insufficient.
The four stages form a funnel of increasing specificity. Compilation checks syntax. Static analysis checks local semantics. Boot testing checks system-level integration. Dynamic analysis checks behavioral correctness under load. Each stage is necessary because each checks something the others cannot.
When we evaluate KDRIFT's patch quality, we report the pass rate at the final stage, not the first. When we say 56% success, we mean 56% of patches survive compilation, static analysis, boot testing, and dynamic analysis. The patches that make it through are patches you can review with confidence that the mechanical correctness has been verified.
The remaining 44% are not failures. They are starting points — partial patches that human reviewers can refine with the knowledge of exactly which stage failed and why. That is a fundamentally different workflow than writing a migration patch from scratch.