Aarno Labs Logo

Aarno Labs Blog

The latest news and research from Aarno Labs

Verified Patching of Predicated Code Without a Trampoline: A Fix for an Infusion Pump

Author: Michael Gordon

13 min read

Posted 1 hour, 20 minutes ago

An infusion pump on a hospital network runs a copy of BusyBox that was already three years behind upstream on the day the firmware was built. One of the fixes it is missing is for CVE-2016-2147, an out-of-bounds write in the DHCP client that any machine on the same network segment can trigger. We patched it directly in the stripped binary, and this post is about what made that patch unusual: the vulnerable code compiles to ARM predicated instructions rather than branches, and our fix is applied in place, with no trampoline, and then proven safe by CodeHawk's binary analysis rather than assumed to be.

The first half of the post is about why predicated code is a problem for both analysis and patching. The second half is about the mechanism we built to handle predicated instructions, and the evidence CodeHawk produces to demonstrate the correctness of the fix.

The vulnerability

BusyBox's DHCP client, udhcpc, accepts a domain search option (option 119) from the DHCP server. The option carries a list of domain names in the compressed wire format of RFC 1035, and dname_dec in networking/udhcp/domain_codec.c expands it into a space-separated string. The function makes two passes over the input: the first computes the output length, and the second, after allocating a buffer, fills it in.

The relevant piece is the handling of a NUL byte, which marks the end of one domain name:

} else {
    /* NUL: end of current domain name */
    if (retpos == 0) {
        crtpos++;
    } else {
        crtpos = retpos;
        retpos = depth = 0;
    }
    if (dst)
        dst[len - 1] = ' ';
}

The write at dst[len - 1] replaces the trailing dot of the name just decoded with a space. If the very first byte of the option is NUL, no name has been decoded yet and len is zero. len is unsigned, so len - 1 wraps, and the store lands somewhere the code never intended. On a 64-bit host the write is four gigabytes past the buffer and the client crashes; the upstream commit message calls it a SEGV. On a 32-bit target like this one, pointer arithmetic wraps and the same statement writes one byte before the output pointer. Either way, a single malformed DHCP response from anyone on the local network reaches a write that should never execute, in a process that runs as root.

Denys Vlasenko fixed it in March 2016, in BusyBox 1.25.0, with a one-clause guard:

-               if (dst)
+               if (dst && len != 0)
                    dst[len - 1] = ' ';

That is the whole fix. It is the kind of change that is trivial to apply when you have the source and a build system, and it is exactly the kind of change that firmware in the field never receives.

Where it lives

The binary we patched is the BusyBox from the firmware of an Alaris infusion pump, analyzed under ARPA-H's DIGIHEALS program. Its version banner reads BusyBox v1.21.0 (2019-01-22 08:51:59 EST). BusyBox 1.21.0 was released in January 2013, so the firmware was built with a six-year-old BusyBox, nearly three years after the fix for this CVE was published.

This is not specific to medical devices. In our earlier study of open-source components in IoT firmware, we released a corpus of about 1,300 binaries extracted from consumer routers, gateways, and cameras. It contains 207 BusyBox binaries. Of those, 201 are versions that predate 1.25.0, and 126 of them shipped in firmware released after the fix was available. The most recent firmware in the corpus carrying a pre-fix BusyBox is dated April 2025. Whether the vulnerable decoder is actually compiled into a given binary depends on the vendor's build configuration, and we have confirmed it is present in ARM BusyBox binaries from that corpus. The pump is one instance of a pattern.

What the compiler did with it

Here is the vulnerable statement as it exists in the pump's binary. R5 holds dst, R4 holds len, and R6 holds retpos:

0x24560   cmp     r5, #0              ; if (dst)          -- sets the NZCV flags
0x24564   addne   r3, r5, r4          ; r3 = dst + len        (only if dst != 0)
0x24568   movne   r2, #0x20           ; r2 = ' '              (only if dst != 0)
0x2456c   movne   r6, #0              ; retpos = 0            (only if dst != 0)
0x24570   moveq   r6, r5              ; retpos = dst, i.e. 0  (only if dst == 0)
0x24574   strbne  r2, [r3, #-1]       ; dst[len - 1] = ' '    (only if dst != 0)
0x24578   ...                         ; fall through, both cases

There is no branch. The compiler if-converted the statement: one compare sets the condition flags, and each of the next five instructions carries its own condition code (ne or eq) and executes only if the flags satisfy it. ARM compilers do this aggressively for short conditional bodies because a predicated instruction that does nothing is cheaper than a mispredicted branch. Embedded code makes use of predicated instructions because predictable timing is often important.

Predication is convenient for the processor and awkward for everyone else.

For analysis, each predicated instruction is a tiny if with its own join immediately after it. Treating them that way is sound but loses precision: five instructions produce five joins, and facts established under the condition are merged away before the next instruction can use them. The alternative, promoting every predicated instruction to real control flow, fragments the control-flow graph and makes the lifted code unreadable.

For patching, a change to the condition has no branch to redirect. The condition lives in the flags, and five separate instructions consume it. Changing the compare changes all of them at once. And if anything after the run also reads those flags, or depends on a register the run wrote, a naive replacement silently breaks it.

Analyzing predicated code: fragments

CodeHawk's original approach to predicated instructions was to recognize common compiler idioms, such as a predicated assignment or a ternary, and fold each into a single statement, falling back to in-block branching for anything else. On desktop binaries that covered most cases. On embedded firmware it did not: the majority of predicated runs match no idiom.

The approach we now use, implemented by Henny Sipma in CodeHawk, is to collect the predicated instructions in a basic block into fragments. A fragment is the set of instructions that depend on the same flag-setting test, split into a then bucket and an else bucket by polarity. For the block above, the fragment is:

fragment
  test        : 0x24560   cmp r5, #0
  opener cc   : NE
  then bucket : 0x24564   addne  r3, r5, r4
                0x24568   movne  r2, #0x20
                0x2456c   movne  r6, #0
                0x24574   strbne r2, [r3, #-1]
  else bucket : 0x24570   moveq  r6, r5

Each bucket is analyzed as a straight-line sequence under its condition, with a single join at the end of the fragment instead of one after every instruction. When a block ends in a conditional branch on the same flags, the branch condition is conceptually hoisted to the start of the fragment and each bucket is connected directly to the successor of matching polarity, which strengthens the invariants in the successor block as well. All of this restructuring happens inside the translation of a single basic block into CodeHawk's internal form. It adds no nodes to the control-flow graph, and invariants are still retrieved by instruction address, so nothing about the soundness of the surrounding analysis changes.

The practical result is that the lifted C for this block reads like the source it came from:

if ((dst != 0)) {
  retpos = 0;
  (*((dst + len) - 1)) = 32;
} else {
  retpos = 0;
}

The retpos = 0 on both sides is faithful to the binary. The compiler set R6 on both paths, and on the eq path R5 is known to be zero, so both assignments are the same.

The fix is one C edit

CodeHawk's patching workflow, which we have described in earlier posts on remediating CVE-2024-12248 and on our BAR publication, is to edit the lifted C and let the patcher work out the binary change. Here the edit is the upstream fix, transcribed:

-              if ((dst != 0)) {
+              if ((dst != 0) && (len != 0)) {
                 retpos = 0;
                 (*((dst + len) - 1)) = 32;
               } else {

No assembly was written. The person making the patch does not need to know that the if was if-converted, which registers hold dst and len, or that the flags are shared by five instructions. Those are the patcher's problems.

Why not a trampoline

The standard way to patch a binary in place is a trampoline: overwrite one or more instructions at the site with a jump to free space, put the new code there along with the displaced instructions and the register spills and restores needed to keep the new code from disturbing anything, and jump back. It is general, and CodeHawk's patcher uses it routinely. It also costs two branches plus the spill traffic on every execution.

This site runs once per decoded domain-name label, inside the inner loop of the decoder. A trampoline here would be paid on every iteration, in a function whose whole purpose is to run in a tight loop over untrusted input. That is a poor trade for a fix whose source-level footprint is three tokens.

The alternative is to change the bytes in place, which is only possible if the new code fits in the space the old code occupied and only safe if it disturbs nothing the surrounding code relies on. For a predicated run, both turn out to be achievable, and the second is provable.

Recompiling the run in place

The mechanism, implemented by Dan Phung in the CodeHawk patcher, treats the feeding compare and the entire predicated run as one unit. The patch description the patcher generates from the C edit records that decision directly:

PatchKind            : Replacement
TrampolineHookStart  : 0x24560           # the compare, not the first predicated instruction
TrampolineHookFootprint : 24 bytes       # compare through end of run, up to 0x24578
NewStmts:
  if (((R5 != 0) && (R4 != 0))) { R6 = 0; (*((R5 + R4) - 1)) = 32; } else { R6 = 0; }
OptimizeBodyAsUnit   : true
PreserveRegs         : R4, R5, R6
RegLiveAtExit        : R4, R5, R6, R8, R9, R10, R11, SP

The body is the modified guard with its then and else code, expressed over the registers the original code used. The patcher compiles it as a single optimizable unit for the target, and the compiler does what the original compiler did: it if-converts the body back into predicated form. The result is written over exactly the bytes the original compare and run occupied, and execution falls through to 0x24578 as before.

ORIGINAL (24 bytes)                     PATCHED (24 bytes)
0x24560  cmp    r5, #0                  0x24560  cmp    r5, #0
0x24564  addne  r3, r5, r4              0x24564  cmpne  r4, #0
0x24568  movne  r2, #0x20               0x24568  addne  r0, r5, r4
0x2456c  movne  r6, #0                  0x2456c  movne  r1, #0x20
0x24570  moveq  r6, r5                  0x24570  strbne r1, [r0, #-1]
0x24574  strbne r2, [r3, #-1]           0x24574  mov    r6, #0

The new condition became a chained compare: cmpne r4, #0 only executes, and only updates the flags, if the first compare found dst non-zero, so the ne instructions after it run only when both dst != 0 and len != 0. The compiler also noticed what we noted earlier: R6 is zero on both paths, so the two predicated moves collapsed into one unconditional mov r6, #0. The patch closes the vulnerability and is one instruction shorter than the code it replaces.

Eighteen of the twenty-four bytes changed. Nothing else in the binary was touched.

What makes an in-place patch safe

A trampoline's spills and restores are what make it safe to write arbitrary new code. An in-place patch has none, so its safety has to be established. The patcher refuses to emit an in-place patch unless three conditions hold, and it takes the facts it needs from CodeHawk's analysis of the original binary rather than from any assumption about the compiler's output.

The compare and the run must be one straight-line block. The in-place path treats the physically preceding compare as the run's test. If a branch target or a branch lies inside the footprint, that assumption is false, and the patch is refused.

No condition flag may be live at the exit. Replacing the compare overwrites the NZCV flags it set. If any later instruction reads those flags, replacing the compare changes its behavior. CodeHawk now derives flag liveness per instruction address from its reaching-definition facts and exports it with the lifting. The patcher checks that all four flags are dead at 0x24578. If they are not, or if the analysis could not establish it, the in-place patch is refused.

No live register may be clobbered. The recompiled body wrote R0 and R1 where the original wrote R2 and R3. That is fine only if none of those registers is needed after the run. CodeHawk exports register liveness the same way it exports flag liveness. The exit-live set for this site is R4, R5, R6, R8, R9, R10, R11, SP. R6 is a declared output of the body, and R0 through R3 are not live, so the patch passes. Had the compiled body clobbered a live register, the patcher would have fallen back to a trampoline that saves and restores it.

Each of these is a sound fact about the original binary, produced by an analyzer with a formal semantics for every instruction, not a heuristic about what compilers usually do. This is the difference between an in-place patch that happens to work and one that is known to work.

Verifying the patched binary

The last step is independent of everything above. CodeHawk analyzes the patched binary from scratch and lifts it again, and we compare the new lifting to the original. The only change is the intended one:

-              if ((dst != 0)) {
-                retpos = 0;
+              if (((dst != 0) && (len != 0))) {
                 (*((dst + len) - 1)) = 32;
-              } else {
-                retpos = 0;
               }
+              retpos = 0;

The guard now includes len != 0, so the store at dst[len - 1] cannot execute when len is zero. The retpos = 0 moved out of the conditional, which is the lifting faithfully reporting the unconditional mov r6, #0 the compiler produced. Semantically it is the same assignment on both paths, as it was before.

Why this matters

Binary patching tools generally get one of two things wrong. Some emit patches that work on the examples their authors tried, with no argument for why the patch is safe beyond the fact that the tests passed. Others are careful to the point of being unusable, insisting on trampolines and full spills everywhere, which is safe but makes the patch expensive.

The approach here takes a third position. The patch is minimal, twenty-four bytes with no trampoline and no new code anywhere else in the binary, and its safety is not assumed. Every fact it depends on, which flags are live, which registers are live, whether the region is straight-line, comes from a sound analysis of the original binary. And the result is checked afterward by analyzing and lifting the patched binary independently, so the argument does not rest on the patcher having been correct. Furthermore, the application of this patch does not require reverse engineering experience, as the fix was defined in our CodeHawk’s verified C lifting.

If you are responsible for devices running code you cannot rebuild, this is what we do. Start the conversation at [email protected].

Appendix: the patch, byte for byte

The patched region is at virtual address 0x24560 in the dname_dec function (entry 0x2449c), file offset 0x1c560. The original binary's SHA-256 is e15ad25829bc3e121aaa3948ccec8e61ddbde75d5f015f98b6eff416494d9dd2.

VAOriginal bytesOriginal instructionPatched bytesPatched instruction
0x24560000055e3cmp r5, #0000055e3cmp r5, #0 (unchanged)
0x2456404308510addne r3, r5, r400005413cmpne r4, #0
0x245682020a013movne r2, #0x2004008510addne r0, r5, r4
0x2456c0060a013movne r6, #02010a013movne r1, #0x20
0x245700560a001moveq r6, r501104015strbne r1, [r0, #-1]
0x2457401204315strbne r2, [r3, #-1]0060a0e3mov r6, #0

Bytes are shown as stored in the file (little-endian words). Result: the store to dst[len - 1] executes only when dst != 0 && len != 0, which is the upstream fix for CVE-2016-2147.