A secret the machine never has

The agent reads an API key from its environment, sends it, and the call succeeds. The key was never in the machine: it is swapped in on the way out, 37 bytes at a time.

Laurentiu CiobanuLaurentiu Ciobanu8 min readDeep dive
A creature in a lit room posts a small blank card through a slot in a thick wall; on the dark far side a mechanical arm catches it and swaps it for a heavy wax-sealed envelope that carries on into the night.
On this page · 8 min

Deep dive is a highly technical series where I go deeper into the technology that makes boxd tick. This is not for the faint of heart!

An agent running on one of our machines wants to call an API. To do that it needs a key, so you give it one. It goes in an environment variable, or a dotfile, or a config the agent reads at startup.

From that moment the key is in the machine. Not "available to the agent". In the machine. Any process can read it. Anything the agent installs can read it. Anything the agent is talked into running by a web page it fetched can read it. If the machine is a dev box you also share it with an editor, a shell, and whatever you cloned this morning. The blast radius of a leak is not the agent, it is everything that has ever run in there.

The usual answers are scoped keys and short expiry, and they are good answers, and they are answers about how much a stolen key is worth rather than about whether it can be stolen. I wanted the other thing: the request arrives at the API with the real key in it, and the machine never held the real key at any point.

Where a secret can hide

There is exactly one place. The key has to reach the API, so it has to cross the network, so it has to pass through something we control on the way out. Every other location is inside the blast radius by definition.

So the guest gets a fake. A dummy, of a fixed and unmistakable shape:

Rust
/// The `bxds_` dummy is `bxds_` + 32 hex chars.
pub const DUMMY_LEN: usize = 5 + 32;

The agent reads bxds_ and thirty-two hex characters from its environment, puts it in an Authorization header, and sends the request. On the way out of the machine, that string is replaced with the real one.

Nothing in the guest changes. No SDK, no proxy configuration, no code that knows about any of this. The agent believes it has a key. It has thirty-seven bytes of nothing.

Getting in front of the traffic

On a machine with an egress policy, the worker installs an nftables rule that REDIRECTs the guest's outbound TCP 80 and 443 to a listener on the bridge gateway. The guest is not configured to use a proxy and cannot decline to.

The listener then has to answer three questions in order.

Which machine is this? The source IP of the connection, looked up against the VM table.

Where is it going? From the TLS ClientHello's SNI, or the Host header on plaintext. Note this is the name the guest asked for, which is the only thing worth having.

What are we allowed to do about it? That is a three-way verdict, and the middle one is the interesting one:

Rust
pub enum Verdict {
    Deny,
    Splice,
    Substitute(Substitutor),
}

Deny is a host not on the machine's allowlist. The connection ends.

Splice is an allowed host with no secrets attached to it. The bytes are forwarded without being decrypted at all. This matters more than it looks: it would be much simpler to terminate everything and inspect all of it, and we deliberately do not, because there is no reason for us to be able to read a guest's traffic to a host we are not editing. Intercepting is a capability we take only where we are using it.

Substitute is a host with secrets bound to it. Only then do we terminate the TLS, and only for that flow.

Being the API, briefly

To read a TLS stream you have to be the far end of it, and to be the far end of it the guest's TLS stack has to accept you.

The guest trusts one cluster CA, installed into its trust store at boot. When a flow needs substitution, the proxy mints a leaf certificate for that hostname on the spot, signed by the CA, cached per host. The guest validates it and is satisfied, because from the guest's point of view nothing is unusual: the name matches, the chain is trusted, the connection is TLS.

Then we dial the real host ourselves, and there is one line in that path that is worth the whole post:

Rust
/// Resolve `host` and refuse anything that is not a public IPv4 address:
/// the upstream is always the NAME the guest asked for, never the address
/// it dialled.

We re-resolve the hostname and connect to that. We do not connect to the address the guest was heading for.

Consider what happens if you skip this. The guest controls both the SNI it presents and the address it connects to, and they do not have to agree. It says api.stripe.com and points the connection at a host it owns. We would terminate TLS, look up "which secrets apply to api.stripe.com", inject the real key, and forward it to the attacker's server. The interception layer becomes a machine for handing out the very secrets it exists to protect, and it does so on request.

Resolving the name ourselves closes that. Refusing any resolution to a non-public IPv4 closes the other half, where the guest points a name it controls at 10.x and uses us to knock on the internal network from outside.

The substitution itself

Now the boring part, which is where the correctness lives. The engine is deliberately pure (no async, no sockets, no network types) so it can be exhaustively unit tested in isolation.

The core is a literal byte scan. In its simplest form, which is the one worth reading, it is this:

Rust
    fn replace_literals(&self, input: &[u8]) -> Vec<u8> {
        let mut out = Vec::with_capacity(input.len());
        let mut i = 0;
        'outer: while i < input.len() {
            for (dummy, value) in &self.tokens {
                if input[i..].starts_with(dummy) {
                    out.extend_from_slice(value);
                    i += dummy.len();
                    continue 'outer;
                }
            }
            out.push(input[i]);
            i += 1;
        }
        out
    }

No HTTP parsing, no header awareness, no structure at all. It finds the dummy anywhere: in a header, in a query string, in a JSON body, in a form field, in a place I have not thought of.

In practice it gets considerably more complicated than that. A byte-at-a-time loop is fine when you are reading it and less fine sitting in the path of every request every machine makes, so the shipped scanner is vectorised: same semantics, same guarantees, a vector at a time instead of a byte at a time. The version above is the one worth reading, because it is the one that says what the engine means. The fast one only says how quickly.

Doing it this crudely is only safe because of how the dummy is shaped. It is 37 bytes, it is unguessable, and it appears in the guest for exactly one reason. There is no legitimate byte sequence it can collide with. Had we picked something structured (substitute only inside Authorization headers, say) we would be shipping a feature that works until the day an SDK decides to put the key in a query parameter, and then silently sends the fake one.

The one place a literal scan cannot see the secret is when something has wrapped it:

Rust
    /// Scan CRLF-delimited lines for `Authorization: Basic <b64>` (case-
    /// insensitive header name) and, if the decoded credential contains a dummy,
    /// decode → substitute → re-encode. Leaves everything else untouched. Safe
    /// to run after `replace_literals` (the base64 hides the dummy from it).

HTTP Basic auth base64s the credential, so the dummy is not there as bytes any more. Decode, substitute, re-encode. Note the ordering argument in the comment: this runs after the literal pass and cannot conflict with it, precisely because base64 hid the dummy from the first pass.

Two details that only show up in production

A secret can be split across two reads. The dummy is 37 bytes; TCP does not care. A read() can end in the middle of one and the rest arrives in the next buffer, and a naive scanner passes both halves through untouched and sends the fake key to the API. So the streaming path holds back a carry of dummy_len - 1 bytes at the end of each buffer (the longest prefix that could still turn out to be the start of a match) and prepends it to the next one. Thirty-six bytes of hesitation.

Substituting changes the length. The dummy is 37 bytes. A real key might be 51, or 108. If the secret was in the body, the body just changed size, and the Content-Length header the guest wrote is now a lie:

Rust
    let new_head = if new_body.len() != body.len() {
        set_content_length(&new_head, new_body.len())
    } else {
        new_head
    };

Get that wrong and the upstream either truncates the request or waits forever for bytes that are not coming, and it does so intermittently, depending on whether the secret happened to land in the body of that particular call.

Both bugs share a shape worth naming: they do not fail when you test them, they fail when a buffer boundary or a body length happens to land somewhere specific. Which is why the engine is a pure function over byte slices with no I/O in it; that is the only version of this you can actually test to exhaustion.

What this does and does not buy

It does not make a compromised machine safe. An agent that has been taken over can still use the dummy, and its requests will still be signed with the real key on the way out, for as long as the machine is running and the host is on its allowlist. This is not a substitute for scoping keys.

What it removes is the durable artifact. There is no string in that machine which is worth stealing, exfiltrating, or finding later in a disk image, a snapshot, a core dump, or a forked child. An attacker who gets in gets the use of a credential while they are there, and nothing they can take away with them. That is a genuinely different exposure from "the key is in .env and now it is in a paste bin".

There is a through-line with the rest of this series that I did not plan. The first post was about not trusting anything the guest tells us. This one is about not giving the guest anything worth taking. Both fall out of the same starting position: the machine belongs to somebody else, we cannot see inside it, and we should not need to.

Laurentiu CiobanuLaurentiu Ciobanu
PostShare
Published
Sep 28, 2026
Reading time
8 min
Words
1,633
Topic
Deep dive

Read next

Field notes

Subscribe for release notes and architecture write-ups

No spam, ever. Unsubscribe anytime.

Your inbox