For GPUs behind NAT and firewalls

Run inference on machines you can’t reach.

Install the agent and the machine connects out to Reachback. Requests come back over that same connection, so you never have to open a port or set up a VPN.

How machines connect to ReachbackThree machines, behind NAT, carrier-grade NAT and a firewall, each connect out over port 443 to the Reachback control plane. Requests travel back to them over those same connections. Nothing connects in to the machines.Reachbackyour control planeOfficeNAT↑ 443gpu-01RTX 4090connectedLabCGNAT↑ 443lab-ws-22× A6000connectedCustomer sitefirewall↑ 443edge-7L4connected
the machine connects outrequests come back over it

How it works

Getting a machine into your fleet

The machine makes the connection, so you don’t have to change anything on the network it’s on. If it can load a website, it can join.

  1. 1

    Create an environment

    This is your own control plane. It isn’t shared with other customers, and it has its own certificate authority.

  2. 2

    Install the agent

    Start it with a join token from the dashboard. It connects to your environment on port 443 and runs models with llama.cpp.

  3. 3

    Send requests

    The API is OpenAI-compatible, so the client libraries you already use will work. Reachback sends each request to a machine that’s serving the model.

# on each machine
reachback-agent \
  -controller k3v9q2m8xa.edge.reachbackhq.com:443 \
  -join-token rb_… \
  -node-id "$(hostname)"

# in your app
from openai import OpenAI

client = OpenAI(
    base_url="https://k3v9q2m8xa.gw.reachbackhq.com/v1",
    api_key=token,
)

Why Reachback

Other control planes need you to own the network

NVIDIA Dynamo, llm-d, Ray and KServe are built for clusters, where every machine is on a network you run and can reach. A lot of useful hardware isn’t set up like that. It might be a workstation in an office, a server at a customer’s site, or a box behind carrier-grade NAT. Reaching it usually takes a VPN or a firewall change, and often you can’t get either.

Reachback only needs the machine to be able to connect out.

RequirementCluster-basedReachback
Machines on a network you can route intoRequiredNot needed
Inbound firewall rules at each siteRequiredNone
Machines behind NAT or carrier-grade NATNeed a VPN or tunnelWork as they are
What the site’s network has to allowInbound connectionsOutbound HTTPS

Networks

Designed for slow, unreliable links

Home and office connections are slower and much less predictable than a datacenter network. We measured these on residential links, and Reachback is built with numbers like them in mind.

round-trip time
154–324 ms
jitter
74–84 ms

Machines don’t need a reachable address

In one test we set a machine to advertise 203.0.113.99, an address reserved for documentation that nothing on the internet can reach. Inference still worked. The control plane never connects to a machine. It only uses the connection the machine opened.

  • Offices

    GPU workstations sitting behind the office router.

  • Customer sites

    Servers on a network you don’t manage and can’t open ports on.

  • Labs and edge sites

    Machines behind carrier-grade NAT, with no public IP.

  • Other people’s networks

    Hardware you’re allowed to use, on a network you can’t change.

Security

What your security team will ask about

Nothing on your machines listens for incoming connections. The agent connects out, checks that it’s talking to your environment, and only then accepts work.

Open source agent
The agent and the protocol it speaks are Apache 2.0 licensed, so you can read the code that runs on your machines.
Your own control plane
Each environment runs on its own, with its own certificate authority. Other customers’ traffic never goes through it.
Certificate pinning
Join tokens include your environment’s CA fingerprint. The agent won’t connect to a server that can’t prove it holds that CA.
Instant revocation
Revoking a token disconnects any machine using it straight away. Join tokens can only register machines, not make admin calls.
Audit log
Every request is logged with the machine, model, token counts and timing. The log never includes prompts or responses.
What passes through us
Your control plane terminates TLS so it can route requests, which means prompts pass through it. It doesn’t store them.

Tell us about your machines

We’re starting with a small number of teams. Email us with where your hardware is and what you’d like to run on it, and we’ll get back to you.