For GPUs behind NAT and firewalls
Run inference on machines you can’t reach.
Install the agent and the machine connects out to Reachback. Requests come back over that same connection, so you never have to open a port or set up a VPN.
How it works
Getting a machine into your fleet
The machine makes the connection, so you don’t have to change anything on the network it’s on. If it can load a website, it can join.
- 1
Create an environment
This is your own control plane. It isn’t shared with other customers, and it has its own certificate authority.
- 2
Install the agent
Start it with a join token from the dashboard. It connects to your environment on port 443 and runs models with llama.cpp.
- 3
Send requests
The API is OpenAI-compatible, so the client libraries you already use will work. Reachback sends each request to a machine that’s serving the model.
# on each machine
reachback-agent \
-controller k3v9q2m8xa.edge.reachbackhq.com:443 \
-join-token rb_… \
-node-id "$(hostname)"
# in your app
from openai import OpenAI
client = OpenAI(
base_url="https://k3v9q2m8xa.gw.reachbackhq.com/v1",
api_key=token,
)Why Reachback
Other control planes need you to own the network
NVIDIA Dynamo, llm-d, Ray and KServe are built for clusters, where every machine is on a network you run and can reach. A lot of useful hardware isn’t set up like that. It might be a workstation in an office, a server at a customer’s site, or a box behind carrier-grade NAT. Reaching it usually takes a VPN or a firewall change, and often you can’t get either.
Reachback only needs the machine to be able to connect out.
| Requirement | Cluster-based | Reachback |
|---|---|---|
| Machines on a network you can route into | Required | Not needed |
| Inbound firewall rules at each site | Required | None |
| Machines behind NAT or carrier-grade NAT | Need a VPN or tunnel | Work as they are |
| What the site’s network has to allow | Inbound connections | Outbound HTTPS |
Networks
Designed for slow, unreliable links
Home and office connections are slower and much less predictable than a datacenter network. We measured these on residential links, and Reachback is built with numbers like them in mind.
- round-trip time
- 154–324 ms
- jitter
- 74–84 ms
Machines don’t need a reachable address
In one test we set a machine to advertise 203.0.113.99, an address reserved for documentation that nothing on the internet can reach. Inference still worked. The control plane never connects to a machine. It only uses the connection the machine opened.
Offices
GPU workstations sitting behind the office router.
Customer sites
Servers on a network you don’t manage and can’t open ports on.
Labs and edge sites
Machines behind carrier-grade NAT, with no public IP.
Other people’s networks
Hardware you’re allowed to use, on a network you can’t change.
Security
What your security team will ask about
Nothing on your machines listens for incoming connections. The agent connects out, checks that it’s talking to your environment, and only then accepts work.
- Open source agent
- The agent and the protocol it speaks are Apache 2.0 licensed, so you can read the code that runs on your machines.
- Your own control plane
- Each environment runs on its own, with its own certificate authority. Other customers’ traffic never goes through it.
- Certificate pinning
- Join tokens include your environment’s CA fingerprint. The agent won’t connect to a server that can’t prove it holds that CA.
- Instant revocation
- Revoking a token disconnects any machine using it straight away. Join tokens can only register machines, not make admin calls.
- Audit log
- Every request is logged with the machine, model, token counts and timing. The log never includes prompts or responses.
- What passes through us
- Your control plane terminates TLS so it can route requests, which means prompts pass through it. It doesn’t store them.
Tell us about your machines
We’re starting with a small number of teams. Email us with where your hardware is and what you’d like to run on it, and we’ll get back to you.