Connect a GPU machine
One command per machine. After that, the machine runs whatever the pool assigns it and keeps itself up to date.
Before you start
The machine needs:
- Linux, on x86-64.
- An NVIDIA GPU of a model the region accepts, with a recent driver. Check with nvidia-smi.
- Docker Engine with the Compose plugin.
- The NVIDIA container toolkit, so containers can use the GPU. Check with docker run --rm --gpus all ubuntu nvidia-smi: it should list your GPUs.
- Disk space for the models: tens of gigabytes per app.
- An outbound internet connection. No ports need opening.
The steps
-
Apply, and wait for approval
On the member site, sign in with your wallet, pick a region, and list each machine with its GPUs. The operator approves the application and assigns each machine its apps. Your overview shows when.
-
Create the machine's credential
On your machines page, press Create credential for the machine. The credential is shown once, inside the command to run. Each machine has its own; creating a new one stops the old one working.
-
Run the command on the machine
It looks like this. Use the one from your machines page, which has your credential in it:
docker run -d --name openpool-agent-<region> --restart unless-stopped --gpus all \ -v /var/run/docker.sock:/var/run/docker.sock \ -e OPENPOOL_CREDENTIAL=<shown once on your machines page> \ <the agent image> --pool <the region's address> --key <the region's signing key>With rootless Docker, mount the socket
DOCKER_HOSTnames in place of/var/run/docker.sock. -
Watch it come up
Within a minute your machines page shows what the agent found: its version, the GPUs, the containers. The first start downloads the app's image and model, which can take a while. When the app is up, its runner appears on your runners page on probation, is tested, and goes active.
What the agent does
Every half minute it does three things:
- Asks the pool what this machine should run. The answer is a Docker Compose file, signed by the pool.
-
Checks and applies it. A file not signed with the key the agent was started with is refused. So is a
correctly signed file that asks for more than app containers: privileged mode, the host's network or processes, added
capabilities, devices, published ports, or mounts of your disks. Then it runs
docker compose up. - Reports which containers are running, the GPUs it found, and any error.
If the pool cannot be reached, the agent keeps running what it was last told. If the pool stops accepting the machine's credential, because the machine was removed or the credential replaced, it stops the pool's containers after an hour.
Common questions
How do I see what is happening on the machine?
docker logs openpool-agent-<region> shows the agent's log. docker ps shows the app and tunnel containers it started.
How do I update an app?
You do not. When the operator releases a new version, machines are moved to it a few at a time, each after finishing the sessions it is serving.
How do I take a machine out of the pool?
Remove it on your machines page. Its runners finish their sessions and are removed, and the agent stops the pool's containers. Then docker rm -f openpool-agent-<region> removes the agent. Model files stay in Docker volumes until you delete them.
My runner is suspended. Why?
Your runners page says why, with the result of every test. A runner suspended for failing tests comes back by itself once it passes again.
Can one machine join two regions?
Yes: apply in both and run each region's command. The two agents keep their containers apart. They share the GPUs, so make sure the machine can serve what both regions assign it.