Compute Mesh
GPU-pool job coordination for submitting model jobs, managing workers, and tracking execution.
- Software
Compute Mesh coordinates GPU jobs across a private group of computers. The coordinator, worker, and client paths handle job state, model availability, failure handling, and usage accounting.
- When
- April 2026
- Where
- Independent project
- Role
- Designer and developer
- Coordinator
- FastAPI with SQLite for jobs, workers, and a credit ledger
- Execution
- Each job runs on one worker through Ollama
- Clients
- React control UI in Swift and Tauri desktop shells
- Tools
- Python, FastAPI, SQLite, Ollama, React, TypeScript, Swift, Rust, Tauri
Job lifecycle: from submit to payout
A consumer submits a job for a model, and the coordinator checks the balance and reserves 10 credits before queueing it. Workers poll, say which models they run, and claim the oldest job they can take under their job cap. Each claim happens inside one locked transaction, so two workers polling at once never get the same job. Done pays the worker 8 credits, a rejection refunds the consumer, and a job with no reply for 3 ticks goes back to the queue.
Synthetic simulation of the coordinator's job lifecycle. No real jobs, machines, or models.
Tick 4. Balance 30 credits, 20 reserved. Press Step to run the next tick.
The controls need JavaScript. The board shows the starting state after 4 ticks.
Worker A
- Advertises
- small-model, medium-model
- Local allowlist
- small-model, medium-model
- Job cap
- 1
Polling Holding 1 of 1: Job 3, medium-model Earned 8 credits
Worker B
- Advertises
- small-model
- Local allowlist
- small-model
- Job cap
- 2
Polling Holding 1 of 2: Job 4, small-model Earned 0 credits
Worker C
- Advertises
- small-model, large-model
- Local allowlist
- small-model
- Job cap
- 1
Polling Holding 0 of 1: No job Earned 0 credits
Queued0
- None
Assigned2
- Job 3medium-modelWorker A, tick 1 of 2
- Job 4small-modelClaimed by Worker B this tick
Done1
- Job 1small-modelOutput from Worker A
Failed1
- Job 2large-modelRejected by Worker C, refunded
Consumer credits
- Balance
- 30
- Reserved
- 20
- Spent
- 10
| Tick | Entry | Credits |
|---|---|---|
| 3 | Reserve for job 4 | −10 |
| 3 | Payout to Worker A for job 1 | 8 from reserve |
| 2 | Refund for job 2 | +10 |
| 0 | Reserve for job 3 | −10 |
| 0 | Reserve for job 2 | −10 |
| 0 | Reserve for job 1 | −10 |
Event log
- Tick 4 Worker C polled in the same tick, but job 4 was already claimed.
- Tick 4 Worker B claimed job 4.
- Tick 3 Job 4 submitted for small-model. 10 credits reserved.
- Tick 3 Worker A claimed job 3.
- Tick 3 Worker A posted output for job 1 and was paid 8 credits.
- Tick 2 Worker C rejected job 2: large-model is not on its local allowlist. 10 credits refunded.
Problem
A few computers with GPUs can share model jobs, but only if something keeps track of which jobs are waiting, which machine has each one, and what happens when a machine disappears or refuses a job. Compute Mesh is that coordinator.
My role
My role covered design and implementation of the coordinator, worker, command-line client, and desktop shells.
Job lifecycle
- Submit. The coordinator checks the requester’s credit balance and reserves the job’s cost in an append-only ledger. A job for a model that is not on the allowlist is refused here.
- Queued. Jobs wait in order of arrival.
- Assigned. Workers poll the coordinator and say which models they can run. A worker under its job limit claims the oldest job it supports. The claim happens inside a locked database transaction with a status check, so two workers polling at the same moment can never take the same job.
- Done. The worker posts its output and is paid from the reserved credits.
- Failed and refunded. If the worker’s own allowlist rejects the model, the job fails and the requester gets a full refund.
- Back to the queue. A job held by a worker for longer than the timeout returns to the queue for another worker, and its credits stay reserved.
Completions and failures are only accepted from the worker that holds the job.
Decisions
- One job, one worker. Each job runs on a single machine, so every job has exactly one owner at a time.
- Pull, not push. Workers ask for work. The coordinator never connects to a worker, and a worker that goes quiet stops claiming jobs.
- Check the model three times. The allowlist is checked at submission, at assignment, and again on the worker.
- One interface, two shells. The control UI is a single React app. On macOS a Swift shell hosts it in a WKWebView and starts the local processes, and a Rust and Tauri shell does the same on Windows and Linux.
Cluster planning
A placement step I wrote plans how one model could run across a group of machines. It checks group size, network latency between members, and model support, then records a plan for how the model’s layers would be divided. Jobs in this mode are refunded rather than executed.