Projects

Compute Mesh

GPU-pool job coordination for submitting model jobs, managing workers, and tracking execution.

  • Software

Compute Mesh coordinates GPU jobs across a private group of computers. The coordinator, worker, and client paths handle job state, model availability, failure handling, and usage accounting.

When
April 2026
Where
Independent project
Role
Designer and developer
Coordinator
FastAPI with SQLite for jobs, workers, and a credit ledger
Execution
Each job runs on one worker through Ollama
Clients
React control UI in Swift and Tauri desktop shells
Tools
Python, FastAPI, SQLite, Ollama, React, TypeScript, Swift, Rust, Tauri

Job lifecycle: from submit to payout

A consumer submits a job for a model, and the coordinator checks the balance and reserves 10 credits before queueing it. Workers poll, say which models they run, and claim the oldest job they can take under their job cap. Each claim happens inside one locked transaction, so two workers polling at once never get the same job. Done pays the worker 8 credits, a rejection refunds the consumer, and a job with no reply for 3 ticks goes back to the queue.

Synthetic simulation of the coordinator's job lifecycle. No real jobs, machines, or models.

Tick 4. Balance 30 credits, 20 reserved. Press Step to run the next tick.

The controls need JavaScript. The board shows the starting state after 4 ticks.

  • Worker A

    Advertises
    small-model, medium-model
    Local allowlist
    small-model, medium-model
    Job cap
    1

    Polling Holding 1 of 1: Job 3, medium-model Earned 8 credits

  • Worker B

    Advertises
    small-model
    Local allowlist
    small-model
    Job cap
    2

    Polling Holding 1 of 2: Job 4, small-model Earned 0 credits

  • Worker C

    Advertises
    small-model, large-model
    Local allowlist
    small-model
    Job cap
    1

    Polling Holding 0 of 1: No job Earned 0 credits

Three synthetic workers. Worker C advertises large-model, but its local allowlist refuses it.

Queued0

  • None

Assigned2

  • Job 3medium-modelWorker A, tick 1 of 2
  • Job 4small-modelClaimed by Worker B this tick

Done1

  • Job 1small-modelOutput from Worker A

Failed1

  • Job 2large-modelRejected by Worker C, refunded
Every job is in exactly one state. Assigned is the running phase.

Consumer credits

Balance
30
Reserved
20
Spent
10
Ledger, newest first. Payouts come out of the job's reserve.
TickEntryCredits
3Reserve for job 4−10
3Payout to Worker A for job 18 from reserve
2Refund for job 2+10
0Reserve for job 3−10
0Reserve for job 2−10
0Reserve for job 1−10

Event log

  1. Tick 4 Worker C polled in the same tick, but job 4 was already claimed.
  2. Tick 4 Worker B claimed job 4.
  3. Tick 3 Job 4 submitted for small-model. 10 credits reserved.
  4. Tick 3 Worker A claimed job 3.
  5. Tick 3 Worker A posted output for job 1 and was paid 8 credits.
  6. Tick 2 Worker C rejected job 2: large-model is not on its local allowlist. 10 credits refunded.

Problem

A few computers with GPUs can share model jobs, but only if something keeps track of which jobs are waiting, which machine has each one, and what happens when a machine disappears or refuses a job. Compute Mesh is that coordinator.

My role

My role covered design and implementation of the coordinator, worker, command-line client, and desktop shells.

Job lifecycle

  1. Submit. The coordinator checks the requester’s credit balance and reserves the job’s cost in an append-only ledger. A job for a model that is not on the allowlist is refused here.
  2. Queued. Jobs wait in order of arrival.
  3. Assigned. Workers poll the coordinator and say which models they can run. A worker under its job limit claims the oldest job it supports. The claim happens inside a locked database transaction with a status check, so two workers polling at the same moment can never take the same job.
  4. Done. The worker posts its output and is paid from the reserved credits.
  5. Failed and refunded. If the worker’s own allowlist rejects the model, the job fails and the requester gets a full refund.
  6. Back to the queue. A job held by a worker for longer than the timeout returns to the queue for another worker, and its credits stay reserved.

Completions and failures are only accepted from the worker that holds the job.

Decisions

  • One job, one worker. Each job runs on a single machine, so every job has exactly one owner at a time.
  • Pull, not push. Workers ask for work. The coordinator never connects to a worker, and a worker that goes quiet stops claiming jobs.
  • Check the model three times. The allowlist is checked at submission, at assignment, and again on the worker.
  • One interface, two shells. The control UI is a single React app. On macOS a Swift shell hosts it in a WKWebView and starts the local processes, and a Rust and Tauri shell does the same on Windows and Linux.

Cluster planning

A placement step I wrote plans how one model could run across a group of machines. It checks group size, network latency between members, and model support, then records a plan for how the model’s layers would be divided. Jobs in this mode are refunded rather than executed.