Private distributed AI, directly in the browser

RallyCompute divides an open LLM model across the WebGPU of computers you choose directly in the browser.

Runtime ready 4 workers · 48 layers
01
Layers 1–12 Ready · 4.8 GiB
02
Layers 13–24 Ready · 4.6 GiB
03
Layers 25–36 Ready · 5.1 GiB
04
Layers 37–48 Ready · 4.9 GiB
PROMPT

Explain the future of distributed inference.

Generating across 4 browsers
01Browser native
02Private inference
03Peer-to-peer execution

Frequently Asked Questions

What is RallyCompute?

RallyCompute is a browser-based distributed AI runtime. It divides an open model into pieces (shards) and executes those shards across a group of computers working together directly in the browser using WebGPU.

Are you crazy!?

Probably. But as the guy who stuffed, GPT-2 into a spreadsheet (Spreadsheets-are-all-you-need) and VanillaJS, I seem to have a deep attraction to impractically implemented inference. The distributed-systems background just means I can now do it across several machines at once.

Can worker machines be spread out across the internet?

RallyCompute can connect workers across the public internet when WebRTC can establish a direct connection. The current preview has no TURN relay, so some restrictive NAT, carrier-grade NAT, and firewall combinations may not connect. Even when it does connect the latency will be high going over the public internet. For the best speed and reliability, keep workers on the same low-latency LAN when possible. Broader network compatibility is on the roadmap.

Which models are supported?

Current model routes include SmolLM2 360M Instruct and Gemma 4 12B, both using Q4_K_M quantization. More models are coming—contact us if there is a specific model you want to run.

Does source code stay local?

The coding workflow reads only the files you select. Proposed changes are staged in browser storage and require approval before writing. Selected context is processed by the trusted computers participating in your room.

Is this faster or cheaper than a cloud API?

RallyCompute's primary advantages are control, privacy, and using existing compute. Faster? No. Cheaper? Possibly.

What hardware and software do I need?

Participating machines need a Chrome browser and WebGPU-capable hardware. The amount of machine will depend on your selected model and each machine's memory limits for WebGPU. RallyCompute will measure your machine's WebGPU limits.

Contact

Are you a team with private data, recurring AI workloads, or underused computers? Drop us a line.

Please enter your name.
Please enter a valid email.
Please add a short note.

You're on the list.

Thanks for telling us what you want to build. We'll be in touch.