Open-weight language models delivered to the browser from a content-addressed CDN and run on WebGPU, on a server, or split between the two at a chosen layer. One run() call picks the mode for this device and network, streams the answer, and reports the speed, the server's share of the cost, and what left the device.
Live demos
Summariser — paste a document, watch the planner choose local, split or server, and compare all three.
Split slider — every split point N with live estimates against the measured run.
Cache demo A / demo B — two sites, one model download (cross-site cache, Chrome only).
Numbers
Stats dashboard — load times and generation sessions reported by the SDK (no prompt content is ever sent).
Write-up — what was built, the two positive and the two negative results, with every number linked to its results file.