AI Playground by hjLabs.in

How big a model can a browser tab actually run?

The honest answer, measured on this site in October 2026, is about four billion parameters — Qwen3 4B at 2.7 GB, quantised to four bits. That is not where the hardware gives out. It is where the published model files give out: above roughly 3B, the ONNX exports that a browser runtime can address stop being produced at all, and what exists instead is packaged for ONNX Runtime GenAI, which a web page cannot load. So the ceiling you hit first is a packaging ceiling, and a GPU with 24 GB of memory runs into it at exactly the same place as a GPU with 8 GB.

At a glance

Largest model that runs here
Qwen3 4B — 2.7 GB at q4f16, needs shader-f16
Largest with no special GPU feature
Llama 3.2 3B Instruct — 3.2 GB at q4
What stops it going higher
No browser-loadable ONNX export above ~4B exists
Why Qwen3 8B is not here
Published only in ONNX Runtime GenAI layout, which pipeline() cannot address
Rough free VRAM for a 4B
5–6 GB, weights plus attention cache
iPhone ceiling
~100 MB of weights — a WebKit limit, not a WebGPU one
Runtime
transformers.js 4.2.0 on WebGPU, falling back to WebAssembly

Below that ceiling, three things decide what your own machine reaches. The first is download size, which is the only one most people think about: 3.2 GB has to arrive once, and it is cached afterwards. The second is the WebGPU "shader-f16" feature, which is a property of your GPU and driver — present or absent, with nothing in between and no setting that changes it. Every ONNX build of a 4B model is a 16-bit-float build, so an adapter without that feature cannot run one at all, whatever its memory. The machine this catalogue is built on is such an adapter, which is why the largest non-f16 model here — Llama 3.2 3B at 3.2 GB — is listed separately from the two that need the feature.

The third is video memory, and it is the one nobody can check for you. WebGPU deliberately exposes no VRAM figure — it is a fingerprinting surface, and there is no API for it — so no web page can tell you in advance whether a 4B model will fit. What a page CAN read is the adapter’s maximum buffer and binding sizes, which scale with card class and are the best available proxy. This site shows you those numbers and warns rather than blocks, because the failure is recoverable: an out-of-memory on a desktop surfaces as a clear error, not a lost tab.

Weights are not the whole appetite. A model also needs its attention cache resident while it writes, and that grows with the length of the conversation — so a 2.7 GB download wants perhaps 5 GB of free GPU memory in practice, and a long chat wants more than a short one. The practical rule: if the download is under half your card’s memory you will be comfortable, and if it is over two thirds you should expect to restart the conversation occasionally.

None of this applies to phones, and the gap is larger than people expect. WebKit on iOS kills a tab that holds much more than 100 MB of weights, without raising an error the page can catch, so the iPhone ceiling is roughly thirty times lower than a desktop’s and is enforced here rather than warned about. Android has no such wall and scales with the RAM the device reports.

biggest llm in browserrun 4b model in browserlargest model webgpuhow much vram to run llm in browsershader-f16 webgpubrowser llm size limitqwen3 4b browserllama 3.2 3b webgpuonnx model size limit browser

Frequently asked questions

What is the largest model I can run in a browser tab?

On this site, Qwen3 4B — four billion parameters, 2.7 GB downloaded, quantised to four bits. It needs a GPU that exposes the WebGPU shader-f16 feature. Without that feature the largest is Llama 3.2 3B Instruct at 3.2 GB, which runs on any WebGPU adapter with the memory for it.

How much VRAM do I need to run a 4B model in the browser?

Roughly 5 to 6 GB free: the 2.7 GB of weights plus the attention cache, which grows with the length of the conversation. No web page can measure your VRAM — WebGPU does not expose it — so this is an estimate, and the honest test is to try it, because an out-of-memory on a desktop is a clear error rather than a crash.

Why can I not run a 7B or 8B model in a browser?

Not because of memory. Above about 4B, nobody publishes an ONNX export in the layout a browser runtime can load. Qwen3 8B, for example, exists as ONNX and even ships a WebGPU-targeted build — but in ONNX Runtime GenAI’s directory layout with a genai_config.json, which transformers.js cannot address. A bigger graphics card does not move that ceiling; a different runtime would.

What is shader-f16 and how do I know if I have it?

It is an optional WebGPU feature that lets the GPU compile half-precision shaders. Plenty of current, powerful desktop GPUs do not expose it. This site probes for it on load and shows the result in the hardware panel; models that cannot run without it are blocked before the download starts rather than failing part-way through.

Does a bigger model always answer better?

For open-ended writing and multi-step reasoning, yes, noticeably — a 3B model will usually get a two-step arithmetic question right that a 135M model confidently gets wrong. For classification, embeddings, transcription and most vision tasks the small specialised models are already at the quality ceiling, and a large general model is slower for no gain.