WebGPU LLM — Ask your questions across all site pages with an in-browser LLM — https://www.pirahansiah.com/webgpu-llm/
🔒 100% local & private — runs in your browser
This page runs a small LLM entirely in your browser — no data, question, or document
ever leaves your device. Pick a model from ~15 MB (Nano) up to ~1.5 GB (Large);
each option shows its estimated total RAM (weights + KV cache + context). The
Method selector chooses how it runs: Auto picks the best backend for your
system, WebGPU uses your GPU, and WASM is the universal fallback that runs
on any device (its SIMD → plain ladder is the modern successor to asm.js). The first load
downloads the model once, then it's cached by the browser. All answers are English only.
Questions run on your own hardware — no account, no tracking.
Try:
checking GPU…
💡 Topic review
📊 Visualizations
Categories in results
Top tags & hashtags
🔗 References
🌐 Find more on the web
Opens a new tab with a search for this topic — combine with the tags above.
🐦 Ready to post on X
🔗 Connection map
Top results and their neighbors from the knowledge graph — click a node to open
💬 Chat with your site
Conversational follow-ups — the model remembers this conversation. Same local & private engine.Hi! Ask me anything about the site — papers, courses, CV, code, anything. I answer from your pages, locally in your browser. (First question may need the model to load once.)
Runs fully in your browser: 9 models from ~15 MB (Nano) to ~1.5 GB (Large) with per-model RAM estimates (weights + KV cache + context), via transformers.js. Method: Auto / WebGPU / WASM — WASM runs on every device (the modern asm.js-style universal fallback). All answers are English only. BM25 retrieval over every page; first run downloads the model (cached).