WebGPU LLM — Ask your questions across all site pages with an in-browser LLM — https://www.pirahansiah.com/webgpu-llm/

WebGPU LLM

Loading…
🔒 100% local & private — runs in your browser
This page runs a small LLM entirely in your browser — no data, question, or document ever leaves your device. Pick a model from ~15 MB (Nano) up to ~1.5 GB (Large); each option shows its estimated total RAM (weights + KV cache + context). The Method selector chooses how it runs: Auto picks the best backend for your system, WebGPU uses your GPU, and WASM is the universal fallback that runs on any device (its SIMD → plain ladder is the modern successor to asm.js). The first load downloads the model once, then it's cached by the browser. All answers are English only. Questions run on your own hardware — no account, no tracking.
Try:
checking GPU…
0%

💡 Topic review

📊 Visualizations
Categories in results
Top tags & hashtags
🔗 References
    🌐 Find more on the web
    Opens a new tab with a search for this topic — combine with the tags above.
    🐦 Ready to post on X
    Open in X

    🔗 Connection map

    Top results and their neighbors from the knowledge graph — click a node to open

    💬 Chat with your site

    Conversational follow-ups — the model remembers this conversation. Same local & private engine.
    Hi! Ask me anything about the site — papers, courses, CV, code, anything. I answer from your pages, locally in your browser. (First question may need the model to load once.)

    Runs fully in your browser: 9 models from ~15 MB (Nano) to ~1.5 GB (Large) with per-model RAM estimates (weights + KV cache + context), via transformers.js. Method: Auto / WebGPU / WASM — WASM runs on every device (the modern asm.js-style universal fallback). All answers are English only. BM25 retrieval over every page; first run downloads the model (cached).