The app ships zero model weights. No
.ggufis bundled, embedded, or auto-downloaded. The user voluntarily picks a model; weights travel HuggingFace/CDN → user's browser directly, never through our Pi server.The inference runtime (not weights) is vendored:
@wllama/wllama 2.3.1(MIT, seestatic/vendor/wllama/LICENCE) ships inside the app shell, so Local AI works fully offline with zero CDN requests.
bash
# from llama.cpp releases: llama-gguf-split --split-max-size 512M
./llama-gguf-split --split-max-size 512M ./my_model.gguf ./my_model
# → my_model-00001-of-0000N.gguf …; paste the FIRST shard URL,
# wllama fetches the rest automatically (parallelDownloads: 3 default)
For file-picker flow the shards must be re-joined locally first
(cat shard-* > model.gguf, must stay <2 GB) — or host the shards and
use the URL flow.rust-embed / the Pi binary contain only the app shell (~MBs).Analyze → Enable local AI… opens a modal with
a license checkbox. localStorage["ede.ai.consent"] gates everything in
static/ai.js. No consent → deterministic insights.js only.model_id + revision +
license + date below. Not legal advice — consult counsel for rollout.static/ai.js), never raw PII rows.Source of truth: static/ai-models.js (sizes verified 2026-09-04 via HF
content-range). RAM ≈ 1.2× file size in-browser.
| Model (Q4_K_M) | Size | License | Notes |
|---|---|---|---|
| Qwen2.5 0.5B Instruct — recommended default | 469 MB | Apache-2.0 | Fast on laptop CPU; NL→SQL + short summaries |
| Llama 3.2 1B Instruct | 770 MB | Llama 3.2 Community | Better following, 128K ctx; custom license |
| Qwen2.5 1.5B Instruct | 1.04 GB | Apache-2.0 | Smarter, ~2× slower; comfort ceiling |
| SmolLM2 1.7B Instruct | 1.00 GB | Apache-2.0 | Strong per-param; slow on old CPUs |
| Gemma 2 2B IT | 1.59 GB | Gemma Terms of Use | Heavy; accept license on HF first |
| Custom… | — | per file | Any GGUF URL/file (BYOM, must fit 2 GB) |
Do not set the 3.2 GB Gemma as the default URL — first-load drop-off and Pi-adjacent bandwidth make it a conversion killer. Keep it as BYOM option.
Option A — I already have a .gguf (recommended for Gemma):
1. Open the model's HuggingFace page, accept its license.
2. Download the .gguf with your browser / scripts/download-model.sh.
3. In the app: Analyze → Enable local AI… → accept → Pick .gguf file.
4. Press Start inference. Refresh the page later — re-pick the file
(browsers don't let sites keep file handles silently; OPFS caching for URLs only).
Option B — paste a model URL:
1. Accept the license on the HF model page first.
2. Copy the direct https://huggingface.co/…/model.gguf link.
3. In the app: paste into the URL field → Use model URL → Start inference.
4. The file caches in your browser (OPFS/Cache) for offline reuse.
| Date | model_id | revision/commit | license (as shown on card) | sha256 (short) | default? |
|---|---|---|---|---|---|
| unfilled | e.g. Qwen/Qwen3-0.6B-GGUF | commit | Apache-2.0 | … | yes |
| 2026-04-10 | local google_gemma-4-E2B-it-Q4_K_M.gguf |
n/a (local file) | verify on source repo before ship | n/a | no (BYOM only) |
scripts/download-model.sh downloads only what you explicitly pass
(MODEL_URL=… ./scripts/download-model.sh --yes). It never runs on build,
never runs in the browser, and refuses gated URLs without a HF token you provide.