MODELS — local-AI weights policy (BYOM, voluntary only)

The app ships zero model weights. No .gguf is bundled, embedded, or auto-downloaded. The user voluntarily picks a model; weights travel HuggingFace/CDN → user's browser directly, never through our Pi server.

The inference runtime (not weights) is vendored: @wllama/wllama 2.3.1 (MIT, see static/vendor/wllama/LICENCE) ships inside the app shell, so Local AI works fully offline with zero CDN requests.

Hard limits (wllama / browser)

Policy

  1. No redistribution by us. We do not host, proxy, or re-upload weights. rust-embed / the Pi binary contain only the app shell (~MBs).
  2. Explicit consent first. Analyze → Enable local AI… opens a modal with a license checkbox. localStorage["ede.ai.consent"] gates everything in static/ai.js. No consent → deterministic insights.js only.
  3. License follows the file, not the brand. Before commercial use, read the exact model card + LICENSE of the checkpoint you download (Gemma 4 = Apache-2.0 per 2026 docs; Gemma 1–3 = custom Terms + Prohibited Use; quant repacks add the quantizer's terms). Record model_id + revision + license + date below. Not legal advice — consult counsel for rollout.
  4. Gated repos stay gated. If HF shows an Accept-license gate, accept it in the browser on huggingface.co first; never script around the gate.
  5. Privacy. File-picker models stay in tab memory; URL models cache in the user's OPFS/Cache. Prompts sent to the local model contain schema + aggregates only (see static/ai.js), never raw PII rows.

Supported models (curated catalog — also the in-app dropdown)

Source of truth: static/ai-models.js (sizes verified 2026-09-04 via HF content-range). RAM ≈ 1.2× file size in-browser.

Model (Q4_K_M) Size License Notes
Qwen2.5 0.5B Instruct — recommended default 469 MB Apache-2.0 Fast on laptop CPU; NL→SQL + short summaries
Llama 3.2 1B Instruct 770 MB Llama 3.2 Community Better following, 128K ctx; custom license
Qwen2.5 1.5B Instruct 1.04 GB Apache-2.0 Smarter, ~2× slower; comfort ceiling
SmolLM2 1.7B Instruct 1.00 GB Apache-2.0 Strong per-param; slow on old CPUs
Gemma 2 2B IT 1.59 GB Gemma Terms of Use Heavy; accept license on HF first
Custom… — per file Any GGUF URL/file (BYOM, must fit 2 GB)

Do not set the 3.2 GB Gemma as the default URL — first-load drop-off and Pi-adjacent bandwidth make it a conversion killer. Keep it as BYOM option.

How to install (user-facing)

Option A — I already have a .gguf (recommended for Gemma): 1. Open the model's HuggingFace page, accept its license. 2. Download the .gguf with your browser / scripts/download-model.sh. 3. In the app: Analyze → Enable local AI… → accept → Pick .gguf file. 4. Press Start inference. Refresh the page later — re-pick the file (browsers don't let sites keep file handles silently; OPFS caching for URLs only).

Option B — paste a model URL: 1. Accept the license on the HF model page first. 2. Copy the direct https://huggingface.co/…/model.gguf link. 3. In the app: paste into the URL field → Use model URL → Start inference. 4. The file caches in your browser (OPFS/Cache) for offline reuse.

Maintainer log (fill per release)

Date model_id revision/commit license (as shown on card) sha256 (short) default?
unfilled e.g. Qwen/Qwen3-0.6B-GGUF commit Apache-2.0 … yes
2026-04-10 local google_gemma-4-E2B-it-Q4_K_M.gguf n/a (local file) verify on source repo before ship n/a no (BYOM only)

Script

scripts/download-model.sh downloads only what you explicitly pass (MODEL_URL=… ./scripts/download-model.sh --yes). It never runs on build, never runs in the browser, and refuses gated URLs without a HF token you provide.


← back to Local Analyst