Offline AI on a Phone: Thread Pinning, Lazy Loading, and a Safety Gate That Knows When to Stay Quiet
aiandroidllmedgeai
TL;DR
- HeliGO is Ginigen's offline disaster and survival app. In airplane mode it still gives you maps, contour lines, altitude, a compass, rescue coordinates, and an AI chat, all running from data and a model stored on the phone.
- The model is Edge-4B-TELL, a 4-bit GGUF build based on Google's Gemma-4 E4B, published on Hugging Face (1,526 downloads and 35 likes at the time of writing).
- Thread count is not a tuning detail on Android. With the screen off, our process was allowed only 4 cores. Whisper with 6 threads took 49.4 s; with 4 threads it took 2.1 s on the same phone (Galaxy S25). Read
Cpus_allowed_listand size your thread pool from it. - Lazy-load rare features. Making the 0.92 GB vision projector always resident pushed model load time from about 10 s to about 1 minute, and slowed plain text chat too.
- Do not trust a model's stated confidence. Asked to self-report confidence, the model was more confident on wrong answers (AUROC 0.441, worse than a coin flip). HeliGO uses a three-stage safety gate, and our TELL signal reads the model's internal state instead of its opinion.
- Bonus from our sister company VIDRAFT: POCKET-Darwin-180B-GGUF shows the same "quantize without losing quality" idea at 180B scale: 111.3 GB, MMLU-Pro 87.65% vs 87.65% for the BF16 original.
Why build an AI app that never touches the network?
In a disaster, the network is usually the first thing to go. Cell towers fail in an earthquake. A hiker who loses the trail in the mountains loses signal at the same time. Yet the information you need right then (where to evacuate, how to treat an injury, which plant or mushroom not to eat) mostly lives on the internet.
HeliGO flips that assumption. Everything the app needs ships inside the phone. If you turn on airplane mode, nothing degrades. That is not a fallback mode; it is the only mode we design for.
Here is what is stored on the device:
| Dataset | Records |
|---|---|
| Terrain tiles with elevation (auto-switching contour lines) | 4,998 |
| Summit elevations | 16,568 |
| Shelters, medical, water supply, and fire facilities | 1,255 |
| Korea Ministry of the Interior and Safety (MOIS) public action guidelines | 75 |
| Venomous species warnings | 12 |
On top of that: nearest high ground for flood situations, sunrise and sunset, moon phase, and breadcrumb backtracking along the path you walked.
The AI is the newest layer, and it is also the one that needs the most care. Maps do not hallucinate. Language models do.
How do you run an LLM fully offline on a phone?
The short answer: a quantized GGUF model, llama.cpp compiled for ARM64, and a lot of discipline about memory and threads.
Our stack for the on-device model:
- Model: Edge-4B-TELL. The inference weights are a verbatim mirror of Google's
gemma-4-E4B-it-qat-q4_0-gguf. We host our own copy so that a disaster app does not break if an upstream path moves. - Files: a 4.8 GB text-generation GGUF and a 0.92 GB vision and audio projector (mmproj).
- Runtime: llama.cpp, running as a native process under the app.
- Speech: whisper.cpp for voice input.
Quantization-aware training (QAT) at 4 bits is what makes a 4B-class multimodal model fit a phone at all. A GGUF file packs the weights and the metadata llama.cpp needs, so the app can load it straight from local storage with no conversion step.
A generic way to try the same pattern on a development device (not our app's code, just the standard llama.cpp workflow):
# Build llama.cpp for Android with the NDK (on your workstation)
cmake -B build-android \
-DCMAKE_TOOLCHAIN_FILE=$ANDROID_NDK/build/cmake/android.toolchain.cmake \
-DANDROID_ABI=arm64-v8a -DANDROID_PLATFORM=android-28 \
-DGGML_OPENMP=OFF
cmake --build build-android -j --target llama-server llama-cli
# Push binary and model to the device
adb push build-android/bin/llama-server /data/local/tmp/
adb push gemma-4-E4B_q4_0-it.gguf /data/local/tmp/
# Run fully offline; bind to localhost only
adb shell "cd /data/local/tmp && ./llama-server \
-m gemma-4-E4B_q4_0-it.gguf -c 4096 -t 4 \
--host 127.0.0.1 --port 8080"
That flag -t 4 looks innocent. It is the single most important number in this article.
Why does thread count matter for on-device inference on Android?
Here is the incident. Our voice pipeline ran whisper with 4 threads. The Galaxy S25 has more cores than that, so we raised it to 6, expecting a speedup.
Translation latency jumped to about 56 seconds.
The cause: when the phone screen is off, Android moves a foreground-service app and its native child processes into a restricted cpuset. On the S25 that set was 4 cores. ggml's thread pool uses spin-waiting, and when you run more worker threads than you have cores, threads that are spinning steal time from threads that are doing real work. The pool does not degrade gracefully. It collapses.
Measured on the same phone, same 4 allowed cores, screen off:
| whisper threads | Time |
|---|---|
-t 6 |
49.4 s |
-t 4 |
2.1 s |
That is roughly a 23x difference from one integer.
Why our benchmarks missed it
We benchmarked through adb shell run-as <package>. A process started that way inherits the shell's cpuset, which allows all cores. So the benchmark said 6 threads was fine. The real app, with the screen off, saw only 4. adb benchmarks hide this class of bug.
The fix: read the allowed cores at runtime
Do not hardcode thread counts, and do not use the total core count. Ask the kernel what this process is actually allowed to use:
# Inside the app's own process context
grep Cpus_allowed_list /proc/self/status
# Cpus_allowed_list: 0-1,4-5 (example: 4 cores)
A small helper that turns that into a thread count:
fun allowedCoreCount(): Int {
val line = java.io.File("/proc/self/status").readLines()
.firstOrNull { it.startsWith("Cpus_allowed_list") } ?: return 4
val spec = line.substringAfter(":").trim()
return spec.split(",").sumOf { part ->
val r = part.split("-").map { it.trim().toInt() }
if (r.size == 2) r[1] - r[0] + 1 else 1
}
}
val threads = minOf(allowedCoreCount(), 4)
// pass "-t $threads" to whisper / llama.cpp
Two practical rules we now follow:
- Compute threads from
Cpus_allowed_list, and cap them. Re-check when the app moves between foreground and background if your workload can run in both. - Benchmark under the real restriction. Pin the process to the same cores the app will get (for example
taskset 33for cores 0, 1, 4, 5) and test with the screen off, not only from an adb shell.
# Reproduce the restricted cpuset during a benchmark (mask 0x33 = cores 0,1,4,5)
adb shell "taskset 33 /data/local/tmp/whisper-cli -m ggml-base.bin -f test.wav -t 4"
Why should rare features be lazy-loaded on mobile?
The second lesson came from a fix that made things worse.
The vision projector (the mmproj file, 0.92 GB) is what lets the model look at a photo. We had a bug in the photo feature, and while fixing it we made the projector always load with the model. The photo feature worked again. But model load time went from about 10 seconds to about 1 minute, and ordinary text questions slowed down too.
What went wrong was not the code. It was the decision. We looked only at the broken feature and never measured the main path. Users chat with text all the time. They send a photo occasionally. "Keep it loaded so it is always ready" sounds safe, but it charges every user, every session, for a feature most sessions never touch.
The rule we adopted:
- Before adding anything to global initialization, ask how often it is used.
- Rare features load when first used (lazy loading), and can be released afterwards under memory pressure.
- Measure the main path (cold start to first token for a text question) before and after every change. If you did not measure it, you did not fix it.
# Lazy pattern with llama.cpp: start text-only by default
./llama-server -m model.gguf -t $THREADS # fast cold start
# Only when the user opens the camera feature, start (or switch to) a
# vision-capable instance with the projector:
./llama-server -m model.gguf --mmproj mmproj.gguf -t $THREADS
On a phone, memory and load time are the product. Every megabyte you keep resident is a megabyte the OS can use as a reason to kill you in the background.
How do you make an on-device AI refuse safely?
In a disaster app, a wrong answer can hurt someone. So HeliGO never passes the model's output straight to the user. Every answer goes through a three-stage safety gate:
- Block life-threatening answers, regardless of confidence. An answer like "this is safe to eat" is stopped even if the model sounds certain.
- Answer from official text when it applies. If the question matches one of the 75 MOIS public action guidelines, the app answers with the guideline text instead of the model's paraphrase.
- Abstain when uncertain. If neither of the above settles it and the answer is not reliable, the app says it does not know.
The asymmetry principle
The gate is deliberately lopsided: it never says "safe", and it always says "poisonous".
The two possible errors do not cost the same. If the app wrongly withholds "safe", you skip a meal and stay hungry. If it wrongly says "safe" about something venomous or toxic, someone can die. When silence and speech carry such different costs, the design should not treat them symmetrically. This is less a machine learning decision than a product ethics decision, and we think every safety-critical on-device AI needs one written down.
Can you trust an LLM when it says how confident it is?
No, and our measurements were sharper than we expected.
We prompted the model to state its confidence with each answer on 665 Korean disaster-procedure questions. The average stated confidence was 0.863. When we used that number to separate correct from incorrect answers, the AUROC was 0.441. Anything below 0.5 means the signal points the wrong way: the model tended to sound more certain on the answers it got wrong.
So "Are you sure?" is not a weak safety check. It is a misleading one.
That is why we built TELL. Instead of asking the model for its opinion, TELL reads the model's internal state during generation and estimates whether the answer is likely to be correct. The readout ships alongside the weights in the Edge-4B-TELL repository, and the public numbers on the model card (held-out cross-validation on the same 665 questions) are:
| Signal | AUROC |
|---|---|
| Self-reported confidence | 0.441 |
| Surface features (length, digits, formatting) | 0.736 |
| TELL (internal-state readout) | 0.759 |
We publish the surface-feature baseline on purpose. TELL beats it by a real but modest margin, and we would rather show that honestly than claim a bigger win. Inside the app, the comparison that matters is TELL versus having no calibration signal at all.
Does 4-bit quantization destroy model quality?
Not if you are careful, and the evidence goes well beyond 4B.
Our sister company VIDRAFT recently released POCKET-Darwin-180B-GGUF, a 4-bit GGUF build of the 180B mixture-of-experts model Darwin-180B-RSI. Public facts from its model card:
- Size: 111.3 GB in 4 files, down from 360 GB for the BF16 original.
- Base layout: built on the widely used UD-Q4_K_XL GGUF layout.
- Quality: MMLU-Pro, 2,000 paired questions: 87.65% vs 87.65% for the BF16 original.
- Shorter answers: on the same 2,000 questions, answers averaged 3,694 tokens vs 4,322 for the same-format parent build.
- Hardware: runs with llama.cpp (b11048 or newer), including a gaming laptop with an 8 GB GPU and 32 GB RAM (4.17 tokens/s) and a CPU-only server (18 to 21 tokens/s).
- Sparsity: about 3B active parameters per token, with experts read on demand from SSD through memory mapping.
The phone and the 180B laptop build share one idea: edge AI is a systems problem, not just a model problem. Quantization makes the model fit. Threads, memory mapping, and load paths decide whether it is actually usable.
What is the practical checklist for shipping on-device LLMs?
- Quantize with a format your runtime loads natively (GGUF for llama.cpp), and verify quality against the full-precision model on paired questions.
- Size thread pools from
/proc/self/statusCpus_allowed_list, not from the core count. - Benchmark with the screen off, under the real cpuset, not only through
adb shell run-as. - Lazy-load rare modalities (vision projectors, extra adapters). Measure cold start to first token before and after every change.
- Never ship raw model output in a safety-critical domain. Put a gate in front: block, defer to official text, or abstain.
- Do not use self-reported confidence as a safety signal. Validate any confidence signal against a simple surface baseline.
FAQ
Does HeliGO really work in airplane mode? Yes. Maps, contours, altitude, compass, rescue coordinates, facility lookups, and AI chat all run from data and a model stored on the phone. HeliGO does not replace emergency services; it is a tool for when you cannot reach them.
Which model runs inside HeliGO? Edge-4B-TELL on Hugging Face. The inference weights are an unmodified mirror of Google's Gemma-4 E4B QAT Q4_0 GGUF; Ginigen's contribution is the TELL calibration readout shipped with it.
How many threads should I give llama.cpp or whisper.cpp on Android?
No more than the number of cores in your process's Cpus_allowed_list. On a Galaxy S25 with the screen off that was 4, and going to 6 made whisper about 23x slower.
Why not just ask the model if it is sure? Because on our disaster-procedure questions the model's stated confidence ranked answers worse than chance (AUROC 0.441). Wrong answers came with higher stated confidence.
Can POCKET-Darwin-180B-GGUF run without a GPU? According to its model card, yes: 18 to 21 tokens/s on one server CPU socket with llama.cpp b11048 or newer, and it also runs on a laptop with an 8 GB GPU and 32 GB RAM.
Why does the safety gate never say "safe"? Because the costs are asymmetric. Wrongly withholding "safe" costs a meal; wrongly saying "safe" can cost a life.
Links
- Ginigen: https://www.ginigen.net
- Edge-4B-TELL model: https://huggingface.co/ginigen-ai/Edge-4B-TELL
- POCKET-Darwin-180B-GGUF (VIDRAFT): https://huggingface.co/FINAL-Bench/POCKET-Darwin-180B-GGUF