a16z Just Funded Typed, Machine-Native AI. Here Is How to Run a Typed Decision On-Device
aiedgeaillmopensource
TL;DR
Andreessen Horowitz led an 870 million dollar round into typed, machine-native AI, where models return structured decisions that software consumes directly instead of chat text. The part the headline misses: a typed decision is small, deterministic, and private, which makes it the ideal thing to run on-device. No cloud round trip, no data leaving the machine, no per-request bill.
Here is how to do it today on a plain CPU, with open models.
Why a typed decision wants to be on-device
A chat answer is big and open-ended, so it leans on the cloud. A typed decision is the opposite. The output is a label, a class, or a score. The compute is a single forward pass. There is nothing to stream and nothing to parse. That profile fits a laptop or a phone better than a datacenter.
On-device also fixes the three things enterprises worry about: latency, privacy, and cost. The decision runs where the data already is, so it is fast, it stays local, and it is free per call.
Zero tokens makes it cheaper still
Our decision method, ZTC (Zero-Token Confidence), generates no tokens at all. It reads the problem in one forward pass, takes the final hidden state, and applies a calibrated probe. On a CPU that is the difference between a usable local feature and a slow one, because there is no decoding loop to grind through.
Run it on a CPU
POCKET is our family of on-device models, quantized to run locally with no GPU. The GGUF builds have been downloaded more than 800,000 times.
# CPU, no GPU, with llama.cpp
llama-cli -hf FINAL-Bench/POCKET-26B-GGUF -p "..."
# or with Ollama
ollama run hf.co/FINAL-Bench/POCKET-26B-GGUF
For the decision layer, the ZTC code and the S1MB number one model are open as well.
The open stack
- POCKET on-device CPU models: github.com/final-bench/pocket
- ZTC method and inference code: github.com/final-bench/ztc
- Darwin-27B-ZTC-v2, number one of 102 on S1MB: github.com/final-bench/s1mb
FAQ
Can I run a typed decision model without a GPU? Yes. POCKET ships CPU-friendly GGUF builds that run with llama.cpp or Ollama on a normal laptop, and the ZTC decision is a single forward pass.
Why run decisions on-device instead of the cloud? A typed decision is small and deterministic, so on-device gives you lower latency, full privacy, and no per-request cost.
What is a zero-token decision? It is a decision produced in one forward pass with no generated text, by applying a calibrated probe to the model's final hidden state.
Is any of this open source? Yes. POCKET, ZTC, and the S1MB number one model are all public under Apache-2.0.
Built by Ginigen. If this helps, a star on the repositories makes it easier for others to find.