An assistant that runs entirely on your phone. Ask it anything — on a plane, on a train, off-grid — and nothing you say ever leaves the device.
Store links are placeholders while we finish review. Android first; iPhone follows.
Ordinary AI does its sums in fine-grained numbers — expensive to store, expensive to move, expensive in energy. ternii's models think in just three: +1 0 −1. That makes the whole assistant small enough, and frugal enough, to run offline on a phone — no data centre, no bill, no waiting on a network.
A whole assistant in about a gigabyte. It loads into your phone's memory and answers there — nothing to upload, nothing to download per question.
Three-state maths moves far less data per word. On an iPhone 17 Pro, ternii's own model writes at up to 118 tokens per second in the app, and the first token arrives in 190 ms. On a Galaxy S25+, 61–74 tokens per second. Measured on the phones, in the shipping app — not projected.
iPhone 17 Pro typical 108–118 tok/s (bench 112); time to first token measured end-to-end from the model call to the first token emitted. Short runs on a cool phone; ±8 % run to run, so we quote ranges.
The model is on the device. Turn off Wi-Fi and mobile data and it still answers — on a train, on a plane, off-grid.
One phone, one app build, one prompt, only the model changed. Ours is ternary from training and runs on the GPU; a 4-bit model is a compressed copy of a bigger one and has no GPU path in our app — so this is an as-it-ships comparison, and we say so.
Llama-3.2-3B in the same app: 12–18. BitNet 2.4B (ternary, same GPU path): 53–56.
Best in-app run; typical 108–118. Model to model on the same backend: 3.1× Llama.
Llama-3.2-3B in the same app: 6.3 s. Qwen3.5-4B: 8.3 s.
33× and 44× lower time to first token, measured end-to-end.
Llama-3.2-3B: 596 MB. Qwen3.5-4B: 600 MB. From a file half to a third the size.
The weights are memory-mapped, not loaded — they cost nothing the OS charges you for.
iPhone X (2017): 12.6–16 tokens a second in 240 MB. On a 2022 Galaxy S22, a 4-billion-parameter model manages 0.22 tokens a second — it doesn't usably run. Ours: 17–20.
One 975 MB model file, same hash on every phone.
Not on this page: answer quality. ternii's model is behind cloud-class 4B models on knowledge benchmarks — faster, smaller, runs where they don't, not yet as accurate. Every number above is a short run on a cool phone; we quote ranges, and say when a number is a best run — 118 is the best in-app run; typical is 108–118.
Most phone AI is a big model squeezed down to fit — and it loses accuracy in the squeeze. ternii's model was trained in three states, so the model on your phone is the model we trained. There is no compression step, and nothing to lose in one.
Every weight was −1, 0 or +1 from the first token of training. Nothing is rounded off afterwards to make it fit a phone — the 1 GB file on your phone gives the same answers as the 7 GB file it was packed from.
Measured: the model file on your phone scores the same as the model we trained — held-out perplexity within 0.25 % of the training checkpoint, and identical HellaSwag (40.4 % vs 40.4 % on the same 2,000 tasks) to the 7× larger uncompressed file. 2026-09-08, same texts, same scoring.
Published: a natively-ternary 2-billion-parameter model trained on 4 trillion tokens lands within about one point of a comparable 16-bit model on a ten-task average (54.2 vs 55.2) — Microsoft's published result, not ours.
Microsoft, BitNet b1.58 2B4T technical report (2025), vs Qwen2.5-1.5B. That is the recipe ternii's model follows — ours is smaller, earlier in training, and English-only for now, and we say so.
A tiny model doesn't know everything — so we don't ask it to. ternii's trick is augmentation over size: routing and retrieval put curated facts and live tools in front of the model exactly when a question needs them, and the model writes from those facts only.
Exact arithmetic is computed, not guessed. Questions about your private data, your history, or the future are declined rather than invented. Every grounded answer is cited.
ternii's own model is the fast default. A larger open ternary model is one tap away when you'd trade some speed for depth — and you always see which model answered.
Every reply can reach out to the web for a cited summary — but only when you tap it. Offline is the default; online is opt-in, and only your query ever leaves.
ternii is private by construction, not by promise. There's no account to create, no profile to build, and no server that sees your conversations, your notes, or your files. In airplane mode it works exactly the same.
Everything runs on the device. A single toggle turns web lookups on, with a clear prompt — and only the query leaves, never your chat or your data.
Download and run. No sign-up, no ads, no analytics profile. What's on your phone stays on your phone.
Your session, notes and settings are stored encrypted on the device. You can wipe them any time.
Knowledge lives in small, curated expert packs — one per subject. The app stays light and you add the packs you care about; each is 55–334 KB and installs in a tap. New packs appear over time, with no app update needed.
With the larger model on, several experts can share one question — the model weaves their facts into a single answer and shows you which packs it drew on. A whole panel, deliberating on your device.
A routine gathers from the tools you've connected — calendar, weather, news, prices, your notes — and hands the model only the real data to summarise. No invented meetings, no made-up headlines: if it isn't there, it isn't in the briefing. One tap, a useful answer.
Today's calendar, weather and headlines, in a few lines — before you're out the door.
Tomorrow's schedule and any loose ends from your inbox, tied off for the night.
"Set an alarm for 6:30 tomorrow" — ternii asks once, then hands it to your clock or calendar. You confirm every action.
Top headlines from your chosen feed, condensed to the bullets that matter.
Your watchlist prices with a one-line read — no dashboards to open.
A genuine nugget pulled from one of your experts, explained simply.
Every part of the name is literally true of what's running on your phone.
Ternary. The models think in three states — +1, 0, −1 — which is the reason a real assistant fits in a gigabyte and runs on a battery.
The bird. The Arctic tern weighs about 100 grams and flies pole to pole every year, with no infrastructure at all. Small, light, goes anywhere. That's on-device AI.
Two rails. Your phone's chip is binary, so each three-state value rides on a pair of binary signals — two rails. ternii is three states, carried on two rails, on the phone you already own.
Want a pack we don't have yet — beekeeping, tax, a language, your trade? Tell us. Because experts are just downloadable packs, we can add a new one for everyone without an app update.
Found a bug? Want a new routine, or a feature? ternii gets better from what real people ask for. Tell us what happened or what you'd love to see.
A private AI in your pocket — and a first look at what three states can do. Download, and see what a phone can do on its own.