An assistant that runs entirely on your phone. Ask it anything — on a plane, on a train, off-grid — and nothing you say ever leaves the device.
ternii and its Brain are in beta. Try it, and use it with caution: it gets things wrong. Brain needs a lot more training, and training means renting GPUs by the hour — that is what Founding Members fund.
Beta. Store links are placeholders while we finish review. Android first; iPhone follows.
Ordinary AI does its sums in fine-grained numbers — expensive to store, expensive to move, expensive in energy. ternii's models think in just three: +1 0 −1. That makes the whole assistant small enough, and frugal enough, to run offline on a phone — no data centre, no bill, no waiting on a network.
A whole assistant in about a gigabyte. It loads into your phone's memory and answers there — nothing to upload, nothing to download per question.
Three-state maths moves far less data per word. On an iPhone 17 Pro, ternii's own model writes at up to 118 tokens per second in the app, and the first token arrives in 190 ms. On a Galaxy S25+, 61–74 tokens per second. Measured on the phones, in the shipping app — not projected.
iPhone 17 Pro typical 108–118 tok/s (bench 112); time to first token measured end-to-end from the model call to the first token emitted. Short runs on a cool phone; ±8 % run to run, so we quote ranges.
The model is on the device. Turn off Wi-Fi and mobile data and it still answers — on a train, on a plane, off-grid.
One phone, one app build, one prompt, only the model changed. Ours is ternary from training and runs on the GPU; a 4-bit model is a compressed copy of a bigger one and has no GPU path in our app — so this is an as-it-ships comparison, and we say so.
Llama-3.2-3B in the same app: 12–18. BitNet 2.4B (ternary, same GPU path): 53–56.
Best in-app run; typical 108–118. Model to model on the same backend: 3.1× Llama.
Llama-3.2-3B in the same app: 6.3 s. Qwen3.5-4B: 8.3 s.
33× and 44× lower time to first token, measured end-to-end.
Llama-3.2-3B: 596 MB. Qwen3.5-4B: 600 MB. From a file half to one-third the size.
The weights are memory-mapped, not loaded — they cost nothing the OS charges you for.
iPhone X (2017): 12.6–16 tokens a second in 240 MB. On a 2022 Galaxy S22, a 4-billion-parameter model manages 0.22 tokens a second — it doesn't usably run. Ours: 17–20.
One 975 MB model file, same hash on every phone.
Not on this page: answer quality. ternii's model is behind cloud-class 4B models on knowledge benchmarks — faster, smaller, runs where they don't, not yet as accurate. Every number above is a short run on a cool phone; we quote ranges, and say when a number is a best run — 118 is the best in-app run; typical is 108–118.
🧠Your name, in the weights. Founding Members put a name or a message on the Founding Roll; if you opt in, it is trained into the next Brain — part of the model, in every copy, on every phone that runs it. From £10, and £10 is 500 million training tokens.→Live screen recordings. Every number traces to a bench.
0:36coming soon
0:25coming soon
0:22coming soon
0:21coming soon
0:22coming soon
0:50coming soonNothing loads until you press play; the video then loads from our servers on Cloudflare. Recorded on the phones named on screen — short runs on a cool phone, as everywhere on this page.
Most phone AI is a big model squeezed down to fit — and it loses accuracy in the squeeze. ternii's model was trained in three states, so the model on your phone is the model we trained. There is no compression step, and nothing to lose in one.
Every weight was −1, 0 or +1 from the first token of training. Nothing is rounded off afterwards to make it fit a phone — the 1 GB file on your phone gives the same answers as the 7 GB file it was packed from.
Measured: the model file on your phone scores the same as the model we trained — held-out perplexity within 0.25 % of the training checkpoint, and identical HellaSwag (40.4 % vs 40.4 % on the same 2,000 tasks) to the 7× larger uncompressed file. 2026-09-08, same texts, same scoring.
Published: a natively-ternary 2-billion-parameter model trained on 4 trillion tokens lands within about one point of a comparable 16-bit model on a ten-task average (54.2 vs 55.2) — Microsoft's published result, not ours.
Microsoft, BitNet b1.58 2B4T technical report (2025), vs Qwen2.5-1.5B. That is the recipe ternii's model follows — ours is smaller, earlier in training, and English-only for now, and we say so.
A tiny model doesn't know everything — so we don't ask it to. ternii's trick is augmentation over size: routing and retrieval put curated facts and live tools in front of the model exactly when a question needs them, and the model writes from those facts only.
Exact arithmetic is computed, not guessed. Questions about your private data, your history, or the future are declined rather than invented. Every grounded answer is cited.
ternii's own model is the fast default. A larger open ternary model is one tap away when you'd trade some speed for depth — and you always see which model answered.
Every reply can reach out to the web for a cited summary — but only when you tap it. Offline is the default; online is opt-in, and only your query ever leaves.
ternii is private by construction, not by promise. There's no account to create, no profile to build, and no server that sees your conversations, your notes, or your files. In airplane mode it works exactly the same.
Everything runs on the device. A single toggle turns web lookups on, with a clear prompt — and only the query leaves, never your chat or your data.
Download and run. No sign-up, no ads, no analytics profile. What's on your phone stays on your phone.
Your session, notes and settings are stored encrypted on the device. You can wipe them any time.
Knowledge lives in small, curated expert packs — one per subject. The app stays light and you add the packs you care about; each is 55–334 KB and installs in a tap. New packs appear over time, with no app update needed.
With the larger model on, several experts can share one question — the model weaves their facts into a single answer and shows you which packs it drew on. A whole panel, deliberating on your device.
A routine gathers from the tools you've connected — calendar, weather, news, prices, your notes — and hands the model only the real data to summarise. No invented meetings, no made-up headlines: if it isn't there, it isn't in the briefing. One tap, a useful answer.
Today's calendar, weather and headlines, in a few lines — before you're out the door.
Tomorrow's schedule and any loose ends from your inbox, tied off for the night.
"Set an alarm for 6:30 tomorrow" — ternii asks once, then hands it to your clock or calendar. You confirm every action.
Top headlines from your chosen feed, condensed to the bullets that matter.
Your watchlist prices with a one-line read — no dashboards to open.
A genuine nugget pulled from one of your experts, explained simply.
Every tool is listed with what it actually does and where the data goes. Green means nothing leaves the phone. Only the "open web" tools touch the internet, and each sends the minimum — a city name, a feed address, a ticker. Anything that acts (an alarm, a message, a call) is handed to the phone's own app for you to confirm. This is the Android list; iPhone matches except where noted.
The 3.4B ternary model, offline. English at launch.
Subject packs of 55–334 KB each — maths, physics, chemistry, biology, medicine, legal, code and more. Install with a tap.
Maths goes to a sandboxed code runner, not a guess — sums, percentages, conversions, dates, parsing data. Nothing leaves the phone.
The clock, here or in a named city. The one thing a model can only guess at, so it doesn't.
Tap Attach to read a PDF, document or image into the chat. Read on the phone.
Reads the text in your recent photos (OCR). It does not describe scenes or identify objects.
Your events, read locally — all calendars on the device.
Looks up a name to a number or email, on request.
Notes you keep, saved on the phone and retrieved by search — never sent anywhere.
ternii can save a verified fact to your own notes so it is found next time. Writes only to the phone.
Any answer can be read aloud by the phone, copied, or exported as PDF or CSV.
Routines are portable files: back one up, move it to a new phone, hand it to a friend.
Talk instead of type, using the phone's own speech recognition.
Morning briefing, evening wind-down, news digest, market check and more — gather real data, summarise it.
Android has no system to-do store to read, so this one exists only on iPhone.
Not in this build — ternii has no Health Connect access yet.
Your own IMAP server with an app password; read on the phone, never through an AI cloud. Providers that only allow OAuth (Outlook, Hotmail, Live) cannot be connected yet.
No app may read these directly, so you share a chat into ternii from the share sheet.
Gated search with readable extracts and citations. Only with Web switched on.
Fetches one page you name and reads its main text. Only with Web switched on.
Keyless forecast for your city (Open-Meteo). Only the city name leaves the phone.
Top headlines from your RSS feed (BBC by default). Only the feed address is fetched.
Keyless quotes for your tickers (Yahoo Finance). Only the symbols leave.
ternii fills in the time; your clock app asks you to confirm.
"Dentist Tuesday 3pm" — your calendar opens with it filled in; you save it.
Handed to the phone's reminders or clock; you confirm.
ternii drafts it to the address you give; your mail app opens; you press send.
No app on Android or iPhone may read, edit or cancel an alarm that already exists. ternii says so rather than pretending, and opens your Clock instead.
Runs a named Shortcut with an input — Apple Shortcuts is iOS-only.
Opens your maps app at the place you named.
Drafts a text or opens the dialler with the number — you press send or call.
ternii drafts, the share sheet opens, you post. It never posts for you.
Opens the link in your browser.
Switch it on and ternii plans multi-step jobs: it chooses a tool, calls it, reads the result, and keeps going — up to a fixed step budget. "Show its work" lists every call it made.
Green tools run on the phone with no question. Amber tools touch the web and run only with Web switched on. Red tools act in the world and always ask you first — every time.
Before answering, the agent checks its draft against what the tools actually returned, and withholds claims a tool did not back.
Tool results are handed to the model on the phone. No conversation is sent to a server to be planned.
Optionally point ternii at an OpenAI-compatible endpoint with your own key. Off by default; every cloud answer is badged so you always know.
Pick short, medium or long answers per chat.
Every part of the name is literally true of what's running on your phone.
Ternary. The models think in three states — +1, 0, −1 — which is the reason a real assistant fits in a gigabyte and runs on a battery.
The bird. The Arctic tern weighs about 100 grams and flies pole to pole every year, with no infrastructure at all. Small, light, goes anywhere. That's on-device AI.
Two rails. Your phone's chip is binary, so each three-state digit rides on a pair of binary signals — two rails. ternii is three states, carried on two rails, on the phone you already own.
Put your name — or a message in your words — on the Founding Roll. If you opt in, it is trained into the next Brain: part of the model, in every copy, on every phone that runs it. The roll's fingerprint goes into every model file we ship, and your entry is issued a certificate on a public blockchain that names the model back. £10 buys 500 million training tokens.
£10 = 500 million tokens
One-off. Worldwide. Priced in pounds.
Names and messages are reviewed by a person before they go on the roll or a certificate is issued: nothing sexual, violent or abusive, no profanity, no other people's personal data. Our decision is final. We'll ask you to change it or refund you.
The name can be anyone's — a gift for a child, a grandparent, a pet. Pseudonyms welcome. The person paying must be 18 or over.
The Founding Roll opens with you.
Every 200 members = 100 billion more training tokens.
Payments open shortly. Everything else is ready — pick your entry now and it's remembered on this device.
Your chosen name — or a message in your words — on the roll, in roll order, with everyone who joined. The roll's fingerprint is written into every model file we ship, and your entry is trained into the next Brain if you opt in.
Two separate opt-ins, both yours to give or keep: shown on the public roll; used in training. Reviewed by a person before it goes on the roll.
Every approved Founding Roll entry is issued a certificate: a unique token on a public blockchain that records your roll number, a fingerprint of your entry, the fingerprint of the roll, and the fingerprint of the ternii model release the roll is bound to. The model file records the roll and the certificate contract in return. We hold it for you; ask and we send it to a wallet you nominate.
It is a certificate of membership, nothing more — not a currency, not a share, not something we sell or trade. Once issued it is public and permanent: it cannot be edited or deleted, by us or by anyone, so choose a name or message you are happy to keep (Terms, 6a).
Every Android and iPhone build, and every model file, reaches Founding Members before anyone else.
The technical film, training notes as they land, the requests board, and the ledger of what your money trained. What's inside →
If we ever open a round, Founding Members hear about it first.
Subject to eligibility and the law.
A ternii model ships as one file — a GGUF, the same format most on-device AI uses. Besides the weights, a GGUF carries a few lines of metadata. Ours will carry the Founding Roll: the roll's fingerprint, how many entries it has, where to read it, the address of the certificate contract, and which chain it is on.
Each approved entry is issued a certificate — a unique token on a public blockchain. It records your roll number, a fingerprint of your entry, the fingerprint of the roll, and the fingerprint of the model file the roll is bound to.
Open the model file: it names the roll and the contract. Open the certificate: it names the roll and the model. Two fingerprints, matching in both directions — checkable by anyone with the file and a browser, with no need to trust us. We looked for another model file that does this and did not find one. That is the first we claim: the two-way binding, as a combination — not the idea of a name in a model.
A fingerprint is a SHA-256 hash: change one byte of the file and it changes completely. The binding covers the shipped model files whose fingerprints are recorded in the contract; each release is bound separately. Every claim on this page is scoped to those files.
We count every £10 as 500 million training tokens — a rate set below what a billion tokens has cost us to train, so it covers failed runs, evaluation, data preparation and price rises. We publish the running total of tokens funded.
Why tokens: the models ternii is compared with learned from trillions — BitNet b1.58 2B4T: 4 trillion; SmolLM3: 11 trillion. Every 200 members = 100 billion more training tokens.
Fees are revenue of Middletech Limited. We intend to spend them on training compute and we publish what we trained, but we do not promise a particular model, release date or result. Early-access builds are pre-release software provided as is.
Payments by Stripe, one-off, in GBP from anywhere; VAT included where it applies.
Founding Membership is a paid membership with the benefits above. It is not a share, a loan, an investment or a donation. It carries no ownership, no voting rights and no financial return. 14-day full refund. Middletech Limited, company no. 16955822. Terms · Privacy.
Every approved entry, in roll order, shown the way its member chose. This is the roll whose fingerprint goes into the model file — the same list, on every phone that runs it.
The roll opens with you — your name here.
a message in your words, up to 500 characters — for someone, about something, or just becauseexample#4 for a grandparent
These four are placeholders to show the shape of the roll, not members. Real entries replace them as they are reviewed and approved.
Brain is a 3.4-billion-parameter model, trained from scratch by a small team, on far fewer tokens than the models it is measured against — they learned from trillions. Brain is faster and smaller than they are, and it runs where they don't; it is not yet as accurate, and we say so on this page.
Founding Members change the arithmetic. Your £10 buys 500 million tokens of training, and you are written into what it buys — on the roll, in the model file, and, if you opt in, in the weights themselves. We publish what we trained, what it cost and what we measured, in the members' area, as it happens. That is the whole deal: no cloud, no account, a small team, and your name on the thing we make.
— Ray, Middletech
Founding Members get a sign-in link by email — no password. Behind it: the work, as it happens, and a board where you tell us what to build.
Not a member yet? Join the Founding Roll — from £10.
The requests board is where members tell us what ternii should do next — a routine, an expert pack, a language, a fix — and vote on each other's. Each request carries a status that moves as we work:
It lives behind sign-in on purpose: members talking to a small team, not a comments section. Ten requests a day each, real names or roll names, and we read every one.
The board is new; what is on it comes from members, not from us. We do not write example requests here — you'll see the real ones once you're in.
Want a pack we don't have yet — beekeeping, tax, a language, your trade? Tell us. Because experts are just downloadable packs, we can add a new one for everyone without an app update.
Found a bug? Want a new routine, or a feature? ternii gets better from what real people ask for. Tell us what happened or what you'd love to see.
A private AI in your pocket — and a first look at what three states can do. Download, and see what a phone can do on its own.