An assistant that runs entirely on your phone. Ask it anything — on a plane, on a train, off-grid — and nothing you say ever leaves the device. Free, no account.
Brain is not fully trained yet. The speed, the size and the privacy above are finished work — test all of it today. What is not finished is how much it knows: it is not yet as accurate as bigger models, and it does get things wrong, so please bear with us and check anything that matters. More training is the whole fix, and training means renting GPUs by the hour — which is what Founding Members pay for.
Ordinary AI does its sums in fine-grained numbers — expensive to store, expensive to move, expensive in energy. ternii's models think in just three: +1 0 −1. That makes the whole assistant small enough, and frugal enough, to run offline on a phone — no data centre, no bill, no waiting on a network. The speed comes from a second thing as well: the model is built as a panel of experts, and only the ones a question needs wake up.
A whole assistant in about a gigabyte. It loads into your phone's memory and answers there — nothing to upload, nothing to download per question.
Two things make it quick. Three-state maths moves far less data per word — and only about a billion of Brain's 3.4 billion parameters run for any one word, because it is a panel of experts and a question wakes only the ones it needs. The models it is measured against are dense: every parameter, every word. On an iPhone 17 Pro, ternii's own model writes at up to 118 tokens per second in a chat turn in the app, and the first token arrives in 190 ms on that phone. On a Galaxy S25+, 61–74 tokens per second, also a chat turn in the app. Measured on the phones, in the shipping app — not projected.
iPhone 17 Pro typical 108–118 tok/s (bench 112); time to first token measured end-to-end from the model call to the first token emitted. Short runs on a cool phone; ±8 % run to run, so we quote ranges.
The model is on the device. Turn off Wi-Fi and mobile data and it still answers — on a train, on a plane, off-grid.
One phone, one app build, one prompt, only the model changed. Ours is ternary from training and runs on the GPU; a 4-bit model is a compressed copy of a bigger one and has no GPU path in our app — so this is an as-it-ships comparison, and we say so.
Llama-3.2-3B in the same app: 12–18. BitNet 2.4B (ternary, same GPU path): 53–56.
Best in-app run on an iPhone 17 Pro, in a chat turn; typical 108–118. Model to model on the same backend: 3.1× Llama. BitNet is ternary too and takes the same GPU path — it is dense, where ternii's model wakes only the experts a question needs.
Llama-3.2-3B in the same app: 6.3 s. Qwen3.5-4B: 8.3 s.
33× and 44× lower time to first token, measured end-to-end. Why: reading your question is the part that takes the time, and Brain reads about 1,500 tokens a second of it on this phone — a bench figure, kernel only. On a Galaxy S25+, where we have all four models on one bench, it reads 1.6× faster than Llama-3.2-3B and 2.6× faster than Qwen3.5-4B.
Llama-3.2-3B: 596 MB. Qwen3.5-4B: 600 MB. From a file half the size of Llama's and 2.7× smaller than Qwen's.
The weights are memory-mapped, not loaded — they cost nothing the OS charges you for.
iPhone X (2017): 12.6–16 tokens a second in 240 MB, in a chat turn. On a 2022 Galaxy S22, a 4-billion-parameter model manages 0.22 tokens a second on the bench — it doesn't usably run. Ours, on that same bench: 25–32.
One 975 MB model file, same hash on every phone.
Not on this page: answer quality. ternii's model is behind cloud-class 4B models on knowledge benchmarks — faster, smaller, runs where they don't, not yet as accurate. Every number above is a short run on a cool phone; we quote ranges, and say when a number is a best run — 118 is the best in-app run; typical is 108–118.
🧠Your name, in the weights. Founding Members put a name or a message on the Founding Roll; if you opt in, it goes into the data the next Brain is trained on. From £10, and £10 is 500 million training tokens.→Live screen recordings. Every number traces to a bench. The figures on these clips are an iPhone 17 Pro in the app, in a chat turn, unless the clip names another phone .
0:23coming soon
0:23coming soon
0:19coming soon
0:19coming soon
0:18coming soon
0:39coming soonNothing loads until you press play; the video then loads from our servers on Cloudflare. Recorded on the phones named on screen — short runs on a cool phone, as everywhere on this page.
Most phone AI is a big model squeezed down to fit — and it loses accuracy in the squeeze. ternii's model was trained in three states, so the model on your phone is the model we trained. There is no compression step, and nothing to lose in one.
Every weight was −1, 0 or +1 from the first token of training. Nothing is rounded off afterwards to make it fit a phone — the 1 GB file on your phone gives the same answers as the 7 GB file it was packed from.
Measured: the model file on your phone scores the same as the model we trained — held-out perplexity within 0.25 % of the training checkpoint, and identical HellaSwag (40.4 % vs 40.4 % on the same 2,000 tasks) to the uncompressed training file, nearly 7× larger. 2026-09-08, same texts, same scoring.
Published: a natively-ternary 2-billion-parameter model trained on 4 trillion tokens lands within about one point of a comparable 16-bit model on a ten-task average (54.2 vs 55.2) — Microsoft's published result, not ours.
Microsoft, BitNet b1.58 2B4T technical report (2025), vs Qwen2.5-1.5B. That is the recipe ternii's model follows — ours is smaller, earlier in training, and English-only for now, and we say so.
A tiny model doesn't know everything — so we don't ask it to. ternii's trick is augmentation over size: routing and retrieval put curated facts and live tools in front of the model exactly when a question needs them, and the model writes from those facts only.
Exact arithmetic is computed, not guessed. Questions about your private data, your history, or the future are declined rather than invented. Every grounded answer is cited.
ternii's own model is the fast default. A larger open ternary model is one tap away when you'd trade some speed for depth — and you always see which model answered.
Every reply can reach out to the web for a cited summary — but only when you tap it. Offline is the default; online is opt-in, and only your query ever leaves.
ternii is private by construction, not by promise. There's no account to create, no profile to build, and no server that sees your conversations, your notes, or your files. In airplane mode it works exactly the same.
Everything runs on the device. A single toggle turns web lookups on, with a clear prompt — and only the query leaves, never your chat or your data.
Download and run. No sign-up, no ads, no analytics profile. What's on your phone stays on your phone.
Your session, notes and settings are stored encrypted on the device. You can wipe them any time.
Knowledge lives in small, curated expert packs — one per subject. The app stays light and you add the packs you care about; each is 55–334 KB and installs in a tap. New packs appear over time, with no app update needed.
With the larger model on, several experts can share one question — the model weaves their facts into a single answer and shows you which packs it drew on. A whole panel, deliberating on your device.
A routine gathers from the tools you've connected — calendar, weather, news, prices, your notes — and hands the model only the real data to summarise. No invented meetings, no made-up headlines: if it isn't there, it isn't in the briefing. One tap, a useful answer.
Today's calendar, weather and headlines, in a few lines — before you're out the door.
Tomorrow's schedule and any loose ends from your inbox, tied off for the night.
"Set an alarm for 6:30 tomorrow" — ternii asks once, then hands it to your clock or calendar. You confirm every action.
Top headlines from your chosen feed, condensed to the bullets that matter.
Your watchlist prices with a one-line read — no dashboards to open.
A genuine nugget pulled from one of your experts, explained simply.
Every tool is listed with what it actually does and where the data goes. Green means nothing leaves the phone. Only the "open web" tools touch the internet, and each sends the minimum — a city name, a feed address, a ticker. Anything that acts (an alarm, a message, a call) is handed to the phone's own app for you to confirm. Each card says which phones it runs on today; the iPhone build trails Android, and where a tool differs it is marked.
The 3.4B ternary model, offline. English at launch.
Subject packs of 55–334 KB each — maths, physics, chemistry, biology, medicine, legal, code and more. Install with a tap.
Maths goes to a sandboxed code runner, not a guess — sums, percentages, conversions, dates, parsing data. Nothing leaves the phone.
The clock, here or in a named city. The one thing a model can only guess at, so it doesn't.
Tap Attach to read a PDF, document or image into the chat. Read on the phone.
Reads the text in your recent photos (OCR). It does not describe scenes or identify objects.
Your events, read locally — all calendars on the device.
Looks up a name to a number or email, on request.
Notes you keep, saved on the phone and retrieved by search — never sent anywhere.
ternii can save a verified fact to your own notes so it is found next time. Writes only to the phone.
Any answer can be read aloud by the phone, copied, or exported as PDF or CSV.
Routines are portable files: back one up, move it to a new phone, hand it to a friend.
Talk instead of type, using the phone's own speech recognition.
Morning briefing, evening wind-down, news digest, market check and more — gather real data, summarise it.
On iPhone, your Reminders list. On Android there is no system reminders store, so ternii reads your calendar's alarmed events for the next 7 days instead.
Not in this build — ternii has no Health Connect access yet.
Your own IMAP server with an app password; read on the phone, never through an AI cloud. Providers that only allow OAuth (Outlook, Hotmail, Live) cannot be connected yet.
No app may read these directly, so you share a chat into ternii from the share sheet.
Gated search with readable extracts and citations. Only with Web switched on.
Fetches one page you name and reads its main text. Only with Web switched on.
Keyless forecast for your city (Open-Meteo). Only the city name leaves the phone.
Top headlines from your RSS feed (BBC by default). Only the feed address is fetched.
Keyless quotes for your tickers (Yahoo Finance). Only the symbols leave.
ternii fills in the time; your clock app asks you to confirm.
"Dentist Tuesday 3pm" — your calendar opens with it filled in; you save it.
Handed to the phone's reminders or clock; you confirm.
ternii drafts it to the address you give; your mail app opens; you press send.
No app on Android or iPhone may read, edit or cancel an alarm that already exists. ternii says so rather than pretending, and opens your Clock instead.
Runs a named Shortcut with an input — Apple Shortcuts is iOS-only.
Opens your maps app at the place you named.
Drafts a text or opens the dialler with the number — you press send or call.
ternii drafts, the share sheet opens, you post. It never posts for you.
Opens the link in your browser.
Switch it on and ternii plans multi-step jobs: it chooses a tool, calls it, reads the result, and keeps going — up to a fixed step budget. "Show its work" lists every call it made.
Green tools run on the phone with no question. Amber tools touch the web and run only with Web switched on. Red tools act in the world and always ask you first — every time.
Before answering, the agent checks its draft against what the tools actually returned, and withholds claims a tool did not back.
Tool results are handed to the model on the phone. No conversation is sent to a server to be planned.
Optionally point ternii at an OpenAI-compatible endpoint with your own key. Off by default; every cloud answer is badged so you always know.
Pick short, medium or long answers per chat.
Every part of the name is literally true of what's running on your phone.
Ternary. The models think in three states — +1, 0, −1 — which is the reason a real assistant fits in a gigabyte and runs on a battery.
The bird. The Arctic tern weighs about 100 grams and flies pole to pole every year, with no infrastructure at all. Small, light, goes anywhere. That's on-device AI.
Two rails. Your phone's chip is binary, so each three-state digit rides on a pair of binary signals — two rails. ternii is three states, carried on two rails, on the phone you already own.
We would rather tell you this plainly than have you find it out. ternii is fast, small and private today, and you can put every one of those claims to the test yourself. What it does not yet have is breadth of knowledge — because Brain simply has not read enough yet.
Hold us to every one of these. They are measured, and they are on the phone in your hand. Figures are from a cool device on a short run — like any phone, it slows down once it warms up.
This is the honest half. It is a training problem, not a design problem, and it is fixed by more tokens.
Reading is the only thing that fixes the right-hand column, and reading costs GPU hours. That is the whole reason we are asking for support: every £10 buys 500 million training tokens for the next Brain, and the more of them we can buy, the sooner ternii knows enough to match how well it already runs. Help us speed up its training — and keep testing the rest of it while we do.
Put your name — or a message in your words — on the Founding Roll. If you opt in, it goes into the data the next Brain is trained on. From the next release on, the roll's fingerprint will be written into the model files we ship, and your entry gets a signed certificate that names that model back. £10 buys 500 million training tokens.
Your chosen name — or a message in your words — on the roll, in roll order, with everyone who joined. The roll's fingerprint will be written into the model files we ship from the next release, and your entry is trained into the next Brain if you opt in.
Two separate opt-ins, both yours to give or keep: shown on the public roll; used in training. Reviewed by a person before it goes on the roll.
Every approved Founding Roll entry gets a signed certificate — number, your entry's fingerprint, the roll's fingerprint, and the fingerprint of the ternii model release it is bound to. The model files we ship from the next release will record the roll in return. Signed by Middletech, anchored publicly, and checkable by anyone at ternii.com/founders/verify — including against the model file itself.
If you would like it as a token in your own wallet, just ask and we'll mint one — no wallet needed otherwise.
It is a certificate of membership, nothing more — not a currency, not a share, not something we sell or trade. What is anchored publicly is a fingerprint of each roll version, never your name; if you withdraw, your certificate is marked withdrawn in the registry (Terms, 6a).
Every Android and iPhone build, and every model file, reaches Founding Members before anyone else.
The technical film, training notes as they land, the requests board, and the ledger of what your money trained. What's inside →
4.3 billion parameters, and quicker than the one you have. Only about one in five of them fires for any one word — that is what mixture-of-experts buys you — so the model can grow while the work per word goes down.
We have built it and measured it: it is quicker than Brain 1.0 on the same phone, in the same app, and it holds four times the conversation in mind. Decode speed does not depend on what the weights have learned, which is why this can be measured before the model is trained.
On an iPhone 17 Pro, in the app with agent mode and tools switched on — the heaviest thing we ask of it — Brain 2.0 writes at 126–140 tokens a second where Brain 1.0 manages 80. On the bench, away from the app, 2.0 reached 147 and 168 tok/s on two runs. Two reasons it pulls ahead: it reads about 17 % fewer bytes for every word it writes, and it can keep its memory of the conversation in a compressed form that Brain 1.0's shape cannot — a design decision taken months ago, paying off. And when the iPhone 18 Pro arrives we want to know whether 200 is within reach.
Both figures are from the same phone in the same session, warm, generating a real answer — the only thing changed between them is the model. Speed figures move with the length of the prompt and the answer, so we quote ranges. The model is trained only far enough to measure how fast it runs; teaching it is what Founding Members pay for. The 4.3 B file is bigger than today's and will cost more memory to run than Brain 1.0's 195 MB on that phone — we will publish that number when we measure it, like every other one here.
If we ever open a round, Founding Members hear about it first.
Subject to eligibility and the law.
A ternii model ships as one file — a GGUF, the same format most on-device AI uses. Besides the weights, a GGUF carries a few lines of metadata. Ours will carry the Founding Roll: the roll's fingerprint, how many entries it has, where to read it, and the certificate registry (its public key and anchor).
Each approved entry gets a signed certificate — signed by Middletech with a published key, and anchored publicly for each roll version. It records your roll number, a fingerprint of your entry, the fingerprint of the roll, and the fingerprint of the model file the roll is bound to.
Open the model file: it names the roll and the certificate registry. Open the certificate: it names the roll and the model. Two fingerprints, matching in both directions — checkable by anyone with the file and a browser, with no need to trust us. We looked for another model file that does this and did not find one. That is the first we claim: the two-way binding, as a combination — not the idea of a name in a model.
A fingerprint is a SHA-256 hash: change one byte of the file and it changes completely. The binding covers the shipped model files whose fingerprints are recorded in the certificate registry; each release is bound separately. Every claim on this page is scoped to those files. Check a certificate at ternii.com/founders/verify.
One-off. Worldwide. Priced in pounds. The amount shown on the button is what you pay. Where local tax applies, training tokens are credited on the amount before tax.
Names and messages are reviewed by a person before they go on the roll or a certificate is issued: nothing sexual, violent or abusive, no profanity, no other people's personal data. Our decision is final. We'll ask you to change it — or, if you would rather not, refund you in full, whatever stage your membership has reached.
The name can be anyone's — a gift for a child, a grandparent, a pet. Pseudonyms welcome. The person paying must be 18 or over.
We count every £10 as 500 million training tokens — a rate set below what a billion tokens has cost us to train, so it covers failed runs, evaluation, data preparation and price rises. We publish the running total of tokens funded.
Why tokens: the models ternii is compared with learned from trillions — BitNet b1.58 2B4T: 4 trillion; SmolLM3: 11 trillion. Every 200 members = 100 billion more training tokens.
Fees are revenue of Middletech Limited. We intend to spend them on training compute and we publish what we trained, but we do not promise a particular model, release date or result. Early-access builds are pre-release software provided as is.
Payments by Stripe, one-off, in GBP from anywhere; VAT included where it applies.
Founding Membership is a paid membership with the benefits above. It is not a share, a loan, an investment or a donation. It carries no ownership, no voting rights and no financial return. Cancel any time until your entry has been reviewed and your certificate issued; after both, your £10 has bought training tokens and is spent. We always refund a duplicate charge, and an entry we decline. Middletech Limited, company no. 16955822. Terms · Privacy.
Every approved entry, in roll order, shown the way its member chose. This is the roll whose fingerprint will go into the model file from the next release — the same list, on every phone that runs it.
The roll opens with you — your name here.
a message in your words, up to 500 characters — for someone, about something, or just becauseexample#4 for a grandparent
These four are placeholders to show the shape of the roll, not members. Real entries replace them as they are reviewed and approved.
Brain is a 3.4-billion-parameter model, trained from scratch by a small team, on far fewer tokens than the models it is measured against — they learned from trillions. Brain is faster and smaller than they are, and it runs where they don't; it is not yet as accurate, and we say so on this page.
Founding Members change the arithmetic. Your £10 buys 500 million tokens of training, and you are written into what it buys — on the roll, in the model file, and, if you opt in, in the data the next Brain is trained on. We publish what we trained, what it cost and what we measured, in the members' area, as it happens. That is the whole deal: no cloud, no account, a small team, and your name on the thing we make.
— Ray, Middletech
Founding Members get a sign-in link by email — no password. Behind it: the work, as it happens, and a board where you tell us what to build.
Not a member yet? Join the Founding Roll — from £10.
The requests board is where members tell us what ternii should do next — a routine, an expert pack, a language, a fix — and vote on each other's. Each request carries a status that moves as we work:
It lives behind sign-in on purpose: members talking to a small team, not a comments section. Ten requests a day each, real names or roll names, and we read every one.
The board is new; what is on it comes from members, not from us. We do not write example requests here — you'll see the real ones once you're in.
A private version of Brain, trained or tuned on your data, for your own use. It runs on your people's phones or on your own hardware, and nothing leaves your control — no cloud, no account, no third party in the loop.
We train or tune Brain on your corpus and your vocabulary — your manuals, your records, your way of saying things. It is delivered as a model file you hold, and it runs on-device and offline, the same way ternii does.
Expert packs and routines built for your tasks, in your terms. Tools that stay inside your boundary: the model reads what you allow, on the device you allow, and answers there.
Researchers, universities, device and chip makers, app builders. Co-develop, evaluate, publish. If you are working on small models, on-device inference or ternary hardware, we would like to compare notes.
Brain is early and still training; a private version is a project we scope together, not a product off the shelf.
A short, honest case — four things you can check on this page, and one you can check on a phone.
A 3.4-billion-parameter ternary model — a panel of experts, about a billion of them awake per word — writing at 118 tokens a second in 195 MB of memory on an iPhone 17 Pro, measured in the app in a chat turn, not projected. The same file runs on a 2017 iPhone X.
Ternary from the first token of training, so there is no compression step to lose accuracy in: the file on the phone scores the same as the checkpoint we trained, within 0.25 % perplexity. And nothing leaves the phone.
Members fund training tokens directly and, from the next release, are written into the model, with a signed certificate. As far as we can find, it will be the first AI model file to name its own founders and their certificate registry — with each certificate naming the model back.
GPU rental for training. Brain has learned from a fraction of what the models it is measured against have seen; the next Brain needs trillions of tokens. We publish what we trained and what we measured.
Next is Brain 2.0: 4.3 billion parameters, about one in five of them active per word. The geometry is built and measured — quicker to decode than Brain 1.0 on the same phone, at four times the context. In the app on an iPhone 17 Pro, with agent mode and tools loaded — the heaviest path — it writes at 126–140 tok/s where Brain 1.0 manages 80: about 1.6×, same phone, same turn, only the model changed. It still needs teaching, and that training is the line item.
This page is information about Middletech Limited and its work. It is not an offer or invitation to invest, and nothing here is financial advice. If you are a professional investor or a funder and would like to talk, get in touch and we will share more under NDA.
Want a pack we don't have yet — beekeeping, tax, a language, your trade? Tell us. Because experts are just downloadable packs, we can add a new one for everyone without an app update.
Found a bug? Want a new routine, or a feature? ternii gets better from what real people ask for. Tell us what happened or what you'd love to see. Organisations, collaborators and funders: pick your type and it comes straight to the team.
🐛 Bugs · ⚡ routine ideas · ✨ feature requests · 💬 just thoughts — pick a type and tell us.
A private AI in your pocket — and a first look at what three states can do. On Google Play today. The iPhone build is with Apple; leave your email and we will tell you the day it lands.