Language models
Reasoning, generation and agents.
Answers computed where the question was asked — no round trip, no per-token bill, no transcript leaving the device.
One-bit foundation models,
built to think locally.
Reasoning, generation and agents.
Answers computed where the question was asked — no round trip, no per-token bill, no transcript leaving the device.
Perception and visual understanding.
Real-time detection and classification on camera-class silicon — every frame processed where it was captured.
Speech, sound and acoustic intelligence.
Speech and sound understood on the device itself — for the rooms, cabins and floors where audio must never leave.

Network monitoring when the link goes down.

Onboard vision without a cloud connection.

Machine control behind the air gap.
A model that needed 4 GB fits in 0.4 GB — small enough for a phone.
Versus full precision — a 2B-parameter model runs in 0.4 GB, beside the application you already ship.
Multiplies become additions, so each token takes far less arithmetic.
On a single CPU — 29 ms per token from a 2B model, fluent real-time generation with no accelerator involved.
Sixteen bits of precision cut to one, with the model still doing its job.
On the ARM cores embedded products already ship — up to 6.2x on x86 — with optimised 1-bit kernels.
1x = the same model on a standard runtime
Less memory traffic, less energy per token, no data centre required.
At the top end — an estimated 0.028 J per token against 0.186–0.649 J for comparable models.
We’re taking a small number of hardware partners — OEMs, chip vendors and SOM providers — to benchmark Fyvie on real platforms and shape the deployment stack. Integrate once at the platform level; every product family built on it inherits the model.
TALK TO US ↗