New whitepaperNext-Generation NI: A Vision, Philosophy, and Technical Path for Continual LearningRead now

They said weights are frozen. We made them grow with you.

AI Models That Grow for You. On Your Device.

Continual learning happens on your device: the model grows with use, remembers across sessions, and your data never leaves.

“We don't sell tokens. We set them free.”

How it works

Neural Imprint

Memory, learning, ownership — the brain of embodied AI lives in the device, not in someone else's cloud.

01

Generate

The devices around you generate data continuously — some from what you do, some from what they sense.

02

Learn

The device locally understands your patterns, forming a cognition of who you are.

03

Evolve

That cognition is imprinted into the model's internal state — no weight edits, no fine-tuning, no cloud.

Your data physically never leaves your device.

Why it's different

Learning without touching the weights.

Not fine-tuning. Not RAG. A recoverable learning state that travels with the model — on whatever runs it.

Weights stay frozen

Learning is captured as a compact inference state — never a gradient step.

Persists · reverts · verifies

It survives sessions and restarts, is gated before activation, and rolls back in one step.

Any model, any chip

It lives at the inference-state level, so it isn't tied to a model family or a processor.

Validated end to end on consumer hardware.

Evidence and boundaries

Cross-silicon

One architecture, shipping on two silicon platforms

The same learning and inference algorithms; only the execution backend changes. On Apple devices, learning closes the loop on the phone itself. On Intel, our runtime has shipped to a world-leading PC manufacturer.

Apple

Learning loop closed on device

iPhone / iPad / Mac

204,800tokens, 0 failures: 9B ran 200 consecutive turns on each of two iPhones, up to 7.1 hours
0.47 smedian wait before each answer starts, with context grown to 200K tokens (iPhone 17 Pro)
8/8life areas passed in one continuous run: profile questions answered directly, facts via the right tool (simulated data)
5.2 minfor 9B to learn from 361 real spending records on an iPhone 17e
  • Our own inference compute and runtime, with hand-written and fused kernels
  • Neural Imprint learning state captured, restored and incrementally updated on real devices
  • Incremental learning: unchanged records are reused, so re-learning 355 records with 9B on an iPhone Air drops from 117 s to 39–58 s
  • Multimodal: a 9B vision model passed all 5 turns on each of 8 images on an iPhone Air
Evidence and boundaries

Intel

Shipped to a world-leading PC manufacturer

Core Ultra · CPU + Arc iGPU

0.37 sto first sound; the open-source reference composes the whole sentence first, taking 15–20 s
0.73speech synthesis real-time factor (RTF), faster than playback; reference 2.1–2.9
0.45 s9B first token with Neural Imprint state restore; full prefill of the same content 0.97 s
0.4 sspeech recognition on the same 4.2 s Chinese clip; reference 0.9 s
  • Our own inference runtime on OpenVINO: stateful streaming vocoder, split predictor, CPU and iGPU pipeline
  • Speech recognition, a 9B language model and speech synthesis resident on one machine, all inference on the device
  • Neural Imprint learning state saved and restored on OpenVINO, live since the second delivery
  • Two deliveries: fresh install in 5 minutes, usable 50 seconds after a power cycle; 357/357 sentences complete over 30 minutes of continuous synthesis
Read the technical write-up

More silicon platforms on the roadmap

Global
NVIDIAAMDQualcomm
China
CIXHoumo NPURockchipAxeraSophgoHuawei Ascend

No dependency on chip vendors' high-level APIs — we go down to kernels and compute graphs when needed.

Intel figures compare runs on the same machine with the same model, an Intel Core Ultra 5 225H. RTF = generation time ÷ audio duration; speech synthesis figures are medians from a 30-minute resident run.

Build with us

We're looking for researchers and engineers who want to push the boundaries of on-device AI.