Sovereign Mind turns your documents into an auditable wiki and a downloadable RAG that answers on-device — with citations, fully offline. Your data never leaves your infrastructure.
M12 flange bolts are torqued to 85 Nm in a crosswise pattern, in two passes (40 Nm, then 85 Nm).
Clean the threads, apply the spec with a calibrated wrench, and verify each fastener.
Work the cross pattern from the inside out to seat the gasket evenly.
service_manual.pdf · §7.3See the ✈ up top? This is answering in airplane mode — fully offline, with citations.
Two failure modes keep serious organizations from adopting the tools everyone else is racing to deploy.
Dental, medical, and legal teams cannot — and will not — send sensitive documents to US hyperscaler APIs. Compliance ends the conversation before it starts.
Field crews in basements, air-gapped plants, and trains need answers when there's no signal. When the connection drops, cloud AI simply disappears.
A coding agent was recently caught uploading entire Git repos — secrets and all — to a vendor bucket at ~27,800× the data the task needed. The privacy toggle changed nothing.
Three steps take a pile of manuals to a phone that answers offline — without a single byte sent for external processing.
Point the pipeline at your PDFs, regulatory texts, and service manuals. A larger model extracts structure and compiles the raw data — on your own infrastructure, so nothing is sent out for training.
PDFs · logs · source code → structured corpusNobody trusts a black box. Every collection becomes a human-readable wiki your domain experts can review and edit before anyone asks a question — and it's what the model cites when it answers.
EU AI Act Art. 12 / 13 aligned — sovereign by architectureSync the compiled collection to an iPhone and go offline. The phone answers with citations even with the network permanently unavailable. Airplane mode is our favorite demo feature.
On-device · cited · role-scoped accessThe stack is designed to go one step further than answering from the manuals: reading the machine's own log files to write the missing pages itself. (Planned — on our roadmap; see below.)
An open-weights model doing the actual inference on the handset. No API call, no round trip, no data leaving the device.
The same model, scoped to a role. A service engineer and a production manager ask the same machine very different questions — each gets a prompt tuned to the job they are actually doing.
The compiled collection synced to the phone. Every answer is retrieved from it and cites the page it came from — the audit trail travels with the device.
Planned capability. The log-file analysis pipeline is a cloud-side, advanced feature on our roadmap — it studies a deployment's first runs to surface the faults that recur. The on-device query, cited retrieval and expert-reviewed wiki above ship today.
Your manuals, service docs and regulatory texts, already compiled into the human-readable wiki your experts reviewed.
The machine's own logs are mined for the errors that actually recur in the field — then answered against the wiki.
The collection that ships to the phone now has the FAQ built in: the common faults come pre-answered, with citations, offline.
We fed the machine's log files into the pipeline and let a large model mine them for the failure modes that keep coming back — the errors real operators actually hit, not the ones the manual anticipated.
machine logs → recurring error casesFor each recurring error, the large model drafts new wiki pages documenting probable causes and solutions — grounded in both the logs and the existing wiki, so the new material stays consistent with the documentation your experts already signed off on.
grounded in logs + existing wikiThe new pages land in the same human-readable wiki, get reviewed like any other page, and recompile into the collection on the phone. The next engineer in front of that machine gets the answer offline — including for a fault nobody had written up before.
reviewed · recompiled · offline againSame retrieval idea. Opposite guarantees on the four things that actually matter for on-device AI.
Trained on 1,800+ languages. The default powerhouse, running entirely on iOS.
A highly compact alternative for specific licensed commercial deployments.
Private instances scaling up to Apertus 70B for connected enterprise queries.
No hand-waving — here is exactly how sovereignty, offline operation, and grounding work in practice.
Nowhere. Once the knowledge base is built, querying it runs entirely on the device — no cloud call, no telemetry, no upload. That is the submarine test: unplug the network and inference still works.
Two phases, and it is worth being precise. Corpus generation happens on your infrastructure: your documents are processed there to build the wiki / knowledge base. After that, inference — the day-to-day querying — is fully local on the device with no network dependency. The build step is a controlled, one-time process on your side; the ongoing use is offline.
Querying has full functionality. The corpus and the model both live locally, so offline is the normal operating mode, not a degraded one. Air-gapped deployment for daily use is supported by design.
Those keep your documents on someone else's servers under a data-processing agreement you have to trust. Sovereign Mind removes the agreement from the equation for inference — there is no server to query against. The honest trade-off: you get sovereignty and predictability rather than a 400B-parameter frontier model. For a bounded domain with good retrieval, the curated small model is what matters, not raw size.
A curated, tested shortlist: Apertus (1.5B and 4B), Ministral 3B, Liquid AI's LFM2-VL, and Qwen2-VL. Every one is validated against real retrieval tasks on the target hardware, so nothing crashes mid-use or drifts wildly on your domain.
The list is deliberately curated rather than load-anything. Curation is the point: it is what lets us guarantee the thing runs stably on your hardware instead of pushing the support and stability burden onto you. If there is a model you need, we evaluate and validate it into the list rather than letting arbitrary weights load blind.
Every answer is grounded in your retrieved corpus with citations back to the source document, so a user can check the origin instead of trusting the model blind. If retrieval finds no supporting passage, the model is never called — it refuses rather than guesses. Judgment in the weights, facts in the corpus.
Your documents are processed on your infrastructure using a frontier model to generate a structured, cited wiki — an indexed, queryable knowledge base — which an admin reviews and approves before it ships. That generation step is the part simple "chat with a PDF" tools skip: you get a real retrieval layer, not documents stuffed into a context window at query time.
The corpus is regenerable — updated source documents run back through the generation pipeline to refresh the index. Update cadence can be scoped into the arrangement.
Depends on the model. Apertus and Qwen are open weights. Ministral and Liquid AI carry their own terms — notably Liquid's LFM2 is free for commercial use only while your annual revenue stays under USD 10M; above that you would need a paid license from Liquid directly. We pick the model to fit your situation and flag any such clause up front.
Yes. Inference runs locally with cited retrieval, so any answer traces to its source and the pipeline is inspectable. Where EU AI Act transparency and logging obligations apply, the architecture supports them rather than fighting them.
Explore a live compiled collection built from a robot's manuals, logs, and source code. No install required.
Want to see on-device inference before you talk to us? Silicon Oracle is our consumer app running open-weight models entirely on iPhone — and the case study documents what it took.
Curious how it began? Watch our hackathon demo, where Sovereign Mind was first built.
Enterprise trials, investor conversations, and partnership enquiries welcome.