On-device intelligence · Enterprise sovereignty

The best AI demo works with the Wi-Fi off.

Sovereign Mind turns your documents into an auditable wiki and a downloadable RAG that answers on-device — with citations, fully offline. Your data never leaves your infrastructure.

09:33
Field Engineer
Ask a question…
Chat Wiki Settings
Wiki
Intake Manifold — Fasteners
1. Torque Specification

M12 flange bolts are torqued to 85 Nm in a crosswise pattern, in two passes (40 Nm, then 85 Nm).

2. Procedure

Clean the threads, apply the spec with a calibrated wrench, and verify each fastener.

2.1 Sequence

Work the cross pattern from the inside out to seat the gasket evenly.

service_manual.pdf · §7.3
Chat Wiki Settings
Admin Console LIVE
New collection
Attach document
service_manual.pdfuploading…
Assign access
AKAnna KellerField Engineer · read
Settings
AI Model
Ministral 3BQ4 · on-device · 2.5 GB
Apertus 1.5BSwiss AI · experimentalDownload
Qwen 2-VLVision · 2.0 GBDownload
Knowledge base
Robot Ross ATF15 docs · 37 chunks
Persona
Field EngineerConcise, practical troubleshooting
Product Manager
Technical Writer
Chat Wiki Settings

See the up top? This is answering in airplane mode — fully offline, with citations.

Runs on: Apertus Mini 4B· Ministral 3B· iOS / on-device· Sovereign Cloud · Infomaniak · Swisscom
The missing pieces of enterprise AI

Generic AI breaks exactly where it matters most.

Two failure modes keep serious organizations from adopting the tools everyone else is racing to deploy.

The compromise

Regulated data can't leave

Dental, medical, and legal teams cannot — and will not — send sensitive documents to US hyperscaler APIs. Compliance ends the conversation before it starts.

The disconnect

The network isn't always there

Field crews in basements, air-gapped plants, and trains need answers when there's no signal. When the connection drops, cloud AI simply disappears.

The fine print

"Private" is a setting, not a guarantee

A coding agent was recently caught uploading entire Git repos — secrets and all — to a vendor bucket at ~27,800× the data the task needed. The privacy toggle changed nothing.

Your bricks · Your baseplate

Your documents become an audit trail you can hold in your hand.

Three steps take a pile of manuals to a phone that answers offline — without a single byte sent for external processing.

1

Ingest

Point the pipeline at your PDFs, regulatory texts, and service manuals. A larger model extracts structure and compiles the raw data — on your own infrastructure, so nothing is sent out for training.

PDFs · logs · source code → structured corpus
2

Compile the auditable wiki

Nobody trusts a black box. Every collection becomes a human-readable wiki your domain experts can review and edit before anyone asks a question — and it's what the model cites when it answers.

EU AI Act Art. 12 / 13 aligned — sovereign by architecture
3

Sync to device — the submarine test

Sync the compiled collection to an iPhone and go offline. The phone answers with citations even with the network permanently unavailable. Airplane mode is our favorite demo feature.

On-device · cited · role-scoped access
Beyond the manual

The system learns from experience.

The stack is designed to go one step further than answering from the manuals: reading the machine's own log files to write the missing pages itself. (Planned — on our roadmap; see below.)

ON DEVICE 01

The local model

Open weights · runs on the iPhone

An open-weights model doing the actual inference on the handset. No API call, no round trip, no data leaving the device.

ON DEVICE 02

The specialised prompt

Service engineer · production manager

The same model, scoped to a role. A service engineer and a production manager ask the same machine very different questions — each gets a prompt tuned to the job they are actually doing.

ON DEVICE 03

The RAG

Compiled wiki · cited retrieval

The compiled collection synced to the phone. Every answer is retrieved from it and cites the page it came from — the audit trail travels with the device.

The feedback loop · planned

From log file to wiki page.

Planned capability. The log-file analysis pipeline is a cloud-side, advanced feature on our roadmap — it studies a deployment's first runs to surface the faults that recur. The on-device query, cited retrieval and expert-reviewed wiki above ship today.

INPUT

Compiled wiki

Your manuals, service docs and regulatory texts, already compiled into the human-readable wiki your experts reviewed.

PROCESS · PLANNED

Parse the log files

The machine's own logs are mined for the errors that actually recur in the field — then answered against the wiki.

OUTPUT

Enhanced RAG

The collection that ships to the phone now has the FAQ built in: the common faults come pre-answered, with citations, offline.

1

Read what actually went wrong

We fed the machine's log files into the pipeline and let a large model mine them for the failure modes that keep coming back — the errors real operators actually hit, not the ones the manual anticipated.

machine logs → recurring error cases
2

Write the pages that were missing

For each recurring error, the large model drafts new wiki pages documenting probable causes and solutions — grounded in both the logs and the existing wiki, so the new material stays consistent with the documentation your experts already signed off on.

grounded in logs + existing wiki
3

Ship it back down to the device

The new pages land in the same human-readable wiki, get reviewed like any other page, and recompile into the collection on the phone. The next engineer in front of that machine gets the answer offline — including for a fault nobody had written up before.

reviewed · recompiled · offline again
Why Sovereign Mind

Generic RAG vs. Sovereign Mind

Same retrieval idea. Opposite guarantees on the four things that actually matter for on-device AI.

 
Generic RAG
Ω Sovereign Mind
Data location
Sent to US APIs
Stays on your infrastructure
Network reliance
Always online
True offline · passes the submarine test
Auditability
Black-box retrieval
Human-readable wiki
Compliance
Promised privacy
EU AI Act Art. 12/13 aligned · measured groundedness
The engine block · The Swiss stack

Open weights, all the way down.

COMPONENT 01

Apertus Mini 4B

ETH · EPFL · CSCS — Apache 2.0

Trained on 1,800+ languages. The default powerhouse, running entirely on iOS.

COMPONENT 02

Ministral 3B

Compact · licensed commercial

A highly compact alternative for specific licensed commercial deployments.

COMPONENT 03

Hosted scalability

Sovereign Cloud · Infomaniak · Swisscom

Private instances scaling up to Apertus 70B for connected enterprise queries.

Straight answers

Q&A

No hand-waving — here is exactly how sovereignty, offline operation, and grounding work in practice.

Where does our data go at inference time?

Nowhere. Once the knowledge base is built, querying it runs entirely on the device — no cloud call, no telemetry, no upload. That is the submarine test: unplug the network and inference still works.

So the documents never leave our infrastructure?

Two phases, and it is worth being precise. Corpus generation happens on your infrastructure: your documents are processed there to build the wiki / knowledge base. After that, inference — the day-to-day querying — is fully local on the device with no network dependency. The build step is a controlled, one-time process on your side; the ongoing use is offline.

What happens with no internet connection?

Querying has full functionality. The corpus and the model both live locally, so offline is the normal operating mode, not a degraded one. Air-gapped deployment for daily use is supported by design.

How is this different from ChatGPT Enterprise or Copilot?

Those keep your documents on someone else's servers under a data-processing agreement you have to trust. Sovereign Mind removes the agreement from the equation for inference — there is no server to query against. The honest trade-off: you get sovereignty and predictability rather than a 400B-parameter frontier model. For a bounded domain with good retrieval, the curated small model is what matters, not raw size.

Which models do you run?

A curated, tested shortlist: Apertus (1.5B and 4B), Ministral 3B, Liquid AI's LFM2-VL, and Qwen2-VL. Every one is validated against real retrieval tasks on the target hardware, so nothing crashes mid-use or drifts wildly on your domain.

Can we swap in our own models?

The list is deliberately curated rather than load-anything. Curation is the point: it is what lets us guarantee the thing runs stably on your hardware instead of pushing the support and stability burden onto you. If there is a model you need, we evaluate and validate it into the list rather than letting arbitrary weights load blind.

How do you handle hallucination or wrong answers?

Every answer is grounded in your retrieved corpus with citations back to the source document, so a user can check the origin instead of trusting the model blind. If retrieval finds no supporting passage, the model is never called — it refuses rather than guesses. Judgment in the weights, facts in the corpus.

How do you build the knowledge base?

Your documents are processed on your infrastructure using a frontier model to generate a structured, cited wiki — an indexed, queryable knowledge base — which an admin reviews and approves before it ships. That generation step is the part simple "chat with a PDF" tools skip: you get a real retrieval layer, not documents stuffed into a context window at query time.

When our documents change, how does the corpus stay current?

The corpus is regenerable — updated source documents run back through the generation pipeline to refresh the index. Update cadence can be scoped into the arrangement.

What are the licensing terms of the underlying models?

Depends on the model. Apertus and Qwen are open weights. Ministral and Liquid AI carry their own terms — notably Liquid's LFM2 is free for commercial use only while your annual revenue stays under USD 10M; above that you would need a paid license from Liquid directly. We pick the model to fit your situation and flag any such clause up front.

Can we audit it for compliance?

Yes. Inference runs locally with cited retrieval, so any answer traces to its source and the pipeline is inspectable. Where EU AI Act transparency and logging obligations apply, the architecture supports them rather than fighting them.

Try it now

See the wiki and RAG in action — online, before you go offline.

Explore a live compiled collection built from a robot's manuals, logs, and source code. No install required.

iOS app now in public TestFlight beta · Android planned · get the app →
Get in touch

Let's Build Together

Enterprise trials, investor conversations, and partnership enquiries welcome.

Partners

Built with great partners

SourceryKit by Provably SourceryKit by Provably
Blindsight Security Blindsight Security