Auryn models · forged by Potara · certified honest

AI that never leaves your building.
Trained on your work. Tuned to your tasks.

Healthcare, legal, finance, and defense can't send their data to a hosted API. The answer is a model that runs on your own hardware, fine-tuned on your own domain, that never phones home. Potara forges it; Auryn is the model; the QPU certificate is the proof it's honest.

What can my GPU run?
Measured, not marketed AURYN-CLAIMS-1.5B vs BASE Qwen2.5-1.5B · SAME Q4 · HELD-OUT
Runs private
8 GB GPU
your hardware · no API · no PHI leaves
Size
986 MB
Q4 · 2.99× smaller
Speed
90 tok/s
base 72 · +24%
Held-out EDI vs base
on par
12% vs 15% — no domain gain

Straight talk: an earlier scorecard showed 0.36 → 1.00, but that eval was drawn from the same facts the model was trained on — it measured memorization, not skill. On a genuinely held-out 26-item EDI test (leakage-checked), the fused model scores ~12% vs the base's ~15% and both trail a 7B model (~27%): the fusion does not out-know the base on your domain. What it does deliver is real and verifiable: it runs on your hardware (privacy), is ~3× smaller and ~24% faster, doesn't lose general ability, and reliably reproduces the specific facts and behaviors you explicitly train it on. So we train skills and retrieve facts — a retrieval harness ships with every model, and every "better at your job" claim is proven on a held-out set, per customer. Full benchmark: openauryn/honest_benchmark.json.

The pipeline

Four steps. Your data stays put.

01

Interview

A chat builds your calibration set: what you do, your tasks, examples. It splits skills (train) from facts (retrieve).

02

Fuse on your hardware

Merge + calibration-prune + light QLoRA on your own GPU (data never leaves), or our zero-idle cloud. Facts stay fresh via a retrieval harness.

03

Certify the test

Optionally draw the eval sample on real IBM Quantum hardware, so the test set is provably not cherry-picked. Verifiable randomness & provenance only — it does not change the model's quality, size, or speed.

04

Run & re-fuse

Export to Ollama, runs local. When a better base drops, re-apply your profile in one click — your investment compounds.

Free tool

What can your GPU actually run?

Plans

Priced like compliance software, not an API.

On BYO-GPU, fusion runs entirely on your hardware — your data never leaves your building (the compliance play). On the Cloud tier your data goes to our GPU workers for the job (job-based, deleted after) — for teams without that constraint. QPU certification is a +$10 add-on.

Loading plans…
Potara Studio · a product of OpenAuryn

Your Studio

The forge. Build, run, and re-fuse your expert engines — all under your OpenAuryn account.

Forge Build a fused engine
Runner Run the fusion on your GPU

BYO-GPU jobs run through the Potara runner on your machine:

pip install "potara[fusion]"
potara-runner login --passphrase <your key>
potara-runner start   # picks up your queued jobs

Cloud-GPU plan? Jobs route to our GPU workers automatically — no runner needed.

Your fusion jobs & models

No jobs yet.

History & billing

Usage & estimated bill (this cycle)
Loading…
Your profiles & version scorecards
No profiles yet — run an interview to create one.
Studio activity log
Loading…

Chat