Patent Pending MNEMOS — Canadian patent application filed, 2026. A new way to run, edit, and extend large models. Talk to us →
MNEMOS · Model Architecture · Patent Pending

Run a model bigger than your hardware — exactly — and edit what it knows.

Aurevus has developed — and filed a patent on — MNEMOS, an architecture that treats a frozen language model not as one immovable block of memory, but as parts that can be handled differently. On a single frozen model, three things run at once: it executes far larger than device memory with exact, never-quantized weights — bit-identical output wherever a full comparison can run — its knowledge becomes editable without retraining, and new capability attaches across knowledge, skills, and reasoning, without touching the original weights.

device.memory
measured

A 32-billion-parameter model (~60 GB)

running on a 24 GB device

~11 GB resident
0device memory24 GB
Exactexact weights throughout
Editablereversible writes in seconds
Frozenhost weights unchanged
Concurrentall tiers at once
~60 GB → ~11 GB A 32B-class model runs coherently on a 24 GB device — exact weights, no quantization.
Bit-for-bit identical A 13B model (26.4 GB) runs in 14 GB with hash-verified identical output. Not a lossy 4-bit shrink.
Knowledge · skills · reasoning Attached capacity lifts all three — +44 median on a held-out reasoning battery — while the host stays byte-identical.
In seconds Write, correct, or reverse a fact — audited, reversible, zero measured collateral.
The architecture

Partition the frozen model by function.

A trained model's weights normally all sit in expensive accelerator memory, all the time. MNEMOS treats different parts of the frozen model differently — three functional tiers that operate together on one host.

1

Storage-resident execution

Weights are streamed from ordinary storage as the model runs, and the output is bit-for-bit identical to the fully-loaded model. A model far larger than the device executes with only a fraction resident — exact, not approximated.

2

Editable memory

Selected structure inside the model becomes an addressable memory. You can write a new fact, correct a wrong one, or reverse the change — in seconds, with a receipt for every change, measured zero collateral to the rest of the model, and no retraining of the host.

3

Capacity augmentation

Independently trained capacity attaches to the frozen host, with measured uplift across all three dimensions of capability — knowledge, skills, and reasoning. On our capability battery, an augmented 1B host outscored an unaugmented 7B — while the host's original weights remain byte-identical, preserving any certification or review attached to it.

All three run concurrently on one frozen model — exact, editable, extensible — measured together, on consumer hardware.

Measured, not claimed

Reduction to practice, on a 24 GB laptop.

Every figure below is a measured result from the working implementation. The exact optimization procedures are held as trade secrets and are not required to understand the architecture.

60 GB in ~11 GB

A 32-billion-parameter model running on a 24 GB device, generating coherently, output exact by construction.

8.9 GB, all tiers

A 7B host with streaming, editable memory, and attached capacity all active — 8.9 GB versus a 14.6 GB baseline.

Knowledge · skills · reasoning

Attached capacity lifts all three measured dimensions — +44 median on a held-out reasoning battery at 7B; on the same battery an augmented 1B outscored an unaugmented 7B. Host weights unchanged.

Exact — and interactive

At the tier that fits fast memory, exact is also fast: a 7B streams at ~15 tokens/sec with hash-identical output; a 13B (26.4 GB) runs in 14 GB, bit-for-bit, string-identical across every control.

Why it matters

A model's parameters, its editable knowledge, and its new capability no longer have to share one substrate.

Separate the computation from the storage, from the editable knowledge, from the newly trainable capability — and device memory stops being the ceiling on what a model can be. MNEMOS is independent of scale: the same partitioning applies at 7B, at 70B, and in principle to frontier-scale hosts far larger than any single accelerator can hold.

1

Larger than memory

Run models several times bigger than the device — exactly, without shrinking them.

2

Correct without retraining

Fix stale or wrong knowledge in seconds, reversibly, with an audit trail — no GPU-hours.

3

Upgrade a frozen artifact

Add capability to a shipped model without disturbing its certified, byte-identical weights.

4

Exact, not lossy

Unlike 4-bit quantization or lossy offload, the weights are never approximated — verified bit-identical at the fitting tier. The same model, relocated.

Status & protection

Patent-pending. Working implementation. For licensing or partnership.

A Canadian patent application covering the MNEMOS architecture has been filed. The invention is reduced to practice on consumer hardware; the exact optimization and control procedures remain proprietary. This is a licensable method, not a product we have to out-execute anyone to sell.