Run a model bigger than your hardware — exactly — and edit what it knows.
Aurevus has developed — and filed a patent on — MNEMOS, an architecture that treats a frozen language model not as one immovable block of memory, but as parts that can be handled differently. On a single frozen model, three things run at once: it executes far larger than device memory with exact, never-quantized weights — bit-identical output wherever a full comparison can run — its knowledge becomes editable without retraining, and new capability attaches across knowledge, skills, and reasoning, without touching the original weights.
A 32-billion-parameter model (~60 GB)
running on a 24 GB device
Partition the frozen model by function.
A trained model's weights normally all sit in expensive accelerator memory, all the time. MNEMOS treats different parts of the frozen model differently — three functional tiers that operate together on one host.
Storage-resident execution
Weights are streamed from ordinary storage as the model runs, and the output is bit-for-bit identical to the fully-loaded model. A model far larger than the device executes with only a fraction resident — exact, not approximated.
Editable memory
Selected structure inside the model becomes an addressable memory. You can write a new fact, correct a wrong one, or reverse the change — in seconds, with a receipt for every change, measured zero collateral to the rest of the model, and no retraining of the host.
Capacity augmentation
Independently trained capacity attaches to the frozen host, with measured uplift across all three dimensions of capability — knowledge, skills, and reasoning. On our capability battery, an augmented 1B host outscored an unaugmented 7B — while the host's original weights remain byte-identical, preserving any certification or review attached to it.
All three run concurrently on one frozen model — exact, editable, extensible — measured together, on consumer hardware.
Reduction to practice, on a 24 GB laptop.
Every figure below is a measured result from the working implementation. The exact optimization procedures are held as trade secrets and are not required to understand the architecture.
60 GB in ~11 GB
A 32-billion-parameter model running on a 24 GB device, generating coherently, output exact by construction.
8.9 GB, all tiers
A 7B host with streaming, editable memory, and attached capacity all active — 8.9 GB versus a 14.6 GB baseline.
Knowledge · skills · reasoning
Attached capacity lifts all three measured dimensions — +44 median on a held-out reasoning battery at 7B; on the same battery an augmented 1B outscored an unaugmented 7B. Host weights unchanged.
Exact — and interactive
At the tier that fits fast memory, exact is also fast: a 7B streams at ~15 tokens/sec with hash-identical output; a 13B (26.4 GB) runs in 14 GB, bit-for-bit, string-identical across every control.
A model's parameters, its editable knowledge, and its new capability no longer have to share one substrate.
Separate the computation from the storage, from the editable knowledge, from the newly trainable capability — and device memory stops being the ceiling on what a model can be. MNEMOS is independent of scale: the same partitioning applies at 7B, at 70B, and in principle to frontier-scale hosts far larger than any single accelerator can hold.
Larger than memory
Run models several times bigger than the device — exactly, without shrinking them.
Correct without retraining
Fix stale or wrong knowledge in seconds, reversibly, with an audit trail — no GPU-hours.
Upgrade a frozen artifact
Add capability to a shipped model without disturbing its certified, byte-identical weights.
Exact, not lossy
Unlike 4-bit quantization or lossy offload, the weights are never approximated — verified bit-identical at the fitting tier. The same model, relocated.
Patent-pending. Working implementation. For licensing or partnership.
A Canadian patent application covering the MNEMOS architecture has been filed. The invention is reduced to practice on consumer hardware; the exact optimization and control procedures remain proprietary. This is a licensable method, not a product we have to out-execute anyone to sell.
