繁中

Research Paper

Allocating a fixed memory budget in visually rich document RAG

Abstract

Three models in one machine’s memory

Retrieval-augmented generation (RAG) over visually rich documents needs several vision-language models: a page embedder, a page-image reranker, and a generator. We study how a fixed unified-memory budget is shared among these models and their caches on a stock low-memory desktop computer running a 12B generator, two retrieval models of about 1.7B parameters, and an application stack in a virtual machine. We varied twelve allocation levers and measured memory, swap, memory pressure, time to first token (TTFT), throughput for 1–4 users, and answer quality on 73 questions.

The models and the virtual machine took 19.6–20.4 GiB, and a 26B mixture-of-experts generator with 4B active parameters could not run alongside the retrieval models. Once the 12B generator fit, the largest memory growth came from caches with large or unbounded defaults. Generator prefix snapshots grew by 7.6 GiB, swap reached 12.5 GiB, and an upload run reached critical memory pressure; the allocator cache grew by 13.8 GiB. The workload rarely reused a prefix (6.7% of prompt tokens). In single runs per setting, capping the snapshots showed no detectable cost; capping the allocator cache added 2% to reranking time.

Per-page vision gating cut a 75-page slide deck from 75 to 11 page images and its ingestion from 93 to 28 s. MLX runtimes lowered median TTFT by 2.1–2.65 s for the generator and a further 3.3–3.5 s for the retrieval models. Cutting reranking candidates from 50 to 6 was faster but lost 5 of 34 and 6 of 39 questions, while unseeded sampling alone moved pass counts by up to two.

In Practice

From the paper to open models

The MLX versions of NVIDIA’s Nemotron embedding and reranking models for visually rich documents come from this research. Both run on Apple silicon with results that closely match the originals, and both are published on Hugging Face for anyone building document retrieval on a single machine.

See the models on Hugging Face