from the notebook ✦

Running a Local LLM for Document Q&A: A Zero-Trust Privacy Setup

Some documents can't go to an API — contracts, audit reports, HR records. Here is the fully offline retrieval pipeline I use, what it costs in latency on ordinary hardware, and why chunking matters more than the model.

— Neel, Jun 2026

3min read

a quick one

AI × Web3Jun 4, 2026

Why "Private AI" Usually Isn't

Legal contracts, audit reports, HR records — a lot of useful documents can't be sent to a hosted model. Not for performance reasons, for the simple reason that the data isn't allowed to leave the building.

Most "private AI" products still phone home somewhere: embeddings are computed by an API, or the model is hosted, or telemetry sneaks out. The only setup I trust for this class of document is one where nothing makes a network call after the models are downloaded.

The Pipeline

It is a standard retrieval-augmented generation (RAG) pipeline, with every stage running locally:

text
Documents (PDF, TXT, DOCX)
    → Text chunking (LangChain RecursiveCharacterTextSplitter)
    → Local embeddings (sentence-transformers/all-MiniLM-L6-v2)
    → Chroma vector store (local SQLite)

Query
    → Embed the query with the same local model
    → Cosine similarity search → top-k chunks
    → Prompt assembly
    → LlamaCpp inference (Mistral 7B, Q4_K_M GGUF)
    → Answer

Ingestion: Where the Quality Actually Comes From

python
from langchain.text_splitter import RecursiveCharacterTextSplitter
from langchain.embeddings import HuggingFaceEmbeddings
from langchain.vectorstores import Chroma

splitter = RecursiveCharacterTextSplitter(chunk_size=500, chunk_overlap=50)
embeddings = HuggingFaceEmbeddings(model_name="all-MiniLM-L6-v2")
db = Chroma(persist_directory="./db", embedding_function=embeddings)

for doc in load_documents("./source_docs"):
    chunks = splitter.split_documents([doc])
    db.add_documents(chunks)
db.persist()

I settled on 500-token chunks with a 50-token overlap. Smaller chunks make retrieval more precise for specific facts; the overlap stops a key sentence being cut in half at a chunk boundary.

This is the part to obsess over. RAG quality is dominated by chunking, not by the model. Chunks that are too large, no overlap, or sloppy document parsing produce worse answers than a smaller model fed good chunks. If I had to split tuning time, I'd spend 80% of it on ingestion.

Inference on the CPU

python
from llama_cpp import Llama

llm = Llama(
    model_path="./models/mistral-7b-instruct-v0.2.Q4_K_M.gguf",
    n_ctx=4096,
    n_threads=8,
    n_gpu_layers=0,  # CPU-only, so it also runs on air-gapped machines
)

def answer(query: str, context_chunks: list[str]) -> str:
    context = "\n\n".join(context_chunks)
    prompt = f"[INST] Answer based only on the context below.\n\n{context}\n\nQuestion: {query} [/INST]"
    return llm(prompt, max_tokens=512)["choices"][0]["text"]

Mistral 7B at Q4_K_M quantisation gave me the best quality-to-speed ratio for CPU-only inference: it fits in 8 GB of RAM and runs roughly four times faster than full precision, for about a 1% accuracy cost.

The prompt matters too. "Answer based only on the context below" is doing real work: it keeps the model from filling gaps with confident guesses.

What It Costs in Latency

HardwareIngestion (100 pages)Query latency
MacBook M1 (CPU)~45s~8s
8-core Intel (CPU)~90s~18s
RTX 3080 (GPU)~12s~2s

CPU-only is fine for occasional questions. If a team is going to query all day, put it on a GPU.

Where This Fits

  • Legal teams searching contract archives
  • HR teams searching policy documents
  • Financial institutions with data-residency rules
  • Air-gapped government environments

When to Outgrow It

Chroma's local SQLite backend is comfortable up to around ten million chunks. Beyond that, move to a dedicated vector database such as Qdrant or Milvus — and accept that you lose the zero-dependency simplicity that made this setup attractive in the first place.

The broader point: for sensitive documents, "private" should be a property you can verify by unplugging the network cable. If the pipeline still works, it's private.

Keep reading

Free · Weekly ✦

Enjoyed this one?

Get The Architect's Brief — weekly insights on blockchain architecture, AI × Web3, and engineering leadership.

Subscribe free →