Skip to content

AI & Machine Learning·9 min read

RAG, Fine-Tuning or Long Context: How We Choose

Start with the simplest setup that passes an eval on your own questions: a cached long prompt for a small, stable corpus, RAG when knowledge changes, must be cited or is access-restricted, and fine-tuning only for behaviour. Here is the order we check things in and where each option breaks.

Anatoli NavahrodskiFounder & CEO, GlanitPublished 7 October 2026

What is the difference between RAG, fine-tuning and long context?

Long context puts the source material straight into the prompt. RAG (retrieval-augmented generation) searches an index at request time and sends the model only the passages that match the question. Fine-tuning changes the model's weights with examples, which makes it good at changing how the model answers and poor at telling it what is true. Material small enough to send every time suits long context; material too large, too fresh or too restricted to send whole needs retrieval; right facts with wrong behaviour is a tuning job.

The question got harder this year. Context windows of 1M tokens are now ordinary on frontier models (a May 2026 benchmark tested five, all advertising 1M), and every major API discounts repeated prompt prefixes. "Just put everything in the prompt" stopped being a joke.

It is still not free, and it is still not always accurate.

How do we choose for a client project?

We ask four questions in a fixed order. Does the knowledge change often, or must every answer cite its source? Does the whole corpus fit comfortably in the context window? Is the real problem behaviour (format, tone, labels) and not missing facts? Is access scoped per user or per tenant? Then we build the simplest candidate and run it against 50 to 100 real questions collected from the client before we add a single component.

  1. Knowledge changes weekly, or answers need citations: RAG. A long prompt can quote its sources too, but every edit to the corpus invalidates the cached prefix.
  2. Corpus fills well under half the window and rarely changes: long context with prompt caching. No index, no ingestion pipeline.
  3. Answers are factually right but come out in the wrong shape: better prompts and few-shot examples first, fine-tuning only if that fails at volume.
  4. Different users may see different documents: RAG with permission filters, whatever the corpus size. This one overrides question 2, because a prompt has no access control.

"Under half the window" is our rule of thumb, not a vendor figure. We picked it after watching accuracy fall well before the advertised limit, and we move it whenever the client's eval set disagrees.

When is long context enough and RAG unnecessary?

Long context is enough when the corpus fits comfortably in the window, changes rarely, and the same prefix is reused across many requests so prompt caching pays off. A product handbook or the contracts for one deal are typical. Send it whole, cache it, and skip embeddings, chunking and the vector store. You also avoid the classic RAG failure, a retriever that misses the passage with the answer. The price is accuracy that drifts as the prompt grows, so we never trust it before testing on the client's own documents.

Caching is what changed the economics. On Anthropic's API a cache read costs 0.1x the base input price (0.05x or 0.025x on a few of the newest models), a 5-minute cache write costs 1.25x and a 1-hour write costs 2x. OpenAI caches automatically and describes cached input as "discounted up to 95%"; on its newest models reads cost 0.1x, writes 1.25x, and an entry lives 30 minutes after its last use. Google bills cached Gemini 2.5 Pro input at $0.125 per million tokens against $1.25 uncached (prompts up to 200K tokens), plus storage at $4.50 per million tokens per hour. That storage fee bites: a big cache kept alive overnight with no traffic still costs money.

resp = client.messages.create(
    model=MODEL,
    max_tokens=1024,
    system=[
        {"type": "text",
         "text": "Answer only from the handbook. Quote the section number."},
        {"type": "text", "text": handbook,
         "cache_control": {"type": "ephemeral", "ttl": "1h"}},
    ],
    messages=[{"role": "user", "content": question}],
)
print(resp.usage.cache_read_input_tokens)  # 0 means full price

Log that last value in production. One changing byte in the prefix (a timestamp in the system prompt is the usual culprit) silently turns every request into a full-price cache write, and nobody notices until the invoice.

Accuracy is the less comfortable part. Chroma's "Context Rot" report from July 2025 tested 18 models and every one of them got worse as input grew; on LongMemEval questions, models did markedly better with a focused prompt of about 300 tokens than with the full one of about 113K. A May 2026 study of five 1M-token models found single-fact lookup solved (100% for the three strongest) while three-hop reasoning split apart: Gemini and Claude stayed above 80% up to 512K tokens, GPT-5.5 and Qwen3.6-plus fell sharply between 512K and 1M, and DeepSeek V4 Pro declined across the whole range.

When does RAG beat fine-tuning and long context?

RAG wins when knowledge changes, when the corpus is far larger than any window, when users have different access rights, or when every answer must point to the passage it came from. Re-indexing one edited document takes seconds; re-training a model or rebuilding a cached prompt does not. Retrieval also keeps prompts short, which keeps the model in the range where the context-rot research found it most accurate. The cost moves into engineering instead: ingestion, chunking, metadata, access filters and a retrieval eval that you maintain like any other test suite.

Permissions decide more projects than corpus size does. Whatever sits in a prompt, the model can repeat to whoever is asking. With retrieval the filter runs before the model sees anything, so a document the user cannot open never reaches it.

Building it well is a separate topic. Our RAG go-live checklist covers evals and operations, the hybrid search write-up explains why we pair keyword search with embeddings, and the pgvector article covers where the vectors can live.

When is fine-tuning the right choice?

Fine-tune when the model knows enough but behaves wrong. It drifts from your JSON schema under load, labels tickets inconsistently, ignores a house style that even a long prompt cannot hold, or you want a small, cheap model to do one narrow task at high volume. Do not fine-tune to teach facts. A tuned model keeps knowledge with no pointer to where it came from, goes stale the day a document changes, and can produce text that looks like a citation with no evidence behind it.

Recent research points the same way. A CoLM 2026 paper by Kaplan, Gekhman and co-authors (first posted in April 2026) identifies new facts introduced through supervised fine-tuning as a key source of hallucinations: the model picks them up and forgets overlapping facts it already knew, and forgetting grows with that overlap. Their fixes (self-distillation, freezing parameter groups) are research, not a vendor checkbox.

The vendor side moved as well. On 7 May 2026 OpenAI closed self-serve fine-tuning to organisations that had never run it, and active existing customers can no longer create new jobs from 6 January 2027; fine-tuned models keep serving until their base model is deprecated. So a fine-tuning plan now starts with a written answer to "what happens when the base model is retired?"

Five setups compared

Read the last column first. It is the condition under which each setup becomes our starting point, and no single row is the starting point for everything.

RAG, fine-tuning and long context side by side (qualitative; figures with sources are in the text)
ApproachBest forWeak spotFreshness of knowledgeCitationsMain cost driverOur default when
Long context (+ prompt caching)Small, stable corpus; questions about whole documentsAccuracy drops as the prompt grows; latency on big promptsAs fresh as the last prompt rebuild; every edit resets the cachePossible by quoting; no access controlTokens per request, cache hit rate, cache storage on GeminiCorpus well under half the window, rarely changes, one access level
RAGLarge or changing corpora, per-user accessA missed retrieval produces a confident wrong answerCurrent as soon as the document is re-indexedNative: each answer points to passagesIngestion, index upkeep, retrieval evalsData changes often, or citations or permissions are required
Fine-tuningFormat, tone, labels; a small model at volumeFrozen knowledge, forgetting, citation-shaped text without evidenceFrozen at training timeNot reliableLabelled examples, training and eval runs, re-training on changePrompts and examples fail on behaviour at volume
RAG + long contextAnswers that need the text around a passageMore to tune; bigger prompts than plain RAGSame as RAGNativeRetrieval plus larger promptsRAG finds the right document but answers miss its context
RAG + fine-tuningCurrent facts with strict output behaviourTwo pipelines to maintain and re-evaluateSame as RAGNative, if the tuned model is trained to quote passagesBoth of the aboveWorking RAG still shows a behaviour gap that prompting cannot close
RAG, fine-tuning and long context side by side (qualitative; figures with sources are in the text) — "Half the window" is our rule of thumb, not a vendor figure.

What drives the cost of each approach?

For long context the bill is prompt size times request count, reduced by the cache hit rate. For RAG it is the pipeline: ingestion, embeddings, index hosting, re-indexing on change and an eval set someone keeps current. Fine-tuning costs training and evaluation runs, repeated whenever requirements or the base model change. The item most often missing from estimates is shared by all three: engineering time for the evaluation that tells you whether any of it works.

  • Long context: prompt size multiplied by traffic, and the hit rate. Requests arriving more than 5 minutes apart miss a default Anthropic cache unless you pay the 2x price for 1-hour writes. On Gemini, add the hours a cache is kept alive.
  • RAG: document volume and change rate, the number of source systems to connect, whether permissions have to be mirrored.
  • Fine-tuning: how many labelled examples already exist (writing them is usually the expensive part), how often requirements change, and the migration when the base model goes away.

A low-traffic internal tool over a 200-page manual is cheapest as a cached prompt. Put the same manual behind a busy public assistant with a document set per customer, and it becomes a RAG system.

How do the three fail in production?

Each approach has one failure that tends to appear only after launch. Long context degrades quietly as documents get appended and the prompt grows, and latency climbs with it. RAG fails when retrieval misses: the model answers fluently from the wrong passage, and users believe it because it cites something. Fine-tuning fails slowly, as knowledge freezes on training day and new facts pushed in through training erode facts the model already had.

The fix is the same in all three cases, and it is boring. Keep a set of real questions with known correct answers and re-run it on every change of prompt, model, index or document set. For RAG, score retrieval separately from the final answer so you know which half broke.

Leaks are their own category. A shared long prompt and a set of fine-tuned weights both hold everything they were given, for everyone who can query them. Retrieval with filters is the only one of the three where "this user may not see this document" can be enforced in code. See our OWASP Top 10 for LLM applications walkthrough, and how AI agents work for what changes once the model calls tools.

Can you combine them, and when should you?

Yes. The two hybrids worth knowing are RAG plus long context and RAG plus fine-tuning. In the first, retrieval narrows a large corpus to a candidate set of whole documents or long sections (not 300-token chunks), and the large window lets the model read them with their surroundings. In the second, retrieval supplies current evidence while a tuned model handles the strict behaviour: schema, labels, tone. We add a layer only when the eval set shows a gap that the simpler setup cannot close.

Order matters more than the choice of hybrid. Start with a prompt. Add caching when the prefix repeats, retrieval when the corpus, its change rate or its permissions demand it, and fine-tuning last, for behaviour only. Tuning first, then finding out the facts change weekly, means paying for training runs the next document update makes obsolete.

Not sure which one fits your data?

Send us the use case, a sample of the documents and twenty questions your users actually ask. We will tell you which setup above we would start with and which eval would confirm or kill it. Details on the AI and machine learning services page, or write to us through the contact form.

Frequently asked questions