Finetuning vs RAG: Diagnose the Failure Before You Change the Model
A product team can spend months finetuning a model to answer a question it could never know. Another can build a retrieval system around a model that already has the facts, but cannot reliably produce the required output. Both teams have selected an impressive mechanism before naming the failure.
Finetuning changes a model's learned behavior. Retrieval changes the evidence available for one request. Those are not interchangeable upgrades. They have different freshness properties, cost profiles, security boundaries, and ways of failing.
The useful question is not "should we use RAG or finetuning?" It is: what must change for this output to become trustworthy?
Start With an Evaluation Set, Not a Technique
The decision is only defensible when a team can show a representative set of inputs, expected outcomes, and the errors that matter. "It feels unreliable" cannot tell an engineer whether the defect is missing information, weak instructions, poor retrieval, malformed output, or a product flow that conceals uncertainty.
Build a small evaluation set from real work before changing the architecture. Each case should record:
- the user request and the user or tenant context that affects authorization;
- the authoritative information needed to answer it, if any;
- the desired response, schema, action, or refusal;
- the cost of a wrong answer, a stale answer, an invalid structure, and an unnecessary escalation;
- enough slices to expose variation: new versus old data, ordinary versus ambiguous wording, and routine versus high-consequence requests.
This is especially important in a JavaScript application because the interface is often where an AI error becomes consequential. A response that appears while a user is editing an invoice, approving a change, or viewing private account data needs more than a good aggregate score. The product needs a traceable result for that exact state, user, and source version.
Without this baseline, finetuning can turn into a costly way to improve examples that were already easy, while retrieval can become an elaborate system for a prompt that was underspecified.
Information Failures Need an Evidence Path
An information failure occurs when the model needs a fact it does not have, or when the fact may have changed after the model was trained. Internal policies, current inventory, account state, recently revised documentation, and a customer's own records are all obvious examples.
Training those facts into weights is a poor freshness contract. It requires a new dataset, a new training run, an evaluation cycle, and a deployment whenever the source changes. It also makes it difficult to show where a particular answer came from or to remove a record everywhere it might influence future output.
Retrieval gives the request a bounded set of current, authorized evidence instead. That evidence can come from lexical search, vector search, a document store, or an application API. The mechanism matters less than the contract:
- identify the source of truth;
- filter it using the caller's authorization before it enters model context;
- select enough evidence to answer the question without flooding the model;
- preserve source identity, revision, and retrieval time with the answer;
- make a no-answer or clarification path legitimate when the evidence is missing or conflicting.
For an internal support assistant, this may mean retrieving the current plan rules and the caller's account state rather than embedding an entire knowledge base into a prompt. For a frontend engineer, the result should influence the UI: cite the policy revision, distinguish an empty retrieval from a model refusal, and do not render an answer as final while the evidence step is still unresolved.
Retrieval is not automatically safe because it is current. A vector database, chunk cache, observability trace, and browser response can each retain derived copies of sensitive material. Access control must happen before context construction, and deletion and retention policies must cover those derived representations rather than only the original document.
Behavior Failures Need a Stable Contract
A behavior failure looks different. The model may have the necessary facts, but it does not consistently use them in the form the product needs. Common cases include:
- a tool call that omits a required field or invents an enum value;
- a domain-specific syntax that does not parse reliably;
- a technical specification that is accurate but consistently misses decisions engineers need;
- a response style that violates a stable safety or editorial policy;
- a small, cheap model that needs to imitate a stronger model on a narrow repeated task.
Finetuning can help because the target is a recurring mapping from input to behavior. High-quality examples adjust the model so the desired structure, style, or task-specific transformation is more likely without repeating a long collection of examples in every request.
That benefit comes with a boundary: a finetuned model can still hallucinate, and a training example is not a live source. Low-quality or contradictory examples can make factual behavior worse. A model that is trained aggressively for one workflow can also lose quality on other requests. Do not treat an improved demo for one intent as proof that a shared assistant is better overall.
For structured output, start with the cheaper controls first: explicit schema instructions, constrained decoding or tool schemas where available, runtime validation, and repair or escalation paths. A TypeScript application should validate model output at the boundary, not cast it because the prompt said "return JSON."
const result = TicketSchema.safeParse(modelOutput)
if (!result.success) {
return { status: 'needs-review', issues: result.error.issues }
}
return { status: 'ready', ticket: result.data }
Finetuning is useful when this contract still fails after systematic prompt and interface work, and when the examples represent the production inputs the model will actually see.
Prompting Is the Baseline, Not an Embarrassing First Draft
Finetuning has a real up-front bill: data curation, training compute, model selection, evaluation, hosting, monitoring, and repeated comparison against newer base models. It also requires enough ML understanding to notice overfitting, unstable training, or a misleading loss curve.
That is why a disciplined progression usually begins with prompting:
- define task metrics and test cases;
- try clear instructions and representative examples;
- add retrieval when failures are caused by missing or changing evidence;
- improve retrieval only when the evidence path remains the defect;
- finetune when the remaining failures are stable behavioral ones;
- combine retrieval and finetuning only when each fixes a distinct observed problem.
Prompt caching changes part of the old economic argument for finetuning. Repeated examples may no longer carry the same per-request cost, though they still consume context and constrain how many examples fit. That makes measurements more important than inherited rules such as "few-shot is always too slow" or "training always makes the prompt shorter."
RAG and Finetuning Can Be Complementary
The most capable product may need both. Imagine an assistant that writes a change plan from current infrastructure records. Retrieval supplies the current topology, ownership, and policy constraints. Finetuning or a carefully engineered contract helps the model produce a plan in the team's required format. One supplies facts; the other shapes behavior.
The order matters. Start with the source path when facts are missing. A finetuned model that writes a perfect-looking answer from stale information is more dangerous than a model that clearly asks for the missing record.
There is also an operational distinction. More sophisticated retrieval adds components to every inference: indexing, searching, reranking, authorization filters, and context assembly. Finetuning makes model development more complex, but can leave the serving request simpler. Neither is universally cheaper. The right comparison includes data update frequency, latency targets, traffic, hardware, privacy constraints, and the people who will own failures after launch.
A Decision Rule That Survives Contact With Production
Use retrieval when a correct response depends on information that is private, current, frequently revised, or must be cited. Use finetuning when the evidence is already available but the model repeatedly fails a stable behavior or output contract. Use neither until a prompt experiment has shown what is actually inadequate.
That rule is deliberately less glamorous than "RAG versus finetuning." It is also easier to test. For every proposed change, ask which evaluation cases it should improve and which it might damage. If a training run cannot answer that question, it is not yet an engineering decision. It is an expensive guess.
