Skip to content
All articles

·4 min read

RAG, Fine-Tuning or Just a Better Prompt? How to Choose the Right Approach

"ChatGPT, but with our data" can mean three very different things. A practical guide to prompting, RAG and fine-tuning: what each one fixes, what it costs and in which order to try them.

#ai#rag#architecture

“We want ChatGPT, but with our data.” Almost every AI conversation with a company starts like this. Behind the sentence are three very different technical approaches: a good prompt, retrieval-augmented generation (RAG) and fine-tuning. They differ in cost, effort and what they can actually fix. Picking the wrong one is the most common reason AI pilots stall.

Start with the question, not the technique

Before choosing, write down what the model gets wrong today. The answer usually falls into one of three buckets:

  • It doesn’t know something (your products, contracts, internal processes) → a knowledge problem.
  • It knows, but answers in the wrong form (tone, structure, length, format) → a behaviour problem.
  • It doesn’t understand the task → usually a prompt problem.

Knowledge problems are solved with RAG. Behaviour problems can be solved with fine-tuning. And a surprising number of problems disappear with a better prompt.

Level 1: a better prompt

Clear instructions, a few good examples and a defined output format (for example JSON) are free and take hours, not weeks. Modern models follow instructions well. Before investing in anything else, build a small test set of 20–50 real questions with expected answers and see how far prompting gets you.

Good for: classification, extraction, rewriting, summaries, any task where the necessary knowledge is in the input itself. Limit: the model can only use what fits into the context. It cannot know your 4,000-page document archive.

Level 2: RAG — give the model the right pages

With RAG, your documents are split into passages and indexed. For each question, the system searches for the most relevant passages and puts them into the prompt. The model answers from those sources and can cite them.

Good for: internal knowledge assistants, support bots, contract and policy questions, product catalogues — anything where answers must be based on documents that change.

Why it is usually the right choice:

  • Up to date: a new document is searchable minutes after upload. No retraining.
  • Traceable: every answer can show its source. That matters for trust, audits and the EU AI Act’s transparency expectations.
  • Permissions: search can respect who may see which document. A fine-tuned model cannot “forget” data for a particular user.
  • Works with local models: RAG runs well with open-source models on your own hardware, so data never leaves your network.

Most of the effort in RAG is not the AI but the data: cleaning PDFs, splitting documents sensibly, keeping the index in sync and measuring whether the right passages are found. Search quality decides answer quality.

Level 3: fine-tuning — change how the model behaves

Fine-tuning continues training a model on your own examples (usually hundreds to thousands of input/output pairs). It is good at teaching style and format: your tone of voice, a strict report structure, domain-specific classification labels, or making a small, cheap model perform a narrow task as well as a large one.

It is a poor way to teach facts. The model does not reliably remember details from training data, it cannot cite sources, and every change in your knowledge means a new training run. It also needs clean, representative training data, which is often the hardest part.

Good for: consistent formats at high volume, narrow repetitive tasks, reducing cost and latency by moving a task to a smaller model.

A quick decision guide

Situation Approach
Answers must come from your documents RAG
Documents change weekly RAG
Users need to see the source RAG
Output format or tone is inconsistent Prompt first, then fine-tuning
Same narrow task, very high volume Fine-tuning (smaller model)
Data must stay on-premise RAG or fine-tuning with a local model

In practice the answer is often a combination: a well-designed prompt, RAG for knowledge and, only later, fine-tuning to make a proven workflow cheaper or more consistent.

What it costs, roughly

A prompt-based solution with an evaluation set can be live in days. A production RAG system (document pipeline, search, permissions, evaluation, UI) is typically a few weeks of work. Fine-tuning adds data preparation, training runs and ongoing re-training — worth it only once you know exactly which task you want to optimise.

The order that works

  1. Write down 20–50 real questions with good answers.
  2. Try a strong prompt and measure.
  3. If knowledge is missing, add RAG and measure again.
  4. If behaviour is still off at scale, consider fine-tuning.

Skipping step 1 is the most expensive mistake: without a test set you cannot tell whether any change made things better.

If you are not sure which level your use case needs, I’m happy to look at it with you in a free 30-minute call — usually that’s enough to name the approach and a realistic budget.

Get the AI Readiness Checklist

10 questions to answer before you start an AI project, plus occasional notes on applied AI and SaaS. Free.

Double opt-in, unsubscribe anytime with one click. Details in the privacy policy.

More articles