RAG or Fine-Tuning on AWS: Which Customization Fits, and What It Costs
Published
A policy changes, a catalogue entry is replaced, or an approved document is withdrawn—and an assistant must use the new information in its very next answer. Fine-tuning can sound like the thorough response, but it solves the wrong problem when the missing ingredient is current knowledge. The practical choice is structural: use retrieval-augmented generation with an appropriate AWS vector store when the model needs changing facts or traceable sources; fine-tune only when its behaviour remains wrong after effective prompting. This distinction also explains the cost ladder: retrieval adds request-time costs, while training creates an ongoing obligation to train, host, evaluate, and maintain a customized model.
The first decision is whether facts or behaviour are wrong
A foundation model can fail in two fundamentally different ways. It may lack information, or it may handle available information badly. Those failures can produce equally disappointing answers, but they require different remedies.
A facts problem occurs when the model has not seen your material or cannot be trusted to remember its current version. Internal policies, approval records, product details, and operational procedures all fit this pattern. Supplying those facts at request time addresses the gap without changing the model.
A behaviour problem occurs when the model has the necessary information but responds in the wrong format, tone, notation, or domain style. That is a candidate for better prompting first and fine-tuning only if prompting genuinely cannot produce the required behaviour.
Consider an assistant that answers from an organization’s procedures. If it cites an obsolete procedure, the system has a content and retrieval problem. If it uses the correct procedure but consistently writes in an unsuitable style, it has a behaviour problem. Retraining for the obsolete document would confuse content with conduct; retrieving another style guide would confuse conduct with content.
Missing facts call for retrieval; persistent behaviour problems may justify fine-tuning.
A useful rule is simple: facts are retrieved; behaviour is tuned. If neither problem has been established, improve the prompt before adding infrastructure or training.
RAG supplies current evidence without changing the model
Retrieval-augmented generation, or RAG, fetches relevant passages from your sources when a question arrives and places them in the prompt. The foundation model answers using that material, but its weights—the learned parameters that determine its behaviour—remain unchanged.
Think of RAG as an open-book exam. Giving someone the correct page does not alter what they have learned; it changes what evidence is available while they answer. The same principle makes RAG suitable for information that changes too often to embed safely in a trained model.
A RAG system has two distinct pipelines. During ingestion, source documents are split into passages, converted into numerical representations called embeddings, and indexed in a vector store. Ingestion runs when the source material changes. During retrieval, the user’s question is embedded, similar passages are found, and those passages are assembled into the prompt. Retrieval runs for each request.
Documents are indexed when sources change; questions retrieve passages when requests arrive.
This structure gives RAG three important advantages over fine-tuning for factual knowledge. New material becomes available as soon as the index is updated. Retrieved passages can accompany the answer as citations. The model does not acquire a separate trained version that must be retrained whenever the source changes.
RAG still requires evaluation. A citation proves provenance—where the supporting passage came from—not correctness. If an obsolete document remains indexed, the answer may confidently cite the wrong policy. Retrieval quality, source approval, and index freshness therefore remain operational responsibilities.
Amazon Bedrock Knowledge Bases reduces retrieval plumbing
Building RAG involves document splitting, embedding, indexing, retrieval, prompt assembly, and citation handling. Amazon Bedrock Knowledge Bases provides these functions as a managed RAG service.
You supply documents, select an embedding model, and use a vector store or managed default. The service handles the mechanics of preparing passages, keeping the index current, retrieving relevant material, assembling it with the question, and returning citations.
This option is especially suitable when the requirement is “answer from our documents without building a retrieval pipeline.” It changes who operates the plumbing, not the underlying distinction between retrieval and training. The retrieved text still joins the prompt, and the foundation model’s weights still do not change.
Managed RAG handles the retrieval pipeline while the organization supplies sources and choices.
The managed service does not eliminate architectural choices. In particular, the vector store should match the existing data estate and the shape of the queries. The question is not which store is universally best. It is which one avoids unnecessary duplication while supporting the required retrieval pattern.
The existing data estate usually chooses the vector store
AWS provides several services that can store embeddings and perform vector search. They are not interchangeable entries in a quality ranking. Each fits a different data context.
Amazon OpenSearch Service fits a search-first workload, especially when a large corpus needs both keyword and vector retrieval. Amazon Neptune fits cases in which relationships between entities are part of the answer and graph traversal must work alongside similarity search.
Amazon Aurora and Amazon RDS for PostgreSQL fit relational estates. Both can support vectors alongside relational data, but the existing footprint is the deciding signal. If an application already uses Aurora, keeping an item’s embedding beside its price, status, or approval metadata can avoid introducing another store. If the existing estate is RDS for PostgreSQL, adding vector support there may achieve the same consolidation without a migration.
Query shape and the existing data estate select the vector store.
For example, suppose approved procedures are already held in Aurora beside revision status. Storing their embeddings in Aurora allows similarity search and relational filters to work over the same governed records. A separate vector database would add data synchronization, security, backup, and cost responsibilities without necessarily improving the result.
That does not make Aurora the default vector store. If the workload is fundamentally search-first, OpenSearch Service may be the stronger fit. If relationships determine the answer, Neptune addresses a requirement that similarity alone cannot. The sound decision begins with the data and query, not a preferred service name.
Cost increases when customization changes the model
Customization forms a ladder. Each rung adds commitment and should be chosen only when the cheaper rung cannot meet a stated requirement.
In-context learning supplies instructions and examples in the prompt. It has no training cost but consumes prompt tokens on each request.
RAG adds embedding, retrieval infrastructure, vector storage, and larger prompts. It buys access to current facts and citations without training. Its costs grow through ingestion and request-time use, but the model itself remains untouched.
Fine-tuning adapts the model’s weights using examples. It adds a training job, custom hosting, evaluation, and maintenance. It is justified when required behaviour—such as format or domain style—remains unattainable after effective prompting. It is not a sensible way to keep frequently changing facts current.
Distillation trains a smaller student model to reproduce the behaviour of a larger teacher model. It spends more up front so that each later request can be cheaper and faster. Volume, rather than missing capability, is the condition that justifies it.
Pre-training builds a model from scratch. It demands extensive data, compute, and expertise and is appropriate only when no suitable existing model fits the domain.
Each higher customization rung needs a condition strong enough to justify greater commitment.
Think of this ladder as tailoring. You do not commission an entirely new garment because an existing one needs a minor adjustment. Likewise, you should not create a customized model because a prompt needs examples or because a document belongs in retrieval.
The decisive cost difference is ownership. RAG creates infrastructure and per-request costs, but the underlying foundation model remains shared and unchanged. Fine-tuning and the higher rungs create a model version that must be hosted, tested, monitored, and refreshed. Training expense is therefore only the beginning of their cost.
Combined requirements may need more than one technique
Real applications often contain both factual and behavioural requirements. Choosing RAG does not mean every problem in the application is a retrieval problem.
Imagine an assistant that answers questions about approved procedures stored in Aurora. The procedures change regularly, every answer must cite an approved source, and responses are too informal.
The first two requirements point to RAG: the facts change, and traceability is mandatory. Aurora is a natural vector-store candidate because the procedure text already sits beside approval and revision metadata. The tone issue should be handled separately with better instructions and examples. Fine-tuning becomes reasonable only if that behaviour remains wrong after prompting has been tested properly.
This layered design is often cheaper than forcing one technique to solve everything. Retrieval supplies evidence; prompting controls the immediate response; fine-tuning is reserved for a demonstrated behavioural gap. The techniques can coexist, but their responsibilities should not blur.
A particularly costly failure mode is training a model on changing documents and then treating retraining as a synchronization mechanism. The trained model represents a snapshot, cannot naturally cite the exact approved passage, and creates another artifact to maintain. Increasing the training frequency does not repair the category error.
Key takeaways
- Use RAG when the model lacks private or changing facts, or when answers must cite approved sources.
- Use fine-tuning for persistent problems with behaviour, format, tone, or domain style—not for factual freshness.
- Try in-context learning before fine-tuning; climb the customization ladder only when a specific condition requires it.
- Use Amazon Bedrock Knowledge Bases when managed retrieval is preferable to building the pipeline.
- Select a vector store from the existing data estate and query shape: OpenSearch Service for search, Aurora or RDS for PostgreSQL for relational data, and Neptune for graph relationships.
- Compare total obligations, not only initial expense: RAG adds retrieval and token costs, while fine-tuning adds training, hosting, evaluation, and maintenance.
- Treat citations as evidence of provenance, not proof that an answer is correct.
- Consider distillation when a capable model’s per-request cost is the problem at volume, not when capability is missing.
This framework settles which customization family fits a stated problem and how their cost structures differ. It does not settle the final architecture or operating cost, which still depends on corpus size, request volume, model choice, retrieval quality, and the system’s evaluation requirements.