Concept Specification
ai-ml2025-08-01

Architectures of Intelligence: Advanced RAG and Context Engineering

A tiered framework for production RAG systems — chunking strategies, query transformation, two-stage re-ranking, prompting patterns, and agentic self-correction loops (CRAG, SELF-RAG) — organized around 'context failures, not model failures.'

Overview

Production-grade AI systems have matured from simple prompt engineering to context engineering — the discipline of architecting what an LLM knows, not just what it's told. Retrieval-Augmented Generation (RAG) solves foundational models' static knowledge, domain-specificity gaps, and expensive retraining by dynamically injecting external knowledge at inference time. Most RAG failures are context failures (a systems design problem), not model failures.

Key Concepts

  • Context engineering vs. prompt engineering — prompt engineering designs a single textual input; context engineering designs the whole system providing the LLM with information, tools, and memory: system instructions, chat history, retrieved information (RAG), and tool definitions/API responses.
  • Chunking tradeoff — balancing precision (small chunks, specific queries) against context (large chunks, preserved meaning). The five main strategies (fixed-size, recursive, document-based, semantic, agentic) trade implementation simplicity for semantic coherence, roughly in that order.
  • The “vocabulary gap” — user queries are often ambiguous or keyword-poor relative to the documents that would answer them; query transformation techniques (Multi-Query, RAG-Fusion, Step-Back Prompting, HyDE, Query Routing) use an LLM to bridge this gap before retrieval.
  • The “Lost in the Middle” problem — LLMs recall information at the start/end of a context window far better than the middle; a re-ranker combats this by reordering retrieved documents to place the most relevant ones at the attentional “spotlight” edges.

Pre-Retrieval: Chunking Strategies

StrategyMechanismBest For
Fixed-SizeSplit by fixed character/token countQuick prototyping, unstructured text
RecursiveSplit via prioritized separator list (paragraphs → sentences)General-purpose documents; often the best default
Document-BasedSplit along document structure (Markdown, HTML, code)Highly structured technical/legal documents
SemanticGroup semantically similar sentences via embeddingsNarrative or user-generated content
AgenticLLM decides the split, simulating human reasoningHigh-value docs where indexing cost is justified

Contextual Embeddings: an LLM summarizes a chunk's surrounding context before embedding, enriching the vector and dramatically improving retrieval.

Query Transformation Techniques

TechniquePrincipleBest For
Multi-Query RetrievalGenerate multiple query variations, merge resultsComplex, multifaceted questions
RAG-FusionMulti-query + Reciprocal Rank Fusion re-rankingWhen relevance ordering is critical
Step-Back PromptingGenerate a more general question, retrieve with bothHighly specific/jargon-heavy queries
HyDEEmbed a hypothetical answer, not the query itselfShort, ambiguous, keyword-poor queries
Query RoutingLLM selects the right data source (vector DB, SQL, etc.)Multi-source enterprise systems

Post-Retrieval: Re-ranking

Two-stage retrieval balances speed and accuracy: a fast first-stage retriever optimizes for recall (find all candidates), then a slower second-stage re-ranker optimizes for precision (sort the truly relevant ones to the top).

ArchitecturePerformanceCostExamples
Cross-EncodersVery HighHighBGE-Reranker, sentence-transformers
Late Interaction (ColBERT)HighMediumColBERT
LLM-basedVery HighVery HighRankGPT, RankZephyr, RankT5
Private APIsHigh–Very HighMediumCohere Rerank, Jina AI

Augmentation: Prompting Patterns for RAG

  • Direct Retrieval Pattern — answer only from provided context; maximizes grounding but can produce overly cautious “I don't know” responses.
  • Chain-of-Thought Inspired — identify key points → outline → write; improves reasoning but adds latency/tokens.
  • Persona-Based — sets tone and domain focus; risk of an oversimplified persona omitting expert nuance.
  • Error Handling / “escape hatch” — give the model an explicit “insufficient information” response option, reducing hallucination when grounding data is missing.
  • Multi-Pass Refinement — generate, self-review for factual consistency, refine; better accuracy at the cost of processing time.

Agentic RAG: Beyond Linear Pipelines

Advanced RAG moves from a linear retrieve→augment→generate pipeline to a cyclical, agentic state machine (e.g., built with LangGraph) that can loop back and self-correct.

  • Corrective RAG (CRAG) — a lightweight “retrieval evaluator” quality-gates results; triggers a corrective action (e.g., web search) if retrieved docs are irrelevant.
  • SELF-RAG — the LLM generates “reflection tokens” to decide if retrieval is needed, grade document relevance, and critique its own output for factual support.
  • Generate → Critique → Refine loop — the core iterative pattern behind reliable agentic AI: draft, evaluate (by the LLM or an external tool), improve.

A Tiered Framework for Practitioners

  1. Level 1 (Baseline RAG) — simple pipeline for proofs-of-concept.
  2. Level 2 (Optimized Retrieval RAG) — adds a re-ranker and query transformation; ideal for most production systems.
  3. Level 3 (Advanced Context RAG) — adds context compression and Chain-of-Thought reasoning patterns.
  4. Level 4 (Agentic RAG) — implements self-correction loops for mission-critical, maximum-reliability applications.

Key Takeaways

  • The framing of "context failures, not model failures" is the article's organizing thesis — every technique covered (chunking, query transformation, re-ranking, prompting patterns, agentic loops) is a systems-design lever, not a model-capability upgrade.
  • The tiered practitioner framework is explicitly a cost/complexity ladder, not a "always use the most advanced technique" recommendation — Level 2 (re-ranker + query transformation) is called out as sufficient for most production systems, with Levels 3-4 reserved for cases that specifically justify the added latency and cost.
  • Re-ranking and chunking strategy selection both hinge on the same underlying tradeoff (precision vs. recall/context), suggesting the whole pipeline should be tuned as a coupled system rather than optimizing each stage independently.

Related Reading

Companion Research Article

Advanced RAG and Context Engineering

From prompt engineering to context engineering: advanced RAG architectures, query transformation, re-ranking, and the agentic systems behind production AI.

Comments

Disclaimer: This application is a personal proof of concept created for study and research purposes only. All analysis, suggestions, and content are generated by AI models using publicly available data and tools, and should not be considered as financial advice. Past performance is not indicative of future results. Always conduct your own research and consult with qualified financial professionals before making investment decisions. The app's AI models may have limitations and may not account for all market factors or recent developments. Users are solely responsible for their investment decisions and should understand that all investments involve risk.