Mini RAG

A modular Retrieval-Augmented Generation (RAG) system built from scratch to explore document indexing, semantic retrieval, prompt engineering, streaming responses, and production-ready AI architectures.
Media






Overview
Mini RAG is a modular, production-style Retrieval-Augmented Generation (RAG) system built from scratch using FastAPI, Qdrant, HuggingFace Embeddings, Groq, RQ, and Docker.
Rather than relying entirely on high-level frameworks, the project implements each stage of the RAG pipeline as an independent, replaceable service—from document ingestion and chunking to embeddings, vector storage, retrieval, prompt construction, conversation memory, and streaming responses. The architecture is designed to be modular, making it easy to swap providers and extend the system with new capabilities.
Why I Built It
I originally started this project while building another terminal-based AI assistant. To implement document-based question answering, I needed a deeper understanding of how Retrieval-Augmented Generation works beyond simply using existing libraries.
What began as a small learning exercise gradually evolved into a complete RAG system with asynchronous indexing, custom vector retrieval, conversation-aware prompting, and a modular architecture. The project now serves as both a production-style foundation for future AI applications and a practical exploration of modern RAG system design.
Key Features
Document Processing
- Asynchronous document indexing using FastAPI, RQ, and background workers
- Configurable document chunking
- Rich metadata generation for every chunk
- Batch embedding generation for efficient indexing
- Batch vector uploads to Qdrant
Retrieval
- Native Qdrant retrieval implementation
- Optional LangChain-based retrieval
- Retrieval score threshold filtering