completed · featured
VoxRAG is a voice- and text-enabled Retrieval-Augmented Generation system built on MSMARCO-XI. Speak or type in English, हिन्दी, मराठी, বাংলা, മലയാളം, ગુજરાતી, অসমীয়া — the pipeline transcribes, retrieves grounded passages from ~3.2M (32 lakh) embedded records, and streams a cited answer back, targeting <200 ms retrieval / time-to-first-token.

VoxRAG is a voice- and text-enabled Retrieval-Augmented Generation system built on MSMARCO-XI. Speak or type in English, हिन्दी, मराठी, বাংলা, മലയാളം, ગુજરાતી, অসমীয়া — the pipeline transcribes, retrieves grounded passages from ~3.2M (32 lakh) embedded records, and streams a cited answer back, targeting <200 ms retrieval / time-to-first-token.
📦 Ingestion deep-dive: how the corpus was minified, sharded across 5 T4 workers, chunked, checkpointed and indexed is documented in ingestion/README.md.
saaras:v3, ElevenLabs fallback)halfvec(384) storage — FP16 vectors, half the bandwidth of float32openai/gpt-oss-120b| Layer | Technology |
|---|---|
| Frontend | Streamlit (custom "Quiet Studio" theme) |
| Speech-to-Text | Sarvam AI Saaras v3 (ElevenLabs fallback) |
| Embeddings | intfloat/multilingual-e5-small (FP16, 384-d) |
| Vector DB | PostgreSQL + pgvector (halfvec, IVFFlat) |
| LLM | Groq openai/gpt-oss-120b (OpenAI-compatible client) |
| Corpus | MSMARCO-XI (minified JSONL via Kaggle amankroot) |
| Tooling | uv, Python 3.12, ffmpeg |
text. ├── README.md # This file ├── analytics.py # Rolling-window P50/P70/P100 latency tracker ├── app.py # Streamlit frontend ("Quiet Studio") ├── config.py # Pydantic settings + language map ├── latency_logs.jsonl # Telemetry log (rolling window source) ├── packages.txt # System deps (ffmpeg) ├── pyproject.toml / uv.lock # uv-managed dependencies ├── ingestion/ │ ├── README.md # 👈 Ingestion deep-dive (read this!) │ ├── download_corpus.py # Pulls minified MSMARCO-XI JSONL from Kaggle │ ├── multi_embed_marco_chunk.py # Distributed multi-T4 embedding worker │ └── semantic_chunker.py # Token-aware chunking (E5 offset_mapping) ├── pipeline/ │ ├── orchestrator.py # Synchronous harness: retries, timing, events │ ├── stt.py # Sarvam / ElevenLabs clients │ ├── retriever.py # pgvector parent–child retrieval │ ├── llm.py # Groq streaming client │ └── guardrails.py # Input filter + citation verifier └── src/voice_rag_project/ # Package stub
halfvec)uv and ffmpegbashsudo apt-get update && sudo apt-get install -y ffmpeg # system deps uv sync # python deps
Create .env:
envDATABASE_URL=postgresql://user:password@host:5432/voxrags_db GROQ_API_KEY=gsk_... SARVAM_API_KEY=... ELEVENLABS_API_KEY=... # optional fallback STT_PROVIDER=sarvam VECTOR_BASE=halfvec TABLE_NAME=corpus_embeddings_chunks GROQ_MODEL=openai/gpt-oss-120b
sqlCREATE EXTENSION IF NOT EXISTS vector; CREATE TABLE IF NOT EXISTS corpus_embeddings_chunks ( source_file text NOT NULL, line_no bigint NOT NULL, chunk_no integer NOT NULL, id text, query_id text, chunk_start integer, chunk_end integer, text text, embedding halfvec(384) NOT NULL, updated_at timestamptz NOT NULL DEFAULT now(), PRIMARY KEY (source_file, line_no, chunk_no) ); -- ANN index: IVFFlat, sized for a 4 GB RAM instance SET maintenance_work_mem = '2GB'; CREATE INDEX corpus_embeddings_chunks_embedding_idx ON corpus_embeddings_chunks USING ivfflat (embedding halfvec_cosine_ops) WITH (lists = 1800); -- ≈ sqrt(row count) ANALYZE corpus_embeddings_chunks;
Why IVFFlat and not HNSW? Building HNSW over ~3.2 M
halfvec(384)rows needs 10–12 GB RAM; this project runs on a 4 GB instance. IVFFlat builds in minutes within that budget and still delivers sub-50 ms retrieval.
Full details, iteration history and worker sharding live in ingestion/README.md.
bashuv run python ingestion/download_corpus.py # Kaggle JSONL FILE_START=1 FILE_END=500 uv run python ingestion/multi_embed_marco_chunk.py # worker 1 FILE_START=501 FILE_END=1000 uv run python ingestion/multi_embed_marco_chunk.py # worker 2
bashuv run streamlit run app.py
Chunking. semantic_chunker.py uses the E5 tokenizer's offset_mapping for token-aware boundaries (400 tokens + 50 overlap), preserving metadata (source_file, line_no, chunk_no, char offsets) on every chunk.
Retrieval. The query embeds with the query: prefix; pgvector returns the nearest chunks; a self-join expands ±2 chunks around each match so the LLM sees continuous context, not isolated fragments.
Latency budget. "200 ms" targets retrieval + TTFT, not full generation (physically >500 ms for a paragraph). FP16 halfvec + IVFFlat + Groq streaming + a synchronous (no-asyncio) orchestrator keep the perceived response instant.
[1]…[n] in the answer are verified against the actual retrieved chunk IDs; invented sources flip the matrix to Review.August 2026