ARIA
RAG assistant for PDF, YouTube and GitHub ingestion — embeddings, Chroma, semantic retrieval, memory and Groq/LangChain orchestration.
01 — System Overview & Primary Intent
ARIA is a specialized system engineered by Team Paradox in Gorakhpur, led by Om Abhishek Tripathi. Teams learn from scattered PDFs, videos and repos. Needed one memory-aware assistant grounded in their own sources.
02 — Architecture & Flow
Ingest pipeline: chunk → embed (HF) → Chroma
Query: embed → retrieval → Groq synthesis with citations
Memory store for follow-ups
Source-grounded answers, no invented URLs
03 — Engineering Decisions & Rationale
- Chroma for local-first vector store
- Groq for low-latency synthesis
- Strict citation framing; fallback to 'not found in sources'
04 — Known Limitations & Failure Modes
Chunking heuristic for scanned PDFs is lossy
No fine-tuning — retrieval quality bounds answers
YouTube transcript dependent on captions
Observed Tech Stack
Next.jsFastAPILangChainGroqChromaHuggingFace
Engineering Governance
All systems undergo internal peer review with strict claims verification. We do not publish simulated benchmarks, fake testimonial quotes, or obfuscated AI wrappers without deterministic fallbacks.