Home Work Enterprise RAG assistant for technical documents
Project
Enterprise RAG assistant for technical documents
I built a retrieval-augmented generation assistant over technical PDFs and internal documents: ingestion and chunking, embedding generation, vector storage for semantic retrieval, and citation-grounded answering with hybrid retrieval and relevance ranking to reduce hallucinations.
The problem
Technical documentation is where answers go to hide. The information exists, but it is spread across PDFs and internal documents, and a plain language model asked about it will produce something fluent and unverifiable. For an enterprise reader, an answer that cannot be traced back to a page is worse than no answer.
The pipeline
The system is a straight sequence, and each step exists to protect the one after it:
- Ingest: PDFs and internal documents are read into a common text representation.
- Chunk: content is split into retrievable units that stay semantically whole.
- Embed: each chunk is turned into a vector.
- Store: vectors are indexed for semantic retrieval.
- Retrieve and rank: hybrid retrieval pulls candidates, relevance ranking decides what actually reaches the model.
- Answer with citations: every response is grounded in the retrieved passages it came from.
Why hybrid retrieval
Pure semantic search is good at paraphrase and bad at exact terms: part numbers, flags, function names, the vocabulary technical documents are made of. Combining it with lexical matching and then re-ranking for relevance keeps both: the question can be asked in plain language, and the retrieval still lands on the passage that uses the precise term.
Grounding over fluency
Citation-grounded answering is the design constraint the whole pipeline serves. Narrowing the context to ranked, relevant passages and attaching citations to the output is what reduces hallucination. It gives the reader a way to check the answer rather than trust it.