AI / Retrieval (RAG) 2026
VivoAssist
A retrieval-augmented assistant that answers questions strictly from PDF manuals, with page-level citations.
- Year
- 2026
- Status
- Prototype
- Platform
- Command line
- Scope
- Ingestion, Retrieval, Terminal chat
- 01PDF manualsLoaded page by page, OCR where needed
- 02ChunkingBig, mid and small, with page metadata
- 03EmbeddingsStored in a persistent vector index
- 04Manual scopeOne manual, held across follow-ups
- 05RetrievalOnly from the selected manual
- 06AnswerWith page citations, or “Not found”
Overview
VivoAssist answers questions about technical product manuals using only what the selected manual says. It is a working prototype with a terminal chat interface; a web interface is listed as future work.
Challenge
When several long manuals share one index, a general-purpose assistant will happily blend them — or fill gaps from outside knowledge. For manuals, a confident wrong answer is worse than no answer.
Direction
Scope retrieval to one manual at a time and keep that scope through follow-up questions. Chunk each manual at three levels so broad and precise questions both find context, cite page numbers, and answer “Not found in the manual.” when the manual doesn't say.
What was built
- Ingestion of multiple large PDF manuals, page by page, with OCR for scanned pages
- Hierarchical chunking at three levels, each chunk carrying file and page metadata
- Persistent vector store, rebuilt on demand when manuals change
- Manual selection with fuzzy matching for misspelled names
- Manual-scoped retrieval that persists across follow-up questions
- Answers with page-level citations, and a strict not-found response
- Debug mode showing chunk counts and retrieval scores
Technology
- Python
- LlamaIndex
- ChromaDB
- Azure OpenAI
- pypdf
- Tesseract OCR
Outcome
A prototype that keeps answers inside the selected manual and shows where each answer came from.
Technical notes
Described at a high level
The source manuals and any organisation-specific content are not reproduced here; only the retrieval architecture is described.