Every team that tries “chat with your PDFs” has the same first reaction: it feels like magic. The second reaction comes a week later, when someone asks where an answer came from and nobody can say.
We built DocuMind to answer questions about PDFs, Word files and slide decks. The model was the easy part. Trust came from everything around it.
Parse each format on its own terms
A PDF, a DOCX and a PPTX store text in very different ways. Push them all through one generic extractor and you lose exactly what makes an answer checkable: page numbers, slide titles, table rows.
- PDF: extract page by page and keep the page number on every chunk.
- DOCX: walk paragraphs and tables in order, so a table row stays one row.
- PPTX: read slide titles, body text and speaker notes, tagged with the slide number.
Chunk with overlap, and keep the metadata
Chunks are the unit of retrieval, so their size decides what the model gets to see. Too small, and an answer is split across two chunks. Too large, and the one relevant sentence drowns. We split on paragraphs first, then by length, with an overlap so a sentence that straddles a boundary survives in both halves.
splitter = RecursiveCharacterTextSplitter(chunk_size=800, chunk_overlap=120)
chunks = [
Document(page_content=text, metadata={"source": file.name, "page": page})
for page, page_text in pages
for text in splitter.split_text(page_text)
]
embeddings = HuggingFaceEmbeddings(model_name="sentence-transformers/all-mpnet-base-v2")
index = FAISS.from_documents(chunks, embeddings)Retrieve first, and set a floor
Embeddings come from all-mpnet-base-v2, an open sentence-transformer model, and live in a FAISS index on the same machine as the app. Each question returns the top few passages with a similarity score, and anything below a threshold is dropped before the model ever sees it.
Make the citation part of the contract
The prompt hands the model numbered passages and one rule: answer only from these, and cite every passage you use. After generation, we check that each citation points to a passage we actually sent, and discard answers that cite nothing.
Answer using only the numbered passages below.
Cite every passage you use, like [2].
If the passages don't contain the answer, say you can't find it.Measure what people actually feel
- Build a small evaluation set: fifty real questions, each with the page that answers it.
- Track retrieval hit rate: did the right page appear in the top results?
- Track citation accuracy: does the cited passage really support the answer?
- Review refusals: was the assistant right not to answer?
If the model can’t point to a passage, it shouldn’t answer.
None of this needs a frontier model. Mistral 7B, an open-weight model, writes perfectly good answers when retrieval does its job — and it can run on infrastructure you control, which matters more than any benchmark when the documents are contracts or HR files.



