growthGrid

Applied AI

Retrieval-augmented generation is easy to demo and hard to trust. How we parse, chunk, retrieve and cite, so every answer can be checked in one click.

2 min read

Retrieval inspector showing the pipeline steps, scored passages on either side of a threshold and a cited answer
On this page

Every team that tries “chat with your PDFs” has the same first reaction: it feels like magic. The second reaction comes a week later, when someone asks where an answer came from and nobody can say.

We built DocuMind to answer questions about PDFs, Word files and slide decks. The model was the easy part. Trust came from everything around it.

Parse each format on its own terms

A PDF, a DOCX and a PPTX store text in very different ways. Push them all through one generic extractor and you lose exactly what makes an answer checkable: page numbers, slide titles, table rows.

  • PDF: extract page by page and keep the page number on every chunk.
  • DOCX: walk paragraphs and tables in order, so a table row stays one row.
  • PPTX: read slide titles, body text and speaker notes, tagged with the slide number.

Chunk with overlap, and keep the metadata

Chunks are the unit of retrieval, so their size decides what the model gets to see. Too small, and an answer is split across two chunks. Too large, and the one relevant sentence drowns. We split on paragraphs first, then by length, with an overlap so a sentence that straddles a boundary survives in both halves.

python
splitter = RecursiveCharacterTextSplitter(chunk_size=800, chunk_overlap=120)

chunks = [
    Document(page_content=text, metadata={"source": file.name, "page": page})
    for page, page_text in pages
    for text in splitter.split_text(page_text)
]

embeddings = HuggingFaceEmbeddings(model_name="sentence-transformers/all-mpnet-base-v2")
index = FAISS.from_documents(chunks, embeddings)
Every chunk carries its file and page. That’s what makes a citation possible later.

Retrieve first, and set a floor

Embeddings come from all-mpnet-base-v2, an open sentence-transformer model, and live in a FAISS index on the same machine as the app. Each question returns the top few passages with a similarity score, and anything below a threshold is dropped before the model ever sees it.

Make the citation part of the contract

The prompt hands the model numbered passages and one rule: answer only from these, and cite every passage you use. After generation, we check that each citation points to a passage we actually sent, and discard answers that cite nothing.

prompt
Answer using only the numbered passages below.
Cite every passage you use, like [2].
If the passages don't contain the answer, say you can't find it.
The entire instruction. Short prompts are easier to test and harder to misread.

Measure what people actually feel

  1. Build a small evaluation set: fifty real questions, each with the page that answers it.
  2. Track retrieval hit rate: did the right page appear in the top results?
  3. Track citation accuracy: does the cited passage really support the answer?
  4. Review refusals: was the assistant right not to answer?
If the model can’t point to a passage, it shouldn’t answer.
— growthGrid AI practice

None of this needs a frontier model. Mistral 7B, an open-weight model, writes perfectly good answers when retrieval does its job — and it can run on infrastructure you control, which matters more than any benchmark when the documents are contracts or HR files.

  • RAG
  • LLMs
  • Vector search
  • Python

Share

Insights

Put it into practice

Let’s turn fresh perspectives into practical solutions built around your product and business goals.

Booking new projects for Q4 2026

We reply to every enquiry within one business day.