LlamaIndex Hub
Flat isometric illustration of two pink cube stacks flanking a glass case holding a glowing amber column, arrows pointing inward across a teal grid platform.
troubleshooting

LlamaIndex Retrieval Troubleshooting: Fix Bad Answers

Nine failure modes in LlamaIndex retrieval, how to tell them apart from source nodes and scores, and the documented fix for each one, in diagnosis order.

By LlamaIndex Hub Editorial · · 7 min read

Almost every “LlamaIndex is giving me wrong answers” problem is one of two completely different faults wearing the same costume. Either the retriever never returned the right passage, or it did and the synthesis step mishandled it. Those have opposite fixes, and guessing wastes days.

There is one diagnostic that separates them, and it should be the first thing you run.

Step zero: look at the source nodes

response = query_engine.query("your question")
for node in response.source_nodes:
    print(round(node.score, 3), node.node_id)
    print(node.node.get_content()[:300])
    print("---")

Read the output before changing anything, and ask one question: is the answer present in this text?

  • The right passage is not there. You have a retrieval problem. Everything in the next section applies.
  • The right passage is there but the answer is still wrong. You have a synthesis problem. Skip ahead.
  • The passage is there but cut in half. You have a chunking problem, which is a retrieval problem created at index time.

If standalone questions work but follow-ups fail, check query engine vs chat engine before rebuilding the index. In condense_question and condense_plus_context modes, inspect the query rewritten from conversation history; context uses the current message for retrieval, and simple does not query the index.

The symptom table

SymptomLikely causeFirst fix to try
Zero nodes returnedNo usable indexed content or filters excluded candidatesCheck loaded text, indexed nodes, and retrieval filters
Nodes returned but unrelated to the questionRetrieval selected the wrong evidenceInspect node content and compare it with the expected passage
Answer exists in a PDF but never retrievedScanned image, no text layerUse a parser that does OCR before indexing
Semantic questions work, exact IDs failVector-only retrievalAdd keyword or hybrid retrieval
Right chunk retrieved, answer still wrongToo many nodes packed into the promptLower similarity_top_k, add a reranker
Answer contradicts the retrieved textSynthesis or model choiceRun a faithfulness evaluator
Answers reflect an old version of a documentIndex never rebuiltRe-ingest, check persistence logic
Quality collapsed after a config changeEmbedding model swappedReindex from scratch
Filters appear to be ignoredMetadata missing on nodesAttach metadata at ingestion, not at query time
Every query is slowRe-embedding on each startPersist the index and load it

The rest of this piece walks the same list in the order the causes actually appear.

1. The documents never loaded

Start by checking what actually loaded. SimpleDirectoryReader picks a parser by extension and returns extracted content as documents. For scanned PDFs, check whether the selected parser extracted usable text. Diagnose zero retrieved nodes separately from irrelevant retrieved passages; inspect loading and filtering for the former, and the returned evidence for the latter.

Print the document count, then print the first few hundred characters of two or three documents. If the text is empty or mangled, stop: no amount of retrieval tuning fixes a corpus that was never read. Route those files through a parser that handles OCR or layout before indexing.

2. The chunk boundary destroyed the answer

Retrieval returns whole nodes. If the sentence that answers the question was split across two nodes, neither node scores well and neither would fully answer even if retrieved.

Symptoms are distinctive: retrieved passages are topically correct but truncated, or the top result stops one line before the useful part. Split along document structure such as headings, sections, and code blocks rather than a fixed character count, and give adjacent nodes some overlap so a boundary-spanning sentence survives.

Chunk size is a genuine tradeoff, not a value to maximise. Small nodes embed precisely and arrive without context. Large nodes carry context and produce diffuse embeddings that match many queries weakly. Change Settings.chunk_size, reindex, and measure, rather than reasoning about it.

3. Exact strings are invisible to embeddings

Semantic similarity is bad at tokens that carry no semantics: order numbers, error codes, SKUs, function names, version strings. The embedding of ERR_4471 sits near the embedding of every other error code.

If your failing queries are the ones containing an identifier, the fix is not a better embedding model. Combine keyword search with vector search and merge the candidate sets, or filter by metadata when the identifier is a field rather than free text. This is the single highest-yield change for technical corpora, and it is a retrieval change, not a prompt change.

4. Metadata was never attached, so filters do nothing

Metadata filters can only match fields that exist on the nodes. If source, section, document type, or date were not attached at ingestion, a filtered query silently matches nothing or, worse, matches everything.

Attach metadata during loading, then confirm it survived by printing node.metadata on a retrieved node. Filtering to the right subset of a corpus is frequently more effective than any embedding improvement, because it removes the wrong answers rather than trying to outrank them.

5. Too many nodes, buried evidence

When similarity_top_k is high, the correct passage is retrieved and then drowned. The synthesizer receives a long prompt in which the relevant paragraph is one of fifteen, and models attend unevenly across a long context.

The node postprocessor layer exists for exactly this stage. SimilarityPostprocessor drops nodes below a score cutoff so weak matches never reach the prompt. A reranker, whether SentenceTransformerRerank, LLMRerank, or CohereRerank, reorders the candidate set by relevance to the actual question rather than raw embedding proximity, which usually improves precision more than any tweak to the embedding stage. LongContextReorder rearranges the surviving nodes before packing. MetadataReplacementPostProcessor swaps a narrowly retrieved sentence for its surrounding window, which is the clean fix for “retrieved the right sentence, lost the context around it”.

A good default: retrieve wide, rerank hard, send few.

6. The answer contradicts the retrieved text

If the correct passage is in source_nodes and the answer still disagrees with it, investigate the synthesis step. Keep the retrieved evidence fixed when comparing model and prompt settings so you can tell whether a change improves the answer.

Then measure instead of arguing. LlamaIndex ships response evaluators for faithfulness, whether the answer reflects the retrieved context without hallucinating, plus context relevancy, answer relevancy, and correctness against a reference. Faithfulness is the one that distinguishes “the model invented this” from “the retrieved text really did say that”.

7. Stale answers

An index is a snapshot. If documents changed and nobody re-ingested, the pipeline confidently serves the old version, and nothing in the output signals it.

Check whether the persisted index is newer than the corpus. Then decide the refresh strategy deliberately: full reindex on a schedule, or incremental ingestion that skips unchanged documents. Rebuilding everything on every process start is the other failure of the same rule and shows up as uniformly slow startup rather than wrong answers.

8. Someone changed the embedding model

Vectors from different embedding models are not comparable. Swapping the model without reindexing leaves you comparing new query vectors against an index of old document vectors, and the result is retrieval that looks random rather than broken.

The tell is a sudden, across-the-board quality collapse right after a config change. Record the embedding model alongside every index, treat a model change as a full reindex, and check the dimension count matches what the vector store collection was created with. The chunking and vector store sizer shows how much the dimension choice changes index size, which is worth knowing before you commit to a model you cannot afford to rerun.

9. The tutorial you copied is out of date

An import error mentioning ServiceContext means the snippet predates version 0.10. ServiceContext was deprecated in 0.10 in favour of the global Settings object and removed in 0.11. Likewise, imports from a flat llama_index namespace belong to the pre-0.10 layout; current code imports from llama_index.core plus per-integration packages.

Stop guessing: build an eval set

Every fix above is a hypothesis, and you cannot tell whether one helped by asking a few questions by hand. Retrieval evaluation with hit rate and mean reciprocal rank answers “did the right node come back”, and the framework can synthetically generate question and context pairs from your own text, so building the set costs an afternoon rather than a labelling project.

Measure retrieval and generation separately. A single end-to-end score cannot tell you which half regressed, which is how teams end up tuning prompts to compensate for a chunking bug.

Sources

  1. LlamaIndex: Node postprocessor modules
  2. LlamaIndex: Evaluating
  3. LlamaIndex: Starter tutorial
  4. LlamaIndex: Migrating from ServiceContext to Settings
  5. LlamaIndex 0.11 release notes: ServiceContext removed
  6. LlamaIndex: Installation and package structure
  7. LlamaIndex: Chat engine usage pattern

Related