LlamaIndex Chunking & Vector Calculator
Estimate how many chunks a corpus produces, how much RAM those vectors occupy, and how many tokens one full index build has to embed.
Assumptions, so you can check the arithmetic: an average source document is treated as 2,000 tokens; vectors are float32, so raw bytes = chunks × dimensions × 4; the index-type selector applies a flat overhead multiplier (1.2× flat, 1.5× HNSW graph) to approximate structure held alongside the vectors; chunk overlap is ignored, so a splitter configured with overlap embeds proportionally more than the token figure shown. Real memory depends on your store's HNSW parameters, payload size and quantisation settings, and re-embedding the whole corpus costs that token figure again every time the embedding model changes.
Using these numbers
- Chunk size is the input that moves everything else: how documents become nodes explains the tradeoff behind the selector above.
- Ready to build the index these numbers describe? Follow the LlamaIndex quickstart, including persistence so you only pay the embedding cost once.
- Changing the dimension selector means a full reindex, which is failure mode eight in retrieval troubleshooting.