Graph retrieval answers questions vector search cannot. It also costs more at every stage: more LLM calls to build the index, more work to update it, and a schema decision you cannot defer. GraphRAG vs RAG is therefore not really a quality comparison โ both work โ but a question of whether your corpus produces enough of the questions only a graph can answer to justify what the graph costs to build and keep.
Standard RAG, in the form Lewis et al. introduced in 2020, chunks your documents, embeds the chunks, and retrieves the ones closest to your query. GraphRAG extracts entities and the relationships between them into a knowledge graph, then retrieves by traversing that structure.
The short version:
Use standard RAG when answers live inside individual passages. "What is our refund policy?" is answered by the paragraph containing the refund policy.
Consider GraphRAG when answers require connecting facts that are never stated together. "Which of our suppliers depend on the same upstream manufacturer?" is not in any single passage. It has to be assembled.
The question that decides it
Ask whether a competent human could answer your typical question by reading one page of your corpus.
If yes, chunk-based retrieval will do, and a graph adds cost without adding capability.
If they would need to read several documents and hold the connections in their head, that is the multi-hop case, the shape that benchmarks like HotpotQA were built to test, and it is precisely where vector similarity struggles. Retrieving the top five chunks by similarity gets you five passages about your query. It does not get you the chain of relationships linking them.
What graphs cost
This is the part usually skipped, and the reason most projects should not start here.
Ingestion is dramatically more expensive. Building the graph means running entity and relationship extraction over every chunk. Microsoft's documented indexing pipeline makes this concrete: separate LLM passes for extracting entities and relationships, summarising them, generating community reports, and then summarising those reports. A further claims-extraction workflow exists but ships disabled by default. Standard RAG makes one embedding call per chunk. Cheaper models can do the extraction, but the gap is structural.
Updates are harder. Re-embedding a changed document is trivial: replace its vectors. Updating a graph means repairing entities and relationships that document contributed to, some shared with other documents. Implementations vary a lot here, so test it early if your corpus changes often.
There is a schema decision you cannot avoid. What counts as an entity? Which relationships matter? Get this wrong and you have an expensive index that answers the wrong questions. Chunking has parameters. Graph construction has design.
Debugging is different. When vector retrieval returns the wrong chunk, you inspect similarity scores. When graph retrieval fails, the extraction may have missed an entity, merged two names into one, or built an edge that does not exist.
Scale, and what it does not tell you
Star counts and the week's change, read from our daily snapshots on 14 August 2026:
| Project | Stars | Gained 8โ14 Aug |
|---|---|---|
| LightRAG | 38.9k | +209 |
| Microsoft GraphRAG | 35.5k | +169 |
| RAG-Anything | 22.9k | +92 |
What this cannot tell you, despite being the obvious thing to ask it: whether anyone is migrating. A star records a moment of discovery, and nobody un-stars a framework when they stop using it. A team that moved off one of these last quarter still shows in its star count and produces no signal here at all. Any comparison article inferring adoption from star charts is over-reading the instrument.
It will not tell you which of these is pulling ahead either, and six days is too short a window to try. The ordering has a confound built into it: RAG-Anything is built directly on LightRAG, so anyone installing it installs a graph engine too, and stars accrue to the thing you install rather than to its dependencies. The two also share a lab, which means a release arrives with an existing audience. Nothing in the decision below depends on which of them is growing faster โ the table is here to tell you these are all real, maintained projects at comparable scale, and no more than that.
What we have not measured, and it is the one this article is named after: the cost. The section above describes the shape of Microsoft's pipeline from its own documentation. More LLM passes per chunk points in an obvious direction, but shape gives you a direction and never a magnitude, and a direction is not enough to budget against. We have not priced either of them. Ingestion cost per document is the measurement that would settle this, and it is the next thing worth doing rather than something we can hand you today. We have also not benchmarked retrieval quality for any of these.
The implementations
Microsoft GraphRAG describes itself as "a data pipeline and transformation suite that is designed to extract meaningful, structured data from unstructured text using the power of LLMs." Microsoft Research introduced it on 13 February 2024, with the method paper following that April. It is thorough and heavy on ingestion.
LightRAG (published in Findings of EMNLP 2025; preprint at arXiv 2410.05779) takes the position that this is not an either/or choice at all. Its README describes a system that "adopts a dual-layer architecture to manage both knowledge graphs (KGs) and vector embeddings," which is the hybrid design most teams arrive at independently.
It also advertises the specific capability the cost section above flags as the hard one: incremental updating and selective deletion, with automatic knowledge-graph regeneration when a document is removed. If your corpus changes frequently, that is the feature to evaluate, and the claim to test against your own data rather than accept.
RAG-Anything extends LightRAG to documents that are not purely text: images, tables and equations alongside prose. What it adds is not better graph traversal โ it is ingestion. If your corpus is scanned reports, spreadsheets or papers full of equations, check whether document parsing rather than retrieval topology is what is actually blocking you. It is a common misdiagnosis, and much the cheaper of the two to fix.
Hybrid is usually the real answer
The leading implementations do not treat this as either/or, and neither should you. Graph systems in practice keep vector retrieval alongside: vector search for passages that directly answer the question, graph traversal for the relationships between them.
So the practical question is rarely "GraphRAG or RAG." It is "should I add a graph layer to my existing retrieval, and for which subset of queries?" Framed that way it becomes an incremental decision with a small failure cost rather than an architecture bet.
A sequence that works
- Build standard RAG first. You need the baseline anyway, and it will handle more of your queries than you expect.
- Collect the failures. Keep the questions it answers badly.
- Look at them. If they fail because a passage was missed, fix chunking, add hybrid keyword search, or add a reranker. If they fail because the answer required joining facts across documents, keep going.
- Try the cheaper multi-hop fixes first. Query decomposition โ break the question into sub-questions, retrieve for each, then compose โ and iterative retrieval both handle a large share of multi-hop cases with no index to build and no schema to design. Skipping straight from "vector search failed" to "build a knowledge graph" skips the options that cost a prompt change.
- Only then decide whether the volume of genuine multi-hop queries justifies the ingestion and maintenance cost.
Teams that jump to the last step usually build a graph for a corpus whose questions were answerable from single passages, or from a second retrieval round.
Star figures read 14 August 2026 from our own daily snapshots. Live rankings: RAG and vector search.