IP / READ / 2026
Similar is not correct
Improving retrieval is at least five different jobs. A method that fixes one of them can leave the other four exactly where they were.
Sooner or later, a conversation about an AI assistant for a business runs into a simple question. Where would it know our own material from?
Not what contracts usually say. What our contract says. Not how the equipment normally works, but whether this part fits this model. Not a convincing answer about internal process, but the answer the current instruction gives, ideally with a pointer to the right paragraph.
One answer to that question is RAG: a language model that is handed the relevant material before it answers. A short acronym over a lot of engineering, and over even more ways to improve it.
Our summer reading list turned into a map of those ways without anyone planning it. We added 136 papers to the starting bibliography, sifted out of an arXiv corpus covering May to early September, and sorted them by method, by the property of the system they claim to improve, and by how they check the result. Selection by abstract was enough to see the shape of the field. The claims we lean on were read in full text separately.
Instead of postcards from the sea we came back with a knowledge graph. Not a bad haul, although the tan is questionable.
The most useful result was not a list of new acronyms. It was that “improve the RAG” turns out to be several different jobs, and they are easy to mistake for one.
What this is, before anything else
Picture a colleague who writes well and knows a lot, but has never read your internal documents. You can ask them to answer from memory. Or you can hand them the right pages first.
The second option is roughly what RAG does. Retrieval-Augmented Generation: generating an answer from material that was found for the question. The usual shape is simple. Documents are cut into fragments, prepared for search, matched against a particular question, and passed to the model. There is a survey of the field if you want the long version.
The difficulty starts at the word “right”.
Take the question “can a product be returned after installation?” The search brings back an excellent paragraph about returns. The model summarises it accurately. The exception for installed goods stayed in the next section.
The model invented nothing. The answer is wrong anyway.
So the quality of such a system does not rest on how well it writes. It rests on what it found, what it lost, and whether what it found is enough to reach a conclusion.
Where symbolic AI comes in
Neural networks are good with the variety of human phrasing. “How do I book leave?” and “I want to take a week off” are different words and a close intent.
The symbolic approach adds things and relations stated explicitly: this document is an annex to that contract; this instruction applies to version 4; this rule has that exception.
The simplest case is a rule. If the order is paid and the item is in stock, it can ship. The system checks facts and applies a condition. The harder cases are knowledge graphs, logical inference, planning, and searching for a solution that satisfies a set of constraints.
Symbolic AI does not have to mean an enormous rule base that people fill in by hand for years. Part of the structure can come from a database you already have, part can be pulled out of documents by a model. An extracted fact can also be wrong, though. Structure does not turn a guess into a truth.
That combination is what GraphRAG, HippoRAG 2 and KAG are built on. Different mechanisms, one shared idea: similarity of language is a useful instrument and not the only way to reach the knowledge you need.
Not all of this is summer news. Understanding the new papers meant going back to the methods they grew out of.
Finding the exact thing, not only the similar thing
“Pump AX-417” and “pump AX-471” look very close. For the person ordering the part, they are two different purchases.
Search by meaning is what you need when people phrase one thought in many ways. Search by words and designations is what you need when the exact designation is the point. The two can be combined: one mechanism finds the rephrased question, the other keeps the part number, the surname or the error code from being lost.
There is interesting ground between those poles. SPLADE learns to expand the search representation with relevant terms. ColBERTv2 matches parts of the question against parts of the text rather than only their summed-up representations, and works on bringing the cost of such an index down. Different ways of making the match more precise.
For a business the question is short. Has the system stopped confusing the things that are, for us, fundamentally different?
The trade-off is short too. A strict filter on a code can remove false matches and discard the correct document with a typo in that code. Two searches then have to be merged into one list. The fact that a system has become “hybrid” does not by itself show that the list got better.
Not shredding the document into confetti
A document is not a long ribbon of words. It has headings, tables, footnotes, versions and references to other sections.
Cut it mechanically and you can separate an answer from the condition under which it is correct. A table row without its column header is sometimes worse than no row at all: it looks clear and it means something else.
So part of the work happens before the search. Keep the heading next to the fragment, read the table as a table, know the version of the document, be able to pull in the neighbouring paragraph.
Another level is reading at several scales. RAPTOR builds a tree of texts and summaries, so the system can look for detail or for the wider context. Rather like being able to open a specific page, or to check the table of contents first.
These approaches matter when a correct fact without its context becomes a wrong answer. A summary can lose the detail, though, and added context can bring noise. “Give the model more text” and “give it what it was missing” are different actions.
Joining facts that live in different places
Take the question “which customers are affected if this supplier stops?”
One document has suppliers and components. Another has components and products. A third has products and customers. No single paragraph holds the answer.
This is where a knowledge graph helps: a map of objects and the relations between them. You can walk from supplier to component, from component to product, and from there to the customer.
But “graph RAG” is not one method either. HippoRAG 2 uses graph structure to find related pieces of evidence. GraphRAG builds, among other things, summaries of groups of related entities for broad questions about a whole collection. Tracing a chain of dependencies and summarising the main themes of a large archive are not the same job.
The gain is access to relations that are hard to reach with a single search query. The price is building and maintaining the structure. Entities have to be recognised correctly, two different people with the same name must not be merged, dependencies have to be updated after changes.
A neatly drawn wrong arrow is still a wrong arrow.
Picking the instrument that fits the question
Not every question should be answered by reading excerpts.
“How many unpaid invoices in August?” is closer to a database query. “What changed in the returns policy?” is a comparison of documents. “Why do these terms apply to this contract?” is work with relations and rules.
Text-to-SQL translates a human question into a query over tables. Methods like RSL-SQL help determine which tables and columns are the right ones. That does not remove the need to check the meaning: a correctly executed query against the wrong field returns a tidy wrong number.
A different direction is choosing how many steps to take. Adaptive-RAG separates questions that need no retrieval, one retrieval, or several in sequence. For an internal fact the model does not know, an answer “from memory” does not become acceptable just because the question was short.
More steps sometimes assemble the evidence. Sometimes they only add time, cost and further opportunities to go the wrong way. Which makes one more ability useful in an assistant: knowing when to stop searching, and when the data for an answer is not there at all.
Surviving Monday
In a demo the knowledge base is usually ready. On Monday somebody changes a price list, adds an instruction and cancels the previous one.
That is a different job: how the search structure gets updated, what it costs, and how quickly new facts reach the answers. Incremental updating, where only the affected part is rebuilt, is one of the directions in LightRAG.
For our own reading the useful distinction was between “cheap to build” and “cheap to keep”. A paper can measure the first convincingly and say nothing about the second. And success on a small controlled base does not answer the question about daily changes in a large heterogeneous archive.
A complicated method does not become a bad method because it is hard to adopt. But its benefit has to be compared not only against the quality of the baseline search, and also against the cost of the system’s whole life.
The most useful part of the reading was the fine print
Our map gradually stopped being a list of “method to advantage”. Each advantage needed two more fields: how it was measured, and under what conditions it showed up.
“Found the right document” is not the same as “collected all the documents needed”. “The answer matches the source” is not the same as “the source is correct and current”. “Works well on this set of questions” is not the same as “will work well on ours”.
Even a number is easy to retell wrongly. Checking against full texts, we found comparisons of different metrics, results for one particular scale that had turned into a general rule, and a build cost that had been read as a cost of operation.
This does not mean the papers proved nothing. It means they proved something more specific than the short retelling suggested.
So we now keep three questions apart. Did we read the protocol in the primary source? Which part of the property we care about does it actually test? And how was the result established: by a number, by comparing groups, by switching a component off, or by a qualitative analysis?
A number is not automatically stronger than a qualitative assessment. What matters is whether the check supports the conclusion we are drawing from it.
So which RAG is best
Our selection gave no grounds for naming a winner. It gave something more useful: a map for choosing.
If the system confuses part numbers, the thing to test is the precision of matching designations. If it loses exceptions, structure and completeness of context. If the answer is spread across documents, the search over relations. If it answers correctly but too slowly and too expensively, the number of steps and the routing. If it answers from yesterday’s data, the update path.
Not every improvement is necessarily bought by degrading something else: you can remove wasted work and win on quality and speed at once. Behind a new method name, though, you still have to see what exactly it changes, which assumptions it adds, and what it did not test.
The practical next step for us is a measurement on our own corpus. Real questions, exact codes, exceptions, outdated versions, and the cases where the right answer is simply not in the base. On the same data, compare the baseline search against one specific improvement: which answers got better, which got worse, what happened to time, cost and updates.
That is how the research summer went. It started as a search for promising technology and ended with more precise questions.
Not “which RAG do we add next?”, but “which mistake is it supposed to remove, and how would we notice that it actually did?”