RAG knowledge retrieval method based on B + tree structure
Through the hierarchical semantic indexing and dynamic update mechanism based on B+ tree, the problem of time-consuming and high storage cost in high-dimensional vector indexing in RAG technology is solved, efficient and real-time knowledge retrieval and generation are realized, and semantic matching accuracy and system stability are improved.
Patent Information
- Application Number
- CN202510854780.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-25
AI Technical Summary
The construction of high-dimensional vector database indexes in the existing RAG technology is time-consuming and has high storage costs. Traditional inverted indexes are difficult to deal with semantic similarity retrieval and lack the ability to organize hierarchical knowledge.
Using a hierarchical semantic indexing mechanism based on B+ trees, the dimensionality reduction to low-dimensional space is adopted through PCA, and the B+ tree index is constructed to replace the high-dimensional vector similarity calculation, and combined with dynamic incremental updates and multi-level retrieval strategies, the search efficiency and storage performance are improved.
It significantly reduces the retrieval complexity and storage costs, supports real-time updates, improves semantic matching accuracy and generation quality, and ensures the continuity of knowledge fragments and system flexibility.
Smart Images

Figure CN120371839A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing and information retrieval, and particularly to a RAG knowledge retrieval method based on a B+ tree structure. Background Art
[0002] In the wave of the booming development of artificial intelligence, the RAG technology, with its innovative architecture that deeply couples a retrieval system with a generative model, has shown significant advantages in knowledge-intensive tasks. By dynamically retrieving an external knowledge base to supplement the generation context, RAG effectively alleviates the "hallucination" problem of traditional large models and achieves more fact-based outputs in scenarios such as question-answering systems and content creation. How to efficiently retrieve the external knowledge base has become a research hotspot.
[0003] Currently, the research on knowledge retrieval methods in RAG mainly falls into two categories:
[0004] Similarity retrieval: Similarity retrieval converts text into high-dimensional vectors, calculates the similarity scores between the query vector and all stored vectors, and returns the records with high scores. Common similarity calculation methods include: cosine similarity, Euclidean distance, Manhattan distance, etc. However, the construction of the vector database index is time-consuming, global index reconstruction may be triggered during incremental insertion, and the complexity of vector similarity calculation increases exponentially with the dimension, and the storage cost of high-dimensional vectors is high.
[0005] Full-text retrieval: The inverted index is the core data structure of full-text retrieval, which constructs the mapping relationship of "term → document list". When constructing the inverted index, the document is first tokenized and split into individual terms. Then, for each term, the document list containing the term and the position information of the term in the document are recorded. For example, for the document set {D1: "Helloworld", D2: "HelloPython"}, the inverted index may be {Hello: [D1, D2], world: [D1], Python: [D2]}. When querying, the documents containing the query term can be quickly located through the inverted index. However, the inverted index is difficult to handle semantic similarity retrieval and lacks hierarchical knowledge organization ability. Summary of the Invention
[0006] Aiming at the defects and deficiencies of the existing technology, the present invention provides a RAG knowledge retrieval method and system based on a B+ tree hierarchical semantic index, and its innovative design includes:
[0007] Low-dimensional semantic index construction mechanism: By fusing the TF-IDF statistical features with the semantic vectors of a pre-trained model (such as Sentence-BERT) and reducing the dimension to a low-dimensional space (cumulative variance contribution rate ≥ 85%) through PCA, the high-dimensional semantic retrieval is converted into an efficient range query of low-dimensional vectors;
[0008] Construct a B+ tree index with the low-dimensional vector as the key value, and utilize its ordered structure and range query ability to replace the traditional high-dimensional vector similarity calculation, significantly reducing the storage cost and retrieval complexity.
[0009] Dynamic incremental update design: When a new knowledge block is added, a low-dimensional key value is generated in real time and inserted into the B+ tree. The index structure is dynamically adjusted through a node splitting algorithm (such as splitting a leaf node into two nodes when it exceeds the capacity and promoting the intermediate key value to the parent node) to avoid global reconstruction;
[0010] Adjacent leaf nodes are merged regularly to optimize the index balance and ensure long-term retrieval efficiency.
[0011] Multi-level retrieval enhancement strategy: Query optimization: Weigh the corresponding dimensions of the keywords in the user query (the preset weight is greater than 1) to improve the semantic matching accuracy;
[0012] Two-stage screening: First, quickly locate the candidate set in the B+ tree through the Euclidean distance (≤ the preset threshold δ), and then finely screen the Top-k results through the cosine similarity (implemented based on the maximum heap);
[0013] Generation enhancement: Concatenate the retrieval results with the user's question to form a prompt, and input it into the large model to generate a streaming answer.
[0014] Knowledge chunking and feature fusion:
[0015] Overlapping chunking (preset length token) and syntactic boundary detection are adopted to ensure semantic continuity;
[0016] Fuse TF-IDF and deep semantic features, taking into account both local keywords and global context information.
[0017] The present invention specifically adopts the following technical solutions:
[0018] A RAG knowledge retrieval method based on the B+ tree structure:
[0019] Convert the user query into a low-dimensional vector obtained by dimensionality reduction through principal component analysis;
[0020] Perform a range query through a pre-constructed B+ tree index to retrieve similar knowledge fragments, where the B+ tree index uses the low-dimensional vector of the knowledge document as the key value;
[0021] Generate a final answer in combination with the retrieval results.
[0022] Furthermore, the construction of the B+ tree index includes:
[0023] Preprocess the knowledge document to generate a low-dimensional vector as the key value, and the preprocessing includes: document segmentation processing and feature extraction processing;
[0024] Construct a B+ tree to store the key values and associated pointers.
[0025] Furthermore, the document segmentation process includes:
[0026] Segment the document into fixed-length text blocks, with adjacent blocks overlapping by a preset length of tokens;
[0027] Adjust the block boundaries using syntactic boundary detection.
[0028] Furthermore, the document segmentation process includes:
[0029] Extract TF-IDF features;
[0030] Generate semantic vectors for the pre-trained model;
[0031] Fuse the TF-IDF features and semantic vectors, and reduce the dimension to a low dimension through principal component analysis, with the cumulative variance contribution rate not less than 85%.
[0032] Furthermore, during the process of performing range queries through the pre-constructed B+ tree index to retrieve similar knowledge fragments:
[0033] Perform weighted processing on the corresponding dimensions of the keywords in the query, with the preset weight greater than 1;
[0034] Screen the candidate set through the Euclidean distance, with the distance not exceeding the preset threshold;
[0035] Screen the Top-k results through cosine similarity.
[0036] Furthermore, it also includes a dynamic update step:
[0037] Preprocess the new knowledge block to generate a low-dimensional vector as the key value;
[0038] Insert the key value into the B+ tree index. If the leaf node capacity exceeds the threshold, trigger the split algorithm to adjust the tree structure.
[0039] And, a RAG knowledge retrieval system based on the B+ tree structure, including:
[0040] A query conversion module for converting a user query into a low-dimensional vector obtained by reducing the dimension through principal component analysis;
[0041] A retrieval module for performing range queries through the pre-constructed B+ tree index to retrieve similar knowledge fragments, where the B+ tree index uses the low-dimensional vectors of knowledge documents as key values;
[0042] A generation module for generating a final answer in combination with the retrieval results.
[0043] Furthermore, it also includes an index construction module for:
[0044] Preprocess the knowledge document to generate low-dimensional vectors as key values. The preprocessing includes: document segmentation processing and feature extraction processing;
[0045] Construct a B+ tree to store the key values and associated pointers.
[0046] And, a computer device, characterized in that it includes a processor and a memory; the memory stores a computer program, and when the processor executes the computer program, it implements the RAG knowledge retrieval method based on the B+ tree structure as described above.
[0047] A non-transitory computer-readable storage medium, characterized in that the storage medium stores a computer program, and when the computer program is executed by a processor, it implements the RAG knowledge retrieval method based on the B+ tree structure as described above.
[0048] Compared with the prior art, the present invention and its preferred solutions at least include the following beneficial effects:
[0049] Significantly improve retrieval efficiency and storage performance:
[0050] By constructing a retrieval mechanism based on the hierarchical semantic index of the B+ tree, convert the high-dimensional semantic similarity calculation into a range query in the low-dimensional space (complexity O(log n)), significantly reducing the retrieval time;
[0051] Combined with PCA dimensionality reduction (cumulative variance ≥ 85%) to significantly compress the vector dimension, reduce the storage overhead, and avoid the storage bottleneck of traditional high-dimensional vector databases.
[0052] Support dynamic incremental update and long-term stability:
[0053] Use the split algorithm of the B+ tree to realize the real-time insertion of new knowledge blocks, without global index reconstruction, ensuring the real-time update ability of the knowledge base;
[0054] Optimize the index structure by periodically merging low-capacity leaf nodes to maintain the stability of long-term retrieval performance.
[0055] Enhance semantic matching accuracy and generation quality:
[0056] Fuse the TF-IDF statistical features and the deep semantic vectors of the pre-trained model, taking into account the keyword weights and context relevance, and improving the accuracy of semantic understanding;
[0057] Adopt a two-stage screening mechanism (Euclidean distance fast initial screening + cosine similarity fine screening) to ensure the accuracy of the Top-k results;
[0058] Optimize the reliability and fluency of the content generated by the large model through the intelligent splicing of the retrieval results and the user questions.
[0059] Ensure knowledge continuity and system flexibility:
[0060] The overlapping block and syntactic boundary detection technology avoids semantic fragmentation and ensures the integrity of knowledge fragments;
[0061] The parameter-configurable design (such as text block length, distance threshold, etc.) adapts to the needs of multiple scenarios and enhances the universality of the system. Brief Description of the Drawings
[0062] The present invention will be further described in detail below in conjunction with the drawings and specific embodiments:
[0063] Figure 1 is the flowchart of the RAG knowledge retrieval implementation based on the B+ tree structure in the embodiment of the present invention, where the left figure is the process of scheme construction and the right figure is the process of retrieval implementation;
[0064] Figure 2 is the schematic diagram of the B+ tree index structure in the embodiment of the present invention. Specific Embodiments
[0065] In the following, specific embodiments of the present application will be described in detail with reference to the drawings. According to these detailed descriptions, those skilled in the art can clearly understand the present application and can implement the present application. Without departing from the principle of the present application, the features in different embodiments can be combined to obtain new implementation manners, or some features in certain embodiments can be replaced to obtain other preferred implementation manners.
[0066] To make the features and advantages of the present invention more obvious and understandable, specific embodiments are given below and described in detail in conjunction with the drawings as follows:
[0067] The embodiment of the present invention provides a RAG knowledge retrieval scheme based on the B+ tree structure, which is applicable to scenarios such as enterprise knowledge base management, intelligent question answering systems, and professional field document retrieval. Aiming at the problems of high high-dimensional storage cost of vector databases, time-consuming index reconstruction, and lack of hierarchical knowledge organization ability in existing RAG systems, the scheme includes: knowledge preprocessing and feature extraction; B+ tree index construction; retrieval enhanced generation; a dynamic update mechanism to reflect the latest knowledge in real time, improve the retrieval efficiency, and optimize the B+ tree structure regularly; by constructing a hierarchical semantic index and a dynamic update mechanism, efficient knowledge retrieval, dynamic update, and comprehensive optimization of storage costs are realized. As Figure 1 shown, the specific implementation process includes the following steps:
[0068] Step S1, knowledge preprocessing and feature extraction.
[0069] As a preferred solution of this embodiment, step S1 specifically includes the following steps:
[0070] Step S11: Split the original knowledge document into text blocks of a fixed length, with adjacent blocks overlapping by 100 tokens to ensure cross-block semantic continuity. Use syntactic boundary detection to adjust the block boundaries to avoid splitting a complete sentence into different blocks.
[0071] Step S12: Calculate TF-IDF features. The calculation formula is:
[0072]
[0073]
[0074] Among them, the meaning of TF is Term Frequency, and the meaning of IDF is Inverse Document Frequency. The subscript t represents the word, and d represents the document block; represents the product of the local frequency ( ) and the global distinctiveness ( ); is the number of occurrences of word t in document block d; is the total number of words in the block. The subscript k in this formula is the traversal subscript, indicating the summation over all word terms.
[0075]
[0076] Among them, N is the total number of documents, is the number of document blocks containing word t.
[0077] Step S13: Use models such as Sentence-BERT to generate 768-dimensional semantic vectors to capture context semantics.
[0078] Step S14: After concatenating the TF-IDF vector and the semantic vector, reduce the dimension to k dimensions through principal component analysis. The calculation formula is:
[0079]
[0080] Among them, represents the feature vector after PCA dimensionality reduction. PCA(, k) is a dimensionality reduction operator, indicating performing principal component analysis on the input data, specifically projecting the data onto the principal component direction by orthogonal transformation of the k eigenvectors with the largest eigenvalues, represents the TF-IDF vector of document block d, represents the semantic vector generated by a pre-trained model (such as Sentence-BERT), represents the concatenation operation.
[0081] It is required that the cumulative variance contribution rate ≥ 85% to ensure that the main semantic information is retained after dimensionality reduction.
[0082] Step S2, B+ tree index construction
[0083] As a preferred solution of this embodiment, step S2 specifically includes the following steps:
[0084] Step S21, define the B+ tree node structure, as Figure 2 shown: The internal nodes of the B+ tree contain pointers and key values, and the leaf nodes contain vector key values of knowledge document fragments, storage location pointers, and pointers to the next leaf node. The internal nodes are used for quick positioning of range queries, and the leaf nodes are used to store the links of actual knowledge fragment information and the links of adjacent leaf nodes for convenient sequential traversal.
[0085] Step S22, take the k-dimensional principal component vector after dimensionality reduction as the key value of the B+ tree.
[0086] Step S23, node splitting and merging: When the key values stored in the node exceed the threshold, the splitting algorithm of the B+ tree is used to maintain the balance of the tree.
[0087] Step S3, retrieval enhancement generation
[0088] As a preferred solution of this embodiment, step S3 specifically includes the following steps:
[0089] Step S31, convert the user's question into a 768-dimensional semantic vector through the same pre-trained model, and then obtain the k-dimensional query key value through PCA dimensionality reduction .
[0090] Step S32, weight the corresponding dimensions of the keywords in the query, and the calculation formula is:
[0091] ( = 1.5 when the word is a keyword)
[0092] Step S33, starting from the root node, select the closest child node to recursively traverse by comparing the query key value with the internal node key values until the leaf node. Collect all the leaf nodes with the Euclidean distance of the key values and to form the candidate set C.
[0093] Step S34, calculate the cosine similarity between each knowledge block vector b in the candidate set C and the original query vector q. The calculation formula is:
[0094]
[0095] Among them, CosineSimilarity is the standard cosine similarity function.
[0096] Use the maximum heap to filter out the Top-k results.
[0097] Step S35: Concatenate the retrieved knowledge block text with the user's question to form a prompt, and input the prompt into the large model to generate an answer, supporting streaming output to improve the interaction experience.
[0098] Step S4: Preprocess the newly added knowledge block to generate a k-dimensional principal component key value . Starting from the root node, find a suitable leaf node to insert the key value and storage pointer. If the leaf node capacity exceeds max_size, trigger the splitting algorithm to adjust the tree structure.
[0099] As a preferred solution of this embodiment, in the dynamic update mechanism:
[0100] When the number of key values in the leaf node exceeds the preset threshold, execute the B+ tree splitting algorithm:
[0101] Split the current node into two new nodes;
[0102] Extract the middle key value and promote it to the parent node;
[0103] Recursively check the capacity of the parent node until the root node is adjusted.
[0104] As a preferred solution of this embodiment, perform regular optimization of the B+ tree structure:
[0105] Periodically traverse the B+ tree, merge adjacent leaf nodes with the number of key values below the threshold, and reconstruct the index balance.
[0106] The following parameters are preset values:
[0107] Fixed length of the text block: L (unit: token)
[0108] Node capacity threshold: max_size
[0109] Euclidean distance threshold: δ
[0110] Number of Top-k results: k.
[0111] The technology provided in this embodiment has the following advantages and effects:
[0112] (1) Hierarchical semantic indexing: Through the ordered structure of the B+ tree and PCA dimensionality reduction, convert high-dimensional semantic retrieval into range queries in a low-dimensional space, improving the retrieval efficiency.
[0113] (2) Dynamic update optimization: Utilize the split and merge mechanisms of the B+ tree to support incremental insertion, eliminating the need to reconstruct the global index and enhancing real-time update performance.
[0114] Multi-dimensional retrieval enhancement: Improve the accuracy of semantic matching through the feature fusion of TF-IDF and pre-trained models.
[0115] Based on the same inventive concept, the present invention also provides a computer device, which includes: one or more processors, and a memory for storing one or more computer programs; the program includes program instructions, and the processor is configured to execute the program instructions stored in the memory. The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and is used to implement one or more instructions. Specifically, it is used to load and execute one or more instructions in the computer storage medium to implement the above method.
[0116] It should be further noted that, based on the same inventive concept, the present invention also provides a computer storage medium, on which a computer program is stored, and the computer program, when run by a processor, executes the above method. The storage medium may be any combination of one or more computer-readable media. The computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium may, for example, but not be limited to, an electrical, magnetic, optical, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a Random Access Memory (RAM), a Read-Only Memory (ROM), an Erasable Programmable Read-Only Memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or in combination with an instruction execution system, apparatus, or device.
[0117] In the description of this specification, the descriptions referring to the terms "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0118] The above shows and describes the basic principles, main features and advantages of the present disclosure. Those skilled in the art should understand that the present disclosure is not limited by the above embodiments. The above embodiments and the descriptions in the specification only illustrate the principles of the present disclosure. Without departing from the spirit and scope of the present disclosure, the present disclosure will have various changes and improvements, and these changes and improvements all fall within the scope of the present disclosure claimed.
[0119] The present invention is not limited to the above best implementation manner. Anyone can derive other various forms of RAG knowledge retrieval methods based on the B+ tree structure under the inspiration of the present invention. All equal changes and modifications made according to the scope of the patent application of the present invention shall fall within the coverage scope of the present invention.
Claims
1. A RAG knowledge retrieval method based on the B+ tree structure, characterized in that: Convert the user query into a low-dimensional vector obtained by dimensionality reduction through principal component analysis; Perform a range query through a pre-constructed B+ tree index to retrieve similar knowledge fragments, where the B+ tree index uses the low-dimensional vector of the knowledge document as the key value; Generate a final answer in combination with the retrieval results; The construction of the B+ tree index includes: Preprocess the knowledge document to generate a low-dimensional vector as the key value, and the preprocessing includes: document segmentation processing and feature extraction processing; Construct a B+ tree to store the key value and the associated pointer.
2. The RAG knowledge retrieval method based on the B+ tree structure according to claim 1, characterized in that: The document segmentation processing includes: Segment the document into fixed-length text blocks, and adjacent blocks overlap by a preset length of tokens; Adjust the block boundary by using syntactic boundary detection.
3. The RAG knowledge retrieval method based on the B+ tree structure according to claim 1, characterized in that: The document segmentation processing includes: Extract TF-IDF features; Generate semantic vectors of the pre-trained model; Fuse the TF-IDF features and the semantic vectors, and reduce the dimension to a low dimension through principal component analysis, and the cumulative variance contribution rate is not less than 85%.
4. The RAG knowledge retrieval method based on the B+ tree structure according to claim 1, characterized in that: In the process of performing a range query through the pre-constructed B+ tree index to retrieve similar knowledge fragments: Perform weighted processing on the corresponding dimensions of the keywords in the query, and the preset weight is greater than 1; Screen the candidate set through the Euclidean distance, and the distance does not exceed the preset threshold; Screen the Top-k results through the cosine similarity.
5. The RAG knowledge retrieval method based on the B+ tree structure according to claim 1, characterized in that: It also includes a dynamic update step: Preprocess the newly added knowledge block to generate a low-dimensional vector as the key value; Insert the key value into the B+ tree index. If the leaf node capacity exceeds the threshold, trigger a splitting algorithm to adjust the tree structure.
6. A RAG knowledge retrieval system based on the B+ tree structure, characterized in that, Includes: A query conversion module for converting the user query into a low-dimensional vector obtained by dimensionality reduction through principal component analysis; A retrieval module for performing a range query through a pre-constructed B+ tree index to retrieve similar knowledge fragments, where the B+ tree index uses the low-dimensional vector of the knowledge document as the key value; A generation module for generating a final answer in combination with the retrieval results; An index construction module for: Preprocess the knowledge document to generate a low-dimensional vector as the key value, and the preprocessing includes: document segmentation processing and feature extraction processing; Construct a B+ tree to store the key value and the associated pointer.
7. A computer device, characterized in that, Includes a processor and a memory; the memory stores a computer program, and when the processor executes the computer program, it implements the RAG knowledge retrieval method based on the B+ tree structure according to any one of claims 1-5.
8. A non-transitory computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by the processor, it implements the RAG knowledge retrieval method based on the B+ tree structure according to any one of claims 1-5.
Citation Information
Patent Citations
Retrieval enhancement method and device, electronic equipment and storage medium
CN118210908A
Hybrid retrieval method and system for RAG question-answering system
CN118656482A
Knowledge base storage retrieval system and method based on retrieval enhancement generation
CN118779429A
Large model retrieval enhancement generation method and device, equipment, storage medium and product
CN119848234A
Multi-mode-based data retrieval enhancement method
CN119961461A