Data reordering retrieval method and system based on RAG

By dynamic semantic chunking processing of the original document and building a hierarchical index, combined with the cross-modal attention mechanism of multi-source search and generative reordering model, the problem of low retrieval accuracy caused by a single search mode is solved, and efficient and accurate retrieval effect is achieved.

CN120086307AInactive Publication Date: 2025-06-03TAIJI COMPUTER CORPORATION LIMITED

Patent Information

Application Number
CN202510560897.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-06-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the search accuracy caused by a single search mode is low, especially in multimodal and cross-language scenarios, it is difficult to achieve high-precision search.

Method used

The original document is dynamic semantic chunking process to generate document blocks, the document blocks are stored in the vector database and a hierarchical index is constructed, the query request is parsed, and the Boolean keyword matching and semantic search is performed through the hierarchical index and vector database, the initial candidate set is generated by fusing the multi-source search results, and the candidate set is reordered using the cross-modal attention mechanism in the generative reordering model.

Benefits of technology

It improves search efficiency and accuracy, reduces unnecessary search operations, improves overall search speed, meets diversified application needs, and achieves comprehensive improvements in efficiency, accuracy, semantic understanding and system expansion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086307A_ABST
    Figure CN120086307A_ABST
Patent Text Reader

Abstract

The invention provides a data reordering retrieval method and system based on RAG, and the method comprises the steps: carrying out the dynamic semantic partitioning processing of an original document, generating corresponding document blocks, storing the document blocks in a vector database, constructing a hierarchical index, analyzing a received query request, extracting keywords in the query request, and carrying out the retrieval of the query request. Boolean keyword matching is carried out through an inverted index in the hierarchical index, semantic retrieval of keywords is carried out in a vector database, an initial candidate set is generated after multi-source retrieval results are fused, a query request and the initial candidate set are input into a generative reordering model, and a query result is obtained; according to the method, the initial candidate set is selected, the corresponding score is generated through the cross-modal attention mechanism in the generative reordering model, the initial candidate set is reordered according to the score, the final retrieval result is obtained, and the reordered document blocks are displayed, so that the overall retrieval speed is increased, the limitation of a single retrieval mode is avoided, and the accuracy of the retrieval result is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a data re-ranking retrieval method and system based on RAG. Background Art

[0002] With the rapid development of information technology, data retrieval systems play an increasingly important role in all walks of life. Traditional data retrieval methods mainly rely on keyword matching and rule sorting. Although they can meet the retrieval needs of users to a certain extent, when faced with multi-source heterogeneous data, complex semantic queries, and dynamically changing user needs, they often show problems such as insufficient retrieval accuracy, slow response speed, and poor flexibility. In recent years, the retrieval method based on Retrieval-Augmented Generation (RAG) has gradually become a hot topic. RAG is a technical architecture or method that combines retrieval technology with natural language generation (NLG). The core idea of RAG is that when generating natural language text, it not only depends on the knowledge and parameters pre-trained by the model itself, but also retrieves relevant information from external knowledge sources or data in real time according to the input content, and integrates the retrieved information into the generation process to improve the quality, accuracy, and relevance of the generated text.

[0003] However, traditional RAG usually relies on a single retrieval mode and is difficult to achieve high-precision retrieval in multi-modal and cross-language scenarios. For example, in cross-language retrieval, due to the semantic gap and language differences, the relevance of retrieval results is often low. Existing re-ranking methods mostly use rule sorting or simple machine learning models, which are difficult to capture the complex semantic relationships between queries and retrieval results. In addition, the re-ranking process usually consumes a large amount of computing resources, resulting in an extended system response time. Summary of the Invention

[0004] The present invention provides a data re-ranking retrieval method and system based on RAG to solve the problem of low retrieval accuracy caused by a single retrieval mode in the prior art.

[0005] The present invention provides a data re-ranking retrieval method based on RAG, including: Performing dynamic semantic chunking processing on the original document to generate corresponding document chunks; Storing the document chunks into a vector database and constructing a hierarchical index; Parsing the received query request and extracting the keywords in the query request, performing Boolean keyword matching through the inverted index in the hierarchical index, and performing semantic retrieval of the keywords in the vector database, and generating an initial candidate set after fusing multi-source retrieval results; Input the query request and the initial candidate set into the generative re-ranking model, and generate corresponding scores through the cross-modal attention mechanism in the generative re-ranking model; Re-rank the initial candidate set according to the scores, obtain the retrieval results, and display the re-ranked document blocks.

[0006] According to the data re-ranking retrieval method based on RAG provided by the present invention, the dynamic semantic chunking process of the original document to generate corresponding document blocks specifically includes: Preprocess the original document to unify the text format of the original document; Segment the original document through a sliding window algorithm; Perform logical unit segmentation on the segmented original document through an abstract syntax tree to generate the document blocks; Evaluate the length of the document blocks. If the length of the document block exceeds the first set threshold, split it. If the length of the document block is shorter than the second set threshold, merge it.

[0007] According to the data re-ranking retrieval method based on RAG provided by the present invention, the logical unit segmentation of the segmented original document through the abstract syntax tree to generate the document blocks specifically includes: Process the segmented original document through a parser to construct the abstract syntax tree; Identify relatively independent semantic logical units from the abstract syntax tree through predefined syntax rules; Extract corresponding text content from the original document through the identified logical units to generate the document blocks.

[0008] According to the data re-ranking retrieval method based on RAG provided by the present invention, the storing the document blocks into a vector database and constructing a hierarchical index specifically includes: Input the document blocks into a bidirectional encoder representation model to generate corresponding document block semantic vectors; Store the document block semantic vectors into a Milvus vector database; Construct a first-level summary index through the semantic summaries of the document blocks, construct a second-level vector index through the document block semantic vectors, and construct a third-level relationship index through the relationship network between the document blocks; Establish an association relationship between the first-level summary index, the second-level vector index, and the third-level relationship index, so that the first-level summary index, the second-level vector index, and the third-level relationship index can be switched and work collaboratively.

[0009] According to the RAG-based data re-ranking and retrieval method provided by the present invention, before storing the document blocks into the vector database and constructing a hierarchical index, it includes: Adding corresponding metadata tags to the document blocks, where the metadata tags include timestamps.

[0010] According to the RAG-based data re-ranking and retrieval method provided by the present invention, parsing the received query request and extracting the keywords in the query request, performing boolean keyword matching through the inverted index in the hierarchical index, and performing semantic retrieval of the keywords in the vector database, and generating an initial candidate set after fusing multi-source retrieval results, specifically including: Using a BERT tokenizer to tokenize and perform part-of-speech tagging on the query request and extract the keywords; Expanding the keywords through a pre-trained thesaurus, performing boolean logical combination on the expanded keywords, and calculating the relevance score of the document blocks through the BM25 algorithm; Converting the query request into a 768-dimensional semantic vector through a Sentence-BERT model, and performing vector similarity search in the HNSW graph index to obtain the cosine similarity; Performing weighted fusion on the relevance score and the cosine similarity to generate the initial candidate set.

[0011] According to the RAG-based data re-ranking and retrieval method provided by the present invention, inputting the query request and the initial candidate set into a generative re-ranking model and generating corresponding scores through the cross-modal attention mechanism in the generative re-ranking model, specifically including: Converting the query request into a query request vector, and converting the document blocks in the initial candidate set into document block vectors; Combining the query request vector and the document block vectors to form an input data pair, and performing formatting processing on the input data pair; Inputting the formatted input data pair into the generative re-ranking model; Fusing different modal features of the document blocks through the cross-modal attention mechanism to generate a comprehensive feature vector; The input data pair performs forward propagation in the generative re-ranking model, and the forward propagation outputs the corresponding scores based on the comprehensive feature vector, and the scores are used to represent the relevance between the document blocks in the initial candidate set and the query request.

[0012] According to the RAG-based data re-ranking and retrieval method provided by the present invention, fusing different modal features of the document blocks through the cross-modal attention mechanism to generate a comprehensive feature vector, specifically including: Extract different modality features of the document block through the generative re-ranking model, where the different modality features include text modality, image modality, and audio modality; Calculate the self-attention of the feature vectors of the text modality, the feature vectors of the image modality, and the feature vectors of the audio modality respectively to obtain corresponding self-attention results; Calculate the cross-attention between the feature vectors of the text modality, the feature vectors of the image modality, and the feature vectors of the audio modality respectively to obtain corresponding cross-attention results; Assign corresponding attention weights to the feature elements of the text modality, the feature elements of the image modality, and the feature elements of the audio modality according to the self-attention results and the cross-attention results; Perform weighted summation on the feature vectors of the text modality, the feature vectors of the image modality, and the feature vectors of the audio modality according to the attention weights to obtain the comprehensive feature vector.

[0013] According to the RAG-based data re-ranking and retrieval method provided by the present invention, before converting the query request into a query request vector and converting the document blocks in the initial candidate set into document block vectors, it includes: Select at least one of the document blocks from the initial candidate set as a candidate document block and generate a corresponding context representation for the candidate document block, where the context representation is a joint embedding vector generated by a transformer encoder; Obtain the document block corresponding to the query request by processing the context representation.

[0014] The present invention also provides a RAG-based data re-ranking and retrieval system for implementing the RAG-based data re-ranking and retrieval method. The RAG-based data re-ranking and retrieval system includes: An original document processing module for performing dynamic semantic chunking on an original document to generate corresponding document blocks; An index construction module connected to the original document processing module, where the index construction module is used to store the document blocks in a vector database and construct a hierarchical index; An initial candidate set generation module connected to the index construction module, where the initial candidate set generation module is used to parse a received query request, extract keywords in the query request, perform boolean keyword matching through an inverted index in the hierarchical index, and perform semantic retrieval of the keywords in the vector database, and generate an initial candidate set after fusing multi-source retrieval results; A generative re-ranking module, the initial candidate set generation module is connected to the generative re-ranking module, and the generative re-ranking module is used to input the query request and the initial candidate set into a generative re-ranking model and generate corresponding scores through the cross-modal attention mechanism in the generative re-ranking model; A scoring module, the scoring module is connected to the generative re-ranking module, and the scoring module is used to re-rank the initial candidate set according to the scores, obtain the retrieval result and display the re-ranked document blocks.

[0015] The present invention provides a data re-ranking retrieval method and system based on RAG. By performing dynamic semantic chunking on the original document to generate corresponding document chunks, storing the document chunks in a vector database and constructing a hierarchical index, parsing the received query request and extracting the keywords in the query request, performing boolean keyword matching through the inverted index in the hierarchical index, and performing semantic retrieval of the keywords in the vector database, fusing multi-source retrieval results to generate an initial candidate set, inputting the query request and the initial candidate set into a generative re-ranking model, and generating corresponding scores through the cross-modal attention mechanism in the generative re-ranking model, re-ranking the initial candidate set according to the scores, obtaining the final retrieval result and displaying the re-ranked document blocks. The present invention generates document chunks through dynamic semantic chunking, constructs a hierarchical index and combines multi-source retrieval to quickly generate an initial candidate set, improves the retrieval efficiency, and uses the cross-modal attention mechanism in the generative re-ranking model to improve the retrieval accuracy and flexibility, greatly reducing unnecessary retrieval operations, improving the overall retrieval speed, giving full play to the advantages of different retrieval methods, avoiding the limitations of a single retrieval method, improving the accuracy of the final retrieval result, and meeting diverse application requirements, achieving an overall improvement in efficiency, accuracy, semantic understanding and system scalability. Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a flowchart of the data re-ranking retrieval method based on RAG provided by the embodiment of the present invention; Figure 2 It is a schematic structural diagram of the data re-ranking retrieval system based on RAG provided by the embodiment of the present invention. Detailed Embodiments

[0018] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts belong to the scope of protection of the present invention.

[0019] The present invention provides a data re - ranking retrieval method based on RAG, including: Performing dynamic semantic chunking on the original document to generate corresponding document chunks; Storing the document chunks in a vector database and constructing a hierarchical index; Parsing the received query request and extracting keywords in the query request, performing Boolean keyword matching through the inverted index in the hierarchical index, and performing semantic retrieval of the keywords in the vector database, and generating an initial candidate set after fusing multi - source retrieval results; Inputting the query request and the initial candidate set into a generative re - ranking model and generating corresponding scores through the cross - modal attention mechanism in the generative re - ranking model; Re - ranking the initial candidate set according to the scores, obtaining the final retrieval result, and displaying the re - ranked document chunks.

[0020] Among them, the performing dynamic semantic chunking on the original document to generate corresponding document chunks specifically includes: Pre - processing the original document to unify the text format of the original document; Dividing the original document through a sliding window algorithm; Performing logical unit division on the divided original document through an abstract syntax tree to generate the document chunks; Evaluating the length of the document chunks, splitting if the length of the document chunk exceeds a first set threshold, and merging if the length of the document chunk is shorter than a second set threshold.

[0021] Among them, the performing logical unit division on the divided original document through an abstract syntax tree to generate the document chunks specifically includes: Processing the divided original document through a parser to construct the abstract syntax tree; Identifying relatively independent semantic logical units from the abstract syntax tree through predefined syntax rules; Extracting corresponding text content from the original document through the identified logical units to generate the document chunks.

[0022] Among them, storing the document blocks into the vector database and constructing a hierarchical index specifically includes: Inputting the document blocks into a bidirectional encoder representation model to generate corresponding document block semantic vectors; Storing the document block semantic vectors into the Milvus vector database; Constructing a first-level summary index through the semantic summary of the document blocks, constructing a second-level vector index through the document block semantic vectors, and constructing a third-level relationship index through the relationship network between the document blocks; Establishing an association relationship between the first-level summary index, the second-level vector index, and the third-level relationship index, enabling switching and collaborative work among the first-level summary index, the second-level vector index, and the third-level relationship index.

[0023] Among them, before storing the document blocks into the vector database and constructing a hierarchical index: Adding corresponding metadata tags to the document blocks, where the metadata tags include timestamps.

[0024] Among them, parsing the received query request and extracting the keywords in the query request, performing Boolean keyword matching through the inverted index in the hierarchical index, and performing semantic retrieval of the keywords in the vector database, and generating an initial candidate set after fusing multi-source retrieval results specifically includes: Using the BERT tokenizer to tokenize and perform part-of-speech tagging on the query request and extract the keywords; Expanding the keywords through a pre-trained thesaurus, performing Boolean logical combination on the expanded keywords, and calculating the relevance scores of the document blocks through the BM25 algorithm; Converting the query request into a 768-dimensional semantic vector through the Sentence-BERT model, and performing vector relevance search in the HNSW graph index to obtain the cosine similarity; Performing weighted fusion on the relevance scores and the cosine similarity to generate the initial candidate set.

[0025] Among them, inputting the query request and the initial candidate set into a generative re-ranking model and generating corresponding scores through the cross-modal attention mechanism in the generative re-ranking model specifically includes: Converting the query request into a query request vector, and converting the document blocks in the initial candidate set into document block vectors; Combining the query request vector and the document block vectors to form an input data pair, and performing formatting processing on the input data pair; Inputting the formatted input data pair into the generative re-ranking model; Fuse the different modality features of the document block through the cross-modal attention mechanism to generate a comprehensive feature vector; The input data pair performs forward propagation in the generative re-ranking model, and the forward propagation outputs the corresponding score based on the comprehensive feature vector, and the score is used to represent the relevance between the document block in the initial candidate set and the query request.

[0026] Among them, the fusing the different modality features of the document block through the cross-modal attention mechanism to generate a comprehensive feature vector specifically includes: Extract the different modality features of the document block through the generative re-ranking model, and the different modality features include text modality, image modality and audio modality; Calculate the self-attention of the feature vectors of the text modality, the image modality and the audio modality respectively to obtain the corresponding self-attention results; Calculate the interactive attention between the feature vectors of the text modality, the image modality and the audio modality respectively to obtain the corresponding interactive attention results; Allocate corresponding attention weights to the feature elements of the text modality, the image modality and the audio modality according to the self-attention results and the interactive attention results; Perform weighted summation on the feature vectors of the text modality, the image modality and the audio modality according to the attention weights to obtain the comprehensive feature vector.

[0027] Among them, before converting the query request into a query request vector and converting the document block in the initial candidate set into a document block vector, it includes: Select at least one document block from the initial candidate set as a candidate document block and generate a context representation corresponding to the candidate document block, and the context representation is a joint embedding vector generated by a transformer encoder; Obtain the document block corresponding to the query request by processing the context representation.

[0028] The present invention also provides a RAG-based data re-ranking retrieval system for implementing the RAG-based data re-ranking retrieval method, and the RAG-based data re-ranking retrieval system includes: An original document processing module, which is used to perform dynamic semantic chunking processing on the original document to generate corresponding document blocks; Build an indexing module, which is connected to the original document processing module. The indexing module is used to store the document chunks into a vector database and build a hierarchical index; An initial candidate set generation module, which is connected to the indexing module. The initial candidate set generation module is used to parse the received query request and extract the keywords in the query request, perform Boolean keyword matching through the inverted index in the hierarchical index, and perform semantic retrieval of the keywords in the vector database, and generate an initial candidate set after fusing multi-source retrieval results; A generative re-ranking module, the initial candidate set generation module is connected to the generative re-ranking module. The generative re-ranking module is used to input the query request and the initial candidate set into a generative re-ranking model and generate corresponding scores through the cross-modal attention mechanism in the generative re-ranking model; A scoring module, the scoring module is connected to the generative re-ranking module. The scoring module is used to re-rank the initial candidate set according to the scores, obtain the final retrieval result and display the re-ranked document chunks.

[0029] The present invention provides a data re-ranking retrieval method and system based on RAG. By performing dynamic semantic chunking on the original document to generate corresponding document chunks, storing the document chunks into a vector database and building a hierarchical index, parsing the received query request and extracting the keywords in the query request, performing Boolean keyword matching through the inverted index in the hierarchical index, and performing semantic retrieval of the keywords in the vector database, generating an initial candidate set after fusing multi-source retrieval results, inputting the query request and the initial candidate set into a generative re-ranking model and generating corresponding scores through the cross-modal attention mechanism in the generative re-ranking model, re-ranking the initial candidate set according to the scores, obtaining the final retrieval result and displaying the re-ranked document chunks. The present invention generates document chunks through dynamic semantic chunking, constructs a hierarchical index and combines multi-source retrieval to quickly generate an initial candidate set, improves the retrieval efficiency, and uses the cross-modal attention mechanism in the generative re-ranking model to improve the retrieval accuracy and flexibility, greatly reducing unnecessary retrieval operations, improving the overall retrieval speed, giving full play to the advantages of different retrieval methods, avoiding the limitations of a single retrieval method, improving the accuracy of the final retrieval result, and meeting diverse application requirements, achieving a comprehensive improvement in terms of efficiency, accuracy, semantic understanding and system scalability.

[0030] In this embodiment, please refer to Figure 1 and Figure 2, which are respectively the flowchart of the RAG-based data re-ranking retrieval method provided by the embodiments of the present invention and the schematic structural diagram of the RAG-based data re-ranking retrieval system provided by the embodiments of the present invention.

[0031] Specifically, in this embodiment, please refer to Figure 1 , the embodiments of the present invention provide a RAG-based data re-ranking retrieval method, including: S1: Perform dynamic semantic chunking on the original document to generate corresponding document chunks; S2: Store the document chunks in a vector database and construct a hierarchical index; S3: Parse the received query request and extract the keywords in the query request, perform boolean keyword matching through the inverted index in the hierarchical index, and perform semantic retrieval of the keywords in the vector database, and generate an initial candidate set after fusing multi-source retrieval results; S4: Input the query request and the initial candidate set into a generative re-ranking model and generate corresponding scores through the cross-modal attention mechanism in the generative re-ranking model; S5: Re-rank the initial candidate set according to the scores, obtain the final retrieval result and display the re-ranked document chunks.

[0032] In this embodiment, S1 specifically includes: (1) Preprocess the original document to unify the text format of the original document, unify all letters in the original document into uppercase or lowercase forms, remove unnecessary spaces and special characters, perform half-width conversion on full-width characters, etc., so as to ensure the consistency of the original document in format and facilitate subsequent processing; (2) Split the original document through a sliding window algorithm. Among them, the sliding window algorithm will set a window with a fixed size, start from the starting position of the unified original document, and slide the window step by step according to a predetermined step length. At each slide, the text segment covered by the window is regarded as a segmentation unit. In this way, the original document is initially divided into multiple smaller text segments to prepare for further fine processing; (3)The original document after segmentation is subjected to logical unit segmentation through an abstract syntax tree to generate the document blocks, specifically including: processing the segmented original document through a parser to construct the abstract syntax tree, then identifying logical units with relatively independent semantics from the abstract syntax tree through predefined syntax rules, and then extracting corresponding text content from the original document through the identified logical units to generate the document blocks. Specifically, first, the segmented original document is processed through a parser to construct an abstract syntax tree. The parser will parse the original document into a tree structure according to corresponding syntax rules, and each node in the tree represents different syntax elements. After constructing the abstract syntax tree, according to the predefined syntax rules, logical units with relatively independent semantics are identified from this tree structure. The logical units can express complete meanings. Finally, according to the identified logical units, corresponding text content is extracted from the original document to generate document blocks; (4)The lengths of the document blocks are evaluated. If the length of a document block exceeds a first set threshold, it is split. If the length of a document block is shorter than a second set threshold, it is merged. In this embodiment, a first set threshold and a second set threshold are preset. If the length of a certain document block exceeds the first set threshold, it means that the document block may contain too much information and the semantics are not focused enough. At this time, it will be split to make its length more in line with the requirements of subsequent processing to improve the retrieval efficiency and accuracy. On the contrary, if the length of a document block is shorter than the second set threshold, it indicates that it may contain too little information and the semantics are incomplete. Then it will be merged with adjacent document blocks to make the generated document blocks reach a more reasonable state in terms of length and semantic integrity.

[0033] Among them, S2 specifically includes: (1)The document blocks are input into a bidirectional encoder representation model to generate corresponding document block semantic vectors. The bidirectional encoder representation model (BERT) can encode the text of the input document blocks from both the forward and backward directions. When processing document blocks, it will perform in-depth semantic understanding on each vocabulary and its context in the document blocks. Through neural network calculation, the document blocks are transformed into corresponding document block semantic vectors. The semantic information specifically corresponding to the document block semantic vectors, each dimension represents the quantitative performance of the document block in a certain semantic feature, providing a numerical representation convenient for subsequent storage and retrieval; (2)The document block semantic vectors are stored in the Milvus vector database. The Milvus vector database has efficient vector indexing and query functions. Storing the document block semantic vectors in the Milvus vector database can ensure that in subsequent large-scale data retrieval scenarios, the target vectors can still be quickly and accurately located, greatly improving the efficiency of data storage and reading; (3) Construct a first-level abstract index through the semantic abstract of the document block. The semantic abstract is a generalization of the content of the document block. Starting from the semantic abstract to construct the index can quickly locate the set of potentially relevant document blocks during the preliminary screening. Construct a second-level vector index through the semantic vectors of the document blocks. Since the semantic vectors comprehensively and accurately reflect the semantic features of the document blocks, the second-level vector index constructed based on the semantic vectors can achieve more accurate semantic retrieval and can find the target that best matches the query request semantically from a large number of document blocks. Construct a third-level relationship index through the relationship network between the document blocks. The document blocks do not exist in isolation, and there are various logical associations between the document blocks, such as: topic relevance, citation relationships. By mining and sorting out the relationships existing between the document blocks, constructing a relationship network and forming an index, it is possible to not only consider the semantic information of the document blocks themselves during the retrieval process, but also utilize the association relationships between the document blocks to further expand the retrieval scope and improve the comprehensiveness and accuracy of the retrieval; (4) Establish an association relationship between the first-level abstract index, the second-level vector index, and the third-level relationship index, so that the first-level abstract index, the second-level vector index, and the third-level relationship index can be switched and work collaboratively. Among them, by establishing the association, the first-level abstract index, the second-level vector index, and the third-level relationship index are not independent individuals, but can cooperate with each other. During the actual retrieval process, according to the query characteristics and needs of the user, it is possible to flexibly switch between the first-level abstract index, the second-level vector index, and the third-level relationship index. For example: in the initial stage, first use the first-level abstract index for quick screening to narrow the retrieval scope; then, based on the screening results, switch to the second-level vector index for more accurate semantic matching; further expand the retrieval results, and with the help of the third-level relationship index, mine other document blocks associated with the current retrieval results. At the same time, the first-level abstract index, the second-level vector index, and the third-level relationship index can also work collaboratively to analyze and process the retrieval request from multiple dimensions, thereby providing a more efficient retrieval.

[0034] In this embodiment, before storing the document block into the vector database and constructing the hierarchical index, it includes: Add corresponding metadata tags to the document block. The metadata tags include timestamps. Specifically, the metadata tags can provide rich and diverse dimensions for subsequent retrieval. As part of the metadata tags, the timestamp records the precise time information when the document block is generated or modified. It records the position of each document block on the timeline in a standardized time format, accurate to seconds or even milliseconds. In addition to the timestamp, other types of metadata tags can be added according to actual needs, such as the document block source information tag for indicating which original document or data source the document block is derived from; the theme classification tag for indicating the core theme area involved in the document block, whether it is in the categories of technology, culture, economy, etc. By comprehensively adding these metadata tags, the management efficiency of the document block and the flexibility and accuracy of retrieval can be greatly improved during subsequent storage, index construction, and retrieval processes.

[0035] In this embodiment, S3 specifically includes: (1) Use the BERT tokenizer to tokenize, perform part-of-speech tagging on the query request, and extract the keywords. When a query request is received, first process it through the BERT tokenizer. BERT stands for Bidirectional Encoder Representations from Transformers, and its Chinese name is Bidirectional Encoder Representation Method. The query request is split into individual tokens by the BERT tokenizer, and part-of-speech tagging is performed on each token. For example, distinguish nouns, verbs, and adjectives. After tokenization and part-of-speech tagging, extract the keywords in the query request according to the set rules. The rules can be based on part of speech, such as preferentially selecting content words like nouns and verbs, or can be based on factors such as word frequency and the importance of the word in the text. For example, for the query request "Find recent popular articles on artificial intelligence technology", the BERT tokenizer will tokenize it into "Find", "recent", "popular", "of", "artificial", "intelligence", "technology", "article", and then extract keywords such as "artificial intelligence", "technology", "article" according to the rules; (2)Expand the keywords through a pre-trained thesaurus, perform Boolean logical combination on the expanded keywords, and calculate the relevance score of the document chunks through the BM25 algorithm. The full name of the BM25 algorithm is Best Matching 25 algorithm, and its Chinese name is the Best Matching 25 algorithm. After obtaining the keywords, the pre-trained thesaurus will be used to expand these keywords. The pre-trained thesaurus contains synonyms and near-synonyms information. By querying the pre-trained thesaurus, words with similar semantics corresponding to each keyword can be found. For example, for the keyword "artificial intelligence", related words such as "AI" and "machine learning" will be expanded. This not only expands the retrieval scope but also avoids missing relevant documents. Perform Boolean logical combination on the expanded keywords. For example, use logical operators such as "AND", "OR", and "NOT" to connect the keywords to form a retrieval expression. For example: for the keywords "artificial intelligence", "technology", and "article", combine them into an expression like "(artificial intelligence OR AI) AND technology AND article". Then use the BM25 (Best Matching 25) algorithm to calculate the relevance score of each document chunk with the query request according to the retrieval expression after Boolean logical combination. The BM25 algorithm will calculate a quantitative relevance score based on factors such as the frequency of occurrence of the keywords in the document chunk and the length of the document chunk. The higher the relevance score, the stronger the relevance between the document chunk and the query request; (3)Convert the query request into a 768-dimensional semantic vector through the Sentence-BERT model, and perform vector relevance search in the HNSW graph index to obtain the cosine similarity. The full name of the Sentence-BERT model is Sentence-Bidirectional Encoder Representations from Transformers, and its Chinese name is the Sentence-Bidirectional Encoder Representation Technology Model. It can map the text sentence of the query request into a 768-dimensional vector space, so that each dimension of the vector contains the semantic information of the query request. The full name of the HNSW graph index is Hierarchical Navigable Small World graph index, and its Chinese name is the Hierarchical Navigable Small World Graph Index. Perform vector relevance search in the HNSW graph index. Through the HNSW graph index, the vector most similar to the query request vector can be quickly found among a large number of document chunk vectors. By calculating the cosine similarity between the query request vector and the document chunk vector, it reflects the semantic similarity between the query request and the document chunk. The closer the cosine similarity is to 1, the higher the similarity; (4) Weightedly fuse the relevance score and the cosine similarity to generate the initial candidate set. Herein, weighted fusion is to comprehensively consider the advantages of the two retrieval methods. According to different application scenarios and requirements, different weights are assigned to the relevance score and the cosine similarity. For example, the weight of the relevance score can be set to 0.6, and the weight of the cosine similarity can be set to 0.4. The two are weighted and summed according to the weights to obtain a comprehensive score. Then, all document chunks are sorted according to the comprehensive score, and a part of the document chunks with higher scores are selected as the initial candidate set for subsequent further processing and reordering.

[0036] In this embodiment, S4 specifically includes: (1) Convert the query request into a query request vector, and convert the document chunks in the initial candidate set into document chunk vectors. Herein, the query request is input into the BERT model, and the BERT model will encode the query request and output a vector with a fixed dimension. This vector is the query request vector, and the query request vector has the semantic features of the query request. Each document chunk in the initial candidate set is processed by the BERT model, and the document chunks are vectorized one by one to obtain the corresponding document chunk vectors. These document chunk vectors are in the same vector space as the query request vector, facilitating subsequent operations such as similarity calculation. (2) Combine the query request vector and the document chunk vectors to form input data pairs, and perform formatting processing on the input data pairs. Herein, each input data pair represents the corresponding relationship between the query request and a document chunk, providing a clear data structure for subsequent processing. To ensure that the input data pairs meet the input requirements of the generative reordering model, formatting processing is performed so that the generative reordering model can correctly identify and process the input data pairs. (3) Input the formatted input data pairs into the generative reordering model, and the generative reordering model can perform in-depth feature extraction and analysis on the input data pairs. (4) Through the cross-modal attention mechanism, fuse different modal features of the document chunks to generate a comprehensive feature vector. Herein, the comprehensive feature vector fuses important information of different modalities of the document chunks and comprehensively represents the features of the document chunks. (5) The input data pair undergoes forward propagation in the generative re-ranking model. The forward propagation outputs the corresponding score based on the comprehensive feature vector. The score is used to represent the relevance between the document block in the initial candidate set and the query request. During the forward propagation process, the generative re-ranking model will, according to the information of the comprehensive feature vector and the query request vector, perform a linear transformation and non-linear processing and output a score. The score is used to represent the relevance between the document block in the initial candidate set and the query request. The higher the score, the closer the semantics between the document block and the query request, and the stronger the relevance. The lower the score, the weaker the relevance. Through this score, the document blocks in the initial candidate set can be re-ranked to provide a retrieval result that better meets the requirements.

[0037] Among them, the cross-modal attention mechanism fuses different modal features of the document block to generate a comprehensive feature vector, specifically including: (1) The generative re-ranking model extracts different modal features of the document block. The different modal features include text modality, image modality, and audio modality. For the text modality, semantic features are extracted through a pre-trained language model. For the image modality, visual features are extracted through a convolutional neural network (CNN). For the audio modality, audio features are extracted through a recurrent neural network (RNN); (2) Calculate the self-attention of the feature vectors of the text modality, the image modality, and the audio modality respectively to obtain the corresponding self-attention results. Among them, self-attention allows the generative re-ranking model to process the relationships between the internal elements of different modal features. For example: in the text modality, the generative re-ranking model identifies key words or phrases through self-attention, and in the image modality, it can highlight important visual regions; (3) Calculate the cross-attention between the feature vectors of the text modality, the image modality, and the audio modality respectively to obtain the corresponding cross-attention results. Among them, the data of different modal features may have a complementary and corroborative relationship. The cross-attention mechanism can dynamically determine the relevance between different modal features. For example: in a document block containing text descriptions and relevant images, the generative re-ranking model determines which parts of the text are most relevant to which features of the image through cross-attention; (4) Assign corresponding attention weights to the feature elements of the text modality, the image modality, and the audio modality according to the self-attention results and the cross-attention results; (5) Weighted sum of the feature vectors of the text modality, the feature vectors of the image modality, and the feature vectors of the audio modality according to the attention weights to obtain the comprehensive feature vector, where the comprehensive feature vector integrates the important information of different modality features of the document block and more comprehensively represents the features of the document block.

[0038] Wherein, before converting the query request into a query request vector and converting the document blocks in the initial candidate set into document block vectors, it includes: Select at least one of the document blocks from the initial candidate set as a candidate document block and generate a context representation corresponding to the candidate document block, where the context representation is a joint embedding vector generated by the candidate document block through a transformer encoder. That is, select at least one of the document blocks from the initial candidate set as a candidate document block, and the candidate document block generates a joint embedding vector as the context representation through a transformer encoder for subsequent processing; Obtain the document block corresponding to the query request by processing the context representation. That is, analyze the context representation and match the analyzed context representation with the query request. According to the matching result, screen out the document blocks in the candidate document blocks that are most relevant to the query request. These document blocks are the document blocks corresponding to the query request obtained after processing. Before converting the query request and the document blocks into vectors, the document blocks are screened and the context information is extracted and processed, improving the accuracy and efficiency of subsequent vector calculation and reordering, thereby providing a retrieval result that better meets the requirements.

[0039] The present invention also provides a RAG-based data reordering retrieval system 1 for implementing the RAG-based data reordering retrieval method described above. Please refer to Figure 2, the RAG-based data re-ranking retrieval system 1 includes: an original document processing module 10, which is used to perform dynamic semantic chunking on the original document to generate corresponding document chunks; a building index module 11, which is connected to the original document processing module 10, and the building index module 11 is used to store the document chunks in a vector database and build a hierarchical index; an initial candidate set generation module 12, which is connected to the building index module 11, and the initial candidate set generation module 12 is used to parse the received query request and extract the keywords in the query request, perform Boolean keyword matching through the inverted index in the hierarchical index, and perform semantic retrieval of the keywords in the vector database, and generate an initial candidate set after fusing multi-source retrieval results; a generative re-ranking module 13, the initial candidate set generation module 12 is connected to the generative re-ranking module 13, and the generative re-ranking module 13 is used to input the query request and the initial candidate set into a generative re-ranking model and generate corresponding scores through the cross-modal attention mechanism in the generative re-ranking model; a scoring module 14, the scoring module 14 is connected to the generative re-ranking module 13, and the scoring module 14 is used to re-rank the initial candidate set according to the scores, obtain the final retrieval result and display the re-ranked document chunks. In this embodiment, the specific content in the RAG-based data re-ranking retrieval system 1 can be referred to the above description and will not be elaborated one by one.

[0040] The present invention provides a RAG-based data re-ranking retrieval method and system, which generates corresponding document chunks by performing dynamic semantic chunking on the original document, stores the document chunks in a vector database and builds a hierarchical index, parses the received query request and extracts the keywords in the query request, performs Boolean keyword matching through the inverted index in the hierarchical index, and performs semantic retrieval of the keywords in the vector database, generates an initial candidate set after fusing multi-source retrieval results, inputs the query request and the initial candidate set into a generative re-ranking model and generates corresponding scores through the cross-modal attention mechanism in the generative re-ranking model, re-ranks the initial candidate set according to the scores, obtains the final retrieval result and displays the re-ranked document chunks. The present invention generates document chunks through dynamic semantic chunking, constructs a hierarchical index and combines multi-source retrieval to quickly generate an initial candidate set, improves the retrieval efficiency, and uses the cross-modal attention mechanism in the generative re-ranking model to improve the retrieval accuracy and flexibility, greatly reduces unnecessary retrieval operations, improves the overall retrieval speed, gives full play to the advantages of different retrieval methods, avoids the limitations of a single retrieval method, improves the accuracy of the final retrieval result, and meets the diverse application requirements, achieving an overall improvement in efficiency, accuracy, semantic understanding and system scalability.

[0041] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0042] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0043] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data reordering and retrieval method based on RAG, characterized in that: include: Perform dynamic semantic segmentation on the original document to generate corresponding document blocks; Storing the document blocks in a vector database and constructing a hierarchical index; Parsing the received query request and extracting keywords in the query request, performing Boolean keyword matching through the inverted index in the hierarchical index, and performing semantic search of the keywords in the vector database, and generating an initial candidate set after fusing multi-source search results; Inputting the query request and the initial candidate set into a generative re-ranking model and generating corresponding scores through a cross-modal attention mechanism in the generative re-ranking model; The initial candidate set is reordered according to the score, a retrieval result is obtained, and the reordered document blocks are displayed.

2. The RAG-based data reordering and retrieval method according to claim 1, characterized in that: The step of performing dynamic semantic segmentation processing on the original document to generate corresponding document blocks specifically includes: Preprocessing the original document to unify the text format of the original document; Segmenting the original document by a sliding window algorithm; Performing logical unit segmentation on the segmented original document through an abstract syntax tree to generate the document blocks; The length of the document block is evaluated, and if the length of the document block exceeds a first set threshold, the document block is split; if the length of the document block is shorter than a second set threshold, the document block is merged.

3. The RAG-based data reordering and retrieval method according to claim 2, characterized in that: The step of performing logical unit segmentation on the segmented original document through an abstract syntax tree to generate the document blocks specifically includes: Processing the segmented original document through a parser to construct the abstract syntax tree; Identifying logical units with relatively independent semantics from the abstract syntax tree through predefined grammar rules; The document block is generated by extracting corresponding text content from the original document through the identified logical unit.

4. The RAG-based data reordering and retrieval method according to claim 1, characterized in that: The storing the document block into the vector database and constructing a hierarchical index specifically includes: Inputting the document block into a bidirectional encoder representation model to generate a corresponding document block semantic vector; Storing the document block semantic vector in the Milvus vector database; Constructing a primary summary index through the semantic summary of the document block, constructing a secondary vector index through the semantic vector of the document block, and constructing a tertiary relationship index through the relationship network between the document blocks; An association relationship is established among the first-level summary index, the second-level vector index and the third-level relationship index, so that the first-level summary index, the second-level vector index and the third-level relationship index can be switched and work in coordination with each other.

5. The RAG-based data reordering and retrieval method according to claim 1, characterized in that: The step of storing the document block in the vector database and constructing a hierarchical index includes: A corresponding metadata tag is added to the document block, where the metadata tag includes a timestamp.

6. The RAG-based data reordering and retrieval method according to claim 1, characterized in that: The step of parsing the received query request and extracting keywords in the query request, performing Boolean keyword matching through the inverted index in the hierarchical index, and performing semantic search of the keywords in the vector database, and generating an initial candidate set after fusing multi-source search results, specifically includes: Use the BERT word segmenter to segment and tag the query request and extract the keywords; Expand the keywords through a pre-trained synonym library, perform Boolean logic combination on the expanded keywords, and calculate the relevance score of the document block through a BM25 algorithm; The query request is converted into a 768-dimensional semantic vector through the Sentence-BERT model, and a vector relevance search is performed in the HNSW graph index to obtain the cosine similarity; The relevance score and the cosine similarity are weightedly fused to generate the initial candidate set.

7. The RAG-based data reordering and retrieval method according to claim 1, characterized in that: The step of inputting the query request and the initial candidate set into a generative re-ranking model and generating corresponding scores through a cross-modal attention mechanism in the generative re-ranking model specifically includes: Converting the query request into a query request vector, and converting the document blocks in the initial candidate set into document block vectors; Combining the query request vector and the document block vector to form an input data pair, and formatting the input data pair; Inputting the formatted input data pair into the generative reordering model; The different modal features of the document block are fused to generate a comprehensive feature vector through the cross-modal attention mechanism; The input data pairs are forward propagated in the generative re-ranking model, and the forward propagation outputs the corresponding score based on the comprehensive feature vector, and the score is used to represent the relevance of the document block in the initial candidate set to the query request.

8. The RAG-based data reordering and retrieval method according to claim 7, characterized in that: The step of fusing the different modal features of the document block through the cross-modal attention mechanism to generate a comprehensive feature vector specifically includes: Extracting different modal features of the document block through the generative re-ranking model, wherein the different modal features include text modality, image modality and audio modality; Respectively calculating the self-attention of the feature vector of the text modality, the feature vector of the image modality, and the feature vector of the audio modality to obtain corresponding self-attention results; Respectively calculating the interactive attention between the feature vector of the text modality, the feature vector of the image modality, and the feature vector of the audio modality to obtain corresponding interactive attention results; assigning corresponding attention weights to the feature elements of the text modality, the feature elements of the image modality, and the feature elements of the audio modality according to the self-attention result and the interactive attention result; The comprehensive feature vector is obtained by performing weighted summation on the feature vector of the text modality, the feature vector of the image modality, and the feature vector of the audio modality according to the attention weight.

9. The RAG-based data reordering and retrieval method according to claim 7, characterized in that: The step of converting the query request into a query request vector and converting the document blocks in the initial candidate set into document block vectors includes: Select at least one of the document blocks from the initial candidate set as a candidate document block and generate a context representation corresponding to the candidate document block, wherein the context representation is a joint embedding vector generated by a transformer encoder; The document block corresponding to the query request is obtained by processing the context representation.

10. A RAG-based data reordering and retrieval system, used to implement the RAG-based data reordering and retrieval method according to any one of claims 1 to 9, characterized in that: The RAG-based data reordering and retrieval system comprises: An original document processing module, the original document processing module is used to perform dynamic semantic block processing on the original document to generate corresponding document blocks; An index building module, the index building module is connected to the original document processing module, and the index building module is used to store the document block into a vector database and build a hierarchical index; An initial candidate set generation module, the initial candidate set generation module is connected to the index building module, the initial candidate set generation module is used to parse the received query request and extract the keywords in the query request, perform Boolean keyword matching through the inverted index in the hierarchical index, perform semantic search of the keywords in the vector database, and generate an initial candidate set after fusing multi-source search results; A generative re-ranking module, wherein the initial candidate set generation module is connected to the generative re-ranking module, and the generative re-ranking module is used to input the query request and the initial candidate set into a generative re-ranking model and generate corresponding scores through a cross-modal attention mechanism in the generative re-ranking model; A scoring module, the scoring module is connected to the generative re-ranking module, and the scoring module is used to re-rank the initial candidate set according to the score, obtain the retrieval result and display the re-ranked document block.

Citation Information

Patent Citations

  • Search result diversification method based on self-attention network

    CN112182439A

  • RAG system optimization method and system, electronic equipment and storage medium

    CN118917305A

  • Zero sample cross-language reordering method based on large language model, electronic equipment and medium

    CN118939757A

  • Domain speech recognition method and system based on RAG

    CN119296516A

Cited By

  • Retrieval optimization method and device in retrieval system, equipment, medium and product

    CN120407516A

  • Search optimization method, device, equipment, medium and product in search system

    CN120407516B

  • Intelligent Agent collaborative decision-making system for environment monitoring

    CN120509614A

  • Design demand result generation method and system and computer medium

    CN120671546A

  • Resource recommendation method and system based on hybrid retrieval RAG

    CN120780916A