Retrieval enhancement generation method and system based on large model multi-level sorting

By introducing multi-level sorting technology based on large models into the traditional generative AI model, the knowledge limitations and accuracy problems of traditional models when dealing with complex scenarios are solved, and more efficient and accurate information retrieval is achieved.

CN120045690APending Publication Date: 2025-05-27AISINO CORPORATION

Patent Information

Application Number
CN202411928229.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Traditional generative AI models face problems such as knowledge limitations, insufficient real-time update capabilities, hallucinations, outdated knowledge, opaque reasoning and inability to trace sources when dealing with complex and changing real scenarios, resulting in low retrieval accuracy and inability to ensure that users obtain the most relevant and useful information.

Method used

The search enhancement generation method based on multi-level sorting of large models is adopted, and multiple rounds of rewriting and disassembly of user inputs are performed, intention recognition optimization model is used for intent analysis, multi-source recall is performed in combination with the built knowledge base, and search results are optimized through multi-level sorting (fine sorting, satisfaction sorting, usability sorting).

Benefits of technology

Improve the accuracy and quality of generated content, ensure that users can obtain the most relevant and useful information, effectively deal with complex and professional problems, and meet practical application needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045690A_ABST
    Figure CN120045690A_ABST
Patent Text Reader

Abstract

The invention discloses a retrieval enhancement generation method and system based on large model multi-level sorting, and the method comprises the steps: carrying out the multi-round rewriting and disassembly operation of a question inputted by a user, and carrying out the analysis and summarization of the disassembled data, so as to obtain the disassembled data; performing intention analysis by utilizing an intention recognition optimization model on the basis of the disassembly data, and obtaining weight data corresponding to each intention; based on the weight data corresponding to each intention and the constructed knowledge base, performing multi-source recall by using a retrieval enhancement generation technology to obtain a multi-source recall result; and performing multi-level sorting on the multi-source recall result to obtain a retrieval result. According to the method, the user intention recognition model, the satisfaction sorting based on the thinking chain technology and the availability sorting based on the large language model are introduced, so that the real requirements of the user are accurately understood, the accuracy and quality of the generated content are improved, the complex and professional problems are solved, and the actual application requirements are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of information retrieval, and more specifically, to a retrieval augmented generation method and system based on multi-level sorting of large models. Background Art

[0002] With the rapid development of artificial intelligence technology, especially the continuous breakthroughs in the field of natural language processing (NLP), generative AI technology has become an important force driving industry innovation and transformation. Traditional generative models have made significant progress in text generation, but they are often limited by the model's own knowledge limitations and lack of the ability to update in real time, making it difficult to handle complex and changing real-world scenarios and facing challenges such as hallucinations, outdated knowledge, opaque reasoning, and inability to trace the source. Against this background, retrieval augmented generation (RAG) technology has emerged. This technology integrates the internal knowledge of large models with external databases, thereby improving the accuracy and credibility of the generated content and bringing prospects for knowledge-intensive tasks.

[0003] Traditional retrieval augmented generation mainly achieves an efficient retrieval process through the combination of a rough ranking and a fine ranking. Based on the rough ranking, relevant fragments are quickly screened out from a large knowledge base, and then the candidate results screened out in the rough ranking stage are further sorted and optimized through the fine ranking to improve the relevance and accuracy of the final results. Although traditional RAG technology has made significant progress in improving the accuracy and credibility of the content generated by the model, there are still problems such as low retrieval accuracy and inability to ensure that users can obtain the most relevant and useful information.

[0004] Therefore, a retrieval augmented generation method based on multi-level sorting of large models is needed. Summary of the Invention

[0005] The present invention proposes a retrieval augmented generation method and system based on multi-level sorting of large models to solve the problem of how to achieve multi-level sorting of large models.

[0006] To solve the above problems, according to one aspect of the present invention, there is provided a retrieval augmented generation method based on multi-level sorting of large models, the method comprising:

[0007] Performing multi-round rewriting and disassembling operations on the problem input by the user respectively, and analyzing and summarizing the disassembled data to obtain disassembled data;

[0008] Based on the disassembled data, using an intent recognition optimization model to perform intent analysis to obtain weight data corresponding to each intent;

[0009] Based on the weight data corresponding to each intent and the constructed knowledge base, using retrieval augmented generation technology to perform multi-source recall to obtain multi-source recall results;

[0010] Perform multi-level sorting on the multi-source recall results to obtain retrieval results.

[0011] Preferably, the analyzing and summarizing the disassembled data to obtain disassembled data includes:

[0012] After the problem is disassembled, for comparison, statistics, and list problems, call the Agent to plan the tasks, execute the tasks in parallel, and perform summary analysis to obtain disassembled data.

[0013] Preferably, the method further includes:

[0014] Construct an initial intent recognition model and perform training to obtain the optimized intent recognition model, including: collecting text data of different intent categories and expression forms, and performing intent annotation and data preprocessing; constructing an initial intent recognition model based on the Transform architecture, and performing training and optimization based on the intent annotation data and the preprocessed data to obtain the optimized intent recognition model.

[0015] Preferably, the method further includes:

[0016] According to the type of data, split the input data according to the preset document parsing and chunking strategy, and use the RoBERTa-based text embedding model to convert the split text sequence into a 768-dimensional vector sequence and store it in the ES vector database, and establish multiple indexes in the ES vector database to construct the database.

[0017] Preferably, the splitting the input data according to the type of data according to the preset document parsing and chunking strategy includes:

[0018] When the input data is document-type data, obtain multi-level headings and paragraphs by extracting the tree structure;

[0019] When the input data is unstructured long text, use full stops and line breaks to separate the data or perform brute-force chunking on the data by window length;

[0020] When the input data is FAQ-type data, place it in a unified library and perform question-to-question Q-Q matching to perform chunking.

[0021] Preferably, the performing multi-source recall using retrieval augmented generation technology based on the weight data corresponding to each intent and the constructed knowledge base to obtain multi-source recall results includes:

[0022] Based on the constructed knowledge base, perform recall in 4 ways: keyword retrieval, vector retrieval, network retrieval, and structured information recall to obtain the initial recall results;

[0023] Determine the actual recall result for each path based on the product of the initial recall result and the weight data corresponding to the corresponding intent.

[0024] Preferably, the method further includes:

[0025] When using the keyword retrieval method for recall, use the BM25 or TF IDF information retrieval algorithm according to the requirements; when using the vector retrieval method for recall, convert the input statement into a vector representation through the BGE-M3 text embedding model, calculate the similarity with the vector set in the knowledge base, sort in descending order according to the similarity results, and return in the form of a list.

[0026] Preferably, the multi-source recall results are sorted at multiple levels to obtain the retrieval results, including:

[0027] Sort the multi-source recall results successively through fine ranking, satisfiability ranking, and usability ranking to obtain the retrieval results;

[0028] Among them, use the BGE-ReRank re-ranking model for fine ranking; use the chain of thought technology for satisfiability ranking; perform usability ranking based on the method of writing and splicing Prompts in order.

[0029] According to another aspect of the present invention, there is provided a retrieval enhanced generation system based on large model multi-level ranking, the system includes:

[0030] A question rewriting module, which is used to perform multi-round rewriting and disassembling operations on the user input question respectively, and analyze and summarize the disassembled data to obtain disassembled data;

[0031] An intent recognition module, which is used to perform intent analysis based on the disassembled data by using an intent recognition optimization model to obtain weight data corresponding to each intent;

[0032] A multi-source recall module, which is used to perform multi-source recall by using retrieval enhanced generation technology based on the weight data corresponding to each intent and the constructed knowledge base to obtain multi-source recall results;

[0033] A multi-level ranking module, which is used to perform multi-level ranking on the multi-source recall results to obtain the retrieval results.

[0034] Preferably, the question rewriting module analyzes and summarizes the disassembled data to obtain the disassembled data, including:

[0035] After the question is disassembled, for comparison, statistics, and list questions, call the Agent to plan the task, execute the task in parallel, and perform summary analysis to obtain the disassembled data.

[0036] Preferably, the system further includes:

[0037] An intention recognition optimization model determination module, configured to construct an initial intention recognition model and perform training to obtain the intention recognition optimization model, including: collecting text data of different intention categories and expression ways, and performing intention annotation and data preprocessing; constructing an initial intention recognition model based on the Transform architecture, and performing training and optimization based on the intention annotation data and the preprocessed data to obtain the intention recognition optimization model.

[0038] Preferably, the system further includes:

[0039] A knowledge base construction module, configured to split the input data according to a preset document parsing and chunking strategy according to the type of data, and use a RoBERTa-based text embedding model to convert the split text sequence into a 768-dimensional vector sequence and store it in the ES vector database, and establish multiple indexes in the ES vector database to construct the database.

[0040] Preferably, the knowledge base construction module splits the input data according to a preset document parsing and chunking strategy according to the type of data, including:

[0041] When the input data is document-type data, obtain multi-level headings and paragraphs by extracting the tree structure;

[0042] When the input data is unstructured long text, use full stops and line breaks to separate the data or perform brute-force chunking on the data by the window length;

[0043] When the input data is FAQ-type data, place it in a single library and perform question-question (Q-Q) matching to perform chunking.

[0044] Preferably, the multi-source recall module performs multi-source recall using retrieval augmented generation technology based on the weight data corresponding to each intention and the constructed knowledge base to obtain multi-source recall results, including:

[0045] Based on the constructed knowledge base, perform recall in 4 ways of keyword retrieval, vector retrieval, network retrieval, and structured information recall to obtain an initial recall result;

[0046] Determine the actual recall result of each path based on the product of the initial recall result and the weight data corresponding to the corresponding intention.

[0047] Preferably, the multi-source recall module is further configured to:

[0048] When using keyword retrieval for recall, the BM25 or TF IDF information retrieval algorithm is adopted according to requirements; when using vector retrieval for recall, the input statement is converted into a vector representation through the BGE-M3 text embedding model, the similarity is calculated with the vector set in the knowledge base, sorted in descending order according to the similarity results, and returned in the form of a list.

[0049] Preferably, the multi-level sorting module performs multi-level sorting on the multi-source recall results to obtain retrieval results, including:

[0050] Performing fine sorting, satisfiability sorting, and usability sorting on the multi-source recall results in sequence to obtain retrieval results;

[0051] Among them, fine sorting is performed using the BGE-ReRank re-ranking model; satisfiability sorting is performed using the chain of thought technique; usability sorting is performed based on the method of writing and splicing Prompts in order.

[0052] Based on another aspect of the present invention, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of any one of the retrieval enhanced generation methods based on large model multi-level sorting.

[0053] Based on another aspect of the present invention, the present invention provides an electronic device, including:

[0054] The above-mentioned computer-readable storage medium; and

[0055] One or more processors for executing the program in the computer-readable storage medium.

[0056] The present invention provides a retrieval enhanced generation method and system based on large model multi-level sorting, including: performing multiple rounds of rewriting and disassembling operations on the problem input by the user respectively, and analyzing and summarizing the disassembled data to obtain disassembled data; based on the disassembled data, using an intent recognition optimization model to perform intent analysis to obtain weight data corresponding to each intent; based on the weight data corresponding to each intent and the constructed knowledge base, using retrieval enhanced generation technology to perform multi-source recall to obtain multi-source recall results; performing multi-level sorting on the multi-source recall results to obtain retrieval results. The present invention realizes the accurate understanding of the user's real needs, improves the accuracy and quality of the generated content, and meets the actual application needs by introducing a user intent recognition model, satisfiability sorting based on the chain of thought technique, and usability sorting based on a large language model on the basis of the traditional "coarse sorting + fine sorting" retrieval strategy. Description of the Drawings

[0057] The exemplary embodiments of the present invention can be more fully understood by referring to the following drawings:

[0058] Figure 1 FIG. is a flowchart of a retrieval enhanced generation method 100 based on large model multi-level sorting according to an embodiment of the present invention;

[0059] Figure 2 FIG. is a schematic diagram of a problem rewriting process according to an embodiment of the present invention;

[0060] Figure 3 FIG. is a flowchart of a knowledge base construction according to an embodiment of the present invention;

[0061] Figure 4 FIG. is a schematic diagram of a multi-source recall and multi-level sorting process according to an embodiment of the present invention;

[0062] Figure 5 FIG. is a schematic diagram of the structure of a retrieval enhanced generation system 500 based on large model multi-level sorting according to an embodiment of the present invention;

[0063] Figure 6 FIG. is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Specific Embodiments

[0064] Now, exemplary embodiments of the present invention will be described with reference to the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. These embodiments are provided to disclose the present invention in detail and completely, and to fully convey the scope of the present invention to those skilled in the art. The terms in the exemplary embodiments shown in the drawings are not intended to limit the present invention. In the drawings, the same units / components are denoted by the same reference numerals.

[0065] Unless otherwise specified, the terms (including scientific and technical terms) used herein have the ordinary meaning understood by those skilled in the art. Additionally, it can be understood that terms defined in a commonly used dictionary should be construed to have a meaning consistent with the context of their relevant fields, and should not be construed as having an idealized or overly formal meaning.

[0066] Figure 1 FIG. is a flowchart of a retrieval enhanced generation method 100 based on large model multi-level sorting according to an embodiment of the present invention. As Figure 1As shown in the figure, the retrieval enhanced generation method based on large model multi-level sorting provided by the embodiment of the present invention realizes the accurate understanding of the user's real needs, improves the accuracy and quality of the generated content to cope with complex and professional problems and meet the actual application requirements by introducing a user intention recognition model, a satisfaction sorting based on the thought chain technology, and an availability sorting based on the large language model on the basis of the traditional "coarse ranking + fine ranking" retrieval strategy. The retrieval enhanced generation method 100 based on large model multi-level sorting provided by the embodiment of the present invention starts from step 101. At step 101, the problem input by the user is rewritten and disassembled in multiple rounds, and the disassembled data is analyzed and summarized to obtain the disassembled data.

[0067] Preferably, the analyzing and summarizing the disassembled data to obtain the disassembled data includes:

[0068] After the problem is disassembled, for comparison, statistics and list problems, call the Agent to plan the task, execute the task in parallel, and conduct summary analysis to obtain the disassembled data.

[0069] The solution of the present invention can be divided into five parts: problem rewriting, intention recognition, knowledge base construction, multi-source recall, and multi-level sorting.

[0070] In the present invention, the user input is rewritten based on the problem rewriting module. Combining Figure 2 As shown in the figure, in this module, the user problem is preferentially rewritten in multiple rounds, and then the problem is disassembled. After the problem is disassembled, for comparison, statistics, list and other problems, call the Agent to plan the task, execute the task in parallel, and finally summarize. For example: the input problem: Please list the fiscal and tax-related policies and regulations since March this year. The Agent will call NL2SQL to query the relational database.

[0071] At step 102, based on the disassembled data, use the intention recognition optimization model to perform intention analysis to obtain the weight data corresponding to each intention.

[0072] Preferably, the method further includes:

[0073] Construct an initial intention recognition model and perform training to obtain the intention recognition optimization model, including: collecting text data of different intention categories and expression ways, and performing intention annotation and data preprocessing; constructing an initial intention recognition model based on the Transform architecture, and performing training and optimization based on the intention annotation data and the preprocessed data to obtain the intention recognition optimization model.

[0074] In the present invention, based on the intention recognition module, use the intention recognition optimization model to perform intention analysis on the rewritten information and give the classification weight.

[0075] Among them, the process of determining the optimized intent recognition model includes: collecting a large amount of text data covering as many intent categories and expression ways as possible, and performing intent annotation and data preprocessing on it; secondly, training and tuning based on pre-trained models (such as Bert, RoBERTa, etc.) of the Transform architecture; finally, evaluating the model performance, and adjusting the model structure or hyperparameters according to the evaluation results, and re-training the model to determine the optimized intent recognition model.

[0076] In step 103, based on the weight data corresponding to each intent and the constructed knowledge base, use retrieval-enhanced generation technology for multi-source recall to obtain multi-source recall results.

[0077] Preferably, the method further includes:

[0078] According to the type of data, split the input data according to the preset document parsing and chunking strategy, and use the RoBERTa-based text embedding model to convert the split text sequence into a 768-dimensional vector sequence and store it in the ES vector database, and establish multiple indexes in the ES vector database to build the database.

[0079] Preferably, the splitting of the input data according to the type of data according to the preset document parsing and chunking strategy includes:

[0080] When the input data is document-type data, obtain multi-level headings and paragraphs by extracting the tree structure.

[0081] When the input data is unstructured long text, use full stops and line breaks to separate the data or perform brute-force chunking on the data by the window length.

[0082] When the input data is FAQ-type data, place it in a unified library and perform question-question (Q-Q) matching to perform chunking.

[0083] Preferably, the use of retrieval-enhanced generation technology for multi-source recall based on the weight data corresponding to each intent and the constructed knowledge base to obtain multi-source recall results includes:

[0084] Based on the constructed knowledge base, perform recall in 4 ways of keyword retrieval, vector retrieval, online retrieval, and structured information recall to obtain initial recall results.

[0085] Determine the actual recall result of each path based on the product of the initial recall result and the weight data corresponding to the corresponding intent.

[0086] Preferably, the method further includes:

[0087] When using keyword retrieval for recall, the BM25 or TF IDF information retrieval algorithm is used according to requirements; when using vector retrieval for recall, the input statement is converted into a vector representation through the BGE-M3 text embedding model, and the similarity is calculated with the vector set in the knowledge base, sorted in descending order according to the similarity results, and returned in the form of a list.

[0088] In the present invention, the knowledge base construction module processes the original data by splitting it into items, stores it in the database, and builds an index based on semantic representation to construct the knowledge base. Combining Figure 3 As shown, first, the original data is split into data blocks. This module adopts different strategies for different types of original data. For document data, by extracting the tree structure, multi-level headings and paragraphs are obtained. For unstructured long texts, the data is separated by full stops and line breaks or violently chunked by window length; for FAQ type data, it is unified in one library and Q-Q matching is performed. Through document parsing and unified chunking strategies, the input data is split into multiple paragraphs with a length less than 512, and the text sequence is converted into a 768-dimensional vector sequence by the RoBERTa-based text embedding model and stored in the ES vector database. Secondly, the knowledge base index construction is completed as needed. Among them, the HNSW algorithm is used to establish multiple indexes in the vector knowledge base, such as document data indexes, unstructured text indexes, FAQ data indexes, etc. The metadata is stored in a relational database.

[0089] In the present invention, the multi-source recall module contains as much relevant knowledge as possible. This module uses four recall methods, namely keyword retrieval, vector retrieval, network retrieval, and structured information recall, to recall the Top500, and no similarity scoring threshold is set for the overall module. Among them, keyword retrieval can use information retrieval algorithms such as BM25 or TF IDF according to requirements. Vector retrieval mainly converts the input statement into a vector representation through the BGE-M3 text embedding model, calculates the similarity with the vector set in the knowledge base, sorts in descending order according to the similarity results, and returns in the form of a list. In addition, the actual recall result of each path should be the product of the current recall result and the intent recognition weight. For example: the input question is "Who is the president of the United States this year?" After traditional retrieval, the ranking is 1-(ranking / total number). After intent recognition, the weights for various retrieval libraries are 0.1, 0.7, and 0.2 respectively. Therefore, the original recall results are sorted after weighting.

[0090] The recall result TopKnewi = (1 - Pi / N) * Wi

[0091] Where Pi is the i-th ranking after traditional retrieval, N is the total number of recalls, which is 500 in this patent, that is, 500 rough-rank recalls, and Wi is the intent recognition weight of the library corresponding to the i-th ranking.

[0092] In step 104, multi-level sorting is performed on the multi-source recall results to obtain retrieval results.

[0093] Preferably, the multi-level sorting of the multi-source recall results to obtain retrieval results includes:

[0094] Performing fine sorting, satisfiability sorting, and usability sorting on the multi-source recall results in sequence to obtain retrieval results;

[0095] Among them, the BGE-ReRank re-ranking model is used for fine sorting; the chain of thought technique is used for satisfiability sorting; and usability sorting is performed based on the method of writing and splicing Prompts in order.

[0096] Combined Figure 4 As shown, in the present invention, a multi-level sorting module is utilized to achieve knowledge retrieval through three-level retrieval of "fine sorting + satisfiability sorting + usability sorting". First, based on the Top 500 results recalled by the multi-source recall module, fine sorting is performed to select the Top 100 results. The fine sorting module adopts the BGE-ReRank re-ranking model, and through deep learning technology and a fine semantic matching algorithm, further optimization and refinement of the rough sorting retrieval results are realized, improving the relevance and accuracy of the retrieval results. Secondly, based on the fine sorting, satisfiability sorting is performed based on the chain of thought technique to select the Top 20 to achieve precise sorting of the retrieval results. Specifically, for the Top 20 retrieval results of the fine sorting, Prompts can be written and spliced as needed (for example, if there are policy results from different sources and different times in the retrieval results, Prompts can be considered to be written in combination with the authority and timeliness of the sources), to achieve effects such as time-limited sorting or region-priority sorting, and the spliced results are input into the large model, and step-by-step reasoning is performed through the chain of thought method to generate a more accurate and reasonable sorting. For example:

[0097]

[0098]

[0099] Finally, Prompts are spliced for the Top 20 results after satisfiability sorting and input into the large model. The large model discriminates the relevance between the Top 20 retrieval results (corresponding to the "paragraphs" in the following Prompts) and the user input, and gives the score of each recall result. Finally, sorting can be performed from high to low according to the scores to obtain the Top 5 retrieval results, further narrowing the scope and improving the retrieval accuracy.

[0100]

[0101] The key points of the present invention are:

[0102] 1. Introduce a user intention recognition model to accurately understand the user's real needs and improve the accuracy in the rough ranking stage;

[0103] 2. On the basis of the traditional "rough ranking + fine ranking" retrieval strategy, introduce a multi-level ranking strategy including satisfaction ranking based on the chain of thought technology and usability ranking based on the large language model to improve the accuracy and quality of the generated content.

[0104] The multi-level ranking method provided by the present invention introduces a user intention recognition model, satisfaction ranking based on the chain of thought technology, and usability ranking strategy based on the large language model, which can quickly analyze the user's input, accurately understand the user's real needs, significantly improve the accuracy and quality of the model-generated content, greatly enhance the user experience, and can well handle complex and professional problems to meet the actual application requirements.

[0105] Figure 5 FIG. 500 is a schematic structural diagram of a retrieval enhanced generation system based on large model multi-level ranking according to an embodiment of the present invention. As Figure 5 shown, the retrieval enhanced generation system 500 based on large model multi-level ranking provided by the embodiment of the present invention includes: a question rewriting module 501, an intention recognition module 502, a multi-source recall module 503, and a multi-level ranking module 504.

[0106] Preferably, the question rewriting module 501 is used to perform multiple rounds of rewriting and disassembling operations on the question input by the user, and analyze and summarize the disassembled data to obtain the disassembled data.

[0107] Preferably, the question rewriting module 501 analyzes and summarizes the disassembled data to obtain the disassembled data, including:

[0108] After the question is disassembled, for comparison, statistics, and list questions, call the Agent to plan the task, execute the task in parallel, and perform summary analysis to obtain the disassembled data.

[0109] Preferably, the intention recognition module 502 is used to perform intention analysis based on the disassembled data by using an intention recognition optimization model to obtain weight data corresponding to each intention.

[0110] Preferably, the system further includes:

[0111] An intention recognition optimization model determination module, which is used to construct an initial intention recognition model and train it to obtain the intention recognition optimization model, including: collecting text data of different intention categories and expression ways, and performing intention annotation and data preprocessing; constructing an initial intention recognition model based on the Transform architecture, and performing training and optimization based on the intention annotation data and the preprocessed data to obtain the intention recognition optimization model.

[0112] Preferably, the multi-source recall module 503 is used to perform multi-source recall based on the weight data corresponding to each intention and the constructed knowledge base by using retrieval-augmented generation technology to obtain a multi-source recall result.

[0113] Preferably, the system further includes:

[0114] A knowledge base construction module, which is used to split the input data according to the preset document parsing and chunking strategy according to the type of data, and use a text embedding model based on RoBERTa to convert the split text sequence into a 768-dimensional vector sequence and store it in the ES vector database, and establish multiple indexes in the ES vector database to construct the database.

[0115] Preferably, the knowledge base construction module splits the input data according to the preset document parsing and chunking strategy according to the type of data, including:

[0116] When the input data is document-type data, obtain multi-level headings and paragraphs by extracting the tree structure;

[0117] When the input data is unstructured long text, separate the data using full stops and line breaks or perform brute-force chunking on the data by the window length;

[0118] When the input data is FAQ-type data, place it in a single library and perform question-to-question Q-Q matching for chunking.

[0119] Preferably, the multi-source recall module 503 performs multi-source recall based on the weight data corresponding to each intention and the constructed knowledge base by using retrieval-augmented generation technology to obtain a multi-source recall result, including:

[0120] Based on the constructed knowledge base, perform recall in a 4-way recall manner of keyword retrieval, vector retrieval, online retrieval, and structured information recall to obtain an initial recall result;

[0121] Determine the actual recall result of each path based on the product of the initial recall result and the weight data corresponding to the corresponding intention.

[0122] Preferably, the multi-source recall module 503 is further used for:

[0123] When using keyword retrieval to recall, the BM25 or TF IDF information retrieval algorithm is used according to requirements; when using vector retrieval to recall, the input statement is converted into a vector representation through the BGE-M3 text embedding model, and the similarity is calculated with the vector set in the knowledge base, sorted in descending order according to the similarity results, and returned in the form of a list.

[0124] Preferably, the multi-level sorting module 504 is used to perform multi-level sorting on the multi-source recall results to obtain the retrieval results.

[0125] Preferably, the multi-level sorting module 504 performs multi-level sorting on the multi-source recall results to obtain the retrieval results, including:

[0126] Performing fine sorting, satisfiability sorting, and usability sorting on the multi-source recall results in sequence to obtain the retrieval results;

[0127] Among them, fine sorting is performed using the BGE-ReRank re-ranking model; satisfiability sorting is performed using the chain of thought technique; usability sorting is performed based on the method of writing and splicing Prompts in order.

[0128] The retrieval enhanced generation system 500 based on large model multi-level sorting in the embodiments of the present invention corresponds to the retrieval enhanced generation method 100 based on large model multi-level sorting in another embodiment of the present invention, and will not be elaborated here.

[0129] On the other hand of the present invention, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of any one of the retrieval enhanced generation methods based on large model multi-level sorting.

[0130] On the other hand of the present invention, the present invention provides an electronic device, including:

[0131] The above-mentioned computer-readable storage medium; and

[0132] One or more processors for executing the program in the computer-readable storage medium.

[0133] Figure 6 is the structure of the electronic device provided by an exemplary embodiment of the present invention. The electronic device can be any one or both of the first device and the second device, or a stand-alone device independent of them, and the stand-alone device can communicate with the first device and the second device to receive the input signals collected from them. Figure 6 Illustrates a block diagram of an electronic device according to an embodiment of the present disclosure. As Figure 6As shown, the electronic device 600 includes one or more processors 601 and a memory 602.

[0134] The processor 601 can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device to perform desired functions.

[0135] The memory 602 can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage media, and the processor 601 can run the program instructions to implement the retrieval-enhanced generation method based on large model multi-level sorting of the software programs of the various embodiments of the present disclosure described above and / or other desired functions. In one example, the electronic device may further include: an input device 603 and an output device 604, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0136] In addition, the input device 603 may further include, for example, a keyboard, a mouse, and so on.

[0137] The output device 604 can output various information to the outside. The output device 604 can include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, and so on.

[0138] Of course, for simplicity, Figure 6 only some of the components related to the present disclosure in the electronic device are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device may further include any other appropriate components.

[0139] In addition to the above methods and devices, the embodiments of the present disclosure may also be computer program products, which include computer program instructions that, when run by a processor, cause the processor to execute the steps in the retrieval-enhanced generation method based on large model multi-level sorting according to various embodiments of the present disclosure described in the "Exemplary Method" section above of this specification.

[0140] The computer program product may be written in any combination of one or more programming languages for executing the program code of the operations of the embodiments of the present disclosure. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0141] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are run by a processor, the processor is caused to execute the steps in the retrieval enhanced generation method based on large model multi-level sorting according to various embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.

[0142] The computer-readable storage medium may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0143] The basic principles of the present disclosure have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. In addition, the above-mentioned specific details are only for the purpose of illustration and easy understanding, and are not limitations. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0144] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments may be referred to each other. For system embodiments, since they basically correspond to method embodiments, the description is relatively simple, and the relevant parts may refer to the partial description of the method embodiments.

[0145] The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any manner. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The word "or" and "and" used herein refer to the word "and / or", and can be used interchangeably with each other, unless the context clearly indicates otherwise. The word "such as" used herein refers to the phrase "such as but not limited to", and can be used interchangeably with each other.

[0146] The methods and apparatuses of this disclosure can be implemented in many ways. For example, the methods and apparatuses of this disclosure can be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of the steps for the methods is for illustrative purposes only, and the steps of the methods of this disclosure are not limited to the specific order described above, unless otherwise specifically stated. In addition, in some embodiments, this disclosure can also be implemented as a program recorded in a recording medium, and these programs include machine-readable instructions for implementing the methods according to this disclosure. Therefore, this disclosure also covers the recording medium storing the programs for executing the methods according to this disclosure.

[0147] It should also be noted that in the apparatuses, equipment, and methods of this disclosure, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this disclosure. The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects are very obvious to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

[0148] The above description has been given for purposes of illustration and description. In addition, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.

Claims

1. A retrieval enhancement generation method based on large model multi-level ranking, characterized in that: The method comprises: Perform multiple rounds of rewriting and disassembly operations on the questions input by the user, and analyze and summarize the disassembled data to obtain disassembled data; Based on the disassembly data, an intent recognition optimization model is used to perform intent analysis to obtain weight data corresponding to each intent; Based on the weight data corresponding to each intent and the constructed knowledge base, multi-source recall is performed using retrieval enhancement generation technology to obtain multi-source recall results; The multi-source recall results are sorted in multiple levels to obtain retrieval results.

2. The method according to claim 1, characterized in that The disassembled data is analyzed and summarized to obtain disassembled data, including: After the problem is broken down, for comparison, statistics and list problems, the Agent is called to plan the tasks, execute the tasks in parallel, and perform summary analysis to obtain the broken down data.

3. The method according to claim 1, characterized in that The method further comprises: Constructing an initial intent recognition model and performing training to obtain the intent recognition optimization model, including: collecting text data of different intent categories and expressions, and performing intent annotation and data preprocessing; constructing an initial intent recognition model based on the Transform architecture, and performing training and tuning based on the intent annotation data and preprocessed data to obtain the intent recognition optimization model.

4. The method according to claim 1, characterized in that: The method further comprises: According to the data type and the preset document parsing and segmentation strategy, the input data is split, and the split text sequence is converted into a 768-dimensional vector sequence using the RoBERTa-based text embedding model and stored in the ES vector database. In addition, multiple indexes are established in the ES vector database to construct the database.

5. The method according to claim 4, characterized in that The input data is split according to the type of data and the preset document parsing and block splitting strategy, including: When the input data is document data, multi-level titles and paragraphs are obtained by extracting the tree structure; When the input data is a long text without structure, use period or line break to separate the data or divide the data into blocks by window length; When the input data is of FAQ type, it is placed in a library and divided into blocks by matching questions with question QQ.

6. The method according to claim 1, characterized in that The multi-source recall is performed based on the weight data corresponding to each intent and the constructed knowledge base using the retrieval enhancement generation technology to obtain the multi-source recall results, including: Based on the constructed knowledge base, keyword search, vector search, network search and structured information recall are used to retrieve the initial recall results in a 4-way recall mode. Based on the product of the initial recall result and the weight data corresponding to the corresponding intent, the actual recall result of each path is determined.

7. The method according to claim 6, characterized in that The method further comprises: When recalling by keyword retrieval, the BM25 or TFIDF information retrieval algorithm is used according to the needs; when recalling by vector retrieval, the input sentence is converted into a vector representation through the BGE-M3 text embedding model, and the similarity is calculated with the vector set in the knowledge base. The similarity results are sorted in descending order and returned in the form of a list.

8. The method according to claim 1, characterized in that The multi-source recall results are sorted in multiple levels to obtain retrieval results, including: The plurality of multi-source recall results are sequentially sorted, sorted by satisfaction and sorted by availability to obtain a search result; Among them, the BGE-ReRank re-ranking model is used for precise sorting; the thinking chain technology is used for satisfaction sorting; and the availability sorting is based on the method of writing and splicing prompts in sequence.

9. A retrieval enhancement generation system based on large model multi-level ranking, characterized in that: The system comprises: The question rewriting module is used to perform multiple rounds of rewriting and disassembly operations on the questions input by the user, and analyze and summarize the disassembled data to obtain the disassembled data; An intention recognition module is used to perform intention analysis based on the disassembly data using an intention recognition optimization model to obtain weight data corresponding to each intention; The multi-source recall module is used to perform multi-source recall based on the weight data corresponding to each intent and the constructed knowledge base, using the retrieval enhancement generation technology to obtain multi-source recall results; The multi-level sorting module is used to perform multi-level sorting on the multi-source recall results to obtain retrieval results.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

11. An electronic device, characterized in that: include: The computer readable storage medium as claimed in claim 10; as well as One or more processors are used to execute the program in the computer-readable storage medium.

Citation Information

Patent Citations

  • Search result sorting method, device and equipment and computer readable storage medium

    CN111949898A

  • Aggregation retrieval method and device based on multivariate data, equipment and storage medium

    CN112182150A

  • Purchase scene vertical search method, device and system

    CN114756570A

  • Multi-intention recognition method and system and readable storage medium

    CN118861200A

Cited By

  • Scientific and technological achievement and enterprise technology demand matching method and system

    CN120631940A

  • Mixed information retrieval generation method, system, equipment and medium

    CN120763229A

  • Intelligent question and answer generation method based on triple retrieval query optimization

    CN120950648A

  • Multi-document industry knowledge assistant system based on multi-modal knowledge graph

    CN121350265A