Retrieval enhancement generation method and retrieval enhancement generation system

Through semantic vectorization processing and similarity analysis of the knowledge base and historical query results, combined with user evaluation feedback, the problem of inaccurate search results in the existing technology is solved, and efficient and low-cost search enhanced generation is achieved to meet the needs of users in professional fields.

CN120336467APending Publication Date: 2025-07-18JIANGNAN SHIPYARD (GRP) CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510401396.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing search enhancement generation technology is difficult to meet users' accuracy requirements for high-frequency causes in professional fields, and relying on pre-trained models leads to increased system complexity and cost, so it is impossible to make full use of historical query records to optimize Q&A content.

Method used

By semantic vectorization of the existing knowledge base content, a historical relationship diagram is established, and similarity analysis is performed based on the historical query results confirmed by users and the direct query results, reordering and generating query results that are more in line with user needs, and introducing a user evaluation feedback mechanism to optimize the historical query library.

Benefits of technology

It realizes more accurate search results, reduces system complexity and cost, improves search efficiency, and continuously enriches the historical query library through user feedback to improve the accuracy of subsequent searches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336467A_ABST
    Figure CN120336467A_ABST
Patent Text Reader

Abstract

The invention provides a retrieval enhancement generation method and a retrieval enhancement generation system. The retrieval enhancement generation method comprises the following steps: segmenting basic knowledge text segments and carrying out semantic vectorization processing; establishing a historical relation graph; carrying out vectorization processing on nodes of the historical relation graph; obtaining a direct query result vector and a corresponding direct matching degree according to user query content; obtaining similar historical query vectors and corresponding similar historical query result vectors according to user query contents; calculating a historical matching degree through the direct query result vector and the similar historical query result vector; and calculating the overall matching degree, reordering the direct query result vectors, and then generating cue words and inputting the cue words into the large model. According to the technical scheme, the similarity analysis can be performed on the historical query result confirmed by the user and the direct query result, and the direct query result is reordered, so that the target query result which is more accurate and better meets the user requirement is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information processing, and more particularly, to a retrieval augmented generation method and a retrieval augmented generation system. Background Art

[0002] With the development of large pre-trained language models, the importance of retrieval augmented generation technology has become increasingly prominent. The retrieval augmented generation technology aims to improve the effects of information retrieval and text generation by combining the advantages of retrieval systems and generation models. Currently, the existing common retrieval augmented generation technology solutions can be divided into two types. The first solution expands the retrieval content by introducing a knowledge graph and using the entity and attribute relationships in the graph to expand the relevant text segments that can be retrieved by the user query. Moreover, by adopting a combination of sparse retrieval and dense retrieval, more comprehensive relevant texts can be obtained. This solution performs well in enhancing the richness of retrieval results and can provide users with a wider range of information sources. The second solution mainly uses a pre-trained model to assist in answer generation, that is, constructing a question-answering system based on the pre-trained model. This solution constructs a pre-trained model and uses additional knowledge to assist in generating more accurate question-answering content. Its core lies in improving the accuracy and fluency of the question-answering system through the deep semantic understanding and generation ability of the pre-trained model.

[0003] However, the existing retrieval augmented generation technology solutions often have the problem of insufficient accuracy. Specifically, in some professional fields (such as the field of ship maintenance), users are more concerned about the high-frequency causes of problems rather than comprehensive and generalized answers. For example, there may be multiple causes for a ship's main engine to leak oil, but users hope that the system can quickly locate the high-frequency causes (such as seal ring aging, improper oil pressure setting, etc.) to improve the maintenance efficiency. Historical query records usually contain sufficient user behavior and preference information, which can provide important reference bases for retrieval. However, the existing retrieval augmented generation technology solutions only adopt the methods of expanding retrieval content or generating answers using a pre-trained model, and fail to fully utilize historical query records to optimize question-answering content. Therefore, it is difficult to meet the accuracy requirements of users. In addition, for the retrieval augmented generation technology solutions that require the use of a pre-trained model, since the pre-trained model needs to be additionally expanded, the complexity and cost of the system will be greatly increased, and the pre-trained model cannot fully combine the specific needs and historical interaction information of users when generating answers, which may lead to the lack of pertinence of the generated answers. Summary of the Invention

[0004] The purpose of the embodiments of this application is to provide a retrieval augmented generation method and a retrieval augmented generation system. The retrieval augmented generation method can perform similarity analysis on the historical query results confirmed by the user and the direct query results, and re-rank all the direct query results according to the analysis conclusion, so as to obtain a more accurate target query result that better meets the user's needs.

[0005] In a first aspect, the present application provides a retrieval-augmented generation method, including the following steps:

[0006] S1. Split the content of the existing knowledge base into multiple basic knowledge text segments and perform semantic vectorization processing on each of them respectively to obtain the corresponding basic knowledge vectors of each basic knowledge text segment;

[0007] S2. Obtain all the historical query texts of the user and the corresponding historical query result texts, establish historical query nodes and historical query result nodes respectively, and then establish edges between the historical query nodes and the historical query result nodes according to the correspondence between the historical query texts and the historical query result texts to form a historical relationship graph;

[0008] S3. Perform semantic vectorization processing on each historical query node to obtain the corresponding historical query vectors, and then perform semantic vectorization processing on each historical query result node to obtain the corresponding historical query result vectors;

[0009] S4. Perform semantic vectorization processing on the query content directly input by the user currently to obtain a direct query vector, then calculate the vector similarity between the direct query vector and each basic knowledge vector respectively, mark the basic knowledge vectors whose vector similarity meets the first similarity standard as direct query result vectors, and then mark the vector similarity between the direct query vector and each direct query result vector as the direct matching degree;

[0010] S5. Calculate the vector similarity between the direct query vector and each historical query vector respectively, mark the historical query vectors whose vector similarity meets the second similarity standard as similar historical query vectors, and then combine the historical relationship graph to obtain the corresponding similar historical query result vectors of each similar historical query vector;

[0011] S6. Select a direct query result vector, then calculate the vector similarity between the direct query result vector and each similar historical query result vector respectively, and add the calculation results to obtain the historical matching degree; Repeat this step for each direct query result vector respectively to obtain its corresponding historical matching degree;

[0012] S7. For each direct query result vector, add the direct matching degree and the historical matching degree to obtain the overall matching degree, and organize all the overall matching degrees to form a scoring set;

[0013] S8. Select all the overall matching degrees that meet the preset matching degree standard in the scoring set and mark the corresponding direct query result vectors as target query result vectors respectively, and then add the basic knowledge text segments corresponding to the target query result vectors to the prompt words and input them into the large model to generate the reply content for presenting to the user.

[0014] In an implementable solution, it further includes step S9: The user evaluates the reply content. If the user's evaluation result is satisfactory, the query content input by the user this time is used as the historical query text, and the basic knowledge text segments corresponding to each target query result vector are used as the historical query result text. A corresponding historical query node, historical query result node, and edge are established to update the historical relationship graph, and the retrieval process ends; if the user's evaluation result is unsatisfactory, the overall matching degree corresponding to the current target query result vector is excluded from the scoring set, and then the preset matching degree standard is adjusted and steps S8 - S9 are repeated.

[0015] In an implementable solution, when calculating the vector similarity, the methods adopted at least include the cosine similarity calculation method, Euclidean distance calculation method, and Manhattan distance calculation method.

[0016] In an implementable solution, in step S7, the overall matching degree is calculated for each direct query result vector through the following formula:

[0017] M i =W c ·M di +W h ·M hi ;

[0018] where, N i is the historical matching degree, N di is the direct matching degree, M hi is the historical matching degree, W c is the weight coefficient of the direct matching degree M di is the weight coefficient of the historical matching degree M h is the weight coefficient of the historical matching degree M hi .

[0019] In an implementable solution, in steps S1 and S4, the semantic vectorization models adopted include one or more of the M3E model, BGE model, bag - of - words model, and BERT model.

[0020] In an implementable solution, in step S3, the Node2Vec algorithm is used to perform semantic vectorization processing on the historical query nodes and historical query result nodes.

[0021] In an implementable solution, in step S4, it further includes a pre - processing step for the query content input by the user, and the pre - processing step at least includes word segmentation, part - of - speech tagging, and stop - word removal.

[0022] In an implementable solution, in step S8, it further includes a post - processing step for the reply content, and the post - processing step at least includes text refinement, text polishing, and text formatting.

[0023] In an implementable solution, in the historical relationship graph, each historical query node is connected to one or more historical query result nodes, and each historical query result node is connected to one or more historical query nodes.

[0024] In a second aspect, the present application further provides a retrieval - enhanced generation system, including a vector database module, a vector generation module, a vector retrieval module, a graph database module, a vector sorting module, and a large - model reply generation module. Among them, the vector database module is used to store all basic knowledge text segments formed by segmenting the content of the existing knowledge base and the corresponding basic knowledge vectors. The vector generation module is used to perform vectorization processing on the basic knowledge text segments and the query content input by the user. The vector retrieval module is used to retrieve among all the basic knowledge vectors to obtain direct query result vectors whose vector similarity with the direct query vector meets the requirements. The graph database module is used to store the historical relationship graph formed by historical query nodes and historical query result nodes, and construct corresponding historical query vectors and historical query result vectors for the historical query nodes and historical query result nodes respectively. The vector sorting module is used to comprehensively score and sort the direct query result vectors retrieved by the vector retrieval module, and return some of the direct query result vectors with the top rankings. The large - model reply generation module is used to combine the query content input by the user and the basic knowledge text segments corresponding to the direct query result vectors returned by the vector sorting module, and generate a reply content for presenting to the customer based on the large model.

[0025] Compared with the prior art, the beneficial effects of the present application at least include:

[0026] The present application provides a retrieval - enhanced generation method, which performs similarity analysis on the historical query results that have been confirmed by the user and the direct query results, and re - sorts all the direct query results according to the analysis results, so as to obtain more accurate and more user - demand - compliant target query results, which helps to better assist users in solving their actual problems. Compared with the traditional method that needs to rely on a pre - trained model, the retrieval - enhanced generation method of the present application does not require a pre - trained model, can effectively reduce the complexity and cost of retrieval, and improve the retrieval efficiency. Further, the retrieval - enhanced generation method of the present application introduces an evaluation feedback mechanism for the user's retrieval results, so as to continuously expand and enrich the user's historical query library, provide more references for subsequent retrieval work, and thus improve the accuracy of subsequent retrieval results. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] To more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the accompanying drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present application and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0028] Figure 1 It is a schematic flowchart of a retrieval-enhanced generation method shown according to an embodiment of the present application;

[0029] Figure 2 It is a partial schematic diagram of a historical relationship diagram;

[0030] Figure 3 It is a basic structural schematic diagram of a retrieval-enhanced generation system. Detailed implementation manners

[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Usually, the components of the embodiments of the present application described and shown in the accompanying drawings here can be arranged and designed in various different configurations.

[0032] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application to be protected, but merely represents the selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0033] As Figure 1 shown, the present application provides a retrieval-enhanced generation method, including the following steps:

[0034] S1. Split the content of the existing knowledge base into multiple basic knowledge text segments and perform semantic vectorization processing on each of them to obtain the corresponding basic knowledge vectors of each basic knowledge text segment.

[0035] S2. Obtain all the historical query texts of the user and the corresponding historical query result texts, and respectively establish historical query nodes q i and historical query result nodes o i , where i is the node serial number. Then, according to the correspondence between the historical query text and the historical query result text, establish an edge between the historical query node q i and the historical query result node o i to form a historical relationship diagram, as Figure 2 shown.

[0036] S3. For each historical query node q i perform semantic vectorization processing to obtain the corresponding historical query vector, and then for each historical query result node o i perform semantic vectorization processing to obtain the corresponding historical query result vector.

[0037] S4. Perform semantic vectorization processing on the query content directly input by the user currently to obtain a direct query vector, and then calculate the vector similarity between the direct query vector and each basic knowledge vector respectively. Mark the basic knowledge vectors that meet the first similarity standard as direct query result vectors, and then mark the vector similarities between the direct query vector and each direct query result vector as direct matching degrees.

[0038] S5. Calculate the vector similarity between the direct query vector and each historical query vector respectively. Mark the historical query vectors that meet the second similarity standard as similar historical query vectors, and then combine with the historical relationship graph to obtain the corresponding similar historical query result vectors for each similar historical query vector.

[0039] S6. Select a direct query result vector, and then calculate the vector similarity between the direct query result vector and each similar historical query result vector respectively, and superimpose the calculation results to obtain a historical matching degree; repeat this step for each direct query result vector respectively to obtain its corresponding historical matching degree.

[0040] S7. For each direct query result vector, superimpose the direct matching degree and the historical matching degree to obtain an overall matching degree, and organize all the overall matching degrees to form a scoring set.

[0041] S8. Select all the overall matching degrees that meet the preset matching degree standard in the scoring set and mark the corresponding direct query result vectors as target query result vectors respectively, and then add the corresponding basic knowledge text segments of the target query result vectors to the prompt words and input them into the large model to generate a reply content for presenting to the user. Usually, the large model will organize and appropriately interpret each prompt word, and optimize the language expression method to make the reply content more convenient for reading and understanding.

[0042] S9. The user evaluates the reply content. If the user evaluation result is satisfied, use the query content input by the user this time as the historical query text, and use the corresponding basic knowledge text segments of each target query result vector as the historical query result text, and establish the corresponding historical query node q i and historical query result node o iSum the edges to update the historical relationship graph and end the retrieval process. If the user evaluation result is unsatisfactory, exclude the overall matching degree corresponding to the current target query result vector from the scoring set, and then repeat steps S8 - S9 after adjusting the preset matching degree criterion. Specifically, when the user evaluation result is satisfactory, a historical query node q is newly created according to the direct query vector and each target query result vector in this retrieval i and a historical query result node o i .

[0043] The retrieval enhancement generation method of the present application performs similarity analysis on the historical query results and direct query results that have been confirmed by the user, and re - ranks all direct query results according to the analysis conclusion, so as to obtain more accurate target query results that better meet the user's needs, which helps to better assist the user in solving their actual problems. Compared with the traditional method that needs to rely on a pre - trained model, the retrieval enhancement generation method of the present application does not require a pre - trained model, which can effectively reduce the complexity and cost of retrieval and improve the retrieval efficiency. Further, the retrieval enhancement generation method of the present application introduces a user evaluation feedback mechanism for retrieval results, so that the user's historical query library can be continuously expanded and enriched, providing more references for subsequent retrieval work, thereby improving the accuracy of subsequent retrieval results.

[0044] In one embodiment, when calculating the vector similarity, the methods adopted at least include the cosine similarity calculation method, the Euclidean distance calculation method, and the Manhattan distance calculation method. The cosine similarity calculation method only focuses on the direction and is not affected by the vector length, which is suitable for high - dimensional data such as text vectors; the Euclidean distance calculation method calculates the straight - line distance between two vectors in a multi - dimensional space, and the smaller the distance, the more similar the two vectors are; the Manhattan distance calculation method calculates the total distance along the coordinate axes of two vectors in a multi - dimensional space, and the smaller the distance, the more similar the two vectors are. When calculating the vector similarity between the direct query vector and each basic knowledge vector, the similarity in direction of different vectors is mainly considered, so it is preferably to use the cosine similarity calculation method. Specifically, the vector similarity between the direct query vector V q and the basic knowledge vector V ki can be calculated by the following formula:

[0045]

[0046] Among them, S i is the vector similarity, V q is the direct query vector, V kiLet the basic knowledge vector be \(V_0\), and \(i\) be the label. If the calculated vector similarity reaches the first similarity standard, it can be considered that the corresponding basic knowledge vector meets the requirements, that is, the basic knowledge text has a large semantic association with the query content directly input by the user currently, and may be helpful for solving the user's problem. At this time, the aforementioned basic knowledge vector can be marked as the direct query result vector, and the corresponding vector similarity can be marked as the direct matching degree. Generally, when setting the first similarity standard, it can be taken at fixed intervals within a certain range (for example, 0.6 - 1) according to engineering experience, and then by testing the overall performance of some labeled data under different first similarity standard values, the actual first similarity standard value to be adopted is determined. Similarly, it is also possible to preferably use the cosine similarity calculation method to calculate and screen the vector similarity between the direct query vector and the historical query vector to obtain similar historical query vectors, and for each direct query result vector, calculate the vector similarity with all similar historical query result vectors respectively to obtain the corresponding historical matching degree for each direct query result vector.

[0047] In one embodiment, in step S7, the overall matching degree \(M\) can be calculated for each direct query result vector through the following formula i :

[0048] \(M\) i =\(W_1\) c \(\cdot M_1\) di +\(W_2\) h \(\cdot M_2\) hi ;

[0049] where, \(M_1\) i is the historical matching degree, \(M_2\) di is the direct matching degree, \(M_3\) hi is the historical matching degree, \(W_1\) c is the weight coefficient of the direct matching degree \(M_2\) di , \(W_2\) h is the weight coefficient of the historical matching degree \(M_3\) hi , and \(i\) is the label. In different application scenarios, the retrieval effect can be optimized by adjusting the weight coefficients \(W_1\) c and \(W_2\) h . For example, if the content of the user's historical query library is very small or of little reference value, the value of \(W_1\) h can be appropriately reduced and the value of \(W_2\) c can be increased. As the retrieval process continues, the user's historical query library will be expanded and enriched. At this time, the value of \(W_2\) h can be appropriately increased.

[0050] In one embodiment, in steps S1 and S4, the semantic vectorization model adopted includes one or more of the M3E model, the BGE model, the bag-of-words model, and the BERT model. Among them, the M3E model (Moka Massive Mixed Embedding) is a vectorization model focusing on Chinese text processing, supporting homogeneous text similarity calculation and heterogeneous text retrieval for Chinese-English bilingual texts. It uses large-scale mixed embedding technology to improve the expression ability and generalization ability of word vectors, and is suitable for private deployment and resource-constrained scenarios. The BGE model (BAAI General Embedding) is a multilingual text embedding model, which supports many languages, can efficiently implement retrieval tasks at different granularities, and performs well in Chinese-English semantic retrieval accuracy and overall semantic representation ability. The bag-of-words model is a classic text representation method that represents text as a word frequency vector, ignoring the order and grammatical structure of words. Its implementation method is simple and intuitive, and it can be combined with various machine learning algorithms and is suitable for various text analysis tasks. The BERT model (Bidirectional Encoder Representations from Transformers) is a bidirectional language model based on the Transformer architecture, which can capture the deep semantic information of text and performs well in a variety of natural language processing tasks. The specific model selection can be determined according to requirements and conditions, and there is no limitation here.

[0051] In one embodiment, in step S3, the Node2Vec algorithm is used to perform semantic vectorization processing on the historical query nodes and the historical query result nodes. The Node2Vec algorithm is a graph embedding algorithm based on random walk and Word2Vec. It generates the context sequence of nodes by simulating random walks, and then uses these sequences to train the low-dimensional vector representation of nodes. Specifically, after forming the historical relationship graph, the Node2Vec algorithm performs random walks on each node in the historical relationship graph to generate node sequences, and then trains these sequences to obtain the low-dimensional vector representation of each node. The Node2Vec algorithm has many advantages. For example, it can achieve a good balance between depth-first search (DFS) and breadth-first search (BFS), so it can capture the local and global structural information of nodes and has good flexibility; the node embedding vectors generated by the Node2Vec algorithm can intuitively reflect the similarity between nodes, which helps to understand the structural characteristics of the network and facilitates researchers and developers to better interpret the output of the model; in addition, Node2Vec is also applicable to various types of graphs (including undirected graphs, directed graphs, and weighted graphs), and only depends on the structural information of the graph without the need for node attribute information, and has wide applicability in many practical application scenarios.

[0052] In one embodiment, the preprocessing step of the query content input by the user may further be included in step S4. The preprocessing step at least includes word segmentation, part-of-speech tagging, and stop word removal. Among them, word segmentation is the basic step of natural language processing, which can convert continuous text into discrete lexical units for subsequent processing. Part-of-speech tagging tags the part of speech of each lexical unit (such as noun, verb, adjective, etc.), which helps to understand the grammatical role of the word in the sentence and thus better process semantic information. Stop word removal can remove common words with little semantic contribution in the query text (such as "de", "shi", "he", etc.), thereby reducing noise words and improving the efficiency and accuracy of text processing.

[0053] In one embodiment, in step S8, the postprocessing step of the reply content may further be included. The postprocessing step at least includes text extraction, text polishing, and text formatting. Among them, text extraction can extract key information and core content from the generated reply content, remove redundant parts, and thus generate a concise and clear reply for the user to quickly obtain key information. Text polishing optimizes the generated reply content linguistically to make it more fluent and natural, thereby improving the readability and professionalism of the reply and enhancing the user experience. Text formatting adjusts the format of the generated reply content to meet specific output requirements (such as paragraph format, punctuation, etc.), thereby ensuring the unified and standardized format of the reply content for easy user reading.

[0054] In one embodiment, as Figure 2 shown, in the historical relationship graph, each historical query node can be connected to one or more historical query result nodes, and each historical query result node can be connected to one or more historical query nodes. This is because in the reply content for a historical query, multiple basic knowledge text segments may be involved. For example, the user's previous query content was "main engine oil leakage", and among the retrieved historical query results, the following three basic knowledge text segments were involved, namely: Text segment 1 "In common problems of ship main engine oil leakage, an aging sealing ring is a common cause"; Text segment 2 "If there are cracks in the engine components, it will also cause oil leakage"; Text segment 3 "Main engine oil leakage may be due to improper installation or damage of the oil filter". Each historical query result node corresponds to a basic knowledge text segment, so each historical query node can be connected to one or more historical query result nodes. Similarly, the retrieval results of multiple different historical queries may involve the same basic knowledge text segment. For example, the historical query text segments "main engine oil leakage" and "hazards of engine cracks" may both retrieve the basic knowledge text segment "If there are cracks in the engine components, it will also cause oil leakage", so each historical query result node can be connected to one or more historical query nodes.

[0055] The present application also provides a retrieval enhanced generation system, including a vector database module, a vector generation module, a vector retrieval module, a graph database module, a vector ranking module, and a large model response generation module. As Figure 3 shown, the vector generation module is connected to the vector database module, the vector database module and the vector generation module are jointly connected to the vector retrieval module, the vector retrieval module and the graph database module are jointly connected to the vector rearrangement module, the vector rearrangement module is connected to the large model response generation module, and the large model response generation module is connected to the graph database module.

[0056] Among them, the vector database module is used to store all basic knowledge text segments formed by segmenting the content of the existing knowledge base and the corresponding basic knowledge vectors. The vector generation module is used to perform vectorization processing on the basic knowledge text segments and the query content input by the user. The vector retrieval module is used to retrieve among all the basic knowledge vectors to obtain a direct query result vector whose vector similarity with the direct query vector meets the requirements. The graph database module is used to store the historical relationship graph formed by the historical query nodes and the historical query result nodes, and construct corresponding historical query vectors and historical query result vectors for the historical query nodes and the historical query result nodes respectively; after each retrieval, the graph database module can sort out the retrieval results confirmed by the user, and then create corresponding historical query nodes, historical query result nodes and new edges in the historical relationship graph, so as to update the historical relationship graph. The vector ranking module is used to comprehensively score and rank the direct query result vectors retrieved by the vector retrieval module, and return some of the direct query result vectors with the top rankings. The large model response generation module is used to combine the query content input by the user and the basic knowledge text segments corresponding to the direct query result vectors returned by the vector ranking module, and generate a response content for presenting to the customer based on the large model.

[0057] The foregoing are only the preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

1. A retrieval-augmented generation method, characterized in that, Including: S1. Split the content of the existing knowledge base into multiple basic knowledge text segments, perform semantic vectorization processing on each of them respectively, and obtain the corresponding basic knowledge vectors for each basic knowledge text segment; S2. Obtain all the historical query texts of the user and the corresponding historical query result texts, establish historical query nodes and historical query result nodes respectively, and then, according to the corresponding relationship between the historical query texts and the historical query result texts, establish edges between the historical query nodes and the historical query result nodes to form a historical relationship graph; S3. Perform semantic vectorization processing on each historical query node to obtain the corresponding historical query vector, and then perform semantic vectorization processing on each historical query result node to obtain the corresponding historical query result vector; S4. Perform semantic vectorization processing on the query content directly input by the user currently to obtain a direct query vector, then calculate the vector similarity between the direct query vector and each basic knowledge vector respectively, mark the basic knowledge vectors whose vector similarity meets the first similarity standard as direct query result vectors, and then mark the vector similarity between the direct query vector and each direct query result vector as the direct matching degree; S5. Calculate the vector similarity between the direct query vector and each historical query vector respectively, mark the historical query vectors whose vector similarity meets the second similarity standard as similar historical query vectors, and then, in combination with the historical relationship graph, obtain the corresponding similar historical query result vectors for each similar historical query vector; S6. Select a direct query result vector, then calculate the vector similarity between the direct query result vector and each similar historical query result vector respectively, and superimpose the calculation results to obtain the historical matching degree; Repeat this step for each direct query result vector respectively to obtain its corresponding historical matching degree; S7. For each direct query result vector, superimpose the direct matching degree and the historical matching degree to obtain the overall matching degree, and organize all the overall matching degrees to form a scoring set; S8. Select all the overall matching degrees that meet the preset matching degree standard in the scoring set, mark the corresponding direct query result vectors as target query result vectors respectively, then add the corresponding basic knowledge text segments of the target query result vectors to the prompt words and input them into the large model to generate the reply content for presenting to the user.

2. The retrieval enhancement generation method according to claim 1, wherein Also including: S9. The user evaluates the reply content. If the user evaluation result is satisfactory, use the query content input by the user this time as the historical query text, and use the corresponding basic knowledge text segments of each target query result vector as the historical query result text, establish the corresponding historical query nodes, historical query result nodes and edges to update the historical relationship graph, and end the retrieval process; If the user evaluation result is unsatisfactory, exclude the overall matching degree corresponding to the current target query result vector in the scoring set, and then repeat steps S8 - S9 after adjusting the preset matching degree standard.

3. The retrieval-enhanced generation method according to claim 1, wherein When calculating the vector similarity, the methods adopted at least include the cosine similarity calculation method, the Euclidean distance calculation method and the Manhattan distance calculation method.

4. The retrieval enhanced generation method according to claim 1, wherein In step S7, the overall matching degree is calculated for each direct query result vector by the following formula: M i = W c · M di + W h * M hii ; Among them, M i is the historical matching degree, M di is the direct matching degree, M hi is the historical matching degree, W c is the direct matching degree M di 's weight coefficient, W h is the weight coefficient of the historical matching degree M hi .

5. The retrieval enhanced generation method according to claim 1, wherein In steps S1 and S4, the semantic vectorization models adopted include one or more of the M3E model, BGE model, bag-of-words model, and BERT model.

6. The retrieval enhanced generation method according to claim 1, wherein In step S3, the Node2Vec algorithm is used to perform semantic vectorization processing on the historical query nodes and historical query result nodes.

7. The retrieval enhancement generation method according to claim 1, wherein In step S4, it also includes a preprocessing step for the query content input by the user, and the preprocessing step at least includes word segmentation, part-of-speech tagging, and stop word removal.

8. The retrieval enhanced generation method according to claim 1, wherein In step S8, it also includes a postprocessing step for the reply content, and the postprocessing step at least includes text refinement, text polishing, and text formatting.

9. The retrieval enhanced generation method according to claim 1, wherein In the historical relationship graph, each historical query node is connected to one or more historical query result nodes, and each historical query result node is connected to one or more historical query nodes.

10. A retrieval-augmented generation system, characterized in that, It includes: A vector database module for storing all basic knowledge text segments formed by splitting the content of the existing knowledge base and the corresponding basic knowledge vectors; A vector generation module for performing vectorization processing on the basic knowledge text segments and the query content input by the user; A vector retrieval module for retrieving among all basic knowledge vectors to obtain direct query result vectors whose vector similarity with the direct query vector meets the requirements; A graph database module for storing the historical relationship graph formed by historical query nodes and historical query result nodes, and constructing corresponding historical query vectors and historical query result vectors for the historical query nodes and historical query result nodes respectively; A vector sorting module for comprehensively scoring and sorting the direct query result vectors retrieved by the vector retrieval module, and returning some of the direct query result vectors with the top rankings; A large model reply generation module for combining the query content input by the user and the basic knowledge text segments corresponding to the direct query result vectors returned by the vector sorting module, and generating a reply content for presenting to the customer based on the large model.

Citation Information

Cited By

  • Query result obtaining method and device based on large model and medium

    CN121388150A

  • A large model-based query result acquisition method, device and medium

    CN121388150B