Retrieval problem optimization method, medium and system based on multi-layer knowledge base
By building a multi-layer knowledge base and using two-round query rewriting methods to optimize user queries, the problem that RAG system is difficult to accurately understand user intentions when processing natural language queries is solved, and more accurate and diversified content is achieved, providing a more efficient solution for complex tasks.
Patent Information
- Application Number
- CN202510082760.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-01-20
AI Technical Summary
The existing search-enhanced large language model (RAG) system is difficult to accurately understand user intentions when processing user natural language queries, resulting in the search results being inaccurate enough to meet user needs, especially in complex tasks.
By building a multi-layer knowledge base and adopting two-round query rewriting methods, user queries are optimized. First, based on user queries, the initial vector block set is generated in the multi-layer knowledge base, and the query is rewrite for the first time; then, the update vector block set is generated again, and the source statistical results of the initial and updated vector block sets are determined, and the second query is rewrite for the second query, and the query is rewrite for outputting more accurate query problems.
It significantly reduces the ambiguity of the query, enhances the accuracy and diversity of generated content, and provides a more efficient and accurate solution for the application of RAG systems in complex tasks.
Smart Images

Figure CN119494411B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of large language model processing, and in particular to a retrieval problem optimization method, medium and system based on a multi-layer knowledge base. Background Art
[0002] In the application of Retrieval-Augmented Generation (RAG), users’ question-and-answer interaction through natural language has become the mainstream. However, due to the colloquial expression, vague description and ambiguous contextual references of user queries, RAG systems often find it difficult to accurately understand user intent, resulting in inaccurate retrieval results and inability to generate answers that meet user needs. This is particularly true in complex tasks, such as scenarios that require multiple rounds of information integration, deep analysis or domain expertise. To address this problem, query rewriting technology has been widely used, aiming to eliminate ambiguity and clarify information needs by optimizing and refining user queries. However, existing methods usually only perform a single query rewrite, which cannot fully capture the multi-level characteristics of user semantics, and it is difficult to balance the professionalism and comprehensiveness of query questions in multiple dimensions.
[0003] Therefore, how to optimize the diversity and professionalism of answers and improve the relevance of search results through multiple rounds of precise query rewriting has become a key challenge that needs to be solved urgently. Summary of the invention
[0004] In order to solve at least one of the above technical problems, an embodiment of the present application provides a retrieval problem optimization method based on a multi-layer knowledge base, the method comprising:
[0005] S1: Manage knowledge data in a hierarchical manner and build a multi-layer knowledge base;
[0006] S2: According to the user query, the multi-layer knowledge base is retrieved to generate an initial vector block set; then the query is rewritten for the first time, and the multi-layer knowledge base is retrieved again to generate an updated vector block set;
[0007] S3: Determine the strong and weak directions of the retrieval according to the source statistics of the initial vector block set and the updated vector block set in the multi-layer data block;
[0008] S4: According to the strong and weak directions of the retrieval, the query is rewritten for the second time and the query question is output.
[0009] Further, step S3 includes:
[0010] S31: Determine the source position of each vector block in the initial vector block set and the updated vector block set in the multi-layer knowledge base;
[0011] S32: Counting the number of vector blocks included in each source position;
[0012] S33: According to the statistical results, sort and determine the strong and weak directions of the retrieval.
[0013] Further, step S33 includes:
[0014] S331: determining the position coefficient at each source position according to the number of vector blocks included at each source position;
[0015] S332: If the position coefficient is greater than the first set threshold or the position coefficient is ranked in the first set position, it is determined as a strong direction of the search; otherwise, it is determined as a weak direction.
[0016] Furthermore, step S332 further includes:
[0017] Among the source positions determined to be weak directions, determine whether their position coefficients are less than the second set threshold; or whether their rankings are greater than the second set rank, if so, screen out the corresponding source positions; the second set threshold is less than the first set threshold; the second set rank is greater than the first set rank.
[0018] Furthermore, step S31 further includes:
[0019] S31a: comparing the vector blocks in the initial vector block set and the vector blocks in the updated vector block set to determine the same vector blocks and different vector blocks;
[0020] S32b: configure a higher number coefficient for the same vector block than for different vector blocks;
[0021] Step S32 specifically includes: according to the quantity coefficient, counting the number of vector blocks included in each source position according to the weight.
[0022] Further, step S331 includes:
[0023] S3311: in the multi-layer knowledge base, determining the position coefficient of each source position of the first level according to the number of vector blocks included in each source position of the first level;
[0024] S3312: For each source position except the first level, determine the position coefficient of each source position except the first level according to the number of vector blocks included in the source position and the position coefficients of all its upper levels.
[0025] Furthermore, step S3312 further includes:
[0026] S33121: for each source position except the first level, count the total number of vector blocks included in the same level;
[0027] S33122: determining a level coefficient at the source position according to a ratio of the number of vector blocks included at the source position to the total number of vector blocks included at the same level;
[0028] S33123: Determine the position coefficients of each source position except the first level according to the number of vector blocks included in the source position and the position coefficients of all upper levels, as well as the level coefficients of the source position.
[0029] Furthermore, it also includes:
[0030] P1: Use any of the above retrieval problem optimization methods to optimize the query problem;
[0031] P2: Determine the weight of the multi-layer knowledge base according to the strong and weak directions of the retrieval determined in any of the above retrieval problem optimization methods;
[0032] P3: Generate the final vector block set based on the optimized query question and the multi-layer knowledge base with weight settings.
[0033] In a second aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, any of the above methods is implemented.
[0034] In a third aspect, an embodiment of the present application provides an electronic system, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements any of the above methods when executing the computer program.
[0035] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0036] The present invention provides a retrieval problem optimization method, medium and system based on a multi-layer knowledge base. The key is to identify repeated and independent knowledge bases, determine the strong and weak directions of the retrieval, perform a second query rewrite on the query, and output a more accurate query problem by counting the initial vector block sets and update vector block sets of two retrievals. On the one hand, for the strong direction determined by the retrieval, see which knowledge base the retrieval results are repeated from, not all knowledge bases are treated uniformly, and the professionalism of the query problem can be taken into account; on the other hand, for the weak direction determined by the retrieval, see which knowledge base the retrieval results are independently from, instead of directly discarding them, and the comprehensiveness of the query problem can be taken into account; therefore, the strong and weak directions of the retrieval are positioned, and the focus of the retrieval and the weight of the knowledge base are preferably adjusted, so that query problems with both professionalism and comprehensiveness can be generated. Through two rounds of optimization, this method significantly reduces the ambiguity of the query, enhances the accuracy and diversity of the generated content, and provides a more efficient and accurate solution for the application of the RAG system in complex tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0038] Figure 1 A flowchart of an embodiment of the search question optimization method provided by the present application;
[0039] Figure 2 A flowchart of another embodiment of the search question optimization method provided by the present application;
[0040] Figure 3 A schematic diagram of an embodiment of a multi-layer knowledge base provided by the present application;
[0041] Figure 4 A schematic diagram of the structure of an embodiment of the electronic system provided in the present application. DETAILED DESCRIPTION
[0042] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0043] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0044] It should also be understood that the term “and / or” used in the specification and appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0045] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0046] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0047] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "include", "including", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0048] For ease of understanding, the technical solution of the present application will be described in detail below with reference to the accompanying drawings.
[0049] The present invention provides a retrieval problem optimization method based on a multi-layer knowledge base, such as Figure 1-2 For the sake of convenience, only the parts related to this embodiment are shown. The method provided in this embodiment includes the following steps:
[0050] S1: Manage knowledge data in a hierarchical manner and build a multi-layer knowledge base;
[0051] Specifically, it is possible to manage knowledge data in a hierarchical manner through scientific classification to ensure the independence and clarity of data sources, and divide the knowledge base into multiple levels, such as agriculture, medical care, education and other fields, each of which includes independent data sources to build a multi-layer knowledge base. For example, it is possible to divide it into 3 levels:
[0052] Level 1: Coarse classification: Main fields, i.e. major categories, such as "agriculture" or "medical care";
[0053] Level 2: Segmentation: sub-fields, such as "crop diseases" and "fertilization" for example, and "medical care" for example, such as "medical diagnosis";
[0054] Level 3: Specific: topics or issues, such as "crop diseases", such as "types of rice diseases"; taking "medical care" as an example, such as "diabetes diagnosis methods".
[0055] It is worth noting that the number of levels and content of the specific levels can be arbitrarily set according to the amount of data contained in the knowledge base, the scope, the accuracy requirements of the problem index, etc., and are not limited to this example; and each level may or may not have a next level. For example, the types of rice diseases can also be divided into a next level; and the diabetes diagnosis method does not include a next level; that is, the relationship between the levels can be set symmetrically or asymmetrically, depending on its content and the classification of the knowledge base. More specifically, it is optional to store knowledge in a layered manner based on predefined tags to achieve data grading and index construction. The specific steps can optionally include:
[0056] S11: Clustering algorithm: Automatically classify unstructured text using topic modeling method (LDA);
[0057] S12: Index construction: Elasticsearch can be used to build a reverse index to support efficient retrieval;
[0058] S13: Dynamic update: Use the real-time data update mechanism to regularly synchronize new data sources and rebuild indexes.
[0059] S2: According to the user query, the multi-layer knowledge base is retrieved to generate an initial vector block set; then the query is rewritten for the first time, and the multi-layer knowledge base is retrieved again to generate an updated vector block set;
[0060] Specifically, after the user raises a query, the multi-layer knowledge base established based on the retrieval-augmented generation (RAG) large language model can be used to generate an initial vector block set A, and then the query can be rewritten for the first time using a small language model, zero-shot learning prompt engineering technology, etc. to preliminarily optimize the query, reduce colloquial expressions and ambiguity, and the established multi-layer knowledge base can be retrieved again to generate an updated vector block set B.
[0061] More specifically, the algorithm for the first rewrite can be:
[0062] Input: User’s original query Q.
[0063] Optimize rules: Design prompts and rewrite queries for specific tasks.
[0064] Prompt example:
[0065] The following user queries were optimized to reduce ambiguity and more clearly express the information need: {Q}
[0066] Output: The optimized query Q' has a clearer structure and more explicit semantics.
[0067] Algorithm pseudo code:
[0068] from transformers import pipeline
[0069] model = pipeline('text2text-generation', model='t5-small')
[0070] query = "What are the diseases of rice?"
[0071] prompt = f"Optimize the following query to reduce ambiguity: {query}"
[0072] optimized_query = model(prompt)[0]['generated_text']
[0073] Generation and comparison of vector block sets A and B:
[0074] 1. Method description
[0075] After the first rewrite, the original query and the optimized query are retrieved to generate the initial vector block set A and the updated vector block set B. By comparing the knowledge base sources and answer contents of the two, the duplicate and independent sources are determined.
[0076] 2. Algorithm Implementation
[0077] Generate an initial vector block set A and an updated vector block set B, and use vector retrieval technology to retrieve the most relevant knowledge base source to save the score and answer content.
[0078] Algorithm pseudo code:
[0079] from sentence_transformers import SentenceTransformer, util
[0080] model = SentenceTransformer('all-MiniLM-L6-v2')
[0081] embeddings_A = {entry["source"]: model.encode(entry["content"]) for entry in data_A}
[0082] embeddings_B = {entry["source"]: model.encode(entry["content"]) for entry in data_B}
[0083] repeated_sources = []
[0084] for source_A, emb_A in embeddings_A.items():
[0085] for source_B, emb_B in embeddings_B.items():
[0086] if util.cos_sim(emb_A, emb_B)>0.85:
[0087] repeated_sources.append((source_A, source_B))
[0088] S3: Determine the strong and weak directions of the retrieval according to the source statistics of the initial vector block set and the updated vector block set.
[0089] Specifically, a statistical comparison can be performed on the two generated results to determine which field direction the current query is more inclined to. The field with more relevant documents is the strong direction, and the field with fewer relevant documents is the weak direction.
[0090] Preferably, step S3 may optionally include but is not limited to:
[0091] S31: Determine the source positions of the vector blocks in the initial vector block set and the updated vector block set in the multi-layer knowledge base;
[0092] S32: Counting the number of vector blocks included in each source position;
[0093] S33: According to the statistical results, sort and determine the strong and weak directions of the retrieval.
[0094] Specifically, the specific source position of each vector block element in the initial vector block set and the updated vector block set in the multi-layer knowledge base constructed in step S1 is identified to identify repeated and independent knowledge base sources and determine the strong and weak directions of the retrieval.
[0095] Taking the above-mentioned fields of agriculture, medical treatment, education, etc. as examples, suppose that there are K search results in each of the initial vector block set A and the updated vector block set B, that is, the initial vector block set A includes relevant documents related to the query question ; Update vector block set B to include relevant documents related to the query question , then determine the position of each relevant document in the multi-layer data block and count the number of vector blocks included at each position. More specifically, you can choose to build a tree structure for clear display, such as Figure 3 As shown, taking K=13 as an example, assuming there are 13 relevant documents, in the first-level large field, 8 documents come from agriculture, 3 documents come from medical care, and 2 documents come from education; in the second-level sub-field, among the 8 documents coming from agriculture, 5 documents come from crop diseases and 3 documents come from fertilization; among the 3 documents coming from medical care, 3 all come from diabetes diagnosis; among the 2 documents coming from education, 1 comes from adult education and 1 comes from children's education; in the third-level sub-field, among the 5 documents coming from crop diseases, 4 come from rice disease types and 1 comes from cabbage disease types; among the 3 documents coming from fertilization, 3 all come from fertilization time; among the 3 documents coming from medical diagnosis, 3 all come from diabetes diagnosis; in the fourth level, among the 3 documents coming from rice diseases, 3 come from insect pests and 1 comes from drought. It is worth noting that in actual retrieval generation, the number of relevant documents K retrieved by the target may be tens of thousands or even more, and the differences in their fields may not be very large. Most of them are concentrated in a large field or even a specific problem. Depending on the question asked and the related prompts, they may be inclined to a certain field or a certain problem. Figure 3 The equilibrium shown, Figure 3 This is for illustrative purposes only and is not intended to be limiting.
[0096] In addition, the source position of each vector block in the updated vector block set B can also be determined; according to the source positions of the 2K vector blocks in the two vector block sets, statistics and sorting are performed, and the strong and weak directions of the retrieval are determined from top to bottom. The strong direction is determined by looking at which knowledge bases the repeated and identical ones come from, and the weak direction is determined by which knowledge bases the scattered and independent ones come from.
[0097] Preferably, step S33 may optionally include but is not limited to:
[0098] S331: determining the position coefficient at each source position according to the number of vector blocks included at each source position;
[0099] S332: If the position coefficient is greater than the first set threshold or the position coefficient is ranked in the first set position, it is determined as a strong direction of the search; otherwise, it is determined as a weak direction.
[0100] In this embodiment, step S33 is given, which is a preferred embodiment of how to sort and determine the strong and weak directions of the retrieval. Of course, there are many other ways to sort and determine the strong and weak directions. As long as the sorting is based on the statistical results, any method of determining the strong and weak directions is acceptable. It is not necessary to calculate the position coefficient of each position, and it is also possible to sort directly according to the number of statistical vector blocks.
[0101] More preferably, step S332 may also optionally include but is not limited to:
[0102] Among the source positions determined to be weak directions, determine whether their position coefficients are less than the second set threshold; or whether their rankings are greater than the second set rank, if so, screen out the corresponding source positions; the second set threshold is less than the first set threshold; the second set rank is greater than the first set rank.
[0103] In this embodiment, a further preferred embodiment of step S332 is given, which filters out directions with particularly small position coefficients or particularly low rankings, and can delete weak directions with particularly few relevant documents, thereby avoiding some related documents that are mistakenly entered but are actually irrelevant, thereby further improving the efficiency and accuracy of subsequent retrieval generation. Such particularly weak directions will not be included in the database that needs to be retrieved subsequently.
[0104] More preferably, step S31 further includes:
[0105] S31a: comparing the vector blocks in the initial vector block set and the vector blocks in the updated vector block set to determine the same vector blocks and different vector blocks;
[0106] S32b: configure a higher number coefficient for the same vector block than for different vector blocks;
[0107] Step S32 specifically includes: according to the quantity coefficient, counting the number of vector blocks included in each source position according to the weight.
[0108] In this embodiment, a preferred embodiment of steps S31 and S32 is provided. The example first compares the relevant documents in the initial vector block set A and the updated vector block set B. , , check whether there are identical vector blocks, and configure a higher number coefficient for them than for different vector blocks. That is to say, when counting the number of identical vector blocks, they have a higher weight than different vector blocks. For example, the number coefficient of identical vector blocks is 1, and the number coefficient of different vector blocks is 0.5. Figure 3In the example of crop disease statistics, 5 relevant documents are counted. If, when compared with the updated vector block set B, 2 vector blocks are the same and 3 vector blocks are different, then after the quantity coefficient is updated according to the weight during statistics, the relevant documents of crop disease statistics should be updated to 2✖1+3✖0.5=3.5. Based on this preferred embodiment, the retrieval efficiency and accuracy can be more effectively improved, because if a relevant document is repeatedly hit in two retrievals, it must be a document with very high relevance. When counting the number, a higher quantity coefficient is configured for it, which can more fully reflect the relevance of the relevant documents.
[0109] More preferably, step S331 may optionally include but is not limited to:
[0110] S3311: in the multi-layer knowledge base, determining the position coefficient of each source position of the first level according to the number of vector blocks included in each source position of the first level;
[0111] S3312: For each source position except the first level, determine the position coefficient of each source position except the first level according to the number of vector blocks included in the source position and the position coefficients of all its upper levels.
[0112] In this embodiment, a preferred embodiment of step S331 is provided, and the example is also as follows: Figure 3 Taking the example of agriculture as an example, for each source position of the first level, the position coefficient of the position can be directly determined according to the number of vector blocks included in the position, and the position coefficient can be represented by 8 or 8 / 12. For each source position other than the first level, such as rice disease, the position coefficient of the position can be determined according to the number of vector blocks included in the source position and the position coefficients of all its upper levels, and the position coefficient can be represented by 8+5+4=17 or 8 / 12+5 / 12+4 / 12.
[0113] In this embodiment, a preferred embodiment of step S331 is provided. For each source position except the first level, the number of vector blocks counted at its own position and the number of vector blocks counted at all its upper levels are comprehensively considered, and the position coefficient at the source position is sorted to determine the strong and weak direction of the retrieval, which can more accurately express the strength of each source position. For example, Figure 3 As shown in the figure, the number of vector blocks at the three source positions of rice disease, fertilization time, and diabetes diagnosis is 3. It is assumed that the number of vector blocks at the upper-level medical field, medical diagnosis sub-field, and lower-level pests are also 3. If only the number of vector blocks at their own positions is considered, their strengths are the same, but from Figure 3From the example shown, the strength of the agricultural field is obviously stronger than that of the medical field. According to steps S3311-S3321, the number of vector blocks at its own position and the number of vector blocks at all its upper levels can be comprehensively considered. Then, the position coefficient of rice disease is 17, the position coefficient of fertilization time is 14, and the position coefficient of diabetes diagnosis is 9. The direction of strength can be determined as rice disease>fertilization time>diabetes diagnosis. The determination of the direction of strength is obviously more accurate, especially for thousands of big data. The advantage is obvious.
[0114] More preferably, step S3312 may also optionally include but is not limited to:
[0115] S33121: for each source position except the first level, count the total number of vector blocks included in the same level;
[0116] S33122: determining a level coefficient at the source position according to a ratio of the number of vector blocks included at the source position to the total number of vector blocks included at the same level;
[0117] S33123: Determine the position coefficients of each source position except the first level according to the number of vector blocks included in the source position and the position coefficients of all upper levels, as well as the level coefficients of the source position.
[0118] In this embodiment, a preferred embodiment of step S3312 is provided. For each source position except the first level, the number of vector blocks included in the source position and the position coefficients of all its upper levels, as well as the level coefficient of the source position, are comprehensively considered, and then the strong and weak directions of the retrieval are determined by sorting, which can further accurately express the strength of each source position. For example, Figure 3 As shown in FIG. 1 , the education field includes 2 documents, of which there is 1 document for adult education and 1 document for children's education in the second level, and there is no third level. Therefore, when counting the position coefficients of each source position of the third level, if the position coefficients are still calculated based on the total number of vector blocks of 12, it is obviously unreasonable. In order to more accurately characterize the position coefficients of each position, the present invention re-counts the sum of the number of vector blocks included in the same level for all levels except the first level. According to steps S33121-S33123, the position coefficient of rice disease should be changed from 8+5+4=17 to , which can effectively avoid the situation where a certain level has no sub-levels, further improve the representation accuracy of the position coefficient, and improve the accuracy of retrieval generation. Similarly, for thousands of large data, the advantages are even more obvious.
[0119] S4: According to the strong and weak directions of the retrieval, the query is rewritten for the second time and the query question is output.
[0120] Specifically, similar to step S2, based on the determined strong and weak directions of the retrieval, a small language model, a zero-sample prompting project, etc. may be used to rewrite the query for the second time to further output a more accurate query question.
[0121] 1. Method description
[0122] Based on the weighted knowledge base, the generated query questions are focused on high-weight sources, while retaining information from low-weight sources
[0123] 2. Algorithm Implementation
[0124] Input: Optimized query Q''.
[0125] Optimization rules: Design prompts to limit the scope
[0126] Generate professional answers from high-weight knowledge bases and integrate low-weight sources for the following queries: {Q''}
[0127] Output: Final optimized query Q'''
[0128] pseudocode:
[0129] prompt = f"Optimize the following query based on the source of the knowledge base: {optimized_query}"
[0130] final_query = model(prompt)[0]['generated_text']
[0131] The present invention provides a retrieval problem optimization method based on a multi-layer knowledge base, the key of which is to identify repeated and independent knowledge bases, determine the strong and weak directions of the retrieval, perform a second query rewrite on the query, and output a more accurate query problem by counting the initial vector block sets and update vector block sets of two retrievals. On the one hand, for the strong direction determined by the retrieval, see which knowledge base the retrieval results are repeated from, not all knowledge bases are treated uniformly, and the professionalism of the query problem can be taken into account; on the other hand, for the weak direction determined by the retrieval, see which knowledge base the retrieval results are independently from, not directly abandoned, and the comprehensiveness of the query problem can be taken into account; therefore, the strong and weak directions of the retrieval are positioned, and the focus of the retrieval and the weight of the knowledge base are preferably adjusted, so that query problems with both professionalism and comprehensiveness can be generated. Through two rounds of optimization, this method significantly reduces the ambiguity of the query, enhances the accuracy and diversity of the generated content, and provides a more efficient and accurate solution for the application of the RAG system in complex tasks.
[0132] In summary, it is particularly important to improve the performance of the Retrieval-Augmented Generation (RAG) system when facing the common problems of colloquialism, omission, ambiguity and context dependency in user queries. To this end, the present invention proposes an optimization framework based on two rounds of query rewriting to improve the query understanding and generation quality of the RAG system. First, we constructed a multi-level knowledge base architecture, and divided the data into three levels according to fields, sub-fields and specific topics through scientific classification methods, ensuring that the modules within the knowledge base are independent of each other and the data sources are clear. This hierarchical structure helps to accurately match user queries with relevant knowledge base content and improve the accuracy and controllability of retrieval. In the first round of query rewriting, the RAG system generates the corresponding initial vector block set A for the query questions raised by the user. Then, the user's initial query questions are rewritten by small language models (SLMs) combined with prompt engineering technology, aiming to eliminate the ambiguity in the query and optimize the expression of the query. Subsequently, based on the user query, the RAG system generates an updated vector block set B. In the second round of optimization, by comparing the initial vector block set A and the updated vector block set B, the knowledge base corresponding to more vector blocks is determined as a strong direction, and the knowledge base corresponding to fewer vector blocks is determined as a weak direction. More preferably, the strong direction knowledge base and the weak direction knowledge base can be weighted, for example: strong: weak = 4:1, or weighted in sequence according to the ranking. This weighting strategy helps to guide the query generation process to pay more attention to the professionalism and indicativeness of the content of the high-weight knowledge base, while taking into account the diversity and comprehensiveness of the content of the low-weight knowledge base. Finally, based on these knowledge bases, combined with prompt engineering technology, the query question is rewritten for the second time to generate the final query question, so that this question is more professional and comprehensive while fully including the user's intention. The generated query question will optimize the expression and depth based on the weighted information of each knowledge base, and improve the generation quality and efficiency of the system.
[0133] On the other hand, the retrieval question optimization method can be further applied to retrieval generation, including:
[0134] P1: Use any of the above retrieval problem optimization methods to optimize the query problem;
[0135] P2: Determine the weight of the multi-layer knowledge base according to the strong and weak directions of the retrieval determined in any of the above retrieval problem optimization methods;
[0136] P3: Generate the final vector block set based on the multi-layer knowledge base with optimized query questions and weight settings.
[0137] Specifically, in the comparison results, duplicate sources are given high weights, and independent sources are given low weights (the default ratio is 4:1);
[0138] 2. Algorithm Implementation
[0139] Weight formula: Repeat source weight W_r = 0.8; independent source weight W_i = 0.2.
[0140] Calculate the comprehensive score of each source in dataset C according to the weight:
[0141]
[0142] Algorithm pseudo code:
[0143] data_C = []
[0144] for source in repeated_sources:
[0145] data_C.append({"source": source,"weight":0.8,"score": original_score[source] ×0.8})
[0146] for source in independent_sources:
[0147] data_C.append({"source": source,"weight":0.2,"score": original_score[source]× 0.2})
[0148] data_C = sorted(data_C, key=lambda x: x['score'], reverse=True)
[0149] on the other hand, Figure 4 This is a schematic diagram of the structure of an electronic system provided in one embodiment of the present application. Figure 4 As shown, the electronic system 4 of this embodiment includes: at least one processor 40 ( Figure 4 Only one is shown in the figure), a memory 41, and a computer program 42 stored in the memory 41 and executable on at least one processor 40, the processor 40 executes the computer program 42 to implement the above Figure 3 The steps in the method embodiment, or the implementation of the above Figure 3 Functions of each module / unit in the device embodiment.
[0150] The electronic system 4 may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The electronic system 4 may include but is not limited to a processor 40 and a memory 41. Those skilled in the art will appreciate that Figure 4It is only an example of the electronic system 4 and does not constitute a limitation on the electronic system 4. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, etc.
[0151] The processor 40 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0152] In some embodiments, the memory 41 may be an internal storage unit of the electronic system 4, such as a hard disk or memory of the electronic system 4. In other embodiments, the memory 41 may also be an external storage device of the electronic system 4, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic system 4. Further, the memory 41 may also include both an internal storage unit and an external storage device of the electronic system 4. The memory 41 is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as program codes of a computer program. The memory 41 may also be used to temporarily store data that has been output or is to be output.
[0153] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0154] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the electronic system, a recording medium, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electrical carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a disk or an optical disk.
[0155] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0156] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0157] In the embodiments provided in the present application, it should be understood that the disclosed devices / equipment and methods can be implemented in other ways. For example, the device / equipment embodiments described above are merely schematic, for example, the division of the modules or units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0158] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0159] The above-mentioned electronic system and storage medium are created based on the above-mentioned retrieval problem optimization method, which will not be described in detail here. The above embodiments are only used to illustrate the technical solution of the present application, rather than to limit it; although the present application is described in detail with reference to the above-mentioned embodiments, a person of ordinary skill in the art should understand that: it is still possible to modify the technical solutions recorded in the above-mentioned embodiments, or to replace some of the technical features therein by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A retrieval problem optimization method based on a multi-layer knowledge base, characterized in that: The method comprises: S1: Manage knowledge data in a hierarchical manner and build a multi-layer knowledge base; S2: According to the user query, the multi-layer knowledge base is retrieved to generate an initial vector block set; then the query is rewritten for the first time, and the multi-layer knowledge base is retrieved again to generate an updated vector block set; S3: Determine the strong and weak directions of the retrieval according to the source statistics of the initial vector block set and the updated vector block set in the multi-layer data block; S4: Perform a second query rewrite on the query based on the strong and weak directions of the retrieval and output the query question; Step S3 includes: S31: Determine the source position of each vector block in the initial vector block set and the updated vector block set in the multi-layer knowledge base; S32: Counting the number of vector blocks included in each source position; S33: According to the statistical results, sort and determine the strong and weak directions of the retrieval; Step S33 includes: S331: determining the position coefficient at each source position according to the number of vector blocks included at each source position; S332: If the position coefficient is greater than the first set threshold or the position coefficient is ranked in the first set position, it is determined as a strong direction of the search; otherwise, it is determined as a weak direction; Step S332 further includes: Among the source positions determined to be weak directions, determine whether their position coefficients are less than the second set threshold; or whether their rankings are greater than the second set rank, if so, screen out the corresponding source positions; the second set threshold is less than the first set threshold; the second set rank is greater than the first set rank.
2. The search problem optimization method according to claim 1, characterized in that: Step S31 further includes: S31a: comparing the vector blocks in the initial vector block set and the vector blocks in the updated vector block set to determine the same vector blocks and different vector blocks; S32b: configure a higher number coefficient for the same vector block than for different vector blocks; Step S32 specifically includes: according to the quantity coefficient, counting the number of vector blocks included in each source position according to the weight.
3. The search problem optimization method according to claim 2, characterized in that: Step S331 includes: S3311: in the multi-layer knowledge base, determining the position coefficient of each source position of the first level according to the number of vector blocks included in each source position of the first level; S3312: For each source position except the first level, determine the position coefficient of each source position except the first level according to the number of vector blocks included in the source position and the position coefficients of all its upper levels.
4. The search problem optimization method according to claim 3, characterized in that: Step S3312 further includes: S33121: for each source position except the first level, count the total number of vector blocks included in the same level; S33122: determining a level coefficient at the source position according to a ratio of the number of vector blocks included at the source position to the total number of vector blocks included at the same level; S33123: Determine the position coefficients of each source position except the first level according to the number of vector blocks included in the source position and the position coefficients of all upper levels, as well as the level coefficients of the source position.
5. The search question optimization method according to any one of claims 1 to 4, characterized in that: include: P1: Use the retrieval problem optimization method to optimize the query problem; P2: Determine the weight of the multi-layer knowledge base according to the strong and weak directions of the retrieval determined in the retrieval problem optimization method; P3: Generate the final vector block set based on the optimized query question and the multi-layer knowledge base with weight settings.
6. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.
7. An electronic system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 5 is implemented.