Proposition combination optimization multi-hop retrieval method and device for multi-domain knowledge hinge
Through the multi-hop search method oriented to multi-domain knowledge hinge proposition combination optimization, the existing RAG framework has solved the problems of redundant information, long response time and inaccurate reasoning in the multi-hop question-and-answer question-and-answer question-and-answer task, efficient and accurate cross-domain knowledge fusion and reasoning are achieved, and high-quality answers are generated.
Patent Information
- Application Number
- CN202510485784.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-17
AI Technical Summary
The existing RAG framework has problems such as redundant information, long response time, lots of noise information and inaccurate reasoning in multi-hop question-and-answer tasks. Especially in the complex problems of multi-domain knowledge intersection, traditional methods are difficult to effectively capture the inherent connections of cross-domain knowledge.
A multi-hop search method for multi-domain knowledge hinges is proposed. By obtaining cross-domain questions and non-structural documents, proposition extraction and network construction are carried out, multi-domain knowledge hinges are identified, candidate proposition combinations are optimized, combination scores are dynamically updated, and answers are finally generated.
It improves the accuracy and relevance of the reasoning chain, reduces redundant information, and improves the system efficiency and reliability of answers. Especially in the multi-domain knowledge cross-domain knowledge scenario, it can effectively capture the inherent connections of cross-domain knowledge and generate high-quality and accurate content.
Smart Images

Figure CN119988690A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a proposition combination optimization multi-hop retrieval method and device for multi-domain knowledge hinges. Background Technology
[0002] With the development of large language model (LLM) technology, multi-hop question answering, as an important application of natural language processing, has received extensive attention in the fields of intelligent question answering systems, knowledge reasoning, and information retrieval. Multi-hop question answering requires the system to be able to extract and combine information from multiple documents and construct answers through multi-step reasoning. In order to improve the accuracy of the answers, the retrieval-augmented generation (RAG) framework is currently widely used to enhance the generation quality by combining external knowledge bases.
[0003] However, the existing RAG framework has obvious shortcomings when dealing with multi-hop question answering tasks. Due to the document-level retrieval method, the retrieval results often contain a lot of redundant information, which reduces the efficiency of the system. At the same time, in order to obtain a complete reasoning link, traditional methods usually adopt a multi-round retrieval strategy, which requires repeated calls to large language models for query rewriting, which not only increases the system response time, but also easily introduces noise information, affecting the accuracy of reasoning. In addition, due to the strong contextual dependence between document fragments, the reasoning process is difficult to track and verify, affecting the reliability of the system in practical applications.
[0004] Traditional methods face greater challenges, especially in complex problems involving the intersection of knowledge in multiple fields. Since most existing retrieval frameworks tend to focus on knowledge retrieval in a single field, they fail to effectively capture the intrinsic connections between cross-domain knowledge, which leads to broken reasoning or wrong answers in the system when it comes to multi-hop reasoning in fields such as medicine and law. In addition, in these fields, the implicit relationships between knowledge are more complex, and existing frameworks often lack the ability to identify and utilize these complex relationships, which makes them limited in effectiveness in practical applications.
[0005] Therefore, traditional methods still face significant challenges, especially when dealing with cross-domain reasoning tasks. It is necessary to find more efficient, accurate and explainable retrieval and reasoning methods to meet application scenarios with high requirements for reasoning accuracy and explainability, so as to ensure the reliability of model output content. SUMMARY OF THE INVENTION
[0006] Based on this, it is necessary to provide a proposition combination optimization multi-hop retrieval method and device for multi-domain knowledge hinges in response to the above technical problems.
[0007] A proposition combination optimization multi-hop retrieval method for multi-domain knowledge hinges, the method comprising: Obtain cross-domain issues and domain-related non-structured documents; the cross-domain issues include at least two different domains; the non-structured documents include domain knowledge; Extract propositions from each unstructured document to obtain an atomic proposition set containing multiple atomic propositions; Use each atomic proposition in the atomic proposition set as a node, and construct a proposition network according to the entity co-occurrence relationship between the atomic propositions; the path in the proposition network is a multi-domain knowledge hinge; Initialize several candidate proposition combinations, select neighbor propositions of each atomic proposition in each candidate proposition combination in the proposition network through a preset pruning method, and obtain several new candidate proposition combinations according to each candidate proposition combination and the corresponding neighbor proposition; each candidate proposition combination is a multi-domain knowledge hinge; The combination score is calculated based on the similarity between the cross-domain problem and the proposition combination. If the combination score of the new candidate proposition combination is higher than the current candidate proposition combination, the current candidate proposition combination is updated, and each candidate proposition combination is updated iteratively until the condition for the iteration to stop is met, and the optimal proposition combination set is output; The optimal proposition combination set is input into the pre-trained large language model to generate answers corresponding to cross-domain questions.
[0008] In one embodiment, it also includes: using a sliding window to semantically segment the unstructured document to obtain multiple segmented contents; using a pre-trained large language model to convert the segmented contents into independent atomic propositions.
[0009] In one embodiment, it also includes: if the atomic propositions share the same entity, then there is an entity co-occurrence relationship between the atomic propositions, and the atomic propositions with the entity co-occurrence relationship are connected to construct a proposition network.
[0010] In one embodiment, it also includes: splicing the atomic propositions in the proposition combination, and text encoding the splicing result to obtain a proposition combination encoding vector; text encoding the cross-domain problem to obtain a cross-domain problem encoding vector; obtaining a combination score of the proposition combination according to the similarity between the proposition combination encoding vector and the cross-domain problem encoding vector; the proposition combination includes a candidate proposition combination and a new candidate proposition combination.
[0011] In one of the embodiments, it also includes: obtaining a set of neighbor propositions in the proposition network for each atomic proposition in each candidate proposition combination; obtaining a direct correlation based on the similarity between each atomic proposition and the cross-domain problem, and obtaining a neighbor correlation based on the similarity between each neighbor proposition in the neighbor proposition set and the cross-domain problem; calculating a pruning score for each atomic proposition based on the direct correlation of each atomic proposition and the corresponding maximum neighbor correlation; if the pruning score corresponding to the atomic proposition is higher than a threshold, retaining the neighbor proposition corresponding to the maximum neighbor correlation.
[0012] In one embodiment, the pruning score is: Score(p) = α·R(p) + (1-α)·maxR(v); Where p represents an atomic proposition, R(p) represents the direct relevance of p to cross-domain issues, v represents p's neighbor proposition, v∈N(p), N(p) is the set of p's neighbor propositions, α is the weight coefficient, and R(v) represents the relevance of v to the neighbors of cross-domain issues.
[0013] In one embodiment, it also includes: sorting and organizing the optimal proposition combination set to obtain a reasoning link, processing the reasoning link through a pre-trained large language model, and generating an answer corresponding to the cross-domain question.
[0014] In one embodiment, it also includes: performing an early stopping check when iteratively updating each candidate proposition combination; the early stopping check includes: obtaining a preset early stopping round number, and if the candidate proposition combination is not updated in multiple consecutive iterations, and the number of iteration rounds meets the early stopping round number, terminating the search in advance.
[0015] A proposition combination optimization multi-hop retrieval device for multi-domain knowledge hinges, the device comprising: Data acquisition module, used to acquire cross-domain issues and domain-related non-structured documents; the cross-domain issues include at least two different domains; the non-structured documents include domain knowledge; The proposition extraction module is used to extract propositions from each unstructured document to obtain an atomic proposition set containing multiple atomic propositions; Network construction module, used to take each atomic proposition in the atomic proposition set as a node, and construct a proposition network according to the entity co-occurrence relationship between the atomic propositions; the path in the proposition network is a multi-domain knowledge hinge; A pruning module is used to initialize several candidate proposition combinations, select neighbor propositions of each atomic proposition in each candidate proposition combination in the proposition network through a preset pruning method, and obtain several new candidate proposition combinations according to each candidate proposition combination and the corresponding neighbor proposition; The update module is used to calculate the combination score based on the similarity between the cross-domain problem and the proposition combination. If the combination score of the new candidate proposition combination is higher than the current candidate proposition combination, the current candidate proposition combination is updated, and each candidate proposition combination is updated iteratively until the condition for stopping the iteration is met, and the optimal proposition combination set is output; Answer generation module, which is used to input the optimal proposition combination set into the pre-trained large language model to generate answers corresponding to cross-domain questions.
[0016] A computer device comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented: Obtain cross-domain issues and domain-related non-structured documents; the cross-domain issues include at least two different domains; the non-structured documents include domain knowledge; Extract propositions from each unstructured document to obtain an atomic proposition set containing multiple atomic propositions; Use each atomic proposition in the atomic proposition set as a node, and construct a proposition network according to the entity co-occurrence relationship between the atomic propositions; the path in the proposition network is a multi-domain knowledge hinge; Initialize several candidate proposition combinations, select neighbor propositions of each atomic proposition in each candidate proposition combination in the proposition network through a preset pruning method, and obtain several new candidate proposition combinations according to each candidate proposition combination and the corresponding neighbor proposition; The combination score is calculated based on the similarity between the cross-domain problem and the proposition combination. If the combination score of the new candidate proposition combination is higher than the current candidate proposition combination, the current candidate proposition combination is updated, and each candidate proposition combination is updated iteratively until the condition for the iteration to stop is met, and the optimal proposition combination set is output; The optimal proposition combination set is input into the pre-trained large language model to generate answers corresponding to cross-domain questions.
[0017] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the following steps: Obtain cross-domain issues and domain-related non-structured documents; the cross-domain issues include at least two different domains; the non-structured documents include domain knowledge; Extract propositions from each unstructured document to obtain an atomic proposition set containing multiple atomic propositions; Use each atomic proposition in the atomic proposition set as a node, and construct a proposition network according to the entity co-occurrence relationship between the atomic propositions; the path in the proposition network is a multi-domain knowledge hinge; Initialize several candidate proposition combinations, select neighbor propositions of each atomic proposition in each candidate proposition combination in the proposition network through a preset pruning method, and obtain several new candidate proposition combinations according to each candidate proposition combination and the corresponding neighbor proposition; The combination score is calculated based on the similarity between the cross-domain problem and the proposition combination. If the combination score of the new candidate proposition combination is higher than the current candidate proposition combination, the current candidate proposition combination is updated, and each candidate proposition combination is updated iteratively until the condition for the iteration to stop is met, and the optimal proposition combination set is output; The optimal proposition combination set is input into the pre-trained large language model to generate answers corresponding to cross-domain questions.
[0018] The above-mentioned proposition combination optimization multi-hop retrieval method and device for multi-domain knowledge hinges first obtains cross-domain problems and related unstructured documents, extracts propositions, generates a set of atomic propositions, provides structured information for subsequent reasoning, and constructs a proposition network by identifying the relationship between atomic propositions. It can capture the semantic connection between different fields, especially the role of cross-domain knowledge hinges, enhance the coherence of knowledge fusion and reasoning, initialize several candidate proposition combinations and optimize the search space through pruning methods, reduce redundant propositions, ensure retrieval efficiency, perform combination scoring based on the similarity between cross-domain problems and proposition combinations, dynamically update candidate proposition combinations, and finally output the optimal proposition combination set. This process improves the accuracy and relevance of the reasoning chain through iterative optimization. Finally, the optimal proposition combination is input into the pre-trained large language model to generate answers, ensuring the accuracy and interpretability of the answers. The embodiment of the present invention can improve the reliability of the answer content output by the large language model in scenarios involving complex problems in multiple fields, efficiently generate high-quality and accurate content, and meet practical application needs. Brief Description of the Figures
[0019] Figure 1 A flowchart of a proposition combination optimization multi-hop retrieval method for multi-domain knowledge hinges in one embodiment; Figure 2 A schematic diagram of proposition-oriented network construction in one embodiment; Figure 3 is a flowchart of the overall framework of the method of the present invention in one embodiment; Figure 4 It is a structural block diagram of a proposition combination optimization multi-hop retrieval device for multi-domain knowledge hinges in one embodiment; Figure 5 is a diagram of the internal structure of a computer device in one embodiment. Specific implementation method
[0020] In order to make the purpose, technical solutions and advantages of this application more clear, the following is a further detailed description of this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described here are only used to explain this application and are not used to limit this application.
[0021] In one embodiment, if Figure 1 As shown in the figure, a proposition combination optimization multi-hop retrieval method for multi-domain knowledge hinges is provided, including the following steps: Step 102, obtaining cross-domain issues and domain-related non-structured documents.
[0022] Cross-domain problems include at least two different domains. Unstructured documents include domain knowledge. This invention has wide application value in professional fields that require rigorous reasoning, especially in scenarios where multi-domain knowledge needs to be connected. Its efficient information retrieval and explainable reasoning process provide reliable technical support for complex question-answering tasks.
[0023] In an embodiment of the present invention, a cross-domain problem may be the risk factors for a certain disease and its association with lifestyle. In the multi-domain attribution analysis scenario of chronic disease risk factors, the present invention constructs a dual-domain proposition network of medical clinical knowledge and health management data. Through atomization parsing technology, the system standardizes the representation of clinical entities such as pathological mechanisms and biomarkers in medical literature and life entities such as behavioral patterns and environmental factors in health monitoring data. Based on the cross-domain knowledge hinge recognition algorithm, the system automatically discovers associated propositions connecting medical diagnostic standards and lifestyle assessment indicators (such as the interactive relationship between metabolic syndrome and work and rest disorders) and constructs multi-hop reasoning links. When dealing with complex health consultations, the system can quickly generate an integrated decision path covering pathological explanations, behavioral intervention recommendations and prognosis assessments, significantly improving the efficiency of cross-domain knowledge collaboration.
[0024] In addition, it can also be applied to handle complex legal document retrieval and reasoning tasks. In response to the application of multi-field provisions in cross-border legal scenarios, a complex proposition map across civil, commercial and administrative supervision fields is established. Through the legal requirements deconstruction engine, the system converts the text of clauses in different legal systems into atomic propositions with logical associations, and identifies the implicit associations between cross-domain legal requirements. When dealing with legal reasoning tasks involving multi-domain collaboration, the system can automatically build a three-dimensional reasoning network connecting substantive law clauses, procedural law provisions and regulatory regulations, effectively solving the problem of missing requirements caused by traditional single-field retrieval, and achieving accurate matching and traceability analysis of multi-dimensional legal requirements.
[0025] Step 104, extracting propositions from each unstructured document to obtain an atomic proposition set containing multiple atomic propositions.
[0026] By splitting unstructured documents into atomic propositions, the information in each sentence or paragraph can be converted into more basic and structured units, which helps to eliminate information redundancy and improve the clarity and accuracy of the reasoning process. Atomic propositions can express a complete fact or opinion alone, and they have semantic integrity, context independence and atomicity.
[0027] Step 106, taking each atomic proposition in the atomic proposition set as a node, and constructing a proposition network based on the entity co-occurrence relationship between the atomic propositions.
[0028] The nodes of a proposition network are atomic propositions, and the edges are semantic connections between atomic propositions. A proposition network is an undirected graph G=(V,E), where the node set V corresponds to the proposition set, and the edge set E represents the semantic connection between propositions. By establishing a proposition network, related propositions are organized into a graph structure (undirected graph) according to their semantic relationships.
[0029] The paths in the proposition network are multi-domain knowledge hinges. By constructing a proposition network, it is helpful to identify multi-domain knowledge hinges. Multi-domain knowledge hinges refer to key propositions or nodes that can connect knowledge in different fields. Each field has its own specific knowledge system and terminology. There may be some shared concepts, entities or relationships between these fields. Multi-domain knowledge hinges are bridges that can establish connections between different fields. They connect the knowledge of various fields, so that the method of the present invention can cross field barriers and perform effective information fusion and reasoning.
[0030] The system first deconstructs document data from different fields into a series of atomic knowledge propositions, and then constructs a proposition network based on the entity co-occurrence relationship between propositions, so that any path in the proposition network is a multi-field knowledge hinge. Among so many potential paths, find several knowledge hinges that are most relevant to the user's question. Specifically, through the BeamSearch (beam search) iterative proposition retrieval method, these knowledge hinges are sorted by relevance according to the user's question to obtain the most relevant knowledge hinges. By identifying multi-field knowledge hinges, the present invention can focus on propositions related to specific cross-field knowledge hinges, avoid interference from irrelevant propositions, greatly improve the efficiency of retrieval, and avoid broken links and redundant information in reasoning. By effectively identifying and utilizing these hinges, the method of the present invention can establish connections between multiple fields, perform deep reasoning, and ultimately generate reliable and explainable answers.
[0031] Step 108, initialize several candidate proposition combinations, select neighbor propositions of each atomic proposition in each candidate proposition combination in the proposition network through a preset pruning method, and obtain several new candidate proposition combinations based on each candidate proposition combination and the corresponding neighbor proposition.
[0032] Each candidate proposition combination is a multi-domain knowledge hinge. Through the iterative proposition retrieval algorithm of the present invention, an iterative search is performed in the proposition network to select the proposition set most relevant to the problem. The algorithm will dynamically adjust the search path according to the characteristics of the input problem to avoid the overhead of multiple LLM calls in traditional methods. During the search process, the proposition combination can be gradually expanded, and the pre-set pruning mechanism can be used to improve efficiency and accuracy.
[0033] Step 110, calculate the combination score based on the similarity between the cross-domain problem and the proposition combination. If the combination score of the new candidate proposition combination is higher than the current candidate proposition combination, update the current candidate proposition combination, iteratively update each candidate proposition combination until the condition for stopping the iteration is met, and output the optimal proposition combination set.
[0034] By evaluating the contribution of propositions in different combinations, the most critical propositions can be identified. The top K proposition combinations with the highest similarity to the question are the most relevant knowledge hinges. In this way, the system can identify the "knowledge hinges" that connect different fields and obtain hinge propositions. These hinge propositions may play a vital role in the cross-domain reasoning process, helping the system to effectively combine knowledge from different fields.
[0035] Step 112, input the optimal proposition combination set into the pre-trained large language model to generate answers corresponding to cross-domain questions.
[0036] After performing proposition retrieval, the optimal proposition combination set is obtained, and finally the answer is generated based on the optimal proposition combination set. Because these propositions are extracted from unstructured documents and form a complete proposition network, the large language model can effectively combine multiple propositions into a coherent reasoning chain to obtain accurate answers.
[0037] In the above-mentioned proposition combination optimization multi-hop retrieval method for multi-domain knowledge hinges, first, cross-domain problems and related unstructured documents are obtained, propositions are extracted, and a set of atomic propositions is generated to provide structured information for subsequent reasoning. By identifying the relationship between atomic propositions and building a proposition network, the semantic connection between different fields, especially the role of cross-domain knowledge hinges, is captured, and the coherence of knowledge fusion and reasoning is enhanced. Several candidate proposition combinations are initialized and the search space is optimized through pruning methods to reduce redundant propositions and ensure retrieval efficiency. Combination scoring is performed based on the similarity between cross-domain problems and proposition combinations, and candidate proposition combinations are dynamically updated. Finally, the optimal proposition combination set is output. This process improves the accuracy and relevance of the reasoning chain through iterative optimization. Finally, the optimal proposition combination is input into the pre-trained large language model to generate answers, ensuring the accuracy and interpretability of the answers. The embodiment of the present invention can improve the reliability of the answer content output by the large language model in scenarios involving complex problems in multiple fields, efficiently generate high-quality and accurate content, and meet practical application needs.
[0038] In one embodiment, proposition extraction is performed on each unstructured document to obtain an atomic proposition set containing multiple atomic propositions, including: using a sliding window to semantically segment the unstructured document to obtain multiple segmented contents; using a pre-trained large language model to convert the segmented contents into independent atomic propositions. In this embodiment, a sliding window mechanism is first used to semantically segment long documents to maintain the coherence of the context, and then an extraction strategy based on a large language model is adopted to convert the unstructured document into an atomic proposition set to ensure that each proposition is an independent and complete statement of fact. In this process, pronoun resolution and entity normalization are automatically performed to ensure that each proposition has semantic integrity, context independence and atomicity.
[0039] In one embodiment, a proposition network is constructed based on the entity co-occurrence relationship between atomic propositions, including: if the atomic propositions share the same entity, then there is an entity co-occurrence relationship between the atomic propositions, and the atomic propositions with the entity co-occurrence relationship are connected to construct the proposition network.
[0040] In this embodiment, if Figure 2 The diagram of proposition network construction shown in the figure shows the basic structure of proposition network, where nodes represent atomic propositions, and the lines between nodes represent the entity co-occurrence relationship between propositions. The bold nodes and lines represent multi-domain knowledge hinges. Through this connection method based on entity co-occurrence, the system can establish associations between propositions and support subsequent multi-step reasoning processes.
[0041] Constructing a proposition network based on entity co-occurrence relationships means establishing a connection between two propositions when they share the same entity. The proposition network is constructed through the connection relationship between propositions, and multi-domain knowledge hinge nodes are identified to provide structural support for subsequent multi-hop reasoning.
[0042] In one embodiment, the combined score is obtained based on the similarity between the cross-domain problem and the proposition combination, including: splicing the atomic propositions in the proposition combination, and text encoding the splicing result to obtain the proposition combination encoding vector; text encoding the cross-domain problem to obtain the cross-domain problem encoding vector; obtaining the combined score of the proposition combination based on the similarity between the proposition combination encoding vector and the cross-domain problem encoding vector; the proposition combination includes a candidate proposition combination and a new candidate proposition combination.
[0043] In this embodiment, the dynamic update mechanism of proposition scoring identifies key propositions by capturing the contribution of propositions in different combinations. For the proposition combination PC, its score is calculated as follows: score(PC) = similarity(encode(concat(PC)), encode(q)); Among them, PC is a proposition combination, concat is a concatenation function, encode is a text encoding function, and similarity is a semantic similarity calculation function. This mechanism can effectively identify propositions that show semantic emergence in the combination, especially hinge propositions that connect knowledge in multiple fields.
[0044] In one embodiment, the neighbor propositions of each atomic proposition in each candidate proposition combination in the proposition network are selected by a preset pruning method, including: obtaining the neighbor proposition set of each atomic proposition in each candidate proposition combination in the proposition network; obtaining the direct relevance according to the similarity between each atomic proposition and the cross-domain problem, and obtaining the neighbor relevance according to the similarity between each neighbor proposition in the neighbor proposition set and the cross-domain problem; calculating the pruning score of each atomic proposition according to the direct relevance of each atomic proposition and the corresponding maximum neighbor relevance; if the pruning score corresponding to the atomic proposition is higher than the threshold, retaining the neighbor proposition corresponding to the maximum neighbor relevance.
[0045] In this embodiment, the pruning mechanism reduces the search space by evaluating the local and neighborhood relevance of propositions. This design ensures that even if a proposition has a low direct relevance to the problem, as long as its neighboring propositions have a high relevance, the proposition may still become an important reasoning intermediate node or knowledge hinge.
[0046] Among them, the candidate proposition combination requires a reliable initial value. Initializing the candidate proposition combination includes: First, using a text encoding model (such as BERT) to calculate the vector representation of the input question q to provide a basis for subsequent proposition scoring. According to the semantic similarity between the question and each proposition, the relevance score of each proposition to the question is calculated. Based on the initial relevance score, the system selects several most relevant proposition combinations as candidates. These initial combinations will provide a starting point for subsequent iterations. Record the current best proposition combination score to ensure that the optimal solution can be tracked during the iteration process. Monitor whether there is improvement during the iteration process. If there is no improvement in the score for several consecutive rounds, terminate early. The above process is the initialization stage, and the combination results of the initialization stage are used to initialize the candidate proposition combination.
[0047] In one embodiment, the pruning score is: Score(p) = α·R(p) + (1-α)·maxR(v); Where p represents an atomic proposition, R(p) represents the direct relevance of p to cross-domain issues, v represents p's neighbor proposition, v∈N(p), N(p) is the set of p's neighbor propositions, α is the weight coefficient, and R(v) represents the relevance of v to the neighbors of cross-domain issues.
[0048] In one embodiment, the method further includes: performing an early stopping check when iteratively updating each candidate proposition combination; the early stopping check includes: obtaining a preset number of early stopping rounds, and if the candidate proposition combination is not updated in multiple consecutive iterations, and the number of iteration rounds meets the number of early stopping rounds, terminating the search in advance. In this embodiment, by monitoring the convergence of the search process, the search is terminated in advance when a proposition combination with a higher score fails to be found after r consecutive iterations. The early stopping mechanism avoids over-exploration, significantly reduces the computational overhead, and ensures the quality of the retrieval results.
[0049] In one embodiment, the optimal proposition combination set is input into a pre-trained large language model to generate answers corresponding to cross-domain questions, including: sorting and organizing the optimal proposition combination set to obtain reasoning links, processing the reasoning links through the pre-trained large language model, and generating answers corresponding to cross-domain questions.
[0050] In this embodiment, the optimal proposition combination retrieved is obtained, and the final answer is generated using a large language model. First, the optimal proposition combination is sorted and organized to build a complete reasoning chain. By identifying the knowledge hinges related to the question, that is, extracting the most relevant propositions from multiple fields, the system enhances the efficiency and accuracy of cross-domain reasoning. This method is more efficient than traditional text paragraph retrieval in terms of information density, and enables the proposition combination to be directly processed by the large language model for reasoning, thereby generating more targeted answers. This technical approach selects the top K most relevant proposition combinations as "knowledge hinges", and then inputs these propositions together with the input questions into the large language model, using its powerful reading comprehension ability to generate accurate and explanatory answers. In this way, the system improves the efficiency of cross-domain multi-hop question answering while ensuring the accuracy and explainability of the answers. This embodiment pays special attention to the role of multi-domain knowledge hinges, and then generates accurate and explainable answers based on these structured evidence. In this way, the system not only ensures the accuracy of the answers, but also provides a clear basis for reasoning.
[0051] In a specific embodiment, such as Figure 3 As shown in the figure, a flowchart of the overall framework of the method of the present invention is provided, in which: the input part includes unstructured documents and questions; the proposition processing part shows the proposition extraction module and the proposition network construction module; the retrieval optimization part includes the Beam Search iterative proposition retrieval module and its three key mechanisms (proposition score dynamic update, pruning and early stopping); and finally the answer generation part generates the final answer based on the retrieved optimal proposition combination. The Beam Search iterative proposition retrieval algorithm of the present invention can be represented by the following pseudo code:
[0052] Through the synergy of these three mechanisms, the present invention achieves efficient multi-hop retrieval: dynamic score update ensures the accuracy of reasoning, pruning mechanism reduces the search space, and early stopping mechanism optimizes computational efficiency. The entire retrieval process not only maintains the integrity of reasoning, but also significantly improves system performance, especially in scenarios that require cross-domain knowledge connections.
[0053] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 1 At least part of the steps in may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0054] In one embodiment, if Figure 4 As shown, a proposition combination optimization multi-hop retrieval device for multi-domain knowledge hinges is provided, including: Data acquisition module 402, used to acquire cross-domain issues and domain-related non-structured documents; cross-domain issues include at least two different domains; non-structured documents include domain knowledge; The proposition extraction module 404 is used to extract propositions from each unstructured document to obtain an atomic proposition set containing multiple atomic propositions; Network construction module 406, used to take each atomic proposition in the atomic proposition set as a node, and construct a proposition network according to the entity co-occurrence relationship between the atomic propositions; the path in the proposition network is a multi-domain knowledge hinge; Pruning module 408, used to initialize several candidate proposition combinations, select neighbor propositions of each atomic proposition in each candidate proposition combination in the proposition network through a preset pruning method, and obtain several new candidate proposition combinations according to each candidate proposition combination and the corresponding neighbor proposition; each candidate proposition combination is a multi-domain knowledge hinge; Update module 410, used to calculate the combination score based on the similarity between the cross-domain problem and the proposition combination. If the combination score of the new candidate proposition combination is higher than the current candidate proposition combination, the current candidate proposition combination is updated, and each candidate proposition combination is updated iteratively until the condition for stopping the iteration is met, and the optimal proposition combination set is output; Answer generation module 412 is used to input the optimal proposition combination set into the pre-trained large language model to generate answers corresponding to cross-domain questions.
[0055] For the specific definition of the proposition combination optimization multi-hop retrieval device for multi-domain knowledge hinges, please refer to the definition of the proposition combination optimization multi-hop retrieval method for multi-domain knowledge hinges in the above text, which will not be repeated here. Each module in the above-mentioned proposition combination optimization multi-hop retrieval device for multi-domain knowledge hinges can be fully or partially implemented by software, hardware and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the corresponding operations of the above modules.
[0056] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in Figure 5 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a proposition combination optimization multi-hop retrieval method for multi-domain knowledge hinges is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.
[0057] Those skilled in the art will understand that Figure 5 The structure shown in is only a block diagram of a part of the structure related to the present application solution, and does not constitute a limitation on the computer device to which the present application solution is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0058] In one embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the method in the above embodiment when executing the computer program.
[0059] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method in the above embodiment are implemented.
[0060] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0061] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0062] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention. It should be pointed out that, for ordinary technicians in this field, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.
Claims
1. A proposition combination optimization multi-hop retrieval method for multi-domain knowledge hinges, characterized by: The method comprises: Acquire cross-domain issues and domain-related unstructured documents; the cross-domain issues include at least two different domains; the unstructured documents include domain knowledge; Extract propositions from each unstructured document to obtain an atomic proposition set containing multiple atomic propositions; Taking each atomic proposition in the atomic proposition set as a node, a proposition network is constructed according to the entity co-occurrence relationship between the atomic propositions; the path in the proposition network is a multi-domain knowledge hinge; Initialize a number of candidate proposition combinations, select neighbor propositions of each atomic proposition in each candidate proposition combination in the proposition network through a preset pruning method, and obtain a number of new candidate proposition combinations according to each candidate proposition combination and the corresponding neighbor proposition; each candidate proposition combination is a multi-domain knowledge hinge; The combination score is calculated based on the similarity between the cross-domain problem and the proposition combination. If the combination score of the new candidate proposition combination is higher than the current candidate proposition combination, the current candidate proposition combination is updated, and each candidate proposition combination is updated iteratively until the condition for stopping the iteration is met, and the optimal proposition combination set is output; The optimal proposition combination set is input into the pre-trained large language model to generate answers corresponding to cross-domain questions.
2. The method according to claim 1, characterized in that The proposition extraction is performed on each unstructured document to obtain an atomic proposition set containing multiple atomic propositions, including: Use sliding windows to semantically segment unstructured documents and obtain multiple segmented contents; Use the pre-trained large language model to transform segmented content into independent atomic propositions.
3. The method according to claim 1, characterized in that Construct a proposition network based on the entity co-occurrence relationship between atomic propositions, including: If atomic propositions share the same entity, then there is an entity co-occurrence relationship between the atomic propositions. The atomic propositions with entity co-occurrence relationships are connected to construct a proposition network.
4. The method according to claim 1, characterized in that: The combined score is calculated based on the similarity between the cross-domain questions and proposition combinations, including: Concatenate the atomic propositions in the proposition combination, and perform text encoding on the concatenated result to obtain a proposition combination encoding vector; Perform text encoding on cross-domain issues to obtain a cross-domain issue encoding vector; A combined score of the proposition combination is obtained according to the similarity between the proposition combination coding vector and the cross-domain problem coding vector; the proposition combination includes a candidate proposition combination and a new candidate proposition combination.
5. The method according to claim 1, characterized in that The neighbor propositions of each atomic proposition in each candidate proposition combination in the proposition network are selected by a preset pruning method, including: Obtaining a set of neighbor propositions of each atomic proposition in each candidate proposition combination in the proposition network; The direct correlation is obtained based on the similarity between each atomic proposition and the cross-domain problem, and the neighbor correlation is obtained based on the similarity between each neighbor proposition in the neighbor proposition set and the cross-domain problem; The pruning score of each atomic proposition is calculated according to the direct relevance of each atomic proposition and the corresponding maximum neighbor relevance; If the pruning score corresponding to the atomic proposition is higher than the threshold, the neighbor proposition corresponding to the maximum neighbor correlation is retained.
6. The method according to claim 5, characterized in that The pruning score is: Score(p) = α·R(p) + (1-α)·maxR(v); Among them, p represents an atomic proposition, R(p) represents the direct relevance of p to cross-domain issues, v represents the neighbor proposition of p, v∈N(p), N(p) is the set of neighbor propositions of p, α is the weight coefficient, and R(v) represents the neighbor relevance of v to cross-domain issues.
7. The method according to claim 1, characterized in that The optimal proposition combination set is input into the pre-trained large language model to generate answers corresponding to cross-domain questions, including: The optimal proposition combination set is sorted and organized to obtain a reasoning chain, which is processed by a pre-trained large language model to generate answers corresponding to cross-domain questions.
8. The method according to claim 1, characterized in that The method further comprises: When iteratively updating each candidate proposition combination, an early stopping check is performed; the early stopping check includes: Get the preset number of early stopping rounds. If the candidate proposition combination is not updated in multiple consecutive iterations and the number of iteration rounds meets the said number of early stopping rounds, terminate the search early.
9. A proposition combination optimization multi-hop retrieval device for multi-domain knowledge hinges, characterized in that: The device comprises: A data acquisition module, used to acquire cross-domain issues and domain-related non-structured documents; the cross-domain issues include at least two different domains; the non-structured documents include domain knowledge; A proposition extraction module is used to extract propositions from each unstructured document to obtain an atomic proposition set containing multiple atomic propositions; A network construction module, used to take each atomic proposition in the atomic proposition set as a node and construct a proposition network according to the entity co-occurrence relationship between the atomic propositions; the path in the proposition network is a multi-domain knowledge hinge; A pruning module is used to initialize several candidate proposition combinations, select neighbor propositions of each atomic proposition in each candidate proposition combination in the proposition network through a preset pruning method, and obtain several new candidate proposition combinations according to each candidate proposition combination and the corresponding neighbor proposition; each candidate proposition combination is a multi-domain knowledge hinge; An updating module is used to calculate the combination score based on the similarity between the cross-domain problem and the proposition combination. If the combination score of the new candidate proposition combination is higher than the current candidate proposition combination, the current candidate proposition combination is updated, and each candidate proposition combination is updated iteratively until the condition for the iteration to stop is met, and the optimal proposition combination set is output; The answer generation module is used to input the optimal proposition combination set into the pre-trained large language model to generate answers corresponding to cross-domain questions.
Citation Information
Patent Citations
Method for representing semanteme and relation between semanteme in natural language, reasoning method and artificial intelligence system
CN115168560A
Knowledge question and answer method based on large language model and knowledge graph
CN117874180A
Inference method and device for multi-hop question and answer of knowledge graph and storage medium
CN118503366A
Data intelligent question and answer method and system fusing domain knowledge
CN118779438A
Knowledge graph auxiliary question and answer method suitable for electric power field
CN119322853A
Cited By
Scalable expert foundry system using hierarchical supervisory networks and geometric manifold architectures for multi-domain cognitive processing
US12572748B1
Evolutionary thought caching for multi-stage language model systems
US12585882B1
Persistent cognitive machine with curated long term memory
US12602549B1
Geometric Multi-Modal Sensor Fusion for Infrastructure Monitoring
US20260228438A1