Graph-based retrieval enhancement generation method and device, equipment and storage medium

By constructing target knowledge graphs and community division in the graph-based search enhancement generation method, the problems of incoherence and inaccurate search are solved, and more efficient and accurate search results are achieved.

CN120011528APending Publication Date: 2025-05-16XIAN SECLOVER INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411945535.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The graph-based search enhancement generation method has problems of incoherence and inaccurate search when processing complex queries.

Method used

By dividing the input text into multiple analytical units, the relevant information is extracted using the preset large language model and the preset entity disambiguation method, the target knowledge graph is constructed, and the community is divided and abstract generation is generated through the hierarchical clustering method, and the query result search is finally carried out based on the target knowledge graph and abstract information.

Benefits of technology

It improves the accuracy and efficiency of retrieval, enhances the consistency of semantic information, and improves the accuracy of user query results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011528A_ABST
    Figure CN120011528A_ABST
Patent Text Reader

Abstract

The invention discloses a graph-based retrieval enhancement generation method and device, equipment and a storage medium, and the method comprises the steps: segmenting an input text into a plurality of analyzable units, extracting related information from the plurality of analyzable units through a preset large language model by adopting a preset entity disambiguation method, constructing a target knowledge graph, and performing community division on the target knowledge graph through a preset hierarchical clustering method, generating summary information corresponding to each community, and searching a target query result corresponding to the current user query information based on the target knowledge graph and the summary information. According to the scheme, when the target knowledge graph is constructed, the entities and the relationships are extracted by adopting the preset entity disambiguation method, and the constructed structured information is beneficial to improving the retrieval precision and efficiency, so that the retrieval accuracy and the semantic information continuity can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a graph-based retrieval enhancement generation method, device, equipment and storage medium. Background Art

[0002] With the rapid development of Large Language Model (LLM) and related technologies, graph-based retrieval enhancement generation method has become an emerging technology, which combines traditional retrieval enhancement generation and knowledge graph technology to improve the performance and accuracy of large language models in processing complex queries.

[0003] At present, graph-based retrieval enhancement generation methods can better understand and retrieve complex information by building and utilizing knowledge graphs. They solve the limitations of traditional retrieval enhancement generation methods in processing certain complex queries, such as the inability to effectively process queries that require cross-document understanding or high-level semantic understanding. In addition, by introducing knowledge graphs, it can more accurately capture the relationship between entities and provide richer contextual information, thereby generating answers that better meet user needs.

[0004] However, graph-based retrieval enhancement generation methods still have many problems, such as information incoherence and inaccurate retrieval. Summary of the invention

[0005] The present application aims to at least solve the technical problems existing in the prior art. To this end, the present application proposes a graph-based retrieval enhancement generation method in a first aspect, the method comprising:

[0006] The input text is divided into multiple analyzable units, and the preset entity disambiguation method is used to extract relevant information from the multiple analyzable units through the preset large language model to construct a target knowledge graph; wherein the relevant information includes multiple entities, relationships between entities and key statements, and the preset entity disambiguation method is obtained through a preset disambiguation level and a corresponding preset disambiguation strategy, the preset disambiguation level is determined according to the sensitivity and acceptance of the current scenario to entity ambiguity, and the preset disambiguation strategy is determined based on at least one preset entity detection technology;

[0007] The target knowledge graph is divided into communities through a preset hierarchical clustering method, and summary information corresponding to each community is generated;

[0008] Based on the target knowledge graph and summary information, search for the target query results corresponding to the current user query information.

[0009] In a possible implementation, a preset large language model is used to extract relevant information from multiple analyzable units using a preset entity disambiguation method to construct a target knowledge graph, including:

[0010] By using a preset large language model and a preset entity disambiguation method, relevant information is extracted from multiple analyzable units to obtain an initial knowledge graph;

[0011] The time relationship information corresponding to each entity and event is obtained, and the initial knowledge graph is updated based on the time relationship information to obtain the target knowledge graph.

[0012] In a possible implementation, a preset large language model is used to extract relevant information from multiple analyzable units using a preset entity disambiguation method to obtain an initial knowledge graph, including:

[0013] By presetting a large language model based on the sensitivity and acceptance of entity ambiguity in the current scenario, the target disambiguation level corresponding to each entity is determined;

[0014] Determining a target disambiguation strategy corresponding to a target disambiguation level based on a preset correspondence between the disambiguation level and the disambiguation strategy;

[0015] Based on the target disambiguation strategy, the relevant information extracted from multiple analyzable units is disambiguated, and an initial knowledge graph is generated based on the processing results.

[0016] In a possible implementation, obtaining the time relationship information corresponding to each entity and event, and updating the initial knowledge graph based on the time relationship information to obtain the target knowledge graph includes:

[0017] The preset word segmentation technology is used to obtain the time relationship information corresponding to each entity and event; wherein the time relationship information includes entity appearance time, entity change time, entity state change time, event occurrence time, event duration, and event end time;

[0018] The temporal relationship information is added to the nodes of the initial knowledge graph to update the initial knowledge graph and obtain the target knowledge graph.

[0019] In a possible implementation, a preset large language model is used to extract relevant information from multiple analyzable units using a preset entity disambiguation method to construct a target knowledge graph, and the method further includes:

[0020] By using a preset large language model and a preset entity disambiguation method, relevant information is extracted from multiple analyzable units to obtain an initial knowledge graph;

[0021] Obtain the weight information corresponding to each entity and the relationship between entities, and generate the target knowledge graph based on the weight information and the initial knowledge graph.

[0022] In a possible implementation, obtaining weight information corresponding to each entity and the relationship between entities, and generating a target knowledge graph based on the weight information and the initial knowledge graph includes:

[0023] Obtain the weight information corresponding to each entity and the relationship between entities;

[0024] Based on the weight information corresponding to each entity, a first weight set of all nodes in the initial knowledge graph is obtained, and based on the weight information corresponding to the relationship between each entity, a second weight set of all edges in the initial knowledge graph is obtained;

[0025] Based on the first weight set, the second weight set and the initial knowledge graph, a target knowledge graph is generated.

[0026] In a possible implementation, based on the target knowledge graph and summary information, searching for a target query result corresponding to the current user query information includes:

[0027] Based on the target knowledge graph, summary information, and the first weight set, the target query result corresponding to the current user query information is searched from the target knowledge graph.

[0028] The second aspect of the present application provides a graph-based retrieval enhancement generation device, the device comprising:

[0029] A construction module is used to divide the input text into multiple analyzable units, and extract relevant information from the multiple analyzable units by using a preset large language model and a preset entity disambiguation method to construct a target knowledge graph; wherein the relevant information includes multiple entities, relationships between entities, and key statements, and the preset entity disambiguation method is obtained by a preset disambiguation level and a corresponding preset disambiguation strategy, the preset disambiguation level is determined according to the sensitivity and acceptance of entity ambiguity in the current scenario, and the preset disambiguation strategy is determined based on at least one preset entity detection technology;

[0030] A generation module is used to divide the target knowledge graph into communities using a preset hierarchical clustering method and generate summary information corresponding to each community;

[0031] The search module is used to search for target query results corresponding to the current user query information based on the target knowledge graph and summary information.

[0032] In a possible implementation manner, the above-mentioned building blocks are specifically used for:

[0033] By using a preset large language model and a preset entity disambiguation method, relevant information is extracted from multiple analyzable units to obtain an initial knowledge graph;

[0034] The time relationship information corresponding to each entity and event is obtained, and the initial knowledge graph is updated based on the time relationship information to obtain the target knowledge graph.

[0035] In a possible implementation, the above building blocks are further used to:

[0036] By presetting a large language model based on the sensitivity and acceptance of entity ambiguity in the current scenario, the target disambiguation level corresponding to each entity is determined;

[0037] Determining a target disambiguation strategy corresponding to a target disambiguation level based on a preset correspondence between the disambiguation level and the disambiguation strategy;

[0038] Based on the target disambiguation strategy, the relevant information extracted from multiple analyzable units is disambiguated, and an initial knowledge graph is generated based on the processing results.

[0039] In a possible implementation, the above building blocks are further used to:

[0040] The preset word segmentation technology is used to obtain the time relationship information corresponding to each entity and event; wherein the time relationship information includes entity appearance time, entity change time, entity state change time, event occurrence time, event duration, and event end time;

[0041] The temporal relationship information is added to the nodes of the initial knowledge graph to update the initial knowledge graph and obtain the target knowledge graph.

[0042] In a possible implementation, the above building blocks are further used to:

[0043] By using a preset large language model and a preset entity disambiguation method, relevant information is extracted from multiple analyzable units to obtain an initial knowledge graph;

[0044] Obtain the weight information corresponding to each entity and the relationship between entities, and generate the target knowledge graph based on the weight information and the initial knowledge graph.

[0045] In a possible implementation, the above building blocks are further used to:

[0046] Obtain the weight information corresponding to each entity and the relationship between entities;

[0047] Based on the weight information corresponding to each entity, a first weight set of all nodes in the initial knowledge graph is obtained, and based on the weight information corresponding to the relationship between each entity, a second weight set of all edges in the initial knowledge graph is obtained;

[0048] Based on the first weight set, the second weight set and the initial knowledge graph, a target knowledge graph is generated.

[0049] In a possible implementation manner, the search module is specifically used to:

[0050] Based on the target knowledge graph, summary information, and the first weight set, the target query result corresponding to the current user query information is searched from the target knowledge graph.

[0051] The third aspect of the present application proposes an electronic device, comprising a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the graph-based retrieval enhancement generation method as described in the first aspect.

[0052] The fourth aspect of the present application proposes a computer-readable storage medium, which stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set are loaded and executed by a processor to implement the graph-based retrieval enhancement generation method as described in the first aspect.

[0053] The embodiments of the present application have the following beneficial effects:

[0054] The embodiment of the present application provides a graph-based retrieval enhancement generation method, which includes: dividing the input text into multiple analyzable units, and extracting relevant information from the multiple analyzable units using a preset large language model using a preset entity disambiguation method, constructing a target knowledge graph, dividing the target knowledge graph into communities using a preset hierarchical clustering method, and generating summary information corresponding to each community, and searching for target query results corresponding to the current user query information based on the target knowledge graph and the summary information. This solution uses a preset entity disambiguation method to extract entities and relationships when constructing the target knowledge graph, and the structured information constructed helps to improve the precision and efficiency of retrieval, thereby improving the accuracy of retrieval and the coherence of semantic information. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 A block diagram of a computer device provided in an embodiment of the present application;

[0056] Figure 2 A flowchart of the steps of a graph-based retrieval enhancement generation method provided in an embodiment of the present application;

[0057] Figure 3 A flowchart of the steps for constructing a target knowledge graph provided in an embodiment of the present application;

[0058] Figure 4 A flowchart of the steps for obtaining an initial knowledge graph provided in an embodiment of the present application;

[0059] Figure 5 A flowchart of the steps for obtaining a target knowledge graph provided in an embodiment of the present application;

[0060] Figure 6 Another flowchart of steps for obtaining a target knowledge graph provided in an embodiment of the present application;

[0061] Figure 7 A flowchart of another step of obtaining a target knowledge graph provided in an embodiment of the present application;

[0062] Figure 8 A structural block diagram of a graph-based retrieval enhancement generation device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0063] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0064] In the following, the terms "first" and "second" are used for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present disclosure, unless otherwise specified, "multiple" means two or more. In addition, the use of "based on" or "according to" means openness and inclusiveness, because the process, steps, calculations or other actions "based on" or "according to" one or more of the conditions or values ​​may be based on additional conditions or values ​​beyond the described values ​​in practice.

[0065] The graph-based retrieval enhancement generation method provided in the present application can be applied to a computer device (electronic device), which can be a server or a terminal, wherein the server can be a single server or a server cluster composed of multiple servers. The embodiments of the present application do not specifically limit this. The terminal can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices.

[0066] Take the computer device as a server as an example. Figure 1 A block diagram of a server is shown, such as Figure 1As shown, the server may include a processor and a memory connected via a system bus. The processor of the server is used to provide computing and control capabilities. The memory of the server includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. When the computer program is executed by the processor, a graph-based retrieval enhancement generation method is implemented.

[0067] Those skilled in the art will understand that Figure 1 The structure shown in the figure is only a block diagram of a partial structure related to the solution of the present application, and does not constitute a limitation on the server to which the solution of the present application is applied. Optionally, the server may include more or fewer components than shown in the figure, or combine certain components, or have a different arrangement of components.

[0068] It should be noted that the execution subject of the embodiments of the present application may be a computer device or a graph-based retrieval enhancement generation device. The following method embodiments will be described using a computer device as the execution subject.

[0069] Figure 2 A flowchart of a graph-based search enhancement generation method provided in an embodiment of the present application. Figure 2 As shown, the method comprises the following steps:

[0070] Step 202: Segment the input text into multiple analyzable units, and extract relevant information from the multiple analyzable units using a preset large language model and a preset entity disambiguation method to construct a target knowledge graph.

[0071] Among them, the graph-based search enhancement generation method is also called GraphRAG. The traditional graph-based search enhancement generation method mainly uses a large language model for disambiguation, which is costly, destructive, and prone to information loss. In order to reduce costs, introduce non-destructive strategies, and improve the accuracy of retrieval and the coherence of semantic information, this application uses a preset large language model to perform disambiguation using a preset entity disambiguation method.

[0072] Optionally, the input text may be first divided into multiple analyzable units, and relevant information may be extracted from the multiple analyzable units using a preset entity disambiguation method through a preset large language model to construct a target knowledge graph. The analyzable unit, as a basic unit for subsequent processing, contains the extracted specific information.

[0073] Related information includes multiple entities, relationships between entities, and key statements. Entities can represent people, places, events, etc., while relationships describe the connections between these entities. The preset entity disambiguation method is obtained through the preset disambiguation level and the corresponding preset disambiguation strategy. The preset disambiguation level is determined according to the sensitivity and acceptance of the current scenario to entity ambiguity.

[0074] The preset disambiguation strategy is determined based on at least one preset entity detection technology, and can be achieved by integrating multiple preset entity detection technologies. For example, different weights can be set for different preset entity detection technologies, and a threshold can be set based on the weighted calculation of the final result, and whether the detection results need to be merged can be determined based on the detection results and the threshold.

[0075] In some optional embodiments, Figure 3 As shown, Figure 3 A flowchart of steps for constructing a target knowledge graph provided in an embodiment of the present application includes:

[0076] Step 302: extract relevant information from multiple analyzable units using a preset large language model and a preset entity disambiguation method to obtain an initial knowledge graph.

[0077] Among them, in some optional embodiments, such as Figure 4 As shown, Figure 4 A flowchart of steps for obtaining an initial knowledge graph provided in an embodiment of the present application includes:

[0078] Step 402: Determine the target disambiguation level corresponding to each entity by presetting a large language model based on the sensitivity and acceptance level of entity ambiguity in the current scenario.

[0079] Step 404: Determine a target disambiguation strategy corresponding to the target disambiguation level based on a preset correspondence between the disambiguation level and the disambiguation strategy.

[0080] Step 406: Disambiguate the relevant information extracted from the multiple analyzable units based on the target disambiguation strategy, and generate an initial knowledge graph based on the processing results.

[0081] Among them, the disambiguation level corresponding to the different sensitivity and acceptance levels of entity ambiguity in the current scene can be pre-set, and the corresponding disambiguation strategy can be set for each disambiguation level. For example, if the disambiguation level is L1, only the disambiguation strategy of the Natural Language Toolkit (NLTK) logic is used for disambiguation; if the disambiguation level is L2, the disambiguation strategy of NLTK and the Chinese word segmentation library, namely jieba, is used for disambiguation; if the disambiguation level is L3, the disambiguation strategy of the large language model is used for disambiguation.

[0082] Before setting up the disambiguation strategy, you can first collect public domain entity information, including external knowledge bases and multimodal data, to help identify and link the same entities mentioned in different texts. It should be noted that the collection here does not only refer to storing the data locally, but should also include the link relationship management of related data sets. Then collect proprietary entity information. Specifically, you can build an entity disambiguation library, including but not limited to different descriptions, references, and other information of the same entity, possible identical descriptions and references of different entities, their common contexts, and the specific objects of their entities in the context. During use, the entity disambiguation library can be updated in real time. Therefore, before the disambiguation process, the public domain entity information and proprietary entity information are integrated into the disambiguation knowledge set in the form of different weights for use in determining the disambiguation strategy.

[0083] Therefore, after determining the target disambiguation level corresponding to each entity and the target disambiguation strategy corresponding to the target disambiguation level, the relevant information extracted from multiple analyzable units can be disambiguated based on the target disambiguation strategy, and an initial knowledge graph can be generated based on the processing results.

[0084] Step 304: Obtain the time relationship information corresponding to each entity and event, and update the initial knowledge graph based on the time relationship information to obtain the target knowledge graph.

[0085] Among them, in some optional embodiments, such as Figure 5 As shown, Figure 5 A flowchart of steps for obtaining a target knowledge graph provided in an embodiment of the present application includes:

[0086] Step 502: Use a preset word segmentation technology to obtain the time relationship information corresponding to each entity and event.

[0087] Step 504: Add the time relationship information to the nodes of the initial knowledge graph to update the initial knowledge graph and obtain the target knowledge graph.

[0088] Among them, the traditional graph-based retrieval enhancement generation method has a deeper data understanding and generation capability, but does not yet have temporal reasoning capabilities. Temporal reasoning capabilities enable models or systems to understand and analyze time series data, or identify and understand time-related clues in text, which can more effectively track the evolution of events or ideas and improve information association and retrieval capabilities. Therefore, the target knowledge graph can be obtained by increasing temporal reasoning capabilities.

[0089] Specifically, the preset word segmentation technology can be Jieba word segmentation technology, so that the preset word segmentation technology can be used to extract the time relationship information corresponding to each entity and event. The time relationship information includes entity appearance time, entity change time, entity state change time, event occurrence time, event duration, and event end time.

[0090] Next, the time relationship information can be added to the nodes of the initial knowledge graph, and the events can be sorted in chronological order to form a timeline. Finally, the time dimension can be added to the initial knowledge graph, and the time sequence, duration and other relationships between events can be updated. This helps to understand the sequence and historical context of events, so as to update the initial knowledge graph and obtain the target knowledge graph.

[0091] In some other optional embodiments, Figure 6 As shown, Figure 6 Another flowchart of steps for obtaining a target knowledge graph provided in an embodiment of the present application includes:

[0092] Step 602: extract relevant information from multiple analyzable units using a preset large language model and a preset entity disambiguation method to obtain an initial knowledge graph.

[0093] Step 604: Obtain weight information corresponding to each entity and the relationship between entities, and generate a target knowledge graph based on the weight information and the initial knowledge graph.

[0094] Among them, the traditional graph-based retrieval enhancement generation method itself does not have the recall and retrieval considerations for different entity and relationship weights. In actual scenarios, it is possible to set different weights for weights and events of different importance to ensure that the possibility of recalling and retrieving the information during retrieval is increased.

[0095] Therefore, after obtaining the initial knowledge graph, we can first obtain the weight information corresponding to each entity and the relationship between entities, and generate the target knowledge graph based on the weight information and the initial knowledge graph.

[0096] Optionally, in some optional embodiments, such as Figure 7 As shown, Figure 7 Another flowchart of steps for obtaining a target knowledge graph provided in an embodiment of the present application includes:

[0097] Step 702: Obtain weight information corresponding to each entity and the relationship between entities.

[0098] Step 704: obtain a first weight set of all nodes in the initial knowledge graph based on the weight information corresponding to each entity, and obtain a second weight set of all edges in the initial knowledge graph based on the weight information corresponding to the relationship between each entity.

[0099] Step 706: Generate a target knowledge graph based on the first weight set, the second weight set and the initial knowledge graph.

[0100] The weight information of each entity may be determined based on the properties, importance, etc. of different entities. The weight information corresponding to the relationship between entities may be determined based on the strength, frequency, or other business logic of the relationship.

[0101] Thus, the first weight set of all nodes in the initial knowledge graph can be obtained based on the weight information corresponding to each entity, that is, the first weight set includes multiple node weights, and the first weight set can be recorded as w_node. The second weight set of all edges in the initial knowledge graph is obtained based on the weight information corresponding to the relationship between each entity, and the second weight set can be recorded as w_edge.

[0102] Then, a target knowledge graph can be generated based on the first weight set, the second weight set and the initial knowledge graph. In addition, a suitable storage method can be selected to store and manage the target knowledge graph with weights, for example, a graph database can be used to store, perform weight-based queries and path analysis, etc.

[0103] Step 204: divide the target knowledge graph into communities using a preset hierarchical clustering method, and generate summary information corresponding to each community.

[0104] The preset hierarchical clustering method may be the Leiden algorithm, so that the target knowledge graph can be divided into communities and the summary information corresponding to each community can be generated by the preset hierarchical clustering method, revealing the hierarchical structure in the graph, and at the same time, the graph embedding technology is used to enhance the expressiveness of the graph. Community detection helps to discover potential patterns and structures in the graph, while the graph embedding technology further improves the representation ability and interpretability of the graph.

[0105] Step 206: Based on the target knowledge graph and summary information, search for target query results corresponding to the current user query information.

[0106] Among them, in some optional embodiments, the target query results corresponding to the current user query information can be searched from the target knowledge graph based on the target knowledge graph, summary information, and the first weight set.

[0107] Optionally, after community division, each node belongs to a community, the node has a weight, and the community should also be given a corresponding weight. Community weight w_code(j) represents the weight of the jth community, w_code is the harmonic mean of the weights of all nodes in the community, and w_node(k) is the weight value of the kth node in the jth community.

[0108] The community weight is then used to correct the weights of each node within it. For example, the corrected weight of node k in community j is w_node_e(k)=w_code(j)*a0+w_node(k)*(1-a0), where a0 is the correction coefficient, which can be customized in advance.

[0109] Finally, the query capabilities of the graph database can be used to perform retrieval with weighted considerations. For example, graph algorithms such as Dijkstra or A* can be used to find the path with the highest weight, or the PageRank algorithm can be used to identify the most important nodes in the graph to obtain the target query results corresponding to the current user query information.

[0110] The present application provides a graph-based retrieval enhancement generation method, which includes: dividing the input text into multiple analyzable units, and extracting relevant information from the multiple analyzable units using a preset large language model using a preset entity disambiguation method, constructing a target knowledge graph, dividing the target knowledge graph into communities using a preset hierarchical clustering method, and generating summary information corresponding to each community, and searching for target query results corresponding to the current user query information based on the target knowledge graph and the summary information. This solution uses a preset entity disambiguation method to extract entities and relationships when constructing the target knowledge graph, and the structured information constructed helps to improve the accuracy and efficiency of retrieval, thereby improving the accuracy of retrieval and the coherence of semantic information.

[0111] Figure 8 A structural block diagram of a graph-based retrieval enhancement generation device provided in an embodiment of the present application.

[0112] like Figure 8 As shown, the graph-based retrieval enhancement generation device 800 includes:

[0113] Construction module 802 is used to divide the input text into multiple analyzable units, and extract relevant information from the multiple analyzable units by using a preset large language model and a preset entity disambiguation method to construct a target knowledge graph; wherein the relevant information includes multiple entities, relationships between entities and key statements, and the preset entity disambiguation method is obtained by a preset disambiguation level and a corresponding preset disambiguation strategy, the preset disambiguation level is determined based on the sensitivity and acceptance of the current scenario to entity ambiguity, and the preset disambiguation strategy is determined based on at least one preset entity detection technology.

[0114] The generation module 804 is used to divide the target knowledge graph into communities using a preset hierarchical clustering method and generate summary information corresponding to each community.

[0115] The search module 806 is used to search for target query results corresponding to the current user query information based on the target knowledge graph and summary information.

[0116] Regarding the device in the above embodiment, the specific manner in which each module performs the operation has been described in detail in the embodiment of the method, and will not be elaborated here. Each module in the above-mentioned graph-based retrieval enhancement generation device can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations of the above modules.

[0117] In one embodiment of the present application, a computer device is provided, the computer device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the following steps when executing the computer program:

[0118] The input text is divided into multiple analyzable units, and the preset entity disambiguation method is used to extract relevant information from the multiple analyzable units through the preset large language model to construct a target knowledge graph; wherein the relevant information includes multiple entities, relationships between entities and key statements, and the preset entity disambiguation method is obtained through a preset disambiguation level and a corresponding preset disambiguation strategy, the preset disambiguation level is determined according to the sensitivity and acceptance of the current scenario to entity ambiguity, and the preset disambiguation strategy is determined based on at least one preset entity detection technology;

[0119] The target knowledge graph is divided into communities through a preset hierarchical clustering method, and summary information corresponding to each community is generated;

[0120] Based on the target knowledge graph and summary information, search for the target query results corresponding to the current user query information.

[0121] In one embodiment of the present application, when the processor executes the computer program, the processor further implements the following steps:

[0122] By using a preset large language model and a preset entity disambiguation method, relevant information is extracted from multiple analyzable units to obtain an initial knowledge graph;

[0123] The time relationship information corresponding to each entity and event is obtained, and the initial knowledge graph is updated based on the time relationship information to obtain the target knowledge graph.

[0124] In one embodiment of the present application, when the processor executes the computer program, the processor further implements the following steps:

[0125] By presetting a large language model based on the sensitivity and acceptance of entity ambiguity in the current scenario, the target disambiguation level corresponding to each entity is determined;

[0126] Determining a target disambiguation strategy corresponding to a target disambiguation level based on a preset correspondence between the disambiguation level and the disambiguation strategy;

[0127] Based on the target disambiguation strategy, the relevant information extracted from multiple analyzable units is disambiguated, and an initial knowledge graph is generated based on the processing results.

[0128] In one embodiment of the present application, when the processor executes the computer program, the processor further implements the following steps:

[0129] The preset word segmentation technology is used to obtain the time relationship information corresponding to each entity and event; wherein the time relationship information includes entity appearance time, entity change time, entity state change time, event occurrence time, event duration, and event end time;

[0130] The temporal relationship information is added to the nodes of the initial knowledge graph to update the initial knowledge graph and obtain the target knowledge graph.

[0131] In one embodiment of the present application, when the processor executes the computer program, the processor further implements the following steps:

[0132] By using a preset large language model and a preset entity disambiguation method, relevant information is extracted from multiple analyzable units to obtain an initial knowledge graph;

[0133] Obtain the weight information corresponding to each entity and the relationship between entities, and generate the target knowledge graph based on the weight information and the initial knowledge graph.

[0134] In one embodiment of the present application, when the processor executes the computer program, the processor further implements the following steps:

[0135] Obtain the weight information corresponding to each entity and the relationship between entities;

[0136] Based on the weight information corresponding to each entity, a first weight set of all nodes in the initial knowledge graph is obtained, and based on the weight information corresponding to the relationship between each entity, a second weight set of all edges in the initial knowledge graph is obtained;

[0137] Based on the first weight set, the second weight set and the initial knowledge graph, a target knowledge graph is generated.

[0138] In one embodiment of the present application, when the processor executes the computer program, the processor further implements the following steps:

[0139] Based on the target knowledge graph, summary information, and the first weight set, the target query result corresponding to the current user query information is searched from the target knowledge graph.

[0140] The computer device provided in the embodiment of the present application has similar implementation principles and technical effects to those of the above-mentioned method embodiment, and will not be described in detail here.

[0141] In one embodiment of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0142] The input text is divided into multiple analyzable units, and the preset entity disambiguation method is used to extract relevant information from the multiple analyzable units through the preset large language model to construct a target knowledge graph; wherein the relevant information includes multiple entities, relationships between entities and key statements, and the preset entity disambiguation method is obtained through a preset disambiguation level and a corresponding preset disambiguation strategy, the preset disambiguation level is determined according to the sensitivity and acceptance of the current scenario to entity ambiguity, and the preset disambiguation strategy is determined based on at least one preset entity detection technology;

[0143] The target knowledge graph is divided into communities through a preset hierarchical clustering method, and summary information corresponding to each community is generated;

[0144] Based on the target knowledge graph and summary information, search for the target query results corresponding to the current user query information.

[0145] In one embodiment of the present application, when the computer program is executed by a processor, the following steps are further implemented:

[0146] By using a preset large language model and a preset entity disambiguation method, relevant information is extracted from multiple analyzable units to obtain an initial knowledge graph;

[0147] The time relationship information corresponding to each entity and event is obtained, and the initial knowledge graph is updated based on the time relationship information to obtain the target knowledge graph.

[0148] In one embodiment of the present application, when the computer program is executed by a processor, the following steps are further implemented:

[0149] By presetting a large language model based on the sensitivity and acceptance of entity ambiguity in the current scenario, the target disambiguation level corresponding to each entity is determined;

[0150] Determining a target disambiguation strategy corresponding to a target disambiguation level based on a preset correspondence between the disambiguation level and the disambiguation strategy;

[0151] Based on the target disambiguation strategy, the relevant information extracted from multiple analyzable units is disambiguated, and an initial knowledge graph is generated based on the processing results.

[0152] In one embodiment of the present application, when the computer program is executed by a processor, the following steps are further implemented:

[0153] The preset word segmentation technology is used to obtain the time relationship information corresponding to each entity and event; wherein the time relationship information includes entity appearance time, entity change time, entity state change time, event occurrence time, event duration, and event end time;

[0154] The temporal relationship information is added to the nodes of the initial knowledge graph to update the initial knowledge graph and obtain the target knowledge graph.

[0155] In one embodiment of the present application, when the computer program is executed by a processor, the following steps are further implemented:

[0156] By using a preset large language model and a preset entity disambiguation method, relevant information is extracted from multiple analyzable units to obtain an initial knowledge graph;

[0157] Obtain the weight information corresponding to each entity and the relationship between entities, and generate the target knowledge graph based on the weight information and the initial knowledge graph.

[0158] In one embodiment of the present application, when the computer program is executed by a processor, the following steps are further implemented:

[0159] Obtain the weight information corresponding to each entity and the relationship between entities;

[0160] Based on the weight information corresponding to each entity, a first weight set of all nodes in the initial knowledge graph is obtained, and based on the weight information corresponding to the relationship between each entity, a second weight set of all edges in the initial knowledge graph is obtained;

[0161] Based on the first weight set, the second weight set and the initial knowledge graph, a target knowledge graph is generated.

[0162] In one embodiment of the present application, when the computer program is executed by a processor, the following steps are further implemented:

[0163] Based on the target knowledge graph, summary information, and the first weight set, the target query result corresponding to the current user query information is searched from the target knowledge graph.

[0164] The computer-readable storage medium provided in this embodiment has similar implementation principles and technical effects to those of the above method embodiments, and will not be described in detail here.

[0165] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0166] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.

[0167] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A graph-based retrieval enhancement generation method, characterized in that: The method comprises: The input text is divided into a plurality of analyzable units, and a preset entity disambiguation method is used to extract relevant information from the plurality of analyzable units through a preset large language model to construct a target knowledge graph; wherein the relevant information includes a plurality of entities, relationships between entities, and key statements, and the preset entity disambiguation method is obtained through a preset disambiguation level and a corresponding preset disambiguation strategy, the preset disambiguation level is determined according to the sensitivity and acceptance of entity ambiguity in the current scenario, and the preset disambiguation strategy is determined based on at least one preset entity detection technology; Divide the target knowledge graph into communities using a preset hierarchical clustering method, and generate summary information corresponding to each community; Based on the target knowledge graph and the summary information, a target query result corresponding to the current user query information is searched.

2. The method according to claim 1, characterized in that The method of extracting relevant information from the plurality of analyzable units by using a preset large language model and a preset entity disambiguation method to construct a target knowledge graph includes: Extracting relevant information from the plurality of analyzable units using a preset large language model and a preset entity disambiguation method to obtain an initial knowledge graph; The time relationship information corresponding to each entity and event is obtained, and the initial knowledge graph is updated based on the time relationship information to obtain the target knowledge graph.

3. The method according to claim 2, characterized in that The extracting relevant information from the plurality of analyzable units by using a preset large language model and a preset entity disambiguation method to obtain an initial knowledge graph includes: Determine the target disambiguation level corresponding to each entity by presetting a large language model based on the sensitivity and acceptance of entity ambiguity in the current scenario; Determining a target disambiguation strategy corresponding to the target disambiguation level based on a preset correspondence between the disambiguation level and the disambiguation strategy; Based on the target disambiguation strategy, disambiguation processing is performed on the relevant information extracted from the multiple analyzable units, and the initial knowledge graph is generated according to the processing results.

4. The method according to claim 2 or 3, characterized in that: The acquiring of the time relationship information corresponding to each entity and event, and updating the initial knowledge graph based on the time relationship information to obtain the target knowledge graph, includes: The preset word segmentation technology is used to obtain the time relationship information corresponding to each entity and event; wherein the time relationship information includes entity appearance time, entity change time, entity state change time, event occurrence time, event duration, and event end time; The time relationship information is added to the nodes of the initial knowledge graph to update the initial knowledge graph to obtain the target knowledge graph.

5. The method according to claim 1 or 2, characterized in that: The extracting relevant information from the plurality of analyzable units by using a preset large language model and a preset entity disambiguation method to construct a target knowledge graph also includes: Extracting relevant information from the plurality of analyzable units using a preset large language model and a preset entity disambiguation method to obtain an initial knowledge graph; Obtain weight information corresponding to each of the entities and the relationships between the entities, and generate the target knowledge graph based on the weight information and the initial knowledge graph.

6. The method according to claim 5, characterized in that The obtaining of weight information corresponding to each of the entities and the relationship between the entities, and generating the target knowledge graph based on the weight information and the initial knowledge graph, includes: Obtaining weight information corresponding to each of the entities and the relationship between the entities; Based on the weight information corresponding to each of the entities, a first weight set of all nodes in the initial knowledge graph is obtained; based on the weight information corresponding to the relationship between each of the entities, a second weight set of all edges in the initial knowledge graph is obtained; Based on the first weight set, the second weight set and the initial knowledge graph, the target knowledge graph is generated.

7. The method according to claim 6, characterized in that The step of searching for a target query result corresponding to the current user query information based on the target knowledge graph and the summary information includes: Based on the target knowledge graph, the summary information, and the first weight set, the target query result corresponding to the current user query information is searched from the target knowledge graph.

8. A graph-based retrieval enhancement generation device, characterized in that: The device comprises: A construction module, used to divide the input text into multiple analyzable units, and extract relevant information from the multiple analyzable units by using a preset large language model and a preset entity disambiguation method to construct a target knowledge graph; wherein the relevant information includes multiple entities, relationships between entities, and key statements, and the preset entity disambiguation method is obtained by a preset disambiguation level and a corresponding preset disambiguation strategy, the preset disambiguation level is determined according to the sensitivity and acceptance of the current scenario to entity ambiguity, and the preset disambiguation strategy is determined based on at least one preset entity detection technology; A generation module, used to divide the target knowledge graph into communities by a preset hierarchical clustering method, and generate summary information corresponding to each community; A search module is used to search for a target query result corresponding to the current user query information based on the target knowledge graph and the summary information.

9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the graph-based retrieval enhancement generation method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the graph-based retrieval enhancement generation method as described in any one of claims 1-7.

Citation Information

Cited By

  • Index system creation method and device, equipment, medium and program product

    CN120822883A

  • Enhanced retrieval generation optimization method fusing knowledge graph

    CN121350280A