Knowledge Graph-Based Information Retrieval Method, Apparatus, Electronic Device, and Medium

By constructing the target knowledge graph in the operation and maintenance scenario of new energy power generation equipment, using edge frequency and event description information, combining vector matching technology and improved PageRank algorithm, the problem of multi-hop reasoning and information overload of the Q&A system is solved, and high-precision information retrieval is achieved.

CN120196732BActive Publication Date: 2025-07-29NANJING HUADUN ELECTRIC POWER INFORMATION SAFETY EVALUATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510677184.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-07-29
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

The existing question-and-answer system is difficult to deal with multi-hop reasoning problems in the operation and maintenance scenarios of new energy power generation equipment, resulting in insufficient accuracy and correlation of answers, and the inability to dynamically adjust the search results, resulting in information overload.

Method used

Build a target knowledge graph, including edge frequency and event description information, form triples through entity recognition and relationship extraction, combine vector matching technology and improved PageRank algorithm to determine the China Unicom sub-graph and edge weight, and accurately retrieve information.

Benefits of technology

Effectively dealing with the problems of multi-hop reasoning and dispersed knowledge fragments has improved the accuracy of information retrieval and the overall performance of the question-and-answer system, and improved the accuracy of information retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196732B_ABST
    Figure CN120196732B_ABST
Patent Text Reader

Abstract

The present invention discloses an information retrieval method, device, electronic device and medium based on a knowledge graph. The method includes: constructing a target knowledge graph, which is constructed in the form of triples and includes edge frequencies and event description information; the edge frequency is the number of occurrences of the same edge relationship; the event description information is the semantic coherence description information for the triples; in response to an input operation of preset information, extracting at least one keyword and at least one target entity of the preset information; determining a connected subgraph from the target knowledge graph according to the association relationship between the target entity and each entity in the target knowledge graph; and determining retrieval information for the preset information according to the keyword, the edge frequency of the connected subgraph and the event description information of the connected subgraph. The present invention combines the edge frequency of the connected subgraph and the event description information of the connected subgraph to perform information retrieval on the preset information, improving the accuracy of information retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of knowledge graph construction and query, and particularly to an information retrieval method, device, electronic device and medium based on a knowledge graph. Background Art

[0002] In the knowledge Q&A scenarios of documents such as equipment manuals and operation manuals in the operation and maintenance Q&A scenarios of new energy power generation equipment, it is a key link to improve equipment operation efficiency and maintenance capabilities.

[0003] However, most of the existing Q&A systems rely on large language models (LLMs) to directly generate answers, and this method has obvious limitations in dealing with complex questions. For example, for multi-hop reasoning questions, the knowledge fragments are scattered and difficult to integrate, resulting in insufficient accuracy and relevance of the answers. In addition, the existing Q&A systems cannot dynamically adjust the retrieval results, are prone to information overload, and cannot meet the requirements of accurate information retrieval in actual application scenarios. Summary of the Invention

[0004] The present invention provides an information retrieval method, device, electronic device and medium based on a knowledge graph to improve the accuracy of information retrieval.

[0005] According to one aspect of the present invention, there is provided an information retrieval method based on a knowledge graph, the method comprising:

[0006] Constructing a target knowledge graph, the target knowledge graph being constructed in the form of triples, the target knowledge graph including edge frequency and event description information; the edge frequency being the number of times the same edge relationship appears; the event description information being semantic coherence description information for the triples;

[0007] In response to an input operation of preset information, extracting at least one keyword and at least one target entity of the preset information; the preset information has an associated relationship with the target knowledge graph; the keyword is information describing the key information of the target entity and / or information describing the relationship between the target entities;

[0008] Determining a connected subgraph from the target knowledge graph according to the associated relationship between the target entity and each entity in the target knowledge graph;

[0009] Determining retrieval information for the preset information according to the keyword, the edge frequency of the connected subgraph and the event description information of the connected subgraph.

[0010] According to another aspect of the present invention, there is provided an information retrieval device based on a knowledge graph, the device comprising:

[0011] A graph construction module for constructing a target knowledge graph, where the target knowledge graph is constructed in the form of triples, and the target knowledge graph includes edge frequencies and event description information; the edge frequency is the number of times the same edge relationship appears; the event description information is semantic coherence description information for the triples.

[0012] An information extraction module for extracting at least one keyword and at least one target entity of the preset information in response to an input operation of the preset information; the preset information has an association relationship with the target knowledge graph; the keyword is information for describing the key information of the target entity and / or information for describing the relationship between the target entities.

[0013] A graph determination module for determining a connected subgraph from the target knowledge graph according to the association relationship between the target entity and each entity in the target knowledge graph.

[0014] An information determination module for determining retrieval information for the preset information according to the keyword, the edge frequency of the connected subgraph, and the event description information of the connected subgraph.

[0015] According to another aspect of the present invention, there is provided an electronic device, which includes:

[0016] At least one processor; and

[0017] A memory communicatively connected to the at least one processor; wherein,

[0018] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor can execute the knowledge graph-based information retrieval method according to any embodiment of the present invention.

[0019] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the knowledge graph-based information retrieval method according to any embodiment of the present invention when executed.

[0020] In the technical solution of the embodiment of the present invention, a target knowledge graph is constructed. The target knowledge graph is constructed in the form of triples and includes edge frequencies and event description information. The edge frequency is the number of occurrences of the same edge relationship, and the event description information is the semantic coherence description information for the triples. The target knowledge graph of the present invention includes edge frequencies and event description information, so that when the target knowledge graph is used for information retrieval subsequently, the corresponding triples can be located more accurately. In response to the input operation of preset information, at least one keyword and at least one target entity of the preset information are extracted. There is an association relationship between the preset information and the target knowledge graph. The keyword is the information describing the key information of the target entity and / or the information describing the relationship between each target entity. The extraction of the target entity facilitates accurately determining the connected subgraph from the target knowledge graph according to the association relationship between the target entity and each entity in the target knowledge graph. Further, according to the keyword, the edge frequency of the connected subgraph, and the event description information of the connected subgraph, the retrieval information for the preset information is determined, which can effectively handle the problems of multi-hop reasoning and scattered knowledge fragments, improve the overall performance of the question-and-answer system, and improve the accuracy of information retrieval.

[0021] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0023] Figure 1 is a flowchart of an information retrieval method based on a knowledge graph according to an embodiment of the present invention;

[0024] Figure 2 is an example diagram of a knowledge graph applicable to an embodiment of the present invention;

[0025] Figure 3 is a flowchart of another information retrieval method based on a knowledge graph according to an embodiment of the present invention;

[0026] Figure 4 is a flowchart of another information retrieval method based on a knowledge graph according to an embodiment of the present invention;

[0027] Figure 5It is a schematic structural diagram of an information retrieval device based on a knowledge graph according to an embodiment of the present invention;

[0028] Figure 6 It is a schematic structural diagram of an electronic device for implementing the information retrieval method based on a knowledge graph according to an embodiment of the present invention. Detailed implementation manners

[0029] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0030] It should be noted that the terms "first", "second", "third", "reference", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0031] Embodiment 1

[0032] Figure 1 It is a flowchart of an information retrieval method based on a knowledge graph according to an embodiment of the present invention. This embodiment is applicable to the situation of combining a large language model and a knowledge graph for information retrieval. This method can be executed by an information retrieval device based on a knowledge graph. The information retrieval device based on a knowledge graph can be implemented in the form of hardware and / or software, and the information retrieval device based on a knowledge graph can be configured in any electronic device with network communication functions. As Figure 1 shown, the information retrieval method based on a knowledge graph of the present invention may include the following steps:

[0033] S110. Construct a target knowledge graph. The target knowledge graph is constructed in the form of triples. The target knowledge graph includes edge frequency and event description information. The edge frequency is the number of times the same edge relationship appears. The event description information is the semantic coherence description information for the triples.

[0034] Among them, a triple consists of a head entity, an edge relation, and a tail entity, and the relation is the relationship between the head entity and the tail entity. For example, as Figure 2 shown, G can be the head entity, C is the tail entity, and relation 1 is the edge relation of the triple containing G and C; the connection between G and C is the edge of this triple. For Figure 2 a knowledge graph with three edges for the triple of relation 1, the edge frequency of the edge corresponding to relation 1 is 3. The event description information is used to enhance the semantic coherence of the triple and strengthen the accuracy of vector retrieval. The edge relation is determined by the edge attribute, and the edge relation can be understood as being used to describe the relationship between the head entity and the tail entity in the triple; for example, the edge relation is "belong to", which can be understood as the head entity belonging to the tail entity.

[0035] Specifically, obtain the data to be processed. The data to be processed can be text information, structured data, etc. in various scenarios that can be used to construct a knowledge graph. One type of scenario corresponds to one knowledge graph. For example, for the scenario of new energy power generation equipment operation and maintenance Q&A, the data to be processed can be information such as equipment manuals, operation manuals, equipment documents, and equipment parameter tables. The equipment documents can be PDF files and / or plain text. Perform entity recognition and relation extraction on the data to be processed to obtain structured triples, and identify the edge frequency and event description information of the triples, so as to construct a target knowledge graph based on the structured triples and record the edge frequency and event description information of the triples in the target knowledge graph for subsequent use.

[0036] Furthermore, constructing the target knowledge graph can also include: performing entity recognition and relation extraction on the data to be processed based on a data extraction model to obtain structured triples; the data extraction model is a large language model; specifically, the data to be processed can be preprocessed first to obtain the preprocessed data to be processed. Preprocessing can be understood as performing basic semantic screening on the information. For example, preprocessing can include text extraction and word segmentation operations, etc. Further, the data extraction model can be guided by prompt words to perform entity recognition and relation extraction on the preprocessed data to be processed, so as to accurately obtain structured triples, so as to construct a target knowledge graph based on the structured triples and record the edge frequency and event description information of the triples in the target knowledge graph. Among them, the prompt words can be guiding information for guiding the content of this data extraction. For example, if extracting entities, the prompt words can be "extract entities"; for example, for the scenario of new energy power generation equipment operation and maintenance Q&A, the entities can be specific things or abstract concepts such as equipment names, equipment types, operation parameters, and fault codes. If extracting the relationship between entities, the prompt words can be "extract the relationship between entities".

[0037] For example, in the scenario of Q&A for the operation and maintenance of new energy power generation equipment, the process of constructing the target knowledge graph can be as follows: For the obtained data to be processed, for the equipment documents in the data to be processed, the equipment documents are preprocessed to obtain the document content, so as to ensure that the text content can be effectively understood and processed by the data extraction model. Subsequently, according to the predefined prompt template, the document content is input into the data extraction model. This prompt template is optimized and designed to guide the data extraction model to identify entities in the document, such as equipment names, equipment types, operating parameters, fault codes, etc., and further analyze the relationships between these entities. For example, through the guidance of a specific prompt template, the data extraction model can identify the "connection relationship" or "operation sequence relationship" between "Equipment A" and "Equipment B", and output these relationships in a structured form.

[0038] At the same time, for the structured data (such as equipment parameter tables) in the data to be processed, it can be converted into a unified text format and comprehensively analyzed in combination with the information in the equipment documents. Through the guidance of the prompt template, the data extraction model can extract the key parameters of the equipment (such as model, power, rated voltage, etc.) from the structured data as node attributes and associate them with the entities extracted from the document. For example, if the structured data records that the "rated power of Equipment A" is "10kW", this information will be extracted and stored as an attribute of the "Equipment A" node.

[0039] Furthermore, after the entity and edge relationship extraction is completed, the entities and edge relationships are combined into triples, thereby constituting the target value knowledge graph in the scenario of Q&A for the operation and maintenance of new energy power generation equipment according to the triples. The attributes of the nodes and the edge relationships can be used to determine the event description information of the triples to enhance the semantic coherence of the triples and strengthen the accuracy of vector retrieval.

[0040] S120. In response to the input operation of the preset information, extract at least one keyword and at least one target entity of the preset information; the preset information has an associated relationship with the target knowledge graph; the keyword is information that describes the key information of the target entity and / or information that describes the relationship between each target entity.

[0041] Among them, the preset information can be the prompt content of the information that the user needs to retrieve, for example, the preset information can be a question input by the user; in the scenario of Q&A for the operation and maintenance of new energy power generation equipment, the preset information can be questions such as querying the operating status of a certain equipment. The preset information corresponds to the target knowledge graph. For example, in the scenario of Q&A for the operation and maintenance of new energy power generation equipment, both the preset information and the target knowledge graph are related to the scenario of Q&A for the operation and maintenance of new energy power generation equipment.

[0042] Specifically, the data extraction model extracts entities and keywords from the preset information, so as to obtain at least one keyword and at least one target entity.

[0043] S130. Determine a connected subgraph from the target knowledge graph according to the association relationship between the target entity and each entity in the target knowledge graph.

[0044] Specifically, the vector matching technology is used to determine the first semantic matching degree between the target entity and each entity in the target knowledge graph, and the entity in the target knowledge graph with the first semantic matching degree greater than the preset matching degree is used as a reference entity, so as to form a connected subgraph from each reference entity in the target knowledge graph, ensuring that each entity in the connected subgraph is highly relevant to the preset information, thereby ensuring the accuracy of information retrieval.

[0045] In this embodiment, optionally, determining a connected subgraph from the target knowledge graph according to the association relationship between the target entity and each entity in the target knowledge graph includes steps A1 - A3:

[0046] Step A1. Determine the first association index between the target entity and each entity in the target knowledge graph.

[0047] Among them, the first association index is used to describe the association degree between the target entity and each entity in the target knowledge graph.

[0048] Specifically, the vector matching technology is used to determine the first association index between the target entity and each entity in the target knowledge graph.

[0049] Step A2. Use the entity in the target knowledge graph corresponding to the first association index exceeding the preset association index in the first association index as the starting node, and set the initial node weight of the starting node to the first preset weight.

[0050] Among them, the first preset weight can be set according to the actual requirements of the scenario. For example, the first preset weight can be set to 1. The number of starting nodes is associated with the number of target entities; generally, the number of starting nodes is the same as the number of target entities. Because the number of starting nodes is associated with the number of target entities, the determination of the preset association index can be based on the number of target entities. The initial node weight can be understood as the node weight initially set for the node.

[0051] Step A3. On the target knowledge graph, start from the starting node and expand a preset number of times to determine the connected subgraph, and set the initial node weight of other nodes except the starting node in the connected subgraph to the second preset weight; each entity in the connected subgraph is a different node; the first preset weight is greater than the second preset weight.

[0052] Among them, the preset level can be understood as the level expanded starting from the corresponding starting node on the target knowledge graph; that is, the starting node is connected to the preset layer nodes, the preset layer is equal to the prediction level, and the preset layer includes one less node between the starting node and the farthest node; that is, the connected subgraph includes all the nodes between the starting node and the farthest node on the target knowledge graph, and the nodes other than the starting node in the connected subgraph are directly or indirectly connected to the starting node. The preset level can be set according to the actual needs of the scenario. If the preset level is 3, then the starting node is connected to 3 layers of nodes.

[0053] For example, Figure 2 As shown, node B is the starting node and the preset level is 2. Then the first-layer nodes associated with node B are the nodes directly connected to node B, that is, the first-layer nodes associated with node B are node A, node G, node K, node L, and node M; the second-layer nodes associated with node B are the nodes with one node between them and node B, that is, the second-layer nodes associated with node B are node D, node E, node F, node R, node Q, node N, node O, and node P. The nodes included in the connected subgraph are node B, node A, node G, node K, node L, node M, node D, node E, node F, node R, node Q, node N, node O, and node P.

[0054] In the technical solution of this embodiment, the first association index between the target entity and each entity in the target knowledge graph is determined, and the association degree between the target entity and each entity in the target knowledge graph is accurately determined. Then, the entity in the target knowledge graph corresponding to the first association index that exceeds the preset association index in the first association index is used as the starting node, and the initial node weight of the starting node is set to the first preset weight to ensure a strong association between the starting node and the target entity; then, on the target knowledge graph, starting from the starting node, expand the preset level to determine the connected subgraph to ensure that the nodes in the connected subgraph have as high an association degree with the target entity as possible; the initial node weights of the other nodes in the connected subgraph except the starting node are set to the second preset weight. The first preset weight is greater than the second preset weight. The setting of the first preset weight and the second preset weight is to initialize the node weights of each node to facilitate the subsequent accurate retrieval requirements. The first preset weight is greater than the second preset weight because the starting node has the highest association degree with the target entity to ensure accuracy.

[0055] S140. Determine the retrieval information for the preset information according to the keyword, the edge frequency of the connected subgraph, and the event description information of the connected subgraph.

[0056] Specifically, determine the second semantic matching degree between the keyword and each event description information of the connected subgraph. Based on the second semantic matching degree and the edge frequency of the connected subgraph, determine the correlation degree of each edge in the connected subgraph. Take the edges in the connected subgraph whose correlation degree is greater than the preset correlation degree as target edges, and the triples corresponding to the target edges as the retrieval information of the preset information.

[0057] S130 and S140 of the present invention can be the matching of integrating the correlation relationship between the target entity and each entity in the target knowledge graph into the PageRank algorithm; the process of calculating keywords, edge frequencies of edges, and event description information, thereby forming an improved PageRank algorithm. Thus, in steps S130 and S140, the improved PageRank algorithm can be used for calculation to accurately determine the retrieval information for the preset information.

[0058] The technical solution of the embodiment of the present invention constructs a target knowledge graph, which is constructed in the form of triples. The target knowledge graph includes edge frequency and event description information; the edge frequency is the number of occurrences of the same edge relationship; the event description information is the semantic coherence description information for the triples. The target knowledge graph of the present invention includes edge frequency and event description information, so that when using the target knowledge graph for information retrieval later, it can more accurately locate the corresponding triples. In response to the input operation of the preset information, extract at least one keyword and at least one target entity of the preset information; there is an association relationship between the preset information and the target knowledge graph; the keyword is the information describing the key information of the target entity and / or the information describing the relationship between each target entity; the extraction of the target entity facilitates accurately determining the connected subgraph from the target knowledge graph according to the association relationship between the target entity and each entity in the target knowledge graph. Further, according to the keyword, the edge frequency of the connected subgraph, and the event description information of the connected subgraph, determine the retrieval information for the preset information, which can effectively handle the problems of multi-hop reasoning and scattered knowledge fragments, improve the overall performance of the question-and-answer system, and improve the accuracy of information retrieval.

[0059] Embodiment 2

[0060] Figure 3 It is a flowchart of another information retrieval method based on a knowledge graph provided by the embodiment of the present invention. The technical solution of this embodiment further optimizes the process of S140 in the foregoing embodiment on the basis of the foregoing embodiment. This embodiment can be combined with each optional solution in the above one or more embodiments. As Figure 3 shown, the information retrieval method based on a knowledge graph of the present invention may include the following steps:

[0061] S210. Construct a target knowledge graph, which is constructed in the form of triples and includes edge frequencies and event description information. The edge frequency is the number of occurrences of the same edge relationship. The event description information is the semantic coherence description information for the triples.

[0062] S220. In response to the input operation of preset information, extract at least one keyword and at least one target entity of the preset information. There is an associated relationship between the preset information and the target knowledge graph. The keyword is the information that describes the key information of the target entity and / or the information that describes the relationship between each target entity.

[0063] S230. According to the association relationship between the target entity and each entity in the target knowledge graph, determine the connected subgraph from the target knowledge graph.

[0064] S240. Determine the second association index between the keyword and each event description information of the connected subgraph. Each edge in the connected subgraph corresponds to an event description information. The second association index is used to describe the association degree between the keyword and each event description information of the connected subgraph.

[0065] Specifically, convert the keyword into a first vector through a language model, and convert each event description information of the connected subgraph into a second vector through the language model. The language model can be understood as a model that converts text into a vector form. Further, based on the first vector and the second vectors corresponding to each event description information, determine the second association index for each edge.

[0066] Correspondingly, based on the first vector and the second vectors corresponding to each event description information , determine the second association index CR for each edge, which can be determined by the following formula:

[0067] ;

[0068] Among them, the normalization processing of the softmax function can make the edge weights of each node determined subsequently form a probability distribution, which not only retains the semantic association strength but also conforms to the probability characteristics of the random walk algorithm.

[0069] S250. Based on the edge frequency and the second association index corresponding to each edge of the connected subgraph, determine the edge weight of each edge in the connected subgraph.

[0070] Specifically, establish a corresponding relationship between the edge frequency and the second association index and the edge weight. After obtaining the edge frequency and the second association index corresponding to each edge, the edge weight corresponding to the edge frequency and the second association index corresponding to each edge can be queried from the corresponding relationship, so that the edge weight of each edge in the connected subgraph can be accurately obtained.

[0071] In this embodiment, optionally, based on the edge frequency corresponding to each edge of the connected subgraph and the second association index, the edge weight of each edge in the connected subgraph is determined, including: determining the first proportion of the edge frequency corresponding to each edge of the connected subgraph, and the second proportion of the second association index corresponding to each edge of the connected subgraph; the sum of the first proportion and the second proportion is 1; based on the first proportion α, the second proportion β, the edge frequency and the second association index CR corresponding to each edge of the connected subgraph, the edge weight w of each edge in the connected subgraph is determined ij , which can be determined by the following formula:

[0072] .

[0073] In the technical solution of this embodiment, the first proportion and the second proportion can control the contribution ratio of the frequency factor and the semantic relevance. The second association index is dynamically adjusted according to the strength of the semantic association with the keyword, so that the result is related to the preset information, thus ensuring the accuracy of the edge weight determination.

[0074] S260. Based on the edge weights of each edge in the connected subgraph, the retrieval information for the preset information is determined.

[0075] Specifically, the edges corresponding to the edge weights greater than the preset weight value in the edge weights are used as target edges, and the triples corresponding to the target edges are used as the retrieval information for the preset information.

[0076] In the technical solution of the embodiment of the present invention, after determining the connected subgraph from the target knowledge graph according to the association relationship between the target entity and each entity in the target knowledge graph, the second association index between the keyword and each event description information of the connected subgraph is further determined; each edge in the connected subgraph corresponds to an event description information; the second association index is used to describe the association degree between the keyword and each event description information of the connected subgraph, and accurately determine the strength of the semantic association between the keyword and each event description information of the connected subgraph, so that the subsequent information retrieval result is related to the preset information. Then, based on the edge frequency corresponding to each edge of the connected subgraph and the second association index, the edge weight of each edge in the connected subgraph is determined, ensuring that the edge weights of the edges with higher edge frequencies and second association indexes are higher, and finding the edges that are more suitable for the preset information. Therefore, based on the edge weights of each edge in the connected subgraph, the retrieval information for the preset information can be accurately determined, effectively dealing with the problems of multi-hop reasoning and scattered knowledge fragments, improving the overall performance of the question-answering system, and improving the accuracy of information retrieval.

[0077] Embodiment III

[0078] Figure 4The flowchart of another information retrieval method based on a knowledge graph provided by an embodiment of the present invention. The technical solution of this embodiment further optimizes the process of determining retrieval information for preset information based on the edge weights of each edge in the connected subgraph on the basis of the foregoing embodiment. This embodiment can be combined with various alternative solutions in one or more of the foregoing embodiments. As Figure 4 shown, the information retrieval method based on a knowledge graph of the present invention may include the following steps:

[0079] S310. Construct a target knowledge graph, which is constructed in the form of triples. The target knowledge graph includes edge frequencies and event description information; the edge frequency is the number of times the same edge relationship appears; the event description information is the semantic coherence description information for the triples.

[0080] S320. In response to an input operation of preset information, extract at least one keyword and at least one target entity of the preset information; there is an association relationship between the preset information and the target knowledge graph; the keyword is information that describes the key information of the target entity and / or information that describes the relationship between each target entity.

[0081] S330. Determine the first association index between the target entity and each entity in the target knowledge graph, and use the entity in the target knowledge graph corresponding to the first association index that exceeds the preset association index as the starting node, and set the initial node weight of the starting node to the first preset weight.

[0082] S340. On the target knowledge graph, expand a preset number of times starting from the starting node to determine a connected subgraph, and set the initial node weights of other nodes in the connected subgraph except the starting node to the second preset weight; each entity in the connected subgraph is a different node; the first preset weight is greater than the second preset weight.

[0083] S350. Determine the second association index between the keyword and each event description information of the connected subgraph, and determine the edge weight of each edge in the connected subgraph based on the edge frequency corresponding to each edge of the connected subgraph and the second association index; each edge in the connected subgraph corresponds to an event description information; the second association index is used to describe the degree of association between the keyword and each event description information of the connected subgraph.

[0084] S360. Normalize the edge weights of each edge in the connected subgraph to obtain the normalized edge weights, and obtain the dynamic transition matrix of the connected subgraph based on the normalized edge weights.

[0085] Among them, the dynamic transition matrix can be understood as the correlation score between each node. The dynamic transition matrix is essentially a stochastic matrix, that is, the dynamic transition matrix will also be updated synchronously during the process of updating the node weights of subsequent nodes. The dynamic transition matrix can be expressed as:

[0086] .

[0087] S370. Based on the dynamic transition matrix of the connected subgraph, update the node weight of the first node in the connected subgraph, and update the dynamic transition matrix of the connected subgraph according to the cumulative value of the edge weights of the edges associated with the first node; the first node is the node whose node weight is updated in the current round; the node weight of one node is updated in each round.

[0088] Specifically, traverse the nodes directly connected to the first node in the connected subgraph as the third nodes associated with the first node, and use the normalized edge weight of the edge connecting the third node and the first node determined from the dynamic transition matrix as the reference edge weight; further determine the node weight of the first node according to the reference edge weight.

[0089] Further, updating the dynamic transition matrix of the connected subgraph according to the cumulative value of the edge weights of the edges associated with the first node may include: determining the reference edge whose cumulative weight value of the edge weight of the reference edge is greater than the second preset edge weight as the first reference edge; the cumulative process of the edge weights of the reference edges is accumulated in the order of the edge weights from large to small; the edge weight of the first reference edge remains unchanged, and update the edge weight of the second reference edge to the third preset edge weight; the second reference edge is the reference edge except the first reference edge among the reference edges; the third preset edge weight is lower than the edge weight of the second reference edge before update; update the dynamic transition matrix of the connected subgraph according to the edge weight of the first reference edge and the edge weight of the second reference edge.

[0090] As an optional embodiment, updating the dynamic transition matrix of the connected subgraph according to the cumulative value of the edge weights of the edges associated with the first node further includes: determining the edges associated with the first node as reference edges, and obtaining the edge weights of the reference edges; the reference edges are all the edges directly connected to the first node; determining the reference edge whose edge weight is greater than the first preset edge weight among the edge weights of the reference edges as the first reference edge, and / or determining the reference edge whose cumulative weight value of the edge weight of the reference edge is greater than the second preset edge weight as the first reference edge; the cumulative process of the edge weights of the reference edges is accumulated in the order of the edge weights from large to small; further, the edge weight of the first reference edge remains unchanged, and update the edge weight of the second reference edge to the third preset edge weight; the second reference edge is the reference edge except the first reference edge among the reference edges; the third preset edge weight is lower than the edge weight of the second reference edge before update; finally, update the dynamic transition matrix of the connected subgraph according to the edge weight of the first reference edge and the edge weight of the second reference edge.

[0091] Specifically, according to the edge weights of the first reference edge and the second reference edge, updating the dynamic transition matrix of the connected subgraph may include: normalizing the edge weights of the first reference edge and the second reference edge in the connected subgraph to obtain the normalized edge weights corresponding to the first reference edge and the second reference edge, and obtaining the dynamic transition matrix of the connected subgraph based on the normalized edge weights.

[0092] In the embodiment of the present invention, a reference edge with an edge weight greater than the first preset edge weight among the edge weights of the reference edges is determined as the first reference edge, and / or a reference edge with a cumulative weight value of the edge weights of the reference edges greater than the second preset edge weight is determined as the first reference edge, ensuring the accuracy of the determination of the first reference edge and avoiding omission.

[0093] In this embodiment, optionally, based on the dynamic transition matrix of the connected subgraph, updating the node weight of the first node in the connected subgraph includes steps B1 - B3:

[0094] Step B1: Determine the first weight of the first node according to the cosine similarity between the preset information and the first node.

[0095] Specifically, the preset information is converted into a third vector through a language model, and the first node is converted into a fourth vector through the language model. The language model can be understood as a model that converts text into a vector form; further, the cosine similarity between the third vector and the fourth vector is determined to determine the first weight of the first node.

[0096] Step B2: Determine the third nodes associated with the first node in the connected subgraph, and the number of nodes of the third nodes, and determine the second weight of each third node; the third node is a node directly connected to the first node.

[0097] Specifically, determining the second weight of each third node includes: if the node weight of the third node has been updated, then using the updated node weight of the third node as the second weight of the third node; if the node weight of the third node has not been updated, then using the initial node weight of the third node as the second weight of the third node.

[0098] Step B3: Update the node weight of the first node in the connected subgraph according to the first weight of the first node, the dynamic transition matrix of the connected subgraph, the number of nodes of the third nodes, and the second weight of the third nodes.

[0099] Specifically, the updated node weight of the first node in the connected subgraph can be determined by the following formula:

[0100] ;

[0101] Wherein, is the updated node weight of the first node, v iis the first weight of the first node, is the normalized edge weight of the edge connecting the first node and the third node in the dynamic transition matrix of the connected subgraph, Score(T j ) is the second weight of the third node, C(T j ) is the number of nodes of the third node, and d is the proportional adjustment coefficient.

[0102] In the embodiment of the present invention, according to the cosine similarity between the preset information and the first node, the first weight of the first node is determined. The third node associated with the first node and the number of nodes of the third node are determined from the connected subgraph, and the second weight of each third node is determined; the third node is the node directly connected to the first node. According to the first weight of the first node, the dynamic transition matrix of the connected subgraph, the number of nodes of the third node, and the second weight of the third node, the node weight of the first node in the connected subgraph is updated to ensure that the semantic features of the preset information can be accurately adapted during the retrieval process, so as to improve the accuracy of the retrieval.

[0103] S380. When the dynamic transition matrix converges, obtain the node weights of each node in the connected subgraph; use the first node corresponding to the node weight greater than the preset node weight in the node weights as the second node, and use the triple corresponding to the second node as the retrieval information.

[0104] In the technical solution of the embodiment of the present invention, after determining the edge weights of each edge in the connected subgraph, the edge weights of each edge in the connected subgraph are normalized to obtain the normalized edge weights, and the dynamic transition matrix of the connected subgraph is obtained based on the normalized edge weights. Further, based on the dynamic transition matrix of the connected subgraph, the node weight of the first node in the connected subgraph is updated, and the dynamic transition matrix of the connected subgraph is updated according to the cumulative value of the edge weights of the edges associated with the first node; the first node is the node whose node weight is updated in the current round; the node weight of one node is updated in each round; so as to update the dynamic transition matrix in real time, reduce unnecessary calculations, improve the operation efficiency of the system, and make it more suitable for the real-time retrieval scenario of large-scale knowledge graphs. When the dynamic transition matrix converges, that is, after multiple rounds of iteration, the retrieval information that better matches the preset information is queried. Further, obtain the node weights of each node in the connected subgraph; use the first node corresponding to the node weight greater than the preset node weight in the node weights as the second node, and use the triple corresponding to the second node as the retrieval information, effectively handling the problems of multi-hop reasoning and scattered knowledge fragments, improving the overall performance of the question-answering system, and improving the accuracy of information retrieval.

[0105] Embodiment 4

[0106] Figure 5The figure is a schematic structural diagram of an information retrieval device based on a knowledge graph provided by an embodiment of the present invention. This embodiment is applicable to the situation of combining a large language model and a knowledge graph for information retrieval. The information retrieval device based on the knowledge graph can be implemented in the form of hardware and / or software, and can be configured in any electronic device with network communication functions. As Figure 5 shown, the information retrieval device based on the knowledge graph of the present invention includes:

[0107] A graph construction module 410, configured to construct a target knowledge graph, where the target knowledge graph is constructed in the form of triples, and the target knowledge graph includes edge frequencies and event description information; the edge frequency is the number of times the same edge relationship appears; the event description information is semantic coherence description information for the triples;

[0108] An information extraction module 420, configured to extract at least one keyword and at least one target entity of the preset information in response to an input operation of the preset information; the preset information has an association relationship with the target knowledge graph; the keyword is information describing the key information of the target entity and / or information describing the relationship between the target entities;

[0109] A graph determination module 430, configured to determine a connected subgraph from the target knowledge graph according to the association relationship between the target entity and each entity in the target knowledge graph;

[0110] An information determination module 440, configured to determine retrieval information for the preset information according to the keyword, the edge frequency of the connected subgraph, and the event description information of the connected subgraph.

[0111] Based on the above embodiment, optionally, the graph construction module is configured to: obtain data to be processed, perform entity recognition and relationship extraction on the data to be processed based on a data extraction model, and obtain structured triples; the triples include edge relationships and event description information; the data extraction model is a large language model; construct the target knowledge graph based on the structured triples, and record the edge frequency corresponding to the edge relationship in the target knowledge graph.

[0112] Based on the above embodiments, optionally, a graph determination module is configured to: determine a first association index between the target entity and each entity in the target knowledge graph; use the entity in the target knowledge graph corresponding to the first association index that exceeds a preset association index as a starting node, and set the initial node weight of the starting node to a first preset weight; on the target knowledge graph, expand a preset number of times starting from the starting node to determine a connected subgraph, and set the initial node weights of other nodes in the connected subgraph except the starting node to a second preset weight; each entity in the connected subgraph is a different node; the first preset weight is greater than the second preset weight.

[0113] Based on the above embodiments, optionally, the information determination module includes: an index determination unit, an edge weight determination unit, and a retrieval information determination unit; the index determination unit is configured to determine a second association index between the keyword and each event description information in the connected subgraph; each edge in the connected subgraph corresponds to an event description information; the second association index is used to describe the association degree between the keyword and each event description information in the connected subgraph; the edge weight determination unit is configured to determine the edge weight of each edge in the connected subgraph based on the edge frequency corresponding to each edge in the connected subgraph and the second association index; the retrieval information determination unit is configured to determine retrieval information for the preset information based on the edge weight of each edge in the connected subgraph.

[0114] Based on the above embodiments, optionally, the index determination unit is configured to: determine a first proportion of the edge frequency corresponding to each edge in the connected subgraph and a second proportion of the second association index corresponding to each edge in the connected subgraph; the sum of the first proportion and the second proportion is 1; determine the edge weight of each edge in the connected subgraph based on the first proportion, the second proportion, the edge frequency, and the second association index corresponding to each edge in the connected subgraph.

[0115] Based on the above embodiments, optionally, the retrieval information determination unit includes: a matrix determination subunit, an update subunit, and a retrieval information determination subunit; the matrix determination subunit is configured to normalize the edge weights of each edge in the connected subgraph to obtain the normalized edge weights, and obtain the dynamic transition matrix of the connected subgraph based on the normalized edge weights; the update subunit is configured to update the node weight of the first node in the connected subgraph based on the dynamic transition matrix of the connected subgraph and the edge weights of each edge in the connected subgraph, and update the dynamic transition matrix of the connected subgraph according to the cumulative value of the edge weights of the edges associated with the first node; the first node is the node whose node weight is updated in the current round; the node weight of one node is updated in each round; the retrieval information determination subunit is configured to, when the dynamic transition matrix converges, obtain the node weights of each node in the connected subgraph; use the first node corresponding to the node weight greater than the preset node weight in the node weights as the second node, and use the triple corresponding to the second node as the retrieval information.

[0116] Based on the above embodiments, optionally, the update subunit is configured to: update the node weight of the first node in the connected subgraph based on the dynamic transition matrix of the connected subgraph and the edge weights of each edge in the connected subgraph, including: determining the first weight of the first node according to the cosine similarity between the preset information and the first node; determining the third nodes associated with the first node and the number of nodes of the third nodes, and determining the second weight of each third node; the third node is the node directly connected to the first node; updating the node weight of the first node in the connected subgraph according to the first weight of the first node, the dynamic transition matrix of the connected subgraph, the number of nodes of the third nodes, and the second weight of the third nodes.

[0117] Based on the above embodiments, optionally, the update subunit is further configured to: if the node weight of the third node has been updated, use the updated node weight of the third node as the second weight of the third node; if the node weight of the third node has not been updated, use the initial node weight of the third node as the second weight of the third node.

[0118] Based on the above embodiments, optionally, the update subunit is further configured to: determine the edges associated with the first node as reference edges, and obtain the edge weights of the reference edges; the reference edges are all the edges directly connected to the first node; determine the reference edges whose edge weights are greater than the first preset edge weight among the edge weights of the reference edges, and / or, determine the reference edges whose cumulative weight value of the edge weights of the reference edges is greater than the second preset edge weight as the first reference edges; the cumulative process of the reference edges is performed in the order of decreasing edge weights; keep the edge weights of the first reference edges unchanged, and update the edge weights of the second reference edges to the third preset edge weight; the second reference edges are the reference edges other than the first reference edges among the reference edges; the third preset edge weight is lower than the edge weight of the second reference edge before update; update the dynamic transition matrix of the connected subgraph according to the edge weights of the first reference edges and the edge weights of the second reference edges.

[0119] The information retrieval device based on a knowledge graph provided by an embodiment of the present invention can execute the information retrieval method based on a knowledge graph provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0120] Embodiment 5

[0121] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.

[0122] Figure 6 The structural schematic diagram of an electronic device that can be used to implement the information retrieval method based on a knowledge graph according to an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0123] As Figure 6As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0124] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0125] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the information retrieval method based on a knowledge graph.

[0126] In some embodiments, the information retrieval method based on a knowledge graph can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the information retrieval method based on a knowledge graph described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the information retrieval method based on a knowledge graph by any other appropriate means (e.g., by means of firmware).

[0127] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0128] The computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer program can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0129] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0130] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0131] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0132] A computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0133] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0134] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An information retrieval method based on a knowledge graph, characterized in that, The method includes: Constructing a target knowledge graph, which is constructed in the form of triples and includes edge frequencies and event description information; the edge frequency is the number of occurrences of the same edge relationship; the event description information is semantic coherence description information for the triples; In response to an input operation of preset information, extracting at least one keyword and at least one target entity of the preset information; the preset information has an associated relationship with the target knowledge graph; the keyword is information describing the key information of the target entity and / or information describing the relationship between each target entity; Determining a connected subgraph from the target knowledge graph according to the association relationship between the target entity and each entity in the target knowledge graph; Determining retrieval information for the preset information according to the keyword, the edge frequency of the connected subgraph, and the event description information of the connected subgraph; Among them, determining a connected subgraph from the target knowledge graph according to the association relationship between the target entity and each entity in the target knowledge graph includes: Determining a first association index between the target entity and each entity in the target knowledge graph; Taking the entities in the target knowledge graph corresponding to the first association indices that exceed the preset association index in the first association indices as starting nodes, and setting the initial node weight of the starting nodes to a first preset weight; On the target knowledge graph, expanding a preset number of times starting from the starting nodes, determining a connected subgraph, and setting the initial node weights of other nodes except the starting nodes in the connected subgraph to a second preset weight; each entity in the connected subgraph is a different node; the first preset weight is greater than the second preset weight; Among them, determining retrieval information for the preset information according to the keyword, the edge frequency of the connected subgraph, and the event description information of the connected subgraph includes: Determining a second association index between the keyword and each event description information of the connected subgraph; each edge in the connected subgraph corresponds to an event description information; the second association index is used to describe the association degree between the keyword and each event description information of the connected subgraph; Based on the edge frequency corresponding to each edge of the connected subgraph and the second association index, determining the edge weight of each edge in the connected subgraph; Based on the edge weights of each edge in the connected subgraph, determining retrieval information for the preset information.

2. The method according to claim 1, wherein Constructing a target knowledge graph includes: Obtaining data to be processed, performing entity recognition and relationship extraction on the data to be processed based on a data extraction model, and obtaining structured triples; the triples include edge relationships and event description information; the data extraction model is a large language model; Constructing the target knowledge graph based on the structured triples, and recording the edge frequencies corresponding to the edge relationships in the target knowledge graph.

3. The method according to claim 1, characterized in that, Based on the edge frequency corresponding to each edge of the connected subgraph and the second association index, determining the edge weight of each edge in the connected subgraph includes: Determine the first proportion of the edge frequency corresponding to each edge of the connected subgraph, and the second proportion of the second association index corresponding to each edge of the connected subgraph; the sum of the first proportion and the second proportion is 1; Based on the first proportion, second proportion, edge frequency, and second association index corresponding to each edge of the connected subgraph, determine the edge weight of each edge in the connected subgraph.

4. The method according to claim 1, characterized in that, Based on the edge weights of each edge in the connected subgraph, determine the retrieval information for the preset information, including: Normalize the edge weights of each edge in the connected subgraph to obtain the normalized edge weights, and obtain the dynamic transition matrix of the connected subgraph based on the normalized edge weights; Based on the dynamic transition matrix of the connected subgraph, update the node weight of the first node in the connected subgraph, and update the dynamic transition matrix of the connected subgraph according to the cumulative value of the edge weights of the edges associated with the first node; the first node is the node whose node weight is updated in the current round; one node's node weight is updated in each round; When the dynamic transition matrix converges, obtain the node weights of each node in the connected subgraph; use the first node corresponding to the node weight greater than the preset node weight in the node weights as the second node, and use the triple corresponding to the second node as the retrieval information.

5. The method according to claim 4, characterized in that, Based on the dynamic transition matrix of the connected subgraph, updating the node weight of the first node in the connected subgraph includes: Determine the first weight of the first node according to the cosine similarity between the preset information and the first node; Determine the third nodes associated with the first node and the number of nodes of the third nodes, and determine the second weight of each third node; the third node is the node directly connected to the first node; Update the node weight of the first node in the connected subgraph according to the first weight of the first node, the dynamic transition matrix of the connected subgraph, the number of nodes of the third nodes, and the second weight of the third nodes.

6. The method according to claim 5, wherein Determining the second weight of each third node includes: If the node weight of the third node has been updated, use the updated node weight of the third node as the second weight of the third node; If the node weight of the third node has not been updated, use the initial node weight of the third node as the second weight of the third node.

7. The method according to claim 4, characterized in that, Updating the dynamic transition matrix of the connected subgraph according to the cumulative value of the edge weights of the edges associated with the first node includes: Determine the edges associated with the first node as reference edges, and obtain the edge weights of the reference edges; the reference edges are all the edges directly connected to the first node; Determine the reference edges with edge weights greater than the first preset edge weight among the edge weights of the reference edges, and / or determine the reference edges with the cumulative weight value of the edge weights of the reference edges greater than the second preset edge weight as the first reference edges; the cumulative process of the edge weights of the reference edges is carried out in the order of the edge weights from large to small; The edge weight of the first reference edge remains unchanged, and the edge weight of the second reference edge is updated to a third preset edge weight; the second reference edge is the reference edge other than the first reference edge among the reference edges; the third preset edge weight is lower than the edge weight of the second reference edge before the update. Update the dynamic transition matrix of the connected subgraph according to the edge weight of the first reference edge and the edge weight of the second reference edge.

8. An information retrieval device based on a knowledge graph, characterized in that, The device includes: A graph construction module for constructing a target knowledge graph, the target knowledge graph being constructed in the form of triples, the target knowledge graph including edge frequencies and event description information; the edge frequency is the number of occurrences of the same edge relationship; the event description information is semantic coherence description information for the triples. An information extraction module for extracting at least one keyword and at least one target entity of the preset information in response to an input operation of the preset information; the preset information has an association relationship with the target knowledge graph; the keyword is information describing the key information of the target entity and / or information describing the relationship between the target entities. A graph determination module for determining a connected subgraph from the target knowledge graph according to the association relationship between the target entity and each entity in the target knowledge graph. An information determination module for determining retrieval information for the preset information according to the keyword, the edge frequency of the connected subgraph, and the event description information of the connected subgraph. Among them, the graph determination module is used to: determine a first association index between the target entity and each entity in the target knowledge graph; use the entity in the target knowledge graph corresponding to the first association index exceeding the preset association index among the first association indexes as the starting node, and set the initial node weight of the starting node to a first preset weight; on the target knowledge graph, expand a preset number of times starting from the starting node, determine the connected subgraph, and set the initial node weight of other nodes in the connected subgraph except the starting node to a second preset weight; each entity in the connected subgraph is a different node; the first preset weight is greater than the second preset weight. Among them, the information determination module includes: an index determination unit, an edge weight determination unit, and a retrieval information determination unit; the index determination unit is used to determine a second association index between the keyword and each event description information of the connected subgraph; each edge in the connected subgraph corresponds to an event description information; the second association index is used to describe the association degree between the keyword and each event description information of the connected subgraph; the edge weight determination unit is used to determine the edge weight of each edge in the connected subgraph based on the edge frequency corresponding to each edge of the connected subgraph and the second association index; the retrieval information determination unit is used to determine the retrieval information for the preset information based on the edge weight of each edge in the connected subgraph.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; where The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, enables the at least one processor to execute the knowledge graph-based information retrieval method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions which, when executed by a processor, implement the knowledge graph-based information retrieval method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Information prediction method and terminal

    CN107358315A

  • Knowledge Canvassing Using a Knowledge Graph and a Question and Answer System

    US20160378851A1