Knowledge graph multi-hop intelligent search algorithm based on subgraph construction
Through the knowledge graph multi-hop intelligent search algorithm built on sub-graphs, the problems of high time complexity and incomplete inference paths in large knowledge graphs in the existing technology are solved, and deeper semantic information acquisition and inference ability are enhanced, supporting complex and in-depth knowledge query.
Patent Information
- Application Number
- CN202510549782.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When processing large knowledge graphs, existing multi-hop query methods have problems such as high time complexity, high computational cost, and incomplete graph construction, resulting in incomplete inference paths.
A knowledge graph multi-hop intelligent search algorithm built on sub-graphs is used to construct a local graph structure containing related entities and relationships through steps such as user problem processing, constraint detection, main entity prediction, knowledge graph embedding, entity and relationship detection, sub-graph construction and entity prediction, which is used to limit the search space and guide the reasoning process.
It realizes deeper semantic information acquisition, enhances reasoning ability, provides accurate result feedback, meets the needs of complex and in-depth knowledge query, and supports multi-level correlation query and personalized recommendations.
Smart Images

Figure CN120067345A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of knowledge graphs, and specifically to a knowledge graph multi-hop intelligent search algorithm based on subgraph construction. Background Art
[0002] Current multi-hop query methods include methods based on semantic parsing, methods based on information extraction, and methods based on knowledge embedding. These multi-hop query methods all have certain drawbacks:
[0003] The method based on semantic parsing converts the problem input by the user into a series of logical forms and then executes the query on the knowledge graph. However, for multi-hop query problems, the time complexity of semantic analysis will increase exponentially, and the inference time of large knowledge graphs will also increase exponentially.
[0004] The method based on information extraction first determines a main entity, then extracts a subgraph related to the main entity from the knowledge graph, and finally encodes the problem into a vector to obtain the answer through inference and knowledge embedding. Although this method does not require manual setting of problem templates, due to the huge number of triples in the knowledge graph and the large number of triples covered by the extracted subgraph, the computational cost will increase. At the same time, due to the incompleteness of the graph construction, an incomplete subgraph may be obtained, resulting in the lack of correct inference paths in the subgraph.
[0005] The method based on knowledge embedding embeds triple knowledge into a low-dimensional vector space and calculates the similarity between entities and relationships by designing different scoring functions. The higher the similarity, the closer the distance in the projected space. However, this method only considers triple facts and ignores the semantic relationship between inference paths and multi-relationship problems. Summary of the Invention
[0006] The purpose of the present invention is to provide a knowledge graph multi-hop intelligent search algorithm based on subgraph construction to solve the problems raised in the above background art.
[0007] To achieve the above purpose, the present invention provides the following technical solutions:
[0008] A knowledge graph multi-hop intelligent search algorithm based on subgraph construction includes the following steps:
[0009] S1. User problem processing:
[0010] The user inputs a problem, which is passed into the problem embedding module. In the problem embedding module, a pre-trained model is used to convert the problem into a vector representation of a fixed dimension. This vector is processed through three fully connected linear layers, and this vector is used to represent the semantic information of the problem.
[0011] S2. Constraint detection stage:
[0012] Compare the problem vector with the entity embedding to detect the constraints, which can help narrow down the scope of entities and improve the search efficiency. Filter the entities according to the constraints and use them as candidate entities for subsequent steps;
[0013] S3. Main entity prediction stage:
[0014] Use the relationship scoring function between the problem embedding and the entity embedding to predict the node of the main entity. This scoring function is used to measure the semantic association degree between the problem and the entity. According to the score ranking, select the node of the main entity as the starting point for subsequent search;
[0015] S4. Knowledge graph embedding:
[0016] The entities and relationships in the knowledge graph are processed by the knowledge graph embedding module to obtain their embedding representations. Use the complex embedding technology to convert the entities and relationships into vector representations, and define a scoring function to evaluate the correlation between entities and relationships;
[0017] S5. Entity detection:
[0018] Detect the entities related to the problem from the knowledge graph to determine the answer candidate set;
[0019] S6. Relationship detection stage:
[0020] Use the sorting method to sort the relationships and select the relationships related to the problem. This step helps to determine the relationship type between the problem and the entities. According to the score ranking, select the relationships as the subsequent search paths;
[0021] S7. Subgraph construction stage:
[0022] Construct a subgraph according to the main entity and relationship detection. When constructing the subgraph, consider the constraints to ensure that the subgraph can reflect the specific requirements of the problem;
[0023] S8. Entity prediction:
[0024] After the subgraph construction is completed, use the graph neural network method to calculate the representation of each node in the subgraph. Then, through the classification operation, predict which entity nodes can answer the question. Finally, the entity node with the highest score is selected as the answer and returned to the user;
[0025] S9. Result output:
[0026] Output the predicted answer to the user. The answer is an entity or a list composed of multiple entities, depending on the type of the problem and the constraints.
[0027] As a further solution of the present invention: the pre-trained model in S1 is the ROBERTa model, and the dimension of the vector is 768.
[0028] As a further solution of the present invention: each linear layer in S1 has ReLU activation and dropout with a probability of 0.1.
[0029] As a further solution of the present invention: during the subgraph construction process of S7, at least one node with the highest probability is selected in each iteration.
[0030] As a further solution of the present invention: the algorithm flow of subgraph construction in S7 is as follows:
[0031] S701. Initialize the problem graph with the question and the question entity , where ; ;
[0032] S702. Classify the entity nodes in the graph and select the entity nodes with a probability greater than ;
[0033] S703. The calculation formula for the entity node is: ; where ; for all the selected entity nodes, perform retrieval in the knowledge graph and retrieval in the external knowledge base respectively to obtain and ;
[0034] S704. For all , extract entities from the new document nodes; for all , extract the head and tail of the new fact nodes;
[0035] S705. Add new nodes and edges ( ) to the problem graph, and use the answer classifier to select the entity node of the best answer in the final graph.
[0036] As a further solution of the present invention: the algorithm description of the subgraph construction is as follows:
[0037] represents the question, represents the question subgraph, represents the set of vertices, also called nodes, represents all the selected entity nodes, represents the initial entity nodes, represents the edges between nodes, represents the edges between the initial nodes, represents the triple, denotes the iteration round, denotes a classifier that uses LSTM and sigmoid for classification, denotes the entities in the document, denotes the entities in the facts;
[0038] Subgraph construction is carried out in T iterations. In each iteration, entities with a probability greater than are expanded. Then, for each selected entity, a set of relevant documents and a set of relevant facts are retrieved. The new documents are passed through the entity linking system to identify the entities appearing therein, and the head and tail entities of each fact are extracted. The last stage of each iteration is to update the problem graph by adding all these new edges. After the t-th iteration of expansion, an additional classification step is applied to the final problem subgraph to predict the answer entity;
[0039] Graph embedding: For a given fact r, if the scoring function is defined as the multilinear product of the embedding vectors of s, r, and o, since the dot product calculation between real number vectors is commutative, the dot product cannot handle asymmetric relationships. When using complex vectors, that is, vectors represented in the complex space C, it is ordered, meaning that asymmetric relationships can be handled and different scores can be obtained according to the order of the entities involved. The complex vector space can describe asymmetric relationships while retaining the efficiency advantages of the dot product, that is, both the space and time complexity are linear.
[0040] As a further solution of the present invention: The subgraph in S7 is a local graph structure containing relevant entities and relationships, which is used to limit the search space and guide the subsequent reasoning process.
[0041] As a further solution of the present invention: The graph neural network method in S8 is a graph CNN.
[0042] As a further solution of the present invention: The classification operation in S8 includes retrieving information from the knowledge graph and retrieving information from the corpus.
[0043] Compared with the prior art, the beneficial effects of the present invention are:
[0044] The present invention combines feature search subgraph, constraint detection, and knowledge graph complex embedding representation to achieve multi-hop intelligent search, obtain deeper semantic information, enhance the reasoning ability, and provide accurate result feedback to users, meeting the user's needs for more complex and in-depth knowledge queries. The algorithm supports complex queries and complex tasks. Through a large number of experiments, it is verified that the algorithm can achieve multi-level association queries and provide accurate and comprehensive knowledge;
[0045] Among them, multi-level association query: Multi-hop query can achieve multi-level association query by connecting the relationships between multiple entities. By specifying multiple relationship paths, it can cross the relationships between multiple entities to obtain richer and deeper knowledge;
[0046] Inference and path inference: Multi-hop query can utilize the inference technology and path inference algorithm in the knowledge graph to perform path inference and reasoning based on the existing knowledge and relationships. Through reasoning, hidden associations and semantic relationships between entities can be discovered, helping users to understand knowledge more comprehensively;
[0047] Multi-dimensional information aggregation: Multi-hop query can aggregate the information between different entities to provide multi-dimensional knowledge. Through multiple relationship connections, relevant information can be obtained from different perspectives, helping users to establish a comprehensive knowledge view;
[0048] Personalized recommendation and question answering: Multi-hop query can support personalized recommendation and question answering. By analyzing the user's query requirements and association relationships, personalized recommendation results and accurate question answers can be provided according to the user's interests and context. Description of the Drawings
[0049] Figure 1 It is a flow chart of a multi-hop intelligent search algorithm for a knowledge graph based on subgraph construction;
[0050] Figure 2 It is a flow chart of the relationship detection stage in a multi-hop intelligent search algorithm for a knowledge graph based on subgraph construction;
[0051] Figure 3 It is an algorithm flow chart of subgraph construction in a multi-hop intelligent search algorithm for a knowledge graph based on subgraph construction. Detailed Implementation Manner
[0052] Please refer to Figures 1 - 3 , in the embodiment of the present invention, a multi-hop intelligent search algorithm for a knowledge graph based on subgraph construction includes the following steps:
[0053] S1. User question processing:
[0054] The user inputs a question, and the question is passed into the question embedding module. In the question embedding module, the ROBERTa model is used to convert the question into a vector representation with a fixed dimension of 768. This vector is processed through three fully connected linear layers, each linear layer having ReLU activation and dropout with a probability of 0.1. This vector is used to represent the semantic information of the question;
[0055] For a user question , the main entity , the guest entity , the answer set , the learned question embedding form is as follows:
[0056]
[0057]
[0058] represents a specific scoring function or triple evaluation function that takes three arguments: the head entity , the relation or query and the tail entity , and returns a scalar value as the score;
[0059] is the universal quantifier, meaning "for all";
[0060] This is an element in the set , representing a tail entity;
[0061] is the membership symbol, meaning "belongs to";
[0062] is a set that contains multiple tail entities;
[0063] The above formula: for all tail entities in the set , the score of the triple composed of the head entity , the relation or query and the tail entity is greater than 0;
[0064] The following formula: for all tail entities in the set , the score of the triple composed of the head entity , the relation or query and the tail entity is less than 0;
[0065] S2. Constraint Detection Phase:
[0066] Compare the question vector with the entity embedding to detect constraints that can help narrow down the scope of entities and improve the search efficiency. Filter the entities according to the constraints and use them as candidate entities for the subsequent steps;
[0067] S3. Main Entity Prediction Phase:
[0068] Predict the node of the main entity using the relationship scoring function between the problem embedding and the entity embedding. This scoring function is used to measure the semantic association degree between the problem and the entity. According to the score ranking, select the node of the main entity as the starting point for subsequent search;
[0069] S4. Knowledge graph embedding:
[0070] Entities and relationships in the knowledge graph are processed by the knowledge graph embedding module to obtain their embedding representations. The complex embedding technology is used to convert entities and relationships into vector representations, and a scoring function is defined to evaluate the correlation between entities and relationships;
[0071] A knowledge graph is a multi-relational graph with entities as nodes and relationships between entities as edges, stored in the form of triples. Given an incomplete knowledge graph, path prediction is to predict valid unknown links, and this task can be achieved through a knowledge graph embedding model, which assigns scores to the predicted links , and verify whether the link is true. Complex embedding is a potential decomposition method for learning a large number of symmetric and anti-symmetric relationships in a complex space. Complex vectors can retain the advantages of the dot product and are linear in both space and time complexity. Generate entity and relationship embeddings in the knowledge graph using complex embedding and define a scoring function. When the value is greater than 0, it indicates that the obtained triple is correct; this scoring function is:
[0072]
[0073] is a scoring function or score function used to evaluate the rationality or correlation of a triple 、relationship and object entity ; ;
[0074] is used to extract the real part of a complex expression;
[0075] is a summation expression that iterates over all integers from 1 to and sums the products of each corresponding 、 、 and ;
[0076] 、 and represent the -th, relationship -th and object entity -th elements of the embedding vectors of the main entity component;
[0077] S5. Entity Detection:
[0078] Detect entities related to the question from the knowledge graph and determine the answer candidate set;
[0079] S6. Relationship Detection Phase:
[0080] Use a sorting method to sort the relationships, select the relationships related to the question. This step helps to determine the type of relationship between the question and the entities. According to the score ranking, select the relationships as the subsequent search paths;
[0081] Method 1: Use an IDF-based retrieval system and assume that all sentences in the corpus have been entity-linked before indexing. Rank them according to their IDF similarity to the user's question, and only return the top five documents;
[0082] Method 2: For retrieving multiple facts about entities from the knowledge graph, the retrieved facts are restricted to the subject and the object, and according to the relationship and the question between the similarity s( , ) for ranking. Since it is not obvious how to evaluate the relevance of the facts to the question , learning is required( , ). Let be the embedding of the relationship found from the embedding table, and let be the word sequence of the question . The similarity is defined as the dot product of the final state LSTM representation of and . Pass the dot product through the sigmoid function to bring it into the range of [0, 1], and train this similarity function into a classifier that predicts which retrieved facts are relevant to the question . The final sorting method for the facts is as follows:
[0083]
[0084]
[0085] After sorting using this sorting method, select the relationship;
[0086] is the output after processing the sequence , where represents the length of the sequence;
[0087] is a special type of Recurrent Neural Network (RNN) that can capture long-term dependencies in a sequence. Each is a word embedding vector representing the -th word in the sequence;
[0088] denotes the output of the network which is a -dimensional vector, where is the dimension of the network's output layer, which determines the size of the hidden state vector;
[0089] represents a scoring function or similarity function used to evaluate the and correlation or matching degree between them;
[0090] is an activation function, also known as the logistic function or sigmoid function, which maps any real-valued number to the interval and is often used in the output layer of binary classification problems;
[0091] is a dot product operation, where and are two vectors representing certain representations (such as embedding vectors) of and respectively; is the "matching score" between and , while the function is used to map this score to the interval, thus obtaining a more interpretable similarity or probability value.
[0092] S7. Subgraph Construction Phase:
[0093] Construct a subgraph based on the main entity and relation detection. The subgraph is a local graph structure containing relevant entities and relations, which is used to restrict the search space and guide the subsequent reasoning process; when constructing the subgraph, consider the constraints to ensure that the subgraph can reflect the specific requirements of the problem. During the subgraph construction process, select at least one node with the highest probability in each iteration;
[0094] The construction of the subgraph is based on the questions raised by the user. It learns to retrieve content from the corpus, knowledge graph, or a combination of both, and combines this heterogeneous information into a single data structure, enabling the system to reason and find the best answer. The final result is a learning iteration process for subgraph construction, which starts from a small subgraph containing only the question text and the entities it contains, and gradually expands the subgraph to include useful information from the knowledge graph and corpus. At the same time, the information increment problem leads to a smaller high-recall subgraph generated by the subgraph construction process compared to heuristically created subgraphs, making the final answer search process easier. The current model only considers each individual fact and ignores the internal relationships, so it cannot capture deeper semantics for better embedding. To prevent the length of the relationship path from growing exponentially, graph embedding is used on some search paths to add constraints and explore the next path segment, which helps reduce the multi-hop search space to generate a more flexible query graph;
[0095] S8, Entity Prediction:
[0096] After the subgraph construction is completed, the graph CNN neural network method is used to calculate the representation of each node in the subgraph. Then, through two classification operations of retrieving information from the knowledge graph and retrieving information from the corpus, it is predicted which entity nodes can answer the question. Finally, the entity node with the highest score is selected as the answer and returned to the user;
[0097] S9, Result Output:
[0098] The predicted answer is output to the user. The answer is an entity or a list composed of multiple entities, depending on the type of the question and the constraints.
[0099] Preferably, the algorithm process of subgraph construction in S6 is as follows:
[0100] S601. Initialize the question graph with the question and the question entity , where ;
[0101] S602. Classify the entity nodes in the graph and select the entity nodes with a probability greater than ;
[0102] S603. The calculation formula for the entity node is: ; where ; For all the selected entity nodes, perform retrieval in the knowledge graph and retrieval in the external knowledge base respectively to obtain and ;
[0103] S604. For all , extract entities from the new document nodes; for all , extract the head and tail of the new fact nodes;
[0104] S605. Add new nodes and edges to the problem graph ( ), and use the answer classifier to select the entity nodes of the best answer in the final graph.
[0105] Its algorithm description is as follows:
[0106] represents the question, represents the question sub-graph, represents the set of vertices, also called nodes, represents all the selected entity nodes, represents the initial entity nodes, represents the edges between nodes, represents the edges between the initial nodes, represents the triple, represents the iteration round, represents the classifier using LSTM and sigmoid for classification, represents the entities in the document, represents the entities in the fact;
[0107] The sub-graph construction is carried out in T iterations. In each iteration, entities with a probability greater than are selected for expansion. Then, for each selected entity, a set of relevant documents and a set of relevant facts are retrieved. The new documents are passed through the entity linking system to identify the entities appearing in them, and the head and tail entities of each fact are extracted. The last stage of each iteration is to update the problem graph by adding all these new edges. After the t-th iteration of expansion, an additional classification step is applied to the final question sub-graph to predict the answer entity;
[0108] Graph Embedding: For a given fact r, if the scoring function is defined as the multilinear product of the embedding vectors of s, r, and o, since the dot product calculation between real number vectors is commutative, the dot product cannot handle asymmetric relations. When using complex vectors, that is, vectors represented in the complex space C, it is ordered, meaning that asymmetric relations can be handled and different scores can be obtained according to the order of the entities involved. The complex vector space can describe asymmetric relations while retaining the efficiency advantages of the dot product, that is, both the space and time complexities are linear.
[0109] The knowledge graph implementation manner of the present invention involves multiple steps, mainly including data preprocessing, knowledge extraction, knowledge fusion, and knowledge application, etc.
[0110] Data preprocessing: Remove incorrect, duplicate, and inconsistent data, fill in missing values, handle outliers, and convert the data into a unified format, such as date of birth, time, etc.;
[0111] Knowledge extraction: Knowledge extraction is the process of identifying and extracting structured information from data sources. In this invention, the named entity algorithm is mainly used to identify entities, relationships, and their attributes from a large amount of text, such as: person names, organization names, ages, belong to, etc.;
[0112] Knowledge fusion: Knowledge fusion is the integration of knowledge extracted from different data sources to solve data redundancy and inconsistency problems. First, it is to identify different expressions in different data sources that refer to the same entity, solve the ambiguity problem of homonymous entities, and determine whether different names point to the same entity. Secondly, it is to merge the information in different data sources to form a complete entity profile, handle information conflicts between different data sources, and select the most reliable or latest data;
[0113] Knowledge application: Knowledge application is to apply the constructed knowledge graph to actual business scenarios and applications. In this invention, the information provided by the knowledge graph is mainly used to build an intelligent question answering system to answer users' questions, and to provide personalized recommendations based on the relationship between users and entities.
[0114] Comparative test
[0115] The traditional mysql database storage method mainly stores source data, and the data search method is vector similarity. Neo4j is a graph database storage method that mainly stores triples, and the data search method is the subgraph construction search algorithm proposed by this invention. The accuracy rate is the average accuracy rate of 1000 test data, and the time consumption is the average time consumption per data.
[0116] Table 1 - Comparison table between the traditional mysql database storage method and the neo4j graph database storage method;
[0117]
[0118] It can be concluded from Table 1 that: under the condition of equal number of entities, the accuracy rate of the neo4j storage method is higher than that of the mysql storage method, and the time consumption of the neo4j storage method is lower than that of the mysql storage method, that is, the search algorithm of this invention is preferably the traditional algorithm.
[0119] Enter an entity name, and you can obtain the source of the entity, the city to which the entity belongs, the type of the entity, and the platform to which the entity belongs; that is, through multiple queries, multi-level association queries can be realized. By specifying multiple relationship paths, the relationships between multiple entities can be crossed to obtain richer and deeper knowledge. Through multiple relationship connections, relevant information can be obtained from different perspectives to help users establish a comprehensive knowledge view.
[0120] To further illustrate the technical effects of the present invention, it is verified through the following cases:
[0121] Scenario setting
[0122] We have a knowledge graph in the field of movies, which contains entities such as movies, actors, directors, etc. and the relationships between them; the user asks a question: "Who is the director of movie A?" We will use the proposed knowledge graph multi-hop intelligent search algorithm based on subgraph construction (hereinafter referred to as the "new algorithm") to answer this question and compare it with the traditional keyword matching-based search method (hereinafter referred to as the "traditional method").
[0123] Data preparation
[0124] 1. Knowledge graph: contains information about 10,000 movies, 5,000 actors, and 1,000 directors, as well as the relationships between them;
[0125] 2. User question: "Who is the director of movie A?";
[0126] 3. Pre-trained model: Use the ROBERTa model for question embedding, with a vector dimension of 768, and each linear layer has ReLU activation and dropout with a probability of 0.1.
[0127] Algorithm execution
[0128] Steps of the new algorithm
[0129] 1. User question processing: Convert the question into a 768-dimensional vector representation;
[0130] 2. Constraint detection stage: Narrow the entity scope by comparing the question vector with the entity embedding;
[0131] 3. Main entity prediction stage: Predict movie A as the main entity;
[0132] 4. Knowledge graph embedding: Use complex embedding technology to process entities and relationships in the knowledge graph;
[0133] 5. Entity detection: Detect the movie A entity from the knowledge graph;
[0134] 6. Relationship detection stage: Select the relationship related to "director";
[0135] 7. Subgraph construction stage:
[0136] 701. Initialize the question graph , containing the movie A entity;
[0137] 702. In T = 3 iterations, each time select the node with the highest probability for expansion;
[0138] 703. Retrieve relevant documents and facts and update the problem graph;
[0139] 704. Finally, obtain a subgraph containing director information;
[0140] 705. Entity prediction: Use graph CNN to calculate the representation of each node in the subgraph and predict the director entity;
[0141] 706. Result output: Return the director entity "Director X".
[0142] Steps of the traditional method
[0143] Keyword matching: Search for entities and relationships containing "Movie A" and "Director X" in the knowledge graph;
[0144] Result output: Return the matched director entity.
[0145] Result comparison
[0146] Table 2 - Comparison table of the new algorithm and the traditional algorithm;
[0147]
[0148] It can be analyzed from Table 2 that:
[0149] By constructing a subgraph, the new algorithm significantly reduces the search space, thereby reducing the search time. The average search time is shortened from 7.8 seconds to 2.3 seconds, improving the efficiency by about 70%.
[0150] By restricting the search space, the new algorithm reduces the unnecessary number of searches. The traditional method requires multiple searches to find the answer, while the new algorithm only needs 3 searches.
[0151] By constructing a subgraph, the new algorithm reduces the search space from 10,000 nodes to 500 nodes, greatly reducing the complexity of the search.
[0152] In summary, the new algorithm is superior to the traditional method in terms of search efficiency, accuracy, user satisfaction, and system resource consumption. Specifically, the new algorithm shortens the search time by about 70%, improves the accuracy to 100%, has higher user satisfaction, and lower system resource consumption, which proves the effectiveness and advantages of the new algorithm in dealing with complex queries and large knowledge graphs.
[0153] It is obvious to those skilled in the art that the present invention is not limited to the details of the above-described exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention. Therefore, in any aspect, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.
[0154] In addition, it should be understood that although this specification is described in terms of embodiments, not every embodiment only contains an independent technical solution. This narrative manner of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A multi-hop intelligent search algorithm for knowledge graphs based on subgraph construction, characterized in that: The following steps are involved: S1. User problem handling: The user inputs a question and passes it into the question embedding module. In the question embedding module, the question is converted into a fixed-dimensional vector representation using a pre-trained model. This vector is processed through three fully connected linear layers. This vector is used to represent the semantic information of the question. S2, constraint detection phase: S3, main entity prediction stage: S4. Knowledge Graph Embedding: The entities and relations in the knowledge graph are processed by the knowledge graph embedding module to obtain their embedded representations. The entities and relations are converted into vector representations using complex embedding techniques, and a scoring function is defined to evaluate the relevance between entities and relations. S5. Entity Detection: Detect entities related to the question from the knowledge graph and determine the set of answer candidates; S6, relationship detection stage: Use the sorting method to sort the relations and select the relations related to the question. This step helps determine the type of relations between the question and the entity. According to the score ranking, the relations are selected as the path for subsequent search. S7, subgraph construction phase: Construct subgraphs based on main entity and relationship detection. When constructing subgraphs, consider constraints to ensure that the subgraphs reflect the specific requirements of the problem. S8. Entity Prediction: S9. Result output: The predicted answer is output to the user. The answer is a single entity or a list of multiple entities, depending on the type of question and the constraints.
2. According to claim 1, a knowledge graph multi-hop intelligent search algorithm based on subgraph construction is characterized in that: The pre-trained model in S1 is a ROBERTa model, and the dimension of the vector is 768; each linear layer in S1 has ReLU activation and dropout with a probability of 0.
1.
3. According to claim 1, a knowledge graph multi-hop intelligent search algorithm based on subgraph construction is characterized in that: S2, the constraint detection phase is specifically as follows: compare the question vector with the entity embedding to detect constraints, which can help narrow the scope of entities and improve search efficiency, filter entities according to the constraints, and use them as candidate entities for subsequent steps.
4. According to claim 1, a knowledge graph multi-hop intelligent search algorithm based on subgraph construction is characterized in that: During the subgraph construction process of S7, at least one node with the highest probability is selected in each iteration.
5. According to claim 1, a knowledge graph multi-hop intelligent search algorithm based on subgraph construction is characterized in that: The algorithm flow of subgraph construction in S7 is as follows: S701, Use Questions and the problem entity Initialize the problem graph ,in ; S702: Classify the entity nodes in the graph and select Entity nodes; S703, the calculation formula of the entity node is: ;in, ; For all selected entity nodes, perform retrieval in the knowledge graph and in the external knowledge base respectively, and obtain and ; S704, for all , extract entities from the new document node; for all , extract the head and tail of the new fact node; S705. Add new nodes and edges to the problem graph ( ), and use the answer classifier to select the entity node with the best answer in the final graph.
6. According to claim 5, a knowledge graph multi-hop intelligent search algorithm based on subgraph construction is characterized in that: The algorithm description of the subgraph construction is as follows: Indicates the problem, represents the problem subgraph, represents a collection of vertices, also called nodes, Represents all selected entity nodes, represents the initial entity node, represents the edges between nodes, represents the edges between the initial nodes, represents a triple, represents the iteration round, Represents a classifier that uses LSTM and sigmoid for classification. Represents an entity in a document, Represents an entity in fact; The subgraph construction is performed in T iterations, and in each iteration, the subgraph with probability greater than The entities are expanded, and then for each selected entity, a set of related documents and a set of related facts are retrieved. The new document is passed through the entity linking system to identify the entities that appear in it, and the head and tail entities of each fact are extracted. The last stage of each iteration is to update the question graph by adding all these new edges. After the tth iteration of expansion, an additional classification step is applied to the final question subgraph to predict the answer entity; Graph embedding: For a given fact r, if the score function is defined as the multilinear product of the embedding vectors of s, r, and o, since the dot product calculation between real vectors is commutative, the dot product cannot handle asymmetric relationships. When using complex vectors, that is, vectors represented in the complex space C, it is ordered, which means that antisymmetric relationships can be handled, and different scores are obtained according to the order of the entities involved. The complex vector space can describe asymmetric relationships while retaining the efficiency advantage of the dot product, that is, both the space and time complexity are linear.
7. According to claim 1, a knowledge graph multi-hop intelligent search algorithm based on subgraph construction is characterized in that: The subgraph in S7 is a local graph structure containing related entities and relationships, which is used to limit the search space and guide the subsequent reasoning process.
8. According to claim 1, a knowledge graph multi-hop intelligent search algorithm based on subgraph construction is characterized in that: S3, the main entity prediction stage is as follows: Use the relationship score function between question embedding and entity embedding to predict the node of the main entity. This score function is used to measure the degree of semantic association between the question and the entity. According to the score ranking, the node of the main entity is selected as the starting point for subsequent search.
9. According to claim 1, a knowledge graph multi-hop intelligent search algorithm based on subgraph construction is characterized in that: S8, entity prediction is specifically as follows: after the subgraph is constructed, the graph neural network method is used to calculate the representation of each node in the subgraph, and then, through classification operations, it is predicted which entity nodes can answer the question. Finally, the entity node with the highest score is selected as the answer and returned to the user; the classification operation includes retrieving information from the knowledge graph and retrieving information from the corpus; the graph neural network method in S8 is graph CNN.
Citation Information
Patent Citations
Question-answering method based on knowledge graph completion
CN112015868A
Complex question multi-hop intelligent question answering method based on knowledge graph representation learning
CN115757715A
Multi-hop retrieval method and system based on knowledge graph embedding and path information
CN116662478A
Complex multi-hop knowledge base question and answer design method based on semantic sorting framework
CN117056487A
Knowledge graph question-answering method for sub-graph retrieval optimization
CN117149974A