Medical entity retrieval method based on knowledge graph embedding and key, and system therefor
Through dependent syntax analysis, query mark trees are generated, and neighbor information screening and sorting of medical knowledge graphs is used to solve the problem that the search results in the existing technology do not meet user expectations, and more accurate and efficient medical entity retrieval is achieved.
Patent Information
- Application Number
- PCT/CN2023/134706
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-27
- Filing Date
- 2023-11-28
- Publication Date
- 2025-06-05
AI Technical Summary
Existing medical entity search methods are difficult to effectively identify and incorporate keywords in user query intentions, resulting in the search results not meeting user expectations.
The query mark tree is generated through dependent syntax analysis, and the keywords that cannot be generated are filtered. The mark tree is filtered using the neighbor information of the medical knowledge graph, and the results are sorted using the vectorized knowledge graph.
Improve the accuracy and user experience of search results, reduce useless information, and improve the search effect, especially on advanced knowledge graph embedding models.
Smart Images

Figure CN2023134706_05062025_PF_FP_ABST
Abstract
Description
Medical entity retrieval method and system based on knowledge graph embedding and keywords Technical Field
[0001] The present invention belongs to the field of computer application technology and relates to a medical entity retrieval method and system based on knowledge graph embedding and keywords. Background Art
[0002] A knowledge graph is a graph structure used to represent and organize the relationships between knowledge, enabling computers to understand and reason about relationships between entities. This allows for better support for applications in natural language processing, such as question answering, recommendation, and retrieval. Basic concepts in a knowledge graph include entities, attributes, relationships, nodes, edges, graphs, and triples. The various entities within a neighborhood are collectively referred to as an ontology. An attribute refers to an entity's name or description, while a relationship is a directional connection between two entities. Triples represent the basic units in a knowledge graph and consist of a subject, a predicate, and an object.
[0003] In the medical field, knowledge graphs can be used to represent diseases, symptoms, medications, doctors, and patients as entities. Relationships between different entities can include treatment, symptoms, and belonging to. (Disease A, symptoms, symptom B) can constitute a triple in a medical knowledge graph. Compared to traditional knowledge bases, using a knowledge graph embedding model to vectorize the medical knowledge graph can integrate, understand, and apply large amounts of medical information. This allows computers to efficiently process and analyze knowledge through vector-to-vector computations. This can help doctors better understand the relationship between different diseases and their symptoms, supporting clinical decision-making and diagnosis. Medical students, patients, and their families can more quickly obtain relevant medical entity information, providing a better user search experience.
[0004] When searching on vectorized medical knowledge graphs, existing methods primarily rely on named entity recognition and relationship extraction to understand the user's search intent, limiting the search scope and clarifying the search purpose. This often involves adding multiple constraints to the search statement. This can result in some constraints that are neither named entities nor relationships being unrecognized. Consequently, these constraints are not incorporated into the search formula generated for the knowledge graph embedding, resulting in potentially unsatisfactory search results. Therefore, special strategies are needed to address this situation and better meet user search needs.
[0005] Summary of the Invention
[0006] In response to the shortcomings of the existing technology, the present invention proposes a medical entity retrieval method and system based on knowledge graph embedding and keywords, which can be effectively applied to the retrieval of knowledge graphs in the medical field.
[0007] In a first aspect, the present invention provides a medical entity retrieval method based on knowledge graph embedding and keywords, including query tag tree generation, query tag tree screening and query result sorting.
[0008] The query tag tree is generated by parsing the natural statement of the query and generating a corresponding knowledge graph embedded search formula according to the parsed results. The results are then calculated in the vectorized knowledge graph according to the formula and the search results are saved in a tree structure. The query tag tree filters the query intent that cannot be generated into the knowledge graph embedded search formula as a keyword, and further filters the query tag tree through the neighbors of the keyword in the medical knowledge graph. The result generation ranking is sorted by the spatial distance between the nodes in the tag tree in the vectorized knowledge graph.
[0009] The specific steps are as follows:
[0010] Step 1: Input question analysis
[0011] For natural language questions input by users, the dependency relationship between words in the natural language questions is analyzed using dependency syntax to obtain different dependency structures. First, for the two components in the "follow-dependency" type that have a parallel relationship, a dependency structure with overlapping components is found from the remaining dependency structures. Using the relationship in this dependency structure, a new dependency structure is constructed with the two components in the "follow-dependency" type. The dependency structures are classified. The dependency structures of the three types of "subject," "object," and "sub-or-obj" are traversed, and the two dependency structures with overlapping components are combined into triples. For dependency structures that cannot be combined into triples, the components that are not combined into triples are used as keywords and placed in the keyword container. The dependency structure of the "question" type is traversed, and the second part is used to generate the question item; the remaining dependency structures are discarded.
[0012] Step 2: Triple screening
[0013] For the triples obtained in step 1, use the forward maximum matching and reverse maximum matching methods to entity link each element in the triple to the components in the knowledge graph. Then obtain the schema-layer ontology of each entity and ontology in the knowledge graph. Traverse all triples and filter out triples that do not have the structure of <entity, relationship, entity>, as well as triples with the structure of <entity, relationship, entity> but where the ontologies of the two entities cannot be connected by a relationship in the schema layer. Discard these filtered triples.
[0014] Step 3: Triple conversion
[0015] Among the triples retained after screening in step 2, there are two types of triples: the first type of triples contains a named entity and a pattern layer ontology, and the second type of triples contains two pattern layer ontologies.
[0016] Use the first type of triples to generate a knowledge graph embedded in the retrieval formula for search. When the found content belongs to a certain pattern layer ontology in the second type of triples, substitute the found content into the position of the pattern layer ontology in the second type of triples, so that the second type of triples becomes a triple containing a named entity and a pattern layer ontology, that is, converting the second type of triples into the first type of triples.
[0017] Step 4: Query tag tree generation
[0018] After completing the transformation of the second type of triples into the first type of triples, all the found contents are saved in a tree structure through a recursive method, where tag is used to save the entity name represented by the current node, children is used to save the child nodes of the current node, parent is used to save the parent node of the current node, and value is used to save the distance between the entity represented by the current node and the entity represented by its parent node, forming a labeled tree.
[0019] Step 5: Query the tag tree filter
[0020] In the symbolic knowledge graph, search for the neighbors of the keyword stored in the keyword container in step 1. Then, in the query tag tree generated in step 4, search for the same node as the keyword's neighbor and use it as the tag node. At the same time, delete the other sibling nodes of the tag node in the query tag tree and all nodes under the sibling nodes. This continues until all keyword neighbors are found. Save the query tag tree after deleting some nodes.
[0021] Step 6: Sort query results
[0022] Normalize the distances stored in the nodes retained after step 5 in the query token tree. Target nodes are assigned to entities matching the question type in the token tree. Calculate the distance from each target node to the root node as the final distance for the entity in the query token tree. For nodes that appear repeatedly in the token tree, take the average of their distances to the root node as the final distance.
[0023] Sort all target nodes in the labeled tree by the number of occurrences from largest to smallest. For entities with the same number of target nodes, sort them by their final distance from smallest to largest. Return the sorted result to the user as the retrieval result.
[0024] In a second aspect, the present invention provides a medical entity retrieval system for implementing the above method, the system comprising:
[0025] The query token tree generation module parses the natural query statement and generates the corresponding knowledge graph embedded search formula according to the parsing results. It then calculates the results in the vectorized knowledge graph according to the formula and saves the search results into a tree structure to obtain the query token tree.
[0026] The query token tree screening module uses query intent that cannot be generated into the knowledge graph embedded in the retrieval formula as a keyword, and filters the nodes in the query token tree based on the keyword's neighbor information in the medical knowledge graph;
[0027] The query result sorting module filters the query tag tree and sorts the query results. It calculates the spatial distance of the remaining nodes in the query tag tree after filtering in the vectorized representation of the knowledge graph and sorts the nodes.
[0028] The present invention has the following beneficial effects:
[0029] 1. Use keyword containers to save some word information that cannot be directly used to generate search formulas, so as to avoid the loss of key information.
[0030] 2. During the screening process of the query tag tree, using the keywords stored in the keyword container to screen nodes can reduce a large amount of useless information, thereby further refining the retrieval results and making it easier for users to quickly obtain the required information from the returned content.
[0031] 3. By utilizing the characteristics of vectorized knowledge graph representation and sorting the search results using the spatial distance of entities in the vectorized knowledge graph, the search effect can be effectively improved. The more advanced the knowledge graph embedding model, the more significant the effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] FIG1 is a component syntax tree in an embodiment;
[0033] FIG2 is a schematic diagram of dependency syntax analysis in an embodiment;
[0034] FIG3 is a schematic diagram of a dependency structure obtained by analysis in the embodiment;
[0035] Figure 4 is a schematic diagram of the ternary combination;
[0036] FIG5 is a schematic diagram of a recursive method in an embodiment;
[0037] FIG6 is a schematic diagram of a query tag tree generated in an embodiment. DETAILED DESCRIPTION
[0038] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0039] The terms "including," "having," and any variations thereof, as used in the embodiments of this application, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not limited to the listed steps or units but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to the process, method, product, or apparatus.
[0040] The present invention will be further explained below with reference to the accompanying drawings:
[0041] The present invention provides a medical entity retrieval method based on knowledge graph embedding and keywords, specifically:
[0042] Step 1: Input question analysis
[0043] For the natural language question input by the user, the dependency relationship between the words in the natural language question is analyzed through dependency syntax, and the sentence structure is constructed based on the dependency relationship.
[0044] The dependency syntax first segments the sentence, marks the part of speech of each word, then selects a verb as the central word, establishes a dependency structure between other words and the central word based on the subject-predicate relationship, verb-object relationship, attributive-adverbial relationship, etc., and uses different types of dependency structures to represent the grammatical functions between words and the semantic structure of the sentence.
[0045] As shown in Figure 1, for the natural question "What medications do patients who underwent craniotomy take?", the sentence is first segmented to obtain "patients who underwent craniotomy took medications." Based on the segmentation results, each word is tagged with part-of-speech (POS): "<received / VV>, <done / AS>, <craniotomy / NN>, <of / DEC>, <patient / NN>, <took / VV>, <of / DEC>, <medication / NN>." Finally, the component syntactic tree is constructed using the POS tagging results.
[0046] The comparison of the parts of speech in Figure 1 is shown in Table 1.
[0047] Table 1
[0048] The component syntax tree reassembles the word segmentation results of a sentence into a tree structure based on the annotated parts of speech, allowing subsequent operations to more accurately capture the semantic relationships between words and the composition of sentences. The dependency structure between each word obtained based on the component syntax tree is shown in Figure 2, and the corresponding Chinese interpretations are shown in Table 2:
[0049] Table 2
[0050] First, we classify the various dependency structures according to the rules in Table 3. Based on these dependency structures, we construct triples, achieving similar results to the named entity recognition and relation extraction tasks. In addition to named entities and relations, dependency syntax also analyzes all components in a sentence, avoiding information omission.
[0051] Table 3
[0052] Among them, under the "subject" type, there are various dependency structure types related to the subject, which may include the head entity in the triple; under the "object" type, there are various dependency structure types related to the object, which may include the tail entity in the triple; under the "sub-or-obj" type, there are various dependency structure types related to adjectives or attributives, which may include the head entity or the tail entity in the triple; under the "question" type, there are dependency structure types that may include the purpose of the query, "constant" is a dependency structure type that can form a phrase, and "follow-dependency" is a dependency structure type that combines the two.
[0053] For example, for a natural question "who is the chief physician who is good at craniotomy", dependency syntactic analysis can obtain the five dependency structures shown in Figure 3. Among them, dobj(good at -1, craniotomy -2) is classified as object type, acl(doctor -5, good at -1) is classified as sub-or-obj type, root(root -0, doctor -5) is classified as question type, compound(doctor -5, director -4) is classified as constant type, and mark(good at -1, of -3) does not contain query intent. Therefore, no triple is generated and it is not stored in the keyword container and is directly discarded.
[0054] After classification is complete, for each set of dependency structures under the "subject" type, search for dependency structures with overlapping components in the dependency structures under the "object" or "sub-or-obj" types, and combine the two into a triple. Then, for each set of dependency structures under the "object" type, search for dependency structures with overlapping components in the dependency structures under the "sub-or-obj" type, and combine the two into a triple, as shown in Figure 4. For dependency structures where no overlapping parts can be found, the components that do not form a triple are placed in the keyword container. For dependency structures under the "question" type, question items are generated. Dependency structures that cannot be classified are directly discarded.
[0055] According to the above rules, dobj(expertise-1, craniotomy-2) and acl(doctor-5, expertise-1) in Figure 3 can be combined into a triple <doctor, expertise, craniotomy> due to the overlapping component "expertise-1". For the "question" type root(root-0,doctor-5), the question item "doctor" is generated. compound(doctor-5, director-4) cannot find other dependency structures that can be combined into a triple, and "doctor" has already been generated and included in the triple <doctor, expertise, craniotomy>, so "director" is stored in the keyword container. mark(expertise-1, de-3) does not contain any query intent, so it is neither generated into a triple nor stored in the keyword container.
[0056] Step 2: Triplet screening
[0057] Ideally, the triples obtained by combining in step 1 should have a standard <entity, relationship, entity> structure, and the corresponding ontologies of the two "entities" in the triples in the knowledge graph's schema layer can be connected through the "relationship" recorded in the triples. However, when faced with some complex sentences, the combined triples do not meet the above conditions. For example, although the format is <entity, relationship, entity>, there is actually no connection relationship in the knowledge graph schema layer, and other triples that do not conform to <entity, relationship, entity>, such as <relationship, entity, relationship>. These triples that do not meet the conditions are redundant information generated in the process of understanding the query statement and do not participate in subsequent retrieval.
[0058] The triple screening process involves entity linking each element of all combined triples to the components in the knowledge graph using forward and reverse maximum matching. Each entity in the triple is then retrieved, along with the corresponding ontology in the schema layer. All triples are then traversed, and triplets with a structure other than <entity, relationship, entity> and where the ontologies of two entities cannot be connected by a relationship in the schema layer are discarded.
[0059] Step 3: Triple conversion
[0060] Among the triples retained after screening in step 2, there are two types of triples: the first type of triples contains a named entity and a pattern layer ontology, and the second type of triples contains two pattern layer ontologies.
[0061] Use the first type of triples to generate a knowledge graph embedded in the retrieval formula for search. When the found content belongs to a certain pattern layer ontology in the second type of triples, substitute the found content into the position of the pattern layer ontology in the second type of triples, so that the second type of triples becomes a triple containing a named entity and a pattern layer ontology, that is, converting the second type of triples into the first type of triples.
[0062] Step 4: Query tag tree generation
[0063] After completing the transformation of the second type of triples into the first type of triples, all the found contents are saved in a tree structure through a recursive method, where tag is used to save the entity name represented by the current node, children is used to save the child nodes of the current node, parent is used to save the parent node of the current node, and value is used to save the distance between the entity represented by the current node and the entity represented by its parent node, forming a labeled tree.
[0064] As shown in Figure 5, {a, b, c, d} represents four different entities, {A, B, C, D} is the logical layer ontology of these entities in the knowledge graph, and v represents the distance between the current node and its parent node in space. First, traverse all triple sets and find the triples related to entity a1 corresponding to the logical layer ontology A.<a1,rAB,B> , generate node a1 as the child node of the root node, then calculate the spatial position of entity a1 after the relationship rAB changes by the formula a1×rAB, find the first k entities b1, b2, ..., bk closest to the position, and use these k entities as the child nodes of the node corresponding to entity a1 in the label tree, and then add the triple<B,rBC,C> Converted into a set of triples of length k<b1,rBC,C> ,<b2,rBC,C> ,...,<bk,rBC,C> , and then repeat the above steps to find the child nodes of entities b1, b2, ..., bk until all triples are calculated. The final label tree is shown in Figure 6. Among them, × is the operation of using relations for spatial movement in the knowledge graph embedding model.
[0065] Step 5: Query the tag tree filter
[0066] In the field of medical entity retrieval, the query intent implied by some search queries is difficult to generate into the knowledge graph embedded search algorithm. As a result, some query intent cannot be fully reflected in the knowledge graph search path, and the search results ultimately returned may cover a relatively wide range. This requires users to search through the numerous returned results for information that matches their true query intent, consuming additional time and effort.
[0067] For example, when a user enters "acute diseases requiring vascular surgery" as a query condition, existing methods will generate the search formula Edisease×R=Esurgery, where Edisease, Esurgery, and R represent the corresponding vectors of "disease," "vascular surgery," and "acceptance" in the knowledge graph vector space, respectively. However, the query intent of "acute" is difficult to generate into the search formula, resulting in the returned results including not only acute diseases such as aneurysm rupture and acute arterial embolism, but also chronic diseases such as vascular stenosis and varicose veins, requiring users to spend extra time screening the returned results.
[0068] In the knowledge graph, there are some attributes that only apply to specific entities. Although these attributes are also represented by triples, incorporating triples containing these attributes into the training of the knowledge graph embedding will not help improve the embedding effect. On the contrary, doing so not only wastes training resources, but may also have a negative impact on the spatial structure of the final vectorized knowledge graph. Although the symbolic knowledge graph does not allow the machine to understand the semantic information of entities and relationships like the vectorized knowledge graph, it is very suitable for preserving the personalized attributes of these entities. Therefore, when training the knowledge graph embedding, these attributes belonging to specific entities are usually retained in the symbolic knowledge graph. In the retrieval of medical knowledge graphs, there are many query intents that cannot be generated into the knowledge graph embedding retrieval formula, which are the attributes of these symbolically represented entities. Therefore, for the screening of the label tree, this method proposes to search for the neighbors of the keywords saved in the keyword container in step 1 in the symbolic knowledge graph, and use these neighbors to screen the label tree generated in step 4 to retain results that are more in line with the user's query intent. The specific steps are as follows:
[0069] Step 5.1: For the keywords stored in the keyword container, first use the forward maximum matching and reverse maximum matching methods to link these keywords with the components in the symbolic knowledge graph and find their neighbors in the knowledge graph. For a keyword key, it may be any of the knowledge graph triples, so different situations should be discussed separately:
[0070] ①When the key is subject, the neighbor information of the key is {(predicate,object)|(key,predicate,object)};
[0071] ②When the key is a predicate, the neighbor information of the key is {(subject, object)|(subject, key, object)};
[0072] ③When the key is object, the neighbor information of the key is {(subject, predicate)|(subject, predicate, key)};
[0073] ④ When the key is the ontology in the pattern layer, the neighbor information of the key is {(subject)|(subject,type,key)}.
[0074] Step 5.2: Based on the neighbor information for the keyword key, search the query tag tree generated in step 4 for nodes that are neighbors of the keyword. If all of the neighbors of a keyword cannot be found in the query tag tree, the keyword is considered unconnectable to the current query tag tree and is discarded. If a node corresponding to a neighbor of the keyword is found in the query tag tree, the node is marked. Meanwhile, all sibling nodes of the marked node and all nodes under the sibling node are deleted. This process continues until all neighbors of the keyword are found. The query tag tree after deleting some nodes is saved.
[0075] Step 6: Sort query results
[0076] Ranking entity retrieval results is crucial for improving retrieval system performance and user experience. By ranking the most relevant entities first, users can more quickly find information that meets their needs, reducing browsing time and effort during the retrieval process. Furthermore, clustering similar or duplicate entities together allows users to better understand the relationships and similarities between different entities in the search results, thus avoiding the waste of duplicate information.
[0077] Step 6.1: Measure the relevance of an entity to the user's search target by looking up the number of occurrences of the node representing the entity in the tag tree and the distance from the node to the root node. Before sorting begins, to prevent the average distance of a single jump from being too large and thus dominating the entire search result sort, it is necessary to normalize the distances stored in each node in the tag tree. This is done by traversing the entire tag tree starting from the root node. If the child node list of the current node is not empty, the distances of the nodes in the list are normalized. The specific formula is as follows:
[0078] Among them, v i is the distance from the i-th node in a child node list to its parent node, and N is the number of all child nodes in the child node list.
[0079] Step 6.2: After completing distance normalization, find the node corresponding to the entity that matches the question item type in the labeled tree and use it as the target node. Calculate the distance from each target node to the root node as the final distance of the entity. For nodes that appear repeatedly in the labeled tree, take the average of the sum of the distances from these repeated nodes to the root node as the final distance:
[0080] Among them, M represents the number of occurrences of the node of the current entity in the label tree, v k is the distance from the kth node to the root node, and distance is the final distance from the node representing the current entity to the root node.
[0081] Step 6.3: Sort the target nodes in the labeled tree by the number of occurrences from largest to smallest. For pairs with the same number of occurrences, sort them by their final distance from smallest to largest. Return the sorting result to the user as the search result.
[0082] To validate the effectiveness of this method, this example used electronic medical records of patients with cerebrovascular diseases, provided by a partner hospital, as experimental data to generate a corresponding knowledge graph. 609 search queries were designed, all of which were complex questions involving multiple triples. The search targets included medications, diseases, departments, and surgeries.
[0083] Entity retrieval aims to retrieve entities related to a user's query from a database containing entity information. Since entity retrieval typically returns a variable number of results, directly using the evaluation metrics Precision, Recall, and F1, which are designed for a fixed number of returned results, to judge the performance of the method may be problematic. Therefore, considering the impact of different numbers of returned results on the metrics, this example uses Precision@10, Recall@10, and F1@10 to characterize the test results.
[0084] Precision@10: The full name is the precision of the first 10 returned results. It is calculated by finding the ratio of the number of correct entities in the first 10 results of the query to the number of returned entities:
[0085] Where TP@10 is the number of correct entities in the top 10 results, and FP@10 is the number of incorrect entities in the top 10 results. P@10∈[0,1], where a higher value indicates that the entities returned in the top 10 results are more relevant.
[0086] Recall@10: The full name is the recall rate of the first 10 returned results. It is calculated by finding the ratio of the number of correct entities in the first 10 results of the query to the total number of correct entities:
[0087] FN@10 is the number of correct entities that do not appear in the first 10 returned results. R@10 also ranges from [0, 1]. A higher value indicates that more relevant entities are retrieved in the first 10 results.
[0088] F1@10: The full name is the balanced F score of the first 10 returned results. It is calculated by finding the harmonic mean of Precision@10 and Recall@10:
[0089] F1@10∈[0,1] combines the precision and recall of the top 10 results and is suitable for comprehensively evaluating the performance of the retrieval system.
[0090] This example selected Xiao's path fusion algorithm and KGEKS retrieval algorithm in the prior art and conducted comparative experiments on TransE, TransD, RotatE, and the optimized RotatE model. The experimental results are shown in Table 4:
[0091] Table 4
[0092] In order to verify the effectiveness of the retrieval tag tree screening step in this method, ablation experiments were also conducted. These included retrieval experiments with and without the tag tree screening process. Ablation experiments of the tag tree screening algorithm were also conducted on TransE, TransD, RotatE, and the optimized RotatE model. The experimental results are shown in Table 5.
[0093] Table 5
[0094] Experimental results show that the retrieval tag tree filtering algorithm can effectively improve the retrieval effect, and the retrieval tag tree filtering algorithm is usually more effective when the retrieval is performed on the more advanced knowledge graph embedding model.
[0095] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above-described embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-described embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be subject to the scope of protection of the claims.
Claims
1. A medical entity retrieval method based on knowledge graph embedding and keywords, characterized in that: It includes query tag tree generation, query tag tree screening, and query result sorting; The query tag tree generation parses the natural statement of the query and generates a corresponding knowledge graph embedding retrieval formula according to the parsing result, then calculates the result in the vectorized knowledge graph and saves the retrieval result in a tree structure; the query tag tree screening uses the query intention that cannot be generated into the knowledge graph embedding retrieval formula as a keyword, and screens the nodes in the query tag tree through the neighbor information of the keyword in the medical knowledge graph; the query result sorting calculates the spatial distance of the remaining nodes in the screened query tag tree in the vectorized knowledge graph, sorts the nodes, and returns the query result.
2. The medical entity retrieval method based on knowledge graph embedding and keywords according to claim 1, characterized in that: The specific steps are as follows: Step 1, Input question analysis For the natural language question input by the user, analyze the dependency relationship between words in the natural language question through dependency syntax analysis to obtain different dependency structures, and generate triples and question items according to the dependency structures; Step 2, Triple screening Use the forward maximum matching and reverse maximum matching methods to perform entity linking between the elements in the triples and the components in the knowledge graph; traverse all triples, screen out the triples whose structure is not <entity, relationship, entity>, and the triples whose structure is <entity, relationship, entity> but the two entities' ontologies cannot be connected through the relationships within the group in the schema layer, and discard them; the remaining triples include the first type of triples containing one named entity and one schema layer ontology, and the second type of triples containing two schema layer ontologies; Step 3, Triple transformation Use the first type of triples to generate an embedding retrieval formula to search in the knowledge graph. When the found content belongs to the schema layer ontology in the second type of triples, substitute the found content into the position of the schema layer ontology in the triple, so that the second type of triples becomes a triple containing one named entity and one schema layer ontology; repeat the above process until all the second type of triples are transformed into triples containing one named entity and one schema layer ontology; Step 4, Query tag tree generation Save the content found in Step 3 in a tree structure through a recursive method to form a query tag tree; Step 5, Query tag tree screening Find the neighbors of the keywords saved in the keyword container in Step 1 in the symbolically represented knowledge graph, and then find the nodes in the query tag tree generated in Step 4 that are the same as the neighbors of the keywords as the marked nodes, and at the same time delete other sibling nodes of the marked nodes in the query tag tree and all nodes under the sibling nodes; after finding the neighbors of all keywords, save the query tag tree after deleting some nodes; Step 6, Query result sorting Take the nodes corresponding to the entities that meet the question item type in the marked tree as the target nodes, sort the entities in descending order according to the number of times the target nodes corresponding to the entities appear, and return the query result.
3. The medical entity retrieval method based on knowledge graph embedding and keywords according to claim 2, characterized in that: In step 1, first, for two components with a parallel relationship in the "follow-dependency" type, find the dependency structures with the same components from the remaining dependency structures, and use the relationships in these dependency structures to construct a new dependency structure with the two components in the "follow-dependency" type respectively; Then classify the dependency structures, traverse the dependency structures of the three types of "subject", "object", and "sub-or-obj", and combine the two dependency structures with overlapping components into triples; for the dependency structures that cannot be combined into triples under these types, use the components that have not been combined into triples as keywords and put them into the keyword container; traverse the dependency structures of the "question" type and use the second component in them to generate question items; discard the remaining dependency structures.
4. The medical entity retrieval method based on knowledge graph embedding and keywords according to claim 2 or 3, characterized in that: Classify the dependency structure according to the records in the following table:
5. The medical entity retrieval method based on knowledge graph embedding and keywords according to claim 2, characterized in that: In the tree structure, use tag to save the entity name represented by the current node, children to save the child nodes of the current node, parent to save the parent node of the current node, and value to save the distance between the entity represented by the current node and the entity represented by its parent node.
6. The medical entity retrieval method based on knowledge graph embedding and keywords according to claim 2, characterized in that: The specific method for querying and marking tree screening is: Step 5.1: For the keywords saved in the keyword container, first use the forward maximum matching and reverse maximum matching methods to link with the components in the symbolically represented knowledge graph to find the neighbors of the keywords in the knowledge graph; Step 5.2: According to the neighbor information of the keywords, find the nodes in the query marking tree generated in step 4 that are the same as the neighbors of the keywords. If all the neighbors of a keyword cannot be found in the query marking tree, then it is considered that the keyword cannot be connected to the current query marking tree and it is discarded; If the nodes corresponding to the neighbors of the keywords are found in the query marking tree, then mark these nodes, and at the same time delete the other sibling nodes of the marked nodes and all the nodes under the sibling nodes until all the neighbor searches of the keywords are completed, and save the query marking tree after deleting some nodes.
7. The medical entity retrieval method based on knowledge graph embedding and keywords according to claim 2 or 6, characterized in that: The neighbor information of the keyword key is related to the type of the keyword: ① When key is subject, the neighbor information of key is {(predicate, object)|(key, predicate, object)}; ② When the key is a predicate, the neighbor information of the key is {(subject, object)|(subject, key, object)}; ③ When the key is an object, the neighbor information of the key is {(subject, predicate)|(subject, predicate, key)}; ④ When the key is an ontology in the schema layer, the neighbor information of the key is {(subject)|(subject, type, key)}.
8. The medical entity retrieval method based on knowledge graph embedding and keywords according to claim 2, characterized in that: After the screening in step 5, the distances remaining in the query marker tree are normalized.
9. The medical entity retrieval method based on knowledge graph embedding and keywords according to claim 2 or 8, characterized in that: For entities with the same number of occurrences of the target node, calculate the average value of the distances between the corresponding multiple target nodes and the root node respectively as the final distance of the entity, and sort the entities with the same number of occurrences of the target node in ascending order of the final distance.
10. A medical entity retrieval system for implementing the method according to any one of claims 1-9, characterized in that the system includes: A query marker tree generation module that generates a corresponding knowledge graph embedding retrieval formula by parsing the natural statement of the query, then calculates the result in the vectorized knowledge graph according to the formula, and saves the retrieval result into a tree structure to obtain a query marker tree; A query marker tree screening module that uses the query intent that cannot be generated into the knowledge graph embedding retrieval formula as a keyword, and screens the nodes in the query marker tree through the neighbor information of the keyword in the medical knowledge graph; A query result sorting module that calculates the spatial distances of the remaining nodes in the screened query marker tree in the vectorized knowledge graph and sorts the nodes.
Citation Information
Patent Citations
Knowledge graph construction method and system and information query method and system
CN112765288A
Ontology label knowledge graph-oriented sample query method
CN113569057A
Commodity relevance query system and method based on keyword search
CN114372156A
Intelligent matching system with ontology-aided relation extraction
US20180232443A1
Cited By
Data query method for medical experiment system
CN120849678A
Intelligent retrieval and reasoning generation method and system based on knowledge graph and geoscience
CN121212352A
Medical question and answer method based on knowledge graph and retrieval enhancement generation
CN121390321A