Medical knowledge graph completion method based on low-rank multi-order linear pooling representation

Through the low-rank multi-order linear pooling representation method, combined with the first-order information and attention mechanism of the entity, a medical knowledge graph link prediction model is constructed, which solves the shortcomings of the existing model in entity feature representation and graph structure understanding, and improves the performance and accuracy of link prediction.

CN120258111APending Publication Date: 2025-07-04JIANGSU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510327612.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing medical knowledge graph model has shortcomings in entity feature representation and graph structure understanding, resulting in limited link prediction performance and inability to fully and accurately characterize entity semantics and understand complex relationships between entities.

Method used

A low-rank multi-order linear pooling representation method is adopted, combined with the entity's first-order information and attention mechanism, a link prediction model is built, and the model parameters are trained using binary cross entropy loss function to complete the medical knowledge graph.

Benefits of technology

It enhances the model's ability to model feature interaction relationships, improves the performance of link prediction, ensures a comprehensive understanding and accurate prediction of entity semantics, and adapts to the complex structure of the medical knowledge graph.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258111A_ABST
    Figure CN120258111A_ABST
Patent Text Reader

Abstract

The invention provides a medical knowledge graph completion method based on low-rank multi-order linear pooling representation, and the method comprises the steps: carrying out the preprocessing of a medical knowledge graph data set, and obtaining an entity embedding matrix and a relation embedding matrix; introducing first-order information and an attention mechanism of an entity to construct a low-rank multi-order linear pooling representation link prediction model; using a binary cross entropy loss function to train link prediction model parameters; performing a knowledge graph link prediction experiment by using the trained link prediction model; and complementing the medical knowledge graph by using the trained low-rank multi-order linear pooling representation link prediction model. According to the method, the modeling capability of the link prediction model for the interaction relationship between the features is enhanced, so that the link prediction model is more suitable for the complex interaction relationship between the features in actual data, and the link prediction performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to knowledge graph completion technology, and specifically, to a medical knowledge graph completion method based on low-rank multi-order linear pooling representation. Background Art

[0002] Knowledge Graphs (KGs) have broad applications in both industrial and academic fields. However, the construction of knowledge graphs usually depends on specific data sources and construction methods, so it may not cover all fields and topics of knowledge. At the same time, its maintenance involves a large amount of data processing and analysis work, requiring a large amount of manpower, material resources and time. The link prediction (LP) task on the knowledge graph aims to infer missing facts from existing facts, generally by scoring relationship and entity triples to predict their authenticity, thus avoiding the cost and time of manually completing the knowledge graph.

[0003] Currently, in the field of medical knowledge graph link prediction technology, the mainstream models cover linear models and non-linear models. The bilinear model has been widely used in the link prediction task due to its unique expression properties. This is mainly because the entities and relationships in the medical knowledge graph can be regarded as features from different modalities, and their effective fusion is crucial for improving the model performance. Although the existing bilinear models have shown great potential in both practical applications and theoretical research, it is undeniable that these methods still have certain limitations:

[0004] On the one hand, as the basic unit of the medical knowledge graph, the intrinsic features of entities play an indispensable role in comprehensively and accurately depicting entity semantics. Existing methods often ignore the intrinsic features of entities themselves, which results in insufficient feature representation ability after fusion, and thus restricts the final performance of link prediction.

[0005] On the other hand, the entities in the medical knowledge graph are interconnected through relationships, forming a complex network structure. Under this structure, the semantics of each entity are not only determined by its own attributes, but also deeply affected by its neighboring entities. However, the existing embedding models often only focus on the fitting degree between triples (head entity, relationship, tail entity), but ignore the position of entities in the entire graph structure and the association information with other entities, resulting in a certain degree of loss of semantic information and affecting the model's comprehensive understanding and accurate prediction of entity semantics. Summary of the Invention

[0006] Aiming at the deficiencies existing in the prior art, the present invention provides a medical knowledge graph completion method based on low-rank multi-order linear pooling representation.

[0007] The present invention achieves the above technical objectives through the following technical means.

[0008] A medical knowledge graph completion method based on low-rank multi-order linear pooling representation, comprising:

[0009] Preprocessing the medical knowledge graph dataset to obtain an entity embedding matrix and a relationship embedding matrix; introducing the first-order information of entities and an attention mechanism to construct a low-rank multi-order linear pooling representation link prediction model; training the parameters of the link prediction model using a binary cross-entropy loss function; conducting knowledge graph link prediction experiments using the trained link prediction model; and completing the medical knowledge graph using the trained low-rank multi-order linear pooling representation link prediction model.

[0010] Further, the entity embedding matrix includes a head entity embedding vector and a tail entity embedding vector.

[0011] Further, the link prediction model includes a feature fusion layer, an aggregation layer, and a link prediction layer.

[0012] Furthermore, the feature fusion layer adds the first-order information of entities on the basis of a bilinear model to obtain the fusion feature e of the head entity and the relationship, and the expression form is as follows: h+r where

[0013]

[0014] SumPooling is a feature pooling method, W′, U′, and V′ are parameter matrices obtained by concatenating and reshaping i Ws, i Us, and i Vs respectively, W is a first-order parameter matrix, U and V are two low-rank matrices obtained by decomposing the second-order parameter matrix W1, e h is the embedding vector of the head entity, e r is the embedding vector of the relationship, is the Hadamard product, and k is the decomposition rank.

[0015] Furthermore, the aggregation layer is based on the attention mechanism, uses the fusion feature to calculate the neighbor embedding vector of the head entity, and aggregates the neighbor embedding vector and the head entity embedding vector to obtain the aggregation feature f GCN and the expression form is as follows:

[0016]

[0017] where is the neighbor embedding vector of the head entity; π′(h,r,t) is the decay factor that controls the propagation from t to h on the triple (h,r,t) conditioned on the relation r; π′(h,r′,t′) is the decay factor that controls the propagation from t′ to h on the triple (h,r′,t′) conditioned on the relation r′, which can be obtained in the same way as π′(h,r,t); (h,r′,t′) and (h,r,t) are different triples sharing the same head entity; e t is the embedding vector of the tail entity, N h is the set of triples sharing the same head entity, LeakyReLU is the activation function set, and M is the trainable weight matrix.

[0018] Furthermore, the link prediction layer uses the aggregated features obtained by the aggregation layer to obtain the probability p of link prediction through the knowledge graph link prediction scoring function, and the expression form is as follows:

[0019]

[0020] Among them, g(f GCN ,e r ) is a function about f GCN ,e r , and f(e h ,e r ,e t ) is the scoring function.

[0021] Furthermore, the formula of the binary cross-entropy loss function is as follows:

[0022]

[0023] Among them, m is the number of batches, B is a batch of data, n t is the number of tail entities, y i is the target label of the given entity relation pair (e i ,e h ,e r ) of the tail entity e is the predicted probability value.

[0024] Furthermore, the link prediction model can predict the head entity according to the relation and the tail entity, and at the same time the relation becomes the original inverse relation.

[0025] Furthermore, the scoring metrics of the link prediction model are MRR and HITS@n.

[0026] Furthermore, the steps for completing the medical knowledge graph are as follows: First, calculate the scores of the missing triples in the medical knowledge graph through the trained link prediction model, and obtain the completed triples by threshold screening; Second, the completed triples are verified by domain experts to obtain the verified triples; Finally, the verified triples are added to the existing medical knowledge graph to update the medical knowledge graph.

[0027] Advantages of the present invention:

[0028] (1) The model designed in the present invention adds the first-order information of entities on the basis of the original model, ensuring that the fusion features of entity relationships are presented in a multi-order form, enhancing the model's ability to model the interaction relationships between features, and making the model more adaptable to the complex interaction relationships between features in actual data.

[0029] (2) The present invention introduces an attention mechanism to aggregate the neighbor features of nodes, assigns different weights to the features of different neighbor nodes, and obtains a node embedding representation that fully integrates neighbor information, improving the performance of link prediction. Description of the drawings

[0030] Figure 1 It is a flowchart for constructing a low-rank multi-order linear pooling representation knowledge graph link prediction model described in this embodiment.

[0031] Figure 2 It is a flowchart for completing the medical knowledge graph described in this embodiment. Detailed implementation manners

[0032] The present invention will be further described below in conjunction with the drawings and specific embodiments, but the protection scope of the present invention is not limited thereto.

[0033] A medical knowledge graph completion method based on low-rank multi-order linear pooling representation includes the following steps:

[0034] Step 1: Prepare medical knowledge graph data, preprocess it to obtain a medical knowledge graph dataset G, and initialize all entities in the dataset as d e dimensional embedding vectors, and initialize the relationships as d r dimensional embedding vectors. Concatenate all entity embedding vectors to obtain an entity embedding matrix Concatenate all relationship embedding vectors to obtain a relationship embedding matrix where n e and n r are the numbers of entities and relationships respectively, and d e and d r are the dimensions of entities and relationships respectively. Given an input triple (h, r, t), the embedding vector e of the head entity h can be obtained through the entity embedding matrix hand the embedding vector e of the tail entity t t , the embedding vector e of the relation r can be obtained through the relation embedding matrix r .

[0035] Step 2: Construct a low-rank multi-order linear pooling representation knowledge graph link prediction model, as Figure 1 shown.

[0036] The low-rank multi-order linear pooling representation knowledge graph link prediction model includes three parts: a feature fusion layer, an aggregation layer, and a link prediction layer. Among them, the feature fusion layer adds the first-order information of entities to the bilinear model to obtain the fusion features of entities and relations; the aggregation layer, based on the attention mechanism, uses the fusion features to calculate the neighbor embedding vectors of entities, and then aggregates the neighbor embedding vectors and entity embedding vectors to obtain the aggregation features; the link prediction layer uses the aggregation features obtained by the aggregation layer to obtain the probability of link prediction through the knowledge graph link prediction scoring function. The processing processes of each part in the model are as follows:

[0037] (1) Processing process of the feature fusion layer

[0038] Use e h+r to represent the fusion features of the head entity and the relation, and express them using the low-rank multi-order linear pooling method:

[0039]

[0040] Among them, 1 T is a vector of all 1s, W is the first-order parameter matrix, U and V are two low-rank matrices obtained by decomposing the second-order parameter matrix W1, the decomposition rank is k, is the Hadamard product.

[0041] At this time, the dimension of e h+r is one-dimensional. In order to adjust the dimension of the fusion feature e h+r to i dimensions, i Ws, i Us, and i Vs are respectively concatenated and reshaped to obtain the parameter matrices W′, U′, and V′, and then the sum of non-overlapping windows of size k on the Hadamard product is calculated, which not only enhances the modeling ability of the bilinear model for the interaction relationship between features, but also makes the bilinear model more adaptable to the complex interaction relationship between features in the actual data, so as to obtain the final fusion features. The expression form of the final fusion features is as follows:

[0042]

[0043] Among them, SumPooling is a feature pooling method.

[0044] (2) Processing process of the aggregation layer

[0045] First, to enhance the expressive ability of the head entity embedding vector, this embodiment makes full use of the neighbor information of the head entity, introduces an attention mechanism, and obtains a neighbor embedding vector of the head entity that fully integrates the neighbor information. The neighbor embedding vector of the head entity is represented as the weighted sum of the tail entities of the triples in the dataset that share the same head entity. The formula is as follows:

[0046]

[0047] Wherein, is the neighbor embedding vector of the head entity; π′(h,r,t) is the attenuation factor that controls the propagation from t to h on the triple (h,r,t) conditional on the relation r, and π′(h,r′,t′) is the attenuation factor that controls the propagation from t′ to h on the triple (h,r′,t′) conditional on the relation r′; (h,r′,t′) and (h,r,t) are different triples that share the same head entity, and N h is the set of triples that share the same head entity; π′(h,r,t) is realized through a relation attention mechanism, and its formula is:

[0048] π′(h,r,t) = e t T tanh(e h+r )

[0049] π′(h,r′,t′) can be obtained in the same way. This embodiment selects the tanh function as the non-linear activation function, so that the attention score depends on the distance between the head entity and the tail entity in the relation space. The closer the entities are, the more information they propagate.

[0050] Then, use the graph convolutional network to and e h to aggregate and obtain the aggregated feature f GCN :

[0051]

[0052] Wherein, LeakyReLU is the activation function set, and M is the trainable weight matrix.

[0053] Finally, use the aggregated feature to directly replace the head entity embedding vector in the entity embedding matrix, thereby updating the entity embedding matrix, and use the updated entity embedding matrix for the processing of the link prediction layer.

[0054] (3) Processing process of the link prediction layer

[0055] First, for a given triple (h,r,t), its scoring function is defined as:

[0056] f(e h ,e r ,et ) = g(f GCN , e r ) · e t = g(f GCN , e r ) T e t

[0057] Among them, g(f GCN , e r ) is a function of f GCN , e r , defined as:

[0058]

[0059] Secondly, the scoring function f(e h , e r , e t ) is transformed into a probability p through the sigmod activation function. The formula is as follows:

[0060]

[0061] Among them, is the probability that the triple (h, r, t) is a factual triple. The greater the probability, the higher the likelihood that the triple is a factual triple.

[0062] Step 3: Use the binary cross-entropy loss function as the loss function of the link prediction model. Train the parameters in the model by minimizing the loss function. The training parameters include W, U, V, the entity embedding matrix and the relation embedding matrix, and the weight matrix M. The binary cross-entropy loss function formula is as follows:

[0063]

[0064] Among them, m is the number of batches, B is a batch of data, n t represents the number of tail entities; y i is the target label of the given entity-relation pair (e i , e h , e r ) of the tail entity e h , e r , e i ). If (e i , e h , e r , e i ) is a factual triple, then y i = 1. If (e is the predicted probability value.

[0065] Step 4: Use the trained low-rank multi-order linear pooling representation link prediction model to conduct knowledge graph link prediction experiments.

[0066] For the link prediction experiment, the dataset used is the standard knowledge graph dataset: WN18RR, FB15k-237, WN18, FB15k. The head entity embedding vector and the relation embedding vector are fed into the model of this embodiment through the entity embedding matrix and the relation embedding matrix, and the probabilities of all entities in the knowledge graph as the tail entities corresponding to the current head entity and relation entity are calculated, and the entity with the highest probability is found and output as the result of link prediction. At the same time, the head entity can also be predicted according to the relation and the tail entity, and the relation becomes the original inverse relation, that is, the triple (t, r -1 , h) is input into the model of the present invention, where the relation r -1 's embedding vector is represented as -e r .

[0067] In addition, compare the prediction results obtained by the method of this embodiment and the existing 4 mainstream methods (HypER, DistMult, ComplEx, and TuckER). The scoring metrics used for comparison are MRR (Mean Reciprocal Ranking) and HITS@n (hit rate in the top n), and their formulas are as follows:

[0068]

[0069] Among them, G is the set of triples, |G| is the number of triples, rank i refers to the link prediction ranking of the i-th triple, and the larger the value of this metric, the better. is the indicator function (the function value is 1 if the condition is true, and the function value is 0 if the condition is false). Generally, n is taken to be equal to 1, 3, and 10.

[0070] The comparison results are shown in Table 1, and the results prove the effectiveness of the method of this embodiment in link prediction and can significantly enhance the quality of link prediction in the field of knowledge graphs.

[0071] Table 1 Comparison of experimental effects obtained by the method of this embodiment and 4 mainstream methods in link prediction experiments on WN18RR, FB15k-237, WN18, and FB15k datasets through MRR and Hit@k scoring metrics

[0072]

[0073] Step 5: Use the trained low-rank multi-order linear pooling representation link prediction model to complete the medical knowledge graph, as Figure 2 shown.

[0074] First, use the trained link prediction model to predict the missing links in the medical knowledge graph: for each possible triple (head entity, relation, tail entity) that may be missing in the medical knowledge graph, the link prediction model outputs a score indicating the likelihood of this triple being a missing triple in the medical knowledge graph; according to the scores output by the link prediction model, set a threshold to determine whether the triple is added to the medical knowledge graph.

[0075] Second, due to the particularity of the medical field, the completed triples ultimately need to be verified by domain experts to ensure their accuracy and reliability.

[0076] Finally, add the verified completed triples to the existing medical knowledge graph to update the medical knowledge graph. The updated medical knowledge graph can be used for various downstream tasks, such as disease diagnosis, drug recommendation, clinical decision support, etc.

[0077] The described embodiments are the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Without departing from the essence of the present invention, any obvious improvements, substitutions or modifications that those skilled in the art can make all fall within the protection scope of the present invention.

Claims

1. A medical knowledge graph completion method based on low-rank multi-order linear pooling representation, characterized in that: Preprocess the medical knowledge graph dataset to obtain an entity embedding matrix and a relationship embedding matrix; introduce the first-order information of entities and an attention mechanism to construct a low-rank multi-order linear pooling representation link prediction model; use a binary cross-entropy loss function to train the parameters of the link prediction model; use the trained link prediction model to conduct knowledge graph link prediction experiments; use the trained low-rank multi-order linear pooling representation link prediction model to complete the medical knowledge graph.

2. The medical knowledge graph completion method based on low-rank multi-order linear pooling representation according to claim 1, characterized in that The entity embedding matrix includes a head entity embedding vector and a tail entity embedding vector.

3. The medical knowledge graph completion method based on low-rank multi-order linear pooling representation according to claim 1, characterized in that, The link prediction model includes a feature fusion layer, an aggregation layer, and a link prediction layer.

4. The medical knowledge graph completion method based on low-rank multi-order linear pooling representation according to claim 3, characterized in that, The feature fusion layer adds the first-order entity information to the bilinear model to obtain the fused features e of the head entity and the relation h+r , and the expression form is as follows: Among them, SumPooling is a feature pooling method, W′, U′, and V′ are parameter matrices obtained by concatenating and reshaping i Ws, i Us, and i Vs respectively. W is a first-order parameter matrix, and U and V are two low-rank matrices obtained by decomposing the second-order parameter matrix W1, and e h is the embedding vector of the head entity, e r is the embedding vector of the relation, is the Hadamard product, and k is the decomposition rank.

5. The medical knowledge graph completion method based on low-rank multi-order linear pooling representation according to claim 4, wherein The aggregation layer is based on the attention mechanism, uses the fused features to calculate the neighbor embedding vector of the head entity, and aggregates the neighbor embedding vector and the head entity embedding vector to obtain the aggregated feature f GCN , and the expression form is as follows: Among them, is the neighbor embedding vector of the head entity; π′(h,r,t) is the attenuation factor that controls the propagation from t to h on the triple (h,r,t) conditioned on the relation r; π′(h,r′,t′) is the attenuation factor that controls the propagation from t′ to h on the triple (h,r′,t′) conditioned on the relation r′, which can be obtained in the same way as π′(h,r,t); (h,r′,t′) and (h,r,t) are different triples sharing the same head entity; e t is the embedding vector of the tail entity, N h is the set of triples sharing the same head entity, LeakyReLU is the activation function set, and M is the trainable weight matrix.

6. The medical knowledge graph completion method based on low-rank multi-order linear pooling representation according to claim 5, wherein The link prediction layer uses the aggregated features obtained by the aggregation layer, and obtains the probability p of link prediction through the knowledge graph link prediction scoring function. The expression form is as follows: Among them, g(f GCN ,e r ) is a function with respect to f GCN ,e r , and f(e h ,e r ,e t ) is a scoring function.

7. The medical knowledge graph completion method based on low-rank multi-order linear pooling representation according to claim 6, characterized in that The formula of the binary cross-entropy loss function is as follows: Among them, m is the number of batches, B is the data of one batch, and n t is the number of tail entities, and y i is the given entity relationship pair (e i , e h , e r ) of the tail entity e, and is the predicted probability value.

8. The medical knowledge graph completion method based on low-rank multi-order linear pooling representation according to claim 6, wherein The link prediction model can predict the head entity according to the relationship and the tail entity, and at the same time the relationship becomes the original inverse relationship.

9. The medical knowledge graph completion method based on low-rank multi-order linear pooling representation according to claim 6, wherein The scoring metrics of the link prediction model are MRR and HITS@n.

10. The medical knowledge graph completion method based on low-rank multi-order linear pooling representation according to claim 6, characterized in that, The steps for completing the medical knowledge graph are as follows: First, calculate the scores of the missing triples in the medical knowledge graph through the trained link prediction model, and obtain the completed triples through threshold screening; Second, the completed triples are verified by domain experts to obtain verified triples; Finally, add the verified triples to the existing medical knowledge graph to update the medical knowledge graph.