Electronic medical record knowledge graph similar doctor seeing sub-graph identification method based on graph neural network

By applying graph neural networks in the electronic medical record knowledge graph, integrating the characteristics of nodes and edges and multi-layer attention mechanisms, we can identify and eliminate redundant similar medical subgraphs, and solve the problems of waste of storage space and reduced statistical reasoning accuracy caused by redundant subgraphs in the knowledge graph, and achieve more efficient identification and elimination effects.

CN120072167APending Publication Date: 2025-05-30BEIJING UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510212806.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

There are redundant similar medical subgraphs in the electronic medical record knowledge graph, resulting in waste of storage space and reduced statistical and inference accuracy based on EMR-KG. It is difficult for existing traditional methods to effectively identify these redundant subgraphs.

Method used

A method for identifying subgraphs of similar medical records based on graph neural networks was designed. By integrating the attribute characteristics of nodes and edges and a multi-layer attention mechanism, the similarity of medical subgraphs was comprehensively compared to improve the accuracy and efficiency of recognition of similar subgraphs.

Benefits of technology

Effectively identify and eliminate redundant subgraph data, improving the storage efficiency of electronic medical record knowledge graphs and the accuracy of statistics and reasoning based on this graph.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072167A_ABST
    Figure CN120072167A_ABST
Patent Text Reader

Abstract

The invention provides an electronic medical record knowledge graph similar doctor seeing sub-graph identification method based on a graph neural network. According to the method, the attribute text features of the graph are fused by using the graph attention network, so that the comprehensiveness of graph feature expression is improved; while feature fusion is carried out, aggregation of graph similarity key information is realized through a multi-level attention mechanism, and the problem of difficult recognition caused by graph topology diversity and text expression diversity is solved; and finally, comprehensively measuring the similarity of the graph by adopting two levels of similarity information of the graph and the node during graph similarity calculation, and solving the problem that the model loses detail attribute text difference in the convergence process of expressing the similarity key information to influence graph similarity judgment. Experimental results show that the method can effectively identify the redundant similar sub-graphs in the electronic medical record knowledge graph, provides effective technical support for redundancy elimination and knowledge fusion of the medical knowledge graph, and has wide application prospects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Based on the graph neural network, the present invention designs a method for identifying similar visit subgraphs in an electronic medical record knowledge graph, which is used to identify redundant similar visit subgraphs in the electronic medical record knowledge graph, and belongs to both the technical field of medical knowledge graphs and the technical field of graph neural networks. Background Art

[0002] The electronic medical record knowledge graph (EMR-KG) is a multi-level and multi-dimensional knowledge graph structure constructed based on patients' electronic medical record data, which is used to represent the relationships between medical information such as symptoms, diseases, and treatments. It represents different medical entities through nodes, such as diseases, symptoms, treatment methods, etc., and represents the relationships between them through edges, which can provide support for medical decision-making and data resource management, and improve the efficiency and accuracy of medical services. However, the repeated descriptions of the same experience during multiple patient visits will lead to many redundant similar visit subgraphs in the EMR-KG, wasting storage space and reducing the accuracy of statistics and reasoning based on the EMR-KG. Therefore, a method is needed to identify similar visit subgraphs in order to eliminate redundant subgraph data subsequently.

[0003] The traditional methods for identifying similar graph structures mainly include graph similarity calculation methods based on graph edit distance and maximum common subgraph. The calculation method based on graph edit distance measures the similarity between graphs by taking the lowest cost of the edit operations required to transform the input graph into the target graph as the distance between the graphs. The calculation method based on the maximum common subgraph measures the similarity between two graphs by comparing the sizes of the maximum common subgraphs in the two graphs. Compared with the graph edit distance, its advantage is that it does not require defining edit operations and their costs. Although the efficiency of these two traditional graph similarity calculation methods has been greatly improved after continuous optimization by predecessors, they still have exponential time complexity and are difficult to judge the similarity relationships between diverse text attributes in the graph, and are not applicable to the identification of similar visit subgraphs.

[0004] In recent years, the rapidly developing graph neural network (Graph Neural Network, GNN) has brought new methods for graph similarity calculation. GNN is a deep learning model specialized for processing graph-structured data. Different from traditional neural networks, GNN can directly operate on graphs, utilize the topological structure and node attributes of the graphs, and update the features of each node iteratively to aggregate the information of node neighbors to learn the representations of nodes, edges, or the entire graph. Currently, many scholars have applied graph neural networks to graph similarity calculation, but most of them are for graph structures with relatively single node features, without considering the diverse text features in the graph and the impact of graph topological heterogeneity on similarity recognition, and without fully considering the impact of edge features on graph similarity calculation, and are not suitable for the similarity recognition of visit subgraphs.

[0005] In summary, the present invention will design a method for identifying similar visit subgraphs based on a graph neural network to complete the task of identifying similar visit subgraphs, so as to eliminate redundant subgraph data subsequently. Summary of the Invention

[0006] The purpose of the present invention is to propose a method for identifying similar visit subgraphs in an electronic medical record knowledge graph based on a graph neural network. By fusing the attribute features of nodes and edges and a multi-layer attention mechanism, the similarity of visit subgraphs is comprehensively compared, and the accuracy and efficiency of similar subgraph identification are improved.

[0007] To achieve the above object, the technical solution adopted by the present invention includes the following steps:

[0008] Step 1: Extract visit subgraph information from the electronic medical record knowledge graph according to the definition of the visit subgraph ontology model;

[0009] Step 2: Use the pre-trained text embedding model Sentence-BERT and unique encoding to generate node features and edge features for each visit subgraph;

[0010] Step 3: Introduce a graph attention network (EGAT) that can fuse edge features. During the feature aggregation process, use the attention mechanism to learn the importance of adjacent nodes and associated edges, and integrate the edge features into the node features; let the node and edge feature matrices pass through a multi-layer graph attention network (EGATs), and iteratively update the visit subgraph node feature matrix to a node full-feature matrix;

[0011] Step 4: Extract global key features from the node full-features of each visit subgraph through the attention mechanism, and generate a graph-level feature vector by weighted summation;

[0012] Step 5: Use the graph-level feature vector to calculate graph-level similarity score information in combination with a neural tensor network; obtain node-level similarity score information by using the node full-features and through histogram analysis;

[0013] Step 6: Input the above two types of score information into a multi-layer perceptron, and finally output the comprehensive similarity determination result of two visit subgraphs; use a full-supervised learning strategy to train and optimize the entire model, and use the optimized model for the actual task of identifying similar visit subgraphs;

[0014] Among them, in Step 2, the node feature matrix is represents the feature vector of node v i , where n is the number of nodes; the edge feature matrix is represents the feature vector of edge e j , where m is the number of edges; F is the number of feature dimensions of nodes and edges;

[0015] In Step 3, the visit subgraph node full-feature matrix is It is obtained by iteratively updating the node feature H by integrating the edge feature E through a multi-layer graph attention network (EGATs). Denotes the node v i The feature vector after fusing the neighbor nodes and the connected edge features. The attention coefficient α in a single-layer EGAT ij The calculation formula is as follows:

[0016]

[0017] Where: α ij Is the attention coefficient, denoting the connected edge e ij And the adjacent node v j To the node v i Of importance; Denotes the node v i 's feature vector; N(v i ) Denotes the set of neighbor nodes of the node v i Both the node v j And the node v k Are all nodes among them; Denotes the node v j 's feature vector; Denotes the node v k 's feature vector; Denotes v i And v j The feature vector of the edge between them; Denotes v i And v k The feature vector of the edge between them; || Denotes the vector concatenation operation; Is a learnable weight matrix. Q is the number of feature dimensions of the attention hidden layer of the EGAT model. In this method, its value is specified as 64; Is the weight vector used to parameterize the attention;

[0018] The updated node feature in EGAT is the full feature vector The feature aggregation calculation formula is as follows:

[0019]

[0020] Where: Is a learnable weight matrix. D represents the number of output dimensions of the EGAT; Is the zero vector;

[0021] The graph-level feature vector generated by using the attention mechanism in step 4 can represent the entire subgraph feature and contains key information for graph similarity calculation;

[0022] In step 5, the neural tensor network can measure the similarity relationship between the graph-level feature vectors of two subgraphs from multiple levels;

[0023] In step 6, a fully supervised learning strategy is adopted, and the cross-entropy loss function is used as the loss function of the model. The cross-entropy between the output result of the multi-layer perceptron and the similarity label between the two medical visit subgraphs is calculated to guide the update and adjustment of the parameters of each module of the entire model.

[0024] The method of the present invention is different from the traditional method of calculating graph similarity based on graph neural networks. It aims at the medical visit subgraphs of the electronic medical record knowledge graph containing a large number of text attributes, rather than the simple graph structure of node features. The present invention can effectively capture the subtle differences between the structural information and text attributes in the graph, and accurately integrate key similarity features through feature fusion and multi-level attention mechanisms. The graph similarity is calculated comprehensively from both the overall graph level and the node detail level, realizing the identification of similar medical visit subgraphs and effectively improving the accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is the model architecture diagram of the present invention;

[0026] Figure 2 is the processing flow diagram of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0027] The technical solution of the present invention can be implemented by software technology to realize the automatic process operation. The technical solution of the present invention will be further described in detail below with reference to the drawings and embodiments. Refer to Figure 1 、 Figure 2 , the specific steps of the embodiment of the present invention are as follows:

[0028] Step 1, according to the definition of the medical visit subgraph ontology model, extract the medical visit subgraph information from the electronic medical record knowledge graph;

[0029] Step 1.1, according to the definition of the medical visit subgraph ontology model, extract the node and relationship data in two medical visit subgraphs from the electronic medical record knowledge graph respectively;

[0030] Step 1.2, based on the node and relationship data of the two medical visit subgraphs, construct two medical visit subgraphs G 1 (V 1 , E 1 , A 1 ) = G 2 (V 2 , E 2 , A 2 ), where V 1 , V 2 represent the node information of the subgraph, including the type and text attributes of the nodes, E 1 , E 2Represents the edge information of the subgraph, including the type and text attributes of the edge, A 1 , A 2 Represents the structural information of the subgraph, represented by an adjacency matrix;

[0031] Step 2, Use the pre-trained text embedding model Sentence-BERT and unique encoding to generate node features and edge features for each visit subgraph;

[0032] Step 2.1, According to the definition of the visit subgraph ontology model, use Sentence-BERT to encode and splice the text attributes on the nodes and edges of each visit subgraph respectively to generate the node text attribute feature matrix Edge text attribute feature matrix Where: Represents the text attribute feature vector of node v i ; Represents the text attribute feature vector of edge e j ; F p Represents the number of dimensions of the text attribute features of nodes and edges. In this method, its value is determined to be 128 through sensitivity experiments; n and m are the number of nodes and edges in the visit subgraph respectively;

[0033] Step 2.2, Use a unique encoding method to generate a node type feature matrix for each visit subgraph Edge type feature matrix Where: Represents the type feature vector of node v i ; Represents the type feature vector of edge e j ; F t Represents the number of dimensions of the type features of nodes and edges, which is the sum of the number of types of nodes and edges in the visit subgraph. In this method, the value is 16;

[0034] Step 2.3, Concatenate the node text attribute feature matrix G of each visit subgraph p With the node type feature matrix G t To generate the node feature matrix of the visit subgraph Where Represents the feature vector of node v i ; F = F p + F t , Represents the number of dimensions of the features of nodes and edges; Concatenate the edge text attribute feature matrix E of each visit subgraph p With the edge type feature matrix E t To generate the edge feature matrix of the visit subgraph Where Represents edge ej Eigenvector;

[0035] Step 3: Introduce a graph attention network (EGAT) that can fuse edge features, and use the attention mechanism to learn the importance of adjacent nodes and associated edges during the feature aggregation process, and integrate the edge features into the node features; let the node and edge feature matrices pass through multiple graph attention networks (EGATs), and iteratively update the node feature matrix of the visit subgraph to the node full feature matrix.

[0036] Step 3.1: Let the node feature matrix H and the edge feature matrix E pass through a single-layer EGAT to update the node feature matrix H of the visit subgraph to the new node feature matrix H ′ The following is the specific processing flow in a single-layer EGAT:

[0037] First, N(v i ) represents the set of neighbor nodes of node v i , calculate the importance of node v i and the connected edge e j in N(v ij ), that is, the attention coefficient c i . Its calculation formula is as follows: ij

[0038]

[0039] where: c ij is the attention coefficient, indicating the importance of adjacent node v j and the connected edge e ij to node v i ; represents the eigenvector of node v i ; N(v i ) represents the set of neighbor nodes of node v i , and both node v j and node v k are nodes in it; represents the eigenvector of v j ; represents the eigenvector of v k ; represents the eigenvector of the edge between v i and v j ; represents the eigenvector of the edge between v i and v k ; || represents the vector concatenation operation; is a learnable weight matrix, Q is the number of feature dimensions of the attention hidden layer of the EGAT model, and its value is specified as 64 in this method; is the weight vector used to parameterize the attention;

[0040] Then, for c ij Use the LeakyReLU and Softmax functions to perform activation and normalization operations on it, and finally generate the attention coefficient α ij , and the calculation formula is as follows:

[0041]

[0042] Finally, for each node v i , after obtaining the attention parameter α ij , for each node v in N(i) j and the connected edge e ij , splice the feature vectors of them and use weighted summation to perform feature aggregation with the feature vector of node v i , and then update the feature vector of each node v i to the node feature vector The calculation formula is as follows: The calculation formula is as follows:

[0043]

[0044] Among them: is a learnable weight matrix, and D represents the number of output dimensions of EGAT; is a zero vector;

[0045] After passing through one layer of EGAT, all nodes in the subgraph are processed according to the above feature aggregation method, and then a new node feature matrix is generated is the new feature vector of node v i incorporating the features of neighbor nodes and connected edges;

[0046] Step 3.2, let the node feature matrix H and the edge feature matrix E pass through 3 layers of graph attention networks (EGATs) in turn to iteratively update the node feature matrix. The number of output dimensions D 1 , D 2 , D 3 are 256, 128, and 64 in turn, and finally obtain the full feature matrix of the nodes in the visit subgraph represents the full feature vector of node v i ;

[0047] Step 4, extract the global key features from the full features of each node in the visit subgraph through the attention mechanism, and generate the graph-level feature vector by weighted summation;

[0048] Step 4.1, initialize the graph-level embedding vector z using the full feature matrix U of the nodes in the visit subgraph, and the calculation formula is as follows:

[0049]

[0050] Among them: is a learnable weight matrix, where n is the number of full feature vectors of nodes in U, that is, the number of nodes in each subgraph.

[0051] Step 4.2, perform an inner product operation on each row of node full feature vectors in the subgraph node full feature matrix U with z and perform a non-linear transformation using the sigmoid function: The result of this operation is used as the attention coefficient to represent the importance of nodes in the calculation of graph similarity.

[0052] Step 4.3, perform a weighted sum of all node full feature vectors in U with the corresponding attention coefficients to generate a graph-level feature vector The calculation formula is as follows:

[0053]

[0054] Step 5, use the graph-level feature vector to calculate the graph-level similarity score information in combination with the neural tensor network; use the node full features and obtain the node-level similarity score information through histogram analysis;

[0055] Step 5.1, let the graph-level feature vectors of two subgraphs G 1 and G 2 be g 1 , g 2 respectively. Generate the graph-level similarity score information s(g 1 , g 2 ) using the neural tensor network. The calculation formula is as follows:

[0056]

[0057] Among them: is a learnable weight tensor, regarded as a matrix weight matrix with K D 3 ×D 3 . represents calculating the similarity score between g , g 1 , g 2 using each slice ; represents the vector concatenation operation; is the weight matrix; b is the bias vector, with the initial value set to zero; f is the Relu activation function;

[0058] Step 5.2, let the full feature matrices of the two subgraphs G 1 and g 2 be U 1 and U 2 respectively. Use the histogram to extract the similarity feature in U 1 and U 2 to calculate the node-level similarity score information as s(U 1 , U 2 ). The calculation process is as follows:

[0059] First, use the inner product of the full feature matrices U 1 and U 2 of the two subgraphs to generate a pairwise similarity score matrix between the node features of the two graphs In this process, "dummy nodes" with zero vector feature expressions need to be added to the graph with fewer nodes to ensure that the number of nodes in the two subgraphs is the same.

[0060] Then, use the histogram to extract the similar features in S U to obtain the node-level feature similarity score information: To avoid the influence caused by the change of node order in the subgraph. Here, B is a hyperparameter that controls the number of intervals in the histogram, and its value is determined to be 16 through sensitivity experiments in this method.

[0061] Step 6, input the above two types of score information into a multi-layer perceptron, and finally output the comprehensive similarity determination result of the two visit subgraphs; adopt a fully supervised learning strategy to train and optimize the entire model, and use the optimized model for the actual similar visit subgraph recognition task;

[0062] Step 6.1, after obtaining the graph-level similarity score information s(g 1 , g 2 ) and the node fusion feature similarity score information s(U 1 , U 2 ), splice the two as s(g 1 , g 2 )||s(U 1 , U 2 ), and input it into the multi-layer perceptron to obtain the final comprehensive similarity determination result S of the visit subgraphs. S represents the probability that the two visit subgraphs are similar. In this method, the multi-layer perceptron contains 2 hidden layers, where the feature dimensions are 32 and 16 respectively, and the sigmoid function is used as the activation function.

[0063] Step 6.2, adopt a fully supervised learning strategy, use the cross-entropy loss function L, and calculate the similarity determination result S output by the multi-layer perceptron and the two visit subgraphs G 1 and g2 The cross-entropy between similar tags among targets is used to guide the update and adjustment of the parameters of each module of the entire model. The calculation formula of the loss function is as follows:

[0064]

[0065] Where: X is the set of graph pairs in the training dataset, and (G 1 , G 2 ) represents a sub-graph pair therein, is the reciprocal of the quantity in the set of graph pairs in the training dataset; S represents the comprehensive similarity determination result of the visit sub-graphs G 1 , G 2 predicted by the method of the present invention, that is, the probability of predicting the similarity of two sub-graphs. target(G 1 , G 2 ) is a tag indicating whether the visit sub-graphs G 1 , G 2 are similar. Each similar tag is determined according to the true similarity information in the electronic medical record knowledge graph. 0 indicates that the two are not similar, and 1 indicates that the two are similar.

[0066] This method splits the dataset of similar visit sub-graphs of the electronic medical record knowledge graph independently constructed by the laboratory into a training set and a test set, and trains and tests the model respectively. When training the model, the Adam algorithm is used for model optimization, and the batch size is set to 100. When the number of iterations reaches 450, the loss value of the test set reaches the lowest and the model performance reaches the best. Then, it can be applied to the task of identifying similar sub-graphs in the electronic medical record knowledge graph. Input two visit sub-graphs into the model, and the similarity probability of the two visit sub-graphs can be output. When the probability is greater than 0.8, the two visit sub-graphs can be considered similar.

[0067] The applicant runs on a machine equipped with an Intel(R) i9-13900 CPU and an NVIDIA GeForve RTX 3090, and uses the method of this embodiment on the dataset of similar visit sub-graphs of the electronic medical record knowledge graph independently constructed by the laboratory. Comparing it with the cutting-edge baseline method, the recognition accuracy, recall rate, and F1-score have been greatly improved, proving that the method of the present invention can effectively complete the task of identifying similar sub-graphs in the electronic medical record knowledge graph.

Claims

1. A method for identifying similar consultation subgraphs in electronic medical record knowledge graphs based on graph neural networks, characterized in that: The steps include: Step 1: Extract the medical consultation subgraph information from the electronic medical record knowledge graph according to the medical consultation subgraph ontology model definition; Step 2: Use the pre-trained text embedding model Sentence-BERT and unique encoding to generate node features and edge features of each visit subgraph; Step 3: Introduce a graph attention network (EGAT) that can fuse edge features. In the feature aggregation process, use the attention mechanism to learn the importance of adjacent nodes and associated edges, and integrate edge features into node features. Pass the node and edge feature matrices through a multi-layer graph attention network (EGATs), and iteratively update the node feature matrix of the visit subgraph to the node full feature matrix. Step 4: Extract global key features from all features of each visit subgraph node through the attention mechanism, perform weighted summation and generate a graph-level feature vector; Step 5: Calculate the graph-level similarity score information by using the graph-level feature vector and the neural tensor network; obtain the node-level similarity score information by using the full node features and histogram analysis; Step 6: Input the above two types of score information into the multi-layer perceptron, and finally output the comprehensive similarity judgment result of the two consultation subgraphs; The whole model is trained and optimized using a fully supervised learning strategy, and the optimized model is used for the actual similar medical visit subgraph recognition task; Among them, the node feature matrix in step 2 is Represents node v i The feature vector of is, n is the number of nodes; the edge feature matrix is Represents edge e j The feature vector of , m is the number of edges; F is the number of feature dimensions of nodes and edges; The full feature matrix of the nodes in the consultation subgraph in step 3 is It is obtained by iteratively updating the edge feature E into the node feature H through a multi-layer graph attention network (EGATs). Represents node v i The feature vector after fusing the features of neighboring nodes and connected edges, the attention coefficient α in the single-layer EGAT ij The calculation formula is as follows: Where: α ij is the attention coefficient, indicating the connected edge e ij and adjacent nodes v j For node v i The importance of Represents node v i The characteristic vector of i ) represents node v i The neighbor node set of node v j and node v k They are all nodes; Represents node v j The eigenvector of Represents node v k The eigenvector of Indicates v i With v j The eigenvector of the edge between them; Indicates v i With v k The feature vector of the edge between them; || represents the vector concatenation operation; is a learnable weight matrix, Q is the number of feature dimensions of the EGAT model attention hidden layer, which is specified as 64; is a weight vector used to parameterize attention; Update node features in EGAT to full feature vector The feature aggregation calculation formula is as follows: in: is the learnable weight matrix, D represents the number of output dimensions of EGAT; is the zero vector; The graph-level feature vector generated by the attention mechanism in step 4 represents the features of the entire subgraph and contains key information for graph similarity calculation; In step 5, the neural tensor network can measure the similarity relationship between the graph-level feature vectors of two subgraphs from multiple levels; In step 6, a fully supervised learning strategy is adopted, and the cross entropy loss function is used as the loss function of the model. The cross entropy between the output of the multi-layer perceptron and the similarity labels between the two visit subgraphs is calculated to guide the update and adjustment of the parameters of each module of the entire model.

Citation Information

Cited By

  • Multidisciplinary fusion consultation method and system based on knowledge graph

    CN120452756A

  • Multidisciplinary fusion consultation method and system based on knowledge graph

    CN120452756B