ICD automatic coding method and system based on knowledge graph

Through the knowledge graph-based method, ICD knowledge graph is constructed and graph neural network and ConvE model is used to solve the shortcomings of the existing ICD automatic encoding method in utilizing multiple information sources, and improve the accuracy and quality of ICD encoding.

CN120072155APending Publication Date: 2025-05-30BEIJING UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510118323.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-24
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When the existing ICD automatic encoding method uses the medical course written by doctors to record information, the accuracy of ICD encoding is affected by the doctor's different writing habits and the lack of effective use of examination and laboratory results and medical order information.

Method used

Using a knowledge graph-based method, ICD knowledge graph is constructed, and multiple source information such as doctor diagnosis description, examination test results, and medical order data are mapped to the graph neural network for aggregation of node relationships, neighbor nodes and path information, and ICD encoding prediction is combined with the ConvE model.

Benefits of technology

It improves the accuracy and quality of ICD encoding, makes full use of various information sources in medical record records, and enhances the ability to explore the relationship between ICD encoding and medical record data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120072155A_ABST
    Figure CN120072155A_ABST
Patent Text Reader

Abstract

The invention discloses an ICD automatic coding method and system based on a knowledge graph. The method comprises the following steps: S1, constructing an ICD knowledge graph based on case data to obtain a triple set; the case data comprises doctor handwritten diagnosis, patient examination results, doctor's advice data and ICD codes of a medical record home page; s2, mapping the triple set into a graph neural network, and respectively carrying out node relation information aggregation, neighbor node and path information aggregation and triple semantic information aggregation; s3, performing hierarchical information aggregation based on the relationship of each entity and the structure embedding information of the neighbor nodes and the triple; and S4, based on the final entity embedding representation and the relationship embedding representation, ICD coding prediction is carried out through a ConvE model, and a final ICD coding list is obtained. According to the method, various source information such as doctor diagnosis description, examination and test results and the like in the medical record is fully explored and utilized, the incidence relation between each ICD code and the information is explored, and the accuracy of the final ICD code is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field, and in particular to an ICD automatic encoding method and system based on a knowledge graph. Background Art

[0002] With the in-depth development of medical information technology, the accuracy and standardization of the medical record homepage, as a concentrated expression of the patient's medical activities during hospitalization, directly affect the quality of medical and disease coding and the accurate use of medical data. As an internationally accepted medical industry standard, the International Classification of Diseases (ICD) coding is not only the key to achieving the standardization of medical record homepage information, but also an important structured data element in the hospital information system. The national hospital high-quality development strategy needs to improve the hospital's service capabilities, and has put forward higher requirements for the efficiency and quality of hospital ICD coding work. However, the huge contrast is the shortage of coding personnel and the uneven level of coding personnel. This demand gap has pointed out the direction for the research of ICD automatic coding methods.

[0003] At present, the research on ICD automatic coding mainly focuses on rule-based, case-based and deep learning-based methods: (1) In terms of rule-based methods, the classifier is based on a set of mapping relationships from medical terms to ICD codes. On the basis of eliminating negative words, the diagnosis names that doctors usually write are summarized, and the items with higher frequency are accurately coded by coding experts and then maintained as rules. (2) In terms of case-based learning, the research focuses on analyzing historical data in order to find the relationship between medical document data and ICD coding schemes. Various machine learning algorithms are applied in this process, such as decision trees, support vector machines, naive Bayes, k-nearest neighbors, etc. (3) In the field of deep learning, the research focuses on applying deep learning frameworks such as neural networks, LSTM, Transformer, Bert, etc., starting from the text in medical documents, further capturing the associated information therein and improving the quality of ICD coding.

[0004] A series of existing solutions provide theoretical basis and methodological support for the present invention, but the current main research direction focuses on the medical record information recorded in medical documents. However, there are differences in doctors' habits of writing medical records, as well as lack of utilization of information such as test results and medical orders, which affects the accuracy of the final ICD coding. Summary of the invention

[0005] In view of the deficiencies in the prior art, the present invention fully explores and utilizes multiple source information such as doctor's diagnosis descriptions, examination and test results in medical records, explores the correlation between each ICD code and the above information, and improves the accuracy of the final ICD code.

[0006] The present invention provides an ICD automatic coding method based on a knowledge graph, and the method includes:

[0007] S1. Based on case data, construct an ICD knowledge graph to obtain a triple set composed of patient numbers, relationship numbers, and entity numbers; the case data includes: doctor's handwritten diagnosis, patient examination results, order data, and ICD codes on the front page of the medical record;

[0008] S2. Map the triple set into a graph neural network, and perform node relationship information aggregation, neighbor node and path information aggregation, and triple semantic information aggregation respectively to obtain the relationship, neighbor nodes, and structural embedding information of triples for each entity in the graph neural network;

[0009] S3. Based on the relationship, neighbor nodes, and structural embedding information of triples for each entity, perform hierarchical information aggregation to obtain the final entity embedding representation and relationship embedding representation;

[0010] S4. Based on the final entity embedding representation and relationship embedding representation, perform ICD coding prediction through the ConvE model to obtain the final ICD coding list.

[0011] Preferably, step S1 includes:

[0012] S11. Extract and encode the information of entities in the case data to obtain a list of entity numbers; when encoding the extracted entities, use the same sequence number starting from 1 and there are no empty items in the number sequence;

[0013] S12. Extract and encode the information of relationships in the case data to obtain a list of relationship numbers;

[0014] S13. Based on the list of entity numbers and the list of relationship numbers, encode the medical record data to be classified, and assemble to obtain a triple set composed of patient numbers, relationship numbers, and entity numbers.

[0015] Preferably, step S11 includes:

[0016] Segment the doctor's handwritten diagnosis according to the writing habit, count the quantity in groups, sort from high to low, manually check the items with too low occurrence frequency, eliminate the items with non-standard writing and errors, and sequentially number the remaining items as entity names;

[0017] For the numerical type results in the patient's examination results, if they are determined to be too high or too low according to the reference values of the examination index items, they are sequentially numbered with the entity name being "index name + high" or "index name + low". For the text type results in the patient's examination results, they are sequentially numbered with the entity name being "index name + result".

[0018] Classify and summarize all the doctor's order items for a patient's one hospitalization in the doctor's order data by item name, and sequentially number the names of the doctor's order items as entity names.

[0019] Sequentially number the codes of the ICD codes on the front page of the medical record as entity names, and create a dictionary for restoring the numbers in the final output result back to ICD codes.

[0020] Step S12 includes:

[0021] According to the organizational characteristics of the medical record data, establish relationships and sequentially number the doctor's handwritten diagnosis, the patient's examination results, and the doctor's order data to obtain a relationship number list; the relationship names in the relationship number list include: gender, age, inpatient department, the main diagnosis and other diagnoses in the doctor's handwritten diagnosis, the names of the examination index items in the patient's examination results, the doctor's order categories in the doctor's order data, and the main diagnosis and other diagnoses in the ICD codes on the front page of the medical record.

[0022] Preferably, in step S2, mapping the triple set to the graph neural network includes:

[0023] Take the head entity, i.e., the patient number, and the tail entity, i.e., the entity number, of each triple in the triple set as nodes in the graph neural network, and take the relationship number as the feature of the edge connecting the head entity and the tail entity in the graph neural network.

[0024] In step S2, performing node relationship information aggregation includes:

[0025] Perform a weighted sum of the relationships between each entity in the graph neural network and all its neighbor entities, and perform a linear transformation on the result of the weighted sum to obtain the relationship embedding representation of each entity; the calculation formula is:

[0026]

[0027] Where S i rel is the relationship embedding representation of entity e i , e i is the embedding representation of the entity corresponding to the i-th node in the graph neural network, N i is the set of all neighbor entities of entity e i and rj For entity e i The relationship with the j-th neighbor entity, α i,j rel For relationship r j The importance attention weight for entity e i , W rel For the weight matrix of the relationship embedding, used to map relationship r j To a new representation space, σ is a non-linear activation function, exp(r j T e i ) is the matching degree of relationship r j With entity e i . i and j are positive integers.

[0028] Preferably, in step S2, aggregating neighbor node and path information includes:

[0029] Performing single-layer aggregation on the path information of each entity in the graph neural network and its directly connected neighbor nodes to obtain the single-layer aggregation result of each entity; the calculation formula is:

[0030]

[0031] Among them, S i ent Represents the aggregated representation of entity e i . After one layer of aggregation, this representation has incorporated the path information of all neighbor nodes directly connected to entity e i . e i Is the embedding representation of the entity corresponding to the i-th node in the graph neural network, e j Is the embedding representation of the j-th neighbor entity of entity e i , N i Is the set of all neighbor entities of entity e i , α i,j ent Is the importance attention weight of neighbor entity e j For entity e i , W ent Is a linear transformation matrix, used to map the embedding representation of neighbor entity e j To a new feature space, σ is a non-linear activation function, exp(e j T e i ) is the matching degree of neighbor entity e j With the central entity e i . i, j, and k are positive integers;

[0032] Based on the single-layer aggregation results of each entity, iterative multi-layer aggregation is performed until the corresponding number of iterations is reached, and the neighbor node embedding representations of each entity are obtained.

[0033] Preferably, in step S2, the triple semantic information aggregation includes:

[0034] Mapping each neighbor entity of each entity in the graph neural network and the corresponding relationship into a fusion vector, and performing triple semantic information aggregation on all fusion vectors of each entity to obtain the triple embedding representation of each entity; the calculation formula is:

[0035]

[0036] Where S i tri Is the triple embedding representation of entity e i , e i Is the embedding representation of the entity corresponding to the i-th node in the graph neural network, e j Is the embedding representation of the j-th neighbor entity of entity e i , N i Is the set of all neighbor entities of entity e i , r j Is the relationship between entity e i And neighbor entity e j , Is the fusion representation of each neighbor entity-relationship pair (e i , r j , r j ) of entity e i,j tri Is the importance attention weight of the fusion representation To entity e i , W tri Is the weight matrix, used to map the fusion representation To the target space, σ is the activation function, Is the similarity between the fusion representation And the central entity e i , i and j are positive integers;

[0037] The fusion representation Is a composite function, used to perform a composite operation on entity e i And relationship r j , The composite operation includes: addition operation, dot multiplication operation and multi-layer perceptron MLP splicing operation.

[0038] Preferably, step S3 includes:

[0039] Iteratively aggregate the structural embedding information of each entity, its relationships, neighbor nodes, and triples in the graph neural network, initialize different relationship embedding vectors for each layer until the corresponding number of iterations is reached, concatenate the relationship embedding vectors of all layers, and perform a linear transformation on the concatenated result to obtain the final entity embedding representation and relationship embedding representation; the calculation formula for iterative multi-layer aggregation is:

[0040]

[0041] where, e i ' l+1 is the entity embedding vector of entity e i after aggregation in the (l + 1)-th layer, e i l is the entity embedding representation of entity e i after aggregation in the l-th layer, is the relationship embedding representation of entity e i after aggregation in the l-th layer, is the triple embedding representation of entity e i after aggregation in the l-th layer, and i, l are positive integers;

[0042] The calculation formulas for the final entity embedding representation and relationship embedding representation are:

[0043] e out = e k

[0044] r out = W out Concat({r l | l = 1,..., K})

[0045] where, e out is the final entity embedding representation, e k is the entity embedding representation after aggregation in the last K-th layer, r out is the final relationship embedding representation, r l is the relationship embedding vector of the l-th layer, and W out is the transformation matrix, and l, K are positive integers.

[0046] Preferably, step S4 includes:

[0047] S41. Input the final entity embedding representation and relationship embedding representation into the ConvE model;

[0048] S42. Generate an embedding vector for each entity and relationship through the ConvE model, reshape the embedding vectors of each entity and relationship into matrices, concatenate the matrix obtained by reshaping the embedding vector of the head entity with the matrix obtained by reshaping the embedding vector of the relationship, perform a convolution operation on the concatenated matrix, perform non-linear activation, feature flattening, and full connection layer mapping on the feature map generated after convolution to obtain a feature vector with a fixed dimension, and perform a dot product operation on this feature vector and the embedding vectors of all entities to obtain the score of each entity as a candidate tail entity;

[0049] S43. Sort all entities according to the scores output by the ConvE model, take the entity with the highest score as the predicted result of the tail entity, predict the possible ICD codes corresponding to each head entity and relationship, and obtain the final ICD code list.

[0050] Based on the same inventive concept, the present invention also provides a knowledge graph-based ICD automatic coding system, and the system includes:

[0051] A case information extraction module, configured to construct an ICD knowledge graph based on case data to obtain a set of triples composed of patient numbers, relationship numbers, and entity numbers; the case data includes: doctor's handwritten diagnosis, patient examination results, medical order data, and ICD codes on the medical record front page;

[0052] A node information aggregation module, configured to map the set of triples into a graph neural network, perform node relationship information aggregation, neighbor node and path information aggregation, and triple semantic information aggregation respectively, to obtain the relationship, neighbor node, and structural embedding information of triples of each entity in the graph neural network;

[0053] A hierarchical information aggregation module, configured to perform hierarchical information aggregation based on the relationship, neighbor node, and structural embedding information of each entity to obtain the final entity embedding representation and relationship embedding representation;

[0054] A prediction module, configured to perform ICD code prediction through the ConvE model based on the final entity embedding representation and relationship embedding representation to obtain the final ICD code list.

[0055] Preferably, the case information extraction module includes:

[0056] An entity coding module, configured to extract and code the information of entities in the case data to obtain a list of entity numbers; when coding the extracted entities, the same sequence number starting from 1 is used and there are no empty items in the number sequence;

[0057] A relationship coding module, configured to extract and code the information of relationships in the case data to obtain a list of relationship numbers;

[0058] An assembly module for encoding medical record data to be classified based on the list of entity numbers and the list of relationship numbers, and assembling a set of triples composed of patient numbers, relationship numbers, and entity numbers.

[0059] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0060] Based on case data, the present invention constructs an ICD knowledge graph and obtains a set of triples composed of patient numbers, relationship numbers, and entity numbers. The case data includes: doctor's handwritten diagnosis, patient examination results, doctor's orders data, and ICD codes on the front page of the medical record. The present invention fully explores and utilizes various sources of information such as doctor's diagnosis descriptions, examination and test results, and surgery, treatment, medication, and consumables in the medical record, discovers the correlation between each ICD code and the above information, and improves the quality of ICD coding.

[0061] The present invention maps the set of triples into a graph neural network, and performs node relationship information aggregation, neighbor node and path information aggregation, and triple semantic information aggregation respectively to obtain the relationship, neighbor node, and structural embedding information of triples of each entity in the graph neural network. Through node relationship information aggregation, not only can the direct relationship between entities be captured, but also the more important neighbor relationships of entities can be highlighted in a weighted manner to capture more complex relationship structures. Through neighbor node and path information aggregation, by aggregating the representations of neighbor entities at one time, the model can capture all paths of length 1, that is, directly connected neighbor relationships. The present invention can capture the information of longer paths, that is, the possible connection path situations of length 2 or longer, through an iterative multi-layer aggregation method. Each layer of aggregation operation not only integrates the information of the entity's direct neighbors, but also further integrates the information of the neighbors of the neighbors. This multi-layer aggregation mechanism enables the model to gradually expand the perception range outward layer by layer, so as to obtain deeper connection relationships at different levels. Through triple semantic information aggregation, the semantic information in the medical record data is extracted, and the information of both its entity and relationship parts is integrated. The neighbor nodes and relationships are jointly incorporated into the feature representation to capture the similarity features of the neighbor structure.

[0062] Based on the relationship, neighbor node, and structural embedding information of each entity, the present invention performs hierarchical information aggregation to obtain the final entity embedding representation and relationship embedding representation. Through hierarchical information aggregation, deeper structural information can be gradually integrated to effectively capture potential complex relationships in the multi-hop neighborhood.

[0063] Based on the final entity embedding representation and relationship embedding representation, the present invention predicts possible ICD coding items for patients through the ConvE model to obtain the final ICD coding list. The ConvE model combines traditional knowledge graph embedding methods with convolutional neural networks to improve the model's expressive ability and reasoning ability. The core idea of ConvE is to reshape entity and relationship embeddings into matrix form, extract complex interaction features between entities and relationships through convolutional neural networks, and map the convolved features back to the entity space through fully connected layers for the prediction of tail entities. Compared with previous linear models, ConvE introduces non-linear operations (convolution and activation functions), which can better capture complex patterns between entities and relationships. This convolution operation allows the model to perform a more in-depth analysis of local interactions, thereby enhancing the model's expressive ability. Through convolutional neural networks and non-linear operations, ConvE effectively enhances the interaction characteristics of entity and relationship embeddings in the knowledge graph and improves the performance of the knowledge graph completion task. Description of the Drawings

[0064] Figure 1 It is a schematic flowchart of an ICD automatic coding method based on a knowledge graph provided by the present invention;

[0065] Figure 2 It is an overall architecture diagram of an ICD automatic coding method based on a knowledge graph provided by the present invention;

[0066] Figure 3 It is a schematic sub-flowchart of an ICD automatic coding method based on a knowledge graph provided by the present invention;

[0067] Figure 4 It is a schematic sub-flowchart of an ICD automatic coding method based on a knowledge graph provided by the present invention;

[0068] Figure 5 It is a schematic structural diagram of an ICD automatic coding system based on a knowledge graph provided by the present invention;

[0069] Figure 6 It is a schematic structural diagram of another ICD automatic coding system based on a knowledge graph provided by the present invention. Detailed Embodiments

[0070] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0071] The present invention will be further described in detail below with reference to the accompanying drawings.

[0072] As Figure 1-2 shown, an embodiment of the present invention provides an ICD automatic coding method based on a knowledge graph. The method includes:

[0073] S1. Based on case data, construct an ICD knowledge graph to obtain a triple set composed of patient numbers, relationship numbers, and entity numbers; the case data includes: doctor's handwritten diagnosis, patient examination results, order data, and ICD codes on the medical record homepage;

[0074] S2. Map the triple set into a graph neural network, and perform node relationship information aggregation, neighbor node and path information aggregation, and triple semantic information aggregation respectively to obtain the relationship, neighbor nodes, and structural embedding information of triples of each entity in the graph neural network;

[0075] S3. Based on the relationship, neighbor nodes, and structural embedding information of triples of each entity, perform hierarchical information aggregation to obtain the final entity embedding representation and relationship embedding representation;

[0076] S4. Based on the final entity embedding representation and relationship embedding representation, perform ICD coding prediction through the ConvE model to obtain the final ICD coding list.

[0077] Figure 2 Describes the overall architecture diagram of an ICD automatic coding method based on a knowledge graph in an embodiment of the present invention. The embodiment of the present invention mainly realizes the ICD automatic coding method through the following four modules: Module 1 Case Information Extraction Module, Module 2 Node Information Aggregation Module, Module 3 Hierarchical Information Aggregation Module, and Module 4 Prediction Module.

[0078] Module 1 Medical Record Information Extraction Module

[0079] As Figure 3 shown, step S1 includes:

[0080] (1) Entity Encoding

[0081] S11. Extract and encode the information of entities in the case data to obtain a list of entity numbers; when encoding the extracted entities, the same sequence number starting from 1 is used and there are no empty items in the number sequence.

[0082] The medical record data mainly involves four aspects of data, doctor's handwritten diagnosis, patient examination results, doctor's orders, and ICD codes. The specific processing methods for these four aspects of data are as follows:

[0083] Step S11 includes: splitting the doctor's handwritten diagnosis according to writing habits, grouping and counting the quantity, sorting from high to low, manually checking the items with too low occurrence frequency, removing the items with non-standard writing or errors, and sequentially numbering the remaining items as entity names; for the numerical type results in the patient's examination results, determining whether they are too high or too low according to the reference values of the examination index items, and sequentially numbering them with "index name + high" or "index name + low" as entity names, for the text type results in the patient's examination results, sequentially numbering them with "index name + result" as entity names; classifying and summarizing all the doctor's order items during a patient's hospitalization in the doctor's order data by item name, and sequentially numbering the names of the doctor's order items as entity names; sequentially numbering the codes of the ICD codes on the medical record front page as entity names and making a dictionary for restoring the numbers in the final output result back to ICD codes.

[0084] Medical record data is complex and all-inclusive and cannot be directly used to generate ICD codes. In the present invention, the medical record data needs to be preprocessed before actual automatic coding. This method is used to construct an ICD knowledge graph and preprocess each medical record data before sending it into the system.

[0085] Doctor's handwritten diagnosis: Split the doctor's handwritten diagnoses of all patients in the dataset according to writing habits such as commas, semicolons, line breaks, etc., group and count the quantity, and sort from high to low. Manually check the items with too low occurrence frequency, remove the items with non-standard writing or errors, and then sequentially number them.

[0086] Examination results: The examination results are relatively complex, including various types such as numerical values and texts. For numerical type data, determine whether it is too high or too low according to the reference value of the index item, name the entity with "index name + high" or "index name + low", for text type results, use "index name + result" as the entity name, and then sequentially number them.

[0087] Doctor's order data: Doctor's orders are issued daily, and the same doctor's order may appear repeatedly in the medical records of many days. For example, some patients with respiratory infections use antibiotics for several consecutive days. Therefore, in this study, all the doctor's order items during a patient's hospitalization are classified and summarized by item name, the doctor's order name is numbered as the entity name, and the types of doctor's order items, such as drugs, surgeries, consumables, etc., are sequentially numbered as the types of relationships.

[0088] ICD dictionary coding: The ICD dictionary coding contains codes and their corresponding dictionary item names, and its codes are unique. Therefore, the codes are sequentially numbered to make a dictionary. This dictionary will also be used after the model's final output result to restore the numbers in the result back to ICD codes.

[0089] It should be noted that the above entities all use the same sequence number starting from 1, and there are no allowed empty items in the number sequence, that is, there is no phenomenon of "skipping numbers".

[0090] (2) Relationship coding

[0091] S12. Extract and code the information of the relationships in the case data to obtain a list of relationship numbers;

[0092] Step S12 includes:

[0093] According to the organizational characteristics of the medical record data, establish relationships and sequentially number the doctor's handwritten diagnosis, patient examination results, and doctor's order data to obtain a list of relationship numbers; the relationship names in the list of relationship numbers include: gender, age, inpatient department, main diagnosis and other diagnoses in the doctor's handwritten diagnosis, name of the examination index item in the patient examination results, doctor's order category in the doctor's order data, and main diagnosis and other diagnoses in the ICD coding of the medical record home page. According to the organizational characteristics of the medical record data, establish relationships among the diagnosis, examination and test, and doctor's order content, as shown in Table 1 below:

[0094]

[0095]

[0096] Table 1

[0097] (3) Assemble triples

[0098] S13. Based on the list of entity numbers and the list of relationship numbers, code the medical record data to be classified, and assemble to obtain a set of triples composed of patient numbers, relationship numbers, and entity numbers.

[0099] After the previous step, a list of entities and relationship numbers is obtained. Using this list, code the medical record data to be classified to obtain a set of triples of (patient number, relationship number, entity number). For example, the triple (10001, gender, female) can be coded as (1, 1, 2), where 10001 is the patient number, the relationship number of the gender relationship is 1, and the number of the entity "female" is 2.

[0100] (4) Load the neural network

[0101] When constructing a graph neural network, the numbers of the head and tail entities of the triples will become the nodes in the network, and the number of the relationship will be used as the feature of the edge in the network.

[0102] In the embodiment of the present invention, in step S2, mapping the set of triples to the graph neural network includes:

[0103] Take the head entity of each triple in the triple set, i.e., the patient number, and the tail entity, i.e., the entity number, as nodes in the graph neural network, and take the relationship number as the feature of the edge connecting the head entity and the tail entity in the graph neural network.

[0104] Module 2 Node Information Aggregation Module

[0105] Such as Figure 2 As shown, in the automatic ICD coding task, relying on a single specific triple is difficult to provide sufficient information to define and distinguish different ICD coding dictionary items. It is necessary to extract the content that affects the coding from complex medical record data and evaluate the influence degree of various contents (such as doctor's description, examination and test results, medical orders, etc.) on the coding result. Based on the triple set obtained in Module 1, Module 2 maps the triples into the graph neural network constructed by the knowledge graph, extracts the information of relationships, nodes, and the triples themselves from the triples respectively, calculates the eigenvalue of the triples, and completes the construction of the medical record data structure dataset. The main functions are as follows:

[0106] (1) Relationship aggregation: Different diseases are recorded in different ways in medical documents, and the supported examination and test results and treatment methods are different. Based on this characteristic, the information of the data category in the medical record is used as the relationship information between entities. Through weighted summation and then through linear transformation, the relationship-level representation is obtained, and the relationship-level representation is constructed. The present invention captures these differences through the attention mechanism and autonomously learns the weights of different relationships. The specific method is: in the graph neural network, assume that the entity e i is a node in the network. The relationships between the entity e i and all its neighbor entities are weighted and summed, and then through a linear transformation, the final relationship-level representation is obtained. The aggregation process can not only capture the direct relationships between entities, but also highlight the more important neighbor relationships for the entity e i through weighted means.

[0107] In the embodiment of the present invention, in step S2, performing node relationship information aggregation includes:

[0108] Weightedly sum the relationships between each entity in the graph neural network and all its neighbor entities, and perform a linear transformation on the result of the weighted sum to obtain the relationship embedding representation of each entity; the calculation formula is:

[0109]

[0110] Among them, S i rel is the relationship embedding representation of the entity e i , e i is the embedding representation of the entity corresponding to the i-th node in the graph neural network, and N i is the entity ei The set of all neighbor entities of r j For entity e i The relationship with the j-th neighbor entity, α i,j rel Is the relationship r j For entity e i The importance attention weight of, W rel Is the weight matrix of the relationship embedding, used to map the relationship r j To a new representation space, σ is a non-linear activation function, exp(r j T e i ) Is the matching degree of the relationship r j With entity e i , where i and j are positive integers.

[0111] Specifically, in formula (1), N i Represents the neighbor set of entity e i , that is, N i Contains all entities that have a certain relationship with entity e i ; α i,j rel Is the importance attention parameter of entity e j Based on the relationship r i , which represents the importance of relationship r j For entity e i Among all neighbor relationships. This attention mechanism enables the model to weight according to the importance of different relationships; W rel Is the weight matrix of the relationship embedding, used to map the relationship r j To a new representation space. This linear transformation allows the model to re-encode neighbor information in different relationship contexts, thus capturing more complex relationship structures; σ is a non-linear activation function to increase the expressive power of the model.

[0112] Specifically, in formula (2), e i Is the embedding representation of the entity. The model uses the dot product operation to dynamically calculate the attention importance of the neighbor relationship r j Relative to the central entity e i . The core of the attention mechanism is that it measures the relative importance of each relationship by calculating the similarity (i.e., dot product) between the central entity and its neighbor relationships. Here, the dot product represents the relationship r j And entity e iThe matching degree. Subsequently, the model uses the SoftMax function for normalization to ensure that the sum of the attention weights of all neighbor relationships is 1, thereby obtaining a probability distribution. This normalized weight distribution allows the model to more effectively focus on the current entity e i The most important relationship.

[0113] (2) Neighbor node and path aggregation: In the application of graph neural networks, the feature representation of entity e i not only depends on the relationships connected to it, but also on the entities connected to it. Considering the actual situation, patients with the same disease have many commonalities in their symptoms, signs, examination and test results, as well as treatment and medication. Therefore, by aggregating the information of different neighbor entities, the target entity can integrate a lot of common knowledge. For example, in the test reports of patients with diabetes, the fasting blood glucose index is significantly higher than that of normal people, and drugs such as insulin and metformin are commonly used in treatment, which can all be used as influencing factors for predicting ICD codes. In the aggregated representation at the entity level, through the path connection information aggregation mechanism between entities, the model can capture the patterns of neighbor entities, thereby obtaining richer relationship information about the connection paths between entities.

[0114] Specifically, by aggregating the representations of neighbor entities at one time, the model can capture all paths of length 1, that is, directly connected neighbor relationships. This method provides a basis for constructing the initial representation of entities. Considering the individual differences among patients in reality, even for the same disease, the doctor's description of the condition, the diagnosis and treatment process, and the treatment plan are different. In graph neural networks, there will be a phenomenon that different patients reach the node of the same ICD coding entity through different entity nodes and paths. For this situation of connection paths of length 2 or longer that may occur in the graph, the model captures the information of longer paths through iterative multi-layer aggregation. Each layer of aggregation operation not only integrates the information of the entity's direct neighbors, but also further integrates the information of the neighbors' neighbors. This multi-layer aggregation mechanism enables the model to gradually expand the perception range outward layer by layer, thereby obtaining deeper connection relationships at different levels.

[0115] In the embodiment of the present invention, in step S2, performing neighbor node and path information aggregation includes:

[0116] Performing single-layer aggregation on the path information of each entity in the graph neural network and its directly connected neighbor nodes to obtain the single-layer aggregation result of each entity; based on the single-layer aggregation result of each entity, performing iterative multi-layer aggregation until the corresponding number of iterations is reached to obtain the neighbor node embedding representation of each entity. Here, a single-layer aggregation formula is shown, and the calculation formula is:

[0117]

[0118] Among them, S i ent represents the aggregated representation of entity e i After one layer of aggregation, this representation has incorporated the path information of entity e i and all its directly connected neighbor nodes. e i is the embedding representation of the entity corresponding to the i-th node in the graph neural network, and e j is the embedding representation of the j-th neighbor entity of entity e i , N i is the set of all neighbor entities of entity e i , α i,j ent is the importance attention weight of neighbor entity e j to entity e i . W ent is a linear transformation matrix used to map the embedding representation of neighbor entity e j to a new feature space. σ is a non-linear activation function, and exp(e j T e i ) represents the matching degree between neighbor entity e j and central entity e i . i and j are positive integers.

[0119] Specifically, in formula (3), S i ent represents the aggregated representation of entity e i . After one layer of aggregation, this representation has incorporated the path information of the entity and its neighbor nodes; N i is the neighbor set of entity e i , containing all entities directly related to e i ; α i,j ent is the attention weight, representing the importance of neighbor entity e j to central entity e i . Through this attention mechanism, the model can assign different weights to different neighbors, and thus perform weighted summation according to their importance to the central entity; W ent is a linear transformation matrix used to map the embedding representation of neighbor entity e j to a new feature space. This transformation enables the model to learn more abstract features and enhances the model's expressive ability; σ is a non-linear activation function to increase the non-linearity of the model, so as to better capture complex relationship patterns. The calculation method of the attention weight is similar to the previous one. α i,j ent in formula (4) is a value used to represent the attention weight, used to measure entity ej For entity e i 's importance. Specifically, α i,j ent is calculated through the attention mechanism and defines the attention weight of entity e j for entity e i This attention weight can reflect the contribution degree of the features of entity e j in the given context to entity e i .

[0120] (3) Triple semantic information aggregation: Extract the semantic information in the medical record data, comprehensively combine the information of the corresponding entities and relationships, and use the aggregation function to perform weighted summation on the weight information obtained by the attention mechanism for each neighbor node and relationship, so as to integrate the neighbor nodes and relationships into the feature representation together and capture the similarity features of the neighbor structure. In this solution, an aggregation function is designed to aggregate the neighbor entity and relationship information to obtain the final triple representation.

[0121] In the embodiment of the present invention, in step S2, performing triple semantic information aggregation includes:

[0122] Map each neighbor entity and the corresponding relationship of each entity in the graph neural network to a fusion vector, perform triple semantic information aggregation on all fusion vectors of each entity, and obtain the triple embedding representation of each entity; The calculation formula is:

[0123]

[0124] where S i tri is the triple embedding representation of entity e i , e i is the embedding representation of the entity corresponding to the i-th node in the graph neural network, e j is the embedding representation of the j-th neighbor entity of entity e i , N i is the set of all neighbor entities of entity e i , r j is the relationship between entity e i and neighbor entity e j , is the fusion representation of each neighbor entity-relationship pair (e i , r j , r j ) of entity e i,j tri is the importance attention weight of the fusion representation for entity e i , W triis a weight matrix used to map the fused representation to the target space, and σ is an activation function. is the fused representation and the similarity with the central entity e i , where i, j, and k are positive integers.

[0125] Specifically, the numerator part exp(φ(e j , r j )e i ) in formula (6) represents the similarity between the combined representation of the neighbor entity e j and the relation r j and the central entity e i . This similarity is measured through an inner product operation and converted to a non - negative value through an exponential function. The denominator part normalizes the similarities of all neighbor entity - relation pairs to ensure that the sum of the attention weights is 1.

[0126] The fused representation is a composite function used to perform a composite operation on the entity e i and the relation r j . The composite operation includes: an addition operation, a dot - product operation, and a multi - layer perceptron (MLP) concatenation operation. The fused representation φ(e i , r j ) is used to combine the entity information e and the relation information r. There are multiple ways to select this composite operation:

[0127] ① Addition function: That is, by simply adding the information of the entity and the relation through the + addition operation, a fused feature representation is obtained. This method is simple and efficient and is suitable for use when the dimensions of the entity and relation vectors are the same.

[0128] ② Multiplication function: That is, by performing a * dot - product operation to combine the entity and relation information, a new vector representation is generated. The dot - product operation can capture the mutual dependence and interaction characteristics between the entity and the relation.

[0129] ③ Multi - layer perceptron (MLP): That is, where [e||r] represents concatenating the entity and relation vectors. The MLP extracts the high - order interaction information between the entity and the relation more flexibly through multiple non - linear transformations of the neural network.

[0130] Module Three - level Information Aggregation Module

[0131] In the embodiments of the present invention, after obtaining the relationship, neighbor nodes, and triple embedding information of the entity in the graph neural network from Module Two, Model Three obtains the final representation of the entity by addition.

[0132] e i ' = e i + s i rel + s i ent + s i tri (7)

[0133] In formula (7), e i ' represents the entity embedding vector after aggregation, which not only contains the embedding information e i of the original entity, but also incorporates the structural embedding information at the relation level, entity level, and triple level. This aggregation method can be regarded as a single-layer aggregation operation in the graph neural network, that is, only the information of the 1-hop neighborhood is captured, which means only the influence of the nodes directly adjacent to the entity is considered.

[0134] To obtain the information of multi-hop neighbors and model the deeper interactions between the structural embedding components, the study introduced a multi-layer version of the structural embedding aggregation. This multi-layer structure enables the model to perform aggregation at a deeper level, thereby obtaining richer neighbor information and enhancing the model's ability to model complex coding relationships. In the multi-layer aggregation version, the output e i ' of the previous layer is used as the input of the next layer, and aggregation is continuously performed in an iterative manner.

[0135] In the embodiment of the present invention, step S3 includes:

[0136] Performing iterative multi-layer aggregation on the structural embedding information of each entity, its relation, neighbor nodes, and triples in the graph neural network, initializing different relation embedding vectors for each layer until the corresponding number of iterations is reached, concatenating the relation embedding vectors of all layers, and performing a linear transformation on the concatenated result to obtain the final entity embedding representation and relation embedding representation; the calculation formula for iterative multi-layer aggregation is:

[0137]

[0138] where e i ' l+1 is the entity embedding vector of entity e i after aggregation at the (l + 1)-th layer, e i l is the entity embedding representation of entity e i after aggregation at the l-th layer, is the relation embedding representation of entity e i after aggregation at the l-th layer, is the triple embedding representation of entity e i after aggregation at the l-th layer, and i and l are positive integers.

[0139] In formula (8), e i ' l+1 represents the entity embedding of the (l + 1)-th layer, which is updated based on the entity embedding e i of the l-th layer and the structural embedding information of the l-th layer. Through this layer-by-layer iterative method, the model can gradually integrate deeper structural information, thereby effectively capturing potential complex relationships in the multi-hop neighborhood. This method not only improves the expressive power of the embedding but also helps the model better generalize to unseen data.

[0140] Considering the different degrees of influence of entities and relationships at different levels on the final ICD coding, in order to adapt to the specific requirements of each layer, we initialize different relation embeddings r l for each layer. This is because in different layers, relationships may play different roles and participate in different interaction patterns. Therefore, by defining separate embedding vectors r l for relationships in each layer, the differences between layers can be captured more flexibly, enabling the model to better learn the hierarchical features of relationships. In the processing of the final output relation embedding, the relation embeddings r l of all layers are concatenated. Through the concatenation operation, the specific information of each layer can be retained, enabling the final relation embedding to contain the information of each layer, thereby more comprehensively describing the characteristics of the relationship.

[0141] To transform the concatenated embedding vector into the final output format, a transformation matrix W out is introduced to perform a linear transformation on the concatenated result.

[0142] The calculation formulas for the final entity embedding representation and relation embedding representation are as follows:

[0143] e out =e k (9)

[0144] r out =W out Concat({r l |l=1,...,K})(10)

[0145] where, e out is the final entity embedding representation, e k is the entity embedding representation after aggregation through the last layer of the K-th layer, r out is the final relation embedding representation, r l is the relation embedding vector of the l-th layer, W out is the transformation matrix, and l and K are positive integers.

[0146] In formulas (9) and (10), e outDenote the final output embedding of the entity, which is the entity embedding e of the last layer K k to represent, r out is the final relation embedding, by concatenating the relation embeddings r of each layer l and then applying the transformation matrix W out to obtain.

[0147] This design method makes full use of the characteristics of multi-layer relation embeddings, enabling the output relation embeddings to synthesize the structural information of each layer, so that the model can more flexibly handle the reasoning tasks of ICD coding for different hierarchical relations. Under the framework of graph neural networks, this aggregation method of multi-level relation embeddings can effectively improve the generalization ability and expressive ability of the model.

[0148] Module Four: Prediction Module

[0149] Based on the content and weight pair relationships of each node (entity) and edge (relation) in the graph neural network constructed by Modules Two and Three, Module Four uses ConvE as the decoder to perform convolutional processing on the information obtained from the previous modules to achieve the prediction of the final ICD coding list. ConvE (Convolutional Knowledge Graph Embedding) is a model for knowledge graph completion. It learns the embedding representations of entities and relations in the knowledge graph through a convolutional neural network CNN, so as to predict missing triples, combining traditional knowledge graph embedding methods with convolutional neural networks to improve the expressive ability and reasoning ability of the model. The core idea of ConvE is to reshape the entity and relation embeddings into matrix forms, extract the complex interaction features between entities and relations through a convolutional neural network, and map the convolutional features back to the entity space through a fully connected layer for the prediction of the tail entity. Compared with previous linear models, ConvE introduces non-linear operations (convolution and activation functions), which can better capture the complex patterns between entities and relations. This convolutional operation allows the model to perform a more in-depth analysis of local interactions, thus enhancing the expressive ability of the model.

[0150] The first step of the ConvE model is to generate an embedding vector for each entity and relationship in the knowledge graph. Suppose there are N entities and M relationships in the knowledge graph, and each entity and relationship is mapped to a fixed low-dimensional vector space through an embedding matrix. Specifically, the embedding of the head entity h is represented as a vector h ∈ Rd, the embedding of the relationship r is represented as a vector r ∈ Rd, and the embedding of the tail entity t is represented as a vector t ∈ Rd, where d is the dimension of the embedding; the embedding vectors of the entity and the relationship are reshaped into a matrix, usually reshaping a vector of dimension d into a matrix of k×k. For example, an embedding vector of dimension 100 can be reshaped into a 10×10 matrix. The reshaped head entity embedding h′ and relationship embedding r′ are concatenated into a 2k×k matrix: H r = concat(h', r'), and the concatenated matrix contains the embedding information of the head entity and the relationship. Then, the concatenated matrix is input into a convolutional neural network, and a convolutional kernel is used to sense and extract features from the local area. The convolutional layer can capture the local interaction features between the head entity and the relationship through the local weight sharing mechanism. This step uses the convolutional layer to learn the complex feature patterns between entities and relationships. The output after convolution is processed through a non-linear activation function such as ReLU, enabling the model to capture more non-linear features. After the convolution operation, the generated feature map is converted into a one-dimensional vector through a flatten operation. This vector is then mapped through a fully connected layer to output a feature vector of a fixed dimension (usually the same as the embedding dimension d). The output vector of this layer can be regarded as the convolutional features of the head entity and the relationship, which are used to calculate the similarity with all entities in the knowledge graph. Finally, the model calculates the dot product of this feature vector and all entity embedding vectors to obtain the score of each entity as a candidate tail entity. The larger the dot product, the higher the probability that the entity is the tail entity. Finally, by sorting the scores, the entity with the highest score is selected as the predicted result of the tail entity. Given the head entity and relationship, the model outputs the scores of each entity as the tail entity after processing. All entities are sorted according to the scores to find the most matching tail entity. Similarly, given the tail entity and relationship, the model outputs the scores of each entity as the head entity after processing. All entities are sorted according to the scores to find the most matching head entity. ConvE effectively enhances the interaction characteristics of entity and relationship embeddings in the knowledge graph through convolutional neural networks and non-linear operations, improving the performance of the knowledge graph completion task.

[0151] As Figure 4 shown, in the embodiment of the present invention, step S4 includes:

[0152] S41. Input the final entity embedding representation and relationship embedding representation into the ConvE model;

[0153] S42. Generate an embedding vector for each entity and relationship through the ConvE model. Reshape the embedding vectors of each entity and relationship into matrices. Concatenate the matrix reshaped from the embedding vector of the head entity with the matrix reshaped from the embedding vector of the relationship. Perform a convolution operation on the concatenated matrix. Non-linearly activate, flatten the features, and map through a fully connected layer on the feature map generated after convolution to obtain a feature vector with a fixed dimension. Perform a dot product operation between this feature vector and the embedding vectors of all entities to obtain the scores of each entity as a candidate tail entity;

[0154] S43. Sort all entities according to the scores output by the ConvE model, take the entity with the highest score as the predicted result of the tail entity, predict the possible ICD codes corresponding to each head entity and relationship, and obtain the final list of ICD codes.

[0155] Input the final entity embedding representation and relationship embedding representation into the ConvE model to predict the missing triples. Generate an embedding vector for each entity and relationship respectively through the ConvE model. Among them, the embedding vector of the head entity h is h ∈ Rd, the embedding vector of the relationship r is r ∈ Rd, and the embedding vector of the tail entity t is t ∈ Rd, where d is the embedding dimension. Reshape the embedding vectors of each entity and relationship into matrices respectively. Concatenate the matrix reshaped from the embedding vector of the head entity (patient number) with the matrix reshaped from the embedding vector of the relationship of this head entity (such as gender, age, main diagnosis, other diagnoses, etc.) to obtain the concatenated matrix, and the concatenated matrix contains the embedding information of the head entity and its relationship. Perform a convolution operation on the concatenated matrix to sense and extract features from the local area to generate a feature map. Non-linearly activate, flatten the features, and map through a fully connected layer on the feature map in sequence to obtain a feature vector with a fixed dimension. Perform a dot product operation between this feature vector and the embedding vectors of all entities, and output the scores of each entity as a candidate tail entity. Sort all entities according to the scores output by the ConvE model, take the entity with the highest score as the predicted result of the tail entity, and predict the possible tail entity (ICD code) corresponding to each head entity (patient number) and relationship to obtain the final list of ICD codes.

[0156] As Figure 5 shown, based on the same inventive concept, an embodiment of the present invention further provides an ICD automatic coding system based on a knowledge graph. The system includes:

[0157] A case information extraction module 100, configured to construct an ICD knowledge graph based on case data, and obtain a set of triples composed of patient numbers, relationship numbers, and entity numbers; the case data includes: doctor's handwritten diagnosis, patient examination results, medical order data, and ICD codes on the medical record homepage;

[0158] The node information aggregation module 200 is used to map the triple set to the graph neural network, perform node relationship information aggregation, neighbor node and path information aggregation, and triple semantic information aggregation, and obtain the structural embedding information of the relationship, neighbor node, and triple of each entity in the graph neural network;

[0159] A hierarchical information aggregation module 300 is used to perform hierarchical information aggregation based on the structural embedding information of the relationship, neighbor nodes and triples of each entity to obtain the final entity embedding representation and relationship embedding representation;

[0160] The prediction module 400 is used to perform ICD code prediction based on the final entity embedding representation and relationship embedding representation through the ConvE model to obtain the final ICD code list.

[0161] like Figure 6 As shown, the case information extraction module 100 includes:

[0162] The entity coding module 101 is used to extract and code the entity information in the case data to obtain an entity number list; when coding the extracted entities, the same sequence number starting from 1 is used and there is no empty item in the number sequence;

[0163] The relationship encoding module 102 is used to extract and encode the relationship information in the case data to obtain a relationship number list;

[0164] The assembly module 103 is used to encode the medical record data to be classified based on the entity number list and the relationship number list, and assemble a set of triples consisting of patient numbers, relationship numbers and entity numbers.

[0165] Compared with the prior art, the present invention has the following advantages:

[0166] 1. The present invention fully explores and utilizes the doctor's diagnosis description, test results, and multiple sources of information such as surgery, treatment, medication, and consumables in the medical record, explores the correlation between each ICD code and the above information, and improves the quality of ICD coding. At present, most of the research in the field of ICD automatic coding is centered around the medical records written by doctors. However, due to the different habits of doctors in writing medical records and the deviation between the description content and the actual disease, relying solely on the medical record content for ICD coding does not fully utilize the information in the patient's medical record. The medical record data extraction method in the present invention solves this problem.

[0167] 2. The present invention uses an attention mechanism to capture the characteristics of the graph structure constructed by medical record information, which helps to explore the correlation between ICD coding results and actual medical record data.

[0168] 3. The present invention maps medical record data into structural features in a graph neural network, assigns real-world meanings to each item in the ICD coding dictionary, and completes the automatic generation of ICD codes.

[0169] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An ICD automatic encoding method based on knowledge graph, characterized in that: The method comprises: S1. Based on case data, an ICD knowledge graph is constructed to obtain a set of triples consisting of patient number, relationship number and entity number; the case data includes: doctor's handwritten diagnosis, patient examination results, medical advice data and ICD code on the front page of the medical record; S2, mapping the triple set to the graph neural network, performing node relationship information aggregation, neighbor node and path information aggregation, and triple semantic information aggregation respectively, to obtain the structural embedding information of the relationship, neighbor node, and triple of each entity in the graph neural network; S3, based on the structural embedding information of each entity’s relationship, neighbor nodes, and triples, hierarchical information aggregation is performed to obtain the final entity embedding representation and relationship embedding representation; S4. Based on the final entity embedding representation and relationship embedding representation, ICD code prediction is performed through the ConvE model to obtain the final ICD code list.

2. According to claim 1, the ICD automatic encoding method based on knowledge graph is characterized in that: Step S1 includes: S11, extracting and encoding the information of entities in the case data to obtain an entity number list; when encoding the extracted entities, the same sequence number starting from 1 is used and there is no empty item in the number sequence; S12, extracting and encoding the relationship information in the case data to obtain a relationship number list; S13. Based on the entity number list and the relationship number list, the medical record data to be classified is encoded, and a triple set consisting of a patient number, a relationship number and an entity number is assembled.

3. According to claim 2, the ICD automatic encoding method based on knowledge graph is characterized in that: Step S11 includes: The doctor's handwritten diagnosis is divided according to the writing habits, the number of groups is counted, and the order is from high to low. The items with too low frequency are manually checked, and the items with irregular writing and errors are eliminated. The remaining items are sequentially numbered as entity names; For the numerical results in the patient examination results, if they are determined to be too high or too low according to the reference value of the examination index item, "index name + high" or "index name + low" are used as the entity name for sequential numbering; for the text type results in the patient examination results, "index name + result" are used as the entity name for sequential numbering; Classify and summarize all medical order items of a patient's hospitalization in the medical order data by item name, and use the names of the medical order items as entity names to perform sequential numbering; The ICD codes on the front page of the medical record are sequentially numbered as entity names, and a dictionary is prepared to restore the numbers in the final output results to ICD codes; Step S12 includes: According to the organizational characteristics of the medical record data, a relationship is established between the doctor's handwritten diagnosis, the patient's examination results and the medical order data and they are numbered in sequence to obtain a relationship number list; the relationship names in the relationship number list include: gender, age, hospitalization department, the main diagnosis and other diagnoses in the doctor's handwritten diagnosis, the name of the examination indicator item in the patient's examination results, the medical order category in the medical order data and the main diagnosis and other diagnoses in the ICD code on the medical record front page.

4. According to claim 1, the ICD automatic encoding method based on knowledge graph is characterized in that: In step S2, mapping the triple set to the graph neural network includes: The head entity, i.e., the patient number, and the tail entity, i.e., the entity number, of each triple in the triple set are used as nodes in the graph neural network, and the relationship number is used as a feature of the edge connecting the head entity and the tail entity in the graph neural network; In step S2, aggregating node relationship information includes: The relationship between each entity and all its neighbor entities in the graph neural network is weighted summed, and the result of the weighted summation is linearly transformed to obtain the relationship embedding representation of each entity; the calculation formula is: Among them, S i rel For entity e i The relation embedding representation, e i is the embedding representation of the entity corresponding to the i-th node in the graph neural network, N i For entity e i The set of all neighbor entities of j For entity e i The relationship between the jth neighbor entity, α i,j rel For the relationship j For entity e i The importance of attention weight, W rel is the weight matrix of relation embedding, which is used to embed relation r j Mapped to a new representation space, σ is a nonlinear activation function, exp(r j T e i ) is the relationship r j With entity e i The matching degree of , i and j are positive integers.

5. According to claim 4, the ICD automatic encoding method based on knowledge graph is characterized in that: In step S2, aggregating neighbor node and path information includes: The path information of each entity in the graph neural network and its directly connected neighbor nodes is aggregated in a single layer to obtain the single layer aggregation result of each entity; the calculation formula is: Among them, S i ent Represents entity e i After one layer of aggregation, the representation has integrated the entity e i The path information of all neighboring nodes directly connected to it, e i is the embedding representation of the entity corresponding to the i-th node in the graph neural network, e j For entity e i The embedding representation of the jth neighbor entity of i For entity e i The set of all neighbor entities of i,j ent is the neighbor entity e j For entity e i The importance of attention weight, W ent is a linear transformation matrix used to transform the neighbor entity e j The embedding representation is mapped to a new feature space, σ is a nonlinear activation function, exp(e j T e i ) is the neighbor entity e j With the central entity e i The matching degree, i and j are positive integers; Based on the single-layer aggregation results of each entity, iterative multi-layer aggregation is performed until the corresponding number of iterations is reached to obtain the neighbor node embedding representation of each entity.

6. The ICD automatic encoding method based on knowledge graph according to claim 5 is characterized in that: In step S2, performing triple semantic information aggregation includes: Each neighbor entity and the corresponding relationship of each entity in the graph neural network are mapped into a fusion vector, and the triple semantic information of all fusion vectors of each entity is aggregated to obtain the triple embedding representation of each entity; the calculation formula is: Among them, S i tri For entity e i The triple embedding representation of i is the embedding representation of the entity corresponding to the i-th node in the graph neural network, e j For entity e i The embedding representation of the jth neighbor entity of i For entity e i The set of all neighbor entities of j For entity e i With neighbor entity e j The relationship between For entity e i Each neighbor entity-relation pair (e j ,r j ) is a fusion representation of i,j tri For fusion representation For entity e i The importance of attention weight, W tri is the weight matrix, which is used to represent the fusion Mapped to the target space, σ is the activation function, For fusion representation With the central entity e i Similarity, i, j, k are positive integers; The fusion representation is a composite function used to convert entity e i and the relationship j A composite operation is performed, wherein the composite operation includes: an addition operation, a dot product operation, and a multi-layer perceptron MLP concatenation operation.

7. The ICD automatic encoding method based on knowledge graph according to claim 6 is characterized in that: Step S3 includes: Iteratively perform multi-layer aggregation on the structural embedding information of each entity and its relationship, neighbor nodes, and triples in the graph neural network, initialize different relationship embedding vectors for each layer, until the corresponding number of iterations is reached, concatenate the relationship embedding vectors of all layers, perform linear transformation on the concatenated results, and obtain the final entity embedding representation and relationship embedding representation; the calculation formula for iterative multi-layer aggregation is: Among them, e i ' l+1 For entity e i The entity embedding vector after the l+1th layer aggregation, e i l For entity e i The entity embedding representation after the l-th layer aggregation is: For entity e i The relation embedding representation after the l-th layer aggregation is: For entity e i The triple embedding representation after the l-th layer aggregation, i and l are positive integers; The calculation formula for the final entity embedding representation and relationship embedding representation is: And out =and k r out =W out Concat({r l |l=1,...,K}) Among them, e out is the final entity embedding representation, e k is the entity embedding representation after the last layer of aggregation in the Kth layer, r out is the final relation embedding representation, r l is the relation embedding vector of the lth layer, W out is the transformation matrix, l and K are positive integers.

8. The ICD automatic encoding method based on knowledge graph according to any one of claims 4 to 7, characterized in that: Step S4 includes: S41, inputting the final entity embedding representation and relationship embedding representation into the ConvE model; S42. Generate an embedding vector for each entity and relationship through the ConvE model, reshape the embedding vector of each entity and relationship into a matrix, concatenate the reshaped matrix of the embedding vector of the head entity and the reshaped matrix of the embedding vector of the relationship, perform a convolution operation on the concatenated matrix, perform nonlinear activation, feature flattening, and full connection layer mapping on the feature map generated after the convolution to obtain a feature vector of a fixed dimension, perform a dot product operation on the feature vector and the embedding vectors of all entities, and obtain a score for each entity as a candidate tail entity; S43. Sort all entities according to the scores output by the ConvE model, take the entity with the highest score as the prediction result of the tail entity, predict the ICD code that may correspond to each head entity and relationship, and obtain the final ICD code list.

9. An ICD automatic encoding system based on knowledge graph, used to implement an ICD automatic encoding method based on knowledge graph according to any one of claims 1 to 8, characterized in that: The system comprises: The case information extraction module is used to construct an ICD knowledge graph based on case data to obtain a set of triples consisting of patient numbers, relationship numbers, and entity numbers; the case data includes: doctor's handwritten diagnosis, patient examination results, medical advice data, and the ICD code on the front page of the medical record; A node information aggregation module is used to map the triple set into the graph neural network, perform node relationship information aggregation, neighbor node and path information aggregation, and triple semantic information aggregation, and obtain the structural embedding information of the relationship, neighbor node, and triple of each entity in the graph neural network; The hierarchical information aggregation module is used to aggregate hierarchical information based on the structural embedding information of each entity’s relationship, neighbor nodes, and triples to obtain the final entity embedding representation and relationship embedding representation; The prediction module is used to perform ICD code prediction through the ConvE model based on the final entity embedding representation and relationship embedding representation to obtain a final ICD code list.

10. The ICD automatic coding system based on knowledge graph according to claim 9 is characterized in that: The case information extraction module includes: An entity coding module is used to extract and encode the information of entities in the case data to obtain an entity number list; when encoding the extracted entities, the same serial number starting from 1 is used and there is no empty item in the number sequence; A relationship encoding module, used to extract and encode the relationship information in the case data to obtain a relationship number list; An assembly module is used to encode the medical record data to be classified based on the entity number list and the relationship number list, and assemble a set of triples consisting of patient numbers, relationship numbers and entity numbers.

Citation Information

Cited By

  • Diagnosis code suggestion method and diagnosis code suggestion system

    TWI921237B