A medical record analysis method, device, and computer-readable storage medium

The method improves diagnostic accuracy by combining initial and authority likelihood ratios with graph convolutional networks to analyze patient data, addressing inefficiencies in existing diagnostic rule generation methods.

CN114328953BActive Publication Date: 2025-07-15ANHUI IFLYHEALTH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111564774.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-20
Publication Date
2025-07-15
Estimated Expiration
2041-12-20

AI Technical Summary

Technical Problem

In the prior art, the medical record analysis method based on logic and statistics has problems such as time-consuming and labor-intensive, low recall, insufficient capture of key information features in the patient's medical records, and low inference accuracy.

Method used

A medical record analysis method is used to obtain the initial likelihood ratio and authoritative likelihood ratio of clinical manifestations relative to the disease type, combine with authoritative statistical results, calculate the final likelihood ratio, and encode the medical record information using graph convolutional neural network and BERT network to generate diagnostic results.

Benefits of technology

It improves the accuracy of medical record analysis and the effectiveness of diagnostic reasoning, and the generated results are more accurate, which improves the accuracy of diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114328953B_ABST
    Figure CN114328953B_ABST
Patent Text Reader

Abstract

The present invention discloses a medical record analysis method, device and computer-readable storage medium. The medical record analysis method includes: obtaining an initial likelihood ratio of clinical manifestations relative to a disease type; obtaining an authoritative likelihood ratio of clinical manifestations relative to a disease type; combining the initial likelihood ratio and the authoritative likelihood ratio to determine a final likelihood ratio of clinical manifestations relative to the disease type; obtaining medical record information of a disease to be diagnosed; and using the final likelihood ratio and the medical record information to obtain the correlation between the clinical manifestations and the disease to be diagnosed, so as to obtain an analysis result of the disease to be diagnosed. By the above method, the present invention can improve the accuracy of the medical record analysis result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent medical treatment, and particularly to a medical record analysis method, device, and computer-readable storage medium. Background Art

[0002] In recent years, due to the support of medical health big data-related policies such as the "Guiding Opinions on Promoting and Standardizing the Application and Development of Health Medical Big Data" in China and the popularization of artificial intelligence technology in the medical field, artificial intelligence-assisted diagnosis and treatment has become an important part of the auxiliary decision-making system, helping multiple medical institutions and hundreds of millions of doctors improve the efficiency of clinical operations and reduce the pain of patients.

[0003] In the prior art, based on logical and statistical methods, knowledge in medical guidelines is extracted to automatically generate rule diagnosis conditions for diagnostic reasoning. Such methods are highly dependent on the formed rule conditions. Often, after the machine extraction, manual inspection is also required, which is time-consuming and laborious and has a low recall rate. Through a large amount of labeled data for sequence learning, the shallow semantic information of medical records is simply learned for diagnostic reasoning. However, such methods still have the problems of insufficient capture of the characteristics of key information in patients' medical records and low reasoning accuracy. Summary of the Invention

[0004] The main technical problem to be solved by the present invention is to provide a medical record analysis method, device, and computer-readable storage medium, which can improve the accuracy of medical record analysis.

[0005] To solve the above technical problem, a technical solution adopted by the present invention is: to provide a medical record analysis method, which includes: obtaining an initial likelihood ratio of clinical manifestations relative to disease types; obtaining an authoritative likelihood ratio of clinical manifestations relative to disease types; combining the initial likelihood ratio and the authoritative likelihood ratio to determine a final likelihood ratio of clinical manifestations relative to disease types; obtaining medical record information of a disease to be diagnosed; using the final likelihood ratio and the medical record information to obtain the correlation between the clinical manifestations and the disease to be diagnosed, and obtaining an analysis result of the disease to be diagnosed.

[0006] Among them, combining the initial likelihood ratio and the authoritative likelihood ratio to determine the final likelihood ratio includes: calculating the average value of the initial likelihood ratio and the authoritative likelihood ratio to obtain the final likelihood ratio.

[0007] Among them, obtaining the authoritative likelihood ratio of clinical manifestations includes: obtaining the importance level of each clinical manifestation; setting different likelihood ratios as the authoritative likelihood ratio for different importance levels, and the higher the importance level, the greater the corresponding authoritative likelihood ratio.

[0008] Among them, after calculating the average value of the initial likelihood ratio and the authoritative likelihood ratio, it includes: normalizing each average value.

[0009] Among them, the correlation between the clinical manifestations and the disease to be diagnosed is obtained by using the final likelihood ratio and the medical record information. The analysis results of the disease to be diagnosed include: obtaining a knowledge graph of the clinical manifestations and the final likelihood ratio, where each node of the knowledge graph includes a diagnosis type and clinical manifestations; encoding the knowledge graph to obtain a knowledge graph encoding matrix; encoding the medical record information to obtain a medical record information encoding matrix; calculating a correlation matrix between the knowledge graph encoding matrix and the medical record information encoding matrix; calculating the probability of the diagnosis type in the knowledge graph by using the correlation matrix, and taking the diagnosis type with the highest probability value as the analysis result.

[0010] Among them, the medical record information includes a medical record text and a clinical manifestation relationship graph. Encoding the medical record information to obtain a medical record information encoding matrix includes: extracting the clinical manifestation information in the medical record information, and obtaining the clinical manifestation relationship graph of the clinical manifestation information; encoding the medical record text to obtain a medical record text encoding matrix; encoding the clinical manifestation relationship graph to obtain a clinical manifestation relationship graph encoding matrix; the medical record information encoding matrix includes the medical record text encoding matrix and the clinical manifestation relationship graph encoding matrix.

[0011] Among them, calculating the correlation matrix between the knowledge graph encoding matrix and the medical record information encoding matrix includes: calculating a first correlation matrix between the knowledge graph encoding matrix and the medical record text encoding matrix; calculating a second correlation matrix between the knowledge graph encoding matrix and the clinical manifestation relationship graph encoding matrix; fusing the first correlation matrix and the second correlation matrix to obtain the correlation matrix.

[0012] Among them, encoding the knowledge graph includes:

[0013] Obtaining the graph matrix of the knowledge graph, where the graph matrix includes an adjacency matrix, a degree matrix, and a likelihood ratio matrix;

[0014] Encoding the graph matrix by using a graph convolutional neural network.

[0015] Among them, calculating the probability of the diagnosis type in the knowledge graph by using the correlation matrix includes: extracting the diagnosis type correlation matrix containing the diagnosis type in the correlation matrix; calculating the probability of the diagnosis type by using a multi-classification model and the diagnosis type correlation matrix.

[0016] Among them, the multi-classification model is trained by using binary cross-entropy.

[0017] To solve the above technical problems, another technical solution adopted by the present invention is: providing a medical record analysis device, which includes a processor for executing to implement the above medical record analysis method.

[0018] To solve the above technical problems, another technical solution adopted by the present invention is: to provide a computer-readable storage medium for storing instructions / program data, and the instructions / program data can be executed to implement the above-mentioned medical record analysis method.

[0019] The beneficial effects of the present invention are as follows: Different from the prior art, the present invention proposes a medical record analysis method, which calculates the likelihood ratio based on a large amount of data and authoritative statistical results, and uses the fusion of the two as the final likelihood ratio, so that the generated result is more accurate. At the same time, the likelihood ratio and medical record information are interacted to learn the relevant important features to obtain the final analysis result, which can greatly improve the effectiveness of diagnostic reasoning and improve the diagnostic accuracy. Description of the Drawings

[0020] Figure 1 is a schematic flowchart of a medical record analysis method in an embodiment of the present application;

[0021] Figure 2 is a schematic flowchart of another medical record analysis method in an embodiment of the present application;

[0022] Figure 3 is a schematic diagram of a knowledge graph of the present application;

[0023] Figure 4 is a schematic structural diagram of a medical record analysis model in an embodiment of the present application;

[0024] Figure 5 is a schematic structural diagram of a medical record analysis device in an embodiment of the present application;

[0025] Figure 6 is a schematic structural diagram of a medical record analysis device in an embodiment of the present application;

[0026] Figure 7 is a schematic structural diagram of a computer-readable storage medium in an embodiment of the present application. Detailed Embodiments

[0027] To make the objectives, technical solutions and effects of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples.

[0028] At present, there are already a large number of research results on automatic diagnosis based on a large amount of narrative documents of electronic medical records available for clinical use. Among them, the likelihood ratio is a statistical index often used in epidemiology and an important index for evaluating diagnostic tests in evidence-based medicine. It is an index reflecting authenticity and a composite index reflecting both sensitivity and specificity. It is the ratio of the probability of obtaining a certain screening test result among those with the disease to the probability of obtaining this result among those without the disease.

[0029] Since clinical manifestations can be divided into positive and negative ones, the likelihood ratio can be correspondingly divided into the positive likelihood ratio (LR+) and the negative likelihood ratio (LR-). Specifically, the positive likelihood ratio includes the likelihood ratios of symptoms, signs, primary medical history, causative factors, positive auxiliary examination results, positive laboratory test results, complications, and risk factors; the negative likelihood ratio includes the likelihood ratios of non-abnormal findings, negative signs, denied medical history, negative auxiliary examination results, and negative laboratory test results. If the positive likelihood ratio of a certain clinical manifestation for a certain disease is larger, it indicates that the contribution of this clinical manifestation to the diagnosis of this disease is greater; on the contrary, if the negative likelihood ratio of a certain clinical manifestation for a certain disease is larger, it indicates that this clinical manifestation can better rule out this disease.

[0030] In medicine, when calculating the likelihood ratio of a clinical manifestation compared to a disease, it is generally calculated through sensitivity and specificity. Sensitivity and specificity are calculated based on the number of true positives (TP), false positives (FP), true negatives (TN), and false negatives (FN) in the test sample, and the formulas are as follows:

[0031]

[0032] The method for calculating the likelihood ratio using sensitivity and specificity is as follows:

[0033]

[0034] In practice, for specific single or several diseases, it is relatively easy to manually perform statistical calculations of Sen and Spe on the test sample results. However, in reality, there are thousands or even tens of thousands of diseases and clinical manifestations, and manual statistics are too time-consuming and laborious. Therefore, this application provides a medical record analysis method that calculates the likelihood ratio based on a large amount of data and authoritative statistical results for further medical record analysis. The medical record analysis method proposed in this application extracts and learns the information in the medical record book recording the patient during the treatment and diagnosis process for subsequent medical record management and treatment.

[0035] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a medical record analysis method in an embodiment of this application. It should be noted that if there are substantially the same results, this embodiment is not limited to the Figure 1 process sequence shown. As Figure 1 shown, this embodiment includes:

[0036] S110: Obtain the initial likelihood ratio of the clinical manifestation relative to the disease type and obtain the authoritative likelihood ratio of the clinical manifestation relative to the disease type.

[0037] Obtain information such as medical records and examination reports of tens of millions of patients. Each piece of information contains some symptoms, signs, examination results, etc., which are called clinical manifestations. Calculate the patient sensitivity and specificity for these pieces of information to obtain the initial likelihood ratio of the clinical manifestations relative to the disease type.

[0038] Obtain the importance levels of each clinical manifestation. Obtain the diagnostic guidelines for each disease and the authoritative statistical results of epidemiology, and classify the importance levels of the clinical manifestations related to each disease. In this application, the importance levels are divided into ten levels, namely super-typical A, typical B, common C, and non-specific D. Among them, super-typical means that this clinical manifestation can directly serve as the gold standard for diagnosing the disease; typical means that when this clinical manifestation exists, it is highly likely to be this disease; common means that this clinical manifestation is a symptom that often accompanies the occurrence of this disease, but is not unique to this disease; non-specific means that this clinical manifestation may be a symptom that accompanies the occurrence of this disease, but does not often appear.

[0039] Set different likelihood ratios as authoritative likelihood ratios for different importance levels. The higher the importance level, the greater the corresponding authoritative likelihood ratio. Based on the importance levels of the clinical manifestations statistically determined, different levels are respectively corresponding to different scores, and this score is the authoritative likelihood ratio. In this embodiment, the corresponding relationship between the importance level and the authoritative likelihood ratio is: super-typical A is 0.9, typical B is 0.7, common C is 0.5, and non-specific D is 0.2.

[0040] S130: Combine the initial likelihood ratio and the authoritative likelihood ratio to determine the final likelihood ratio of the clinical manifestations relative to the disease type.

[0041] Calculate the average value of the initial likelihood ratio and the authoritative likelihood ratio to obtain the final likelihood ratio. In one embodiment, for the convenience of subsequent learning and calculation of the data information of the final likelihood ratio, perform normalization processing on the final likelihood ratio. Please refer to Table 1, which is the calculation table of the final likelihood ratio of chronic tonsillitis.

[0042] Table 1 Calculation Table of the Final Likelihood Ratio of Chronic Tonsillitis

[0043]

[0044] S150: Obtain the medical record information of the disease to be diagnosed.

[0045] Obtain the medical record text information of the disease to be diagnosed. The medical record text information includes the chief complaint, present illness history, past history, physical examination, auxiliary examination, and clinical manifestations, etc. Each part contains some symptoms, signs, examination results, etc. of the clinical manifestations, which are called clinical manifestations.

[0046] S170: Use the final likelihood ratio and the medical record information to obtain the correlation between the clinical manifestations and the disease to be diagnosed, and obtain the analysis result of the disease to be diagnosed.

[0047] The final likelihood ratio is an important indicator reflecting the importance of clinical manifestations relative to the disease type. By combining the final likelihood ratio with the patient's medical record information, the correlation between the two is determined to infer the disease to be diagnosed for the patient and obtain the analysis result of the disease to be diagnosed.

[0048] There are complex correlation relationships between clinical manifestations, between clinical manifestations and disease types. For example, one disease may cause multiple clinical manifestations, and one clinical manifestation may be caused by multiple diseases; there are also correlation relationships between different clinical manifestations in the same disease. The patient's medical record information is generally text information. In one embodiment, the medical record text information can be directly combined with the final likelihood ratio information to infer the analysis result of the disease to be diagnosed. In actual situations, a pure text sequence learning model cannot effectively understand the structural correlation information between clinical manifestations and between clinical manifestations and disease types in the original medical record text. At the same time, when the same clinical manifestation appears in different diseases, its contribution degree to the diagnosis of different diseases is also different, that is, the importance degree is also different. For example, cough (clinical manifestation) is a typical symptom in bronchitis (disease type), but it is only an accompanying symptom in rhinitis (disease type). The typical symptoms of rhinitis are nasal congestion, runny nose, decreased sense of smell and other nasal symptoms. It is difficult for the model to pay attention to these typical information through pure text classification or sequence learning. Therefore, in another embodiment, named entity recognition and relation extraction techniques are used to extract the relationships between clinical manifestations in the medical record information, obtain a clinical manifestation relationship graph of the medical record information, and combine the clinical manifestation relationship graph with the final likelihood ratio information to infer the analysis result of the disease to be diagnosed. In another embodiment, the medical record text information, the clinical manifestation relationship graph and the final likelihood ratio information of the medical record information are combined to infer the analysis result of the disease to be diagnosed.

[0049] In the embodiments of the present application, the final likelihood ratio information is respectively encoded into a graph, the medical record text is encoded into text, and the clinical manifestation relationship graph is encoded into a graph. The interactive attention mechanism is used to enhance the features of the three groups of encoded features with each other, and finally the final decision is obtained through the decision layer for the features after interaction.

[0050] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of another medical record analysis method in the embodiments of the present application. Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of the medical record analysis model in the embodiments of the present application. It should be noted that if there are substantially the same results, this embodiment is not limited to Figure 2 the process sequence shown. As Figure 2 shown, this embodiment includes:

[0051] S210: Obtain the knowledge graph of clinical manifestations and final likelihood ratios.

[0052] Extract data information from the final likelihood ratios of clinical manifestations relative to disease types, and construct a knowledge graph using the extracted data information. The obtained knowledge graph is G=(V, E). There are N KG nodes in this knowledge graph, and the nodes are v i ∈V. The edges are (v i , v j ) ∈E. Among them, each node in the obtained knowledge graph includes a diagnosis type and a clinical manifestation. The weight of each edge in the knowledge graph is the final likelihood ratio between the two connected nodes, that is, the final likelihood ratio between a diagnosis type and a clinical manifestation, or the final likelihood ratio between diagnosis types, or the final likelihood ratio between clinical manifestations. For example, the diagnosis types include gastritis, acute gastritis, and chronic gastritis. There are connections between gastritis and acute gastritis and chronic gastritis respectively, that is, there are final likelihood ratios. For example, there is a connection between the clinical manifestation related to the upper limb and the clinical manifestation related to the left upper limb, that is, there is a final likelihood ratio. For ease of explanation, please refer to Figure 3 , Figure 3 which is a schematic diagram of a knowledge graph in this application. In Figure 3 , there are 4 nodes, namely nodes 1-4. Among them, nodes 1, 2, and 4 are clinical manifestations, and node 3 is a diagnosis type. Each edge represents the final likelihood ratio between the two connected nodes. Among them, the final likelihood ratio between node 1 and node 2 is 0.8, and the final likelihood ratio between node 1 and node 2 is 0.2.

[0053] S230: Encode the knowledge graph to obtain a knowledge graph encoding matrix.

[0054] Obtain the graph matrix of the knowledge graph. The graph matrix includes an adjacency matrix A, a degree matrix D, and a likelihood ratio matrix L. Among them, the adjacency matrix A is used to represent whether there is a connection between every two nodes. To ensure that its own information is not lost, it is default that there is a connection between each node itself. The degree matrix D is used to represent how many nodes each node is connected to. The likelihood ratio matrix L is used to represent the final likelihood ratio between each node. It is default that the final likelihood ratio of each node to its own node is 1. For ease of explanation, please refer to Figure 3 , Figure 3 the adjacency matrix A of the knowledge graph in 、the degree matrix D= and the likelihood ratio matrix L= .

[0055] Please refer to Figure 4, the knowledge graph is encoded using a Graph Convolutional Network (GCN) to obtain a knowledge graph encoding matrix. In an embodiment of the present application, based on the GCN encoding, the weight information of the edges in the knowledge graph with the final likelihood ratio is incorporated, and the encoding method in the present application is referred to as LR-GCN. The graph matrix of the knowledge graph in the present application is encoded using LR-GCN, and the specific encoding method is as follows:

[0056] LR-GCN(X, A, L) = ALReLU(ALXW0)W1,

[0057] where X represents the input features of each node in the knowledge graph, and W0 and W1 are training parameters obtained during the training of the LR-GCN model, where ReLU is the activation function.

[0058] In the present application, a graph convolutional network is used to perform graph encoding on the disease knowledge graph to obtain graph features. In addition to encoding the nodes and edges in the graph, the likelihood ratio of each clinical manifestation node to the disease node, as the importance of the edge, is also encoded and fused with the graph encoding features of the disease knowledge graph to obtain knowledge graph features. Adding the likelihood ratio encoding features in the medical knowledge graph encoding can greatly improve the effectiveness of diagnostic reasoning.

[0059] S250: Encode the medical record information to obtain a medical record information encoding matrix.

[0060] The medical record information includes medical record text and a clinical manifestation relationship graph. The medical record information encoding matrix includes a medical record text encoding matrix and a clinical manifestation relationship graph encoding matrix.

[0061] Please refer to Figure 4 , in an embodiment of the present application, the BERT network is used to encode the medical record text. First, training set data is obtained. The training set data is a large amount of unlabeled medical domain data, where the medical domain data includes outpatient medical records, inpatient medical records, medical guidelines, and medical textbooks, etc. The BERT network is trained without supervision using the training set data to obtain a trained BERT network.

[0062] To enable the BERT network to better learn the overall features of medical records. Therefore, identifiers are first added to the text fields in the medical record text. The identifiers include the CLS identifier and the SEP identifier. The CLS identifier indicates the start of the text, and the SEP identifier indicates the end of the text or the separation of two text fields. For example, in an embodiment of the present application, the medical record text information includes the chief complaint, current medical history, past medical history, physical examination, auxiliary examination, and clinical manifestations. First, the six text fields are concatenated and need to be processed into: [CLS] chief complaint [SEP] current medical history [SEP] past medical history [SEP] physical examination [SEP] auxiliary examination [SEP] clinical manifestations [SEP]. The processed medical record text is input into the trained BERT network for encoding. During the encoding process, different field type encodings are added to different text fields, and the dimension of the encoded matrix obtained is N MR x F, where N MR is the length of the medical record text, and F is the length of the hidden layer. The matrix representing the medical record features of the medical record text is extracted using the BERT network, and the medical record text encoding matrix is obtained as the result output. The size of the medical record text encoding matrix is (N MR , F). In a specific embodiment, the length of the hidden layer is 512, so the size of the medical record text encoding matrix is (N MR , 512).

[0063] In the medical record information, there are certain associations between the texts. The named entity recognition and relation extraction technologies are used to extract the relationships between the clinical manifestations in the medical record information, extract the associations of the clinical manifestation information in the medical record information, and obtain the clinical manifestation relationship diagram of the clinical manifestation information. The clinical manifestation relationship diagram contains multiple nodes, and the nodes are each clinical manifestation. There may be associations between the clinical manifestations. When there is an association between two clinical manifestations, the two corresponding nodes are connected by an edge. It should be noted that in this embodiment, the clinical manifestation relationship diagram is different from the knowledge graph of the above-mentioned final likelihood ratio. The knowledge graph of the final likelihood ratio not only has nodes and edges, but each edge has a weight representing the final likelihood ratio, while the clinical manifestation relationship diagram only has nodes and the edges representing the node association relationships between the nodes.

[0064] The relationship diagram matrix of the clinical manifestation relationship diagram is obtained. The relationship diagram matrix includes the adjacency matrix B and the degree matrix C. Among them, the adjacency matrix B is used to indicate whether there is a connection between every two nodes. To ensure that its own information is not lost, it is default that there is a connection between each node itself; the degree matrix C is used to indicate how many nodes each node is connected to. Please refer to Figure 4 , and the graph convolutional neural network is used to encode the clinical manifestation relationship diagram to obtain the clinical manifestation relationship diagram encoding matrix. The specific encoding method is as follows:

[0065] GCN(X, B) = BReLU(BXW2)W3,

[0066] where X represents the input features of each node in the clinical manifestation relationship graph, W2 and W3 are training parameters obtained during the training of the GCN model, and ReLU is the activation function.

[0067] In this application, through semantic structured extraction of the input medical records, the medical record graph information is obtained and the medical record graph encoding features are obtained through the graph convolutional network. Secondly, in order to ensure that the original information is not lost, the original medical record text is also encoded, and then through the gating mechanism, the medical record text features and the medical record graph features are fused to obtain the medical record features. This solves the problem of the insensitivity of pure text models to structured information.

[0068] S270: Calculate the correlation matrix between the knowledge graph encoding matrix and the medical record information encoding matrix.

[0069] The medical record information encoding matrix includes the medical record text encoding matrix and the clinical manifestation relationship graph encoding matrix. The knowledge graph encoding is interacted with the medical record text encoding and the knowledge graph encoding is interacted with the clinical manifestation relationship graph encoding, and the respective correlation scores are calculated. Calculate the first correlation matrix between the knowledge graph encoding matrix and the medical record text encoding matrix. First, through dot product calculation, the correlation between the knowledge graph encoding and the medical record text encoding is obtained, and the calculation method is as follows:

[0070] ,

[0071] where, E MR is the medical record text encoding matrix, with a dimension of , N MR is the length of the medical record sequence, F is the output dimension of the encoding layer; E KG is the knowledge graph encoding matrix, with a dimension of , N KG is the number of nodes in the knowledge graph, F is the output dimension of the encoding layer; W MK , W KG are trainable training parameters, both with a dimension of , M MR,KG is the calculated first weight, with a dimension of .

[0072] Then, use the softmax function and the first weight to calculate the first correlation matrix, and the specific calculation method is as follows:

[0073] ,

[0074] where, is the first correlation matrix, with a dimension of .

[0075] Calculate the second correlation matrix of the knowledge graph encoding matrix and the clinical manifestation relationship graph encoding matrix. First, through dot product calculation, obtain the correlation between the knowledge graph encoding and the clinical manifestation relationship graph encoding. The calculation method is as follows:

[0076] ,

[0077] where E F is the clinical manifestation relationship graph encoding matrix, with a dimension of , N MR is the number of nodes in the clinical manifestation relationship graph, F is the output dimension of the encoding layer; E KG is the knowledge graph encoding matrix, N KG is the number of nodes in the knowledge graph, F is the output dimension of the encoding layer, and the dimension is ; W KG , W F are trainable training parameters, both with a dimension of , M F,KG is the calculated second weight, with a dimension of .

[0078] Then, use the softmax function and the second weight to calculate the second correlation matrix. The specific calculation method is as follows:

[0079] ,

[0080] where is the second correlation matrix, with a dimension of .

[0081] Fuse the first correlation matrix and the second correlation matrix to obtain the correlation matrix between the medical record encoding and the knowledge graph encoding. The specific fusion method is as follows:

[0082] ,

[0083] where T KG is the correlation matrix, with a dimension of .

[0084] In this application, the graph features, medical record text features, and clinical manifestation relationship graph features mutually enhance their features through the interactive attention mechanism of the interaction layer, and obtain important features relative to each other, and then perform the final decision through the decision layer.

[0085] S290: Calculate the probability of the diagnosis type in the knowledge graph using the correlation matrix, and take the diagnosis type with the highest probability value as the analysis result.

[0086] In the embodiments of the present application, each node of the knowledge graph includes a diagnosis type and clinical manifestations. The correlation matrix includes the correlation between the diagnosis type and medical record information and the correlation between clinical manifestations and medical record information. Only the diagnosis type can be used as the final analysis result, that is, the medical record of the patient is analyzed to obtain the result of the patient's diagnosis type. Therefore, first extract the diagnosis type correlation matrix containing the diagnosis type in the correlation matrix. Among the nodes of the knowledge graph, there are a total of N D nodes that are diagnosis type nodes, and the diagnosis type correlation matrix is , and the dimension is .

[0087] Use a multi-classification model and the diagnosis type correlation matrix to perform multi-classification on the diagnosis type correlation matrix, and calculate the probabilities of each diagnosis type. The specific calculation method is as follows:

[0088] ,

[0089] where W is a trainable training parameter with a dimension of , and Score is the probability of the output diagnosis type with a dimension of .

[0090] Select the diagnosis type with the highest probability as the diagnosis result of the disease to be diagnosed in this application.

[0091] Among them, before using the multi-classification model, first use the loss function to train the multi-classification model. Select the binary cross-entropy as the loss function. During the training process, use the Adam optimizer to optimize the model parameters, and complete the training optimization by minimizing the cross-entropy loss. The specific calculation of the binary cross-entropy is as follows:

[0092] ,

[0093] where C represents the number of training samples.

[0094] In this embodiment, by combining a large amount of data and statistical results for the clinical manifestation likelihood ratio statistical method, knowledge is incorporated into the process of likelihood ratio statistics; a graph convolutional network incorporating likelihood ratio is proposed to improve the effectiveness of representation learning; using the medical record text, clinical manifestation relationship graph, and disease knowledge graph as inputs, reasoning is performed on the patient's diagnosis. A graph convolutional network is introduced to perform representation learning on the disease knowledge graph and the clinical manifestation relationship graph; an interactive attention mechanism is used to enhance the representation of the graph through the clinical manifestation relationship graph, and then the representation of the graph is used to enhance the representation of the clinical manifestation relationship graph; the mutual attention mechanism between the pure text and the graph is retained to prevent the loss of original information. Finally, the two interactive features are fused to improve the accuracy of diagnostic reasoning.

[0095] Please refer to Figure 5, Figure 5 is a schematic structural diagram of a medical record analysis device in an embodiment of the present application. In this embodiment, the medical record analysis device includes a first acquisition module 51, a combination module 52, a second acquisition module 53, and an analysis module 54.

[0096] Among them, the first acquisition module 51 is used to obtain the initial likelihood ratio of clinical manifestations relative to disease types; obtain the authoritative likelihood ratio of clinical manifestations relative to disease types; the combination module 52 is used to combine the initial likelihood ratio and the authoritative likelihood ratio to determine the final likelihood ratio of clinical manifestations relative to disease types; the second acquisition module 53 is used to obtain the medical record information of the disease to be diagnosed; the analysis module 54 is used to use the final likelihood ratio and the medical record information to obtain the correlation between clinical manifestations and the disease to be diagnosed, and obtain the analysis result of the disease to be diagnosed. This medical record analysis device is used to calculate the likelihood ratio based on a large amount of data and authoritative statistical results, and use the fusion of the two as the final likelihood ratio, and the generated result is more accurate. At the same time, the likelihood ratio and medical record information are interacted to learn relevant important features to obtain the final analysis result, which can greatly improve the effectiveness of diagnostic reasoning and improve diagnostic accuracy.

[0097] Please refer to Figure 6 , Figure 6 is a schematic structural diagram of a medical record analysis device in an embodiment of the present application. In this embodiment, the medical record analysis device 61 includes a processor 62.

[0098] The processor 62 can also be called a CPU (Central Processing Unit, central processing unit). The processor 62 may be an integrated circuit chip with signal processing capabilities. The processor 62 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor 62 can also be any conventional processor, etc.

[0099] The medical record analysis device 61 may further include a memory (not shown in the figure) for storing instructions and data required for the operation of the processor 62.

[0100] The processor 62 is used to execute instructions to implement the methods provided by any embodiment and any non-conflicting combination of the medical record analysis methods of the present application described above.

[0101] Please refer to Figure 7 , Figure 7It is a schematic structural diagram of a computer-readable storage medium in an embodiment of the present application. The computer-readable storage medium 71 of the embodiment of the present application stores instruction / program data 72, and when the instruction / program data 72 is executed, it implements the methods provided by any embodiment of the medical record analysis method of the present application and any non-conflicting combination. Among them, the instruction / program data 72 can form a program file and be stored in the above storage medium 71 in the form of a software product, so that a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor can execute all or part of the steps of the methods of various embodiments of the present application. The foregoing storage medium 71 includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, or terminal devices such as computers, servers, mobile phones, and tablets.

[0102] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0103] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0104] The above is only the embodiment of the present invention, and it does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be included in the patent protection scope of the present invention by the same token.

Claims

1. A medical record analysis method, characterized in that, The method includes: Obtaining an initial likelihood ratio of clinical manifestations with respect to a disease type; obtaining an authoritative likelihood ratio of the clinical manifestations with respect to the disease type, where the authoritative likelihood ratio is a set likelihood ratio that matches the importance level of the clinical manifestations, and the higher the importance level, the greater the corresponding authoritative likelihood ratio; Combining the initial likelihood ratio and the authoritative likelihood ratio to determine a final likelihood ratio of the clinical manifestations with respect to the disease type; Obtaining medical record information of a disease to be diagnosed; Using the final likelihood ratio and the medical record information to obtain the correlation between the clinical manifestations and the disease to be diagnosed, and obtaining an analysis result of the disease to be diagnosed; Among them, the step of obtaining the analysis result of the disease to be diagnosed includes: obtaining a knowledge graph of the clinical manifestations and the final likelihood ratio, where the nodes of the knowledge graph include diagnosis types and clinical manifestations, and the weights of the edges of the knowledge graph are the final likelihood ratios between the two connected nodes; encoding the knowledge graph to obtain a knowledge graph encoding matrix; encoding the medical record information to obtain a medical record information encoding matrix; calculating a correlation matrix between the knowledge graph encoding matrix and the medical record information encoding matrix; extracting a diagnosis type correlation matrix containing the diagnosis type from the correlation matrix, where the diagnosis type correlation matrix includes the correlation between the diagnosis type and the medical record information, and using a multi-classification model and the diagnosis type correlation matrix to calculate the probability of the diagnosis type in the knowledge graph, and taking the diagnosis type with the highest probability value as the analysis result.

2. The medical record analysis method according to claim 1, wherein The combining the initial likelihood ratio and the authoritative likelihood ratio to determine the final likelihood ratio of the clinical manifestations with respect to the disease type includes: Calculating the average value of the initial likelihood ratio and the authoritative likelihood ratio to obtain the final likelihood ratio.

3. The medical record analysis method according to claim 2, characterized in that, After calculating the average value of the initial likelihood ratio and the authoritative likelihood ratio, it includes: Normalizing each of the average values.

4. The medical record analysis method according to claim 1, wherein, The medical record information includes a medical record text and a clinical manifestation relationship diagram, and the encoding the medical record information to obtain a medical record information encoding matrix includes: Extracting clinical manifestation information from the medical record information and obtaining a clinical manifestation relationship diagram of the clinical manifestation information; Encoding the medical record text to obtain a medical record text encoding matrix; encoding the clinical manifestation relationship diagram to obtain a clinical manifestation relationship diagram encoding matrix; the medical record information encoding matrix includes the medical record text encoding matrix and the clinical manifestation relationship diagram encoding matrix.

5. The medical record analysis method according to claim 4, wherein The calculating the correlation matrix between the knowledge graph encoding matrix and the medical record information encoding matrix includes: Calculating a first correlation matrix between the knowledge graph encoding matrix and the medical record text encoding matrix; calculating a second correlation matrix between the knowledge graph encoding matrix and the clinical manifestation relationship diagram encoding matrix; Fusing the first correlation matrix and the second correlation matrix to obtain the correlation matrix.

6. The medical record analysis method according to claim 1, wherein The encoding the knowledge graph includes: Obtaining a graph matrix of the knowledge graph, where the graph matrix includes an adjacency matrix, a degree matrix, and a likelihood ratio matrix; Encode the spectral matrix using a graph convolutional neural network.

7. The medical record analysis method according to claim 1, wherein The multi-classification model is trained using binary cross-entropy.

8. A medical record analysis device, characterized in that, It includes a processor, and the processor is used to execute instructions to implement the medical record analysis method according to any one of claims 1-7.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store instructions / program data, and the instructions / program data can be executed to implement the medical record analysis method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Methods and apparatus for phenotype-driven clinical genomics using a likelihood ratio paradigm

    CN113272912A