Disease auxiliary diagnosis system based on multi-source information fusion

By building a subgraph and timing module combined with attention network, the problem of multi-source medical data fusion is solved, more accurate disease-assisted diagnosis and interpretability is achieved, the problem of insufficient data fusion in the existing technology is solved, and the accuracy and interpretability of disease prediction is improved.

CN120299682APending Publication Date: 2025-07-11SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510427028.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing disease-assisted diagnosis system is difficult to effectively integrate multi-source medical data with strong heterogeneity, fails to fully explore the correlation between different source information, and lacks utilization of current medical information, resulting in insufficient accuracy and interpretability of prediction results.

Method used

The pre-training module is used to extract medical entity data and construct sub-graphs, and pre-training is performed by combining external and local knowledge; the timing module modeles the timing relationship of multiple visits to patients, and uses a two-layer attention network to measure the importance of visits; the similar visits module mines similar visits to provide reference for disease diagnosis; the current visits module extracts the current visits data and provides direct evidence after encoding through the pre-training model.

Benefits of technology

It improves the accuracy of disease-assisted diagnosis, reduces dependence on past historical medical records, enhances the quality of model encoder, provides a reference for future disease development, and improves the interpretability of predicted results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299682A_ABST
    Figure CN120299682A_ABST
Patent Text Reader

Abstract

The invention discloses an auxiliary disease diagnosis system based on multi-source information fusion, which is used for extracting medical entity information from an electronic health record and constructing a related local sub-graph in combination with internal statistical data and external knowledge so as to mine correlation among different entities and model a disease development condition of a patient. And similar doctor-seeing information is mined as a reference for future disease development, and meanwhile, the high dependence on historical doctor-seeing records in the past is reduced. According to the method and the system, the dependence on past historical treatment records is reduced after the symptom information of the current treatment is considered. Therefore, better entity representation is obtained, and good support is provided for subsequent downstream tasks. The method accords with the general practice of doctors in disease diagnosis, can provide certain interpretation for the result of auxiliary disease diagnosis, and is beneficial to the development and application of an intelligent medical system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of disease assisted diagnosis, and in particular to a disease assisted diagnosis system based on multi-source information fusion. Background Art

[0002] With the rapid development of medical informatization, electronic health record systems have been widely used in medical institutions at all levels. The electronic health record system stores various medical information of patients in an electronic manner, including demographic characteristics (such as age, gender, race, etc.), laboratory test results, diagnosis records, and medication information, etc. In recent years, the electronic health record system has accumulated a large amount of medical data, providing an important data basis for the development of intelligent medical technologies. Among them, disease assisted diagnosis is a key task. By analyzing the historical medical data of patients, it is possible to predict the diseases that may occur in the future and provide assistance for doctors' clinical decisions. This is of great significance for the early prevention and intervention of certain serious diseases (such as heart failure). The existing disease prediction technologies are mainly divided into the following categories: One type of method is based on the historical disease diagnosis information of patients. By analyzing the disease codes in past medical records, the disease development rules are mined and transformed into patient feature representations. This method is simple to implement and has good interpretability. However, since only the disease diagnosis information is considered and other important medical data (such as test results, medication records, etc.) are ignored, the accuracy and reliability of the prediction results are limited. Another type of method is a rule-based expert system. The disease risk is evaluated through the experience rules summarized by medical experts. This method has strong professionalism and interpretability. However, since it cannot effectively utilize the rules hidden in the large amount of medical data, and the formulation and update of the rules require a large amount of manual intervention, its adaptability and scalability are poor. In recent years, with the development of deep learning technologies, intelligent disease prediction methods based on multi-source information fusion have gradually become a research hotspot. This method not only considers the historical disease information of patients, but also comprehensively analyzes multi-source heterogeneous data such as demographic characteristics, test results, and medication records. Through deep learning models, the complex associations between different types of medical information can be automatically learned, thereby obtaining a more comprehensive patient feature representation. However, this type of method also faces some challenges: First, the medical data from different sources has heterogeneity and sparsity, and how to effectively fuse these data is still a difficult point; Second, the problems of missing and noise generally exist in medical data, affecting the robustness of the model; Finally, the "black box" characteristic of deep learning models makes the interpretability of the prediction results poor, which is an important issue in the medical field. Summary of the Invention

[0003] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a disease auxiliary diagnosis system based on multi-source information fusion to solve the problems in the existing disease auxiliary diagnosis system, that is, it is difficult to effectively fuse heterogeneous multi-source medical data, and fails to fully explore the relationship between different source information, resulting in the encoded representation being unable to properly model the patient's condition; at the same time, when discovering and utilizing important similar patient information, the solution is relatively simple and does not consider many aspects, but is difficult to provide useful reference for downstream tasks; and it is too dependent on information such as historical medical conditions. Without medical record data algorithms, auxiliary diagnosis cannot be provided, and there is a lack of further utilization of current medical information. These technical problems.

[0004] To achieve the above object, the technical solution provided by the present invention is: a disease auxiliary diagnosis system based on multi-source information fusion, comprising:

[0005] The pre-training module is used to extract medical entity data and build subgraphs based on external and local knowledge. It uses single visit data to complete pre-training and enhance the model's ability to encode medical entities.

[0006] The time series module is used to model the time series relationship between multiple visits of patients. It uses a two-layer attention network to measure the importance of different visits and different medical entities to model the progression of the disease.

[0007] Similar visit module, which is used to mine each visit of each patient and find similar visit groups under multiple metrics. By combining different types of visit groups, it provides a reference for the patient's disease diagnosis;

[0008] The current visit module is used to extract the patient's current visit data, which is encoded through a pre-trained model to provide direct evidence for disease diagnosis reference.

[0009] Furthermore, the similar medical consultation module includes a similar disease condition module, a similar symptom module, and a cross-screening module, wherein:

[0010] The similar disease status module extracts the disease status information of each patient's visit from the local medical health data, performs similarity measurement calculation, selects a group of visit information that is most similar to it, called similar visits, and extracts the next visit of the similar visit as a reference for the disease development status at the next visit of the current visit;

[0011] The similar symptom module extracts the symptom information of each patient's current visit from the local medical health data, performs similarity measurement calculation, selects a group of visit information that is most similar to it, called similar visits, and extracts the disease condition information of the similar visits as a reference for the current disease diagnosis;

[0012] The cross-screening module extracts the similar medical visits extracted by the similar disease condition module and the similar symptom module for the current medical visit, and screens out a batch of medical visit information with the most reference value for the current medical visit by taking the intersection.

[0013] Furthermore, the pre-training module includes the following steps:

[0014] 1) Extract three types of medical entities from the electronic health record data, including: disease entities, drug entities, and symptom entities, and construct three subgraphs respectively, including a knowledge graph subgraph, a co-occurrence graph subgraph, and a hierarchical structure graph subgraph;

[0015] 2) Use the graph neural network GNN to encode the three subgraphs constructed in step 1) to obtain the representations of the corresponding subgraphs, and fuse the representations of the three subgraphs into the representation of the medical entity by weighted summation. The obtained representation will be used as the basis for subsequent medical visit representation calculation. Taking one medical visit as a unit, input the obtained representation of the medical entity into the Transformer model to obtain the medical visit representation;

[0016] 3) Make full use of the single medical visit data as the training data and perform two types of reconstruction tasks. The reconstruction tasks include: splicing the representation of the symptom entity and the representation of the drug entity to predict the disease entity; splicing the representation of the symptom entity and the representation of the disease entity to predict the disease entity; during the pre-training process, use the binary cross-entropy loss function L BCE Calculate the error between the predicted disease entity and the real disease entity.

[0017] Furthermore, the time series module includes a time series information extraction module and an attention fusion module, where:

[0018] The time series information extraction module is used to model the time interval information between different medical visits of the patient, and use the RNN module to integrate a series of representations within the medical visit sequence into one representation, including the following steps:

[0019] 1) Encode the medical entities in each medical visit record, and encode and splice the medical entities through the pre-trained Transformer model and the graph neural network to obtain the medical visit representation. Among them, the Transformer model is used to capture the semantic relationship between medical entities, and the graph neural network is used to model the structured information between medical entities;

[0020] 2) According to the medical visit representations obtained in step 1), input the representation sequences of all historical medical visits into a temporal model with a gated recurrent unit to model the temporality between medical visit records, thereby capturing the temporal dependencies in historical medical visit records and obtaining historical medical visit representations. Through the recursive calculation of the temporal model with a gated recurrent unit, a hidden state sequence at each time step is obtained, where the hidden state at each time step accumulates all historical information from the first step to the current time step;

[0021] The attention fusion module uses a two-layer attention mechanism to fuse different entities in the same medical visit and information from different medical visits of the same patient, including the following steps:

[0022] 1) Weightedly fuse the information of different entities through the first-layer attention mechanism to measure the importance of medical entities in the same medical visit and generate an overall representation of a single medical visit;

[0023] 2) Weightedly fuse the medical visit records at different time steps through the second-layer attention mechanism to measure the importance of different medical visit records and generate an overall representation of historical medical visits. Specifically: calculate the matching scores between the hidden state at each time step and the representation of the current medical visit record to generate corresponding attention weights.

[0024] Furthermore, the current medical visit module includes the following steps:

[0025] 1) Extract the set of symptom entities of the last medical visit from the electronic health record data. The set of symptom entities is denoted as where pi represents the patient number, t represents the t-th medical visit of the patient, and e s1 ,e s2 ,...,e sp represent the symptom entities extracted from the discharge record of this medical visit, and construct a knowledge graph subgraph G kg , and encode the knowledge graph subgraph G kg through a graph neural network to obtain the graph representation of each symptom entity

[0026] 2) Input the graph representation of each symptom entity obtained in step 1) into a pre-trained Transformer model. The Transformer model is used to encode a number of symptom entities in a medical visit. The specific process is as follows: Assume that the symptom entity sequence of a certain medical visit visit pi,t is {e s1 ,e s2 ,...,e sp}, and its corresponding graph representation sequence is Encode the graph of the symptom entity through the pre-trained Transformer

[0027] The characterization sequences are fused and encoded to generate an overall characterization of the symptom entity sequence for the current medical visit. In this process, the self-attention mechanism of the Transformer model is used to capture the global dependencies between symptom entities, thereby fusing the information of each symptom entity and generating a more comprehensive characterization of the symptom entity sequence.

[0028] 3) Concatenate the historical medical visit characterization with the overall characterization of the symptom entity sequence for the current medical visit to obtain a new current medical visit characterization that incorporates symptom information, which is used to assist in disease diagnosis.

[0029] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0030] The present invention reduces the dependence on past historical medical visit records. After considering the symptom information of the current medical visit, it realizes a more accurate effect of disease-assisted diagnosis. At the same time, it fully explores the relationships between different medical entities and uses the pre-training module to improve the quality of the encoder, thereby obtaining better entity characterizations and providing good support for subsequent downstream tasks. Moreover, it focuses on mining similar medical visit information as a reference for the future development of diseases, which conforms to the common practice of doctors in disease diagnosis and also provides a certain degree of interpretability for the results of disease-assisted diagnosis, contributing to the development and application of intelligent medical systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a schematic flowchart of the system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the embodiments of the present invention are not limited thereto.

[0033] As Figure 1 shown, this embodiment discloses a disease-assisted diagnosis system based on multi-source information fusion, including:

[0034] A pre-training module for extracting medical entity data, constructing a subgraph by combining external and local knowledge, and completing pre-training using single-visit data to enhance the model's ability to encode medical entities.

[0035] A temporal module for modeling the temporal relationship between multiple medical visits of a patient, and adopting a two-layer attention network to measure the importance of different medical visits and different medical entities to model the disease development status.

[0036] A similar medical visit module for mining similar medical visit groups for each medical visit of each patient under multiple measurement criteria, and providing a reference for the disease diagnosis of the patient by combining different types of medical visit groups.

[0037] The current visit module is used to extract data of the patient's current visit. After being encoded by a pre-trained model, it provides direct evidence for reference in disease diagnosis.

[0038] Specifically, the similar visit module includes a similar disease condition module, a similar symptom module, and a cross-screening module, where:

[0039] The similar disease condition module extracts the disease condition information of each patient's each visit from local medical and health data, performs similarity metric calculations, selects a group of visits that are most similar to it, called similar visits. At the same time, it extracts the next visit of the similar visit as a reference for the disease development situation at the next visit of the current visit;

[0040] The similar symptom module extracts the symptom information of each patient's current visit from local medical and health data, performs similarity metric calculations, selects a group of visits that are most similar to it, called similar visits. At the same time, it extracts the disease condition information of the similar visit as a reference for the current disease diagnosis;

[0041] The cross-screening module extracts the similar visits extracted by the similar disease condition module and the similar symptom module respectively for the current visit, and screens out a batch of visit information that is most valuable for reference to the current visit by taking the intersection.

[0042] Specifically, the pre-training module includes the following steps:

[0043] 1) Extract three types of medical entities from electronic health record data, including: disease entities, drug entities, and symptom entities, and construct three subgraphs respectively, including a knowledge graph subgraph, a co-occurrence graph subgraph, and a hierarchical structure graph subgraph;

[0044] 2) Use the graph neural network GNN to encode the three subgraphs constructed in step 1) to obtain corresponding graph representations. The representations of the three subgraphs are fused into a medical entity representation by weighted summation. The obtained representation will be used as the basis for subsequent visit representation calculation. Taking one visit as a unit, the representation of the medical entity obtained is input into the Transformer model to obtain the visit representation;

[0045] 3) Make full use of single-visit data as training data and perform two types of reconstruction tasks. The reconstruction tasks include: concatenating the representation of the symptom entity and the representation of the drug entity to predict the disease entity; concatenating the representation of the symptom entity and the representation of the disease entity to predict the disease entity. During the pre-training process, use the binary cross-entropy loss function L BCE Calculate the error between the predicted disease entity and the true disease entity.

[0046] The timing module includes a timing information extraction module and an attention fusion module, where:

[0047] The timing information extraction module is used to model the time interval information between different patient visits, and uses an RNN module to integrate a series of representations within the visit sequence into an overall representation, including the following steps:

[0048] 1) Encode the medical entities in each visit record, and encode and splice the medical entities through a pre-trained Transformer model and a graph neural network to obtain a visit representation. Among them, the Transformer model is used to capture the semantic relationship between medical entities, and the graph neural network is used to model the structured information between medical entities;

[0049] 2) According to the visit representation obtained in step 1), input the representation sequence of all historical visit records into a timing model with a gated recurrent unit to model the temporality between visit records, so as to capture the time-dependent relationship in historical visit records and obtain a historical visit representation. Through the recursive calculation of the timing model with a gated recurrent unit, a hidden state sequence at each time step is obtained, where the hidden state at each time step accumulates all historical information from the first step to the current time step;

[0050] The attention fusion module uses a two-layer attention mechanism to fuse different entities in the same visit and information from different visits of the same patient, including the following steps:

[0051] 1) Weightedly fuse the information of different entities through the first-layer attention mechanism to measure the importance of medical entities in the same visit and generate an overall representation of a single visit;

[0052] 2) Weightedly fuse the visit records at different time steps through the second-layer attention mechanism to measure the importance of different visit records and generate an overall representation of historical visits. Specifically: calculate the matching score between the hidden state at each time step and the representation of the current visit record to generate corresponding attention weights.

[0053] The current visit module includes the following steps:

[0054] 1) Extract the set of symptom entities of the last visit from the electronic health record data, and denote the set of symptom entities as where pi represents the patient number, t represents the t-th visit of the patient, and e s1 ,e s2 ,...,e sp represents the symptom entity extracted from the discharge record of this visit, and construct a knowledge graph subgraph G kg , and perform graph neural network on the knowledge graph subgraph G kgEncode to obtain the graphical representation of each symptom entity

[0055] 2) Input the graphical representation of each symptom entity obtained in step 1) into a pre-trained Transformer model, where the Transformer model is used to encode a number of symptom entities in a single medical visit. The specific process is as follows: Assume a medical visit visit pi,t with a symptom entity sequence of {e s1 , e s2 ,..., e sp}, and its corresponding graphical representation sequence is Fuse and encode the graphical representation sequence of this symptom entity through the pre-trained Transformer to generate the overall representation of the symptom entity sequence for the current medical visit During this process, the self-attention mechanism of the Transformer model is used to capture the global dependencies between symptom entities, thereby fusing the information of each symptom entity to generate a more comprehensive representation of the symptom entity sequence;

[0056] 3) Concatenate the historical medical visit representation with the overall representation of the symptom entity sequence for the current medical visit to obtain a new current medical visit representation that incorporates symptom information for assisting in disease diagnosis.

[0057] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited by the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A disease assisted diagnosis system based on multi-source information fusion, characterized in that, Including: A pre-training module, which is used to extract medical entity data, construct a subgraph by combining external and local knowledge, and complete pre-training using single-visit data to enhance the model's ability to encode medical entities; A temporal module, which is used to model the temporal relationship between a patient's multiple visits, and adopts a two-layer attention network to measure the importance of different visits and different medical entities to model the disease development status; A similar visit module, which is used to mine the similar visit groups of each patient's each visit under multiple measurement criteria, and provide a reference for the patient's disease diagnosis by combining different types of visit groups; The current visit module, which is used to extract the data of the patient's current visit, and after being encoded by the pre-trained model, provide direct evidence for the reference of disease diagnosis.

2. The disease-assisted diagnosis system based on multi-source information fusion according to claim 1, wherein The similar visit module includes a similar disease status module, a similar symptom module, and a cross-screening module, where: The similar disease status module extracts the disease status information of each patient's each visit from the local medical and health data, performs similarity measurement calculations, and selects the most similar set of visit information, which is called a similar visit. At the same time, it extracts the next visit of this similar visit as a reference for the disease development situation at the next visit of the current visit; The similar symptom module extracts the symptom information of the patient's current visit from the local medical and health data, performs similarity measurement calculations, and selects the most similar set of visit information, which is called a similar visit. At the same time, it extracts the disease status information of this similar visit as a reference for the current disease diagnosis; The cross-screening module extracts the similar visits extracted by the similar disease status module and the similar symptom module respectively for the current visit, and screens out a batch of visit information with the most reference value for the current visit by taking the intersection.

3. The disease-assisted diagnosis system based on multi-source information fusion according to claim 1, characterized in that, The pre-training module includes the following steps: 1) Extract three types of medical entities from the electronic health record data, including: disease entities, drug entities, and symptom entities, and construct three types of subgraphs respectively, including a knowledge graph subgraph, a co-occurrence graph subgraph, and a hierarchical structure graph subgraph; 2) Use the graph neural network GNN to encode the three subgraphs constructed in step 1) to obtain the representations of the corresponding subgraphs, and fuse the representations of the three subgraphs into the representation of the medical entity by weighted summation. The obtained representation will be used as the basis for subsequent visit representation calculation. Taking one visit as a unit, input the obtained representation of the medical entity into the Transformer model to obtain the visit representation; 3) Make full use of the single-visit data as training data and perform two types of reconstruction tasks. The reconstruction tasks include: concatenating the representations of symptom entities and drug entities to predict disease entities; concatenating the representations of symptom entities and disease entities to predict disease entities. During the pre-training process, use the binary cross-entropy loss function L BCE Calculate the error between the predicted disease entities and the true disease entities.

4. The disease-assisted diagnosis system based on multi-source information fusion according to claim 1, characterized in that The temporal module includes a temporal information extraction module and an attention fusion module, where: The temporal information extraction module is used to model the time interval information between different visits of the patient, and uses the RNN module to integrate a series of representations in the visit sequence into one representation, including the following steps: 1) Encode the medical entities in each visit record, and encode and splice the medical entities through the pre-trained Transformer model and the graph neural network to obtain the visit representation. Among them, the Transformer model is used to capture the semantic relationship between medical entities, and the graph neural network is used to model the structured information between medical entities; 2) Based on the medical visit representations obtained in step 1), input the representation sequences of all historical medical visits into a temporal model with a gated recurrent unit to model the temporality between medical visit records, thereby capturing the temporal dependencies in the historical medical visit records and obtaining historical medical visit representations. Through the recursive calculation of the temporal model with a gated recurrent unit, obtain the hidden state sequence at each time step, where the hidden state at each time step accumulates all historical information from the first step to the current time step; The attention fusion module uses a two-layer attention mechanism to fuse different entities in the same medical visit and different medical visit information of the same patient, including the following steps: 1) Weightedly fuse the information of different entities through the first-layer attention mechanism to measure the importance of medical entities in the same medical visit and generate an overall representation of a single medical visit; 2) Weightedly fuse the medical visit records at different time steps through the second-layer attention mechanism to measure the importance of different medical visit records and generate an overall representation of historical medical visits. Specifically: calculate the matching scores between the hidden state at each time step and the representation of the current medical visit record, and generate corresponding attention weights.

5. The disease assisted diagnosis system based on multi-source information fusion according to claim 1, characterized in that, The current medical visit module includes the following steps: 1) Extract the set of symptom entities of the last visit from the electronic health record data, and denote the set of symptom entities as where pi represents the patient number, t represents the t-th visit of the patient, and e s1 , e s2 ,..., e sp represents the symptom entity extracted from the discharge record of this visit, and construct a subgraph G of the knowledge graph kg , and encode the subgraph G of the knowledge graph through a graph neural network kg to obtain the graph representation of each symptom entity i = 1, 2,..., p; 2) Input the chart representation of each symptom entity obtained in step 1) into a pre-trained Transformer model, which is used to encode several symptom entities in a single medical visit. The specific process is as follows: Assume a certain medical visit visit pi,t with a symptom entity sequence of {e s1 , e s2 ,..., e sp}, and its corresponding chart representation sequence is Fuse and encode the chart representation sequence of this symptom entity through the pre-trained Transformer to generate the overall representation of the symptom entity sequence for the current medical visit In this process, the self-attention mechanism of the Transformer model is used to capture the global dependencies between symptom entities, thereby fusing the information of each symptom entity to generate a more comprehensive representation of the symptom entity sequence; 3) Concatenate the historical medical visit representation with the overall representation of the symptom entity sequence of the current medical visit to obtain a new current medical visit representation that integrates symptom information, which is used to assist in disease diagnosis.

Citation Information

Cited By

  • Slow obstructive pulmonary disease patient screening system for respiratory medicine department

    CN120954687A

  • A system for screening patients with chronic obstructive pulmonary disease in pneumology

    CN120954687B