Clinical data analysis method and device, electronic equipment and storage medium
By using multimodal encoders and feature fusion technology, the challenge of integrating multimodal clinical data has been solved, improving diagnostic and treatment efficiency and analytical accuracy, and reducing the risk of medical errors.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ANHUI IFLYHEALTH CO LTD
- Filing Date
- 2025-12-03
- Publication Date
- 2026-04-28
AI Technical Summary
Existing clinical data processing technologies struggle to effectively integrate and utilize multimodal clinical data, resulting in information not being fully explored and utilized, increasing the risk of medical errors and reducing diagnostic and treatment efficiency.
The system encodes clinical data from multiple modalities using a multimodal encoder. It trains the encoder by utilizing the similarity between clinical features of samples from multiple modalities and the similarity between entity features of associated entity pairs. By incorporating knowledge from the medical domain, it achieves cross-modal feature fusion and decoding, and generates analysis results.
It enables the effective integration and utilization of multimodal clinical data, improves diagnostic and treatment efficiency, reduces the risk of medical errors, and ensures the accuracy and reliability of analysis results.
Smart Images

Figure CN121938536A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical information processing technology, and in particular to a clinical data analysis method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the development of medical informatization, doctors need to process large amounts of complex clinical data, including medical history, symptoms, and test results, in a short period of time. Processing massive amounts of data poses a huge challenge to doctors, not only increasing the risk of medical errors but also consuming a significant amount of their time and reducing overall diagnostic and treatment efficiency.
[0003] Furthermore, clinical data typically includes multiple modalities such as text, numerical data, and images. Current technologies such as natural language processing and computer vision are mostly designed for single modalities, and are significantly insufficient in effectively integrating and utilizing multimodal clinical data to support clinical decision-making. Summary of the Invention
[0004] This invention provides a clinical data analysis method, apparatus, electronic device, and storage medium to address the deficiency in related technologies of lacking an effective method for integrating multimodal clinical data.
[0005] This invention provides a clinical data analysis method, comprising: Acquire clinical data from patients across multiple modalities; Based on a multimodal encoder, clinical data from multiple modalities are encoded separately to obtain clinical features of multiple modalities. The multimodal encoder is trained based on the similarity between sample clinical features of multiple modalities and the similarity between entity features of associated entity pairs in sample clinical features of multiple modalities. The associated entity pairs are related entity pairs determined based on medical domain knowledge. By fusing the clinical features of the multiple modalities, a fused feature is obtained; The fusion features are decoded to obtain the analysis results of the clinical data.
[0006] According to a clinical data analysis method provided by the present invention, the training steps of the multi-modal encoder include: Based on a multimodal encoder, sample clinical data from multiple modalities are encoded separately to obtain sample clinical features from multiple modalities; Based on the similarity between the clinical features of the samples from the multiple modalities and the patient and time period classifications of the clinical data from the samples from the multiple modalities, the contrast loss is determined. The associated entity loss is determined based on the similarity between entity features of associated entity pairs in the clinical features of the samples from the multiple modalities. Based on the contrast loss and the associated entity loss, the encoders of the multiple modalities are subjected to parameter iteration.
[0007] According to a clinical data analysis method provided by the present invention, the step of iterating the parameters of the encoder of the multiple modalities based on the contrast loss and the associated entity loss includes: Obtain the clinical association conditions of the associated entity pairs, wherein the clinical association conditions are determined based on the medical domain knowledge; When the clinical data of the samples from the multiple modalities meet the clinical correlation conditions, the encoders of the multiple modalities are subjected to parameter iteration based on the contrast loss and the correlation entity loss.
[0008] According to a clinical data analysis method provided by the present invention, the sample clinical data of the multiple modalities are local data; The step of iterating the parameters of the encoder for the multiple modalities based on the contrast loss and the associated entity loss includes: Based on the contrast loss and the associated entity loss, local update parameters are determined and uploaded to the central server to trigger the central server to determine global update parameters based on the received local update parameters from various locations. The global update parameters returned by the central server are received, and the encoders of the multiple modalities are iterated based on the global update parameters.
[0009] According to a clinical data analysis method provided by the present invention, the fusion of clinical features from multiple modalities to obtain fused features includes: Based on the cross-modal similarity between the clinical features of the multiple modalities, cross-modal feature enhancement is performed on the clinical features of the multiple modalities to obtain the first enhanced features of the multiple modalities; Based on the element similarity of the first enhanced feature of each modality within the modality, intra-modal feature enhancement is performed on the first enhanced feature of each modality to obtain the second enhanced feature of each modality; The second enhancement features of the multiple modalities are fused to obtain the fused features.
[0010] According to a clinical data analysis method provided by the present invention, the acquisition of patient clinical data in multiple modalities includes: Acquire raw patient data across multiple modalities; Based on the data attributes of the original data, the credibility of the original data is determined. The data attributes include at least one of data quality, source authority, standardization degree, compliance attributes, and application relevance. The clinical data are determined based on the reliability of the original data.
[0011] According to a clinical data analysis method provided by the present invention, determining the reliability of the original data based on its data attributes includes: The credibility of the original data is determined based on its data attributes and the attribute weights of each data attribute in the application scenario of the original data.
[0012] According to a clinical data analysis method provided by the present invention, the step of performing feature decoding on the fused features to obtain the analysis results of the clinical data includes: The fused features are decoded to obtain the patient's medical condition summary information as the analysis result.
[0013] According to a clinical data analysis method provided by the present invention, the step of performing feature decoding on the fused features to obtain the analysis results of the clinical data further includes: Based on the disease summary information and / or the fusion features, the patient's first and second medical information are determined. The first medical information is obtained based on a clinical guideline rule base, and the second medical information is obtained based on a medical recommendation model, which is obtained based on reinforcement learning. Based on the first and second diagnostic information, the patient's diagnostic information is determined as the analysis result.
[0014] The present invention also provides a clinical data analysis device, comprising: The data acquisition unit is used to acquire clinical data from patients across multiple modalities. The feature encoding unit is used to encode clinical data from multiple modalities using a multi-modal encoder to obtain clinical features from multiple modalities. The multi-modal encoder is trained based on the similarity between entity features of associated entity pairs in the sample clinical features of sample clinical data from multiple modalities. The associated entity pairs are entity pairs that are associated in different modalities, as determined by knowledge in the medical domain. A feature fusion unit is used to fuse the clinical features of the multiple modalities to obtain fused features; The feature decoding unit is used to perform feature decoding on the fused features to obtain the analysis results of the clinical data.
[0015] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the clinical data analysis method as described above.
[0016] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the clinical data analysis method as described above.
[0017] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the clinical data analysis method as described above.
[0018] The clinical data analysis method, apparatus, electronic device, and storage medium provided by this invention train a multi-modal encoder based on the similarity between clinical features of samples from multiple modalities and the similarity between entity features of associated entity pairs within the clinical features of samples from multiple modalities. This transforms clinical data from multiple modalities into a shared feature space while introducing medical domain knowledge during the data encoding process to enhance the professionalism and accuracy of the semantic representation of clinical features. Based on this, feature fusion and feature decoding are performed, enabling cross-modal clinical data analysis. This ensures the accuracy and reliability of clinical data analysis while maintaining the utilization rate of multimodal information. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in this invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is one of the flowcharts of the clinical data analysis method provided by the present invention.
[0021] Figure 2 This is a flowchart illustrating the diagnostic information output method provided by the present invention.
[0022] Figure 3 This is the second flowchart of the clinical data analysis method provided by the present invention.
[0023] Figure 4 This is a schematic diagram of the structure of the clinical data analysis device provided by the present invention.
[0024] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0026] With the development of medical informatization, clinical data is constantly accumulating. Doctors need to handle a large amount of complex clinical data in a short period of time, which not only places extremely high demands on their professional experience and energy, but also increases the risk of medical errors due to information overload or human negligence. At the same time, doctors spend a lot of time collecting and organizing information, leading to a decrease in the efficiency of doctor-patient communication and overall diagnosis and treatment.
[0027] Furthermore, current clinical data typically exhibits significant multimodal characteristics, including not only unstructured text data such as medical records and treatment notes, but also structured numerical data such as laboratory indicators, and often imaging data such as CT (Computed Tomography) and MRI (Magnetic Resonance Imaging). Existing technologies usually process data on a single modality, often making it difficult to uniformly and efficiently process data from diverse sources and with varying structures, resulting in a large amount of valuable information not being fully explored and utilized.
[0028] How to effectively integrate and utilize multimodal clinical data, and then achieve cross-modal collaborative analysis to support clinical decision-making, is a technical challenge that urgently needs to be solved in this field.
[0029] Figure 1 This is one of the flowcharts illustrating the clinical data analysis method provided by the present invention, such as... Figure 1 As shown, the method includes: Step 110: Obtain the patient's clinical data across multiple modalities.
[0030] Specifically, for any given patient, clinical data across multiple modalities can be obtained. That is, the clinical data obtained here from multiple modalities belongs to the same patient; for example, clinical data from multiple modalities can be identified by the same patient ID (IdentityDocument).
[0031] The multiple modalities referred to herein may include at least two of the following: text modality, numerical modality, and image modality. Specifically, text modality clinical data may be medical records, surgical records, discharge summaries, etc., or transcribed text obtained by transcribing a user's spoken description of their discomfort. Numerical modality clinical data may be laboratory test reports, such as blood routine tests and biochemical indicators, or patient information, such as the patient's age, height, weight, body temperature, heart rate, and blood pressure. Image modality clinical data may be visual images acquired through various medical imaging devices, such as CT images, MRI images, and ultrasound images; however, this embodiment of the invention does not specifically limit the types of images acquired.
[0032] In some embodiments, after obtaining clinical data from multiple modalities, the clinical data from each modality can be mapped to entities in a knowledge graph using entity linking technology. This mapping of knowledge graph entities can be implemented based on a medical domain knowledge graph. Examples of medical domain knowledge graphs include Hetionet and UMLS. A medical domain knowledge graph can contain entities such as diseases, symptoms, drugs, tests, and anatomical locations, along with their rich relationships. These relationships can include causes, treatments, associations, and examinations. The medical domain knowledge graph is shareable, does not involve any patient privacy, and only provides public medical knowledge. Furthermore, the medical domain knowledge graph can be dynamically updated and locally adapted, thereby avoiding knowledge obsolescence or alignment failures.
[0033] By using a medical knowledge graph, entity links can be established for clinical data across different modalities. This enables terminology standardization across modalities, facilitating entity alignment when encoding clinical data from multiple modalities separately. For example, "WBC" in numerical modality clinical data can be mapped to the "white blood cell count" entity in the knowledge graph; "fever" in text modality clinical data can be mapped to the "fever" entity. Through entity links, clinical data can be represented as a set of entities within the knowledge graph, incorporating attributes such as time and numerical values.
[0034] Step 120: Based on a multi-modal encoder, the clinical data of multiple modalities are encoded to obtain clinical features of multiple modalities. The multi-modal encoder is trained based on the similarity between sample clinical features of multiple modalities and the similarity between entity features of associated entity pairs in the sample clinical features of multiple modalities. The associated entity pairs are related entity pairs determined based on medical domain knowledge.
[0035] Specifically, for each modality, an encoder for that modality can be pre-trained. The resulting encoder is used to convert clinical data in that modality into clinical features. Furthermore, each modality's encoder is dedicated to converting clinical data in its own modality into clinical features in a shared feature space, thereby achieving alignment of clinical data from different modalities in the feature space.
[0036] For example, there are encoders for text modalities, encoders for numerical modalities, and encoders for image modalities. The encoders for text modalities, numerical modalities, and image modalities are used to encode features of clinical data in text modalities, numerical modalities, and image modalities, respectively, thereby obtaining and outputting the clinical features of text modalities, numerical modalities, and image modalities.
[0037] For example, a text modality encoder can encode data from a patient's chief complaint to obtain clinical features; a numerical modality encoder can encode data from entity sequences obtained by mapping test reports to knowledge graph entities using entity linking technology to obtain clinical features; and an image modality encoder can encode data from a patient's CT images, MRI images, etc., to obtain clinical features. Here, the image modality encoder can be a pre-trained model based on ResNet (Residual Network) or ViT (Vision Transformer).
[0038] Furthermore, the clinical features of the text modality, numerical modality, and image modality are all clinical features under a shared feature space. Under the shared feature space, the clinical features of the text modality, numerical modality, and image modality all have consistent semantic representations.
[0039] In other words, by using encoders of multiple modalities to encode features of clinical data from multiple modalities, the clinical data from multiple modalities can be forcibly projected into the same shared feature space. This makes the clinical features of clinical data from different modalities semantically consistent in feature representation, and thus allows for the fusion of clinical features from multiple modalities to achieve clinical data analysis. This greatly enhances the information utilization rate of clinical data analysis and ensures the accuracy and reliability of clinical data analysis.
[0040] In order to enable encoders of multiple modalities to project clinical data of their respective modalities into a shared feature space, the similarity between sample clinical features of multiple modalities and the similarity between entity features of associated entity pairs in the sample clinical features of multiple modalities can be applied when training encoders of multiple modalities.
[0041] In this context, associated entity pairs refer to combinations of entities that are related in the medical field. For example, the entities "fever" and "white blood cell count" are highly correlated in the medical field, and "fever" and "white blood cell count" can be considered as associated entity pairs. Similarly, the entities "ground-glass opacity in the lungs" and "early-stage lung cancer" are highly correlated in the medical field, and "ground-glass opacity in the lungs" and "early-stage lung cancer" can be considered as associated entity pairs.
[0042] Related entity pairs can be obtained through medical domain knowledge analysis, for example, two connected nodes in a medical domain knowledge graph can be considered as a pair of related entity pairs.
[0043] Furthermore, encoders for multiple modalities can be trained based on the idea of contrastive learning.
[0044] Specifically, a set of sample clinical data from multiple modalities can be collected and input into the encoder of the corresponding modality, thereby obtaining the clinical features of the sample clinical data of each modality output by the encoder of each modality, which are referred to here as sample clinical features.
[0045] After obtaining the clinical features of samples from each modality, the similarity between clinical features from different modalities can be calculated. For cases where clinical data from two modalities belong to the same patient and represent the same patient's clinical information from different modalities, the greater the similarity between the clinical features of the two modalities, the closer their semantic representations are in the shared feature space, and the smaller the loss of the encoders in feature encoding. Conversely, the smaller the similarity between the clinical features of the two modalities, the more different their semantic representations are in the shared feature space, and the greater the loss of the encoders in feature encoding. Similarly, for cases where clinical data from two modalities belong to different patients and represent different clinical information from different modalities, the greater the similarity between the clinical features of the two modalities, the closer their semantic representations are in the shared feature space, and the greater the loss of the encoders in feature encoding. Conversely, the smaller the similarity between the clinical features of the two modalities, the more different their semantic representations are in the shared feature space, and the smaller the loss of the encoders in feature encoding. Therefore, the loss value of the encoder for different modalities can be determined based on the similarity between the clinical features of samples from different modalities, and then the parameters of the encoder for different modalities can be iterated.
[0046] Furthermore, for cases where there are associated entity pairs in a set of multimodal sample clinical data, the similarity between entity features of associated entity pairs in the sample clinical features can be calculated.
[0047] For example, in a set of multimodal clinical sample data, the text modality sample clinical data contains "fever" and the numerical modality sample clinical data contains "white blood cell count". "Fever" and "white blood cell count" constitute an associated entity pair. The similarity between the entity feature of "fever" in the text modality sample clinical data and the entity feature of "white blood cell count" in the numerical modality sample clinical data can be calculated.
[0048] For example, in a set of multimodal clinical sample data, the text modality sample clinical data contains "early lung cancer" and the image modality sample clinical data contains "ground-glass opacity in the lungs". "Early lung cancer" and "ground-glass opacity in the lungs" constitute an associated entity pair. The similarity between the entity features of "early lung cancer" in the text modality sample clinical features and the entity features of "ground-glass opacity in the lungs" in the image modality sample clinical features can be calculated.
[0049] Understandably, the higher the similarity between these two modalities, the more consistent the semantic representations of the entity representations of two highly related entities in the shared feature space of the clinical features obtained from feature encoding in different modalities, and the smaller the loss of the encoders in feature encoding. Conversely, the lower the similarity between these two modalities, the more different the semantic representations of the entity representations of two highly related entities in the shared feature space of the clinical features obtained from feature encoding in different modalities, and the greater the loss of the encoders in feature encoding. Therefore, the loss value of the encoders in different modalities can be determined based on the similarity between the entity features of related entity pairs in the clinical features of samples from different modalities, and then the parameters of the encoders in different modalities can be iterated.
[0050] In this embodiment of the invention, in addition to using the similarity between sample clinical features from multiple modalities to train the multi-modal encoder, the similarity between entity features of associated entity pairs within the sample clinical features from multiple modalities is also used to train the multi-modal encoder. The application of the similarity between sample clinical features from multiple modalities can continuously improve the semantic consistency of the clinical features encoded by encoders from different modalities during training. The application of the similarity between entity features of associated entity pairs not only improves the semantic consistency of the clinical features encoded by encoders from different modalities at the entity granularity level, but also introduces medical domain knowledge into the encoder training process. This allows entities that are related in the medical domain, even if they belong to different modalities, to exhibit semantic correlation in the encoded clinical features. The resulting multi-modal encoder can project clinical data from multiple modalities into the same shared feature space and exhibits superior semantic representation capabilities for medical domain-related entities in the clinical data.
[0051] Step 130: Fuse the clinical features of the multiple modalities to obtain the fused features.
[0052] Specifically, after obtaining the clinical features of each modality, the clinical features of each modality can be fused to obtain a fused feature that incorporates the clinical features of all modalities. Here, the feature fusion method can be feature concatenation, weighted summation, weighted fusion based on attention mechanisms, etc., and the embodiments of the present invention do not specifically limit this.
[0053] Step 140: Decode the fused features to obtain the analysis results of the clinical data.
[0054] Specifically, after obtaining the fusion features, feature decoding can be performed based on the fusion features, thereby enabling data analysis of clinical data across multiple modalities. Here, the analysis result can be the result obtained from feature decoding, specifically, it can be a summary text of the patient's condition generated from clinical data across multiple modalities, or it can be a text of treatment suggestions for the patient generated from clinical data across multiple modalities. This embodiment of the invention does not specifically limit this.
[0055] In the method provided in this invention, a multi-modal encoder is trained based on the similarity between clinical features of samples from multiple modalities and the similarity between entity features of associated entity pairs within the clinical features of samples from multiple modalities. This transforms clinical data from multiple modalities into a shared feature space while introducing medical domain knowledge during data encoding to enhance the professionalism and accuracy of the semantic representation of clinical features. Based on this, feature fusion and feature decoding are performed, enabling cross-modal clinical data analysis. This ensures the accuracy and reliability of clinical data analysis while maintaining the utilization rate of multi-modal information.
[0056] In the method provided in this invention, data analysis based on multiple modalities of clinical data can quickly organize the patient's scattered clinical data into comprehensive and concise information for doctors. This allows doctors to quickly obtain comprehensive patient information, greatly improving diagnostic efficiency. Furthermore, considering the subjective nature of patients' verbal descriptions during consultations, which may omit key medical history (such as drug allergies, chronic disease medications, etc.), and the possibility that doctors, due to insufficient time or experience, may overlook potentially related diseases (such as the risk of retinopathy in diabetic patients), data analysis based on multiple modalities of clinical data can ensure the completeness and comprehensiveness of the analyzed data. Introducing medical knowledge during the data analysis process can further avoid information omissions, thereby providing doctors with more comprehensive and reliable diagnostic and treatment reference information and reducing the risk of misdiagnosis.
[0057] Based on the above embodiments, the training steps of the encoder for the multiple modalities include: Based on a multimodal encoder, sample clinical data from multiple modalities are encoded separately to obtain sample clinical features from multiple modalities; Based on the similarity between the clinical features of the samples from the multiple modalities and the patient and time period classifications of the clinical data from the samples from the multiple modalities, the contrast loss is determined. The associated entity loss is determined based on the similarity between entity features of associated entity pairs in the clinical features of the samples from the multiple modalities. Based on the contrast loss and the associated entity loss, the encoders of the multiple modalities are subjected to parameter iteration.
[0058] Specifically, for encoders with multiple modalities, comparative learning can be performed in the following way: First, positive and negative samples needed for training can be collected.
[0059] Positive samples are clinical data from multiple modalities of the same patient within the same time period. This means that the multiple modalities of clinical data in the positive samples belong to the same patient and the same time period. Clinical data from multiple modalities belonging to the same patient and the same time period can reflect the same disease, symptoms, and other information from different modalities.
[0060] A negative sample is a set of multimodal clinical data from different patients within the same time period, or a set of multimodal clinical data from the same patient within different time periods, or a set of multimodal clinical data from different patients within different time periods. Essentially, the multimodal clinical data in a negative sample represents different patient types and / or different time periods. Multiple modalities of clinical data belonging to different patients and / or different time periods can reflect different diseases, symptoms, and other information from different modalities.
[0061] It is understandable that both positive and negative samples can be represented as clinical data from multiple modalities.
[0062] Clinical data from multiple modalities can be input into the encoder of each modality, thereby obtaining the corresponding modal clinical features output by the encoder of each modality.
[0063] After obtaining the clinical features of samples from multiple modalities, the similarity between these features can be calculated. Furthermore, based on the patient and time period affiliations of the clinical data from multiple modalities, it can be determined whether the clinical data from these modalities are positive or negative samples. Based on this, a contrastive loss can be calculated using the similarity between the clinical features of the samples from multiple modalities and whether the clinical data from these modalities are positive or negative. Here, the contrastive loss is used to narrow the distance between clinical features of samples from different modalities that are positive samples, and to widen the distance between clinical features of samples from different modalities that are negative samples, within the shared feature space.
[0064] Furthermore, when the clinical data from multiple modalities are positive samples, the greater the similarity between the clinical features of the samples from multiple modalities, the smaller the contrast loss; conversely, the smaller the similarity between the clinical features of the samples from multiple modalities, the larger the contrast loss. Similarly, when the clinical data from multiple modalities are negative samples, the greater the similarity between the clinical features of the samples from multiple modalities, the larger the contrast loss; and the smaller the similarity between the clinical features of the samples from multiple modalities, the smaller the contrast loss. Here, the contrast loss can be calculated based on a loss function such as InfoNCE (Information Noise Contrastive Estimation).
[0065] Furthermore, the similarity between entity features of associated entity pairs in the clinical features of samples from multiple modalities can be calculated, and the associated entity loss can be determined accordingly. Here, the associated entity loss is used to bring together the feature representations of entities belonging to different modalities but related in medical domain knowledge within a shared feature space.
[0066] Furthermore, regardless of whether the clinical data samples from multiple modalities are positive or negative, the greater the similarity between the entity features of associated entity pairs in the clinical features of samples from different modalities, the smaller the associated entity loss; conversely, the smaller the similarity between the entity features of associated entity pairs in the clinical features of samples from different modalities, the greater the associated entity loss. Here, the associated entity loss can be calculated based on the regularization term of knowledge distillation or the constraint term of relation awareness.
[0067] After obtaining the contrastive loss and associated entity loss, the parameters of the encoders for multiple modalities can be iterated, thereby enabling the training of encoders for multiple modalities. In this process, the contrastive loss guides the encoders of multiple modalities to bring together the clinical features of clinical data representing the same disease, symptoms, etc., during data encoding. The associated entity loss guides the encoders of multiple modalities to ignore whether clinical data from different modalities reflect the same disease or symptoms, bringing together the feature representations of entities associated in medical domain knowledge, and penalizing entities that are associated in medical domain knowledge but far apart in the shared vector space.
[0068] The multimodal encoders trained in this way can not only project clinical data from multiple modalities into the same shared feature space, but also exhibit superior semantic representation capabilities for medical-related entities in the clinical data.
[0069] Based on any of the above embodiments, in the training step of the encoder for the multiple modalities, the step of iterating the parameters of the encoder for the multiple modalities based on the contrastive loss and the associated entity loss includes: Obtain the clinical association conditions of the associated entity pairs, wherein the clinical association conditions are determined based on the medical domain knowledge; When the clinical data of the samples from the multiple modalities meet the clinical correlation conditions, the encoders of the multiple modalities are subjected to parameter iteration based on the contrast loss and the correlation entity loss.
[0070] Specifically, when related entity pairs exist in clinical data across multiple modalities, the clinical association conditions for these pairs can be obtained. Here, for any related entity pair, the clinical association condition reflects the preconditions under which the two entities in the pair are associated; it can also be understood as the pair being considered valid only if the clinical association conditions are met. For example, for the related entity pair "fever" and "white blood cell count," the clinical association condition is "bacterial infection," meaning that under the condition of "bacterial infection," "fever" and "white blood cell count" are associated.
[0071] Here, the acquisition of clinical association conditions for associated entity pairs can be based on medical domain knowledge. For example, for two connected nodes in a medical domain knowledge graph, if there is a conditional relationship on the edge connecting the two nodes, the two connected nodes can be regarded as associated entity pairs, and the conditional relationship can be regarded as clinical association conditions.
[0072] When iterating parameters for an encoder across multiple modalities, it can be determined whether sample clinical data from multiple modalities satisfy the clinical association conditions for associated entity pairs. Specific determination methods could include whether the sample clinical data from the text modality contains segments that are close to or consistent with the clinical association conditions, or whether inputting sample clinical data from one modality into the data analysis model of that modality yields a conclusion consistent with the clinical association conditions. This embodiment of the invention does not impose specific limitations on these methods. For example, information such as patient age, underlying diseases, and onset time in the sample clinical data can be used to infer whether the clinical association conditions are met.
[0073] It is understandable that if the clinical data of multiple modalities satisfy the clinical association conditions of associated entity pairs, it means that the association relationship of associated entity pairs in the clinical data of multiple modalities is valid in clinical medical logic. Therefore, when iterating parameters for encoders of multiple modalities, not only contrastive loss but also associated entity loss can be applied.
[0074] If the clinical data samples from multiple modalities do not meet the clinical association conditions for associated entity pairs, it indicates that the association relationship between the associated entity pairs in the clinical data samples from multiple modalities is not valid in clinical medical logic. Therefore, when iterating parameters for the encoder of multiple modalities, only contrastive loss is applied, and associated entity loss is not applied.
[0075] For example, when multiple modalities of clinical data meet the "bacterial infection" condition, in addition to applying contrast loss, the associated entity loss generated for "fever" and "white blood cell count" is also applied when iterating parameters for the encoder of multiple modalities.
[0076] In the method provided in this embodiment of the invention, clinical association conditions for associated entity pairs are introduced to control the application of associated entity loss when iterating parameters of encoders of multiple modalities. This makes the training of encoders of multiple modalities more in line with clinical medical logic, thereby optimizing the semantic reasoning ability of the encoder.
[0077] Based on any of the above embodiments, the sample clinical data of the multiple modalities are local data; In the training steps of the encoders for the multiple modalities, the step of iterating the parameters of the encoders for the multiple modalities based on the contrastive loss and the associated entity loss includes: Based on the contrast loss and the associated entity loss, local update parameters are determined and uploaded to the central server to trigger the central server to determine global update parameters based on the received local update parameters from various locations. The global update parameters returned by the central server are received, and the encoders of the multiple modalities are iterated based on the global update parameters.
[0078] Specifically, the clinical data analysis methods can be implemented using physician-side devices in various hospitals. Since the clinical data samples from each hospital involve patient privacy, a federated learning approach is adopted when training encoders for multiple modalities to avoid privacy leaks.
[0079] That is, for a hospital as the implementing entity, the sample clinical data of multiple modalities used for encoder training is the hospital's local data, stored in the hospital's local memory, and is not transmitted to other hospitals or institutions.
[0080] In the process of federated learning, hospitals can be included as participants, thereby leveraging the characteristics of clinical data from different hospitals to conduct joint training. The algorithms used can be FedProx, FedOpt, etc.
[0081] For each hospital's implementing entity, during training, local data and locally deployed multi-modal encoders can be used for data encoding to obtain clinical features of samples from multiple modalities, thereby determining the contrast loss and entity association loss. After obtaining the contrast loss and entity association loss, local update parameters can be calculated based on them. These local update parameters can be understood as the model update parameters calculated by the participants based on local data during the federated learning process; for example, they could be gradients or updated model parameters. After obtaining the local update parameters, each hospital's implementing entity can send them to the central server. Furthermore, when sending the local update parameters, they can be encrypted, for example, using the Paillier homomorphic encryption method, or adding Gaussian noise to the local update parameters based on model differential privacy methods; this embodiment of the invention does not specifically limit this approach.
[0082] Accordingly, the central server can receive local update parameters sent by the execution entities of multiple hospitals. Therefore, the central server can perform secure aggregation based on the local update parameters of each hospital, thereby summarizing the global update parameters. This secure aggregation can be implemented using algorithms such as FedAvg. The global update parameters here can be understood as the model update parameters that reflect the global situation in federated learning, obtained by the central server summarizing the local update parameters uploaded by each participant. Afterwards, the central server can send the global update parameters to the execution entity of each hospital.
[0083] Therefore, each hospital's implementing entity can receive global update parameters issued by the central server, and perform parameter iteration on the encoders of multiple modalities in each hospital based on the global update parameters, thereby realizing horizontal federated learning for encoders of multiple modalities.
[0084] In the method provided in the embodiments of the present invention, by performing federated learning on encoders of multiple modalities, the dual challenges of data silos and uneven quality of private sample clinical data of various hospitals can be solved without disclosing the private sample clinical data of each hospital.
[0085] Based on any of the above embodiments, in step 130, fusing the clinical features of the multiple modalities to obtain the fused features includes: Based on the cross-modal similarity between the clinical features of the multiple modalities, cross-modal feature enhancement is performed on the clinical features of the multiple modalities to obtain the first enhanced features of the multiple modalities; Based on the element similarity of the first enhanced feature of each modality within the modality, intra-modal feature enhancement is performed on the first enhanced feature of each modality to obtain the second enhanced feature of each modality; The second enhancement features of the multiple modalities are fused to obtain the fused features.
[0086] Specifically, after obtaining clinical features from multiple modalities, feature fusion of these multimodal clinical features can be achieved based on a hierarchical attention mechanism. Here, hierarchical attention specifically includes two layers of attention mechanisms: the first is a cross-modal attention mechanism, and the second is an intra-modal attention mechanism. It can be understood that the attention mechanism can mimic the human cognitive ability of "selective attention," thereby dynamically assigning different weights to information during processing, focusing on key parts of the information and ignoring irrelevant content. For example, when doctors review laboratory reports, they focus on abnormal indicators (such as "elevated blood sugar") while ignoring normal values.
[0087] In this process, a cross-modal attention mechanism can first be implemented. Specifically, the similarity between clinical features of different modalities can be calculated, i.e., cross-modal similarity can be obtained. Based on the cross-modal similarity, cross-modal attention transformation can be performed on the clinical features of different modalities, thereby achieving cross-modal feature enhancement of clinical features of different modalities. Here, the clinical features obtained after cross-modal feature enhancement are denoted as the first enhanced feature. The cross-modal feature enhancement referred to here can enhance the parts of the clinical features with high cross-modal similarity and weaken the parts with low cross-modal similarity. The resulting first enhanced feature better highlights the common information expressed by each modality.
[0088] For example, during the execution of cross-modal attention, cross-modal similarity can be calculated between clinical features in the text modality and clinical features in the numerical modality. The cross-modal similarity between the clinical features "cough" and "chest pain" in the text modality and the clinical features "history of pneumonia", "smoking history" and "chest CT" in the numerical modality is relatively high, thereby enhancing the feature representation of "history of pneumonia", "smoking history" and "chest CT" in the numerical modality.
[0089] Secondly, an intramodal attention mechanism can be implemented. Specifically, for each modality, the element similarity between elements within the first enhanced feature of the modality can be calculated. Based on the element similarity, intramodal attention transformation is performed on the first enhanced feature, thereby achieving feature enhancement of clinical features within the modality. Here, the clinical features obtained after intramodal feature enhancement are denoted as the second enhanced feature. The intramodal feature enhancement referred to here can be to enhance the parts with high element similarity and weaken the parts with low element similarity within the first enhanced feature. The resulting second enhanced feature more highlights the key information expressed by a single modality.
[0090] For example, during the execution of in-modal attention, the first enhancement feature for the text modality, which contains elements with high element similarity between "cough" and "chest pain", can enhance the feature representation of "cough" and "chest pain".
[0091] For example, during the attention process within a modality, for the first enhancement feature of the patient's complaint in the text modality, a query, key, and value can be generated for each token. The query represents the token's "inquiry intent," the key represents the token's "identity identifier," and the value represents the token's "actual information." Attention weights are calculated by evaluating the element similarity between the query of each token and the keys of all tokens. Understandably, higher element similarity results in a larger attention weight; for example, the tokens "palpitation" and "heart palpitations" have high element similarity. The dimension of the key is used to adjust gradient stability. Based on this, the values of each token can be weighted and summed according to the attention weights to obtain the output for each token. It can be understood that the outputs of all tokens constitute the second enhancement feature of the patient's complaint. For example, when performing in-modal attention on the chief complaint "palpitations, blood pressure 160 / 100", the similarity between the query for "blood pressure" and the keys for "palpitations" and "blood pressure" can be calculated. This results in "blood pressure" having the highest element similarity, corresponding to an attention weight of 0.9, while "palpitations" has an attention weight of 0.1. Based on this, a second enhanced feature reflecting "blood pressure" can be output.
[0092] After obtaining the second augmentation features for each modality, the second augmentation features for each modality can be fused to obtain the fused features. Here, the fusion can be performed in an early fusion manner, such as directly concatenating the second augmentation features of each modality and inputting them into a DNN (Deep Neural Network), or in a late fusion manner, such as processing the temporal features of the second augmentation features of the text modality through an LSTM (Long Short-Term Memory network) and inputting the second augmentation features of the numerical modality into an MLP (Multilayer Perceptron) for feature processing before fusion. This embodiment of the invention does not specifically limit the specific method used.
[0093] In the method provided in this embodiment of the invention, a hierarchical attention mechanism enables multimodal clinical features to automatically learn the complex relationships between different modalities, thereby achieving the fusion of multimodal clinical features and enhancing the interpretability of clinical data analysis.
[0094] Based on any of the above embodiments, in step 110, acquiring the patient's clinical data in multiple modalities includes: Acquire raw patient data across multiple modalities; Based on the data attributes of the original data, the credibility of the original data is determined. The data attributes include at least one of data quality, source authority, standardization degree, compliance attributes, and application relevance. The clinical data are determined based on the reliability of the original data.
[0095] Specifically, the patient's raw data across multiple modalities can be health-related data obtained from various data sources. For example, it can be electronic medical records, test reports, imaging data, and prescription records obtained from hospital systems; heart rate and sleep data obtained from wearable devices; blood pressure data obtained from home smart blood pressure monitors; genetic testing reports obtained from external institutions; health questionnaires, dietary records, and symptom complaints provided by the patient themselves; and medical insurance records, vaccination history, and epidemiological database data obtained from public health systems.
[0096] Raw data from different sources, of different types, and in different modalities may reflect varying degrees of reliability. If low-reliability data is misused in clinical data analysis, it is highly likely to affect the reliability of the analysis results. To address this, this embodiment of the invention performs a reliability assessment on the raw data.
[0097] Specifically, credibility assessment can be achieved based on at least one of the following: data quality of the original data, source authority, degree of standardization, compliance attributes, and application relevance.
[0098] The quality of raw data is used to assess its accuracy, completeness, consistency, and timeliness. For example, biochemical indicators from hospital laboratories and imaging diagnostic reports generally have high data quality, while sleep data from consumer fitness trackers and questionnaires with significant data gaps generally have lower quality. Generally, higher-quality raw data has higher reliability. Data quality is a core indicator of data attributes, and in application scenarios such as acute diagnosis, data quality has a significant impact on reliability assessment.
[0099] The authority of the source of raw data is used to assess the professionalism of the source and the reliability of the equipment. For example, medical-grade monitoring equipment and EMR data entered by clinicians have higher source authority, while patient subjective reports and data from uncertified health apps have lower source authority. Source authority is a key indicator of data attributes; generally, the higher the source authority of the raw data, the higher its credibility.
[0100] The degree of standardization of raw data is used to assess the normalization of the data's structure and terminology. For example, diagnostic and laboratory data coded using ICD (International Classification of Diseases) / LOINC (Logical Observation Identifiers Names and Codes) have a higher degree of standardization, while physician's free text notes and patient's voice reports have a lower degree of standardization. Generally, the higher the degree of standardization of raw data, the higher its reliability.
[0101] The compliance attributes of raw data are used to assess its privacy protection, informed consent, and legal compliance. For example, research data that has undergone ethical review and strict de-identification possesses compliance attributes, while data from unknown sources or without authorization does not. Raw data with compliance attributes generally has higher credibility. Conversely, raw data lacking compliance attributes is considered unusable by default, and its credibility can be set to 0.
[0102] The application relevance of raw data is used to assess the degree of relevance of the raw data to the current application scenario. For example, heart rate data has a high relevance in the application scenario of heart failure management, blood glucose values have a high relevance in the application scenario of diabetes monitoring, and genetic data has a low relevance in the application scenario of acute trauma diagnosis. Generally, the higher the application relevance of raw data, the higher its credibility. If the raw data is not applicable to the current application scenario, then regardless of the data quality, source authority, or other data attributes, the credibility of the raw data is set to low.
[0103] All of the above data attributes can be used to assess the credibility of the original data. For example, based on pre-defined weights for each type of data attribute, a weighted sum can be calculated for each type of data attribute in the original data, and the result of the weighted sum can be used as the credibility of the original data. For example, the formula for calculating credibility could be: In the formula, For credibility; For the first i Scoring of each data attribute For the first i The weights of each data attribute are specified. The scoring criteria and weight ranges for each data attribute can be found in Table 1.
[0104] Table 1. Credibility Scoring Criteria
[0105] Furthermore, the reliability can be dynamically adjusted. For example, if home blood pressure measurement data is consistent with hospital measurement results over a long period, the reliability of home blood pressure measurement can be increased.
[0106] After completing the credibility assessment of the raw data, clinical data can be determined based on the credibility of the raw data. For example, raw data with a credibility below a preset threshold can be discarded, and only raw data with a credibility above or equal to the preset threshold can be used as clinical data; or, for example, credibility can be used as the weight of the raw data, and raw data of the same type can be weighted and merged. This embodiment of the invention does not specifically limit this.
[0107] Furthermore, raw data from different sources may possess different data attributes and correspond to varying degrees of reliability. These data sources can include hospital systems, approved home medical devices, consumer wearable devices, patient-provided data, and public health data. The data attributes of some of the raw data from these sources can be found in Table 2. Table 2. Examples of data attributes from various data sources
[0108] Before conducting data analysis on clinical data, preprocessing can be performed on the clinical data. For example, unit unification and data completion can be performed on the clinical data, and warnings can be issued for conflicting data in the clinical data. This embodiment of the invention does not specifically limit this.
[0109] Furthermore, during the process of determining clinical data, the scores of each data attribute, the weight of each data attribute, and the confidence score of the original data can all be output and logs generated through Tableau, Power BI visualization reports, and API (Application Programming Interface) to facilitate traceability and auditing.
[0110] In the method provided in this embodiment of the invention, the credibility of the data is evaluated through multiple data attributes, thereby obtaining the clinical data to be analyzed, ensuring the reliability of the clinical data itself, and thus ensuring the reliability of the clinical data analysis.
[0111] Based on any of the above embodiments, determining the credibility of the original data based on its data attributes includes: The credibility of the original data is determined based on its data attributes and the attribute weights of each data attribute in the application scenario of the original data.
[0112] Specifically, when assessing the credibility of raw data, the data attributes of the raw data can be weighted and summed to obtain the credibility score. Here, the weights used in the weighted summation can be adjusted to suit the application scenario of the raw data. That is, the weights applied in the credibility assessment can be different in different application scenarios.
[0113] Furthermore, different weights can be set for each data attribute for different application scenarios. For example, the weight settings can be referenced in the table below: Table 3. Weight Adaptation Table for Different Application Scenarios
[0114] The resulting credibility score, besides being used to confirm clinical data, can also be used for subsequent clinical data analysis. For example, a prompt can be generated along with the analysis results, such as "The current diagnostic conclusion is based on data with a credibility score of <60; please be cautious." Alternatively, credibility can be used as input during data analysis or as feature weights during model training, allowing for a greater focus on high-quality, high-credibility data during the analysis process.
[0115] For example, regarding a patient submitting "average weekly blood pressure readings (135 / 85 mmHg) from a home blood pressure monitor for hypertension management," the credibility assessment process for this data is as follows: Source Authority: This home blood pressure monitor has obtained NMPA Class II certification, but it was measured by the patient themselves, score S1=3 points; Data Quality: One day of data is missing in a week, the rest are normal. Score S2=4 points; Standardization: The patient stated that the measurement was performed according to the instructions, but there was no log record, score S3=3 points; Application Relevance: It is basically consistent with the hospital measurement value (138 / 88 mmHg) 3 months ago, score S4=4 points; Compliance: The patient has signed an authorization for the data to be used for health management, score S5=5 points.
[0116] The weighting for health management scenarios is as follows: W1=25%, W2=25%, W3=20%, W4=15%, W5=15%.
[0117] Therefore, the credibility is calculated: T S = (S1×W1)+(S2×W2)+(S3×W3)+(S4×W4)+(S5×W5) T S = (3×0.25)+(4×0.25)+(3×0.20)+(4×0.15)+(5×0.15) T S = 0.75 + 1.00 + 0.60 + 0.60 + 0.75 T S = 3.7 (out of 5), or equivalent to 74 out of 100. Therefore, the data demonstrates good reliability in health management scenarios. It can be adopted as clinical data; however, when used for high-risk decisions such as adjusting medication regimens, it should be noted that "confirmation in conjunction with recent clinical measurements is required."
[0118] Based on any of the above embodiments, step 140, which involves decoding the fused features to obtain the analysis results of the clinical data, includes: The fused features are decoded to obtain the patient's medical condition summary information as the analysis result.
[0119] Specifically, after obtaining the fused features, these features can be input into a decoder. The decoder performs feature decoding on the fused features, thereby obtaining the patient's disease summary information, which is then used as the analysis result of the clinical data. The decoder here can be a Transformer decoder, or it can be a CRF layer, MLP layer, etc.
[0120] Here, feature decoding of fusion features can output keywords or key information from multiple modalities of clinical data, and use these keywords or key information as disease summary information; or, feature decoding of fusion features can output a structured medical summary, such as the following format: Main symptoms: cough, chest pain, suspected diagnosis: community-acquired pneumonia, recommended examinations: chest X-ray, CRP test.
[0121] Based on any of the above embodiments, step 140, which involves decoding the fused features to obtain the analysis results of the clinical data, further includes: Based on the disease summary information and / or the fusion features, the patient's first and second medical information are determined. The first medical information is obtained based on a clinical guideline rule base, and the second medical information is obtained based on a medical recommendation model, which is obtained based on reinforcement learning. Based on the first and second diagnostic information, the patient's diagnostic information is determined as the analysis result.
[0122] Specifically, patient diagnosis and treatment information can also be obtained based on fusion features and / or disease summary information decoded from fusion features. This diagnosis and treatment information can reflect recommendations for the diagnosis and treatment of the patient's disease.
[0123] In this embodiment of the invention, first and second diagnostic and treatment information can be obtained. Both the first and second diagnostic and treatment information are diagnostic and treatment information specific to the patient, but they differ in the methods by which the information is obtained.
[0124] The primary diagnostic and treatment information can be obtained by matching the disease summary information and / or fusion features with a clinical guideline rule base. For example, it can be obtained by matching the disease summary information and / or fusion features with rules in the clinical guideline rule base, thereby obtaining a clinical guideline that matches the disease summary information and / or fusion features. The primary diagnostic and treatment information is the result of referencing clinical guidelines and has strong authority and reliability.
[0125] The second set of medical information can be output by a medical recommendation model based on reinforcement learning. For example, a medical recommendation model can be pre-trained using reinforcement learning, and then the patient's medical summary information and / or fusion features can be input into the model. The model then makes medical recommendations based on the input patient summary information and / or fusion features, thus producing the second set of medical information output by the model. Since the medical recommendation model itself acquires a large number of historical samples through reinforcement learning, the second set of medical information output by the model is a result obtained with reference to historical samples and can serve as a supplement to the first set of medical information.
[0126] After obtaining the primary and secondary diagnostic information, these two can be integrated to obtain the final diagnostic information for the patient, which is also used as the result of clinical data analysis.
[0127] For example, if the first diagnostic information recommends "glycated hemoglobin testing" and the second diagnostic information recommends "renal function testing (due to metformin use)," the following diagnostic information can be obtained: [Urgent] Glycated hemoglobin test (mandatory requirement in guidelines) [High Priority] Serum creatinine test (model predicts +23% risk of kidney dysfunction) For example, Figure 2 This is a flowchart illustrating the diagnostic information output method provided by the present invention, as shown below. Figure 2 As shown, disease summary information and / or fused features can be considered as patient data. For the clinical guideline rule base, a rule engine like Drools can be used to match the primary diagnostic and treatment information associated with the patient data from the clinical guideline rule base. Alternatively, the patient data can be input into a reinforcement learning model to obtain the secondary diagnostic and treatment information output by the model. The primary and secondary diagnostic and treatment information can be aggregated to form a recommendation list. Based on this, the recommendations can be sorted by credibility, thus prioritizing recommendations to doctors for diagnostic and treatment information with higher credibility.
[0128] In some embodiments, the resulting medical information can be pushed in real time through a floating window embedded in the doctor's end of the HIS (Hospital Information System); in addition, a simplified version of the medical report containing medication illustrations, symptom monitoring reminders, and other information can be generated and pushed to the patient's end.
[0129] Based on any of the above embodiments Figure 3 This is the second flowchart of the clinical data analysis method provided by the present invention, as shown below. Figure 3 As shown in the figure, an embodiment of the present invention provides a clinical data analysis method, including the following steps: First, acquire clinical data from patients across multiple modalities: It can obtain raw data related to patient health. This raw data may contain both structured and unstructured data.
[0130] For structured data, such as electronic medical records, test reports, and medical insurance records with fixed fields, database storage and analysis can be used directly. For example, a sample of structured data could be represented as: Time, Blood Glucose (mmol / L), Blood Pressure (mmHg) 2024-05-01 08:00, 5.2, 120 / 80. For unstructured data, such as data from wearable devices (heart rate, sleep), patient self-reported voice recordings, and social media health logs (e.g., dietary records), which lack a fixed format, natural language processing and computer vision technologies are needed to extract structured information. For example, unstructured data could be text-based data like handwritten doctor's records or a patient's self-report of "recently experiencing dizziness, which worsens after meals"; image-based data like CT scan DICOM files or photos of rashes taken with a smartphone; numerical-based data like waveform data from an electrocardiogram monitor; or audio-based data like a patient's recording describing their symptoms. For unstructured data, entities such as diseases, drugs, and symptoms can be extracted using natural language processing models such as BERT (Bidirectional Encoder Representations from Transformers) or large language models.
[0131] Furthermore, the credibility of raw data can be assessed by considering its data quality, source authority, standardization, compliance attributes, and application relevance. This helps determine whether to use the raw data as clinical data for analysis, and whether to combine raw data from different sources as clinical data.
[0132] Furthermore, clinical data can be preprocessed before being applied to subsequent data analysis. For example, ETL tools such as DataStage and Flink can be used to perform data cleaning operations such as deduplication and outlier filtering for clinical data of various modalities, as well as standardization operations such as ICD-10 and LOINC encoding.
[0133] For example, clinical data in the voice modality can be transcribed into the text modality; or, UMLS or MetaMap can be used to annotate medical terms in the text modality of clinical data as entities in the clinical data.
[0134] For example, for clinical data in numerical modalities, structural annotation can be performed, normalization can be performed, and discrete labels (such as "hemoglobin ↓") can be generated by combining medical reference ranges.
[0135] Secondly, based on a multi-modal encoder, clinical data from multiple modalities are encoded separately to obtain clinical features of multiple modalities: The multimodal encoders here can be obtained through federated learning, thereby ensuring the privacy of each hospital's patient data. Furthermore, contrastive learning is employed during the training of these multimodal encoders to achieve matching and correspondence between clinical data from different modalities at the feature, semantic, and representation levels. In addition, during the contrastive learning process, medical domain knowledge graphs can be used to establish associated entity pairs. By calculating the similarity between the entity features of associated entity pairs from different modalities, the loss of the contrastive learning is adjusted, thereby aligning the feature representations of associated entities.
[0136] Subsequently, based on the cross-modal attention mechanism, attention enhancement was performed on the clinical features of multiple modalities, thereby obtaining the first enhanced feature for each modality.
[0137] Next, based on the intramodal attention mechanism, attention enhancement is performed on the first enhancement feature of each modality, thereby obtaining the second enhancement feature of each modality.
[0138] Then, feature fusion is performed on the second enhancement feature of each modality to obtain the fused feature.
[0139] Finally, clinical data analysis was conducted based on fusion features: Specifically, correlation scoring, summary generation, and diagnostic information generation can be performed based on fusion features. Correlation scoring refers to calculating the matching degree between clinical data from different modalities, such as the matching degree between a text-based complaint and a numerical-based test result; specifically, it could be the matching degree between the complaint of "dry mouth" and the test result of blood glucose level. Diagnostic information can be output through a large language model (e.g., Med-PaLM 2), for example, "The patient complains of polydipsia and polyuria; combined with a blood glucose level of 15 mmol / L, it is recommended to rule out diabetes."
[0140] The above clinical data analysis methods can be deployed on the doctor's end. For example, when a general practitioner sees a patient, the doctor's end can automatically prompt: "The patient has a history of diabetes and is complaining of fatigue. It is recommended to prioritize the investigation of ketoacidosis (blood ketones need to be checked urgently)." As another example, specialists can view the medical history timeline generated on the doctor's end and quickly locate past surgical records (such as "cholecystectomy in 2019").
[0141] The clinical data analysis device provided by the present invention is described below. The clinical data analysis device described below and the clinical data analysis method described above can be referred to in correspondence.
[0142] Figure 4 This is a schematic diagram of the structure of the clinical data analysis device provided by the present invention, as shown below. Figure 4 As shown, the device includes: The data acquisition unit 410 is used to acquire clinical data of patients in multiple modalities; The feature encoding unit 420 is used to encode clinical data of multiple modalities based on a multi-modal encoder to obtain clinical features of multiple modalities. The multi-modal encoder is trained based on the similarity between entity features of associated entity pairs in the sample clinical features of sample clinical data of multiple modalities. The associated entity pairs are entity pairs that are associated in different modalities based on medical domain knowledge. The feature fusion unit 430 is used to fuse the clinical features of the multiple modalities to obtain fused features; The feature decoding unit 440 is used to perform feature decoding on the fused features to obtain the analysis results of the clinical data.
[0143] The apparatus provided in this invention trains a multi-modal encoder based on the similarity between clinical features of samples from multiple modalities and the similarity between entity features of associated entity pairs within the clinical features of samples from multiple modalities. This transforms clinical data from multiple modalities into a shared feature space while introducing medical domain knowledge during the data encoding process to enhance the professionalism and accuracy of the semantic representation of clinical features. Based on this, feature fusion and feature decoding are performed, enabling cross-modal clinical data analysis. This ensures the accuracy and reliability of clinical data analysis while maintaining the utilization rate of multimodal information.
[0144] Based on any of the above embodiments, the device further includes a training unit, used for: Based on a multimodal encoder, sample clinical data from multiple modalities are encoded separately to obtain sample clinical features from multiple modalities; Based on the similarity between the clinical features of the samples from the multiple modalities and the patient and time period classifications of the clinical data from the samples from the multiple modalities, the contrast loss is determined. The associated entity loss is determined based on the similarity between entity features of associated entity pairs in the clinical features of the samples from the multiple modalities. Based on the contrast loss and the associated entity loss, the encoders of the multiple modalities are subjected to parameter iteration.
[0145] Based on any of the above embodiments, the training unit is specifically used for: Obtain the clinical association conditions of the associated entity pairs, wherein the clinical association conditions are determined based on the medical domain knowledge; When the clinical data of the samples from the multiple modalities meet the clinical correlation conditions, the encoders of the multiple modalities are subjected to parameter iteration based on the contrast loss and the correlation entity loss.
[0146] Based on any of the above embodiments, the sample clinical data of the multiple modalities are local data; The training unit is specifically used for: Based on the contrast loss and the associated entity loss, local update parameters are determined and uploaded to the central server to trigger the central server to determine global update parameters based on the received local update parameters from various locations. The global update parameters returned by the central server are received, and the encoders of the multiple modalities are iterated based on the global update parameters.
[0147] Based on any of the above embodiments, the feature fusion unit is specifically used for: Based on the cross-modal similarity between the clinical features of the multiple modalities, cross-modal feature enhancement is performed on the clinical features of the multiple modalities to obtain the first enhanced features of the multiple modalities; Based on the element similarity of the first enhanced feature of each modality within the modality, intra-modal feature enhancement is performed on the first enhanced feature of each modality to obtain the second enhanced feature of each modality; The second enhancement features of the multiple modalities are fused to obtain the fused features.
[0148] Based on any of the above embodiments, the data acquisition unit is specifically used for: Acquire raw patient data across multiple modalities; Based on the data attributes of the original data, the credibility of the original data is determined. The data attributes include at least one of data quality, source authority, standardization degree, compliance attributes, and application relevance. The clinical data are determined based on the reliability of the original data.
[0149] Based on any of the above embodiments, the data acquisition unit is specifically used for: The credibility of the original data is determined based on its data attributes and the attribute weights of each data attribute in the application scenario of the original data.
[0150] Based on any of the above embodiments, the feature decoding unit is specifically used for: The fused features are decoded to obtain the patient's medical condition summary information as the analysis result.
[0151] Based on any of the above embodiments, the feature decoding unit is further configured to: Based on the disease summary information and / or the fusion features, the patient's first and second medical information are determined. The first medical information is obtained based on a clinical guideline rule base, and the second medical information is obtained based on a medical recommendation model, which is obtained based on reinforcement learning. Based on the first and second diagnostic information, the patient's diagnostic information is determined as the analysis result.
[0152] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5 As shown, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a clinical data analysis method, which includes: Acquire clinical data from patients across multiple modalities; Based on a multimodal encoder, clinical data from multiple modalities are encoded separately to obtain clinical features of multiple modalities. The multimodal encoder is trained based on the similarity between sample clinical features of multiple modalities and the similarity between entity features of associated entity pairs in sample clinical features of multiple modalities. The associated entity pairs are related entity pairs determined based on medical domain knowledge. By fusing the clinical features of the multiple modalities, a fused feature is obtained; The fusion features are decoded to obtain the analysis results of the clinical data.
[0153] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to related technologies, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0154] On the other hand, the present invention also provides a computer program product, the computer program product comprising a computer program that can be stored on a non-transitory computer-readable storage medium, wherein when the computer program is executed by a processor, the computer is able to perform the clinical data analysis method provided by the above methods, the method comprising: Acquire clinical data from patients across multiple modalities; Based on a multimodal encoder, clinical data from multiple modalities are encoded separately to obtain clinical features of multiple modalities. The multimodal encoder is trained based on the similarity between sample clinical features of multiple modalities and the similarity between entity features of associated entity pairs in sample clinical features of multiple modalities. The associated entity pairs are related entity pairs determined based on medical domain knowledge. By fusing the clinical features of the multiple modalities, a fused feature is obtained; The fusion features are decoded to obtain the analysis results of the clinical data.
[0155] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the clinical data analysis methods provided by the methods described above, the method comprising: Acquire clinical data from patients across multiple modalities; Based on a multimodal encoder, clinical data from multiple modalities are encoded separately to obtain clinical features of multiple modalities. The multimodal encoder is trained based on the similarity between sample clinical features of multiple modalities and the similarity between entity features of associated entity pairs in sample clinical features of multiple modalities. The associated entity pairs are related entity pairs determined based on medical domain knowledge. By fusing the clinical features of the multiple modalities, a fused feature is obtained; The fusion features are decoded to obtain the analysis results of the clinical data.
[0156] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0157] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the parts that contribute to the related technology, can be embodied in the form of software products. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0158] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A clinical data analysis method, characterized in that, include: Acquire clinical data from patients across multiple modalities; Based on a multimodal encoder, clinical data from multiple modalities are encoded separately to obtain clinical features of multiple modalities. The multimodal encoder is trained based on the similarity between sample clinical features of multiple modalities and the similarity between entity features of associated entity pairs in sample clinical features of multiple modalities. The associated entity pairs are related entity pairs determined based on medical domain knowledge. By fusing the clinical features of the multiple modalities, a fused feature is obtained; The fusion features are decoded to obtain the analysis results of the clinical data.
2. The clinical data analysis method according to claim 1, characterized in that, The training steps for the multimodal encoder include: Based on a multimodal encoder, sample clinical data from multiple modalities are encoded separately to obtain sample clinical features from multiple modalities; Based on the similarity between the clinical features of the samples from the multiple modalities and the patient and time period classifications of the clinical data from the samples from the multiple modalities, the contrast loss is determined. The associated entity loss is determined based on the similarity between entity features of associated entity pairs in the clinical features of the samples from the multiple modalities. Based on the contrast loss and the associated entity loss, the encoders of the multiple modalities are iterated.
3. The clinical data analysis method according to claim 2, characterized in that, The step of iterating the parameters of the encoder for the multiple modalities based on the contrast loss and the associated entity loss includes: Obtain the clinical association conditions of the associated entity pairs, wherein the clinical association conditions are determined based on the medical domain knowledge; When the clinical data of the samples from the multiple modalities meet the clinical correlation conditions, the encoders of the multiple modalities are subjected to parameter iteration based on the contrast loss and the correlation entity loss.
4. The clinical data analysis method according to claim 2, characterized in that, The clinical data from the multiple modalities are local data; The step of iterating the parameters of the encoder for the multiple modalities based on the contrast loss and the associated entity loss includes: Based on the contrast loss and the associated entity loss, local update parameters are determined and uploaded to the central server to trigger the central server to determine global update parameters based on the received local update parameters from various locations. The global update parameters returned by the central server are received, and the encoders of the multiple modalities are iterated based on the global update parameters.
5. The clinical data analysis method according to any one of claims 1 to 4, characterized in that, The fusion of clinical features from the multiple modalities yields fused features, including: Based on the cross-modal similarity between the clinical features of the multiple modalities, cross-modal feature enhancement is performed on the clinical features of the multiple modalities to obtain the first enhanced features of the multiple modalities; Based on the element similarity of the first enhanced feature of each modality within the modality, intra-modal feature enhancement is performed on the first enhanced feature of each modality to obtain the second enhanced feature of each modality; The second enhancement features of the multiple modalities are fused to obtain the fused features.
6. The clinical data analysis method according to any one of claims 1 to 4, characterized in that, The acquisition of patient clinical data across multiple modalities includes: Acquire raw patient data across multiple modalities; Based on the data attributes of the original data, the credibility of the original data is determined. The data attributes include at least one of data quality, source authority, standardization degree, compliance attributes, and application relevance. The clinical data are determined based on the reliability of the original data.
7. The clinical data analysis method according to claim 6, characterized in that, Determining the credibility of the original data based on its data attributes includes: The credibility of the original data is determined based on its data attributes and the attribute weights of each data attribute in the application scenario of the original data.
8. The clinical data analysis method according to any one of claims 1 to 4, characterized in that, The step of decoding the fused features to obtain the analysis results of the clinical data includes: The fused features are decoded to obtain the patient's medical condition summary information as the analysis result.
9. The clinical data analysis method according to claim 8, characterized in that, The step of decoding the fused features to obtain the analysis results of the clinical data further includes: Based on the disease summary information and / or the fusion features, the patient's first and second medical information are determined. The first medical information is obtained based on a clinical guideline rule base, and the second medical information is obtained based on a medical recommendation model, which is obtained based on reinforcement learning. Based on the first and second diagnostic information, the patient's diagnostic information is determined as the analysis result.
10. A clinical data analysis device, characterized in that, include: The data acquisition unit is used to acquire clinical data from patients across multiple modalities. The feature encoding unit is used to encode clinical data from multiple modalities using a multi-modal encoder to obtain clinical features from multiple modalities. The multi-modal encoder is trained based on the similarity between entity features of associated entity pairs in the sample clinical features of sample clinical data from multiple modalities. The associated entity pairs are entity pairs that are associated in different modalities, as determined by knowledge in the medical domain. A feature fusion unit is used to fuse the clinical features of the multiple modalities to obtain fused features; The feature decoding unit is used to perform feature decoding on the fused features to obtain the analysis results of the clinical data.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the clinical data analysis method as described in any one of claims 1 to 9.
12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the clinical data analysis method as described in any one of claims 1 to 9.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the clinical data analysis method as described in any one of claims 1 to 9.