Diagnosis method, system and equipment based on multi-modal medical data and medium
Through the combined method of joint learning and neural networks, combined with blockchain and encryption technology, the privacy protection and efficient integration of multimodal medical data is achieved, and the security and accuracy of multimodal data in diagnosis and efficacy prediction are solved, improving diagnostic accuracy and prediction accuracy.
Patent Information
- Application Number
- CN202510374678.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-08-05
AI Technical Summary
The prior art is difficult to efficiently utilize multimodal medical data while protecting the security and privacy of the security and privacy of the disease, resulting in limited diagnostic accuracy and efficacy prediction accuracy.
The method of combining joint learning and neural network is adopted to protect the multimodal medical data through a multimodal data fusion model, extract and splice fusion features, use diagnostic prediction models to diagnose and efficacy prediction, and verify data integrity through blockchain storage model parameters and encrypted digests to realize multimodal data fusion across institutions.
On the premise of ensuring data security and privacy protection, the diagnostic accuracy and efficacy prediction accuracy of multimodal medical data are improved, the problems of data heterogeneity and inconsistency are solved, and efficient data utilization and collaborative modeling are achieved.
Smart Images

Figure CN120432124A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical data processing technology, and in particular to a diagnostic method, system, device and medium based on multimodal medical data. Background Art
[0002] With the development of artificial intelligence (AI) technology, significant breakthroughs have been achieved in its application to medical innovation. Analyzing medical data using technologies such as big data, deep learning, and machine learning can enable medical diagnosis and efficacy prediction, providing effective insights for the development of clinical practice. Currently, an increasing number of solutions utilize multimodal medical data fusion for diagnosis and efficacy prediction. However, the complexity and sensitivity of multimodal medical data make it difficult to efficiently utilize the data while protecting its security, limiting the accuracy of diagnostics and efficacy predictions. Summary of the Invention
[0003] In order to solve the above technical problems or at least partially solve the above technical problems, the present invention provides a diagnosis method, system, device and medium based on multimodal medical data.
[0004] In a first aspect, the present invention provides a diagnostic method based on multimodal medical data, comprising:
[0005] Performing privacy protection processing on the multimodal medical data collected from each patient, wherein the multimodal medical data includes multiple single-modal medical data;
[0006] Utilizing a multimodal data fusion model, extracting features of each of the processed single-modal medical data of each patient, and concatenating and fusing the features of the single-modal medical data of the same patient, to obtain fused features of the multimodal medical data of each patient represented in a unified feature space; wherein the multimodal data fusion model is obtained by: obtaining model parameters trained by each participant in model training, the model parameters trained by the participant being obtained by training a first neural network using local multimodal medical data as training samples; and determining model parameters of the multimodal data fusion model based on the model parameters trained by each participant;
[0007] A diagnostic prediction model is used to perform diagnostic prediction based on the fusion features of each patient to obtain diagnostic information and efficacy prediction information for each patient; wherein the diagnostic prediction model is obtained by training the second neural network based on the fusion feature samples.
[0008] According to a multimodal medical data-based diagnostic method provided by the present invention, extracting features of each of the processed single-modal medical data of each patient, and concatenating and fusing the features of each of the single-modal medical data of the same patient to obtain fused features of the multimodal medical data of each patient represented in a unified feature space, includes:
[0009] Extracting features of each of the processed unimodal medical data of each patient to obtain a feature matrix of each of the unimodal medical data; the feature matrix of the unimodal medical data includes a first dimension and a second dimension; the first dimension represents different patients; and the second dimension represents features of the unimodal medical data of each patient;
[0010] The feature matrices of each of the single-modal medical data are fused to obtain a unified feature space; the unified feature space includes a third dimension and a fourth dimension; the third dimension represents different patients; the fourth dimension represents the fusion features of each patient's multimodal medical data; the fusion features are obtained by splicing the features of each of the single-modal medical data of the same patient along the second dimension.
[0011] According to a diagnostic method based on multimodal medical data provided by the present invention, the multiple single-modal medical data include medical imaging data, genomic data and clinical record data;
[0012] The extracting the features of each of the single-modality medical data of each patient after the processing to obtain a feature matrix of each of the single-modality medical data includes:
[0013] extracting morphological features of the tumor from the processed medical imaging data of each patient to obtain a feature matrix of the medical imaging data;
[0014] extracting cancer-related gene variation features from the processed genomic data of each patient to obtain a feature matrix of the genomic data;
[0015] Extracting features of time series information from the processed clinical record data of each patient to obtain a feature matrix of the clinical record data;
[0016] The step of fusing the feature matrices of the single-modal medical data to obtain a unified feature space includes:
[0017] The features of the medical imaging data, the genomic data, and the clinical record data of the same patient in the feature matrix of the medical imaging data, the feature matrix of the genomic data, and the feature matrix of the clinical record data are spliced along the second dimension to obtain the unified feature space.
[0018] A diagnostic method based on multimodal medical data provided by the present invention further includes:
[0019] The model parameters of the multimodal data fusion model are stored in the blockchain and updated when the model parameters of the multimodal data fusion model change.
[0020] A diagnostic method based on multimodal medical data provided by the present invention further includes:
[0021] Generating an encrypted summary of the collected multimodal medical data for each patient, and storing the encrypted summary in a blockchain;
[0022] regenerating an encrypted summary of the target patient's multimodal medical data;
[0023] comparing the regenerated encrypted digest with the encrypted digest of the target patient's multimodal medical data stored in the blockchain;
[0024] Based on the comparison results, the integrity of the multimodal medical data of the target patient is verified.
[0025] A diagnostic method based on multimodal medical data provided by the present invention further includes:
[0026] When a data access request is received, the requester’s permissions are verified based on the smart contract in the blockchain;
[0027] If the permission verification is passed, the requesting party is authorized to access the corresponding data.
[0028] According to a multimodal medical data-based diagnostic method provided by the present invention, privacy protection processing is performed on the collected multimodal medical data of each patient, including:
[0029] Annotate and temporally align the collected multimodal medical data for each patient;
[0030] De-identifying the aligned multimodal medical data of each patient using differential privacy technology;
[0031] The de-identified multimodal medical data of each patient is symmetrically encrypted.
[0032] According to a multimodal medical data-based diagnostic method provided by the present invention, the differential privacy technology is used to de-identify the aligned multimodal medical data of each patient, including:
[0033] Determine the intensity of de-identification currently required;
[0034] Determining parameter values in the differential privacy technology based on the strength of the de-identification;
[0035] The aligned multimodal medical data of each patient is de-identified based on the parameter values using differential privacy technology.
[0036] According to a diagnostic method based on multimodal medical data provided by the present invention, symmetrically encrypting the de-identified multimodal medical data of each patient includes:
[0037] Determine the encryption strength currently required;
[0038] Determining an encryption algorithm for symmetric encryption based on the encryption strength;
[0039] The de-identified multimodal medical data of each patient is symmetrically encrypted using the encryption algorithm.
[0040] In a second aspect, the present invention provides a diagnostic system based on multimodal medical data, comprising:
[0041] a data processing module, configured to perform privacy protection processing on the multimodal medical data collected from each patient, wherein the multimodal medical data includes a plurality of single-modal medical data;
[0042] a multimodal data fusion module, configured to utilize a multimodal data fusion model to extract features of each of the processed single-modal medical data of each patient, and to concatenate and fuse the features of each of the single-modal medical data of the same patient, thereby obtaining fused features of the multimodal medical data of each patient represented in a unified feature space; wherein the multimodal data fusion model is obtained by: obtaining model parameters trained by each participant in model training, the model parameters trained by each participant being obtained by training a first neural network using local multimodal medical data as training samples; and determining model parameters of the multimodal data fusion model based on the model parameters trained by each participant;
[0043] The diagnosis prediction module is used to use the diagnosis prediction model to perform diagnosis prediction based on the fusion features of each patient to obtain the diagnosis information and efficacy prediction information of each patient; wherein, the diagnosis prediction model is obtained by training the second neural network based on the fusion feature samples.
[0044] In a third aspect, the present invention provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the processor implements a diagnostic method based on multimodal medical data as described above.
[0045] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described diagnostic methods based on multimodal medical data.
[0046] The technical solution provided by the embodiment of the present invention has the following advantages compared with the existing technology:
[0047] In the solution of the present invention, after the patient's multimodal medical data is privacy-protected to improve data security and privacy, a multimodal data fusion model based on joint learning and neural networks can efficiently fuse the features of the patient's multimodal medical data (such as medical imaging data, genomic data, and clinical record data), realize the integration of multimodal data in a unified feature space, establish consistent semantic relationships, solve the problems of data heterogeneity and inconsistency, and enable subsequent diagnostic prediction models to simultaneously utilize the features of multimodal data, thereby obtaining more comprehensive, accurate, and expressive disease descriptions and diagnostic information, thereby improving the accuracy of diagnosis and the accuracy of efficacy prediction. In addition, the method of combining joint learning and neural networks is used to achieve multi-party collaborative data modeling, without the need for multimodal medical data exchange between the participating parties, ensuring a balance between privacy protection, data security, and accuracy improvement of multimodal medical data during the fusion process. In this way, multimodal medical data is efficiently utilized while ensuring data security and privacy protection, thereby improving the accuracy of diagnosis and efficacy prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0050] Figure 1 This is one of the flowcharts of the multimodal medical data-based diagnostic method provided by the present invention;
[0051] Figure 2 This is the second flow chart of the diagnostic method based on multimodal medical data provided by the present invention;
[0052] Figure 3 Schematic diagram of the structure of the multimodal medical data-based diagnostic system provided by the present invention;
[0053] Figure 4It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0054] In order to more clearly understand the above-mentioned objectives, features and advantages of the present invention, the scheme of the present invention will be further described below. It should be noted that, in the absence of conflict, the embodiments of the present invention and the features therein can be combined with each other.
[0055] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present invention, rather than all the embodiments.
[0056] The following combination Figures 1 to 4 The present invention describes a multimodal medical data-based diagnosis method, system, device, and medium.
[0057] This embodiment provides a diagnosis method based on multimodal medical data, which can be executed by a diagnosis system based on multimodal medical data, such as Figure 1 As shown, it at least includes the following steps:
[0058] S11. Perform privacy protection processing on the collected multimodal medical data of each patient, where the multimodal medical data includes a plurality of single-modal medical data.
[0059] S12. Utilizing a multimodal data fusion model, extract the features of each of the single-modal medical data of each patient after the processing, and concatenate and fuse the features of each of the single-modal medical data of the same patient to obtain the fusion features of the multimodal medical data of each patient represented in a unified feature space; wherein, the multimodal data fusion model is obtained in the following manner: obtaining model parameters trained by each participant in the model training, the model parameters trained by the participant are obtained by the participant training a first neural network using local multimodal medical data as training samples; and determining the model parameters of the multimodal data fusion model based on the model parameters trained by each of the participants.
[0060] S13. Using a diagnostic prediction model, a diagnostic prediction is performed based on the fusion features of each patient to obtain diagnostic information and efficacy prediction information for each patient; wherein the diagnostic prediction model is obtained by training a second neural network based on the fusion feature samples.
[0061] In specific implementation, multimodal medical data of one or more patients may be collected, where the multimodal medical data includes multiple single-modal medical data, such as medical imaging data, genomic data, and clinical record data.
[0062] Medical imaging data is used to detect the morphological characteristics of tumors, and may include computed tomography (CT) data or magnetic resonance imaging (MRI) data.
[0063] Genomic data is used to analyze genetic variations associated with cancer and may include deoxyribonucleic acid (DNA) sequencing results.
[0064] Clinical record data, including patient medical history, treatment plan, drug response and other information.
[0065] To protect data privacy and security, privacy protection can be first applied to each patient's collected multimodal medical data. The processed unimodal medical data for each patient are then fed into a multimodal data fusion model to extract features from the privacy-protected unimodal medical data. This extracts the features of each patient's unimodal medical data. The features of the unimodal medical data for the same patient are then concatenated and fused to obtain fused features of each patient's multimodal medical data represented in a unified feature space. Finally, the fused features are fed into a diagnostic prediction model to obtain diagnostic information and efficacy prediction information for each patient.
[0066] To achieve efficient model training while protecting data privacy, this paper uses a federated learning approach to train the multimodal data fusion model. The goal of federated learning is to enable multiple participants to collaboratively train a shared multimodal data fusion model without exchanging original data.
[0067] First, determine the participants in the model training. Let the i-th participant hold a local multimodal medical data dataset D i The model update process of joint learning includes:
[0068] The participants use the dataset D i The multimodal medical data in is used as the training sample to train the first neural network, and the model parameters w of the multimodal data fusion model are initially obtained. i .
[0069]
[0070] in, represents the model parameters after the tth iteration, represents the model parameters after the t+1th iteration, represents the loss function and η is the learning rate.
[0071] The server aggregates the model parameters trained by all participants Get the global model parameter w t+1 , which is the model parameter w of the final multimodal data fusion model.
[0072]
[0073] Where N is the number of participants.
[0074] The first neural network may be a deep neural network (DNN) or other neural networks.
[0075] Through the above-mentioned joint learning process, there is no need to exchange multimodal medical data between the participating parties. While ensuring data privacy, it is possible to achieve cross-institutional multimodal data fusion modeling and improve the generalization ability of the model.
[0076] Based on feature fusion and joint learning, a pre-trained diagnostic prediction model is used to diagnose diseases (such as cancer) and predict therapeutic effects to obtain diagnostic information and therapeutic effect prediction information.
[0077] During training, DNN can be selected as the second neural network. The second neural network is trained using the fusion feature samples. Let the vector of fusion features be represented as F, then the output diagnosis information and efficacy prediction information prediction result y can be expressed as:
[0078] y= D NN( F )(3)
[0079] Among them, DNN represents a deep neural network, which consists of multiple layers of fully connected layers and activation functions. The model is trained by minimizing the error of the prediction results.
[0080] In order to train the diagnosis prediction model, this embodiment uses the cross entropy loss function to measure the difference between the model prediction result and the true label. Let the true label be y and the prediction result be Then the cross entropy loss function L is defined as:
[0081]
[0082] Among them, s is the number of categories, y j and denote the true label and predicted probability of the jth class respectively.
[0083] Furthermore, in order to dynamically evaluate the treatment effect of patients, we can also combine time series prediction models, such as Long Short Term Memory (LSTM), to analyze the changing trend of patients during treatment. Suppose the characteristics of the patient at time step t′ are F t′ , the output of LSTM is the predicted efficacy indicator
[0084]
[0085] LSTM can capture long-term dependencies in time series through the memory and forget gate mechanisms, thereby dynamically evaluating the patient's treatment progress and helping doctors adjust treatment plans based on the patient's real-time condition.
[0086] In order to evaluate the performance of the cancer diagnosis and efficacy prediction model, this embodiment uses multiple evaluation indicators, including accuracy, precision, recall, and F1 score.
[0087] Assume that the number of true positives is TP, the number of false positives is FP, and the number of false negatives is FN, then the definitions of these evaluation indicators are as follows:
[0088] Accuracy:
[0089]
[0090] Accuracy:
[0091]
[0092] Recall:
[0093]
[0094] F1 score:
[0095]
[0096] The above model evaluation metrics comprehensively measure the performance of diagnostic prediction models, including their classification accuracy and ability to distinguish between different categories. By continuously optimizing the structure and training methods of diagnostic prediction models, we can further enhance their performance, providing strong support for accurate diagnosis and personalized treatment of diseases.
[0097] In this embodiment, after the patient's multimodal medical data is privacy-protected to improve data security and privacy, a multimodal data fusion model based on joint learning and deep neural networks can efficiently fuse the features of the patient's multimodal medical data (such as medical imaging data, genomic data, and clinical record data), realize the integration of multimodal data in a unified feature space, establish consistent semantic relationships, solve the problems of data heterogeneity and inconsistency, and enable subsequent diagnostic prediction models to simultaneously utilize the features of multimodal data, thereby obtaining more comprehensive, accurate, and expressive disease descriptions and diagnostic information, thereby improving the accuracy of diagnosis and the accuracy of efficacy prediction. In addition, the method of combining joint learning and deep neural networks is used to achieve multi-party collaborative data modeling, without the need for multimodal medical data exchange between the participating parties, ensuring the balance of privacy protection, data security, and accuracy improvement of multimodal medical data during the fusion process. In this way, while ensuring data security and privacy protection, multimodal medical data is efficiently utilized, improving diagnostic accuracy and efficacy prediction accuracy.
[0098] In some embodiments, extracting the features of each of the single-modal medical data of each patient after the processing, and concatenating and fusing the features of each of the single-modal medical data of the same patient to obtain the fused features of the multimodal medical data of each patient represented in a unified feature space, may specifically include:
[0099] The first step is to extract the features of each of the single-modal medical data of each patient after the processing to obtain a feature matrix of each of the single-modal medical data; the feature matrix of the single-modal medical data includes a first dimension and a second dimension; the first dimension represents different patients; the second dimension represents the features of the single-modal medical data of each patient.
[0100] In practical applications, batch processing can be used to process the multimodal medical data of each patient. The number of patients processed in each batch can be set according to actual needs. The number of patients processed in each batch is the batch size (Batch Size), represented by B. The size of the first dimension in the feature matrix of the single-modal medical data is B, and the size of the second dimension is the vector length of the features of the single-modal medical data of each patient.
[0101] The second step is to fuse the feature matrices of each of the single-modal medical data to obtain a unified feature space; the unified feature space includes a third dimension and a fourth dimension; the third dimension represents different patients; the fourth dimension represents the fusion features of each patient's multimodal medical data; the fusion features are obtained by splicing the features of each of the single-modal medical data of the same patient along the second dimension.
[0102] The third dimension in the unified feature space is the same as the second dimension, and its size is B. The size of the fourth dimension is the sum of the vector lengths of the features of each of the single-modal medical data.
[0103] In this embodiment, before feature fusion is performed, feature extraction is performed on each single modality medical data in the multimodal medical data respectively, so as to convert the different modal medical data into a unified feature space representation, and directly splicing is performed in the feature dimension to achieve feature-level fusion. Through this feature-level fusion operation, the integration of different modal medical data is achieved in a unified feature space, so that the subsequent diagnostic prediction model can simultaneously utilize the features of multimodal medical data, thereby obtaining more comprehensive, accurate and expressive disease descriptions and diagnostic information.
[0104] In the case where the multiple single-modality medical data include medical imaging data, genomic data, and clinical record data, the features of each single-modality medical data of each patient after the processing are extracted to obtain a feature matrix of each single-modality medical data, such as Figure 2 Specifically, it may include:
[0105] The first step is to extract the morphological features of the tumor from the processed medical imaging data of each patient to obtain a feature matrix of the medical imaging data.
[0106] Specifically, for each patient, convolutional neural networks (CNN) are used to extract the spatial features (i.e., morphological features) of the tumor from medical imaging data.
[0107] Suppose the patient's medical image data is I, with dimensions (B, H, W, D1, N1), where B is the batch size, H is the height, W is the width, D1 is the depth (if it is a 3D image), and N1 is the number of channels.
[0108] Through multi-layer convolution and pooling operations, feature extraction is performed to obtain the feature representation f of medical imaging data. I , the dimension is (B, F I ), where F I Is the feature dimension after the convolution operation. Here F I The specific value of depends on the configuration of the convolutional layer, the pooling layer, and the number of output channels:
[0109]
[0110] in, is the feature matrix of medical imaging data.
[0111] The second step is to extract cancer-related gene variation features from the processed genomic data of each patient to obtain a feature matrix of the genomic data;
[0112] Specifically, for each patient, the genomic data is subjected to a one-dimensional convolutional layer to extract cancer-related gene variation features.
[0113] Let the genomic data be G, with dimensions (B, L, N2), where B is the batch size, L is the length of the gene sequence, and N2 is the number of channels (which may be 1 or contain more variation information).
[0114] The feature extraction result of genomic data is f G , the dimension is (B, F G ), where F G is the length of the sequence after convolution kernel processing, then:
[0115]
[0116] in, is the feature matrix of genomic data.
[0117] The third step is to extract the features of the time series information of the processed clinical record data of each patient to obtain a feature matrix of the clinical record data.
[0118] Specifically, for each patient, a recurrent neural network (RNN) or LSTM is used to process the time series information in the clinical record data. The dimensions of the input clinical record data C are (B, T, F C ), where B is the batch size, T is the number of time steps, and F C is the input feature dimension at each time step. For example, the feature representation f of clinical record data is obtained C , the dimension is (B, H C ), where H C is the hidden layer feature dimension of LSTM, that is, the hidden state of the final time step represents the information of the entire sequence:
[0119]
[0120] in, is the feature matrix of clinical record data.
[0121] Accordingly, the fusion of the feature matrices of the single-modal medical data to obtain a unified feature space may specifically include:
[0122] The features of the medical imaging data, the genomic data, and the clinical record data of the same patient in the feature matrix of the medical imaging data, the feature matrix of the genomic data, and the feature matrix of the clinical record data are spliced along the second dimension to obtain the unified feature space.
[0123] When the characteristics of medical imaging data are f I , the dimension is (B, F I ), the characteristics of genomic data are f G , the dimension is (B, F G ), the clinical record data is characterized by f C , the dimension is (B, F C ), then the fused feature vector is represented as F, and the dimension is (B, F I +F G +F C ), so the feature-level splicing operation is used to directly merge the above three feature vectors in the feature dimension (i.e., the second dimension) to achieve splicing. The splicing operation can be expressed by the following formula:
[0124]
[0125] Among them, [f I , f G , f C ] represents horizontal concatenation of the feature matrix of medical imaging data, the feature matrix of genomic data, and the feature matrix of clinical record data along the feature dimension. Specifically: for each patient, the feature matrix of their medical imaging data f I The length is F I vector, the feature f of genomic data G The length is F G vector, the features f of clinical record data C The length is F C After concatenation, each patient corresponds to a vector of length (F I +F G +F C )’s unified feature vector F, i.e., fusion feature:
[0126]
[0127] In the batch processing case (B>1), this operation performs the same splicing for each patient in the batch, thus generating a (B×(F I +F G +F C ))’s unified characteristic matrix.
[0128] Finally, the diagnostic prediction model is used to perform diagnostic prediction based on the fusion features of the patients in the unified feature matrix to obtain the diagnostic information and efficacy prediction information of each patient.
[0129] Through the above-mentioned feature-level fusion operation, this embodiment realizes the integration of three types of modal medical data: medical imaging data, genomic data, and clinical record data in a unified feature space, so that the subsequent diagnostic prediction model can simultaneously utilize the features of multimodal medical data, thereby obtaining more comprehensive, accurate, and expressive disease descriptions and diagnostic information.
[0130] In some embodiments, the diagnostic method based on multimodal medical data may further include: storing model parameters of the multimodal data fusion model in the blockchain, and updating the model parameters of the multimodal data fusion model when the model parameters of the multimodal data fusion model change.
[0131] Blockchain is decentralized and immutable. In this embodiment, the model parameters of the multimodal data fusion model are uploaded to the blockchain for storage and updated in a timely manner, which can ensure the security and transparency of the model parameters.
[0132] In some embodiments, the multimodal medical data-based diagnostic method may further include: generating an encrypted digest of the collected multimodal medical data for each patient, storing the encrypted digest in a blockchain, regenerating an encrypted digest of the target patient's multimodal medical data, comparing the regenerated encrypted digest with the encrypted digest of the target patient's multimodal medical data stored in the blockchain, and verifying the integrity of the target patient's multimodal medical data based on the comparison result.
[0133] In specific implementation, a corresponding encrypted digest (Hash) can be generated for each patient's multimodal medical data set to ensure the uniqueness and security of the data. Assuming that the original data of the patient's multimodal medical data is D, the encrypted digest H(D) can be expressed as:
[0134] H(D)=Hash(D)(15)
[0135] Hash(·) represents a cryptographic hash function, such as SHA-256. This encrypted summary is stored on the blockchain, with each patient assigned a unique hash value, ensuring the data is tamper-proof.
[0136] When data is requested for access or transmission, the encrypted summary of the dataset D″ of the multimodal medical data of the accessed patient (i.e., the target patient) can be regenerated and compared with the original summary H(D) stored in the blockchain. The verification process can be expressed as:
[0137] H(D″)=?H(D)(16)
[0138] If the regenerated summary H(D″) is the same as H(D) stored in the blockchain, it proves that the data has not been tampered with during storage or transmission, and the integrity of the data is verified.
[0139] In addition, blockchains have consensus mechanisms, such as Proof of Work (PoW) or Proof of Stake (PoS). Through the blockchain's consensus mechanism, each node reaches a consensus on the consistency of data, ensuring the reliability of each transaction record and data storage, thereby further guaranteeing the authenticity and integrity of the data, making any attempt to tamper with the data difficult to achieve.
[0140] In this embodiment, by combining the encryption summary and the blockchain consensus mechanism, an efficient and secure data integrity verification method is implemented, ensuring the reliability and non-tamperability of multimodal medical data throughout its entire life cycle.
[0141] In some embodiments, the diagnostic method based on multimodal medical data may further include: when a data access request is received, verifying the authority of the requester based on the smart contract in the blockchain; if the authority verification passes, authorizing the requester to access the corresponding data.
[0142] A smart contract is an automated protocol running on a blockchain that defines the permissions and conditions for data sharing. A smart contract specifies the permissions and conditions for data access. Let's denote the smart contract C, whose inputs are the requester's ID and the access request R. The execution result P of the smart contract can be expressed as:
[0143] P=C(ID,R)(17)
[0144] Among them, P is the permission verification result. Only when the request meets the predefined conditions in the smart contract can it pass the permission verification and be authorized to access the corresponding data.
[0145] In this embodiment, refined management of data access rights is achieved through smart contracts. The data access rights and conditions are specified in the smart contract to ensure that only the authorized party can access sensitive data, thereby achieving secure data sharing. This allows participants to share data without having to directly exchange original data, thereby achieving efficient collaborative diagnosis and efficacy prediction while protecting data privacy, and enabling automated and transparent management of the data sharing process.
[0146] In some embodiments, the privacy protection processing of the collected multimodal medical data of each patient may specifically include:
[0147] The first step is to annotate the multimodal medical data collected for each patient and align them in time.
[0148] Since different data sources may have data missing and time asynchrony problems, the data needs to be labeled and interpolated. Through interpolation, multimodal medical data can be aligned in time.
[0149] Assume that the data set is X = {x1, x2, ..., x n}, where x i′ Represents the i′th data point. For missing data points, the nearest neighbor interpolation method is used to ensure the continuity of the data on the time axis.
[0150] The interpolated data can be expressed as:
[0151]
[0152] Among them, N(i′) represents the i′ The k most similar data.
[0153] The second step is to use differential privacy technology to de-identify the multimodal medical data of each patient after alignment.
[0154] In order to protect the privacy of patients, differential privacy technology can be used to de-identify multimodal medical data.
[0155] Differential privacy technology hides individual information by adding noise to the original data, thereby statistically protecting the privacy of the data. For the original dataset D and the de-identified dataset D′ of the patient's multimodal medical data, the differential privacy mechanism can be expressed as:
[0156]
[0157] Where f(D) is the statistical function of the data set, Indicates that the mean is 0 and the variance is σ 2 Gaussian noise.
[0158] The third step is to symmetrically encrypt the de-identified multimodal medical data of each patient.
[0159] To further ensure the security of data during transmission and sharing, a symmetric encryption algorithm can be used to encrypt the de-identified multimodal medical data of each patient.
[0160] Assuming the original data is m and the encryption key is k, the encrypted data c can be expressed as:
[0161] c=E k (m)(20)
[0162] Among them, E k It is a symmetric encryption algorithm. The encrypted data can only be decrypted and obtained when the key k is possessed.
[0163] In the above embodiment, when processing data for privacy protection, the data is first de-identified using differential privacy technology, and then encrypted using a symmetric encryption algorithm. That is, by combining multi-level protection of symmetric encryption and differential privacy technology, comprehensive protection of patient sensitive information is ensured during data transmission and sharing, thereby ensuring the privacy and security of the data. This multi-layer security mechanism can effectively reduce the risk of data leakage, and while de-identifying the data, it can minimize computing overhead and improve data sharing efficiency.
[0164] Specifically, the use of differential privacy technology to de-identify the multimodal medical data of each patient after annotation may include: determining the currently required de-identification strength; determining the parameter value in the differential privacy technology based on the de-identification strength; and using differential privacy technology to de-identify the multimodal medical data of each patient after alignment based on the parameter value.
[0165] The parameter value of the differential privacy technology may include the variance value of Gaussian noise. Different de-identification strengths correspond to different parameter values of the differential privacy technology.
[0166] In actual applications, the strength of identification can be adjusted according to different user needs and data usage scenarios, dynamic privacy protection strategies can be implemented, and different privacy protection needs can be flexibly responded to.
[0167] Similarly, symmetrically encrypting the de-identified multimodal medical data of each patient may specifically include: determining the currently required encryption strength; determining an encryption algorithm for symmetric encryption based on the encryption strength; and symmetrically encrypting the de-identified multimodal medical data of each patient using the encryption algorithm.
[0168] Different encryption strengths correspond to different encryption algorithms. Symmetric encryption algorithms can include DES (Data Encryption Standard), a block algorithm that uses secret key encryption, and AES (Advanced Encryption Standard), an advanced encryption standard.
[0169] In actual applications, the encryption strength can be adjusted according to different user needs and data usage scenarios, so as to select encryption algorithms with different encryption strengths, implement dynamic privacy protection strategies, and flexibly respond to different privacy protection needs.
[0170] Through dynamic privacy protection strategies, data utilization efficiency is maximized while ensuring data security.
[0171] The solution of the embodiment of the present invention significantly improves the accuracy, security and privacy protection of cancer diagnosis and efficacy prediction through the integration of multimodal medical data, blockchain storage management and smart contract-driven data sharing, providing an efficient and reliable solution to address many challenges faced in multimodal medical data processing and collaboration in the existing technology.
[0172] The following describes a diagnostic system based on multimodal medical data provided by the present invention. The diagnostic system based on multimodal medical data described below and the diagnostic method based on multimodal medical data described above can refer to each other.
[0173] This embodiment provides a diagnosis system based on multimodal medical data, such as Figure 3 As shown, including:
[0174] A data processing module 301 is configured to perform privacy protection processing on the multimodal medical data collected from each patient, wherein the multimodal medical data includes multiple single-modal medical data;
[0175] The multimodal data fusion module 302 is configured to utilize a multimodal data fusion model to extract features of each of the processed single-modal medical data of each patient, and to concatenate and fuse the features of each of the single-modal medical data of the same patient, thereby obtaining fused features of the multimodal medical data of each patient represented in a unified feature space. The multimodal data fusion model is obtained by: obtaining model parameters trained by each participant in model training, the model parameters trained by each participant being obtained by training a first neural network using local multimodal medical data as training samples; and determining model parameters of the multimodal data fusion model based on the model parameters trained by each participant.
[0176] The diagnosis prediction module 303 is used to use the diagnosis prediction model to perform diagnosis prediction based on the fusion features of each patient to obtain the diagnosis information and efficacy prediction information of each patient; wherein, the diagnosis prediction model is obtained by training the second neural network based on the fusion feature samples.
[0177] In some embodiments, the multimodal data fusion module 302 is specifically configured to:
[0178] Extracting features of each of the processed unimodal medical data of each patient to obtain a feature matrix of each of the unimodal medical data; the feature matrix of the unimodal medical data includes a first dimension and a second dimension; the first dimension represents different patients; and the second dimension represents features of the unimodal medical data of each patient;
[0179] The feature matrices of each of the single-modal medical data are fused to obtain a unified feature space; the unified feature space includes a third dimension and a fourth dimension; the third dimension represents different patients; the fourth dimension represents the fusion features of each patient's multimodal medical data; the fusion features are obtained by splicing the features of each of the single-modal medical data of the same patient along the second dimension.
[0180] In some embodiments, the plurality of single-modality medical data includes medical imaging data, genomic data, and clinical record data;
[0181] The multimodal data fusion module 302 is specifically configured to:
[0182] extracting morphological features of the tumor from the processed medical imaging data of each patient to obtain a feature matrix of the medical imaging data;
[0183] extracting cancer-related gene variation features from the processed genomic data of each patient to obtain a feature matrix of the genomic data;
[0184] Extracting features of time series information from the processed clinical record data of each patient to obtain a feature matrix of the clinical record data;
[0185] The step of fusing the feature matrices of the single-modal medical data to obtain a unified feature space includes:
[0186] The features of the medical imaging data, the genomic data, and the clinical record data of the same patient in the feature matrix of the medical imaging data, the feature matrix of the genomic data, and the feature matrix of the clinical record data are spliced along the second dimension to obtain the unified feature space.
[0187] In some embodiments, a blockchain storage management module is further included, configured to:
[0188] The model parameters of the multimodal data fusion model are stored in the blockchain and updated when the model parameters of the multimodal data fusion model change.
[0189] In some embodiments, the blockchain storage management module is further configured to:
[0190] Generating an encrypted summary of the collected multimodal medical data for each patient, and storing the encrypted summary in a blockchain;
[0191] regenerating an encrypted summary of the target patient's multimodal medical data;
[0192] comparing the regenerated encrypted digest with the encrypted digest of the target patient's multimodal medical data stored in the blockchain;
[0193] Based on the comparison results, the integrity of the multimodal medical data of the target patient is verified.
[0194] In some embodiments, the blockchain storage management module is further configured to:
[0195] When a data access request is received, the requester’s permissions are verified based on the smart contract in the blockchain;
[0196] If the permission verification is passed, the requesting party is authorized to access the corresponding data.
[0197] In some embodiments, the data processing module 301 is specifically configured to:
[0198] Annotate and temporally align the collected multimodal medical data for each patient;
[0199] De-identifying the aligned multimodal medical data of each patient using differential privacy technology;
[0200] The de-identified multimodal medical data of each patient is symmetrically encrypted.
[0201] In some embodiments, the data processing module 301 is specifically configured to:
[0202] Determine the intensity of de-identification currently required;
[0203] Determining parameter values in the differential privacy technology based on the strength of the de-identification;
[0204] The aligned multimodal medical data of each patient is de-identified based on the parameter values using differential privacy technology.
[0205] In some embodiments, the data processing module 301 is specifically configured to:
[0206] Determine the encryption strength currently required;
[0207] Determining an encryption algorithm for symmetric encryption based on the encryption strength;
[0208] The de-identified multimodal medical data of each patient is symmetrically encrypted using the encryption algorithm.
[0209] Figure 4 An example of a physical structure diagram of an electronic device is shown below. Figure 4 As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other via the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute a diagnostic method based on multimodal medical data, which includes:
[0210] Performing privacy protection processing on the multimodal medical data collected from each patient, wherein the multimodal medical data includes multiple single-modal medical data;
[0211] Utilizing a multimodal data fusion model, extracting features of each of the processed single-modal medical data of each patient, and concatenating and fusing the features of the single-modal medical data of the same patient, to obtain fused features of the multimodal medical data of each patient represented in a unified feature space; wherein the multimodal data fusion model is obtained by: obtaining model parameters trained by each participant in model training, the model parameters trained by the participant being obtained by training a first neural network using local multimodal medical data as training samples; and determining model parameters of the multimodal data fusion model based on the model parameters trained by each participant;
[0212] A diagnostic prediction model is used to perform diagnostic prediction based on the fusion features of each patient to obtain diagnostic information and efficacy prediction information for each patient; wherein the diagnostic prediction model is obtained by training the second neural network based on the fusion feature samples.
[0213] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0214] On the other hand, the present invention further provides a computer program product, comprising a computer program stored on a non-transitory computer-readable storage medium, wherein the computer program comprises program instructions. When the program instructions are executed by a computer, the computer is capable of performing the diagnostic method based on multimodal medical data provided by the above methods, the method comprising:
[0215] Performing privacy protection processing on the multimodal medical data collected from each patient, wherein the multimodal medical data includes multiple single-modal medical data;
[0216] Utilizing a multimodal data fusion model, extracting features of each of the processed single-modal medical data of each patient, and concatenating and fusing the features of the single-modal medical data of the same patient, to obtain fused features of the multimodal medical data of each patient represented in a unified feature space; wherein the multimodal data fusion model is obtained by: obtaining model parameters trained by each participant in model training, the model parameters trained by the participant being obtained by training a first neural network using local multimodal medical data as training samples; and determining model parameters of the multimodal data fusion model based on the model parameters trained by each participant;
[0217] A diagnostic prediction model is used to perform diagnostic prediction based on the fusion features of each patient to obtain diagnostic information and efficacy prediction information for each patient; wherein the diagnostic prediction model is obtained by training the second neural network based on the fusion feature samples.
[0218] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the above-mentioned diagnostic method based on multimodal medical data, the method comprising:
[0219] Performing privacy protection processing on the multimodal medical data collected from each patient, wherein the multimodal medical data includes multiple single-modal medical data;
[0220] Utilizing a multimodal data fusion model, extracting features of each of the processed single-modal medical data of each patient, and concatenating and fusing the features of the single-modal medical data of the same patient, to obtain fused features of the multimodal medical data of each patient represented in a unified feature space; wherein the multimodal data fusion model is obtained by: obtaining model parameters trained by each participant in model training, the model parameters trained by the participant being obtained by training a first neural network using local multimodal medical data as training samples; and determining model parameters of the multimodal data fusion model based on the model parameters trained by each participant;
[0221] A diagnostic prediction model is used to perform diagnostic prediction based on the fusion features of each patient to obtain diagnostic information and efficacy prediction information for each patient; wherein the diagnostic prediction model is obtained by training the second neural network based on the fusion feature samples.
[0222] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0223] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0224] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A diagnostic method based on multimodal medical data, characterized in that: include: Performing privacy protection processing on the multimodal medical data collected from each patient, wherein the multimodal medical data includes multiple single-modal medical data; Utilizing a multimodal data fusion model, extracting features of each of the processed single-modal medical data of each patient, and concatenating and fusing the features of the single-modal medical data of the same patient, to obtain fused features of the multimodal medical data of each patient represented in a unified feature space; wherein the multimodal data fusion model is obtained by: obtaining model parameters trained by each participant in model training, the model parameters trained by the participant being obtained by training a first neural network using local multimodal medical data as training samples; and determining model parameters of the multimodal data fusion model based on the model parameters trained by each participant; A diagnostic prediction model is used to perform diagnostic prediction based on the fusion features of each patient to obtain diagnostic information and efficacy prediction information for each patient; wherein the diagnostic prediction model is obtained by training the second neural network based on the fusion feature samples.
2. The diagnostic method based on multimodal medical data according to claim 1, characterized in that: The extracting the features of each of the single-modal medical data of each patient after the processing, and concatenating and fusing the features of each of the single-modal medical data of the same patient to obtain the fused features of the multimodal medical data of each patient represented in a unified feature space, includes: Extracting features of each of the processed unimodal medical data of each patient to obtain a feature matrix of each of the unimodal medical data; the feature matrix of the unimodal medical data includes a first dimension and a second dimension; the first dimension represents different patients; and the second dimension represents features of the unimodal medical data of each patient; The feature matrices of each of the single-modal medical data are fused to obtain a unified feature space; the unified feature space includes a third dimension and a fourth dimension; the third dimension represents different patients; the fourth dimension represents the fusion features of each patient's multimodal medical data; the fusion features are obtained by splicing the features of each of the single-modal medical data of the same patient along the second dimension.
3. The diagnostic method based on multimodal medical data according to claim 2, characterized in that: The multiple single-modality medical data include medical imaging data, genomic data, and clinical record data; The extracting the features of each of the single-modality medical data of each patient after the processing to obtain a feature matrix of each of the single-modality medical data includes: extracting morphological features of the tumor from the processed medical imaging data of each patient to obtain a feature matrix of the medical imaging data; extracting cancer-related gene variation features from the processed genomic data of each patient to obtain a feature matrix of the genomic data; Extracting features of time series information from the processed clinical record data of each patient to obtain a feature matrix of the clinical record data; The step of fusing the feature matrices of the single-modal medical data to obtain a unified feature space includes: The features of the medical imaging data, the genomic data, and the clinical record data of the same patient in the feature matrix of the medical imaging data, the feature matrix of the genomic data, and the feature matrix of the clinical record data are spliced along the second dimension to obtain the unified feature space.
4. The diagnostic method based on multimodal medical data according to claim 1, characterized in that: Also includes: The model parameters of the multimodal data fusion model are stored in the blockchain and updated when the model parameters of the multimodal data fusion model change.
5. The diagnostic method based on multimodal medical data according to claim 1, characterized in that: Also includes: Generating an encrypted summary of the collected multimodal medical data for each patient, and storing the encrypted summary in a blockchain; regenerating an encrypted summary of the target patient's multimodal medical data; comparing the regenerated encrypted digest with the encrypted digest of the target patient's multimodal medical data stored in the blockchain; Based on the comparison results, the integrity of the multimodal medical data of the target patient is verified.
6. The diagnostic method based on multimodal medical data according to claim 1, characterized in that: Also includes: When a data access request is received, the requester’s permissions are verified based on the smart contract in the blockchain; If the permission verification is passed, the requesting party is authorized to access the corresponding data.
7. The diagnostic method based on multimodal medical data according to claim 1, characterized in that: The privacy protection processing of the collected multimodal medical data of each patient includes: Annotate and temporally align the collected multimodal medical data for each patient; De-identifying the aligned multimodal medical data of each patient using differential privacy technology; The de-identified multimodal medical data of each patient is symmetrically encrypted.
8. The diagnostic method based on multimodal medical data according to claim 7, characterized in that: De-identifying the aligned multimodal medical data of each patient using differential privacy technology includes: Determine the intensity of de-identification currently required; Determining parameter values in the differential privacy technology based on the strength of the de-identification; The aligned multimodal medical data of each patient is de-identified based on the parameter values using differential privacy technology.
9. The diagnostic method based on multimodal medical data according to claim 7, characterized in that: Symmetrically encrypting the de-identified multimodal medical data of each patient includes: Determine the encryption strength currently required; Determining an encryption algorithm for symmetric encryption based on the encryption strength; The de-identified multimodal medical data of each patient is symmetrically encrypted using the encryption algorithm.
10. A diagnostic system based on multimodal medical data, characterized in that: include: a data processing module, configured to perform privacy protection processing on the multimodal medical data collected from each patient, wherein the multimodal medical data includes a plurality of single-modal medical data; a multimodal data fusion module, configured to utilize a multimodal data fusion model to extract features of each of the processed single-modal medical data of each patient, and to concatenate and fuse the features of each of the single-modal medical data of the same patient, thereby obtaining fused features of the multimodal medical data of each patient represented in a unified feature space; wherein the multimodal data fusion model is obtained by: obtaining model parameters trained by each participant in model training, the model parameters trained by each participant being obtained by training a first neural network using local multimodal medical data as training samples; and determining model parameters of the multimodal data fusion model based on the model parameters trained by each participant; The diagnosis prediction module is used to use the diagnosis prediction model to perform diagnosis prediction based on the fusion features of each patient to obtain the diagnosis information and efficacy prediction information of each patient; wherein, the diagnosis prediction model is obtained by training the second neural network based on the fusion feature samples.
11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the diagnostic method based on multimodal medical data as claimed in any one of claims 1 to 9 is implemented.
12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the diagnosis method based on multimodal medical data according to any one of claims 1 to 9 is implemented.