Parkinson's disease combined diagnosis method based on cross-modal consistency fusion and modal dynamic equilibrium mechanism
By adopting the method of cross-modal consistency fusion and modal dynamic balance mechanism in Parkinson's disease diagnosis, the intermodal heterogeneity and information imbalance when neuroimaging data and hematologic transcriptomics data are fusion, and the accuracy and reliability of the diagnosis are improved.
Patent Information
- Application Number
- CN202411925294.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-06-03
AI Technical Summary
When the prior art fuses neuroimaging data with hematologic transcriptomics data, it faces problems of heterogeneity between modals, information imbalance and modal imbalance, which affects the accuracy and reliability of Parkinson's disease diagnosis.
Using a method based on the cross-modal consistency fusion and modal dynamic equilibrium mechanism, the characteristics of nuclear magnetic resonance imaging and hematologic transcriptomics data are extracted through a multi-scale multi-channel convolutional neural network and a multi-layer perceptron. Combined with the cross-modal consistency fusion module and the modal dynamic equilibrium mechanism, the specificity and consistency information between modals is coordinated, and the cross-modal joint representation is generated to improve diagnostic accuracy.
It effectively solves the problem of information imbalance between modals, improves the accuracy and reliability of Parkinson's disease diagnosis, and enhances the ability to dynamically model the disease.
Smart Images

Figure CN120089327A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of multimodal data fusion and Parkinson's disease diagnosis, and particularly relates to a combined diagnosis method for Parkinson's disease based on cross-modal consistency fusion and modal dynamic balance mechanism. Background Art
[0002] Parkinson's disease is a common neurodegenerative disease, mainly manifested as motor dysfunction, tremors, muscle stiffness, and bradykinesia. As the disease progresses, Parkinson's disease can also cause non-motor symptoms such as cognitive decline and depression. The pathogenesis of this disease is complex, involving multiple biological pathways and genetic factors. Due to the complex etiology, traditional diagnostic methods mostly rely on clinical symptoms and imaging examinations, and it is usually difficult to accurately diagnose the disease in the early stage. Therefore, the early diagnosis of Parkinson's disease has always been a major challenge.
[0003] With the progress of molecular biology and high-throughput sequencing technology, more and more research has focused on the analysis of the molecular mechanism of Parkinson's disease. Transcriptomics technologies such as RNA sequencing provide new means for studying the molecular level of neurodegenerative diseases. By detecting changes in gene expression levels, molecular markers related to the disease state can be identified. At the same time, neuroimaging technologies, such as magnetic resonance imaging, can directly reflect changes in brain structure and function. These two types of data provide key information about the disease from the molecular and macroscopic levels respectively. However, due to the large heterogeneity between the two, how to effectively combine these different modal data and conduct joint analysis has become one of the key issues in the current field of Parkinson's disease diagnosis.
[0004] Existing multimodal learning methods have tried to fuse neuroimaging data with blood transcriptomics data, but still face many challenges. First, there are significant heterogeneities between different modal data. Neuroimaging data usually shows the spatial structure information of the brain, while blood transcriptomics data reflects gene expression levels. These two types of data have significant differences in spatio-temporal scale and biological significance, resulting in great inconsistencies in their joint analysis. Second, existing methods often use simple feature splicing or decision-level fusion in modal fusion, ignoring the interaction information between different modal data. Such a method is difficult to fully explore the potential value of multimodal data. Finally, due to the imbalance problem of data volume and information content between modalities, some dominant modalities may inhibit the learning of other modalities, leading to the fusion result being biased towards the modality with larger information content, thereby affecting the generalization ability of the model and the reliability of diagnosis.
[0005] To address the above problems, some researchers have proposed solutions for multi-modal data fusion using deep learning models. For example, convolutional neural networks are used to process imaging data, fully connected neural networks are used to process blood transcriptomics data, and specific fusion strategies are employed to combine this modal information. However, these methods still fail to effectively solve the problem of information imbalance between modalities. Especially when dealing with heterogeneous data, it is prone to modal imbalance, which affects the accuracy and robustness of diagnosis.
[0006] Therefore, how to effectively coordinate and interact information between modalities and thereby improve the accuracy and reliability of the Parkinson's disease diagnosis model has become a core challenge in current multi-modal data analysis. For this reason, the present invention proposes a combined Parkinson's disease diagnosis method based on cross-modal consistency fusion and modal dynamic balance mechanism. By fusing the specific and consistent information in magnetic resonance imaging data and blood transcriptomics data, the imbalance problem between modalities is solved, thereby enhancing the accuracy and reliability of Parkinson's disease diagnosis. Summary of the Invention
[0007] To overcome the above problems, the present invention proposes a combined Parkinson's disease diagnosis method based on cross-modal consistency fusion and modal dynamic balance mechanism. This method utilizes a multi-modal learning framework, integrates magnetic resonance imaging and blood transcriptomics data, and maintains the balance of specific and consistent information between modalities through a dynamic balance mechanism, thereby improving the accuracy and reliability of disease dynamic modeling.
[0008] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0009] A combined Parkinson's disease diagnosis method based on cross-modal consistency fusion and modal dynamic balance mechanism, the method comprising the following steps:
[0010] A. Preprocess the magnetic resonance imaging data, and use a multi-scale multi-channel convolutional neural network to extract features from the magnetic resonance imaging data
[0011] to extract pathological features at different spatial and channel scales;
[0012] B. Preprocess the blood transcriptomics data, and use a multi-layer perceptron model to extract features from the blood transcriptomics data
[0013] to obtain specific features related to gene expression levels;
[0014] C. Use a cross-modal consistency fusion module to fuse the magnetic resonance imaging features and blood transcriptomics features, and extract the consistency between modalities through a multi-layer attention mechanism;
[0015]
[0016] D. Use the modal dynamic balance mechanism to dynamically calculate the contribution weights between modalities, balance the influence of different modalities in the joint representation, generate a cross-modal joint representation, and obtain the final Parkinson's disease diagnosis result.
[0017] Preferably, in the steps A and B, the data sources specifically include but are not limited to:
[0018] (1) Magnetic resonance imaging data of the Parkinson's Progression Markers Initiative project PPMI;
[0019] (2) Magnetic resonance imaging data of the Parkinson's Data and Biomarker Program PDBP;
[0020] (3) Transcriptomic data of PPMI and PDBP in the Accelerating Medicines Partnership - Parkinson's Disease AMP-PD.
[0021] Preferably, in the steps A and B, the data labels specifically include but are not limited to:
[0022] (1) Parkinson's disease healthy control group, preclinical Parkinson's disease, Parkinson's disease with normal scans, Parkinson's disease; (2) Stages 0 - 5 in the Hoehn - Yahr grading disease classification.
[0023] Preferably, in step A, use a multi-scale multi-channel convolutional neural network to extract features from magnetic resonance imaging.
[0024] The specific steps are as follows:
[0025] (1) Define the magnetic resonance imaging data input as X M ;
[0026] (2) Extraction of multi-scale features; in this step, let the input matrix X M be P 0 in order to obtain feature information at multiple scales.
[0027] The corresponding convolution operation is expressed as:
[0028] Y t = P i * K t , t = 1, 2,..., T
[0029] where P i is the result after applying the multi-scale multi-channel convolutional neural network module i times, the symbol * represents the convolution operation, t represents the index of the convolution kernel function, and there are T convolution kernels in total; K t represents the t-th convolution kernel function; Y t represents the feature map obtained after passing through the t-th convolution kernel function; subsequently, the convolution feature maps are concatenated together to obtain a multi-scale feature map YM ;
[0030] Y M = Concat(Y 1 , Y 2 , …, Y T )
[0031] where Concat represents the feature concatenation operation;
[0032] (3) Extraction of multi-channel features; In this step, the SE module models the dependencies between channels and adaptively recalibrates the channel feature responses, thereby enhancing the representation ability of the network;
[0033] The SE module includes Squeeze and Excitation; The Squeeze operation compresses the spatial information of each channel into a scalar by applying global average pooling:
[0034]
[0035] where H and W are the height and width of the feature map respectively, C is the number of channels, c represents the channel, and h, w represent the pixel positions; Subsequently, the Excitation operation recalibrates the weights of each channel through two fully connected layers and the ReLU activation function:
[0036] s c = σ(W 2 δ(W 1 )z c ), c = 1, 2, …, C
[0037] where W 1 and W 2 are weight matrices, δ represents the ReLU activation function, and σ represents the Sigmoid activation function;
[0038] After the scaling and weighting operation, the learned weight s c of each channel is multiplied by the corresponding channel of the previous feature map to obtain:
[0039]
[0040] where, represents the part belonging to channel c in the multi-scale feature map Y M ;
[0041] (5) After the above processing, the output p i+1 is flattened and input into a fully connected layer to obtain a neuroimage-specific representation
[0042]
[0043] Among them, Fc represents the fully connected layer, Flatten represents the flattening operation, GAP refers to the global average pooling operation, and p I is the output result of the last layer of the multi-scale multi
[0044] channel convolutional neural network module.
[0045] Preferably, in step B, a multi-layer perceptron model is used to extract features from the blood transcriptomics data, and its specific
[0046] formula is:
[0047] h 1 = Dropout 1 (BatchNorm 1 (ReLU(W 1 X T + b 1 )))
[0048] h 2 = Dropout 2 (BatchNorm 2 (ReLU(W 2 h 1 + b 2 )))
[0049] h 3 = Dropout 3 (BatchNorm 3 (ReLU(W 3 h 2 + b 3 )))
[0050]
[0051] Among them, X T is the input matrix of the blood transcriptomics data, W i and b i respectively represent the weight matrix and the bias vector of the i-th layer, ReLU represents the Leaky ReLU activation function; BatchNorm i represents the batch normalization operation of the i-th layer; Dropout i represents the Dropout operation of the i-th layer; Softmax represents the Softmax function.
[0052] Preferably, in step C, a cross-modal consistency fusion module is used for magnetic resonance imaging features and blood transcriptomics
[0053] Fuse the features, and extract the consistency between modalities through a multi-layer attention mechanism. The specific steps are as follows:
[0054] (1) Define a one-sided single-layer cross-modal consistency fusion module, and its specific formula is:
[0055]
[0056]
[0057] Among them, LN(·) is a linear layer, MultiHead(·) is a multi-head attention mechanism, and FFN is a feed-forward network. is the intermediate result; is the result obtained by the one-sided single-layer cross-modal consistency fusion module; the nuclear magnetic resonance imaging modality data representation is used as the query, that is, the blood transcriptomics modality data representation is used as the key and values, that is, The multi-head attention mechanism in it divides the input into multiple sub-spaces, applies attention calculation to each sub-space, then splices the results, and finally obtains the output through a linear transformation; integrating and defining the two formulas, we get:
[0058]
[0059] Among them, CMCFusion represents integrating the one-sided single-layer cross-modal consistency fusion module into a function operation.
[0060] (2) Similarly, use the blood transcriptomics modality data representation as the query and the nuclear magnetic resonance imaging modality data representation
[0061] as the key and values, and the process and result X M→T are expressed as:
[0062]
[0063] (3) Then, integrate the two parts through concatenation to obtain the result of the cross-modal consistency representation:
[0064] X consensus = Concat[X T→M , X M→T
[0065] Preferably, in step D, a modality dynamic balance mechanism is used to dynamically calculate the contribution weights between modalities and balance the influences of different
[0066] modalities in the joint representation. The specific steps are as follows:
[0067] (1) Define the joint representation as:
[0068]
[0069] (2) Calculate the difference rate of the contribution of each modality to the learning objective:
[0070]
[0071]
[0072]
[0073] Among them, F 1 score represents the F1 score, y i represents the true label of the sample, f(·) represents the linear classifier, represents the class index j with the largest predicted probability among the outputs of all classifiers, B t is the batch during training, is defined as the reciprocal of;
[0074] (3) Use to dynamically adjust the contribution difference between the blood transcriptomics modality and the magnetic resonance imaging modality, and adaptively modulate the loss function through the following formula:
[0075]
[0076] Among them, u represents any one of the modalities in M and T.
[0077] Preferably, in step D, the step of constructing the dynamic loss function is as follows:
[0078] (1) Define two modality-specific feature extractors as
[0079] (2) Use the single-modal datasets to pre-train the two modality-specific feature extractors respectively, so that the model can fully
[0080] learn the features of a single modality, and obtain the pre-trained modality-specific feature extractors:
[0081]
[0082]
[0083] Among them, and represent the single-modal-specific feature extractors after migration.
[0084] (3) Definition and represent the parameters of the final linear classifier, W M and W T represent the projection matrices respectively, represents the set of real numbers in set theory, and are the dimensions of the features of modalities M and T respectively, and Class represents the number of classes; the logits output of the multi-modal model is as follows:
[0085]
[0086] where, θ M and θ T represent the parameters to be optimized during training, and b represents the bias.
[0087] (4) Calculate the cross-entropy loss of the joint representation:
[0088]
[0089] where, Class represents the number of classes, represents the cross-entropy loss of the joint representation, X joint represents the joint representation, and N represents the number of samples.
[0090] (5) Define the specificity loss of modality u as:
[0091]
[0092] where, represents the cross-entropy loss of the specific representation of modality u; finally, obtain the overall dynamic loss function:
[0093]
[0094] where, γ is an adjustable hyperparameter, and L M 、L T represent the specificity losses of modalities M and T respectively;
[0095] Compared with the prior art, the advantages of the present invention are:
[0096] (1) The present invention provides a combined diagnosis method for Parkinson's disease based on cross-modal consistency fusion and modal dynamic balance mechanism. A multi-scale multi-channel convolutional neural network is used to extract the modality-specific features of nuclear magnetic resonance imaging data, and a multi-layer perceptron is used to extract the modality-specific features of blood transcriptomics data. Then, consensus feature fusion between modalities is performed through a cross-modal consistency fusion module. Finally, the modality dynamic balance mechanism coordinates the specificity and consistency information between the two modalities to generate a combined representation for disease diagnosis. A more comprehensive perspective is utilized for the combined diagnosis of Parkinson's disease.
[0097] (2) The cross-modal consistency fusion mechanism designed in the present invention combines blood transcriptomics data and nuclear magnetic resonance imaging information to better understand the distribution characteristics of Parkinson's disease. Blood transcriptomics data can capture the characteristics of data at the microscopic scale, while nuclear magnetic resonance imaging data can identify atrophy or lesions generated in relevant brain regions at the apparent scale. Co-training can help the present invention reveal the internal connections between data modalities and deeply understand the multi-scale features of the data.
[0098] (3) The modality dynamic balance mechanism designed in the present invention alleviates the problem of modality imbalance to a certain extent. Using the technique of single-modal distillation, the representation patterns of single modalities are fully learned. During joint training, the difference rate between modalities is calculated to dynamically adjust the weights between modalities. It can help the present invention fully extract the information in single-modal data and perform reasonable dynamic coordination. BRIEF DESCRIPTION OF THE DRAWINGS
[0099] The present invention will be further described below with reference to the accompanying drawings.
[0100] Figure 1 is a flowchart of the diagnosis method based on cross-modal consistency fusion and modal dynamic balance mechanism of the present invention;
[0101] Figure 2 is the trajectory of the F1 score and modality difference rate during multi-modal learning in the PPMI case in the embodiment of the present invention;
[0102] Figure 3 is a comparison of the evaluation indexes obtained by the combined diagnosis method for Parkinson's disease based on cross-modal consistency fusion and modal dynamic balance mechanism and the SOTA multi-modal learning framework in the embodiment of the present invention;
[0103] Figure 4 is the ROC curve and PR curve obtained from the ablation experiment in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0104] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0105] The hardware environment for the operation of this embodiment: one laptop, CPU: 6.00GHz, GPU: RTX 4090 (64G); software environment: Python 3.8, Pytorch 1.11.0, Cuda 11.3; operating platform: ubuntu20.04.
[0106] This instance tested the effectiveness of the proposed method on the Parkinson's disease dataset, which includes the Parkinson's Progression Markers Initiative database (PPMI), the Parkinson's Disease Biomarker Program database (PDBP), and the Accelerating Medicines Partnership - Parkinson's (AMP - PD). Two types of labels were used to divide Parkinson's disease subtypes: the first was divided into Healthy Control, Prodromal Parkinson's disease, Scans Without Evidence of Dopamine Deficit (SWEDD), and Parkinson's disease (PD) according to the Parkinson's type. The second was divided into 6 stages from stage 0 to 5 according to the Hoehn - Yahr disease staging. The detailed information of the experimental data is shown in Table 1.
[0107] Table 1 Detailed information of the experimental dataset
[0108]
[0109] The combined Parkinson's disease diagnosis method based on cross - modal consistency fusion and modal dynamic balance mechanism provided by the present invention has a flowchart as Figure 1 shown, and specifically includes the following steps:
[0110] A. Pre - process the nuclear magnetic resonance imaging data, and use a multi - scale multi - channel convolutional neural network to extract features from the nuclear magnetic resonance imaging, extracting pathological features at different spatial and channel scales.
[0111] Furthermore, the steps for using a multi - scale multi - channel convolutional neural network to extract features from the nuclear magnetic resonance imaging are as follows:
[0112] (1) Define the input of the nuclear magnetic resonance imaging data as X M .
[0113] (2) Extraction of multi - scale features. In this step, let the input matrix X M be P 0 , and the convolution kernel function is designed as K 1 , K 2 ,..., K T so as to obtain feature information at multiple scales. The corresponding convolution operation is expressed as:
[0114] Y t =P i *K t, t = 1, 2, ..., T
[0115] where P i is the result after applying the multi-scale multi-channel convolutional neural network module i times. The symbol * represents the convolution operation, t represents the index of the convolution kernel function, and there are a total of T convolution kernels. K t represents the t-th convolution kernel function. Y t represents the feature map obtained after passing through the t-th convolution kernel function; subsequently, the convolution feature maps are concatenated together to obtain a multi-scale feature map Y M ;
[0116] Y M = Concat(Y 1 , Y 2 , …, Y T )
[0117] where Concat represents the feature concatenation operation.
[0118] (3) Extraction of multi-channel features. In this step, the SE module models the dependencies between channels and adaptively recalibrates the channel feature responses, thereby enhancing the representation ability of the network. The SE module consists of two main operations: Squeeze and Excitation. The Squeeze operation compresses the spatial information of each channel into a scalar by applying global average pooling:
[0119]
[0120] where H and W are the height and width of the feature map respectively, C is the number of channels, c represents the channel, and h, w represent the pixel positions.
[0121] Subsequently, the Excitation operation recalibrates the weights of each channel through two fully connected layers and the ReLU activation function:
[0122] s c = σ(W 2 δ(W 1 )z c ), c = 1, 2, …, C
[0123] where W 1 and W 2 are weight matrices, δ represents the ReLU activation function, and σ represents the Sigmoid activation function;
[0124] After the scaling and weighting operation, the learned weight s c of each channel is multiplied by the corresponding channel of the previous feature map to obtain:
[0125]
[0126] Among them, represents the part belonging to channel c in the multi-scale feature map Y M .
[0127] (4) After the above processing, output p i+1 is flattened and input into the fully connected layer to obtain a neuroimage-specific representation
[0128]
[0129] Among them, Fc represents the fully connected layer, Flatten represents the flattening operation, GAP refers to the global average pooling operation, and pI is the output result of the last layer of the multi-scale multi-channel convolutional neural network module.
[0130] B. Preprocess the blood transcriptomics data, and use a multi-layer perceptron model to extract features from the blood transcriptomics data to obtain specific features related to gene expression levels.
[0131] Furthermore, use a multi-layer perceptron model to extract features from the blood transcriptomics data, and its specific formula is:
[0132] h 1 = Dropout 1 (BatchNorm 1 (ReLU(W 1 X T + b 1 )))
[0133] h 2 = Dropout 2 (BatchNorm 2 (ReLU(W 2 h 1 + b 2 )))
[0134] h 3 = Dropout 3 (BatchNorm 3 (ReLU(W 3 h 2 + b 3 )))
[0135]
[0136] Among them, X T is the input matrix of the blood transcriptomics data, W i and b irespectively represent the weight matrix and bias vector of the i-th layer. ReLU represents the Leaky ReLU activation function. BatchNorm i represents the batch normalization operation of the i-th layer. Dropout i represents the Dropout operation of the i-th layer. Softmax represents the Softmax function.
[0137] C. Use a cross-modal consistency fusion module to fuse nuclear magnetic resonance imaging features and blood transcriptomics features, and extract the consistency between modalities through a multi-layer attention mechanism.
[0138] Furthermore, use a cross-modal consistency fusion module to fuse nuclear magnetic resonance imaging features and blood transcriptomics features, and extract the consistency between modalities through a multi-layer attention mechanism. The specific steps are as follows:
[0139] (1) Define a one-sided one-layer cross-modal consistency fusion module, and its specific formula is:
[0140]
[0141]
[0142] where LN(·) is a linear layer, MultiHead(·) is a multi-head attention mechanism, FFN is a feed-forward network, is the intermediate result; is the result obtained by the one-layer cross-modal consistency fusion module; the nuclear magnetic resonance imaging modality data representation is used as the query, that is the blood transcriptomics modality data representation is used as the key and values, that is The multi-head attention mechanism in it divides the input into multiple subspaces, applies attention calculation to each subspace, then splices the results, and finally obtains the output through a linear transformation; integrating and defining the two formulas, we get:
[0143]
[0144] where CMCFusion represents integrating the one-sided one-layer cross-modal consistency fusion module into a function operation.
[0145] (2) Similarly, when the blood transcriptomics modality data representation is used as the query and the nuclear magnetic resonance imaging modality data representation is used as the key and values, the process and result X M→T can be expressed as:
[0146]
[0147] (3) Then, integrate the two parts through concatenation to obtain the result of cross-modal consistency representation:
[0148] X consensus = Concat[X T→M , X M→T
[0149] D. Use the modal dynamic balance mechanism to dynamically calculate the contribution weights between modalities, balance the influence of different modalities in the joint representation, and generate a cross-modal joint representation to obtain the final Parkinson's disease diagnosis result.
[0150] Furthermore, use the modal dynamic balance mechanism to dynamically calculate the contribution weights between modalities and balance the influence of different modalities in the joint representation. The specific steps are as follows:
[0151] (1) Define the joint representation as:
[0152]
[0153] (2) Calculate the difference rate of the contribution of each modality to the learning objective:
[0154]
[0155]
[0156]
[0157] Among them, F 1 score represents the F1 score, y i represents the true label of the sample, f(·) represents a linear classifier, ) represents, among the outputs of all classifiers, taking out the class index j with the largest predicted probability, and B t is the batch during training. is defined as the reciprocal of.
[0158] (3) Use to dynamically adjust the contribution difference between the blood transcriptomics modality and the magnetic resonance imaging modality, and adaptively modulate the loss function through the following formula:
[0159]
[0160] Among them, u represents any one of M and T modalities.
[0161] Furthermore, the above steps construct a dynamic loss function as follows:
[0162] (1) Define two modality-specific feature extractors as
[0163] (2) Pre-train the two modality-specific feature extractors using single-modal datasets respectively, so that the model can fully learn the features of a single modality, and obtain the pre-trained modality-specific feature extractors:
[0164]
[0165]
[0166] Among them, and represent the single-modal specific feature extractors after migration.
[0167] (3) Define and to represent the parameters of the final linear classifier, W M and W T represent the projection matrices respectively, and are the dimensions of the features of modalities M and T respectively, and Class represents the number of classes; the logits output of the multi-modal model is as follows:
[0168]
[0169] Among them, θ M and α T represent the parameters to be optimized during training. In this embodiment, the loss function is dynamically adjusted according to the contribution differences between modalities, and the loss function is fed back to the training process of the joint representation. The final diagnosis result is obtained from the formula of the logits output of the multi-modal model through the joint representation.
[0170] (4) Calculate the cross-entropy loss of the joint representation:
[0171]
[0172] Among them, Class represents the number of classes.
[0173] (5) Define the specific loss of modality u as:
[0174]
[0175] Among them, Class represents the number of classes, represents the cross-entropy loss of the joint representation, X joint represents the joint representation, and N represents the number of samples.
[0176] (6) Finally, obtain the overall loss function:
[0177]
[0178] Among them, γ is an adjustable hyperparameter, L M and L T respectively represent the specificity losses of modalities M and T.
[0179] Six types of metrics, namely accuracy (ACC), F1-score (F1), recall, precision, area under the receiver operating characteristic curve (AUROC), and area under the precision-recall curve (AUPR), are used to evaluate the accuracy of the model, and the larger these metrics are, the better. F1 is the harmonic mean of precision and recall, comprehensively considering the accuracy and integrity of the model. Precision represents the proportion of samples correctly predicted as positive by the classifier among the samples predicted as positive, focusing on the accuracy of the model's prediction of positive samples. The formulas are as follows:
[0180]
[0181]
[0182]
[0183]
[0184]
[0185]
[0186]
[0187]
[0188] Among them, TP is true positive, TN is true negative, FP is false positive, and FN is false negative. TPR represents the true positive rate, which is the proportion of positive samples correctly predicted by the model among the positive samples. FPR represents the false positive rate, which is the proportion of positive samples among the samples predicted as negative by the model.
[0189] Example 2
[0190] The main contribution of the present invention is to propose a joint diagnosis method for Parkinson's disease that integrates multi-scale information based on cross-modal consistency fusion and modal dynamic balance mechanism, and identifies Parkinson's disease through blood transcriptome data and magnetic resonance imaging information.
[0191] During the data fusion process, there is still a problem in the late fusion strategy of insufficient learning of the unique representation of each modality, which is called modality imbalance. This modality imbalance phenomenon indicates that some modalities may play a dominant role, resulting in reduced prediction accuracy and poor model generalization. For the two cases in AMP-PD, there is a modality imbalance phenomenon in multi-modal learning.
[0192] Although the cross-modal consistency fusion strategy makes it possible to capture cross-scale interactions, the cross-modal consistency fusion strategy embedded in the joint diagnosis method of Parkinson's disease based on cross-modal consistency fusion and modality dynamic balance mechanism still has the problem of modality imbalance.
[0193] To explore the problem of modality imbalance, the prediction performance obtained from the single-modal feature representation was compared. It seems that some modalities have a greater impact on the diagnostic results, indicating that heterogeneous modalities have different degrees of influence on decision-making. This verifies the necessity of the modality dynamic balance mechanism. The trajectory of the meta-parameters in the modality dynamic balance mechanism during data fusion is as Figure 2 shown.
[0194] Due to the inherent imbalance of heterogeneous representations, the contributions of blood transcriptomics and nuclear magnetic resonance imaging representations to multi-modal learning are different.
[0195] To analyze the function of the modality dynamic balance mechanism, the validation experiment further monitored the trajectory of the prediction performance and the difference ratio of single-modal features during the multi-modal learning process. As Figure 2 (a)(b) shows, the performance of the weak modality after modulation has been improved.
[0196] In addition, it is observed from Figure 2 that after applying the modality dynamic balance mechanism, the difference rate significantly decreases. This is the specific evidence of the effectiveness of the modality dynamic balance mechanism.
[0197] In the validation experiment, the Transformer method, the multi-modal attention-based method (MADDI), and the graph-based machine learning architecture (MOGONET) were compared. The proposed joint diagnosis method of Parkinson's disease based on cross-modal consistency fusion and modality dynamic balance mechanism achieved excellent performance in six evaluation aspects, especially in terms of accuracy and F1-score. The detailed information of the evaluation metrics during the method comparison has been listed in Table 2.
[0198] Table 2 Comparison of evaluation metrics of the joint diagnosis method of Parkinson's disease based on cross-modal consistency fusion and modality dynamic balance mechanism and candidate methods
[0199]
[0200]
[0201] In Table 2, the cross-scale learning method of the Parkinson's disease joint diagnosis method based on cross-modal consistency fusion and modal dynamic balance mechanism outperforms two state-of-the-art multi-modal learning architectures in the Parkinson's diagnosis task.
[0202] This phenomenon indicates that the cross-scale interaction between blood transcriptomics and neuroimaging-specific representations is valuable for improving diagnostic accuracy.
[0203] As SOTA multi-modal learning architectures, Transformer, MADDi, and MOGONET have demonstrated the basic ability to integrate blood transcriptomics and nuclear magnetic resonance imaging data to explore the cellular dynamics of neurodegeneration. However, these integrated frameworks still suffer from the problem of modal imbalance.
[0204] One possible explanation is that the modal imbalance between the two data modalities at the time scale increases the uncertainty of clinical decision-making. Using a hybrid information source, the Parkinson's disease joint diagnosis method based on cross-modal consistency fusion and modal dynamic balance mechanism outperforms MADDi and MOGONET in the Parkinson's staging task, verifying the advantages of the cross-modal consistency fusion strategy. This phenomenon indicates the necessity of considering cross-scale interactions.
[0205] The model performance is evaluated from six perspectives, including precision, F1 score, etc., as Figure 3 shown.
[0206] In the radar chart, the red line in each radar sub-chart represents the performance of the Parkinson's disease joint diagnosis method based on cross-modal consistency fusion and modal dynamic balance mechanism in the Parkinson's staging task. In addition, the shaded area in each sub-chart corresponds to the performance of the prediction model.
[0207] To discuss the efficiency of the cross-modal consistency fusion strategy and the modal dynamic balance mechanism, an ablation study was conducted on the PPMI and PDBP cases.
[0208] The cross-modal consistency fusion strategy combines modal-specific and modal-consistency components between the two representations. In the PPMI case, using Transformer as the backbone, the F1 score of the Parkinson's disease joint diagnosis method based on cross-modal consistency fusion and modal dynamic balance mechanism increased by 7.1% and 9.6% respectively compared with the scores obtained through transcriptome-specific and nuclear magnetic resonance imaging-specific representations. This proves the positive role of cross-scale interactions in exploring deep disease dynamics.
[0209] For the PDBP case, compared with the two unimodal scenarios, the Parkinson's disease joint diagnosis method based on cross-modal consistency fusion and modal dynamic balance mechanism improved the F1 score by 12.2% and 4.2%, further verifying the advantages of hybrid information sources. The evaluation metrics obtained in multiple scenarios are shown in Table 3.
[0210] Table 3 Ablation experiments of cross-modal consistency fusion strategy and modal dynamic balance mechanism in the Parkinson's disease joint diagnosis method based on cross-modal consistency fusion and modal dynamic balance mechanism
[0211]
[0212] The cross-modal consistency fusion strategy and the modal dynamic balance mechanism have been incorporated into the cross-scale Transformer backbone model. A significant performance improvement was observed after introducing the modal dynamic balance mechanism, with the F1 score increasing by 9.4% and 5.9% in the PPMI and PDBP cases, respectively. Both the cross-modal consistency fusion strategy and the modal dynamic balance mechanism improved the prediction accuracy, with the F1 score increasing by 13.6% and 10.8% on the PPMI and PDBP benchmarks, respectively. After determining the disease stage, Figure 4 The ROC and PR curves in multiple scenarios were compared. These scenarios cover two unimodal-based and three multimodal-based predictions.
[0213] According to Figure 4 , the combination of the cross-modal consistency fusion strategy and the modal dynamic balance mechanism showed the highest accuracy in the Parkinson's staging task.
[0214] According to the diagnostic results, the F1 score resulting from the direct concatenation strategy was even lower than that of the unimodal scheme. This phenomenon indicates that the modal imbalance between blood transcriptomics and magnetic resonance imaging specific characterizations leads to the uncertainty of joint learning. In this case, cross-scale interaction can obtain more reliable predictions.
[0215] The above embodiments are used to explain the present invention, rather than to limit the present invention. Any modifications and changes made within the spirit and scope of the claims of the present invention fall within the protection scope of the present invention.
Claims
1. A joint diagnosis method for Parkinson's disease based on cross-modal consistency fusion and modality dynamic balance mechanism, characterized in that: The method comprises the following steps: A. Preprocess the MRI data and use a multi-scale and multi-channel convolutional neural network to extract features from the MRI data and extract pathological features at different spatial and channel scales; B. Preprocess the blood transcriptomics data and use the multi-layer perceptron model to extract features from the blood transcriptomics data to obtain specific features related to gene expression levels; C. Use the cross-modal consistency fusion module to fuse MRI features and blood transcriptomics features, and extract the consistency between modalities through a multi-layer attention mechanism; D. Use the modal dynamic balance mechanism to dynamically calculate the contribution weights between modalities, balance the influence of different modalities in the joint representation, generate a cross-modal joint representation to obtain the final Parkinson's disease diagnosis result.
2. The Parkinson's disease joint diagnosis method based on cross-modality consistency fusion and modality dynamic balance mechanism according to claim 1 is characterized in that: In steps A and B, the data sources include but are not limited to: (1) MRI data from the Parkinson's Disease Progression Biomarkers Initiative (PPMI); (2) MRI data from the Parkinson's Disease Data and Biomarker Project PDBP; (3) Transcriptomic data of PPMI and PDBP in Accelerated Drug Partnership Parkinson's Disease AMP-PD.
3. The Parkinson's disease joint diagnosis method based on cross-modality consistency fusion and modality dynamic balance mechanism according to claim 1 is characterized in that: In steps A and B, data tags specifically include but are not limited to: (1) Parkinson's disease healthy control group, prodromal Parkinson's disease, Parkinson's disease with no abnormalities on scans, Parkinson's disease; (2) Stages 0-5 in the Hoehn-Yahr disease classification.
4. The Parkinson's disease joint diagnosis method based on cross-modality consistency fusion and modality dynamic balance mechanism according to claim 1 is characterized in that: In step A, a multi-scale multi-channel convolutional neural network is used to extract features from nuclear magnetic resonance imaging, and the specific steps are as follows: (1) Define the MRI data input as X M ; (2) Extraction of multi-scale features; in this step, let the input matrix X M P 0 , in order to obtain feature information at multiple scales; the corresponding convolution operation is expressed as: Y t =P i *K t ,t=1,2,…,T Among them, P i is the result of applying the multi-scale multi-channel convolutional neural network module i times. The symbol * represents the convolution operation, t represents the index of the convolution kernel function, and there are T convolution kernels in total. K t represents the tth convolution kernel function; Y t represents the feature map obtained after the tth convolution kernel function; then, the convolved feature maps are spliced together to obtain a multi-scale feature map Y M ; AND M =Concat(Y1,Y2,…,Y T ) Where Concat represents the feature concatenation operation; (3) Extraction of multi-channel features; in this step, the SE module models the dependencies between channels and adaptively recalibrates the channel feature responses, thereby enhancing the representation capability of the network; The SE module includes Squeeze compression and Excitation excitation; the Squeeze operation compresses the spatial information of each channel into a scalar by applying global average pooling: Among them, H and W are the height and width of the feature map respectively, C is the number of channels, c represents the channel, and h and w represent the pixel positions; Subsequently, the Excitation operation recalibrates the weights of each channel through two fully connected layers and the ReLU activation function: s c =σ(W2δ(W1)z c ),c=1,2,…,C Among them, W1 and W2 are weight matrices, δ represents the ReLU activation function, and σ represents the Sigmoid activation function; After the scaling and weighting operation, the learned weight s for each channel is c Multiply the corresponding channel of the previous feature map to get: in, Represents the multi-scale feature map Y M The part belonging to channel c; (5) After the above processing, the output is p i+1 is flattened and fed into a fully connected layer to obtain a neuroimaging-specific representation Among them, Fc represents the fully connected layer, Flatten represents the flattening operation, GAP refers to the global average pooling operation, and p I It is the output result of the last layer of the multi-scale multi-channel convolutional neural network module.
5. The Parkinson's disease joint diagnosis method based on cross-modality consistency fusion and modality dynamic balance mechanism according to claim 1 is characterized in that: In step B, a multi-layer perceptron model is used to extract features from blood transcriptomics data, and the specific formula is: h1=Dropout1(BatchNorm1(ReLU(W1X T +b1))) h2=Dropout2(BatchNorm2(ReLU(W2h1+b2))) h3=Dropout3(BatchNorm3(ReLU(W3h2+b3))) Among them, X T is the input matrix of blood transcriptomics data, W i and b i Represent the weight matrix and bias vector of the i-th layer respectively, ReLU represents the Leaky ReLU activation function; BatchNorm i Represents the batch normalization operation of the i-th layer; Dropout i represents the Dropout operation of the i-th layer; Softmax represents the Softmax function.
6. The Parkinson's disease joint diagnosis method based on cross-modality consistency fusion and modality dynamic balance mechanism according to claim 1 is characterized in that: In step C, a cross-modality consistency fusion module is used to fuse the MRI features and the blood transcriptomics features, and the consistency between the modalities is extracted through a multi-layer attention mechanism. The specific steps are as follows: (1) Define a one-side layer cross-modal consistency fusion module, whose specific formula is: Among them, LN(·) is a linear layer, MultiHead(·) is a multi-head attention mechanism, FFN is a feedforward network, is the intermediate result; This is the result obtained by a layer of cross-modal consistency fusion module; Magnetic Resonance Imaging Modality Data Representation As a query, Blood transcriptomics modality data characterization As keys and values, that is, The multi-head attention mechanism divides the input into multiple subspaces, applies attention calculation to each subspace, and then concatenates the results, and finally obtains the output through linear transformation; integrating the two formulas, we get: CMCFusion means integrating a single-layer cross-modal consistency fusion module into a function operation. (2) Similarly, characterize the blood transcriptomics modality data As query and MRI modality data representation As key and value, process and result X M→T It is expressed as: (3) Then, the two parts are integrated through cascading to obtain the result of cross-modal consistent representation: X consensus =Concat[X T→M ,X M→T ] 7. The Parkinson's disease joint diagnosis method based on cross-modality consistency fusion and modality dynamic balance mechanism according to claim 1 is characterized in that: In step D, a modal dynamic balance mechanism is used to dynamically calculate the contribution weights between modalities and balance the influence of different modalities in the joint representation. The specific steps are as follows: (1) Define the joint representation as: (2) Calculate the difference rate of each modality’s contribution to the learning objective: Among them, F1score represents the F1 score, y i represents the true label of the sample, f(·) represents the linear classifier, Indicates that among all the classifier outputs, take out the category index j with the largest predicted probability, B t is the batch size during training, Defined as The reciprocal of (3) Use The contribution difference between the blood transcriptomics modality and the MRI modality is dynamically adjusted to adaptively modulate the loss function by the following formula: Where u represents any mode of M and T.
8. The Parkinson's disease joint diagnosis method based on cross-modality consistency fusion and modality dynamic balance mechanism according to claim 1 is characterized in that: In step D, the step constructs a dynamic loss function as follows: (1) Define two modality-specific feature extractors as (2) Use the single-modal dataset to pre-train the two modality-specific feature extractors respectively, so that the model can fully learn the features of a single modality and obtain the pre-trained modality-specific feature extractors: in, and Represents the transferred unimodal-specific feature extractor. (3) Definition and represents the parameters of the final linear classifier, W M and W T They represent the projection matrix, represents the set of real numbers in set theory, and are the dimensions of the modal M and T features, respectively, and Class represents the number of categories; the logits output of the multimodal model is as follows: Among them, θ M and θ T It represents the parameters that need to be optimized during training, and b represents the bias. (4) Calculate the cross entropy loss of the joint representation: Among them, Class represents the number of categories, represents the cross entropy loss of the joint representation, X joint represents the joint representation, and N represents the number of samples. (5) Define the specificity loss of mode u as: in, represents the cross entropy loss of the specific representation of modality u; finally, the overall dynamic loss function is obtained: Among them, γ is an adjustable hyperparameter, L M , L T denote the specificity loss of modalities M and T respectively.