Brain disease prediction method based on double encoders and diffusion model
Through a triple contrastive learning framework based on dual encoders and diffusion models, the problems of insufficient data enhancement and feature extraction in resting-state functional magnetic resonance imaging in small sample scenarios are solved, the model generalization ability and cross-site adaptability of brain disease detection are improved, and stable prediction and visual diagnosis support is provided.
Patent Information
- Application Number
- CN202511277273.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-09-09
AI Technical Summary
Existing brain disease detection methods based on resting-state functional magnetic resonance imaging have problems such as a single data augmentation strategy, overfitting, and insufficient feature extraction capabilities in small sample scenarios. In particular, when the data distribution varies greatly across sites, the model's migration ability and adaptability are insufficient.
A triple contrastive learning framework based on dual encoders and diffusion models is adopted. The diffusion model is introduced to generate diverse samples, and the dual encoder is combined to extract multimodal features in the time and spatial dimensions. A triple contrastive learning mechanism is designed to optimize feature representation, thereby improving the model's generalization ability and classification performance.
It effectively alleviates the small sample problem, improves the adaptability and accuracy of the model in cross-site brain disease classification tasks, provides stable prediction results and reliable diagnostic basis, reduces the risk of misdiagnosis, and supports unified analysis of cross-institutional data and intuitive understanding of visualization results.
Smart Images

Figure CN120766980A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical artificial intelligence technology, and specifically to a brain disease prediction method based on a dual encoder and a diffusion model. Background Art
[0002] Brain disease detection is an important research area in neuroscience and clinical medicine. Constructing a brain functional network (BFN) based on blood oxygenation level-dependent (BOLD) signals from resting-state functional magnetic resonance imaging (fMRI) and extracting features from the original BOLD signal time series has become a key approach to brain disease diagnosis. Calculating the functional connectivity strength between brain regions using the Pearson correlation coefficient allows for the construction of brain networks and analysis of brain diseases. However, due to the high cost of acquiring brain disease data and the difficulty of labeling them, existing methods are prone to overfitting in small sample scenarios, resulting in insufficient model generalization. Furthermore, traditional feature extraction methods typically focus on single-dimensional information in either the spatial or temporal dimensions, making it difficult to fully capture the complex characteristics of brain diseases.
[0003] In recent years, contrastive learning, as a self-supervised learning method, has shown significant advantages in alleviating the lack of labeled data and improving feature representation capabilities. However, the application of existing contrastive learning methods in brain disease detection still faces many challenges. First, the design of data augmentation strategies is often not sophisticated enough, making it difficult to generate semantically diverse samples, which affects the robustness and generalization performance of the model. Secondly, current methods mostly focus on single-dimensional feature extraction and lack in-depth exploration of cross-dimensional feature interactions between the topological structure of brain functional networks and dynamic time series. Finally, the solution to the small sample problem is still imperfect, especially when the data distribution varies greatly across sites, the migration ability and adaptability of the model still need to be improved.
[0004] In order to solve the above problems, there is an urgent need for a new method that can extract brain disease features from multiple dimensions and enhance the generalization ability of the model. The present invention proposes a triple contrast learning framework based on dual encoders and diffusion models. By introducing the diffusion model, semantically preserved data enhancement is achieved, and dual encoders are used to extract the spatial features of the brain functional network and the temporal features of the BOLD signal respectively, thereby comprehensively capturing the multidimensional information of brain diseases. At the same time, by designing a triple contrast learning mechanism, the learning process of feature representation is optimized, and the performance of the model in downstream tasks is further improved. This method not only effectively alleviates the small sample problem, but also significantly improves the adaptability and accuracy of the model in cross-site brain disease classification tasks. Summary of the Invention
[0005] This paper addresses the challenges of existing resting-state functional magnetic resonance imaging-based brain disease detection methods, which suffer from a single data augmentation strategy, overfitting to small samples, and insufficient feature extraction. By doing so, we propose a brain disease detection method based on triple contrastive learning using a dual encoder and a diffusion model. This method introduces a diffusion model to generate diverse samples, combines the dual encoder with a method to extract multimodal features in both temporal and spatial dimensions, and designs a triple contrastive learning mechanism to optimize these features, thereby improving the model's generalization and classification performance.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a brain disease prediction method based on dual encoders and diffusion model, comprising the following steps:
[0007] S1. The brain functional images collected by functional magnetic resonance imaging equipment were preprocessed using the DPARSF toolbox in MATLAB as follows: time points; perform head motion and time slice correction on the image; remove the influence of ventricular, white matter signals and high-order effects of head motion; align the corrected image to the standard space; Temporal bandpass filtering to reduce the effects of heartbeat and respiration.
[0008] S2. Based on the automatic anatomical labeling atlas, the brain is divided into several brain regions. The BOLD signal of a given subject is in matrix form. ,in Indicates the The BOLD signal of each ROI, Indicates the number of time nodes, Indicates the number of ROIs.
[0009] S3. Calculate the resting-state functional connectivity strength between all ROIs using the Pearson correlation coefficient:
[0010]
[0011] in, Indicates the Brain area and Pearson correlation coefficient between brain regions. Thus, the functional connectivity matrix between brain regions was obtained. For a given node, the Pearson correlation coefficient between it and other nodes is regarded as its feature, so the node feature matrix of the graph can be expressed as ,in The matrix Row represents the brain regions In order to maintain strong functional connectivity and remove weak connections in the brain network, the matrix The elements greater than 0.3 are set to 1, and the remaining elements are set to 0. This gives the edge feature matrix of the brain network So the brain network is constructed. .
[0012] S4. Diffusion data enhancement module: First, let’s look at the noise addition process, which gradually adds Gaussian noise to the input data through a series of Markov chain steps, and finally Then it is gradually converted into a stable random noise distribution. Diffusion operations are performed independently on node features and edge features. Given the original brain network , A node matrix representing a brain network, represents the edge matrix of the brain network. Specifically, for each node, the node noise matrix is defined as , for each edge, the edge noise matrix is defined as .in, Here is the total number of diffusion steps. In order to directly Add noise to the final noise state , so the noise adding process can be defined as:
[0013]
[0014]
[0015] Here ,in The noise matrix is not an arbitrary matrix and needs to meet the following conditions: First, the sum of all rows of the noise matrix should be 1. Second, the sum of each column of the noise matrix is equal to 1. This means that the matrix It is also a Markov transition matrix (row sum is 1) and has a uniform steady-state distribution (column sum is 1). A common practice is to use a doubly random matrix to define the noise matrix :
[0016]
[0017] in, is the identity matrix, Represents the last dimension of the node or edge feature, is a column vector with all elements equal to 1, Following cosine scheduling:
[0018]
[0019] here It is a very small value. , It can be calculated according to the following formula:
[0020]
[0021]
[0022]
[0023] Next is the denoising process, in order to For denoising, a denoising neural network based on GraphTransformer was trained , the network predicts and verifies the clean image The network is composed of an input MLP, a GraphTransformer block, and an output MLP. The GraphTransformer module consists of a self-attention module, two fully connected layers, a layer normalization layer, and a ReLU activation function. Since simple node and edge features can only represent the local structural information of the graph, in order to enhance the expressive power of the model, a global feature that can represent the topological characteristics of the entire graph is introduced. , which is composed of the spectral features of the graph Laplacian matrix, specifically including: the number of connected components of the graph and the first five non-zero eigenvalues. Based on this design, the module can simultaneously process noise node features , noise edge features And global features First, the original features are mapped into , and , and then input these high-order features into the GraphTransformer module. The GraphTransformer module uses the self-attention mechanism to dynamically update the node features, and at the same time realizes the mutual fusion of node features, edge features and global features based on the FiLM layer, and updates the global features by aggregating node features and edge features, and finally generates new features , , These features can be input to the next GraphTransformer module for deeper feature abstraction, and finally generated by the output MLP .
[0024] Based on the prediction graph and a real graph with nodes and edges The mean square error between them is used to optimize the diffusion loss, and the loss function is defined as follows of:
[0025]
[0026] where, represents the true node features, represents the predicted node features, represents the true edge features, represents the predicted edge features.
[0027] Once the network is trained, it can generate new brain networks using it. Specifically, first a completely random graph is sampled, and then the trained denoising neural network is used to predict a clean graph. The predicted graph is then processed using the posterior distribution, which is used as input to the denoising network for the next time step, iteratively generating the final graph. To achieve this, the reverse diffusion iterations need to be estimated based on the denoising neural network . This distribution is represented as the product of the node and edge distributions:
[0028]
[0029] Using the denoising neural network to predict each node, for example, the formula is as follows:
[0030] ;
[0031] ;
[0032] The sampling process for edges is similar, with the distribution to edges These distributions are used to sample discrete , which will be the input to the denoising network for the next time step. After iterations, a new graph is generated.
[0033] S5. Dual encoder part, two different encoders are used to extract features of different data dimensions of resting state functional magnetic resonance imaging data, in order to extract the biological characteristics of the subject's BOLD signal in time, Mamba encoder is used to extract , and three matrices, the hidden state is calculated by recursive formula:
[0034]
[0035]
[0036] The output of all time steps is spliced along the time dimension to form a complete output sequence , and then the linear projections of the original input are fused via skip connections:
[0037]
[0038] After that, the extracted feature vector can be obtained through the linear layer .
[0039] The GIN encoder is used to extract the data features of the spatial dimension of the resting state functional magnetic resonance imaging data, respectively, with the original brain network and the brain network generated after diffusion model enhancement As input. For the original brain network The nodes Indicates the brain regions, side Represents a collection of edges. Similarly, the enhanced brain network is similar. To enhance the expressiveness of the model, a graph isomorphism network (GIN) is used as the basic feature encoder. The node feature update of each layer is completed by aggregating its neighbor information:
[0040]
[0041] in, represents the MLP operation, is a parameter, Indicates the Layer Node The eigenvector of Representation node The neighbor set of . In the model, the number of GIN layers is set to 2. Taking a single branch as an example, the original graph Input to GIN for feature learning, so the node feature matrix updated after two stacked GIN layers is It can be expressed as:
[0042]
[0043] in, yes activation function, and Represent the weight matrices of the first and second GIN layers learned by MLP respectively. Finally, based on the feature matrix obtained Perform average pooling operation to obtain graph-level feature vector . Similar will enhance the view It is also fed into the parameter-sharing GIN encoder to obtain the feature vector .
[0044] S6. Next is the triple contrastive learning module. In contrastive learning, the rational construction of positive and negative sample pairs is crucial. A triple contrastive learning mechanism was designed to extract features based on the different dimensions of the subject's fMRI data (original spatial features and temporal dynamic features, and enhanced spatial features and temporal dynamic features), as well as the same-dimensional enhanced contrast (original spatial features and enhanced spatial features).
[0045] Taking the original spatial features and temporal dynamic features as an example: the feature vectors of different dimensions of the same subject are regarded as positive sample pairs, and the feature vectors of different dimensions of different subjects are regarded as negative sample pairs. Specifically, according to the following formula:
[0046]
[0047]
[0048]
[0049] Here and The feature vectors representing the original brain network and BOLD signal, as well as the enhanced spatial features and temporal dynamic features, as well as the original spatial features and enhanced spatial features are the same as the above formulas:
[0050]
[0051]
[0052] Therefore, the overall loss function of the pre-trained model is:
[0053]
[0054] Here , , , is a hyperparameter.
[0055] S7. After the pre-training phase is completed, the trained spatial feature encoder (GIN) and temporal feature encoder (Mamba) are transferred to the downstream brain disease classification task. and BOLD signals are input into GIN and Mamba encoders respectively to obtain feature vectors and , then concatenate the two eigenvectors just obtained and record them as Then, MLP and Softmax layers are used to classify brain diseases, and the model is updated by the cross-entropy loss function.
[0056] The present invention provides a brain disease prediction method based on dual encoders and diffusion models. It has the following beneficial effects:
[0057] 1. This invention uses a diffusion model to generate semantically preserved, diverse, and enhanced samples, addressing the challenges of difficult data annotation and limited sample sizes for brain disease. Therefore, even when case data is limited, it can still provide doctors with stable and reliable prediction results, reducing the risk of misdiagnosis due to insufficient samples and providing an objective basis for preliminary screening.
[0058] 2. This invention uses dual encoders to extract the spatial topological features of brain functional networks and the temporal dynamic features of BOLD signals, respectively. Triple contrast learning enables cross-dimensional feature interaction. This simultaneously presents doctors with quantitative information on abnormal brain connectivity and temporal changes in neural activity, assisting them in comprehensively assessing the condition from multiple perspectives and avoiding missed diagnoses from a single perspective.
[0059] 3. The enhanced samples generated by this invention through the diffusion model retain the semantic information of the original data; the brain functional network directly reflects the strength of brain region connections. Physicians can combine the generated enhanced sample visualization results with the brain network connectivity map to intuitively understand the pathological basis predicted by the model. By integrating the algorithm results with clinical experience, the credibility of diagnostic decisions can be improved.
[0060] 4. This invention optimizes feature representation through diffusion enhancement and contrastive learning, significantly improving the model's generalization capabilities for data from different medical sites. It supports the integration of heterogeneous brain imaging data from multiple hospitals and devices, facilitating unified case analysis for doctors in cross-institutional consultations and avoiding model failures due to device differences.
[0061] 5. This method improves the discriminability and robustness of features through a triple contrastive learning mechanism. On the ABIDE dataset, metrics such as ACC and AUC significantly outperform traditional methods. It provides objective, quantitative classification results, reducing subjective judgment bias due to differences in physician experience. It is particularly suitable for borderline cases with atypical symptoms.
[0062] In summary, this method generates high-quality samples, integrates spatiotemporal features, improves model generalization, and enhances interpretability of results. It provides physicians with: a more reliable early screening tool; multi-dimensional pathology analysis views; cross-institutional data collaboration capabilities; and traceable quantitative diagnostic evidence. Ultimately, it assists physicians in making more comprehensive and accurate diagnostic decisions in scenarios with limited data, complex symptoms, or when multi-center collaboration is required. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 This is a diagram of the overall process framework of a brain disease prediction method based on dual encoders and diffusion models disclosed in this application;
[0064] Figure 2A process diagram for data enhancement of the diffusion model of the method in the embodiment of the present application is shown in the figure.
[0065] Figure 3 A structure diagram of the triple contrast learning mechanism of the method in the embodiment of the present application is shown in the figure.
[0066] Figure 4 A whole structure diagram of the denoising neural network in the embodiment of the present application is shown in the figure.
[0067] Figure 5 An internal structure diagram of the GraphTransformer module in the denoising neural network in the embodiment of the present application is shown in the figure.
[0068] Figure 6 A diagram of the global feature updating mechanism in the denoising neural network in the embodiment of the present application is shown in the figure.
[0069] Figure 7 A time dynamic feature extraction process diagram of the Mamba encoder in the embodiment of the present application is shown in the figure.
[0070] Figure 8 A whole flow diagram of the embodiment of the present application is shown in the figure. DETAILED DESCRIPTION
[0071] The technical solutions of the present application will be described clearly and completely below with reference to the accompanying drawings of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0072] Please refer to the accompanying drawings of the present application Figure 1 - the accompanying drawings of the present application Figure 8 The embodiment of the present application provides a brain disease prediction method based on double encoders and diffusion model, which specifically comprises:
[0073] Selecting a data set: in this example, a public autism resting-state functional magnetic resonance imaging (ABIDE) data set is used, and the data of the two largest sites, NYU and UM116, are selected; the NYU data set has a total of 184 subjects, including 79 autism patients and 105 healthy subjects; the UM116 data set has a total of 145 subjects, including 68 autism patients and 77 healthy subjects.
[0074] The brain functional images collected by the functional magnetic resonance imaging device were preprocessed using the DPARSF toolbox in Matlab as follows: the previous time point of the magnetic resonance image was removed; the image was corrected for head motion and time slices; the influence of ventricular and white matter signals and high-order effects of head motion was removed; the corrected image was registered to the standard space; and the subject image was subjected to temporal bandpass filtering to reduce the influence of heartbeat and breathing.
[0075] Divide the brain into several brain regions: Divide the brain into N regions of interest (ROIs) based on the standard brain network spectrum, and each ROI corresponds to a node in the graph , the node content is a matrix of Pearson correlation coefficients between the node and each other node. Then the functional connectivity calculation is performed: the BOLD signal of a given subject is in matrix form ,in Indicates the The BOLD signal of each ROI, Indicates the number of time nodes, Indicates the number of ROIs.
[0076] The Pearson correlation coefficient was used to calculate the resting-state functional connectivity strength between all ROIs:
[0077] in represents the Pearson correlation coefficient between the i-th brain region and the j-th brain region. Thus, the functional connectivity matrix between brain regions is obtained For a given node, the Pearson correlation coefficient between it and other nodes is regarded as its feature, so the node feature matrix of the graph can be expressed as ,in The matrix Row represents the brain regions In order to maintain strong functional connectivity and remove weak connections in the brain network, the matrix The elements greater than 0.3 are set to 1, and the remaining elements are set to 0. This gives the edge feature matrix of the brain network So the brain network is constructed. .
[0078] Next is the diffusion data enhancement module. First, let’s look at the noise addition process, which gradually adds Gaussian noise to the input data through a series of Markov chain steps. Then it is gradually converted into a stable random noise distribution. Diffusion operations are performed independently on node features and edge features. Given the original brain network , A node matrix representing a brain network, represents the edge matrix of the brain network. Specifically, for each node, the node noise matrix is defined as , for each edge, the edge noise matrix is defined as .in, Here is the total number of diffusion steps. In order to directly Add noise to the final noise state , so the noise adding process can be defined as: ,
[0079] Here ,in The noise matrix is not an arbitrary matrix and needs to meet the following conditions: First: the sum of all rows of the noise matrix should be 1. Second: the sum of each column of the noise matrix is equal to 1. This means that the matrix It is also a Markov transition matrix (row sum is 1) and has a uniform steady-state distribution (column sum is 1). A common practice is to use a doubly random matrix to define the noise matrix : in, is the identity matrix, Represents the last dimension of the node or edge feature, is a column vector with all elements equal to 1, Following cosine scheduling: ,here It is a very small value. , It can be calculated according to the following formula: , , .
[0080] Next is the denoising process, in order to For denoising, a denoising neural network based on GraphTransformer was trained , the network predicts and verifies the clean image Similar graphs are used to achieve denoising. Figure 4-Figure 6 As shown in Figure 1, the network consists of an input MLP, a GraphTransformer block, and an output MLP. The GraphTransformer module consists of a self-attention module, two fully connected layers, a layer normalization layer, and a ReLU activation function. Since simple node and edge features can only represent the local structural information of the graph, in order to enhance the expressive power of the model, a global feature that can represent the topological characteristics of the entire graph is introduced. , which is composed of the spectral features of the graph Laplacian matrix, specifically including: the number of connected components of the graph and the first five non-zero eigenvalues. Based on this design, the module can simultaneously process noise node features , noise edge features And global features First, the original features are mapped into , and , and then input these high-level features into the GraphTransformer module. Figure 5 As shown in , this module uses the self-attention mechanism to dynamically update node features, and at the same time realizes the mutual fusion of node features, edge features and global features based on the FiLM layer. Figure 6 As shown, the global features are updated by aggregating node features and edge features, and finally new features are generated. , , These features can be input to the next GraphTransformer module for deeper feature abstraction, and finally generated by the output MLP .
[0081] Based on the prediction graph and the real graph with nodes and edges The mean square error between them is used to optimize the diffusion loss, and the loss function is defined as follows of ,in It represents the node features of the real graph. It represents the predicted node features. represents the edge features of the real graph, It represents the predicted edge features.
[0082] Once the network is trained, sampling can be used to generate new brain networks Specifically, we first sample a completely random graph and then use the trained denoising neural network to predict a clean graph. We then use the posterior distribution to process the predicted graph, which serves as the input of the denoising network at the next time step, and iteratively generates the final graph. To achieve this, we need to Estimating backdiffusion iterations . Express this distribution as the product of the node and edge distributions: Taking nodes as an example, a denoising neural network is used to predict each node , , the edge sampling process is similar, the distribution of the edge These distributions are used to sample discrete , which will be the input of the denoising network at the next time step. Iterations, finally generating a new graph .
[0083] The next step is the dual encoder part. Two different encoders are used to extract features of different data dimensions of the resting-state functional magnetic resonance imaging data. In order to extract temporal features, the Mamba encoder is used to extract the BOLD signal biometric features of the subjects, such as Figure 7 As shown, created , and Three matrices, cyclically calculating the hidden state according to the recursive formula: , , the output of all time steps Splicing along the time dimension to form a complete output sequence , and then the linear projections of the original input are fused via skip connections: After that, the extracted feature vector can be obtained through the linear layer .
[0084] The GIN encoder is used to extract data features of the spatial dimension of resting-state functional magnetic resonance imaging data, such as Figure 1-Figure 3 As shown, the original brain network and the brain network generated after diffusion model enhancement As input. For the original brain network The nodes Indicates the brain regions, side Represents a collection of edges. Similarly, the enhanced brain network is similar. To enhance the expressiveness of the model, a graph isomorphism network (GIN) is used as the basic feature encoder. The node feature update of each layer is completed by aggregating its neighbor information: ,in represents the MLP operation, is a parameter, Indicates the Layer Node The eigenvector of Representation node The neighbor set of . In the model, the number of GIN layers is set to 2. Taking a single branch as an example, the original graph Input to GIN for feature learning, so the node feature matrix updated after two stacked GIN layers is It can be expressed as: ,in yes activation function, and Represent the weight matrices of the first and second GIN layers learned by MLP respectively. Finally, based on the feature matrix obtained Perform average pooling operation to obtain graph-level feature vector . Similar will enhance the view It is also fed into the parameter-sharing GIN encoder to obtain the feature vector .
[0085] Next is the triple contrastive learning module. In contrastive learning, the reasonable construction of positive and negative sample pairs is the key. Figure 1 As shown in the figure, a triple contrast learning mechanism is designed to extract features from different dimensions of the subject's fMRI data (original spatial features and temporal dynamic features, enhanced spatial features and temporal dynamic features) and the same-dimensional enhanced contrast (original spatial features and enhanced spatial features).
[0086] Taking the original spatial features and temporal dynamic features as an example: the feature vectors of different dimensions of the same subject are regarded as positive sample pairs, and the feature vectors of different dimensions of different subjects are regarded as negative sample pairs. Specifically, according to the following formula: , , , here and The feature vectors representing the original brain network and BOLD signal, as well as the enhanced spatial features and temporal dynamic features, as well as the original spatial features and enhanced spatial features are the same as the above formulas: , , therefore, the overall loss function of the pre-trained model is: , here , , , is a hyperparameter.
[0087] After contrastive learning, the trained Mamba encoder and GIN encoder are transferred to downstream tasks, and the trained spatial feature encoder (GIN) and temporal feature encoder (Mamba) are transferred to downstream brain disease classification tasks. Figure 3 As shown, only the original brain network and BOLD signals are input into GIN and Mamba encoders respectively to obtain feature vectors and , then concatenate the two eigenvectors just obtained and record them as Then, MLP and Softmax layers are used to classify brain diseases, and the model is updated by the cross-entropy loss function.
[0088] Specifically: All experiments in this application are implemented based on the PyTorch 2.5.1 deep learning framework. The hardware configuration of the experimental platform includes NVIDIA GeForce RTX 4060Ti (32GB video memory), 13th Gen Intel (R) Core (TM) i5-13500. Some of the important parameters are set as follows: the number of diffusion steps of the pre-trained model The number of GraphTransformerBlocks in the denoising neural network is set to 2. In the feature extractor, the GIN hyperparameter configuration is as follows: the input dimension is 116, which is consistent with the number of brain functional regions (ROIs), the hidden layer dimension is 64, and the number of GIN layers is 2. The main dimension of the Mamba module is 116, and the state space dimension is 14. The parameter group used in the final loss function is To train the pretrained model, a five-fold cross-validation strategy was used to evaluate the performance of the task-specific model. The dataset was randomly divided into five mutually exclusive subsets. In each round, one subset was selected as the validation set, and the remaining four subsets were used as the training set. This process was repeated five times, ensuring that each sample participated in validation at least once. The final performance metric was the average of the five experiments. The task-specific training was performed using a stochastic gradient descent (SGD) optimizer with an initial learning rate of 0.004, a weight decay of 0.0001, and a momentum of 0.9. The training batch size was 16, and the training epochs were 40. In this work, cross-site brain disease classification was performed. CL-MambaGIN was first pretrained using fMRI data from one site and then fine-tuned on other sites. ASD vs. HC classification was performed on the ABIDE dataset. CL-MambaGIN was pretrained using site NYU, and the model was fine-tuned and tested using sites UM116 and LEUVEN, respectively. Four metrics were used to measure the performance of each method: Accuracy (ACC), Area Under the Horizon (AUC), Sensitivity Enrichment (SEN), and SPE. The experimental results and the comparison methods are shown in Tables 1 and 2:
[0089] Table 1 Comparative experimental results of NYU →UM116 on the ABIDE dataset
[0090] Model ACC AUC SEN SPE Window_sliceWindow_warp 0.560.57 0.590.56 0.530.56 0.660.65 Node DroppingGCA 0.590.58 0.600.61 0.650.53 0.590.70 AD-GCLAuto-GCL 0.600.61 0.640.63 0.560.51 0.630.66 No Pretrain 0.56 0.56 0.54 0.61 CL-MambaGIN 0.67 0.69 0.63 0.64
[0091] Table 2 Comparative experimental results of NYU →LEUVEN on the ABIDE dataset
[0092] Model ACC AUC SEN SPE Window_sliceWindow_warp 0.590.57 0.600.59 0.510.53 0.660.54 Node DroppingGCA 0.610.61 0.650.63 0.520.58 0.650.62 AD-GCLAuto-GCL 0.630.64 0.600.63 0.560.64 0.580.54 No Pretrain 0.59 0.61 0.51 0.65 CL-MambaGIN 0.69 0.73 0.65 0.63
[0093] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A brain disease prediction method based on dual encoders and diffusion model, characterized in that: The following steps are involved: S1. Preprocess functional brain images acquired by fMRI, including removing the first few time points, correcting for head motion, correcting for temporal slices, removing signal interference, performing spatial registration, and performing bandpass filtering. S2. Delineate brain regions based on anatomical atlases and construct the subject's BOLD signal matrix. S3. Calculate the resting-state functional connectivity strength between brain regions, generate the node feature matrix and edge feature matrix of the brain functional network, and construct the original brain network; S4. Data augmentation of the original brain network by diffusion model: Perform the noise addition process: Gaussian noise is added to node features and edge features independently, and converted into a random noise distribution through the Markov chain step; Perform denoising: Use a GraphTransformer-based neural network to denoise the noisy graph structure and generate an enhanced brain network; S5. Use dual encoder to extract features: Extract temporal dynamic features from BOLD signals using the Mamba encoder; A graph isomorphism network (GIN) encoder is used to extract spatial topological features from the original and enhanced brain networks. S6. Design of triple contrastive learning mechanism: The first level: cross-dimensional comparison of original spatial features and temporal dynamic features; The second level: enhance the cross-dimensional comparison of spatial features and temporal dynamic features; The third level: the same-dimensional comparison between the original spatial features and the enhanced spatial features; S7. Transferring the trained dual encoder to downstream classification tasks: concatenating spatial features with temporal dynamic features to achieve brain disease prediction through the classification layer.
2. The brain disease prediction method based on dual encoder and diffusion model according to claim 1, characterized in that: In the step S3: Functional connectivity strength was calculated using the Pearson correlation coefficient; The functional connectivity matrix is binarized: connections with strength greater than 0.3 are retained as 1, and the rest are set to 0.
3. The brain disease prediction method based on dual encoder and diffusion model according to claim 1, characterized in that: The noise addition process in step S4 satisfies: The noise matrix is a doubly random matrix, whose row sum and column sum are both 1; Noise scheduling uses a cosine attenuation strategy to control noise intensity.
4. The brain disease prediction method based on dual encoder and diffusion model according to claim 1, characterized in that: The denoising process of step S4: The neural network includes an input MLP, a GraphTransformer module and an output MLP; Global features, including the number of connected components of the graph and the first five non-zero eigenvalues, are introduced and input into the Graph Transformer module together with node features and edge features.
5. The brain disease prediction method based on dual encoder and diffusion model according to claim 4, characterized in that: The Graph Transformer module: Update node features through self-attention mechanism; Fusion of node features, edge features, and global features based on the FiLM layer; Update the global features through the aggregation operation.
6. The brain disease prediction method based on dual encoder and diffusion model according to claim 1, characterized in that: In the step S5: The GIN encoder uses a two-layer graph isomorphism network to update features by aggregating neighbor node information; The GIN encoder parameters are shared between the original and enhanced brain networks.
7. The brain disease prediction method based on dual encoder and diffusion model according to claim 1, characterized in that: In the triple contrastive learning mechanism of step S6: Positive samples correspond to feature vectors of different dimensions of the same subject; Negative sample pairs are feature vectors of different subjects; The overall loss function is the weighted sum of the triple contrast losses.
8. The brain disease prediction method based on dual encoders and diffusion model according to claim 1, characterized in that: The downstream classification task of step S7 is: Only the original brain network and BOLD signal are used as input to the trained dual encoder; The output spatial feature vector is concatenated with the temporal feature vector and classified through the MLP and Softmax layers.
9. The brain disease prediction method based on dual encoders and diffusion model according to claim 1, characterized in that: The brain disease is autism spectrum disorder, and the dataset used is a resting-state functional magnetic resonance imaging dataset.
10. The brain disease prediction method based on dual encoders and diffusion model according to claim 1, characterized in that: Supports classification of brain imaging data across medical sites, including unified analysis of heterogeneous data collected by different devices.
Citation Information
Patent Citations
Small sample medical image segmentation method, system and device based on self-supervised learning and medium
CN116681667A
Autism disease prediction technology based on self-supervised graph convolution model
CN116797817A
Brain network feature classification system based on double collaborative learning and training method thereof
CN117918817A
Multi-behavior recommendation method based on diffusion contrast learning
CN118821839A
Brain network analysis method and system based on adaptive graph collaborative contrast learning
CN119418172A
Cited By
Neurological disease diagnosis system based on space-time attention and dynamic domain self-adaption
CN120932823A