Depression image classification method based on multi-modal dynamic graph convolutional neural network

By using a multimodal dynamic graph convolutional neural network, a sliding window method and a multi-channel spatial attention network are used to extract high-order spatiotemporal features from fMRI data. By combining self-attention mechanism and LSTM for spatiotemporal feature fusion, the problem of insufficient utilization of multimodal information in existing technologies is solved, and high-precision and robust depression classification is achieved.

CN120877005AActive Publication Date: 2025-10-31ZHEJIANG UNIV OF TECH
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511383645.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-10-31
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

Existing brain network analysis methods fail to adequately consider multimodal complementary information and dynamic spatiotemporal characteristics in the classification of depression, resulting in insufficient classification accuracy and poor robustness.

Method used

A multimodal dynamic graph convolutional neural network is employed to segment fMRI data using the sliding window method. High-order spatiotemporal topological features are extracted using a multi-channel spatial attention network. Spatiotemporal feature fusion is achieved by combining self-attention mechanism and LSTM. Multimodal features are aggregated through bilateral graph convolution. The consistency of multimodal predictions is constrained by an auxiliary classifier and KL divergence, thus achieving effective fusion of cross-modal information.

Benefits of technology

It improved the accuracy and robustness of depression classification, effectively identified important brain regions related to depression, and achieved a classification accuracy of 80.32% and an F1 score of 92.86%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877005A_ABST
    Figure CN120877005A_ABST
Patent Text Reader

Abstract

A depression image classification method based on a multi-modal dynamic graph convolutional neural network belongs to the field of medical image processing and artificial intelligence, and comprises the following steps: dividing a time sequence after functional magnetic resonance data preprocessing into a plurality of time windows by using a sliding time window method, and constructing a functional connection matrix of each window; capturing time information of a cross-time window by using a plurality of spatial attention contrast networks to obtain a high-order dynamic function brain network; and the cross-modal graph neural network and the cross-modal knowledge distillation are utilized to realize complementary information transmission between modals and modal fusion. According to the method, the dynamic high-order information of the multi-mode brain network can be fully utilized to realize effective classification of depression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical image processing and artificial intelligence, and specifically refers to an image classification method for depression based on a multimodal dynamic graph convolutional neural network. Background Technology

[0002] With the development of artificial intelligence and neuroimaging, the analysis and processing of magnetic resonance imaging data using neuroimaging analysis technology and brain network modeling can enable the early diagnosis and classification of depression. This can assist doctors in the qualitative and even quantitative analysis of depression-specific biomarkers, thereby improving the accuracy and reliability of depression diagnosis and classification.

[0003] Brain network analysis, based on graph theory, models the brain and reveals the complex relationships between brain regions, becoming an important tool in neuroscience research and the diagnosis of neurological diseases. Previous studies have shown that altered brain network connectivity patterns are associated with depression. Therefore, brain network analysis can provide an effective means for diagnosing depression and identifying biomarkers.

[0004] Currently, commonly used brain network analysis methods mainly include functional brain network analysis and structural brain network analysis. Most current functional brain network-based depression classifications consider static functional connectivity, ignoring the dynamic characteristics of connectivity fluctuations over time. This means they are limited to extracting isolated features within discrete time slices, making it difficult to comprehensively capture high-order spatiotemporal topological features over continuous time intervals. On the other hand, most brain network-based depression classification methods only consider single-modality brain networks, failing to consider the significant impact of multimodal complementary information in brain network analysis. Summary of the Invention

[0005] To address the issues of insufficient classification accuracy and poor robustness in existing brain network-based depression classification methods due to inadequate consideration of multimodal complementary information and dynamic spatiotemporal characteristics, this invention proposes a depression image classification method based on a multimodal dynamic graph convolutional neural network. This method can fully utilize the dynamic high-order information of multimodal brain networks to achieve effective classification of depression.

[0006] The technical solution adopted by this invention to solve its technical problem is: A method for classifying depression images based on a multimodal dynamic graph convolutional neural network includes the following steps: Step 1: The time series of functional magnetic resonance imaging (fMRI) data is divided into multiple non-overlapping time windows using the sliding window method. For each window, a functional connectivity matrix is ​​generated by calculating the correlation of time series between brain regions, thereby quantifying the functional connectivity between brain regions. Step 2: Extract high-order spatiotemporal topological features using a multi-channel spatial attention contrast network; Step 3: Introduce contrastive learning constraints. The fusion of spatiotemporal features is achieved through self-attention mechanism and LSTM. High-order dynamic functional networks (DBNs) are computed by scaling dot product attention. Step 4: Independently encode the input DBNs and structural brain network to generate modality-specific embedding features; Step 5: Aggregate the modal features of DBNs and the structure on the dynamic graph structure using bilateral graph convolution; Step 6: Apply auxiliary classifiers to the DBNs modalities, structural modalities, and multimodal joint features respectively, and output the single-modal prediction probability distribution. Then, use KL divergence to constrain the consistency between multimodal prediction and single-modal prediction; the model adopts progressive weight adjustment.

[0007] Furthermore, in step three, the in-window contrast loss By maximizing feature similarity across different augmented views within the same time window while minimizing similarity between different brain regions, these contrastive constraints enable the model to learn discriminative spatiotemporal features; embedding features from each window... The mapping is represented as a Query, Key, and Value matrix.

[0008] Furthermore, in step four, for the features corresponding to DBNs... Corresponding features of structural networks Calculate the cross-modal mapping matrix of the model ; then The matrix is ​​normalized to become a double random matrix, and finally the cross-modal correlation weight matrix is ​​obtained. The normalized matrix It is used for cross-modal feature mapping, which reflects the weight allocation during cross-modal fusion.

[0009] Furthermore, in step six, the auxiliary classifier of the DBNs modality receives the features after graph convolution. The structural modality auxiliary classifier receives features after graph convolution. The auxiliary classifier for multimodal joint features receives the features after graph convolution. The three modes output the single-mode prediction probability distribution respectively. , and The auxiliary classifier consists of a fully connected layer and a SoftMax function.

[0010] Preferably, in step six, the model employs a gradual weight adjustment process as follows: weight of distillation loss With training rounds Linear growth, in the early stages of training , Approaching 0, the model depends on calculating multimodal prediction results. With real labels Cross-entropy loss Come for training and learning.

[0011] In step two, for each time window, two views are generated using a random mask. and Time window and Attention coefficient The calculation is as follows: ; in, The weight matrix is ​​a learnable matrix. For attention vectors, The feature concatenation operation involves processing the attention coefficients using a random mask. Finally, the node features are updated through weighted aggregation: ; in, The weight matrix is ​​a learnable matrix. This is the attention vector.

[0012] The technical concept of this invention is as follows: the time series of preprocessed fMRI data is divided into multiple time windows using a sliding time window, a functional brain network is constructed for each window, multiple spatial attention contrast networks are used to capture time information across time windows to obtain a dynamic functional brain network; and cross-modal graphical neural networks and cross-modal knowledge distillation are used to realize complementary information transmission and modality fusion between modalities, thereby improving the accuracy of depression classification.

[0013] The beneficial effects of this invention are as follows: By utilizing multimodal brain networks, a depression image classification method based on spatiotemporal contrastive learning and multimodal brain network fusion is proposed. This method is a robust, information-rich, and highly accurate depression classification method, which improves the classification accuracy of depression and effectively identifies important brain regions related to depression. Attached Figure Description

[0014] Figure 1 This is a flowchart of an image classification method for depression based on a multimodal dynamic graph convolutional neural network. Detailed Implementation

[0015] The invention will now be further described with reference to the accompanying drawings.

[0016] Reference Figure 1A method for classifying depression images based on a multimodal dynamic graph convolutional neural network includes the following steps: Step one involves using a sliding window method to segment the time series of functional magnetic resonance imaging (fMRI) data into multiple non-overlapping time windows. For each window, a functional connectivity matrix is ​​generated by calculating the correlation between time series data between brain regions, thereby quantifying the functional connectivity between brain regions. This step transforms continuous fMRI signals into a dynamic functional connectivity matrix. Indicates within the time window The whole-brain functional connectivity pattern within the brain, For the number of brain regions, Total number of windows; Step two involves extracting high-order spatiotemporal topological features using a multi-channel spatial attention contrastive network. The core of this network is a graph attention mechanism (GAT), designed to adaptively aggregate information from surrounding nodes. For each time window, two views are generated using a random mask. and Time window and Attention coefficient The calculation is as follows: ; in, The weight matrix is ​​a learnable matrix. For attention vectors, This represents the feature concatenation operation. Attention coefficients are processed using a random mask. Finally, the node features are updated through weighted aggregation: ; in, The weight matrix is ​​a learnable matrix. It serves as an attention vector; by aggregating dynamic brain functional networks from different views, it not only preserves the topological structure of the brain network, but also dynamically adjusts the importance between nodes through an attention mechanism.

[0017] Step 3: To further improve the discriminative power of the features, a contrastive learning constraint is introduced, using in-window contrastive loss. By maximizing feature similarity across different augmented views within the same time window, while minimizing similarity between different brain regions, the intra-window contrast loss is defined as follows: ; in, To different views and The cosine similarity between the embedding vectors corresponding to the time window. Cosine similarity between embedding vectors of different brain regions in the same view. The constant is then used; subsequently, inter-window contrast loss is used to further capture the spatiotemporal dependencies across time windows, for the same brain region. In different windows and The inter-window comparison loss constraint is as follows: ; in, and brain region These contrastive constraints, centered on the embedding nodes corresponding to different time windows, enable the model to learn discriminative spatiotemporal features. The fusion of these features is achieved through a self-attention mechanism and LSTM. The embedding features of each window are then... The mapping is done as a Query, Key, and Value matrix, and finally, higher-order dynamic function networks (DBNs) are computed using scaled dot product attention. ; Step four involves independently encoding the input DBNs and the structural brain network to generate modality-specific embedding features for the DBNs. Corresponding features of structural networks The cross-modal mapping matrix of the model The calculation is as follows: ; In the cross-modal mapping matrix, elements DBNs nodes were quantized. and structural nodes The strength of the association. To ensure that the mapping relationship between DBNs and structural connections conforms to the probability distribution characteristics, then... Normalization is performed to make it a double random matrix (the sum of each row and each column is 1), finally yielding the cross-modal correlation weight matrix. The normalized matrix It is used for cross-modal feature mapping, reflecting the weight allocation during cross-modal fusion, based on and The calculation process for the mapping feature from one mode to another is as follows: ; Step 5: Aggregate the modal features of DBNs and the structure on the dynamic graph structure using bilateral graph convolution. The graph convolution operation utilizes normalized... As a dynamic adjacency matrix to transmit cross-modal information, for the ... The modality update formula for convolutional layers is as follows: ; in, For learnable matrices, As the activation function, through multiple stacked graph convolutions, the model progressively aggregates multi-hop neighbor information, capturing the local and global dependencies between functional and structural networks. Multimodal joint features are defined as the top-level features corresponding to the DBN modalities and structural modalities. and splicing: ; Step 6: Apply auxiliary classifiers to the three modalities respectively. The auxiliary classifier for the DBN modality receives the features after graph convolution. The structural modality auxiliary classifier receives features after graph convolution. The auxiliary classifier for multimodal joint features receives the features after graph convolution. The three modes output the single-mode prediction probability distribution respectively. , and Classifiers typically consist of fully connected layers and a SoftMax function. ; Subsequently, the consistency between multimodal and single-modal predictions is constrained by KL divergence. Specifically, multimodal joint prediction is first enforced. Predictions that approximate DBN modes and structural modes: ; ; Subsequently, cross-modal alignment is used to further constrain the predictive similarity between DBN modes and structural modes: ; The total distillation loss is: ; To reduce the impact of randomness in single-modal prediction results during the initial training phase, the model employs progressive weight adjustment. Specifically, the weights of the distillation loss are adjusted accordingly. With training rounds Linear growth, in the early stages of training ( ), The value is close to 0, and the model mainly relies on calculating multimodal prediction results. With real labels Cross-entropy loss To train and learn, cross-entropy loss The calculation process is as follows: ; in This represents the number of participants. As training progresses... As distillation losses gradually increase, they become more dominant, forcing single-mode predictions to align with multi-mode predictions. The self-distillation loss function of the wheel is expressed as: ; During model training, the total loss function The final representation is as follows: .

[0018] The implementation process of this embodiment is as follows: The dataset used in this invention is the Guangzhou Medical University dataset. All subjects' multimodal magnetic resonance imaging data and demographic data were collected in the Department of Radiology, Affiliated Brain Hospital of Guangzhou Medical University. All subjects included basic demographic data, Hamilton Depression Rating Scale scores, and magnetic resonance imaging data of three modalities: structural magnetic resonance imaging data, resting-state functional magnetic resonance imaging data, and diffusion tensor imaging data. The dataset was divided into training, test, and validation sets using a five-fold cross-validation method. All samples were randomly divided into five non-repeating subsets, of which three subsets were used as the training set for model parameter learning, and the remaining two subsets were used for model hyperparameter tuning and the test set, respectively. In this invention, the model was trained on the training set, and the model hyperparameters were adjusted on the validation set using two evaluation indicators: classification accuracy and F1 score. Finally, the multi-window functional connectivity matrix and structural brain network of each subject were input into the test set to assist in determining whether each subject was ill and to calculate the classification accuracy and F1 score of the test set. To evaluate the classification performance of the model, the proposed method was compared with common machine learning methods, deep learning methods that have achieved good performance in other studies, and multimodal feature fusion methods. Machine learning methods included Support Vector Machine (SVM) and Multilayer Perceptron (MLP), while the deep learning method used was BrainGNN. The multimodal feature fusion method used for comparison was M-GCN. In a multimodal self-collected depression dataset, the proposed method outperformed the other methods, achieving a classification accuracy of 80.32% and an F1 score of 92.86%.

[0019] The embodiments described in this specification are merely examples of implementations of the inventive concept and are for illustrative purposes only. The scope of protection of this invention should not be considered limited to the specific forms described in these embodiments; rather, it extends to equivalent technical means conceived by those skilled in the art based on the inventive concept.

Claims

1. A method for classifying images of depression based on a multimodal dynamic graph convolutional neural network, characterized in that, Includes the following steps: Step 1: The time series of functional magnetic resonance imaging data is divided into multiple non-overlapping time windows using the sliding window method. For each window, the functional connectivity matrix between brain regions is generated by calculating the correlation of time series between brain regions, thereby quantifying the functional connectivity between brain regions. Step 2: Extract high-order spatiotemporal topological features using a multi-channel spatial attention contrast network; Step 3: Introduce contrastive learning constraints. The fusion of spatiotemporal features is achieved through self-attention mechanism and LSTM. High-order dynamic function networks (DBNs) are computed by scaling dot product attention. Step 4: Independently encode the input DBNs and structural brain network to generate modality-specific embedding features; Step 5: Aggregate the modal features of DBNs and the structure on the dynamic graph structure using bilateral graph convolution; Step 6: Apply auxiliary classifiers to the DBNs modalities, structural modalities, and multimodal joint features respectively, and output the single-modal prediction probability distribution. Then, use KL divergence to constrain the consistency between multimodal prediction and single-modal prediction; the model adopts progressive weight adjustment.

2. The image classification method for depression based on a multimodal dynamic graph convolutional neural network as described in claim 1, characterized in that, In step three, the in-window contrast loss By maximizing feature similarity across different augmented views within the same time window while minimizing similarity between different brain regions, these contrastive constraints enable the model to learn discriminative spatiotemporal features. Embedded features of each window The mapping is represented as a Query, Key, and Value matrix.

3. The image classification method for depression based on a multimodal dynamic graph convolutional neural network as described in claim 1 or 2, characterized in that, In step four, for the features corresponding to DBNs Corresponding features of structural networks Calculate the cross-modal mapping matrix of the model ; then The matrix is ​​normalized to become a double random matrix, and finally the cross-modal correlation weight matrix is ​​obtained. The normalized matrix It is used for cross-modal feature mapping, which reflects the weight allocation during cross-modal fusion.

4. The image classification method for depression based on a multimodal dynamic graph convolutional neural network as described in claim 1 or 2, characterized in that, In step six, the auxiliary classifier of the DBNs modality receives the features after graph convolution. The structural modality auxiliary classifier receives features after graph convolution. The auxiliary classifier for multimodal joint features receives the features after graph convolution. The three modes output the single-mode prediction probability distribution respectively. , and The auxiliary classifier consists of a fully connected layer and a SoftMax function.

5. The image classification method for depression based on a multimodal dynamic graph convolutional neural network as described in claim 4, characterized in that, In step six, the model employs a gradual weight adjustment process as follows: weighting the distillation loss. With training rounds Linear growth, in the early stages of training , Approaching 0, the model depends on calculating multimodal prediction results. With real labels Cross-entropy loss Come for training and learning.

6. The image classification method for depression based on a multimodal dynamic graph convolutional neural network as described in claim 1 or 2, characterized in that, In step two, for each time window, two views are generated using a random mask. and Calculate the time window and Attention coefficient Attention coefficients are processed using random masks, and finally, node features are updated through weighted aggregation.

Citation Information

Patent Citations

  • Depression classification method based on graph embedding and multi-modal brain network

    CN113255728A

  • Multi-modal emotion recognition method based on spatial-temporal feature fusion

    CN113935435A

  • Method for identifying depression of dwarf weasel through optimization algorithm based on music-induced electrocardio

    CN117796807A

  • Multi-source remote sensing image collaborative classification method and system

    CN118365924A

  • Multi-modal depression intelligent analysis method oriented to modal deficiency

    CN118503852A