Dynamic graph convolution electroencephalogram depression detection method based on spatial-temporal feature fusion
Through the GRU and TSCN dual-branch parallel processing technology, combined with the attention mechanism, the spatiotemporal characteristics of EEG signals are adaptively integrated to solve the problem of insufficient modeling of dynamic interactions between brain regions and improve the accuracy and robustness of depression detection.
Patent Information
- Application Number
- CN202510938966.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-10-28
AI Technical Summary
Existing EEG depression recognition methods have problems such as insufficient modeling of dynamic interactions between brain regions, insufficient fusion of multi-scale features, and weak model generalization ability.
The GRU and TSCN dual-branch parallel processing technology is adopted, combined with the attention mechanism, to adaptively fuse the spatiotemporal features of EEG signals, and capture the dynamic changes of brain functional networks through dynamic graph convolutional networks.
It significantly improves the accuracy and robustness of depression detection, enhances the ability to extract depression-related information, and improves the generalization ability of the model.
Smart Images

Figure CN120853885A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of brain-computer interfaces, and more particularly to a dynamic graph convolutional EEG method for detecting depression based on spatiotemporal feature fusion. Background Technology
[0002] Major Depressive Disorder (MDD) is a prevalent mental disorder that severely impairs patients' social functioning. Accurate diagnosis is crucial for developing effective treatment strategies and improving prognosis. However, current clinical practice primarily relies on clinical interviews and patient self-report scales (such as the Hamilton Depression Rating Scale and the Beck Depression Rating Scale) based on the Diagnostic and Statistical Manual of Mental Disorders (DSM) or the International Classification of Diseases (ICD). This diagnostic model has significant limitations: it is highly subjective and easily influenced by physician experience, patient communication and cooperation, and subjective reporting bias. Electroencephalography (EEG) technology records the electrophysiological activity of the cerebral cortex non-invasively, providing an objective means of monitoring the state of the nervous system. EEG signals contain rich physiological and psychological information, making them a promising biomarker and diagnostic tool for neurological diseases such as depression, epilepsy, and Alzheimer's disease, as well as in the field of emotion analysis.
[0003] Traditional machine learning methods heavily rely on manual feature engineering, meaning researchers must manually design and select features based on domain knowledge and experience. This process is not only time-consuming and labor-intensive, but also makes it difficult to guarantee that the extracted features can fully capture the complex nonlinear patterns in EEG signals that are potentially associated with depression. Limitations in feature design often lead to insufficient model generalization ability, posing challenges to the model's stability and accuracy in real-world clinical scenarios.
[0004] In recent years, deep learning has demonstrated tremendous potential in the field of biomedical signal analysis due to its powerful automatic feature learning capabilities, its ability to handle high-dimensional data, and its flexible model architecture. However, existing deep learning-based EEG studies on depression often focus on extracting features from a single modality (such as focusing only on one of the time-domain waveform, frequency-domain power, or spatial topology information), failing to effectively integrate multi-dimensional information. This single-perspective feature representation is insufficient to comprehensively characterize the complex psychological state of depression, limiting the model's recognition performance and hindering a deeper understanding of the disease mechanisms. Summary of the Invention
[0005] This invention provides a dynamic graph convolutional EEG depression detection method based on spatiotemporal feature fusion, aiming to address the problems of insufficient dynamic interaction modeling of brain regions, inadequate multi-scale feature fusion, and weak model generalization ability in existing EEG depression recognition methods. This method uses extracted power spectral density as input features and employs a dual-branch parallel processing technique of GRU and TSCN to extract sequence history information features and multi-scale depth information features from EEG signals, respectively, and utilizes an attention mechanism to achieve adaptive fusion. The fused features are then used as node inputs to a dynamic graph convolutional network, which adaptively captures the dynamic changes of brain functional networks through a trainable adjacency matrix. This method comprehensively enhances the extraction capability of depression-related information in the temporal, frequency, and spatial dimensions, effectively improving the accuracy and robustness of depression detection.
[0006] The technical solution of the present invention includes the following steps:
[0007] Step 1: Preprocess the input raw EEG data, including signal truncation, feature extraction, and standardization.
[0008] Step 2: The gated recurrent unit (GRU) model is used to capture the long-term temporal dependence of EEG signals, and EEG historical information is learned through its gating mechanism;
[0009] Step 3: Introduce a separable convolutional network (TSCN) into the residual block, combining dilated convolution and separable convolution to capture multi-scale depth information;
[0010] Step 4: Introduce an attention mechanism to dynamically adjust the feature weights extracted by GRU and TSCN, focusing on the feature expressions that are most discriminative for depression classification;
[0011] Step 5: Use a trainable adjacency matrix to parameterize brain region connectivity patterns, and use the gradient descent algorithm to adaptively optimize connection weights to learn dynamic brain region interaction patterns related to depression.
[0012] Step 6: Perform global max pooling on the output of the graph convolutional layer, then perform feature transformation and dimensionality compression through a fully connected layer, and finally generate the final depression classification result through a softmax layer.
[0013] Furthermore, step 1 involves preprocessing the input raw EEG data, including signal truncation, feature extraction, and standardization; specifically:
[0014] Signal slicing technology is used to divide the original EEG signal into segments of fixed duration (1 second) at integer multiples of the sampling rate, generating new samples. The segmented EEG signal is then further divided into L overlapping sub-segments using the Welch method. The sub-segment length N is taken as the maximum allowable value of the signal; the overlap length between adjacent sub-segments is fixed at M = N / 2 (i.e., 50% overlap); the starting position of the l-th segment is (l-1)(NM), and its data sequence is defined as:
[0015] x l [n]=x[n+(l-1)(NM)],n=0,...,N-1,l=1,2,......,L,
[0016] A Hamming window function ω(n) is applied to each segment to reduce spectral leakage, and the normalization factor of the window function is calculated:
[0017]
[0018] Calculate the periodicity graph based on windowed segment data:
[0019]
[0020] Where k = 0, 1, ..., N-1 represents the frequency index. The power spectral density estimate is obtained by averaging the periodograms of all segments:
[0021]
[0022] Four depression-related frequency bands—delta (1-4Hz), theta (4-8Hz), alpha (8-14Hz), beta (14-30Hz), and gamma (30-48Hz)—were extracted, and the extracted frequency band power features were normalized using the z-score method.
[0023] Furthermore, step 2 involves using a gated recurrent unit (GRU) model to capture the long-term temporal dependencies of EEG signals and learning historical EEG information through its gating mechanism; specifically:
[0024] Input the current time feature x t Compared to the previous hidden state h t-1 Input GRU cells and compute update gate z using learnable parameters. t and reset door r t :
[0025] z t =σ(W z x t +U z h t-1 )
[0026] r t =σ(W r x t +U r h t-1 )
[0027] Where σ(·) is the sigmoid activation function, W z W r ∈R c×d U is the input transformation weight matrix; z U r ∈R c×c The weight matrix is passed to the state; b z ,b r ∈R c This is the bias vector.
[0028] Generate candidate hidden states by combining the output of the reset gate:
[0029]
[0030] Where ⊙ represents element-wise multiplication, W t ∈R c×d ,U∈R c×c b∈R c The candidate state generation parameter matrix is constructed.
[0031] By updating the gate fusion historical state and candidate state:
[0032]
[0033] When z t →At time 0, all historical information is fully preserved (h) t ≈h t-1 When z t →The new state is fully adopted at 1 o'clock. Achieve adaptive information flow control.
[0034] Map the hidden state to the output feature:
[0035] f t GRU =ReLU(W o h t +b o )
[0036] Where W0∈R c×c b0∈R c For the output layer parameters, ReLU activation enhances the nonlinear characterization capability.
[0037] Furthermore, step 3 involves introducing a separable convolutional network (TSCN) into the residual block, combining dilated convolution with separable convolution to capture multi-scale depth information; specifically:
[0038] The TSCN consists of multi-layered cascaded improved residual modules, each containing three levels of feature processing units. The first level employs a causal dilated convolution unit with 64 5×1 convolution kernels and a dilation coefficient of 1. The output is processed by a batch normalization layer and then activated by a ReLU function, with Dropout regularization applied at a preset probability. The second level is a separable dilated convolution unit, comprising two processing stages: depthwise convolution and pointwise convolution. The depthwise convolution uses 32 3×1 convolution kernels for channel-independent processing with a dilation coefficient of 2, while the pointwise convolution uses a 1×1 convolution kernel to achieve cross-channel feature fusion. The third level is a residual connection unit, which performs dimensionality transformation using a 1×1 convolution kernel and adds it to the main path features, finally outputting after activation by a ReLU function.
[0039] In terms of parameter configuration, the network adopts a three-layer residual module cascade structure, with the expansion coefficient of each module increasing geometrically in the order of 1, 2, and 4. The number of network filters adopts a decreasing configuration strategy of 64-32-32, and the convolution kernel size decreases sequentially with the network depth to 5×1, 3×1, and 2×1.
[0040] Furthermore, step 4 involves introducing an attention mechanism to dynamically adjust the feature weights extracted by GRU and TSCN, focusing on the feature expressions most discriminative for depression classification; specifically:
[0041] The attention layer receives GRU features f t GRU and TSCN features f t TSCN splicing result Calculate attention weights:
[0042]
[0043] Among them, W a It is a learnable weight matrix, b a It is the bias vector, η T is the parameter vector obtained during training, and tanh is the activation function.
[0044] The calculated attention weights are used to dynamically weight and fuse the features of the two branches:
[0045]
[0046] Furthermore, step 5 involves: using a trainable adjacency matrix to parameterize brain region connectivity patterns, adaptively optimizing connection weights using a gradient descent algorithm, and learning dynamic brain region interaction patterns related to depression; specifically:
[0047] Define a loss function based on cross-entropy cost, and iteratively update the network parameters using the adjacency matrix through backpropagation (BP) until the optimal solution A is obtained:
[0048] Loss = cross_entropy(l,l) p )+α||Θ||
[0049] Where l represents the true label vector, l p This represents the label vector predicted by the model. An L2 regularization term α‖Θ‖ is added to prevent overfitting, where Θ includes the adjacency matrix parameters and other model parameters, and α is a weighting coefficient.
[0050] Calculate the partial derivative of the loss function with respect to A. Update A using gradient descent:
[0051]
[0052] Where ρ is the learning rate, and the update process continues until the loss converges or the preset number of iterations is reached. The input to GCN is the feature matrix f. A Given the adjacency matrix A, the interlayer propagation of GCN is as follows:
[0053]
[0054] Where σ is a nonlinear activation function, and W is the weight matrix to be trained. yes The degree matrix, I is the identity matrix.
[0055] Furthermore, step 6 involves performing global max pooling on the output of the graph convolutional layer, then implementing feature transformation and dimensionality compression through a fully connected layer, and finally generating the final depression classification result via a softmax layer; specifically:
[0056] For feature f A+1 Perform global max pooling to generate a vector of maximum values for node feature dimensions:
[0057] f pool =maxpool(f A+1 )
[0058] The pooled feature vectors are passed through a two-layer fully connected network and a softmax activation function to output the depression classification probability:
[0059]
[0060] Where W1 and W2 are weight matrices, and b1 and b2 are bias vectors.
[0061] The advantages and beneficial effects of this invention are as follows:
[0062] This invention employs a dual-branch parallel processing technique using GRU and TSCN to extract sequence history information and multi-scale deep features from EEG signals, respectively, and utilizes an attention mechanism to achieve adaptive fusion. This multi-scale feature fusion strategy can fully exploit the rich spatiotemporal features in EEG signals, significantly improving feature expressiveness and providing more accurate feature input for depression detection. Simultaneously, it effectively addresses the problem of insufficient dynamic interaction modeling of brain regions in existing EEG-based depression identification methods. The model can adaptively adjust and optimize, thereby better capturing changes in brain functional networks. Furthermore, this invention comprehensively enhances the extraction capability of depression-related information across temporal, frequency, and spatial dimensions, making the model more robust under different individuals and experimental conditions, and improving its generalization ability. Attached Figure Description
[0063] Figure 1 The flowchart of a dynamic graph convolutional EEG depression detection method based on spatiotemporal feature fusion provided by the present invention.
[0064] Figure 2 This is a diagram of the TSCN network structure. Detailed Implementation
[0065] The technical solutions of the embodiments of the present invention will be clearly and thoroughly described below with reference to the accompanying drawings. The described embodiments are merely some embodiments of the present invention.
[0066] As shown in the figure, the dynamic graph convolutional EEG depression detection method based on spatiotemporal feature fusion provided in this embodiment includes the following steps:
[0067] Step 1: Preprocess the input raw EEG data, including signal truncation, feature extraction, and standardization.
[0068] Step 2: The gated recurrent unit (GRU) model is used to capture the long-term temporal dependence of EEG signals, and EEG historical information is learned through its gating mechanism;
[0069] Step 3: Introduce a separable convolutional network (TSCN) into the residual block, combining dilated convolution and separable convolution to capture multi-scale depth information;
[0070] Step 4: Introduce an attention mechanism to dynamically adjust the feature weights extracted by GRU and TSCN, focusing on the feature expressions that are most discriminative for depression classification;
[0071] Step 5: Use a trainable adjacency matrix to parameterize brain region connectivity patterns, and use the gradient descent algorithm to adaptively optimize connection weights to learn dynamic brain region interaction patterns related to depression.
[0072] Step 6: Perform global max pooling on the output of the graph convolutional layer, then perform feature transformation and dimensionality compression through a fully connected layer, and finally generate the final depression classification result through a softmax layer.
Claims
1. A dynamic graph convolutional EEG method for detecting depression based on spatiotemporal feature fusion, characterized in that, Includes the following steps: Step 1: Preprocess the input raw EEG data, including signal truncation, feature extraction, and standardization. Step 2: The gated recurrent unit (GRU) model is used to capture the long-term temporal dependence of EEG signals, and EEG historical information is learned through its gating mechanism; Step 3: Introduce a separable convolutional network (TSCN) into the residual block, combining dilated convolution and separable convolution to capture multi-scale depth information; Step 4: Introduce an attention mechanism to dynamically adjust the feature weights extracted by GRU and TSCN, focusing on the feature expressions that are most discriminative for depression classification; Step 5: Use a trainable adjacency matrix to parameterize brain region connectivity patterns, and use the gradient descent algorithm to adaptively optimize connection weights to learn dynamic brain region interaction patterns related to depression. Step 6: Perform global max pooling on the output of the graph convolutional layer, then perform feature transformation and dimensionality compression through a fully connected layer, and finally generate the final depression classification result through a softmax layer.
2. The dynamic graph convolutional EEG depression detection method based on spatiotemporal feature fusion according to claim 1, characterized in that, Step 1: Preprocessing the input raw EEG data, including signal truncation, feature extraction, and standardization; specifically: Signal slicing technology is used to divide the original EEG signal into segments of fixed duration (1 second) at integer multiples of the sampling rate, generating new samples. The segmented EEG signal is then further processed by dividing each segment into L overlapping sub-segments using the Welch method. The sub-segment length N is taken as the maximum allowable value of the signal; the overlap length between adjacent sub-segments is fixed at M = N / 2 (i.e., 50% overlap); the starting position of the l-th segment is (l-1)(NM), and its data sequence is defined as: x l [n]=x[n+(l-1)(N-M)],n=0,......,N-1,l=1,2,......,L, A Hamming window function ω(n) is applied to each segment to reduce spectral leakage, and the normalization factor of the window function is calculated: Calculate the periodicity graph based on windowed segment data: Where k = 0, 1, ..., N-1 represents the frequency index. The power spectral density estimate is obtained by averaging the periodograms of all segments. Four depression-related frequency bands—delta (1-4Hz), theta (4-8Hz), alpha (8-14Hz), beta (14-30Hz), and gamma (30-48Hz)—were extracted, and the extracted frequency band power features were normalized using the z-score method.
3. The dynamic graph convolutional EEG depression detection method based on spatiotemporal feature fusion according to claim 1, characterized in that, Step 2: A gated recurrent unit (GRU) model is used to capture the long-term temporal dependencies of EEG signals, and historical EEG information is learned through its gating mechanism; specifically: Input the current time feature x t Compared to the previous hidden state h t-1 Input GRU cells and compute update gate z using learnable parameters. t and reset door r t : z t =σ(W z x t +U z h t-1 ) r t =σ(W r x t +U r h t-1 ) Where σ(·) is the sigmoid activation function, W z W r ∈R c×d U is the input transformation weight matrix; z U r ∈R c×c The weight matrix is passed to the state; b z ,b r ∈R c For bias vectors, Generate candidate hidden states by combining the output of the reset gate: Where ⊙ represents element-wise multiplication, W t ∈R c×d ,U∈R c×c b∈R c Construct the candidate state generation parameter matrix. By updating the gate fusion historical state and candidate state: When z t →At time 0, all historical information is fully preserved (h) t ≈h t-1 When z t →The new state is fully adopted at 1 o'clock. To achieve adaptive information flow control, Map the hidden state to the output feature: f t GRU =ReLU(W o h t +b o ) Where W0∈R c×c b0∈R c For the output layer parameters, ReLU activation enhances the nonlinear characterization capability.
4. The dynamic graph convolutional EEG depression detection method based on spatiotemporal feature fusion according to claim 1, characterized in that, Step 3: Introducing a separable convolutional network (TSCN) into the residual block, combining dilated convolution and separable convolution to capture multi-scale depth information; specifically: The TSCN consists of multi-layered cascaded improved residual modules. Each residual module contains three levels of feature processing units. The first level uses a causal dilated convolution unit with 64 5×1 convolution kernels and a dilation coefficient of 1. The output is processed by a batch normalization layer and then activated by a ReLU function, with Dropout regularization performed at a preset probability. The second level is a separable dilated convolution unit, which includes two processing stages: depthwise convolution and pointwise convolution. The depthwise convolution uses 32 3×1 convolution kernels for independent channel processing with a dilation coefficient of 2, while the pointwise convolution uses a 1×1 convolution kernel to achieve cross-channel feature fusion. The third level is a residual connection unit, which performs dimensionality transformation using a 1×1 convolution kernel and then adds it to the main path features, finally outputting after activation by a ReLU function. In terms of parameter configuration, the network adopts a three-layer residual module cascade structure, with the expansion coefficient of each module increasing geometrically in the order of 1, 2, and 4. The number of network filters adopts a decreasing configuration strategy of 64-32-32, and the convolution kernel size decreases sequentially with the network depth to 5×1, 3×1, and 2×1.
5. The dynamic graph convolutional EEG depression detection method based on spatiotemporal feature fusion according to claim 1, characterized in that, Step 4: Introduce an attention mechanism to dynamically adjust the feature weights extracted by GRU and TSCN, focusing on the feature expressions that are most discriminative for depression classification; Specifically: The attention layer receives GRU features f t GRU and TSCN features f t TSCN splicing result Calculate attention weights: Among them, W a It is a learnable weight matrix, b a It is the bias vector, η T This is the parameter vector obtained during training, and tanh is the activation function. The calculated attention weights are used to dynamically weight and fuse the features of the two branches:
6. The dynamic graph convolutional EEG depression detection method based on spatiotemporal feature fusion according to claim 1, characterized in that, Step 5: A trainable adjacency matrix is used to parameterize brain region connectivity patterns, and the gradient descent algorithm is used to adaptively optimize connection weights to learn dynamic brain region interaction patterns related to depression; specifically: Define a loss function based on cross-entropy cost, and iteratively update the network parameters using the adjacency matrix through backpropagation (BP) until the optimal solution A is obtained: Loss=cross_entropy(l,l p )+a||Θ|| Where l represents the true label vector, l p This represents the label vector predicted by the model. An L2 regularization term α||Θ|| is added to prevent overfitting, where Θ includes the adjacency matrix parameters and other model parameters, and α is a weighting coefficient. Calculate the partial derivative of the loss function with respect to A. Update A using gradient descent: Where ρ is the learning rate, the update process continues until the loss converges or the preset number of iterations is reached, and the input to the GCN is the feature matrix f. A Given the adjacency matrix A, the interlayer propagation of GCN is as follows: Where σ is a nonlinear activation function, and W is the weight matrix to be trained. yes The degree matrix, I is the identity matrix.
7. The dynamic graph convolutional EEG depression detection method based on spatiotemporal feature fusion according to claim 1, characterized in that, Step 6: Perform global max pooling on the output of the graph convolutional layer, then perform feature transformation and dimensionality compression through a fully connected layer, and finally generate the final depression classification result through a softmax layer; specifically: For feature f A+1 Perform global max pooling to generate a vector of maximum values for node feature dimensions: f pool =maxpool(f A+1 ) The pooled feature vectors are passed through a two-layer fully connected network and a softmax activation function to output the depression classification probability: Where W1 and W2 are weight matrices, and b1 and b2 are bias vectors.
Citation Information
Cited By
Hyperspectral image building extraction method, system and program product
CN121053544A
Rice blast incubation period diagnosis and prediction method and model based on double-flow data fusion
CN121582748A
Rice blast latent period diagnosis and prediction method and model based on double-flow data fusion
CN121582748B
ADHD brain function connection dynamic representation system and method based on space-time diagram
CN121687549A
Electroencephalogram emotion recognition method and system based on frequency domain dynamic graph and domain generalization
CN122074990A