Depression eeg signal detection method based on multi-scale SELSTMNet model
By using the multi-scale SELSTMNet model, combined with multi-scale convolutional feature extraction and channel attention mechanism, the problems of insufficient feature expression and insufficient modeling of temporal dependencies in EEG signal detection in existing technologies are solved, and efficient and accurate detection of depression is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIAN UNIV OF POSTS & TELECOMM
- Filing Date
- 2026-03-30
- Publication Date
- 2026-06-26
AI Technical Summary
Existing methods for detecting depression based on electroencephalogram (EEG) signals suffer from limitations in feature scale, key feature expression capabilities, and time-series dependency modeling, making it difficult to achieve accurate and stable detection.
We employ a multi-scale SELSTMNet model, combining multi-scale convolutional feature extraction, channel attention mechanism, and bidirectional long short-term memory network to perform end-to-end EEG signal processing, thereby achieving multi-scale feature extraction and long-term temporal dependency modeling.
It improves the accuracy and stability of depression detection, enhances the robustness and generalization ability of the model, and reduces the reliance on manual feature design.
Smart Images

Figure CN122286253A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of biomedical signal processing and artificial intelligence technology, specifically relating to a method for detecting depression based on electroencephalogram (EEG) signals. More specifically, this invention relates to a method for detecting EEG signals in depression that integrates multi-scale feature extraction, attention mechanisms, and temporal modeling. Background Technology
[0002] Depression is a common mental disorder characterized by high incidence and recurrence rates, and has become a significant public health problem posing a serious threat to human physical and mental health. Patients with depression typically exhibit symptoms such as persistent low mood, decreased interest and pleasure, and cognitive impairment; in severe cases, it can even lead to extreme behaviors such as suicide. Therefore, early identification and accurate diagnosis of depression are of significant clinical and social value for timely intervention and effective treatment.
[0003] Currently, the clinical diagnosis of depression mainly relies on clinical interviews with psychiatrists and psychological assessment tools such as depression scales. However, these diagnostic methods are highly dependent on patients' subjective statements and doctors' clinical experience in practice, making them susceptible to subjective interference. They suffer from problems such as insufficient diagnostic consistency and difficulty in quantifying results, failing to meet the practical needs for objective assessment and accurate diagnosis of depression. Therefore, exploring more objective and reliable auxiliary diagnostic methods has become an urgent need in the field of depression research.
[0004] Electroencephalography (EEG) signals are physiological signals that reflect the electrical activity of neuronal populations in the brain, offering advantages such as non-invasive acquisition and high temporal resolution. Existing research has shown that the EEG signals of patients with depression differ from those of healthy individuals. Therefore, EEG-based auxiliary detection of depression is gradually becoming an important direction in the objective diagnosis of mental illnesses.
[0005] However, EEG signals are characterized by strong non-stationarity, low signal-to-noise ratio, and complex temporal structure. Existing EEG-based methods for detecting depression mostly focus on feature extraction at a single time scale, making it difficult to fully utilize both local and global features of EEG signals simultaneously. Furthermore, some methods fail to effectively model the short-term dynamic changes and long-term temporal dependencies of EEG signals.
[0006] In summary, there is an urgent need for a method to detect EEG signals of depression that can achieve multi-timescale feature extraction and fusion and effectively capture complex temporal dependencies, in order to overcome the limitations of existing technologies, improve the accuracy, stability and generalization ability of detection results, and thus provide more reliable technical support for the objective auxiliary diagnosis of depression. Summary of the Invention
[0007] This invention addresses the shortcomings of existing EEG-based methods for detecting depression, such as limited feature scale, insufficient ability to express key features, and limited capacity for modeling temporal dependencies. It provides a multi-scale SELSTMNet model-based EEG signal detection method for depression. This method constructs an end-to-end deep learning network that integrates multi-scale feature extraction, attention mechanisms, and sequence temporal modeling. This enables the collaborative modeling and efficient extraction of multi-scale features and their long-term temporal dependencies in EEG signals, thereby improving the accuracy, stability, and robustness of depression detection and providing a reliable technical means for the objective auxiliary diagnosis of depression.
[0008] To achieve the above-mentioned objectives, the present invention adopts the following technical solution.
[0009] A method for detecting EEG signals in depression based on a multi-scale SELSTMNet model includes the following steps:
[0010] The raw multichannel EEG signals were preprocessed, including bandpass filtering and notch filtering to remove noise, and independent component analysis (ICA) was used to remove artifacts. Subsequently, the processed signals were segmented according to fixed time windows, ultimately forming the input dataset for the neural network.
[0011] The EEG signal obtained in step 1 is input into the multi-scale convolutional feature extraction module. The module consists of multiple convolutional layers with different kernel lengths along the time dimension. Each convolutional layer performs convolution operations only in the time dimension to extract dynamic features of EEG at different time scales, thereby achieving joint modeling of short-term, medium-term, and long-term signal features.
[0012] After extracting convolutional features at various scales, a channel attention mechanism is introduced to weight the channel features of the convolutional output. By learning the importance weights of different EEG channels, the key channel features related to the depressive state are enhanced, while redundant and noisy information is suppressed.
[0013] Adaptive pooling is used to align the sizes of convolutional outputs at different scales to ensure consistent temporal and spatial resolution. The aligned features are then concatenated along the channel dimension to form a multi-scale fused feature representation. Subsequently, the fused features are further integrated and enhanced to improve the discriminative power of the feature representation.
[0014] Adaptive average pooling is applied to the fused features obtained in step 4 to map them into a fixed-size time window feature representation to eliminate the influence of input length differences; and the dimensions of the pooled features are rearranged to construct a feature sequence that meets the requirements of temporal modeling.
[0015] The feature sequences obtained in step 5 are input into a bidirectional long short-term memory network to model the long-term temporal dependencies of EEG signals and extract high-level feature representations containing forward and backward temporal information.
[0016] The temporal features output from step 6 are input into the fully connected classification module, which outputs the corresponding depression state classification results, thereby realizing the automatic identification and detection of the subject's depression state.
[0017] Compared with the prior art, the present invention has at least the following beneficial effects:
[0018] By employing a multi-scale convolutional structure, the dynamic features of EEG signals at different time scales can be fully explored, improving feature representation capabilities. The introduction of a channel attention mechanism effectively enhances channel information in key brain regions, increasing the model's sensitivity to depression-related spatial features. Combined with a bidirectional long short-term memory network, long-term temporal dependencies of EEG signals can be modeled, improving the stability and accuracy of classification results. The overall method is an end-to-end structure, reducing reliance on manual feature design and exhibiting good generalization ability and practical application value. Attached Figure Description
[0019] Figure 1 is a schematic diagram of the overall process of a method for detecting EEG signals of depression based on a multi-scale SELSTMNet model as described in this embodiment;
[0020] Figure 2 is a schematic diagram of the overall network structure of the multi-scale SELSTMNet model described in this embodiment;
[0021] Figure 3 This is a schematic diagram of the internal structure of the basic module in the multi-scale SELSTMNet model described in this embodiment. Detailed Implementation
[0022] like Figure 1 As shown in the figure, this embodiment provides a method for detecting EEG signals of depression based on a multi-scale SELSTMNet model. The method mainly includes steps such as EEG signal preprocessing, multi-scale feature extraction, channel attention weighting, multi-scale feature fusion, temporal modeling, and classification decision.
[0023] First, the acquired raw multi-channel EEG signals are preprocessed. This preprocessing includes bandpass filtering and notch filtering to remove irrelevant frequency noise and power frequency interference. Then, independent component analysis (ICA) is used to decompose the EEG signals, identifying and removing electrooculography (EOG) and electromyography (EMG) artifacts, thus obtaining purified EEG signals. The preprocessed EEG signals are then segmented according to a preset time window, with each time window corresponding to one EEG signal sample, to construct the input dataset for model training and testing.
[0024] Then, the above-mentioned EEG signal samples are input into the multi-scale convolutional feature extraction module. The multi-scale convolutional feature extraction module includes multiple cascaded two-dimensional convolutional layers. Each two-dimensional convolutional layer has a convolution kernel of different lengths along the time dimension. Convolution operations are performed on the EEG signals only in the time dimension to extract the dynamic features of EEG at different time scales, thereby realizing the joint modeling of the multi-scale temporal characteristics of EEG signals.
[0025] After completing multi-scale convolutional feature extraction, a channel attention mechanism is introduced to adaptively weight features corresponding to different EEG channels. This mechanism learns the importance weights of each EEG channel, enhancing channel features in key brain regions closely related to depressive states while suppressing channel features that contribute less to classification, thereby improving the discriminative power of feature representation.
[0026] Subsequently, the convolutional features obtained at different time scales are size-aligned using adaptive pooling to ensure consistent feature lengths across all scales in the time dimension. The aligned features are then concatenated along the channel dimension to form a multi-scale fused feature representation. Further, the multi-scale fused features are subjected to adaptive average pooling to map them to fixed-length feature representations, thus meeting the input requirements of subsequent temporal modeling networks.
[0027] Next, the pooled and rearranged feature sequences are input into a bidirectional long short-term memory network. The bidirectional long short-term memory network performs temporal modeling on the feature sequences in both forward and backward directions, thereby effectively capturing the long-term temporal dependencies of EEG signals and obtaining a high-level temporal feature representation containing complete temporal context information.
[0028] Finally, the high-level temporal features output by the bidirectional long short-term memory network are input into the fully connected classification module to output the corresponding depression state classification results, thereby realizing the automatic identification and detection of the subject's depression state.
[0029] In this embodiment, a systematic performance verification experiment was conducted on the proposed detection method. The results show that the method has good detection performance in tasks such as depression state classification. Table 1 shows the comparison results of the proposed method with other methods in detecting EEG signals of depression.
[0030] Table 1. Performance Comparison of EEG Detection Methods for Depression
[0031] .
Claims
1. A method for detecting EEG signals in depression based on a multi-scale SELSTMNet model, characterized in that, The method extracts and classifies features from EEG signals based on a deep neural network model, which includes at least a multi-scale SE convolution module and a BiLSTM-based temporal modeling module.
2. The method for detecting EEG signals of depression based on a multi-scale SELSTMNet model according to claim 1, characterized in that, The multi-scale SE convolutional module is composed of multiple convolutional sub-modules connected in sequence, with each sub-module using a convolutional kernel of a different size. The input EEG signal is first fed into the first convolutional sub-module for feature extraction, and the feature map output by that sub-module is retained. Subsequently, the output of the previous convolutional sub-module is used as the input of the next convolutional sub-module, and so on, to each convolutional sub-module. The output of each convolutional sub-module passes through the Squeeze-and-Excitation (SE) attention module, which compresses the feature map in the channel dimension to obtain global statistical information, and generates channel weight coefficients accordingly. The feature map is then recalibrated to obtain a weighted feature representation. As EEG signals are transmitted layer by layer between sub-modules of different convolutional kernel sizes, the network's receptive field gradually expands. The feature maps of each convolutional sub-module, compressed by the SE attention module, are spliced together at the output, thus forming a multi-scale weighted feature representation that integrates information from multiple time scales.
3. The method for detecting EEG signals of depression based on a multi-scale SELSTMNet model according to claim 1, characterized in that, The BiLSTM-based temporal modeling module receives the multi-scale weighted feature representation output by the multi-scale SE convolution module and performs temporal modeling of EEG signal features through a bidirectional long short-term memory network (Bi-LSTM). The Bi-LSTM network includes an LSTM sub-module running forward in time and an LSTM sub-module running backward in time, used to simultaneously capture the forward and backward dependencies of EEG signals in the temporal dimension. Finally, the features output by the Bi-LSTM network are further mapped through a fully connected layer to complete the classification and discrimination of EEG signals related to depression.
4. The method for detecting EEG signals of depression based on a multi-scale SELSTMNet model according to any one of claims 1 to 3, characterized in that, The multi-scale SE convolution module generates multi-scale weighted feature representations through the synergistic effect of features extracted by convolution kernels of different scales and channel attention weighting mechanism. The multi-scale weighted feature representations are then input into the BiLSTM-based temporal modeling module, enabling BiLSTM to utilize contextual information from multiple time scales simultaneously during temporal modeling, thereby improving the ability to express EEG signal features and the accuracy of depression detection.