Depression detection method based on emotion electroencephalogram signal joint characteristics and related equipment
By extracting the differential entropy features of emotional EEG signals and using a two-dimensional convolutional neural network and a Transformer encoder for joint modeling, the low accuracy problem of existing depression diagnosis methods is solved, and efficient and accurate automatic identification of depression is achieved.
Patent Information
- Application Number
- CN202510940207.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-09-26
AI Technical Summary
Existing clinical diagnostic methods for depression rely on subjective descriptions and self-rating scales, which have problems of low accuracy and poor consistency.
By acquiring emotional EEG signals, extracting differential entropy features, and using a two-dimensional convolutional neural network and Transformer encoder to learn the joint deep features of emotional EEG signals in the spatial, frequency, and temporal dimensions, combined with a multi-layer perceptron classifier, the healthy and depressive states are classified.
It improves the diagnostic accuracy and generalization ability of depression, and achieves efficient, accurate and automated identification of depression.
Smart Images

Figure CN120694646A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of auxiliary diagnosis of mental illness, and more specifically, to a depression detection method based on the joint features of emotional EEG signals and related equipment. Background Art
[0002] Major depressive disorder (MDD) is a common mental disorder characterized by persistent low mood, anhedonia, and decreased cognitive function. MDD severely impacts patients' daily lives and social functioning, and has become a leading cause of disability worldwide.
[0003] Current clinical diagnosis of MDD primarily relies on subjective patient descriptions, self-assessment scales (such as the PHQ-9), and structured interviews conducted by psychiatrists. These methods are subject to high subjectivity, poor consistency, and diagnostic delays, resulting in low clinical diagnostic accuracy for MDD.
[0004] Therefore, how to improve the accuracy of clinical diagnosis of MDD is an urgent problem to be solved in this application. Summary of the Invention
[0005] In view of this, the present application discloses a depression detection method and related equipment based on the combined features of emotional EEG signals, aiming to improve the accuracy of clinical diagnosis of depression.
[0006] In order to achieve the above purpose, the disclosed technical solutions are as follows:
[0007] The first aspect of the present application discloses a method for detecting depression based on joint features of emotional EEG signals, the method comprising:
[0008] Acquire an original emotional EEG signal; wherein the original emotional EEG signal is an emotional EEG signal that has not been subjected to feature extraction;
[0009] Extracting differential entropy features from the original emotional EEG signal;
[0010] By using the differential entropy features, pre-trained two-dimensional convolutional neural network and Transformer encoder, the local and global dependencies of emotional EEG signals in the spatial and frequency dimensions, as well as the dynamic evolution of emotional EEG signals in the time dimension, are learned to construct joint deep features of emotional EEG signals; wherein, the time dimension is the dimension of dynamic changes in the time series of emotional EEG signals; the spatial dimension is the dimension of dynamic correlation of the spatial distribution of emotional EEG signals; and the frequency dimension is the dimension of dynamic interaction of frequency components of emotional EEG signals;
[0011] The joint deep features are classified into healthy and depressed states through a multi-layer perceptron classifier to obtain and output the classification results.
[0012] Preferably, the obtaining of the original emotional EEG signal includes:
[0013] By using a preset dry electrode EEG cap and a preset acquisition frequency, the original emotional EEG signals generated by emotional stimulation of the target under the preset scenario are collected;
[0014] The emotional stimulation includes at least happy emotional stimulation, sad emotional stimulation and neutral emotional stimulation.
[0015] Preferably, the step of extracting differential entropy features from the original emotional EEG signal includes:
[0016] performing bandpass filtering on the original emotional EEG signal;
[0017] Divide non-overlapping time windows into preset time units;
[0018] Applying a Hanning window to the emotional EEG signal after bandpass filtering through the non-overlapping time window and performing a fast Fourier transform to obtain a frequency domain energy distribution;
[0019] Calculating the total energy within each frequency band according to the frequency energy distribution and a preset number of frequency bands;
[0020] Based on the assumption that the short-time stationary signal obeys the Gaussian distribution and the total energy in each frequency band, the differential entropy feature is calculated.
[0021] Preferably, the method of learning the local and global dependencies of emotional EEG signals in the spatial and frequency dimensions, as well as the dynamic evolution of emotional EEG signals in the time dimension, through the differential entropy features, the pre-trained two-dimensional convolutional neural network and the Transformer encoder, and constructing the joint deep features of emotional EEG signals includes:
[0022] Inputting the differential entropy features into a two-dimensional convolutional neural network trained with a Focal Loss function as a loss function and a Transformer encoder with a multi-head self-attention mechanism;
[0023] Through the two-dimensional convolutional neural network and the Transformer encoder with a multi-head self-attention mechanism, emotional EEG signals are jointly modeled in the time dimension, spatial dimension and frequency dimension to learn the local and global dependencies of emotional EEG signals in the spatial dimension and frequency dimension, as well as the dynamic evolution process of emotional EEG signals in the time dimension, thereby completing the process of constructing the joint deep features of emotional EEG signals.
[0024] Preferably, the classifying the joint deep features into healthy and depressed states by a multi-layer perceptron classifier, obtaining and outputting the classification results, includes:
[0025] Inputting the joint deep features into a multi-layer perceptron classifier to output a predicted probability distribution of a healthy category and a depressed category through the multi-layer perceptron classifier;
[0026] The classification result is determined from the predicted probability distribution of the healthy category and the depressed category by using the maximum probability criterion and is output.
[0027] The second aspect of the present application discloses a depression detection system based on joint features of emotional EEG signals, the system comprising:
[0028] An acquisition unit, configured to acquire an original emotional EEG signal; wherein the original emotional EEG signal is an emotional EEG signal that has not been subjected to feature extraction;
[0029] An extraction unit, configured to extract differential entropy features from the original emotional EEG signal;
[0030] A construction unit is used to learn the local and global dependencies of emotional EEG signals in the spatial and frequency dimensions, as well as the dynamic evolution process of emotional EEG signals in the time dimension, through the differential entropy features, the pre-trained two-dimensional convolutional neural network and the Transformer encoder, to construct a joint deep feature of the emotional EEG signals; wherein the time dimension is the dimension of the dynamic change of the emotional EEG signals in the time series; the spatial dimension is the dimension of the dynamic correlation of the spatial distribution of the emotional EEG signals; and the frequency dimension is the dimension of the dynamic interaction of the frequency components of the emotional EEG signals;
[0031] The classification unit is used to classify the joint deep features into healthy and depressed states through a multi-layer perceptron classifier, obtain a classification result and output it.
[0032] Preferably, the acquisition unit is specifically used to collect the original emotional EEG signals generated by emotional stimulation of the target under test in a preset scenario through a preset dry electrode EEG cap and a preset acquisition frequency; wherein the emotional stimulation includes at least happy emotional stimulation, sad emotional stimulation and neutral emotional stimulation.
[0033] Preferably, the extraction unit comprises:
[0034] A filtering module, configured to perform bandpass filtering on the original emotional EEG signal;
[0035] A division module is used to divide non-overlapping time windows into preset time units;
[0036] an execution module, configured to apply a Hanning window to the bandpass filtered emotional EEG signal through the non-overlapping time window, and perform a fast Fourier transform to obtain a frequency domain energy distribution;
[0037] A first calculation module, configured to calculate the total energy within each frequency band according to the frequency energy distribution and a preset number of frequency bands;
[0038] The second calculation module is used to calculate the differential entropy feature based on the assumption that the short-time stationary signal obeys the Gaussian distribution and the total energy within each frequency band.
[0039] A third aspect of the present application discloses a storage medium, which includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the depression detection method based on the joint features of emotional EEG signals as described in any one of the first aspects.
[0040] The fourth aspect of the present application discloses an electronic device comprising a memory and one or more instructions, wherein the one or more instructions are stored in the memory and configured to be executed by one or more processors to perform the depression detection method based on the combined features of emotional EEG signals as described in any one of the first aspects.
[0041] It can be seen from the above technical solution that the present application discloses a depression detection method and related equipment based on the joint features of emotional EEG signals to obtain original emotional EEG signals; wherein, the original emotional EEG signals are emotional EEG signals without feature extraction, and differential entropy features are extracted from the original emotional EEG signals. Through differential entropy features, pre-trained two-dimensional convolutional neural networks and Transformer encoders, the local and global dependencies of emotional EEG signals in spatial and frequency dimensions, as well as the dynamic evolution process of emotional EEG signals in time dimension are learned to construct joint deep features of emotional EEG signals, wherein the time dimension is the dimension of dynamic changes of emotional EEG signals in time series; the spatial dimension is the dimension of dynamic correlation of emotional EEG signals in spatial distribution; the frequency dimension is the dimension of dynamic interaction of frequency components of emotional EEG signals, and the joint deep features are classified into healthy and depressive states through a multi-layer perceptron classifier to obtain and output the classification results.
[0042] The beneficial effects of this application are: this scheme uses a joint high-order feature extraction mechanism in the time dimension, space dimension and frequency dimension, combined with a two-dimensional convolutional neural network trained with Focal Loss as the loss function and an encoder of the Transformer structure with a multi-head self-attention mechanism, to perform high-order multi-dimensional joint modeling of emotional EEG signals, and to mine the joint deep features of emotional EEG signals in the three dimensions of time series (time domain), brain area distribution (spatial domain) and frequency component (frequency domain), thereby improving the accuracy and generalization ability of depression detection, realizing efficient, accurate and automated identification of depression, and thus improving the accuracy of clinical diagnosis of depression. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0044] Figure 1 A flowchart of a method for detecting depression based on combined features of emotional EEG signals disclosed in an embodiment of the present application is shown;
[0045] Figure 2 A schematic diagram of the overall model structure disclosed in the embodiments of this application;
[0046] Figure 3 This is a schematic diagram of the structure of a depression detection system based on joint features of emotional EEG signals disclosed in an embodiment of the present application;
[0047] Figure 4 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present application. DETAILED DESCRIPTION
[0048] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0049] In this application, the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0050] As can be seen from the background, current clinical diagnosis of MDD mainly relies on subjective descriptions by patients, self-rating scales, and structured interviews conducted by psychiatrists. These methods suffer from strong subjectivity, poor consistency, and diagnostic delays, resulting in low accuracy in clinical diagnosis of MDD.
[0051] In order to solve the above problems, the present application discloses a depression detection method and related equipment based on the joint features of emotional EEG signals. This solution uses a joint high-order feature extraction mechanism of time dimension, space dimension and frequency dimension, combined with a two-dimensional convolutional neural network trained with Focal Loss as the loss function and an encoder of a Transformer structure with a multi-head self-attention mechanism, to perform high-order multi-dimensional joint modeling of emotional EEG signals, and mine the joint deep features of emotional EEG signals in the three dimensions of time series, brain area distribution and frequency component, thereby improving the accuracy and generalization ability of depression detection, achieving efficient, accurate and automatic recognition of depression, and thus improving the accuracy of clinical diagnosis of depression. The specific implementation method is specifically described through the following examples.
[0052] It should be noted that the depression detection method and related equipment based on the combined features of emotional EEG signals provided in this application can be used in technical fields such as brain-computer interfaces, artificial intelligence, and auxiliary diagnosis of mental illness. These technologies fall within the cutting-edge intersection of affective computing and auxiliary diagnosis of mental illness. The above is merely illustrative and does not limit the application areas of the depression detection method and related equipment based on the combined features of emotional EEG signals provided in this application.
[0053] refer to Figure 1 FIG. 1 is a flow chart of a method for detecting depression based on the combined features of emotional EEG signals disclosed in an embodiment of the present application. The method for detecting depression based on the combined features of emotional EEG signals mainly includes the following steps:
[0054] S101: Acquire original emotional EEG signals; wherein the original emotional EEG signals are emotional EEG signals that have not been subjected to feature extraction.
[0055] In S101, a preset dry electrode EEG cap and a preset acquisition frequency are used to acquire original emotional EEG signals generated by emotional stimulation of a target under a preset scenario, wherein the emotional stimulation includes at least happy emotional stimulation, sad emotional stimulation, and neutral emotional stimulation.
[0056] The preset dry electrode EEG cap includes but is not limited to the DSI-24 dry electrode EEG cap. The preset dry electrode EEG cap of the present application is preferably the DSI-24 dry electrode EEG cap.
[0057] The preset acquisition frequency includes but is not limited to 300 Hz. The preset acquisition frequency of the present application is preferably 300 Hz.
[0058] Preset scenarios include scenarios involving different video clips of emotional experience material. For example, 73 subjects (33 with depression and 40 healthy subjects) were recruited and asked to watch nine video clips of emotional experience material, including three each of happy (HAP), sad (SAD), and neutral (NEU) emotional stimuli. The videos were presented to the subjects in the order of HAP, NEU, SAD, SAD, NEU, HAP, HAP, NEU, and SAD.
[0059] In practical applications, a DSI-24 dry electrode EEG cap can be used to collect raw emotional EEG signals from subjects while they watch videos at a frequency of 300Hz. Using the international 10-20 system to position the sensor, with the Pz electrode as a reference, data analysis is performed on the other 18 channels to collect raw emotional EEG signals generated by emotional stimulation of the subject in pre-set scenarios.
[0060] S102: Extract differential entropy features from the original emotional EEG signal.
[0061] In S102 , the original emotional EEG signal is preprocessed, and the preprocessed original emotional EEG signal is subjected to window function processing using a Hanning window, a fast Fourier transform is performed, and a differential entropy feature is calculated.
[0062] Preprocessing includes but is not limited to bandpass filtering, time window segmentation, etc.
[0063] The differential entropy features are extracted from the original emotional EEG signals through the EEG signal preprocessing and feature extraction module, and a feature tensor containing a three-dimensional structure of time, space and frequency is constructed through the differential entropy features.
[0064] The specific process of extracting differential entropy features from the original emotional EEG signal is shown in A1-A5.
[0065] A1: Bandpass filter the original emotional EEG signal.
[0066] In A1, the raw EEG signals of each channel were band-pass filtered (1–45 Hz).
[0067] A2: Divide the time windows into non-overlapping windows using the preset time units.
[0068] The preset time may be 1 second, 3 seconds, etc., and the preset time in this application is not specifically limited. For example, 1 second is used as a unit to divide the time windows into non-overlapping time windows.
[0069] A3: Apply a Hanning window to the bandpass filtered emotional EEG signal using non-overlapping time windows and perform a fast Fourier transform to obtain the frequency domain energy distribution.
[0070] A4: Calculates the total energy within each frequency band based on the frequency energy distribution and a preset number of frequency bands.
[0071] Among them, the preset number is generally five.
[0072] According to the five commonly used EEG frequency bands, namely delta (1-4Hz), theta (4-8Hz), alpha (8-14Hz), beta (14-31Hz) and gamma (31-45Hz), the total energy E within each frequency band is calculated.
[0073] A5: Based on the assumption that short-term stationary signals follow a Gaussian distribution and the total energy within each frequency band, the differential entropy feature is calculated.
[0074] In A5, based on the assumption that short-term stationary signals obey Gaussian distribution, the total energy of the current time window, frequency band, and channel is approximately the signal variance, that is, , calculate the differential entropy (DE) feature of the EEG signal x. The calculation formula of the differential entropy feature is shown in formula (1).
[0075] (1)
[0076] in, is the differential entropy feature of the EEG signal x; is the circumference of a circle, which is 3.14; E is the total energy within each frequency band; is a natural constant.
[0077] S103: Through differential entropy features, a pre-trained two-dimensional convolutional neural network (CNN) and a Transformer encoder, we learn the local and global dependencies of emotional EEG signals in the spatial and frequency dimensions, as well as the dynamic evolution of emotional EEG signals in the temporal dimension, to construct joint deep features of emotional EEG signals. Among them, the temporal dimension is the dimension of dynamic changes in the time series of emotional EEG signals; the spatial dimension is the dimension of dynamic correlation in the spatial distribution of emotional EEG signals; and the frequency dimension is the dimension of dynamic interaction of frequency components of emotional EEG signals.
[0078] The two-dimensional convolutional neural network and the Transformer encoder jointly consider the three dimensions of time, spatial channels, and frequency of the EEG signal when extracting features to obtain higher-order EEG feature representations with stronger discriminability (i.e., joint deep features).
[0079] The two-dimensional convolutional neural network and Transformer encoder were trained using the Focal Loss function. During training, three-fold cross-validation was employed, and the EEG data was divided into training, validation, and test sets according to a preset ratio, such as 8:2:5. By maintaining the isolation of the data of the tested subjects, the risk of data leakage was minimized, ensuring objective evaluation of the model's performance on unseen subjects.
[0080] During the training phase of the two-dimensional convolutional neural network and the Transformer encoder, the parameters of each module are jointly optimized by minimizing the loss function until the model converges.
[0081] The preset ratio may be 8:2:5, 8:3:4, etc. The specific preset ratio is not specifically limited in this application.
[0082] The specific process of constructing the joint deep features of emotional EEG signals is shown in B1-B2.
[0083] B1: Input the differential entropy features into a two-dimensional convolutional neural network trained with Focal Loss as the loss function and a Transformer encoder with a multi-head self-attention mechanism.
[0084] B2: Through a two-dimensional convolutional neural network and a Transformer encoder with a multi-head self-attention mechanism, emotional EEG signals are jointly modeled in the time, space, and frequency dimensions to learn the local and global dependencies of emotional EEG signals in the spatial and frequency dimensions, as well as the dynamic evolution of emotional EEG signals in the time dimension, thereby completing the process of constructing joint deep features of emotional EEG signals.
[0085] In B2, the DE features of all time windows, all frequency bands, and all channels are aggregated to jointly construct a three-dimensional feature tensor of emotional EEG signals.
[0086] Two-dimensional convolutional neural networks are used to model the local and global dependencies of EEG signals in spatial channel and frequency dimensions.
[0087] The Transformer encoder with multi-head self-attention is used to capture the dynamic evolution characteristics of EEG signals in time series.
[0088] The expression of the joint deep feature is shown in formula (2).
[0089] (2)
[0090] in, is a three-dimensional feature tensor; is the real number domain; N is the number of time sampling points; Fr is the number of frequency bands after frequency domain decomposition; Ch is the number of spatial electrode channels.
[0091] This three-dimensional feature tensor fully retains the time-space-frequency dynamic evolution information of the EEG signal and is the basis for subsequent joint deep feature learning.
[0092] In view of the coupling characteristics of EEG signals in frequency dynamics, spatial topology and temporal evolution, the multi-dimensional deep features of EEG signals are jointly extracted through the fusion structure of two-dimensional convolutional neural network and Transformer encoder.
[0093] The 2D convolutional neural network module aims to jointly model the local and global features of EEG signals in both frequency and spatial dimensions. The specific steps are as follows:
[0094] First, slide sampling along the time axis to generate batch subsample tensors of DE features , where T is the adjustable time window length, T≤N; to ensure the smooth transition and feature integrity of time series information, overlapping sampling is adopted between adjacent subsamples, the overlap rate, the sliding step size, and the subsample can be expressed as shown in formula (3).
[0095] (3)
[0096] in, is the three-dimensional feature tensor of the t-th subsample; is a three-dimensional feature tensor; t is the subsample number; s is the sliding step size; T is the adjustable time window length; N is the number of time sampling points.
[0097] Then, the subsample tensor is input into the two-dimensional CNN, and the convolution kernel slides jointly in the frequency domain and the spatial domain to capture the local and global dependencies between different brain regions in different frequency bands, and enhance the expression modeling ability of brain region co-activation and frequency resonance in the depressive state; specifically, the two-dimensional CNN module contains two layers of convolution operations, the convolution kernel size is (3, 7), the step size is (1, 1), and the same padding is used to keep the size of the output feature map in the frequency and spatial dimensions consistent with the input. The number of input and output feature maps is consistent with the time series length T to maintain the temporal structure of the data throughout the convolution process.
[0098] make Represents the input features at position (i, j) and feature map t, after layer l ( ) After the convolution operation, the hidden representation in the output feature map k can be expressed as formula (4):
[0099] (4)
[0100] in, is the hidden feature of the kth feature map, the fth frequency band, and the cth electrode channel in the lth layer of the two-dimensional convolutional network; is the corresponding network weight; is the tth feature map in the l-1th layer of the two-dimensional convolutional network, the f+ i The hidden features of the frequency band and the c+jth electrode channel; is the bias constant.
[0101] In order to enhance the model’s ability to capture complex patterns and prevent overfitting, an average pooling layer is added after each convolutional layer (the average pooling layer is added to reduce the spatial dimension of the feature map, highlight the main features and reduce noise interference), a ReLU nonlinear activation layer (the ReLU nonlinear activation layer is used to introduce nonlinear transformations into the neural network to improve the model’s fitting ability) and a Dropout regularization layer (randomly discarding some neuron outputs with a certain probability to reduce the risk of overfitting during the training process). The corresponding calculation formula is shown in Formula (5).
[0102] (5)
[0103] in, is the pooling range; is the output feature of the kth feature map, fth frequency band, and cth electrode channel after processing by the two-dimensional convolutional network; For the mask; is the frequency band; is the electrode channel; For hidden features.
[0104] After stacking two layers of convolution-pooling-ReLU-Dropout modules, the module finally outputs the feature tensor ,This feature strengthens the space-frequency coupling expression and preserves the T-dimensional time series structure, laying the foundation for subsequent time domain feature modeling.
[0105] in, is the three-dimensional feature tensor after processing by the two-dimensional convolutional network; is the real number domain; T is the adjustable time window length; Fr is the number of frequency bands after frequency domain decomposition; Ch is the number of spatial electrode channels;
[0106] The Transformer encoder with multi-head self-attention is designed to extract deep features of EEG signals in the temporal dimension. The specific steps are as follows:
[0107] First, the three-dimensional feature tensor output by the two-dimensional CNN, namely the differential entropy feature, is flattened and reconstructed into a time series matrix , so that the space-frequency features of each time point are spliced into a vector and injected into the position encoding based on sine-cosine recursion , get the input features of this stage ;in, is a time series matrix; PE is the position code of sine-cosine recursion; is the input feature;
[0108] Then, the input features of this stage The input is fed into the Transformer encoder to model the dynamic changes and long-range dependencies of neural activity in the temporal dimension. The Transformer encoder consists of L identical blocks, each of which contains a multi-head self-attention layer and a fully connected feedforward layer. Each layer combines residual connections and layer normalization to stabilize the training process.
[0109] The attention layer captures the dependencies between feature vectors at different time points in the sequence, which is calculated using formula (6).
[0110] (6)
[0111] Where Q is the query vector; K is the key vector; V is the numerical vector; T is the transpose operation; is the scaling factor.
[0112] Query ,key Sum , by combining the input feature vector with three learnable weight matrices , and Multiplying them together, These matrices represent different projections of the same input sequence, allowing the model to capture various aspects of temporal dynamics. Each query is dot-producted with all keys to obtain attention weights, scaling factors This ensures stable gradient computation. The attention weights are used to compute the weighted sum of the value vectors, resulting in a context-enhanced feature representation.
[0113] in, 、 、 is the dimension of the corresponding vector.
[0114] In order to enhance the representation ability of the model, a multi-head attention mechanism is introduced into the Transformer encoder, which enables the Transformer encoder to learn from multiple subspaces at the same time. Using different learnable projections, queries, keys, and values are projected into h smaller dimensions. These projections are performed in parallel and the resulting representation feature vectors are concatenated to produce a shape of The global timing characteristics of are shown in formulas (7) and (8).
[0115] (7)
[0116] (8)
[0117] in, Feature tensor output by multi-head attention; is the weight; is the i-th attention head; is the i-th projection of the query vector; is the i-th projection of the key vector; is the i-th projection of the numeric vector.
[0118] The feedforward network of the Transformer encoder uses two layers of linear transformation and ReLU activation to expand the feature transformation space.
[0119] The overall Transformer encoder module outputs a high-order temporal feature tensor with rich contextual associations, improving the model's ability to discriminate depressive states.
[0120] S104: Classify the joint deep features into healthy and depressed states through a multi-layer perceptron classifier, obtain the classification results and output them.
[0121] In S104, the extracted high-order features, i.e., the joint deep features, are input into a multilayer perceptron classifier to output the subject's health or depression status. Specifically, the joint deep features are input into the multilayer perceptron classifier to output a predicted probability distribution of the health category versus the depression category. A classification result is determined from the predicted probability distribution of the health category versus the depression category using a maximum probability criterion and output.
[0122] For each emotional experience, the EEG features of the training set are input into the deep neural network module to learn the joint deep features, which are then input into the multi-layer perceptron classifier to output the predicted probability distribution of the healthy category and the depressed category. The classification result is determined and output from the predicted probability distribution of the healthy category and the depressed category using the maximum probability (argmax) criterion.
[0123] Iterative training yields the final model. The EEG features of the test set are input into the model, which outputs the predicted probability distribution for the healthy-depressed category, resulting in the predicted category label. This predicted category label is then compared with the true label for performance evaluation, which yields performance results. These results include at least accuracy, F1 score, AUC, sensitivity, and specificity. The F1 score represents a composite indicator of precision and recall, and the AUC (Area Under Curve) represents the area under the receiver operating characteristic (ROC) curve.
[0124] The model EMOCT in this application presents the three-fold cross-average prediction results for health-depression classification on three emotional EEG datasets, and compares its performance with the following five baseline models:
[0125] The SVM baseline model is used to operate directly on the flattened frequency bands and channels, i.e., d model =5×18=90, using the "Linear" or "Polygonal" kernel;
[0126] The MLP baseline model operates directly on the flattened frequency bands and channels, i.e., d model =5×18=90;
[0127] CNN1D-T benchmark model, used for one-dimensional convolution and multi-layer perceptron operations on time series, i.e. d model =T(adjustable);
[0128] CNN1D-S benchmark model for one-dimensional convolution and multi-layer perceptron operations on flattened frequency bands and channels, i.e. d model =5×18=90;
[0129] CNN2D-S benchmark model, used to perform two-dimensional convolution and multi-layer perceptron operations on the reconstructed features in frequency bands and channels, i.e. d model=(5, 18).
[0130] The results show that the EMOCT model, based on EEG data for happy, neutral, and sad emotional experiences, achieved AUCs of 87.18%, 84.91%, and 89.38% for health-depression detection, respectively, outperforming the baseline method. The feature extraction techniques described in this application are effective for objectively detecting depression.
[0131] Table 1. Accuracy, F1 score, AUC, sensitivity, and specificity (mean, %) of healthy-depressed classification based on three types of emotional EEG signals.
[0132] Table 1
[0133]
[0134] The specific steps of classification by the multi-layer perceptron classifier are as follows:
[0135] Expand and concatenate the temporal features output by the Transformer encoder into a vector , the vector Input into a multi-layer perceptron classifier composed of a fully connected layer, a LeakyReLU nonlinear activation function, a Dropout layer and a softmax logarithmic normalization function to achieve accurate automatic judgment of health and depression states. Specifically, the vector The calculation formula input to the multi-layer perceptron classifier is shown in formula (9).
[0136] (9)
[0137] in, 、 are all weight matrices of the fully connected layers; 、 are bias vectors; is the final healthy-depressed probability distribution.
[0138] In order to further improve the classification performance of the depression detection model in this application, especially to enhance the model's ability to recognize difficult-to-classify samples when the category distribution is unbalanced or the differences between samples are not significant, the Focal Loss function is introduced as an optimization objective in the model training stage.
[0139] The Focal Loss function in this application replaces the traditional cross-entropy loss function (Cross-Entropy Loss), and its definition is shown in formula (10).
[0140] (10)
[0141] in, is the Focal Loss function; is the predicted probability of the sample being correctly classified. The larger the probability, the more accurate the classification and the higher the confidence level. α is the importance weight factor for balancing positive and negative samples. , The model’s learning of the depression category can be emphasized; is the loss weight adjustment factor, set This allows the model to focus on incorrectly predicted samples.
[0142] Compared to the cross-entropy loss function, the Focal Loss function maintains the same loss for inaccurately classified samples, but reduces the loss for accurately classified samples. Therefore, it effectively increases the weight of inaccurately or difficult-to-classify samples in the loss function, making the model focus more on "marginal samples" that are difficult to correctly classify, thereby improving the model's discriminative ability. At the same time, by reducing the focus on easily classified samples, the risk of overfitting is reduced, and the model's generalization ability is enhanced.
[0143] In order to facilitate the understanding of the process of depression detection based on emotional EEG signals, combined with Figure 2 Provide explanation. Figure 2 A schematic diagram of the overall model structure is shown.
[0144] Figure 2 In the process, the original emotional EEG signal is band-pass filtered;
[0145] Divide non-overlapping time windows into preset time units;
[0146] Applying Hanning window to the bandpass filtered emotional EEG signal through non-overlapping time windows and performing fast Fourier transform to obtain the frequency domain energy distribution;
[0147] Calculate the total energy within each frequency band based on the frequency energy distribution and a preset number of frequency bands;
[0148] Based on the assumption that the short-term stationary signal obeys the Gaussian distribution and the total energy within each frequency band, the differential entropy feature is calculated;
[0149] The differential entropy features are input into a two-dimensional convolutional neural network trained with Focal Loss as the loss function and a Transformer encoder with a multi-head self-attention mechanism;
[0150] To enhance the ability of the two-dimensional convolutional neural network to capture complex patterns and prevent overfitting, each convolutional layer is followed by an average pooling layer (the average pooling layer is used to reduce the spatial dimension of the feature map, highlight the main features, and reduce noise interference), a ReLU nonlinear activation layer (the ReLU nonlinear activation layer is used to introduce nonlinear transformations to the neural network to improve the model's fitting ability), and a Dropout regularization layer (randomly discarding some neuron outputs with a certain probability to reduce the risk of overfitting during training).
[0151] The three-dimensional feature tensor output by the two-dimensional CNN, namely the differential entropy feature, is flattened and reconstructed into a time series matrix , so that the space-frequency features of each time point are spliced into a vector and injected into the position encoding based on sine-cosine recursion , get the input features of this stage ;in, is a time series matrix; PE is the position code of sine-cosine recursion; is the input feature;
[0152] The input features of this stage The input is fed into a Transformer encoder to model the dynamic changes and long-range dependencies of neural activity in the temporal dimension. The Transformer encoder consists of L identical blocks, each of which contains a multi-head self-attention layer and a feed-forward fully connected layer. Each layer combines residual connections and layer normalization to stabilize the training process.
[0153] Expand and concatenate the temporal features output by the Transformer encoder into a vector , the vector The data is input into a multi-layer perceptron classifier consisting of a fully connected layer, a LeakyReLU nonlinear activation function, a Dropout layer, and a softmax logarithmic normalization function to obtain the classification results and output them to achieve accurate and automatic judgment of healthy and depressive states.
[0154] This approach achieves efficient, accurate, and automated identification of depression by constructing a deep model framework that combines EEG features in the time, space, and frequency dimensions, thus overcoming the shortcomings of existing methods in terms of feature expression and detection accuracy. It also enables automatic classification and identification of individuals' healthy and depressive states. Unlike existing methods that only utilize a single dimension for analysis, this approach proposes a joint high-order feature extraction mechanism based on time, space, and frequency. Combining a two-dimensional convolutional neural network with a Transformer structure with a multi-head self-attention mechanism, this approach exploits the joint deep features of emotional EEG signals across three dimensions: time series (time domain), brain region distribution (spatial domain), and frequency component (frequency domain), significantly improving the accuracy and generalization of depression detection.
[0155] This solution comprises an EEG signal preprocessing and feature extraction module, a two-dimensional convolutional neural network module, a Transformer encoder module with a multi-head self-attention mechanism, and a multi-layer perceptron classifier module. This solution first collects EEG signals while the subject is viewing emotion-inducing material. The total frequency energy is calculated using a fast Fourier transform (FFT) and differential entropy features are extracted. This is then fed into a neural network model. The EEG signals are then jointly modeled across three dimensions: time series, spatial channels, and frequency components, using the convolutional neural network and the Transformer module. High-order joint spatiotemporal and frequency features are extracted, and a classifier is used to discriminate between healthy and depressed states. This approach achieves higher classification accuracy than traditional classification methods. This method employs Focal Loss to train a deep neural network. This solution effectively improves the accuracy and robustness of depression detection, providing a novel, non-invasive solution for the objective identification of mood disorders.
[0156] The beneficial effects of the embodiments of the present application: This solution uses a joint high-order feature extraction mechanism in the time dimension, space dimension and frequency dimension, combined with a two-dimensional convolutional neural network trained with Focal Loss as the loss function and an encoder of a Transformer structure with a multi-head self-attention mechanism, to perform high-order multi-dimensional joint modeling of emotional EEG signals, and to explore the joint deep features of emotional EEG signals in the three dimensions of time series, brain area distribution and frequency component, thereby improving the accuracy and generalization ability of depression detection, realizing efficient, accurate and automated identification of depression, and thus improving the accuracy of clinical diagnosis of depression.
[0157] Based on the above embodiment Figure 1 A depression detection method based on the joint features of emotional EEG signals is disclosed. The embodiment of the present application also discloses a depression detection system based on the joint features of emotional EEG signals. Figure 3 As shown in FIG, the depression detection system based on the joint features of emotional EEG signals includes:
[0158] The acquisition unit 301 is used to acquire an original emotional EEG signal; wherein the original emotional EEG signal is an emotional EEG signal that has not been subjected to feature extraction;
[0159] Extraction unit 302, used to extract differential entropy features from the original emotional EEG signal;
[0160] The construction unit 303 is used to learn the local and global dependencies of the emotional EEG signals in the spatial and frequency dimensions, as well as the dynamic evolution of the emotional EEG signals in the time dimension, through differential entropy features, a pre-trained two-dimensional convolutional neural network, and a Transformer encoder, to construct a joint deep feature of the emotional EEG signals. The time dimension is the dimension of the dynamic change of the emotional EEG signals in the time series; the spatial dimension is the dimension of the dynamic correlation of the emotional EEG signals in the spatial distribution; and the frequency dimension is the dimension of the dynamic interaction of the frequency components of the emotional EEG signals.
[0161] The classification unit 304 is used to classify the joint deep features into healthy and depressed states through a multi-layer perceptron classifier, obtain a classification result, and output it.
[0162] Furthermore, the acquisition unit 301 is specifically used to collect the original emotional EEG signals generated by emotional stimulation of the target under test in a preset scenario through a preset dry electrode EEG cap and a preset acquisition frequency; wherein the emotional stimulation includes at least happy emotional stimulation, sad emotional stimulation and neutral emotional stimulation.
[0163] Furthermore, the extraction unit 302 includes:
[0164] The filtering module is used to perform bandpass filtering on the original emotional EEG signal;
[0165] A division module is used to divide non-overlapping time windows into preset time units;
[0166] an execution module, configured to apply a Hanning window to the bandpass filtered emotional EEG signal through a non-overlapping time window and perform a fast Fourier transform to obtain a frequency domain energy distribution;
[0167] A first calculation module is used to calculate the total energy within each frequency band according to the frequency energy distribution and a preset number of frequency bands;
[0168] The second calculation module is used to calculate the differential entropy feature based on the assumption that the short-time stationary signal obeys the Gaussian distribution and the total energy within each frequency band.
[0169] Furthermore, the construction unit 303 includes:
[0170] The input module is used to input the differential entropy features into a two-dimensional convolutional neural network trained with Focal Loss as the loss function and a Transformer encoder with a multi-head self-attention mechanism;
[0171] The joint modeling module is used to jointly model emotional EEG signals in the time dimension, spatial dimension and frequency dimension through a two-dimensional convolutional neural network and a Transformer encoder with a multi-head self-attention mechanism, so as to learn the local and global dependencies of emotional EEG signals in the spatial dimension and frequency dimension, as well as the dynamic evolution process of emotional EEG signals in the time dimension, thereby completing the process of constructing the joint deep features of emotional EEG signals.
[0172] Furthermore, the classification unit 304 includes:
[0173] An input-output module, used to input the joint deep features into a multi-layer perceptron classifier, so as to output a predicted probability distribution of a health category and a depression category through the multi-layer perceptron classifier;
[0174] An output module is determined, which is used to determine and output a classification result from the predicted probability distribution of the health category and the depression category through a maximum probability criterion.
[0175] The beneficial effects of the embodiments of the present application: This solution uses a joint high-order feature extraction mechanism in the time dimension, space dimension and frequency dimension, combined with a two-dimensional convolutional neural network trained with Focal Loss as the loss function and an encoder of a Transformer structure with a multi-head self-attention mechanism, to perform high-order multi-dimensional joint modeling of emotional EEG signals, and to explore the joint deep features of emotional EEG signals in the three dimensions of time series, brain area distribution and frequency component, thereby improving the accuracy and generalization ability of depression detection, realizing efficient, accurate and automated identification of depression, and thus improving the accuracy of clinical diagnosis of depression.
[0176] An embodiment of the present application further provides a storage medium, which includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the above-mentioned depression detection method based on the joint features of emotional EEG signals.
[0177] The present application also provides an electronic device, the structure of which is shown in FIG. Figure 4 As shown, it specifically includes a memory 401 and one or more instructions 402, wherein one or more instructions 402 are stored in the memory 401 and are configured to be executed by one or more processors 403 to execute the one or more instructions 402 to perform the above-mentioned depression detection method based on the combined features of emotional EEG signals.
[0178] For the sake of simplicity, the aforementioned method embodiments are described as a series of action combinations. However, those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0179] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similarities between the various embodiments can be referred to in conjunction with each other. For system-related embodiments, since they are generally similar to method-related embodiments, their description is relatively simple. For relevant details, refer to the description of the method-related embodiments.
[0180] The steps in the methods of the various embodiments of the present application can be adjusted in sequence, combined, and deleted according to actual needs.
[0181] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are merely used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations.
[0182] The above description of the disclosed embodiments will enable those skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
[0183] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A depression detection method based on the joint features of emotional EEG signals, characterized in that: The method comprises: Acquire an original emotional EEG signal; wherein the original emotional EEG signal is an emotional EEG signal that has not been subjected to feature extraction; Extracting differential entropy features from the original emotional EEG signal; By using the differential entropy features, pre-trained two-dimensional convolutional neural network and Transformer encoder, the local and global dependencies of emotional EEG signals in the spatial and frequency dimensions, as well as the dynamic evolution of emotional EEG signals in the time dimension, are learned to construct joint deep features of emotional EEG signals; wherein, the time dimension is the dimension of dynamic changes in the time series of emotional EEG signals; the spatial dimension is the dimension of dynamic correlation of the spatial distribution of emotional EEG signals; and the frequency dimension is the dimension of dynamic interaction of frequency components of emotional EEG signals; The joint deep features are classified into healthy and depressed states through a multi-layer perceptron classifier to obtain and output the classification results.
2. The method according to claim 1, characterized in that The obtaining of the original emotional EEG signal includes: By using a preset dry electrode EEG cap and a preset acquisition frequency, the original emotional EEG signals generated by emotional stimulation of the target under the preset scenario are collected; The emotional stimulation includes at least happy emotional stimulation, sad emotional stimulation and neutral emotional stimulation.
3. The method according to claim 1, characterized in that The step of extracting differential entropy features from the original emotional EEG signal includes: performing bandpass filtering on the original emotional EEG signal; Divide non-overlapping time windows into preset time units; Applying a Hanning window to the emotional EEG signal after bandpass filtering through the non-overlapping time window and performing a fast Fourier transform to obtain a frequency domain energy distribution; Calculating the total energy within each frequency band according to the frequency energy distribution and a preset number of frequency bands; Based on the assumption that the short-time stationary signal obeys the Gaussian distribution and the total energy in each frequency band, the differential entropy feature is calculated.
4. The method according to claim 1, wherein The method uses the differential entropy features, pre-trained two-dimensional convolutional neural network, and Transformer encoder to learn the local and global dependencies of emotional EEG signals in the spatial and frequency dimensions, as well as the dynamic evolution of emotional EEG signals in the time dimension, to construct joint deep features of emotional EEG signals, including: Inputting the differential entropy features into a two-dimensional convolutional neural network trained with a Focal Loss function as a loss function and a Transformer encoder with a multi-head self-attention mechanism; Through the two-dimensional convolutional neural network and the Transformer encoder with a multi-head self-attention mechanism, emotional EEG signals are jointly modeled in the time dimension, spatial dimension and frequency dimension to learn the local and global dependencies of emotional EEG signals in the spatial dimension and frequency dimension, as well as the dynamic evolution process of emotional EEG signals in the time dimension, thereby completing the process of constructing the joint deep features of emotional EEG signals.
5. The method according to claim 1, wherein The method of classifying the joint depth features into healthy and depressed states by a multi-layer perceptron classifier, obtaining and outputting the classification results, includes: Inputting the joint deep features into a multi-layer perceptron classifier to output a predicted probability distribution of a healthy category and a depressed category through the multi-layer perceptron classifier; The classification result is determined from the predicted probability distribution of the healthy category and the depressed category by using the maximum probability criterion and is output.
6. A depression detection system based on joint features of emotional EEG signals, characterized by: The system comprises: An acquisition unit, configured to acquire an original emotional EEG signal; wherein the original emotional EEG signal is an emotional EEG signal that has not been subjected to feature extraction; An extraction unit, configured to extract differential entropy features from the original emotional EEG signal; A construction unit is used to learn the local and global dependencies of emotional EEG signals in the spatial and frequency dimensions, as well as the dynamic evolution process of emotional EEG signals in the time dimension, through the differential entropy features, the pre-trained two-dimensional convolutional neural network and the Transformer encoder, to construct a joint deep feature of the emotional EEG signals; wherein the time dimension is the dimension of the dynamic change of the emotional EEG signals in the time series; the spatial dimension is the dimension of the dynamic correlation of the spatial distribution of the emotional EEG signals; and the frequency dimension is the dimension of the dynamic interaction of the frequency components of the emotional EEG signals; The classification unit is used to classify the joint deep features into healthy and depressed states through a multi-layer perceptron classifier, obtain a classification result and output it.
7. The system according to claim 6, characterized in that The acquisition unit is specifically used to collect the original emotional EEG signals generated by emotional stimulation of the target under test in a preset scenario through a preset dry electrode EEG cap and a preset acquisition frequency; wherein the emotional stimulation includes at least happy emotional stimulation, sad emotional stimulation and neutral emotional stimulation.
8. The system according to claim 6, wherein: The extraction unit comprises: A filtering module, configured to perform bandpass filtering on the original emotional EEG signal; A division module is used to divide non-overlapping time windows into preset time units; an execution module, configured to apply a Hanning window to the bandpass filtered emotional EEG signal through the non-overlapping time window, and perform a fast Fourier transform to obtain a frequency domain energy distribution; A first calculation module, configured to calculate the total energy within each frequency band according to the frequency energy distribution and a preset number of frequency bands; The second calculation module is used to calculate the differential entropy feature based on the assumption that the short-time stationary signal obeys the Gaussian distribution and the total energy within each frequency band.
9. A storage medium, characterized in that: The storage medium includes stored instructions, wherein when the instructions are executed, the device where the storage medium is located is controlled to execute the depression detection method based on the joint features of emotional EEG signals according to any one of claims 1 to 5.
10. An electronic device, characterized in that: It includes a memory and one or more instructions, wherein one or more instructions are stored in the memory and are configured to be executed by one or more processors to perform the depression detection method based on the joint features of emotional EEG signals as described in any one of claims 1 to 5.
Citation Information
Cited By
Depression recognition system based on electroencephalogram signals
CN121489484A