An interpretable depression recognition method and system based on feature decoupling
By using a domain adversarial neural network model with feature decoupling, the joint features of audio features and EEG responses are extracted, which solves the problems of high computational complexity and insufficient transparency in depression detection. This achieves highly accurate and robust depression identification and provides multi-level interpretive analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- LANZHOU UNIV
- Filing Date
- 2025-07-08
- Publication Date
- 2026-05-22
Smart Images

Figure CN120837102B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio signal recognition technology, and in particular to an interpretable depression recognition method and system based on feature decoupling. Background Technology
[0002] Electroencephalography (EEG) signals, as an objective assessment tool, have been widely used in depression detection due to their ability to provide objective measurements of brain activity. Numerous studies have confirmed the link between brain activity and depression, especially the crucial role of the prefrontal cortex in affective dysfunction in depression. With the development of portable EEG devices, automated depression detection methods combined with machine learning technologies have shown great potential. Researchers have developed detection methods based on portable three-channel EEG devices, typically collecting EEG signals from the prefrontal cortex (Fp1, Fp2, and Fpz), ensuring both convenient signal acquisition and focusing on key brain regions related to depression. Furthermore, detection methods combining EEG with audio stimulation have also shown special value, because patients with depression exhibit asymmetric activity in the left and right hemispheres of the prefrontal cortex, and mood-inducing paradigms (such as audio stimulation) can more effectively activate neural circuits associated with depression.
[0003] Effective feature selection is crucial when developing machine learning (ML)-based depression detection systems using electroencephalogram (EEG) signals. Researchers have employed various feature selection methods, such as Fisher vector clustering, sparse coding, principal component analysis, and correlation filtering, to identify the most discriminative and relevant neural patterns from high-dimensional and complex EEG features, thereby reducing feature dimensionality and improving the efficiency of depression detection. Traditional feature selection strategies can be broadly categorized into filtering, wrapping, and embedding methods. Filtering methods rank features based on their inherent properties, improving computational efficiency; wrapping methods evaluate feature subsets based on the performance of a specific model, typically achieving better predictive accuracy; and embedding methods integrate the feature selection process into model training to balance performance and efficiency.
[0004] Meanwhile, with the widespread application of deep learning models in the healthcare field, explainable artificial intelligence (XAI) methods have emerged. XAI methods are mainly divided into two categories: ex-ante explanation and ex-post explanation. Ex-ante explanation methods directly construct interpretable model structures, such as linear regression; ex-post explanation methods provide explanations for the trained model, such as permutation feature importance and SHAP.
[0005] However, existing technologies still have many shortcomings. On the one hand, although feature selection methods are diverse, traditional feature selection strategies have limitations. For example, filtering methods may ignore feature interactions related to the classifier; wrapping methods have high computational complexity, especially inefficient in large feature spaces; and the effectiveness of embedding methods may be closely related to the assumptions of the chosen learning algorithm. On the other hand, many complex machine learning and deep learning models are criticized as "black boxes," lacking transparency and interpretability, which limits their application in medical diagnosis. While existing interpretable artificial intelligence methods can estimate the contribution of each feature to the model output, they lack in-depth explanations of feature importance, making it difficult to identify robust, situation-specific biomarkers, thus hindering a deeper understanding of the neurophysiology of depression. Summary of the Invention
[0006] The purpose of this invention is to provide an interpretable depression identification method and system based on feature decoupling, which solves the problems of high computational complexity, low accuracy and poor applicability of existing technologies.
[0007] To achieve the above objectives, this invention provides an interpretable depression identification method based on feature decoupling, comprising the following steps:
[0008] S1. Obtain the audio feature information of the currently playing audio and the EEG response information of the person to be identified, and obtain comprehensive features based on the audio feature information and EEG response information;
[0009] Among them, the comprehensive features include emotion recognition delay features, pitch sensitivity features, rhythm perception features, and emotion transition features;
[0010] S2. Construct a feature-enhanced domain adversarial network model based on the domain adversarial neural network mechanism and feature decoupling mechanism;
[0011] S3. Based on the feature-enhanced domain adversarial network model, feature extraction, domain adversarial and adaptive processes, depression state classification, and feature orthogonality constraints are performed sequentially.
[0012] S4. The gradient descent method is used to optimize the feature-enhanced domain adversarial network model, and the depression identification of the subject is performed based on the optimized feature-enhanced domain adversarial network model.
[0013] In some embodiments of this application, in S1, obtaining audio feature information of the currently playing audio and EEG response information of the person to be identified, and obtaining comprehensive features based on the audio feature information and EEG response information includes:
[0014] S11. Calculate the emotion recognition delay features, specifically including:
[0015] The expression for obtaining the emotion change points identified from the audio signal is:
[0016] ;
[0017] in, A set representing points of emotional change in audio. and They represent the first Frame and the Mel-spectral coefficients of the frame and These are the mean and standard deviation of the Mel-frequency cepstral coefficients, respectively. These are the points where the audio's emotional tone changes.
[0018] For each detected change in audio emotion, its corresponding response signal in the electroencephalogram (EEG) signal is obtained, expressed as:
[0019] ;
[0020] in, In response to the signal, This is the response intensity function of the electroencephalogram (EEG) signal. and These represent the mean and standard deviation of the Mel-frequency cepstral coefficients of the EEG signal, respectively. These are the points where the EEG signal response changes.
[0021] If an EEG response signal is detected within the maximum delay time, the time difference between the earliest significant EEG response and the stimulus is calculated, expressed as:
[0022] ;
[0023] in, For time difference;
[0024] If no EEG response signal is detected, the delay at that point is set to the maximum delay.
[0025] S12. Calculate pitch sensitivity features, specifically including:
[0026] The fundamental frequency of each audio frame is estimated using the autocorrelation method, and then significant fundamental frequency changes are detected:
[0027] ;
[0028] in, and The first Frame and the The estimated fundamental frequency of the frame. This is the set of points with significant fundamental frequency changes. and These are the mean and standard deviation of the Gaussian function, respectively.
[0029] Calculate the response intensity of the EEG signal at these points of change:
[0030] ;
[0031] in, and The points of change are respectively The average energy of the pre- and post-electroencephalogram (EEG) signals, For time sets, The response intensity of the EEG signal at the point of change;
[0032] S13. Calculate the Rhythm Perception Feature (RP), which specifically includes:
[0033] The main rhythmic period is calculated based on the peak position of the envelope autocorrelation function, and the change points of the envelope are detected. The expression is as follows:
[0034] ;
[0035] in, It is the first-order difference of the audio envelope. It is the set of points where the envelope changes significantly;
[0036] The EEG at the point of change was calculated. α The change in wave energy is expressed as:
[0037] ;
[0038] in, and The points of change are respectively Forehead and back EEG signals α Average power of the band, for α Wave energy change value;
[0039] S14. Calculate the characteristics of emotion transition, specifically including:
[0040] The change in sentiment parameters between adjacent windows is calculated using the following expression:
[0041] ;
[0042] in, and These represent the arousal and valence of audio, respectively, with subscripts. and These are before and during audio playback, respectively. The Euclidean distance for changes in emotional state;
[0043] The corresponding brainwave frequency band energy change value is obtained, and the expression is:
[0044] ;
[0045] in, They are respectively θ Wave, α wave and β Average power in the frequency band, This represents the relative rate of change of power in the corresponding frequency band;
[0046] The correlation coefficient between the intensity of emotional changes and the intensity of EEG response was calculated, and the characteristics of emotional transition were obtained based on the correlation coefficient.
[0047] In some embodiments of this application, in S2, constructing a feature-enhanced domain adversarial network model based on a domain adversarial neural network mechanism and a feature decoupling mechanism includes:
[0048] A domain adversarial neural network mechanism is embedded in the feature decoupling network architecture, and a feature-enhanced domain adversarial network model is constructed by combining it with a recognition algorithm module. The recognition algorithm module includes a feature extraction module, a domain adversarial adaptation module, multiple label prediction modules, and a feature orthogonality constraint module.
[0049] In some embodiments of this application, S3, the sequential performance of feature extraction, domain adversarial analysis, and adaptation based on the feature-enhanced domain adversarial network model includes:
[0050] S31. Use the feature extraction module in the feature enhancement domain adversarial network model to extract shared and private features from the input EEG feature vector;
[0051] Among them, shared features are domain-invariant features of the subject group, and private features are features of the subject group that contain information on the discrimination of depressive state.
[0052] S32. The domain adversarial adaptation module performs domain adversarial operations on shared features. Based on the domain classifier and gradient inversion layer, it performs domain classification and gradient inversion on the shared features, making the shared features indistinguishable in different domains.
[0053] In some embodiments of this application, S3, the sequential classification of depressive states and constraint of feature orthogonality based on the feature-enhanced domain adversarial network model includes:
[0054] S33. The label prediction module uses the extracted shared features and private features to classify the depression state. The private feature classifier performs the initial classification based on the private features and obtains specific information directly related to the depression state. The combined feature classifier combines the shared features and private features to perform the secondary classification and outputs the depression state classification result.
[0055] S34. The feature orthogonality constraint module applies orthogonality constraints to shared features and private features, minimizing the absolute cosine similarity between shared features and private features, and separating shared features and private features in the feature space.
[0056] In some embodiments of this application, in step S4, optimizing the feature enhancement domain adversarial network model using gradient descent includes:
[0057] The feature enhancement domain adversarial network model is optimized and trained based on the combined loss function and gradient descent method, as expressed in:
[0058] ;
[0059] in, For the combined loss function, and All of these are preset hyperparameters. For orthogonality loss function, The loss function for the shared feature classifier, The loss function is the characteristic orthogonality constraint. This is the loss function for the private feature classifier.
[0060] In some embodiments of this application, an interpretable depression recognition system based on feature decoupling is also disclosed, comprising:
[0061] The data acquisition module is used to acquire the audio feature information of the currently playing audio and the EEG response information of the person to be identified, and to obtain comprehensive features based on the audio feature information and the EEG response information;
[0062] Among them, the comprehensive features include emotion recognition delay features, pitch sensitivity features, rhythm perception features, and emotion transition features;
[0063] The model building module is used to build feature-enhanced domain adversarial network models based on domain adversarial neural network mechanisms and feature decoupling mechanisms.
[0064] The model update module is used to perform feature extraction, domain adversarial and adaptive processes, depression state classification, and feature orthogonality constraints sequentially based on the feature-enhanced domain adversarial network model.
[0065] The depression identification module is used to optimize the feature-enhanced domain adversarial network model using gradient descent, and then uses the optimized feature-enhanced domain adversarial network model to identify the depression of the subject.
[0066] The beneficial effects of this invention are:
[0067] 1. This invention extracts joint features of specific physical properties of audio and EEG responses to explore how EEG signals dynamically reflect the physical characteristics of sound in real time. This invention detects changes in audio physical features (spectrum, pitch, and rhythm, etc.) and quantifies the immediate EEG responses triggered by these changes. Compared with traditional EEG features, audio-based joint features not only improve recognition accuracy but also provide multimodal evidence for research on the neural mechanisms of depression.
[0068] 2. This invention proposes a novel feature decoupling framework based on domain adversarial neural networks for depression detection. This framework separates a pre-extracted EEG feature space into shared (domain-invariant) and private (domain-specific) components. Through adversarial training, shared features become indistinguishable across different emotional domains, while orthogonality constraints ensure that private features capture domain-specific variations. Compared to traditional methods, this approach achieves effective feature decoupling, improving the accuracy and robustness of depression identification. Furthermore, it provides a structured framework for understanding the commonalities and specific differences in the pathophysiological mechanisms of depression.
[0069] 3. This invention employs multi-level interpretability analysis to reveal the decision-making mechanism of the E-DANN model. Based on the SHAP method, this invention simultaneously provides explanations at three levels: global importance, local interpretation, and feature interaction. This multi-level interpretation framework compensates for the lack of transparency in deep learning models and provides a paradigm for constructing interpretable EEG analysis systems. Furthermore, this invention, combined with an audio paradigm, reveals the differential activation mechanisms of specific audio stimuli and the complexity of feature interaction patterns.
[0070] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0071] Figure 1 This is a schematic diagram illustrating the steps of an interpretable depression identification method based on feature decoupling in an embodiment of the present invention;
[0072] Figure 2 This is a structural diagram of an interpretable depression recognition system based on feature decoupling, according to an embodiment of the present invention.
[0073] Figure 3 This is a flowchart illustrating the construction of the multimodal feature space in an embodiment of the present invention.
[0074] Figure 4 This is a statistical result graph of the 10-fold cross-validation results of the Wilcoxon model in an embodiment of the present invention;
[0075] Figure 5A comparison chart of ROC curves for the feature-enhanced domain adversarial network model, machine learning (RF) model, and deep learning (DANN) model in embodiments of the present invention;
[0076] Figure 6 This is a graph showing the feature decoupling performance verification results of an embodiment of the present invention;
[0077] Figure 7 This is a diagram showing the interpretability analysis results of the E-DANN model in the depression recognition task according to an embodiment of the present invention. Detailed Implementation
[0078] In the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product is in use. They are used only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," and "connect" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0079] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0080] like Figure 1 As shown, this invention provides an interpretable depression identification method based on feature decoupling, comprising the following steps:
[0081] S1. Obtain the audio feature information of the currently playing audio and the EEG response information of the person to be identified, and obtain comprehensive features based on the audio feature information and EEG response information;
[0082] Among them, the comprehensive features include emotion recognition delay features, pitch sensitivity features, rhythm perception features, and emotion transition features.
[0083] S2. Construct a feature-enhanced domain adversarial network model based on the domain adversarial neural network mechanism and feature decoupling mechanism.
[0084] S3. Based on the feature-enhanced domain adversarial network model, feature extraction, domain adversarial and adaptive processes, depression state classification, and feature orthogonality constraints are performed sequentially.
[0085] S4. The gradient descent method is used to optimize the feature-enhanced domain adversarial network model, and the depression identification of the subject is performed based on the optimized feature-enhanced domain adversarial network model.
[0086] Specifically, this invention further enhances features based on Domain Adversarial Networks (DANNs) to form a unique Feature-Enhanced Domain Adversarial Network (E-DANN). To address the challenges of effective feature selection under different conditions and the need for enhanced interpretability to support biomarker discovery, this invention proposes a depression recognition model based on Enhanced Domain Adversarial Network (E-DANN). This invention first constructs a multimodal space, then uses a feature decoupling network to separate features and achieve depression detection.
[0087] Specifically, in order to effectively distinguish and utilize the shared neural patterns (domain-invariant shared features) and the unique neural patterns (domain-specific private features) of healthy control groups and patients with depression when processing emotional audio, this invention proposes an enhanced domain adversarial network (E-DANN) architecture.
[0088] In some embodiments of this application, in S1, obtaining audio feature information of the currently playing audio and EEG response information of the person to be identified, and obtaining comprehensive features based on the audio feature information and EEG response information includes:
[0089] S11. Calculate the Emotion Recognition Delay Feature (ERD), which specifically includes:
[0090] The expression for obtaining the emotion change points identified from the audio signal is:
[0091] ;
[0092] in, A set representing points of emotional change in audio. and They represent the first Frame and the Mel-spectral coefficients of the frame and These are the mean and standard deviation of the Mel-frequency cepstral coefficients, respectively. These are the points where the audio's emotional tone changes.
[0093] For each detected change in audio emotion, its corresponding response signal in the electroencephalogram (EEG) signal is obtained, expressed as:
[0094] ;
[0095] in, In response to the signal, This is the response intensity function of the electroencephalogram (EEG) signal. and These represent the mean and standard deviation of the Mel-frequency cepstral coefficients of the EEG signal, respectively. These are the points where the EEG signal response changes.
[0096] If an EEG response signal is detected within the maximum delay time, the time difference between the earliest significant EEG response and the stimulus is calculated, expressed as:
[0097] ;
[0098] in, For time difference;
[0099] If no EEG response signal is detected, the delay at that point is set to the maximum delay.
[0100] S12. Calculate the pitch sensitivity feature (PS), which specifically includes:
[0101] The fundamental frequency of each audio frame is estimated using the autocorrelation method, and then significant fundamental frequency changes are detected:
[0102] ;
[0103] in, and The first Frame and the The estimated fundamental frequency of the frame. This is the set of points with significant fundamental frequency changes. and These are the mean and standard deviation of the Gaussian function, respectively.
[0104] Calculate the response intensity of the EEG signal at these points of change:
[0105] ;
[0106] in, and The points of change are respectively The average energy of the pre- and post-electroencephalogram (EEG) signals, For time sets, The response intensity of the EEG signal at the point of change;
[0107] S13. Calculate the Rhythm Perception Feature (RP), which specifically includes:
[0108] The main rhythmic period is calculated based on the peak position of the envelope autocorrelation function, and the change points of the envelope are detected. The expression is as follows:
[0109] ;
[0110] in, It is the first-order difference of the audio envelope. It is the set of points where the envelope changes significantly;
[0111] The EEG at the point of change was calculated. α The change in wave energy is expressed as:
[0112] ;
[0113] in, and The points of change are respectively Forehead and back EEG signals α Average power of the band, for α Wave energy change value;
[0114] S14. Calculate the Emotion Transition Feature (ET), which specifically includes:
[0115] The change in sentiment parameters between adjacent windows is calculated using the following expression:
[0116] ;
[0117] in, and These represent the arousal and valence of audio, respectively, with subscripts. and These are before and during audio playback, respectively. The Euclidean distance for changes in emotional state;
[0118] The corresponding brainwave frequency band energy change value is obtained, and the expression is:
[0119] ;
[0120] in, They are respectively θ Wave, α wave and β Average power in the frequency band, This represents the relative rate of change of power in the corresponding frequency band;
[0121] The correlation coefficient between the intensity of emotional changes and the intensity of EEG response was calculated, and the characteristics of emotional transition were obtained based on the correlation coefficient.
[0122] In some embodiments of this application, in S2, constructing a feature-enhanced domain adversarial network model based on a domain adversarial neural network mechanism and a feature decoupling mechanism includes:
[0123] A domain adversarial neural network mechanism is embedded in the feature decoupling network architecture, and a feature-enhanced domain adversarial network model is constructed by combining it with a recognition algorithm module. The recognition algorithm module includes a feature extraction module, a domain adversarial adaptation module, multiple label prediction modules, and a feature orthogonality constraint module.
[0124] In some embodiments of this application, S3, the sequential performance of feature extraction, domain adversarial analysis, and adaptation based on the feature-enhanced domain adversarial network model includes:
[0125] S31. Use the feature extraction module in the feature enhancement domain adversarial network model to extract shared and private features from the input EEG feature vector;
[0126] Among them, shared features are domain-invariant features of the subject group, and private features are features of the subject group that contain information on the discrimination of depressive state.
[0127] S32. The domain adversarial adaptation module performs domain adversarial operations on shared features. Based on the domain classifier and gradient inversion layer, it performs domain classification and gradient inversion on the shared features, making the shared features indistinguishable in different domains.
[0128] In some embodiments of this application, S3, the sequential classification of depressive states and constraint of feature orthogonality based on the feature-enhanced domain adversarial network model includes:
[0129] S33. The label prediction module uses the extracted shared features and private features to classify the depression state. The private feature classifier performs the initial classification based on the private features and obtains specific information directly related to the depression state. The combined feature classifier combines the shared features and private features to perform the secondary classification and outputs the depression state classification result.
[0130] S34. The feature orthogonality constraint module applies orthogonality constraints to shared features and private features, minimizing the absolute cosine similarity between shared features and private features, and separating shared features and private features in the feature space.
[0131] In some embodiments of this application, in step S4, optimizing the feature enhancement domain adversarial network model using gradient descent includes:
[0132] The feature enhancement domain adversarial network model is optimized and trained based on the combined loss function and gradient descent method, as expressed in:
[0133] ;
[0134] in, For the combined loss function, and All of these are preset hyperparameters. For orthogonality loss function, The loss function for the shared feature classifier, The loss function is the characteristic orthogonality constraint. This is the loss function for the private feature classifier.
[0135] In some embodiments of this application, such as Figure 2 As shown, an interpretable depression recognition system based on feature decoupling is also disclosed, including:
[0136] The data acquisition module is used to acquire the audio feature information of the currently playing audio and the EEG response information of the person to be identified, and to obtain comprehensive features based on the audio feature information and EEG response information;
[0137] Among them, the comprehensive features include emotion recognition delay features, pitch sensitivity features, rhythm perception features, and emotion transition features;
[0138] The model building module is used to build feature-enhanced domain adversarial network models based on domain adversarial neural network mechanisms and feature decoupling mechanisms.
[0139] The model update module is used to perform feature extraction, domain adversarial and adaptive processes, depression state classification, and feature orthogonality constraints sequentially based on the feature-enhanced domain adversarial network model.
[0140] The depression identification module is used to optimize the feature-enhanced domain adversarial network model using the gradient descent method, and then uses the optimized feature-enhanced domain adversarial network model to identify depression in the subject.
[0141] The beneficial effects of this invention are:
[0142] 1. This invention extracts joint features of specific physical properties of audio and EEG responses to explore how EEG signals dynamically reflect the physical characteristics of sound in real time. This invention detects changes in audio physical features (spectrum, pitch, and rhythm, etc.) and quantifies the immediate EEG responses triggered by these changes. Compared with traditional EEG features, audio-based joint features not only improve recognition accuracy but also provide multimodal evidence for research on the neural mechanisms of depression.
[0143] 2. This invention proposes a novel feature decoupling framework based on domain adversarial neural networks for depression detection. This framework separates a pre-extracted EEG feature space into shared (domain-invariant) and private (domain-specific) components. Through adversarial training, shared features become indistinguishable across different emotional domains, while orthogonality constraints ensure that private features capture domain-specific variations. Compared to traditional methods, this approach achieves effective feature decoupling, improving the accuracy and robustness of depression identification. Furthermore, it provides a structured framework for understanding the commonalities and specific differences in the pathophysiological mechanisms of depression.
[0144] 3. This invention employs multi-level interpretability analysis to reveal the decision-making mechanism of the E-DANN model. Based on the SHAP method, this invention simultaneously provides explanations at three levels: global importance, local interpretation, and feature interaction. This multi-level interpretation framework compensates for the lack of transparency in deep learning models and provides a paradigm for constructing interpretable EEG analysis systems. Furthermore, this invention, combined with an audio paradigm, reveals the differential activation mechanisms of specific audio stimuli and the complexity of feature interaction patterns.
[0145] The embodiments of the present invention will be described in detail below with reference to specific examples.
[0146] For the three-channel EEG signals collected under audio stimulation by the portable EEG acquisition system developed by the Gansu Provincial Key Laboratory of Wearable Computing at Lanzhou University, this invention uses a 500th-order finite pulse filter to filter the signals to 0.1-45Hz. This avoids low-frequency drift, as well as the effects of 50Hz power line interference and high-frequency electromyography noise, while preserving the main frequency bands for EEG signal analysis. Subsequently, this invention uses discrete wavelet transform combined with Kalman filtering to remove electrooculography artifacts. This step is important because the acquisition electrodes of this invention are very close to the eye region. In addition, to explore the differences in participants' responses to emotional audio stimulation, this invention extracts EEG data under audio stimulation.
[0147] Multimodal feature space construction: This invention extracts the respective attribute features of audio signals and EEG signals, including audio features (Mel-spectral coefficients, fundamental frequency, rhythmic period, and emotional valence) and EEG features (frequency domain, linear, and nonlinear features). Audio valence is shown in Table 1. The complete construction process is as follows: Figure 3As shown in section 3.1, joint features of audio-specific physical properties and EEG responses (emotion recognition delay features, pitch sensitivity features, rhythm perception features, and emotion transition features) were first extracted. Then, EEG-related features (frontal alpha asymmetry (FAA), Lempel-Ziv Complex (LZC), Correlation Dimension (CD), Sample Entropy (SamEn), and Shannon Entropy (ShanEn)) were added. 3*8 EEG features were obtained from the Fp1, Fpz, and Fp2 electrodes. These features include four traditional EEG features and four joint features. Furthermore, the merging of the FAA indicators resulted in a total of 25 extracted features, enabling a more nuanced and comprehensive analysis of the EEG signals.
[0148] Table 1. Detailed description of audio stimuli
[0149]
[0150] Depression Recognition Based on E-DANN: This invention compares the performance and effectiveness of E-DANN using baseline models and validates the models. This includes five classic machine learning (ML) methods (Support Vector Machine (SVM), K-Nearest Neighbor (KNN), Logistic Regression (LR), Random Forest (RF), and Extreme Gradient Boosting (XGBOOST)) and three deep learning (DL) methods (Multilayer Perceptron (MLP), EEGNet, LSDD_EEGNe, and DANN). Table 2 lists details related to the parameters of each model. Experiments were conducted in a Python 3.9 environment, with an Intel® Core™ i5-10400F CPU and an NVIDIA GeForce RTX 4090 GPU (24GB VRAM). The DL models were implemented using the PyTorch 2.2 framework supporting CUDA 12.4. To prevent data leakage and imbalanced samples from affecting model performance evaluation, this invention uses StratifiedGroupKFold for 10-fold cross-validation.
[0151] Table 2. Model Parameter Statistics
[0152]
[0153] To rigorously evaluate the performance differences between E-DANN and the baseline model, this invention performed a comprehensive statistical analysis of the results of 10-fold cross-validation of each model using Wilcoxon's test. The statistical significance of the observed differences is as follows: Figure 4As shown, p-values are explicitly annotated for pairwise comparisons. Compared to most traditional machine learning methods (such as SVM, KNN, and LR), E-DANN achieves significant improvements in sensitivity metrics (p=0.035 and p=0.018, respectively), indicating that the proposed method has a statistical advantage in identifying patients with depression. Compared to deep learning methods, E-DANN demonstrates significant advantages over MLP, EEGNet, and LSDD_EEGNet on all metrics (p-values ranging from 0.005 to 0.041), particularly in accuracy and F1 score.
[0154] To evaluate the classification performance of each model, this invention uses 10-fold cross-validation to generate the Receiver Operating Curve (ROC) curve. Specifically, in each fold of the cross-validation aggregated prediction, this invention collects the true labels and corresponding predicted probabilities from the classifier outputs in the validation set. The aggregated true labels and predicted probabilities form a complete evaluation sequence, which is then used to calculate the ROC curve and quantize the area under the curve (AUC). Figure 5 The ROC curve based on the comprehensive cross-validation results is shown. The model proposed in this invention achieved a maximum area under the curve (AUC) of 98.8%, which is significantly better than the DANN model (AUC=90.1%) and the RF model (AUC=80.2%).
[0155] Furthermore, in order to evaluate the orthogonality constraint in the proposed E-DANN model for decoupling shared features... and private characteristics Regarding effectiveness, this invention calculates the effectiveness of each fold of the validation set in 10-fold cross-validation. and Cosine similarity between vectors. The results are visualized in two ways: showing the overall orthogonality distribution and demonstrating its stability across different folds. For example... Figure 6 As shown, firstly, Figure 6 (a) shows the shared features of all validation set samples during 10-fold cross-validation. With private characteristics The distribution of cosine similarity between the data sets is shown. It also includes a kernel density estimation (KDE) curve to smoothly illustrate the distribution shape. This distribution exhibits an approximately bell-shaped (single-peak) pattern, with the peak value located at the mean of 0.028 (red dashed line), very close to the ideal zero point (gray dashed line). Secondly, to evaluate the stability of the decoupling effect across different data subsets, this invention calculates the mean absolute cosine similarity for each fold validation set and... Figure 6(b) shows a box plot of its distribution. The box plot summarizes the statistical characteristics of these 10 folds (median, quartiles, etc.), with each black dot representing a specific value of a fold. The mean absolute cosine similarity measures the average degree to which the similarity deviates from zero within that fold. The results show that the mean absolute cosine similarity of each fold is generally low (median approximately 0.037), and the differences between folds are small (interquartile range approximately [0.035-0.048]). In addition, all black dots are relatively concentrated in the range of [0.022-0.048]. These results indicate that the feature decoupling effect achieved by the model has good stability and consistency under different data partitions.
[0156] Multi-level interpretability analysis: To delve into the decision-making mechanism of the proposed E-DANN in the depression identification task, this invention employs a method based on SHAP theory for model interpretability analysis, such as... Figure 7 As shown. First, based on the trained E-DANN model, this invention calculates the SHAP values of all features in each sample. At the global interpretation level, this invention visually displays the average importance ranking and influence distribution of all features in the overall sample through feature importance bar charts and SHAP summary charts, showcasing the top features that have a significant impact on the model's prediction results ( Figure 7 (a)). At the local interpretation level, this invention uses SHAP dependency graphs to show how key features influence model predictions and the potential interactions between features. This invention generates SHAP dependency graphs for the top 5 features by global importance (a). Figure 7 (b) To reveal potential feature interaction effects, these dependency graphs use color coding to display the effects of other features that interact most strongly with the main graph feature; the SHAP library automatically selects this feature based on an algorithm (interaction_index='auto'). Furthermore, this invention also overlays feature importance values under different audio stimuli (...). Figure 7 (c)).
[0157] like Figure 7 As shown, the interpretability analysis results of the E-DANN model based on the SHAP method in the depression identification task are presented. Specifically, Figure 7 The global feature importance analysis in (a) shows that ERD_Fp1 is the most influential feature, with an average SHAP value of approximately 0.14, followed by RP_Fp1, PS_Fp3, Theta_Fp1, and ET_v. The right-hand figure visually shows the distribution of SHAP values for each feature, where red dots represent the influence that pushes the prediction toward the depression category, while blue dots represent the influence that favors the non-depression category. In most samples, ERD_Fp1 mainly plays a positive role (more red dots), while PS_Fp3 shows a strong negative effect in the middle range (more blue dots). Figure 7 In (b), the dependency plot of ERD_Fp1 shows that the SHAP value increases significantly when the value of ERD_Fp1 exceeds 2, indicating a clear interaction with Theta_Fp1. The dependency plot of RP_Fp1 highlights its interaction with LZC_Fp3, showing a complex pattern. Similarly, the dependency plot of PS_Fp3 shows a strong interaction with LZC_Fp3, particularly at higher PS_Fp3 levels, manifested as noticeable color changes. The dependency plot of Theta_Fp1 shows its interaction with ERD_Fp1, and its contribution to depression prediction becomes more pronounced once Theta_Fp exceeds a threshold of around 3. Furthermore, the interaction between ET_v and FAA4 is particularly significant, reflected in the noticeable color changes within the ET_v value range. Figure 7 In (c), the first (cow), third (baby crying) and fifth (baby) audio segments were observed to have high feature importance values.
[0158] In this application, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. In case of any inconsistency, the meaning set forth in this specification or derived from the content described herein shall prevail. Furthermore, the terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit the scope of this application.
[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for interpretable depression identification based on feature decoupling, characterized in that, Includes the following steps: S1. Obtain the audio feature information of the currently playing audio and the EEG response information of the person to be identified, and obtain comprehensive features based on the audio feature information and EEG response information; Among them, the comprehensive features include emotion recognition delay features, pitch sensitivity features, rhythm perception features, and emotion transition features; S2. Construct a feature-enhanced domain adversarial network model based on the domain adversarial neural network mechanism and feature decoupling mechanism; In S2, the feature-enhanced domain adversarial network model constructed based on the domain adversarial neural network mechanism and the feature decoupling mechanism includes: A domain adversarial neural network mechanism is embedded in the feature decoupling network architecture, and a feature-enhanced domain adversarial network model is constructed by combining it with a recognition algorithm module. The recognition algorithm module includes a feature extraction module, a domain adversarial adaptation module, multiple label prediction modules, and a feature orthogonality constraint module. S3. Based on the feature-enhanced domain adversarial network model, feature extraction, domain adversarial and adaptive processes, depression state classification, and feature orthogonality constraints are performed sequentially. In S3, the feature-enhanced domain adversarial network model sequentially performs feature extraction, domain adversarial processes, and adaptation, including: S31. Use the feature extraction module in the feature enhancement domain adversarial network model to extract shared and private features from the input EEG feature vector; Among them, shared features are domain-invariant features of the subject group, and private features are features of the subject group that contain information on the discrimination of depressive state. S32. The domain adversarial adaptation module performs domain adversarial operations on shared features. Based on the domain classifier and gradient inversion layer, it performs domain classification and gradient inversion on the shared features, making the shared features indistinguishable in different domains. In S3, the depression state classification and feature orthogonality constraints are performed sequentially based on the feature-enhanced domain adversarial network model, including: S33. The label prediction module uses the extracted shared features and private features to classify the state of depression. The private feature classifier performs the initial classification based on the private features and obtains information directly related to the state of depression. The combined feature classifier combines the shared features and private features to perform the secondary classification and outputs the state of depression classification results. S34. The feature orthogonality constraint module applies orthogonality constraints to shared features and private features, minimizing the absolute cosine similarity between shared features and private features, and separating shared features and private features in the feature space. S4. The gradient descent method is used to optimize the feature-enhanced domain adversarial network model, and the depression identification of the subject is performed based on the optimized feature-enhanced domain adversarial network model.
2. The interpretable depression identification method based on feature decoupling according to claim 1, characterized in that, In step S1, the audio feature information of the currently playing audio and the EEG response information of the person to be identified are obtained, and the comprehensive features obtained based on the audio feature information and the EEG response information include: S11. Calculate the emotion recognition delay features, specifically including: The expression for obtaining the emotion change points identified from the audio signal is: ; in, A set representing points of emotional change in audio. and They represent the first Frame and the Mel-spectral coefficients of the frame and These are the mean and standard deviation of the Mel-frequency cepstral coefficients, respectively. These are the points where the audio's emotional tone changes. For each detected change in audio emotion, its corresponding response signal in the electroencephalogram (EEG) signal is obtained, expressed as: ; in, In response to the signal, This is the response intensity function of the electroencephalogram (EEG) signal. and These represent the mean and standard deviation of the Mel-frequency cepstral coefficients of the EEG signal, respectively. These are the points where the EEG signal response changes. If an EEG response signal is detected within the maximum delay time, the time difference between the earliest significant EEG response and the stimulus is calculated, expressed as: ; in, For time difference; If no EEG response signal is detected, the delay at that point is set to the maximum delay. S12. Calculate pitch sensitivity features, specifically including: The fundamental frequency of each audio frame is estimated using the autocorrelation method, and then significant fundamental frequency changes are detected: ; in, and The first Frame and the The estimated fundamental frequency of the frame. This is the set of points with significant fundamental frequency changes. and These are the mean and standard deviation of the Gaussian function, respectively. Calculate the response intensity of the EEG signal at these points of change: ; in, and The points of change are respectively The average energy of the pre- and post-electroencephalogram (EEG) signals, For time sets, The response intensity of the EEG signal at the point of change; S13. Calculate rhythm perception features, specifically including: The main rhythmic period is calculated based on the peak position of the envelope autocorrelation function, and the change points of the envelope are detected. The expression is as follows: ; in, It is the first-order difference of the audio envelope. It is the set of points where the envelope changes significantly; The EEG at the point of change was calculated. α The change in wave energy is expressed as: ; in, and The points of change are respectively Forehead and back EEG signals α Average power of the band, for α Wave energy change value; S14. Calculate the characteristics of emotion transition, specifically including: The change in sentiment parameters between adjacent windows is calculated using the following expression: ; in, and These represent the arousal and valence of audio, respectively, with subscripts. and These are before and during audio playback, respectively. The Euclidean distance for changes in emotional state; The corresponding brainwave frequency band energy change value is obtained, and the expression is: ; in, They are respectively θ Wave, α wave and β Average power in the frequency band, This represents the relative rate of change of power in the corresponding frequency band; The correlation coefficient between the intensity of emotional changes and the intensity of EEG response was calculated, and the characteristics of emotional transition were obtained based on the correlation coefficient.
3. The interpretable depression identification method based on feature decoupling according to claim 2, characterized in that, In step S4, optimizing the feature enhancement domain adversarial network model using gradient descent includes: The feature enhancement domain adversarial network model is optimized and trained based on the combined loss function and gradient descent method, as expressed in: ; in, For the combined loss function, and All of these are preset hyperparameters. For orthogonality loss function, The loss function for the shared feature classifier, The loss function is the characteristic orthogonality constraint. This is the loss function for the private feature classifier.
4. A feature-decoupling-based interpretable depression recognition system for implementing the method of any one of claims 1 to 3, characterized in that, include: The data acquisition module is used to acquire the audio feature information of the currently playing audio and the EEG response information of the person to be identified, and to obtain comprehensive features based on the audio feature information and the EEG response information; Among them, the comprehensive features include emotion recognition delay features, pitch sensitivity features, rhythm perception features, and emotion transition features; The model building module is used to build feature-enhanced domain adversarial network models based on domain adversarial neural network mechanisms and feature decoupling mechanisms. The model update module is used to perform feature extraction, domain adversarial and adaptive processes, depression state classification, and feature orthogonality constraints sequentially based on the feature-enhanced domain adversarial network model. The depression identification module is used to optimize the feature-enhanced domain adversarial network model using gradient descent, and then uses the optimized feature-enhanced domain adversarial network model to identify the depression of the subject.