A Deep Learning-Based Multi-State EEG Fusion Method for Identifying Monopolar and Bipolar Depression
By employing a deep learning-based multi-state EEG fusion method, utilizing continuous wavelet transform and ConvNeXt model to extract features, and combining them with the KAN model for classification, the problem of insufficient data collection complexity and accuracy in the identification of unipolar and bipolar depression is solved, achieving efficient automated identification and early intervention.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HEBEI UNIV OF TECH
- Filing Date
- 2026-02-10
- Publication Date
- 2026-05-26
AI Technical Summary
Existing technologies for identifying unipolar and bipolar depression suffer from insufficient analysis of single-modal data, resulting in incomplete feature extraction. Furthermore, the complex acquisition of multimodal data affects the convenience and accuracy of practical applications.
A deep learning-based multi-state EEG fusion method is adopted. By performing continuous wavelet transform on EEG signals in both open and closed states, a time-frequency map is generated. Features are extracted using the ConvNeXt model and combined with the KAN model for classification and recognition, thus achieving end-to-end automated assisted diagnosis.
It improves the classification and identification accuracy of unipolar and bipolar depression, simplifies the data collection process, reduces costs, is suitable for initial screening in clinical settings, and increases the possibility of early intervention for the disease.
Smart Images

Figure CN122074985A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of biomedical engineering, specifically to a method for identifying monopolar and bipolar depression based on deep learning-based multi-state EEG fusion. Background Technology
[0002] In recent years, the automatic identification of mental states and disorders has received widespread attention from the computer vision and artificial intelligence communities. Deep learning models excel at extracting complex features and learning complex patterns from data, making them well-suited for tasks involving multidimensional and sequential data analysis and image classification. However, while deep learning techniques currently perform well in mental disorder identification tasks, traditional methods often use a single neural network for training and classification, leaving room for improvement in accuracy.
[0003] Furthermore, previous studies have mostly analyzed the EEG characteristics of patients with unipolar depression and bipolar disorder in a single-modal context. Limited by single-modal data, the extracted features cannot fully reflect changes in brain neural electrical activity. In recent years, combining data from two or more modalities to classify depression has yielded promising results. For example, the paper "Ma,W. , Qiu,S. , Miao,J. , Li,M. , Tian,Z. , Zhang,B. , Li,W. , Feng,R. , Wang,C. , & Cui , Y. (2023). Detecting depression tendency based on deep learning and multi-sources data. Biomedical Signal Processing and Control,86,105226." collected multi-source linked datasets, including social network text data, human gait image data, and human gait keypoint data, and established three deep learning models: a BERT-based social network model, a 3D convolution-based gait image model, and an LSTM-based gait keypoint model, to predict depressive tendencies. The three models were then integrated using a majority voting method, achieving an accuracy of 91.51%. However, while multimodal fusion can improve classification accuracy, it requires the simultaneous collection of data from multiple modalities, and the methods and difficulties of collecting data from different modalities may vary, which limits its practical application.
[0004] While multimodal fusion can improve classification accuracy, it requires the simultaneous collection of data from multiple modalities, and the methods and difficulties of collecting data from different modalities may vary, which limits its practical application. For example, the paper "Yang, J. , Zhang, Z. , Fu, Z. , Li, B. , Xiong, P. , & Liu , X. (2023). Cross-subject classification of depression by using multiparadigm EEG feature fusion. Computer Methods and Programs in Biomedicine, 233, 107360." proposes a multiparadigm fusion method that extracts the Lempel-Ziv complexity feature matrix of EEG signals, fuses the feature matrices extracted from open-eye and closed-eye EEG signals, and uses support vector machines, K-nearest neighbors, and decision tree classifiers to classify depression under open-eye, closed-eye, and fusion paradigms, achieving an accuracy of 94.03%. However, this method still relies on manual feature design and extraction, and the model cannot uncover other useful information after fixing the features, lacking the convenience of automated assisted diagnosis. Therefore, a method is needed that combines the characteristics of multi-state EEG fusion with the advantages of deep learning in automatically extracting deeper, non-linear relationship features, to achieve end-to-end automated assisted identification of monopolar and bipolar depression. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention provides a method for identifying monopolar and bipolar depression based on deep learning-based multi-state EEG fusion.
[0006] The technical solution of this invention to solve the aforementioned technical problem is to provide a method for identifying monopolar and bipolar depression based on multi-state EEG fusion using deep learning, characterized in that the method includes the following steps: Step 1: Perform continuous wavelet transform on the EEG signals in the open-eye and closed-eye states respectively to obtain the open-eye time-frequency diagram and the closed-eye time-frequency diagram respectively. Step 2: Use a deep learning model to extract features from the open-eye and closed-eye time-frequency maps obtained in Step 1, and obtain feature vectors for the open-eye and closed-eye states. Then, fuse the feature vectors for the open-eye and closed-eye states to obtain a multi-state fused feature vector. Finally, use a deep learning classification network to classify and recognize the multi-state fused feature vector to obtain the recognition result and complete the recognition.
[0007] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) This invention provides a novel method for tri-classifying unipolar depression, bipolar disorder, and healthy individuals through multi-state fusion of electroencephalogram (EEG) signals. This method has good classification accuracy, and the acquisition of EEG signals in both open and closed eye states is simple and easy to implement in a clinical setting. In the future, it can also be applied to initial screening in physical examinations. Due to the convenience and low cost of its EEG signal data acquisition, it can be included in routine physical examinations for adolescents, thereby enabling intervention measures to be taken in the early stages of the disease and improving treatment effectiveness.
[0008] (2) This invention proposes a new classification and recognition method based on deep learning. Based on EEG signals, a deep learning model is used to perform three-class classification and recognition of unipolar depression, bipolar disorder and healthy people, thereby improving classification performance.
[0009] (3) The present invention uses continuous wavelet transform to convert one-dimensional EEG signals into two-dimensional time-frequency maps, which helps to more accurately reflect the frequency characteristics of EEG activity. The obtained time-frequency maps can capture details that may be ignored in the original signals.
[0010] (4) This invention fuses the features of resting-state closed-eye EEG signals and resting-state open-eye EEG signals, combining the advantages of the features of the two states to achieve accurate classification of unipolar depression, bipolar disorder, and healthy individuals. At the same time, it reduces the cost and complexity of data collection in practical applications. Attached Figure Description
[0011] Figure 1 This is an overall flowchart of the present invention; Figure 2 This is a flowchart illustrating the architecture and processing flow of the deep learning model of the present invention. Detailed Implementation
[0012] Specific embodiments of the present invention are given below. These specific embodiments are only used to further illustrate the present invention in detail and do not limit the scope of protection of the present invention.
[0013] This invention provides a method for identifying monopolar and bipolar depression based on deep learning-based multi-state EEG fusion (hereinafter referred to as the method), characterized by the following steps: Step 1: Perform continuous wavelet transform on the EEG signals in the open-eye and closed-eye states respectively to obtain the open-eye time-frequency diagram and the closed-eye time-frequency diagram respectively. Preferably, step 1 specifically includes: S11. Acquire resting-state EEG signals in both open-eye and closed-eye states; Preferably, in step S11, the length of the resting-state EEG signal is a five-minute acquisition time during the entire EEG signal acquisition process, in which the impedance of the EEG cap electrodes is kept below 10kΩ and the sampling rate is 1000Hz.
[0014] S12. Preprocess the resting-state EEG signal to obtain the preprocessed EEG signal; Preferably, in step S12, the preprocessing operations include, in sequence, electrode positioning, downsampling, bad segment removal, electrode rereference, independent component analysis, and artifact removal.
[0015] Preferably, in step S12, the preprocessing operation includes the following steps: S121. Electrode localization: Mark the electrode positions of resting-state EEG signals, establish a spatial coordinate system of a standard 10-20 system or an extended electrode system, determine the precise three-dimensional position of each channel electrode on the scalp, and provide a geometric reference system for subsequent spatial analysis and rereference. S122. Downsampling: The original high-sampling-rate EEG signal is reduced to 256Hz through anti-aliasing filtering and decimation algorithm. This reduces data redundancy, improves the computational efficiency of subsequent processing, and standardizes the data format while ensuring the Nyquist theorem. S123. Bad Segment Removal: Based on multiple criteria such as signal amplitude threshold detection, flat line segment identification, and smoothness test, abnormal data segments are automatically screened and removed to remove signal contamination segments caused by poor electrode contact, muscle tension, or other physiological artifacts, ensuring that the data quality meets the analysis requirements. S124. Electrode Re-reference: The signal reference point is reset by means of average reference, bilateral mastoid reference or specific electrode reference to eliminate the measurement deviation caused by the original reference electrode, optimize the spatial distribution characteristics of the signal and improve the accuracy of functional connectivity analysis. S125 Independent Component Analysis: Decomposes multi-channel EEG signals and converts them into a statistically independent set of components. Each component contains a specific temporal process and spatial distribution pattern, providing a mathematical basis for further artifact recognition and removal. S126. Artifact Removal: Based on the results of independent component analysis, combined with the morphological characteristics of component topography, power spectrum distribution patterns and time series characteristics, artifact components caused by eye movements, blinking, electrocardiogram interference and electromyographic activity are manually or automatically identified and removed to reconstruct pure neurogenic EEG signals.
[0016] S13. Apply continuous wavelet transform to the preprocessed EEG signal to decompose the EEG signal at different scales and time positions, thereby realizing time-frequency analysis of the signal and obtaining a time-frequency diagram.
[0017] Preferably, in step S13, the calculation expression for the continuous wavelet transform is as follows: (1) In equation (1), x(t) represents the resting-state EEG signal to be processed; t represents time. (a,b) represents the wavelet transform result of signal x(t) under scale parameter a and translation parameter b; These are wavelet basis functions used for signal analysis, and their complex conjugates are... The scaling parameter 'a' controls the scaling of the wavelet function, affecting the frequency range of the analysis; the translation parameter 'b' controls the translation of the wavelet function on the time axis, determining the temporal position of the image analysis.
[0018] Preferably, in step 1, the continuous wavelet transform adopts the Morlet wavelet; the Morlet wavelet is a symmetrical, periodic wavelet that can reliably detect oscillations, and can detect both amplitude and phase. Substituting the Morlet wavelet function into equation (1), that is... Set to the Morlet wavelet function. The Morlet wavelet function is shown in equation (2): (2) In equation (2), ω is a frequency parameter that controls the frequency of the wavelet basis; i is the imaginary unit; and π is the value of pi.
[0019] Step 2: Use a deep learning model to extract features from the open-eye and closed-eye time-frequency maps obtained in Step 1, and obtain feature vectors for the open-eye and closed-eye states. Then, fuse the feature vectors for the open-eye and closed-eye states to obtain a multi-state fused feature vector. Finally, use a deep learning classification network to classify and recognize the multi-state fused feature vector to obtain the recognition result and complete the recognition.
[0020] Preferably, in step 2, the deep learning model is the ConvNeXt model; the ConvNeXt model is a pure convolutional network architecture, which adopts a four-stage feature extraction architecture, and each stage includes a downsampling module and multiple ConvNeXt Block units.
[0021] Preferably, in step 2, feature extraction includes the following steps: A21. The open-eye time-frequency map and closed-eye time-frequency map obtained in step 1 are pre-processed by block partitioning base layer, and downsampling is achieved by convolution operation with 4×4 convolution kernel and stride of 4. A22. The time-frequency image is processed in each stage using the following structure: (1) Layer Normalization: normalize the time-frequency image; (2) Depthwise Separable Convolution: use 7×7 depthwise convolution kernels for spatial feature extraction; (3) Fully Connected Layer: transform the channel dimension through 1×1 convolution; (4) Layer Scale: perform trainable scaling operation on each channel with a scaling factor of 1e. -6 ; A23. The time-frequency graph undergoes spatial dimension compression between stages using a downsampling module; the downsampling module is a combination of LayerNormalization and a 2×2 convolution kernel, performing convolution operations with a stride of 2; A24. The final feature map is reduced in dimensionality using global average pooling, as shown in equation (3): (3) In equation (3), H is the height of the feature map obtained after step S213 of the time-frequency graph, and W is the width. Let represent the feature values of all channels at position (i,j). This is the output after global average pooling.
[0022] Preferably, in step 2, feature fusion is performed using linear addition; the linear addition method is as follows: the extracted feature vectors for the open-eye state and the closed-eye state are respectively represented as... and , where i represents the i-th feature column vector in the feature matrix; the linear combination of the feature matrices is as follows: , where Ap represents the feature matrix obtained by directly adding the feature matrices in the open and closed eye states.
[0023] Preferably, in step 2, the deep learning classification network is the KAN model; the KAN model uses the Kolmogorov-Arnold representation theorem to construct the neural network architecture and achieves nonlinear feature mapping through learnable spline functions.
[0024] Preferably, in step 2, the classification and identification includes the following steps: B21. Construct a network architecture based on the Kolmogorov-Arnold representation theorem: According to the Kolmogorov-Arnold representation theorem, any multivariable continuous function can be represented as a finite combination of single-variable nonlinear functions; B22. Perform learnable nonlinear edge mapping: During data forward propagation, input features are transformed through learnable activation functions at the network edges; for the first... l Layer, let its input be The output is Then the calculation formula for the i-th output neuron is shown in equation (4): (4) In equation (4), Indicates the first l The value of the i-th neuron in layer +1 (i.e., the output layer of the current layer); Indicates the first l The total number of neurons in the layer (i.e., the input layer of the current layer). Indicates the firstl The value of the j-th input neuron in the layer; Indicates the connection of the first l The learnable univariate function between the j-th input node and the i-th output node of a layer. Unlike traditional networks, there are no fixed linear weights here; each connection is itself a non-linear mapping process.
[0025] B23. Function Fitting Based on B-Splines: To achieve learnable functions To enhance flexibility and fitting ability, a combination of B-spline curves and basis functions is used; the connection function Parameterized as: (5) In equation (5), x represents the input value on the connection; The weight parameters (scaling factors) of the basis functions are represented. The basis function is represented by the SiLU function, which is preferred in this embodiment. The weight parameters (scaling factors) of the spline function are represented. Represents a learnable spline function; spline function The linear combination of B-spline basis functions is shown in equation (6): (6) In equation (6), The coefficient of the k-th control point is the main learnable parameter during network training. This represents the k-th B-spline basis function defined on a preset grid. With this design, the model can dynamically adjust the shape of the spline curve during training, thereby accurately approximating the complex nonlinear relationship between input features and classification labels; B24. Classification Probability Output: After feature transformation and abstraction through multiple KAN layers, the output vector of the last layer is mapped through the classification layer; the probability distribution of each category is calculated using the Softmax function as shown in equation (7): (7) In equation (7), This represents the predicted probability that the input feature X belongs to the k-th category. This represents the logical value corresponding to the k-th class output of the last layer of the KAN network. This represents the normalized term representing the sum of the logical exponents of all categories; B25. Determine the category label to which the feature belongs based on the maximum probability to complete the classification process.
[0026] Any aspects not covered in this invention are applicable to existing technologies.
Claims
1. A method for identifying monopolar and bipolar depression based on deep learning-based multi-state EEG fusion, characterized in that, The method includes the following steps: Step 1: Perform continuous wavelet transform on the EEG signals in the open-eye and closed-eye states respectively to obtain the open-eye time-frequency diagram and the closed-eye time-frequency diagram respectively. Step 2: Use a deep learning model to extract features from the open-eye and closed-eye time-frequency maps obtained in Step 1, and obtain feature vectors for the open-eye and closed-eye states. Then, fuse the feature vectors for the open-eye and closed-eye states to obtain a multi-state fused feature vector. Finally, use a deep learning classification network to classify and recognize the multi-state fused feature vector to obtain the recognition result and complete the recognition.
2. The method for identifying monopolar and bipolar depression based on deep learning-based multi-state EEG fusion according to claim 1, characterized in that, Step 1 is as follows: S11. Acquire resting-state EEG signals in both open-eye and closed-eye states; S12. Preprocess the resting-state EEG signal to obtain the preprocessed EEG signal; S13. Apply continuous wavelet transform to the preprocessed EEG signal to decompose the EEG signal at different scales and time positions, thereby realizing time-frequency analysis of the signal and obtaining a time-frequency diagram.
3. The method for identifying monopolar and bipolar depression based on deep learning-based multi-state EEG fusion according to claim 2, characterized in that, In step S12, the preprocessing operations include electrode localization, downsampling, bad segment removal, electrode rereference, independent component analysis, and artifact removal performed sequentially.
4. The method for identifying monopolar and bipolar depression based on deep learning multi-state EEG fusion according to claim 2 or 3, characterized in that, In step S12, the preprocessing operation includes the following steps: S121. Electrode localization: Mark the electrode positions of resting-state EEG signals, establish a spatial coordinate system of a standard 10-20 system or an extended electrode system, determine the precise three-dimensional position of each channel electrode on the scalp, and provide a geometric reference system. S122. Downsampling: The sampling frequency of the original high-sampling-rate EEG signal is reduced by anti-aliasing filtering and decimation algorithm to reduce data redundancy while ensuring the Nyquist theorem. S123, Bad Segment Removal: Automatically screens abnormal data segments and removes signal contamination segments based on multiple criteria including signal amplitude threshold detection, flat line segment identification, and smoothness verification; S124. Electrode Re-reference: The signal reference point is reset by using an average reference, bilateral mastoid reference, or specific electrode reference to eliminate the measurement deviation caused by the original reference electrode. S125 Independent Component Analysis: Decomposes multi-channel EEG signals and converts them into a statistically independent set of components, each of which contains a specific temporal progression and spatial distribution pattern. S126. Artifact Removal: Based on the results of independent component analysis, combined with the morphological characteristics of component topography, power spectrum distribution patterns and time series characteristics, artifact components are identified and eliminated to reconstruct pure neurogenic EEG signals.
5. The method for identifying monopolar and bipolar depression based on deep learning-based multi-state EEG fusion according to claim 2, characterized in that, In step S13, the calculation expression for the continuous wavelet transform is as follows: (1) In equation (1), x(t) represents the resting-state EEG signal to be processed; t represents time. (a,b) represents the wavelet transform result of signal x(t) under scale parameter a and translation parameter b; These are wavelet basis functions used for signal analysis, and their complex conjugates are... The scaling parameter 'a' controls the scaling of the wavelet function, affecting the frequency range of the analysis; the translation parameter 'b' controls the translation of the wavelet function on the time axis, determining the temporal position of the image analysis.
6. The method for identifying monopolar and bipolar depression based on deep learning-based multi-state EEG fusion according to claim 1, characterized in that, In step 1, the continuous wavelet transform adopts the Morlet wavelet; the Morlet wavelet function is shown in equation (2): (2) In equation (2), ω is a frequency parameter that controls the frequency of the wavelet basis; i is the imaginary unit; and π is the value of pi.
7. The method for identifying monopolar and bipolar depression based on deep learning-based multi-state EEG fusion according to claim 1, characterized in that, In step 2, the deep learning model is the ConvNeXt model; the ConvNeXt model adopts a four-stage feature extraction architecture, each stage containing a downsampling module and multiple ConvNeXt Block units; In step 2, the deep learning classification network is the KAN model; the KAN model uses the Kolmogorov-Arnold representation theorem to construct the neural network architecture and achieves nonlinear feature mapping through learnable spline functions.
8. The method for identifying monopolar and bipolar depression based on deep learning-based multi-state EEG fusion according to claim 1, characterized in that, In step 2, feature extraction includes the following steps: A21. The open-eye time-frequency map and closed-eye time-frequency map obtained in step 1 are pre-processed by block partitioning base layer, and downsampling is achieved by convolution operation with 4×4 convolution kernel and stride of 4. A22. The time-frequency graph is processed in each stage using the following structure: Layer Normalization: Normalizes the time-frequency graph; Depthwise Separable Convolution: Extracts spatial features using 7×7 depthwise convolution kernels; Fully Connected Layer: Transforms the channel dimensions using 1×1 convolutions; Layer Scale: Performs trainable scaling on each channel with a scaling factor of 1e. -6 ; A23. The time-frequency plot undergoes spatial dimension compression between stages using a downsampling module; the downsampling module is a combination of Layer Normalization and a 2×2 convolution kernel, performing convolution operations with a stride of 2; A24. The final feature map is reduced in dimensionality using global average pooling, as shown in equation (3): (3) In equation (3), H is the height of the feature map obtained after step S213 of the time-frequency graph, and W is the width. Let represent the feature values of all channels at position (i,j). This is the output after global average pooling.
9. The method for identifying monopolar and bipolar depression based on deep learning-based multi-state EEG fusion according to claim 1, characterized in that, In step 2, feature fusion is performed using linear addition; the linear addition method is as follows: the extracted feature vectors for the open-eye and closed-eye states are respectively represented as... and , where i represents the i-th feature column vector in the feature matrix; the linear combination of the feature matrices is as follows: A p This represents the feature matrix obtained by directly adding the feature matrices of the open and closed eye states.
10. The method for identifying monopolar and bipolar depression based on deep learning-based multi-state EEG fusion according to claim 1, characterized in that, Step 2, the classification and identification includes the following steps: B21. Construct a network architecture based on the Kolmogorov-Arnold representation theorem: According to the Kolmogorov-Arnold representation theorem, any multivariable continuous function can be represented as a finite combination of single-variable nonlinear functions; B22. Perform learnable nonlinear edge mapping: During data forward propagation, input features are transformed through learnable activation functions at the network edges; for the first... l Layer, let its input be The output is Then the calculation formula for the i-th output neuron is shown in equation (4): (4) In equation (4), Indicates the first l The value of the i-th neuron in layer +1; Indicates the first l The total number of neurons in the layer, Indicates the first l The value of the j-th input neuron in the layer; Indicates the connection of the first l A learnable univariate function between the j-th input node and the ith output node of a layer; B23. Function fitting based on B-splines: To achieve learnable functions To enhance flexibility and fitting ability, a combination of B-spline curves and basis functions is used; the connection function Parameterized as: (5) In equation (5), x represents the input value on the connection; The weight parameters of the basis functions are represented. The basis function is represented by the SiLU function, which is preferred in this embodiment. The weight parameters of the spline function are represented. Represents a learnable spline function; spline function The linear combination of B-spline basis functions is shown in equation (6): (6) In equation (6), The coefficient of the k-th control point is the main learnable parameter during network training. This represents the k-th B-spline basis function defined on a preset grid. With this design, the model can dynamically adjust the shape of the spline curve during training, thereby accurately approximating the complex nonlinear relationship between input features and classification labels; B24. Classification Probability Output: After feature transformation and abstraction through multiple KAN layers, the output vector of the last layer is mapped through the classification layer; the probability distribution of each category is calculated using the Softmax function as shown in equation (7): (7) In equation (7), This represents the predicted probability that the input feature X belongs to the k-th category. This represents the logical value corresponding to the k-th class output of the last layer of the KAN network. This represents the normalized term representing the sum of the logical exponents of all categories; B25. Determine the category label to which the feature belongs based on the maximum probability to complete the classification process.