Depression state prediction method and system based on prefrontal lobe electroencephalogram feature fusion selection
By using a hybrid feature selection model combining RankSearch and Genetic Algorithm, which integrates time-domain, frequency-domain, and nonlinear features, the problem of low efficiency in EEG signal recognition in existing technologies is solved, and efficient identification of depressive states is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANJING UNIV OF POSTS & TELECOMM
- Filing Date
- 2026-03-11
- Publication Date
- 2026-06-12
AI Technical Summary
Existing technologies for identifying depressive states based on EEG signals suffer from insufficient feature utilization, low model efficiency, and redundancy and noise during high-dimensional feature fusion, making it difficult to effectively select the most discriminative features.
A two-layer hybrid feature selection model, including RankSearch and Genetic Algorithm, is adopted, which combines time domain, frequency domain and nonlinear features. Through preprocessing and multi-feature fusion, the XGBoost classifier is used to identify the depressive state.
While achieving efficient reduction of computational complexity, it improves the accuracy and recognition efficiency of feature fusion and enhances the ability to identify depressive states.
Smart Images

Figure CN122182033A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electroencephalogram (EEG) signal processing technology, specifically to a method and system for predicting depressive states based on the fusion and selection of prefrontal EEG features. Background Technology
[0002] With the development of brain science and signal processing technology, electroencephalogram (EEG) signals, due to their high temporal resolution and ability to directly reflect the brain's neural electrical activity, have become an important approach for studying the neural mechanisms of depression and achieving objective identification. Studies have shown that the EEG activity of patients with depression exhibits specific changes in the frequency domain, time domain, and nonlinear dynamic characteristics in brain regions such as the prefrontal cortex, providing a theoretical basis for objective assessment based on EEG.
[0003] Research on signal feature recognition based on EEG signals is mainly limited by insufficient feature utilization and low model efficiency. Some studies, in order to control model complexity, use only single-type features such as time-domain or frequency-domain features, supplemented by traditional filtering or encapsulation-based feature selection methods. At the feature engineering level, traditional methods heavily rely on manually designed features and domain knowledge, such as geometric features based on phase space reconstruction or features based on functional connectivity of specific brain regions; their generalization ability is limited by prior assumptions. To improve feature representation capabilities, researchers have turned to fusing multi-dimensional (time-domain, frequency-domain, nonlinear) features, but this directly leads to a dramatic expansion of feature dimensions, reaching tens or even higher, introducing a large amount of redundancy and noise. At the feature dimensionality reduction and selection level, existing methods generally suffer from the problem of difficulty in balancing efficiency and effectiveness. Traditional filtering methods, such as channel selection based on kernel target alignment, are computationally efficient but cannot evaluate complex interactions between features; while encapsulation-based methods, especially metaheuristic algorithms, although capable of global search, are prone to getting trapped in local optima in ultra-high-dimensional spaces, have high computational costs, and their randomly initialized populations contain a large number of invalid features, severely limiting search efficiency. While recent studies have attempted to combine various optimization strategies, such as the ALO-MARL algorithm, they are essentially still single-stage search frameworks and have failed to effectively address the core issues of severe noise interference and poor quality of the search starting point in high-dimensional initial spaces. Therefore, effectively fusing and selecting the most discriminative multi-dimensional features from prefrontal EEG signals to construct a recognition model that is both accurate and efficient has become the current mainstream requirement. Summary of the Invention
[0004] The purpose of this invention is to provide a method and system for predicting depressive states based on the fusion of prefrontal EEG features, so as to solve the problems mentioned in the background art.
[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0006] A method for predicting depressive states based on the fusion of prefrontal EEG features, characterized by the following steps:
[0007] S100. Acquire EEG signals from multiple users to be tested, wherein the EEG signals include resting-state closed-eye EEG signals, resting-state open-eye EEG signals, and evoked-state EEG signals.
[0008] S200: Screen out the prefrontal cortex EEG signals and preprocess the prefrontal cortex EEG signals;
[0009] S300. Calculate and extract the time-domain features, frequency-domain features, and nonlinear and complex features of the prefrontal cortex EEG signal, and then fuse the extracted features to construct a multi-feature fusion vector.
[0010] S400. A two-layer hybrid feature selection model is used to perform feature selection and dimensionality reduction on multi-feature fusion vectors; wherein, the two-layer hybrid feature selection model includes the RankSearch model and the Genetic Algorithm model;
[0011] S500: Import the multi-domain feature fusion vector after feature selection into the XGBoost classifier and output the result of the depression status of the user to be detected.
[0012] Preferably, step S200 specifically includes:
[0013] S201. Electrode localization was performed using the EEGLAB toolbox. Three prefrontal electrode channels, Fp1, Fz, and Fp2, were selected. 50Hz common frequency interference was removed by filtering, and high-frequency signal interference was removed by bandpass filtering from 0.1 to 45Hz. Bad segments were removed, and low-frequency interference was removed from the original signal by independent component analysis (ICA). After ICA was performed, the artifacts of electrooculography (EOG), electromyography (EMG), and electrocardiography (ECG) were identified and removed based on the energy distribution and power spectrum analysis of the topographic map.
[0014] S202. Perform wavelet threshold denoising on the processed signal; the wavelet threshold denoising includes wavelet decomposition, threshold processing, and wavelet reconstruction; select... Wavelet denoising is performed using a fixed threshold calculation formula: ,in The signal length is represented; then sample generation is performed, and a 4-second sliding window is used to segment the data with an overlap rate of 50%, resulting in a sample set.
[0015] Preferably, the time-domain features and frequency-domain features in step S300 specifically include:
[0016] The temporal domain characteristics of prefrontal cortex EEG signals include:
[0017] Standard deviation That is, the arithmetic square root of the variance: ;in, express point signal sequence, Represents the mean of a signal sequence;
[0018] Peak-to-peak value This refers to the difference between the maximum and minimum values of a signal within a given time window. ;in, and These represent the maximum and minimum values of the sequence within the analysis window, respectively.
[0019] Root Mean Square , is the square root of the arithmetic mean of the squares of the values at each point in the signal sequence: ;
[0020] Hjorth activity parameter , refers to the variance of the EEG signal, i.e., the signal power, and represents the width: ;in, The standard deviation of the signal;
[0021] Hjorth mobility parameter Used to estimate the average frequency of EEG signals: ;in, The standard deviation of the first derivative of a signal;
[0022] Hjorth complexity parameter This refers to the mobility of the first derivative relative to the EEG itself. ;in, It represents the standard deviation of the second derivative of the signal.
[0023] Frequency domain characteristics of prefrontal cortex EEG signals include:
[0024] Waves and Wave, Waves and Wave, Waves and Wave, Waves and Wave power ratio:
[0025] ; ; ; ;
[0026] in, Indicates frequency range power, express Waves and The ratio of band power, express Waves and The ratio of band power, express Waves and The ratio of band power, express Waves and The ratio of band power;
[0027] Wave, Wave, Wave, Prefrontal asymmetry of waves:
[0028] ; ; ; ;
[0029] in, and These represent the left and right hemispherical symmetrical electrodes in the frequency range, respectively. power, express Wave asymmetry, express Wave asymmetry, express Wave asymmetry, express Wave asymmetry;
[0030] Wave, Wave, Wave, Spectral entropy of a wave:
[0031] ;in, Indicates the signal at frequency The power spectral density at that location.
[0032] Preferably, the nonlinear and complex features in step S300 include:
[0033] Differential entropy : ;in, The standard deviation of the signal;
[0034] Sample Entropy It is used to describe the complexity of time series and the probability that a time series will generate new patterns when its dimension changes;
[0035] Permutation Entropy A nonlinear method for detecting time series complexity or dynamic abrupt changes, capable of quantitatively assessing random noise contained in a signal;
[0036] Lempel-Ziv complexity : ;in, For complexity, The number of loops. To normalize the Lempel-Ziv complexity;
[0037] Higuchi fractal dimension: The Higuchi algorithm is used to calculate the fractal dimension of the original time series. Sub-segments and calculate the average curve length of the sub-segments. ,according to and A linear relationship between them is fitted to obtain a straight line, the slope of which is the Higuchi fractal dimension. .
[0038] Preferably, the sample entropy The acquisition process includes:
[0039] Group the dimensions according to the sequence number. vector sequence ,in , ;
[0040] definition and The distance between them is ,but , ;
[0041] Given a threshold ,in ,statistics and The distance between them is less than or equal to The quantity, denoted as ;for ,definition Seek its effect on all The average value is ;
[0042] Increase the dimension to ,statistics and The distance between them is less than or equal to The quantity, denoted as ;for ,definition Seek its effect on all The average value is ;
[0043] in, and This indicates that the two sequences are matched under the similarity tolerance. points and The probability of a point; then for a finite length The sequence has the following sample entropy estimate: ;in, Indicates the dimension of the pattern. Indicates similarity tolerance, take ,in is the standard deviation of the sequence.
[0044] Preferably, the permutation entropy The acquisition process includes:
[0045] For a length of time series Given the embedding dimension With delay time The original sequence is reconstructed in phase space to obtain the reconstruction matrix. :
[0046] , Each row of the matrix is called a reconstructed component;
[0047] For each reconstructed component Sort the elements in ascending order of their numerical values, and record the index of each element in the original sequence after sorting, forming a sequence of symbols: ,in , The maximum number of different arrangements of elements is [number]. kind;
[0048] Statistical analysis of each symbol sequence Calculate the relative frequency of occurrence based on the frequency of occurrence: , For sequence The number of times it appears in the reconstructed matrix; calculate the permutation entropy according to the Shannon entropy formula: ;
[0049] The maximum value of the permutation entropy is Then, the permutation entropy is normalized: .
[0050] Preferably, step S400 specifically includes:
[0051] S401. Based on the RankSearch model, mutual information is used as a measure of feature importance. The statistical dependency between each feature and the classification label is evaluated. All features are sorted from high to low according to their mutual information scores, and a subset of candidate features is selected from the top-ranked features.
[0052] S402. Based on the Genetic Algorithm model, a genetic algorithm is introduced to perform heuristic search in the feature space, capture the high-order interaction relationships and synergistic effects between features, and perform combinatorial optimization of the feature subset structure.
[0053] Preferably, step S401 specifically includes:
[0054] If we obtain each feature and classification label, then for feature X and label Y, their mutual information is... for: ;
[0055] in, For joint probability distribution, and It is distributed at the edge;
[0056] After calculating the mutual information scores of all features, they are sorted from high to low. A threshold is set according to the cumulative contribution curve of the features, and only the top features whose cumulative contribution rate is required to reach the threshold are retained to form a subset of candidate features.
[0057] A depression state prediction system based on prefrontal EEG feature fusion includes:
[0058] The EEG signal acquisition module is used to acquire EEG signals from multiple users to be tested.
[0059] The preprocessing module is used to filter out and preprocess the prefrontal EEG signals.
[0060] The feature extraction module is used to calculate and extract the time-domain features, frequency-domain features, and nonlinear and complex features of the prefrontal cortex EEG signal, respectively; the nonlinear and complex features include differential entropy, sample entropy, permutation entropy, Lempel-Ziv complexity, and Higuchi fractal dimension.
[0061] The multi-feature fusion module is used to fuse time-domain features, frequency-domain features, and nonlinear and complex features to construct a multi-feature fusion vector;
[0062] The feature selection module is used to perform feature selection and dimensionality reduction on multi-feature fusion vectors using a two-layer hybrid feature selection model of RankSearch model and Genetic Algorithm model.
[0063] The state decision module is used to import the multi-domain feature fusion vector after feature selection into the XGBoost classifier and output the depression state discrimination result of the user to be detected.
[0064] Compared with the prior art, the beneficial effects achieved by the present invention are:
[0065] This invention improves upon traditional single-feature recognition methods by preprocessing EEG signals to calculate time-domain, frequency-domain, and complexity features, and then combining the extracted features into a multi-domain feature fusion vector for classification. Furthermore, it proposes a "RankSearch-Genetic Algorithm" hybrid selection model to achieve complementary advantages and disadvantages. The resulting multi-domain feature fusion vector significantly reduces computational complexity while more efficiently fusing features, making it more beneficial for medical personnel to identify depressive states based on feature data. Attached Figure Description
[0066] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0067] Figure 1 This is a flowchart of the depressive state prediction method based on the fusion of prefrontal EEG features of the present invention;
[0068] Figure 2 This is the international 10-20 EEG acquisition electrode diagram of the EEG acquisition device of this invention;
[0069] Figure 3 This is a schematic diagram of the EEG signal preprocessing process of the present invention;
[0070] Figure 4 This is a schematic diagram showing the results before and after using wavelet denoising in this invention;
[0071] Figure 5 This is a schematic diagram illustrating the generation of sample data according to the present invention;
[0072] Figure 6 This is a schematic diagram of the frequency bands obtained using a bandpass filter in this invention;
[0073] Figure 7 This is a schematic diagram of the multi-domain fusion feature vector combination constructed by the present invention;
[0074] Figure 8 This is the evolutionary trend of feature subset size after using the hybrid feature selection algorithm in this invention;
[0075] Figure 9 This paper compares the classification performance of the feature selection method proposed in this invention with other classic feature selection methods. Detailed Implementation
[0076] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0077] Please see Figures 1-9 The present invention provides the following technical solution:
[0078] Example 1: A method for predicting depressive states based on the fusion of prefrontal EEG features, the flowchart of which is as follows: Figure 1 As shown, the EEG signal is first preprocessed, and then time-domain features, frequency-domain features, and complexity features are calculated separately. The extracted features are combined into a multi-domain feature fusion vector. A "RankSearch-Genetic Algorithm" hybrid selection model is proposed to achieve complementary advantages and disadvantages. The resulting multi-domain feature fusion vector significantly reduces computational complexity while more efficiently identifying depressive states. Specifically, the steps are explained below:
[0079] Step 1. Prepare a dataset of EEG signals for depression;
[0080] The EEG dataset for depression prepared in this embodiment is selected from the HUSM dataset from Universiti Sains Malaysia Kelantan Hospital. This dataset provides data from 34 MDD patients (17 women and 17 men, mean age: The MDD patients met the clinical diagnostic criteria for major depressive disorder, namely the Diagnostic and Statistical Manual of Mental Disorders, Fourth Edition (DSM-4), and 30 healthy individuals (9 women and 21 men, mean age: The physiological signals of the subjects included two different experiments: resting-state EEG acquisition and visually evoked task-based EEG acquisition. The resting-state task in the experiment included 5 minutes of eyes closed (EC) and 5 minutes of eyes open (EO). The data samples of each subject included 19 channels of resting-state EEG signals with eyes closed, resting-state EEG signals with eyes open, and evoked EEG signals in a 10-20 system.
[0081] Step 2. Preprocess the raw EEG signals in the dataset;
[0082] This embodiment combines EEG signals with depressive mood recognition. The EEG signals of the HUSM depression dataset are preprocessed. Since the frontal lobe is more sensitive to changes in emotional EEG signals, the EEG signals of the three electrode channels Fp1, Fz, and Fp2 in the frontal lobe are selected.
[0083] Electrode channel distribution diagram of EEG acquisition device as shown below Figure 2 As shown, this invention selects EEG channels based on the distribution map. The 19-lead electrode distribution conforms to the international 10-20 system standard electrode placement method, specifically including frontal electrodes Fp1 and Fp2, frontal lobe electrodes F3, F4, F7, F8, and Fz, central area electrodes C3, C4, and CZ, parietal lobe electrodes P3, P4, and PZ, occipital area electrodes O1 and O2, temporal lobe electrodes T3, T4, T5, and T6, and reference electrodes A1 and A2.
[0084] The flowchart of EEG signal preprocessing in this embodiment is shown below. Figure 3 As shown. Preprocessing is based on the EEGLAB toolbox and Matlab software, and consists of three parts: EEGLAB toolbox processing, wavelet thresholding denoising, and sample set generation.
[0085] Step 2.1: Use the EEGLAB toolbox for electrode localization. Select three prefrontal electrode channels: Fp1, Fz, and Fp2. Remove 50Hz common frequency interference through a filter and remove high-frequency signal interference through a bandpass filter of 0.1-45Hz. Remove bad segments and then remove low-frequency interference from the original signal through independent component analysis (ICA). After performing ICA, identify and remove artifacts of electrooculography (EOG), electromyography (EMG), and electrocardiography (ECG) based on the energy distribution and power spectrum analysis of the topographic map.
[0086] Step 2.2, as follows Figure 4 As shown, wavelet thresholding denoising is performed on the signal processed in step 2.1. Wavelet thresholding denoising is a widely used signal processing method to remove noise components from signals. It has low entropy and multi-resolution characteristics and includes wavelet decomposition, thresholding, and wavelet reconstruction. In this embodiment, wavelet thresholding is selected. Wavelet denoising is performed using a fixed threshold calculation formula. ,in The signal length;
[0087] Step 2.3, Sample generation, refers to segmenting the signal obtained after wavelet threshold denoising to obtain a sample set. Figure 5 Schematic diagram of sample set generation. Samples are generated from the processed signals. A time window with a duration of 4 seconds and a step size of 2 seconds is selected. That is, the EEG signal is segmented into segments with a 4-second sliding window and an overlap rate of 50%, resulting in a sample set with a single sample length of 4 seconds.
[0088] Step 3. Calculate the time-domain characteristics, frequency-domain characteristics, and a series of nonlinear and complex characteristics of the prefrontal cortex EEG signal, such as... Figure 6 The image shows waveforms for the Delta, Theta, Alpha, Beta, and Gamma frequency bands obtained using a bandpass filter. Each feature contains parameters from multiple electrodes, and the extracted feature parameters are fused to construct a multi-feature fusion vector.
[0089] Figure 7 A schematic diagram of the combination of multi-domain feature fusion vectors. To effectively identify depressive states, this invention fuses extracted features to construct a multi-domain feature fusion vector, resulting in an 81-dimensional (3 channels * 27 feature parameters) feature vector F. In the diagram, F1 to F27 represent the following feature parameters, respectively: standard deviation (Std), peak-to-peak value (PP), root mean square (RMS), Hjorth activity parameter (HA), Hjorth mobility parameter (HM), Hjorth complexity parameter (HC), and power (…). , , , ), power ratio ( , , , ), frontal wave asymmetry ( , , , ), spectral entropy ( , , , The features include differential entropy (DE), sample entropy (SampEn), permutation entropy (Pec), Lempel-Ziv complexity (LZC), and Higuchi fractal dimension (HFD), all of which are feature vectors containing 3 electrode channels.
[0090] Step 4. Feature selection and dimensionality reduction: Use a feature selection model to perform feature selection on the multi-domain feature vectors obtained in Step 3;
[0091] To reduce redundancy in high-dimensional, multi-domain feature fusion vectors, a two-layer hybrid feature selection model, "RankSearch - Genetic Algorithm," is proposed to achieve complementary advantages between different feature selection methods and filter out the feature subset that best performs for the classifier. The principles of the RankSearch and Genetic Algorithm algorithms are as follows:
[0092] Step 4.1, RankSearch Algorithm Principle:
[0093] Mutual information is used as a measure of feature importance to evaluate the statistical dependency between each feature and the classification label. Mutual information quantifies the amount of information about the class label contained in a feature. Its advantage lies in its ability to capture both linear and nonlinear relationships. For feature X and label Y, their mutual information... Defined as: ,in, For joint probability distribution, and For marginal distributions, after calculating the mutual information scores of all features, they are sorted from highest to lowest score. Then, a threshold is set based on the cumulative contribution curve of the features, and only the top features whose cumulative contribution rate needs to reach the threshold are retained, forming a subset of candidate features.
[0094] Step 4.2, Principle of the Genetic Algorithm:
[0095] In this invention, the genetic algorithm is used as the optimizer for the second-level feature selection. Its parameter settings follow the general design principles of genetic algorithms and refer to common practical experience in the field of machine learning feature selection.
[0096] 1) Chromosome encoding: Binary encoding is used, which is the most direct and universal representation method for feature selection problems. Each gene position corresponds to the selection status of a candidate feature (1 selected, 0 not selected), and a chromosome represents a complete subset of features;
[0097] 2) Initializing the Population: Randomly generate 50 initial feature subsets (chromosomes) to form the first-generation population, serving as the starting point for the search. The population size affects the algorithm's global search capability. Too small a population may lead to insufficient diversity and getting stuck in local optima; too large a population will increase unnecessary computational overhead. 50 is a commonly used empirical value in small to medium-sized feature selection problems, ensuring sufficient search diversity while maintaining high iteration efficiency.
[0098] 3) Fitness evaluation: The fitness function is set to the accuracy of 5-fold cross-validation based on random forest. The quality of each feature subset is directly measured by its classification accuracy on the random forest model. The higher the accuracy, the higher the fitness.
[0099] 4) Selection operation: Based on fitness scores, superior individuals are selected as parents using mechanisms such as roulette or tournaments to produce the next generation;
[0100] 5) Elite preservation: The individual with the highest fitness in each generation is directly copied to the next generation, ensuring that the algorithm does not lose the current optimal solution and accelerating convergence;
[0101] 6) Crossover operation: Chromosomal segments are exchanged between selected parent individuals with an 80% probability to generate new offspring, promoting the fusion of superior trait patterns;
[0102] 7) Mutation operation: Randomly flip a certain gene locus in the offspring with a 5% probability (randomly adding or removing features), introduce diversity, and help escape local optima;
[0103] 8) Iteration Termination: Repeat steps 3)-7) for 20 generations. Dynamic random seeds ensure the reproducibility of the experiment and avoid single-shot random bias.
[0104] Figure 8 To further analyze the evolutionary trend of feature subset size after initial screening using RankSearch, the optimal feature subset size gradually converged from approximately 40 features to around 30 features with increasing iterations. Simultaneously, the average subset size of the population also showed a synchronous decreasing trend, indicating that the algorithm effectively identifies and eliminates redundant features, achieving a simplified and optimized feature space. This feature selection process not only reduces feature dimensionality and model complexity but also improves the model's generalization ability by removing irrelevant and redundant features. The simplified feature set is more conducive to revealing the intrinsic structure of the data, enhancing the model's interpretability, and significantly improving computational efficiency, laying a solid foundation for subsequent machine learning modeling.
[0105] Step 5. Import the multi-domain feature fusion vector after feature selection into the XGBoost classifier and output the result of the depression status of the user to be detected;
[0106] Figure 9 To compare the classification performance of different feature selection methods and further demonstrate the effectiveness of our proposed method, we conducted a comprehensive comparative analysis with three classic feature selection methods using the same dataset, such as... Figure 9 As shown, this multi-layered comparison can more comprehensively evaluate the performance of different methods under different conditions. Experimental results show that although the genetic algorithm increases computational cost, RankSearch's pre-screening step reduces the feature space by 62%, and its overall performance is still higher than metaheuristic algorithms such as ACO.
[0107] This invention achieves efficient feature classification and recognition by providing accurate local optimal features, offering multiple possible directions for future EEG medical research.
[0108] Example 2: A depressive state prediction system based on prefrontal EEG feature fusion, comprising:
[0109] The EEG signal acquisition module is used to acquire EEG signals from multiple users to be tested.
[0110] The preprocessing module is used to filter out and preprocess the prefrontal EEG signals.
[0111] The feature extraction module is used to calculate and extract the time-domain features, frequency-domain features, and nonlinear and complex features of the prefrontal cortex EEG signal, respectively; the nonlinear and complex features include differential entropy, sample entropy, permutation entropy, Lempel-Ziv complexity, and Higuchi fractal dimension.
[0112] The multi-feature fusion module is used to fuse time-domain features, frequency-domain features, and nonlinear and complex features to construct a multi-feature fusion vector;
[0113] The feature selection module is used to perform feature selection and dimensionality reduction on multi-feature fusion vectors using a two-layer hybrid feature selection model of RankSearch model and Genetic Algorithm model.
[0114] The state decision module is used to import the multi-domain feature fusion vector after feature selection into the XGBoost classifier and output the depression state discrimination result of the user to be detected.
[0115] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for predicting depressive states based on the fusion of prefrontal EEG features, characterized in that, Includes the following steps: S100. Acquire EEG signals from multiple users to be tested, wherein the EEG signals include resting-state closed-eye EEG signals, resting-state open-eye EEG signals, and evoked-state EEG signals. S200: Screen out the prefrontal cortex EEG signals and preprocess the prefrontal cortex EEG signals; S300. Calculate and extract the time-domain features, frequency-domain features, and nonlinear and complex features of the prefrontal cortex EEG signal, and then fuse the extracted features to construct a multi-feature fusion vector. S400. A two-layer hybrid feature selection model is used to perform feature selection and dimensionality reduction on multi-feature fusion vectors; wherein, the two-layer hybrid feature selection model includes the RankSearch model and the Genetic Algorithm model; S500: Import the multi-domain feature fusion vector after feature selection into the XGBoost classifier and output the result of the depression status of the user to be detected.
2. The method for predicting depressive states based on the fusion of prefrontal EEG features as described in claim 1, characterized in that, Step S200 specifically includes: S201. Electrode localization was performed using the EEGLAB toolbox. Three prefrontal electrode channels, Fp1, Fz, and Fp2, were selected. 50Hz common frequency interference was removed by filtering, and high-frequency signal interference was removed by bandpass filtering from 0.1 to 45Hz. Bad segments were removed, and low-frequency interference was removed from the original signal by independent component analysis (ICA). After ICA was performed, the artifacts of electrooculography (EOG), electromyography (EMG), and electrocardiography (ECG) were identified and removed based on the energy distribution and power spectrum analysis of the topographic map. S202. Perform wavelet threshold denoising on the processed signal; the wavelet threshold denoising includes wavelet decomposition, threshold processing, and wavelet reconstruction; select... Wavelet denoising is performed using a fixed threshold calculation formula: ,in The signal length is represented; then sample generation is performed, and a 4-second sliding window is used to segment the data with an overlap rate of 50%, resulting in a sample set.
3. The method for predicting depressive states based on the fusion of prefrontal EEG features as described in claim 1, characterized in that, The time-domain and frequency-domain features in step S300 specifically include: Temporal characteristics of prefrontal EEG signals include: standard deviation, peak-to-peak value, root mean square, Hjorth activity parameter, Hjorth mobility parameter, and Hjorth complexity parameter. Frequency domain characteristics of prefrontal cortex EEG signals include: 1) Waves and Wave, Waves and Wave, Waves and Wave, Waves and Wave power ratio: ; ; ; ; in, Indicates frequency range power, express Waves and The ratio of band power, express Waves and The ratio of band power, express Waves and The ratio of band power, express Waves and The ratio of band power; 2) Wave, Wave, Wave, Prefrontal wave asymmetry: ; ; ; ; in, and These represent the left and right hemispherical symmetrical electrodes in the frequency range, respectively. power, express Wave asymmetry, express Wave asymmetry, express Wave asymmetry, express Wave asymmetry; 3) Wave, Wave, Wave, Spectral entropy of a wave: ;in, Indicates the signal at frequency The power spectral density at that location.
4. The method for predicting depressive states based on the fusion of prefrontal EEG features as described in claim 1, characterized in that, The nonlinear and complex features in step S300 include: differential entropy, sample entropy, permutation entropy, Lempel-Ziv complexity, and Higuchi fractal dimension; Differential entropy : ;in, The standard deviation of the signal; Sample Entropy It is used to describe the complexity of time series and the probability that a time series will generate new patterns when its dimension changes; Permutation Entropy A nonlinear method for detecting time series complexity or dynamic abrupt changes, capable of quantitatively assessing random noise contained in a signal; Lempel-Ziv complexity : ;in, For complexity, The number of loops. To normalize the Lempel-Ziv complexity; Higuchi fractal dimension: The Higuchi algorithm is used to calculate the fractal dimension of the original time series. Sub-segments and calculate the average curve length of the sub-segments. ,according to and A linear relationship between them is fitted to obtain a straight line, the slope of which is the Higuchi fractal dimension. .
5. The method for predicting depressive states based on the fusion of prefrontal EEG features as described in claim 4, characterized in that, The sample entropy The acquisition process includes: Group the dimensions according to the sequence number. vector sequence ,in , ; definition and The distance between them is ,but , ; Given a threshold ,in ,statistics and The distance between them is less than or equal to The quantity, denoted as ;for ,definition Seek its effect on all The average value is ; Increase the dimension to ,statistics and The distance between them is less than or equal to The quantity, denoted as ;for ,definition Seek its effect on all The average value is ; in, and This indicates that the two sequences are matched under the similarity tolerance. points and The probability of a point; then for a finite length The sequence has the following sample entropy estimate: ;in, Indicates the dimension of the pattern. Indicates similarity tolerance, take ,in is the standard deviation of the sequence.
6. The method for predicting depressive states based on the fusion of prefrontal EEG features as described in claim 4, characterized in that, The permutation entropy The acquisition process includes: For a length of time series Given the embedding dimension With delay time The original sequence is reconstructed in phase space to obtain the reconstruction matrix. : , Each row of the matrix is called a reconstructed component; For each reconstructed component Sort the elements in ascending order of their numerical values, and record the index of each element in the original sequence after sorting, forming a sequence of symbols: ,in , The maximum number of different arrangements of elements is [number]. kind; Statistical analysis of each symbol sequence Calculate the relative frequency of occurrence based on the frequency of occurrence: , For sequence The number of times it appears in the reconstructed matrix; calculate the permutation entropy according to the Shannon entropy formula: ; The maximum value of the permutation entropy is Then, the permutation entropy is normalized: .
7. The method for predicting depressive states based on the fusion of prefrontal EEG features as described in claim 1, characterized in that, Step S400 specifically includes: S401. Based on the RankSearch model, mutual information is used as a measure of feature importance. The statistical dependency between each feature and the classification label is evaluated. All features are sorted from high to low according to their mutual information scores, and a subset of candidate features is selected from the top-ranked features. S402. Based on the Genetic Algorithm model, a genetic algorithm is introduced to perform heuristic search in the feature space, capture the high-order interaction relationships and synergistic effects between features, and perform combinatorial optimization of the feature subset structure.
8. The method for predicting depressive states based on the fusion of prefrontal EEG features as described in claim 7, characterized in that, Step S401 specifically includes: If we obtain each feature and classification label, then for feature X and label Y, their mutual information is... for: ; in, For joint probability distribution, and It is distributed at the edge; After calculating the mutual information scores of all features, they are sorted from high to low. A threshold is set according to the cumulative contribution curve of the features, and only the top features whose cumulative contribution rate is required to reach the threshold are retained to form a subset of candidate features.
9. A depression state prediction system for implementing the depression state prediction method based on prefrontal EEG feature fusion selection as described in any one of claims 1-8, characterized in that, include: The EEG signal acquisition module is used to acquire EEG signals from multiple users to be tested. The preprocessing module is used to filter out and preprocess the prefrontal EEG signals. The feature extraction module is used to calculate and extract the time-domain features, frequency-domain features, and nonlinear and complex features of the prefrontal cortex EEG signal, respectively; the nonlinear and complex features include differential entropy, sample entropy, permutation entropy, Lempel-Ziv complexity, and Higuchi fractal dimension. The multi-feature fusion module is used to fuse time-domain features, frequency-domain features, and nonlinear and complex features to construct a multi-feature fusion vector; The feature selection module is used to perform feature selection and dimensionality reduction on multi-feature fusion vectors using a two-layer hybrid feature selection model of RankSearch model and Genetic Algorithm model. The state decision module is used to import the multi-domain feature fusion vector after feature selection into the XGBoost classifier and output the depression state discrimination result of the user to be detected.