Adaptive frequency domain enhanced electroencephalogram classification method and system
By using multi-scale frequency domain feature extraction and attention fusion processing, the problem of insufficient frequency domain feature extraction in traditional EEG classification technology is solved, achieving higher EEG signal classification accuracy and robustness, and adapting to different individual and task requirements.
Patent Information
- Application Number
- CN202511709281.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-20
- Publication Date
- 2026-02-06
AI Technical Summary
Existing EEG classification technologies face bottlenecks in the frequency domain feature extraction stage. Traditional methods cannot effectively capture the dynamic changes and stable low-frequency components of EEG signals, resulting in low classification accuracy and hindering the practical application of brain-computer interface systems.
An adaptive frequency domain enhancement EEG classification method is adopted. Through a multi-scale collaborative frequency domain feature extraction mechanism, including multi-scale short-time Fourier transform, frequency mask generation and attention fusion processing, frequency domain features at different scales are obtained and deep supervision is performed to improve classification accuracy.
It significantly improves the accuracy and robustness of EEG signal classification, can adapt to different individual and task requirements, enhances the interpretability of frequency domain features, and improves the reliability of classification results.
Smart Images

Figure CN121465610A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of bioelectric signal processing, and particularly relates to an electroencephalogram classification method and system with adaptive frequency domain enhancement. BACKGROUND
[0002] Electroencephalogram signals are weak electrophysiological signals generated by brain nerve activity, and contain rich neural function information. Classification technology based on electroencephalogram signals is the core of a brain-computer interface system. Among them, the steady-state visual evoked potential (SSVEP) as a common electroencephalogram signal type is widely used in the fields of disabled person auxiliary control and intelligent interaction due to high information transmission rate and convenient acquisition. However, the existing electroencephalogram classification technology has a core bottleneck in the frequency domain feature extraction link. Traditional methods generally use single-scale short-time Fourier transform (STFT) to extract frequency domain features, while electroencephalogram signals have the essential characteristics of “non-stationary and time-varying”. The SSVEP instantaneous response under a short time scale is easily smoothed, and the stable low-frequency component under a long time scale is difficult to accurately capture the feature difference due to insufficient frequency resolution, ultimately resulting in insufficient representation of frequency domain features and low feature discrimination. This problem directly leads to the difficulty of the subsequent classification network in accurately identifying different types of SSVEP signals, and further restricts the accuracy improvement of the electroencephalogram classification technology, becoming a key obstacle to the practicalization and landing of the brain-computer interface system. SUMMARY
[0003] To solve the above technical problems, the application provides an electroencephalogram classification method and system with adaptive frequency domain enhancement. The application realizes comprehensive capture of the time-frequency characteristics of electroencephalogram signals by innovatively designing a multi-scale collaborative frequency domain feature extraction mechanism, breaks through the limitation of feature representation of traditional methods, and improves the classification accuracy of SSVEP signals.
[0004] To achieve the above purpose, the application provides an electroencephalogram classification method with adaptive frequency domain enhancement, comprising: Collecting electroencephalogram waves induced by a user's gaze on a preset target in a specific scenario; Pretreating the electroencephalogram waves to obtain standardized electroencephalogram signals; Extracting multi-scale frequency domain features from the standardized electroencephalogram signals to obtain frequency domain features of different scales; Performing projection processing on the frequency domain features of different scales respectively; Stacking the frequency domain features after the projection processing into a feature sequence, and performing attention fusion processing on the feature sequence to obtain the fusion features; Inputting the fusion features into a classification model to obtain a classification result, wherein the classification model is a multi-layer fully connected network.
[0005] Optionally, the pretreatment of the electroencephalogram waves to obtain standardized electroencephalogram signals comprises: The average reference method was used to process the acquired EEG waves to eliminate common-mode noise; The brainwaves after common-mode noise elimination are filtered using a Butterworth filter to retain signals in a preset frequency band. Notch filtering is applied to the Butterworth filtered signal to remove preset power frequency interference. The signals after removing power frequency interference are normalized according to channel to obtain standardized EEG signals.
[0006] Optionally, multi-scale frequency domain feature extraction can be performed on the standardized EEG signal to obtain discriminative frequency domain features, including: Short-time Fourier transforms of standardized EEG signals were performed at short, medium, and long scales to obtain the amplitude spectra corresponding to each scale. K target frequency bands are preset, and a frequency mask is generated based on each target frequency band; The frequency band power is calculated based on the frequency mask and amplitude spectrum to obtain the frequency domain features at different scales.
[0007] Optionally, performing projection processing on the frequency domain features at different scales includes: The frequency domain features at different scales are projected separately to obtain the projected features; The projected features are respectively connected to the corresponding auxiliary classification heads, which are used to perform auxiliary classification tasks during the training phase to achieve deep supervision.
[0008] Optionally, the frequency domain features after projection processing are stacked into a feature sequence, and attention fusion processing is performed on the feature sequence to obtain the fused features, including: The projected features at different scales are stacked into a sequence to form a feature sequence of a preset dimension; The feature sequence is input into one or more converter encoder modules to interact with the features across scales. Each encoder module includes a multi-head self-attention layer, a feedforward network layer, and layer normalization. The interactive feature sequence is input into an attention pooling layer containing a learnable query vector, and the attention relationship between the query vector and the feature sequence is calculated through a multi-head attention mechanism to obtain the fused features.
[0009] Optionally, during the training of the classification model, the total loss function is obtained by weighted summing of the main classification loss of the classification model and the auxiliary classification losses of all auxiliary classification heads. : ; in, For scale quantity, The classification loss is the classification loss of the main classification head corresponding to the fused features. For the first The classification loss of the auxiliary classification head corresponding to the frequency domain features at each scale. This is a balancing coefficient used to adjust the proportion of auxiliary classification loss in the total loss.
[0010] Optionally, the brainwave classification method further includes: mapping and generating control commands for the corresponding scenario based on the classification results.
[0011] Optionally, generating control commands for the corresponding scenario based on the classification results includes: Based on the classification results, a category probability distribution is obtained, wherein the probability distribution corresponds to a preset target stimulus frequency; The maximum value in the probability distribution of the category is selected. If the maximum value is higher than the threshold, a control command is generated; if the maximum value is lower than the threshold, no control command is generated.
[0012] The present invention also provides an adaptive frequency domain enhanced EEG classification system, comprising: an EEG signal acquisition module, a data preprocessing module, a frequency domain feature extraction module, a feature projection and fusion module, and a classification module; The EEG signal acquisition module is used to acquire the brain waves induced by the user staring at a preset target in a specific scenario. The data preprocessing module is used to preprocess the brain waves to obtain standardized brain waves; The frequency domain feature extraction module is used to extract multi-scale frequency domain features from standardized EEG signals to obtain frequency domain features at different scales. The feature projection and fusion module is used to sequentially perform projection processing on the frequency domain features, stack the frequency domain features after projection processing into a feature sequence, and perform attention fusion processing on the feature sequence to obtain fused features; The classification module is used to input the fused features into the classification model to obtain the classification result, wherein the classification model is a multi-layer fully connected network.
[0013] Compared with the prior art, the present invention has the following advantages and technical effects: (1) This invention can more effectively extract discriminative information from EEG signals by adaptively enhancing the frequency domain features related to specific tasks, thereby significantly improving classification accuracy. Compared with traditional time domain or frequency domain feature extraction methods, this invention can better adapt to the needs of different individuals and different tasks, thus achieving better classification performance.
[0014] (2) This invention can adaptively select appropriate time-frequency parameters, such as window length and window function, to meet the needs of different individuals and tasks. Compared with methods that require manual selection of time-frequency parameters, this invention can automatically optimize parameters, thereby achieving better classification performance. The invention employs a multi-scale STFT analysis method, which can extract frequency domain features of EEG signals at different time resolutions, thereby improving the robustness of classification. Even in the presence of noise and artifacts, this invention can effectively extract useful information from EEG signals, thus ensuring the accuracy of classification.
[0015] (3) This invention uses a frequency domain enhancement method to highlight the frequency components in the EEG signal that are relevant to a specific task, thereby improving the interpretability of the classification results. By analyzing the enhanced frequency domain features, we can understand the contribution of different frequency components to the classification results, thereby gaining a deeper understanding of the generation mechanism of EEG signals. Attached Figure Description
[0016] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of an adaptive frequency domain enhanced EEG classification method according to an embodiment of the present invention; Figure 2 This is a confusion matrix diagram of EEG classification results according to an embodiment of the present invention. Detailed Implementation
[0017] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0018] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.
[0019] This embodiment proposes an adaptive frequency domain enhanced EEG classification method, such as... Figure 1 As shown, the specific steps include: Collect brainwaves induced by users staring at a preset target in a specific scenario; Preprocessing of brain waves yields standardized brainwave signals; Multi-scale frequency domain feature extraction was performed on standardized EEG signals to obtain frequency domain features at different scales; Projection processing is performed on frequency domain features at different scales respectively; The frequency domain features after projection processing are stacked into a feature sequence, and attention fusion processing is performed on the feature sequence to obtain fused features; The fused features are input into the classification model to obtain the classification result. The classification model is a multi-layer fully connected network.
[0020] Specifically, this embodiment includes: collecting brainwaves induced by a user gazing at a specific target in a smart manufacturing scenario; The acquired EEG signals are preprocessed by methods such as dimension verification, filtering, and artifact removal. Distinguishing frequency domain features are extracted through multi-scale short-time Fourier transform, target frequency band mask generation, and frequency band power calculation. Multi-scale frequency domain features are spliced and fused after being processed by flattening and fully connected mapping, and then optimized and fused by consistency constraints and attention mechanisms; The fused features are input into a multi-layer fully connected network for classification. Output the classification results and map them to control commands for specific scenarios.
[0021] The acquisition and generation of electroencephalogram (EEG) signals include: By constructing an EEG-evoked environment in a smart manufacturing scenario and collecting signals, and by presenting a target with superimposed sine wave encoded transparency on a display screen, it is possible to effectively induce users to generate EEG signals related to the gaze target.
[0022] Stimulus source setup: At least three target objects (corresponding to different operational intentions, such as "grab target 1", "grab target 2", and "grab target 3") are displayed on a 27-inch screen (1920×1080 resolution, 60Hz refresh rate). Each target object is a black square (10cm×10cm) distributed on the screen at three equal intervals (20cm). A sinusoidal encoded transparency effect is superimposed on each target object to stably induce steady-state visual evoked potentials (SSVEPs). The transparency formula is... ,in The stimulation frequencies are matched with the target frequency bands extracted in the subsequent frequency domain, namely 10Hz, 12Hz, and 14Hz. The phase difference is 120° between different targets to avoid frequency confusion; the distance between the display screen and the user is 600mm, and the ambient light intensity is controlled at 50 lux to reduce visual fatigue and environmental interference.
[0023] EEG acquisition equipment and parameter configuration: An 8-channel EEG acquisition device with a sampling rate of 1000Hz is used. The reference electrode is placed in the left ear, and the ground electrode is placed on the top of the forehead. The signal acquisition channel focuses on the occipital region (corresponding to the main area where SSVEP signals are generated). The skin resistance of each electrode is ensured to be ≤10kΩ to ensure signal quality. The device synchronously records the EEG signal during the user's gaze. The single acquisition time is 3 seconds (corresponding to 3000 time points). Each target is repeatedly acquired 50 times to generate multiple trial EEG data.
[0024] Data Format and Storage: The acquired EEG signals are stored in real time in a three-dimensional (N,C,T) format (N is the number of acquisition trials, C=8 is the number of channels, and T is the number of time points). At the same time, the target label corresponding to each trial is recorded (e.g., 10Hz stimulus corresponds to the label "grasp target 1", 12Hz corresponds to the label "grasp target 2"). The data file is named "subject ID-scene number-target frequency" to facilitate subsequent reading and verification by group.
[0025] Furthermore, the brainwaves are preprocessed to obtain standardized brainwave signals, including: The average reference method was used to process the acquired EEG waves to eliminate common-mode noise; The brainwaves after common-mode noise elimination are filtered using a Butterworth filter to retain signals in a preset frequency band. Notch filtering is applied to the Butterworth filtered signal to remove preset power frequency interference. The signals after removing power frequency interference are normalized according to channel to obtain standardized EEG signals.
[0026] Specifically, the acquired EEG signals are preprocessed, including: first, common-mode noise of the EEG signals is eliminated by the average reference method; then, the Butterworth filter is used to retain the signals in the 1.0-49.5Hz frequency band; then, notch filtering is performed to remove 50Hz power frequency interference; finally, the processed signals are normalized by channel to make the signals of each channel more regular and convenient for subsequent analysis.
[0027] More specifically, to eliminate noise and artifacts in EEG signals and improve signal quality, this invention employs a series of preprocessing steps. Averaged reference eliminates common-mode interference between electrodes, Butterworth bandpass filtering extracts frequency components relevant to the target task, notch filtering eliminates power frequency interference, artifact removal removes artifacts from eye movements, electromyography, etc., downsampling reduces data volume, and Z-score standardization eliminates individual differences.
[0028] Re-reference operation: The average reference method is used, and the formula is as follows: ,in This represents the sampled value of the c-th channel of the original EEG signal at time t. To eliminate common-mode noise by averaging the signal value of channel c at time t after referencing; Notch filtering operation: Removes 50Hz power frequency interference, the formula is as follows ,in This represents the signal value at time t in channel c after notch filtering; subsequently, normalization is performed using the following formula: ; in This represents the sampled value of the c-th channel at time t after filtering. The mean of the sampled values at all time points of the c-th channel is... Let be the standard deviation of the sampled values at all time points of the c-th channel. The normalized signal value; Butterworth filtering operation: Bandpass filtering is applied to the signal, retaining the signal in the 1.0 - 49.5Hz frequency band. The formula is as follows: ,in The signal value at time t of channel c after performing Butterworth bandpass filtering; Power frequency filtering and baseline correction: for raw EEG signals ( C For the number of channels, T (Time point number), using a 50Hz notch filter to suppress power frequency interference, and using baseline correction (Baseline( X Eliminating signal drift yields a preprocessed signal. ; Artifact Removal and Downsampling: Using Independent Component Analysis (ICA) Algorithm The signal is decomposed, and artifacts such as electrooculography and electromyography are identified and removed based on the component topology map and time series features. The signal is downsampled to 256Hz to reduce data dimensionality and computational complexity.
[0029] Standardization: Z-score standardization is used to normalize the signals of each channel, ensuring that the feature distribution has a mean of 0 and a standard deviation of 1. The formula is: ; in, μ The mean of the characteristic signal, σ The standard deviation is used to ensure consistent input feature scale and optimize subsequent model training efficiency.
[0030] Furthermore, multi-scale frequency domain feature extraction is performed on the standardized EEG signals to obtain discriminative frequency domain features, including: Short-time Fourier transforms of standardized EEG signals were performed at short, medium, and long scales to obtain the amplitude spectra corresponding to each scale. K target frequency bands are preset, and a frequency mask is generated based on each target frequency band; The frequency band power is calculated based on the frequency mask and amplitude spectrum, and the frequency domain characteristics at different scales are obtained.
[0031] Specifically, multi-scale band power extraction uses short-time Fourier transform parameters of short, medium, and long scales to extract band power from the preprocessed EEG signal, including: The amplitude spectrum was obtained by performing a short-time Fourier transform on the preprocessed EEG signal. ( For the number of channels, This refers to the number of frequency points, i.e., the number of sampling points on the frequency axis after the short-time Fourier transform. (This refers to the number of time frames, i.e., the number of frames on the time axis after the short-time Fourier transform), and then for the preset... target frequency band Generate frequency mask ,in: ; in, The frequency value corresponding to the sampling point on the frequency axis. For the first The lower limit frequency of the target frequency band The upper limit frequency of the k-th target frequency band; if a frequency band has no matching frequency point (i.e. If so, the center frequency of the frequency band is adaptively selected. The corresponding frequency points are used to construct the mask; finally, the bandwidth power is calculated using the following formula: ; in, Let be the bandwidth power value of the k-th target frequency band in the c-th channel. For the preprocessed EEG signals Channel, frequency point Time frame The short-time Fourier transform amplitude at point is obtained, with dimension . The frequency band power characteristics.
[0032] The short-time Fourier transform parameters for short, medium, and long scales include: The short-time Fourier transform parameters for short scales are (The frame length of the short-time Fourier transform, i.e., the length of the signal used for each transformation). (Frame shift of short-time Fourier transform, i.e., the overlap length between adjacent frames). (Length of Fast Fourier Transform); Mesoscale Short-Time Fourier Transform parameters are , , The parameters of the long-scale short-time Fourier transform are: , , ; The short-time Fourier transform uses a Hanning window for windowing and enables trailing zero padding; the multi-scale frequency band power extraction unit is implemented as a custom layer and integrated into the model as a non-training parameter, supporting complete model saving and loading.
[0033] Generate frequency mask The preset target frequency bands (10Hz, 12Hz, 14Hz) are set to (9.4-10.6Hz), (11.4-12.6Hz), and (13.4-14.6Hz).
[0034] More specifically, the system employs multi-scale STFT parameter adaptive configuration: It innovatively sets three different scale STFT parameter combinations: short-scale (frame length 256, frame shift 128, FFT length 256), medium-scale (frame length 512, frame shift 256, FFT length 512), and long-scale (frame length 1024, frame shift 512, FFT length 1024), respectively used to capture the rapid time-varying features, medium-time-scale features, and stable low-frequency features of EEG signals. The short-scale focuses on the instantaneous changes in the signal, while the long-scale emphasizes capturing stable frequency domain components. Multi-scale collaboration achieves a comprehensive characterization of the time-frequency properties of EEG signals. All three scales employ Hanning windows with trailing zero padding to reduce spectral leakage and ensure the effectiveness of features at each scale. Adaptive frequency mask generation: For the preset target frequency band of the SSVEP signal (e.g., 10Hz corresponds to 9.4-10.6Hz, 12Hz corresponds to 11.4-12.6Hz, 14Hz corresponds to 13.4-14.6Hz), a binary frequency mask is generated according to the frequency resolution of each scale STFT (short scale Δf=1Hz, medium scale Δf=0.5Hz, long scale Δf=0.25Hz). When a target frequency band has no matching frequency point due to signal sampling rate limitations, the system will automatically calculate the center frequency of that frequency band. And select the closest frequency point to construct a mask to ensure the completeness and accuracy of frequency domain feature extraction; Fixed-dimensional feature output: A three-level calculation process—mask weighting, frequency averaging, and time averaging—is performed on the STFT amplitude spectrum at each scale to obtain fixed-dimensional (N, C, 3) frequency band power features (N is the number of samples, C is the number of channels, and 3 is the number of target frequency bands). Specifically, mask weighting filters the target frequency band energy, and frequency and time averaging eliminates interference fluctuations. The final fixed-dimensional feature output completely solves the problem of traditional feature dimensions depending on signal length, ensuring that multi-scale features can be directly adapted to subsequent classification networks, achieving a comprehensive and accurate representation of the time-frequency characteristics of EEG signals.
[0035] Furthermore, projection processing is performed on the frequency domain features at different scales, including: The frequency domain features at different scales are projected separately to obtain the projected features; The projected features are connected to the corresponding auxiliary classification heads. The auxiliary classification heads are used to perform auxiliary classification tasks during the training phase to achieve deep supervision.
[0036] Furthermore, the frequency domain features after projection processing are stacked into a feature sequence, and attention fusion processing is performed on the feature sequence to obtain fused features including: Projected features at different scales are stacked into a sequence to form a feature sequence of a preset dimension; The feature sequence is input into one or more transformer encoder modules to interact with the features across scales. Each encoder module includes a multi-head self-attention layer, a feedforward network layer, and layer normalization. The interactive feature sequence is input into an attention pooling layer containing a learnable query vector, and the attention relationship between the query vector and the feature sequence is calculated through a multi-head attention mechanism to obtain the fused features.
[0037] Specifically, the multi-scale frequency domain features are flattened and fully connected before being spliced and fused, including: Feature projection stage: The frequency band power features at short, medium, and long scales (each with dimensions (C,K), where C is the number of EEG acquisition channels and K is the number of target frequency bands) are processed separately: First, a flattening operation is performed to convert the two-dimensional (C,K) frequency band power features into a one-dimensional vector (dimension (C×K,)); then, through fully connected mapping, the weight matrix W is used... The one-dimensional vector is mapped to a 64-dimensional feature; then ReLU activation (activation function is a = max(0,z), z is the output of the fully connected layer, and a is the output after activation) introduces nonlinearity; finally, random deactivation regularization (retention probability 0.7, randomly retaining 70% of the neuron output and setting 30% of the neuron output to zero) is used to suppress overfitting. Feature fusion stage: The 64-dimensional features after projection processing at three scales are sequentially concatenated to form a 192-dimensional global feature. At the same time, a multi-scale frequency domain consistency loss function is introduced to constrain the consistency of the multi-scale frequency domain features in representation. ; in, This is the intra-scale consistency loss, used to constrain the consistency of features across channels within a single scale. Cross-scale consistency loss is used to constrain the alignment of feature distributions across different scales. and The balancing coefficients are used to adjust the proportions of intra-scale consistency loss and cross-scale consistency loss in the total loss. After experimental optimization, they are set to 0.3 and 0.7.
[0038] Then, through optimization using a fusion module based on attention mechanism and converter, weighted features are obtained by first calculating scale weights through global average pooling and multilayer perceptron: ; in, This is the result of the k-th scale feature after global average pooling (GAP: compressing the feature space dimension by averaging the features across the channel and frequency band dimensions). No. n The first trial, the first c The first channel, the first k Characteristic values of each frequency band.
[0039] ; in, It represents the attention weight of the k-th scale feature, used to measure the importance of the feature at that scale. , The activation function maps the output of the multilayer perceptron to a range of 0 to 1, used to generate attention weights. The MLP (Multilayer Perceptron) performs a non-linear transformation on the features after average pooling. ; Original features at different scales According to their respective attention weights By performing a weighted summation, we obtain the initial characteristics of the fusion. The fused features are then input into the converter encoder, in the self-attention layer. It will be converted into a matrix of query Q, key K, and value V, through The calculation, capture Deep connections between different internal parts, followed by the feedforward network layer. The output of the self-attention layer undergoes a non-linear transformation to further enhance the expressive power of the features. After processing by these two key layers, the final output is an optimized 192-dimensional fused feature.
[0040] More specifically, feature fusion and classification include: First, the optimized 192-dimensional fused features are input into a multi-layer fully connected network. The first layer performs batch normalization, and the normalization formula is as follows: ; Where x is the characteristic signal, The mean of the characteristic signal, For characteristic variance, To prevent the small constant of division by zero and accelerate training convergence; then it passes through a 256-dimensional fully connected layer in sequence: ; 128-dimensional fully connected layer: ; Each layer is equipped with random inactivation regularization (retention probability 0.5) and L2 regularization (regulation term is...). The network output layer uses the Softmax activation function. , (For the logits output of class i), the 128-dimensional features are mapped to the probability distribution of each class; Automatic class weight calibration: Class weights are calculated only based on the training set data to avoid information leakage; a "balanced" strategy is used to balance the loss contribution of samples from different classes, solving the accuracy deviation caused by class imbalance; Multi-level regularization: 0.3 dropout is used in the feature projection stage to suppress local overfitting, and 0.5 dropout + batch normalization is used in the classification network stage to balance training convergence speed and generalization ability. Dynamic training control: An early stopping mechanism (patience value 20) is introduced to prevent model overfitting. Combined with a learning rate decay mechanism (decay factor 0.5, patience value 8), stable convergence is ensured in the later stages of training. In addition, an optimal model saving mechanism (only saving the model parameters with the highest accuracy on the validation set) is used to reduce storage redundancy while ensuring classification performance.
[0041] Furthermore, projection processing is performed on frequency domain features at different scales, including: First, the frequency domain features at different scales are flattened into a two-dimensional feature matrix to obtain the flattened features; The frequency domain features after projection are obtained by performing a fully connected mapping on the flat features using a weight matrix.
[0042] Furthermore, a fusion process is performed on the frequency domain features after constraint optimization to obtain fused features including: First, the pooling result of the frequency domain features after constraint optimization is calculated by global average pooling. Then, the attention weight of the frequency domain features after constraint optimization is calculated by multilayer perceptron and sigmoid activation function. The frequency domain features after constraint optimization are weighted according to the attention weights to obtain the weighted features; The weighted features are input into the converter encoder and processed sequentially through a self-attention layer and a feedforward network layer to obtain the fused features.
[0043] Furthermore, during the training of the classification model, a multi-scale frequency domain consistency loss function is used. With classification loss function Weighted summation as the total loss function : in, , This is a balancing coefficient used to adjust the proportion of classification accuracy and feature consistency in the total loss.
[0044] Specifically, deep processing of the fused features to achieve classification includes: The optimized 192-dimensional fused features are first input into a multi-layer fully connected network. The first layer performs batch normalization (standardizing the features to accelerate network training convergence). Subsequently, the features pass through a 256-dimensional fully connected layer and a 128-dimensional fully connected layer. Each layer is equipped with a ReLU activation function (introducing non-linear expressive power), random deactivation regularization (retaining a probability of 0.5), and L2 regularization (constraining the weight parameter size to further suppress overfitting). The network output layer uses a Softmax activation function to map the 128-dimensional features to probability distributions for each category (the sum of probabilities is 1, corresponding to the classification results of different operational intentions). During training, the multi-scale frequency domain consistency loss function is used. With classification loss (Cross-entropy loss, which measures the difference between the model's predicted class probabilities and the true labels) is weighted and summed to obtain the total loss; After experimental optimization, the following was set 0.7 The learning rate is set to 0.3. The Adam optimizer (initial learning rate 1e-3) is used to minimize the total loss. At the same time, early stopping and learning rate decay mechanisms are enabled to ensure the model's generalization ability. After training, the model with the highest accuracy on the validation set is selected as the optimal model. The fused features of the input are classified and predicted, and the class with the highest probability is output as the final classification result so that it can be mapped to control instructions for specific scenarios in the future.
[0045] Furthermore, the EEG classification method also includes: mapping and generating control commands for the corresponding scenario based on the classification results.
[0046] Furthermore, the control commands generated based on the classification results for the corresponding scenario include: Based on the classification results, obtain the category probability distribution, where the probability distribution corresponds to the preset target stimulus frequency; Select the maximum value in the category probability distribution. If the maximum value is higher than the threshold, a control command is generated; if the maximum value is lower than the threshold, no control command is generated.
[0047] Specifically, the classification results are output and mapped to control commands for specific scenarios, including: Obtain the class probability distribution of the output of a multilayer fully connected network unit (e.g., the probability corresponding to stimulation frequencies of 10Hz, 12Hz, and 14Hz). The category corresponding to the maximum probability is selected as the final classification result. If the maximum probability is lower than the preset threshold, "no valid intent" is output and no control command is generated. If the probability is higher than the threshold, the corresponding operation intent is determined based on the preset "frequency-intent-command" mapping table (e.g., 10Hz classification result corresponds to "grab target 1", 12Hz corresponds to "grab target 2" and 14Hz corresponds to "grab target 3").
[0048] The following is a detailed description of an adaptive frequency domain enhanced EEG classification method in this embodiment: The EEG classification system described in this embodiment is built using the Python programming language and the TensorFlow deep learning framework. It adopts a highly modular design, with each module functioning independently yet working closely together to ensure efficient and stable system operation. The system mainly includes a data reading module, a preprocessing module, an adaptive frequency domain enhancement module, a classification model building module, an intelligent training module, and a result management module. The data reading module is responsible for reading EEG data and corresponding label data according to groups; the preprocessing module performs channel-level normalization on the EEG data and adaptively converts the label format; the adaptive frequency domain enhancement module generates fixed-dimensional features based on multi-scale short-time Fourier transform; the classification model building module constructs the FEMT-NET architecture model based on the input shape of the EEG data; the intelligent training module configures the optimizer and loss function and starts training, while using strategies such as early stopping and learning rate decay to optimize the training process; the result management module is responsible for saving various types of data during the training process and summarizing the key performance indicators of all groups.
[0049] This embodiment provides an adaptive frequency domain enhanced EEG classification method and system, including the following steps: Step 1, Signal Preprocessing: 1.1 Power Frequency Filtering and Baseline Correction: raw EEG signals ( C For the number of channels, T (Time point number), using a 50Hz notch filter to suppress power frequency interference, and using baseline correction (Baseline( X Eliminate signal drift to obtain a preprocessed signal. ; 1.2 Artifact Removal and Downsampling: Using Independent Component Analysis (ICA) algorithm The signal is decomposed, and artifacts such as electrooculography and electromyography are identified and removed based on the component topology map and time series features. The signal is downsampled to 256Hz to reduce data dimensionality and computational complexity.
[0050] 1.3 Standardization Processing: The Z-score normalization method is used to normalize the signals of each channel so that the feature distribution has a mean of 0 and a standard deviation of 1. The formula is as follows: ; in, μ The mean of the characteristic signal, σ The standard deviation is used to ensure consistent input feature scale and optimize subsequent model training efficiency.
[0051] Step 2: Adaptive multi-scale frequency domain enhanced feature extraction: This step is completed using a custom STFTBandPowerLayer (implemented based on the TensorFlow / Keras framework), the core of which is the normalization of the signal. (N is the number of samples, N=800 in this embodiment) Multi-scale frequency domain feature extraction is performed, and the specific process is as follows: 2.1 Multi-scale STFT parameter configuration and transformation: Short-time Fourier Transform (STFT) parameters are set for short, medium, and long time scales to adapt to the frequency domain characteristics of EEG signals at different time scales: Short-scale STFT: Frame length = 256, frame shift = 128, FFT length = 256, Hanning window with additional windows, zero padding at the end to the frame length; Mesoscale STFT: Frame length = 512, frame shift = 256, FFT length = 512, Hanning window with additional windows, zero padding at the end to the frame length; Long-scale STFT: Frame length = 1024, frame shift = 512, FFT length = 1024, Hanning window with additional windows, zero padding at the end to the frame length; right Perform STFT transformations at three different scales to obtain the amplitude spectra at each scale: ; ; ; in , , These represent the number of frequency points at each scale (FFT length / 2 + 1). =4、 =2、 =1 represents the number of time frames at each scale (calculated from T' and frame shift).
[0052] 2.2 Adaptive generation of target frequency band mask: For the steady-state visual evoked potential (SSVEP) task, three target frequency bands are preset: B1=(9.4-10.6Hz), B2=(11.4-12.6Hz), B3=(13.4-14.6Hz), and the corresponding frequency mapping FREQ_MAP=[10,12,14]. The frequency resolution is calculated based on the frequency axis of each STFT scale (obtained from the sampling rate of 256Hz and the FFT length). Generate binary frequency mask (k=1,2,3 correspond to 3 frequency bands, F is the number of frequency points at the corresponding scale), the mask generation rule is: Other cases. If a frequency band has no matching frequency point (e.g., no corresponding frequency in B1 in a long-scale STFT), the center frequency of that frequency band will be automatically calculated. Select the closest frequency axis For each frequency point, set its corresponding mask position to 1 to ensure the integrity of the frequency domain enhancement.
[0053] 2.3 Bandwidth power calculation and fixed-dimensional feature output: Bandwidth power is a core feature reflecting the energy distribution of EEG signals within a specific frequency range. This step extracts stable bandwidth power features from the STFT amplitude spectrum at various scales through a three-level calculation process: "frequency mask weighting → frequency dimension weighted mean → time dimension averaging". The specific process is as follows: Frequency mask weighting operation: For STFT amplitude spectra at short, medium, and long scales, weighting is performed using target frequency band masks corresponding to each scale—the masks filter out energy components belonging to the target frequency band in the amplitude spectrum, while blocking interference components from non-target frequency bands. Taking the short scale as an example, for the... The sample, the first Short-scale STFT amplitude spectrum of the channel , and short-scale Mask of each frequency band Performing element-wise multiplication yields the weighted amplitude spectrum: ; Among them, when Belongs to the When there is a target frequency band, Preserve the amplitude information at that frequency point; when Not belonging to the first When there is a target frequency band, This allows for the shielding of amplitude information at a specific frequency point, enabling precise screening of energy in the target frequency band.
[0054] Short-scale bandwidth power calculation (taking short-scale as an example): Short-scale band power calculation requires three steps in sequence: "weighted summation in the frequency dimension → averaging in the time dimension → normalization". The specific formula is as follows: ; The definitions and calculation logic of each parameter in the formula are as follows: : No. The sample, the first Channel, First The short-scale bandwidth power of each target frequency band, in units of This reflects the energy intensity of the sample in the corresponding channel and frequency band; : Number of time frames in short-scale STFT (in this embodiment) ), composed of short-scale frame shift of 128 and normalized signal duration The calculation yields the following formula: ; Short-scale The number of "1"s in a frequency band mask, i.e., the number of frequency points that match the band on the short-scale frequency axis (in this embodiment, the first short-scale frequency band). The number of matched frequency points is 3, therefore ; : No. The sample, the first Channel, First The total energy of each frequency band at a short scale is obtained by summing the weighted amplitude spectrum at all frequency points and all time frames. In the formula, the denominator It is the product of "number of time frames × number of effective frequency points" and is used to normalize the total energy, eliminate the influence of differences in the number of effective frequency points in different frequency bands and the number of time frames at different scales on power calculation, and ensure that the power characteristics of different scales and different frequency bands are comparable.
[0055] Mesoscale and longscale bandwidth power calculation: The calculation logic for medium- and long-scale bandwidth power is completely consistent with that for short-scale power; only the parameters for the corresponding scales need to be replaced. Mesoscale band power: Mesoscale STFT amplitude spectrum was used. Mesoscale mask Mesoscale time frame count (In this embodiment) ), number of mesoscale frequency points (In this embodiment) ); Long-scale bandwidth power: Long-scale STFT amplitude spectrum is used. Long-scale mask Long-scale time frame count (In this embodiment) ), number of long-scale frequency points (In this embodiment) ).
[0056] Fixed-dimensional feature generation: For all samples (total) (number), all channels (total) After performing the above frequency band power calculations sequentially on all target frequency bands (3 in total), the results are organized according to the "sample-channel-frequency band" dimension to obtain fixed-dimensional feature matrices of three scales: Short-scale feature matrix Dimensions are , of which The element corresponds to the first element. The sample, the first Channel, First Short-scale bandwidth power of each frequency band; Mesoscale feature matrix Dimensions are , of which The element corresponds to the first element. The sample, the first Channel, First Midscale band power of each frequency band; Long-scale feature matrix Dimensions are , of which The element corresponds to the first element. The sample, the first Channel, First Long-scale bandwidth power of each frequency band.
[0057] This process, through "fixed target frequency bands + standardized calculation procedures," ensures that the final output feature dimension is only related to the number of samples. Number of channels The target frequency band number is related to 3, and is also related to the duration of the original EEG signal. Complete decoupling completely solves the problem that the dimensions of traditional frequency domain features (such as STFT amplitude spectrum) change with the signal length and are difficult to directly input into the classification network. This ensures that the three scale features can be directly used for subsequent feature projection fusion and model training, improving the system's versatility and stability.
[0058] Step 3: Feature projection fusion and classification model training: 3.1 Feature Projection and Fusion: Feature projection processing is performed on the frequency domain features at the three scales respectively, as follows: Flattening: , , The (N×C×3) three-dimensional tensor is flattened into an (N×(C×3)) two-dimensional feature matrix. (In this embodiment, the number of EEG channels C=8, so the dimension after flattening is N×24). Fully connected mapping: through a fully connected layer (weight matrix) The flattened features are mapped to 64-dimensional features, and the activation function used is ReLU (…). (where z is the output of the fully connected layer). Dropout regularization: Set the dropout probability to 0.3 to randomly block 30% of the neuron outputs, thus suppressing local overfitting; Feature projection and fusion: Flattening is first performed on features at three scales ( ), 64-dimensional fully connected mapping (formula is) ReLU activation and 0.3 random deactivation; then a multi-scale frequency domain consistency loss function is introduced to constrain feature consistency, and then fusion is performed through an attention-transformer two-stage architecture—after global average pooling ( ) and the computational scale weights of the multilayer perceptron ( )and ) to obtain weighted features ), input converter encoder (including self-attention layer) The deep relationships are captured and finally spliced into 192-dimensional global features; 3.2 Classification Model Construction and Training Configuration: A two-level fully connected structure is adopted. The first level uses a weight matrix. The mapping is to 256-dimensional features, where W1 is the weight matrix of the first-level fully connected layer; the second level uses the weight matrix... The mapping is to 128-dimensional features, and W2 is the weight matrix of the second-level fully connected layer; batch normalization is performed sequentially after each fully connected layer: ; z is the output of the fully connected layer. This represents the mean of the data within the batch. The standard deviation of the data within the batch. (The output after batch normalization), ReLU activation and dropout (with a retention probability of 0.5).
[0059] Through the fully connected layer, the weight matrix is: ; W3 is the weight matrix of the fully connected classification layer. For the number of categories, the weight matrix is mapped to... 3D features, activated by softmax: ; For classifying fully connected layers One output, For the first The probability of a class is output, and the probability distribution of the class is displayed. L2 regularization is used during training, and an additional loss function is applied. The term, where λ is the regularization coefficient, This is the weight matrix for each fully connected layer, and the class weights are automatically calculated based on the training set. (N is the total number of samples in the training set, For the number of categories, The contribution of the balance loss is calculated based on the number of samples in class c.
[0060] Optimizer: Adam optimizer, initial learning rate 1e-3, first moment decay rate β1=0.9, second moment decay rate β2=0.999; Loss function: Total loss is , To combine category weights The classification cross-entropy loss; Training strategy: early stop mechanism (patience value 20, training is terminated if there is no improvement), learning rate decay (patience value 8, decay factor 0.5, minimum learning rate 1e-6), training until the accuracy on the validation set is stable, and saving the optimal model (HDF5 format).
[0061] Dataset partitioning: The 800 samples are divided into a training set (560 samples) and a validation set (240 samples) in a 7:3 ratio. The batch size is 32 and the maximum number of training epochs is 120.
[0062] 3.3 Model Training and Convergence Validation: In this embodiment, the model converges after 35-40 training rounds, and the accuracy on the validation set stabilizes at over 93%, which is 9 percentage points higher than the traditional single-scale STFT method (accuracy 84%), proving the effectiveness of multi-scale frequency domain enhancement. After training, the optimal model is saved to the path for subsequent prediction.
[0063] This embodiment also provides an adaptive frequency domain enhanced EEG classification system, including: an EEG signal acquisition module, a data preprocessing module, a frequency domain feature extraction module, a feature projection and fusion module, and a classification module; The EEG signal acquisition module is used to acquire the brain waves induced by the user staring at a preset target in a specific scenario. The data preprocessing module is used to preprocess the brain waves to obtain standardized brain signals; The frequency domain feature extraction module is used to extract multi-scale frequency domain features from standardized EEG signals to obtain frequency domain features at different scales. The feature projection and fusion module is used to sequentially perform projection processing, consistency constraint optimization, and fusion processing on the frequency domain features to obtain fused features. A multi-layer fully connected network module is used to receive fused features, which are then processed by batch normalization, fully connected layers, and loss function optimization to output a class probability distribution. The classification module is used to input the fused features into the classification model and obtain the classification result. The classification model is a multi-layer fully connected network.
[0064] Specifically, this embodiment proposes an adaptive frequency-domain enhanced EEG classification system. This system achieves efficient classification of EEG signals through a process of "real-time EEG signal acquisition and input → pre-trained model processing → predicted classification result output." The real-time EEG signal serves as the system input, acquired by an EEG acquisition device in scenarios such as intelligent manufacturing, collecting EEG waves induced by a user's gaze at a specific target. The acquired signal undergoes preprocessing operations such as data verification (ensuring dimensionality meets requirements), filtering, and artifact removal. The pre-trained model is trained based on the adaptive frequency-domain enhanced EEG classification method, integrating core modules such as multi-scale short-time Fourier transform, target frequency band mask generation, frequency band power calculation, feature projection fusion, and multi-layer fully connected networks. This model enables adaptive frequency domain feature extraction and deep classification processing of the input real-time EEG signal. The predicted classification result is the system output, which can be mapped to specific control commands (such as grasping and placement commands for a robotic arm in a brain-computer interface scenario) or diagnostic reference information (such as abnormal prompts in the auxiliary diagnosis of neurological diseases) depending on the actual application scenario, thus providing a basis for subsequent task execution.
[0065] To verify the effectiveness and implementation of the adaptive frequency domain enhanced EEG classification method and system proposed in this embodiment, the following scientific experiments were conducted.
[0066] (1) Experimental volunteers: Six healthy volunteers (aged 22-30, 3 males and 3 females) were recruited. All volunteers were right-handed and had no history of neurological diseases, visual impairment, or participation in EEG experiments. Before the experiment, volunteers were informed of the experimental procedure and precautions and signed informed consent forms. During the experiment, volunteers were required to keep their heads stable, avoid frequent blinking and limb movements, reduce interference from electrooculography (EOG) and electromyography (EMG) artifacts on EEG signals, and ensure the validity of the collected data.
[0067] (2) Experimental setup: EEG acquisition equipment and parameters: An 8-channel EEG acquisition device (sampling rate 1000Hz, reference electrode placed in the left ear, ground electrode placed on the top of the forehead) was used. The acquisition channels focused on the occipital region (corresponding to the main SSVEP signal generation area, such as the electrode positions O1, O2, Oz, etc.). Before acquisition, it was ensured that the skin contact resistance of each electrode was ≤10kΩ to avoid signal attenuation. EEG data was stored in real time in (N,C,T) format (N is the number of acquisition trials, C=8 is the number of channels, and T is the number of time points). The acquisition time of a single trial was 3 seconds (corresponding to T=3000 time points). Each target frequency band was repeatedly acquired 50 times. The data file was named "Volunteer ID-Trial Number-Target Frequency" for easy subsequent grouping verification and reading.
[0068] Stimuli and Scene Construction: A 27-inch display screen (1920×1080 resolution, 60Hz refresh rate) was used to display three target targets (preset target frequency bands of 10Hz, 12Hz, and 14Hz). The target targets were black squares (10cm×10cm in size) and distributed at three equal intervals (20cm) along the horizontal center line of the screen. The distance between the display screen and the volunteers was 1.5m, and the ambient light intensity was controlled at 50 lux to reduce visual fatigue. Each target is superimposed with a sine wave encoded transparency effect. The transparency calculation formula is as follows: ,in The frequencies are 10Hz, 12Hz, and 14Hz (matching the center frequency of the target frequency band). To avoid EEG signal interference caused by frequency confusion, phase difference (the three targets are 0°, 120°, and 240° respectively).
[0069] Experimental hardware and software environment: Hardware: The local client uses an Intel Core i7-12700H computer (16GB RAM) to handle data acquisition and preprocessing; the remote server uses an NVIDIA RTX 3090 server (64GB RAM) to handle multi-scale frequency domain enhancement and model training. The client and server communicate via the DDS protocol to ensure data transmission latency ≤100ms. Software: The experimental environment was built based on Python 3.9, and custom layers were implemented using the TensorFlow 2.10 framework. NumPy 1.23 was used for numerical computation, and Scikit-learn 1.2 was used for data partitioning and evaluation. Logs were recorded in real time through the logging module to ensure the reproducibility of the experiment.
[0070] (3) Online experimental process: Data collection and verification: Volunteers focused on three target objects in sequence according to the experimental instructions (focusing on each target for 3 seconds, with a 2-second rest interval), and the device simultaneously collected EEG signals and corresponding tags (10Hz → tag 0, 12Hz → tag 1, 14Hz → tag 2). The data verification module automatically checks whether the dimensions of the collected data conform to the (N,C,T) format (N=50, C=8, T=5000 in this stage). If there are missing files or abnormal dimensions, a warning log is recorded in real time and the group of trials is re-collected to ensure the integrity of the input data.
[0071] Data preprocessing: The validated EEG data undergoes a four-step preprocessing step: common-mode noise is eliminated using average reference (formula: Effective frequency bands are preserved through Butterworth bandpass filtering (1.0-49.5Hz); power frequency interference is removed through 50Hz notch filtering. The ICA algorithm is used (maximum 1000 iterations, convergence threshold 1e). -6 The signal was decomposed into 8 independent components. Two electrooculography (EOG) components and one electromyography (EMG) component were identified and removed using scalp topology and time-series features. After reconstructing the artifact-free signal, it was downsampled to 256Hz (T'=1280) and Z-score normalization was performed on each channel (formula: ); The label conversion module automatically recognizes the label format (1D integer), converts it into (N,3)-dimensional one-hot encoding (e.g., label 0 → [1,0,0]), and generates standardized data input.
[0072] Online prediction and intent output: Load the trained optimal model, and the volunteer fixates on the target again (10 trials are randomly selected). The system reads the EEG data in real time and completes preprocessing, frequency domain enhancement and feature fusion, and outputs the classification results (target frequency and corresponding probability). If the maximum classification probability is ≥0.5, the corresponding operation intention is output (e.g., 10Hz → "robotic arm ready to grasp"); if <0.5, "no valid intention" is output and the data is collected again to ensure the reliability of online prediction.
[0073] (4) Experimental results: The experimental results were evaluated from three dimensions: classification performance indicators, feature validity verification, and system stability. All data were based on 800 valid samples from 6 volunteers (150 training samples and 50 validation samples per volunteer).
[0074] Classification performance indicators: Accuracy and F1 score: The average classification accuracy of the method of this invention is 93.5%, and the average F1 score is 0.93. The accuracy of the 10Hz target frequency band is the highest (95.2%), and the accuracy of the 14Hz target frequency band is the lowest (91.8%), but both are significantly higher than the traditional single-scale STFT method (accuracy 84.0%, F1 score 0.83), which proves the effect of multi-scale frequency domain enhancement on improving feature discrimination. To intuitively evaluate the classification performance of the EEG classification model of this invention, a confusion matrix is used for visualization, such as... Figure 2 As shown in the figure, the horizontal axis represents the category predicted by the model, and the vertical axis represents the actual category. The categories correspond to three SSVEP stimulation frequencies: 10Hz, 12Hz, and 14Hz. The values in the matrix represent the number of samples in the corresponding category, and the gray bars on the right are a mapping reference between the values and colors (grayscale). The larger the value, the darker the grayscale. From the confusion matrix, we can see that: for samples that actually occurred at 10Hz, the model accurately predicted 73, and misclassified them as other categories by 0; among samples that actually occurred at 12Hz, 71 were accurately predicted, and only 2 were misclassified as 10Hz; among samples that actually occurred at 14Hz, 69 were correctly predicted, 3 were misclassified as 10Hz, and 1 was misclassified as 12Hz.
[0075] Individual difference analysis: The accuracy of the 6 volunteers ranged from 91.2% to 95.8%, with a standard deviation of only 2.1%, which is much lower than the 5.3% of the traditional method. This shows that the present invention effectively reduces the impact of individual EEG differences on classification results through fixed-dimensional features and adaptive mask design, and has better robustness.
[0076] Training stability: The model converges in an average of 38 rounds, and the validation loss does not fluctuate significantly during training (it eventually stabilizes at around 0.15). The early stopping mechanism effectively avoids overfitting, and the optimal model parameters are only 105,000 (far lower than EEGNet's 148,000), resulting in low storage overhead. Online real-time performance: The average time from data acquisition to intention output in a single trial is 3.5 seconds, including 1.2 seconds for preprocessing, 1.8 seconds for frequency domain enhancement, and 0.5 seconds for model prediction, which meets the requirements of real-time control of brain-computer interface (delay ≤ 10 seconds).
[0077] In summary, the experimental results show that the adaptive frequency domain enhanced EEG classification method and system proposed in this embodiment can effectively solve the problem of insufficient frequency domain feature representation caused by traditional single-scale STFT. It performs excellently in classification accuracy, robustness and real-time performance, verifying the effectiveness and practical value of the method and system.
[0078] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An adaptive frequency domain augmentation EEG classification method, characterized in that, include: Collect brainwaves induced by users staring at a preset target in a specific scenario; The brain waves are preprocessed to obtain standardized brainwave signals; Multi-scale frequency domain feature extraction was performed on standardized EEG signals to obtain frequency domain features at different scales; Projection processing is performed on the frequency domain features at different scales respectively; The frequency domain features after projection processing are stacked into a feature sequence, and attention fusion processing is performed on the feature sequence to obtain fused features; The fused features are input into a classification model to obtain classification results, wherein the classification model is a multi-layer fully connected network.
2. The adaptive frequency domain enhancement EEG classification method according to claim 1, characterized in that, Preprocessing the brainwaves to obtain standardized brainwave signals includes: The average reference method was used to process the acquired EEG waves to eliminate common-mode noise; The brainwaves after common-mode noise elimination are filtered using a Butterworth filter to retain signals in a preset frequency band. Notch filtering is applied to the Butterworth filtered signal to remove preset power frequency interference. The signals after removing power frequency interference are normalized according to channel to obtain standardized EEG signals.
3. The adaptive frequency domain enhancement EEG classification method according to claim 1, characterized in that, Multi-scale frequency domain feature extraction is performed on standardized EEG signals to obtain discriminative frequency domain features, including: Short-time Fourier transforms of standardized EEG signals were performed at short, medium, and long scales to obtain the amplitude spectra corresponding to each scale. K target frequency bands are preset, and a frequency mask is generated based on each target frequency band; The frequency band power is calculated based on the frequency mask and amplitude spectrum to obtain the frequency domain features at different scales.
4. The adaptive frequency domain enhancement EEG classification method according to claim 1, characterized in that, Performing projection processing on the frequency domain features at different scales includes: The frequency domain features at different scales are projected separately to obtain the projected features; The projected features are respectively connected to the corresponding auxiliary classification heads, which are used to perform auxiliary classification tasks during the training phase to achieve deep supervision.
5. The adaptive frequency domain enhancement EEG classification method according to claim 4, characterized in that, The frequency domain features after projection processing are stacked into a feature sequence, and attention fusion processing is performed on the feature sequence to obtain the fused features, including: The projected features at different scales are stacked into a sequence to form a feature sequence of a preset dimension; The feature sequence is input into one or more converter encoder modules to interact with the features across scales. Each encoder module includes a multi-head self-attention layer, a feedforward network layer, and layer normalization. The interactive feature sequence is input into an attention pooling layer containing a learnable query vector, and the attention relationship between the query vector and the feature sequence is calculated through a multi-head attention mechanism to obtain the fused features.
6. The adaptive frequency domain enhanced EEG classification method according to claim 1 or 4, characterized in that, During training, the classification model uses a weighted sum of the main classification loss and the auxiliary classification losses of all auxiliary classification heads as the total loss function. : ; in, For scale quantity, The classification loss is the classification loss of the main classification head corresponding to the fused features. For the first The classification loss of the auxiliary classification head corresponding to the frequency domain features at each scale. This is a balancing coefficient used to adjust the proportion of auxiliary classification loss in the total loss.
7. The adaptive frequency domain enhancement EEG classification method according to claim 1, characterized in that, The brainwave classification method further includes: mapping and generating control commands for the corresponding scenario based on the classification results.
8. The adaptive frequency domain enhancement EEG classification method according to claim 7, characterized in that, The control commands generated based on the classification results for the corresponding scenario include: Based on the classification results, a category probability distribution is obtained, wherein the probability distribution corresponds to a preset target stimulus frequency; The maximum value in the probability distribution of the category is selected. If the maximum value is higher than the threshold, a control command is generated; if the maximum value is lower than the threshold, no control command is generated.
9. An adaptive frequency domain enhanced EEG classification system for implementing the method according to any one of claims 1-8, characterized in that, include: The system includes an EEG signal acquisition module, a data preprocessing module, a frequency domain feature extraction module, a feature projection and fusion module, and a classification module. The EEG signal acquisition module is used to acquire the brain waves induced by the user staring at a preset target in a specific scenario. The data preprocessing module is used to preprocess the brain waves to obtain standardized brain waves; The frequency domain feature extraction module is used to extract multi-scale frequency domain features from standardized EEG signals to obtain frequency domain features at different scales. The feature projection and fusion module is used to sequentially perform projection processing on the frequency domain features, stack the frequency domain features after projection processing into a feature sequence, and perform attention fusion processing on the feature sequence to obtain fused features; The classification module is used to input the fused features into the classification model to obtain the classification result, wherein the classification model is a multi-layer fully connected network.
Citation Information
Patent Citations
Fusion feature enhanced fracture image classification method and system
CN120259780A
Brain-controlled unmanned aerial vehicle group method based on brain-computer deep collaborative fusion
CN120295326A
Epilepsy automatic detection method and system based on EEG
CN120477705A