Electroencephalogram classification method and system based on physiological prior guidance and cross-scale information fusion, electronic equipment and storage medium
By combining the parallel structure of CNN and Transformer, using physiological prior guidance and cross-scale information fusion methods, the lack of classification performance of the electroencephalogram signal classification method in the existing technology under complex auditory stimuli is solved, and higher classification accuracy and ability to understand the brain's processing stimuli are achieved. It is suitable for brain-computer interfaces and neurorehabilitation fields.
Patent Information
- Application Number
- CN202510551377.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-15
AI Technical Summary
The existing EEG signal classification methods lack a deep understanding of brain neural activities, especially in complex auditory stimulation situations, and it is difficult to fully explore the interactions between different brain regions, and fail to effectively integrate physiological priors and brain region diffusion behaviors related to brain region traceability, resulting in insufficient classification performance and generalization ability.
Using the EEG classification method based on physiological prior guidance and cross-scale information fusion, local and global features are extracted through the parallel structure of CNN and Transformer, combined with the multi-level physiological prior feature guidance method, the global feature capture ability of Transformer and the local feature capture ability of CNN are fused to different levels of brain state data as the classification feature set.
It improves the accuracy of EEG signal classification, can better understand how the brain processes and responds to various stimuli, and provides new tools and perspectives for the fields of brain-computer interface, neurorehabilitation and cognitive science, and is suitable for sound signal classification and recognition tasks in complex environments.
Smart Images

Figure CN120493005A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of biometric recognition technology, and relates to an EEG classification method, system, electronic device and storage medium based on physiological prior guidance and cross-scale information fusion, which is suitable for EEG signal classification and recognition, especially EEG signals of non-invasive EEG acquisition signals. Background Art
[0002] The brain response triggered by auditory stimulation is a key area of electroencephalography (EEG) research. How to effectively classify and identify EEG signals stimulated by different types of auditory stimulation has always been an important challenge in scientific research. Traditional EEG classification methods mostly rely on machine learning algorithms for feature extraction and classification. Although they have improved the classification accuracy to a certain extent, they lack a deep understanding of the brain's neural activities. Especially in complex auditory stimulation scenarios, it is difficult to fully explore the interactions between different brain regions. In addition, this type of method does not adequately characterize the physiological characteristics of EEG signals. For example, it fails to effectively integrate physiological priors related to brain region tracing and brain region diffusion behavior, resulting in a vague mapping relationship between channel features and brain region activities, making it difficult to capture the biological mechanism of cross-brain region collaborative processing of auditory stimuli, thereby limiting classification performance and generalization capabilities.
[0003] Existing research attempts to combine deep learning with traditional feature extraction methods to further improve classification results. For example, Chinese patent publication number CN115470821A incorporates a more complex model structure into the neural network to improve classification accuracy; Chinese patent publication number CN118797496A focuses more on the spatiotemporal characteristics of brain signals and optimizes recognition performance by fusing spatiotemporal features outside the network with multi-scale features. However, these methods often underemphasize the biological characteristics of the brain and fail to fully utilize the prior knowledge contributed by brain regions.
[0004] In the processing of visual and auditory information, the brain possesses relatively fixed and functionally distinct neural circuits, with different brain regions working collaboratively across scales and at multiple levels. The primary auditory and visual cortex is primarily responsible for encoding and initially processing the basic features of sensory signals, while the intermediate and higher-level cortex further analyzes and understands the signals based on higher-level feature extraction and information integration. Through multi-level information integration, the brain is able to perceive and synthesize external stimuli in multiple dimensions, ultimately forming a complete perception and cognition of the external environment.
[0005] Therefore, in order to make full use of the brain's cross-scale and multi-level functional integration mechanism and brain region functional specificity to improve the accuracy of EEG signal classification, the present invention proposes an EEG classification method, system, electronic device and storage medium based on physiological prior guidance and cross-scale information fusion to make up for the above-mentioned shortcomings and provide new perspectives and tools for research in fields such as brain-computer interface, neurorehabilitation and cognitive science. Summary of the Invention
[0006] In order to solve the above problems, the present invention provides an EEG classification method, system, electronic device and storage medium based on physiological prior guidance and cross-scale information fusion, which integrates the global feature extraction capability of Transformer and the local feature capture capability of CNN. Through a multi-level physiological prior feature guidance method, brain state data at different levels are extracted and fused as a classification feature set, effectively improving the classification accuracy.
[0007] The technical solution adopted by the present invention is an EEG classification method based on physiological prior guidance and cross-scale information fusion, comprising the following steps:
[0008] S1, obtaining sound stimulation data, and performing matching processing on the sound stimulation data to make them consistent in physical characteristics;
[0009] S2, collecting EEG data of sound stimulation and preprocessing the collected EEG data;
[0010] S3, divide the EEG data into test set and training set, and determine the hyperparameters of model training;
[0011] S4, feature extraction of EEG data of the training set is performed through the parallel structure of CNN and Transformer;
[0012] The CNN branch structure is used to extract local features, including:
[0013] Module 1, wherein the module controls the complexity of the network through bottleneck blocks, reduces the amount of computation and improves the efficiency of feature extraction through depthwise separable convolution, and improves the training stability of the network through residual connections;
[0014] Module 2, wherein the module 2 increases the depth of the network layer by stacking convolutional layers;
[0015] The Transformer model is used to capture global features. The Transformer model includes a Transformer encoder. Channel embedding and time step embedding are added to the position embedding in the Transformer encoder. The channel embedding, time step embedding, and position embedding are added to the original input.
[0016] The global features extracted by the Transformer branch and the local features extracted by the CNN branch are guided and fused through the multi-scale upsampling physiological prior guidance module and the convolutional fusion module;
[0017] S5, classify the extracted features through the classifier module, and train and verify the classifier.
[0018] Furthermore, in S4, module 1 includes:
[0019] Two 1×1 convolutional layers are used to adjust the number of channels and reduce the amount of computation;
[0020] Depthwise separable convolution layer uses a 3×3 convolution kernel and sets groups = number of channels to implement depthwise separable convolution, that is, first performing channel-wise convolution and then point-wise convolution;
[0021] Residual connection is used to add the original input x through the output of the two-dimensional convolution layer with a convolution kernel of 1 to the processed feature map;
[0022] The weights of the batch normalization in the last layer of module one are initialized to 1.
[0023] Furthermore, in S4, module 2 includes:
[0024] Two 1×1 convolutional layers are used to adjust the number of channels and reduce the amount of computation;
[0025] A 3×3 convolutional layer with groups set to 1 to ensure that all input and output channels are the same.
[0026] Residual connection is used to add the input of module 2 directly to the output of the last 1×1 convolutional layer.
[0027] Furthermore, in S4, channel embedding includes:
[0028] First, an embedding layer is used to map the signal of each channel into a higher-dimensional space;
[0029] Then, the shape of the embedding matrix is expanded from [C, d] to [1, C, d, 1] through the expansion dimension operation, and repeated on the batch size and time step through the repeat operation, so that the embedding matrix is expanded in the batch dimension and time dimension to match the dimension of the original input data;
[0030] Finally, the dimension order of the embedding matrix is adjusted through the dimension permutation operation so that it can be added to the input data correctly;
[0031] The time-step embedding includes:
[0032] First, each time step is embedded through the embedding layer;
[0033] Then, the dimension is adjusted by expanding the dimension operation and repeating the operation to align with the dimension of the input data;
[0034] The position embedding includes:
[0035] First, a trainable tensor generated by trainable parameters;
[0036] Then, it is expanded to the same shape as the input data through a repeat operation.
[0037] Furthermore, in S4, the Transformer encoder includes a multi-head self-attention mechanism and a feedforward network;
[0038] The multi-head self-attention mechanism is used to capture the relationship between different positions in the sequence. It splits the query, key, and value into multiple subspaces for parallel calculation. The calculation results of each head are spliced to obtain the final multi-head self-attention output.
[0039] The feedforward neural network includes: a linear transformation layer, an activation layer, a layer normalization, a linear transformation layer, and a residual connection. The input is the output after the multi-head self-attention layer. It first undergoes a layer normalization and then enters the linear transformation layer and then enters the activation layer, followed by another linear transformation layer, and finally a residual connection is performed with the original input.
[0040] Furthermore, in said S4, the multiple scale upsampling physiological prior guidance module includes an upsampling resolution increase module and a brain region contribution guided attention module;
[0041] The upsampling and resolution-enhancing module is used to convert the original input three-dimensional data into four-dimensional data through transposition and reshaping operations, mapping the sampling points to spatial dimensions; performing deeper feature extraction on the four-dimensional data through module 2 of the CNN branch; and performing upsampling using bilinear interpolation to increase the resolution of the feature map to match the spatial dimensions of the high-resolution CNN feature map.
[0042] The attention module for guiding brain contribution is used to guide brain contribution of the input after upsampling and increasing the resolution; and comprises the following steps:
[0043] a. Construct a brain region channel correlation matrix based on anatomical partitioning; transform biophysical constraints into a learnable 64×67 brain region attention weight matrix through FreeSurfer source space discretization and BEM head model lead field calculation;
[0044] b. Register the traced brain region contributions to the neural network channels as a brain region attention weight matrix of shape (C, R), where C is the number of channels and R is the number of brain regions. Use the einsum computational tool to perform a tensor product of the input feature x and the brain region attention weight matrix to generate a spatial attention map with brain region traceability priors. The input feature x is a high-resolution feature after upsampling.
[0045] c. Initialize the learnable regional attention parameters, perform weighted adjustment on the spatial attention map, activate it through the Sigmoid function, normalize it to the interval [0,1], and perform dynamic calibration; multiply the calibrated spatial attention map by the brain region weights through reverse einsum to reconstruct it into a shape aligned with the input feature x channel; use channel-aligned 1×1 convolution to adjust the feature dimension to ensure channel compatibility with the input feature x, and then perform residual fusion with the input feature x to retain the original information while enhancing the feature representation guided by brain region attention.
[0046] Furthermore, in S4, the convolution fusion module is used to fuse the convolution layer output of module 2 of the CNN branch and the output of different layers of the Transformer branch through jump connections, and superimpose the residual output of module 2 of the CNN branch after changing the number of channels through the 1×1 convolution layer.
[0047] An EEG classification system based on physiological prior guidance and cross-scale information fusion adopts the above EEG classification method, including:
[0048] A sound stimulation data acquisition module is used to acquire sound stimulation data and perform matching processing on the sound stimulation data to make them consistent in physical characteristics;
[0049] The sound stimulation EEG data acquisition module is used to collect the sound stimulation EEG data and pre-process the collected EEG data;
[0050] The EEG data feature extraction module is used to extract features from the EEG data of the training set through a parallel structure of CNN and Transformer;
[0051] The classifier module is used to classify the extracted features.
[0052] An electronic device uses the above-mentioned EEG classification method based on physiological prior guidance and cross-scale information fusion to realize EEG data classification.
[0053] A computer storage medium stores at least one program instruction, which is loaded and executed by a processor to implement the above-mentioned EEG classification method based on physiological prior guidance and cross-scale information fusion.
[0054] The beneficial effects of the present invention are:
[0055] 1. This paper is based on a joint representation method of time step embedding (Temporal Embedding) and channel embedding (Channel Embedding). It maps the time step index and channel signal to a high-dimensional vector space respectively through a learnable embedding layer, and fuses position encoding to generate spatiotemporal dynamic features, thereby accurately capturing the cross-channel dependency and temporal evolution law of EEG signals.
[0056] 2. The present invention extracts local time-frequency features and global context-dependent features respectively through a parallel network architecture (CNN-Transformer), and designs a cross-scale dynamic fusion mechanism guided by brain region contribution. The cross-scale dynamic fusion network based on brain source guidance (Brain Physiology-Prior Guided Multiscale Transformer-CNN FusionNetwork, abbreviated as BPP-TCFNet) decouples the contribution weights of multiple brain regions to channel signals (such as the temporal auditory cortex) through inverse mapping, quantifies the functional connection strength in the spatial dimension, and constructs cross-brain region collaborative features with biological interpretability. The weight of the traceable brain region is used as the attention weight prior for feature fusion to achieve feature characterization from "channel signal-brain region activity-cross-brain region collaboration". Through this method, we can better understand how the brain processes and responds to various stimuli, and provide new perspectives and tools for research in fields such as brain-computer interface, neurorehabilitation and cognitive science.
[0057] 3. The present invention cleverly leverages the human brain's bioinformatic characteristics in response to sound stimuli, particularly its advantages in long-term and short-term memory and context-dependent properties, to better identify the semantic features of sound. This effectively overcomes the limitations of existing statistical and machine learning methods in EEG signal classification, particularly the signal-to-noise ratio issues when processing complex auditory stimuli. The present invention not only improves the accuracy of classification and recognition, but also, due to the high temporal resolution and small data size of EEG signals, it can demonstrate certain real-time processing capabilities in future applications, showing broad application prospects in sound target recognition. The method of the present invention is highly simple to operate and has good versatility, making it suitable for a wide range of auditory application scenarios. In addition to underwater applications, models based on the method of the present invention are also adaptable to sound signal classification and recognition tasks in other complex environments, such as urban noise recognition, medical sound monitoring, and natural environment sound analysis. By cleverly combining EEG signal processing methods, confidence is increased. Furthermore, the method of the present invention is highly versatile for processing brain signals with similar characteristics and is not limited to auditory EEG classification tasks, but can be applied to the classification and recognition of neural signals for other tasks. For example, EEG signals can be used to identify different cognitive states, emotional changes, and brain-computer interface systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0059] Figure 1 Schematic diagram of the system structure of an embodiment of the present invention in auditory stimulation EEG classification.
[0060] Figure 2 This is a model structure diagram of each branch module in the embodiment of the present invention; (1) is module one of the CNN branch, (2) is module two of the CNN branch, (3) is the Transformer encoder, (4) is the REU Layer, (5) is the BPD Layer, and (6) is the convolutional fusion module.
[0061] Figure 3 It is a structural diagram of the overall network model of an embodiment of the present invention.
[0062] Figure 4 This is a test set evaluation and single-scale comparison of an ablation experiment in auditory stimulation EEG classification in an embodiment of the present invention.
[0063] Figure 5 This is a test set evaluation and single-level comparison of an ablation experiment in auditory stimulation EEG classification in an embodiment of the present invention.
[0064] Figure 6 This is a test set evaluation of an ablation experiment in auditory stimulation EEG classification in an embodiment of the present invention compared with that without adding physiological priors.
[0065] Figure 7 This is a comparison of the classification accuracy of the embodiment of the present invention in auditory stimulation EEG classification with other deep learning methods. DETAILED DESCRIPTION
[0066] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0067] The embodiment of the present invention provides an EEG classification method based on physiological prior guidance and cross-scale information fusion, such as Figure 1 As shown, follow these steps:
[0068] S1, obtaining a sound stimulation data set, and performing matching processing on the sound stimulation data to make them consistent in physical characteristics.
[0069] Get stimulus data:
[0070] We collected and screened datasets of complex environmental semantic auditory stimuli suitable for the present invention. First, we extensively reviewed relevant literature, including scientific papers, technical reports, and patents. We then performed a systematic search using databases and online resource platforms such as IEEE Xplore, Google Dataset Search, and Kaggle to collect datasets related to complex environmental semantic auditory stimuli, such as ESC-50, Shipsear, DeepShips, and UrbanSound8.
[0071] In some data sets, the coherence of the sound is low, which has a negative impact on the collection of sound-stimulated EEG. In order to ensure the scientific nature and effectiveness of the data set, the embodiment of the present invention further performs a quality assessment on the candidate data set, using a simple envelope analysis method to analyze the amplitude envelope of the signal and detect the active area of the signal. If there are too few active parts of the signal within 1s, it is judged to have low coherence; in this way, suitable samples are screened out from the data set as the stimulus sample set. The embodiment of the present invention also pays special attention to whether the data set contains rich underwater acoustic signal samples and whether it can represent the sound characteristics in a real underwater environment.
[0072] Stimulus matching: The purpose of this step is to ensure that the sound stimuli used in the experiment are consistent in physical properties so that the subjects can focus on the category of the sound source stimulus. To achieve this goal, the sound stimuli were matched in the following aspects:
[0073] S11, Duration Matching: Ensure that all sound stimuli have the same duration to avoid discrimination bias caused by varying durations. To this end, assume the duration of the signal is T, and adjust the length of the stimulus to keep the different stimulus signals consistent in duration. This process is performed by interpolation or signal clipping. The duration T is expressed as: T = t2 - t1; where t1 and t2 are the start and end times of the sound stimulus, respectively.
[0074] S12, Root Mean Square (RMS) Power Matching: Root mean square power is an important indicator for measuring sound intensity. By adjusting the RMS power of sound stimuli to keep it consistent, the influence of power differences of different stimulus signals on the judgment of the subjects is eliminated, and the root mean square (RMS) power of the stimulus is standardized. In this way, the RMS of each signal will be adjusted to RMS avg , to achieve power standardization.
[0075] S13, Temporal Envelope Matching: The temporal envelope describes the amplitude variation trend of the sound signal. By matching the temporal envelope, the dynamic characteristics of the sound stimulus are kept consistent. In temporal envelope matching, the instantaneous envelope of the signal is extracted using the Hilbert Transform.
[0076] S14, Time Contour Matching of Harmonic Structure: To further process the sound stimulus, the time contour of the harmonic structure is matched. This involves analyzing and adjusting the harmonic components of the sound signal. First, the harmonic part of the audio needs to be extracted and then further y-spectrum is used to match the harmonic part. harmonic (t) Perform feature extraction.
[0077] The specific steps are as follows: Extract the harmonic part and use the harmonic / percussion separation (HPS) method to separate the audio signal into the harmonic part and the percussion part. This method can extract the periodic harmonic components of the audio signal and remove the noise and non-periodic percussion components. The formula is as follows:
[0078] y harmonic (t) = HPS(y(t))
[0079] Where y(t) represents the original audio signal, y harmonic (t) represents the harmonic components extracted from the original audio signal, removing the non-periodic percussion sounds and noise, and what remains is a periodic signal with a harmonic structure.
[0080] Once the harmonic part is obtained, the Mel spectrum is calculated. The Mel spectrum is used to capture the spectral characteristics of the audio and thus characterize the harmonic structure of the audio. The calculation formula of the Mel spectrum is as follows:
[0081] S harmonic (f,t)=MelSpectrogram(y harmonic (t))
[0082] Among them, S harmonic (f, t) is the Mel spectrum at frequency f and time t. The MelSpectrogram() function calculates the spectrum of the audio signal on the Mel scale, which is expressed as the intensity distribution of the audio signal at different times and Mel frequencies.
[0083] In order to match the harmonic structure between two signals, their Mel spectra are compared. This process can be accomplished through a variety of similarity measurement methods, such as Euclidean distance, dynamic time warping (DTW), and cross-correlation. The Euclidean distance mainly used in the embodiment of the present invention is a common method for measuring the difference between two vectors. For the Mel spectra of two signals, the target signal and the signal to be matched, the Euclidean distance calculation formula is:
[0084]
[0085] Among them, S target (n, m) and S match (n, m) are the values of the target signal and the signal to be matched at the nth Mel frequency and the mth frame time point on the Mel spectrum respectively. euclidean The Euclidean distance between the Mel spectrum of the target signal and the Mel spectrum of the signal to be matched is the square root of the sum of the differences between the two spectra at all Mel frequencies and time frames.
[0086] S2, EEG data acquisition.
[0087] S21 paradigm design;
[0088] Experimental paradigms based on neuroscience can provide a reliable experimental framework. First, we identified and understood these classic paradigms by extensively reading relevant literature, including scientific papers and technical reports. Then, we sought out work similar to the research objective of the present invention (EEG classification of auditory target stimuli) in order to learn from and adapt their experimental designs.
[0089] While paradigm design utilizes classic task paradigms such as block, event, P300, and RSAP as the main framework, it also focuses on adjusting several parameters, such as stimulus duration, a key factor influencing subject responses. Experimental conditions are controlled by adjusting stimulus duration to ensure subjects have sufficient time to process auditory information while avoiding excessive fatigue. The attention cross is a commonly used visual cue used to direct subjects' attention. During experiments, the timing and duration of the attention cross are adjusted to control the subject's focus. Feedback mechanisms are crucial components of experimental design, influencing subject behavior and the learning process. Designing feedback mechanisms provides subjects with immediate feedback on their performance, thereby enhancing learning outcomes. By fine-tuning these parameters, a range of experimental conditions can be designed to explore the impact of different factors on auditory stimulus EEG classification tasks, ultimately identifying the block paradigm as the optimal paradigm for the task and adopting it as the formal experimental paradigm.
[0090] S22, collecting EEG data of sound stimulation;
[0091] Based on the previous steps, a formal experiment was conducted under the block paradigm using precisely matched stimuli to verify whether the subjects could effectively distinguish sound categories under short-term stimulation of complex environmental semantic sounds and to train them in the network.
[0092] Experimental design description:
[0093] The formal experiment adopted a block design, which divides the experimental conditions into blocks, with each block containing the same type of task or stimulus, so as to facilitate the comparison of differences in brain activity under different conditions. In the experiment, the stimulation duration was set to 1 second to simulate sound recognition under short-term stimulation. After each stimulation, a 1-second attention cross rest period was set to ensure that the subjects could get the necessary rest while concentrating their attention. No feedback mechanism was used in the experiment. Instead, guiding words were used to induce the subjects to distinguish sounds, reduce external interference, and focus on internal cognitive processes. The experiment included 5 sound categories, 80 to 100 trials for each category, and a total of 400 to 500 trials to ensure the statistical power of the results. In order to avoid stimulating the subjects with excessive volume, the volume of each signal was adjusted to ensure the fairness of the experiment and the comfort of the subjects.
[0094] Equipment and experimental environment:
[0095] The experiment used a 64-channel Biosemi ActiveTwo EEG device with a sampling rate of 1024Hz. Its 64 electrodes, arranged according to the international 10-20 system, fully cover the head, capturing electrical activity across various brain regions. Participants wore headphones to control sound input and minimize external noise interference. The experiment was conducted in a quiet laboratory to minimize the impact of ambient noise on EEG signals.
[0096] Experimental operation instructions:
[0097] During the experiment, participants must apply electroencephalogram (EEG) cream and wear the EEG device. The device must be comfortable to wear, ensuring that the subject does not experience discomfort that could affect the experiment. Participants must maintain a good mental state and get adequate rest before the experiment to prevent fatigue from affecting the results. Experimental instructions include choosing a comfortable posture, controlling eye movements, avoiding deep breathing, coughing, swallowing, and head and body movements. Upon hearing a sound, participants must carefully and silently identify the sound. No movement feedback is required during this process; all activity occurs in the brain.
[0098] S23, data preprocessing;
[0099] The EEG classification data obtained from the auditory stimulation of the subjects were preprocessed. A series of standardized data preprocessing methods were used for the collected EEG data to improve the signal quality and accuracy of subsequent analysis. Specific preprocessing operations include: downsampling, bandpass filtering, baseline correction, interpolation, and event-based data segmentation. Each preprocessing method aims to eliminate noise interference, remove signal offsets, and standardize the signal, thereby ensuring that in the subsequent signal analysis and processing process, the neural activity signals related to the experimental task can be more accurately extracted. The following is a detailed description of each preprocessing step, and the relevant mathematical expressions and implementation reasons are given.
[0100] (1) Downsampling: Downsampling is the process of reducing the sampling rate of the original EEG data to reduce the amount of data and shorten processing time while retaining important EEG activity information. In this embodiment of the present invention, downsampling is performed to 256 Hz.
[0101] (2) Filtering: In order to remove noise outside the frequency range, the signal is usually filtered. EEG signals usually contain low-frequency (such as motion artifacts) and high-frequency (such as myoelectric interference) noise. By filtering, these noises are eliminated and the frequency range related to neural activity in the signal is retained. In an embodiment of the present invention, a bandpass filter of 0.5Hz-95Hz is used. The filter design is implemented by a Butterworth filter, and its transfer function H(f) can be expressed by the following formula:
[0102]
[0103] Where f is the frequency of the signal, f c is the cutoff frequency of the filter, n is the order of the filter, and 4 is selected in the embodiment.
[0104] A 50Hz notch filter is also applied. In many areas, the power system frequency is 50Hz, so EEG signals are often interfered with by 50Hz frequency components. The notch filter can accurately remove this frequency component from the signal. The transfer function H(f) of the notch filter can be expressed as:
[0105]
[0106] (3) Baseline correction: The effect of baseline drift is eliminated by removing the mean of the signal or the average value within a certain time period, such as the average value of the 200ms before the event stimulus, and subtracting this mean from the entire signal.
[0107] (4) Interpolation guide:
[0108] During and after the experiment, the EEG recordings were observed. For bad electrodes (i.e., poor contact or detached electrodes), the Cubic Spline Interpolation on a Spherical Model was used to interpolate the bad electrode signals. This method accurately estimates and supplements the bad electrode signals by utilizing the signal information of the neighboring normal electrodes. The Cubic Spline Interpolation on a Spherical Model can be implemented by the following steps: First, define the bad electrode C bad and its adjacent electrode set C neighbours , then the spatial position and signal value of the adjacent electrodes are used to perform spherical cubic spline interpolation. The interpolation formula is:
[0109]
[0110] Among them, x c (t) is the signal of the neighboring electrode c. ω(c) is the interpolation weight, which is usually calculated based on the distance between electrodes. The weight is inversely proportional to the distance. neighbours It is C bad Proximity electrode set.
[0111] (5) Complete data segmentation by event: Divide the data into multiple time periods (i.e. epochs) based on the event markers in the experiment.
[0112] S3, dataset partitioning and hyperparameter design;
[0113] Dataset partitioning: The embodiment of the present invention adopts a data set partitioning method to divide the entire data set into a test set and a training set, where the test set accounts for 10% of the total data set and the training set accounts for 90%. This division ratio is determined based on experiments and theoretical analysis, and is intended to ensure the generalization ability of the model on unseen data while retaining a sufficient amount of data for effective training. The training set further adopts a five-fold cross-validation method to avoid contingency in the training process and enhance the stability and reliability of the model. Cross-validation is a statistical method used to evaluate and improve the predictive performance of a model by dividing the data set into multiple small sets (folds), each of which is used as a validation set in turn, and the rest are used as training sets. Specifically, the data set is randomly divided into 5 subsets, 1 of which is used as a validation set, and the remaining 4 subsets are used as training sets. This cycle is repeated 5 times, using a different subset as a validation set each time, and the rest as training sets. The final result is obtained by averaging. The cross-validation formula is as follows:
[0114]
[0115] Where K=5 means five-fold cross validation, score i is the model performance indicator (such as accuracy) obtained in the i-th validation.
[0116] Hyperparameter design: In terms of hyperparameter design, this implementation case selected appropriate hyperparameter configurations to optimize the model training process, including the selection of hyperparameters such as batch size, learning rate, and optimizer. The specific parameter settings are as follows:
[0117] (1) Batch Size: The batch size is set to 32, which is determined based on the memory capacity of the graphics card (such as NVIDIA GeForce RTX 3090) and the complexity of the model to ensure that each iteration can effectively utilize the graphics card resources while avoiding memory overflow.
[0118] (2) Learning Rate: The learning rate is set to 0.001, which is a commonly used initial learning rate value that can ensure rapid convergence of the model in the early stages of training.
[0119] (3) Learning rate scheduler: To further improve model training results, a learning rate scheduler is used. The learning rate scheduler adjusts the learning rate based on the training progress, reducing the learning rate after every set number of n epochs. This strategy allows for rapid convergence in the early stages of training, while allowing for detailed adjustments in the later stages to improve model accuracy.
[0120] (4) Loss function: This implementation uses the cross-entropy loss function to measure the difference between the model's prediction and the true label. For multi-class classification problems, the cross-entropy loss function It can be expressed as:
[0121]
[0122] Among them, C is the number of categories, y i is the true label of the sample, p i The probability of the class predicted by the model.
[0123] (5) Optimizer: This implementation uses the Adam optimizer (Adaptive Moment Estimation) to optimize model parameters. The Adam optimizer combines gradient descent and momentum methods and has the feature of adaptive learning rate.
[0124] S4, feature extraction through the parallel structure of CNN and Transformer;
[0125] Raw EEG signal data (raw_data) is processed through two paths: a CNN and a Transformer. The Transformer primarily captures global features and can handle long-term dependencies and global context in EEG signals. The CNN, on the other hand, extracts local features and is suitable for capturing details and local variations in EEG signals. To improve the performance and stability of the model, a network guidance layer was designed to coordinate and guide the cross-scale fusion of global and local features.
[0126] (1) CNN branch;
[0127] The CNN branch structure consists of multiple modules, in which the batch normalization layer and the activation layer ReLu are connected after the convolution layer. The specific design is as follows:
[0128] The main purpose of module 1 is to control the complexity of the network through bottleneck blocks, use depthwise separable convolution to reduce the amount of computation and improve the efficiency of feature extraction, and use residual connections to improve the training stability of the network and avoid the gradient vanishing problem. First, a 1×1 convolution layer changes the output channel, then a depthwise separable convolution layer keeps the channel unchanged, then a 1×1 convolution layer changes the channel, and finally a residual connection is made to the input after a two-dimensional convolution, such as Figure 2 (1), Module 1 contains the following important parts:
[0129] Two 1×1 convolutional layers: Used to adjust the number of channels and reduce the amount of computation. By using 1×1 convolution kernels, the network can efficiently transfer information between different channels while controlling the size of the feature map.
[0130] The calculation formula for 1×1 convolution is:
[0131]
[0132] Among them, Y Cout is the Cth output feature map out The value of the channel at position (h,w), is the cth input feature map in The value of the channel at position (h,w), is the weight of the convolution kernel, indicating the cth out Output channel and c in The weight between input channels, b Cout is the output channel c out Bias.
[0133] The essence of 1×1 convolution is a linear transformation across channels, which allows the number of channels of the feature map to be changed without changing the spatial size. In this way, information can be transferred and mixed without introducing a large amount of computation.
[0134] Depthwise Separable Convolution (DWC): Using a 3×3 convolution kernel and setting groups = the number of channels, this implements depthwise separable convolution. This operation breaks down traditional convolution into two steps: first, channel-by-channel convolution (processing each channel separately), then point-by-point convolution (combining across channels). This approach significantly reduces computation while maintaining efficient feature extraction.
[0135] The complete calculation process of depth-wise separable convolution combines channel-by-channel convolution and point-by-point convolution. The calculation formula of depth-wise separable convolution is:
[0136]
[0137] in, is the final output of the depthwise separable convolution, C in is the number of input channels, C out is the number of output channels, K dw,c is the depthwise convolution kernel, It is a point-by-point convolution kernel. The advantages of depth-wise separable convolution are reduced computational effort, fewer parameters, and efficient feature extraction.
[0138] Residual connection: The input x is passed through a two-dimensional convolution layer with a convolution kernel of 1 to change the number of channels and then added to the processed feature map; the processed feature map is Figure 2The 1×1 convolution in (1) is followed by DWC and then the output of the 1×1 convolution. The purpose is to align the feature dimensions so that they can be connected. The purpose of the residual connection is to make the network easier to optimize by skipping certain layers, especially in deep networks, to avoid the gradient vanishing problem, thereby accelerating training and improving performance.
[0139] Initialization: To improve the stability of the model, the weights of the last batch normalization layer in module 1 are initialized to 1. This means that in the initial stage of the network, minimizing the impact of the BN layer on the input distribution in the initial stage helps make the residual block more like an identity mapping, thereby making the training process smoother.
[0140] Initialized as:
[0141]
[0142] Where BN(x) represents batch layer normalization, μ and σ are the mean and variance of the batch, γ and β are the scaling and bias parameters. When initialized, set γ = 1 and β = 0, and ∈ is a small constant set to 10 -5 .
[0143] Module 2: The function of module 2 is to increase the depth of the network and improve the ability to express the characteristics of EEG signals. Module 2 is a simplified version of module 1, which retains the core idea of residual connection but removes the depth-wise separable convolution. It uses standard convolution operation, with a convolution kernel size of 3×3, groups=1, and controls the entire input and output channels to be the same. Direct residual connection: without additional processing of the convolution layer, the input of module 2 is directly added to the output of the main module (that is, directly added to the output after conventional convolution), such as Figure 2 In (2), the design of Module 2 focuses more on increasing the depth of the network layers. Module 2 expands the network depth by adding multiple convolutional layers. Each convolutional layer extracts more complex features as it gradually processes the input signal. By stacking convolutional layers, the network can learn high-level features of the signal at multiple scales, improving its representational capabilities. Increasing the number of network layers can gradually extract higher-dimensional and more abstract features, allowing the network to capture more details of the EEG signal and improving the model's understanding of signal complexity.
[0144] Because EEG signals have high temporal resolution and complex spatiotemporal distribution characteristics, and are subject to noise interference, direct large-scale convolution calculations often result in excessive computational effort. Complex network structures can lead to excessive computation, which in turn affects performance and reduces signal extraction efficiency. Module 1 of this embodiment of the present invention reduces computational effort and improves feature extraction efficiency through bottleneck blocks and depthwise separable convolution (DWC), thereby reducing the demand for computing resources while maintaining information extraction capabilities.
[0145] The brain processes EEG signals not in a linear manner but in a hierarchical and progressively deeper manner. A single feature extraction method cannot capture the full picture of the signal. Module 2 of this embodiment of the present invention uses stacked convolutional layers and skip connections to gradually delve deeper into the complex features of the signal. This is particularly true for the temporal characteristics of EEG signals. Deep networks are able to better capture temporal correlations and cross-channel dependencies in EEG signals, achieving higher accuracy.
[0146] (2) Transformer branch
[0147] The Transformer branch of the implementation case of the present invention adopts the traditional Transformer encoder structure. Unlike the traditional Transformer model that directly uses fixed position embedding (Position Embedding), the embodiment of the present invention introduces data embedding (Data Embedding), that is, channel embedding (Channel Embedding) and time step embedding (Temporal Embedding) are added to the Transformer encoder, so as to better adapt to the processing requirements of EEG signals.
[0148] The following is a detailed description of each embedding process along with the mathematical representation.
[0149] Channel embedding: At the input of the Transformer, channel embedding is first performed. This process maps the signal of each channel to a higher-dimensional space through an embedding layer. Suppose the dimension of the input data x is [batch_size, 1, C, T], where C = 64 represents the number of channels and T is the number of sampling points. The channel embedding process can be expressed as:
[0150] C embed =ChannelEmbedding(x)=E C ·x
[0151] in, is the channel embedding matrix, C is the number of channels, d is the embedding dimension, is the original input data. A tensor representing the number of channels 64*embedding dimension size, Represents the original input data, a tensor of size (B, 1, C, T).
[0152] C embed It is the channel embedding feature, and its physical meaning is to map the signal features of each channel into a higher-dimensional representation space; ChannelEmbedding(x) is the process of mapping the signal of each channel in the input signal x to a high-dimensional space through the embedding matrix.
[0153] The shape of the embedding matrix is then expanded from [C, d] to [1, C, d, 1] using the unsqueeze operation and expanded to the same dimensions as the original input using the repeat operation (that is, repeatedly expanding the original tensor along the specified dimension to fit a larger input shape or align other tensors):
[0154] C embed =E C ·x(repeat along batch size and time steps)
[0155] Repeat along batch size and time steps so that the embedding matrix is expanded in batch size and time steps to match the dimensions of the original input data.
[0156] Finally, a permute operation is used to adjust the dimension order of the embedding matrix so that it can be added correctly to the input data.
[0157] Time-step embedding: Time-step embedding first embeds each time step through the nn.Embedding layer (embedding layer). This is a layer used in PyTorch to implement embeddings, and is commonly used when discrete word indices (such as each word in a vocabulary) need to be mapped to a dense, low-dimensional vector space. In time series, the nn.Embedding layer can also be used to map each time step (i.e., each discrete time index) into an embedding space, thereby providing a vector representation with a fixed dimension for each time step. These vectors are learned during training, allowing the model to capture the relationship between different time steps through the embedding space.
[0158] After that, the dimensions are adjusted through unsqueeze and repeat operations to align with the dimensions of the input data. The specific steps are as follows:
[0159] T embed =E T ·t
[0160] The time-step embedding is reshaped to [1, 1, T, d] by the unsqueeze operation and then expanded to the same shape as the original input [batch_size, C, T, d] using the repeat operation.
[0161] T embed =E T·t(repeat along batch size and channels)
[0162] T embed represents the time embedding feature, E T represents the temporal embedding matrix.
[0163] Position Embedding:
[0164] Positional embedding is a trainable tensor generated by nn.Parameter (trainable parameter). In PyTorch, nn.Parameter is a special tensor type used to identify trainable parameters in the model. It is automatically added to the model's parameter list and optimized during training. To align it with the input data, a repeat operation is used to replicate it in the batch_size (batch size) and channel number dimensions. The mathematical representation is as follows:
[0165] P embed =E P
[0166] Then, through the repeat operation, it is expanded to the same shape as the input data, that is, expanded to [batch_size, C, T, d].
[0167] P embed Represents the position embedding feature, E P Represents the position embedding matrix. The “position embedding” in the embodiment of the present invention is the same as the traditional Transformer model directly using fixed position embedding.
[0168] Addition with embedded operation:
[0169] Finally, all three embeddings (channel embedding, timestep embedding, position embedding) are added to the original input x. In order to perform the addition operation, it is necessary to ensure that the shapes of all tensors are consistent. This has been ensured by operations such as unsqueeze, repeat, and permute, so element-wise addition can be performed:
[0170] x embed =x+C embed +T embed +P embed
[0171] Here, x is the original input data and all embeddings have the shape [batch_size, C, T, d] so that they can be directly added together.
[0172] T embed With Tembed (t) Substantially identical.
[0173] The Transformer encoder consists of a multi-head self-attention mechanism and a feedforward network. The input of the Transformer encoder is the sum of the original input x and the output after data embedding, such as Figure 2 As shown in (3), the specific structure is described as follows:
[0174] After the previous step (channel embedding, time step embedding, position embedding) is added to the original input x, some operations are performed, which can be based on Figure 3 You can see that data_embedding is Flattened (dimensional expansion) and Avgpooled (average pooling) before Transformer_encoder and then cls_token is added to the second dimension.
[0175] The multi-head self-attention mechanism is the core part of the Transformer encoder, which is mainly used to capture the relationship between different positions in the sequence. First, the query, key, and value vectors are calculated through three different linear transformations:
[0176]
[0177] in, is the trainable weight matrix, d k is the dimension of each query / key vector.
[0178] W is a weight tensor that changes during training. Q is the query weight matrix, which is used to transform the input vector Converted to query vector Q; W K is the key weight matrix used to transform the input vector Convert to key vector K; W V is the value weight matrix used to convert the input vector Convert to a value vector V.
[0179] Input Matrix After multiplying with the weight matrix, it is understood as the current The "query vector" is used to determine the most interesting part of the current signal in the global context. K can be understood as The "key vector" is used to provide information in the global context. V represents the actual content that should be obtained from each time point or each channel when the query matches the key. It can be understood as The "value vector" that provides the actual signal values. It is obtained from EEG, so Q, K, and V indirectly represent the characteristic information of EEG.
[0180] The self-attention score is calculated by the dot product between the query and the key and normalized using the softmax function:
[0181]
[0182] in, It is a normalization factor used to prevent the dot product from being too large, causing the gradient to disappear or explode.
[0183] Attention(Q, K, V) represents the attention score calculated from self-attention.
[0184] In order to capture the various relationships in the input, Transformer uses a multi-head mechanism, which splits the query, key, and value into multiple subspaces for parallel computation. The computation results of each head can be concatenated to obtain the final multi-head self-attention output:
[0185] MultiHead(Q,K,V)=Concat(head1,head2,...,head h )W O
[0186] The output of each header is:
[0187]
[0188] It is the output weight matrix, which is used to concatenate the outputs of the multi-head attention mechanism and perform a linear transformation to reduce the dimension and fuse the information of different heads.
[0189] They represent the Q, K, and V weight matrices of the i-th head respectively.
[0190] Feedforward neural network: It mainly consists of a linear transformation layer, an activation layer, a layer normalization, a linear transformation layer, and a residual connection. The input is the output after the multi-head self-attention layer. It first undergoes a layer normalization and then enters the linear transformation layer and then enters the activation layer. After that, it enters another linear transformation layer and finally performs a residual connection with the original input.
[0191] (3) Cross-scale physiological prior network guidance layer:
[0192] The cross-scale physiological prior network guidance layer is mainly composed of two key sub-modules: a multi-level upsampling physiological prior guidance module BioPrior-Driven Upsampling Layer (BPDU Layer) and a convolutional fusion module. The upsampling physiological prior guidance module includes an upsampling resolution increase module and an attention module guided by brain region contribution. Through the multi-level upsampling physiological prior guidance module and the convolutional fusion module, the global features extracted by the Transformer branch and the local features extracted by the CNN branch are effectively guided and cross-scaled, thereby improving the network's feature processing capabilities.
[0193] Upsampling to increase resolution module (REU Layer): First, the original input three-dimensional data (size is BNC, where B is the batch size, N is the number of channels, and C is the number of sampling points) is converted into four-dimensional data through transpose and reshape operations so that the spatial dimension can be expanded in subsequent processing. On this basis, deeper feature extraction is performed through module 2 of the CNN branch; next, the upsampling operation is used to increase the resolution of the feature map. By upsampling the feature map, its spatial dimension is increased so that it can be better aligned with the high-resolution CNN feature map during subsequent fusion. Figure 2 As shown in (4) in the figure, the specific operation uses bilinear interpolation to complete feature upsampling. Bilinear interpolation uses the four nearest pixels in the input image for interpolation. Given a target output position, the interpolation in two directions (horizontally and vertically) in the input image is calculated.
[0194] Brain contribution guided attention module (BPD Layer): guides brain contribution to the input after upsampling and increasing resolution, such as Figure 2As shown in (5). First, the brain region channel correlation matrix is constructed; the first step is to perform channel mapping calibration on the original EEG signal, and use the standard electrode layout template to map the original acquisition channel to the international 10-20 system naming specification to establish the electrode space coordinate model; the second step is source space modeling: the fsaverage standard brain source space is constructed by the FreeSurfer tool, and the cortical discretization method of oct4 spacing is used to generate a set of source points containing the left and right hemispheres, and the source points correspond to the equivalent current dipole positions of the cerebral cortex; the third step is head model construction: the boundary element method (BEM) is used to establish a three-layer head model, including the conductivity parameter configuration of the scalp, skull and cerebrospinal fluid, and the potential distribution is obtained by numerical calculation. The fourth step is lead field calculation: combining the electrode coordinates, source point positions and head model parameters, a forward model is established and a lead matrix is generated. The dimension of the lead matrix is (C×Ns), where C is the number of EEG channels and Ns is the total number of source points. The fifth step is brain region source point mapping: reading the anatomical partition annotation of the fsaverage standard brain, mapping each brain region vertex to the source space index, and establishing a corresponding relationship between the brain region name and the lead matrix column number. The sixth step is contribution matrix generation: for each brain region, extract all source point column vectors corresponding to the brain region in the lead matrix, and generate a brain region channel association matrix with a dimension of (C×R) by calculating the arithmetic average of the contribution values of each channel, where R is the total number of brain regions.
[0195] Secondly, the channel contribution of the brain regions obtained by tracing back to the neural network is registered as a brain region attention weight matrix with the shape of (C, R) where C represents the number of channels (64) and R represents the number of brain regions (67). Using the predefined brain region attention weight matrix (64×67), the tensor operation tool einsum based on the Einstein summation convention is used to calculate the tensor product of the input feature x (the feature map after upsampling to increase the resolution) and the brain region attention weight matrix to generate a spatial attention map with brain region tracing back prior (B, R, 16, 64).
[0196] Again, initialize a regional attention parameter (shape is 1, R, 1, 1), adjust the brain area attention weight matrix (spatial attention map), and use the Sigmoid function to activate and normalize it to [0, 1] for dynamic calibration; through the reverse Einstein summation, multiply the calibrated spatial attention map with the brain area weight and reconstruct it into a shape (B, C = 64, 16, 64) that is channel-aligned with the input feature x (feature map after upsampling to increase the resolution), and use channel-aligned convolution (1×1 convolution) to achieve compatible conversion of feature dimensions; then perform residual fusion with the input feature x.
[0197] The upsampling physiological prior guidance module (BPDU Layer) transfers global features to local features and achieves the guidance effect, and uses the convolution fusion module to assist in feature fusion.
[0198] Convolution fusion module: such as Figure 2 As shown in (6), a fusion jump connection is performed after two convolutional layers (one 1×1 and one convolution kernel size of 3×3). In the second convolutional layer, if the input is two signals, the input feature map will be added, and the output of the first convolutional layer will be residually connected with the output after the two convolutional layers. This approach allows the network to learn richer features, especially in the case of multi-branch input, which can promote the network to learn more meaningful features through residual connections. The convolution fusion module fuses the convolution layer output of module 2 of the CNN branch and the output of the upsampling resolution increase module of the Transformer branch at multiple levels, and uses residual connections and two-dimensional convolutions to guide and fuse them.
[0199] The specific process is as follows: Module 2 of the CNN branch first increases the depth of the network block through convolution operations to better capture high-dimensional features, which serves as an input to the convolution fusion module. The output of the Transformer branch, which increases the resolution through upsampling, serves as another input to the convolution fusion module. This fusion operation helps integrate features at different levels, ensuring that the network can effectively utilize multi-layer information from low-level to high-level, such as Figure 2 As shown in (5), the specific fusion operation is as follows:
[0200] Y fusion =Y CNN+trans former +X residual
[0201] Among them, X residual It represents the residual output of module 2 after the 1×1 convolution layer changes the number of channels, Y CNN+transformer Y is the fusion output of the CNN (specifically CNN module 2) and the Transformer branch after upsampling to increase the resolution module output, fusion The fused feature map is obtained by guiding the fusion through branches of different depths, completing the guidance and fusion of multi-level features. The network can effectively process features at different levels and dimensions, improving the representation and generalization capabilities of the model.
[0202] The core task of the convolutional fusion module is to fuse information from different branches across scales, ensuring that the features of the Transformer branch and the CNN branch can be combined to improve the overall performance of the network. The convolutional fusion module further strengthens the fusion of cross-scale signals, allowing the fusion of feature maps of two input signals and using residual connections to improve the network's feature learning ability. Residual connections enable the direct addition of features from different levels, preserving multi-level information. At the same time, the convolutional fusion layer helps integrate features from different branches. By fusing features at different scales, the network can leverage local and global information to comprehensively capture the spatiotemporal characteristics of EEG signals and improve signal classification accuracy.
[0203] In EEG signal processing, global features (such as the overall trend of the EEG signal) and local features (such as local changes in electrical activity) are interrelated. In a cross-scale and multi-level framework, the complementarity of global and local information is utilized to achieve more efficient signal processing. This embodiment of the present invention implements feature fusion of the Transformer branch and the CNN branch through module three.
[0204] (4) Network data flow
[0205] Network model structure and data flow Figure 3 As shown, the specific structure and data flow are as follows:
[0206] In this embodiment, the input data is a four-dimensional tensor with a shape of (32, 1, 64, 256), where 32 is the BatchSize, 64 is the number of channels, and 256 is the time step of each sample. The input feature of the CNN branch is the original data, which is a two-dimensional data that roughly includes both time and space information; the CNN branch first extracts features through a standard two-dimensional convolution layer, a batch normalization layer (Batch Norm) and a ReLU activation function, and then passes through an average pooling layer (AvgPool) to output a shape of (32, 64, 16, 64). This stage mainly realizes the change in the number of channels and the adjustment of the feature map size through two-dimensional convolution operations, and the pooling operation further reduces the spatial dimension. Then, through module 1 of the CNN branch, the output shape is (32, 256, 16, 64). This module adopts a residual connection structure and combines depth-wise separable convolution (DWC) to extract more complex local features. Subsequently, after module 2 of the CNN branch, the output dimension remains unchanged, but the network depth is increased to capture higher-dimensional local features in the EEG signal.
[0207] In the Transformer branch, data embedding is first performed on the input data, specifically combining channel embedding, temporal embedding, and position embedding. This ensures that each input data point encodes not only its spatial position but also its temporal information. The output shape after data embedding is (32, 64, 256, 768), where 768 is the feature dimension. Next, the input passes through a flattening layer and a one-dimensional average pooling layer for spatial integration and dimensionality compression, resulting in an output shape of (32, 64, 768). The output of the Transformer branch is then concatenated with a [CLS] marker and used as the input to the standard Transformer encoder layer. The Transformer encoder, which consists of three layers of multi-head self-attention mechanisms and a feedforward neural network, maintains the same output shape, successfully capturing the global characteristics of the EEG signal.
[0208] The cross-scale physiological prior network guidance layer includes multiple layers of upsampling physiological prior guidance modules and convolutional fusion modules. Through upsampling, convolution, and residual connections, the two types of features can be fused at the same scale while preserving the validity of their information to the greatest extent possible. The output of the Transformer branch (specifically the output of the Transformer encoder) is increased in resolution by an upsampling resolution increase module (REU layer), resulting in an output shape of (32, 64, 16, 64). Meanwhile, the output of the CNN branch (specifically the output of module 2 in the CNN) enters the convolutional fusion module, where the number of channels is adjusted through a 1×1 convolutional residual block and the output is concatenated with the output of the Transformer branch. Further fusion of the convolutional layers completes the guidance and integration of information from the two branches, resulting in an output of (32, 64, 16, 64). This process is repeated to achieve multi-level information fusion.
[0209] S5, classify the extracted features through the classifier module, and train and verify the classifier.
[0210] First, spatial integration and dimensionality compression are performed through 2D average pooling and flattening layers, with an output shape of (32, 64). Then the data is input into the Softmax layer, and the final output shape is (32, num_classes), completing the classification task, where num_classes is the number of categories.
[0211] The embodiment of the present invention uses automatic feature extraction (ie, performed end-to-end without the need for additional manual feature extraction operations), which greatly reduces the inference time.
[0212] The deep learning model is trained using the training set data, and the model parameters are fine-tuned using the validation set to optimize model performance and prevent overfitting.
[0213] Using the training set data, this implementation uses the backpropagation algorithm to optimize the model weights. The backpropagation algorithm is a gradient descent method commonly used in supervised learning. It calculates the gradient of the loss function with respect to the network parameters and updates the model weights along the direction of gradient descent to minimize the loss function. The mathematical expression of this process can be expressed as:
[0214]
[0215] Among them, θ represents the model parameters, η is the learning rate, J(θ) is the loss function, is the gradient of the loss function with respect to the parameter θ. new ,θ old Represent the model parameters before and after the model parameters change.
[0216] During model training, this implementation uses a validation set to evaluate the model and monitor its performance on unseen data. The validation set is a set of data independent of the training set and is used to adjust model hyperparameters, such as the learning rate and regularization coefficient, and to prevent the model from overfitting to the training set. Overfitting occurs when a model performs well on the training data but generalizes poorly to new data. By regularly evaluating model performance on the validation set, overfitting can be detected and avoided.
[0217] Based on performance feedback from the validation set, this implementation adjusts the model's hyperparameters. Hyperparameters are parameters set before the learning process begins, such as the learning rate, batch size, and regularization parameters. Hyperparameter adjustments are based on validation set performance metrics, such as accuracy and loss, to find the optimal hyperparameter combination, thereby improving the model's generalization ability.
[0218] Regarding the classifier, the implementation example uses the Softmax activation function, which converts the original scores (logits) output by the model into a probability distribution. The calculation formula of the Softmax function is as follows:
[0219]
[0220] Among them, z is a vector logits output by the model, z i is the i-th element of z, z j is the jth element of z, so the denominator is the sum of the exponentials of all the elements of the vector. is the predicted probability of the corresponding category, N is the number of categories, and e is the base of the natural logarithm.
[0221] Model Evaluation:
[0222] The performance of the model on the test set is evaluated using the accuracy evaluation metric. This step involves evaluating the performance of the trained model on an independent test set. The test set contains data that the model has not seen during training, so its performance can truly reflect the model's ability to process new data. This metric can be used to evaluate the classification performance of the model. Specifically, given the true label y on the test set, true and predicted label y pred , the calculation formula of the evaluation index is as follows:
[0223] Classification accuracy: Classification accuracy measures the model's prediction accuracy for all samples. The formula is as follows:
[0224]
[0225] Where N is the number of samples, 1(·) is the indicator function, when y true,i =y pred,i , the value is 1; otherwise, it is 0.
[0226] In an embodiment of the present invention, the generalization ability of the evaluation model is achieved by the following method:
[0227] (1) Test set evaluation: Use a pre-partitioned test set (accounting for 10% of the total dataset) to evaluate the performance of the model, ensure that the evaluation results are well representative, and complete ablation experiments to check the performance of some modules.
[0228] (2) Multi-model comparison: In order to fully understand the performance of different model architectures, we will compare the experimental results based on different neural network architectures and conduct module ablation experiments. The experimental results are shown in Tables 1-2 and Figure 4-7 shown.
[0229] Table 1 Evaluation results of BPP-TCFNet on the test set
[0230]
[0231]
[0232] In Table 1, single-scale CNN refers to the classification capability at the scale of local features captured by the CNN branch alone. Single-level cross-scale fusion refers to single-level fusion that combines local and global feature information extraction. Multi-level fusion refers to multiple (≥1) fusions. Adding brain source guidance refers to the entire BPP-TCFNet network. The data in Table 1 shows the accuracy results for the five-category classification.
[0233] The results show that as the complexity of the fusion strategy increases, the performance of the model is significantly improved. Figure 4 It can be found that from a single scale of CNN to a single-level fusion across scales, the variable lies in the change in scale, and the classification accuracy is improved from 78% to 88%; Figure 5 It can be found that from cross-scale single-level fusion to cross-scale multi-level fusion, where the variable lies in the change at the level, the classification accuracy is improved from 88% to 90.5%; Figure 6 It can be seen that after adding physiological prior guidance, the classification accuracy increased from 90.5% to 92.8%. This result shows that the multi-level fusion of cross-scale physiological prior guidance improves the model's classification ability in terms of precision. The scale, level, and physiological prior guidance all effectively improve classification ability, indicating that they further enhance the ability to capture different EEG signal features. This change may be related to the design of cross-scale physiological prior guidance. This layer can simultaneously fuse multiple levels of features extracted from CNN and Transformer at both local and global scales. By adding physiological prior weights that trace back to brain regions, the model can better understand complex EEG signals.
[0234] Table 2 Comparison of multi-model experimental results
[0235] Resnet110 Eegnet Densenet BPP-TCFNet Accuracy 76% 47% 61% 92.8%
[0236] The results showed that, compared with several existing classic neural network models (such as ResNet110, EEGNet, and DenseNet), BPP-TCFNet significantly improved the accuracy by 16.8% compared with ResNet110 and 43.8% compared with EEGNet. These results show that the BPP-TCFNet of the embodiment of the present invention can handle EEG signal classification tasks more effectively, showing its advantages in complex data structures. In particular, by adding the design of physiological prior weights for brain region tracing, BPP-TCFNet can more accurately extract the classifiable features of EEG signals, thereby achieving higher performance in classification tasks.
[0237] In the embodiment of the present invention, the multi-branch parallel input of Transformer and CNN is used, the resolution increase module guides the global features to flow to the local features through upsampling, and the CNN module is used to assist in completing cross-scale feature fusion, and the global features extracted by Transformer and the local features extracted by CNN are fused at multiple levels.
[0238] Traditional neural network architectures usually focus on capturing local features (such as CNN) or global features (such as Transformer). In such networks, the fusion of local features and global features is usually serial, that is, global information and local information are processed separately in one step, and then simply fused. However, this serial fusion cannot fully utilize the advantages of both. The embodiment of the present invention adopts a cross-scale parallel mechanism, that is, the Transformer branch is used to capture the global context features in the brain signal, and the CNN branch is used to extract local detail features in the EEG signal. Through parallel processing and cross-scale fusion, global information and local information can be captured at the same time, and this information can be fused layer by layer at multiple levels, thereby improving the model's ability to process complex signals.
[0239] When the brain recognizes an object, it typically undergoes a gradual process, in which information at different levels is gradually processed. This invention simulates this process through a multi-level mechanism, enabling a network-wide implementation similar to the brain's step-by-step processing of objects. Traditional neural network methods often lack attention to this gradual processing, focusing more on directly extracting global or local features from signals without effectively simulating the brain's hierarchical thinking model.
[0240] By placing the Transformer encoder within a cross-scale processing framework, the present invention enables the model to process features step by step from low-level to high-level features, just like the brain. This design makes the network more flexible and accurate in handling complex sound stimulus classification tasks, and can efficiently extract key features from raw EEG signals.
[0241] EEG signals have extremely high temporal resolution, but their feature distribution is complex and they are susceptible to noise interference. Traditional methods usually fail to fully utilize the biological information contained in EEG signals, such as dynamic changes in time and space. These signal characteristics should be effectively expressed in the design of the model, but in the prior art, the spatiotemporal characteristics of EEG signals are usually only processed indirectly in the feature extraction stage. For example, the CNN branch can capture spatial features, however, the spatiotemporal characteristics are not fully utilized in the input stage of the model. The present invention uses a data embedding mechanism to directly extract spatiotemporal features of the signal before the input encoding layer (Transformer-encoder) layer of the network. Specifically, time embedding and channel embedding are used to enable the network to capture the timing changes and cross-channel dependencies in the EEG signal at the earliest stage. This approach not only improves the accuracy of signal processing, but also by processing spatiotemporal features at the input stage, the model can capture more biological features and optimize the signal understanding and classification effect through adaptive adjustment of trainable parameters during the training process. Experiments have shown that compared with traditional single CNN or Transformer networks, the model of the embodiment of the present invention improves the accuracy of EEG signal classification tasks by about 10%-15%, especially in cross-scale multi-level fusion, which can effectively improve the classification accuracy.
[0242] If the EEG classification method based on physiological prior guidance and cross-scale information fusion described in the embodiment of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the EEG classification method based on physiological prior guidance and cross-scale information fusion described in the embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, ROM, RAM, a magnetic disk or an optical disk.
[0243] The above description is only a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention are included in the scope of protection of the present invention.
Claims
1. A method for EEG classification based on physiological prior guidance and cross-scale information fusion, characterized in that: The following steps are involved: S1, obtaining sound stimulation data, and performing matching processing on the sound stimulation data to make them consistent in physical characteristics; S2, collecting EEG data of sound stimulation and preprocessing the collected EEG data; S3, divide the EEG data into test set and training set, and determine the hyperparameters of model training; S4, feature extraction of EEG data of the training set is performed through the parallel structure of CNN and Transformer; The CNN branch structure is used to extract local features, including: Module 1, wherein the module controls the complexity of the network through bottleneck blocks, reduces the amount of computation and improves the efficiency of feature extraction through depthwise separable convolution, and improves the training stability of the network through residual connections; Module 2, wherein the module 2 increases the depth of the network layer by stacking convolutional layers; The Transformer model is used to capture global features. The Transformer model includes a Transformer encoder. Channel embedding and time step embedding are added to the position embedding in the Transformer encoder. The channel embedding, time step embedding, and position embedding are added to the original input. The global features extracted by the Transformer branch and the local features extracted by the CNN branch are guided and fused through the multi-scale upsampling physiological prior guidance module and the convolutional fusion module; S5, classify the extracted features through the classifier module, and train and verify the classifier.
2. The EEG classification method based on physiological prior guidance and cross-scale information fusion according to claim 1, characterized in that: In said S4, module 1 includes: Two 1×1 convolutional layers are used to adjust the number of channels and reduce the amount of computation; Depthwise separable convolution layer uses a 3×3 convolution kernel and sets groups = number of channels to implement depthwise separable convolution, that is, first performing channel-wise convolution and then point-wise convolution; Residual connection is used to add the original input x through the output of the two-dimensional convolution layer with a convolution kernel of 1 to change the number of channels and the processed feature map; The weights of the batch normalization in the last layer of module one are initialized to 1.
3. The EEG classification method based on physiological prior guidance and cross-scale information fusion according to claim 1, characterized in that: In said S4, module 2 includes: Two 1×1 convolutional layers are used to adjust the number of channels and reduce the amount of computation; A 3×3 convolutional layer with groups set to 1 to ensure that all input and output channels are the same. Residual connection is used to add the input of module 2 directly to the output of the last 1×1 convolutional layer.
4. The EEG classification method based on physiological prior guidance and cross-scale information fusion according to claim 1, characterized in that: In S4, channel embedding includes: First, an embedding layer is used to map the signal of each channel into a higher-dimensional space; Then, the shape of the embedding matrix is expanded from [C, d] to [1, C, d, 1] through the expansion dimension operation, and repeated on the batch size and time step through the repeat operation, so that the embedding matrix is expanded in the batch dimension and time dimension to match the dimension of the original input data; Finally, the dimension order of the embedding matrix is adjusted through the dimension permutation operation so that it can be added to the input data correctly; The time-step embedding includes: First, each time step is embedded through the embedding layer; Then, the dimension is adjusted by expanding the dimension operation and repeating the operation to align with the dimension of the input data; The position embedding includes: First, a trainable tensor generated by trainable parameters; Then, it is expanded to the same shape as the input data through a repeat operation.
5. The EEG classification method based on physiological prior guidance and cross-scale information fusion according to claim 1, characterized in that: In S4, the Transformer encoder includes a multi-head self-attention mechanism and a feedforward network; The multi-head self-attention mechanism is used to capture the relationship between different positions in the sequence. It splits the query, key, and value into multiple subspaces for parallel calculation. The calculation results of each head are spliced to obtain the final multi-head self-attention output. The feedforward neural network includes: a linear transformation layer, an activation layer, a layer normalization, a linear transformation layer, and a residual connection. The input is the output after the multi-head self-attention layer. It first undergoes a layer normalization and then enters the linear transformation layer and then enters the activation layer, followed by another linear transformation layer, and finally a residual connection is performed with the original input.
6. The EEG classification method based on physiological prior guidance and cross-scale information fusion according to claim 1, characterized in that: In said S4, the multiple scale upsampling physiological prior guidance module includes an upsampling resolution increase module and a brain region contribution guided attention module; The upsampling and resolution-enhancing module is used to convert the original input three-dimensional data into four-dimensional data through transposition and reshaping operations, mapping the sampling points to spatial dimensions; performing deeper feature extraction on the four-dimensional data through module 2 of the CNN branch; and performing upsampling using bilinear interpolation to increase the resolution of the feature map to match the spatial dimensions of the high-resolution CNN feature map. The attention module for guiding brain contribution is used to guide brain contribution of the input after upsampling and increasing the resolution; and comprises the following steps: a. Construct a brain region channel correlation matrix based on anatomical partitioning; transform biophysical constraints into a learnable 64×67 brain region attention weight matrix through FreeSurfer source space discretization and BEM head model lead field calculation; b. Register the traced brain region contributions to the neural network channels as a brain region attention weight matrix of shape (C, R), where C is the number of channels and R is the number of brain regions. Use the einsum computational tool to perform a tensor product of the input feature x and the brain region attention weight matrix to generate a spatial attention map with brain region traceability priors. The input feature x is a high-resolution feature after upsampling. c. Initialize the learnable regional attention parameters, perform weighted adjustment on the spatial attention map, activate it through the Sigmoid function, normalize it to the interval [0,1], and perform dynamic calibration; multiply the calibrated spatial attention map by the brain region weights through reverse einsum to reconstruct it into a shape aligned with the input feature x channel; use channel-aligned 1×1 convolution to adjust the feature dimension to ensure channel compatibility with the input feature x, and then perform residual fusion with the input feature x to retain the original information while enhancing the feature representation guided by brain region attention.
7. The EEG classification method based on physiological prior guidance and cross-scale information fusion according to claim 1, characterized in that: In S4, the convolution fusion module is used to fuse the convolution layer output of module 2 of the CNN branch and the output of different layers of the Transformer branch through jump connections, and superimpose the residual output of module 2 of the CNN branch after changing the number of channels through the 1×1 convolution layer.
8. An EEG classification system based on physiological prior guidance and cross-scale information fusion, characterized by: The method for EEG classification based on physiological prior guidance and cross-scale information fusion as claimed in claim 1 includes: A sound stimulation data acquisition module is used to acquire sound stimulation data and perform matching processing on the sound stimulation data to make them consistent in physical characteristics; The sound stimulation EEG data acquisition module is used to collect the sound stimulation EEG data and pre-process the collected EEG data; The EEG data feature extraction module is used to extract features from the EEG data of the training set through a parallel structure of CNN and Transformer; The classifier module is used to classify the extracted features.
9. An electronic device, characterized in that: The EEG data classification is achieved by using the EEG classification method based on physiological prior guidance and cross-scale information fusion as described in any one of claims 1 to 7.
10. A computer storage medium, characterized in that The storage medium stores at least one program instruction, and the at least one program instruction is loaded and executed by the processor to implement the EEG classification method based on physiological prior guidance and cross-scale information fusion as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Electroencephalogram response deep learning classification identification method based on underwater acoustic signal stimulation
CN115470821A
Classification method for electroencephalogram emotion recognition through multi-scale spatial-temporal feature extraction based on CNN and Transform
CN118797496A
Cited By
Asphalt mixture image classification method and system based on convolutional neural network
CN121074491A
Depression type electroencephalogram signal recognition method, storage medium, product and recognition device
CN122163232A