Module adaptive EEG (electroencephalogram) classification method based on multi-view features

By constructing a multi-domain feature extraction network and a modular transfer adaptation strategy, the problems of individual variability and data scarcity in EEG signal classification are solved, achieving higher classification accuracy and robustness, and making it suitable for cross-individual EEG signal classification tasks.

CN120892918AActive Publication Date: 2025-11-04RES & DEV INST OF NORTHWESTERN POLYTECHNICAL UNIV IN SHENZHEN

Patent Information

Application Number
CN202510951663.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-11-04
Estimated Expiration
2045-07-10

AI Technical Summary

Technical Problem

Existing EEG signal classification methods cannot effectively address individual differences and data scarcity, resulting in insufficient classification accuracy and model generalization performance. In particular, they are difficult to accurately capture individual-specific patterns when applied across individuals. Furthermore, traditional transfer learning methods have failed to fully explore the complementarity between multi-view features and the quality assessment mechanism of transfer features.

Method used

A multi-domain feature extraction network is constructed, and a modular transfer adaptation strategy is adopted. Through multi-view feature collaborative transfer and modular adaptive adjustment, time domain, frequency domain, and spatial domain features are separated. The transfer module is selected through Gaussian mixture model to achieve effective transfer of domain-invariant features and dynamic adaptation of domain-specific features.

Benefits of technology

It significantly improves the robustness and classification accuracy of EEG signal classification tasks and enhances EEG signal classification performance in cross-subject scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120892918A_ABST
    Figure CN120892918A_ABST
Patent Text Reader

Abstract

The invention provides a module adaptive EEG (electroencephalogram) classification method based on multi-view features. Firstly, a multi-domain feature extraction network is constructed, multiple view features such as a time domain, a frequency domain and a space domain can be separated, and electroencephalogram signal classification is achieved; secondly, a modular migration adaptation strategy is adopted to realize fine tuning of network parameters, migration scores are adopted to quantify effectiveness of view features in cross-individual migration, and clustering analysis based on a Gaussian mixture model is adopted to realize adaptive screening of migration modules; through collaborative migration and modular adaptive adjustment of multiple view features, effective transmission of domain invariant features and dynamic adaptation of domain specific features can be realized, and robustness and classification precision of a network model in an electroencephalogram signal classification task in a cross-subject scene are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of medical signal processing, and particularly relates to a module adaptive electroencephalogram (EEG) classification method based on multi-view features. BACKGROUND

[0002] With the help of brain-computer interface technology, the brain, as the main way for users to communicate and control the outside world, can better exert the ability of thought control. Electroencephalogram (EEG) is an important bioelectric signal for recording brain neural activity, which can reflect the cognitive state, emotional changes and brain function under different tasks and environments of individuals. Precise analysis and decoding of EEG is the key to in-depth understanding of brain working mechanism, revealing the principle of neurological diseases and promoting the development of brain-computer interface technology. Non-invasive BCI based on scalp EEG has become the cornerstone of BCI research due to its non-invasiveness, high real-time performance, low cost and rich application ecology. Although there are limitations, its universality in medical, consumer and other fields, as well as the synergistic potential with other technologies, will still be the core driving force for the development of BCI in the foreseeable future.

[0003] EEG is essentially a macroscopic representation of brain neural activity, but the original signal itself cannot directly reflect the user's intention. Researchers need to extract features from motor imagery EEG and classify each extracted feature to convert EEG into executable instructions. However, due to the low signal-to-noise ratio and non-stationary nature of EEG, it is difficult to accurately interpret EEG. In particular, due to significant differences in brain structure, functional connectivity patterns and cognitive task execution strategies among different individuals, EEG shows significant individual differences, further increasing the difficulty of EEG analysis.

[0004] Current mainstream EEG signal classification methods are mainly based on two assumptions: one is to build a general classification model based on multi-subject data, assuming that the EEG features of different individuals are isomorphic; the other is to train a special model for each individual. However, due to the significant individual differences in EEG signals, the general model will have a feature mismatch problem when applied to new individuals. For example, the activation intensity, phase synchrony and other features of the mu rhythm (8-12 Hz) and beta rhythm (18-26 Hz) of EEG signals of different subjects performing the same motor imagery task differ significantly, making it difficult for the general model to accurately capture individual-specific patterns, severely affecting its accuracy in clinical diagnosis and the practical value of personalized brain-computer interface systems. More seriously, the high cost of EEG signal acquisition results in extremely limited data available for individual subjects, and conventional deep learning methods are prone to overfitting under small sample conditions, resulting in a sharp decline in the generalization performance of the model on new data, restricting the practical application of related technologies.

[0005] To address the above challenges, researchers have explored domain adaptation, small sample learning and meta-learning technology paths to alleviate the data scarcity problem through cross-individual knowledge transfer. Traditional transfer learning methods generally use model fine-tuning or feature reuse strategies, but these methods use a unified transfer mechanism to handle different target individuals, which fails to effectively distinguish between domain-invariant features (such as time-frequency response patterns of motor imagery tasks) and domain-specific features (such as signal shifts caused by individual physiological characteristics), resulting in low efficiency of key feature transfer. As the mainstream domain adaptation method, the distribution alignment-based method achieves knowledge transfer by minimizing the feature distribution difference between the source and target domains, but it only aligns the overall distribution, ignoring the differential transfer value of different feature subsets and failing to effectively identify the adjustment of individual-specific features. Existing methods mostly use single-view feature alignment strategies, using only a single type of feature (such as spatial domain features) for transfer learning, which not only fails to fully exploit the complementarity between multi-view features (time-frequency domain, spatial domain, phase features, etc.), but also lacks a quantitative evaluation mechanism for the quality of transferred features, resulting in blindness in the transfer process.

[0006] Overall, even if these existing methods can use the spatio-temporal features of EEG signals to complete the classification task, they cannot be well applied to EEG signals of different individuals. In other words, the above methods cannot fundamentally alleviate the challenges of brain EEG signal classification tasks. SUMMARY

[0007] In order to overcome the shortcomings of the prior art, the present application provides a multi-view feature-based module adaptive electroencephalogram (EEG) classification method. First, in view of the problem of insufficient feature complementarity, a multi-domain feature extraction network is constructed, which can separate multiple view features such as time domain, frequency domain and space domain and realize EEG signal classification; secondly, a modular transfer adaptation strategy is adopted to realize network parameter fine-tuning, wherein the transfer quantization is adopted to quantify the effectiveness of each view feature in cross-subject transfer, and the clustering analysis based on Gaussian mixture model is adopted to realize adaptive filtering of the transfer module. Through the cooperative transfer of multiple view features and the modular adaptive adjustment, the present application can realize the effective transmission of domain-invariant features and the dynamic adaptation of domain-specific features, and significantly improve the robustness and classification accuracy of the network model in the cross-subject EEG signal classification task.

[0008] A multi-view feature-based module adaptive electroencephalogram (EEG) classification method, characterized by the following steps:

[0009] Step 1: EEG signal preprocessing: for the existing EEG signal data containing different subjects performing different imagination tasks, the EEG signal corresponding to the task category is extracted according to the label and subjected to filtering processing, the mean baseline correction method is used for baseline correction of the filtered data, the ICA method is used to remove the eye movement artifact data in the baseline corrected data, and then data normalization processing is performed; all EEG signals of a single subject constitute a target domain data set, and all EEG signals of other subjects constitute a source domain data set; the target domain data set is divided into a training set and a validation set according to the 4:1 ratio of the number of EEG signal samples;

[0010] Step 2: network training: the source domain data set preprocessed in step 1 is input, and the EEG signal classification result is output, and the multi-domain feature extraction network is trained;

[0011] The multi-view feature extraction network mainly includes a frequency domain feature extraction module and a space-time feature extraction module deployed in parallel, and a classifier module, wherein the frequency domain feature extraction module extracts the frequency domain feature map of the input data, the space-time feature extraction module extracts the time domain feature map and the space domain feature map of the input data, the feature maps of different domains are fused and input into the classifier module, the features are mapped to the target category space through the fully connected classification head, and the probability distribution is calculated by using the Softmax function, and the classification result is obtained;

[0012] Step 3: network fine-tuning: the multi-domain feature extraction network trained in step 2 is adjusted by using a modular transfer adaptation strategy;

[0013] Step 4: EEG signal classification: input the EEG signal data set to be classified into the adjusted multi-domain feature extraction network to obtain the classification result.

[0014] Specifically, the filtering processing includes first using a butterworth band-pass filter of 0.5 Hz to 100 Hz to perform band-pass filtering to filter out high-frequency noise interference in the electroencephalogram signal, and then using a 50 Hz notch filter to suppress power frequency noise in the electroencephalogram signal; the data normalization processing adopts z-score normalization operation.

[0015] Specifically, the specific processing process of the frequency domain feature extraction module is as follows: first, the input electroencephalogram signal Band filtering is performed on X to obtain filtered data Wherein, C represents the number of channels, T represents the number of sampling points, N b represents the number of band filters; then, the differential entropy feature of the filtered data is extracted and the power spectral density feature are spliced to obtain the spliced feature Wherein, N s represents the number of time dimension segments, N psd represents the single frequency band power spectral density (PSD) feature length; then, the spliced feature F is input into the linear projection layer, and the z lp =F×W LP is calculated to obtain the feature data Wherein, D represents the feature dimension generated after linear projection; finally, the feature data z lp is input into the feature fusion network composed of a plurality of Transformer encoders connected in series after adding the position encoding to obtain the final frequency domain feature map.

[0016] The differential entropy feature extraction method is: for a continuous random variable X, if its probability density function is f(x), its differential entropy feature is

[0017] The power spectral density feature extraction method is: first, the original signal x(n) is uniformly segmented, each segment has a length of M, and an overlapping part is set between segments, the number of overlapping points overlap is M / 2, and a window function ω(n) is applied to each segment signal; the window function value is smoothly attenuated to zero at the segment endpoints; then, the discrete Fourier transform is performed on each windowed signal, and the power spectral density estimation value is obtained according to the Parseval theorem:

[0018]

[0019] Wherein, P xx,m (k) represents the power spectral density estimation value of the mth segment signal x m (n) at the frequency point k, k=0,1,…,M-1, The complex exponential base used in the calculation of the Fourier transform is U, which is a normalization factor, and the calculation formula is

[0020] Finally, according to The final power spectral density estimate of the signal x(n) is calculated.

[0021] Specifically, the spatial-temporal feature extraction module mainly includes a multi-scale convolution module, a channel attention module and a spatial convolution module, wherein the multi-scale convolution module uses several 1D convolution kernels of different sizes to perform independent convolution operations on the input electroencephalogram signal, adds the feature maps obtained from each convolution operation to obtain a time-domain feature map of the electroencephalogram signal; the time-domain feature map is sent to the spatial convolution module after the channel attention module and batch normalization operation and a nonlinear activation function ReLU; the spatial convolution module uses a convolution kernel of size 3x1x1 to perform convolution operation on the input feature map, and then performs batch normalization operation and nonlinear activation function ReLU to obtain a spatial feature map of the electroencephalogram signal.

[0022] The convolution operation adopts same convolution;

[0023] The channel attention module first uses a global average pooling method to compress the feature map output by the multi-scale convolution module by channel to obtain global features of each channel, then uses two fully connected layers of the same size to learn the relationship between channels, uses a nonlinear activation function Sigmoid to obtain channel attention weights, and multiplies the channel attention weights with the features of the corresponding channel to output the weighted feature map.

[0024] Specifically, the specific implementation process of fusing different domain features in the multi-domain feature extraction network is as follows: first, the feature maps of different domains are respectively decoupled and expanded along the sample dimension, then the channel dimension is spliced and the spatial dimension is aligned to obtain the spliced multi-domain features; then, the spliced multi-domain features are divided into blocks by the image block embedding layer through a sliding window, the feature vectors in each window are mapped to the hidden space through a learnable linear projection to generate an embedded vector sequence with position encoding information, and a class identifier is introduced as a trainable parameter to be inserted into the head of the sequence; finally, the multi-domain feature fusion is realized through a stacked Transformer encoder.

[0025] Specifically, the following total loss function is used in the network training:

[0026] L final =L d +L cls (2)

[0027] Wherein, L finaldenotes the total loss of the network, L d denotes the domain loss, L cls denotes the classification loss;

[0028] The domain loss L d is calculated as follows:

[0029]

[0030] wherein x s denotes a signal sample in the source domain dataset, x t denotes a signal sample in the target domain dataset, M denotes the number of feature maps, M = 3, which are respectively frequency domain feature map, time domain feature map and spatial domain feature map, i, j denote the serial number of signal samples in the source domain dataset, p, q denote the serial number of signal samples in the target domain dataset, and φ(·,·) denotes a multi-kernel function of two signals, which is calculated as follows:

[0031]

[0032] wherein, is the u-th kernel function related to signals x and y, β u is the corresponding coefficient, satisfying k is the number of kernel functions;

[0033] The classification loss L cls is calculated as follows:

[0034]

[0035] wherein x i denotes the i-th signal sample, y i denotes the label of the i-th signal sample, N is the total number of samples, f(·) denotes the multi-domain feature extraction network, and CE(·) denotes the calculation of cross-entropy loss function.

[0036] In the training process, the Adam optimization algorithm is used to minimize the total loss function to obtain the best network model.

[0037] Specifically, the specific process of adjusting the network by using the modular migration adaptation strategy is as follows: first, the migration score vector of the feature extraction module corresponding to different views is calculated, then the migration score vector is taken as a sample, all migration score vectors constitute a new sample set, the samples in the set are analyzed by using a Gaussian mixture model for clustering analysis, and are divided into two classes, the samples with low mean value are taken as a frozen class, and the other samples are taken as a fine-tuning class; finally, the target domain training data set is input into the multi-domain feature extraction network for retraining, and the parameters of the feature extraction module corresponding to the view of the frozen class sample are kept unchanged during training, and the parameters of the feature acquisition module corresponding to the view of the fine-tuning class sample are adjusted, and the network fine-tuning is completed.

[0038] Specifically, the migration score vector is composed of the distribution difference score Score c and the target domain discrimination score Score d The specific calculation process is as follows:

[0039] Step (1): the signal sample distribution difference score of the source domain data set and the target domain data set under the mth feature map V m is calculated according to the following formula

[0040]

[0041] Wherein, x s represents a signal sample in the source domain data set, x t represents a signal sample in the target domain data set, i and j represent the serial numbers of signal samples in the source domain data set, p and q represent the serial numbers of signal samples in the target domain data set, m = 1, 2, …, M, k(·,·) is a positive definite kernel function, and a Gaussian kernel function is adopted x and y represent two signals in the input space, and σ 2 is a bandwidth parameter;

[0042] Step (2): the average inter-class distance between different class signal samples in the target domain data set under the mth feature map V m is calculated according to the following formula and the average intra-class distance is calculated according to the following formula

[0043]

[0044] Wherein, |·| represents L1 distance, C represents the total number of classes of signal samples in the target domain data set, f center (·) represents the calculation of the center feature vector of all signal samples in a certain class, represents all signal samples in class c1, represents all signal samples in class c2, and xc represents all signal samples in class c, represents all samples of class c1 input into the feature extraction module for extracting the mth feature map, the feature vector set of all samples of this class is obtained, represents all samples of class c2 input into the feature extraction module for extracting the mth feature map, the feature vector set of all samples of this class is obtained, m (x c ) represents all samples x c of class c input into the feature extraction module for extracting the mth feature map, c1, c2, c respectively represent class serial numbers, N c represents the total number of signal samples contained in class c, i represents the signal sample serial number, represents the ith signal sample in class c, represents the ith signal sample of class c input into the feature extraction module for extracting the mth feature map, input into the feature extraction module for extracting the mth feature map, the feature of the sample is obtained; f center (·) represents the calculation of the center feature vector of all signal samples in a certain class, for all signal samples of class c The calculation formula of the center feature vector under the mth feature map is:

[0045]

[0046] Step (3): calculate the target domain data set discrimination score under the mth view V m

[0047]

[0048] Step (4): take as the transfer score vector of the feature extraction module corresponding to the mth view.

[0049] Specifically, the implementation process of the clustering analysis of the transfer score vector by the Gaussian mixture model is: the optimal mean μ c and covariance matrix Conv c are found by maximizing the log-likelihood function, so as to cluster the transfer score vector into two categories of freezing and fine-tuning, which is expressed by the formula:

[0050]

[0051] wherein, Scor V represents the set of transfer score vectors under all views, D(Score​V ) represents the clustering result of all migration score vectors, D(Score V )∈{0,1}, if the output is 0, it is judged that the input belongs to the frozen class, if the output is 1, it is judged to belong to the fine-tuning class, μ c is the mean vector of the cluster c, Cov c is the corresponding covariance matrix, and p(c) is the probability of the cluster c, is the migration score vector of the feature extraction module corresponding to the mth view; c=1 or 2 respectively represents the frozen class and the fine-tuning class.

[0052] Specifically, the Transformer encoder is a composite computing unit containing multi-head attention mechanism, layer normalization, residual connection and feedforward network, which establishes long-range dependency relationship between cross-domain features through multi-head self-attention mechanism, and quantifies the correlation strength of frequency domain features and time domain features by using cross-attention weight matrix, and each level of encoder realizes the fusion of multi-domain features through cascaded multi-layer nonlinear transformation.

[0053] The beneficial effects of the present application are: since the multi-domain feature extraction network is constructed, the time domain, frequency domain, spatial domain and other view features can be separated, and the electroencephalogram signal classification is realized; since the network adopts modular design, it is beneficial to adaptively adjust the parameters of the module, obtain a better network model, and thus obtain a better classification result by using the network; since the migration score calculation method based on multi-view features is adopted, the multi-domain distribution difference of the electroencephalogram signal can be better expressed, and then the migration path is determined according to the difference, so that the effective transmission of domain-invariant features and the dynamic adaptation of domain-specific features can be realized, and the robustness and classification accuracy of the network model in the electroencephalogram signal classification task in the cross-subject scenario are significantly improved. BRIEF DESCRIPTION OF DRAWINGS

[0054] Figure 1 is a basic flowchart of the module adaptive electroencephalogram (EEG) classification method based on multi-view features of the present application;

[0055] Figure 2 is a schematic diagram of the network fine-tuning process using the modular migration adaptation strategy of the present application. DETAILED DESCRIPTION

[0056] The present application will be further described below in conjunction with the drawings and examples, and the present application includes but is not limited to the following examples.

[0057] The running environment of this embodiment: the server for model training is configured with two GeForce GTX 3090 GPUs with 24GB of video memory, 256GB of running memory, and Ubuntu 22.04 LTS as the operating system. Pytorch 2.5.0 is used as the deep learning framework, Python 3.8 is used as the programming language, and VS Code is used as the integrated development environment for related experiments.

[0058] The experiment uses brain-computer interface competition dataset (BCI Competition IV 2a / 2b). The BCI Competition IV 2a dataset includes motor imagery electroencephalogram data from 9 subjects. The prompt-based BCI paradigm includes four different motor imagery tasks, namely left hand, right hand, both feet and tongue motor imagery. Each subject performed two experiments on different days, and the experiment was divided into sessions. Each session consists of 6 runs with appropriate rest time between runs. A run includes 48 trials (12 for each of the four imagination categories), and each session contains a total of 288 trials. At the beginning of each session, the subject performed a 5-minute electroencephalogram data collection for EOG (electrooculogram) evaluation and subsequent pre-processing to handle ocular artifacts interference. The subjects of the BCI Competition IV 2b dataset are consistent with those in the BCI Competition IV 2a dataset. The subjects performed left and right hand motor imagery tasks without feedback screening at different times. The experiment is divided into sessions, each containing 6 runs, each run has 10 trials of two types of tasks, each session contains 120 trials, and each person has 120 groups of data for each type of motor imagery task. Three bipolar electrodes are used to collect electroencephalogram signals at three positions (C3, Cz, C4) with a sampling frequency of 250Hz. The experimental paradigm and other experimental settings are consistent with BCI Competition IV 2a.

[0059] As shown in Figure 1 , the specific implementation process of the present application is as follows:

[0060] 1. Electroencephalogram preprocessing

[0061] For the existing electroencephalogram signal data containing different subjects performing different imagination tasks, the electroencephalogram signals of the corresponding task categories are extracted and filtered according to the labels, the filtered data are baseline corrected by using the mean baseline correction method (that is, the electroencephalogram signals of all subjects are taken as the reference with the same zero point, so as to avoid the adverse effects on subsequent analysis due to the different baseline levels of the two groups of subjects), the EOG data in the baseline corrected data are removed by using the ICA method, in order to eliminate the differences between the amplitudes and characteristics of different electroencephalogram signals in each data set and make the electroencephalogram signals comparable, and then the electroencephalogram signals are normalized. Then, all the electroencephalogram signals of a single subject constitute a target domain data set, and all the electroencephalogram signals of other subjects constitute a source domain data set, and the target domain data set is divided into a training set and a verification set according to the number of electroencephalogram signal samples in the ratio of 4:1.

[0062] Specifically, the specific way of filtering processing is: first, a butterworth band-pass filter of 0.5Hz to 100Hz is used for band-pass filtering to filter out high frequency noise interference in the electroencephalogram signal, and then a 50Hz notch filter is used to suppress power frequency noise in the electroencephalogram signal.

[0063] Specifically, the data normalization adopts z-score normalization operation. The z-score normalization makes the mean value of the normalized data 0 and the standard deviation 1 by subtracting the mean value of each sample value of the original signal from the sample value and then dividing by the standard deviation. This method is suitable for the case where the data distribution is approximately normal distribution. In most electroencephalogram signal analysis, if the data does not have obvious skew distribution, z-score normalization can effectively make the data comparable under different conditions. The z-score calculation formula is:

[0064]

[0065] Wherein, x i represents input data, x0 represents normalized output, μ and σ respectively represent the mean and variance of the input data.

[0066] 2. Network training

[0067] In view of the problem that the model generalization ability is limited due to the cross-subject and cross-scene differences in the electroencephalogram signal classification task, the present application first constructs a modularized multi-domain feature extraction network, such as Figure 1As shown, it mainly includes a parallelly deployed frequency domain feature extraction module and a space-time feature extraction module, and a classifier module, wherein the frequency domain feature extraction module extracts a frequency domain feature map of the input data, the space-time feature extraction module extracts a time domain feature map and a space domain feature map of the input data, the feature maps of different domains are fused and input into the classifier module, the features are mapped to a target category space through a fully connected classification head, and a probability distribution is calculated by using a Softmax function to obtain a classification result.

[0068] Specifically, as shown in the figure, Figure 1 the specific processing process of the frequency domain feature extraction module is as follows: first, the input electroencephalogram signal is subjected to frequency band filtering to obtain filtered data , wherein C represents the number of channels, T represents the number of sampling points, N b represents the number of frequency band filters; then, differential entropy features and power spectral density features of the filtered data are extracted and spliced to obtain spliced features , wherein N s represents the number of time dimension segments, N psd represents the length of single frequency band power spectral density (PSD) features; then, the spliced features F are input into a linear projection layer, and the feature data is calculated according to z lp =F×W LP , wherein D represents the feature dimension generated after the projection. lp Finally, the feature data z is input into a feature fusion network composed of a plurality of serially connected Transformer encoders after adding position encoding to obtain the final frequency domain feature map.

[0069] The differential entropy feature extraction method is as follows: for a continuous random variable X, if its probability density function is f(x), the differential entropy (DE) feature of X is

[0070] The power spectral density feature extraction method is as follows: first, the original signal x(n) is uniformly segmented, each segment has a length of M, and an overlap part is arranged between segments, the number of overlap points overlap is M / 2 (which can also be adjusted according to actual conditions), signal segmentation can make the statistical characteristics of each segment more stable, which is beneficial to subsequent accurate estimation, the length of the original signal is N, and the number of segments L is:

[0071]

[0072] A window function ω(n) (such as a Hanning window function) is applied to each segment of the signal, and the window function value is smoothly attenuated to zero at the segment endpoints to suppress spectral leakage caused by signal truncation edge discontinuities. The windowing process can reduce the impact of edge discontinuities on the spectrum. Taking the Hanning window function as an example, The mth windowed segment of the signal is:

[0073] x m (n)=x(n+m(M-overlap))ω(n)

[0074] where n is the sampling point, x m (n) is the signal after applying the window function.

[0075] Then, a discrete Fourier transform is performed on each windowed segment of the signal, and a power spectral density estimate (i.e., a periodogram) is obtained according to the Parseval theorem:

[0076]

[0077] where P xx,m (k) represents the power spectral density estimate (periodogram) of the mth segment of the signal x m (n) at the frequency point k, k = 0, 1, …, M-1, is a complex exponential base used when calculating the Fourier transform, and U is a normalization factor, and the calculation formula is

[0078] Finally, all the periodograms at the frequency point k are averaged, i.e., according to the final power spectral density estimate of the signal x(n) is calculated.

[0079] The Transformer encoder is N serial complex computing units containing multi-head attention mechanisms, layer normalization, residual connections, and feedforward networks, which establish long-range dependencies between cross-domain features through multi-head self-attention mechanisms, quantify the correlation strength between frequency domain features and time domain features using cross-attention weight matrices, and realize the fusion of multi-domain features through cascaded multi-layer nonlinear transformations at each level of the encoder. The calculation process can be represented by the following formula:

[0080]

[0081] where l represents the number of layers of the encoder, for the lth layer of the encoder, represents the input feature data of the lth layer of the encoder, represents the output feature data of the lth layer of the encoder, E pos represents the position encoding, For the intermediate representation after the attention module, MHA(·) represents the multi-head attention module, MLP(·) a feedforward network, and LN(·) a layer normalization operation. First, the input feature is input into the multi-head self-attention module MHA(·), which learns the information relationship between different positions and outputs the result to the original input to form a residual connection, which can alleviate the gradient vanishing problem in deep networks and help the model retain the original information. Then, the added result is subjected to layer normalization LN(·) to stabilize the training process and accelerate the convergence speed. The output of this step is , which is the intermediate representation after this layer of attention module. Next, is input into a feedforward network MLP(·), which is used to perform nonlinear transformation on the feature vector of each position to enhance the expression ability of the model and add the output of the MLP to its input to form a second residual connection. Finally, the added result is subjected to layer normalization again to obtain the output feature data of the lth layer of the encoder

[0082] Specifically, as shown in Figure 1 , the spatio-temporal feature extraction module mainly includes a multi-scale convolution module, a channel attention module, and a spatial domain convolution module. The multi-scale convolution module uses several 1D convolution kernels of different sizes to perform independent convolution operations on the input electroencephalogram signal. The convolution operation adopts same convolution. Further, the feature maps obtained by each convolution operation are added to obtain a multi-scale feature map (i.e., a time domain feature map) of the electroencephalogram signal. Then, the channel attention module is used to assign different channel weights to the feature channels in the feature map obtained by time domain feature extraction, highlighting key features and ignoring redundancies. The channel attention module first uses a global average pooling (GAP) method to compress the feature map output by the multi-scale convolution module by channel to obtain global features of each channel. Then, two fully connected layers of the same size are used to learn the relationship between channels, and a nonlinear activation function Sigmoid is used to obtain channel attention weights. The channel attention weights are multiplied with the features of the corresponding channels, and the weighted feature map is output. The size of the weighted feature map is consistent with that before weighting. The weighted feature map is subjected to batch normalization (Batch Norm), and then a nonlinear activation function ReLU is used to send it to the spatial domain convolution module. The spatial domain convolution module also uses same convolution to obtain a spatial domain feature map of the electroencephalogram signal, and performs batch normalization (Batch Norm) on the obtained spatial domain feature map. A nonlinear activation function ReLU is used to obtain the spatial domain feature map of the electroencephalogram signal.

[0083] Specifically, the specific implementation process of the fusion of the frequency domain feature map, the time domain feature map and the spatial domain feature map is as follows: first, the feature maps in different domains are respectively decoupled and unfolded along the sample dimension, then the channel dimension is spliced, and the spatial dimension is aligned to obtain the spliced multi-domain features; then, the spliced multi-domain features are divided into blocks by a patch embedding layer, the feature vectors in each window are mapped to the latent space by a learnable linear projection, an embedded vector sequence with position encoding information is generated, and a class token is introduced as a trainable parameter and inserted into the head of the sequence; finally, the multi-domain feature fusion is realized by a stacked Transformer encoder.

[0084] The preprocessed source domain data set in step 1 is input into the multi-domain feature extraction network for training, and a pre-trained network model is obtained.

[0085] Specifically, in the network training process, in order to align the feature distributions corresponding to the frequency domain, time domain and spatial domain feature maps, the following domain loss L is introduced d :

[0086]

[0087] wherein x s represents a signal sample in the source domain data set, x t represents a signal sample in the target domain data set, M represents the number of feature maps, here M = 3, i.e. the frequency domain feature map output by the multi-domain feature extraction module, the time domain feature map obtained by the time-space feature extractor module and the spatial domain feature map, i and j represent the serial numbers of signal samples in the source domain data set, p and q represent the serial numbers of signal samples in the target domain data set, and φ(·,·) represents a multi-kernel function of two signals, and the calculation formula is as follows:

[0088]

[0089] wherein, is a different kernel function (such as a Gaussian radial basis kernel function) related to signals x and y, β u is the corresponding coefficient, and By weighting and summing the results of multiple different kernel functions, the characteristics of the data can be more comprehensively captured, thereby improving the estimation accuracy of the distribution difference, and k is the number of kernel functions.

[0090] The classification loss is evaluated using a cross-entropy loss function, i.e.

[0091]

[0092] wherein L cls represents the classification loss, x idenotes the i-th signal sample, y i denotes the label of the i-th signal sample, N is the total number of samples, f(·) denotes a multi-domain feature extraction network, and CE(·) denotes a cross-entropy loss function.

[0093] In order to make full use of the features in the label and the domain, a composite loss function is used as the final loss function of the model:

[0094] L final = L d + L cls (20)

[0095] wherein L final denotes the total loss of the network.

[0096] During the model training process, the composite loss function is minimized by using the Adam optimization algorithm, so that the model can learn a feature representation with good generalization performance, and better align the feature distributions of the source domain and the target domain.

[0097] 3. Network fine-tuning

[0098] The multi-domain feature extraction network trained in step 2 is adjusted using a modular transfer adaptation strategy. Specifically, as shown in Figure 2 , the present application introduces the concept of a transfer score vector to quantitatively evaluate the modules that extract frequency domain feature maps, time domain feature maps and spatial domain feature maps in the multi-domain feature extraction network, and performs clustering analysis on the transfer score vector through a Gaussian mixture model. According to the clustering results, it is determined whether the module parameters need to be fine-tuned, wherein the clustering results are divided into two categories. Low mean clustering represents a good transferability module, which is classified as a frozen class, and other transfer vectors are classified as a fine-tuning class. The target domain training data set is input into the multi-domain feature extraction network for retraining. During training, the parameters of the feature acquisition modules corresponding to the views of the frozen class transfer vectors are kept unchanged, and the parameters of the feature acquisition modules corresponding to the views of the fine-tuning class transfer vectors are adjusted, thereby completing the network fine-tuning.

[0099] Specifically, the transfer score vector is composed of a source domain and target domain distribution difference score Score c and a target domain discrimination score Score d The specific calculation process is as follows:

[0100] Step (1): Calculate the signal sample distribution difference score of the source domain data set and the target domain data set under the m-th feature map V m

[0101]

[0102] wherein x s denotes a signal sample in the source domain data set, and x​t Let i and j represent the signal sample indices in the source domain dataset, and p and q represent the signal sample indices in the target domain dataset. Let m = 1, 2, ..., M, and k(·,·) be the positive definite kernel function, typically a Gaussian kernel function. x and y represent two signals in the input space, σ 2 It is the bandwidth parameter, which controls the width of the function and determines the rate at which the function value decays. The Gaussian kernel function measures the similarity between two input signals in this way. Signals with high similarity will have higher values ​​under the Gaussian kernel function, while signals with low similarity will have lower values ​​under the Gaussian kernel function.

[0103] Step (2): Calculate the m-th view V according to the following formula. m The average interclass distance between different class signal samples in the target domain dataset below and average intra-class distance

[0104]

[0105] Where |·| represents the L1 distance, C represents the total number of classes of signal samples in the target domain dataset, and f center (·) represents calculating the central feature vector of all signal samples in a certain category. This represents all signal samples in category c1. x represents all signal samples in category c2. c This represents all signal samples in category c. This means all samples of category c1 After being input into the m-th view feature extractor, the resulting set of feature vectors for all samples of that class is obtained. This means all samples of category c2. After being input into the m-th view feature extractor, the resulting set of feature vectors for all samples of that class, V, is obtained. m (x c ) represents all samples x of category c. c After being input into the m-th view feature extractor, the resulting set of feature vectors for all samples of that class is obtained, where c1, c2, and c represent the class indices, and N represents the class number. c This represents the total number of signal samples in category c, where i represents the signal sample index. This represents the i-th signal sample in category c. This indicates that the i-th signal sample of category c is... After being input into the m-th view feature extractor, the resulting set of feature vectors for all samples of that class is obtained; f center(·) represents the calculation of the center eigenvector of all signal samples in a certain category, for all signal samples of the category c The calculation formula of the center eigenvector under the mth feature map is:

[0106]

[0107] Step (3): calculate the target domain data set discriminant score of the mth view V m

[0108]

[0109] Step (4): take as the transfer score vector of the feature extraction module corresponding to the mth view.

[0110] Specifically, the implementation process of the clustering analysis of the transfer score vector by the Gaussian mixture model is: the optimal mean μ c and covariance matrix Cov c are found by maximizing the log-likelihood function, so that the transfer score vector is clustered into two categories of freezing and fine-tuning, which is expressed by the formula:

[0111]

[0112] Where Score V represents the set of transfer score vectors under all views, D(Score V ) represents the clustering result of all transfer score vectors, D(Score V ) ∈ {0, 1}, if the output is 0, it is judged that the input belongs to the freezing class, if the output is 1, it is judged to belong to the fine-tuning class, μ c is the mean vector of the cluster c, Cov c is the corresponding covariance matrix, p(c) is the probability of cluster c, is the transfer score vector of the feature extraction module corresponding to the mth view; c = 1 or 2 respectively represents the freezing class and the fine-tuning class.

[0113] 4. Electroencephalogram classification

[0114] The electroencephalogram data set to be classified is input into the adjusted multi-domain feature extraction network to obtain the classification result.

[0115] Based on the data of the present embodiment, the effectiveness of the method of the present application is verified in the following way:

[0116] ​In view of the high cost of electroencephalogram (EEG) acquisition and the problem of overfitting of deep learning model caused by the scarcity of single-subject sample size, EEGMix method is used to enhance the EEG signal data in the training and testing stages, that is, for a given set of EEG signals X of the same category, a pair of EEG signals x i and x j are randomly selected from X new , and the new EEG signal is obtained by mixing using EEGMix:

[0117]

[0118] wherein x i represents the new EEG signal obtained by mixing, β represents the mixing ratio of the pair of EEG signals x j and x j , β∈Beta(α,α), wherein α∈(0,∞), σ(X) represents the global standard deviation of the set of EEG signals X, μ(X) represents the global average value of the set of EEG signals X, μ(x j ) represents the average value of the randomly selected EEG signal x j , and σ(x j ) represents the standard deviation of the randomly selected EEG signal x o .

[0119] The classification accuracy Acc. (Acuraccy) is selected to evaluate the effectiveness of the method, which is defined as follows:

[0120]

[0121] wherein TP (true positive) is the number of samples that are actually positive and are correctly classified, TN (true negative) is the number of samples that are actually negative and are correctly classified, FP (false positive) is the number of samples that are actually negative but are misclassified as positive, and FN (false negative) is the number of samples that are actually positive but are misclassified as negative.

[0122] The Kappa coefficient is selected to evaluate the classification consistency, and the calculation formula is:

[0123]

[0124] wherein p e represents the overall accuracy, and p o represents the expected accuracy.

[0125] The classification accuracy results of the method of the present application and other current optimal methods (SOTA, state of art) on the BCI Competition IV 2a dataset are shown in Table 1, and the classification accuracy results on the BCI Competition IV 2b dataset are shown in Table 2. In the tables, DAN represents a deep adaptation network (DAN) architecture proposed by Long et al., DANN represents a domain-adversarial neural network (DANN) proposed by Yaroslav et al., DRDA represents an end-to-end domain adaptation method based on deep representation (DRDA) proposed by Zhao et al. considering the edge and conditional distribution differences between the source domain and the target domain, DJDAN represents a dynamic joint domain adaptation network (DJDAN) proposed by Hong et al., CLUDA represents a deep domain adaptation framework based on CORAL loss (DDAF-CORAL) proposed by Zhong et al. to solve the distribution difference problem caused by subject-related and time-related changes in electroencephalogram signals, and average / std represents the average / standard deviation of the classification accuracy of the corresponding method on different individuals. It can be seen that the present application can obtain higher classification accuracy compared with the benchmark method.

[0126] Table 1

[0127]

[0128]

[0129] Table 2

[0130]

Claims

1. A modular adaptive EEG signal classification method based on multi-view features, characterized in that... The steps are as follows: Step 1, EEG signal preprocessing: For existing EEG signal data containing different subjects performing different imagination tasks, the EEG signals of the corresponding task categories are extracted and filtered according to the labels. The mean baseline correction method is used to correct the baseline of the filtered data. The ICA method is used to remove eye movement artifacts from the baseline corrected data, and then the data is normalized. The target domain dataset is composed of all EEG signals from a single subject, and the source domain dataset is composed of all EEG signals from other subjects. The target domain dataset is divided into a training set and a validation set according to a 4:1 ratio of the number of EEG signal samples. Step 2, Network Training: The multi-domain feature extraction network is trained using the source domain dataset after preprocessing in Step 1 as input and the EEG signal classification results as output. The multi-view feature extraction network mainly includes a frequency domain feature extraction module and a spatiotemporal feature extraction module deployed in parallel, as well as a classifier module. The frequency domain feature extraction module extracts the frequency domain feature map of the input data, and the spatiotemporal feature extraction module extracts the temporal domain feature map and the spatial domain feature map of the input data. The feature maps from different domains are fused and input into the classifier module. The features are mapped to the target class space through a fully connected classification head, and the probability distribution is calculated using the Softmax function to obtain the classification result. Step 3, Network Fine-tuning: The multi-domain feature extraction network trained in Step 2 is adjusted using a modular transfer adaptation strategy; Step 4, EEG signal classification: Input the EEG signal dataset to be classified into the adjusted multi-domain feature extraction network to obtain the classification result.

2. The modular adaptive EEG signal classification method based on multi-view features as described in claim 1, characterized in that: The filtering process includes first using a Butterworth bandpass filter from 0.5Hz to 100Hz to filter out high-frequency noise interference in the EEG signal, and then using a 50Hz notch filter to suppress power frequency noise in the EEG signal; the data normalization process uses z-score normalization.

3. The modular adaptive EEG signal classification method based on multi-view features as described in claim 1, characterized in that: The specific processing procedure of the frequency domain feature extraction module is as follows: First, input the electroencephalogram (EEG) signal. Perform frequency band filtering on X to obtain the filtered data. Where C represents the number of channels, T represents the number of sampling points, and N b This indicates the number of frequency band filters; then, the differential entropy features of the filtered data are extracted. and power spectral density characteristics Then, the features are spliced ​​together to obtain the spliced ​​features. Where, N s N represents the number of segments in the time dimension. psd This represents the feature length of the single-band power spectral density (PSD); then, the stitched feature F is input into the linear projection layer, according to z... lp =F×W LP Calculate the feature data in, D represents the feature dimension generated after linear projection; finally, the feature data z lp After adding position encoding, the input is a feature fusion network consisting of several Transformer encoders connected in series, resulting in the final frequency domain feature map; The differential entropy feature extraction method is as follows: For a continuous random variable X, if its probability density function is f(x), its differential entropy feature is... The power spectral density feature extraction method is as follows: First, the original signal x(n) is uniformly divided into segments of length M, with overlapping portions between segments. The number of overlap points is M / 2. A window function ω(n) is applied to each segment, and the window function value smoothly decays to zero at the segment endpoints. Then, a discrete Fourier transform is performed on each windowed segment, and the power spectral density estimate is obtained according to Parseval's theorem. Among them, P xx,m (k) represents the m-th signal segment x m (n) The power spectral density estimate at frequency point k, k = 0, 1, ..., M-1. The complex exponential basis used in calculating the Fourier transform, where U is the normalization factor, is calculated using the following formula: Finally, according to The final power spectral density estimate of signal x(n) is calculated.

4. The modular adaptive EEG signal classification method based on multi-view features as described in claim 1, characterized in that: The spatiotemporal feature extraction module mainly includes a multi-scale convolution module, a channel attention module, and a spatial convolution module. The multi-scale convolution module uses several 1D convolution kernels of varying sizes to perform independent convolution operations on the input EEG signal. The feature maps obtained from each convolution operation are summed to obtain the temporal feature map of the EEG signal. The temporal feature map is then processed by the channel attention module to obtain a weighted feature map. This weighted feature map undergoes batch normalization and a ReLU activation function before being fed into the spatial convolution module. The spatial convolution module uses a 3×1×1 convolution kernel to perform convolution operations on the input feature map, followed by batch normalization and a ReLU activation function to obtain the spatial feature map of the EEG signal. The convolution operation described uses same convolution; The channel attention module first uses global average pooling to compress the feature map output by the multi-scale convolution module by channel to obtain the global features of each channel. Then, it uses two fully connected layers of the same size to learn the relationship between channels, uses the non-linear activation function Sigmoid to obtain the channel attention weights, and multiplies the channel attention weights with the features of the corresponding channels to output the weighted feature map.

5. The modular adaptive EEG signal classification method based on multi-view features as described in claim 1, characterized in that: The specific implementation process of fusing features from different domains in the multi-domain feature extraction network is as follows: First, the feature maps of different domains are decoupled and unfolded along their sample dimensions, then concatenated along the channel dimension and aligned along the spatial dimension to obtain the concatenated multi-domain features; then, the concatenated multi-domain features are divided into sliding window blocks by an image block embedding layer, and the feature vectors in each window are mapped to the latent space through a learnable linear projection to generate an embedding vector sequence with positional encoding information, and a class identifier is introduced as a trainable parameter and inserted into the beginning of the sequence; finally, multi-domain feature fusion is achieved through a stacked Transformer encoder.

6. The modular adaptive EEG signal classification method based on multi-view features as described in claim 1, characterized in that: The network training uses the following total loss function: L final =L d +L cls (2) Among them, L final L represents the total network loss. d Representation domain loss, L cls Indicates classification loss; The domain loss L d The calculation formula is as follows: Where, x s x represents a signal sample in the source domain dataset. t Let M represent the signal samples in the target domain dataset, M=3, which are the frequency domain feature map, time domain feature map, and spatial domain feature map, respectively. Let i and j represent the signal sample indices in the source domain dataset, and p and q represent the signal sample indices in the target domain dataset. Let φ(·,·) represent the multi-kernel function of the two signals, and its calculation formula is as follows: in, It is the u-th kernel function with respect to signals x and y, β u The corresponding coefficients satisfy... k is the number of kernel functions; The classification loss L cls The calculation formula is as follows: Where, x i Let y represent the i-th signal sample. i Let f(·) represent the label of the i-th signal sample, N be the total number of samples, f(·) represent the multi-domain feature extraction network, and CE(·) represent the cross-entropy loss function. During training, the best network model is obtained by minimizing the total loss function using the Adam optimization algorithm.

7. The modular adaptive EEG signal classification method based on multi-view features as described in claim 1, characterized in that: The specific process of adjusting the network using the modular transfer adaptation strategy is as follows: First, calculate the transfer score vectors of the feature extraction modules corresponding to different views. Then, using the transfer score vectors as samples, all transfer score vectors form a new sample set. Use a Gaussian mixture model to perform cluster analysis on the samples in this set, dividing them into two classes: the low-mean clustered samples are designated as the frozen class, and the other samples are designated as the fine-tuning class. Finally, input the target domain training dataset into the multi-domain feature extraction network for retraining. During training, keep the parameters of the feature extraction modules for the views corresponding to the frozen class samples unchanged, and adjust the parameters of the feature acquisition modules for the views corresponding to the fine-tuning class samples to complete the network fine-tuning.

8. The modular adaptive EEG signal classification method based on multi-view features as described in claim 1, characterized in that: The transfer score vector is composed of the distribution difference score between the source domain and the target domain. c Score d The composition, and its specific calculation process are as follows: Step (1): Calculate the m-th feature map V using the following formula. m The signal sample distribution difference score between the source domain dataset and the target domain dataset is as follows. Where, x s x represents a signal sample in the source domain dataset. t Let i and j represent the signal sample indices in the source domain dataset, and p and q represent the signal sample indices in the target domain dataset. Let m = 1, 2, ..., M, and k(·,·) be the positive definite kernel function, using a Gaussian kernel function. x and y represent two signals in the input space, σ 2 It is a bandwidth parameter; Step (2): Calculate the m-th feature map V according to the following formula. m The average inter-class distance between different class signal samples in the target domain dataset below and average intra-class distance Where |·| represents the L1 distance, C represents the total number of classes of signal samples in the target domain dataset, and f center (·) represents calculating the central feature vector of all signal samples in a certain category. This represents all signal samples in category c1. x represents all signal samples in category c2. c This represents all signal samples in category c. This means all samples of category c1 The set of feature vectors for all samples of this class is obtained after inputting into the feature extraction module that extracts the m-th feature map. This means all samples of category c2. The feature vector set V, obtained after inputting into the feature extraction module that extracts the m-th feature map, represents the set of feature vectors for all samples of that class. m (x c ) represents all samples x of category c. c The feature vector set of all samples of this class is obtained after inputting into the feature extraction module that extracts the m-th feature map. c1, c2, and c represent the class index, respectively, and N c This represents the total number of signal samples in category c, where i represents the signal sample index. This represents the i-th signal sample in category c. This indicates that the i-th signal sample of category c is... The features of the sample are obtained after being input into the feature extraction module that extracts the m-th feature map; f center (·) represents calculating the central feature vector of all signal samples in a certain category. For all signal samples of category c, · represents the central feature vector of all signal samples in a certain category. The formula for calculating the central eigenvector under the m-th feature map is: Step (3): Calculate the m-th view V using the following formula. m Discriminant scores of the target domain dataset below Step (4): with The transfer score vector of the feature extraction module corresponding to the m-th view.

9. The modular adaptive EEG signal classification method based on multi-view features as described in claim 1, characterized in that: The process of clustering the migration score vector using a Gaussian mixture model is as follows: the optimal mean μ is found by maximizing the log-likelihood function. c Covariance matrix Cov c This clusters the transfer score vectors into two categories: frozen and fine-tuned, as expressed by the formula: Among them, Score V It represents the set of migration score vectors across all views. D(Score V D(Score) represents the clustering result of all transfer score vectors. V For inputs ∈ {0, 1}, if the output is 0, the input is considered to belong to the frozen class; if the output is 1, the input is considered to belong to the fine-tuning class. c It is the mean vector of cluster c, Cov c Let p(c) be the corresponding covariance matrix, and p(c) be the probability of clustering c. It is the transfer score vector of the feature extraction module corresponding to the m-th view; c = 1 or 2 represents the frozen class and the fine-tuned class, respectively.

10. A modular adaptive EEG signal classification method based on multi-view features as described in claim 3 or 5, characterized in that: The Transformer encoder is a composite computational unit that includes a multi-head attention mechanism, layer normalization, residual connections, and a feedforward network. It establishes long-range dependencies between cross-domain features through a multi-head self-attention mechanism, quantifies the correlation strength between frequency domain features and time domain features using a cross-attention weight matrix, and achieves the fusion of multi-domain features through cascaded multi-layer nonlinear transformations at each encoder level.

Citation Information

Patent Citations

  • Auditory attention detection method and system based on time-frequency domain fusion

    CN118121192A

  • Motor imagery electroencephalogram recognition method based on domain adaptation network

    CN119622543A

Cited By

  • Model training method, cross-modal retrieval method, device and equipment

    CN122241162A