SSVEP frequency identification method based on domain adversarial network
By adopting a domain adversarial network architecture in the SSVEP decoding algorithm and combining MMD and MCC losses, the problem of poor generalization of the SSVEP decoding algorithm under cross-participants conditions is solved, achieving higher classification accuracy and better decoding effect.
Patent Information
- Application Number
- CN202510035229.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-09
AI Technical Summary
The existing SSVEP decoding algorithm is difficult to generalize effectively under cross-participants conditions. Individual physiological differences lead to predictive confusion when the model is classified in the target domain, which reduces the accuracy of the classification.
DAF-SSVEP frequency identification framework based on domain adversarial network is adopted to reduce the differences between the source and target domains by introducing MMD losses, and to reduce predicted confusion between correct and fuzzy categories in the target domain using MCC losses, thereby improving migration capabilities.
This improves the classification accuracy of the model in the target domain, reduces prediction confusion, and enhances the performance of the SSVEP brain-computer interface.
Smart Images

Figure CN119961750A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of brain-computer interface, and in particular relates to a SSVEP frequency recognition method based on a domain adversarial network. Background Art
[0002] Brain-computer interface (BCI) is a device that can transmit information between external devices and the human brain. It measures brain activity and converts it into specific instructions to achieve interaction between the brain and external devices (Schirrmeister RT, Springenberg JT, Fiederer LDJ, et al. Human brain mapping, 2017, 38(11): 5391-5420.). SSVEP has become one of the mainstream paradigms of BCI systems due to its stable frequency characteristics and high information transfer rate (ITR) (Lee MH, Kwon OY, KimY J, et al. GigaScience, 2019, 8(5): giz002.).
[0003] Designing an efficient SSVEP decoding algorithm is a key task in developing high-performance SSVEP-BCI systems (Yang ZY, Liu ZH, Zhang YN, et al. Parasites & vectors, 2021, 14: 1-14.). So far, researchers have proposed many SSVEP decoding algorithms. Deep learning has developed rapidly in recent years, and deep learning has also been introduced into the field of SSVEP-BCI. The deep learning decoding algorithms proposed by some researchers have shown excellent classification performance in the field of SSVEP-BCI. For example, the EEGNet model proposed by Waytowich et al. (Waytowich N, Lawhern VJ, Garcia JO, et al. Journal of neural engineering, 2018, 15 (6): 066031.) extracts information from each channel through time domain convolution and spatial convolution. Ravi proposed the CCNN model, which uses complex spectral representation of signals to further explore signal features and improve the classification of SSVEP (Ravi, Aravind, Beni, Nargess Heydari, Manuel, Jacob, & Jiang, Ning (2020). Journal of Neural Engineering, 17 (2), Article 026028.). Chen et al. first applied Transformer to the field of SSVEP-BCI and proposed SSVEPformer and FBSSVEPformer (Chen J, Zhang Y, Pan Y, et al. Neural Networks, 2023, 164: 521-534.). SSVEPformer uses the spectral features of SSVEP data as network input and uses convolutional attention modules and channel MLP modules to encode SSVEP features. FBSSVEPformer filters the input data using multiple filter groups and then uses the SSVEPformer model for feature encoding. Data in different frequency bands can extract more useful information, which greatly improves the recognition performance.
[0004] Under cross-subject conditions, due to individual physiological differences, it is still challenging to design a decoding algorithm that can generalize from one subject to other subjects. To solve this problem, researchers have introduced domain adaptation technology to design various algorithms. Domain adaptation is widely used in the field of image processing, such as domain adversarial training of neural networks (Ganin Y, Ustinova E, Ajakan H, et al. Journal of machine learning research, 2016, 17 (59): 1-35.), which achieves alignment of source domain and target domain features through mutual adversarial training of feature extractors and domain discriminators, thereby improving the performance of target domain tasks. In the field of emotion recognition, Huang et al. proposed DA-Tsnet, which uses convolutional neural networks and domain adversarial training strategies (Huang H, Guan Z, Pan J, et al. IEEE Transactions on Biomedical Engineering, 2024.) In the field of SSVEP, Liu et al. proposed ALPHA, a transfer learning framework, which performs domain adaptation by aligning spatial patterns and covariances (Liu B, Chen X, Li X, et al. IEEE Transactions on Biomedical Engineering, 2021, 69(2): 795-806.).
[0005] Among the existing SSVEP decoding algorithms using domain adaptation, most rely only on domain alignment strategies to reduce the differences between the source domain and the target domain. However, this approach has limitations. They fail to fully consider the problem of prediction confusion between different categories in the target domain. Although domain alignment helps to make the source domain and the target domain more similar in feature space, this is not enough to ensure that the model can clearly distinguish between categories when classifying the target domain, which leads to prediction confusion and reduces the accuracy of classification. Therefore, this patent uses a domain adaptation algorithm to design a new SSVEP frequency recognition framework DAF-SSVEP. Summary of the invention
[0006] The domain adversarial network-based SSVEP frequency recognition framework DAF-SSVEP provided by the present invention uses a domain adversarial network architecture under cross-subject conditions, introduces the MMD loss to further reduce the difference between the source domain and the target domain, and uses the MCC loss to reduce the prediction confusion between the correct category and the fuzzy category in the target domain, thereby improving the migration ability, thereby overcoming the shortcomings of the existing algorithms and achieving better decoding effects.
[0007] In order to achieve the above-mentioned invention object, the technical solution adopted by the present invention is: an SSVEP frequency recognition framework based on a domain adversarial network, comprising the following steps:
[0008] S1. Construct a feature extractor to extract the features of source domain data and target domain data, and use MMD loss to perform domain alignment to reduce the difference in features between the source domain and the target domain.
[0009] S2. Construct a classifier to classify and predict the source domain data features, use the minimum class confusion loss to optimize the target domain, and enhance the classification effect of the target domain data;
[0010] S3. Construct a domain discriminator to classify the target domain and source domain features, introduce a gradient reversal layer to reduce the distribution difference between the target domain and the source domain, and use negative log-likelihood loss to optimize the domain discriminator so that it can distinguish the source domain from the target samples as accurately as possible.
[0011] Furthermore, the feature extractor of S1 can be constructed using SSVEPformer, EEGNet or CCNN model, denoted by g(·).
[0012] Furthermore, in S1, the source domain and target domain data x s 、x t Input g(·) for feature extraction, the size becomes (1,128), and the extracted target domain feature g s and source domain features g t Domain alignment is performed to reduce the difference between the two domains. Specifically, MMD is used to calculate the difference in feature distribution between the two domains, and this is used as the loss function to optimize the model. The alignment effect is achieved by minimizing the MMD loss of source domain features and target domain features as much as possible. The MMD loss is expressed as follows:
[0013]
[0014] The features are projected into the same subspace through the ψ(·) mapping function, and the distribution difference is measured by calculating the mean difference between the two domains in the subspace.
[0015] Furthermore, the classifier of S2 is composed of a Linear layer, a LayerNorm layer, a GELU layer, and a Dropout layer. The feature (1, 128) is passed through the classifier to output the classification prediction probability (1, class_num), where class_num is the number of classifications. In order to optimize the classification performance of the model, the cross entropy loss function is used as the objective function, which is expressed as follows:
[0016]
[0017] Among them, A represents the number of samples, C represents the number of categories, and f ij is the true label of sample i in category j, f' ijThe model is optimized by minimizing the cross entropy loss to make the predicted probability distribution closer to the true label.
[0018] Furthermore, the classifier of S2 uses the minimum class confusion loss (MCC) for the target domain data to increase the inter-class boundary by minimizing the inter-class similarity in the feature space. Specifically, the features of the target domain are output through the classifier. Where N is the number of samples, C is the number of categories, and Q is temperature normalized to obtain T is the temperature hyperparameter, which is then calculated The weighted covariance matrix of and normalize it to get the correlation matrix Q between categories norm , and finally calculate the sum of the off-diagonal elements in the covariance and compare it with the sum of the diagonal elements. The MCC loss is as follows:
[0019]
[0020] Where Tr(·) represents the trace of the matrix.
[0021] Furthermore, the domain discriminator of S3 is composed of two layers of Linear layers, and a gradient reversal layer (GRL) is used between the feature extractor and the domain discriminator. Specifically, the features output by the encoder are converted from (1,128) to (1,72) through a linear layer and ReLU activation function, and then to (1,32) through a linear layer and ReLU, and finally to (1,2) through a linear layer and LogSoftmax to discriminate whether the features belong to the source domain or the target domain. The output of the domain discriminator is the predicted probability of the domain label, and the negative log-likelihood loss (NLLLoss) is used as the optimization target, which is expressed as follows:
[0022]
[0023] Among them, A is the number of samples, y i represents the true domain label of the i-th sample, y i =0 is the source domain, y i =1 is the target domain, P(y i |x i ) is the predicted probability of the correct domain label.
[0024] Furthermore, the trained model is The total losses are jointly supervised and the total losses are as follows:
[0025]
[0026] The test data is extracted features by the feature extractor and then sent to the classifier to output the final classification prediction to obtain the classification result.
[0027] The beneficial effects of the present invention are:
[0028] The DAF-SSVEP framework proposed in the present invention further reduces the difference between the source domain and the target domain by introducing the MMD loss, and uses the MCC loss to reduce the prediction confusion between the correct category and the fuzzy category in the target domain, thereby improving the migration ability. When classifying the target domain, the model can clearly distinguish between the categories, reduce prediction confusion, and improve the accuracy of classification. Thereby overcoming the shortcomings of the existing algorithms and achieving a better decoding effect. In addition, the experimental results of the present invention verify the effectiveness of the DAF-SSVEP framework in the SSVEP frequency recognition method, and can further enhance the performance of the SSVEP brain-computer interface on the existing basis. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 Flowchart of the SSVEP frequency recognition framework based on the domain adversarial network of the present invention.
[0030] Figure 2 It is a schematic diagram of the structure of the DAF-SSVEP framework of the present invention. DETAILED DESCRIPTION
[0031] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.
[0032] Embodiment 1:
[0033] like Figure 1 As shown, in one embodiment of the present invention, the SSVEP frequency recognition framework based on the domain adversarial network includes the following steps:
[0034] S1. Construct a feature extractor to extract the features of source domain data and target domain data, and use MMD loss to perform domain alignment to reduce the difference in features between the source domain and the target domain.
[0035] S2. Construct a classifier to classify and predict the source domain data features, use the minimum class confusion loss to optimize the target domain, and enhance the classification effect of the target domain data;
[0036] S3. Construct a domain discriminator to classify the target domain and source domain features, introduce a gradient reversal layer to reduce the distribution difference between the target domain and the source domain, and use negative log-likelihood loss to optimize the domain discriminator so that it can distinguish the source domain from the target samples as accurately as possible.
[0037] Furthermore, the feature extractor of S1 can be constructed using SSVEPformer, EEGNet or CCNN model, denoted by g(·).
[0038] Furthermore, in S1, the source domain and target domain data x s 、x t Input g(·) for feature extraction, the size becomes (1,128), and the extracted target domain feature g s and source domain features g t Domain alignment is performed to reduce the difference between the two domains. Specifically, MMD is used to calculate the difference in feature distribution between the two domains, and this is used as the loss function to optimize the model. The alignment effect is achieved by minimizing the MMD loss of source domain features and target domain features as much as possible. The MMD loss is expressed as follows:
[0039]
[0040] The features are projected into the same subspace through the ψ(·) mapping function, and the distribution difference is measured by calculating the mean difference between the two domains in the subspace.
[0041] Furthermore, the classifier of S2 is composed of a Linear layer, a LayerNorm layer, a GELU layer, and a Dropout layer. The feature (1, 128) is passed through the classifier to output the classification prediction probability (1, class_num), where class_num is the number of classifications. In order to optimize the classification performance of the model, the cross entropy loss function is used as the objective function, which is expressed as follows:
[0042]
[0043] Among them, A represents the number of samples, C represents the number of categories, and f ij is the true label of sample i in category j, f' ij The model is optimized by minimizing the cross entropy loss to make the predicted probability distribution closer to the true label.
[0044] Furthermore, the classifier of S2 uses the minimum class confusion loss (MCC) for the target domain data to increase the inter-class boundary by minimizing the inter-class similarity in the feature space. Specifically, the features of the target domain are output through the classifier. Where N is the number of samples, C is the number of categories, and Q is temperature normalized to obtain T is the temperature hyperparameter, which is then calculated The weighted covariance matrix of and normalize it to get the correlation matrix Q between categories norm , and finally calculate the sum of the off-diagonal elements in the covariance and compare it with the sum of the diagonal elements. The MCC loss is as follows:
[0045]
[0046] Where Tr(·) represents the trace of the matrix.
[0047] Furthermore, the domain discriminator of S3 is composed of two layers of Linear layers, and a gradient reversal layer (GRL) is used between the feature extractor and the domain discriminator. Specifically, the features output by the encoder are converted from (1,128) to (1,72) through a linear layer and ReLU activation function, and then to (1,32) through a linear layer and ReLU, and finally to (1,2) through a linear layer and LogSoftmax to discriminate whether the features belong to the source domain or the target domain. The output of the domain discriminator is the predicted probability of the domain label, and the negative log-likelihood loss (NLLLoss) is used as the optimization target, which is expressed as follows:
[0048]
[0049] Among them, A is the number of samples, y i represents the true domain label of the i-th sample, y i =0 is the source domain, y i =1 is the target domain, P(y i |x i ) is the predicted probability of the correct domain label.
[0050] Furthermore, the trained model is The total losses are jointly supervised and the total losses are as follows:
[0051]
[0052] The test data is extracted features by the feature extractor and then sent to the classifier to output the final classification prediction to obtain the classification result.
[0053] Embodiment 2:
[0054] This example is a specific experiment provided in Example 1.
[0055] The experimental data set is a public data set. The data set is a 12-classification experiment of 10 subjects. The experiment contains 15 trials, each trial is performed 12 times, and each time a stimulus is selected from 12 visual stimuli. EEG data is collected from 8 electrode areas, each trial lasts 6 seconds, and the collected data is finally downsampled to 256Hz. In the process of evaluating the performance of the model, we use the leave-one-out method for experimental verification. Specifically, the experiment is repeated 10 times, and each time one of the 10 subjects is selected as a test set, and the remaining subjects are used as training sets. Among them, the test set divides the trial number into target domain and test data according to 9 / 6, that is, 9 trial data are target domains, and 6 trial data are test data. Finally, the average value of the 10 classification accuracy rates is used as the final classification performance indicator of the algorithm. In the experiment, we use a time window of 0.5s-1s, a step size of 0.1s, and 6 different experimental data lengths to evaluate the classification performance of the algorithm under different time windows. All experiments of the present invention are completed on a desktop computer Linux system. All models were implemented using Python 3.7.9 and Pytorch1.11.0 deep learning open source frameworks, and the GeForce RTX 3090 (24G) graphics card was used to complete model training and testing. In the experiment, the optimizer used was Adam, and the learning rate was 0.001. The data batch_size was set to 128, the epochs was set to 500 rounds, and the training iterations were carried out until the validation set and the training set loss became stable.
[0056] In order to fully demonstrate the effectiveness of the present invention, the experiment used two indicators, accuracy and information translation rate (ITR), for performance evaluation.
[0057] Accuracy: Accuracy refers to the ratio of the number of samples (T) predicted correctly by the model to the total number of samples (C), which can be expressed as follows:
[0058] Information transmission rate: the amount of information transmitted per unit time, expressed as:
[0059] Where T represents a single target response time, and the interval time of 0.5s between two trials is added to the parameter T. When the data length is 1s, T = 1.5s. B is the accuracy of frequency recognition. C is the number of classification categories, C = 12. According to the above indicators, the classification results under different time window conditions are shown in Table 1. In order to verify the performance and effectiveness of the network model constructed by the framework of the present invention, we use SSVEPformer as a feature extractor and compare it with current mainstream methods such as EEGNet, CCNN and SSVEPformer. The test data of the comparison method uses the test data divided by the present invention. The results in Table 1 show the performance of the DAF-SSVEP framework constructed by the present invention and other methods in terms of accuracy indicators. Table 2 shows the performance of the DAF-SSVEP framework constructed by the present invention and other models in terms of information transmission rate indicators. As can be seen from the table, the framework constructed by the present invention surpasses other models in terms of accuracy and ITR, which proves that the SSVEP frequency recognition framework proposed in the present invention uses domain adaptation technology and effectively utilizes unlabeled target domain EEG data to provide a feasible technical solution to the poor generalization of frequency recognition caused by individual differences, providing effective technical support for the further expansion of SSVEP technology in practical applications.
[0060] Table 1: Classification accuracy of different models (%)
[0061]
[0062] Table 2: Average information transmission rate (bits / min) of different models
[0063]
Claims
1. The SSVEP frequency recognition framework based on domain adversarial network is characterized by: The following steps are involved: S1. Construct a feature extractor to extract the features of source domain data and target domain data, and use MMD loss to perform domain alignment to reduce the difference in features between the source domain and the target domain. S2. Construct a classifier to classify and predict the source domain data features, use the minimum class confusion loss to optimize the target domain, and enhance the classification effect of the target domain data; S3. Construct a domain discriminator to perform domain classification on the target domain and source domain features, introduce a gradient reversal layer to reduce the distribution difference between the target domain and source domain data, and use negative log-likelihood loss to optimize the domain discriminator so that it can distinguish the source domain from the target domain data as accurately as possible.
2. The SSVEP frequency identification framework based on domain adversarial network according to claim 1 is characterized in that: The feature extractor of S1 can be constructed using SSVEPformer, EEGNet or CCNN model, and is represented by g(·).
3. The SSVEP frequency identification framework based on the domain adversarial network architecture according to claim 2 is characterized in that: The source domain and target domain data x s 、x t Input g(·) for feature extraction, the size becomes (1,128), and the extracted target domain feature g s and source domain features g t Domain alignment is performed to reduce the difference between the two domain features. Specifically, MMD is used to calculate the distribution difference of the two domain features, and this is used as the loss function to optimize the model. The alignment effect is achieved by minimizing the MMD loss of the source domain features and the target domain features as much as possible. The MMD loss is expressed as follows: The features are projected into the same subspace through the ψ(·) mapping function, and the distribution difference is measured by calculating the mean difference between the two domains in the subspace.
4. The SSVEP frequency identification framework based on domain adversarial network according to claim 1 is characterized in that: The classifier of S2 is composed of Linear layer, LayerNorm layer, GELU layer and Dropout layer. The feature (1,128) is passed through the classifier to output the classification prediction probability (1,class_num), where class_num is the number of classifications. In order to optimize the classification performance of the model, the cross entropy loss function is used as the target loss function, which is expressed as follows: Among them, A represents the number of samples, C represents the number of categories, and f ij is the true label of sample i in category j, f' ij The probability predicted by the model is optimized by minimizing the cross entropy loss to make the predicted result closer to the true label.
5. The SSVEP frequency identification framework based on domain adversarial network according to claim 1 is characterized in that The classifier of S2 uses the minimum class confusion loss (MCC) for the target domain data to increase the inter-class boundary by minimizing the inter-class similarity in the feature space. Specifically, the features of the target domain are output through the classifier. Where N is the number of samples, C is the number of categories, and Q is temperature normalized to obtain T is the temperature hyperparameter, which is then calculated The weighted covariance matrix of and normalize it to get the correlation matrix Q between categories norm , and finally calculate the sum of the off-diagonal elements in the covariance and compare it with the sum of the diagonal elements. The MCC loss is as follows: Where Tr(·) represents the trace of the matrix.
6. The SSVEP frequency identification framework based on domain adversarial network according to claim 1 is characterized in that: The domain discriminator of S3 is composed of two linear layers. A gradient reversal layer (GRL) is used between the feature extractor and the domain discriminator. Specifically, the features output by the encoder are converted from (1,128) to (1,72) through a linear layer and ReLU activation function, and then to (1,32) through a linear layer and ReLU. Finally, a linear layer and LogSoftmax are converted to (1,2) to discriminate whether the features belong to the source domain or the target domain. The output of the domain discriminator is the predicted probability of the domain label. The negative log-likelihood loss (NLLLoss) is used as the optimization target, which is expressed as follows: Among them, A is the number of samples, y i represents the true domain label of the i-th sample, y i =0 is the source domain, y i =1 is the target domain, P(y i |x i ) is the predicted probability of the correct domain label.
7. Finally, the trained model is The total losses are jointly supervised and the total losses are as follows: The test data is extracted features by the feature extractor and then sent to the classifier to output the final classification prediction to obtain the classification result.