A cross-subject electroencephalogram signal emotion recognition method based on an incremental learning strategy
Patent Information
- Application Number
- CN202411868577.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2044-12-18
AI Technical Summary
尽管CNN具有优势,但是由于EEG的非平稳和动态性质,单一大小的时间内核、浅层的神经网络不能有效地捕获在不同时间尺度和持续时间发生的情绪的神经处理;同时,受试者之间的EEG信号变异性限制了其实际使用和分类器的通用性,特别是在受试者独立的应用中;此外,由于EEG信号通道数量过多可能导致维度过高、数据量过大、消耗大量的计算时间
Smart Images

Figure CN119740077B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electroencephalogram (EEG) signal processing and pattern recognition, and in particular to a cross-individual EEG signal emotion recognition method based on an incremental learning strategy. Background Technology
[0002] Emotion is a fundamental element of daily human life, influencing decision-making, perception, interpersonal communication, and human intelligence. Emotion recognition plays a crucial role in the evaluation of cognitive behavioral therapy (CBT), emotion regulation therapy (ERT) / emotion-focused therapy (EFT), and the pharmacological treatment of emotion-related mental disorders such as generalized anxiety disorder (GAD) and depression. With the potential applications of artificial intelligence (AI) in CBT and EFT, enabling AI to recognize human emotions has attracted increasing interest from researchers in recent years.
[0003] Electroencephalography (EEG) is a widely used brain imaging technique that directly measures human brain activity. Several electrodes are placed on the surface of the human head to collect EEG signals. EEG has high temporal resolution, so it can capture different brain states in sub-second timescales. Brain-computer interface (BCI) systems can recognize human emotions through EEG with the help of machine learning and signal processing techniques.
[0004] Deep learning is a machine learning method based on artificial neural networks. By simulating how the human brain processes information, it can automatically learn features and patterns from large amounts of data, and is used for various complex tasks such as image recognition, speech processing, and natural language processing. Convolutional Neural Networks (CNNs) have shown promising results in BCI (Brain-Induced Computation) due to their ability to learn directly from EEG. Currently, researchers have proposed deep and shallow CNNs, called DCN and SCN, achieving good results. Despite the advantages of CNNs, due to the non-stationary and dynamic nature of EEG, single-size temporal kernels and shallow neural networks cannot effectively capture the neural processing of emotions occurring at different time scales and durations. Furthermore, the variability of EEG signals among subjects limits their practical use and the generality of the classifier, especially in applications involving individual subjects. In addition, an excessive number of EEG signal channels can lead to excessively high dimensionality, large data volume, and significant computational time consumption. Summary of the Invention
[0005] In view of the above problems, the purpose of this invention is to solve some of the problems in the prior art, or at least alleviate these problems.
[0006] The present invention is achieved by at least one of the following technical solutions.
[0007] A cross-individual EEG signal emotion recognition method based on an incremental learning strategy includes the following steps:
[0008] Step 1: Import the collected EEG data and perform preprocessing and data augmentation on the data. The preprocessing includes artifact removal, filtering, and EEG channel selection. The data augmentation includes sample segmentation and time-frequency transformation.
[0009] Step 2: Divide the dataset into the original training set, the incremental training set, and the test set. The original training set is used for basic training, the incremental training set is used for incremental training, and the test set is used to evaluate the performance of the physiological-emotion recognition model.
[0010] Step 3: Train the physiological-emotion recognition model (hereinafter referred to as TSN-LWF). The TSN-LWF model mainly consists of three modules: two parallel backbone networks (TSN1, TSN2), two parallel output branches (N1, N2), and an LWF-based classifier module. Training is divided into two stages. In the basic stage, the original training set is fed into the TSN-LWF model to train the parameters of TSN1 and N1 and save the converged pre-trained model. In the subsequent incremental stage, the output module N2 is randomly initialized, the parameters of the TSN1 module are transferred to TSN2, the TSN1 and N1 modules are fixed, and the incremental training set is input into the model, training only the TSN2 and N2 modules. The LWF-based classifier calculates the loss and optimizes the model hyperparameters, using a public dataset for model training.
[0011] Step 4: Output the results. Based on the TSN-LWF model trained in Step 3, input the test set and output the valence and arousal dimensions of the cross-individual EEG signal emotion recognition results for binary or multi-class tasks. In the binary or multi-class task, select 5 or the corresponding value as the threshold to classify each dimension into high / low levels.
[0012] Furthermore, the preprocessing includes artifact removal, filtering, and EEG channel selection, specifically including:
[0013] First, the baseline of the EEG signal x seconds before the test was removed, and the data was downsampled to 128 Hz. Then, independent component analysis was used to remove confounding electrooculogram (EOG) from the EEG signal. Finally, the EEG signal was filtered with a bandpass filter of 4-45 Hz to remove low-frequency and high-frequency noise interference.
[0014] Existing physiological research shows that the frontal region stores more emotional information and is most closely related to emotions. The FP1-FP2 and F3-F4 electrode pairs in the frontal region are located anteriorly, in an area of the forehead without hair obstruction, which is more conducive to placing dry electrodes to collect EEG signals. Therefore, the four EEG channels with the highest correlation with emotion recognition in the frontal region, namely FP1, FP2, F3, and F4, were selected.
[0015] The data augmentation includes sample segmentation and time-frequency transformation, as detailed below:
[0016] The preprocessed dataset contains n seconds of EEG signals; the EEG signals are segmented in the time domain using a 1-second sliding window to significantly increase the sample size.
[0017] For the input signal after sample segmentation, the short-time Fourier transform is used to transform it to the frequency domain to extract features. The formula for the short-time Fourier transform is as follows:
[0018]
[0019] in It's a window function. It's frequency. It is a time index.
[0020] Furthermore, the specific steps of partitioning the dataset include:
[0021] The dataset is split into an original training set, an incremental training set, and a test set. The original training set is obtained by mixing data from multiple individuals and has a large enough amount of data for pre-training. 5% of the data from the new individual is split to form the incremental training set for incremental training. The remaining 95% of the data from the new individual is split to form the test set to evaluate the model performance.
[0022] Furthermore, in step 3, two parallel backbone networks (TSN1, TSN2) are respectively connected to two parallel output modules (N1, N2). TSN1 and TSN2 have the same structure, and N1 and N2 have the same structure. TSN1 and TSN2 are composed of temporal convolutional blocks, spatial convolutional blocks, and classifier modules. N1 and N2 are both fully connected layers, and the output feature is m. The temporal convolutional blocks are used to fully extract different frequency domain information. The spatial convolutional blocks obtain global feature representation by performing convolution operations in the dimension of EEG channels. The classifier module is used to better obtain emotion classification results.
[0023] Furthermore, the temporal convolutional block consists of three parallel multi-scale temporal convolutional layers. Each layer first performs a convolution operation, with kernel sizes of (1, 0.5*f), (1, 0.25*f), and (1, 0.15*f), where f refers to the number of sampling points. After convolution, the layers pass through a ReLU activation function and an average pooling layer, and then the time-frequency features at different scales are connected along the feature dimension.
[0024] The spatial convolutional block consists of two parallel multi-scale spatial convolutional layers. Each layer first performs a convolution operation, with the kernel sizes being (0.5*c,1) and (1*c,1), where c refers to the number of electrodes. After convolution, the spatial features at different scales are connected along the feature dimension through the ReLU activation function, average pooling layer, and batch normalization.
[0025] The classifier module includes three ordinary convolutional layers, one depthwise separable convolutional layer, and two connection layers. The input vector first passes through three ordinary convolutional layers with the same kernel size, then through the depthwise separable convolutional layer, and finally through a fully connected layer to obtain the output vector. The kernel size of the ordinary convolutional layers is 1×3, with a stride of 1 and padding of 1, and the number of output channels is 32, 64, and 64 respectively. The depthwise separable convolutional layer consists of depthwise convolution and pointwise convolution. The depthwise convolution uses a 1x3 kernel, and the pointwise convolution uses 128 output channels. After each convolution operation and depthwise separable convolution operation in the classifier module, batch normalization, ReLU activation function, and max pooling are used to avoid model overfitting. The connection layer includes two fully connected layers with sizes of 512 and 128 respectively.
[0026] Furthermore, in step 3, the LWF-based classifier module is specifically as follows:
[0027] The LWF (Less Forgetting Learning) method refers to the already trained TSN1 and N1 models as teacher models, and TSN2 and N2 as student models. In the basic training phase, the teacher models are trained. In the subsequent incremental phase, to allow the student models to learn from the teacher models and achieve higher recognition accuracy when learning new data, the samples in the incremental training set must first pass through two parallel backbone networks and two parallel output modules. After prediction, a conventional loss is defined. And a loss due to conventional and distillation loss Combination loss function And update the model using the corresponding loss function at each stage.
[0028] Furthermore, the combined loss function is specifically as follows:
[0029] Combination loss function From conventional losses and distillation loss composition,
[0030]
[0031] in It is a network model. It is a collection of input data samples. These are the corresponding real labels. This is the output of the student model. It is the output of the teacher model. Yes, the relaxation coefficient. It is a parameter set from the teacher model. These are the student model parameters;
[0032] conventional losses To measure the degree to which a sample is correctly classified, a common method is the multinomial logistic loss:
[0033]
[0034] in, These are the corresponding real labels. This is the output of the student model.
[0035] Distillation loss KL divergence (KLD) is used to evaluate the approximation between the two sets of probability distributions. By optimizing KL divergence, the student model output is made closer to the recorded output of the teacher model, thus preserving knowledge and transferring it to the new model.
[0036]
[0037] in and These represent the outputs of the teacher model and the student model, respectively. It is the distillation temperature;
[0038] In the foundational stage, choose As a loss function, the parameters that need to be updated are those in TSN1 and N1. In the incremental phase, use As the loss function, the parameters of the output module N2 are randomly initialized. Migrate TSN1 module parameters To TSN2, fix the parameters of TSN1 and N1. And only update the parameters in TSN2 and N2, i.e. .
[0039] It should be understood that the description in the Summary of the Invention is not intended to limit the key or essential features of the embodiments of the present invention, nor is it intended to restrict the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0040] Figure 1 This is an overall flowchart of the present invention;
[0041] Figure 2 This is a diagram showing the overall framework of the TSN-LWF model of the present invention.
[0042] Figure 3 This is an overall framework diagram of the backbone network of the present invention; Detailed Implementation
[0043] To enable those skilled in the art to better understand the present invention, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0044] like Figure 1 , Figure 2 As shown in the figure, this embodiment presents a cross-individual EEG signal emotion recognition method based on an incremental learning strategy. The method includes the following steps:
[0045] Step 1: Import the collected EEG data and perform preprocessing and data augmentation on the data. The preprocessing includes artifact removal, filtering, and EEG channel selection. The data augmentation includes sample segmentation and time-frequency transformation.
[0046] The preprocessing includes artifact removal, filtering, and EEG channel selection, specifically including:
[0047] First, the baseline of the EEG signal x seconds before the test was removed, and the data was downsampled to 128 Hz. Then, independent component analysis was used to remove confounding electrooculogram (EOG) from the EEG signal. Finally, the EEG signal was filtered with a bandpass filter of 4-45 Hz to remove low-frequency and high-frequency noise interference.
[0048] Existing physiological research shows that the frontal region stores more emotional information and is most closely related to emotions. The FP1-FP2 and F3-F4 electrode pairs in the frontal region are located anteriorly, in an area of the forehead without hair obstruction, which is more conducive to placing dry electrodes to collect EEG signals. Therefore, the four EEG channels with the highest correlation with emotion recognition in the frontal region, namely FP1, FP2, F3, and F4, were selected.
[0049] The data augmentation includes sample segmentation and time-frequency transformation, as detailed below:
[0050] The preprocessed dataset contains n seconds of EEG signals; the EEG signals are segmented in the time domain using a 1-second sliding window to significantly increase the sample size.
[0051] For the input signal after sample segmentation, the short-time Fourier transform is used to transform it to the frequency domain to extract features. The formula for the short-time Fourier transform is as follows:
[0052]
[0053] in It's a window function. It's frequency. It is a time index.
[0054] As an example, in the publicly available DEAP dataset, x is set to 3. The preprocessed EEG signal dataset is 60 seconds long. The final number of samples after segmenting the EEG signal using a sliding window with a 1-second time window is 2400.
[0055] Step 2: Divide the dataset into the original training set, the incremental training set, and the test set. The original training set is used for basic training, the incremental training set is used for incremental training, and the test set is used to evaluate the performance of the physiological-emotion recognition model.
[0056] The original training set is obtained by mixing data from multiple individuals, and has a large enough amount of data for pre-training. 5% of the data from the new individual is split to form the incremental training set for incremental training. The remaining 95% of the data from the new individual is used to form the test set to evaluate the model performance.
[0057] Step 3: Train the physiological-emotion recognition model (hereinafter referred to as TSN-LWF). The TSN-LWF model mainly consists of three modules: two parallel backbone networks (TSN1, TSN2), two parallel output branches (N1, N2), and an LWF-based classifier module. Training is divided into two stages. In the basic stage, the original training set is fed into the TSN-LWF model to train the parameters of TSN1 and N1 and save the converged pre-trained model. In the subsequent incremental stage, the output module N2 is randomly initialized, the parameters of the TSN1 module are transferred to TSN2, the TSN1 and N1 modules are fixed, and the incremental training set is input into the model, training only the TSN2 and N2 modules. The LWF-based classifier calculates the loss and optimizes the model hyperparameters, using a public dataset for model training.
[0058] The system consists of two parallel backbone networks (TSN1 and TSN2) connected to two parallel output modules (N1 and N2). TSN1 and TSN2 have identical structures, as do N1 and N2. TSN1 and TSN2 are composed of temporal convolutional blocks, spatial convolutional blocks, and a classifier module. N1 and N2 are both fully connected layers, with an output feature of m. The temporal convolutional blocks are used to fully extract different frequency domain information. The spatial convolutional blocks obtain global feature representations by performing convolution operations in the dimension of EEG channels. The classifier module is used to better obtain emotion classification results.
[0059] like Figure 3 As shown, the temporal convolutional block consists of three parallel multi-scale temporal convolutional layers. Each layer first performs a convolution operation, with kernel sizes of (1, 0.5*f), (1, 0.25*f), and (1, 0.15*f), where f refers to the number of sampling points. After convolution, the layers pass through the ReLU activation function and an average pooling layer, and then the time-frequency features at different scales are connected along the feature dimension.
[0060] The spatial convolutional block consists of two parallel multi-scale spatial convolutional layers. Each layer first performs a convolution operation, with the kernel sizes being (0.5*c,1) and (1*c,1), where c refers to the number of electrodes. After convolution, the spatial features at different scales are connected along the feature dimension through the ReLU activation function, average pooling layer, and batch normalization.
[0061] The classifier module includes three ordinary convolutional layers, one depthwise separable convolutional layer, and two connection layers. The input vector first passes through three ordinary convolutional layers with the same kernel size, then through the depthwise separable convolutional layer, and finally through a fully connected layer to obtain the output vector. The kernel size of the ordinary convolutional layers is 1×3, with a stride of 1 and padding of 1, and the number of output channels is 32, 64, and 64 respectively. The depthwise separable convolutional layer consists of depthwise convolution and pointwise convolution. The depthwise convolution uses a 1x3 kernel, and the pointwise convolution uses 128 output channels. After each convolution operation and depthwise separable convolution operation in the classifier module, batch normalization, ReLU activation function, and max pooling are used to avoid model overfitting. The connection layer includes two fully connected layers with sizes of 512 and 128 respectively.
[0062] The LWF-based classifier module is as follows:
[0063] The LWF (Less Forgetting Learning) method refers to the already trained TSN1 and N1 models as teacher models, and TSN2 and N2 as student models. In the basic training phase, the teacher models are trained. In the subsequent incremental phase, to allow the student models to learn from the teacher models and achieve higher recognition accuracy when learning new data, the samples in the incremental training set must first pass through two parallel backbone networks and two parallel output modules. After prediction, a conventional loss is defined. And a loss due to conventional and distillation loss Combination loss function And update the model using the corresponding loss function at each stage.
[0064] The combined loss function is as follows:
[0065] Combination loss function From conventional losses and distillation loss composition,
[0066]
[0067] in It is a network model. It is a collection of input data samples. These are the corresponding real labels. This is the output of the student model. It is the output of the teacher model. Yes, the relaxation coefficient. It is a parameter set from the teacher model. These are the student model parameters;
[0068] conventional losses To measure the degree to which a sample is correctly classified, a common method is the multinomial logistic loss:
[0069]
[0070] in, These are the corresponding real labels. This is the output of the student model.
[0071] Distillation loss KL divergence (KLD) is used to evaluate the approximation between the two sets of probability distributions. By optimizing KL divergence, the student model output is made closer to the recorded output of the teacher model, thus preserving knowledge and transferring it to the new model.
[0072]
[0073] in and These represent the outputs of the teacher model and the student model, respectively. It is the distillation temperature;
[0074] In the foundational stage, choose As a loss function, the parameters that need to be updated are those in TSN1 and N1. In the incremental phase, use As the loss function, the parameters of the output module N2 are randomly initialized. Migrate TSN1 module parameters To TSN2, fix the parameters of TSN1 and N1. And only update the parameters in TSN2 and N2, i.e. .
[0075] As one example, in a binary classification task, the dimension of the input vector is... ,in This is the batch number, with a value of 32. This is the number of brainwave channels, with a value of 4. The number of sampling points is 128, therefore the dimension of the input vector is... The dimension of the output vector after the temporal convolution block is The dimension of the output vector after the spatial convolution block is The dimension of the output vector after passing through the classifier module and the output branch is .in, The value is 0.1, and T is 2.
[0076] Step 4: Output the results. Based on the TSN-LWF model trained in Step 3, input the test set and output the valence and arousal dimensions of the cross-individual EEG signal emotion recognition results for binary or multi-class tasks. In the binary or multi-class task, select 5 or the corresponding value as the threshold to classify each dimension into high / low levels.
[0077] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, enabling those skilled in the art to better understand and utilize the invention.
Claims
1. A cross-individual EEG signal emotion recognition method based on an incremental learning strategy, characterized in that, Includes the following steps: Step 1: Import the collected EEG data and perform preprocessing and data augmentation on the data. The preprocessing includes artifact removal, filtering, and EEG channel selection. The data augmentation includes sample segmentation and time-frequency transformation. Step 2: Divide the dataset into the original training set, the incremental training set, and the test set. The original training set is used for basic training, the incremental training set is used for incremental training, and the test set is used to evaluate the performance of the physiological-emotion recognition model. Step 3: Train the physiological-emotion recognition model, namely the TSN-LWF model. The TSN-LWF model consists of three modules: two parallel backbone networks TSN1 and TSN2, two parallel output branches N1 and N2, and an LWF-based classifier module. The two parallel backbone networks TSN1 and TSN2 are connected to the two parallel output modules N1 and N2, respectively. TSN1 and TSN2 have the same structure, as do N1 and N2. TSN1 and TSN2 consist of temporal convolutional blocks, spatial convolutional blocks, and a classifier module. N1 and N2 are both fully connected layers, and the output feature is m. The temporal convolutional blocks are used to fully extract different frequency domain information. The spatial convolutional block obtains global feature representation by performing convolution operations along the dimension of the EEG channels. The classifier module is used to better obtain emotion classification results. Training is divided into two stages. In the basic stage, the original training set is fed into the TSN-LWF model to train the parameters of TSN1 and N1 and save the converged pre-trained model. In the subsequent incremental stage, the output module N2 is randomly initialized, the parameters of the TSN1 module are transferred to TSN2, the TSN1 and N1 modules are fixed, the incremental training set is input into the model, and only the TSN2 and N2 modules are trained. The LWF-based classifier calculates the loss and optimizes the model hyperparameters, and uses a public dataset for model training. Step 4: Output the results. Based on the TSN-LWF model trained in Step 3, input the test set and output the valence and arousal dimensions of the cross-individual EEG signal emotion recognition results for binary or multi-class tasks. In the binary classification task, 5 is selected as the threshold to classify each dimension into high / low levels. In the multi-class task, a combination of each dimension is selected for classification.
2. The cross-individual EEG signal emotion recognition method based on incremental learning strategy according to claim 1, characterized in that: The preprocessing includes artifact removal, filtering, and EEG channel selection, specifically including: First, the baseline of the EEG signal x seconds before the test was removed, and the data was downsampled to 128 Hz. Then, independent component analysis was used to remove the electrooculogram (EOG) contamination from the EEG signal. Finally, a bandpass filter of 4-45 Hz was used to filter the EEG signal to remove low-frequency and high-frequency noise interference. Existing physiological research shows that the frontal region stores more emotional information and is most closely related to emotions. The FP1-FP2 and F3-F4 electrode pairs in the frontal region are located in the front area without hair covering, which is more conducive to placing dry electrodes to collect EEG signals. Therefore, the four electrode positions with the highest correlation with emotion recognition in the frontal region were selected, namely FP1, FP2, F3, and F4.
3. The cross-individual EEG signal emotion recognition method based on incremental learning strategy according to claim 1, characterized in that: The data augmentation includes sample segmentation and time-frequency transformation, as detailed below: The processed dataset consists of n seconds of EEG signals; the EEG signals are segmented in the time domain using a 1-second sliding window to significantly increase the sample size. For the input signal after sample segmentation, the short-time Fourier transform is used to transform it to the frequency domain to extract features. The formula for the short-time Fourier transform is as follows: in It's a window function. It's frequency. It is a time index.
4. The cross-individual EEG signal emotion recognition method based on incremental learning strategy according to claim 1, characterized in that, Step 2, specifically the partitioning of the dataset, includes: The dataset is split into an original training set, an incremental training set, and a test set. The original training set is obtained by mixing data from multiple individuals and has a large enough amount of data for pre-training. 5% of the data from the new individual is split to form the incremental training set for incremental training. The remaining 95% of the data from the new individual is split to form the test set to evaluate the model performance.
5. The cross-individual EEG signal emotion recognition method based on incremental learning strategy according to claim 1, characterized in that, The temporal convolutional block consists of three parallel multi-scale temporal convolutional layers. Each layer first performs a convolution operation with kernel sizes of (1, 0.5*f), (1, 0.25*f), and (1, 0.15*f), where f refers to the time sampling point. After convolution, the layers pass through the ReLU activation function and an average pooling layer, and then the time-frequency features at different scales are connected along the feature dimension. The spatial convolutional block consists of two parallel multi-scale spatial convolutional layers. Each layer first performs a convolution operation, with the kernel sizes being (0.5*c,1) and (1*c,1), where c refers to the number of electrodes. After convolution, the spatial features at different scales are connected along the feature dimension through the ReLU activation function, average pooling layer, and batch normalization. The dimension of the input vector is defined as follows: The dimension of the output vector is 1 is the number of convolutional channels, b is the number of batches, c is the number of electrodes, f is the time sampling point, ch is the number of features after the convolutional layer, and n is the number of features after the pooling operation. First, the temporal convolutional block is used to learn the time-frequency representation, and then the spatial convolutional block works in a cross-channel manner to obtain the spatial representation. The classifier module includes three ordinary convolutional layers, one depthwise separable convolutional layer, and two connection layers; The input vector first passes through three ordinary convolutional layers with the same kernel size, then through a depthwise separable convolutional layer, and finally through a fully connected layer to obtain the output vector. The kernel size of the ordinary convolutional layers is 1×3, with a stride of 1 and padding of 1, and the number of output channels is 32, 64, and 128 respectively. The depthwise separable convolutional layer consists of depthwise convolution and pointwise convolution. The depthwise convolution uses a 1x3 kernel, and the pointwise convolution uses 64 output channels. After each convolution operation and depthwise separable convolution operation in the classifier module, batch normalization, ReLU activation function, and max pooling are used to avoid model overfitting. The connection layer includes two fully connected layers with sizes of 512 and 128 respectively.
6. The cross-individual EEG signal emotion recognition method based on incremental learning strategy according to claim 1, characterized in that, In step 3, the LWF-based classifier module is as follows: The LWF method refers to the already trained TSN1 and N1 models as teacher models, and TSN2 and N2 as student models. In the basic training phase, the teacher models are trained. In the subsequent incremental phase, to allow the student models to learn from the teacher models and achieve higher recognition accuracy when learning new data, the samples in the incremental training set must first pass through two parallel backbone networks and two parallel output modules. After prediction, a conventional loss is defined. And a loss due to conventional and distillation loss Combination loss function And update the model using the corresponding loss function at each stage.
7. The cross-individual EEG signal emotion recognition method based on incremental learning strategy according to claim 6, characterized in that, The combined loss function is as follows: Combination loss function From conventional losses and distillation loss composition, in It is a network model. It is a collection of input data samples. These are the corresponding real labels. This is the output of the student model. It is the output of the teacher model. It is the relaxation coefficient. It is a parameter set from the teacher model. These are the student model parameters; conventional losses To measure the degree to which a sample is correctly classified, a common method is the multinomial logistic loss: in, These are the corresponding real labels. This is the output of the student model. Distillation loss KL divergence is used to evaluate the approximation between the two sets of probability distributions. By optimizing the KL divergence, the student model output is made closer to the recorded output of the teacher model, which will preserve knowledge and transfer it to the new model. in and These represent the outputs of the teacher model and the student model, respectively. It is the distillation temperature; In the foundational stage, choose As a loss function, the parameters that need to be updated are those in TSN1 and N1. In the incremental phase, use As the loss function, the parameters of the output module N2 are randomly initialized. Migrate TSN1 module parameters To TSN2, fix the parameters of TSN1 and N1. And only update the parameters in TSN2 and N2, i.e. .
Citation Information
Patent Citations
Electroencephalogram emotion recognition method and system based on multiple tasks and attention mechanism
CN118576206A
Systems and methods for automatic and incremental learning of patient states from biomedical signals
US20040199482A1