Emotion recognition model training method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202611014917.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]本申请实施例提供一种情绪识别模型的训练方法、装置、电子设备和存储介质,以解决现有情绪识别模型在跨被试脑电情绪识别场景下识别性能不佳的问题
[0016]基于本申请实施例的方法,通过在模型训练过程中实现了对有标签样本情绪类别学习、跨域特征分布对齐以及多分支特征协同性的联合优化,使得第一情绪识别模型不仅能够在样本脑电数据集上准确学习情绪分类知识,还能有效降低待测脑电数据集与样本脑电数据集之间的域偏移影响,并通过强制多分支输出一致性来增强特征表示的鲁棒性。相较于传统的单一损失函数训练方式,显著提升了模型在面对新被试脑电数据时的情绪识别准确率和泛化稳定性,为解决脑电信号情绪识别中的个体差异问题提供了有效的技术途径。
Smart Images

Figure CN122805268A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of deep learning technology, and in particular to a training method, apparatus, electronic device and storage medium for an emotion recognition model. Background Technology
[0002] Electroencephalogram (EEG) signals can directly and objectively reflect an individual's emotions at the level of neural activity. Moreover, the devices are portable and low-cost, making them an important research direction for emotion recognition. They have significant application value in fields such as affective computing, human-computer interaction, and mental health monitoring.
[0003] Cross-subject EEG emotion recognition is key to achieving large-scale deployment in real-world scenarios. While existing emotion recognition models have high accuracy within a single subject, significant individual differences in EEG signals lead to data distribution shifts when the models are directly applied to new subjects, resulting in a substantial decrease in recognition performance. Summary of the Invention
[0004] This application provides a training method, apparatus, electronic device, and storage medium for an emotion recognition model to address the problem of poor recognition performance of existing emotion recognition models in cross-subject EEG emotion recognition scenarios.
[0005] In a first aspect, embodiments of this application provide a method for training an emotion recognition model, comprising: Obtain the EEG dataset to be tested and multiple sample EEG datasets; the EEG dataset to be tested contains the EEG data of the subject to be tested, and the sample EEG datasets contain the EEG data of the subject to be tested and their true emotion labels. The neural network model to be trained is used to process the EEG data in the test EEG dataset and the sample EEG dataset to obtain the multi-branch latent features and emotion category membership distribution of each EEG data. The domain alignment loss between the EEG dataset to be tested and the sample EEG dataset is calculated based on the multi-branch latent features of each EEG dataset, and the branch consistency loss is calculated based on the emotion category membership distribution of each EEG dataset in the EEG dataset to be tested. The classification loss is calculated based on the emotion category membership distribution and true emotion labels of the EEG data in the sample EEG dataset. The first loss function is constructed based on domain alignment loss, branch consistency loss, and classification loss. The neural network model is then iteratively trained using the first loss function to obtain the first emotion recognition model.
[0006] In some embodiments, the EEG data in the test EEG dataset and the sample EEG dataset are processed based on the neural network model to be trained, including: inputting the EEG data in the test EEG dataset and the sample EEG dataset into the feature extraction layer of the neural network model to be trained, respectively, to extract the original features of each EEG data; inputting the original features into the multi-branch feature transformation layer of the neural network model, and performing feature mapping and transformation of the original features in different dimensions through multiple parallel branch networks to obtain the multi-branch latent features of each EEG data; inputting the multi-branch latent features into the classification layer of the neural network model, and processing the latent features of each branch through the softmax function to obtain the emotion category membership distribution of each EEG data corresponding to different emotion categories.
[0007] In some embodiments, the classification loss is calculated based on the emotion category membership distribution of EEG data in the sample EEG dataset and the true emotion label, including: comparing the emotion category membership distribution of each EEG data in the sample EEG dataset with the true emotion label corresponding to the EEG data, calculating the difference between the two using the cross-entropy loss function, and obtaining the classification loss value of each EEG data in the sample EEG dataset; and determining the classification loss based on the classification loss value of each EEG data in the sample EEG dataset.
[0008] In some embodiments, the domain alignment loss between the EEG dataset to be tested and the sample EEG dataset is calculated based on the multi-branch latent features of each EEG dataset, and the branch consistency loss is calculated based on the emotion category membership distribution of each EEG dataset. This includes: calculating the maximum mean difference loss for the latent features of the sample EEG dataset and the latent features of the EEG dataset to be tested corresponding to each branch output by the multi-branch feature transformation layer; determining the domain alignment loss based on the maximum mean difference loss of each branch; determining the distribution difference between each pair of branches based on the emotion category membership distribution of each EEG dataset in the EEG dataset to be tested output by each branch of the classification layer; and summing the distribution differences to obtain the branch consistency loss.
[0009] In some embodiments, acquiring a test EEG dataset and multiple sample EEG datasets includes: acquiring the test EEG dataset and multiple candidate sample EEG datasets; extracting first distribution feature statistics of the test EEG dataset and second distribution feature statistics of each candidate sample EEG dataset; determining the difference between the second distribution feature statistics of each candidate sample EEG dataset and the first distribution feature statistics, as the distribution difference degree corresponding to each candidate sample EEG dataset; and determining the sample EEG dataset from the multiple candidate sample EEG datasets based on the distribution difference degree corresponding to each candidate sample EEG dataset.
[0010] In some embodiments, the method further includes: inputting EEG data from the EEG dataset to be tested into a first emotion recognition model to obtain pseudo-emotion labels and confidence scores for each EEG data output by the first emotion recognition model; filtering the EEG data from the EEG dataset to be tested based on the pseudo-emotion labels and confidence scores for each EEG data, and constructing a target EEG dataset using the filtered EEG data and their corresponding pseudo-emotion labels; determining a global confidence coefficient based on the confidence scores corresponding to each EEG data in the target EEG dataset; performing emotion category matching on the EEG data in the target EEG dataset and the sample EEG dataset based on the pseudo-emotion labels for each EEG data in the target EEG dataset and the real emotion labels for each EEG data in the sample EEG dataset to obtain multiple EEG data subsets; constructing a local difference measure using the multiple EEG data subsets, and constructing a local alignment loss based on the local difference measure and the global confidence coefficient; constructing a second loss function based on the local alignment loss and the classification loss, and iteratively training the first emotion recognition model using the second loss function to obtain a second emotion recognition model.
[0011] In some embodiments, a local difference measure is constructed using multiple subsets of EEG data, and a local alignment loss is constructed based on the local difference measure and the global confidence coefficient. This includes: for any subset of EEG data, constructing a local subset difference measure based on intra-class aggregation constraints and inter-class separation constraints; wherein, the local subset difference measure is used to quantify the difference in feature distribution between the sample EEG dataset and the target EEG dataset under the same emotion category; summarizing the local subset difference measures corresponding to each subset of EEG data to obtain the local difference measure; and multiplying the global confidence coefficient and the local difference measure to obtain the local alignment loss.
[0012] Secondly, embodiments of this application provide a training apparatus for an emotion recognition model, comprising: The dataset acquisition module is used to acquire the EEG dataset to be tested and multiple sample EEG datasets; the EEG dataset to be tested contains the EEG data of the test subjects, and the sample EEG datasets contain the EEG data of the sample subjects and their true emotion labels. The first processing module is used to process the EEG data in the test EEG dataset and the sample EEG dataset based on the neural network model to be trained, and to obtain the multi-branch latent features and emotion category membership distribution of each EEG data. The second processing module is used to calculate the classification loss based on the emotion category membership distribution and true emotion labels of the EEG data in the sample EEG dataset; calculate the domain alignment loss between the EEG dataset to be tested and the sample EEG dataset based on the multi-branch latent features of each EEG data; and calculate the branch consistency loss based on the emotion category membership distribution of each EEG data in the EEG dataset to be tested. The model training module is used to construct a first loss function based on domain alignment loss, branch consistency loss, and classification loss. The first loss function is then used to iteratively train the neural network model to obtain the first emotion recognition model.
[0013] Thirdly, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor implements any of the methods of embodiments of this application when executing the computer program.
[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method of any one of the embodiments of this application.
[0015] Fifthly, embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements any of the implementation methods in the first aspect described above.
[0016] The method based on the embodiments of this application achieves joint optimization of labeled sample emotion category learning, cross-domain feature distribution alignment, and multi-branch feature synergy during model training. This enables the first emotion recognition model to not only accurately learn emotion classification knowledge on the sample EEG dataset but also effectively reduce the domain offset influence between the test EEG dataset and the sample EEG dataset. Furthermore, it enhances the robustness of feature representation by forcing multi-branch output consistency. Compared to traditional single loss function training methods, this significantly improves the model's emotion recognition accuracy and generalization stability when facing new subject EEG data, providing an effective technical approach to address the individual differences problem in EEG signal emotion recognition.
[0017] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application, it can be implemented according to the contents of the specification. In order to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description
[0018] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this application. Wherein: Figure 1 A flowchart illustrating the training method for an emotion recognition model provided in an exemplary embodiment of this application. Figure 1 ; Figure 2 This is a flowchart illustrating the training method of an emotion recognition model provided in an exemplary embodiment of this application. Figure 2 ; Figure 3This is a flowchart illustrating the training method of an emotion recognition model provided in an exemplary embodiment of this application. Figure 3 ; Figure 4 This is a flowchart illustrating the training method of an emotion recognition model provided in an exemplary embodiment of this application. Figure 4 ; Figure 5 This is a flowchart illustrating the training method of an emotion recognition model provided in an exemplary embodiment of this application. Figure 5 ; Figure 6 This is a schematic diagram of a training apparatus for an emotion recognition model provided in an exemplary embodiment of this application; Figure 7 This is a schematic diagram of the internal structure of an electronic device provided in an exemplary embodiment of this application. Detailed Implementation
[0019] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the concept or scope of this application. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.
[0020] To facilitate understanding of the technical solutions of the embodiments of this application, the relevant technologies of the embodiments of this application are described below. The following relevant technologies are optional solutions and can be combined with the technical solutions of the embodiments of this application in any way, and all of them fall within the protection scope of the embodiments of this application.
[0021] Figure 1 This is a flowchart illustrating a training method for an emotion recognition model provided in an exemplary embodiment of this application. This embodiment can be applied to electronic devices, such as... Figure 1 As shown, in some embodiments, the training method for the emotion recognition model provided in this application includes steps S110-S150: Step S110: Obtain the EEG dataset to be tested and multiple sample EEG datasets.
[0022] The test EEG dataset contains the EEG data of the test subjects, while the sample EEG dataset contains the EEG data of the sample subjects and their true emotion labels. The EEG data can be obtained by recording with EEG acquisition devices under different emotional induction conditions, such as by viewing emotional pictures or videos or performing specific emotional tasks. The true emotion labels are emotion category identifiers determined according to the experimental design or the subjective reports of the subjects, such as positive, negative, neutral, etc., or specific emotion types such as joy, calm, sadness, etc.
[0023] In some embodiments, such as Figure 2As shown, step S110 involves acquiring the EEG dataset to be tested and multiple sample EEG datasets, including the following steps S201-S204: Step S201: Obtain the EEG dataset to be tested and multiple candidate sample EEG datasets.
[0024] Among them, the EEG dataset to be tested consists of EEG data collected from new subjects under emotional induction conditions, without emotion labels; the multiple candidate sample EEG dataset consists of EEG data from multiple different non-target subjects, all with corresponding real emotion labels. Both types of data undergo the same preprocessing and feature extraction to form EEG feature samples of the same dimension, ensuring the effectiveness of subsequent distribution comparison.
[0025] For example, the preprocessing of EEG data may include: downsampling the raw EEG data to 200Hz to unify the data sampling scale; performing a 0.75Hz bandpass filter on the downsampled EEG data to retain the effective EEG signal frequency band; sequentially performing power frequency interference suppression and artifact removal on the filtered EEG data to improve the signal-to-noise ratio and stability of the EEG signal; segmenting the preprocessed EEG signal into 1s non-overlapping windows and extracting DE features (differential entropy features); normalizing or standardizing the extracted DE features to construct feature samples of a uniform scale, ultimately forming the EEG dataset to be tested and multiple candidate sample EEG datasets.
[0026] Step S202: Extract the first distribution feature statistics of the EEG dataset to be tested and the second distribution feature statistics of the EEG dataset of each candidate sample.
[0027] Based on the preprocessed EEG feature samples, for the EEG dataset to be tested, statistical information that can characterize its overall feature distribution pattern is extracted, denoted as the first distribution feature statistical information. For each candidate sample EEG dataset, the same extraction method as the dataset to be tested is used to obtain the distribution feature statistical information corresponding to each candidate set, denoted as the second distribution feature statistical information. By unifying the extraction rules, it is ensured that the distribution feature statistical information of the test domain and all candidate source domains are on the same comparison dimension.
[0028] Step S203: Determine the difference between the second distribution feature statistics of each candidate sample EEG dataset and the first distribution feature statistics, and use it as the distribution difference degree corresponding to each candidate sample EEG dataset.
[0029] To quantify the distributional shift between the test EEG dataset and the candidate EEG datasets, the Jensen-Shannon (JS) divergence can be used as a measure of distributional dissimilarity. The JS divergence is calculated by comparing the second distributional feature statistics of each candidate EEG dataset with the first distributional feature statistics of the test EEG dataset. The result represents the distributional dissimilarity of that candidate EEG dataset. The magnitude of the JS divergence value is positively correlated with the distributional dissimilarity; a smaller divergence value indicates a more similar feature distribution between the candidate source domain and the test domain, resulting in better cross-subject fit.
[0030] Step S204: Based on the distribution difference degree corresponding to each candidate sample EEG dataset, determine the sample EEG dataset from multiple candidate sample EEG datasets.
[0031] For example, the candidate sample EEG datasets can be sorted and filtered according to the distribution difference of each candidate sample EEG dataset, and a dual filtering rule can be set: (1) the candidate sets with JS divergence less than the preset threshold, or those ranked in the top few in the distribution difference ranking among all candidate sample datasets, are included in the final sample EEG dataset; (2) the candidate sets with large JS divergence and high distribution difference are eliminated to avoid incompatible sample EEG data from interfering with the cross-domain alignment and emotion recognition performance of the subsequent model.
[0032] In this embodiment, by calculating and filtering the distribution difference of candidate sample EEG datasets, sample EEG data that is closer in distribution to the EEG dataset to be tested can be selected from multiple candidate sample EEG datasets, providing high-quality auxiliary data for subsequent model training and effectively reducing the problem of insufficient model generalization ability caused by excessive data distribution differences.
[0033] Step S120: Based on the neural network model to be trained, process the EEG data in the test EEG dataset and the sample EEG dataset to obtain the multi-branch latent features and emotion category affiliation distribution of each EEG data.
[0034] For example, the neural network model to be trained can adopt an architecture including a shared feature extraction layer and a multi-branch classification layer. The shared feature extraction layer is used for deep feature mining of the input EEG data and can be composed of network structures such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), or Transformers. For instance, a combination of three one-dimensional convolutional layers and two LSTM layers can be used to jointly extract the time-series features and spatial distribution features of the EEG data, obtaining a basic feature vector with strong representational capabilities. The multi-branch classification layer sets up multiple branches in parallel after the shared feature extraction layer. Each branch contains an independent fully connected layer and a softmax activation function, used to map the basic feature vector to the emotion category space.
[0035] In some embodiments, such as Figure 3 As shown, step S120 processes the EEG data in the test EEG dataset and the sample EEG dataset based on the neural network model to be trained, including steps S301-S303: Step S301: Input the EEG data from the EEG dataset to be tested and the sample EEG dataset into the feature extraction layer of the neural network model to be trained, respectively, in order to extract the original features of each EEG data. Step S302: Input the original features into the multi-branch feature transformation layer of the neural network model. Through multiple parallel branch networks, perform feature mapping and transformation of the original features in different dimensions to obtain the multi-branch potential features of each EEG data. Step S303: Input the multi-branch latent features into the classification layer of the neural network model, and process the latent features of each branch through the softmax function to obtain the emotion category membership distribution of each EEG data corresponding to different emotion categories.
[0036] For example, when the EEG data to be tested and the sample EEG data are input into the neural network model, the shared feature extraction layer first performs unified feature encoding on all data to generate corresponding basic feature vectors. Subsequently, the fully connected layer of each branch performs further nonlinear transformation and dimensionality mapping on the basic feature vectors, and outputs the probability distribution of each EEG data belonging to different emotion categories under that branch through the Softmax function, i.e., the emotion category membership distribution. At the same time, the intermediate features generated by each branch during the nonlinear transformation process are the multi-branch latent features of each EEG data. These latent features characterize the emotion-related patterns of the EEG data from different perspectives.
[0037] The method employed in this embodiment, through a multi-branch structure, can extract emotional feature representations from EEG data across different dimensions. This avoids the information loss that may occur with a single feature extraction path, providing a richer feature foundation for subsequent loss calculation and model optimization. Multi-branch latent features can capture emotional association information in EEG signals from multiple perspectives, including temporal dynamic changes, spatial lead distribution, and frequency band energy differences. Furthermore, the emotional category membership distribution provides a direct basis for calculating classification loss and branch consistency loss, enabling the model to simultaneously focus on feature discriminativeness and inter-branch synergy during training.
[0038] Step S130: Calculate the classification loss based on the emotion category membership distribution and the true emotion label of the EEG data in the sample EEG dataset.
[0039] In some embodiments, such as Figure 4As shown, step S130 calculates the classification loss based on the emotion category membership distribution and the true emotion label of the EEG data in the sample EEG dataset, specifically including the following steps S401-S402: Step S401: Compare the emotion category membership distribution of each EEG data in the sample EEG dataset with the real emotion label corresponding to the EEG data, and use the cross-entropy loss function to calculate the difference between the two to obtain the classification loss value of each EEG data in the sample EEG dataset.
[0040] For example, each branch can be set in the feature space. There are 10 adjustable rules, each consisting of a center vector. and scale vector Representation, for the input feature vector The Dimensional components Its Gaussian membership degree It can be represented as:
[0041] For example, the membership degrees of each dimension can be aggregated to obtain the activation strength of each rule. A linear transformation of the rule activation vector followed by the Softmax function can then be applied to output the activation strength of the rule. Distribution of emotion category membership degrees among branches Based on this, for the emotion category membership distribution of each EEG data point in the sample EEG dataset, cross-entropy loss can be used to supervise the training of the classification output. The classification loss... It can be represented as:
[0042] Step S402: Determine the classification loss based on the classification loss value of each EEG data in the sample EEG dataset.
[0043] For example, the classification loss values of all branches can be weighted and summed to obtain the final classification loss. For instance, each branch can be assigned the same weight coefficient, and the cross-entropy loss values of each branch can be directly accumulated and averaged to obtain the overall classification loss; or the weights can be dynamically adjusted based on the initial performance of each branch on the validation set, giving higher weights to branches with better classification performance to enhance the model's focus on learning effective features.
[0044] The method employed in this embodiment calculates cross-entropy loss on a sample-by-sample, multi-branch basis by comparing the emotion category membership distribution of sample EEG data with the actual emotion labels, and then summarizes the results using a branch weighting strategy. This effectively supervises the model's learning of emotion categories for labeled samples, ensuring that the model possesses basic emotion classification capabilities on the sample data. Simultaneously, the classification loss calculation under the multi-branch structure allows each branch to independently receive label supervision, providing a foundation for subsequent feature collaboration and consistency optimization between branches, and avoiding overfitting or feature bias issues that may arise from a single branch.
[0045] Step S140: Calculate the domain alignment loss between the EEG dataset to be tested and the sample EEG dataset based on the multi-branch latent features of each EEG dataset, and calculate the branch consistency loss based on the emotion category membership distribution of each EEG dataset in the EEG dataset to be tested.
[0046] In some embodiments, in step S140, the maximum mean difference loss can be calculated for the latent features of the sample EEG dataset and the latent features of the EEG dataset to be tested corresponding to each branch output by the multi-branch feature conversion layer.
[0047] For example, for the latent feature set of the kth branch of the sample EEG dataset and the latent feature set of the kth branch of the EEG dataset to be tested, the maximum mean discrepancy (MMD) loss is calculated, which is used to measure the distance between the two distributions in the feature space.
[0048] Subsequently, the MMD losses of all K branches are weighted and summed to obtain the domain alignment loss. The domain alignment loss aims to promote the model to learn domain-invariant emotion-related features and improve the model's generalization ability to new subjects (test subjects) by minimizing the distribution difference between the sample EEG data and the test EEG data in the latent feature space of each branch.
[0049] In some embodiments, in step S140, the domain alignment loss can be determined based on the maximum mean difference loss of each branch, the distribution difference between each branch can be determined according to the emotion category membership distribution of each EEG data in the EEG dataset output by each branch of the classification layer, and the branch consistency loss can be obtained by summing the distribution differences.
[0050] For example, since the EEG dataset to be tested has no real emotion labels, it is impossible to directly supervise the output consistency of each branch through labels. Therefore, this is achieved by measuring the difference between the emotion category affiliation distributions of different branches to the same EEG dataset.
[0051] Specifically, for each EEG data sample in the dataset to be tested, the emotion category membership distribution of all branches in the multi-branch classification layer is obtained, and then the difference between the emotion category membership distributions of any two branches is calculated. For example, the L1 norm can be used as an indicator to measure the difference between the two probability distributions. The branch consistency loss is obtained by averaging the L1 norm differences between all pairs of branches for all EEG data samples to be tested.
[0052] In this embodiment, by introducing branch consistency loss, different branches can be forced to make similar emotion category judgments on unlabeled EEG data, thereby prompting the multi-branch network to learn robust and consistent emotion feature representations, reducing feature bias between branches, and improving the overall stability and generalization ability of the model.
[0053] Step S150: Construct a first loss function based on domain alignment loss, branch consistency loss, and classification loss. Use the first loss function to iteratively train the neural network model to obtain the first emotion recognition model.
[0054] For example, the first loss function can be represented as a weighted combination of classification loss, domain alignment loss, and branch consistency loss. The weight coefficients of classification loss, domain alignment loss, and branch consistency loss can be dynamically adjusted according to the focus of model training. In each iteration, the gradient of the first loss function with respect to the parameters of each layer of the neural network model is calculated through the backpropagation algorithm, and the model parameters are updated using a gradient descent optimizer (such as Adam, SGD, etc.) until the model's performance on the validation set no longer improves or reaches the preset number of training rounds. The model obtained at this point is the first emotion recognition model. Based on the distribution characteristics of the EEG data to be tested and the sample EEG data, the first emotion recognition model mines rich emotion-related features through a multi-branch structure and combines classification loss, domain alignment loss, and branch consistency loss for joint optimization. As a result, it has a good generalization ability for emotion recognition on the EEG data of new subjects (test subjects) and can accurately identify the corresponding emotion category from the EEG data to be tested.
[0055] The method employed in this embodiment achieves joint optimization of labeled sample emotion category learning, cross-domain feature distribution alignment, and multi-branch feature synergy during model training. This enables the first emotion recognition model to not only accurately learn emotion classification knowledge on the sample EEG dataset but also effectively reduce the domain offset effect between the test EEG dataset and the sample EEG dataset. Furthermore, it enhances the robustness of feature representation by forcing multi-branch output consistency. Compared to traditional single loss function training methods, this significantly improves the model's emotion recognition accuracy and generalization stability when facing new subject EEG data, providing an effective technical approach to address the individual differences in EEG signal emotion recognition.
[0056] In some embodiments, such as Figure 5 As shown, in Figures 1 to 4 Based on the illustrated embodiment, the training method of the emotion recognition model in this application further includes the following steps S501-S505: Step S501: Input the EEG data in the EEG dataset to be tested into the first emotion recognition model to obtain the pseudo-emotion labels and confidence scores of each EEG data output by the first emotion recognition model; Based on the pseudo-emotion labels and confidence scores of each EEG data, filter the EEG data in the EEG dataset to be tested, and use the filtered EEG data and their corresponding pseudo-emotion labels to construct the target EEG dataset.
[0057] In this dataset, pseudo-emotion labels are generated by the first emotion recognition model when predicting the emotion category of unlabeled EEG data. The confidence score reflects the reliability of the model's prediction. For example, a confidence threshold, such as 0.8, can be set, and EEG data with confidence scores higher than this threshold, along with their corresponding pseudo-emotion labels, can be retained to form the target EEG dataset. The purpose of this step is to select samples from the EEG data that have relatively reliable model predictions, providing additional supervision information for subsequent model optimization.
[0058] By introducing these high-confidence pseudo-labeled data, the lack of real labels in the EEG dataset can be compensated for to some extent, further improving the model's adaptability to new subject data.
[0059] In some embodiments, in step S501, a portion of the EEG data in the dataset to be tested can be selected based on a confidence threshold. This portion of the EEG data and its corresponding pseudo-emotion labels are used as an initial retention set. Further, using the prototype of the sample features with real emotion labels in the sample EEG dataset as a reference, the EEG data that highly matches the features of the same emotion category in the sample EEG dataset is selected from the initial retention set as target EEG data (based on prototype distance selection). The target EEG data and its corresponding pseudo-emotion labels are then used to construct the target EEG dataset.
[0060] Specifically, based on the real emotion labels corresponding to each EEG data in the sample EEG dataset, the sample features can be grouped according to emotion category, and the mean of the sample features of each emotion category can be calculated to obtain the feature prototype vector corresponding to the corresponding emotion category. The feature prototype vector can represent the core features of the corresponding emotion category.
[0061] Furthermore, for each EEG data in the initial retention set, the corresponding feature representation output by the model is extracted, and the Euclidean distance between the feature representation and the feature prototype vector of the emotion category corresponding to the pseudo-emotion label is calculated. The smaller the distance value, the higher the structural consistency between the EEG data in the initial retention set and the emotion category features in the sample EEG dataset.
[0062] Therefore, based on the Euclidean distance calculated from the EEG data in the initial retained set, a prototype distance-based screening can be performed. For example, the 0.8 quantile can be used for screening based on empirical values. This allows for further screening of the target EEG data from the initial retained set, and the target EEG data and their corresponding pseudo-emotion labels can be used to construct the target EEG dataset.
[0063] The method in this embodiment uses a dual screening of confidence level and prototype distance to obtain the target EEG dataset. This ensures both the predictive reliability of the sample pseudo-labels and the structural consistency between the sample feature distribution and the same category features in the sample EEG dataset. This effectively reduces the training risks caused by pseudo-label noise and abnormal feature structure, and can significantly improve the training accuracy and robustness of the emotion recognition model in cross-subject scenarios.
[0064] Step S502: Determine the global confidence coefficient based on the confidence levels of each EEG data point in the target EEG dataset.
[0065] For example, the global confidence coefficient can be obtained by arithmetically averaging or weighted averaging the confidence scores of all EEG data in the target EEG dataset. For instance, using an arithmetic average, the global confidence coefficient is the sum of the confidence scores of all samples divided by the number of samples. If the reliability differences in confidence scores among different samples are considered, a weighted average can also be used, where the confidence score of each sample is used as its own weight, and the weighted average of the confidence scores is calculated as the global confidence coefficient. The global confidence coefficient is used to measure the overall reliability of pseudo-labels in the target EEG dataset, providing a basis for subsequent weight adjustments to the pseudo-label loss.
[0066] Step S503: Based on the pseudo-emotion labels of each EEG data in the target EEG dataset and the real emotion labels of each EEG data in the sample EEG dataset, emotion category matching is performed on the EEG data in the target EEG dataset and the sample EEG dataset to obtain multiple EEG data subsets.
[0067] For example, the target EEG dataset and the sample EEG dataset can be divided according to emotion categories, grouping EEG data belonging to the same emotion category into a subset. For instance, if the emotion categories include "happy," "sad," "calm," and "angry," then samples in the target EEG dataset with the pseudo-label "happy" can be merged with samples in the sample EEG dataset with the true label "happy," forming a subset of EEG data for the "happy" emotion category; similarly, subsets for other emotion categories such as "sad," "calm," and "angry" can be constructed. This emotion category matching ensures that each subset contains both sample EEG data with true labels and target EEG data with high-confidence pseudo-labels.
[0068] Step S504: Construct a local discrepancy measure using multiple subsets of EEG data, and construct a local alignment loss based on the local discrepancy measure and the global confidence coefficient.
[0069] For example, for any subset of EEG data, a local subset difference measure can be constructed based on intra-class aggregation constraints and inter-class separation constraints. The local subset difference measures corresponding to each subset of EEG data are summarized to obtain the local difference measure. The global confidence coefficient is multiplied by the local difference measure to obtain the local alignment loss.
[0070] Among them, the local subset difference measure is used to quantify the difference in feature distribution between the sample EEG dataset and the target EEG dataset under the same emotion category. For example, the ratio of intra-class divergence to inter-class divergence can be used as the local subset difference measure. The smaller the intra-class divergence, the more concentrated the feature distribution within the same category is. The larger the inter-class divergence, the higher the feature distribution discrimination between different categories is. By minimizing this ratio, close alignment of cross-domain features under the same emotion category can be achieved.
[0071] Specifically, for a subset of EEG data belonging to a certain emotion category, the mean vectors of the sample EEG data features and the target EEG data features are calculated. The mixed mean of the two classes of features is used as the center of that class. The intra-class divergence is defined as the sum of the squared distances between the sample EEG data features and their class centers, and the sum of the squared distances between the target EEG data features and their class centers. The inter-class divergence is defined as the sum of the squared distances between the centers of different emotion categories. This local subset difference metric can accurately capture the distributional differences between samples and target data within the same emotion category, providing detailed supervisory signals for subsequent calculation of local alignment loss.
[0072] Step S505: Construct a second loss function based on local alignment loss and classification loss, and use the second loss function to iteratively train the first emotion recognition model to obtain the second emotion recognition model.
[0073] For example, the second loss function can be expressed as a weighted combination of classification loss and local alignment loss, wherein the classification loss still adopts the cross-entropy loss calculated based on the true labels of the sample EEG data in step S130, while the local alignment loss is obtained through step S504.
[0074] During model training, sample EEG data and target EEG data are first input into the first emotion recognition model to extract their respective features and predict the emotion category. For the sample EEG data, the classification loss relative to the true emotion label is calculated; for the target EEG data, combined with its pseudo-emotion label, a local alignment loss is calculated using a local difference metric to further narrow the distance between the sample and target data in the feature space for the same emotion category. Subsequently, the classification loss and the local alignment loss are weighted and summed according to preset weights to obtain the second loss function. In each iteration, the gradient of the second loss function with respect to the model parameters is calculated using the backpropagation algorithm, and the parameters are updated using an optimizer (such as Adam) until the model converges or reaches the preset number of training epochs, ultimately yielding the second emotion recognition model.
[0075] In this embodiment, by introducing high-confidence pseudo-label data and local alignment loss, the model can further utilize the emotional information contained in the EEG data to be tested, and finely optimize the alignment effect of cross-domain features within the same emotion category, thereby significantly improving the model's emotion recognition accuracy and generalization ability on new subject EEG data. Especially when there are large individual differences or domain shifts between the sample EEG data and the EEG data to be tested, this optimization strategy can effectively enhance the robustness of the model.
[0076] The specific settings and implementation methods of the embodiments of this application have been described above from different perspectives. Using the methods provided in the above embodiments, the generalization ability and recognition accuracy of emotion recognition models when faced with new subject EEG data can be effectively improved. By learning features from EEG data through a multi-branch network structure, and combining the joint optimization of classification loss, domain alignment loss, and branch consistency loss, the model can initially overcome the domain shift problem and learn robust emotion features.
[0077] Furthermore, by introducing strategies of high-confidence pseudo-labels and local alignment loss, secondary optimization is performed using the information inherent in the test data itself. This achieves fine alignment of features between sample data and target data within the same emotion category, thus mitigating the impact of individual differences. This phased, multi-objective training method enables the model to not only perform exceptionally well on labeled sample datasets but also adapt to new test subjects, providing strong technical support for the practical application of EEG signal emotion recognition technology. It demonstrates significant advantages and application value, particularly in application scenarios requiring processing multiple sample sources or significant individual differences.
[0078] As an implementation of the above methods, such as Figure 6 As shown in the embodiments of this application, a training device for an emotion recognition model is also provided, which may include:
[0079] The dataset acquisition module 601 is used to acquire the EEG dataset to be tested and multiple sample EEG datasets; wherein, the EEG dataset to be tested contains the EEG data of the subject to be tested, and the sample EEG dataset contains the EEG data of the sample subject and its true emotion label. The first processing module 602 is used to process the EEG data in the test EEG dataset and the sample EEG dataset based on the neural network model to be trained, and to obtain the multi-branch latent features and emotion category membership distribution of each EEG data.
[0080] The second processing module 603 is used to calculate the classification loss based on the emotion category membership distribution and true emotion label of the EEG data in the sample EEG dataset; calculate the domain alignment loss between the EEG dataset to be tested and the sample EEG dataset based on the multi-branch latent features of each EEG data; and calculate the branch consistency loss based on the emotion category membership distribution of each EEG data in the EEG dataset to be tested. The model training module 604 is used to construct a first loss function based on domain alignment loss, branch consistency loss and classification loss, and to iteratively train the neural network model using the first loss function to obtain the first emotion recognition model.
[0081] In some embodiments, the first processing module 602 is configured to: input the EEG data from the EEG dataset to be tested and the sample EEG dataset into the feature extraction layer of the neural network model to be trained, respectively, to extract the original features of each EEG data; input the original features into the multi-branch feature transformation layer of the neural network model, and perform feature mapping and transformation of the original features in different dimensions through multiple parallel branch networks to obtain the multi-branch latent features of each EEG data; input the multi-branch latent features into the classification layer of the neural network model, and process the latent features of each branch through the softmax function to obtain the emotion category membership distribution of each EEG data corresponding to different emotion categories.
[0082] In some embodiments, the second processing module 603 is used to: calculate the maximum mean difference loss for the latent features of the sample EEG dataset and the latent features of the EEG dataset to be tested corresponding to each branch output by the multi-branch feature transformation layer; determine the domain alignment loss based on the maximum mean difference loss of each branch; determine the distribution difference between each pair of branches according to the emotion category membership distribution of each EEG data in the EEG dataset to be tested output by each branch of the classification layer; and summarize the distribution differences to obtain the branch consistency loss.
[0083] In some embodiments, the second processing module 603 is used to: compare the emotion category membership distribution of each EEG data in the sample EEG dataset with the real emotion label corresponding to the EEG data, calculate the difference between the two using the cross-entropy loss function, and obtain the classification loss value of each EEG data in the sample EEG dataset; and determine the classification loss based on the classification loss value of each EEG data in the sample EEG dataset.
[0084] In some embodiments, the dataset acquisition module 601 is configured to: acquire a test EEG dataset and multiple candidate sample EEG datasets; extract first distribution feature statistics of the test EEG dataset and second distribution feature statistics of each candidate sample EEG dataset; determine the difference between the second distribution feature statistics of each candidate sample EEG dataset and the first distribution feature statistics, as the distribution difference degree corresponding to each candidate sample EEG dataset; and determine the sample EEG dataset from the multiple candidate sample EEG datasets based on the distribution difference degree corresponding to each candidate sample EEG dataset.
[0085] In some embodiments, the model training module 604 is further configured to input EEG data from the EEG dataset to be tested into a first emotion recognition model to obtain pseudo-emotion labels and confidence scores for each EEG data output by the first emotion recognition model; filter the EEG data in the EEG dataset to be tested based on the pseudo-emotion labels and confidence scores of each EEG data, and construct a target EEG dataset using the filtered EEG data and their corresponding pseudo-emotion labels; determine a global confidence coefficient based on the confidence scores of each EEG data in the target EEG dataset; perform emotion category matching on the EEG data in the target EEG dataset and the sample EEG dataset based on the pseudo-emotion labels of each EEG data in the target EEG dataset and the real emotion labels of each EEG data in the sample EEG dataset to obtain multiple EEG data subsets; construct a local difference measure using the multiple EEG data subsets, and construct a local alignment loss based on the local difference measure and the global confidence coefficient; construct a second loss function based on the local alignment loss and the classification loss, and iteratively train the first emotion recognition model using the second loss function to obtain a second emotion recognition model.
[0086] In some embodiments, the model training module 604 is further configured to: construct a local subset difference measure based on intra-class aggregation constraints and inter-class separation constraints for any subset of EEG data; wherein the local subset difference measure is used to quantify the feature distribution differences between the sample EEG dataset and the target EEG dataset under the same emotion category; summarize the local subset difference measures corresponding to each subset of EEG data to obtain the local difference measure; and multiply the global confidence coefficient with the local difference measure to obtain the local alignment loss.
[0087] The functions of each unit, module, or sub-module in the various devices of this application embodiment can be found in the corresponding descriptions in the above method embodiments, and they have corresponding beneficial effects, which will not be repeated here.
[0088] like Figure 7 The diagram shown is a schematic representation of the internal structure of an electronic device provided in this embodiment. This electronic device can be a server. It includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements the method of this embodiment.
[0089] Those skilled in the art will understand that Figure 7 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0090] In a specific implementation, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method in any of the above embodiments.
[0091] In a specific implementation, this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method in any of the above embodiments.
[0092] This application also provides a chip, including: an input interface, an output interface, a processor, and a memory. The input interface, output interface, processor, and memory are connected through an internal connection path. The processor is used to execute code in the memory. When the code is executed, the processor is used to execute the method provided in this application.
[0093] It should be understood that the aforementioned processor can be a CPU, or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), FPGAs, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. General-purpose processors can be microprocessors or any conventional processor. It is worth noting that the processor can be a processor supporting Advanced Reduced Instruction Set Machines (ARM) architecture.
[0094] Further, optionally, the aforementioned memory may include read-only memory and random access memory. The memory may be volatile memory or non-volatile memory, or may include both. Non-volatile memory may include read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may include random access memory (RAM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available. Examples include Static Random Access Memory (SRAM), Dynamic Random Access Memory (DRAM), Synchronous DRAM (SDRAM), Double Data Rate SDRAM (DDR SDRAM), Enhanced Synchronous DRAM (ESDRAM), Sync Link DRAM (SLDRAM), and Direct Rambus RAM (DR RAM).
[0095] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.
[0096] Computer program products can be written in any combination of one or more programming languages to perform the operations of embodiments of this disclosure. These programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0097] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. All or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware, the program being stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiments.
[0098] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. This storage medium can be a read-only memory, a disk, or an optical disk, etc.
[0099] The above description is merely an exemplary embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope described in this application, and these should all be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A training method for an emotion recognition model, characterized in that, include: Acquire a test EEG dataset and multiple sample EEG datasets; wherein, the test EEG dataset contains the EEG data of the test subject, and the sample EEG dataset contains the EEG data of the sample subject and its true emotion label; The EEG data in the EEG dataset to be tested and the sample EEG dataset are processed based on the neural network model to be trained, so as to obtain the multi-branch latent features and emotion category membership distribution of each EEG data. The classification loss is calculated based on the emotion category membership distribution of the EEG data in the sample EEG dataset and the true emotion labels. The domain alignment loss between the EEG dataset to be tested and the sample EEG dataset is calculated based on the multi-branch latent features of each EEG dataset, and the branch consistency loss is calculated based on the emotion category membership distribution of each EEG dataset to be tested. A first loss function is constructed based on the domain alignment loss, the branch consistency loss, and the classification loss. The neural network model is then iteratively trained using the first loss function to obtain a first emotion recognition model.
2. The method according to claim 1, characterized in that, The process of the EEG data in the test EEG dataset and the sample EEG dataset based on the neural network model to be trained includes: The EEG data from the EEG dataset to be tested and the sample EEG dataset are respectively input into the feature extraction layer of the neural network model to be trained in order to extract the original features of each EEG data. The original features are input into the multi-branch feature transformation layer of the neural network model. The original features are then mapped and transformed in different dimensions through multiple parallel branch networks to obtain the multi-branch latent features of each EEG data. The multi-branch latent features are input into the classification layer of the neural network model, and the latent features of each branch are processed by the softmax function to obtain the emotion category membership distribution of each EEG data corresponding to different emotion categories.
3. The method according to claim 2, characterized in that, The calculation of the domain alignment loss between the EEG dataset to be tested and the sample EEG dataset based on the multi-branch latent features of each of the EEG datasets, and the calculation of the branch consistency loss based on the emotion category membership distribution of each of the EEG datasets, include: For the latent features of the sample EEG dataset and the latent features of the EEG dataset to be tested corresponding to each branch output by the multi-branch feature transformation layer, the maximum mean difference loss is calculated respectively. The domain alignment loss is determined based on the maximum mean difference loss of each branch; Based on the emotion category membership distribution of each EEG data in the EEG dataset output by each branch of the classification layer, the distribution differences between each pair of branches are determined, and the branch consistency loss is obtained by summing the distribution differences.
4. The method according to claim 1, characterized in that, The step of calculating the classification loss based on the emotion category membership distribution of the EEG data in the sample EEG dataset and the true emotion labels includes: The emotion category membership distribution of each EEG data in the sample EEG dataset is compared with the real emotion label corresponding to the EEG data. The difference between the two is calculated using the cross-entropy loss function to obtain the classification loss value of each EEG data in the sample EEG dataset. The classification loss is determined based on the classification loss value of each EEG data in the sample EEG dataset.
5. The method according to claim 1, characterized in that, The acquisition of the EEG dataset to be tested and multiple sample EEG datasets includes: Obtain the EEG dataset to be tested and multiple candidate sample EEG datasets; The first distribution feature statistical information of the EEG dataset to be tested and the second distribution feature statistical information of each candidate sample EEG dataset are extracted respectively. The difference between the second distribution feature statistics of each candidate sample EEG dataset and the first distribution feature statistics is determined as the distribution difference degree corresponding to each candidate sample EEG dataset. Based on the distribution difference degree corresponding to each of the candidate sample EEG datasets, the sample EEG dataset is determined from the plurality of candidate sample EEG datasets.
6. The method according to any one of claims 1-5, characterized in that, Also includes: Input the EEG data from the EEG dataset to be tested into the first emotion recognition model to obtain the pseudo-emotion labels and confidence scores of each EEG data output by the first emotion recognition model; The EEG data in the EEG dataset to be tested are filtered based on the pseudo-emotion labels and confidence scores of each EEG data, and the target EEG dataset is constructed using the filtered EEG data and their corresponding pseudo-emotion labels. The global confidence coefficient is determined based on the confidence level of each EEG data point in the target EEG dataset. Based on the pseudo-emotion labels of each EEG data in the target EEG dataset and the real emotion labels of each EEG data in the sample EEG dataset, emotion category matching is performed on the EEG data in the target EEG dataset and the sample EEG dataset to obtain multiple EEG data subsets. A local discrepancy metric is constructed using the multiple subsets of EEG data, and a local alignment loss is constructed based on the local discrepancy metric and the global confidence coefficient. A second loss function is constructed based on the local alignment loss and the classification loss. The first emotion recognition model is then iteratively trained using the second loss function to obtain the second emotion recognition model.
7. The method according to claim 6, characterized in that, The step of constructing a local discrepancy metric using the multiple subsets of EEG data, and constructing a local alignment loss based on the local discrepancy metric and the global confidence coefficient, includes: For any of the aforementioned EEG data subsets, a local subset difference measure is constructed based on intra-class aggregation constraints and inter-class separation constraints; wherein, the local subset difference measure is used to quantify the difference in feature distribution between the sample EEG dataset and the target EEG dataset under the same emotion category; The local subset difference measures corresponding to each of the aforementioned EEG data subsets are summarized to obtain the local difference measures. The local alignment loss is obtained by multiplying the global confidence coefficient with the local difference metric.
8. A training device for an emotion recognition model, characterized in that, include: The dataset acquisition module is used to acquire the EEG dataset to be tested and multiple sample EEG datasets; wherein, the EEG dataset to be tested contains the EEG data of the subject to be tested, and the sample EEG dataset contains the EEG data of the sample subject and its true emotion label. The first processing module is used to process the EEG data in the EEG dataset to be tested and the sample EEG dataset based on the neural network model to be trained, and to obtain the multi-branch latent features and emotion category membership distribution of each EEG data. The second processing module is used to calculate the classification loss based on the emotion category membership distribution of the EEG data in the sample EEG dataset and the real emotion label; calculate the domain alignment loss between the EEG dataset to be tested and the sample EEG dataset based on the multi-branch latent features of each EEG data; and calculate the branch consistency loss based on the emotion category membership distribution of each EEG data in the EEG dataset to be tested. The model training module is used to construct a first loss function based on the domain alignment loss, the branch consistency loss, and the classification loss, and to iteratively train the neural network model using the first loss function to obtain a first emotion recognition model.
9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory, wherein the processor, when executing the computer program, implements the method of any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method of any one of claims 1-7.