Emotion prediction system, method, device and medium
Through vector quantization variational autoencoder and flexible actor-criticist model combined with self-attention mechanism, the problem of lack of dynamic emotions modeling in the existing technology is solved, and efficient prediction of continuous emotional states of EEG signals is achieved, which improves the accuracy and application potential of emotion recognition.
Patent Information
- Application Number
- CN202510740712.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The prior art lacks effective methods for dynamic modeling and prediction of continuous emotional states, which limits the application of emotion recognition in brain-computer interfaces and human-computer interactions.
The differential entropy characteristics of the EEG signal were extracted using vector quantized variational autoencoder (VQ-VAE), and the timing dependence relationship of the deep features of dynamic emotions was captured through the flexible actor-criticist model (SAC). Combining the self-attention mechanism and reinforcement learning, regression prediction of continuous emotional state was achieved.
It realizes dynamic emotions modeling with continuous changes in time, improves the accuracy and stability of emotion recognition, and promotes the application of brain-computer interfaces and human-computer interactions.
Smart Images

Figure CN120241071A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data processing, and in particular, to an emotion prediction system, method, device and medium. Background Art
[0002] Human emotion is a continuous dynamic process, which is characterized by reflecting the complex interaction between the human body and the external environment. How to identify emotion segments related to tasks from continuous electroencephalogram (EEG) signals is a major challenge.
[0003] Electroencephalogram (EEG) provides a direct, objective and scientific basis for evaluating emotional states and is a valuable tool in emotion recognition research. In recent years, the potential of EEG-based emotion recognition has attracted increasing attention from researchers.
[0004] By analyzing EEG data, researchers can identify and classify different emotional states, promoting a deeper understanding of human emotions. The prior art has proposed a multi-modal emotion dataset with continuous labels corresponding to emotions, but these labels are only used for the selection of training data for emotion classification, lacking continuous modeling of dynamic emotions and a method for predicting regression of continuous emotional states, thus limiting the application of emotion recognition in brain-computer interfaces and human-computer interactions. Summary of the Invention
[0005] The present application proposes an emotion prediction system, method, device and medium, which can solve one of the problems existing in the background art.
[0006] To achieve the above object, the present application adopts the following technical solutions: In a first aspect, an emotion prediction system is provided. The prediction system includes: A vector quantization variational autoencoder for obtaining EEG signals; extracting differential entropy features of the EEG signals that are continuous in time; mapping the differential entropy features to a latent space to obtain feature vectors; performing vector quantization on the feature vectors to obtain codebook features; and concatenating the feature vectors and the codebook features to obtain dynamic emotion depth features; and A flexible actor-critic model under a Markov decision framework with the goal of maximizing reward and maximizing entropy, including: an actor network and a critic network. The actor network is used to capture the temporal dependence of the dynamic emotion depth features to obtain depth features for prediction; generating an action policy according to the current state. The critic network is used to evaluate the value of the state-action pair. The reward function of the flexible actor-critic model includes: a first part for evaluating the current prediction accuracy, and a second part for evaluating the consistency between the change trend of the predicted values at the previous and current time steps and the change trend of the true values.
[0007] Based on the above technical solution, the emotion prediction system includes a vector quantization variational autoencoder and a flexible actor-critic model. The vector quantization variational autoencoder extracts the differential entropy features that are temporally continuous from the EEG signals, maps them to the latent space to obtain feature vectors, and performs vector quantization to obtain codebook features. The codebook features and the feature vectors are concatenated to obtain dynamic emotion depth features. Then, the flexible actor-critic model captures the temporal dependence relationship of the dynamic emotion depth features to obtain the depth features for prediction. By processing the depth features for prediction, an emotion prediction value is obtained. Based on the set reward function, the flexible actor-critic model can not only consider the prediction accuracy, but also consider whether the change trend of the prediction values at the previous and subsequent time steps is consistent with the change trend of the true values. In this way, the modeling of dynamic emotions that change continuously over time is realized, and at the same time, the regression prediction of continuous emotion states is realized, thereby promoting the application of emotion recognition in brain-computer interfaces and human-computer interactions.
[0008] In a possible design manner of the first aspect, the vector quantization variational autoencoder is specifically configured to: Discretize the feature vectors into discrete vectors; and Based on the discrete vectors, the codewords and indices in the codebook, and using the nearest neighbor search, determine the codebook features.
[0009] In a possible design manner of the first aspect, the loss function adopted by the vector quantization variational autoencoder includes: a reconstruction loss term and a quantization loss term.
[0010] In a possible design manner of the first aspect, when the vector quantization variational autoencoder concatenates the feature vectors and the codebook features, a self-attention mechanism is introduced.
[0011] Since some semantic information of individual differences may be lost during the vector quantization process, based on the above technical solution, by using the self-attention mechanism, important features can be effectively highlighted, the semantic information of the features can be enhanced, and the accuracy of emotion prediction is further improved.
[0012] In a possible design manner of the first aspect, the actor network includes: A long short-term memory network, which is used to capture the temporal dependence relationship of the dynamic emotion depth features to obtain the depth features for prediction; and A fully connected layer, which is used to generate an action policy according to the current state.
[0013] In a possible design manner of the first aspect, the second part Is constructed by the following function: Wherein, Represents the change in the true value, Represents the change in the predicted value, and else represents other cases.
[0014] Based on the above technical solution, the design of the second part in the reward function enables the agent to capture the transfer relationship of the EEG emotional state during the exploration and exploitation processes, thereby ensuring accurate prediction of the emotional state.
[0015] In a second aspect, a training method for an emotion prediction system is provided. The training method trains the emotion prediction system as described above. The training method includes: Obtaining a training data set, where the training data set includes: EEG signals and true emotion values; and Using the training data set to train the emotion prediction system.
[0016] In a third aspect, an emotion prediction method is provided. The prediction method includes: Obtaining the current EEG signal; and Using the emotion prediction system as described above to process the current EEG signal to obtain an emotion prediction result.
[0017] In a fourth aspect, an electronic device is provided. The electronic device includes: a processor and a memory coupled to the processor. The memory is used to store a computer program. The processor is used to execute the computer program stored in the memory so that the electronic device executes the training method as described in the second aspect or executes the prediction method as described in the third aspect.
[0018] In a fifth aspect, a computer-readable storage medium is provided, including a computer program or instruction. When the computer program or instruction runs on a computer, the computer is caused to execute the training method as described in the second aspect or execute the prediction method as described in the third aspect.
[0019] In a sixth aspect, a computer program product is provided, including: a computer program or instruction. When the computer program or instruction runs on a computer, the computer is caused to execute the training method as described in the second aspect or execute the prediction method as described in the third aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings in the following description are only some embodiments of the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1It is a schematic structural diagram of the emotion prediction system provided in the first embodiment of the present application; Figure 2 It is a framework diagram of the dynamic emotion prediction model provided in the second embodiment of the present application; Figure 3 It is a framework diagram of the VQ-VAE model provided in the second embodiment of the present application. Detailed implementation manners
[0022] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0023] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different module division in the device or a different order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0025] Embodiment 1 As Figure 1 shown, this embodiment provides an emotion prediction system 100, and the prediction system includes: A vector quantization variational autoencoder 101, configured to obtain electroencephalogram signals; extract differential entropy features of the electroencephalogram signals that are continuous in time; map the differential entropy features to a latent space to obtain feature vectors; perform vector quantization on the feature vectors to obtain codebook features; and splice the feature vectors and the codebook features to obtain dynamic emotion depth features; and A flexible actor-critic model 102 under the Markov decision framework with the objectives of maximizing rewards and maximizing entropy, including: an actor network 201 and a critic network 202. The actor network 201 is configured to capture the temporal dependence relationship of the dynamic emotion depth features to obtain depth features for prediction; generate an action policy according to the current state. The critic network 202 is configured to evaluate the value of the state-action pair. The reward function of the flexible actor-critic model includes: a first part for evaluating the current prediction accuracy, and a second part for evaluating the consistency between the change trend of the predicted values at the previous and current time steps and the change trend of the true values.
[0026] Specifically, the Vector Quantized Variational Autoencoder (VQ-VAE) is an improvement of the Variational Autoencoder (VAE). By introducing a Vector Quantization (VQ) layer, it maps the continuous encoder output to a discrete latent space. This discrete representation helps capture modalities such as language, reasoning, and planning, and also results in higher-quality generated data.
[0027] The VQ-VAE mainly consists of the following core components: a Vector Quantizer, an encoder, and a decoder. First, it encodes the input into latent vectors through the encoder, then maps these vectors to the discrete latent space through the vector quantizer, and finally reconstructs through the decoder.
[0028] Electroencephalogram (EEG) signals are the overall reflection of the electrophysiological activities of brain nerve cells on the cerebral cortex or the scalp surface. In engineering applications, EEG signals can be used to implement a Brain-Computer Interface (BCI). By taking advantage of the different EEG signals corresponding to different sensory, motor, or cognitive activities of a person, and through the effective extraction and processing of EEG signals, a certain purpose can be achieved, such as the continuous dynamic emotion prediction in this embodiment.
[0029] The differential entropy feature of EEG signals is a feature used to analyze EEG signals, mainly used to describe the complexity and irregularity of EEG signals. Differential entropy is a generalized form of Shannon entropy for continuous variables.
[0030] The extraction of differential entropy features can be mainly achieved through steps such as importing relevant libraries, defining a function to calculate differential entropy, defining a function to calculate the power spectral density, and extracting differential entropy features.
[0031] Differential entropy features have a wide range of applications in EEG signal analysis, especially in the field of emotion recognition. For example, on the DEAP dataset, by segmenting EEG signals in the frequency domain and extracting their differential entropy features, combined with a Convolutional Neural Network (CNN), a Support Vector Machine (SVM), and a Multi-Layer Perceptron (MLP), high-precision emotion recognition can be achieved, and the accuracy rate can usually reach over 90%. Of course, the processing of differential entropy features in this embodiment is different from the above example.
[0032] In this embodiment, after the VQ-VAE maps the differential entropy features to the latent space to obtain feature vectors, it further discretizes the feature vectors into discrete vectors, and uses the discrete vectors, the codewords and indices in the codebook, and the nearest neighbor search to determine the codebook features.
[0033] Specifically, in the latent space, the continuous feature vectors are discretized into discrete vectors. By minimizing the difference between the discrete vectors and the codewords in the Codebook, the index used to determine the codebook features is obtained, and then the codeword corresponding to this index is determined as the codebook feature.
[0034] The loss function adopted by VQ-VAE includes a reconstruction loss term and a quantization loss term. Among them, the reconstruction loss term is used to measure the difference between the encoded input and the decoded output, and the quantization loss term is used to measure the vector quantization loss.
[0035] When VQ-VAE concatenates the feature vector and the codebook feature, a self-attention mechanism can be introduced to achieve weighted feature vectors and codebook features with dynamic weights.
[0036] Markov decision-making (MDP) refers to a method of making decisions using a Markov transition matrix, which belongs to probabilistic decision-making techniques. Its basic principle is that although the decision-maker cannot know the probability of a certain natural state occurring in the near future, but knows the probability distribution change, that is, the transition matrix, between natural states, the stable probability of each natural state in the future environment can be calculated according to the transition matrix, and then the expected value decision-making method or deterministic decision-making technique can be used to select the best solution. MDP is constructed based on a set of interacting objects, namely agents and the environment, and its elements include states, actions, and rewards. In the simulation of MDP, the agent will perceive the current system state, perform actions on the environment according to the strategy, thereby changing the state of the environment and obtaining rewards, and the accumulation of rewards over time is called return.
[0037] The problem solved by the Soft Actor-Critic (SAC) model is the reinforcement learning problem in discrete and continuous action spaces, and it uses an off-policy reinforcement learning algorithm. The core idea of the SAC model is to add a "soft" objective, that is, the maximum entropy objective, on the basis of the standard Actor-Critic model, so as to not only maximize the reward but also maximize the entropy.
[0038] The SAC model optimizes the objective function through the policy gradient method.
[0039] The SAC model includes an Actor network and a Critic network.
[0040] The Actor network includes a Long Short-Term Memory (LSTM) network and a Fully Connected layer (FC). The Long Short-Term Memory network is used to capture the temporal dependencies of dynamic emotion depth features to obtain depth features for prediction, and the Fully Connected layer is used to process the depth features for prediction and generate action policies according to the current state.
[0041] The critic network is used to evaluate the value of state-action pairs. In the SAC algorithm, there are usually two critic networks, which respectively output value estimates for given states and actions. These two value estimates are used to calculate the target value and update the parameters of the critic network by minimizing the Bellman error.
[0042] In the SAC model, target networks can also be set, such as: a target actor network and a target critic network. The target network is a delayed copy, which is used to slow down the fluctuations during the training process.
[0043] In this embodiment, the reward function of the SAC model includes: a first part for evaluating the current prediction accuracy, and a second part for evaluating the consistency between the change trends of the predicted values at the previous and current time steps and the change trend of the true values.
[0044] Specifically, the first part can be defined by the predicted value and the true value at the current time step, and the second part is defined by the change amounts of the predicted values and the true values at the previous and current time steps.
[0045] In this way, in consecutive sequential decisions, the reward function will guide the agent to capture the changes in the emotional state and accurately predict the emotion.
[0046] This embodiment also provides a training method for an emotion prediction system. The training method trains the above-mentioned emotion prediction system, and the training method includes: Obtaining a training data set, the training data set including: electroencephalogram signals and true emotion values; and Using the training data set to train the emotion prediction system.
[0047] This embodiment also provides an emotion prediction method, and the prediction method includes: Obtaining the current electroencephalogram signal; and Using the above-mentioned emotion prediction system to process the current electroencephalogram signal to obtain an emotion prediction result.
[0048] The content of the above training method and prediction method is similar to the content of the prediction system, and will not be elaborated here.
[0049] Embodiment 2 This embodiment provides a method for continuous electroencephalogram signal emotion prediction based on a vector quantization variational autoencoder (VQ-VAE) for deep reinforcement learning.
[0050] 1. Framework description First, as Figure 2 shown, the dynamic emotion prediction model we proposed includes the following modules or sub-models: (1)VQ-VAE-based emotional state representation: By performing vector quantization on the electroencephalogram (EEG) signal features, we can obtain the latent emotional state representation globally while compressing the data distribution. Through nearest-neighbor search, we can obtain the index of the corresponding codebook vector, and then obtain the embedded representation of the corresponding codebook vector through the index, thus obtaining the global semantic feature vector corresponding to the EEG feature.
[0051] (2)Self-Attention mechanism: It is used to concatenate the latent space features of the VQ-VAE encoder with their corresponding codebook vectors, and perform dynamic weighting using the Self-Attention mechanism to obtain the deep features related to the prediction task.
[0052] (3)Soft Actor-Critic (SAC) strategy optimization: Based on the Actor-Critic framework, we use the SAC reinforcement learning algorithm to optimize the model. We carefully design the reward function to guide the agent to be able to predict the emotional state score well according to the changes in the EEG emotional state. The actor is the above-mentioned actor, and the critic is the above-mentioned critic.
[0053] Secondly, we propose a training method for a dynamic emotion prediction model (framework): This training method includes two stages: (1)Initial training stage (unsupervised learning stage), as Figure 3 shown, this initial training stage includes the following steps: Step 1: Map the EEG signal into the latent space through the encoder.
[0054] Step 2: Quantize the embedding into codebook vectors through clustering. The EEG features in the latent space obtain the index of the corresponding codebook vector through nearest-neighbor search, and then obtain the embedded representation of the corresponding codebook vector through the index, thus obtaining the global semantic feature vector corresponding to the EEG feature.
[0055] Step 3: The decoder restores the embedded representation of the codebook vector to the original data.
[0056] Step 4: Optimize the model through reconstruction loss and quantization loss.
[0057] (2)Reinforcement learning stage, as Figure 2 shown, this reinforcement learning stage includes the following steps: Step 1: Only use the pre-trained VQ-VAE encoder and codebook with unsupervised learning, and freeze the encoder and codebook. The EEG features are input into the encoder and mapped into the latent space. The EEG features in the latent space obtain the indices of the corresponding codebook vectors through nearest neighbor search, and then the embedded representations of the corresponding codebook vectors are obtained through the indices. The latent space features are concatenated with their corresponding codebook vectors and dynamically weighted using the self-attention mechanism to obtain the deep features related to the prediction task.
[0058] Step 2: Use LSTM to obtain the state transition relationship of the deep features to get the deep temporal features.
[0059] Step 3: Based on SAC for policy optimization, continuously guide the agent to solve the optimal policy through the reward function designed by us.
[0060] In reinforcement learning, experience replay is a method to improve learning efficiency by storing the states, actions, rewards, and next states (i.e., experiences) that the agent has experienced during training.
[0061] The training of the entire model is divided into two stages: 1. In the first stage, use VQ-VAE for unsupervised learning. Reconstruct the EEG signals by using the encoder and decoder, and discretize the continuous latent space into a codebook, so as to extract the deep features of EEG in an unsupervised manner.
[0062] 2. In the second stage, we conduct reinforcement learning training. In this training process, only the encoder (Encoder) and codebook in VQ-VAE are used to extract EEG features. Then, in order to alleviate the quantization problem, the self-attention mechanism is used. The decoder (Decoder) is not required in the second stage, and the encoder is only used to extract features. SA Fusion is another module that needs to be trained in the second stage.
[0063] In this section, we will introduce a method for continuous EEG signal emotion prediction based on VQ-VAE deep reinforcement learning. During the training of the entire model, we use the differential entropy features of the EEG signal EEG as the input of the model. The overall framework diagram of dynamic emotion prediction is as Figure 2 shown, mainly divided into two stages: (1) the emotion state representation stage based on VQ-VAE unsupervised learning; (2) the reinforcement learning stage based on the emotion state representation.
[0064] In the unsupervised learning emotional state representation stage, based on VQ-VAE, the differential entropy features of the EEG of the subjects are vector quantized and reconstructed to obtain the deep EEG emotional state of the subjects. VQ-VAE maps the data into the latent space through the encoder and discretizes and compresses it into a codebook for representation. All the semantic information of the features of the EEG data can be represented by the codebook vectors therein, and we define the codebook vectors therein as the potential emotional states. In this way, we extract the deep features while reducing the redundancy in the data distribution, and obtain the semantic representation of the potential emotional state that does not change with time.
[0065] In the reinforcement learning stage, the model predicts the emotional score based on the potential emotional state of unsupervised learning. We obtain the deep features by passing the EEG signal through the encoder, and obtain the corresponding codebook vector (codeword) of the features through the nearest neighbor search; in order to enhance the semantic information of the individual differences of the emotional state, we splice the codebook vector with its features and use the self-attention mechanism for dynamic weighting to obtain the deep features related to the prediction task; and capture its temporal dependence through LSTM to obtain the deep features of its corresponding dynamic emotion; finally, we use the Soft Actor-Critic (SAC) reinforcement learning algorithm based on the Actor-Critic framework to optimize the model. We carefully design the reward function to guide the agent to be able to predict the emotional state score well according to the changes in the EEG emotional state. In the model optimization stage, the Actor explores the action space in a trial-and-error manner, and learns the interaction relationship between the local video-level EEG sample features and the global emotional state distribution. The Critic scores the Actor's decision in this round at the end of each interaction round to find the improvement points of the strategy. Through continuous iteration, the Actor finally learns the optimal strategy to accurately predict the emotional state of the subjects.
[0066] 2. Method Introduction 2.1 Markov Decision Process We define the process of the agent predicting the emotional state score of the subject from the continuous EEG signal as a Markov decision process. That is, the Markov decision process (MDP). The MDP is defined as a tuple , where S represents the state space, A represents the action space, and A = [-1, 1] corresponds to the actions that the agent may take. The transition probability function describes the possibility of transitioning from one state to another given a specific action. γ ∈ [0, 1] is the decay factor, which determines the trade-off between the short-term and long-term interests of the agent in the decision-making process. The reward function assigns a numerical reward according to the state-action-state transition, providing feedback for the agent's decision-making.
[0067] 2.2 Emotional state representation To obtain the semantic representation of the time-invariant latent emotional state from the EEG data of all subjects for VQ-VAE, the entire data distribution is vector-quantized and compressed into a codebook. The unsupervised pre-training of VQ-VAE is as follows Figure 1 As shown. Looking at the global data distribution of the signal statically, we believe that each cluster of EEG signal features aggregated in the data distribution is a potential emotional state. We use VQ-VAE to perform vector quantization on the EEG signal data features, map the EEG signal into the latent space through the encoder, and embed it into the codebook vector through clustering quantization. For each encoded vector, the nearest neighbor search is performed on the codebook to quantize it into the corresponding codeword, and then the decoder tries its best to restore the codeword in the latent space to the original data.
[0068] Specifically, we use VQ-VAE to input the differential entropy features of the EEG signal into the encoder, input the EEG signal into the encoder and map it in the continuous latent space to obtain the feature vector ; in the latent space, we discretize a group of neighboring feature vectors into a vector in the codebook , that is, the codeword, where t is the time series.
[0069] (1) Each time, through the nearest neighbor search in the codebook, obtain the index i of the codebook vector corresponding to the EEG feature, and obtain its nearest codebook vector through the index, and use it as the semantic feature of the current EEG feature.
[0070] (2) (3) (4) where the input of the EEG signal is x, and the output restored by the decoder is , sg represents the stop gradient, and β is the weight coefficient.
[0071] During the process of vector quantization of the EEG signal, we use the reconstruction loss and the quantization loss for optimization. The reconstruction loss calculates the mean square error between the EEG signal input x and the output restored by the decoder ; the quantization error It consists of two parts. The former is the vector quantization loss, which is used to measure the difference between the encoded features and the codebook vectors, aiming to minimize the distance between each encoded feature and the codebook vectors; the latter is the commitment loss, whose goal is to encourage the output of the encoder to be as close as possible to its nearest codebook vector.
[0072] We obtained the depth features of EEG signals and removed redundant EEG signal data by performing vector quantization on the features in the latent space, and compressed the global data distribution into codebook vectors for representation. We believe that the codebook vectors represent a certain potential emotional state to a certain extent, and the semantic information of this emotional state is static and does not change over time, just like the emotional states of happiness or sadness of a person have no direct relationship with the time point where the subject is located.
[0073] The emotional state representation based on VQ-VAE loses some semantic information of individual differences through vector quantization. In order to enhance the semantic information of the features and enhance our state representation, we introduced the self-attention mechanism, concatenated the feature vectors in the latent space with the codewords in the latent space, and realized weighted with dynamic weights to obtain dynamic depth features. The global feature template (codewords) and the personalized EEG representation (feature vectors) of the current subject are weighted and concatenated, effectively highlighting important features, and assigning higher weights to the features whose results of the real emotion scores are closer, obtaining the dynamic emotion depth features (Dynamic Emotion feature). We calculate the attention weight A of the current feature fusion through the following formula: (5) where Q is the query vector, K is the key vector, V is the value vector, A is the attention weight matrix, and d is the vector dimension. Weighting through the self-attention mechanism can effectively highlight important features.
[0074] Finally, we use LSTM to capture the temporal dependencies of the dynamic emotion depth features to capture the movement relationships between different potential emotional states. LSTM completes the dependency of temporal features through the forget gate, input gate, and output gate, and its calculation formula is as follows: Forget gate Determines which information to discard from the cell. The input gate consists of and components. After that, the cell state updates the features by combining the information of the forget gate and the input gate. The output gate consists of and components, determining which information to output. LSTM will process according to the temporal feature of the previous step and the current input and output the temporal feature of the current time step.
[0075] (6) where is the output of the forget gate, and are its weight and bias constant term respectively, and σ is the activation function; (7) where is the output of the input gate, and are its weight and bias constant term respectively; (8) where is the output of the candidate memory, and are its weight and bias constant term respectively, and tanh is the activation function; (9) where is the output of the current memory, is the output of the previous step memory; (10) where is the output of the output gate, and are its weight and bias constant term respectively; (11) where is the temporal feature of the current time step. By capturing the temporal dependence of the EEG signal sequence features through LSTM, the transition relationship between states is obtained, which can enhance the temporal semantic information of the dynamic emotion depth features and is beneficial to the accuracy and stability of the sequential decision-making of the agent under the perception state.
[0076] 2.3 Reward Function Design To guide the agent to accurately predict the emotion state score according to the changes in the EEG emotion state transition, considering that we hope the agent can not only accurately predict the state score, but also hope to predict the change trend of the score according to the state transition; for this reason, we designed two process reward functions: MAE reward function and Delta reward function. For the EEG signals at the experimental level of Trial-Level, during the process of the agent making continuous sequential decisions, the reward function will guide the agent to capture the changes in the subject's emotion state while accurately predicting its emotion state score.
[0077] MAE Reward Function: This reward function effectively ensures the prediction accuracy by calculating the squared error between the emotion score currently predicted by the agent and the annotated continuous score.
[0078] (12) where is the true label at the current time step, is the predicted value of the agent at the current time step, is the exponential function .
[0079] Delta Reward Function: This reward function compares the change trend of the state prediction value with the actual change trend of the true labels before and after at each time step.
[0080] (13) (14) Guided by the reward function Reward, the agent can capture the transfer relationship of EEG emotion states during the exploration and exploitation process, thus accurately predicting the emotion state score.
[0081] 2.4 Optimization Process We use the SAC reinforcement learning algorithm to train our model. Throughout the process, at the Trial-level, we predict the real-time continuous EEG emotion states of the subjects. In the continuous action space, the model learns the optimal policy through trial and error. Solving the optimal policy in the continuous action space faces problems such as a large exploration space and low sample efficiency. To make the optimization process more efficient, we use the SAC reinforcement learning algorithm to optimize the model. By introducing the constraint of entropy maximization, the agent always maintains exploration in the face of a stochastic dynamic environment, preventing the model from falling into local optima, thereby enhancing the robustness and generalization ability of the model. The SAC algorithm is an off-policy method based on actor-critic, and its cumulative expectation under the maximum entropy mechanism is expressed as: (15) In the formula: π and are the current policy and the optimal policy of the agent respectively; is the action executed by the agent in state , is the reward obtained by the agent; is the state-action trajectory distribution formed by policy π; is the distribution policy mapped from the current state space to the action space; H is the entropy under policy The action entropy taken below, the larger the entropy value, the more exploration of the environment, avoiding the strategy from converging to the local optimum; α is the temperature coefficient of the action entropy, used to determine the weight ratio of entropy to reward.
[0082] The Q-value function improved based on the entropy value is defined as follows: (16) where γ is the reward discount factor; V(s) is the state value function, and its calculation expression is (17) Combine the Bellman operator with the current Q-value function for iterative update: (18) where is the Bellman operator under the policy π; is the value function at the k-th iteration, and finally Q will converge to the soft Q value function under the fixed policy π.
[0083] Adopt the form of minimizing the KL divergence to realize the update of the agent's strategy: (19) where is the KL divergence; is the set of policy distributions; is the old policy under the Q-value function; is the normalization factor constant.
[0084] In the SAC algorithm, use a neural network to fit the Q-value function and the policy function, and minimize the Bellman residual in the form of mean square error to realize the update of the Q-value network parameters: (20) where θ is the parameter of the Q-value network, is the parameter of the target Q-value network, is the parameter of the policy network; 、 and are the updated functions; D is the experience pool.
[0085] Transform the above formula (20) to obtain the optimization objective of the actor, and the expression is (21) where the actor will output the policy entropy, where a is the temperature coefficient; it is adaptively updated during the training process by minimizing J(a): (22) where is the dimension number of the actions output by the policy network.
[0086] During the training process, the model is divided into two parts: the actor and the critic. The actor continuously interacts with the latent space of the emotional state during the decision-making process and iteratively searches for the optimal policy according to the reward mechanism. The critic guides this exploration process by estimating the cumulative expected return of each state and action, thus assisting the actor to find the optimal policy.
[0087] In this method, we conducted a cross-subject experiment on the SEED database and experimented with the method we proposed on this database. The training method of the method we proposed used an experimental protocol of leave-one-out cross-validation for subjects, preserving the original temporal order of the EEG signal features of the subjects. The entire segment of EEG signals at the video level of each subject was used as training data. We guided the agent to accurately predict the emotional state scores of the subjects through supervised rewards. We will perform sequential inference on the test set to obtain their emotional state scores. Here, we use MSE to evaluate the relationship between the predicted emotional state scores and the true labels, and the range of the true labels is [-1, 1]. Among them, the calculation formula of MSE is: (23) where n is the total number of EEG signal samples at the video level, is the true value of the i-th sample, is the model prediction value of the i-th sample.
[0088] Table 1 Experimental results of leave-one-out cross-validation EEG emotional state regression prediction for subjects in the SEED database According to the results in Table 1, in the leave-one-out cross-validation experimental protocol on the SEED database for the task of predicting emotional states, the mean squared error MSE of the method we proposed in the leave-one-out cross-validation prediction for subjects is 0.037. Compared with the current SOTA method EmoTDMF, the MSE of the regression prediction under the same experimental protocol is 0.10, and the regression prediction performance has been improved by 64%; to a certain extent, the model we proposed can better capture the dynamic changes of emotional states and can accurately and stably complete the prediction regression task.
[0089] In addition, to further verify the effectiveness of the proposed method, we compared the results of the model built using different training methods on the SEED database. Here we use MSE, MAE and as evaluation metrics, and the specific calculation formula of MAE is as follows: (24) where n is the total number of EEG signal samples at the video level, is the true value of the i-th sample, is the model prediction value of the i-th sample.
[0090] Table 2 Performance of the proposed model in regression prediction with subject leave-one-out cross-validation on the SEED dataset under different training methods As shown in Table 2, the experimental results of subject leave-one-out cross-validation of different models with different training methods are compared. Here, we use MSE, MAE, and three evaluation metrics; currently, the SOTA EmoTDMF uses the supervised training method SFT, with the MSE result reaching 0.10, the MAE reaching 0.24, and in reaching 0.59; our method uses the training method of reinforcement learning RL, with the MSE result reaching 0.036, the MAE reaching 0.14, reaching 0.62; the Baseline compared to our proposed method uses supervised learning SFT, with the MSE result reaching 0.25 and the MAE reaching 0.39, reaching 0.43. From the above experimental results, we can know that our proposed method uses the way of reinforcement learning, defines the process of the model predicting the emotional state as a Markov decision process, enables the model to better learn the relationship of EEG feature state transition, captures the subtle dynamic changes of emotions, optimizes the model's sequential prediction ability, and the regression prediction results are greatly improved compared with the existing methods in terms of mean square error and absolute error.
[0091] Reference documents: [1]Zhou, X., Liang, Z., Ye, W., Xue, J., Liu, H., Zhang, M.,&Zhang,Z. (April 2024). EmoTVR: A Hybrid Model to Estimate Continuous-Time andContinuous-Level Emotion from Electroencephalography. In ICASSP 2024-2024IEEE International Conference on Acoustics, Speech and Signal Processing(ICASSP) (pp. 2021-2025). IEEE (Zhou, X., Liang, Z., Ye, W., Xue, J., Liu, H.,Zhang, M.,&Zhang, Z. (April 2024), EmoTVR: A Hybrid Model to Estimate Continuous-Time andContinuous-Level Emotion from Electroencephalography, in ICASSP 2024 - IEEE International Conference on Acoustics, Speech and Signal Processing(ICASSP) (pp. 2021-2025). IEEE); [2]Yang, Z., Dong, W., Li, X., Huang, M., Sun, Y.,&Shi, G. (2023).Vector quantization with self-attention for quality-independentrepresentation learning. In Proceedings of the IEEE / CVF Conference onComputer Vision and Pattern Recognition (pp. 24438-24448) (Yang, Z., Dong, W.,Li, X., Huang, M., Sun, Y.,&Shi, G. (2023), Vector quantization with self-attention for quality-independentrepresentation learning, in Proceedings of the IEEE / CVF Conference onComputer Vision and Pattern Recognition (pp. 24438-24448)).
[0092] In summary, this embodiment mainly reflects the following technical points and advantages: 1. The VQ-VAE is used to perform vector quantization on the EEG signal features, encode them into codebook vectors, learn the potential emotional states in the EEG signal features, and compress the data distribution. The beneficial effects are as follows: By this method, more essential semantic information of the EEG signal is learned, and the potential emotional states are encoded into codebook vectors, obtaining a more robust and computationally efficient representation.
[0093] 2. The self-attention mechanism is used to achieve feature splicing before and after quantization, enhancing the semantic information of individual differences. The beneficial effect is that the important features related to the prediction regression task can be effectively highlighted through the self-attention mechanism.
[0094] 3. The SAC algorithm is used for reinforcement learning training. The beneficial effects are that it can solve the optimal strategy of the continuous action space more efficiently and stably at the same time; reinforcement learning can better learn the relationship of EEG feature state transitions, capture the subtle dynamic changes of emotions, and optimize the model's sequential prediction ability.
[0095] The embodiment of the present application also provides an electronic device, including: a processor, and a memory coupled to the processor. The memory is used to store a computer program; the processor is used to execute the computer program stored in the memory so that the electronic device executes the method described in any one of the above embodiments.
[0096] The electronic device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The electronic device may include, but is not limited to, a processor and a memory.
[0097] The so-called processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc. The processor is the control center of the electronic device, connecting various parts of the entire device through various interfaces and lines.
[0098] The memory can be used to store the computer program. The processor realizes various functions of the electronic device by running or executing the computer program stored in the memory and calling the data stored in the memory.
[0099] The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function, etc.; the data storage area may store data created according to the use of the mobile phone, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, memory, plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.
[0100] An embodiment of the present application also provides a storage medium, which is a computer-readable storage medium, and the computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and flexible media distribution medium, etc.
[0101] An embodiment of the present application also provides a computer program product, including: a computer program or instruction. When the computer program or instruction runs on a computer, the computer is enabled to execute the method of any of the above possible implementation manners.
[0102] The above is the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present application.
Claims
1. An emotion prediction system, characterized in that, The prediction system includes: A vector quantization variational autoencoder for obtaining electroencephalogram (EEG) signals, extracting differential entropy features of the EEG signals that are continuous in time, mapping the differential entropy features to a latent space to obtain feature vectors, performing vector quantization on the feature vectors to obtain codebook features, and concatenating the feature vectors and the codebook features to obtain dynamic emotion depth features; and A flexible actor-critic model under a Markov decision framework with the objectives of maximizing reward and maximizing entropy, including: an actor network and a critic network. The actor network is used to capture the temporal dependencies of the dynamic emotion depth features to obtain depth features for prediction, and generate an action policy according to the current state. The critic network is used to evaluate the value of the state-action pair. The reward function of the flexible actor-critic model includes: a first part for evaluating the current prediction accuracy, and a second part for evaluating the consistency between the change trends of the predicted values at adjacent time steps and the change trend of the true values.
2. The prediction system according to claim 1, wherein The vector quantization variational autoencoder is specifically used for: Discretizing the feature vectors into discrete vectors; and Determining the codebook features by using the discrete vectors, the codewords and indices in the codebook, and performing nearest neighbor search.
3. The prediction system according to claim 1, wherein The loss function adopted by the vector quantization variational autoencoder includes a reconstruction loss term and a quantization loss term.
4. The prediction system according to claim 1, wherein When the vector quantization variational autoencoder concatenates the feature vectors and the codebook features, a self-attention mechanism is introduced.
5. The prediction system according to claim 1, wherein The actor network includes: A long short-term memory network for capturing the temporal dependencies of the dynamic emotion depth features to obtain depth features for prediction; and A fully connected layer for generating an action policy according to the current state.
6. The prediction system according to claim 1, wherein The second part is constructed by the following function: Among them, represents the change in the true value, represents the change in the predicted value, and else represents other cases.
7. A training method for an emotion prediction system, characterized in that, The training method is for training the emotion prediction system according to any one of claims 1-6. The training method includes: Obtaining a training data set, where the training data set includes EEG signals and true emotion values; and Using the training data set to train the emotion prediction system.
8. A method for emotion prediction, characterized in that The prediction method includes: Obtaining the current EEG signal; and Using the emotion prediction system according to any one of claims 1-6 to process the current EEG signal to obtain an emotion prediction result.
9. An electronic device, characterized in that, The electronic device includes: a processor, and a memory coupled to the processor, The memory is used to store a computer program; and The processor is used to execute the computer program stored in the memory so that the electronic device executes the training method according to claim 7, or executes the prediction method according to claim 8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program or instruction. When the computer program or instruction runs on a computer, the computer is caused to execute the training method according to claim 7, or execute the prediction method according to claim 8.
Citation Information
Patent Citations
Emotion recognition method and system based on generative self-supervised learning and electroencephalogram signals
CN115590515A
AI-based pet emotion recognition system
CN119049086A
Unsupervised continuous emotion electroencephalogram analysis method and device based on deep reinforcement learning
CN119474948A
Upsampling of compressed financial time-series data using a jointly trained Vector Quantized Variational Autoencoder neural network
US12229679B1
Monitoring the emotional state of a computer user by analyzing screen capture images
US20140247989A1
Cited By
Continuous dynamic emotion prediction method, training method and equipment
CN122087622A
Continuous dynamic emotion prediction method, training method and device
CN122087622B