Video-to-game cross-scene emotion recognition system

By introducing the dual-stream converter Transformer model (DDCT) into the emotion recognition technology, combining video and game stimulation materials, cross-scene emotion recognition from video to game is achieved, and the problems of single material and simple model in the existing technology are solved, and the generalization and accuracy of emotion recognition are improved.

CN120030413APending Publication Date: 2025-05-23SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510127165.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-31
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In the existing emotion recognition technology, the experimental materials are single and the model is simple, which leads to the generalization and application of emotion recognition being limited, especially in the migration of emotion recognition from video scenes to game scenes.

Method used

A cross-scene emotion recognition system is proposed, using the dual-stream converter Transformer model (DDCT), combining video and games as emotional stimulation materials, and cross-scene emotion recognition from video to game is realized through components such as converter module, source domain discriminator, feature extractor, target domain discriminator and emotion classifier.

Benefits of technology

It effectively improves the generalization of emotion recognition detection, reduces the impact of interference in game scenes on EEG characteristics, better captures emotion-related features, and improves the accuracy of emotion recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030413A_ABST
    Figure CN120030413A_ABST
Patent Text Reader

Abstract

A video-to-game cross-scene emotion recognition system comprises a converter module, a source domain discriminator, a feature extractor, a target domain discriminator and an emotion classifier, the converter module is based on electroencephalogram feature distribution of a source domain and a target domain, a source domain (target domain) converter converts input features into source domain (target domain) features, and the source domain (target domain) features are extracted from the source domain (target domain) features; finally, obtaining features containing source domain information and target domain information at the same time; the source domain discriminator judges whether the output of the source domain converter is a source domain feature according to the distribution of the source domain feature, so that the feature converted by the source domain converter meets the distribution of the source domain feature, and the feature extractor extracts the common information of the two domains according to the feature data which is output by the converter and contains the information of the source domain and the target domain at the same time; more accurate emotion features are obtained; the target domain discriminator judges whether the output of the target domain converter is a target domain feature or not according to the target domain feature distribution, so that the feature converted by the target domain converter meets the target domain feature distribution; the emotion classifier classifies the features according to emotions according to the features extracted by the feature extractor, and an emotion label corresponding to each electroencephalogram sample is obtained. According to the method, the influence of extra interference of a game scene on electroencephalogram features can be reduced, and emotion related features can be better captured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technology in the field of neural network applications, specifically a cross-scene emotion recognition system from video to game based on video and game stimulus data sets. Background Art

[0002] In the field of emotion recognition, existing datasets mainly use passive stimulation materials such as videos, music, and pictures to induce specific emotions. Among them, video stimulation materials are the most effective. However, under passive material stimulation, the subjects' sense of immersion is not strong, and the target emotions may not be effectively stimulated. At the same time, the experimental environment for collecting EEG in existing datasets is generally very ideal (for example, in videos, the subjects need to sit in a specific position and try to stay still), which brings huge challenges to the generalization and application of detection. Summary of the invention

[0003] In view of the shortcomings of the prior art that the experimental materials are still relatively single and the models adopted are relatively simple, the present invention proposes a cross-scene emotion recognition system from video to game, namely, a dual-stream converter Transformer model (DDCT), and simultaneously uses videos and games as emotion stimulation materials to explore the feasibility of emotion recognition migration from video scenes to game scenes, effectively promoting the generalization of emotion recognition detection, and can reduce the impact of additional interference from game scenes on EEG features, so as to better capture emotion-related features.

[0004] The present invention is achieved through the following technical solutions:

[0005] The present invention relates to a cross-scene emotion recognition system from video to game, comprising: a converter module, a source domain discriminator, a feature extractor, a target domain discriminator and an emotion classifier, wherein: the converter module is respectively based on the EEG feature distribution of the source domain and the target domain, the source domain (target domain) converter converts the input feature into the source domain (target domain) feature, and finally obtains the feature containing both the source domain information and the target domain information; the source domain discriminator judges whether the output of the source domain converter is the source domain feature according to the source domain feature distribution, so that the feature converted by the source domain converter satisfies the source domain feature distribution; the feature extractor extracts the information common to the two domains according to the feature data output by the converter containing both the source domain and the target domain information, so as to obtain more accurate emotion features; the target domain discriminator judges whether the output of the target domain converter is the target domain feature according to the target domain feature distribution, so that the feature converted by the target domain converter satisfies the target domain feature distribution; the emotion classifier classifies the features according to the emotions according to the features extracted by the feature extractor, and obtains the emotion label corresponding to each EEG sample.

[0006] The present invention relates to a cross-scenario emotion recognition method based on the above system, comprising:

[0007] Step 1: Collect raw data: According to the international 10-20 system for EEG electrode arrangement, the ESI neural scanning system was used to collect the EEG of the subjects from a 62-channel electrode cap at a sampling rate of 1000 Hz in the video watching scene and the game playing scene.

[0008] Step 2: Data preprocessing, including:

[0009] 2.1 Perform filtering processing on the collected original EEG signals to reduce noise and artifacts;

[0010] 2.2 Use short-time Fourier transform to convert the preprocessed EEG signal in the time domain to the frequency domain, calculate the energy spectrum of the characteristic frequency band in the frequency domain, and then extract the differential entropy (DE) of the energy spectrum;

[0011] 2.3 Use the linear dynamic system (LDS) to remove or weaken EEG signal features that are not related to emotions.

[0012] Step 3: Build a neural network and perform the following six tasks on the preprocessed data to verify the validity of the dataset and the recognition ability of the model:

[0013] Task 1: Emotion recognition in video scenes. In each experiment, three video clips are used as training sets to train the model, and one video clip is used as a validation set to select appropriate hyperparameters.

[0014] Task 2: Emotion recognition in game scenes. In each experiment, three game scenes are used as training sets to train the model, and one game scene is used as a validation set to select appropriate hyperparameters.

[0015] Task 3: Emotion recognition in mixed scenes of video games. In each experiment, 3 video clips and 3 game scenes are used as training sets to train the model, and 1 video clip and 1 game scene are used as validation sets to select appropriate hyperparameters.

[0016] Task 4: Cross-scene emotion recognition from video to game. In each experiment, four video clips are used as training sets to train the model, and one video clip is used as a validation set to select appropriate hyperparameters.

[0017] Task 5: Use domain adaptation method to perform mixed scene emotion recognition. Use 3 video clips and 3 game scenes in each experiment as training sets, and use domain adaptation method for training, that is, add domain labels during training, and use 1 video clip and 1 game scene as validation sets to select hyperparameters. In particular, in order to be consistent with other previous domain adaptation methods, the EEG clips of game scenes in the training set do not use emotion labels, only domain labels.

[0018] Task 6: Use domain adaptation to perform cross-scene emotion recognition from video to game. In each experiment, 4 video clips and 3 game scenes are used as training sets. The domain adaptation method is used for training, and 1 video clip is used as a validation set to select hyperparameters. Similarly, the EEG clips of the game scenes in the training set do not have emotion labels, only domain labels are used.

[0019] Step 4: Use the trained neural network to test the above 6 tasks:

[0020] Task 1: Video scene emotion recognition, using the remaining 1 video clip in each experiment as the test set.

[0021] Task 2: Emotion recognition in game scenes, using one remaining game scene in each experiment as the test set.

[0022] Task 3: Emotion recognition in mixed scenes of video games, using 1 video clip and 1 game scene remaining in each experiment as the test set.

[0023] Task 4: Cross-scene emotion recognition from video to game, using all game scenes in each experiment as the test set.

[0024] Task 5: Use domain adaptation method for mixed scene emotion recognition, using 1 video clip and 1 game scene remaining in each experiment as the test set.

[0025] Task 6: Use domain adaptation methods to perform cross-scene emotion recognition from video to game, using all game scenes in each experiment as the test set.

[0026] The training set, validation set and test set samples in step 3 and step 4 are different from each other, and all contain data from three experiments of the subject. Technical Effects

[0027] The present invention projects the same input to different domains, and at the same time effectively retains the original features of different domains through adversarial training; the features of different domains are used as sequence input, and the Transformer architecture is used to capture the common emotional information between different domains, effectively improving the emotion recognition ability of the model. Compared with the prior art, the present invention effectively reduces the distribution difference of data from different domains; the Transformer architecture is used to capture the common emotional information between different domains, and the common information between the two domain distributions obtained by the previous converter is more accurately captured, that is, the required emotional information, thereby improving the accuracy of the model's emotion recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a schematic diagram of the system of the present invention;

[0029] Figure 2 This is the experimental flow chart;

[0030] Figure 3 This is an example of a game scene; DETAILED DESCRIPTION

[0031] like Figure 1 As shown, this embodiment relates to a cross-scene emotion recognition system from video to game based on video and game stimulus data sets, including: a converter module, a source domain discriminator, a feature extractor, a target domain discriminator and an emotion classifier, wherein: the converter module includes a source domain converter and a target domain converter, the converter module is based on the EEG feature distribution of the source domain and the target domain respectively, the source domain (target domain) converter converts the input feature into the source domain (target domain) feature, and finally obtains the feature containing both the source domain information and the target domain information; the source domain discriminator judges the source domain converter according to the source domain feature distribution. The output of the converter is the source domain feature, so that the feature converted by the source domain converter meets the source domain feature distribution. The feature extractor extracts the information common to the two domains based on the feature data output by the converter that contains both source domain and target domain information, so as to obtain more accurate emotion features. The target domain discriminator judges whether the output of the target domain converter is the target domain feature based on the target domain feature distribution, so that the feature converted by the target domain converter meets the target domain feature distribution. The emotion classifier classifies the features according to emotions based on the features extracted by the feature extractor, and obtains the emotion label corresponding to each EEG sample.

[0032] The converter module includes: a source domain converter and a target domain converter, wherein: the source domain converter projects the input onto the source domain distribution according to the source domain feature distribution to obtain features that conform to the source domain distribution while retaining the original information; the target domain converter projects the input onto the target domain distribution according to the target domain distribution to obtain features that conform to the target domain distribution while retaining the original information.

[0033] The source domain discriminator includes: a fully connected layer, which determines whether the result of the conversion by the source domain converter conforms to the source domain data distribution according to the source domain data distribution, so that the source domain converter meets the desired purpose of projecting the input onto the source domain.

[0034] The feature extractor includes: a position encoding layer, a plurality of multi-head self-attention mechanisms and a plurality of forward propagation layers, wherein: the position encoding layer assigns a specific position encoding to the input according to the information of the two domains, so that the attention mechanism can perceive the specificity of the two domains of the input, the multi-head attention mechanism assigns a separate query (Query), key (Key), and value (Value) to each input according to the specificity of the input data, and calculates according to the previous position encoding to obtain the required features, and the forward propagation layer further calculates through the fully connected layer according to the information obtained by the previous multi-head attention mechanism to further extract features. In particular, the multi-head attention mechanism and the forward propagation layer both use residual connections to maintain gradients for ease of calculation.

[0035] The target domain discriminator includes: a fully connected layer, which determines whether the result of the target domain converter conforms to the target domain data distribution according to the target domain data distribution, so that the target domain converter meets the desired purpose of projecting the input onto the target domain.

[0036] The emotion classifier includes: a fully connected layer, through which the features obtained by the feature extractor are calculated and classified, and finally the emotion label corresponding to the input sample is obtained.

[0037] This embodiment relates to a cross-scenario emotion recognition method based on the above system, including:

[0038] Step 1: Collect raw data, including:

[0039] 1.1 The subjects watched 5 videos designed to stimulate positive or neutral emotions. The videos stimulating different emotions were presented alternately. Each video lasted 3-6 minutes, with a total duration of 20-30 minutes. Between different videos, the subjects had 45 seconds to evaluate the emotion-stimulating effect of the previous video and 15 seconds for a short break.

[0040] 1.2 The subjects entered their own game accounts and completed 5 game scenarios;

[0041] 1.3 The subjects evaluated the emotion arousal effect of the video in step 1 and the game screen recording in step 2.

[0042] Each subject completed three experiments according to the above three steps, with at least one week between each experiment. The video stimulus material was different in each experiment, but the game steps were the same. A total of 15 video clips were used, 8 of which were used to stimulate positive emotions and 7 were used to stimulate neutral emotions.

[0043] Step 2, data preprocessing: Data preprocessing includes two steps, filtering denoising, feature extraction and smoothing. Filtering denoising is to preprocess the collected original EEG signals to reduce noise and remove artifacts; feature extraction is to use short-time Fourier transform to convert the preprocessed EEG signals in the time domain to the frequency domain, calculate the energy spectrum of the characteristic frequency band in the frequency domain, and then extract differential entropy (DE) from the energy spectrum; smoothing is to use the linear dynamic system (LDS) to remove or weaken the EEG signal features that are not related to emotions.

[0044] 2.1 Filtering and denoising The specific operation is to use a bandpass filter with a range of 1 to 75 Hz for filtering.

[0045] 2.2 Feature Extraction The window function used in short-time Fourier transform is Hanning window, and the window size is set to 4 seconds; the selected feature segments include: Delta wave, whose frequency range is 1-4 Hz; Theta wave, whose frequency range is 4-8 Hz; Alpha wave, whose frequency range is 8-14 Hz; Beta wave, whose frequency range is 14-31 Hz; Gamma wave, whose frequency range is 31-50 Hz.

[0046] Step 3: Build a neural network and perform the following six tasks on the preprocessed data to verify the validity of the dataset and the recognition ability of the model:

[0047] Task 1: Video scene emotion recognition (V2V). In each experiment, three video clips are used as training sets to train the model, and one video clip is used as a validation set to select appropriate hyperparameters.

[0048] Task 2: Game scene emotion recognition (G2G). In each experiment, three game scenes are used as training sets to train the model, and one game scene is used as a validation set to select appropriate hyperparameters.

[0049] Task 3: Emotion recognition in mixed scenes of video games (Mix). In each experiment, 3 video clips and 3 game scenes are used as training sets to train the model, and 1 video clip and 1 game scene are used as validation sets to select appropriate hyperparameters.

[0050] Task 4: Video to game cross-scene emotion recognition (V2G). In each experiment, 4 video clips are used as training sets to train the model, and 1 video clip is used as a validation set to select appropriate hyperparameters.

[0051] Task 5: Use domain adaptation method for mixed scene emotion recognition (MDA). In each experiment, 3 video clips and 3 game scenes are used as training sets. The domain adaptation method is used for training, that is, domain labels are added during training, and 1 video clip and 1 game scene are used as validation sets to select hyperparameters. In particular, in order to be consistent with other previous domain adaptation methods, the EEG clips of the game scenes in the training set do not use emotion labels, only domain labels.

[0052] Task 6: Use domain adaptation method to perform video-to-game cross-scene emotion recognition (V2GDA). In each experiment, 4 video clips and 3 game scenes are used as training sets, and 1 video clip is used as a validation set to select hyperparameters. Similarly, the EEG clips of the game scenes in the training set have no emotion labels, only domain labels are used.

[0053] In the tasks five and six, the training steps specifically include:

[0054] 3.1 Source domain training stage: In this stage, only source domain data is used for training. In order to ensure the effect of the source domain converter, the parameters of the source domain converter need to be frozen. During training, the input x passes through the source domain converter and the target domain converter respectively to obtain x s and x t , merge x s and x t , we get x c , x c After the feature extractor and sentiment classifier, the classification loss L can be obtained. c ;x s and x t After passing through the source domain discriminator and the target domain discriminator respectively, the source domain discrimination loss L can be obtained s and target domain discriminative loss L t , these two losses are adversarial losses; x t After the source domain converter, the reconstruction loss L can be obtained. r The total loss in this stage is L = L c +L s +L t +L r ;

[0055] 3.2 Target domain training phase: In this phase, only the target domain data is used for training. In order to ensure the effect of the target domain converter, the parameters of the target domain converter need to be frozen. During training, the input x passes through the source domain converter and the target domain converter respectively to obtain x s and x t , x s and x t After passing through the source domain discriminator and the target domain discriminator respectively, the source domain discrimination loss L can be obtaineds and target domain discriminative loss L t , these two losses are adversarial losses; x s After the target domain converter, the reconstruction loss L can be obtained. r The total loss in this stage is L = L s +L t +L r .

[0056] Step 4: Use the trained neural network to test the above 6 tasks:

[0057] Task 1: Video scene emotion recognition (V2V), using the remaining 1 video clip in each experiment as the test set.

[0058] Task 2: Game scene emotion recognition (G2G), using the remaining 1 game scene in each experiment as the test set.

[0059] Task 3: Emotion recognition in mixed scenes of video games (Mix), using one video clip and one game scene remaining in each experiment as the test set.

[0060] Task 4: Video to game cross-scene emotion recognition (V2G), using all game scenes in each experiment as the test set.

[0061] Task 5: Use domain adaptation method to perform mixed scene emotion recognition (MDA), using 1 video clip and 1 game scene remaining in each experiment as the test set.

[0062] Task 6: Use domain adaptation method for video to game cross-scene emotion recognition (V2GDA), using all game scenes in each experiment as the test set.

[0063] The training set, validation set and test set samples in step 3 and step 4 are different from each other, and all contain data from three experiments of the subject.

[0001] After specific practical experiments, support vector machine (SVM), multi-layer perceptron (MLP), convolutional neural network (CNN), and Transformer were used to test the first four tasks in step 4. The results (mean ± variance) are as follows:

[0002] In the test results of the above four basic models, all the four basic models have good performance in various tasks. Overall, from strong to weak, Transformer is the best, followed by CNN and MLP, and SVM is the worst, which is consistent with the properties of the four models themselves. At the same time, the effects of each model on V2V and G2G are similar, followed by Mix, and V2G is the worst, which is also in line with objective laws. This proves the effectiveness of the experimental paradigm and experimental plan used.

[0003] For the last two tasks, the test results (mean ± variance) are as follows:

[0004] Compared with the prior art, the present invention adopts multi-domain adversarial training and Transformer architecture, which significantly improves the accuracy of emotion recognition in MDA and V2GDA tasks and effectively reduces the variance of subject-dependent emotion recognition, which demonstrates the reliability and generalization of DDCT.

[0005] The above-mentioned specific implementation can be partially adjusted in different ways by those skilled in the art without departing from the principle and purpose of the present invention. The protection scope of the present invention shall be based on the claims and shall not be limited by the above-mentioned specific implementation. Each implementation scheme within its scope shall be subject to the constraints of the present invention.

Claims

1. A cross-scene emotion recognition system from video to game, characterized in that: include: The converter module comprises a source domain discriminator, a feature extractor, a target domain discriminator and an emotion classifier, wherein: the converter module performs conversion based on the EEG feature distribution of the source domain and the target domain respectively to obtain features containing both source domain information and target domain information; the source domain discriminator determines whether the output of the source domain converter is a source domain feature according to the source domain feature distribution, so that the features converted by the source domain converter meet the source domain feature distribution; the feature extractor extracts the information common to the two domains according to the feature data output by the converter containing both source domain and target domain information to obtain more accurate emotion features; the target domain discriminator determines whether the output of the target domain converter is a target domain feature according to the target domain feature distribution, so that the features converted by the target domain converter meet the target domain feature distribution; the emotion classifier classifies the features according to emotions according to the features extracted by the feature extractor to obtain the emotion label corresponding to each EEG sample.

2. The cross-scene emotion recognition system from video to game according to claim 1 is characterized in that: The converter module includes: a source domain converter and a target domain converter, wherein: the source domain converter projects the input onto the source domain distribution according to the source domain feature distribution to obtain features that conform to the source domain distribution while retaining the original information; the target domain converter projects the input onto the target domain distribution according to the target domain distribution to obtain features that conform to the target domain distribution while retaining the original information.

3. The cross-scene emotion recognition system from video to game according to claim 1 is characterized in that: The source domain discriminator includes: a fully connected layer, which determines whether the result of the conversion of the source domain converter conforms to the source domain data distribution according to the source domain data distribution, so that the source domain converter meets the desired purpose of projecting the input onto the source domain.

4. The cross-scene emotion recognition system from video to game according to claim 1 is characterized in that: The feature extractor includes: a position encoding layer, several multi-head self-attention mechanisms and several forward propagation layers, wherein: the position encoding layer assigns a specific position encoding to the input according to the information of the two domains, so that the attention mechanism can perceive the specificity of the two domains of the input; the multi-head attention mechanism assigns a separate query (Query), key (Key), and value (Value) to each input according to the specificity of the input data, and calculates according to the previous position encoding to obtain the required features; the forward propagation layer further calculates through the fully connected layer based on the information obtained by the previous multi-head attention mechanism to further extract features.

5. The cross-scene emotion recognition system from video to game according to claim 4 is characterized in that: The multi-head attention mechanism and forward propagation layer both use residual connections to maintain gradients.

6. The cross-scene emotion recognition system from video to game according to claim 1 is characterized in that: The target domain discriminator includes: a fully connected layer, which determines whether the result of the target domain converter conforms to the target domain data distribution according to the target domain data distribution, so that the target domain converter meets the desired purpose of projecting the input onto the target domain.

7. The cross-scene emotion recognition system from video to game according to claim 1 is characterized in that: The emotion classifier includes: a fully connected layer, through which the features obtained by the feature extractor are calculated and classified, and finally the emotion label corresponding to the input sample is obtained.

8. A cross-scenario emotion recognition method based on the system described in any one of claims 1 to 7, characterized in that: include: Step 1: Collect raw data: According to the international standard 10-20 system for EEG electrode arrangement, the ESI neural scanning system was used to collect the EEG of the subjects from the 62-channel electrode cap at a sampling rate of 1000 Hz in the video watching scene and the game playing scene; Step 2: Data preprocessing, including: 2.1 Perform filtering processing on the collected original EEG signals to reduce noise and artifacts; 2.2 Use short-time Fourier transform to convert the preprocessed EEG signal in the time domain to the frequency domain, calculate the energy spectrum of the characteristic frequency band in the frequency domain, and then extract the differential entropy of the energy spectrum; 2.3 Use linear dynamical systems to remove or weaken EEG signal features that are not related to emotions; Step 3: Build a neural network and perform complex tasks on the preprocessed data to verify the validity of the data set and the recognition ability of the model; Step 4: Use the trained neural network to test complex tasks.

9. The cross-scene emotion recognition method according to claim 8, characterized in that: The complex tasks described include: Task 1: Emotion recognition in video scenes. In each experiment, three video clips are used as training sets to train the model, and one video clip is used as a validation set to select appropriate hyperparameters. Task 2: Emotion recognition in game scenes. In each experiment, three game scenes are used as training sets to train the model, and one game scene is used as a validation set to select appropriate hyperparameters. Task 3: Emotion recognition in mixed scenes of video games. In each experiment, 3 video clips and 3 game scenes are used as training sets to train the model, and 1 video clip and 1 game scene are used as validation sets to select appropriate hyperparameters. Task 4: Cross-scene emotion recognition from video to game. In each experiment, four video clips are used as training sets to train the model, and one video clip is used as a validation set to select appropriate hyperparameters. Task 5: Use domain adaptation method to perform mixed scene emotion recognition. Take 3 video clips and 3 game scenes in each experiment as training sets, and use domain adaptation method for training, that is, add domain labels during training, and use 1 video clip and 1 game scene as validation sets to select hyperparameters; in particular, in order to be consistent with other previous domain adaptation methods, the EEG clips of game scenes in the training set do not use emotion labels, only domain labels; Task 6: Use domain adaptation method to perform cross-scene emotion recognition from video to game. In each experiment, 4 video clips and 3 game scenes are used as training sets. The domain adaptation method is used for training, and 1 video clip is used as the validation set to select hyperparameters. Similarly, the EEG clips of the game scenes in the training set have no emotion labels, and only domain labels are used.

10. The cross-scene emotion recognition method according to claim 8, characterized in that: The training set, validation set and test set samples in step 3 and step 4 are different from each other, and all contain data from three experiments of the subject.