Emotion recognition method and device, computer equipment and readable storage medium
By collecting and analyzing eye movement data when users watch videos in real time, identifying the target objects that users are concerned about and extracting eye movement characteristics, the real-time and precision challenges of emotion recognition in the prior art are solved, and efficient and accurate emotion recognition and dynamic feedback display are achieved.
Patent Information
- Application Number
- CN202510064565.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art has real-time and accuracy challenges in emotion recognition, especially in the case of rapid changes in targets in dynamic videos, and noise problems of EEG signals and high-dimensional data differences lead to high accuracy and cost of emotion recognition.
By collecting eye movement data frames when the user watches video in real time, determining the target object that the user is concerned about, filtering the target eye movement data frame sequence, extracting eye movement characteristics, and inputting a preset emotion classification model to obtain the user's emotion type for the target object.
It realizes real-time and accurate capture of user emotional fluctuations during video viewing, breaks through the limitations of traditional methods, improves the multi-objective recognition ability and dynamic feedback display of emotional recognition, and reduces the complexity and cost of the system.
Smart Images

Figure CN119924836A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of emotion recognition technology, and in particular to an emotion recognition method, device, computer equipment and readable storage medium. Background Art
[0002] At present, eye movement technology has made some progress in the field of target analysis and emotion assessment, but real-time and accuracy are still important challenges in practical applications. The rapid changes of targets in dynamic videos put forward higher requirements on the adaptability and stability of algorithms. At the same time, the complex relationship between emotion and target importance also requires the integration of multimodal data and the support of deep learning algorithms to analyze.
[0003] There are many problems with the related methods in the existing technology. For example, the limitations of emotional state, the limitations of feature selection, and the limited time series modeling capabilities of the model; for example, the accuracy of emotion recognition through EEG signals is greatly affected by irremovable noise, the high maintenance cost of the EEG system, and the great differences in the dimensions and time-frequency characteristics of EEG data, resulting in many fusion problems.
[0004] There is currently no effective solution to the above defects. Summary of the invention
[0005] The object of the present invention is to provide an emotion recognition method, apparatus, computer device and readable storage medium that can solve the above-mentioned technical problems.
[0006] According to one aspect of the present invention, there is provided an emotion recognition method, the method comprising: Collect eye movement data frames of users when they watch videos in real time; each eye movement data frame is a frame of eye movement data collected based on a preset eye movement collection algorithm; Based on the eye movement data frame, determining the target object that the user is paying attention to from the objects displayed in the video; Filtering out a target eye movement data frame sequence for focusing on the target object from the collected eye movement data frames; Extracting eye movement features from the target eye movement data frame sequence; The extracted eye movement features are input into a preset emotion classification model to obtain the emotion type of the user towards the target object.
[0007] Optionally, determining the target object that the user is paying attention to from the objects displayed in the video based on the eye movement data frame includes: detecting an object in each video frame of the video and generating a bounding box around the object; Extracting the coordinates of the user's sight point from the eye movement data frame; When the coordinates of the sight point of N consecutive eye movement data frames are all located in the same bounding box, and N is a positive integer greater than or equal to a preset frame number threshold, determining that the object in the bounding box is the target object; The step of selecting a target eye movement data frame sequence for focusing on the target object from the collected eye movement data frames comprises: The N consecutive eye movement data frames are determined as the target eye movement data frame sequence.
[0008] Optionally, extracting eye movement features from the target eye movement data frame sequence includes: Determining the eye movement behavior type to which each eye movement data frame in the target eye movement data frame sequence belongs; When the target eye movement data frame sequence includes eye movement data frames belonging to the gaze behavior type, extracting features and pupil features for characterizing the gaze behavior from all eye movement data frames belonging to the gaze behavior type; When the target eye movement data frame sequence includes eye movement data frames belonging to the saccadic behavior type, the features for characterizing the saccadic behavior and the pupil features are extracted from all eye movement data frames belonging to the saccadic behavior type.
[0009] Optionally, extracting eye movement features from the target eye movement data frame sequence includes: Determining the eye movement behavior type to which each eye movement data frame in the target eye movement data frame sequence belongs; When the target eye movement data frame sequence includes eye movement data frames belonging to the gaze behavior type, extracting features and pupil features for characterizing the gaze behavior from all eye movement data frames belonging to the gaze behavior type; When the target eye movement data frame sequence includes eye movement data frames belonging to the saccadic behavior type, the features for characterizing the saccadic behavior and the pupil features are extracted from all eye movement data frames belonging to the saccadic behavior type.
[0010] Optionally, determining the eye movement behavior type to which each eye movement data frame in the target eye movement data frame sequence belongs includes: Calculating the eye movement speed between every two adjacent eye movement data frames in the target eye movement data frame sequence; When the eye movement speed is less than a preset speed threshold, determining the last eye movement data frame of two corresponding adjacent eye movement data frames as a gaze behavior type; When the eye movement speed is greater than a preset speed threshold, the last eye movement data frame of the corresponding two adjacent eye movement data frames is determined as a scanning behavior type.
[0011] Optionally, the features used to characterize gaze behavior include at least one of the following: features used to characterize gaze frequency and features used to characterize gaze time; the features used to characterize scanning behavior include at least one of the following: features used to characterize scanning frequency, features used to characterize scanning time, features used to characterize scanning amplitude and features used to characterize scanning speed; the pupil features include: features used to characterize pupil size.
[0012] Optionally, after inputting the extracted eye movement features into a preset emotion classification model to obtain the emotion type of the user towards the target object, the method further comprises: The emotion type is displayed at the target object shown in the video; wherein the emotion type is used to characterize the degree of emotional awakening of the user by the target object.
[0013] Optionally, the method further comprises: Determining a current eye movement data frame sequence from the collected eye movement data frames according to a sliding window of a preset length; Determining in the video a target position mapped by the sight point coordinates of each eye movement data frame in the current eye movement data frame sequence; Add a preset Gaussian template at each target location; Based on the added Gaussian template, the sight track of the user is constructed in the video.
[0014] In order to achieve the above object, the present invention further provides an emotion recognition device, the device comprising: The first acquisition module is used to acquire eye movement data frames of the user when watching the video in real time; wherein each eye movement data frame is a frame of eye movement data acquired based on a preset eye movement acquisition algorithm; A first determination module, configured to determine, based on the eye movement data frame, a target object that the user is paying attention to from objects displayed in the video; A screening module, used to screen out a target eye movement data frame sequence for focusing on the target object from the collected eye movement data frames; An extraction module, used for extracting eye movement features from the target eye movement data frame sequence; The second determination module is used to input the extracted eye movement features into a preset emotion classification model to obtain the emotion type of the user towards the target object.
[0015] In order to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor is used to implement the steps of an emotion recognition method introduced above when executing the computer program.
[0016] In order to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it is used to implement the steps of an emotion recognition method introduced above.
[0017] The emotion recognition method, device, computer equipment and readable storage medium provided by the present invention have the following beneficial effects: First, solving the limitations of traditional methods: This system breaks through the limitations of eye movement data collection and processing in traditional psychological experiments, especially in emotion classification, and can more accurately capture the emotional fluctuations of individuals in specific scenarios.
[0018] Second, real-time collection and analysis: The system can collect eye movement data in real time while the user is watching the video, and perform efficient processing and emotion classification to achieve real-time monitoring of the user's emotional changes.
[0019] Third, multi-target emotion recognition: For different targets in the video, the system can identify the emotional changes of users towards different targets, providing a more refined reference for emotion research and psychological state analysis.
[0020] Fourth, dynamic feedback and visual display: The system can present users’ attention behaviors and emotional changes in real time and dynamically, making the emotional analysis results more vivid and intuitive, and enhancing the interactivity of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings: Figure 1 A flowchart of the emotion recognition method provided in Example 1; Figure 2 A flowchart for determining a target object provided in Example 1; Figure 3 A schematic diagram of the emotion recognition method provided in Example 1; Figure 4A A schematic diagram of a Gaussian template provided in Example 1; Figure 4B A schematic diagram of the superposition of multiple Gaussian templates provided in Example 1; Figure 5 A block diagram of an emotion recognition device provided in Example 2; Figure 6 A block diagram of a computer device suitable for implementing the emotion recognition method provided in Example 3. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical scheme and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0023] At present, one existing technology is to combine decision trees and neural networks for emotion classification based on eye movement data. The main purpose is to use the hierarchical decision-making ability of decision trees and the nonlinear mapping advantages of neural networks to improve classification performance. Through the analysis of experimental data, it is found that the pupil and gaze behaviors of negative emotions (Negative) are significantly different from those of the other two types of emotions (neutral and positive). The behavioral differences between neutral (Neutral) and positive (Positive) are small. Therefore, the model design follows the idea of divide and conquer. First, a binary classification neural network (NN1) is used to classify the emotional state into "negative" or "non-negative". Then, for "non-negative" emotions, another binary classification neural network (NN2) is used to further distinguish them into "neutral" and "positive".
[0024] The above-mentioned emotion assessment based on eye movement features has several shortcomings. First, the limitation of emotional state. Emotions are usually divided into multiple dimensions, such as pleasure, arousal, and dominance. However, pupil size and focus position reflect arousal more, but have limited expressiveness for pleasure and dominance. In addition, the hybrid structure of the model may need to significantly increase the number of network nodes and decision levels when dealing with more emotion categories (such as more than 3 categories), resulting in a rapid increase in complexity. The decision tree structure limits the scalability of the model and cannot flexibly handle more complex classification tasks. Therefore, the emotion assessment in this scheme can only simply divide emotions into three categories: positive, neutral, and negative, and cannot evaluate more delicate and multi-dimensional emotion types such as happiness, sadness, and fear. Second, the limitation of feature selection. The input features only include pupil size and focus position. These features may not fully reflect the emotional state. The change of pupil size will be affected by external stimuli (such as light) and emotional factors. Its signal will be stimulated by the external environment and has strong ambiguity. In addition, some emotional states (such as fear and surprise) may be similar in arousal, but belong to different categories and may be difficult to distinguish only by pupil changes and gaze behavior. Third, the model's time series modeling capabilities are limited. The current hybrid model uses a fixed sliding window to capture time dependencies. The length of the sliding window is manually selected and may not be able to adapt to the optimal time range in different data. This method cannot dynamically adapt to long-term dependencies and cannot effectively capture complex time patterns.
[0025] Another existing technique is to collect EEG and eye movement signals from participants and record these data simultaneously while watching movie clips with rich emotional content, selecting different clips for each emotion category (happy, sad, fearful, disgusted, neutral) for the experiment.
[0026] This technology has the following disadvantages: First, the noise problem of EEG signals. EEG signals are easily affected by factors such as eye movement artifacts, electromyographic activity, and environmental electromagnetic interference. Although bandpass filtering and principal component analysis (PCA) are used to remove noise, these methods cannot completely eliminate artifacts and noise. Therefore, the accuracy of emotion recognition through EEG signals is greatly affected by irremovable noise. Second, EEG equipment is complex and expensive. EEG equipment usually requires multiple electrodes (such as 64, 128 or more) to capture EEG signals. The operation requires professional personnel to install and adjust, and the installation and adjustment process is very complicated and cumbersome. In addition, professional EEG equipment, especially high-density multi-channel EEG systems, is expensive and may cost thousands or even tens of thousands of dollars. Moreover, it is also costly to maintain and calibrate these devices. Third, data dimensions and feature differences. EEG data is usually time series data, and each signal channel records brain activity in different frequency bands (such as δ, θ, α, β, and γ bands). The data of each EEG channel is usually a continuous signal, which makes it highly complex in space and time. Eye movement data, such as pupil diameter, eye movement speed, saccade amplitude, fixation duration, etc., are discrete, and some features may even be statistical information based on events or fixed time periods. Therefore, the data dimensions and time-frequency characteristics of the two are very different, resulting in many problems in their fusion.
[0027] In order to overcome the above-mentioned defects, the present invention patent aims to design a real-time evaluation system for video targets, specifically solving the limitations of traditional methods in obtaining eye movement data in psychological experiment processes. The traditional eye movement data collection and processing methods based on psychological experiment design are general and vague in terms of emotion classification, and cannot accurately capture the emotional fluctuations of individuals in specific scenes. To this end, the present invention innovatively proposes a solution for video objects. Through this system, the eye movement data of users for specific objects during video viewing can be collected in real time, and efficiently processed and analyzed. The system can not only classify the eye movement data for emotions, but also accurately identify the emotional fluctuations generated by users for different objects during viewing, and accurately locate and analyze the user's emotional state in combination with the time point of emotional changes. In addition, the technology can also identify the changes in the user's emotional state for different targets in the scene, thereby providing a more refined and real-time reference basis for emotional research and psychological state analysis. See the following embodiments for details.
[0028] Embodiment 1 The embodiment of the present invention provides an emotion recognition method, such as Figure 1 As shown, the method includes steps S1 to S5, wherein: Step S1, collecting eye movement data frames of a user when watching a video in real time; wherein each eye movement data frame is a frame of eye movement data collected based on a preset eye movement collection algorithm.
[0029] The video includes a plurality of continuous video frames. The preset eye movement acquisition algorithm can be set in an internal system or in an external device, such as an external eye tracker.
[0030] The eye movement data frame includes: timestamp, horizontal and vertical coordinates of the gaze point, average pupil size, etc.
[0031] Step S2: determining the target object that the user is paying attention to from the objects displayed in the video based on the eye movement data frame.
[0032] Each video frame is used to display an image, and the image displayed by each video frame includes several objects. When the eye movement data is used to characterize that the user pays attention to the same object for a long time, the object is determined as the target object. Among them, this embodiment is described by taking one of the target objects that the user pays attention to as an example, and for other target objects that the user pays attention to, operations can be performed according to the steps described in this embodiment.
[0033] Optionally, determining the target object that the user is paying attention to from the objects displayed in the video based on the eye movement data frame includes: detecting an object in each video frame of the video and generating a bounding box around the object; Extracting the coordinates of the user's sight point from the eye movement data frame; When the coordinates of the sight point of N consecutive eye movement data frames are all located in the same bounding box, and N is a positive integer greater than or equal to a preset frame number threshold, determining that the object in the bounding box is the target object; The sight point coordinates include the sight point abscissa and the sight point ordinate. Each object corresponds to a bounding box. When the sight point coordinates are within a bounding box, the eye movement data frame to which the sight point coordinates belong is determined to focus on the object within the bounding box.
[0034] Optionally, when the sight point coordinates of N consecutive eye movement data frames are all located in the same bounding box, and N is a positive integer greater than or equal to a preset frame number threshold, determining that the object in the bounding box is the target object includes: Determining whether the sight point coordinates of the eye movement data frame are within the bounding box; If so, retain the eye movement data frame, and continue to determine whether the sight point coordinates of the next eye movement data frame are within the bounding box, until all eye movement data frames are determined; If not, determine whether before the eye movement data frame, there are N consecutive eye movement data frames whose sight point coordinates are all located in the same bounding box, and N is a positive integer greater than or equal to the preset frame number threshold. If so, determine that the object in the bounding box is the target object, and continue to determine whether the sight point coordinates of the next eye movement data frame are located in the bounding box until all eye movement data frames are determined; otherwise, directly continue to determine whether the sight point coordinates of the next eye movement data frame are located in the bounding box until all eye movement data frames are determined.
[0035] The above operation is performed once for each eye movement data frame.
[0036] It should be noted that, to determine whether there are N consecutive eye movement data frames before the eye movement data frame, whose sight point coordinates are all located in the same bounding box, and N is a positive integer greater than or equal to the preset frame number threshold, specifically: to determine whether there is an eye movement data frame immediately adjacent to the eye movement data frame before the eye movement data frame, and whose sight point coordinates are all located in the same bounding box. If it exists, and the number of frames of these eye movement data frames is greater than or equal to the preset frame number threshold, determine that the object in the bounding box is the target object, and use the determined eye movement data frames as the target eye movement data frame sequence, and continue to determine whether the sight point coordinates of the next eye movement data frame are located in the bounding box until all the eye movement data frames are determined; if it exists, and the number of frames of these eye movement data frames is less than the preset frame number threshold, delete the retained eye movement data frames, and continue to determine whether the sight point coordinates of the next eye movement data frame are located in the bounding box until all the eye movement data frames are determined; if it does not exist, directly continue to determine whether the sight point coordinates of the next eye movement data frame are located in the bounding box until all the eye movement data frames are determined.
[0037] Step S3, filtering out a target eye movement data frame sequence for focusing on the target object from the collected eye movement data frames.
[0038] The target eye movement data frame sequence includes a plurality of continuous eye movement data frames, and the number of frames of the plurality of continuous eye movement data frames is greater than or equal to a preset frame number threshold.
[0039] Optionally, the step of filtering out a target eye movement data frame sequence for focusing on the target object from the collected eye movement data frames includes: determining the N consecutive eye movement data frames as the target eye movement data frame sequence.
[0040] Optionally, filtering out a target eye movement data frame sequence for focusing on the target object from the collected eye movement data frames includes: filtering out from the collected eye movement data frames all eye movement data frames whose gaze point coordinates are within a bounding box of the target object and whose frame numbers are continuous.
[0041] Step S4, extracting eye movement features from the target eye movement data frame sequence.
[0042] The eye movement features include pupil features and features for characterizing gaze behavior and / or features for characterizing scan behavior. When the target eye movement data frame sequence includes eye movement data frames of the gaze behavior type, the eye movement features include pupil features and features for characterizing gaze behavior; when the target eye movement data frame sequence includes eye movement data frames of the scan behavior type, the eye movement features include pupil features and features for characterizing scan behavior.
[0043] The features used to characterize the gaze behavior include at least one of the following: features used to characterize the gaze frequency and features used to characterize the gaze duration. Among them, the features used to characterize the gaze frequency include at least one of the following: the gaze frequency and the gaze frequency; the features used to characterize the gaze duration include at least one of the following: the average gaze duration, the gaze duration standard deviation, the gaze duration kurtosis, the gaze duration skewness, the first gaze time and the first gaze duration.
[0044] The features used to characterize the scanning behavior include at least one of the following: features used to characterize the scanning frequency, features used to characterize the scanning time, features used to characterize the scanning amplitude, and features used to characterize the scanning speed. Among them, the features used to characterize the scanning frequency include at least one of the following: the scanning frequency and the scanning frequency; the features used to characterize the scanning time include at least one of the following: the average scanning time, the scanning time standard deviation, the scanning time kurtosis, and the scanning time skewness; the features used to characterize the scanning amplitude include at least one of the following: the average scanning amplitude, the scanning amplitude standard deviation, the scanning amplitude kurtosis, the scanning amplitude skewness, and the scanning amplitude peak value; the features used to characterize the scanning speed include at least one of the following: the average scanning speed, the scanning speed standard deviation, the scanning speed kurtosis, the scanning speed skewness, and the scanning speed peak value.
[0045] The pupil features include: features for characterizing pupil size. The features for characterizing pupil size include at least one of the following: average pupil size, pupil size standard deviation, pupil size kurtosis and pupil size skewness.
[0046] As shown in Table 1 below, various eye movement features and calculation methods of the eye movement features are listed.
[0047] Table 1 Eye movement characteristics and calculation methods
[0048] Among them, n is the number of fixations, s is the total fixation duration, is the fixation duration of the ith fixation, is the average fixation duration; is the first fixation duration, is the duration of the first fixation; m is the number of glances, w is the total duration of the glance, is the duration of the i-th scan, is the average scanning time; is the scanning amplitude of the ith scanning, is the average scanning time; is the scanning speed of the i-th scanning, is the average scanning speed; x is the number of pupils, is the pupil size of the i-th eye movement data frame, is the average pupil size. The pupil size refers to the pupil diameter.
[0049] It should be noted that the eye movement data frame of a fixation is a plurality of continuous eye movement data frames belonging to the fixation behavior type; the eye movement data frame before the first eye movement data frame in the plurality of continuous eye movement data frames is a saccade behavior type, and the eye movement data frame next to the last eye movement data frame in the plurality of continuous eye movement data frames is also a saccade behavior type. Correspondingly, the eye movement data frame of a saccade is a plurality of continuous eye movement data frames belonging to the saccade behavior type; the eye movement data frame before the first eye movement data frame in the plurality of continuous eye movement data frames is a fixation behavior type, and the eye movement data frame next to the last eye movement data frame in the plurality of continuous eye movement data frames is also a fixation behavior type.
[0050] Optionally, extracting eye movement features from the target eye movement data frame sequence includes: Determining the eye movement behavior type to which each eye movement data frame in the target eye movement data frame sequence belongs; When the target eye movement data frame sequence includes eye movement data frames belonging to the gaze behavior type, extracting features and pupil features for characterizing the gaze behavior from all eye movement data frames belonging to the gaze behavior type; When the target eye movement data frame sequence includes eye movement data frames belonging to the saccadic behavior type, the features for characterizing the saccadic behavior and the pupil features are extracted from all eye movement data frames belonging to the saccadic behavior type.
[0051] Among them, the eye movement behavior type includes a fixation behavior type and a saccade behavior type. When the target eye movement data frame sequence only includes eye movement data frames belonging to the fixation behavior type, the extracted eye movement features include features for characterizing fixation behavior and pupil features; when the target eye movement data frame sequence only includes eye movement data frames belonging to the saccade behavior type, the extracted eye movement features include features for characterizing saccade behavior and pupil features; when the target eye movement data frame sequence includes both eye movement data frames belonging to the fixation behavior type and eye movement data frames belonging to the saccade behavior type, the extracted eye movement features include features for characterizing fixation behavior, features for characterizing saccade behavior and pupil features.
[0052] Optionally, determining the eye movement behavior type to which each eye movement data frame in the target eye movement data frame sequence belongs includes: Calculating the eye movement speed between every two adjacent eye movement data frames in the target eye movement data frame sequence; When the eye movement speed is less than a preset speed threshold, determining the last eye movement data frame of two corresponding adjacent eye movement data frames as a gaze behavior type; When the eye movement speed is greater than a preset speed threshold, the last eye movement data frame of the corresponding two adjacent eye movement data frames is determined as a scanning behavior type.
[0053] After extracting the target eye movement data frame sequence, it is classified by behavior. The classification algorithm used is mainly based on the I-VT (Velocity-Threshold Identification) algorithm. The I-VT algorithm can effectively distinguish different eye movement behaviors by analyzing the changes in eye movement speed. Specifically, when the human eye looks at an object, the sight point vibrates slightly within a small range, and the eye movement speed is low. When the eyes are scanning the screen, the sight point moves greatly in a short period of time, and the eye movement speed increases significantly. Therefore, by calculating the difference in eye movement speed between adjacent frames, the type of eye movement behavior can be identified.
[0054] Step S5: input the extracted eye movement features into a preset emotion classification model to obtain the emotion type of the user towards the target object.
[0055] The emotion classification model is trained using a random forest model. The emotion type is used to characterize the degree of emotional arousal of the target object to the user. Specifically, the emotion types include low arousal, medium arousal, and high arousal. The higher the arousal, the stronger the emotional response and the more excited the performance. By analyzing and classifying the eye movement data, the user's emotional state can be captured more accurately.
[0056] Optionally, after inputting the extracted eye movement features into a preset emotion classification model to obtain the emotion type of the user towards the target object, the method further comprises: The emotion type is displayed at the target object shown in the video; wherein the emotion type is used to characterize the degree of emotional awakening of the user by the target object.
[0057] Specifically, the corresponding emotion type is displayed around the bounding box of the target object.
[0058] Optionally, the method further comprises: Determining a current eye movement data frame sequence from the collected eye movement data frames according to a sliding window of a preset length; Determining in the video a target position mapped by the sight point coordinates of each eye movement data frame in the current eye movement data frame sequence; Add a preset Gaussian template at each target location; Based on the added Gaussian template, the sight track of the user is constructed in the video.
[0059] All eye movement data frames within the sliding window are the current eye movement data frame sequence. Among them, the eye movement data frames in the current eye movement data frame sequence are continuous frames, and the sliding step length of the sliding window is less than the length of the sliding window. When the collected eye movement data frames are slide-detected using a sliding window of a preset length, all eye movement data frames within the current sliding window in the collected eye movement data frames are taken as the current eye movement data frame sequence. After each sliding of the sliding window, the above operation is performed once. The present invention takes the eye movement data frame sequence within a certain sliding window as an example to explain the scheme.
[0060] Among them, the Gaussian template is used to display the target position mapped by the coordinates of the sight point in the picture shown in the video. There are multiple values in the Gaussian template, the highest value is located at the center of the Gaussian template, and the other values are arranged around the center. For example, the Gaussian template is an n-order matrix, n is a positive integer greater than 1, each element in the n-order matrix is a specific value, and the highest value has only one and is located at the center of the n-order matrix. When adding a Gaussian template at the target position, the highest value of the Gaussian template is set at the target position, that is, the highest value of the Gaussian template is located at the position where the user's sight falls. Optionally, the color can be set in the Gaussian template to enhance the presentation effect of the user's sight track, wherein the same value presents the same color, and the higher the value, the darker the corresponding color. As the user's sight moves, the Gaussian template is continuously superimposed. In the superimposed Gaussian template, the value of the overlapping part becomes larger, and the corresponding color becomes darker, indicating that the user has been looking at this position for a long time. Through the added Gaussian template, the user's sight track is formed in the video.
[0061] like Figure 4A As shown, it is assumed that each Gaussian template is a 3-order matrix; Figure 4B As shown, an example is given to explain the effect of superimposing three Gaussian templates, wherein the value 8 of each Gaussian template is set at the corresponding target position. After setting a Gaussian template at each of the three target positions, the overlapping parts of the three Gaussian templates are superimposed on each other, and the values become larger.
[0062] Optionally, when the number of added Gaussian templates reaches a preset number threshold, the Gaussian templates of the preset number threshold are deleted in order of the adding time from the smallest to the largest, so as to avoid affecting the user's viewing of the video.
[0063] This embodiment can superimpose historical sight points on the original video screen to intuitively display the user's visual attention. This function can present the changes of the human eye's sight point over time by drawing trajectory lines on the video frame or using color markings, so that the user can clearly see the movement path of the sight and the hot spots of attention. At the same time, the emotion classification results of the target object will be superimposed on the video in real time. Based on the results of real-time emotion analysis, the area where the target object in the video is located will display the corresponding emotional state, such as the classification results of low arousal, medium arousal or high arousal. The current emotion classification results can be intuitively reflected by displaying text information, color coding or other graphical elements on the target object. This not only enhances the visualization of the results, but also allows users to intuitively perceive and understand the emotion classification judgment of the system.
[0064] like Figure 3 As shown, the present invention is mainly composed of three modules, namely, an eye movement feature calculation module based on a target, an emotion recognition module, and a visualization display module. The present invention first extracts eye movement data related to the object based on the ROI area (region of interest) of the video image and the target. The extracted original eye movement data is processed to calculate and extract eye movement features. These features are then input into the emotion recognition module, and the eye movement features are classified by a pre-trained random forest model to obtain the emotion classification results corresponding to the target object. Finally, through the visualization module, the emotion classification results of the visual salient area and the target can be displayed in real time on the original video. The real-time display function effectively reveals the focus of the human eye and intuitively reflects the target that the subject is paying attention to in the video. Unlike the traditional method that only relies on the analysis of the entire segment of eye movement data, the present invention can accurately locate the specific target that the human eye is paying attention to, and provide real-time feedback on the emotional state of the subject to different targets. This method significantly improves the accuracy and real-time performance of emotion analysis, and can provide more detailed reference data for further emotion research.
[0065] The entire real-time display process of the present invention can provide dynamic feedback, so that users can observe gaze behavior and emotional changes in real time in the video. Through this visual display, it can not only help researchers or application developers better understand the user's behavior and emotional state, but also provide powerful technical support for real-time monitoring, advertising effect evaluation, user experience optimization and other scenarios. This real-time display function greatly improves the interactivity and practicality of the system, making the results of eye movement analysis and emotion classification more vivid and intuitive.
[0066] Embodiment 2 The embodiment of the present invention provides an emotion recognition device, such as Figure 5 As shown, the emotion recognition device 50 specifically includes the following components: The acquisition module 501 is used to acquire eye movement data frames of a user watching a video in real time; wherein each eye movement data frame is a frame of eye movement data acquired based on a preset eye movement acquisition algorithm; A first determination module 502 is used to determine the target object that the user is paying attention to from the objects displayed in the video based on the eye movement data frame; A screening module 503 is used to screen out a target eye movement data frame sequence for focusing on the target object from the collected eye movement data frames; An extraction module 504 is used to extract eye movement features from the target eye movement data frame sequence; The second determination module 505 is used to input the extracted eye movement features into a preset emotion classification model to obtain the emotion type of the user towards the target object.
[0067] Optionally, the first determining module is specifically configured to: detecting an object in each video frame of the video and generating a bounding box around the object; Extracting the coordinates of the user's sight point from the eye movement data frame; When the coordinates of the sight point of N consecutive eye movement data frames are all located in the same bounding box, and N is a positive integer greater than or equal to a preset frame number threshold, determining that the object in the bounding box is the target object; The screening module is specifically used for: The N consecutive eye movement data frames are determined as the target eye movement data frame sequence.
[0068] Optionally, the extraction module is specifically used for: Determining the eye movement behavior type to which each eye movement data frame in the target eye movement data frame sequence belongs; When the target eye movement data frame sequence includes eye movement data frames belonging to the gaze behavior type, extracting features and pupil features for characterizing the gaze behavior from all eye movement data frames belonging to the gaze behavior type; When the target eye movement data frame sequence includes eye movement data frames belonging to the saccadic behavior type, the features for characterizing the saccadic behavior and the pupil features are extracted from all eye movement data frames belonging to the saccadic behavior type.
[0069] Optionally, when determining the eye movement behavior type to which each eye movement data frame in the target eye movement data frame sequence belongs, the extraction module is specifically configured to: Calculating the eye movement speed between every two adjacent eye movement data frames in the target eye movement data frame sequence; When the eye movement speed is less than a preset speed threshold, determining the last eye movement data frame of two corresponding adjacent eye movement data frames as a gaze behavior type; When the eye movement speed is greater than a preset speed threshold, the last eye movement data frame of the corresponding two adjacent eye movement data frames is determined as a scanning behavior type.
[0070] Optionally, the features used to characterize gaze behavior include at least one of the following: features used to characterize gaze frequency and features used to characterize gaze time; the features used to characterize scanning behavior include at least one of the following: features used to characterize scanning frequency, features used to characterize scanning time, features used to characterize scanning amplitude and features used to characterize scanning speed; the pupil features include: features used to characterize pupil size.
[0071] Optionally, the method also includes: a display module, used to display the emotion type at the target object displayed in the video after the extracted eye movement features are input into a preset emotion classification model to obtain the emotion type of the user towards the target object; wherein the emotion type is used to characterize the degree of emotional arousal of the target object to the user.
[0072] Optionally, the method further comprises: A third determination module, configured to determine a current eye movement data frame sequence from the collected eye movement data frames according to a sliding window of a preset length; A fourth determination module is used to determine, in the video, a target position mapped by the sight point coordinates of each eye movement data frame in the current eye movement data frame sequence; Add a module for adding a preset Gaussian template at each target position; A construction module is used to construct the user's sight track in the video based on the added Gaussian template.
[0073] Embodiment 3 This embodiment also provides a computer device, such as a smart phone, tablet computer, laptop computer, desktop computer, rack server, blade server, tower server or cabinet server (including an independent server or a server cluster composed of multiple servers) that can execute programs. Figure 6 As shown, the computer device 60 of this embodiment includes at least but not limited to: a memory 601 and a processor 602 that can communicate with each other via a system bus. It should be noted that Figure 6 Only computer device 60 is shown with components 601 - 602 , but it should be understood that implementing all of the components shown is not a requirement, and more or fewer components may alternatively be implemented.
[0074] In this embodiment, the memory 601 (i.e., readable storage medium) includes flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 601 can be an internal storage unit of the computer device 60, such as a hard disk or memory of the computer device 60. In other embodiments, the memory 601 can also be an external storage device of the computer device 60, such as a plug-in hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card, etc. equipped on the computer device 60. Of course, the memory 601 can also include both the internal storage unit of the computer device 60 and its external storage device. In this embodiment, the memory 601 is generally used to store the operating system and various application software installed on the computer device 60. In addition, the memory 601 can also be used to temporarily store various types of data that have been output or are to be output.
[0075] In some embodiments, the processor 602 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 602 is generally used to control the overall operation of the computer device 60 .
[0076] Specifically, in this embodiment, the processor 602 is used to execute the program of the emotion recognition method stored in the memory 601, and the program of the emotion recognition method implements the following steps when being executed: Collect eye movement data frames of users when they watch videos in real time; each eye movement data frame is a frame of eye movement data collected based on a preset eye movement collection algorithm; Based on the eye movement data frame, determining the target object that the user is paying attention to from the objects displayed in the video; Filtering out a target eye movement data frame sequence for focusing on the target object from the collected eye movement data frames; Extracting eye movement features from the target eye movement data frame sequence; The extracted eye movement features are input into a preset emotion classification model to obtain the emotion type of the user towards the target object.
[0077] The specific implementation process of the above method steps can be found in Example 1, and this embodiment will not be repeated here.
[0078] Embodiment 4 This embodiment also provides a computer-readable storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a disk, an optical disk, a server, an App application store, etc., on which a computer program is stored. When the computer program is executed by a processor, the following method steps are implemented: Collect eye movement data frames of users when they watch videos in real time; each eye movement data frame is a frame of eye movement data collected based on a preset eye movement collection algorithm; Based on the eye movement data frame, determining the target object that the user is paying attention to from the objects displayed in the video; Filtering out a target eye movement data frame sequence for focusing on the target object from the collected eye movement data frames; Extracting eye movement features from the target eye movement data frame sequence; The extracted eye movement features are input into a preset emotion classification model to obtain the emotion type of the user towards the target object.
[0079] The specific implementation process of the above method steps can be found in Example 1, and this embodiment will not be repeated here.
[0080] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0081] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0082] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method.
[0083] The above are only preferred embodiments of the present invention, and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. An emotion recognition method, characterized in that: The method comprises: Collect eye movement data frames of users when they watch videos in real time; each eye movement data frame is a frame of eye movement data collected based on a preset eye movement collection algorithm; Based on the eye movement data frame, determining the target object that the user is paying attention to from the objects displayed in the video; Filtering out a target eye movement data frame sequence for focusing on the target object from the collected eye movement data frames; Extracting eye movement features from the target eye movement data frame sequence; The extracted eye movement features are input into a preset emotion classification model to obtain the emotion type of the user towards the target object.
2. The method according to claim 1, characterized in that: The step of determining the target object that the user is paying attention to from the objects displayed in the video based on the eye movement data frame includes: detecting an object in each video frame of the video and generating a bounding box around the object; Extracting the coordinates of the user's sight point from the eye movement data frame; When the coordinates of the sight point of N consecutive eye movement data frames are all located in the same bounding box, and N is a positive integer greater than or equal to a preset frame number threshold, determining that the object in the bounding box is the target object; The step of selecting a target eye movement data frame sequence for focusing on the target object from the collected eye movement data frames comprises: The N consecutive eye movement data frames are determined as the target eye movement data frame sequence.
3. The method according to claim 1, characterized in that The extracting eye movement features from the target eye movement data frame sequence comprises: Determining the eye movement behavior type to which each eye movement data frame in the target eye movement data frame sequence belongs; When the target eye movement data frame sequence includes eye movement data frames belonging to the gaze behavior type, extracting features and pupil features for characterizing the gaze behavior from all eye movement data frames belonging to the gaze behavior type; When the target eye movement data frame sequence includes eye movement data frames belonging to the saccadic behavior type, the features for characterizing the saccadic behavior and the pupil features are extracted from all eye movement data frames belonging to the saccadic behavior type.
4. The method according to claim 3, characterized in that: Determining the eye movement behavior type to which each eye movement data frame in the target eye movement data frame sequence belongs includes: Calculating the eye movement speed between every two adjacent eye movement data frames in the target eye movement data frame sequence; When the eye movement speed is less than a preset speed threshold, determining the last eye movement data frame of two corresponding adjacent eye movement data frames as a gaze behavior type; When the eye movement speed is greater than a preset speed threshold, the last eye movement data frame of the corresponding two adjacent eye movement data frames is determined as a scanning behavior type.
5. The method according to claim 3, characterized in that: The features used to characterize gaze behavior include at least one of the following: features used to characterize gaze frequency and features used to characterize gaze time; the features used to characterize scan behavior include at least one of the following: features used to characterize scan frequency, features used to characterize scan time, features used to characterize scan amplitude and features used to characterize scan speed; the pupil features include: features used to characterize pupil size.
6. The method according to claim 1, characterized in that After inputting the extracted eye movement features into a preset emotion classification model to obtain the emotion type of the user towards the target object, the method further includes: The emotion type is displayed at the target object shown in the video; wherein the emotion type is used to characterize the degree of emotional awakening of the user by the target object.
7. The method according to claim 1, characterized in that The method further comprises: Determining a current eye movement data frame sequence from the collected eye movement data frames according to a sliding window of a preset length; Determining in the video a target position mapped by the sight point coordinates of each eye movement data frame in the current eye movement data frame sequence; Add a preset Gaussian template at each target location; Based on the added Gaussian template, the sight track of the user is constructed in the video.
8. An emotion recognition device, characterized in that: The device comprises: The acquisition module is used to acquire eye movement data frames of the user when watching the video in real time; wherein each eye movement data frame is a frame of eye movement data acquired based on a preset eye movement acquisition algorithm; A first determination module, configured to determine, based on the eye movement data frame, a target object that the user is paying attention to from objects displayed in the video; A screening module, used to screen out a target eye movement data frame sequence for focusing on the target object from the collected eye movement data frames; An extraction module, used for extracting eye movement features from the target eye movement data frame sequence; The second determination module is used to input the extracted eye movement features into a preset emotion classification model to obtain the emotion type of the user towards the target object.
9. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is used to implement the steps of the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it is used to implement the steps of the method according to any one of claims 1 to 7.