Short video emotion value detection and quantitative evaluation method and system
By collecting and analyzing users' heart rate, blood oxygen and facial data, and combining it with deep learning models, the emotional value of short videos is quantified, solving the problem of inaccurate emotional evaluation in existing technologies, realizing the identification and quantification of emotional changes, and improving user experience and mental health management.
Patent Information
- Application Number
- CN202510803770.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-07
- Publication Date
- 2025-09-26
AI Technical Summary
Existing short video platforms find it difficult to conduct detailed and effective quantitative evaluation of user emotional value, resulting in inaccurate quality evaluation methods, which affects user experience and industry development.
An STM32 microcontroller, a heart rate and blood oxygen sensor, and a digital webcam are used to collect the user's heart rate, blood oxygen, and facial data. A deep learning model is used to analyze facial expressions and blinking frequency. Emotional changes are quantified by calculating the Euclidean distance through a multi-scale sliding window, and the Russell Circle Emotion Classification Theory is used for evaluation.
It realizes the quantitative evaluation of the emotional value of short videos, can identify and quantify emotional changes, provide more in-depth emotional analysis, support personalized recommendations, and improve user experience and mental health management.
Smart Images

Figure CN120711232A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of short video evaluation, and in particular to a system that collects facial expressions and physiological data of people when watching videos through a combination of hardware and software, detects and quantifies the emotional value and concentration of short videos watched by users, and provides an objective basis for short video evaluation systems and collaborative filtering recommendation systems. Background Art
[0002] Advances in mobile internet have boosted the economic development of the short video industry, and short videos have become increasingly integrated into people's lives. Take Douyin, for example, which boasts 880 million daily active users and 30 million short videos uploaded daily. Short videos have become a new source of information and a source of emotional comfort. However, the quality of these short videos varies greatly, and the user experience often differs widely. Therefore, it is necessary to establish a system for measuring the emotional value of short videos and conducting objective quantitative evaluations. This system, which can quantify the sense of immersion to assess short video quality, and establish a targeted, customized, and precise recommendation system, will benefit people's physical and mental health and guide the continuous advancement of the short video industry.
[0003] To quantify the emotional value provided by short videos, we can collect facial expressions and obtain user physiological data, calculate the emotional value quantitative score, and improve the overall usability and classification dimension of the method for distinguishing video quality. Using the emotional value quantification method to extract value information can achieve the distinction and evaluation of the emotional color and emotional value of the video.
[0004] Human expressions can be categorized into various types, often expressing different emotions through different combinations of facial muscles. Psychologist Paul Ekman et al. proposed a theory of six basic emotions, which are considered universal and recognizable through facial expressions. These basic emotions include happiness, sadness, surprise, fear, anger, and disgust. Neutral, which indicates a lack of emotional inclination, is also considered neutral. Beyond these basic emotions, there are more nuanced or complex emotional states, such as embarrassment, curiosity, and contempt. In expression recognition technology, the results are often presented quantitatively. Quantitative methods can, on the one hand, classify recognized expressions into one of the predefined emotion categories using classification labels. On the other hand, confidence scores can be used to assign a numerical value to each possible emotion category, representing the probability or confidence level that the expression belongs to a specific emotion.
[0005] For continuous dimension scoring, expressions are sometimes assessed along several continuous emotional dimensions (such as pleasure and arousal), which allows for a more detailed description of emotional states. Russell (1980) proposed a circumplex model for emotion classification, arguing that emotions can be divided into two dimensions: pleasure and intensity. The Circumplex Model of Affect uses two main dimensions to describe emotions. Pleasure (Valence) is used to represent the change from very unpleasant to very pleasant. This is a dimension that measures the degree of positivity or negativity of an emotion. Arousal (Arousal) is used to represent the change from very calm to very excited. This dimension measures the activity or intensity of an emotion. Based on these two dimensions, a two-dimensional coordinate system can be established to locate different emotional states. For example, high pleasure and high activation levels may indicate positive and intense emotions such as excitement and happiness; low pleasure and high activation levels may correspond to negative but intense emotions such as fear and anger; low pleasure and low activation levels may represent relatively negative and calm emotions such as sadness and depression; and high pleasure and low activation levels may indicate positive and calm states such as satisfaction and relaxation. The continuous dimension scoring method for facial expressions is a way to convert facial expressions into quantifiable numerical values, which is usually based on the concept of emotional space. Quantifying the results of facial expression recognition facilitates further data analysis, statistical processing, or serves as feature input in machine learning algorithms. In practical applications, how to quantify facial expressions depends on the specific application scenario and technical implementation details.
[0006] Heart rate (HR) and blood oxygen saturation (SpO2) are commonly used as physiological indicators to assess physical condition, but they can also reflect an individual's emotional state. Emotional reactions are often accompanied by changes in the autonomic nervous system, which can affect heart activity and breathing patterns, and in turn, alter heart rate and blood oxygen levels. Therefore, measuring heart rate and blood oxygen data can be used to quantify the emotional value of short videos. When people are nervous or anxious, the sympathetic nervous system is activated, causing the heart rate to increase and blood pressure to rise. Therefore, in such emotional states, heart rate is typically elevated. Activation of the parasympathetic nervous system typically occurs during periods of relaxation and calm, when heart rate tends to decrease. Positive emotions such as excitement and happiness can also cause heart rate to increase, as these emotions also activate the sympathetic nervous system. In some cases, sadness or depression can cause heart rate to slow, but this is not a universal phenomenon and varies greatly from person to person. In situations of acute stress or fear, a brief increase in respiratory rate may occur, which helps to increase oxygen levels in the blood and thus temporarily improve blood oxygen saturation. During periods of extreme stress or panic attacks, people may hyperventilate, which can lead to excessive excretion of carbon dioxide, causing respiratory alkalosis and possibly a slight decrease in blood oxygen saturation. Chronic stress or negative emotions can affect lung function and sometimes indirectly affect blood oxygen saturation.
[0007] When people are nervous or anxious, they may experience increased involuntary blinking. When people are highly focused on a task, their blinking rate may decrease; conversely, when they are relaxed, their blinking rate may return to normal. By analyzing blinking rate data over time, we can understand people's concentration levels while watching videos. Therefore, blinking rate data collected from users watching videos can be used as an effective indicator for evaluating the quality of short videos.
[0008] Currently, there is limited research on the impact of short videos on user emotions and their ability to assist with psychological diagnosis and treatment. Short video platforms typically evaluate video quality based on text content, clarity, and number of views, failing to conduct targeted, detailed, and effective evaluation and differentiation of video content. This is detrimental to the long-term healthy development of the short video industry. Therefore, detecting and quantifying the emotional value of short videos is a pressing technical challenge. Summary of the Invention
[0009] The technical problem to be solved by the present invention is to address the deficiencies of the above-mentioned existing technologies and provide a short video emotional value detection, quantitative evaluation system and recommendation method, so as to solve the problem that the emotional value brought by existing short videos is difficult to quantify, and quantify the immersive feeling of watching short videos. As a method of short video evaluation, quantifying the emotional value provided by short videos can help users to guide their emotions.
[0010] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0011] The present invention provides a detection and quantitative evaluation system for the emotional value of short videos, comprising an STM32 single-chip microcomputer, a heart rate and blood oxygen sensor, a digital network camera, and a computer processor;
[0012] The hardware part uses a single-chip microcomputer and a heart rate and blood oxygen sensor to collect the heart rate and blood oxygen data of the user when watching a video, and then sends the collected raw data to the computer processor through the serial port for data preprocessing.
[0013] The software unit includes a video processing unit, a facial information recognition unit, and a data processing and analysis unit.
[0014] The video processing unit uses OpenCV image processing technology to process each frame of the video stream read by the camera, adjust the image size, convert the RGB image into a grayscale image, and perform Gaussian blur processing on the binary image through Gaussian filtering. This can remove noise in the image and improve the image clarity. The pre-processed video data is input into the deep learning model in the computer processor;
[0015] The facial information recognition unit uses a deep convolutional neural network (VGG16) to extract features from facial images. The Keras machine learning library and TensorFlow framework are used for model training and prediction. The feature vectors are input into a fully connected layer and a softmax classifier, which maps them to six different emotion categories: angry, afraid, happy, calm, sad, and surprised. For gender identification of users watching videos, a pre-trained deep convolutional neural network (CNN) model is used to distinguish between different genders. Physical features such as facial contours, skin texture, hairstyle, beard, and Adam's apple are often directly correlated with gender. With the support of extensive training data, the CNN model is able to map frequently occurring features to specific gender categories. Blink frequency collection uses a face detector, keypoint location tools, and eye aspect ratio algorithm based on the dlib machine learning library. By analyzing faces in a camera or video stream, the eye aspect ratio (EAR) is calculated in real time. The data is then fitted to obtain the user's blink frequency over a time series.
[0016] The data processing and analysis unit preprocesses and analyzes the collected physiological data, including heart rate, blood oxygen level, facial expression, gender, and blink rate, removing noise and normalizing the data to ensure accurate and reliable analysis. After preprocessing, heart rate and blood oxygen level data undergo a Gaussian transform, as their impact on emotion is not linear. A multi-scale sliding window approach is used to analyze the data. A fixed-size "window" is defined and slid across the time series. The feature information of heart rate, blood oxygen level, facial expression, gender, and blink rate within the window is calculated, thereby extracting the statistical features of the time series within that window. Because different modal physiological information has different emotional relevance, different weights are assigned to heart rate, blood oxygen level, facial expression, gender, and blink rate within the window to obtain the feature vector within that window. The Euclidean distance between feature vectors in two adjacent windows is then calculated to quantify the emotional value and engagement of the short video to the user. This approach captures trends and fluctuations within local time periods at different time scales, understanding the dynamic characteristics of the data.
[0017] Based on this system, a short video recommendation method is proposed. The detected and quantified emotional value of short videos is uploaded to a short video server. The server, which holds a vast amount of short video data with pre-existing tags, uses the emotional value quantification detection system to detect emotional value. The server then recalls (roughly screens) and sorts the labeled short video data, filtering out short videos that elicit negative emotions in users and recommending positive short videos to guide their emotions. Based on the user's emotional value, the server optimizes the platform's short video recommendation scheme, personalizedly pushing short videos that enhance the user's emotional value, improving the user's viewing experience and guiding their emotional value. The short video server, upon detecting the user's emotional state, pushes funny videos when the user is sad, emotional videos when the user is angry, and inspirational videos when the user is depressed, alleviating the impact of negative emotions and thus promoting the development of the short video platform. Based on the user's emotional value and physiological information, the server pushes targeted short videos that enhance their emotional value, effectively regulating the user's mental health. This method not only achieves intelligent recommendation but also maintains user mental health.
[0018] The implementation of the system specifically includes the following steps:
[0019] Step 1: Use an STM32 microcontroller and a heart rate and blood oxygen sensor to collect heart rate and blood oxygen data for a fragment (taken as 90 seconds) at a certain time interval (Gap) (taken as Gap of 10 minutes). Use a median filter algorithm to filter out noise from the raw data collected by the sensor, and then package the heart rate and blood oxygen sensor and microcontroller into a wearable device.
[0020] Step 2: Use the Haar cascade classifier in the OpenCV library to perform face detection on the input image to obtain the key areas of the face. Use the machine learning library to load a pre-trained deep learning model. The Haar cascade classifier extracts features from the face image and builds a deep learning network model for facial expression classification. The vectors are mapped to seven different emotion categories: happiness, calmness, sadness, surprise, fear, anger, and disgust. This identifies the facial expression information that changes over time while the user watches the video. People of different genders often have directly related different features. The deep convolutional neural network (CNN) model captures features that often appear together and maps them to a specific category. The CNN pre-trained model is used to identify the gender of the video viewer.
[0021] Step 3: Use the dlib machine learning library to obtain the coordinates of the key points of the human eye, calculate the Euclidean distance between the horizontal and vertical key points, and calculate the eye aspect ratio (EAR). During blinking, the eye aspect ratio changes, and each sudden change is identified as a blink. The time length of a cycle is captured at regular intervals, and the blink frequency within each cycle is calculated.
[0022] Step 4: Low-pass filter the heart rate, blood oxygen, facial expression, gender, and blink frequency data collected in steps 1, 2, and 3 to remove some data that deviates greatly from the true value, fit the data to obtain continuous time series data, and then perform Gaussian transformation on the heart rate, blood oxygen, and blink frequency data; set a window with a fixed length of win, and use a sliding window mechanism to slide gradually along the time series; according to the length of the data collected in the time series is N, the moving step of the sliding window is 1, and the sliding window calculates the 1st to win data of the time series, and the 2nd to win+1 data, in sequence, until N data are read; in the time series data windows of heart rate, blood oxygen, and blink frequency, calculate the statistical feature mean, maximum and minimum value, standard deviation and variance, skewness and kurtosis, and autocorrelation coefficient of each window to capture the characteristics of each window;
[0023] Step 5: Compare the Gaussian transformed waveforms of the heart rate, blood oxygen, and blink frequency data in different windows to analyze changes in the user's emotional state. Before calculating the eigenvector of each window, assign different weights wi (representing the weight of the i-th feature) to heart rate, blood oxygen, facial expression, gender, and blink frequency according to the different contributions of different feature dimensions to the measurement of emotional state and its emotional changes. In the time dimension, combined with the detection of the user's facial expression when watching the video, the eigenvalue of each window is calculated and constructed into a multidimensional feature vector within the window. Based on the emotion detection scores in different time windows, and then according to the classification model of Russell's circular emotion classification theory, the direction of the emotion change vector is obtained, thereby obtaining the positive and negative emotional value of the short video.
[0024] Step 6: Quantify the emotional changes caused by the short video to the user by calculating the Euclidean distance between the feature vectors of two adjacent windows;
[0025] The feature vector of the i-th window is: v i =(v i,1 , v i,2 ,…,v i,5 );
[0026] The eigenvector of the i+1th window is: v i+1 =(v i+1,1 , v i+1,2 ,…,v i+I,5 );
[0027] Among them, v in It represents the value of the nth feature in the i-th window (including the normalized values of heart rate, blood oxygen, and blink frequency), reflecting the statistical characteristics of a physiological data in the window. The Euclidean distance between two adjacent windows can be calculated using the following formula:
[0028]
[0029] By accumulating the multi-dimensional feature vector distances in adjacent windows, the emotional value of short videos can be quantitatively evaluated to obtain the emotional value Q brought to users over a period of time:
[0030]
[0031] Compared with the prior art, the advantages of the present invention are:
[0032] The system utilizes an STM32 microcontroller, a standard digital camera, and a heart rate and blood oxygen sensor. These sensors and the microcontroller are packaged into a wearable device, enabling real-time collection of a user's heart rate and blood oxygen data, as well as facial video data. Machine learning libraries such as OpenCV and dlib, along with deep learning models, are used for face detection, expression classification, and blink rate calculation. This combination of hardware and software enables the system to acquire user physiological data at the physical level, enabling multimodal data collection and analysis, providing a basis for subsequent emotional value analysis.
[0033] Most existing technologies focus on the classification and recognition of emotions, such as identifying a user's emotional state (such as happiness, sadness, etc.) through facial expressions. However, these technologies often remain at the level of emotion classification, using a single modality (such as only visual or audio features) for emotion recognition and lacking quantitative analysis of emotional changes. This system quantitatively evaluates the user's emotional fluctuations while watching a video by calculating the Euclidean distance between the feature vectors of adjacent time windows. This method can not only recognize facial expressions, but also quantify the amplitude and frequency of emotional changes, thereby providing a more in-depth and comprehensive analysis of emotional value. In other words, through a comprehensive evaluation system based on detection indicators such as vision, heart rate, and blood oxygen, if it is determined that the user's emotions fluctuate greatly when watching a certain video clip, then it can be inferred that the emotional value provided by the video clip is relatively large. Otherwise, the emotional value is relatively small.
[0034] The system collects physiological data such as heart rate and blood oxygen levels, and also acquires user behavioral data through facial expression recognition and blink rate calculation. This multimodal data collection enables the system to analyze the user's emotional state from multiple dimensions, which is more objective than single-modal recognition and analysis. Data such as heart rate, blood oxygen levels, facial expressions, and blink rate are transformed, standardized, and feature extracted, and the data from different modalities is fused together to form a multidimensional feature vector. This feature fusion method enables the system to comprehensively consider multiple data sources to accurately reflect the user's emotional state.
[0035] Beneficial effects
[0036] This system can be applied to mental health assessments, such as depression screening. Doctors can use the emotional values captured in patients to assist in diagnosing the progression of their condition. This objective, quantifiable monitoring of emotional values can be used to measure the emotional values of individual users over a period of time. If the emotional values are low, more positive, energetic, or entertaining short videos can be pushed. This method of emotion monitoring and regulation not only supports users' mental health in daily life but also provides effective support for psychologists in their patient care. Through quantitative data analysis, doctors can better understand patients' emotional states and develop personalized treatment plans. This not only enhances the scientific nature and effectiveness of psychotherapy but also provides a new approach to emotional management.
[0037] This system can be applied to short video platforms. By analyzing multimodal data, it can derive the emotional value of users watching different short videos, thereby facilitating the evaluation of video content quality. This data-driven analysis not only reveals user preferences for various short videos but also deeply captures the interests of users of different genders in video content. These insights will provide strong support for optimizing short video recommendation algorithms. By analyzing user physiological data, the platform can monitor the effectiveness of the recommendation algorithm in real time and dynamically adjust it based on feedback. This helps improve the platform's retention rate and user experience, allowing users to enjoy content while experiencing more personalized and thoughtful recommendation services. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments.
[0039] Figure 1 This is a schematic diagram of the hardware architecture of the short video emotion value detection system according to an embodiment of the present invention;
[0040] Figure 2 Schematic diagram of the architecture of a multi-source sensor data processing system according to an embodiment of the present invention;
[0041] Figure 3 This is a schematic diagram of the system functional modules of an embodiment of the present invention;
[0042] Figure 4 This is a schematic diagram of the short video recommendation system according to an embodiment of the present invention;
[0043] Figure 5 Schematic diagram of the six key points of the human eye;
[0044] Figure 6 Schematic diagram of the process of detecting the emotional value of short videos according to an embodiment of the present invention;
[0045] Figure 7 Schematic diagram of the circular model of Russell's emotion classification. DETAILED DESCRIPTION
[0046] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0047] The emotional value and concentration detection and quantification method provided in the embodiment of the present application can be applied to Figure 1In the application environment shown, a wearable device can collect physiological data such as heart rate and blood oxygen levels, and communicate with a data analysis device via a network. The terminal can capture facial images of a person watching a video and, using an internally deployed deep learning model, identify facial expressions, gender, and calculate blink frequency. It then establishes communication with the data processing and analysis device over the network and uploads discrete data showing the changes in the recognized facial expressions, gender, and blink frequency over time. The data processing and analysis unit processes the uploaded data, calculating feature vectors using a multi-scale sliding window approach. It then scores the emotional value of short videos by calculating the Euclidean distance between feature vectors in adjacent windows. The analyzed data is then sent to a server, which collects and adjusts existing labeled short video data in real time based on the emotional value collected and analyzed from the user. The server then updates the model to analyze the user's preferences and needs, sorts the existing labeled short video data, recalls the labeled videos required, and continuously adjusts them in real time. The terminal can be, but is not limited to, various personal computers, laptops, smartphones, and tablets. The data collection device can be a wearable health device such as a smart bracelet or smartwatch. The data processing and analysis device can be a computer, server, or other device.
[0048] The system function modules of the present invention are as follows Figure 3 As shown, an STM32 microcontroller controls a heart rate and blood oxygen saturation sensor to acquire a user's heart rate and blood oxygen saturation data. The sensor illuminates the skin with red and infrared light. Utilizing the different absorption rates of these two light waves by redox proteins and oxygenated proteins in the blood, it creates a light source that alternates between infrared and red light on the skin's surface. This light source detects changes in light transmittance as blood vessels pulsate, thereby measuring heart rate and blood oxygen saturation. A low-pass filter is used to ensure data accuracy. The collected data is then processed to remove data that deviates significantly from the normal value. After fitting the data, stable heart rate and blood oxygen saturation data is obtained. Heart rate and blood oxygen data are closely linked to emotional levels. Emotional changes can cause certain variations in heart rate and blood oxygen data. This variation in heart rate and blood oxygen data can be used as an important indicator to quantify a user's emotional value.
[0049] Digital network cameras capture images in real time and perform operations such as grayscale conversion, Gaussian filtering, and smoothing. OpenCV's Haar cascade classifier, a face detection method, is used. The Haar cascade classifier can detect faces in multiple regions and at multiple scales across an entire image. Multi-region detection divides the image into multiple blocks, and detection is performed on each block (detection window). During face detection, if the same face is detected multiple times, regions need to be merged. This process uses non-maximum suppression (NMS) to prevent multiple detections of the same face region. After the camera captures a facial image, face detection is performed on the captured image using Haar features and an AdaBoost cascade classifier. OpenCV face detection calculates the Haar feature values of the detection area. These values are then analyzed by the AdaBoost algorithm to determine whether a face is present. The AdaBoost algorithm combines multiple weak classifiers to form a strong classifier. To improve detection performance, OpenCV uses a cascade classifier, connecting multiple strong AdaBoost-based classifiers in series. During the detection process, an area is identified as a face area only when all classifiers determine that a face exists.
[0050] The Haar cascade classifier from the OpenCV library is used as a feature extractor to extract facial features from the input image to identify the facial region. The expression recognition system uses Keras to load pretrained deep learning models for face detection, emotion, and gender classification. A mini-Xception network model is constructed based on the Xception network for expression classification. This model simplifies the network hierarchy, removes fully connected layers, and replaces traditional convolutional layers with depthwise separable convolutions. The model consists of four residual depthwise separable convolutional blocks, each followed by batch normalization and ReLU activation. Finally, a global average pooling layer and a softmax activation function are used for prediction. The model outputs a confidence score for each emotion category. The program selects the category with the highest score as the current emotion recognition result, classifying it into seven different emotion categories: anger, disgust, fear, happiness, calm, sadness, and surprise. The mini-Xception network model simplifies the Xception network architecture, removing fully connected layers and adopting depthwise separable convolutional layers. This change significantly reduces training parameters and time, enhancing model generalization.
[0051] The present invention uses a deep convolutional neural network (CNN) to achieve gender differentiation. CNN is a deep learning model specifically used to process image data, with its core being the convolution layer and the pooling layer. In face recognition, CNN can automatically learn and extract features of facial images, such as edges, textures, and shapes, to form high-dimensional feature vectors. Deep CNN can capture features at different levels, from local details to global patterns. People of different genders have different preferences for short video types, so a CNN-based face recognition model is constructed to identify a person's gender, providing a reference for the short video recommendation system.
[0052] The coordinates of the six key points of the human eye obtained using the dlib machine learning library are p1, p2, p3, p4, p5, and p6, as shown in Figure 5 , from which we can calculate the horizontal key point Euclidean distance as ||p1-p4||, the vertical key point Euclidean distance as ||p2-p6|| and ||p3-p5||, and calculate the eye aspect ratio EAR;
[0053]
[0054] Since blinking is a short-lived process, we need to analyze data from multiple frames. Setting the number of consecutive frames to 3, that is, if the EAR values calculated for three consecutive frames are all less than the EAR threshold of 30%, this represents a blink. Every 10 seconds is counted as a period T (in seconds). The number of blinks within T is recorded as blink_time, and the user's blink frequency (in times / minute) is calculated as:
[0055]
[0056] The blinking frequency of video users reflects the attractiveness of the video, which can be used to quantify the immersiveness of the video and then evaluate and score the quality of the video. It can intuitively understand the emotional state of the user when watching the video, which is the most intuitive feedback on the quality of the video. Knowing the emotional state of the current user watching the video can better guide the user's emotions. In order to quantify the blinking frequency into the immersive range [0%, 100%], a maximum blinking frequency max_frequency is set, and the average blinking frequency average_eye_frequency of each video is calculated. The immersive index immersion is obtained by calculating the average blinking frequency in a video, and then normalized. The formula is as follows:
[0057]
[0058] After collecting a large amount of data, the collected data will be cleaned and processed, and the filling method based on statistical variables will be used to fill the missing values through the mode to make the fitted data smoother. When a data point in the data deviates significantly from the distribution of other data points or a data point is significantly different from other data points, it will be judged as an outlier (abnormal value). The abnormal data detection method can be used to detect the outlier and remove it.
[0059] After obtaining the heart rate and blood oxygen data in time series, the sliding window method is used to analyze the data. A fixed window length is used to slide the window gradually along the time series. Before extracting statistical features, the heart rate n1, blood oxygen n2, and blink frequency data n5 are Gaussian transformed to obtain f1 = G1(n1), f2 = G2(n2), and f5 = G5(n5). The general form of the Gaussian function is:
[0060]
[0061] For the characteristics of different detection indicators n, the Gaussian function transformation parameters are determined by more than 200 sets of collected data. According to the degree of deviation between the transformed Gaussian function and the normal heart rate and blood oxygen, the feature vector f(n) is constructed. The detection data of the five dimensions of heart rate, blood oxygen, expression, gender, and blinking frequency have different contributions to the description of human emotional changes. Usually, facial expressions are the most intuitive reflection of emotions. Therefore, before constructing the feature vector describing emotions, the expression feature is given the largest weight coefficient. According to the circular model of Russell's emotion classification, the classification score of the expression is used as an indicator of the emotional value (pleasure and arousal) of the expression. For example, if 7 categories are selected, the classification scores are {-3, -2, -1, 0, 1, 2, 3}. Different weights w are assigned to the five features of heart rate, blood oxygen, expression, gender, and blink frequency. m , w m Indicates the weight of the mth feature, m = 1, 2, ..., 5. In a certain sliding window, the multimodal physiological signal data is transformed and weighted to form a feature vector v = {w m ·f m}, calculate the eigenvalue in the window and use F_sim to represent it:
[0062]
[0063] Regarding the impact of focus on emotional activity, many studies suggest that people's emotional responses to external stimuli differ significantly when they are focused and when they are not. The same stimulus often produces different emotional responses in people of different genders. Because the effects of gender and blinking frequency on emotional state values are often nonlinear, for further in-depth analysis of emotional states, gender and blinking frequency detection indicators are set as optional features. For example, w4 = w5 = 0, the gender function is set to gender(), and the focus function associated with blinking frequency f5 is set to gaze(). The detailed emotional statistical features of the time series within this window are extracted as follows:
[0064] F=gaze(f5,gender(f4,F_sim))
[0065] Emotional dimensionality theory posits that different emotions evolve gradually and smoothly. The similarities or differences between different emotions can be expressed based on their distances in dimensional space. Therefore, the distances between different feature dimensions in feature space over time can describe the magnitude of emotional change. Sliding window features can capture trends and volatility within a local time period and are very effective for modeling short-term trends and processing high-frequency data. Given a time series data length of N, a window size of win, and a sliding window step size of 1, the sliding window is calculated sequentially for the first to win data points in the time series, then for the second to win+1 data points, and so on, until the last window. Within each sliding window, various statistical features are calculated to capture local information. For the heart rate, blood oxygen, and blink rate time series data windows, the mean, maximum and minimum values, variance, skewness and kurtosis, and sliding cumulative sum are calculated for each window to capture the characteristics of each window.
[0066] In multi-scale sliding window analysis, standard deviation and variance are statistical measures of data volatility for physiological data such as heart rate and blood oxygen levels, effectively reflecting the magnitude and volatility of data changes. In sliding window analysis, data is segmented into multiple windows based on the time series. Window sizes of varying sizes can reveal volatility at different time scales. For example, smaller windows (e.g., seconds to minutes) can capture short-term fluctuations and are suitable for rapidly responding physiological changes; whereas larger windows (e.g., tens of minutes to hours) can reflect long-term trends and help identify chronic physiological instability. Comprehensive analysis of volatility at these different scales allows for a more comprehensive assessment of a user's health. The data within each window is used to calculate the variance and standard deviation. This approach helps us assess the volatility of data within the window and, in turn, understand the patterns of change in physiological signals such as heart rate and blood oxygen levels over that time period. For example, a stable heart rate typically exhibits minimal fluctuations, while an unstable heart rate may exhibit significant fluctuations. A small standard deviation within a window indicates low data volatility, indicating a relatively stable heart rate or blood oxygen level. Conversely, a large standard deviation indicates significant signal fluctuations, potentially reflecting changes in physical condition or emotional fluctuations. Variance is the degree to which the data deviates from its mean. It measures the dispersion of the data points relative to the mean. The standard deviation is the square root of the variance. It also measures the dispersion of the data. The calculation formula is:
[0067]
[0068] Kurtosis and skewness are used to measure the shape of heart rate, blood oxygen, and blink rate data. Kurtosis is a characteristic number that characterizes the height of the peak at the mean of a probability density distribution curve. Intuitively, kurtosis reflects the sharpness of the peak. A kurtosis of 0 indicates that the population data distribution is as steep as a normal distribution; a kurtosis of >0 indicates that the population data distribution is steeper than a normal distribution, with a peaked peak; a kurtosis of <0 indicates that the population data distribution is flatter than a normal distribution, with a flat-topped peak. Kurtosis includes normal distributions (kurtosis value = 3), thick tails (kurtosis value > 3), and thin tails (kurtosis value < 3).
[0069] Skewness describes the symmetry of the distribution of values in a population. Skewness = 0 means that the data distribution has the same degree of skewness as the normal distribution; skewness > 0 means that the data distribution is positively skewed or right-skewed compared to the normal distribution, with more extreme values on the right end of the data; skewness < 0 means that the data distribution is negatively skewed or left-skewed compared to the normal distribution, with more extreme values on the left end of the data. The definition of skewness includes normal distribution (skewness = 0), right-skewed distribution (also called positively skewed distribution, with skewness > 0), and left-skewed distribution (also called negatively skewed distribution, with skewness < 0). For heart rate and blink rate data, we prefer a sharp peak, and for blood oxygen data, we prefer a right skew. The calculation method for the kurtosis and skewness of a random variable is:
[0070]
[0071] Before capturing each window's features, heart rate, blood oxygen, and blink rate data are normalized using a Gaussian transform. The mean and variance of each window's heart rate, blood oxygen, and blink rate data are calculated, and these data are converted to zero mean and unit standard deviation to prevent bias due to different scales or ranges in subsequent analysis. In the temporal dimension, combined with the user's facial expressions, the eigenvalues of each window are calculated. These features are then used as feature vectors within the window, with each feature representing a dimension, including {heart rate, blood oxygen, facial expression, gender, blink rate}. Based on the emotion detection scores within different time windows and the classification model based on Russell's Circular Emotion Classification Theory, the direction of the emotion change vector is determined, thereby determining the positive and negative emotional value of the short video.
[0072] By calculating the Euclidean distance between two adjacent window vectors, a large Euclidean distance between feature vectors indicates rapid emotional fluctuations, indicating a high absolute value of the emotional value of a short video clip. A small Euclidean distance between adjacent time segments indicates a relatively stable emotional state. By accumulating the Euclidean distance within each window, we can quantify the emotional value people experience while watching a video.
[0073] The feature vector v of the i-th window i =(v i1 , v i2 ,…,v i5 )
[0074] Feature vector of the i+1th window: v i+1 =(v i+1,1 ,v i+1,2 ,…,v i+1,5 )
[0075] Among them, v inrepresents the value of the nth feature in the i-th window (e.g., normalized values of heart rate, blood oxygen, and blink frequency). Each dimension of these feature vectors represents the statistical characteristics of a physiological data in the window. Then, the Euclidean distance between two adjacent windows is:
[0076]
[0077] According to the sliding window mechanism analysis, the Euclidean distance between every two adjacent window feature vectors is accumulated to quantitatively evaluate the emotional value of short videos and obtain the emotional value Q brought to users over a period of time:
[0078]
[0079] The resulting short video emotional value and gender characteristic data (f4, Q) is uploaded to the short video server, where the emotional value evaluation data for the short videos is added and updated. The server, which holds a vast amount of short video data with various evaluation tags, uses (f4, Q) as the emotional value tag for the short video data. The server can recall, sort, and filter the short video data based on the existing tags. Based on user needs, it can filter out short videos that evoke negative emotions in users. Based on user evaluations of the emotional value of short videos, the server optimizes the platform's short video recommendation scheme, delivering personalized short videos that capture emotional appeal.
Claims
1. A method for detecting and quantitatively evaluating the emotional value of short videos, characterized in that: The following steps are involved: Step 1: At a certain time interval (Gap), collect the audience's heart rate and blood oxygen data within a fragment, and use the median filter algorithm to filter out noise from the raw data collected by the sensor; Step 2: Use the Haar cascade classifier in the OpenCV library to perform face detection on the collected images to obtain the key areas of the face. Use the machine learning library to load a pre-trained deep learning model. The Haar cascade classifier extracts features from the facial images and builds a deep learning network model for facial expression classification. The vectors are mapped to seven different emotion categories: happiness, calmness, sadness, surprise, fear, anger, and disgust. This identifies the facial expression information that changes over time while the user is watching the video. To identify the gender of the video viewer, a pre-trained deep convolutional neural network (CNN) model is used to achieve gender differentiation. Step 3: Use the dlib machine learning library to obtain the coordinates of the key points of the human eye, calculate the Euclidean distance between the horizontal and vertical key points, and calculate the eye aspect ratio (EAR). During blinking, the eye aspect ratio changes, and each sudden change is identified as a blink. The time length of a cycle is captured at regular intervals, and the blink frequency within each cycle is calculated. Step 4: Low-pass filter the heart rate, blood oxygen, facial expression, gender, and blink frequency data collected in steps 1, 2, and 3 to remove data that deviates significantly from the true value. Fit the data to obtain continuous time series data. Then, perform Gaussian transformation on the heart rate n1, blood oxygen n2, and blink frequency data n5 to obtain f1 = G1(n1), f2 = G2(n2), and f5 = G5(n5). Set the fixed window length to win, and use a sliding window mechanism to slide gradually along the time series to read data. Based on the length of the data collected in the time series as N, the moving step of the sliding window is 1. The sliding window sequentially calculates the 1st to win data in the time series, and the 2nd to win+1 data, until N data are read. In the time series data windows of heart rate, blood oxygen, and blink frequency, the statistical feature mean, maximum and minimum values, standard deviation and variance, skewness and kurtosis, and autocorrelation coefficient of each window are calculated to capture the characteristics of each window. Step 5: Compare the Gaussian transformed waveforms of heart rate, blood oxygen, and blink rate in different windows to analyze changes in the viewer's emotional state. Before calculating the eigenvector for each window, assign different weights to heart rate, blood oxygen, facial expression, gender, and blink rate based on the different contributions of different feature dimensions to the measurement of emotional state and its emotional changes. In the time dimension, combined with the detection of the viewer's facial expression while watching the video, the eigenvalue of each window is calculated and constructed into a multidimensional feature vector within the window. Based on the emotion detection scores in different time windows, and then according to the classification model of Russell's circular emotion classification theory, the direction of the emotion change vector is obtained, thereby obtaining the positive and negative emotional value of the short video; Step 6: Quantify the emotional changes caused by the short video to the user by calculating the Euclidean distance between the feature vectors of two adjacent windows; The feature vector of the i-th window is: v i =(v i,1 , v i,2 ,…,v i,5 ); The eigenvector of the i+1th window is: v i+1 =(v i+1,1 , v i+1,2 ,…,v i+1,5 ); Among them, v in It represents the value of the nth feature in the i-th window (including the normalized values of heart rate, blood oxygen, and blink frequency), reflecting the statistical characteristics of a physiological data in the window. The Euclidean distance of the multidimensional feature vector between two adjacent windows according to time series analysis is: By accumulating the distances between adjacent windows within a period of time, we can quantify the emotional value Q that the short video brings to the user during that period: Among them, N is the total data length, win is the window length, i is the window number, and n is the feature data dimension.
2. A short video emotional value detection and quantitative evaluation method according to claim 1, characterized in that The multidimensional feature vector constructed in step 5 is used to assign different weight coefficients w to the five features of heart rate, blood oxygen, expression, gender, and blinking frequency. m , w m Represents the weight of the mth feature, m = 1, 2, ..., 5. Facial expression is an intuitive reflection of emotion. Therefore, before constructing the feature vector describing emotion, the expression feature is given the largest weight coefficient. In a single sliding window, the multimodal physiological signal data is transformed and weighted to form a feature vector v = {w m ·f m }, calculate the eigenvalue F_sim in the window as: For further in-depth analysis of emotional states, gender and blink frequency detection indicators are set as optional features, that is, w4 = w5 = 0, and the gender function and concentration function are set. The detailed emotional statistical features F of the time series within this window are extracted as follows: F=gaze(f5,gender(f4,F_sim)) Among them, gender() is the gender function in the feature dimension, and gaze() is the concentration function associated with the blinking frequency.
3. A short video emotional value detection and quantitative evaluation method according to claim 1, characterized in that In step 3, the coordinates of the six key points of the human eye are obtained using the dlib machine learning library, namely p1, p2, p3, p4, p5, and p6. From this, the Euclidean distance of the horizontal key points can be calculated as ||p1-p4||, the Euclidean distance of the vertical key points can be calculated as ||p2-p6|| and ||p3-p5||, and the eye aspect ratio EAR (Eye Aspect Ratio) can be calculated; Since blinking is a short-lived process, we need to analyze the data of multiple frames. We set the number of consecutive frames to be counted to 3, that is, if the EAR values calculated for three consecutive frames are all less than the EAR threshold of 30%, it means one blink. Every 10 seconds is recorded as a period T (in seconds). The number of blinks within T is recorded as blink_time, and the blink frequency (in times / minute) of the user is calculated as: The blinking frequency of video users reflects the attractiveness of the video, which can be used to quantify the immersiveness of the video and then evaluate and score the quality of the video. It can provide the most direct feedback on the quality of the video and better guide the audience's emotions. The blinking frequency is quantified into an immersive range of [0%, 100%], a maximum blinking frequency max_frequency is set, and the average blinking frequency average_eye_frequency of each video is calculated. By calculating the average blinking frequency within a video and then normalizing it, the immersive index is obtained.
4. A short video recommendation method based on the short video emotional value detection and quantitative evaluation method of claim 1, characterized in that: The obtained short video emotional value and gender characteristic data (f4, Q) are uploaded to the short video server, and the short video emotional value evaluation data is added and updated. The server side has a large amount of short video data with multiple evaluation labels, and (f4, Q) is used as the emotional value label of the short video data; the server side can recall, sort and filter the short video data according to the existing labels, and according to user needs, it can filter out short videos that bring negative emotions to users; the server side optimizes the platform's short video recommendation plan based on the user's judgment data on the emotional value of short videos, and personalizes the push of short videos with emotional guidance.
5. A short video emotional value detection and quantitative evaluation system, characterized by The system comprises: An STM32 single-chip computer is used to connect a heart rate sensor and a blood oxygen sensor to the STM32 single-chip computer, and the heart rate sensor and the blood oxygen sensor are packaged into a wearable device; the heart rate and blood oxygen data are collected, the sampling time interval and the working cycle of the sensor are controlled, and data preprocessing is performed. The collected raw data is sent to a computer processor through a serial port to provide basic data for feature fusion; A digital camera is used to collect facial video data, and the collected facial video data is analyzed through a deep learning algorithm to obtain information about the user's facial expression, blinking frequency, and gender; Computers perform data processing and analysis, construct sliding window feature vectors, perform feature fusion on multimodal data, and quantify the emotional value of short videos; the emotional value data is uploaded to the short video platform recommendation system.