Video emotion recognition method, electronic equipment, storage medium and product
By acquiring the content information and timing information of the video frame, using explicit time encoding and smoothing processing, the problem of neglecting the timing relationship in the prior art is solved, and a higher precision and stable video emotion recognition is achieved.
Patent Information
- Application Number
- CN202510495868.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-08-15
Smart Images

Figure CN120495727A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of emotion recognition technology, and in particular to a video emotion recognition method, electronic device, storage medium, and product. Background Art
[0002] Current video emotion recognition technology mainly focuses on identifying the emotional state of characters in videos, such as inferring emotions by analyzing facial expressions and voice signals.
[0003] However, existing emotion recognition methods mainly focus on the modal information of a single frame, resulting in the model being unable to effectively capture emotional changes in the video. Summary of the Invention
[0004] The embodiments of the present application provide a video emotion recognition method, electronic device, storage medium and program product to improve the accuracy of emotion recognition and at least partially solve the above-mentioned technical problems.
[0005] In order to achieve the above-mentioned object, according to a first aspect of the present application, a method for emotion recognition is provided, comprising:
[0006] Processing the target video to obtain content information and timing information of multiple video frames of the target video;
[0007] Emotion recognition is performed on the target video according to the content information and timing information of the multiple video frames to obtain emotion information of the video frames.
[0008] Optionally, the processing the target video to obtain content information and timing information of multiple video frames of the target video includes:
[0009] Performing frame extraction processing on the target video to obtain multiple video frames, and obtaining content information of each video frame; and
[0010] Generating timing information of the video frame according to the time stamp of the video frame.
[0011] Optionally, the timing information includes a timing feature vector.
[0012] Optionally, the temporal feature vector includes a time code corresponding to the video frame.
[0013] Optionally, before identifying the emotional information of the video frame according to the content information and timing information of the video frame in the target video, the method further includes:
[0014] According to the timestamp of the video frame, timing information of the video frame is generated, wherein the timing information includes a time code corresponding to the video frame.
[0015] Optionally, generating the timing information of the video frame according to the timestamp of the video frame includes:
[0016] The timing information of the video frame is generated according to the timestamp of the video frame and a preset first function.
[0017] Optionally, the first function includes a first periodic function and a second periodic function, the vector elements at even positions in the timing information are generated based on the first periodic function and the timestamp, and the vector elements at odd positions in the timing information are generated based on the second periodic function and the timestamp.
[0018] Optionally, generating the timing information of the video frame according to the timestamp of the video frame includes:
[0019] The timing information of the video frame is generated according to the timestamp of the video frame and a preset second function, where the second function is set based on a learnable parameter.
[0020] Optionally, the second function includes a first sub-function set based on the learnable parameters, and a second sub-function set based on the first sub-function; a portion of the continuous vector elements in the timing information are generated based on the first sub-function and the timestamp, and the remaining vector elements are generated based on the second sub-function and the timestamp.
[0021] Optionally, the first sub-function is a linear function, and the second sub-function is a periodic function set based on the linear function.
[0022] Optionally, the vector dimension of the time series information is D, and the vector elements of the first X dimensions in the time series information are generated by a first sub-function, where X is a maximum integer not exceeding D / 2.
[0023] Optionally, the extracting a frame from the target video to obtain a plurality of video frames, and obtaining content information of each video frame includes:
[0024] Performing frame extraction processing on the target video according to a preset frequency to obtain multiple video frames;
[0025] Feature extraction is performed on content data of at least one modality of the video frame to obtain content information.
[0026] Optionally, the content information includes a content feature vector.
[0027] Optionally, the content feature vector of the video frame includes an image feature vector and / or an audio feature vector.
[0028] Optionally, after performing emotion recognition on the target video according to the content information and timing information of the multiple video frames to obtain the emotion information of the video frames, the method further includes:
[0029] The emotional information of the video frame is smoothed to obtain target emotional information of the video frame.
[0030] Optionally, the smoothing process on the emotional information of the video frame to obtain target emotional information includes:
[0031] Based on the emotion information of the reference frame of the video frame, the emotion information of the video frame is smoothed to obtain target emotion information.
[0032] Optionally, the performing smoothing processing on the emotional information of the video frame based on the emotional information of the reference frame of the video frame to obtain target emotional information includes:
[0033] The target emotional information of the video frame is obtained by performing weighted summation based on the emotional information of the reference frame of the video frame and the first weight, and the emotional information of the video frame and the second weight.
[0034] Optionally, the smoothing process on the emotional information of the video frame to obtain target emotional information of the video frame further includes:
[0035] When the difference in emotional information between the video frame and a reference frame of the video frame is less than a preset smoothing threshold, the emotional information of the video frame is smoothed to obtain target emotional information of the video frame.
[0036] Optionally, the method further includes:
[0037] When a difference in emotion information between the video frame and a reference frame of the video frame is not less than a preset smoothing threshold, the emotion information of the video frame is used as target emotion information of the video frame.
[0038] Optionally, performing emotion recognition on the target video according to the content information and timing information of the multiple video frames to obtain the emotion information of the video frames includes:
[0039] For each video frame of the target video, the content information and the timing information corresponding to the video frame are fused to obtain a comprehensive representation vector;
[0040] According to the comprehensive representation vector, emotional information corresponding to the video frame is obtained.
[0041] Optionally, fusing the content information and timing information corresponding to the video frames to obtain a comprehensive representation vector includes:
[0042] The content information and timing information corresponding to the video frames are vector-concatenated to obtain a comprehensive representation vector.
[0043] Optionally, obtaining the emotion information corresponding to the video frame according to the comprehensive representation vector includes:
[0044] The emotion information corresponding to the video frame is obtained based on the comprehensive representation vector using a preset emotion prediction model.
[0045] Optionally, the emotion prediction model training step includes:
[0046] Identifying predicted emotional information of the sample video frame based on a sample comprehensive representation vector of the sample video frame by using a to-be-trained emotion prediction model, wherein the sample comprehensive representation vector is obtained by fusing content information and timing information of the sample video frame;
[0047] The emotion prediction model is trained based on the sample weights corresponding to the sample video frames and the predicted emotion information.
[0048] Optionally, the training of the emotion prediction model based on the sample weights corresponding to the sample video frames and the predicted emotion information includes:
[0049] Calculating the loss of the emotion prediction model based on the sample weights corresponding to the sample video frames and the predicted emotion information;
[0050] Parameters of the sentiment prediction model are adjusted based on the loss.
[0051] Optionally, the emotion prediction model includes a machine learning model or a deep learning model.
[0052] Optionally, for the deep learning model, identifying predicted emotion information of the sample video frame based on the sample comprehensive representation vector of the sample video frame by using the emotion prediction model to be trained includes:
[0053] Inputting the comprehensive representation vectors of the sample video frame and other sample video frames belonging to the same time window as the sample video frame into the emotion prediction model;
[0054] The emotion prediction model obtains emotion information corresponding to the sample video frame based on the input comprehensive representation vector.
[0055] Optionally, the calculating the loss of the emotion prediction model based on the sample weights corresponding to the sample video frames and the predicted emotion information includes:
[0056] Based on the predicted emotion information and emotion label of each sample video frame, obtaining emotion difference information of the sample video frame;
[0057] Calculate the sample error of each sample video frame based on the weight loss and sentiment difference information corresponding to each sample video frame;
[0058] The loss of the emotion prediction model is determined based on the sample error of each sample video frame.
[0059] Optionally, the step of obtaining the sample weights corresponding to the sample video frames includes:
[0060] According to the emotion label corresponding to the sample video frame, a sample weight corresponding to the sample video frame is obtained.
[0061] Optionally, the emotion label includes emotion information of multiple emotion representation dimensions, and obtaining the sample weight corresponding to the sample video frame according to the emotion label corresponding to the sample video frame includes:
[0062] For each sample video frame, bucket processing is performed on the sample video frame according to the emotional information of multiple emotion representation dimensions to obtain the target bucket coordinates to which the sample video frame belongs;
[0063] A sample weight corresponding to the sample video frame is determined according to the target bucket coordinates of the sample video frame.
[0064] Optionally, determining the sample weight corresponding to the sample video frame according to the target bucket coordinates of the sample video frame includes:
[0065] Based on the target bucket coordinates of the sample video frame, obtaining the emotion label distribution density of the sample video frame;
[0066] According to the emotion tag distribution density and the target bucket coordinates, a sample weight corresponding to the sample video frame is obtained.
[0067] Optionally, obtaining the sample weight corresponding to the sample video frame according to the emotion label distribution information and the target bucket coordinates includes:
[0068] According to the density information corresponding to the target bucket coordinates in the emotion tag distribution density, a sample weight corresponding to the sample video frame is obtained.
[0069] Optionally, obtaining the sample weight corresponding to the sample video frame according to the density information corresponding to the target bucket coordinates in the emotion tag distribution density includes:
[0070] The density information corresponding to the target bucket coordinates in the emotion tag distribution density is squared and inverted to obtain a sample weight corresponding to the sample video frame.
[0071] Optionally, before obtaining the sample weight corresponding to the sample video frame according to the density information corresponding to the target bucket coordinates in the emotion tag distribution density, the method further includes:
[0072] The emotion tag distribution density is optimized to obtain an optimized emotion tag distribution density.
[0073] Optionally, optimizing the emotion tag distribution density to obtain an optimized emotion tag distribution density includes:
[0074] The emotion tag distribution density is optimized by using a binary Gaussian kernel density estimation method to obtain an optimized emotion tag distribution density.
[0075] In a second aspect, this embodiment further provides an electronic device, comprising a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.
[0076] In a third aspect, this embodiment further provides a computer-readable storage medium, which includes a computer program. When the computer program is run on an electronic device, the computer program is used to enable the electronic device to execute the steps of the above method.
[0077] In a fourth aspect, this embodiment also provides a computer program product, including a computer program, which is stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device performs the steps of the above method.
[0078] To sum up, the embodiments of the present application, through the above-mentioned technical scheme, can process the target video to obtain the content information and timing information of multiple video frames of the target video. In this way, the emotional information of the video frame can be identified based on the content information and the timing information of the video frame, so as to capture the emotional changes in the video based on the timing relationship between the video frames, and effectively improve the recognition accuracy of the emotional information in the video.
[0079] Other features and advantages of the present application will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0080] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. Those skilled in the art can also derive other drawings based on these drawings without inventive effort.
[0081] In order to more completely understand the present application and its beneficial effects, the following description will be given in conjunction with the accompanying drawings, wherein the same drawing numbers represent the same parts in the following description.
[0082] Figure 1 It is a schematic diagram of the related technology provided by this application;
[0083] Figure 2 This is a first schematic diagram of the emotion recognition process provided in an exemplary embodiment of the present application;
[0084] Figure 3 is a schematic diagram of a data processing flow provided in an exemplary embodiment of the present application;
[0085] Figure 4 is a first schematic diagram of the density distribution of emotion labels provided in an exemplary embodiment of the present application;
[0086] Figure 5 is a second schematic diagram of the density distribution of emotion labels provided in an exemplary embodiment of the present application;
[0087] Figure 6 is a schematic diagram of a sample weighting process provided in an exemplary embodiment of the present application;
[0088] Figure 7 This is a second schematic diagram of the emotion recognition process provided in an exemplary embodiment of the present application;
[0089] Figure 8 is a schematic diagram of an emotion recognition device provided in an exemplary embodiment of the present application;
[0090] Figure 9 Schematic diagram of the electronic device provided in an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0091] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the embodiments described are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.
[0092] In conjunction with the background technology of this application, current video emotion recognition technology mainly focuses on identifying the emotional state of people in videos, such as inferring emotions by analyzing facial expressions and voice signals. Existing methods mainly rely on multimodal fusion models, such as Figure 1As shown in Figure 1, the multimodal fusion model first extracts individual features from the visual and audio modalities in the video data on a frame-by-frame basis. The fusion model then integrates these unimodal features into a comprehensive feature vector that represents the overall content of the video. Finally, a regression model is used to predict the emotional state of each video frame based on the comprehensive feature vector. This approach leverages both visual and audio information to achieve a deep understanding and accurate prediction of video emotion. Although this approach has made some progress in video emotion recognition, it still has some drawbacks:
[0093] First, existing multimodal fusion models primarily focus on the modal information of a single frame during unimodal feature extraction, while ignoring the temporal relationships between frames. Although sequence models are subsequently used to capture temporal changes, this implicit temporal feature extraction method increases the learning difficulty of the model and may result in the model being unable to effectively capture emotional changes in the video.
[0094] Secondly, video data often has an uneven distribution of emotions. Because users tend to display relatively neutral emotions when watching videos, data with fluctuating emotions accounts for a relatively small proportion of the entire dataset. This data imbalance can lead to prediction bias during model training, thus affecting the model's predictive performance.
[0095] In order to solve the above problems, the present application proposes an emotion recognition method, an emotion recognition device, an electronic device, a computer-readable storage medium, a computer program product and a vehicle, aiming to improve the accuracy of video emotion recognition.
[0096] In a specific embodiment, the emotion recognition method in the present application can be used in a terminal device, which can be a mobile phone, tablet, computer, server and other network devices, etc., without specific limitation.
[0097] like Figure 2 As shown, the emotion recognition method in this application may at least include the following steps:
[0098] S10, processing the target video to obtain content information and timing information of multiple video frames of the target video;
[0099] It should be noted that, in this embodiment, the target video may be a video to be subjected to emotion recognition, and this embodiment does not limit parameters such as the source and format of the target video.
[0100] Furthermore, in this embodiment, the content information of the target content may include content feature vectors, such as audio feature vectors, video feature vectors, etc., and the timing information may include timing feature vectors corresponding to video frames, such as time codes, etc.
[0101] S20 , performing emotion recognition on the target video according to the content information and timing information of the multiple video frames to obtain emotion information of the video frames.
[0102] In this embodiment, the terminal device can identify the emotional information of the video frame based on the content information and timing information of the video frame in the target video.
[0103] It should be noted that, in this embodiment, the terminal device can process the video file and its corresponding tag file, and then can frame the video according to the user emotion sampling frequency recorded in the tag file (for example, but not limited to 1 frame per second). For example, on the one hand, the image data of the video in the sampling frame is extracted as the input of the image modality model; on the other hand, the audio data between adjacent sampling frames is captured as the input of the audio modality model, and then these intercepted data can be paired with the corresponding emotion tags in chronological order to construct an ordered frame data sequence S = {s1,…,s T}. Where T is the number of video frames.
[0104] It is worth noting that in this embodiment, the time coding in the timing information can be explicit time coding, which can be directly integrated into the feature extraction process later, so that the model can more accurately capture the temporal dynamics of video emotions, thereby improving the accuracy of emotion recognition.
[0105] Therefore, in the embodiments of the present application, the emotional information of the video frames can be identified based on the content information and timing information of the video frames in the target video. Compared with the existing technology that only focuses on the information of a single frame, the present application can combine the timing information of the video frames with the content information to identify the emotional information of the video frames, so as to capture the emotional changes in the video based on the timing relationship between the video frames, and effectively improve the accuracy of identifying the emotional information in the video.
[0106] In one embodiment, the above S10, “processing the target video to obtain content information and timing information of multiple video frames of the target video”, may further include:
[0107] S101, performing frame extraction processing on the target video to obtain multiple video frames, and obtaining content information of each video frame; and
[0108] S102: Generate timing information of the video frame according to the timestamp of the video frame.
[0109] In this embodiment, the terminal device can perform frame extraction processing on the target video and obtain content information of each video frame.
[0110] In addition, the terminal device can also obtain the timestamp of each video frame, and then generate the timing information of the video frame according to the timestamp of the video frame.
[0111] In one embodiment, the time series information includes a time series feature vector.
[0112] In one embodiment, the temporal feature vector includes a time code corresponding to the video frame.
[0113] It should be noted that in this embodiment, the emotion recognition of video content relies heavily on temporal information. For example, the emotional information conveyed by observing a person change from expressionless to smiling is completely different from that conveyed by observing a person change from smiling to expressionless. In the feature extraction process, the data modal features of each frame are processed independently, but in the subsequent multimodal fusion stage, the traditional machine learning model does not consider the temporal relationship between frames; and for the deep learning recurrent neural network model, although it can learn temporal change information, this information is implicit, there is a certain degree of learning difficulty, and there is a high requirement for the number of samples.
[0114] Therefore, in order to solve the above problems, this application can define time coding and directly integrate the timing information into the feature representation of each frame. This method explicitly enhances the timing change information between frames, thereby improving the performance of the model in predicting video emotions.
[0115] For diverse emotion prediction models, this application can provide different time encoding strategies to adapt to the needs of different model architectures.
[0116] For example, for deep learning models, this application can use a learnable dynamic time encoding mechanism to inject time information, so that the model can adaptively learn complex changes in time series.
[0117] As for traditional machine learning models, given that their characteristics are not learnable, this application proposes two solutions: one is to directly adopt the static position encoding method in the Transformer model to represent time information through fixed position encoding; the other is to use the time encoding module pre-trained in the deep learning model to generate time encoding suitable for the machine learning model.
[0118] On this basis, the above-mentioned “generating the timing information of the video frame according to the timestamp of the video frame” may include:
[0119] The timing information of the video frame is generated according to the timestamp of the video frame and a preset first function.
[0120] In this embodiment, the terminal device may generate a static time code of the video frame according to the timestamp of the video frame and a preset first function.
[0121] In a specific embodiment, the first function includes a first periodic function and a second periodic function, the vector elements at even positions in the timing information are generated based on the first periodic function and the timestamp, and the vector elements at odd positions in the timing information are generated based on the second periodic function and the timestamp.
[0122] Specifically, for example, for any video frame s i , the corresponding timestamp is t i , its static time code TE(i) is a D-dimensional vector, and each element is defined as follows:
[0123]
[0124] Here, MOD(·,·) is a remainder operation. This representation vector simulates the periodicity of temporal information changes through trigonometric functions and does not require learning, making it suitable for scenarios with relatively small amounts of data. However, because this temporal encoding is independent of data content, it struggles to capture relative temporal changes between video frames, and therefore performs inferior to learnable temporal encoding.
[0125] In one embodiment, the above-mentioned “generating the timing information of the video frame according to the timestamp of the video frame” may include:
[0126] The timing information of the video frame is generated according to the timestamp of the video frame and a preset second function, wherein the preset function is set based on a learnable parameter.
[0127] In this embodiment, the terminal device may generate a learnable time code of the video frame according to the timestamp of the video frame and a preset second function.
[0128] Specifically, for example, for any sample s i , the corresponding timestamp is t i , its learnable temporal encoding TE(i) is a D-dimensional vector, defined as follows:
[0129]
[0130] Among them, w j and b j are all learnable parameters. The first half of this time encoding captures the aperiodic nature of the time series, while the second half captures the periodic nature. Furthermore, this encoding mechanism can be directly integrated into deep learning models for training and, after training, used as temporal features to input into machine learning models. After training, this time encoding can fit the training data, providing superior time series feature extraction capabilities compared to static encoding.
[0131] In a specific embodiment, the second function includes a first sub-function set based on the learning parameter, and a second sub-function set based on the first sub-function; a portion of continuous vector elements in the timing information are generated based on the first sub-function and the timestamp, and the remaining vector elements are generated based on the second sub-function and the timestamp.
[0132] In this embodiment, the first sub-function is w j t i +b j , and the second subfunction is SIN(w j t i +b j ), it can be seen that the second sub-function is actually determined based on the first sub-function.
[0133] On this basis, a part of the continuous vector elements in the time series feature vector (vector elements in the range of j≤D / 2) is based on the first sub-function w j t i +b j and timestamp t i Generate, the remaining vector elements (vector elements in the range of j>D / 2) are based on the second sub-function SIN(w j t i +b j ) and timestamp t i generate.
[0134] In one embodiment, the first sub-function is a linear function, and the second sub-function is a periodic function based on the linear function.
[0135] Combined with the above description, the first sub-function w j t i +b j is a linear function, and the second sub-function is based on the first sub-function w j t i +b j The periodic function SIN(w j t i +b j ).
[0136] In one embodiment, the vector dimension of the time series feature vector is D, and the vector elements of the first X dimensions in the time series feature vector are generated by a first sub-function, where X is a maximum integer not exceeding D / 2.
[0137] Combined with the above description, the vector elements of the first X dimensions in the time series feature vector (the vector elements in the range of j≤D / 2) are based on the first sub-function w j t i +b j and timestamp t iGenerate, the remaining vector elements (vector elements in the range of j>D / 2) are based on the second sub-function SIN(w j t i +b j ) and timestamp t i generate.
[0138] In this way, in the embodiment of the present application, by adding explicit time coding, time information is directly integrated into the feature extraction process, so that the subsequent emotion prediction model can more accurately capture the temporal dynamics of video emotions, thereby improving the accuracy of emotion recognition.
[0139] In one embodiment, the above-mentioned “performing frame extraction processing on the target video to obtain multiple video frames, and obtaining content information of each video frame” may further include:
[0140] Performing frame extraction processing on the target video according to a preset frequency to obtain multiple video frames;
[0141] Feature extraction is performed on content data of at least one modality of the video frame to obtain content information.
[0142] In this embodiment, if Figure 3 As shown, the terminal device can process the video file and its corresponding tag file. First, the video is divided into frames according to the user emotion sampling frequency recorded in the tag file (usually 1 frame per second). Specifically, on the one hand, the image data of the video in the sampling frame is extracted as the input of the image modality model; on the other hand, the audio data between adjacent sampling frames is captured as the input of the audio modality model. Subsequently, these intercepted data are paired with the corresponding emotion tags in chronological order to construct an ordered frame data sequence S = {s1,…,s T}. Where T is the number of video frames.
[0143] Furthermore, if Figure 3 As shown in the figure, the image and audio data extracted in the above steps can be preprocessed. For image modal data, they are first scaled to a uniform resolution through methods such as bilinear interpolation to meet the model input requirements; then, the central area of the image is cropped to remove unnecessary edge parts and focus on displaying the core content of the picture; finally, the image is normalized to adjust the pixel values to a predetermined range, reducing the adverse effects of pixel value differences between different images on model training and enhancing the generalization performance of the model.
[0144] For audio modal data, we first unify the audio sampling rate to ensure the consistency of the model input; then, we use deep learning technology to separate the human voice and background noise in the audio to further improve the training effect of the audio modal model.
[0145] Furthermore, the terminal device can perform feature extraction of the corresponding modality on the image data and audio data obtained above.
[0146] Specifically, for example, when performing image feature extraction, for each frame of image data, a deep model pre-trained on a large-scale dataset, such as ResNet, EfficientNet, etc., is used to extract different features of the image data through stacked convolutional neural networks, pooling layers, and fully connected layers, including target object features, scene features, character facial features, character limb features, etc., and finally obtain a representation vector set of the image features of the frame Among them, N v is the number of image features.
[0147] When extracting audio features, for each piece of audio data, we use deep models pre-trained on large-scale datasets, such as VGGish and DistilHubert, to extract different features of the audio data, including human voice features, background music features, etc., through multi-layer convolutional neural networks, recurrent neural networks, and fully connected layers, and finally obtain a set of representation vectors of the audio features. Among them, N a is the number of audio features.
[0148] In this way, the present application can obtain the image feature vector of each video frame and the audio feature vector
[0149] In one embodiment, the content information of the video frame includes an image feature vector and / or an audio feature vector.
[0150] In this embodiment, the content information of the video frame may include an image feature vector and / or an audio feature vector.
[0151] In one embodiment, after the above S20, “performing emotion recognition on the target video according to the content information and timing information of the multiple video frames to obtain emotion information of the video frames”, the following steps may also be performed:
[0152] S30, performing smoothing processing on the emotional information of the video frame to obtain target emotional information of the video frame.
[0153] In this embodiment, considering that viewers' emotions fluctuate slightly most of the time while watching a video, the prediction results for a continuous segment of frames should be smooth in most cases. However, for a machine learning model, there is no correlation between the prediction results of each frame, resulting in some abnormal prediction values affecting the overall prediction performance of the model. To address this problem, the embodiment of the present application can smooth the emotional information of the video frame to obtain the target emotional information of the video frame.
[0154] In a specific embodiment, the above-mentioned “smoothing the emotional information of the video frame to obtain target emotional information” may include:
[0155] S301 , based on the emotion information of the reference frame of the video frame, smoothing the emotion information of the video frame to obtain target emotion information.
[0156] In this embodiment, the terminal device can use the emotional information of the reference frame to smooth the emotional information of the current video frame to obtain target emotional information.
[0157] The reference frame may be, but is not limited to, a previous video frame adjacent to the current video frame, or may be other historical video frames in the target video, which is not specifically limited.
[0158] In a specific embodiment, the above-mentioned “smoothing the emotional information of the video frame based on the emotional information of the reference frame of the video frame to obtain target emotional information” may include:
[0159] S3011, performing weighted summation based on the emotion information of the reference frame of the video frame and the first weight, and the emotion information of the video frame and the second weight, to obtain target emotion information of the video frame.
[0160] In this embodiment, the terminal device can obtain the target emotional information of the current video frame based on the emotional information of the reference frame of the current video frame and the first weight, and the emotional information of the current video frame and the second weight.
[0161] Specifically, for example, Among them, yi-1 is the emotional information corresponding to the reference frame, β is the first weight, yi is the emotional information of the current video frame, and (1-β) is the second weight.
[0162] Among them, β∈(0,1) is the smoothing coefficient, which determines the extent to which past prediction results are retained.
[0163] In another embodiment, when the difference in emotional information between the video frame and the reference frame of the video frame is less than a preset smoothing threshold, the emotional information of the video frame is smoothed based on the emotional information of the reference frame of the video frame to obtain target emotional information.
[0164] When a difference in emotion information between the video frame and a reference frame of the video frame is not less than a preset smoothing threshold, the emotion information of the video frame is used as target emotion information of the video frame.
[0165] In this embodiment, for the recognized emotional information sequence {y1,…,y T}, for the prediction result yi of one of the frames, the filtered prediction result Expressed as:
[0166]
[0167] Among them, β∈(0,1) is the smoothing coefficient, which determines the extent to which past prediction results are retained; β∈(0,1) is the threshold coefficient, which allows the prediction to jump out of the smoothing process when the prediction results change greatly, thereby retaining the model's sensitivity to emotional fluctuations.
[0168] As can be seen, compared to the prior art that directly uses the output of the prediction model, these prediction results often fluctuate greatly, resulting in unstable prediction effects. In the embodiment of the present application, the output of the prediction model is smoothed by a filter, while retaining the prediction results with large emotion changes. This effectively improves the stability of the emotion prediction curve without affecting the prediction of emotion fluctuations, making it closer to the data characteristics of real emotion changes.
[0169] In one embodiment, in S20 above, “performing emotion recognition on the target video according to the content information and timing information of the multiple video frames to obtain emotion information of the video frames” may include:
[0170] S201, for each video frame of the target video, fusing content information and timing information corresponding to the video frame to obtain a comprehensive representation vector;
[0171] S202: Obtain emotion information corresponding to the video frame according to the comprehensive representation vector.
[0172] In this embodiment, after the terminal device obtains the content feature vector and the timing feature vector corresponding to the video frame through feature extraction, it can fuse the content feature vector and the timing feature vector corresponding to the video frame to obtain a comprehensive representation vector, and obtain the emotional information corresponding to the video frame based on the comprehensive representation vector.
[0173] Specifically, for example, the terminal device can perform multimodal feature fusion on the image feature vector, audio feature vector and time series feature vector obtained above. For example, for any single video frame, the image feature vector of its image modality is The audio feature vector of the audio modality is The time code (i.e., the time series feature vector) is TE(i), and its comprehensive representation vector E(i) is:
[0174]
[0175] Among them, CONCAT(·) is a basic vector concatenation operation.
[0176] The terminal device can then input the comprehensive representation vector E(i) into the emotion prediction model to predict the emotional information of the target video. The subsequent embodiments will be described in detail and will not be repeated here.
[0177] In a specific embodiment, the above-mentioned “fusing the content information and timing information corresponding to the video frame to obtain a comprehensive representation vector” may include:
[0178] The content information and timing information corresponding to the video frames are vector-concatenated to obtain a comprehensive representation vector.
[0179] In this embodiment, the image feature vector Audio feature vector of the audio modality When the time code TE(i) is fused, the feature vectors of multiple modes can be directly spliced to obtain the comprehensive representation vector of single frame data.
[0180] In one embodiment, in S102 above, “obtaining emotion information corresponding to the video frame according to the comprehensive representation vector” may include:
[0181] S1021: Obtain emotion information corresponding to the video frame based on the comprehensive representation vector through an emotion prediction model.
[0182] In this embodiment, the comprehensive representation vector E(i) can be input into the emotion prediction model to obtain the emotion information corresponding to the video frame output by the emotion prediction model.
[0183] In one embodiment, the steps of training the emotion prediction model in the present application may include:
[0184] S40, identifying predicted emotion information of the sample video frame based on a sample comprehensive representation vector of the sample video frame using the emotion prediction model to be trained, where the sample comprehensive representation vector is obtained by fusing content information and timing information of the sample video frame;
[0185] S50: Training the emotion prediction model based on the sample weights corresponding to the sample video frames and the predicted emotion information.
[0186] In this embodiment, the sample comprehensive representation vector can be obtained based on the fusion of content information and timing information of the sample video frame. Please refer to the above embodiment and will not be repeated here.
[0187] On this basis, the terminal device can identify the predicted emotional information of the sample video frame based on the sample comprehensive representation vector of the sample video frame through the emotion prediction model to be trained, and train the emotion prediction model based on the sample weights and predicted emotional information corresponding to the sample video frame.
[0188] In a specific embodiment, in the above S50, “training the emotion prediction model based on the sample weights corresponding to the sample video frames and the predicted emotion information” may include:
[0189] S501, calculating the loss of the emotion prediction model based on the sample weights corresponding to the sample video frames and the predicted emotion information;
[0190] S502: Adjust parameters of the emotion prediction model based on the loss.
[0191] In this embodiment, the terminal device may calculate the loss of the emotion prediction model based on the sample weights corresponding to the sample video frames and the predicted emotion information, and may further adjust the parameters of the emotion prediction model based on the loss.
[0192] It should be noted that, in this embodiment, the input of the emotion prediction model is the feature representation vector sequence {E(1),…,E(T)}, and the output is the emotion information of the video in each sampling frame. The emotion information can specifically be an emotion label, and the emotion prediction model can be a machine learning model, a deep learning model, etc., and there is no specific limitation on this.
[0193] In a specific embodiment, the emotion prediction model includes a machine learning model or a deep learning model.
[0194] In an embodiment of the present application, the machine learning model can only model single-frame data. For example, the machine learning model can be an XGBoost model.
[0195] The deep learning model can be an LSTM model, and its model parameter is φ.
[0196] In one embodiment, for the deep learning model, in S40, “identifying predicted emotion information of the sample video frame based on the sample comprehensive representation vector of the sample video frame by the emotion prediction model to be trained” may include:
[0197] S401, inputting the comprehensive representation vector of the sample video frame and other sample video frames belonging to the same time window as the sample video frame into the emotion prediction model;
[0198] S402: Acquire emotion information corresponding to the sample video frame through the emotion prediction model based on the input comprehensive representation vector.
[0199] It should be noted that in this embodiment, the deep learning model is able to model sequence data. Therefore, in order to utilize the sequence modeling capability of the deep learning model, a time window can be maintained on the data sequence, and the data in the time window can be input into the deep learning model as a whole by sliding according to time.
[0200] Specifically, for example, the terminal device can input the comprehensive representation vector of the sample video frame and other sample video frames belonging to the same time window as the sample video frame into the emotion prediction model, and then the emotion prediction model can obtain the emotion information corresponding to the sample video frame based on the input comprehensive representation vector.
[0201] In one embodiment, in S501 above, “calculating the loss of the emotion prediction model based on the sample weights corresponding to the sample video frames and the predicted emotion information” may include:
[0202] S5011, obtaining emotion difference information of each sample video frame based on the predicted emotion information and emotion label of the sample video frame;
[0203] S5012, calculating a sample error for each sample video frame based on the weight loss and sentiment difference information corresponding to each sample video frame;
[0204] S5013: Determine the loss of the emotion prediction model based on the sample error of each sample video frame.
[0205] In this embodiment, the terminal device can obtain the emotional difference information of the sample video frame based on the predicted emotional information and emotional label of each sample video frame, and calculate the sample error of each sample video frame based on the weight loss and emotional difference information corresponding to each sample video frame, and then determine the loss of the emotional prediction model based on the sample error of each sample video frame.
[0206] In a specific embodiment, the input of the emotion prediction model is a feature representation vector sequence {E(1),…,E(T)}, and the output is the emotion information of the video in each sampling frame, which can specifically be an emotion label.
[0207] For machine learning models:
[0208] The machine learning model can only model single-frame data. Taking the XGBoost model as an example, its model parameters are set to . Based on the training samples, the loss function is constructed as follows:
[0209]
[0210] Here, MAE(·,·) is the mean absolute error; R(·) is the regularization function used to prevent model overfitting, and λ is the weight coefficient of the regularization term. Using this loss function, the model parameters θ are trained and optimized. Finally, the trained model is used to predict the sentiment of the test data. Although machine learning models can only model single-frame data, they offer greater real-time performance and lower computing power requirements than deep learning models, making them a cost-effective solution.
[0211] For deep learning models:
[0212] Deep learning models can model sequential data. Therefore, to leverage the sequence modeling capabilities of deep learning models, a time window is maintained on the data sequence and then sliding over time. The data within this time window is input into the deep learning model as a whole. Taking the LSTM model as an example, let its model parameters be, and based on the training samples, the loss function is constructed as follows:
[0213]
[0214] Where MSE(·,·) is the mean square error; R(·) is the regularization function used to prevent the model from overfitting, λ is the weight coefficient of the regularization term; T w is the size of the time window, and the input data is the sequence {E(iT w ),…,E(i)}. This loss function is used to train and optimize the model parameters φ. Finally, the trained model is used to predict the sentiment of the test data. Deep learning models can model sequential patterns and capture the changing characteristics between adjacent frames, thus achieving higher emotion recognition accuracy.
[0215] It is understandable that the above loss calculation method is only based on a single sample video as an example, but the actual final total loss function of this application is the average of the sum of the losses of all video samples.
[0216] In one embodiment, the step of obtaining the sample weights corresponding to the sample video frames may include:
[0217] S60 , obtaining a sample weight corresponding to the sample video frame according to the emotion label corresponding to the sample video frame.
[0218] It should be noted that, in this embodiment, the emotion labels corresponding to the sample video frames may be obtained by manual or machine calibration, and there is no specific limitation on this.
[0219] The terminal device can obtain the sample weight corresponding to the sample video frame based on the emotion label corresponding to the sample video frame. The sample weight can be used to ensure that the model can pay more attention to those samples that are sparsely distributed but more important in the data set, thereby improving the training efficiency and prediction performance of the model.
[0220] In one embodiment, the emotion label includes emotion information of multiple emotion representation dimensions. In the above S60, "obtaining a sample weight corresponding to the sample video frame according to the emotion label corresponding to the sample video frame" may include:
[0221] S601, for each sample video frame, performing bucket processing on the sample video frame according to emotion information of multiple emotion representation dimensions, and obtaining the target bucket coordinates to which the sample video frame belongs;
[0222] S602 : Determine a sample weight corresponding to the sample video frame according to the target bucket coordinates of the sample video frame.
[0223] It should be noted that, in this embodiment, a multi-dimensional bucketing strategy can be used to estimate the joint empirical distribution of pleasure and arousal labels in the emotion tag, wherein the emotion information of multiple emotion representation dimensions contained in the emotion tag can include pleasure and arousal.
[0224] Specifically, for example, for any sample video frame s i , whose sentiment label is y i , and their normalized pleasure and arousal are and
[0225] Then the terminal device can and arousal For sample video frame s i Perform bucket processing to obtain the target bucket coordinate BIN(i) to which the sample video frame belongs:
[0226]
[0227] ROUND(·) is rounded to the nearest integer, and B is the number of buckets.
[0228] Then, the sample weight corresponding to the sample video frame can be determined according to the target bucket coordinates of the sample video frame.
[0229] In a specific embodiment, in S602 above, “determining the sample weight corresponding to the sample video frame according to the target bucket coordinates of the sample video frame” may include:
[0230] S6021, based on the target bucket coordinates of the sample video frame, obtaining the emotion label distribution density of the sample video frame;
[0231] S6022: Obtain a sample weight corresponding to the sample video frame according to the emotion tag distribution density and the target bucket coordinates.
[0232] In this embodiment, based on the above bucketing strategy, the number of samples in the bucket corresponding to each target bucket coordinate is determined, and the emotion tag distribution density can be obtained according to the number of samples.
[0233] Then, the sample weight corresponding to the sample video frame can be obtained according to the above-mentioned emotional label distribution density and the target bucket coordinates.
[0234] In a specific embodiment, in S6022 above, “obtaining a sample weight corresponding to the sample video frame according to the emotion tag distribution density and the target bucket coordinates” may include:
[0235] According to the density information corresponding to the target bucket coordinates in the emotion tag distribution density, a sample weight corresponding to the sample video frame is obtained.
[0236] In this embodiment, combined with the above description, the terminal device can and arousal For sample video frame s i Perform bucket processing to obtain the target bucket coordinate BIN(i) to which the sample video frame belongs:
[0237]
[0238] Then, we can use the density information corresponding to the target bucket coordinates Get the sample weight corresponding to the sample video frame.
[0239] In a specific embodiment, the above-mentioned “obtaining a sample weight corresponding to the sample video frame according to density information corresponding to the target bucket coordinates in the emotion tag distribution density” may include:
[0240] The density information corresponding to the target bucket coordinates in the emotion tag distribution density is squared and inverted to obtain a sample weight corresponding to the sample video frame.
[0241] In order to enhance the importance of sparsely distributed samples during training, in this embodiment, the density information corresponding to the target bucket coordinates can be After performing the square root operation, the inverse is taken and used as the sample weight, which is applied to the loss function in the subsequent training process.
[0242]
[0243] Among them, W(i) represents the sample weight. This strategy ensures that the emotion recognition model can pay more attention to those samples that are sparsely distributed but more important in the dataset, thereby improving the training efficiency and prediction performance of the model.
[0244] In one embodiment, before the above “obtaining the sample weight corresponding to the sample video frame according to the density information corresponding to the target bucket coordinates in the emotion tag distribution density”, the following steps may also be included:
[0245] The emotion tag distribution density is optimized to obtain an optimized emotion tag distribution density.
[0246] In order to reduce the similarity effect of continuous tags between similar values, this embodiment can optimize the distribution density of emotion tags to obtain an optimized distribution density of emotion tags.
[0247] In a specific embodiment, the above-mentioned “optimizing the distribution density of the emotion tags to obtain an optimized distribution density of the emotion tags” may include:
[0248] A binary Gaussian kernel density estimation method is used to optimize the emotion tag distribution information to obtain an optimized emotion tag distribution density.
[0249] In this embodiment, in order to reduce the similarity effect of continuous labels between similar values, a binary Gaussian kernel density estimation can be used to approximate the true distribution density of sample labels.
[0250]
[0251] Among them, p(y 1 ,y 2 ) is its empirical label distribution, CONV(·,·) is a two-dimensional Gaussian convolution, for example, Figure 4 is the distribution density of sentiment labels before Gaussian convolution, Figure 5 is the distribution density of sentiment labels after Gaussian convolution (i.e., after optimization).
[0252] Specifically, if Figure 6As shown in the figure, given that video data often exhibits an imbalance in emotional distribution, users may maintain neutral emotions most of the time while watching videos, resulting in a smaller proportion of data with large emotional fluctuations in the entire dataset. This data imbalance may introduce bias in model predictions. Secondly, in the task of video emotion recognition, to capture subtle changes in emotion between different frames, emotion labels are typically composed of continuous values of pleasure (Valence) and arousal (Arousal). Such continuous-valued labels cause similar values to influence each other, making the balancing method commonly used when processing discrete category labels inapplicable.
[0253] Therefore, this application introduces an innovative data balancing method for continuous multi-dimensional emotion labels. Specifically, a two-dimensional bucketing strategy is first used to estimate the joint empirical distribution of pleasure and arousal labels. For any sample s i , whose sentiment label is y i , and their normalized pleasure and arousal are and The bucket coordinates to which the sample belongs are:
[0254]
[0255] Among them, ROUND(·) is rounded off and B is the number of buckets. Based on this bucketing strategy, according to the number of samples in each bucket, the empirical label distribution can be obtained as follows Figure 4 As shown (i.e., the distribution density of sentiment labels before optimization).
[0256] Then, in order to reduce the similarity effect of continuous labels between similar values, this application uses binary Gaussian kernel density estimation to approximate the true distribution density of sample labels.
[0257]
[0258] Among them, p(y 1 ,y 2 ) is its empirical label distribution, CONV(·,·) is a two-dimensional Gaussian convolution, and the result is as follows Figure 5 Finally, in order to enhance the importance of sparsely distributed samples during training, this application proposes to perform a square root operation on the true distribution density of the samples and then negate it, and use it as the sample weight, which is applied to the loss function in the subsequent training process.
[0259]
[0260] Where W(i) represents sample s iThis strategy ensures that the model can pay more attention to those samples that are sparsely distributed but more important in the dataset, thereby improving the training efficiency and prediction performance of the model.
[0261] Compared to existing techniques that directly optimize models using training data, video data often has an uneven distribution of emotions. Users often display relatively neutral emotions when watching videos, resulting in a relatively small proportion of emotionally fluctuating data in the entire dataset, which can cause model prediction bias. This social situation uses a Gaussian convolution kernel to smooth the emotional distribution based on the distribution of training data and calculate weight coefficients. This guides the model to focus more on the smaller but important emotional fluctuation data during training, thereby improving the model's ability to recognize different emotions and reducing prediction bias.
[0262] In general, the emotion recognition method in this application is as follows Figure 7 As shown, it can at least include:
[0263] (1) Video data frame extraction: Read the video data and extract the frames at a frequency of one frame per second. Each frame includes the image modal data at the time of extraction and the audio modal data contained in the interval between frames.
[0264] (2) Data preprocessing: Image modal data is preprocessed, including operations such as image scaling, center cropping, and normalization. Audio modal data is processed by resampling, segmentation, and voice separation.
[0265] (3) Data balance: Count and analyze the joint distribution of the label data in the dataset, use the Gaussian kernel for smooth convolution, and calculate the weight coefficient to solve the data imbalance problem.
[0266] (4) Unimodal feature extraction: To extract image features, a pre-trained model is used to extract image features through stacked convolutional layers, pooling layers, and fully connected layers, including target object features, scene features, facial features, and body features. To extract audio features, a multi-layer convolutional layer, recurrent network layer, and fully connected layer is used to extract audio features, including human voice and background music features.
[0267] (5) Time coding: Generate corresponding timestamp codes for the relative time information of each frame and add them to the unimodal feature set to enhance the model’s understanding of time information.
[0268] (6) Multimodal fusion: After obtaining the feature representation vectors of each modality, the feature information of different modalities is fused through vector cascading, splicing, or calculating the attention coefficient and weighted summation to obtain a comprehensive representation vector of the video data, preparing for subsequent emotion prediction.
[0269] (7) Sentiment prediction: Use classifiers (such as support vector machines and random forests) or regressors (such as multi-layer perceptrons and long short-term memory networks) to predict sentiment based on the fused features.
[0270] (8) Result filtering: Use custom filters to filter the sentiment prediction results, thereby reducing the gap between the prediction results and making the overall prediction results smoother.
[0271] Through the above steps, the present application can effectively realize video emotion recognition, improve recognition accuracy, and have good robustness.
[0272] In general, compared with the prior art, the present application has at least the following differences:
[0273] (1) Explicit temporal coding: Existing multimodal fusion models primarily focus on the modal information of a single frame during feature extraction, ignoring the temporal relationship between frames. This makes model learning more difficult and prevents effective capture of emotional changes in videos. This invention incorporates explicit temporal coding directly into the feature extraction process, enabling the model to more accurately capture the temporal dynamics of video emotions, thereby improving the accuracy of emotion recognition.
[0274] (2) Data balancing: Existing methods directly use training data to optimize the model. However, video data is usually uneven in terms of emotional distribution. Users tend to show a relatively neutral emotional state when watching videos, resulting in a relatively small proportion of emotional fluctuation data in the entire data set, causing model prediction bias. The present invention uses a Gaussian convolution kernel to smooth the emotional distribution based on the distribution of training data and calculates the weight coefficient. This guides the model to pay more attention to the less important but less important emotional fluctuation data during training, thereby improving the model's ability to recognize different emotions and reducing prediction bias.
[0275] (3) Result smoothing and filtering: Existing methods often directly use the output of the prediction model. However, these prediction results often fluctuate greatly, resulting in unstable prediction results. The present invention smoothes the output of the prediction model through a filter, while retaining prediction results with large emotional fluctuations. This method effectively improves the stability of the emotion prediction curve without affecting the prediction of emotional fluctuations, making it closer to the data characteristics of real emotional changes.
[0276] Accordingly, the embodiment of the present application also provides an emotion recognition device, such as Figure 8 As shown, the device may include:
[0277] The processing module 1001 is used to process the target video to obtain content information and timing information of multiple video frames of the target video;
[0278] The recognition module 1002 performs emotion recognition on the target video according to the content information and timing information of the multiple video frames to obtain emotion information of the video frames.
[0279] Optionally, the processing module 1001 is further configured to:
[0280] Performing frame extraction processing on the target video to obtain multiple video frames, and obtaining content information of each video frame; and
[0281] Generating timing information of the video frame according to the time stamp of the video frame.
[0282] Optionally, the timing information includes a timing feature vector.
[0283] Optionally, the temporal feature vector includes a time code corresponding to the video frame.
[0284] Optionally, the processing module 1001 is further configured to:
[0285] Generating timing information of the video frame according to the time stamp of the video frame.
[0286] Optionally, the temporal feature vector includes a time code corresponding to the video frame.
[0287] Optionally, the processing module 1001 is further configured to:
[0288] The timing information of the video frame is generated according to the timestamp of the video frame and a preset first function.
[0289] Optionally, the first function includes a first periodic function and a second periodic function, the vector elements at even positions in the timing information are generated based on the first periodic function and the timestamp, and the vector elements at odd positions in the timing information are generated based on the second periodic function and the timestamp.
[0290] Optionally, the processing module 1001 is further configured to:
[0291] The timing information of the video frame is generated according to the timestamp of the video frame and a preset second function, where the second function is set based on a learnable parameter.
[0292] Optionally, the second function includes a first sub-function set based on the learnable parameters, and a second sub-function set based on the first sub-function; a portion of the continuous vector elements in the timing information are generated based on the first sub-function and the timestamp, and the remaining vector elements are generated based on the second sub-function and the timestamp.
[0293] Optionally, the first sub-function is a linear function, and the second sub-function is a periodic function set based on the linear function.
[0294] Optionally, the vector dimension of the time series information is D, and the vector elements of the first X dimensions in the time series information are generated by a first sub-function, where X is a maximum integer not exceeding D / 2.
[0295] Optionally, the emotion recognition device in the present application further includes:
[0296] The first smoothing processing module is used to perform smoothing processing on the emotional information of the video frame to obtain target emotional information of the video frame.
[0297] Optionally, the first smoothing processing module is further configured to:
[0298] Based on the emotion information of the reference frame of the video frame, the emotion information of the video frame is smoothed to obtain target emotion information.
[0299] Optionally, the first smoothing processing module is further configured to:
[0300] The step of smoothing the emotional information of the video frame based on the emotional information of the reference frame of the video frame to obtain target emotional information includes:
[0301] The target emotional information of the video frame is obtained by performing weighted summation based on the emotional information of the reference frame of the video frame and the first weight, and the emotional information of the video frame and the second weight.
[0302] Optionally, the emotion recognition device in the present application further includes:
[0303] The second smoothing processing module is further configured to, when the difference in emotional information between the video frame and the reference frame of the video frame is less than a preset smoothing threshold, perform smoothing processing on the emotional information of the video frame based on the emotional information of the reference frame of the video frame to obtain target emotional information. When the difference in emotional information between the video frame and the reference frame of the video frame is not less than the preset smoothing threshold, use the emotional information of the video frame as the target emotional information of the video frame.
[0304] Optionally, the processing module 1001 in the present application is further configured to:
[0305] A frame extraction module is used to extract frames from the target video according to a preset frequency to obtain multiple video frames;
[0306] The feature extraction module is used to extract features from content data of at least one modality of the video frame to obtain content information.
[0307] Optionally, the content feature vector of the video frame includes an image feature vector and / or an audio feature vector.
[0308] Optionally, the identification module 1002 is further configured to:
[0309] For each video frame of the target video, the content information and the timing information corresponding to the video frame are fused to obtain a comprehensive representation vector;
[0310] According to the comprehensive representation vector, emotional information corresponding to the video frame is obtained.
[0311] Optionally, the identification module 1002 is further configured to:
[0312] The content information and timing information corresponding to the video frames are vector-concatenated to obtain a comprehensive representation vector.
[0313] Optionally, the identification module 1002 is further configured to:
[0314] The emotion information corresponding to the video frame is obtained based on the comprehensive representation vector using a preset emotion prediction model.
[0315] Optionally, the training method of the emotion prediction model includes:
[0316] Identifying predicted emotional information of the sample video frame based on a sample comprehensive representation vector of the sample video frame by using a to-be-trained emotion prediction model, wherein the sample comprehensive representation vector is obtained by fusing content information and timing information of the sample video frame;
[0317] The emotion prediction model is trained based on the sample weights corresponding to the sample video frames and the predicted emotion information.
[0318] Optionally, the training of the emotion prediction model based on the sample weights corresponding to the sample video frames and the predicted emotion information includes:
[0319] Calculating the loss of the emotion prediction model based on the sample weights corresponding to the sample video frames and the predicted emotion information;
[0320] Parameters of the sentiment prediction model are adjusted based on the loss.
[0321] Optionally, the emotion prediction model includes a machine learning model or a deep learning model.
[0322] Optionally, for the deep learning model, identifying predicted emotion information of the sample video frame based on the sample comprehensive representation vector of the sample video frame by using the emotion prediction model to be trained includes:
[0323] Inputting the comprehensive representation vectors of the sample video frame and other sample video frames belonging to the same time window as the sample video frame into the emotion prediction model;
[0324] The emotion prediction model obtains emotion information corresponding to the sample video frame based on the input comprehensive representation vector.
[0325] Optionally, the calculating the loss of the emotion prediction model based on the sample weights corresponding to the sample video frames and the predicted emotion information includes:
[0326] Based on the predicted emotion information and emotion label of each sample video frame, obtaining emotion difference information of the sample video frame;
[0327] Calculate the sample error of each sample video frame based on the weight loss and sentiment difference information corresponding to each sample video frame;
[0328] The loss of the emotion prediction model is determined based on the sample error of each sample video frame.
[0329] Optionally, a method of obtaining a sample weight corresponding to a sample video frame includes:
[0330] According to the emotion label corresponding to the sample video frame, a sample weight corresponding to the sample video frame is obtained.
[0331] Optionally, the emotion label includes emotion information of multiple emotion representation dimensions, and obtaining the sample weight corresponding to the sample video frame according to the emotion label corresponding to the sample video frame includes:
[0332] For each sample video frame, bucket processing is performed on the sample video frame according to the emotional information of multiple emotion representation dimensions to obtain the target bucket coordinates to which the sample video frame belongs;
[0333] A sample weight corresponding to the sample video frame is determined according to the target bucket coordinates of the sample video frame.
[0334] Optionally, determining the sample weight corresponding to the sample video frame according to the target bucket coordinates of the sample video frame includes:
[0335] Based on the target bucket coordinates of the sample video frame, obtaining the emotion label distribution density of the sample video frame;
[0336] According to the emotion tag distribution density and the target bucket coordinates, a sample weight corresponding to the sample video frame is obtained.
[0337] Optionally, obtaining the sample weight corresponding to the sample video frame according to the emotion label distribution information and the target bucket coordinates includes:
[0338] According to the density information corresponding to the target bucket coordinates in the emotion tag distribution density, a sample weight corresponding to the sample video frame is obtained.
[0339] Optionally, obtaining the sample weight corresponding to the sample video frame according to the density information corresponding to the target bucket coordinates in the emotion tag distribution density includes:
[0340] The density information corresponding to the target bucket coordinates in the emotion tag distribution density is squared and inverted to obtain a sample weight corresponding to the sample video frame.
[0341] Optionally, before obtaining the sample weight corresponding to the sample video frame according to the density information corresponding to the target bucket coordinates in the emotion tag distribution density, the method further includes:
[0342] The emotion tag distribution density is optimized to obtain an optimized emotion tag distribution density.
[0343] Optionally, optimizing the emotion tag distribution density to obtain an optimized emotion tag distribution density includes:
[0344] The emotion tag distribution density is optimized by using a binary Gaussian kernel density estimation method to obtain an optimized emotion tag distribution density.
[0345] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0346] Accordingly, the embodiment of the present application further provides an electronic device, such as Figure 9 As shown, Figure 9 Schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device 1100 includes a processor 1101 having one or more processing cores, a memory 1102 having one or more computer-readable storage media, and a computer program stored in the memory 1102 and executable on the processor. The processor 1101 is electrically connected to the memory 1102. Those skilled in the art will understand that the vehicle structure shown in the figure does not constitute a limitation of the vehicle, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.
[0347] The processor 1101 is the control center of the electronic device 1100. It connects the various parts of the entire electronic device 1100 using various interfaces and lines. By running or loading software programs and / or units stored in the memory 1102 and calling data stored in the memory 1102, it executes various functions of the electronic device 1100 and processes data, thereby monitoring the electronic device 1100 as a whole. The processor 1101 can be a processor (Central Processing Unit, CPU), a graphics processing unit (GPU), a network processor (Network Processor, NP), etc., and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of this application.
[0348] In the embodiment of the present application, the processor 1101 in the electronic device 1100 loads instructions corresponding to one or more application processes into the memory 1102 according to the following steps, and the processor 1101 runs the application stored in the memory 1102 to implement various functions, such as:
[0349] Processing the target video to obtain content information and timing information of multiple video frames of the target video;
[0350] Emotion recognition is performed on the target video according to the content information and timing information of the multiple video frames to obtain emotion information of the video frames.
[0351] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0352] Optional, such as Figure 9 As shown, the electronic device 1100 further includes: a touch screen 1103, a radio frequency circuit 1104, an audio circuit 1105, an input unit 1106, and a power supply 1107. Among them, the processor 1101 is electrically connected to the touch screen 1103, the radio frequency circuit 1104, the audio circuit 1105, the input unit 1106, and the power supply 1107 respectively. Those skilled in the art will understand that Figure 9 The vehicle structure shown in the figure does not constitute a limitation to the vehicle, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0353] The touch display screen 1103 can be used to display a graphical user interface and receive operation instructions generated by the user acting on the graphical user interface. The touch display screen 1103 may include a display panel and a touch panel. Among them, the display panel can be used to display information input by the user or information provided to the user and various graphical user interfaces of the vehicle, which can be composed of graphics, text, icons, videos and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. The touch panel can be used to collect user touch operations on or near it (such as operations performed by the user using any suitable object or accessory such as a finger, stylus, etc. on or near the touch panel), and generate corresponding operation instructions, and the operation instructions execute corresponding programs. Optionally, the touch panel may include two parts: a touch display system and a touch controller. Among them, the touch display system detects the user's touch direction, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch display system, converts it into touch point coordinates, and then sends it to the processor 1101, and can receive commands sent by the processor 1101 and execute them. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it is transmitted to the processor 1101 to determine the type of touch event. The processor 1101 then provides a corresponding visual output on the display panel according to the type of touch event. In an embodiment of the present application, the touch panel and the display panel can be integrated into the touch display screen 1103 to realize input and output functions. However, in some embodiments, the touch panel and the touch panel can be used as two independent components to realize input and output functions. That is, the touch display screen 1103 can also be used as part of the input unit 1106 to realize the input function.
[0354] The RF circuit 1104 may be used to transmit and receive RF signals, thereby establishing wireless communication with network devices or other vehicles through wireless communication, and transmitting and receiving signals with network devices or other vehicles.
[0355] Audio circuit 1105 can be used to provide an audio interface between the user and the vehicle through a speaker and microphone. Audio circuit 1105 converts received audio data into electrical signals and transmits them to the speaker, which then converts them into sound signals for output. The microphone, on the other hand, converts collected sound signals into electrical signals, which are received by audio circuit 1105 and converted into audio data. This audio data is then output to processor 1101 for processing, then transmitted via RF circuit 1104 to, for example, another vehicle, or to memory 1102 for further processing. Audio circuit 1105 may also include an earphone jack to allow communication between an external headset and the vehicle.
[0356] The input unit 1106 may be configured to receive input digital, character information, or user feature information (such as fingerprint, iris, or facial information), and to generate keyboard, mouse, joystick, optical, or trackball signal input related to user settings and function control.
[0357] Power supply 1107 is used to supply power to various components of electronic device 1100. Optionally, power supply 1107 can be logically connected to processor 1101 via a power management device, thereby enabling the power management device to manage charging, discharging, and power consumption. Power supply 1107 can also include one or more DC or AC power supplies, a recharging device, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.
[0358] although Figure 9 Not shown, the electronic device 1100 may further include a camera, a sensor, a wireless fidelity module, a Bluetooth module, etc., which will not be described in detail here.
[0359] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0360] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0361] To this end, an embodiment of the present application provides a computer-readable storage medium storing a plurality of computer programs, which can be loaded by a processor to execute any of the emotion recognition methods provided in the embodiments of the present application. The computer program can execute the following steps of the emotion recognition method:
[0362] Processing the target video to obtain content information and timing information of multiple video frames of the target video;
[0363] Emotion recognition is performed on the target video according to the content information and timing information of the multiple video frames to obtain emotion information of the video frames.
[0364] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0365] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0366] Since the computer program stored in the computer-readable storage medium can execute any emotion recognition method provided in the embodiments of the present application, the beneficial effects that can be achieved by any emotion recognition method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0367] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0368] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0369] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0370] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0371] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0372] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.
[0373] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media such as modulated communication signals and carrier waves.
[0374] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more features. In the description of this application, "plurality" means two or more, unless otherwise specifically defined.
[0375] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0376] The embodiments, implementation methods and related technical features of the present application can be combined and replaced with each other without conflict.
[0377] The above are merely preferred embodiments of the present application and do not constitute any form of limitation to the present application. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application are still within the scope of the technical solution of the present application.
Claims
1. A video emotion recognition method, characterized in that: The method comprises: Processing the target video to obtain content information and timing information of multiple video frames of the target video; Emotion recognition is performed on the target video according to the content information and timing information of the multiple video frames to obtain emotion information of the video frames.
2. The emotion recognition method according to claim 1, characterized in that The processing of the target video to obtain content information and timing information of multiple video frames of the target video includes: Performing frame extraction processing on the target video to obtain multiple video frames, and obtaining content information of each video frame; and Generating timing information of the video frame according to the time stamp of the video frame.
3. The emotion recognition method according to claim 2, characterized in that The time series information includes a time series feature vector.
4. The emotion recognition method according to claim 3, characterized in that The temporal feature vector includes the time code corresponding to the video frame.
5. The emotion recognition method according to claim 2, characterized in that Generating the timing information of the video frame according to the timestamp of the video frame includes: The timing information of the video frame is generated according to the timestamp of the video frame and a preset first function.
6. The emotion recognition method according to claim 5, characterized in that The first function includes a first periodic function and a second periodic function. The vector elements at even positions in the timing information are generated based on the first periodic function and the timestamp. The vector elements at odd positions in the timing information are generated based on the second periodic function and the timestamp.
7. The emotion recognition method according to claim 2, characterized in that Generating the timing information of the video frame according to the timestamp of the video frame includes: The timing information of the video frame is generated according to the timestamp of the video frame and a preset second function, where the second function is set based on a learnable parameter.
8. The emotion recognition method according to claim 7, characterized in that The second function includes a first sub-function set based on the learnable parameter and a second sub-function set based on the first sub-function; a portion of the continuous vector elements in the timing information are generated based on the first sub-function and the timestamp, and the remaining vector elements are generated based on the second sub-function and the timestamp.
9. The emotion recognition method according to claim 8, characterized in that The first sub-function is a linear function, and the second sub-function is a periodic function based on the linear function.
10. The emotion recognition method according to claim 8, characterized in that The vector dimension of the time series information is D, and the vector elements of the first X dimensions in the time series information are generated by a first sub-function, where X is a maximum integer not exceeding D / 2.
11. The emotion recognition method according to claim 2, characterized in that: The step of extracting a frame from the target video to obtain a plurality of video frames and obtaining content information of each video frame includes: Performing frame extraction processing on the target video according to a preset frequency to obtain multiple video frames; Feature extraction is performed on content data of at least one modality of the video frame to obtain content information.
12. The emotion recognition method according to claim 11, characterized in that: The content information includes a content feature vector.
13. The emotion recognition method according to claim 12, characterized in that: The content feature vector of the video frame includes an image feature vector and / or an audio feature vector.
14. The emotion recognition method according to claim 1, characterized in that After performing emotion recognition on the target video according to the content information and timing information of the plurality of video frames to obtain the emotion information of the video frames, the method further includes: The emotional information of the video frame is smoothed to obtain target emotional information of the video frame.
15. The emotion recognition method according to claim 14, characterized in that: The smoothing process of the emotional information of the video frame to obtain target emotional information includes: Based on the emotion information of the reference frame of the video frame, the emotion information of the video frame is smoothed to obtain target emotion information.
16. The emotion recognition method according to claim 15, characterized in that: The step of smoothing the emotional information of the video frame based on the emotional information of the reference frame of the video frame to obtain target emotional information includes: The target emotional information of the video frame is obtained by performing weighted summation based on the emotional information of the reference frame of the video frame and the first weight, and the emotional information of the video frame and the second weight.
17. The emotion recognition method according to claim 14, characterized in that: The smoothing process of the emotional information of the video frame to obtain target emotional information of the video frame includes: When the difference in emotional information between the video frame and a reference frame of the video frame is less than a preset smoothing threshold, the emotional information of the video frame is smoothed to obtain target emotional information of the video frame.
18. The emotion recognition method according to claim 17, characterized in that: The method further comprises: When a difference in emotion information between the video frame and a reference frame of the video frame is not less than a preset smoothing threshold, the emotion information of the video frame is used as target emotion information of the video frame.
19. The emotion recognition method according to any one of claims 1 to 18, characterized in that: The performing emotion recognition on the target video according to the content information and timing information of the plurality of video frames to obtain the emotion information of the video frames includes: For each video frame of the target video, the content information and the timing information corresponding to the video frame are fused to obtain a comprehensive representation vector; According to the comprehensive representation vector, emotional information corresponding to the video frame is obtained.
20. The emotion recognition method according to claim 19, characterized in that: The fusing of the content information and the timing information corresponding to the video frame to obtain a comprehensive representation vector includes: The content information and timing information corresponding to the video frames are vector-concatenated to obtain a comprehensive representation vector.
21. The emotion recognition method according to claim 19, characterized in that The acquiring, according to the comprehensive representation vector, the emotion information corresponding to the video frame includes: The emotion information corresponding to the video frame is obtained based on the comprehensive representation vector using a preset emotion prediction model.
22. The emotion recognition method according to claim 21, characterized in that: The training steps of the emotion prediction model include: Identifying predicted emotional information of the sample video frame based on a sample comprehensive representation vector of the sample video frame by using a to-be-trained emotion prediction model, wherein the sample comprehensive representation vector is obtained by fusing content information and timing information of the sample video frame; The emotion prediction model is trained based on the sample weights corresponding to the sample video frames and the predicted emotion information.
23. The emotion recognition method according to claim 22, characterized in that: The training of the emotion prediction model based on the sample weights corresponding to the sample video frames and the predicted emotion information includes: Calculating the loss of the emotion prediction model based on the sample weights corresponding to the sample video frames and the predicted emotion information; Parameters of the sentiment prediction model are adjusted based on the loss.
24. The emotion recognition method according to claim 22, characterized in that: The emotion prediction model includes a machine learning model or a deep learning model.
25. The emotion recognition method according to claim 24, characterized in that: For the deep learning model, identifying predicted emotion information of the sample video frame based on the sample comprehensive representation vector of the sample video frame by the emotion prediction model to be trained includes: Inputting the comprehensive representation vectors of the sample video frame and other sample video frames belonging to the same time window as the sample video frame into the emotion prediction model; The emotion prediction model obtains emotion information corresponding to the sample video frame based on the input comprehensive representation vector.
26. The emotion recognition method according to claim 23, characterized in that The calculating the loss of the emotion prediction model based on the sample weights corresponding to the sample video frames and the predicted emotion information includes: Based on the predicted emotion information and emotion label of each sample video frame, obtaining emotion difference information of the sample video frame; Calculate the sample error of each sample video frame based on the weight loss and sentiment difference information corresponding to each sample video frame; The loss of the emotion prediction model is determined based on the sample error of each sample video frame.
27. The emotion recognition method according to claim 22, characterized in that: The step of obtaining the sample weight corresponding to the sample video frame includes: According to the emotion label corresponding to the sample video frame, a sample weight corresponding to the sample video frame is obtained.
28. The emotion recognition method according to claim 27, characterized in that: The emotion label includes emotion information of multiple emotion representation dimensions, and obtaining the sample weight corresponding to the sample video frame according to the emotion label corresponding to the sample video frame includes: For each sample video frame, bucket processing is performed on the sample video frame according to the emotional information of multiple emotion representation dimensions to obtain the target bucket coordinates to which the sample video frame belongs; A sample weight corresponding to the sample video frame is determined according to the target bucket coordinates of the sample video frame.
29. The emotion recognition method according to claim 28, characterized in that: The determining, according to the target bucket coordinates of the sample video frame, a sample weight corresponding to the sample video frame includes: Based on the target bucket coordinates of the sample video frame, obtaining the emotion label distribution density of the sample video frame; According to the emotion tag distribution density and the target bucket coordinates, a sample weight corresponding to the sample video frame is obtained.
30. The emotion recognition method according to claim 29, characterized in that The obtaining, according to the emotion label distribution information and the target bucket coordinates, a sample weight corresponding to the sample video frame, includes: According to the density information corresponding to the target bucket coordinates in the emotion tag distribution density, a sample weight corresponding to the sample video frame is obtained.
31. The emotion recognition method according to claim 30, characterized in that: The obtaining, according to density information corresponding to the target bucket coordinates in the emotion tag distribution density, a sample weight corresponding to the sample video frame includes: The density information corresponding to the target bucket coordinates in the emotion tag distribution density is squared and inverted to obtain a sample weight corresponding to the sample video frame.
32. The emotion recognition method according to claim 30, characterized in that: Before obtaining the sample weight corresponding to the sample video frame according to the density information corresponding to the target bucket coordinates in the emotion tag distribution density, the method further includes: The emotion tag distribution density is optimized to obtain an optimized emotion tag distribution density.
33. The emotion recognition method according to claim 32, characterized in that: The step of optimizing the emotion tag distribution density to obtain an optimized emotion tag distribution density includes: The emotion tag distribution density is optimized by using a binary Gaussian kernel density estimation method to obtain an optimized emotion tag distribution density.
34. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to perform the method according to any one of claims 1 to 33.
35. A computer-readable storage medium, characterized in that It includes a computer program, which is used to make the electronic device perform any one of the methods described in claims 1 to 33 when the computer program is run on the electronic device.
36. A computer program product, characterized in that The method comprises a computer program stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device executes any one of the methods described in claims 1 to 33.
Citation Information
Cited By
Method and device for training expression recognition model and electronic equipment
CN121686543A
Method, device and electronic equipment for training an expression recognition model
CN121686543B