Emotional Prediction Method, Device, Equipment and Readable Storage Medium for Videos

By extracting long-term action features and sound features in video data and integrating them, the problem that the prior art is difficult to predict emotional changes in long videos is solved, and higher emotional prediction accuracy is achieved.

CN114067241BActive Publication Date: 2025-05-27GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111294845.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-03
Publication Date
2025-05-27
Estimated Expiration
2041-11-03

AI Technical Summary

Technical Problem

The prior art is difficult to effectively predict emotional changes in long videos because existing methods are mainly applicable to short videos and cannot adapt to the characteristics of viewer's emotions changing over time.

Method used

By obtaining image frame sequences and audio data in video data, the encoding network and recurrent neural network extract the action and sound features and fuse these features for emotional prediction. The specific steps include: extracting long-term action features from the image frame sequence using the first coding network and the first cyclic neural network, extracting long-term sound features from the audio data using the second coding network and the second cyclic neural network, and fusing the two for emotional prediction.

Benefits of technology

This method can effectively improve the accuracy of video emotional prediction, especially when processing long videos, which can better retain useful information, thereby more accurately reflecting the emotional changes of the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114067241B_ABST
    Figure CN114067241B_ABST
Patent Text Reader

Abstract

The present application discloses an emotion prediction method, apparatus, device, and readable storage medium for videos. The method includes: obtaining video data to be processed, where the video data includes an image frame sequence and audio data; using a first encoding network to extract a first action feature vector from the image frame sequence, and using a first recurrent neural network to extract a second action feature vector from the first action feature vector, the video duration corresponding to the first action feature vector being shorter than the video duration of the second action feature vector; using a second encoding network to extract a first sound feature vector from the audio data, and using a second recurrent neural network to extract a second sound feature vector from the first sound feature vector, the video duration corresponding to the first sound feature vector being shorter than the video duration corresponding to the second sound feature vector; fusing the second action feature vector and the second sound feature vector to obtain a fused feature; and performing emotion prediction based on the fused feature. Through the above method, the present application can improve the accuracy of emotion prediction for videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing, and in particular, to a method, apparatus, device, and readable storage medium for predicting the emotion of a video. Background Art

[0002] Through long-term research, it is found that generally, predicting the emotion of a video is to predict the emotion classification of the entire video, which makes most of the existing technologies limited to the processing of short videos.

[0003] For the case of long videos, the emotions of viewers will change over time. At this time, it is obviously unreasonable to classify the emotion of the entire video. Summary of the Invention

[0004] This application mainly provides a method, apparatus, device, and readable storage medium for predicting the emotion of a video.

[0005] In a first aspect of this application, a method for predicting the emotion of a video is provided, including: obtaining video data to be processed; where the video data includes an image frame sequence and corresponding audio data; using a first encoding network to extract a first action feature vector from the image frame sequence, and using a first recurrent neural network to extract a second action feature vector from the first action feature vector; using a second encoding network to extract a first sound feature vector from the audio data, and using a second recurrent neural network to extract a second sound feature vector from the first sound feature vector; fusing the second action feature vector and the second sound feature vector to obtain a fused feature; and performing emotion prediction on the video data based on the fused feature.

[0006] In a second aspect of this application, a video emotion prediction apparatus is provided, including: an obtaining module, configured to obtain video data to be processed; where the video data includes an image frame sequence and corresponding audio data; an action feature extraction module, configured to use a first encoding network to perform feature extraction on the image frame sequence to obtain a first action feature vector, and use a first recurrent neural network to perform feature extraction on the first action feature vector to obtain a second action feature vector, where the video duration corresponding to the first action feature vector is shorter than the video duration corresponding to the second action feature vector; a sound feature extraction module, configured to use a second encoding network to perform feature extraction on the audio data to obtain a first sound feature vector, and use a second recurrent neural network to perform feature extraction on the first sound feature vector to obtain a second sound feature vector, where the video duration corresponding to the first sound feature vector is shorter than the video duration corresponding to the second sound feature vector; a feature fusion module, configured to fuse the second action feature vector and the second sound feature vector to obtain a fused feature; and an emotion prediction module, configured to perform emotion prediction on the video data based on the fused feature.

[0007] In a third aspect of the present application, an electronic device is provided, including a processor and a memory coupled to each other. A computer program capable of running on the processor is stored in the memory. Wherein, when the processor is used to run the computer program, the emotional prediction method of the video provided in the first aspect above is implemented.

[0008] In a fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores program data, and when the program data is executed by a processor, the emotional prediction method of the video provided in the first aspect above is implemented.

[0009] The beneficial effects of the present application are as follows: Different from the prior art, the present application uses a first encoding network to extract a first action feature vector of an image frame sequence, then uses a first recurrent neural network to extract a second action feature vector from the first action feature vector, uses a second encoding network to extract a first sound feature vector from audio data, and uses a second recurrent neural network to extract a second sound feature vector from the first sound feature vector; the second action feature vector and the second sound feature vector are fused to obtain a fusion feature; emotional prediction is performed on the video data based on the fusion feature. Both the second action feature vector and the second sound feature vector obtained by the above method are long-term features, retaining more useful information. When applied to the emotional prediction level, the accuracy of the emotional prediction result can be effectively improved. Description of the Drawings

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0011] Figure 1 It is a structural schematic block diagram of an embodiment of the electronic device of the present application;

[0012] Figure 2 It is a flowchart of an embodiment of the emotional prediction method of the video of the present application;

[0013] Figure 3 It is a flowchart of an embodiment of step S12 of the present application;

[0014] Figure 4 It is a flowchart of an embodiment of step S13 of the present application;

[0015] Figure 5 It is a flowchart of an embodiment of step S14 of the present application;

[0016] Figure 6It is a flowchart showing an embodiment of the training of the first encoding network and the first recurrent neural network in the present application;

[0017] Figure 7 It is a flowchart showing an embodiment of the training of the second encoding network and the second recurrent neural network in the present application;

[0018] Figure 8 It is a structural schematic diagram of an embodiment of the video emotion prediction network in the present application;

[0019] Figure 9 It is a flowchart showing an embodiment of the training of the regression layer in the present application;

[0020] Figure 10 It is a structural schematic diagram of an embodiment of the emotion prediction device for the video in the present application;

[0021] Figure 11 It is a structural schematic diagram of an embodiment of the computer-readable storage medium in the present application. Detailed implementation manners

[0022] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present application belong to the scope of protection of the present application.

[0023] Referring to "embodiment" herein means that the specific features, structures or characteristics described in conjunction with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various positions in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0024] Please refer to Figure 1 , Figure 1 It is a structural schematic diagram of an embodiment of the electronic device in the present application. The electronic device 100 includes a processor 101 and a memory 102 that are coupled to each other. A computer program capable of running on the processor 101 is stored in the memory 102. Among them, when the processor 101 is used to execute the computer program, the emotion prediction method for the video described in the following embodiments is implemented.

[0025] The memory 102 can be used to store program data and modules. The processor 101 executes various functional applications and data processing by running the program data and modules stored in the memory 102. The memory 102 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the electronic device 100 (such as video data, an image frame sequence, audio data, etc.). In addition, the memory 102 can include high-speed random access memory and can also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 102 can also include a memory controller to provide the processor 101 with access to the memory 102.

[0026] In some specific embodiments, the electronic device 100 is not limited to including a television, a desktop computer, a laptop computer, a handheld computer, a wearable device, a notebook computer.

[0027] For the description of each step of the processing execution, please refer to the description of each step of the emotional prediction method embodiment of the video of the present application below, and details will not be repeated here.

[0028] Please refer to Figure 2 , Figure 2 is a flowchart of an embodiment of the emotional prediction method of the video of the present application. This embodiment includes the following steps:

[0029] Step S11: Obtain video data to be processed; wherein, the video data includes an image frame sequence and corresponding audio data.

[0030] The video to be processed can be obtained from a server through a network, or can be obtained from a storage device such as a USB flash drive or a hard disk through a physical connection module, or can also be obtained by the device main body that issues the processing operation through its own camera function.

[0031] The video to be processed includes an image frame sequence and audio data. The audio data corresponds to the image frame sequence and is the content of the same video data.

[0032] Among them, information such as the duration and resolution of the video to be processed can be set as needed according to the actual network processing performance. The video to be processed can be in formats such as wmv, rmvb, mkv, mp4, etc., and no limitations are imposed here.

[0033] Step S12: Extract a first action feature vector from the image frame sequence using a first encoding network, and extract a second action feature vector from the first action feature vector using a first recurrent neural network.

[0034] Among them, the video duration corresponding to the first action feature vector is shorter than the video duration corresponding to the second action feature vector. That is, the first action feature vector is a short-term action feature, containing information of fewer frames, and the second action feature vector is a long-term action feature, retaining more temporal information in the image frames.

[0035] Among them, the first encoding network is a three-dimensional convolutional neural network or a three-dimensional residual neural network, which can extract action information in addition to image pixel information and retain the correlation information between image frames.

[0036] The first recurrent neural network is, for example, an RNN (Recurrent Neural Network) or an LSTM (long short-term memory).

[0037] Please refer to Figure 3 , Figure 3 which is a flowchart of an embodiment of step S12 of this application. This embodiment includes the following steps:

[0038] Step S121: Segment the image frame sequence to obtain a plurality of frame segments, where each frame segment includes at least two image frames.

[0039] Among them, the segmentation method may be to segment the image frame sequence into a plurality of equally long frame segments. The number of image frames included in each frame segment is greater than or equal to 2, and can be specifically set by comprehensively considering the performance of the network and the total length of the image frame sequence.

[0040] Step S122: Input the frame segment into the first encoding network to obtain a first action feature vector corresponding to the frame segment.

[0041] In this step, a plurality of temporally consecutive frame segments can be input into the first encoding network to output the first action feature vector corresponding to each frame segment.

[0042] Step S123: Input a plurality of first action feature vectors into the first recurrent neural network to obtain a second action feature vector.

[0043] Among them, the plurality of first action feature vectors may specifically include at least 2 first action feature vectors.

[0044] Optionally, the plurality of first action feature vectors correspond to a plurality of temporally consecutive frame segments. Specifically, before each feature extraction operation in step S122, a preset number of consecutive frame segments are selected by using a sliding window, and the selected frame segments are respectively input into the first encoding network to respectively extract the first action feature vectors.

[0045] Among them, the number of selected frame segments can be determined by setting the sliding window parameters.

[0046] The second action feature vector of this embodiment is obtained by processing multiple first action feature vectors, and it contains the action feature information of multiple consecutive frame segments, resulting in higher accuracy for subsequent prediction.

[0047] Step S13: Use the second encoding network to extract the first sound feature vector from the audio data, and use the second recurrent neural network to extract the second sound feature vector from the first sound feature vector.

[0048] Among them, the video duration corresponding to the first sound feature vector is shorter than the video duration corresponding to the second sound feature vector. That is to say, the first sound feature vector is a short-term sound feature, containing the shorter time information in the audio data, and the second sound feature vector is a long-term sound feature, retaining more time information in the audio data.

[0049] Please refer to Figure 4 , Figure 4 which is a flowchart of an embodiment of step S13 of this application. This embodiment specifically includes the following steps:

[0050] Step S131: Corresponding to the segmentation method of the image frame sequence, segment the audio to obtain multiple audio segments.

[0051] Please refer to the segmentation method of the image frame sequence in step S121 of the previous embodiment. The segmentation method of the audio data in this step corresponds to the segmentation method of the image frame sequence, and the audio data is segmented into multiple audio segments, so that each audio segment has a corresponding frame segment, and after extracting the sound features of each audio segment, there is a corresponding action feature.

[0052] Step S132: Input the audio segment into the second encoding network to obtain the first sound feature vector corresponding to the audio segment.

[0053] In this step, multiple temporally consecutive audio segments can be input into the first encoding network to output the first sound feature vector corresponding to each audio segment.

[0054] Optionally, before extracting the first sound feature vector from the audio segment in this step, the Mel spectrogram of the audio segment can be extracted in advance, and the Mel spectrogram can be used as the representation of the audio segment and input into the second encoding network for extracting the first sound feature vector.

[0055] Specifically, a sound signal is originally a one-dimensional time-domain signal, and it is difficult to intuitively see the frequency variation law. A long signal is framed and windowed, then the Fourier transform (FFT) is performed on each frame, and finally the results of each frame are stacked along another dimension to obtain a two-dimensional signal form similar to a picture, that is, a spectrogram. The Mel spectrum is obtained by transforming the spectrogram through a mel-scale filter bank to obtain sound features of a suitable size.

[0056] Step S133: Input multiple first sound feature vectors into a second recurrent neural network to obtain a second sound feature vector.

[0057] Among them, the multiple first sound feature vectors may specifically include at least 2 first sound feature vectors.

[0058] Optionally, the multiple first sound feature vectors correspond to multiple temporally continuous audio segments. Specifically, before each feature extraction operation in step S132, a preset number of continuous audio segments are selected by using a sliding window, and the selected audio segments are respectively input into the first encoding network to respectively extract first sound feature vectors.

[0059] Among them, the number of selected audio segments can be determined by setting the sliding window parameters.

[0060] The second sound feature vector in this embodiment is obtained by processing multiple first sound feature vectors, and it contains the sound feature information of multiple continuous audio segments, and the subsequent prediction accuracy is higher.

[0061] Step S14: Fuse the second action feature vector and the second sound feature vector to obtain a fused feature.

[0062] The fused feature contains features in both aspects of the image and the audio, and can more comprehensively and accurately represent the characteristics of the video, thereby improving the accuracy of video emotion prediction.

[0063] Please refer to Figure 5 , Figure 5 which is a flowchart of an embodiment of step S14 of this application. The specific steps for fusing the second action feature vector and the second sound feature vector in this embodiment may include the following steps:

[0064] Step S141: Perform pooling processing on the second action feature vector and the second sound feature vector respectively to make the second action feature vector and the second sound feature vector of the same dimension.

[0065] Step S142: Concatenate the second action feature vector and the second sound feature vector after pooling processing to obtain a fused feature.

[0066] In the above method, the second action feature vector and the second voice feature vector are concatenated to obtain a fusion feature. In another embodiment, the attention method can also be used to perform weighted fusion on the second action feature vector and the second voice feature vector to obtain a fusion feature.

[0067] Step S15: Perform emotion prediction on the video data based on the fusion feature.

[0068] In this step, the fusion feature is input into an emotion prediction network to obtain the emotion prediction result of the corresponding video segment.

[0069] The above embodiments can perform segmented emotion prediction on video data using the second action feature vector and the second voice feature vector. On the one hand, long-term features are beneficial to improving the accuracy of the prediction result. On the other hand, compared with the method of performing overall emotion prediction on video data, segmented prediction of the video has higher accuracy, and multiple emotion values represent the emotion trend of the overall video more clearly, which is beneficial to the further processing of the video.

[0070] Please refer to Figure 6 , Figure 6 which is a flowchart showing an embodiment of the training of the first encoding network and the first recurrent neural network in the present application. This embodiment may include the following steps:

[0071] Step S21: Connect the first encoding network to the third recurrent neural network, and perform self-supervised training on the first encoding network and the third recurrent neural network using an unlabeled image frame dataset, where the third recurrent neural network is used to predict the next first action feature vector of the first encoding network based on the output result of the current first action feature vector of the first encoding network.

[0072] Predicting the next first action feature vector above means obtaining the predicted value of the first action feature vector corresponding to the next frame segment. At the same time, the first encoding network can extract the first action feature vector of the next frame segment, compare the predicted value with the first action feature vector extracted by the first encoding network to obtain a loss, and then adjust and optimize the parameters of the first encoding network according to the loss.

[0073] In this step, self-supervised training is performed on the first encoding network using an unannotated image frame dataset, which can greatly reduce the cost of data annotation and greatly expand the number of available datasets.

[0074] Step S22: Remove the third recurrent neural network and connect the first encoding network to the first recurrent neural network.

[0075] Step S21: Adjust the parameters of the first encoding network, that is, complete the self-supervised training of the first encoding network to obtain a first encoding network with good performance. Connect the first encoding network to the first recurrent neural network to facilitate the training of the first recurrent neural network.

[0076] Step S23: With the parameters of the first encoding network fixed, use the labeled image frame dataset to train the first encoding network and the first recurrent neural network to adjust the parameters of the first recurrent neural network.

[0077] Among them, the first recurrent neural network performs emotion prediction based on the first action feature vector output by the first encoding network. The labels in the image frame dataset are added according to the emotions of the viewers.

[0078] In one implementation, the first recurrent neural network includes an emotion prediction regression layer for performing emotion prediction. The emotion prediction result is a confidence score corresponding to several emotion categories. The confidence score is between 0 and 1, and the larger the value, the stronger the emotion corresponding to the category.

[0079] Specifically, in this step, the labeled image frames are input into the first encoding network to obtain multiple first action feature vectors. The first recurrent neural network obtains a second action feature vector based on the first action feature vector, predicts the next first action feature vector according to the second action feature vector, calculates the loss according to the prediction result and the label, and continuously adjusts the parameters of the first recurrent neural network according to the loss.

[0080] Step S24: Remove the emotion prediction regression layer of the first recurrent neural network to use the output result of the last layer of the remaining first recurrent neural network as the second action feature vector.

[0081] Specifically, the emotion prediction regression layer is used to perform emotion prediction based on long-term features, and then use the prediction result to adjust the network parameters to complete the training of the first recurrent neural network. The joint network structure of the first encoding network and the first recurrent neural network obtained in this way has good performance in extracting the second action feature vector. After the first recurrent neural network is trained, its role is to output the second action feature vector, and the emotion prediction regression layer is no longer used and is removed.

[0082] Among them, the first encoding network is a ResNet-3D network, and the first recurrent neural network is an LSTM network. The first encoding network and the first recurrent neural network form an E3D-LSTM network (i.e., Eidetic 3D LSTM), which has excellent long-term memory performance and better perception of long-distance information. Therefore, the emotion prediction for videos is also more accurate.

[0083] Please refer toFigure 7 , Figure 7 is a flowchart showing the process of training the second encoding network and the second recurrent neural network in an embodiment of the present application. This embodiment may include the following steps:

[0084] Step S31: Connect the second encoding network to the fourth recurrent neural network, and perform self-supervised training on the second encoding network and the fourth recurrent neural network using an unlabeled audio dataset, where the fourth recurrent neural network is used to predict the next first sound feature vector of the second encoding network based on the output result of the current first sound feature vector of the second encoding network.

[0085] The above prediction of the next first sound feature vector, that is, obtaining the predicted value of the first sound feature vector corresponding to the next audio segment. At the same time, the second encoding network can extract the first sound feature vector of the next sound segment, compare the predicted value with the first sound feature vector extracted by the second encoding network to obtain a loss, and then adjust and optimize the parameters of the second encoding network according to the loss.

[0086] In this step, the second encoding network is self-supervised trained using an unannotated image frame dataset, which can greatly reduce the cost of data annotation and greatly expand the number of available datasets.

[0087] Step S32: Remove the fourth recurrent neural network and connect the second encoding network to the second recurrent neural network.

[0088] In the above manner, the parameters of the second encoding network are adjusted in step S31, that is, the self-supervised training of the second encoding network is completed, and a second encoding network with good performance is obtained. Connecting the second encoding network to the second recurrent neural network facilitates the training of the second recurrent neural network.

[0089] Step S33: With the parameters of the second encoding network fixed, use a labeled audio dataset to train the second encoding network and the second recurrent neural network to adjust the parameters of the second recurrent neural network.

[0090] Among them, the second recurrent neural network performs emotion prediction based on the first sound feature vector output by the second encoding network. The labels in the audio dataset are added according to the emotions of the viewers.

[0091] Optionally, the second recurrent neural network also includes an emotion prediction regression layer for performing emotion prediction. The emotion prediction result is a confidence score corresponding to several emotion categories. The confidence score is between 0 and 1, and the larger the value, the stronger the emotion corresponding to the category.

[0092] Specifically, in this step, the image frame with labels is input into the second encoding network to obtain multiple first sound feature vectors. The second recurrent neural network obtains a second sound feature vector based on the first sound feature vectors, predicts the next first sound feature vector according to the second sound feature vector, calculates the loss based on the prediction result and the labels, and continuously adjusts the parameters of the second recurrent neural network according to the loss.

[0093] Step S34: Remove the emotion prediction regression layer of the second recurrent neural network, and use the output result of the last layer of the retained second recurrent neural network as the second sound feature vector.

[0094] Specifically, the emotion prediction regression layer of the second recurrent neural network is used to perform emotion prediction based on the second sound feature vector, and then use the prediction result to adjust the network parameters to complete the training of the second recurrent neural network. In this way, the combined network structure of the second encoding network and the second recurrent neural network has good performance in extracting the second sound feature vector. After the training of the second recurrent neural network is completed, its function is to output the second sound feature vector, and the emotion prediction regression layer is no longer used and is removed.

[0095] Among them, the second encoding network can be a 3D residual network, and the second recurrent neural network can be an LSTM network. In this way, the second encoding network and the second recurrent neural network form an E3D-LSTM network (i.e., Eidetic 3D LSTM), which has excellent long-term memory performance, better perception of long-distance signals, and advantages in extracting the second sound feature vector. Therefore, the emotion prediction of the video is also more accurate.

[0096] Please refer to Figure 8 , Figure 8 which is a schematic block diagram of the structure of an embodiment of the video emotion prediction network of the present application. This embodiment trains the regression layer based on the Figure 8 shown video emotion prediction network. Figure 8 The shown emotion prediction network includes a first encoding network 10, a second encoding network 20, a first recurrent neural network 30, a second recurrent neural network 40, a feature fusion layer 50, and an emotion prediction regression network 60. The first encoding network 10, the second encoding network 20, the first recurrent neural network 30, and the second recurrent neural network 40 are all trained and the network parameters are fixed. Among them, the first encoding network 10 is connected to the first recurrent neural network 30, the second encoding network 20 is connected to the second recurrent neural network 40, the output layers of the first recurrent neural network 30 and the second recurrent neural network 40 are both connected to the feature fusion layer 50, and the output end of the feature fusion layer is connected to the emotion prediction regression network 60.

[0097] Please refer to Figure 9 , Figure 9It is a flowchart showing an embodiment of training the regression layer in the present application. This embodiment may include the following steps:

[0098] Step S41: After the parameters of the first encoding network 10, the second encoding network 20, the first recurrent neural network 30, and the second recurrent neural network 40 are fixed, connect the first encoding network 10 and the first recurrent neural network 30, and connect the second encoding network 20 and the second recurrent neural network 40.

[0099] Step S42: Input the image frame sequence and audio data of the video data with labels into the first encoding network 10 and the second encoding network 20 respectively. The first short-term feature output by the first encoding network 10 is used as the input of the first recurrent neural network 30, and the first recurrent neural network 30 outputs a first long-term feature according to the first short-term feature. The second short-term feature output by the second encoding network 20 is used as the input of the second recurrent neural network 40, and the second recurrent neural network 40 outputs a second long-term feature according to the second short-term feature.

[0100] Among them, both the first long-term feature and the first short-term feature are action features, and both the second long-term feature and the second short-term feature are sound features.

[0101] Step S43: Fuse the first long-term feature and the second long-term feature to obtain a fused video feature.

[0102] For this step, please refer to the feature fusion method in step S14 of the foregoing embodiment. Use the feature fusion layer 50 to fuse the first long-term feature and the second long-term feature to obtain a fused video feature, which will not be elaborated here.

[0103] Step S44: Input the fused video feature into the emotion prediction regression network 60 to obtain an emotion prediction result. According to the emotion prediction result and the label, adjust the parameters of the emotion prediction regression network 60.

[0104] Continuously adjust the corresponding parameters in the emotion prediction regression network 60 according to the difference between the emotion prediction result and the label, so as to gradually improve the accuracy of the emotion prediction result of the emotion prediction regression network 60 for the video matching the emotion label of the video data, and give a confidence score corresponding to several emotion categories. The confidence score is between 0 and 1, and the larger the value, the stronger the emotion of the corresponding category.

[0105] After adjusting the parameters of the emotion prediction regression network 60, it can be applied Figure 4 The shown emotion prediction network to perform emotion prediction on the video. This network has good emotion prediction effect and high accuracy for long videos.

[0106] Beneficial effects:

[0107] 1. This solution uses self-supervised learning technology to train the first encoder and the second encoder, greatly reducing the cost of data annotation and significantly expanding the number of available datasets.

[0108] 2. This solution defines the video emotion prediction task as densely regressing the confidence levels of multiple emotion categories simultaneously. Compared with classifying the emotion categories of videos, the task in this solution is more suitable for processing long videos. Predicting the confidence levels of multiple emotion categories is more in line with the objective laws of human emotions than classifying videos into a single category.

[0109] 3. This solution uses a 3D residual network to extract short-term features of video segments, an E3D-LSTM network structure composed of a first encoding network and a first recurrent network to obtain long-term action feature vectors of videos, and a long-term feature extraction network composed of a second encoding network and a second recurrent network to obtain long-term sound feature vectors of videos. Compared with the method of using 2D convolution to extract pixel features of images, the method used in this solution can obtain more useful information and has more advantages in performance.

[0110] 4. This solution is applicable to all current categories of videos and is not restricted by video content. In addition, this solution can be easily deployed and applied without any additional wearable devices.

[0111] Please refer to Figure 10 , Figure 10 which is a structural schematic block diagram of an embodiment of the video emotion prediction device of this application. The video emotion prediction device 300 includes: an acquisition module 310, an action feature extraction module 320, a sound feature extraction module 330, a feature fusion module 340, and an emotion prediction module 350.

[0112] Among them, the acquisition module 310 is used to acquire video data to be processed; among them, the video data includes an image frame sequence and corresponding audio data.

[0113] Among them, the action feature extraction module 320 is used to extract features from the image frame sequence using the first encoding network to obtain a first action feature vector, and to extract features from the first action feature vector using the first recurrent neural network to obtain a second action feature vector, where the video duration corresponding to the first action feature vector is shorter than the video duration corresponding to the second action feature vector.

[0114] The sound feature extraction module 330 is used to extract features from the audio data using the second encoding network to obtain a first sound feature vector, and to extract features from the first sound feature vector using the second recurrent neural network to obtain a second sound feature vector, where the video duration corresponding to the first sound feature vector is shorter than the video duration corresponding to the second sound feature vector.

[0115] The feature fusion module 340 is used to fuse the second action feature vector and the second sound feature vector to obtain a fused feature.

[0116] The emotion prediction module 350 is used to perform emotion prediction on the video data based on the fused feature.

[0117] Among them, the action feature extraction module 320 can also be used to segment the image frame sequence to obtain a plurality of frame segments, where each frame segment includes at least two image frames; input the frame segments into the first encoding network to obtain a first action feature vector corresponding to the frame segments; and finally input the plurality of first action feature vectors into the first recurrent neural network to obtain a second action feature vector.

[0118] Among them, the sound feature extraction module 330 can also be used to segment the audio data to obtain a plurality of audio segments; input the audio segments into the second encoding network to obtain a first sound feature vector corresponding to the audio segments; and then input the plurality of first sound feature vectors into the second recurrent neural network to obtain a second sound feature vector.

[0119] Among them, the video emotion prediction device 300 may further include a training module (not shown in the figure). The training module is used to, when the parameters of the first encoding network are fixed, use the labeled image frame data set to train the first encoding network and the first recurrent neural network to adjust the parameters of the first recurrent neural network, where the first recurrent neural network performs emotion prediction based on the first action feature vector output by the first encoding network; remove the emotion prediction regression layer of the first recurrent neural network, and use the output result of the last layer of the remaining first recurrent neural network as the second action feature vector.

[0120] Among them, the training module can also be used to connect the first encoding network to the third recurrent neural network, and perform self-supervised training on the first encoding network and the third recurrent neural network using the unlabeled image frame data set, where the output result of the third recurrent neural network based on the current first action feature vector of the first encoding network is used to predict the next first action feature vector of the first encoding network; remove the third recurrent neural network, and connect the first encoding network to the first recurrent neural network.

[0121] Among them, the training module can also be used to, when the parameters of the second encoding network are fixed, use the labeled audio data set to train the second encoding network and the second recurrent neural network to adjust the parameters of the second recurrent neural network, where the second recurrent neural network performs emotion prediction based on the first sound feature vector output by the second encoding network; remove the emotion prediction regression layer of the second recurrent neural network, and use the output result of the last layer of the remaining second recurrent neural network as the second sound feature vector.

[0122] Among them, the training module can also be used to connect the second encoding network to the fourth recurrent neural network, and perform self-supervised training on the second encoding network and the fourth recurrent neural network using an unlabeled audio dataset, where the fourth recurrent neural network is used to predict the next first sound feature vector of the second encoding network based on the output result of the current first sound feature vector of the second encoding network; remove the fourth recurrent neural network, and connect the second encoding network to the second recurrent neural network.

[0123] Among them, the training module can also be used to output confidence scores corresponding to several emotion categories using the first recurrent neural network and the second recurrent neural network.

[0124] Among them, the feature fusion module 340 can also be used to perform pooling processing on the second action feature vector and the second sound feature vector respectively, so that the second action feature vector and the second sound feature vector are of the same dimension; splice the second action feature vector and the second sound feature vector after pooling processing to obtain a fused feature.

[0125] For the specific execution manners of the steps executed by each module, please refer to the descriptions of the steps in the above-mentioned embodiments of the video emotion prediction method of this application, and will not be elaborated here.

[0126] In some specific embodiments, the video emotion prediction device 300 is not limited to including a television, a desktop computer, a laptop computer, a handheld computer, a wearable device, and a notebook computer.

[0127] In the embodiments of the present application, the disclosed video emotion prediction method and electronic device can be implemented in other ways. For example, the above-described embodiments of the transportation device and the electronic device are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0128] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0129] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0130] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product, and this computer software product is stored in a storage medium.

[0131] Refer to Figure 11 , Figure 11 which is a schematic structural diagram of an embodiment of the computer-readable storage medium of the present application. The computer-readable storage medium 200 stores program data 210, and when the program data 210 is executed, the steps of each embodiment of the above-mentioned emotion prediction method for videos are implemented.

[0132] For the description of each step of the processing execution, please refer to the description of each step of the embodiment of the emotion prediction method for videos of the present application above, and details will not be repeated here.

[0133] The computer-readable storage medium 200 can be various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc.

[0134] The above are only the embodiments of the present application, and do not limit the patent scope of the present application accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A method for emotion prediction of a video, characterized in that, the method includes: Obtain video data to be processed; wherein, the video data includes an image frame sequence and corresponding audio data; Use a first encoding network to extract features from the image frame sequence to obtain a first action feature vector, and use a first recurrent neural network to extract features from the first action feature vector to obtain a second action feature vector, wherein the video duration corresponding to the first action feature vector is shorter than the video duration corresponding to the second action feature vector; Use a second encoding network to extract features from the audio data to obtain a first sound feature vector, and use a second recurrent neural network to extract features from the first sound feature vector to obtain a second sound feature vector, wherein the video duration corresponding to the first sound feature vector is shorter than the video duration corresponding to the second sound feature vector; Fuse the second action feature vector and the second sound feature vector to obtain a fused feature; Perform emotion prediction on the video data based on the fused feature; Before using the first encoding network to extract features from the image frame sequence to obtain a first action feature vector and using a first recurrent neural network to extract features from the first action feature vector to obtain a second action feature vector, the method further includes: With the parameters of the first encoding network fixed, use an image frame dataset with labels to train the first encoding network and the first recurrent neural network to adjust the parameters of the first recurrent neural network, wherein the first recurrent neural network performs emotion prediction based on the first action feature vector output by the first encoding network; Remove the emotion prediction regression layer of the first recurrent neural network, and use the output result of the last layer of the remaining first recurrent neural network as the second action feature vector.

2. The method according to claim 1, characterized in that, using the first encoding network to extract features from the image frame sequence to obtain a first action feature vector, and using a first recurrent neural network to extract features from the first action feature vector to obtain a second action feature vector, includes: Segment the image frame sequence to obtain a plurality of frame segments, wherein each frame segment includes at least two image frames; Input the frame segments into the first encoding network to obtain a first action feature vector corresponding to the frame segments; Input a plurality of the first action feature vectors into the first recurrent neural network to obtain the second action feature vector.

3. The method according to claim 1, characterized in that, using the second encoding network to extract features from the audio data to obtain a first sound feature vector, and using a second recurrent neural network to extract features from the first sound feature vector to obtain a second sound feature vector, includes: Segment the audio data to obtain a plurality of audio segments; Input the audio segments into the second encoding network to obtain a first sound feature vector corresponding to the audio segments; Input the multiple first voice feature vectors into a second recurrent neural network to obtain the second voice feature vector.

4. The method according to claim 1, wherein, before training the first encoding network and the first recurrent neural network by using the labeled image frame data set, the method further includes: connect the first encoding network to a third recurrent neural network, and perform self-supervised training on the first encoding network and the third recurrent neural network by using an unlabeled image frame data set, wherein the third recurrent neural network is used to predict the next first action feature vector of the first encoding network based on the output result of the current first action feature vector of the first encoding network; remove the third recurrent neural network, and connect the first encoding network to the first recurrent neural network.

5. The method according to claim 1, wherein, before using a second encoding network to extract features from the audio data to obtain a first voice feature vector, and using a second recurrent neural network to extract features from the first voice feature vector to obtain a second voice feature vector, it further includes: while the parameters of the second encoding network are fixed, train the second encoding network and the second recurrent neural network by using a labeled audio data set to adjust the parameters of the second recurrent neural network, wherein the second recurrent neural network performs emotion prediction based on the first voice feature vector output by the second encoding network; remove the emotion prediction regression layer of the second recurrent neural network, and use the output result of the last layer of the remaining second recurrent neural network as the second voice feature vector.

6. The method according to claim 5, wherein, before training the second encoding network and the second recurrent neural network by using the labeled audio data set, the method further includes: connect the second encoding network to a fourth recurrent neural network, and perform self-supervised training on the second encoding network and the fourth recurrent neural network by using an unlabeled audio data set, wherein the fourth recurrent neural network is used to predict the next first voice feature vector of the second encoding network based on the output result of the current first voice feature vector of the second encoding network; remove the fourth recurrent neural network, and connect the second encoding network to the second recurrent neural network.

7. The method according to claim 1 or 5, wherein, the emotion prediction results of the first recurrent neural network and the second recurrent neural network are confidence scores corresponding to several emotion categories.

8. The method according to any one of claims 1-6, wherein, the first encoding network is a ResNet-3D network, and the first recurrent neural network is an LSTM network.

9. The method according to claim 1, wherein, the step of fusing the second action feature vector and the second voice feature vector to obtain a fusion feature includes: Perform pooling processing on the second action feature vector and the second voice feature vector respectively, so that the second action feature vector and the second voice feature vector are of the same dimension; Concatenate the second action feature vector and the second voice feature vector after the pooling processing to obtain the fusion feature.

10. A video emotion prediction device, characterized in that, the device includes: An acquisition module, configured to acquire video data to be processed; wherein, the video data includes an image frame sequence and corresponding audio data; An action feature extraction module, configured to use a first encoding network to extract features from the image frame sequence to obtain a first action feature vector, and use a first recurrent neural network to extract features from the first action feature vector to obtain a second action feature vector, wherein, the video duration corresponding to the first action feature vector is shorter than the video duration corresponding to the second action feature vector; A voice feature extraction module, configured to use a second encoding network to extract features from the audio data to obtain a first voice feature vector, and use a second recurrent neural network to extract features from the first voice feature vector to obtain a second voice feature vector, wherein, the video duration corresponding to the first voice feature vector is shorter than the video duration corresponding to the second voice feature vector; A feature fusion module, configured to fuse the second action feature vector and the second voice feature vector to obtain a fusion feature; An emotion prediction module, configured to perform emotion prediction on the video data based on the fusion feature; A training module, before using the first encoding network to extract features from the image frame sequence to obtain a first action feature vector, and using a first recurrent neural network to extract features from the first action feature vector to obtain a second action feature vector, the training module is configured to, with the parameters of the first encoding network fixed, use an image frame data set with labels to train the first encoding network and the first recurrent neural network to adjust the parameters of the first recurrent neural network, wherein the first recurrent neural network performs emotion prediction based on the first action feature vector output by the first encoding network; remove the emotion prediction regression layer of the first recurrent neural network, and use the output result of the last layer of the remaining first recurrent neural network as the second action feature vector.

11. An electronic device, characterized in that, the electronic device includes a processor and a memory coupled to each other, and the memory stores a computer program that can run on the processor, wherein, when the processor runs the computer program, it executes the steps of the video emotion prediction method according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores program data, and when the program data is executed by a processor, it implements the steps of the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Attention fusion-based online short video multi-modal emotion recognition method

    CN111275085A

  • No-reference audio and video quality evaluation method based on gated recurrent neural network

    CN113473117A