A video character emotion recognition method, device, equipment and storage medium

By using hierarchical dynamic temporal fusion and gated temporal neural networks to process multimodal data features, the problems of excessive computational consumption and low recognition accuracy in existing technologies are solved, and efficient emotion recognition is achieved.

CN116824421BActive Publication Date: 2026-04-07CHINA MOBILE (SUZHOU) SOFTWARE TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-18
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing video emotion recognition methods rely on pre-trained models, which consume excessive computation, and multimodal fusion methods cannot deeply mine modal information, resulting in low recognition accuracy and failing to meet the needs of efficient emotion recognition.

Method used

The preprocessing module extracts data features from multiple modalities, and the hierarchical dynamic temporal fusion module and the gated temporal neural module are used to perform data feature fusion and computation. The results of the previous round of computation are combined to perform sentiment classification, which reduces computational overhead and improves recognition accuracy.

Benefits of technology

It improves the accuracy and efficiency of video character emotion recognition while reducing computational overhead, and is suitable for lightweight but high-performance multimodal emotion recognition tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116824421B_ABST
    Figure CN116824421B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and storage medium for recognizing emotions in video characters. The method includes: extracting first data features from each of multiple modalities of a target video using a preprocessing module; fusing the first data features of multiple modalities using a hierarchical dynamic temporal fusion module to obtain second data features; performing computational processing on the second and third data features using a gated temporal neural module to obtain fourth data features, wherein the third data feature is the data feature obtained by the gated temporal neural module through the previous round of computational processing; and classifying the fourth data feature using an emotion classification module to obtain the emotion type of the character in the target video. Thus, by fusing data features from multiple modalities and combining the results of the previous round of computational processing with the current input information for emotion recognition, the accuracy of emotion recognition can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and particularly relates to a video character emotion recognition method and device, equipment and a storage medium. BACKGROUND

[0002] Video character emotion recognition is a research hotspot in the field of artificial intelligence. In order to improve the accuracy of video character emotion recognition, the emotions of characters in a video can be recognized by combining multiple modalities, for example, the emotions of characters in a video can be recognized by combining the text, voice, image and other modalities of the video.

[0003] Most of the current video emotion recognition methods are based on pre-trained models. This method needs to train a model with a huge amount of parameters in a large amount of data set, so the calculation consumption is too much, and it cannot meet the efficient emotion recognition demand. At the same time, the existing emotion recognition method based on the combination of multiple modalities usually realizes the fusion of multiple modalities by simply splicing the features of multiple modalities, so it cannot deeply mine the information of multiple modalities, resulting in low accuracy of emotion recognition. SUMMARY

[0004] Therefore, the embodiments of the present application provide at least a video character emotion recognition method, device, equipment and storage medium.

[0005] The technical solution of the present application is implemented as follows:

[0006] In a first aspect, the embodiments of the present application provide a video character emotion recognition method, which comprises: extracting first data features of each modality in multiple modalities of a target video through a preprocessing module; performing fusion processing on the first data features of the multiple modalities through a hierarchical dynamic temporal fusion module to obtain second data features; performing operation processing on the second data features and third data features through a gated temporal neural module to obtain fourth data features, wherein the third data features are data features obtained by the gated temporal neural module through the last round of operation processing; and performing classification processing on the fourth data features through an emotion classification module to obtain an emotion type of a character in the target video.

[0007] Secondly, embodiments of this application provide a device for recognizing the emotions of people in a video. This device includes a preprocessing module, a hierarchical dynamic temporal fusion module, a gated temporal neural network module, an emotion classification module, and an emotion classification module. The preprocessing module extracts first data features from each of multiple modalities in the target video. The hierarchical dynamic temporal fusion module fuses the first data features from multiple modalities to obtain second data features. The gated temporal neural network module performs calculations on the second and third data features to obtain fourth data features, where the third data feature is the data feature obtained by the gated temporal neural network module through the previous round of calculations. The emotion classification module classifies the fourth data feature to obtain the emotion type of the person in the target video.

[0008] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by at least one processor, implements the method described in the first aspect.

[0009] Fourthly, embodiments of this application provide a computer device including a memory and a processor; wherein the memory is used to store computer-executable instructions; and the processor is connected to the memory and is used to implement the method described in the first aspect by executing the computer-executable instructions.

[0010] In this embodiment, a preprocessing module extracts first data features from each of the multiple modalities of the target video; a hierarchical dynamic temporal fusion module fuses the first data features from multiple modalities to obtain second data features; a gated temporal neural module performs calculations on the second and third data features to obtain fourth data features, where the third data feature is the data feature obtained by the gated temporal neural module through the previous round of calculations; and an emotion classification module classifies the fourth data feature to obtain the emotion type of the person in the target video. Thus, by fusing data features from multiple modalities and combining the results of the previous round of calculations with the current input information for emotion recognition, not only can the accuracy of recognition be improved, but compared to methods based on pre-trained models, computational overhead can also be reduced, thereby improving recognition efficiency.

[0011] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this disclosure. Attached Figure Description

[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.

[0013] Figure 1 A flowchart illustrating a method for recognizing emotions in a video, provided in an embodiment of this application;

[0014] Figure 2 A schematic diagram of a gated temporal neural module provided in an embodiment of this application;

[0015] Figure 3 A schematic diagram illustrating a possible implementation flow of the video character emotion recognition method provided in this application embodiment;

[0016] Figure 4 for Figure 3 A schematic diagram of the composition structure of the hierarchical dynamic temporal fusion module;

[0017] Figure 5 A schematic diagram of the composition structure of a video character emotion recognition device provided in an embodiment of this application;

[0018] Figure 6 This is a schematic diagram of a hardware entity of a computer device in an embodiment of this application. Detailed Implementation

[0019] In order to gain a more detailed understanding of the features and technical content of the embodiments of this application, the implementation of the embodiments of this application will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for reference and illustration only and are not intended to limit the embodiments of this application.

[0020] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.

[0021] In the following description, references to "some embodiments" refer to a subset of all possible embodiments. It is understood that "some embodiments" may be the same or different subsets of all possible embodiments and may be combined with each other without conflict. It should also be noted that the terms "first, second, third" used in the embodiments of this application are merely for distinguishing similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.

[0022] Emotion recognition in videos is a major research hotspot in the field of artificial intelligence. To improve the accuracy of emotion recognition, multiple modalities can be combined to identify the emotions of people in videos. For example, the emotions of people in a video can be identified by combining text, speech, and image modalities.

[0023] Most current video emotion recognition methods are based on pre-trained models. These multimodal pre-trained models often use models with a large number of parameters and are pre-trained on massive datasets. Compared to single-modal pre-training, multimodal pre-training requires datasets with more modalities and more pre-training tasks, often resulting in enormous computational overhead. Furthermore, pre-trained models, in order to have sufficient representational power, are often very large in scale, computationally intensive, and slow in response, making them difficult to deploy in real-world environments. Therefore, achieving lightweight yet high-performance emotion recognition based on multimodal tasks remains a challenge. Simultaneously, existing emotion recognition methods based on combining multiple modalities typically achieve fusion by simply concatenating features from multiple modalities, thus failing to deeply mine information from multiple modalities, resulting in low accuracy in emotion recognition.

[0024] To this end, this application provides a method, apparatus, and computer-readable storage medium for recognizing emotions in video characters. The method can be executed by a processor of a computer device. The computer device can refer to a server, laptop, tablet, desktop computer, or other device with data processing capabilities. This method determines the emotion type of a person in a target video by fusing data features from multiple modalities and combining the processing results from the previous stage with the input information of the current stage. This not only improves the accuracy of emotion recognition but also reduces computational overhead compared to methods based on pre-trained models, thereby improving recognition efficiency.

[0025] The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0026] In one embodiment of this application, see [link to embodiment]. Figure 1 This illustrates a method for recognizing the emotions of people in a video, provided by an embodiment of this application. The method may include steps S101 to S104.

[0027] S101, through the preprocessing module, extract the first data feature for each of the multiple modalities of the target video.

[0028] In S101, the preprocessing module can extract first data features for each of the multiple modalities of the target video. As an example, assuming that the multiple modalities of the target video include two or more modalities, such as the three modalities denoted as the first modal, the second modal, and the third modal, the preprocessing module can extract first data features for the first modal, the second modal, and the third modal respectively.

[0029] In one example, the modality of the target video may include text, speech, image, or any two of these. In this embodiment, extracting a first data feature for a certain modality can also be understood as extracting a first data feature from the data of that modality. For example, extracting a first data feature from text data; similarly, extracting a first data feature from speech data and image data.

[0030] For example, the way to extract the first data features of text through the preprocessing module can be: using global vectors for word representation (Glove) as the feature input word vectors of the text, or using the BERT (Bidirectional Encoder Representations from Transformers) model to extract text features.

[0031] For example, the way to extract the first data features of speech through the preprocessing module can be as follows: the speech information is truncated according to a predefined frequency, and then the speech feature input vector is obtained by a speech analysis framework (such as COVAREP). Alternatively, speech features can be extracted using models such as Coformer and Wav2vec.

[0032] For example, the way to extract the first data features of an image through a preprocessing module can be as follows: capture the current image information of the video at a predefined frequency, and then obtain the feature input vector of the image through a facial expression analysis framework (such as FACET). Alternatively, a convolutional model such as ResNet can be used to extract the facial features of the person in the image.

[0033] It should be understood that the number and type of modalities in the target video described above are merely exemplary and should not constitute any limitation on the implementation process of the embodiments of this application. For example, in some scenarios, the modalities of the target video may also include other types of modalities, such as human actions.

[0034] It should be noted that the preprocessing module can perform multiple rounds of data feature extraction for each of the multiple modalities, where each round of data feature extraction can correspond to a timestamp of the target video. For example, if the timestamp corresponding to the first round of data feature extraction is the 1st second of the target video, it means that the data features extracted in the first round are the data features of the target video between 0 and 1 second; as another example, if the timestamp corresponding to the second round of data feature extraction is the 2nd second of the target video, it means that the data features extracted in the second round are the data features of the video between 1 and 2 seconds. It should be understood that this application does not limit the value of the time interval between adjacent timestamps. In S101, the first data feature can be considered as the data feature extracted by the preprocessing module in the t-th round, where t is a positive integer.

[0035] S102, through the hierarchical dynamic temporal fusion module, the first data features of multiple modalities are fused to obtain the second data features.

[0036] The fusion process can be understood as fusing multiple input data features and then outputting the fused data features.

[0037] To facilitate understanding of the embodiments of this application, the steps performed by the hierarchical dynamic temporal fusion module are illustrated below, taking the modalities of the target video including text, voice, and images as an example.

[0038] For example, the hierarchical dynamic temporal fusion module may include multiple single-modal feature extraction layers, multiple dual-modal fusion layers, a trimodal fusion layer, and an output layer. The process of fusing the first data features of text, speech, and image to obtain the second data features through the hierarchical dynamic temporal fusion module can be implemented, for example, through steps 11 to 14.

[0039] Step 11: Through each single-modal feature extraction layer, perform feature extraction on the first data feature corresponding to that single-modal feature extraction layer to obtain the single-modal data feature.

[0040] For example, when the target video's modalities include text, speech, and image, the unimodal feature extraction layer may include a unimodal feature extraction layer corresponding to text (hereinafter referred to as the text unimodal feature extraction layer), a unimodal feature extraction layer corresponding to speech (hereinafter referred to as the speech unimodal feature extraction layer), and a unimodal feature extraction layer corresponding to image (hereinafter referred to as the image unimodal feature extraction layer). The first data feature corresponding to the text unimodal feature extraction layer is the first data feature extracted from the text; the first data feature corresponding to the speech unimodal feature extraction layer is the first data feature extracted from the speech; and the first data feature corresponding to the image unimodal feature extraction layer is the first data feature extracted from the image.

[0041] In this step, the first data features of the text can be extracted through the text unimodal feature extraction layer to obtain text unimodal features. At the same time, the first data features of the speech can be extracted through the speech unimodal feature extraction layer to obtain speech unimodal features. And the first data features of the image can be extracted through the image unimodal feature extraction layer to obtain image unimodal features.

[0042] In this embodiment, text monomodal features, speech monomodal features, and image monomodal features can be collectively referred to as monomodal features.

[0043] As one implementation, the aforementioned single-modal feature extraction layer can be a Long Short-Term Memory (LSTM) network. In this implementation, let x be the first data features extracted from text, speech, and image, respectively. l x a x v And assume that the text unimodal features, speech unimodal features, and image unimodal features are denoted as h, respectively. l h a h v Then the unimodal feature of the text can be represented as h l =LSTM(x l The single-modal features of speech can be represented as h. a =LSTM(x a The single-modal features of an image can be represented as h. v =LSTM(x v ).

[0044] Step 12: For each bimodal fusion layer, extract features from the sum of the two single-modal features corresponding to the bimodal fusion layer to obtain bimodal data features.

[0045] When the target video's modalities include text, speech, and image, the bimodal feature extraction layer may include a bimodal fusion layer corresponding to text and speech (hereinafter referred to as the text-speech bimodal fusion layer), a bimodal fusion layer corresponding to text and image (hereinafter referred to as the text-image bimodal fusion layer), and a bimodal fusion layer corresponding to image and speech (hereinafter referred to as the image-speech bimodal fusion layer). Specifically, the two single-modal features corresponding to the text-speech bimodal fusion layer are a text single-modal feature and a speech single-modal feature, respectively; the two single-modal features corresponding to the text-image bimodal fusion layer are a text single-modal feature and an image single-modal feature, respectively; and the two single-modal features corresponding to the image-speech bimodal fusion layer are an image single-modal feature and a speech single-modal feature, respectively.

[0046] In this step, a text-speech bimodal fusion layer can be used to extract features from the sum of text monomodal features and speech monomodal features to obtain text-speech bimodal features. At the same time, a text-image bimodal fusion layer can be used to extract features from the sum of text monomodal features and image monomodal features to obtain text-image bimodal features. Finally, an image-speech bimodal fusion layer can be used to extract features from the sum of image monomodal features and speech monomodal features to obtain image-speech bimodal features.

[0047] In this embodiment, text-to-speech bimodal features, text-to-image bimodal features, and image-to-speech bimodal features can be collectively referred to as bimodal features.

[0048] As one implementation, the aforementioned bimodal fusion layer is a Long Short-Term Memory (LSTM) network. In this implementation, let h be the text monomodal features, speech monomodal features, and image monomodal features, respectively. l h a h v And assume that the text-speech bimodal features, text-image bimodal features, and image-speech bimodal features are denoted as h, respectively. la h lv h av Then the bimodal features of text and speech can be represented as h la =LSTM(h l+ h a The text-image bimodal features can be represented as h lv =LSTM(h l+ h v The image-speech bimodal features can be represented as h av =LSTM(h a+ h v ).

[0049] Step 13: Through the trimodal fusion layer, feature extraction is performed on the result of adding the single-modal features and the dual-modal features to obtain trimodal data features.

[0050] When the modalities of the target video include text, speech, and image, the trimodal fusion layer can also be called the text-speech-image trimodal fusion layer, and the resulting trimodal data features can also be called text-speech-image trimodal data features.

[0051] In this step, all single-modal and dual-modal data features can be added together, and the result is then input into the trimodal fusion layer to obtain the desired trimodal data features. For example, in this embodiment, three single-modal data features can be obtained through step 11, namely h l h a hv Three bimodal data features can be obtained through step 12, namely h la h lv h av In step 13, h can be... l h a h v h la h lv h av The data are added together, and the result is then input into the three-modal fusion layer to obtain the required three-modal data features.

[0052] As one implementation, the aforementioned trimodal fusion layer is a Long Short-Term Memory (LSTM) network. Let h be the trimodal data features of text, speech, and image. lav In this implementation, the three-modal data features of text, speech, and image can be represented as h lav =LSTM(h l +h a +h v +h la +h lv +h av ).

[0053] In this embodiment, since the LSTM network not only accepts the input from the previous layer, but also accepts the information of the neurons in the current layer at the previous time step, using the LSTM network as a single-modal feature extraction layer, a dual-modal fusion layer, and a trimodal fusion layer in the hierarchical dynamic temporal fusion module is beneficial to make full use of the temporal information within the modality, thereby improving the accuracy of emotion recognition.

[0054] Step 14: Through the output layer, add the single-modal data features, dual-modal data features, and trimodal data features to obtain the second data feature.

[0055] In this step, the multiple unimodal data features obtained in step 11, the multiple bimodal features obtained in step 12, and the trimodal data features obtained in step 13 can be added together through the output layer to obtain the final second data feature. For example, in the example above, the second data feature (e.g., denoted as O) can be represented as O = h l +h a +h v +h la +h lv +h av +h lav .

[0056] It is understandable that when the target video includes two or more modalities, the first data features of multiple modalities can be fused using a similar method to obtain the desired second data features. For example, when the target video includes two modalities, the hierarchical dynamic temporal fusion module may include two single-modal feature extraction layers, one bimodal fusion layer, and an output layer. Each single-modal feature extraction layer can be used to extract features from the first data features corresponding to that single-modal feature extraction layer to obtain single-modal data features; the bimodal fusion layer can be used to extract features from the sum of the two obtained single-modal features to obtain bimodal data features; and the output layer can be used to add the two single-modal data features and the bimodal data feature to obtain the desired second data features.

[0057] It should be noted that when the preprocessing module performs multiple rounds of data feature extraction, the hierarchical dynamic temporal fusion module can perform multiple rounds of fusion processing on the data features output by the preprocessing module. Similarly, each round of fusion processing can correspond to a timestamp of the target video. In S102, the second data feature can be considered as the data feature output by the hierarchical dynamic temporal fusion module after fusing the data features extracted by the preprocessing module in round t. That is, the second data feature is the data feature output by the hierarchical dynamic temporal fusion module in round t.

[0058] In traditional multimodal fusion methods, most achieve fusion of multiple modalities through simple concatenation of modalities. For example, during fusion, multiple modalities are fused by concatenating input features. This method often focuses only on capturing features within a single modality, with little consideration for mining relationships between modalities. However, multimodal data often exhibit correlation and complementarity, which are typically significant in emotion recognition. Therefore, in the method of this embodiment, the input data features are fused through multiple single-modal feature extraction layers, multiple bimodal fusion layers, a trimodal fusion layer, and an output layer in a hierarchical dynamic temporal fusion module. This approach enhances the deep capture of relationships between modalities, thereby improving the accuracy of emotion recognition in complex human videos.

[0059] S103, through the gated temporal neural module, the second data feature and the third data feature are processed to obtain the fourth data feature, wherein the third data feature is the data feature obtained by the gated temporal neural module through the previous round of processing.

[0060] The third data feature can also be understood as the data feature output by the gated temporal neural network module in the previous iteration. This third data feature is the output result obtained by inputting the data feature output by the hierarchical dynamic temporal fusion module in the previous round (i.e., round t-1) into the gated temporal neural network. Since the data feature output by the hierarchical dynamic temporal fusion module in the previous round can correspond to the previous timestamp of the target video, the third data feature can also correspond to the previous timestamp of the target video. It should be understood that in the first iteration, i.e., when t is 1, the initial value of the third data feature is 0.

[0061] In video character emotion recognition tasks, since the emotional changes of characters in the target video are generally continuous, the data features corresponding to the previous timestamp of the target video may have reference value for the emotion recognition result corresponding to the next timestamp. Therefore, in this method, by combining the historical information of the previous round (i.e. the result of the previous round of calculation) for calculation, it is beneficial to better mine the temporal information of multiple modalities, that is, to consider the influence of the previous time sequence on the result, thereby improving the accuracy of subsequent emotion recognition.

[0062] Figure 2 A schematic diagram of a gated temporal neural module provided in an embodiment of this application is shown.

[0063] like Figure 2 As shown, the gated temporal neural module includes a first gate unit 201, a second gate unit 202, a third gate unit 203, a first operation unit 204, and a second operation unit 205. The process of processing the second and third data features through the gated temporal neural module to obtain the fourth data feature can be achieved, for example, through steps 21 to 26:

[0064] Step 21: Extract features of dimension l from the second data features through the first gating unit 201 to obtain the first processing result.

[0065] Step 22: Extract features of dimension m from the second data features through the second gating unit 202 to obtain the second processing result.

[0066] Step 23: The second data features are extracted with dimension n through the third gating unit 203 to obtain the third processing result.

[0067] exist Figure 2 In the process, the first gating unit 201, the second gating unit 202, and the third gating unit 203 can respectively extract features of dimensions l, m, and n from the second data, where l, m, and n are positive integers.

[0068] It should be understood that the values ​​of l, m, and n can be equal or unequal, and this application does not limit them.

[0069] As one implementation method, the values ​​of l, m, and n are not equal. That is to say, the first gating unit 201, the second gating unit 202, and the third gating unit 203 can respectively extract features of different dimensions from the second data. Since features of different dimensions can be observed from different angles, information mining can be more thorough.

[0070] Step 24: Perform a first operation on the first processing result and the third data feature through the first operation unit to obtain the first operation result.

[0071] Step 25: Perform a second operation on the second processing result and the third processing result through the second operation unit to obtain the second operation result.

[0072] In one implementation, the first and second arithmetic units can be used, for example, to perform element-wise multiplication of the input data. That is, the first arithmetic unit can be used to perform element-wise multiplication of the first processing result and the third data feature to obtain the first arithmetic result; the second arithmetic unit can be used to perform element-wise multiplication of the second processing result and the third processing result to obtain the second arithmetic result.

[0073] Step 26: The result of adding the first and second operation results is determined as the fourth data feature.

[0074] In this step, the first calculation result obtained in step 24 and the second calculation result obtained in step 25 can be added together, and the result of the addition can be determined as the fourth data feature. That is, the fourth data feature = the first calculation result + the second calculation result.

[0075] In some embodiments, before performing the first operation on the first processing result and the third data feature through the first operation unit, the method further includes: performing a nonlinear transformation on the first processing result through the first activation function 206; before performing the second operation on the second processing result and the third processing result through the second operation unit, the method further includes: performing a nonlinear change on the second processing result through the second activation function 207, and performing a nonlinear transformation on the third processing result through the third activation function 208.

[0076] For example, the first activation function 206 and the second activation function 207 can be first-type activation functions, and the third activation function 208 can be a second-type activation function. The first-type activation function and the second-type activation function can be different. For example, the first-type activation function can be a sigmoid activation function; for example, the second-type activation function can be a Tanh activation function.

[0077] It should be noted that the activation functions described above are merely exemplary and should not constitute any limitation on the implementation process of the embodiments of this application. For example, in other scenarios, other activation functions can be used to perform nonlinear transformation processing on the first and second processing results according to actual needs, and this application does not limit this.

[0078] In some embodiments, the gated temporal neural module may further include a temporal memory 209, which can be used to store a third data feature. That is, the data feature obtained by the gated temporal neural module through the previous round of processing can be stored in the temporal memory 209; or, in other words, while outputting the third data feature in the previous iteration, the gated temporal neural module can also update the third data feature in the temporal memory 209. Thus, in the current round of processing, the gated temporal neural module can obtain the second data feature through the hierarchical dynamic temporal fusion module, and simultaneously read the third data feature from the temporal memory 209 of the gated temporal neural module, and then combine the second and third data features to obtain the desired fourth data feature.

[0079] According to the method of this embodiment, the gated temporal neural module can combine the processing results of the previous round with the data features of the current input to obtain the processing result of the current round. Because this method incorporates historical information (i.e., the processing results of the previous round) in obtaining the processing result of the current round, it helps to improve the accuracy of subsequent emotion recognition.

[0080] S104, through the emotion classification module, the fourth data feature is classified to obtain the emotion type of the person in the target video.

[0081] In this embodiment, the fourth data feature can be predicted and processed by the emotion classification module to obtain the probability distribution of the emotions of the characters in the target video, and then the emotion type of the characters in the target video can be determined based on the probability distribution.

[0082] In one possible scenario, if the maximum probability value in the probability distribution is greater than or equal to the first threshold, then the emotion type corresponding to the maximum probability value in the probability distribution can be determined as the emotion type of the person in the target video. In other words, if the maximum probability value in the probability distribution is greater than or equal to the first threshold, the final emotion recognition result can be output.

[0083] The probability distribution of human emotions in the target video includes probability values ​​corresponding to multiple emotion types. For example, the probability distribution can be represented as P = [P1, P2, P3, P4, P5, P6], where P1 to P6 can represent the probability values ​​corresponding to different emotion types. For example, P1 can represent the probability value of the emotion type being anger, P2 can represent the probability value of the emotion type being disgust, P3 can represent the probability value of the emotion type being fear, P4 can represent the probability value of the emotion type being happiness, P5 can represent the probability value of the emotion type being surprise, and P6 can represent the probability value of the emotion type being sadness.

[0084] In the example above, assuming that P4 is greater than other probability values ​​in the probability distribution and P4 is greater than the first threshold, then "happy" can be identified as the emotional type of the person in the target video.

[0085] As one implementation, the emotion classification module may include a feedforward neural network and a softmax activation function. The input fourth data feature can be passed sequentially through the feedforward neural network and the softmax activation function to obtain the probability distribution of the emotions of the characters in the target video.

[0086] It should be understood that in this embodiment, the first threshold can be a numerical value or a range of values, and is not limited thereto.

[0087] It should also be understood that the first threshold can be a pre-configured fixed value or a dynamically configured value. For example, the first threshold can be pre-stored in the sentiment classification module or input into the sentiment classification module via external input; this application does not limit this. In a preferred example, the value of the first threshold is 0.8.

[0088] In another possible scenario, if the maximum probability value in the probability distribution is less than the first threshold, the method further includes: using the emotion classification module to predict the fifth data feature to obtain the probability distribution of the emotions of the characters in the target video, wherein the fifth data feature is the data feature obtained by the gated temporal neural module through the next round of computation.

[0089] In other words, if the maximum probability value in the probability distribution is less than the first threshold, the next iteration can proceed to obtain a new probability distribution. If the maximum probability value in the new probability distribution is greater than or equal to the first threshold, the iteration stops, and the emotion type corresponding to the maximum probability value is determined as the emotion type of the person in the target video. Similarly, if the maximum probability value in the new probability distribution is still less than the first threshold, the next iteration continues.

[0090] In some embodiments, if the maximum probability value in the probability distribution obtained in each iteration is less than the first threshold, the maximum probability value in the probability distribution obtained in the last iteration can be determined as the emotional type of the person in the final target video. Alternatively, the emotional type corresponding to the maximum probability value obtained throughout the iteration process can be determined as the emotional type of the person in the final target video.

[0091] According to the method in this embodiment, when the maximum probability value in the probability distribution is greater than or equal to the first threshold, the current emotion recognition result is considered to have high confidence. At this point, the iteration can be stopped and the emotion recognition result can be output; otherwise, the next iteration continues. This significantly reduces the model's inference time and further increases the model's practicality. In the context of the big data era, this algorithm can meet the parallel needs of customers for rapid recognition and efficient processing of person videos.

[0092] Sentiment analysis has always been a hot research topic in natural language processing. In today's big data era, hundreds of millions of different emotional viewpoints are generated daily on various online platforms, such as product reviews, opinions on events, or even simple posts on social media. Accurately grasping the emotions implied in user-generated data has always been a core focus of sentiment analysis. Currently, text-based sentiment classification is relatively mature. However, for multimodal data, using text alone for sentiment analysis cannot accurately capture the emotions behind the text, resulting in a certain degree of bias. Analyzing human multimodal language is an emerging research area in natural language processing. Multimodal sentiment analysis extends text-based single-modal sentiment analysis to a fusion analysis of multiple modalities. Currently, common modalities include text, speech, and image. Compared to single-modal analysis, the challenge of multimodal analysis lies not only in capturing the feature information within each modality but also in grasping the interaction feature information between modalities.

[0093] Traditional multimodal fusion methods are mostly based on simple modal concatenation and fusion, often focusing only on capturing features within a single modality and giving little consideration to the mining of relationships between modalities. However, multimodal data often have correlations and complementarities, and their importance is no less than that of information within a single modality. Current research is mostly based on pre-trained models, training models with a large number of parameters on massive datasets. Although pre-trained models have excellent performance, their large size and excessive computational cost limit their practicality in real-world applications.

[0094] With the widespread adoption of the internet, existing multimodal fusion methods lack the ability to capture the feature relationships between modalities in the face of massive amounts of data. Furthermore, pre-trained models are too large to deploy effectively and cannot fully meet the demands for efficient emotion recognition. Therefore, how to achieve efficient emotion recognition of multimodal videos in big data has become an urgent problem to solve.

[0095] LSTM is a type of recurrent neural network (RNN), which is suitable for solving sequence problems. Traditional neural networks have fully connected layers, but neurons within the same layer do not communicate with each other. However, in sequence processing, the output of the previous stage influences the output of the next stage. Therefore, recurrent neural networks (RNNs) can be used. RNNs not only accept input from the previous layer but also information from neurons in the current layer at the previous time step. RNNs effectively address the shortcomings of traditional neural networks in solving sequence problems, but they also suffer from gradient explosion or vanishing gradient problems when the network is too deep or the number of time steps is too large. LSTM, or Long Short-Term Memory network, can solve these problems. The key to the widespread application of RNNs lies in LSTM. The input gate receives recently useful information, and the forget gate selectively forgets older, less useful information. The output gate determines the output based on the current state.

[0096] In summary, current multimodal video emotion recognition still has some issues in terms of performance and practicality, such as:

[0097] Early multimodal fusion methods mostly involved simply concatenating input features from multiple modalities, such as Select-Additive Learning (SAL) and convolutional MKL multimodal sentiment analysis. Early fusion can be seen as an initial attempt by multimodal researchers to learn multimodal representations, as it utilized the correlation between the underlying features of each modality. The fusion process was simply a concatenation of input features, fusing multiple independent modal information into a single feature vector without requiring specific model design. However, it often failed to fully utilize the complementarity between multiple modal data, resulting in a large amount of redundant information in the fused data, and its fusion characteristics typically ignored the time factor. Later multimodal fusion methods fused the scores (decisions) from classifiers trained on different modalities, such as information acquisition analysis and deep multimodal fusion. The advantage of this approach is that the errors in the fused model come from different classifiers, and these errors are often uncorrelated and do not affect each other, thus preventing further accumulation of errors. While late-stage fusion is strong in modeling intramodal information within a specific modality, it has significant shortcomings in understanding intermodal feature interactions, as these interactions are usually more complex than decision voting. Previous research focused more on combining intramodal and intermodal information, achieving good results in multimodal emotion classification by effectively fusing intramodal and intermodal features. However, it did not delve deeper into capturing and fusing the temporal characteristics of each modality. Therefore, most traditional multimodal fusion methods are based on simple modal concatenation and fusion, often focusing only on capturing intramodal features and giving little consideration to mining intermodal relationships. In contrast, multimodal data often have correlations and complementarities, and their importance is no less than that of intramodal information.

[0098] Current multimodal pre-trained models mostly use models with a large number of parameters for pre-training on massive datasets. Compared to single-modal pre-training, multimodal pre-training requires datasets with more modalities and more pre-training tasks, often resulting in enormous computational overhead. Furthermore, pre-trained models are often very large in scale to achieve sufficient representational power, leading to slow computation and difficulty in deploying them in real-world environments. Therefore, realizing a lightweight yet high-performance multimodal task video emotion recognition model has become a challenge.

[0099] This application's embodiment of the video character emotion recognition method based on multimodal dynamic temporal fusion graphs improves upon the aforementioned shortcomings:

[0100] For intramodal feature capture, an LSTM network is used to consider historical input information, achieving better depth capture of single-modal temporal information. For intermodal relationship capture, a hierarchical temporal dynamic fusion graph is designed, using a single-modal feature capture layer and a multimodal feature fusion layer to mine deep intermodal relationships. Compared with traditional algorithms, the embodiments of this application can not only deeply mine single-modal information, but also further deeply fuse information between different modalities, resulting in a significant improvement in multimodal video emotion recognition performance.

[0101] Regarding the model's inference time, this embodiment uses only a few LSTM networks and linear layers, significantly reducing training costs and achieving a more lightweight model size compared to current pre-trained models. Simultaneously, a multi-level classification system is employed, combining historical information with the data at each timestamp for classification. If a threshold is exceeded, the classification is directly returned; otherwise, the next timestamp's information is identified. This greatly reduces the model's inference time, resulting in better practicality in real-world applications.

[0102] The above text combined Figure 1 and Figure 2 This application introduces a method for recognizing emotions in video characters, provided by an embodiment of the present application. To facilitate understanding of the embodiments of this application, a specific scenario is used as an example below, combined with... Figure 3 This application describes one possible implementation flow of the video character emotion recognition method provided in its embodiments. In this scenario, the modalities of the target video include text, speech, and images.

[0103] like Figure 3 As shown, the implementation process may include the following steps:

[0104] a) The first data features are extracted from the text data, voice data and image data by the preprocessing module 301.

[0105] It should be noted that the preprocessing module 301 can perform multiple rounds of data feature extraction on text data, voice data, and image data. Each round of data feature extraction can correspond to a timestamp of the target video. For example, if the timestamp corresponding to the first round of data feature extraction is the 1st second of the target video, it means that the data features extracted in the first round are the data features of the target video between 0 and 1 second; as another example, if the timestamp corresponding to the second round of data feature extraction is the 2nd second of the target video, it means that the data features extracted in the second round are the data features of the video between 1 and 2 seconds. It should be understood that this application does not limit the value of the time interval between adjacent timestamps.

[0106] In this embodiment, the first data features extracted from text data, voice data, and image data can be denoted as follows: and Here, the subscript t represents the iteration round, or t can also be understood as the index number of the timestamp.

[0107] In one optional example, text data uses GloVe vectors as feature input word vectors (i.e., the first data feature of the text), with a dimension of 300; speech data is extracted at a frequency of 30 frames per second, and then the COVAREP speech analysis framework is used to obtain feature input vectors (i.e., the first data feature of the speech), with a dimension of 74; image data is extracted at a frequency of 100 frames per second, and then the FACET facial expression analysis framework is used to extract image feature input vectors (i.e., the first data feature of the image), with a dimension of 35.

[0108] b) The first data features extracted from text data, voice data and image data are fused through the hierarchical dynamic temporal fusion module 302 to obtain the second data features.

[0109] like Figure 3 As shown, the first data output by the preprocessing module 301 and The data can be further input into the hierarchical dynamic temporal fusion module 302. Through fusion processing, it outputs the second data feature (or multimodal fusion feature) after fusing the first data features of multiple modalities (text, speech, and image). The second data feature can be denoted as O. t .

[0110] Figure 4 Further shown Figure 3 A schematic diagram of the composition structure of the mid-layer dynamic timing fusion module 302. (See diagram below.) Figure 4 As shown, the hierarchical dynamic temporal fusion module 302 includes three unimodal feature extraction layers, three bimodal fusion layers, one trimodal fusion layer, and an output layer 408. Specifically, the unimodal feature extraction layers include a text unimodal feature extraction layer 401, a speech unimodal feature extraction layer 402, and an image unimodal feature extraction layer 403; the bimodal fusion layers include a text-speech bimodal fusion layer 404, a text-image bimodal fusion layer 405, and an image-speech bimodal fusion layer 406; and the trimodal fusion layer includes a text-speech-image trimodal fusion layer 407.

[0111] Through the hierarchical dynamic time-series fusion module 302, the first data features and The second data feature O is obtained by performing fusion processing. t The specific implementation steps are as follows, which correspond to S102 in the aforementioned method embodiment:

[0112] First, the first data features of the text are extracted through the text single-modal feature extraction layer 401. Feature extraction is performed to obtain the text unimodal features. Simultaneously, the first data features of the speech are extracted through the speech single-modal feature extraction layer 402. Feature extraction is performed to obtain single-modal speech features. The first data features of the image are extracted through the image unimodal feature extraction layer 403. Feature extraction is performed to obtain the image's single-modal features.

[0113] Second, the text-speech dual-modal fusion layer 404 is used to process the text single-modal features. and speech monomodal features The summation results are used for feature extraction to obtain the text-speech bimodal features. Simultaneously, the text-image dual-modal fusion layer 405 is used to process the text single-modal features. and image single-modal features The summation result is used for feature extraction to obtain the bimodal features of the text image. And through the image-speech dual-modal fusion layer 406, the image single-modal features are processed. and speech monomodal features The summation result is used for feature extraction to obtain image-speech bimodal features.

[0114] Third, through the text-speech-image trimodal fusion layer 407, feature extraction is performed on the sum of all single-modal features and all bimodal features to obtain the text-speech-image trimodal data features.

[0115] Fourth, through the output layer 408, all unimodal data features, all bimodal data features, and all trimodal data features are added together to obtain the second data feature O. t .

[0116] As one implementation method, the single-modal feature extraction layer, dual-modal fusion layer and trimodal fusion layer in the embodiments of this application can be an LSTM network.

[0117] c) The second and third data features are processed by the gated temporal neural module 303 to obtain the fourth data feature.

[0118] The third data feature is the data feature obtained from the previous round of computation by the gated temporal neural module 303. Alternatively, the third data feature can also be understood as the data feature corresponding to the previous timestamp, which is obtained by using O t-1 The output obtained after inputting into a gated temporal neural network. It should be understood that in the first iteration, the initial value of the third data feature is 0.

[0119] In this embodiment, the third data feature can be denoted as U. t-1 The fourth data feature output by the gated temporal neural module 303 in the current stage can be denoted as U. t , .

[0120] like Figure 3 As shown, the gated timing neural module 303 consists of three gated units (D1, D2, D...). u It consists of two arithmetic units (first arithmetic unit and second arithmetic unit) and a sequential memory, wherein D1, D2, and D... u It can be used to process the input O. t Perform feature extraction in different dimensions; the time-series memory can be used to store third-party data features U. t-1 .

[0121] In some embodiments, in D1, D2, D u Each of these is followed by an activation function, where σ represents the sigmoid activation function and h represents the Tanh activation function.

[0122] The second data feature O is processed through the gated temporal neural module 303. t and third data feature U t-1 The fourth data feature U is obtained through computational processing. t The specific implementation steps are as follows, which correspond to S103 in the aforementioned method embodiment:

[0123] The first processing result is obtained by performing feature extraction of dimension l on the second data features using D1; the second processing result is obtained by performing feature extraction of dimension m on the second data features using D2; and so on. u The second data feature is subjected to feature extraction of dimension n to obtain the third processing result; the first processing unit performs a first operation on the first processing result and the third data feature to obtain the first operation result; the second processing unit performs a second operation on the second processing result and the third processing result to obtain the second operation result; the result of adding the first operation result and the second operation result is determined as the fourth data feature U. t .

[0124] In some embodiments, before performing the first operation on the first processing result and the third data feature through the first operation unit, the above implementation steps further include: performing nonlinear transformation processing on the first processing result through the sigmoid activation function; before performing the second operation on the second processing result and the third processing result through the second operation unit, the above steps further include: performing nonlinear transformation processing on the second processing result through the sigmoid activation function, and performing nonlinear transformation processing on the third processing result through the Tanh activation function.

[0125] d) The fourth data feature is classified and processed by the emotion classification module 304 to obtain the emotion type of the person in the target video.

[0126] As an example, the sentiment classification module 304 can classify U t Prediction processing is performed to obtain the probability distribution of the emotions of the characters in the target video. If the maximum probability value in the probability distribution is greater than or equal to a first threshold, the emotion type corresponding to the maximum probability value in the probability distribution can be determined as the emotion type of the character in the target video, and the final emotion recognition result is output. If the maximum probability value in the probability distribution is less than the first threshold, the next iteration continues, or in other words, the iteration continues with the next timestamp. That is, if the maximum probability value in the probability distribution is less than the first threshold, the emotion classification module 304 can continue to process U... t+1 A prediction process is performed to obtain a new probability distribution. Similarly, if the maximum probability value in this new probability distribution is still less than the first threshold, the next iteration continues. This step corresponds to S104 in the aforementioned method embodiment.

[0127] According to the method of this embodiment, on the one hand, the reliability of sentiment analysis is increased by fusing data features from multiple modalities; on the other hand, the accuracy of sentiment recognition is further improved by combining the results of the previous round of computation and the current input information. Furthermore, this method only requires one forward propagation to obtain the sentiment recognition result, which reduces computational overhead compared to methods based on pre-trained models, thereby improving recognition efficiency and making it more practical in real-world applications.

[0128] It should be noted that the video character emotion recognition method provided in this application embodiment can also be applied to any other multimodal scenario to realize the computational analysis of multimodal tasks in different environments, thus having strong versatility and scalability.

[0129] The training process of the hierarchical dynamic temporal fusion module and the gated temporal neural module is briefly explained below:

[0130] Before training, the model parameters of the hierarchical dynamic temporal fusion module and the gated temporal neural module are randomly initialized. In this embodiment, text data uses GloVe vectors as feature input word vectors (i.e., the first data feature of the text), with a dimension of 300; speech data is extracted at a frequency of 30 frames per second, and then the COVAREP speech analysis framework is used to obtain feature input vectors (i.e., the first data feature of the speech), with a dimension of 74; image data is extracted at a frequency of 100 frames per second, and then the FACET facial expression analysis framework is used to extract image feature input vectors (i.e., the first data feature of the image), with a dimension of 35.

[0131] During training, the first data features of text, speech, and image are input into a hierarchical dynamic temporal fusion module to obtain the second data features. These second data features are then input into a gated temporal neural module, where they are combined with historical information (i.e., the data features output from the previous round of the gated temporal neural module) to obtain the fourth data feature. Next, the fourth data feature is input into a sentiment classification module to obtain the probability distribution of human emotions in the target video. Afterward, loss is calculated based on the corresponding label of the target video, and the model parameters are optimized through backpropagation. The initial learning rate can be set to 0.0001.

[0132] As one implementation approach, the Adaptive Moment Estimation (Adam) optimizer can be used to optimize the model parameters, and the Cross Entropy Loss function can be used as the model's loss function. The batch size can be set to 64. Furthermore, to prevent overfitting, the Dropout algorithm can be used to randomly ignore some neurons in the fully connected layers.

[0133] During training, the loss function can be minimized using an optimizer based on the training samples to achieve model convergence. The model weights that achieve the best results after a certain number of iterations are saved as the model parameters for the final hierarchical dynamic temporal fusion module and gated temporal neural module. In the experiments of this application, the training samples included: 3443 angry videos, 2720 disgusted videos, 1319 afraid videos, 8147 happy videos, 3906 surprised videos, and 1562 sad videos.

[0134] Table 1 shows the emotion recognition accuracy and F1 score obtained using a multilayer feedforward neural network (MFN) and the method of this application, respectively. Here, WA represents weighted accuracy; the F1 score is related to the model's classification precision and recall, and it comprehensively considers both metrics, making it a good measure.

[0135] Table 1

[0136]

[0137] As shown in Table 1, the method in this application achieves the best results for most test samples compared to the MFN method, with a performance improvement of up to 7.4%. Regarding sentiment recognition efficiency, the test sample time for the MFN method is 4.261 ms, while the test sample time for the method in this application is 2.133 ms. In contrast, the method in this application can stop iterating and output the sentiment recognition result when the confidence level of the current sentiment recognition result is high. This significantly reduces the model's inference time and further increases the model's practicality.

[0138] The hardware environment for the experimental data in this application is shown in Table 2. The experiment was run on a GPU server to deploy the model on the GPU, as detailed below:

[0139] GPU: Tesla K40m*1; CUDA core count: 2880; Double-precision floating-point performance: 1.43 Tflops; Single-precision floating-point performance: 4.29 Tflops; Memory bandwidth: 288GB / s, supports PCI-E 3.0; Power consumption: 235W TDP, passive cooling; Memory: 128GB.

[0140] Table 2

[0141]

[0142] The software environment for the experimental data in this application is shown in Table 3. The key component for implementing this experiment is the Python compiler, specifically PyCharm, version 2021.1.

[0143] Table 3

[0144] Key Components Technical / Software Name Version Number Python Compiling Software Pycharm 2021.1

[0145] This application proposes a video character emotion recognition system based on multimodal dynamic temporal fusion graphs. This method achieves automated recognition of complex character emotions in videos, significantly reducing human resource costs. Because it performs calculations based on multiple modalities, it can adapt to emotion analysis in complex scenarios, increasing the efficiency and reliability of the analysis. Compared with traditional emotion analysis algorithms, this method can greatly improve the accuracy of emotion recognition. Compared with pre-trained models, the model in this application is more lightweight and practical. This application can help customers quickly and in parallel determine the emotion distribution in massive amounts of video data.

[0146] This application improves upon the MFN-DFG algorithm by adding a multimodal temporal dynamic fusion graph framework (corresponding to the hierarchical dynamic temporal fusion module mentioned above). It uses multiple hierarchical LSTM modules, which not only consider temporal information within a modality and further deepen feature mining, but also enhance the deep capture of relationships between modalities, thereby improving the accuracy of emotion recognition in complex human videos.

[0147] This application further incorporates a hierarchical sentiment classification module (corresponding to the sentiment classification module mentioned above). Results with high confidence are returned early, significantly reducing inference time while preserving the performance of matching models unsuitable for multi-level classification, thus further increasing the model's practicality. In the context of the big data era, this algorithm can meet the parallel needs of clients for rapid identification and efficient processing of person videos.

[0148] Current emotion recognition models are mainly divided into traditional methods and pre-trained methods. Traditional methods are mostly based on intra-modal data mining and have a weak ability to capture relationships between modalities. While pre-trained methods have excellent performance, the parameters of pre-trained models are huge in real-world scenarios, making them difficult to deploy in the runtime environment. Furthermore, their computational efficiency is too low to meet the requirements of rapid response. This application proposes a video human emotion recognition system based on multimodal dynamic temporal fusion graphs, which comprehensively solves the above two problems. Under the premise of lightweight model size, it has better accuracy, effectiveness, and robustness, and has certain novelty and innovation.

[0149] The embodiments of this application have the following advantages:

[0150] 1) This application proposes a dynamic temporal fusion graph (corresponding to the hierarchical dynamic temporal fusion module mentioned above) structure and applies it to the emotion recognition process of complex human videos. Compared with traditional methods, it greatly improves the accuracy, effectiveness and robustness of multimodal human emotion recognition and effectively helps staff to quickly locate the emotion distribution of human beings in a large number of human videos in parallel.

[0151] 2) The dynamic temporal fusion graph method proposed in this application can not only perform in-modal feature depth capture and in-depth intermodal relationship mining of text, image and voice in the video of a person, but also seamlessly switch to any other multimodal scene. It can realize the computational analysis of multimodal tasks in various environments and has strong versatility and scalability.

[0152] 3) This application proposes a hierarchical sentiment classification system (corresponding to the sentiment classification module mentioned above), which returns results with high confidence levels in advance, greatly reducing inference time while retaining the performance of matching without using a multi-level classification model, further increasing the practicality of the model.

[0153] In another embodiment of this application, based on the same inventive concept as the foregoing embodiments, see [link to previous embodiment]. Figure 5 This illustration shows a schematic diagram of the composition of a video character emotion recognition device 500 provided in an embodiment of this application. Figure 5 As shown, the device 500 may include: a preprocessing module 501, a hierarchical dynamic temporal fusion module 502, a gated temporal neural module 503, and an emotion classification module 504; wherein,

[0154] The preprocessing module 501 is used to extract first data features for each of the multiple modalities of the target video;

[0155] The hierarchical dynamic temporal fusion module 502 is used to fuse the first data features of multiple modalities to obtain the second data features;

[0156] The gated temporal neural module 503 is used to perform calculations on the second data feature and the third data feature to obtain the fourth data feature. The third data feature is the data feature obtained by the gated temporal neural module through the previous round of calculations.

[0157] The emotion classification module 504 is used to classify the fourth data feature to obtain the emotion type of the person in the target video.

[0158] In some embodiments, the emotion classification module 504 is specifically used to: perform prediction processing on the fourth data feature through the emotion classification module to obtain the probability distribution of the emotions of the characters in the target video; if the maximum probability value in the probability distribution is greater than or equal to the first threshold, then determine the emotion type corresponding to the maximum probability value in the probability distribution as the emotion type of the characters in the target video.

[0159] In some embodiments, if the maximum probability value in the probability distribution is less than the first threshold, the emotion classification module 504 is further used to perform prediction processing on the fifth data feature to obtain the probability distribution of the emotions of the characters in the target video, wherein the fifth data feature is the data feature obtained by the gated temporal neural module through the next round of computation.

[0160] In some embodiments, the gated temporal neural module 503 includes: a first gating unit, a second gating unit, a third gating unit, a first operation unit, and a second operation unit, wherein: the first gating unit is used to perform feature extraction of dimension l on the second data feature to obtain a first processing result; the second gating unit is used to perform feature extraction of dimension m on the second data feature to obtain a second processing result; the third gating unit is used to perform feature extraction of dimension n on the second data feature to obtain a third processing result, wherein l, m, and n are positive integers; the first operation unit is used to perform a first operation on the first processing result and the third data feature to obtain a first operation result; the second operation unit is used to perform a second operation on the second processing result and the third processing result to obtain a second operation result; the gated temporal neural module 503 is further used to determine the result of adding the first operation result and the second operation result as a fourth data feature.

[0161] In some embodiments, before performing the first operation on the first processing result and the third data feature through the first operation unit, the gated temporal neural module 503 is further configured to perform nonlinear transformation processing on the first processing result through the first activation function; before performing the second operation on the second processing result and the third processing result through the second operation unit, the gated temporal neural module 503 is further configured to perform nonlinear change processing on the second processing result through the second activation function, and perform nonlinear transformation processing on the third processing result through the third activation function.

[0162] In some embodiments, the gated timing neural module 503 further includes a timing memory for storing third data features.

[0163] In some embodiments, the hierarchical dynamic temporal fusion module 502 includes multiple single-modal feature extraction layers, multiple bimodal fusion layers, a trimodal fusion layer, and an output layer, wherein: each single-modal feature extraction layer is used to extract features from a first data feature corresponding to the single-modal feature extraction layer to obtain single-modal data features; each bimodal fusion layer is used to extract features from the sum of two single-modal features corresponding to the bimodal fusion layer to obtain bimodal data features; the trimodal fusion layer is used to extract features from the sum of single-modal features and bimodal features to obtain trimodal data features; and the output layer is used to add the single-modal data features, bimodal data features, and trimodal data features to obtain a second data feature.

[0164] The descriptions of the apparatus embodiments above are similar to those of the method embodiments above, and have similar beneficial effects. In some embodiments, the functions or modules included in the apparatus provided in this disclosure can be used to perform the methods described in the method embodiments above. For technical details not disclosed in the apparatus embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0165] It should be noted that, in the embodiments of this application, if the above-described data processing method is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to the related technology, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.

[0166] This application provides a computer device including a memory and a processor. The memory stores a computer program that can run on the processor. When the processor executes the program, it implements some or all of the steps in the above-described method.

[0167] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements some or all of the steps in the above-described method. The computer-readable storage medium can be transient or non-transient.

[0168] This application provides a computer program including computer-readable code, wherein when the computer-readable code is executed in a computer device, a processor in the computer device performs some or all of the steps in the above-described method.

[0169] This application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described method. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium; in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0170] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between them, while their similarities or commonalities can be referred to interchangeably. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have similar beneficial effects. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the descriptions of the method embodiments of this application for understanding.

[0171] It should be noted that, Figure 6 This is a schematic diagram of a hardware entity of a computer device in an embodiment of this application, such as... Figure 6 As shown, the hardware entity of the computer device 600 includes: a processor 601, a communication interface 602, and a memory 603, wherein:

[0172] Processor 601 typically controls the overall operation of computer device 600.

[0173] Communication interface 602 enables computer devices to communicate with other terminals or servers via a network.

[0174] The memory 603 is configured to store instructions and applications executable by the processor 601, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) of the processor 601 and various modules in the computer device 600. It can be implemented using flash memory or random access memory (RAM). Data transfer between the processor 601, the communication interface 602, and the memory 603 can be performed via bus 604.

[0175] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above steps / processes do not imply a sequential order of execution; the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above embodiments of this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0176] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0177] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.

[0178] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.

[0179] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0180] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.

[0181] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.

[0182] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A method for recognizing emotions in video characters, characterized in that, The method includes: The preprocessing module extracts the first data features for each of the multiple modalities in the target video. The first data features of the multiple modalities are fused using a hierarchical dynamic temporal fusion module to obtain the second data features. The second and third data features are processed by the gated temporal neural module to obtain the fourth data feature, wherein the third data feature is the data feature obtained by the gated temporal neural module through the previous round of processing. The emotion classification module classifies the fourth data feature to obtain the emotion type of the person in the target video. The gated temporal neural module includes a first gating unit, a second gating unit, a third gating unit, a first arithmetic unit, and a second arithmetic unit, wherein: The fourth data feature is obtained by processing the second and third data features through a gated temporal neural module, including: The second data feature is dimensionally defined by the first gating unit. l Feature extraction is performed to obtain the first processing result; The second data feature is dimensionally defined by the second gating unit. m Feature extraction is performed to obtain the second processing result; The second data feature is dimensionally defined by the third gating unit. n Feature extraction yields a third processing result, in which... l , m , n It is a positive integer; The first processing unit performs a first operation on the first processing result and the third data feature to obtain a first operation result. The second processing unit performs a second operation on the second processing result and the third processing result to obtain a second operation result. The result of adding the first calculation result and the second calculation result is determined as the fourth data feature.

2. The method according to claim 1, characterized in that, The step of classifying the fourth data feature using the emotion classification module to obtain the emotion type of the person in the target video includes: The emotion classification module performs prediction processing on the fourth data feature to obtain the probability distribution of human emotions in the target video. If the maximum probability value in the probability distribution is greater than or equal to the first threshold, then the emotion type corresponding to the maximum probability value in the probability distribution is determined as the emotion type of the person in the target video.

3. The method according to claim 2, characterized in that, The method further includes: If the maximum probability value in the probability distribution is less than the first threshold, then the emotion classification module performs prediction processing on the fifth data feature to obtain the probability distribution of the character's emotion in the target video. The fifth data feature is the data feature obtained by the gated temporal neural module through the next round of computation.

4. The method according to any one of claims 1 to 3, characterized in that, Before performing the first operation on the first processing result and the third data feature through the first processing unit, the method further includes: The first processing result is subjected to a nonlinear transformation using a first activation function; Before performing the second operation on the second processing result and the third processing result through the second processing unit, the method further includes: The second processing result is subjected to nonlinear transformation processing by a second activation function, and the third processing result is subjected to nonlinear transformation processing by a third activation function.

5. The method according to any one of claims 1 to 3, characterized in that, The gated temporal neural module includes a temporal memory for storing the third data feature.

6. The method according to any one of claims 1 to 3, characterized in that, The hierarchical dynamic temporal fusion module includes multiple single-modal feature extraction layers, multiple dual-modal fusion layers, a trimodal fusion layer, and an output layer, wherein: The second data feature is obtained by fusing the first data features of the multiple modalities through the hierarchical dynamic temporal fusion module, including: Through each of the single-modal feature extraction layers, feature extraction is performed on the first data feature corresponding to the single-modal feature extraction layer to obtain single-modal data features; Through each of the bimodal fusion layers, feature extraction is performed on the result of adding the two single-modal data features corresponding to the bimodal fusion layer to obtain bimodal data features; The trimodal fusion layer extracts features from the sum of the single-modal data features and the dual-modal data features to obtain trimodal data features. The second data feature is obtained by adding the single-modal data feature, the dual-modal data feature, and the trimodal data feature through the output layer.

7. A device for recognizing the emotions of people in a video, characterized in that, The device includes a preprocessing module, a hierarchical dynamic temporal fusion module, a gated temporal neural network module, and an emotion classification module; wherein... The preprocessing module is used to extract first data features for each of the multiple modalities of the target video; The hierarchical dynamic temporal fusion module is used to fuse the first data features of the multiple modalities to obtain the second data features. The gated temporal neural module is used to perform calculations on the second data feature and the third data feature to obtain a fourth data feature, wherein the third data feature is the data feature obtained by the gated temporal neural module through the previous round of calculations. The emotion classification module is used to classify the fourth data feature to obtain the emotion type of the person in the target video; The gated temporal neural module includes: a first gate unit, a second gate unit, a third gate unit, a first arithmetic unit, and a second arithmetic unit, wherein: The first gating unit is used to perform a dimensional adjustment on the second data feature. l Feature extraction is performed to obtain the first processing result; The second gating unit is used to perform a dimensional adjustment on the second data feature. m Feature extraction is performed to obtain the second processing result; The third gating unit is used to perform a dimension-based gating of the second data feature. n Feature extraction yields a third processing result, in which... l , m , n It is a positive integer; The first arithmetic unit is used to perform a first operation on the first processing result and the third data feature to obtain a first operation result; The second arithmetic unit is used to perform a second operation on the second processing result and the third processing result to obtain a second operation result; The gated temporal neural module is further configured to determine the result of adding the first operation result and the second operation result as the fourth data feature.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by at least one processor, implements the method as described in any one of claims 1 to 6.

9. A computer device, characterized in that, The computer device includes: Memory is used to store executable instructions for a computer; A processor, connected to the memory, is configured to implement the method of any one of claims 1 to 6 by executing the computer-executable instructions.

Citation Information

Patent Citations

  • Multi-modal emotion recognition method based on fusion attention network

    CN110188343A

  • Hierarchical multi-modal sentiment analysis method based on multi-task learning

    CN114973045A