A method and system for analyzing emotional video content based on a constrained multi-modal multi-level attention fusion model

Through the constrained multimodal multi-level attention fusion model, the modal feature difference and alignment problems in emotional video content analysis are solved, and more accurate emotional state recognition and more effective multimodal emotion recognition are achieved.

CN116503780BActive Publication Date: 2025-05-27CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310465547.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-26
Publication Date
2025-05-27
Estimated Expiration
2043-04-26

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify the emotional state of the video in the analysis of emotional video content, mainly due to the differences between different modal features and modal alignment problems.

Method used

The constrained multimodal multi-level attention fusion model is adopted to extract multimodal information through the multimodal pre-trained model, perform in-modal feature fusion and intermodal feature integration, and use loss functions of reconstruction loss, difference loss and similarity loss for constraints to improve the feature alignment efficiency of the model.

Benefits of technology

It realizes more accurate identification of the emotional state of the video, alleviates the interference of redundant noise in the modal features, fully extracts key emotional information between different modalities, and improves the effect of multimodal emotion recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116503780B_ABST
    Figure CN116503780B_ABST
Patent Text Reader

Abstract

The present invention is a method and system for analyzing emotional video content based on a constrained multimodal multi-level attention fusion model. First, the global and local features of each modality are combined to help the model extract the overall tone of the video and the local details of the video. Next, the method uses a cross-attention module to combine data from three modalities to further extract emotionally rich features within the multimodal range, and then uses a self-attention module to integrate data from each modality. The applicant proposed a constrained multimodal multi-level Tranformer derivative method based on a standard self-attention mechanism and a cross-attention mechanism, including a multimodal emotional intra-analysis model that gradually fuses features through multiple levels. For the first time, a loss function was used to constrain the learning of Tokens in Tranformer, and good results were achieved. In classification and regression experiments, better results than previous technologies were achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of pattern recognition, and particularly to a method and system for analyzing emotional video content of a constrained multi-modal multi-level attention fusion model. Background Art

[0002] The statements in this section merely provide background technical information related to the present disclosure, and these statements may constitute prior art. In the process of implementing the present invention, the inventors found that at least the following problems exist in the prior art.

[0003] Emotions can affect our well-being, the way we interact with each other, our behavior, and our judgment. Many media that we encounter in our daily lives, such as music and movies, are designed specifically to evoke emotional responses. Many psychological studies have investigated the emotional impact of movies on audiences. According to the scenes, music, and sounds in movies, different emotions of the audience can be triggered. And due to its applications in the fields of human-computer interaction, emotion-based customized content recommendation, personalized video recommendation, violent video recognition, etc., in recent years, emotional video content analysis has received more and more attention. Although encouraging progress has been made in the research on emotion computing, it is still very difficult to accurately interpret the emotions of videos using computer algorithms.

[0004] Emotional video content analysis aims to automatically analyze the emotions of videos, and the emotional content of videos is defined as the intensity and type of emotions that people expect to generate when watching videos. Generally, videos are divided into audio content and video content. Today's mainstream methods for emotional video content analysis first extract video image representations and audio representations through features, and then integrate multi-modal information through a series of alignments and fusions for classification or regression. The biggest problems faced by these fusion technologies are the differences between different modal features and the modal alignment problem, mainly including intra-modal fusion and inter-modal fusion. Intra-modal fusion is divided into the fusion of global and local features and the fusion of temporal information. Intra-modal fusion mainly aligns information from different modalities at different times. With a single and identical overemphasis on the information from all modalities, the feature alignment efficiency of its model is low, and the emotional state of the video cannot be accurately identified. Summary of the Invention

[0005] In view of the above problems, the object of the present invention is to provide a method for analyzing emotional video content of a multi-modal multi-level attention fusion model, in which the relevant convolution kernels can dynamically change with features, and can more flexibly model the interaction between multi-modal features, so as to more accurately identify the emotional state of the video.

[0006] A method for analyzing emotional video content based on a constrained multi-modal multi-level attention fusion model includes:

[0007] Extract samples from the emotional video content analysis database, and extract multimodal information from the samples through a multimodal pre-trained model;

[0008] Extract and splice the multimodal information to obtain the global local information corresponding to each modality;

[0009] Combine the global local information corresponding to each modality in pairs for intra-modal feature fusion to obtain the pairwise modality information corresponding to the modality;

[0010] Add and integrate the obtained pairwise modality information corresponding to the modality to obtain the single modality feature information of each modality;

[0011] Process the single modality feature information of each modality through classification or regression to obtain the final classification regression result; use the corresponding classification or regression loss function to obtain the final classification regression loss; wherein, the final classification regression result is the predicted emotion value of the model;

[0012] Calculate the loss value through the single modality feature information of each modality; the loss value includes difference loss, similarity loss and reconstruction loss;

[0013] Compare the predicted emotion value of the model with the label value of the sample; if the predicted emotion value does not match the label value, update the weight of the model using the backpropagation algorithm, and repeat the above steps until the predicted emotion value matches the label value, and use the model with stable weights for emotion classification and recognition of film and television works.

[0014] The multimodal information includes speech information F a 、key frame information F f and temporal information F composed of extraction from multiple consecutive frames s ; where a is the speech information modality, f is the key frame information modality, and s is the temporal information modality composed of extraction from multiple consecutive frames.

[0015] Furthermore, extracting and splicing the multimodal information to obtain the global local information corresponding to each modality includes the following steps:

[0016] Use a temporal neural network and a fully connected network to extract global information and local information from the multi-information modality respectively;

[0017] Process the global information and local information through a fully connected layer, and splice them after completion to obtain the global local information corresponding to each modality;

[0018] The temporal neural network used for local information extraction uses a long short-term memory network; the fully connected network used for global information extraction uses a single-layer fully connected layer.

[0019] Furthermore, the intra-modal feature fusion adopts an intra-modal feature fusion module, which includes two self-attention layers and two cross-attention layers; the two self-attention layers respectively process two global local information corresponding to the inputs to unify the global local information; the two cross-attention layers interact the two modal information processed by the self-attention layers to transfer pairwise modal information.

[0020] Furthermore, the single-modal feature information of each modality is processed by classification or regression to obtain the final classification and regression results; the corresponding classification or regression loss function is used to obtain the final classification and regression loss, including the following steps:

[0021] The single-modal feature information of each modality is connected in parallel;

[0022] The final classification and regression results are obtained through a self-attention layer and a classification or regression head;

[0023] The corresponding classification or regression loss function is used to obtain the final classification and regression loss; among them, the classification loss function uses the standard cross-entropy loss function; the regression loss function uses the standard mean square error loss function.

[0024] Furthermore, the loss value is calculated through the single-modal feature information of each modality, including the following steps:

[0025] The single-modal feature information of each modality is respectively passed through two linear layers to map the corresponding features to the dimension size of the features when inputting into the model; then the reconstruction loss L is obtained by doing the loss function with the model input rebuild , the formula is:

[0026]

[0027] The difference loss L is calculated by using the single-modal feature information of each modality diff , the formula is:

[0028]

[0029] The similarity loss L is calculated by using the single-modal feature information of each modality sim , the formula is:

[0030]

[0031] Among them, i and j represent one of the three modalities a, f, s, and i, j are different, f i represents the modality i feature obtained by the model, M i is the single-modal feature information of each modality, represents M i the global feature vector in Similarly; represents the local feature vector in M i Among them, similarly, n is the number of segments of the video or audio, s() represents the bitwise summation function, and F i represents the corresponding three multi-modal information inputs. In the formula, T represents the matrix transpose operation, and F is the inherent subscript expressed by the symbol of the F norm.

[0032] Furthermore, if the predicted emotion value does not match the label value, update the weights of the model using the backpropagation algorithm for the loss value, including the following steps:

[0033] Compare the obtained predicted emotion value with the label value of the sample;

[0034] If the comparison result does not match, update the parameters of the temporal neural network and the fully connected network using the backpropagation algorithm for the loss value;

[0035] Extract samples from the emotional video content analysis database again, and repeat the entire process using the updated model until the predicted emotion value matches the label value of the sample, and the model training is completed.

[0036] Furthermore, the database includes LIRIS-ACCEDE and its subsets.

[0037] An emotional video content analysis system based on a constrained multi-modal multi-level attention fusion model, including:

[0038] An extraction module for extracting multi-modal information from samples through a multi-modal pre-training model;

[0039] A global-local module for extracting and splicing multi-modal information to obtain the global-local information corresponding to each modality;

[0040] An intra-modal feature fusion module for pairwise combining the global-local information corresponding to each modality for intra-modal feature fusion to obtain the pairwise modality information corresponding to the modality;

[0041] An inter-modal fusion module for adding and integrating the pairwise modality information corresponding to the modality to obtain the single-modal feature information of each modality;

[0042] An aggregation module for connecting the single-modal feature information of each modality in parallel, and then obtaining the final classification and regression result through a self-attention layer and a classification or regression head, and obtaining the final classification and regression loss using the corresponding classification or regression loss function; among them, the final classification and regression result is the predicted emotion value of the model;

[0043] The auxiliary training module includes a reconstruction module and a similarity and difference loss calculation module; the reconstruction module is used to respectively pass the single-modal feature information of each modality through two linear layers, map the obtained corresponding features to the dimension size of the features when inputting into the model, and then calculate the reconstruction loss with the model input using a loss function; the similarity and difference loss calculation module is used to calculate the similarity loss and the difference loss through the single-modal feature information of each modality;

[0044] The model weight update module is used to compare the predicted emotion value of the model with the label value of the sample; if the predicted emotion value does not match the label value, the reconstruction loss, the similarity loss, and the difference loss are used to update the parameters of the temporal neural network and the fully connected network in the global-local module using the backpropagation algorithm; if the predicted emotion value matches the label value, the model training is completed.

[0045] The calculation formula of the reconstruction loss L rebuild is as follows:

[0046]

[0047] The calculation formula of the difference loss L diff is as follows:

[0048]

[0049] The calculation formula of the similarity loss L sim is as follows:

[0050]

[0051] Where i and j represent one of the three modalities a, f, s, and i and j are different, f i represents the modality i feature obtained by the model, M i is the single-modal feature information of each modality, represents the global feature vector in M i ; Similarly; represents the local feature vector in M i ; Similarly; n is the number of segments of the video or audio, s() represents the bitwise summation function, F i represents the corresponding three multi-modal information inputs, T in the formula represents the matrix transpose operation, and F is the inherent subscript expressed by the symbol of the F norm.

[0052] The present invention has the following beneficial effects:

[0053] 1. Use the loss function with reconstruction loss, difference loss, and similarity loss as constraints to improve the efficiency of the model. Among them, the similarity loss is used to obtain more similar global features, learn the common points in different modalities, and learn modality-invariant representations. The difference loss is used to avoid losing the specific features within each modality in the loss function. The reconstruction loss is to retain the original details in the features extracted by the extraction model, so as to avoid learning useless features after applying the difference loss and similarity loss, rather than just learning unrepresentative vectors according to the loss.

[0054] 2. Propose an efficient module to improve the feature alignment efficiency of the model by multi-level fusion instead of single and identical over-concern synthesis of information from all modalities. Compared with the prior art, based on using the self-attention mechanism after feature linear mapping and adding global and temporal information to the features before attention, it is more suitable for video tasks. The idea of hierarchical fusion is added to make the fusion of features more refined.

[0055] 3. The relevant convolutional kernels can change dynamically with the features, enabling more flexible modeling of the interactions between multi-modal features, thereby more accurately identifying the emotional state of the video.

[0056] 4. By using different branch networks to fuse the features of different modalities, it can effectively alleviate the interference of redundant noise in the features, enable the model to fully extract the key emotional information between different modalities, and more effectively achieve multi-modal emotion recognition. At the same time, it can more flexibly model the interactions between multi-modal features. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 is the model flow chart of the present invention;

[0058] Figure 2 is the overall structure diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0059] The following further describes the present invention with reference to the accompanying drawings. The embodiments of the present invention are only used to illustrate the present invention and do not limit the present invention. Without departing from the technical idea of the present invention, various substitutions and changes made according to the common general knowledge and customary means in the art shall be included within the scope of the present invention.

[0060] As Figure 1 shown in or 2, an emotional video content analysis method based on a constrained multi-modal multi-level attention fusion model includes:

[0061] Extract samples from the emotional video content analysis database, and extract multi-modal information from the samples through a multi-modal pre-training model;

[0062] Extracting and splicing the multimodal information to obtain global local information corresponding to each modality;

[0063] The global local information corresponding to each modality is combined in pairs to perform intra-modal feature fusion to obtain the pairwise modal information of the corresponding modality;

[0064] The obtained pairwise modal information of the corresponding modes is added and the information is integrated to obtain the single modal feature information of each mode;

[0065] The single modal feature information of each modality is processed by classification or regression to obtain the final classification regression result; the corresponding classification or regression loss function is used to obtain the final classification regression loss; wherein the final classification regression result is the predicted emotion value of the model;

[0066] Calculating loss values ​​through single modal feature information of each modality; the loss values ​​include difference loss, similarity loss and reconstruction loss;

[0067] The predicted emotion value of the model is compared with the label value of the sample; if the predicted emotion value does not match the label value, the loss value is used to update the weight of the model using the back propagation algorithm, and the above steps are repeated until the predicted emotion value matches the label value, and the model with stabilized weights is used for emotion classification and recognition of film and television works.

[0068] The multimodal information includes voice information F a , key frame information F f and the timing information F extracted from multiple consecutive frames s ; Among them, a is the speech information mode, f is the key frame information mode, and s is the timing information mode extracted from multiple consecutive frames.

[0069] The database includes LIRIS-ACCEDE and its subsets. The experiments carried out by the present invention are carried out on LIRIS-ACCEDE and its subsets, and the performance of the present invention is evaluated and analyzed.

[0070] Extracting and splicing the multimodal information to obtain global local information corresponding to each modality includes the following steps:

[0071] Use temporal neural network and fully connected network to extract global information and local information from multiple information modalities respectively;

[0072] The global information and local information are processed by a fully connected layer, and after completion, the global and local information corresponding to each mode is obtained by splicing;

[0073] The temporal neural network for local information extraction uses a long short-term memory network; the fully connected network for global information extraction uses a single-layer fully connected layer.

[0074] The in-modal feature fusion adopts an in-modal feature fusion module, which includes two layers of self-attention layers and two layers of cross-attention layers; the two layers of self-attention layers respectively process two global local information corresponding to the inputs to unify the global local information; the two layers of cross-attention layers interact the two modal information processed by the self-attention layers to transfer pairwise modal information.

[0075] The single-modal feature information of each modality is processed by classification or regression to obtain the final classification and regression results; the corresponding classification or regression loss function is used to obtain the final classification and regression loss, including the following steps:

[0076] The single-modal feature information of each modality is connected in parallel;

[0077] The final classification and regression results are obtained through a self-attention layer and a classification or regression head;

[0078] The corresponding classification or regression loss function is used to obtain the final classification and regression loss; among them, the classification loss function uses the standard cross-entropy loss function; the regression loss function uses the standard mean squared error loss function.

[0079] The loss value is calculated through the single-modal feature information of each modality, including the following steps:

[0080] The single-modal feature information of each modality is respectively passed through two linear layers to map the corresponding features to the dimension size of the features when inputting into the model; then the reconstruction loss L is obtained by making a loss function with the model input rebuild , and the formula is:

[0081]

[0082] The difference degree loss L is calculated by using the single-modal feature information of each modality diff , and the formula is:

[0083]

[0084] The similarity loss L is calculated by using the single-modal feature information of each modality sim , and the formula is:

[0085]

[0086] Among them, i and j represent one of the three modalities a, f, s, and i, j are different, f i represents the modality i feature obtained by the model, M i is the single-modal feature information of each modality, represents M i the global feature vector in Similarly; representing M i the local feature vector in Similarly; n is the number of segments of the video or audio, s() represents the bitwise summation function, F i represents the corresponding three multimodal information inputs. In the formula, T represents the matrix transpose operation, and F is the inherent subscript expressed by the symbol of the F-norm.

[0087] The reconstruction loss is to map the obtained feature vectors of each modality back to the shape of the original modality, so as to retain all the model details of the original features and prevent the model from learning irrelevant trivial knowledge.

[0088] If only the above loss function is used to capture the common features of the modalities, the specific features within each modality will be lost. The applicant uses the difference loss to avoid this situation and ensure that the local feature vectors and global feature vectors within a single branch can learn different information.

[0089] The similarity loss constrains the results of the inter-modal fusion module. The applicant calculates the pairwise similarity between the results of this module and the average value of the input features to obtain more similar global features, learn the common points in different modalities, and learn the modality-invariant representation.

[0090] Among all the above loss functions, the classification and regression functions are the performance indicators of the model, and the others are loss functions for auxiliary training.

[0091] If the predicted emotion value does not match the label value, update the weights of the model using the backpropagation algorithm for the loss value, including the following steps:

[0092] Compare the obtained predicted emotion value with the label value of the sample;

[0093] If the comparison result does not match, update the parameters of the temporal neural network and the fully connected network using the backpropagation algorithm for the loss value;

[0094] Extract samples from the emotion video content analysis database again, and repeat the entire process using the updated model until the predicted emotion value matches the label value of the sample, and the model training is completed.

[0095] The label value is the value of the video emotion defined manually in the samples extracted from the database. Comparing it with the predicted emotion value (i.e., the final classification and regression result) is to see the gap between the model's prediction and the true value, so as to judge whether the model is effective and whether it has learned the content to be recognized. It is the judgment criterion for the method performance. The backpropagation algorithm is a prior art.

[0096] The specific steps of the experiment are as follows:

[0097] Step 1: Extract multiple samples from the emotional video content analysis database, and then extract multimodal information F from each sample through the multimodal software development kit i , including speech information F a , key frame information F f and temporal information F composed of extraction from multiple consecutive frames s .

[0098] Step 2: Establish three corresponding groups of temporal neural networks and fully connected network branches respectively to process the corresponding speech information F a , key frame information F f and temporal information F composed of extraction from multiple consecutive frames s . The temporal neural network of each branch generates a local information summary of the corresponding modality where i represents one of the three modalities a, f, and s of the input multimodality. The fully connected network of each branch generates a global information summary of the corresponding modality

[0099] Step 3: Concatenate the local and global information summaries in each modality along the dimension of the number of feature vectors, and input them into a fully connected network layer to fuse the information of the two parts to obtain the feature vectors E i (F i ), where E represents the processing process of the previous steps, L means local, and G means global Global

[0100] The three features described in Steps 2 and 3, namely speech information F a , key frame information F f and temporal information F composed of extraction from multiple consecutive frames s are respectively sent to the global-local network of the corresponding branch to extract the corresponding local and global feature representations. Among them, long short-term memory networks are used for local feature extraction, and single-layer fully connected layers are used for global feature extraction. Finally, the extracted global-local features are passed through a fully connected layer and then concatenated. The specific steps of the concatenation process can be expressed as:

[0101] E i (F i ) = [H i,G ; H i,L

[0102] ​After extracting video or audio features, we can obtain a feature vector with the shape of (b×n×d), where b is the batch size, n is the number of video or audio segments, and d is the hidden dimension size of the audio-visual features within a fixed time domain. For local features, each segment vector is correspondingly sent into a long short-term memory network layer and mapped to the dimension size of D (D is the dimension size after mapping), and the last layer of the result is used as the local feature with the shape of (b×1×D). For global features, we map the d dimension to the D dimension and take the average over the n dimension to obtain global features with the same shape as the local features, which describe the characteristic tone of the entire video. Finally, the shapes of both the global and local vectors are (b×1×D). Among them, is the global information of the i-th modality, represents the local information of the i-th modality, E i (F i ) is the global-local information corresponding to each modality, that is, E a (F a ), E f (F f ), E s (F s ).

[0103] Step 4: Pairwise combine the obtained feature vectors E a (F a ), E f (F f ), E s (F s ) of the three modalities and send them into the intra-modal feature fusion module as shown in Figure 2 . The intra-modal feature fusion module contains a self-attention module and a cross-attention module. The self-attention module processes the two single-modal information corresponding to the input, and the cross-attention module interacts the two-modal information. The intra-modal feature fusion module consists of two layers of self-attention and two layers of cross-attention, and its corresponding formula is

[0104] m i,(i,j) , m j,(i,j) = CA i,j (SA i (E i (F i ))), SA j (E j (F j ))), i, j ∈ (a, f, s), i ≠ j

[0105] Finally, the corresponding outputs (i.e., pairwise modal information) are obtained

[0106] m a,(a,f) , m f,(a,f) , m a,(a,s) , ms,(a,s) , m s,(s,f) , m f,(s,f) .

[0107] Among them, CA i,j and SA j , SA i are cross-attention and self-attention operations. i, j represent input vectors from different branches. m a,(a,f) represents the result of processing the a-modal and f-modal in the a-modal branch, m f,(a,f) represents the result of processing the a-modal and f-modal in the f-modal branch, and so on for the rest; E j (F j ), E i (F i ) represents the combination selected from E a (F a ), E f (F f ), E s (F s ); E a (F a ) represents the result of using the formula for the input feature F a in the network a-modal branch, and so on for the rest.

[0108] It is worth mentioning that we send the same vector into different branches, and finally add the corresponding vectors of these two different branches to obtain the final module output.

[0109] Step 5: Add the pairwise modal information of the corresponding modalities obtained in Step 4 and integrate them through the inter-modal fusion module to obtain a single-modal feature. And use the similarity loss function and the difference loss function on the output M i of the inter-modal fusion module to calculate the similarity loss and the difference loss respectively. These two loss functions are used for subsequent gradient backpropagation.

[0110] Step 6: Concatenate the single-modal feature information M i of each modality obtained in Step 5 in parallel and pass them through a self-attention layer and a classification or regression head to obtain the final classification or regression result (i.e., the predicted emotion value of the model), and use the corresponding classification or regression loss function to obtain the final classification regression loss. These two loss functions are used for subsequent gradient backpropagation.

[0111] Step 7: Pass the single-modal feature information M i of each modality obtained in Step 5 through two linear layers respectively, map the obtained corresponding features to the size extracted by the input model, and then calculate the reconstruction loss with the model input. The reconstruction loss function is used for subsequent gradient backpropagation.

[0112] Step 8. Finally, compare the predicted emotion value of the model with the label value. The loss values (i.e., similarity loss, difference loss, and reconstruction loss) obtained by the loss function in Steps 5, 6, and 7 are used in the backpropagation algorithm to update the attention network and the fully connected layer parameters of the "three corresponding groups of temporal neural networks and fully connected network branches" in Step 2. Then, continuously repeat Steps 1 to 7 until the model can correctly predict the emotion category.

[0113] Step 9. Use the model after the update and convergence of the corresponding branches of different modalities to identify the emotion of the video to be identified.

[0114] Brief operation process:

[0115] First, we extract the feature vectors of three modalities from the given video, namely one audio modality and two visual modalities, corresponding to the extraction module in the figure. Immediately afterwards, we separately isolate the corresponding global features and local features from the obtained feature vectors of each modality and splice them together, corresponding to Figure 2 the global-local feature module in it. Then, we send them in pairs to the cross-modal fusion module for cross-modal information interaction. Among the obtained results, we add the Tokens of the same modality correspondingly and then input them into the corresponding intra-modal information integration modules, corresponding to Figure 2 the intra-modal fusion module in it.

[0116] The experimental results are compared as follows:

[0117] Each experiment aims to prove the effectiveness of the applicant's method under the experimental settings used previously. Table 1 shows the comparison between our method and other state-of-the-art methods in recent years. ACC represents the accuracy of the classification task, and MSE and PCC represent the mean square error and Pearson correlation coefficient in the regression task. Since the features and configurations of the experiments in other papers are different from each other, we only compare their best results here.

[0118] Table 1·Comparison of different methods.

[0119]

[0120] The experimental results of the existing papers are from the following references:

[0121] Ruc at mediaeval 2016emotional impact of movies task:Fusion ofmultimodal features.

[0122] Multi-modal learning for affective content analysis in movies

[0123] Multimodal local-global attention network for affective video contentanalysis,

[0124] Affective video content analysis with adaptive fusion recurrentnetwork

[0125] AttendAffectNet–Emotion Prediction of Movie Viewers Using MultimodalFusion with Self-Attention

[0126] Unified Multi-stage Fusion Network for Affective Video ContentAnalysis

[0127] P2SL:PRIVATE-SHARED SUBSPACES LEARNING FOR AFFECTIVE VIDEO CONTENTANALYSIS

[0128] From the above experimental data, it can be seen that the method for analyzing affective video content of the multi-modal multi-level attention fusion model of the present invention is overall superior to the existing classical methods. This verifies that the present invention can effectively alleviate the interference of redundant noise in multi-modal features, enable the model to fully extract the key affective information between different modalities, and more effectively realize multi-modal affective recognition.

[0129] An affective video content analysis system based on a constrained multi-modal multi-level attention fusion model, comprising:

[0130] An extraction module, configured to extract multi-modal information from samples through a multi-modal pre-trained model;

[0131] A global-local module, configured to extract and splice the multi-modal information to obtain the global-local information corresponding to each modality;

[0132] An intra-modal feature fusion module, configured to perform intra-modal feature fusion on the global-local information corresponding to each modality in pairs to obtain the pairwise modality information corresponding to the modality;

[0133] The inter-modal fusion module is used to add pairwise modal information of corresponding modalities and integrate the information to obtain single-modal feature information of each modality;

[0134] The aggregation module is used to connect the single-modal feature information of each modality in parallel, and then obtain the final classification and regression result through a self-attention layer and a classification or regression head, and use the corresponding classification or regression loss function to obtain the final classification and regression loss; wherein, the final classification and regression result is the predicted emotion value of the model;

[0135] The auxiliary training module includes a reconstruction module and a similarity and difference loss calculation module; the reconstruction module is used to respectively pass the single-modal feature information of each modality through two linear layers, map the obtained corresponding features to the dimension size of the features when inputting into the model, and then calculate the reconstruction loss with the model input using a loss function; the similarity and difference loss calculation module is used to calculate the similarity loss and the difference loss through the single-modal feature information of each modality;

[0136] The model weight update module is used to compare the predicted emotion value of the model with the label value of the sample; if the predicted emotion value does not match the label value, update the parameters of the temporal neural network and the fully connected network in the global-local module using the reconstruction loss, the similarity loss and the difference loss through the backpropagation algorithm; if the predicted emotion value matches the label value, the model training is completed.

[0137] The reconstruction loss L rebuild The calculation formula is:

[0138]

[0139] The difference loss L diff The calculation formula is:

[0140]

[0141] The similarity loss L sim The calculation formula is:

[0142]

[0143] wherein, i and j represent one of the three modalities a, f, s, and i, j are different, f i represents the modality i feature obtained by the model, M i is the single-modal feature information of each modality, represents M i the global feature vector in, similarly; represents M i the local feature vector in, Similarly, n is the number of segments of video or audio, s() represents the bitwise summation function, and F i represents the corresponding three multi-modal information inputs. In the formula, T represents the matrix transpose operation, and F is the inherent subscript expressed by the symbol of using the F-norm.

[0144] The applicant proposed a multi-modal multi-level Tranformer derivative method with constraints based on the standard self-attention mechanism and cross-attention mechanism, including a multi-modal sentiment internal analysis model that gradually fuses features through multiple levels. The loss function was also first used to constrain the learning of Tokens in the Tranformer, and good results were achieved. The research is based on a fusion model of multi-layer multi-modal Transformers, which improves the feature alignment efficiency of the model by multi-level fusion instead of single identical over-concern synthesis of information from all modalities. Compared with existing papers, based on using the self-attention mechanism after feature linear mapping, we want to add global and temporal information to the features before attention, which is more suitable for video tasks. The idea of hierarchical fusion is added to make the fusion of features more refined.

[0145] The Transformer can be said to be a deep learning model completely based on the self-attention mechanism. Because it is suitable for parallel computing and the complexity of its own model, it is higher than the previously popular RNN recurrent neural network in terms of accuracy and performance. Tokens include: class tokens, patch tokens. In natural language tasks, each word is called a token, and then there is a CLS annotation that labels the semantics of the sentence. In computer vision, the image is cut into a sequence of non-overlapping packets (which are actually tokens).

[0146] Through the above method, the method for analyzing the emotional video content of the multi-modal multi-level attention fusion model of the present invention can more accurately identify the emotional state of the movie. In addition, the present invention fuses the features of different modalities through different branch networks, which can effectively alleviate the interference of redundant noise in the features. At the same time, it can more flexibly model the interaction between multi-modal features.

Claims

1. A method for analyzing emotional video content based on a constrained multi-modal multi-level attention fusion model, characterized in that, it includes: Extract samples from the emotional video content analysis database, and extract multi-modal information from the samples through a multi-modal pre-training model; Extract and splice the multi-modal information to obtain the global local information corresponding to each modality; Combine the global local information corresponding to each modality in pairs for intra-modal feature fusion to obtain the pairwise modal information corresponding to the modality; the intra-modal feature fusion uses an intra-modal feature fusion module, which includes two layers of self-attention layers and two layers of cross-attention layers; the two layers of self-attention layers respectively process the two global local information corresponding to the input to unify the global local information; the two layers of cross-attention layers interact the two modal information processed by the self-attention layer to transmit the pairwise modal information; Add and integrate the pairwise modal information corresponding to the modality to obtain the single-modal feature information of each modality; Process the single-modal feature information of each modality through classification or regression to obtain the final classification regression result; use the corresponding classification or regression loss function to obtain the final classification regression loss; among them, the final classification regression result is the predicted emotion value of the model; Calculate the loss value through the single-modal feature information of each modality; the loss value includes a difference loss, a similarity loss, and a reconstruction loss; the difference loss L diff , and the formula is: The similarity loss L sim , the formula is: where i and j represent one of the three modalities a, f, and s, and i and j are different, and M i is the single-modal feature information of each modality, represents the global feature vector in M i ; similarly, represents the local feature vector in M i ; similarly, n is the number of segments of the video or audio, s() represents the bitwise summation function, T represents the matrix transpose operation in the formula, and F is the inherent subscript expressed by the symbol using the F-norm; ​​ Compare the predicted emotion value of the model with the label value of the sample; if the predicted emotion value does not match the label value, update the weight of the model using the backpropagation algorithm, and repeat the above steps until the predicted emotion value matches the label value, and use the model with stable weights for emotional classification and recognition of film and television works.

2. The method for analyzing emotional video content based on a constrained multi-modal multi-level attention fusion model according to claim 1, characterized in that, The multimodal information includes speech information F a , key frame information F f and temporal information F composed of extractions from multiple consecutive frames s ; where a is the speech information modality, f is the key frame information modality, and s is the temporal information modality composed of extractions from multiple consecutive frames 3. The method for analyzing emotional video content based on a constrained multi-modal multi-level attention fusion model according to claim 1 or 2, characterized in that, Extract and splice the multi-modal information to obtain the global local information corresponding to each modality, including the following steps: Use a temporal neural network and a fully connected network to extract global information and local information from the multi-modal information respectively; Process the global information and local information through a fully connected layer, and splice them after completion to obtain the global local information corresponding to each modality; The temporal neural network used for local information extraction uses a long short-term memory network; the fully connected network used for global information extraction uses a single-layer fully connected layer.

4. The method for analyzing emotional video content based on a constrained multi-modal multi-level attention fusion model according to claim 1 or 2, characterized in that, Process the single-modal feature information of each modality through classification or regression to obtain the final classification regression result; use the corresponding classification or regression loss function to obtain the final classification regression loss, including the following steps: Connect the single-modal feature information of each modality in parallel; Obtain the final classification regression result through a self-attention layer and a classification or regression head; Obtain the final classification and regression loss using the corresponding classification or regression loss function; among them, the classification loss function uses the standard cross-entropy loss function; the regression loss function uses the standard mean squared error loss function.

5. The method for analyzing emotional video content based on a constrained multi-modal multi-level attention fusion model according to claim 1 or 2, characterized in that calculate the loss value through the single-modal feature information of each modality, including the following steps: The single-modal feature information of each modality is respectively passed through two linear layers to map the corresponding features obtained to the dimensionality of the features when input into the model; then, a loss function is calculated with the model input to obtain the reconstruction loss L rebuild , and the formula is: The difference loss L is calculated using the single-modal feature information of each modality diff , and the formula is as follows: Calculate the similarity loss L using the single-modal feature information of each modality sim , and the formula is: Among them, f i represents the feature of modality i obtained by the model, and F i represents the corresponding three multi-modal information inputs.

6. The method for analyzing emotional video content based on a constrained multi-modal multi-level attention fusion model according to claim 3, characterized in that if the predicted emotion value does not match the label value, update the weights of the model using the backpropagation algorithm for the loss value, including the following steps: Compare the obtained predicted emotion value with the label value of the sample; if the comparison result does not match, update the parameters of the temporal neural network and the fully connected network using the backpropagation algorithm for the loss value; Extract samples from the emotional video content analysis database again, and repeat the entire process using the updated model until the predicted emotion value matches the label value of the sample, and the model training is completed.

7. The method for analyzing emotional video content based on a constrained multi-modal multi-level attention fusion model according to claim 1, characterized in that the database includes LIRIS-ACCEDE and its subsets.

8. An emotional video content analysis system based on a constrained multi-modal multi-level attention fusion model, characterized in that comprises: an extraction module for extracting multi-modal information from samples through a multi-modal pre-training model; a global-local module for extracting and splicing the multi-modal information to obtain the global-local information corresponding to each modality; an intra-modal feature fusion module for pairwise combining the global-local information corresponding to each modality for intra-modal feature fusion to obtain the pairwise modality information corresponding to the modality; the intra-modal feature fusion module includes two layers of self-attention layers and two layers of cross-attention layers; the two layers of self-attention layers respectively process the two global-local information corresponding to the input to unify the global-local information; the two layers of cross-attention layers interact the two modality information processed by the self-attention layer to transmit the pairwise modality information; an inter-modal fusion module for adding and integrating the pairwise modality information corresponding to each modality to obtain the single-modal feature information of each modality; an aggregation module for connecting the single-modal feature information of each modality in parallel, and then obtaining the final classification and regression result through a self-attention layer and a classification or regression head, and obtaining the final classification and regression loss using the corresponding classification or regression loss function; among them, the final classification and regression result is the predicted emotion value of the model; The auxiliary training module includes a reconstruction module and a similarity and difference loss calculation module; the reconstruction module is used to respectively pass the single-modal feature information of each modality through two linear layers, map the obtained corresponding features to the dimension size of the features when input into the model, and then calculate the reconstruction loss with the model input; the similarity and difference loss calculation module is used to calculate the similarity loss and the difference loss through the single-modal feature information of each modality; the difference loss L diff , and the formula is: The similarity loss L sim , the formula is as follows: where i and j represent one of the three modalities a, f, s, and i and j are different, and M i is the single-modal feature information of each modality, represents the global feature vector in M i ; similarly, represents the local feature vector in M ; similarly, n is the number of segments of the video or audio, s() represents the bitwise summation function, T represents the matrix transpose operation in the formula, and F is the inherent subscript expressed by the symbol using the F-norm; i ​​ a model weight update module for comparing the predicted emotion value of the model with the label value of the sample; if the predicted emotion value does not match the label value, update the parameters of the temporal neural network and the fully connected network in the global-local module using the backpropagation algorithm for the reconstruction loss, similarity loss and difference loss; if the predicted emotion value matches the label value, the model training is completed.

9. The emotional video content analysis system based on the constrained multi-modal multi-level attention fusion model according to claim 8, characterized in that, The reconstruction loss L rebuild is calculated as follows: Among them, f i represents the feature of modality i obtained by the model, and F i represents the corresponding three multi-modal information inputs.