Video Emotion Polarity Analysis Method Based on Multi-modal Deep Feature Hierarchical Fusion

Through the multimodal deep feature hierarchical fusion method, the redundancy and interference problems of multimodal information fusion in video sentiment analysis are solved, the accuracy and recognition rate of emotional polarity recognition are improved, and better emotional feature representation is achieved.

CN116844095BActive Publication Date: 2025-07-25TAIYUAN UNIVERSITY OF TECHNOLOGY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311064915.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-23
Publication Date
2025-07-25
Estimated Expiration
2043-08-23

AI Technical Summary

Technical Problem

The existing video sentiment analysis methods have information redundancy and mutual interference when multimodal information fusion, resulting in noise characteristics and affecting the accuracy of sentiment analysis.

Method used

The multimodal deep feature hierarchical fusion method is adopted to divide the video into segments through the video processing unit, extract the original features of facial expressions, speech and text data, and perform feature processing through the multimodal feature hierarchical interaction fusion unit and emotional polarity discrimination unit, including the hierarchical fusion of the underlying dual-modal and high-level trimodal feature hierarchical fusion, filtering redundant and noise information, and improving the emotional polarity recognition effect.

Benefits of technology

The recognition rate of the speaker's emotional polarity in the video clip is improved, the redundancy and noise in the multimodal characteristics are effectively filtered, the representation ability of multimodal emotional characteristics is improved, and the accuracy of emotional polarity analysis is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116844095B_ABST
    Figure CN116844095B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of intelligent information processing and multimodal emotion technology, and particularly relates to a method for video emotion polarity analysis based on multimodal deep feature hierarchical fusion, which solves the technical problems in the background art. The method includes original feature extraction, constructing a video segment emotion polarity analysis model, and model training and testing; the video segment emotion polarity analysis model includes a multimodal feature hierarchical interaction fusion unit and an emotion polarity discrimination unit. The multimodal feature hierarchical interaction fusion unit includes a bottom-layer bimodal feature interaction module and a top-layer trimodal feature hierarchical fusion module. The top-layer trimodal feature hierarchical fusion module includes a pairwise bilinear gating fusion unit and a trimodal self-attention feed-forward fusion unit. The present invention filters redundancy and noise in the fusion features while fully fusing multimodal features, improves the representation ability of multimodal emotion features, and effectively identifies the positive and negative polarities of the emotions of the speaker in the video segment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent information processing and multimodal emotion technology, and particularly relates to a method for video emotion polarity analysis based on multimodal deep feature hierarchical fusion. Background Art

[0002] With the rapid development of the Internet, intelligent mobile terminals, and domestic and foreign social media platforms such as Bilibili, Douyin, Kuaishou, and YouTube, more and more users tend to share their views on certain events, topics, policies, products, and services in the form of videos on various online social media. A large number of multimedia video resources containing personal emotional attitudes have emerged on the Internet. The video content generated by these users contains a large amount of information, which can reflect the emotions, attitudes, and views of the speakers, and has great commercial value and application value. For example, brand companies can clarify the public's evaluation of a certain brand through social media video analysis, and take corresponding actions and propose optimization plans when negative evaluations increase in the short term. In addition, research shows that the short video contribution behavior of non-professional content creators has increased significantly. Government agencies, social organizations, and other institutions analyze these multimedia video resources containing a large amount of personal emotional attitudes to cope with possible sudden public opinion impacts. In addition, with the advent of ChatGPT, the emotional dialogue technology has triggered a research boom, and accurately understanding and identifying the emotions of the speakers in the video is the primary basis for generating emotional responses.

[0003] Compared with text statements, multimodal videos containing visual, voice, and text information are more in line with the nature of human multi-sensory expression and multi-sensory perception. Users can express and perceive the emotions in the video from multiple dimensions. Emotion has very important social value and environmental adaptation significance and is the core of human cross-cultural communication. Humans naturally rely on identifying the emotions of others to judge their behavioral tendencies, so as to mobilize appropriate brain resources, adjust their own behaviors, and make reasonable decisions.

[0004] As early as the 1960s, the role of emotion in "machine intelligence" has attracted the attention of many scholars. For example, in 1967, Professor Herb Simon proposed that the general theory of thinking and problem-solving must include the influence of emotion; in 1986, Professor Minsky of the Massachusetts Institute of Technology put forward the assertion in "The Society of Mind" that "emotion is an important part of machine intelligence"; in 1997, Professor Picard of the Massachusetts Institute of Technology's Media Laboratory proposed the concept of Affective Computing: "Affective Computing is computing related to, derived from, or capable of influencing emotions"; in 1999, Professor Wang Zhiliang of the University of Science and Technology Beijing proposed the theory of artificial psychology; Professor Hu Baogang of the Institute of Automation of the Chinese Academy of Sciences and others also gave a definition of affective computing in combination with their own research: "The purpose of affective computing is to establish a harmonious human-computer environment by endowing computers with the ability to recognize, understand, express, and adapt to human emotions, and to make computers have higher and more comprehensive intelligence." Affective computing is gradually becoming an emerging research field. With the development of artificial intelligence technology and the continuous accumulation and improvement of the emotional theory system, affective computing will achieve more mature applications in fields such as distance education, healthcare, smart cities, fintech, smart home appliances, online entertainment, psychological construction, people's livelihood services, and natural human-computer interaction.

[0005] Sentiment Analysis (SA) is one of the key technical aspects in the field of affective computing, aiming to automatically process and analyze the physiological data (such as electrocardiogram, electroencephalogram, electromyogram, skin conductance, and respiratory signals) and behavioral data (such as gestures, facial expressions, body postures, and speech intonations) collected by computers, extract relevant emotional features, and accordingly build models to analyze the mapping relationship between the external manifestations and internal states of emotions, so as to predict information such as the current emotional state, emotion type, and emotional intensity value of the speaker.

[0006] In the early stage of research, sentiment analysis only considered the information in the text modality, extracting emotional clues from the internal context of the text and using them for emotion recognition. In recent years, with the popularization of multimodal human-computer interaction devices (Multimodal Human-Computer Interaction, MHCI) and the development of short-video online platforms, new professional identities such as video bloggers, Vloggers, talent streamers, and self-media people have emerged one after another. These professionals and ordinary users can upload and publish various videos at any time and place, and user-generated content has gradually shifted from text form to video form, giving rise to the research on video sentiment analysis.

[0007] There are mainly two emotional analysis theoretical models in the field of psychology, including the discrete emotion classification model and the dimensional emotion classification model. The discrete emotion classification model represents emotions as individual independent labels, and there is no correlation between each emotion. Ekman was the first to demonstrate the correlation between facial expressions and emotions. Through cross-cultural research, it was shown that people in different cultural environments perceive certain basic emotions in the same way, and based on this, a classification model based on 6 basic emotions was proposed. The dimensional emotion model represents more fine-grained emotions through multiple dimensions. The commonly used three-dimensional model defines emotions through axes and poles, and distributes emotions at different positions between the two poles of each axis. The classic models include the PAD model and the inverted cone three-dimensional emotion model. Currently, using the discrete emotion model for emotion analysis is still the most popular method in the field of multimodal emotion computing.

[0008] The target objects of video emotion analysis include multimodal information such as text, audio, and vision separated from the video. In order to use video emotion analysis technology to extract emotion polarity from these multimodal data, existing technical methods usually consider that all multimodal features have a promoting effect on emotion analysis. First, various feature extraction models are used to obtain the original representations of each modality, and then a linear classifier or a multilayer perceptron (MLP) is used to integrate all modality features to identify the emotions in the video; technical methods such as MMMU-BA and DEAN focus on the innovation of multimodal fusion methods, using cross-modal attention mechanisms to fuse aligned multimodal feature sequences to achieve discourse-level emotion analysis; the CIA technical method introduces an autoencoder network to guide the relationship between different modalities of the model.

[0009] However, the above methods can only capture cross-modal interaction information from one angle, without considering that there may be information redundancy and mutual interference between different modalities. Improper multimodal information fusion may bring noise features during emotion analysis. Therefore, the above methods still have limitations. To make up for the deficiencies of the above traditional video emotion analysis methods, this method constructs a suitable multimodal fusion method according to the characteristics of each modality in the video segment and establishes an emotion polarity analysis model, focusing on the data features most relevant to emotion analysis in the multimodal video sequence, and jointly learning the correlation information between modalities. Summary of the Invention

[0010] To overcome the technical defect that the existing methods can only capture cross-modal interaction information from one angle and bring noise features during emotion analysis, the present invention provides a video emotion polarity analysis method based on hierarchical fusion of multimodal deep features.

[0011] The present invention discloses a video emotion polarity analysis method based on hierarchical fusion of multimodal deep features, which is implemented by a video processing unit and an emotion analysis unit, and includes the following steps:

[0012] S1: Original feature extraction:

[0013] The complete video is divided into multiple video segments by the video processing unit, and the multiple video segments are divided into training data, validation data, and test data based on the random sampling method; the facial expression data, voice signal data, and text caption data of the speaker in each video segment are collected and sent to the single-modal original feature extraction unit to obtain the original depth features of the three single-modal data.

[0014] S2: Construct a video segment sentiment polarity analysis model:

[0015] The video segment sentiment polarity analysis model includes a multi-modal feature hierarchical interaction fusion unit and a sentiment polarity discrimination unit. The multi-modal feature hierarchical interaction fusion unit includes a bottom-layer bimodal feature interaction module and a top-layer trimodal feature hierarchical fusion module. The top-layer trimodal feature hierarchical fusion module includes a pairwise bilinear gated fusion unit and a trimodal self-attention feed-forward fusion unit; first, the original depth features of the single-modal data are processed by the bottom-layer bimodal feature interaction module, introducing a pairwise attention mechanism to capture the semantic relationship between any two single-modal data; then, they are respectively processed by the pairwise bilinear gated fusion unit and the trimodal self-attention feed-forward fusion unit of the top-layer trimodal feature hierarchical fusion module, and finally, the multi-modal features after hierarchical interaction fusion are obtained; the multi-modal features are processed by the sentiment polarity discrimination unit, that is, the multi-modal features extracted by the multi-modal feature hierarchical interaction fusion unit are passed through the classification layer to calculate the sentiment probability distribution result of the target video segment, and the category corresponding to the maximum probability is the sentiment polarity type judged for the video segment.

[0016] S3: Model training and testing:

[0017] The constructed video segment sentiment polarity analysis model is trained using the training data; the validation data is used to evaluate the training effect of the video segment sentiment polarity analysis model during the training process. After continuous adjustment and optimization, the optimal video segment sentiment polarity analysis model is obtained; the optimal video segment sentiment polarity analysis model is tested using the test data, and the classification effect index of the final sentiment polarity type is calculated.

[0018] The technical solution provided by the present invention has the following advantages compared with the prior art: the recognition effect of the sentiment polarity of the speaker in the video segment is better and the recognition rate is higher; while fully integrating multi-modal features, redundant and noise information in the fusion features is filtered, which improves the representation ability of multi-modal sentiment features to a certain extent and can effectively identify the positive and negative sentiment polarities of the speaker in the video segment. Description of the Drawings

[0019] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0021] Figure 1 It is a schematic flowchart of the video emotion polarity analysis method based on multi-modal deep feature hierarchical fusion provided by an embodiment of the present invention;

[0022] Figure 2 It is a technical roadmap of the video emotion polarity analysis method based on multi-modal deep feature hierarchical fusion provided by an embodiment of the present invention;

[0023] Figure 3 It is a schematic overall framework diagram of the video emotion polarity analysis method based on multi-modal deep feature hierarchical fusion provided by an embodiment of the present invention;

[0024] Figure 4 It is a schematic diagram of a three-modal self-attention feed-forward fusion unit provided by an embodiment of the present invention;

[0025] Figure 5 It is a schematic diagram of a pairwise bilinear fusion unit provided by an embodiment of the present invention;

[0026] Figure 6 It is a schematic diagram of a multi-modal gating mechanism provided by an embodiment of the present invention. Detailed implementation manners

[0027] In order to be able to more clearly understand the above objects, features, and advantages of the present invention, the following will further describe the solutions of the present invention. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.

[0028] In the description, it should be noted that the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. It should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific situations.

[0029] In the following description, many specific details are set forth in order to provide a thorough understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present invention, rather than all the embodiments.

[0030] The following specifically describes the embodiments of the present invention with reference to the accompanying drawings.

[0031] In an embodiment of the present invention, a video emotion polarity analysis method based on multi-modal depth feature hierarchical fusion is disclosed, which is implemented by a video processing unit and an emotion analysis unit, and includes the following steps:

[0032] S1: Original feature extraction: The complete video is divided into multiple video segments by the video processing unit, and the multiple video segments are divided into training data, validation data, and test data based on the random sampling method; the facial expression data, speech signal data, and text caption data of the speaker in each video segment are collected and sent to the single-modal original feature extraction unit to obtain the original depth features of the three single-modal data.

[0033] S2: Construct a video segment emotion polarity analysis model: The video segment emotion polarity analysis model includes a multi-modal feature hierarchical interaction fusion unit and an emotion polarity discrimination unit. The multi-modal feature hierarchical interaction fusion unit includes a bottom-layer bimodal feature interaction module and a top-layer trimodal feature hierarchical fusion module. The top-layer trimodal feature hierarchical fusion module includes a pairwise bilinear gating fusion unit and a trimodal self-attention feed-forward fusion unit; first, the original depth features of the single-modal data are processed by the bottom-layer bimodal feature interaction module, introducing a pairwise attention mechanism to capture the semantic relationship between any two single-modal data; then, they are respectively processed by the pairwise bilinear gating fusion unit and the trimodal self-attention feed-forward fusion unit of the top-layer trimodal feature hierarchical fusion module, and finally, the multi-modal features after hierarchical interaction fusion are obtained; the multi-modal features are processed by the emotion polarity discrimination unit, that is, the multi-modal features extracted by the multi-modal feature hierarchical interaction fusion unit are passed through a classification layer. In a specific embodiment, the classification layer is a fully connected layer and a Softmax normalization layer to calculate the emotion probability distribution result of the target video segment, and the category corresponding to the maximum probability is the emotion polarity type judged for the video segment.

[0034] S3: Model training and testing: Use the training data to train the constructed video segment emotion polarity analysis model; use the validation data to evaluate the training effect of the video segment emotion polarity analysis model during the training process, and after continuous adjustment and optimization, obtain the optimal video segment emotion polarity analysis model; use the test data to test the optimal video segment emotion polarity analysis model and calculate the classification effect index of the final emotion polarity type.

[0035] Based on the above embodiments, in a preferred embodiment, in step S1, when dividing the complete video, the complete video is divided into units of utterances, where an utterance is defined as multiple speaking segments segmented according to the pauses or sentence breaks in the speech signal data in the video; the facial expression data, speech signal data, and text subtitle data respectively correspond to the visual modality, audio modality, and text modality. The facial expression data of the speaker in the video segment is obtained from the camera, the facial expression data of the speaker in the video segment is obtained from the microphone, and the text subtitle data of the video segment is obtained from the subtitle document. The steps for extracting the original depth features of the three single-modal data are as follows:

[0036] S11. To extract the three-modal features of each segment, align the data of the visual modality and audio modality to the text modality to make the time step lengths of the visual modality, audio modality, and text modality consistent. Since in various languages, the most basic language component is the word, align the subsequences of the three modalities at the word level. Specifically, use the P2FA tool to obtain the start and end timestamps of each word in the text modality, and then align them to the visual modality sequence and audio modality sequence accordingly, so that the subsequences of the visual modality, audio modality, and text modality are aligned at the word level to obtain a video segment with the three single-modal data aligned.

[0037] S12. Use different depth feature extraction methods to extract features from the video segment with the three single-modal data aligned, and convert the heterogeneous multimodal information with different forms and sources into dense feature vectors that can be understood by the computer. That is, use word embedding technology and a CNN network to extract the original depth features of the text modality, use common speech analysis frameworks such as OpenSMILE or COVAREP to extract the original depth features of the audio modality, and use common visual analysis frameworks such as FACET or 3D-CNN to extract the original depth features of the visual modality.

[0038] Based on the above embodiments, in a preferred embodiment, in step S2, the processing process of the underlying bimodal feature interaction module is as follows: capture the internal temporal dependencies of the original depth features of each unimodal data, encode the original depth features respectively, then use the Dense layer to map the original depth features of each unimodal data to a common semantic space to eliminate the semantic gap. Finally, combine the original depth features of each unimodal data in pairs and send them into the pairwise attention mechanism unit to learn the interaction dependencies between bimodals. Independently train the three bimodal combinations of text modality - audio modality, audio modality - visual modality, and text modality - visual modality. Finally, extract the hidden layer features of the three bimodal combinations output by the pairwise attention mechanism unit as the initial input of the high-level module; the processing process of the high-level trimodal feature hierarchical fusion module is as follows: use the trimodal self-attention feed-forward fusion unit to first filter out the noise features of the hidden layer features of the three bimodal combinations, then extract the trimodal features after removing the noise, and then use the pairwise bilinear gating fusion unit to first obtain the dependencies between the trimodal features and then perform feature filtering, only retaining the multimodal features related to sentiment analysis.

[0039] Based on the above embodiments, in a preferred embodiment, dividing multiple video segments into training data, validation data, and test data based on the random sampling method means dividing the divided video segments into total training data and test data according to a ratio of 8:2, and then dividing the total training data into training data and validation data according to a ratio of 8:2.

[0040] Based on the above embodiments, in a preferred embodiment, when the underlying bimodal feature interaction module processes, select the cross-modal context attention unit as the pairwise attention mechanism unit, and its specific steps are as follows:

[0041] S211: Pass the original depth features of each unimodal data through a bidirectional gated recurrent unit to capture the temporal dependency relationship within a single modality, and obtain the segment feature representation containing context information;

[0042] S212: Pass the segment feature representation through a fully connected layer with non-linear activation, project the utterance features of each modality to a common semantic space with a dimension of D, and obtain the vector representations of the text modality, audio modality, and visual modality in the common semantic space;

[0043] S213: Combine the original depth features of each unimodal data in pairs, and use three pairs of cross-modal context attention units to fuse the bimodal information to obtain the feature representation vectors after fusing the three bimodal information;

[0044] S214: Encode the original depth features of each unimodal data using step S211. After concatenating the feature representation vectors of the three fused bimodal information with the original depth features of the encoded unimodal data respectively, use them as the input of the sentiment polarity discrimination unit. After passing through a fully connected layer and a Softmax normalization layer, obtain the emotion probability distribution of the target utterance;

[0045] S215: Apply steps S211 - S214 to independently train the three bimodal combinations of text modality - audio modality, audio modality - visual modality, and text modality - visual modality respectively. Finally, use the three hidden layer features output by the cross - modal context attention unit as the initial input of the high - level module.

[0046] In a specific embodiment, the feature representation vector of step S213 is , and the specific calculation steps of the cross - modal context attention unit are:

[0047]

[0048]

[0049]

[0050]

[0051]

[0052] Among them, , respectively represent two different modalities, represents matrix multiplication, represents the Hadamard product, represents the feature vector after fusing the text modality and the visual modality.

[0053] In a specific embodiment, the specific calculation steps of the emotion probability distribution of the target utterance in step S214 are: ;

[0054] Among them, and are respectively the learnable weight matrix and bias term in the fully connected layer; represents the probability distribution of the emotional type corresponding to the target video segment.

[0055] In a specific embodiment, when training the underlying module, the important network parameter values are set as follows: the number of neurons in the hidden layer is set to 100, the learning rate is set to 1e-3, the batch size is set to 16, a total of 100 epochs are trained, and the dropout rate is set to 0.5; the Adam optimizer is used to train the network. Compared with the stochastic gradient descent method, the Adam optimizer is simple to implement and computationally efficient.

[0056] Based on the above embodiment, in a preferred embodiment, the three-modal self-attention feed-forward fusion unit includes a self-attention filtering layer and a feed-forward network shallow fusion layer connected to each other. The paired bilinear gating fusion unit includes a paired bilinear fusion unit and a multi-modal gating output layer connected to each other. The specific steps for processing the high-level three-modal feature hierarchical fusion module are as follows: the self-attention filtering layer filters the noise features of the hidden layer features of the three bimodal combinations of text modality-audio modality, audio modality-visual modality, and text modality-visual modality respectively, to obtain the self-attention filtered bimodal representation. Then, the feed-forward network shallow fusion layer is used to extract the noise-removed bimodal representation. The self-attention filtered bimodal representation is concatenated with the feed-forward fused bimodal representation to promote gradient backpropagation, and finally the three-modal feature representation after self-attention feed-forward fusion is obtained; the paired bilinear fusion unit is used to obtain the dependencies between the three-modal feature representations, and then the multi-modal gating output layer adaptively learns the proportion of different input features, activates the features useful for sentiment classification, filters the features, and only retains the multi-modal features most relevant to sentiment analysis. The specific strategy for the multi-modal gating layer to assign different weights to different input features is as follows:

[0057] ;

[0058] Among them, respectively represent different input features of the multi-modal gating output layer, 、 and represent the weights of each input information adaptively learned by using three independent double-layer non-linear feed-forward neural networks; after assigning the weights to each input feature, an average operation is performed in the feature dimension. The average operation maximally retains the feature information while reducing the dimension; finally, the three-modal features after bilinear gating fusion are obtained .

[0059] The multi-modal gating layer can adaptively learn the proportions of different input features, and activate the features useful for sentiment classification. Specifically, for the features after bilinear fusion of the input, the multi-modal gating output mechanism of the multi-modal gating layer can assign different weights to different input features, eliminate redundant features and noise features, and improve the discriminability of sentiment features. The multi-modal gating mechanism is widely used in multi-modal sentiment analysis tasks. For example, the DEAN technique uses this mechanism to calculate the importance of different modalities and weighted control the output of each target modality.

[0060] In some embodiments, first, all the features after bilinear fusion are concatenated to obtain , and then three independent two-layer non-linear feed-forward neural networks are used to adaptively learn the weights of each input information. After assigning the weights to each input feature, an average is taken in the feature dimension. The average operation maximally retains the feature information while reducing the dimension. Finally, the three-modal features after bilinear gating fusion are obtained. The specific calculation process is as follows:

[0061] ;

[0062] ;

[0063] ;

[0064] ;

[0065] ;

[0066] Among them, represents the sigmoid activation function, , and , are the weight matrix and bias term of the first-layer feed-forward network. , are the weight matrices of the second-layer feed-forward network.

[0067] This module introduces two levels of three-modal fusion mechanisms, which are abbreviated as Tri-SAFFU and PBGFU respectively. Tri-SAFFU first uses the self-attention mechanism to filter the noise in the bimodal features obtained by training the underlying bimodal feature interaction module, and then uses feed-forward fusion to obtain the three-modal feature representation. PBGFU uses the pairwise bi-linear fusion (PBF) module to obtain the dependencies between three-modal features, and then uses the gated output module (GOM) for feature filtering, only retaining the features most relevant to sentiment analysis.

[0068] In some embodiments,

[0069] S221: Apply steps S211 - S215 to obtain three bimodal input feature matrices. Encode the temporal dependencies of the three bimodal input feature matrices through BiGRU, so that a specific segment contains information about its context before and after; add a fully connected layer with non-linear activation to project the discourse features of each bimodal into a common feature space with a dimension of D;

[0070] S222: Obtain the trimodal fine-grained sentiment representation by passing the bimodal vectors in the common feature space through two levels of fusion units, specifically: a trimodal self-attention feed-forward fusion unit and a pairwise bilinear gated fusion unit.

[0071] In step S221, three groups of bimodal feature matrices are obtained through BiGRU and non-linear fully connected encoding. On this basis, self-attention operations are performed on the three feature matrices respectively to remove redundant components in the bimodal interaction information. The specific process is as follows:

[0072] ;

[0073] ;

[0074] ;

[0075] Combine the self-attentioned bimodal feature representations in pairs, concatenate them in the feature dimension, and then pass through a fully connected layer to achieve the feed-forward fusion of trimodal features. The specific process is as follows:

[0076] ;

[0077] ;

[0078] .

[0079] where , and are the weight matrices of the fully connected layer.

[0080] In some embodiments, inspired by the residual network, the bimodal representation after feed-forward fusion is concatenated with the bimodal representation after self-attention filtering to facilitate gradient backpropagation. Finally, the trimodal feature representation after self-attention feed-forward fusion is obtained , specifically:

[0081] .

[0082] In specific embodiments, obtaining the dependencies between the tri-modal feature representations using paired bilinear fusion units specifically includes that the paired bilinear fusion units use a low-rank bilinear model to embed the input bimodal feature matrix into a new feature space through a non-linear Dense layer; applying the Hadamard product to approximate the bilinear model for sufficient feature interaction, and adding a self-attention mechanism to associate context information to further improve the sentiment feature representation.

[0083] The low-rank bilinear model is widely used in classification tasks. To reduce the computational cost, it uses the Hadamard product to approximate the bilinear model. The specific calculation process is as follows:

[0084] ;

[0085] ;

[0086] ;

[0087] where ⊙ represents the Hadamard product, and all input bimodal pairs share the weight parameters and .

[0088] In some embodiments, to further improve the sentiment feature representation and enhance the model's ability to recognize sentiment, this module uses self-attention after bilinear calculation to associate context information. Taking as an example, the specific calculation process is as follows:

[0089] ;

[0090] ;

[0091] ;

[0092] Similarly, taking and as inputs, we can obtain and .

[0093] In some embodiments, the self-attention-filtered bimodal representation and the feed-forward-fused bimodal representation, that is, and are concatenated. At the same time, inspired by the residual network, to facilitate the backpropagation of the model gradient, the input features , and are skip-connected to the output of this module. Finally, the sentiment feature representation for sentiment discrimination is:

[0094] ;

[0095] For the sentiment polarity classification task, After passing through a fully connected layer and a Softmax normalization layer, the sentiment probability distribution of the target video segment is obtained, as shown in the formula:

[0096] ;

[0097] Where and are the learnable weight matrix and bias term in the fully connected layer, respectively. represents the probability distribution of the emotion label corresponding to the th segment in the video.

[0098] In some embodiments, to fairly compare the performance of the models, all comparison models use the binary cross-entropy (Cross Entropy, CE) loss function to train the sentiment polarity classification model. The specific formula is:

[0099] ;

[0100] Where, represents the number of videos, represents the number of segments in the th video, represents the probability that the model predicts the th segment of the th video as a positive example, represents the probability that the model predicts the th segment of the th video as a negative example. represents the true label of the th segment of the th video. If the segment is a positive example, the value is 1; otherwise, it is 0.

[0101] In some embodiments, to prevent overfitting caused by the mismatch between the data volume and the neural network parameters, a Dropout layer is added to alleviate the overfitting phenomenon;

[0102] In some embodiments, during training, the important network hyperparameter values are set as follows: the learning rate is set to 1e-3, the dropout rate of the Dense layer and the Dropout rate of the BiGRU layer are set to 0.7 and 0.5 respectively, the number of hidden layer units is set to 300, and the batch size is set to 16; in the experiment, the early stop strategy (Early Stop) is adopted, and the training is stopped when the number of times the Loss value on the test set no longer decreases accumulates to the set threshold. The thresholds for the two datasets in the experiment are both set to 10; during the process of training the neural network, the Adam optimizer based on stochastic gradient descent is used to optimize the model parameters.

[0103] Based on the above embodiments, in a preferred embodiment, the model training and testing method includes:

[0104] S41: Use the training data to train the constructed video clip sentiment polarity analysis model, and adjust the model parameters according to the training results;

[0105] S42: Use the validation data to evaluate the training effect of the video clip sentiment polarity analysis model with adjusted parameters during the training process, for adjusting and optimizing the model. Stop training when the number of times the model's performance on the validation data no longer improves accumulates to a set threshold, and obtain the optimal video clip sentiment polarity analysis model;

[0106] S43: Use the test data to test the optimal video clip sentiment polarity analysis model, obtain the sentiment polarity analysis result of the video clip to be analyzed, and calculate the classification effect index of the final sentiment polarity type.

[0107] In some embodiments, to measure the sentiment analysis effect of the present invention, the experiment uses two evaluation indicators, namely classification accuracy Accuracy and Weight-Avg-F1 value, to evaluate the model's effect, which are abbreviated as Acc and F1 respectively. At the same time, a classification result confusion matrix is introduced to intuitively show the quality of the model classification. To reduce the randomness of the experimental process and illustrate the stability of the performance of the method proposed in the present invention, ten fixed random seeds are selected for the experiment. Finally, the average values and standard deviations of Acc and F1 values in the ten experiments are used as the experimental results. A small standard deviation indicates that the model has a relatively consistent performance in multiple experiments, and the model performance is more stable.

[0108] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Although the above embodiments have been described in detail, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the above embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the above embodiments, and they should all be covered by the protection scope of the claims.

Claims

1. A video emotion polarity analysis method based on multi-modal deep feature hierarchical fusion, which is implemented by a video processing unit and an emotion analysis unit, is characterized in that Including the following steps: S1: Original feature extraction: The complete video is divided into multiple video segments by the video processing unit, and the multiple video segments are divided into training data, validation data, and test data based on the random sampling method; the facial expression data, voice signal data, and text caption data of the speaker in each video segment are collected and sent to the single-modal original feature extraction unit to obtain the original depth features of the three single-modal data. S2: Constructing a video segment sentiment polarity analysis model: The video segment sentiment polarity analysis model includes a multi-modal feature hierarchical interaction and fusion unit and a sentiment polarity discrimination unit. The multi-modal feature hierarchical interaction and fusion unit includes a bottom-layer bimodal feature interaction module and a top-layer trimodal feature hierarchical fusion module. The top-layer trimodal feature hierarchical fusion module includes a pairwise bilinear gating fusion unit and a trimodal self-attention feed-forward fusion unit; first, the original depth features of the single-modal data are processed by the bottom-layer bimodal feature interaction module, introducing a pairwise attention mechanism to capture the semantic relationship between any two single-modal data; then, they are processed by the pairwise bilinear gating fusion unit and the trimodal self-attention feed-forward fusion unit of the top-layer trimodal feature hierarchical fusion module respectively, and finally, the multi-modal features after hierarchical interaction and fusion are obtained. The multi-modal features are processed by the sentiment polarity discrimination unit, that is, the multi-modal features extracted by the multi-modal feature hierarchical interaction and fusion unit are passed through the classification layer to calculate the sentiment probability distribution result of the target video segment, and the category corresponding to the maximum probability is the sentiment polarity type judged for the video segment. S3: Model training and testing: The constructed video segment sentiment polarity analysis model is trained using the training data; the training effect of the video segment sentiment polarity analysis model during the training process is evaluated using the validation data. After continuous adjustment and optimization, the optimal video segment sentiment polarity analysis model is obtained; the optimal video segment sentiment polarity analysis model is tested using the test data, and the classification effect index of the final sentiment polarity type is calculated.

2. The video emotion polarity analysis method based on multi-modal depth feature hierarchical fusion according to claim 1, wherein In step S1, when dividing the complete video, the complete video is divided into units of utterances, where an utterance is multiple speaking segments cut according to the pauses or sentence breaks in the voice signal data in the video; the facial expression data, voice signal data, and text caption data correspond to the visual modality, audio modality, and text modality respectively. The extraction steps of the original depth features of the three single-modal data are as follows: S11. Align the data of the visual modality and audio modality to the text modality to make the time step lengths of the visual modality, audio modality, and text modality the same. Use the P2FA tool to obtain the start and end timestamps of each word in the text modality, and then align them to the visual modality sequence and audio modality sequence accordingly, so that the subsequences of the visual modality, audio modality, and text modality are aligned at the word level to obtain the video segments with the three single-modal data aligned. S12. Use different deep feature extraction methods to extract features from the video segments aligned with the three types of unimodal data, that is, use word embedding technology and CNN network to extract the original deep features of the text modality, use the speech analysis framework to extract the original deep features of the audio modality, and use the visual analysis framework to extract the original deep features of the visual modality.

3. The method for video emotion polarity analysis based on multi-modal depth feature hierarchical fusion according to claim 2, wherein The processing process of the underlying bimodal feature interaction module is as follows: capture the internal temporal dependencies of the original deep features of each unimodal data, encode the original deep features respectively, then use the Dense layer to map the original deep features of each unimodal data to a common semantic space. Finally, combine the original deep features of each unimodal data in pairs and send them into the pairwise attention mechanism unit to learn the interaction dependencies between bimodalities. Independently train the three bimodal combinations of text modality - audio modality, audio modality - visual modality, and text modality - visual modality. Finally, take out the hidden layer features of the three bimodal combinations output by the pairwise attention mechanism unit as the initial input of the high-level module; The processing process of the high-level trimodal feature hierarchical fusion module is as follows: use the trimodal self-attention feed-forward fusion unit to first filter the noise features of the hidden layer features of the three bimodal combinations, then extract the trimodal features after removing the noise, and then use the pairwise bilinear gating fusion unit to first obtain the dependencies between the trimodal features, and then perform feature filtering, only retaining the multimodal features related to sentiment analysis.

4. The video emotion polarity analysis method based on multi-modal depth feature hierarchical fusion according to claim 3, characterized in that When the underlying bimodal feature interaction module is processed, select the cross-modal context attention unit as the pairwise attention mechanism unit, and its specific steps are as follows: S211: Pass the original deep features of each unimodal data through a bidirectional gated recurrent unit to capture the temporal dependency relationship within a single modality, and obtain the segment feature representation containing context information; S212: Pass the segment feature representation through a fully connected layer with non-linear excitation, project the discourse features of each modality to a common semantic space with a dimension of D, and obtain the vector representations of the text modality, audio modality, and visual modality in the common semantic space; S213: Combine the original deep features of each unimodal data in pairs, and use three pairs of cross-modal context attention units to fuse the bimodal information to obtain the feature representation vectors after fusing the three bimodal information; S214: Apply step S211 to encode the original deep features of each unimodal data, splice the feature representation vectors after fusing the three bimodal information with the encoded original deep features of the unimodal data in pairs as the input of the sentiment polarity discrimination unit, and pass through a fully connected layer and a Softmax normalization layer to obtain the emotion probability distribution of the target discourse; S215: Apply steps S211 - S214 to independently train the three bimodal combinations of text modality - audio modality, audio modality - visual modality, and text modality - visual modality respectively. Finally, take out the three hidden layer features output by the cross-modal context attention unit as the initial input of the high-level module.

5. The method for video emotion polarity analysis based on multi-modal deep feature hierarchical fusion according to claim 4, characterized in that The three-modal self-attention feedforward fusion unit includes a self-attention filtering layer and a feedforward network shallow fusion layer connected in sequence. The pairwise bilinear gated fusion unit includes a pairwise bilinear fusion unit and a multi-modal gated output layer connected in sequence. The specific steps for processing the high-level three-modal feature hierarchical fusion module are as follows: The self-attention filtering layer filters the noise features of the hidden layer features of the three bimodal combinations of text modality-audio modality, audio modality-visual modality, and text modality-visual modality respectively to obtain the self-attention filtered bimodal representations. Then, the feedforward network shallow fusion layer is used to extract the bimodal representations after removing the noise. The self-attention filtered bimodal representations are concatenated with the feedforward fused bimodal representations to finally obtain the three-modal feature representations after self-attention feedforward fusion; The pairwise bilinear fusion unit is used to obtain the dependencies between the three-modal feature representations. Then, the multi-modal gated output layer adaptively learns the proportions of different input features, activates the features useful for sentiment classification, and filters the features, only retaining the multi-modal features most relevant to sentiment analysis. The specific strategy for the multi-modal gated layer to assign different weights to different input features is as follows: ; Among them, respectively represent different input features of the multi-modal gating output layer, , and represent the weights of each input information adaptively learned by using three independent double-layer non-linear feedforward neural networks; after allocating each weight to each input feature, an average operation is performed in the feature dimension, and the average operation maximally retains the feature information while reducing the dimension; finally, the three-modal features after bilinear gating fusion are obtained .

6. The method for video emotion polarity analysis based on multi-modal depth feature hierarchical fusion according to claim 5, characterized in that When the sentiment polarity discrimination unit processes the multi-modal features, the classification layer is a fully connected layer and a Softmax normalization layer.

7. The method for video emotion polarity analysis based on multi-modal deep feature hierarchical fusion according to claim 6, characterized in that The model training and testing method includes: S41: Use the training data to train the constructed video clip sentiment polarity analysis model, and adjust the model parameters according to the training results; S42: Use the validation data to evaluate the training effect of the video clip sentiment polarity analysis model after adjusting the parameters during the training process, for adjusting and optimizing the model. When the number of times the model's performance on the validation data no longer improves accumulates to a set threshold, stop the training to obtain the optimal video clip sentiment polarity analysis model; S43: Use the test data to test the optimal video clip sentiment polarity analysis model to obtain the sentiment polarity analysis result of the video clip to be analyzed, and calculate the classification effect index of the final sentiment polarity type.

8. The method for video emotion polarity analysis based on multi-modal deep feature hierarchical fusion according to claim 7, wherein Dividing multiple video clips into training data, validation data, and test data based on the random sampling method means dividing the divided video clips into total training data and test data according to a ratio of 8:2, and then dividing the total training data into training data and validation data according to a ratio of 8:2.

Citation Information

Patent Citations

  • Multi-modal sentiment analysis model fusing sentiment resources

    CN115577161A

  • Method and apparatus for automatically generating inference questions and answers

    WO2021184311A1