A Video Emotion Classification Method Based on Gated Fusion and Multi-Task Learning

By introducing a multimodal Transformer architecture with gate mechanism and a multi-task learning network in the video emotion classification method, the problem of modal fusion ignoring single-modal information in the existing technology is solved, and a more efficient video emotion classification effect is achieved.

CN115203409BActive Publication Date: 2025-05-27BEIJING INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210732914.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-27
Publication Date
2025-05-27
Estimated Expiration
2042-06-27

AI Technical Summary

Technical Problem

The existing video emotion classification method ignores single-modal information when modal fusion, and does not fully consider the consistency and differences between modals, resulting in poor emotional classification results.

Method used

The video emotion classification method based on gated fusion and multi-task learning is adopted, and modal fusion is performed through a multi-modal Transformer architecture with gate mechanism, single-modal information is retained, and the multi-task learning network is combined to make full use of the complementarity of information between modes.

Benefits of technology

The integrity of multimodal global representation and the accuracy of singlemodal labels are improved, the effect of video sentiment classification is improved, and the model can better model the consistency and difference between modes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115203409B_ABST
    Figure CN115203409B_ABST
Patent Text Reader

Abstract

The present invention extracts feature vectors of text, pictures, and audio from a video; encodes the feature vectors of each modality using GRU to obtain vector representations of specific dimensions for each modality; uses a Transformer with a gating mechanism to fuse the information of each modality and splices the fused vectors as multi-modal vector representations; encodes the feature vectors of each modality using LSTM and a fully connected network to obtain vector representations of the conversion of each modality; calculates single-modal sentiment labels using the multi-modal vector representations, the vector representations of the conversion of each modality, and multi-modal sentiment labels; combines the multi-modal sentiment labels and single-modal labels for multi-task learning, and simultaneously performs multi-modal sentiment classification and single-modal sentiment classification. The video sentiment classification method provided by the present invention uses the fused multi-modal vector representations to participate in generating single-modal labels, improving the accuracy of single-modal labels; also adopts the method of multi-task learning, and simultaneously performs multi-modal sentiment classification and single-modal sentiment classification, enhancing the effect of video sentiment classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing and deep learning technology, and in particular to a video emotion classification method based on gated fusion and multi-task learning. Background Art

[0002] With the rapid development of the Internet, many users share their opinions on video websites such as Bilibili, Douyin, and Kuaishou by recording videos. For example, some users will upload videos on Bilibili to share their experience of using certain products. In this case, users urgently need a method to obtain the emotional polarity of the video publisher in order to judge the product so that other users can decide whether to buy the product. Different video clips may express different emotions, and the number of videos has increased dramatically over time. It is unrealistic to use manpower to analyze emotional polarity. Therefore, how to mine the implicit emotions from multimodal data has become an urgent problem to be solved.

[0003] Each source or form of information can be called a modality, and text, images, and audio separated from videos constitute the three most common modalities in reality. With the continuous growth of multimodal content on the Internet, the task of multimodal sentiment classification has emerged. Video sentiment classification is a type of multimodal sentiment classification. It is a process of separating text modality, image modality, and audio modality from videos, extracting features from the three modalities, and then using deep neural networks to fuse the three modalities and finally perform sentiment classification. As a core component of multimodal sentiment classification, video sentiment classification is of great significance and has gradually become a research hotspot in recent years.

[0004] Existing video sentiment classification methods include the Self-MM model proposed by Yu et al. and the SPT model proposed by Junyan et al. The Self-MM model uses a self-supervised generation strategy to generate unimodal labels without manual annotation, which solves the shortcoming that most data sets cannot perform multi-task learning and improves the generalization ability of the model. However, when the Self-MM model generates unimodal labels, its multimodal global representation vector only comes from the simple concatenation of three modal vectors, not multimodal information that has been deeply encoded, which makes the generated unimodal labels less accurate, which also limits the effect of sentiment classification. The SPT model reduces the number of parameters and improves the practicality of the model through parameter sharing and factorization. However, it only focuses on the fusion interaction between modalities and ignores the information of unimodal, resulting in incomplete information after fusion, which limits the effect of sentiment classification to a certain extent. Summary of the invention

[0005] In order to solve the problem that existing video emotion classification methods often only focus on the fusion interaction of modalities, ignore the information of a single modality, and do not fully consider the consistency and difference between modalities, the present invention provides a video emotion classification method and system based on gated fusion and multi-task learning.

[0006] To achieve the above object, the present invention proposes a video emotion classification method based on gated fusion and multi-task learning, the method comprising:

[0007] Input video, and obtain sentiment classification through sentiment classification model, wherein the sentiment classification model includes feature extraction module, GRU network, multimodal Transformer architecture based on gate mechanism, LSTM and third fully connected network, single modal label generation module, multi-task learning network and training module. The training method of sentiment classification model includes:

[0008] S1, transmitting the video input into the emotion classification model to the feature extraction module, the feature extraction module separates the video to obtain multimodal data, and then extracts and transforms the multimodal data features into a multimodal initial feature vector;

[0009] S2, input the multimodal initial feature vector into the GRU network to obtain a number of first single-modal feature vectors with a dimension of Z;

[0010] S3, inputting the first single-modal feature vector into the multi-modal Transformer architecture based on the gate mechanism to generate a modal fusion vector, and concatenating the modal fusion vectors to obtain a multi-modal final vector representation;

[0011] S4, inputting the multimodal initial feature vector into the LSTM and the first fully connected network to generate a second unimodal vector representation with a dimension of Z;

[0012] S5, inputting the multimodal final vector representation, the second unimodal feature vector representation and the multimodal emotion label into a unimodal label generation module to generate a unimodal emotion label;

[0013] S6, inputting the multimodal final vector representation and the second unimodal feature vector into a multi-task learning network, and outputting a multimodal sentiment prediction value and a unimodal sentiment prediction value;

[0014] S7. The multimodal emotion prediction value and the unimodal emotion prediction value, the known multimodal emotion label and the calculated unimodal emotion label are input into a training module for comparison, and the training module adjusts the parameters in the emotion classification model.

[0015] Furthermore, the multimodality includes text modality, picture modality and audio modality.

[0016] Furthermore, the Bert model is used to obtain the feature vector representation F of the wordt , use Facet and Openface to obtain the feature vector representation F of the image v , use the COVAREP library to obtain the feature vector representation F of the audio a .

[0017] Furthermore, in step S2, dimension Z is a text modality dimension.

[0018] Furthermore, the multimodal Transformer architecture with a gate mechanism includes a multi-head attention mechanism, a first mapping module, a residual connection and a normalization module, a first memory gate, a second memory gate and a hybrid gate, wherein:

[0019] The multi-head attention mechanism receives the first input vector A1 and the second input vector A2, and outputs a fusion vector M;

[0020] The first mapping module receives the first input vector A1 and the second input vector A2, transforms A1 and A2 into the same vector space and performs pooling, and then inputs the obtained vector into the first memory gate, the second memory gate and the mixing gate;

[0021] The first memory gate, the second memory gate, and the mixed gate respectively obtain the coefficients of the residual connection of the first input vector A1, the coefficients of the residual connection of the second input vector A2, and the coefficients of the residual connection of the fusion vector M through their respective fully connected networks and activation functions;

[0022] The residual connection and normalization module is used for residual connection and normalization calculation of the fusion vector M and the first input vector A1 and the second input vector A2, so as to obtain the modal fusion vector of the first input vector A1 and the second input vector A2.

[0023] Furthermore, in step S3, the method for generating the modality fusion vector includes:

[0024] S31, receiving two first unimodal feature vectors, represented by m1 and m2 respectively;

[0025] S32, the operations of the first layer of the first Transformer architecture and the second Transformer architecture include:

[0026] Input m1 as the main mode and m2 as the auxiliary mode as the first input vector and the second input vector into the first layer of the first Transformer architecture to obtain the fusion vector

[0027] Input m2 as the main mode and m1 as the auxiliary mode as the first input vector and the second input vector into the first layer of the second Transformer architecture to obtain the fusion vector

[0028] S33, the operations after the first layer of the first Transformer architecture and the second Transformer architecture include:

[0029] The fusion vector and As the first input vector and the second input vector input the first Transformer encoder of the i-th layer, get the fusion vector

[0030] The fusion vector and As the first input vector and the second input vector input the second Transformer encoder of the i-th layer, get the fusion vector and

[0031] Repeat the above steps until the modal fusion vector of the two first single-modal feature vectors is obtained.

[0032] Furthermore, the multi-head attention mechanism includes:

[0033] The modal fusion vector with m1 as the main modality output by the first Transformer i-1 layer And the modality fusion vector of the second Transformer layer i-1 with m2 as the main modality Mapping is performed to obtain matrices Q, K, and V, and then the multi-head attention mechanism is used to obtain the fusion vector representation M i , the formula is as follows:

[0034]

[0035]

[0036] The Q matrix is ​​given by The K and V matrices are obtained by mapping the eigenvector of The feature vector mapping is obtained, MH ATT represents the multi-head attention mechanism, Represents the similarity of the sequence elements of the main mode and the auxiliary mode. Q, K, and V are obtained by multiplying the input vector by the parameter matrix. The parameter matrix is ​​adjusted during the training of the sentiment classification model. k is the dimension of matrix K, W 0 , are parameter matrices, K m1 The transposed matrix of , h represents the number of heads of the multi-head attention mechanism, and the value range of j is 1≤j≤4.

[0037] Furthermore, in the residual connection and normalization module, the modality fusion vector representation is calculated The formula is as follows:

[0038]

[0039]

[0040] Among them, MLP represents a fully connected network, and Q represents the mapping from the input vector The matrix K represents the mapping from the input vector The matrix, M i represents the fusion vector, represents the output of the mixing gate, is the output of the first memory gate, is the output of the second memory gate, and LN represents normalization.

[0041] Furthermore, the multimodality is text modality, picture modality and audio modality, and in the step S3, the text first unimodal feature vector and the audio first unimodal feature vector, the text first unimodal feature vector and the picture first unimodal feature vector are fused and encoded.

[0042] Furthermore, the calculation steps of the unimodal emotion label in step S5 are as follows:

[0043] S51, calculate the central value of each single-mode positive state and negative state center value

[0044]

[0045]

[0046] Among them, ∑ represents the sum, N represents the number of samples, I pos,j represents the global representation of positive samples, I neg,j represents the global representation of negative samples, represents the vector representation after gated fusion, i represents each modality, i∈{a,t,v};

[0047] S52, calculating the relative distance between the second unimodal feature vector and the two center values:

[0048]

[0049]

[0050] in, is the second unimodal eigenvector, d i express The dimension of is the square value of the second norm, represents the central value of each modal positive state, represents the central value of each modal negative state, Represents the relative distance between each mode and the central value of the positive state, Indicates the relative distance of each mode from the central value of the negative state;

[0051] S53, calculation and The relative distance between i :

[0052]

[0053] in, Represents the relative distance between each mode and the central value of the positive state, Represents the relative distance between each mode and the center value of the negative state, ε is a constant with a value of 10 -8 , which is used to prevent the divisor from approaching 0 and causing calculation anomalies;

[0054] S54, according to the multimodal label value y m , unimodal label value y i , relative distance α i and known multimodal information α m It is a proportional relationship, and the single-mode label value y is calculated i :

[0055]

[0056] The video emotion classification method based on gated fusion and multi-task learning provided by the present invention has the following beneficial effects compared with the existing video emotion classification method:

[0057] 1. The present invention retains the original information of the main modality and the auxiliary modality while performing modal fusion, so that the obtained multimodal global representation information is more complete, thereby improving the effect of sentiment classification.

[0058] 2. The present invention improves the accuracy of the generated unimodal labels, enables multi-task learning to more fully model the consistency and difference between modalities, fully utilizes the complementarity of information between modalities, and improves the model effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0060] Figure 1 A schematic diagram of a flow chart of a video emotion classification method based on gated fusion and multi-task learning according to an embodiment of the present invention;

[0061] Figure 2 A schematic diagram of a flow chart of a video emotion classification method based on gated fusion and multi-task learning according to another embodiment of the present invention;

[0062] Figure 3 A schematic diagram of the structure of an improved Transformer architecture according to another embodiment of the present invention;

[0063] Figure 4 A schematic diagram of a process for obtaining a final modality fusion vector according to an embodiment of the present invention;

[0064] Figure 5 The figure is a flow chart of obtaining unimodal sentiment classification according to an embodiment of the present invention. DETAILED DESCRIPTION

[0065] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0066] In the modal fusion stage, the present invention uses a cross-modal Transformer with a gate mechanism to fuse the modalities. On the basis of the traditional Transformer encoder, two memory gates and a hybrid gate are added. While performing modal fusion, the information of the single modality is retained, and a complete and rich multimodal joint representation can be obtained. Then the present invention uses the gated fused vector as a global multimodal representation to participate in the generation of the single modal label, so that the generated single modal label is more accurate. Multimodal sentiment classification uses data in multimodal forms such as text, pictures and audio, and inputs it into a deep learning model to obtain the corresponding sentiment category. Single-modal sentiment classification only uses data of a single modality, such as inputting text into a neural network model to determine its corresponding emotional tendency. Then, a multi-task learning method is adopted to simultaneously perform multimodal sentiment analysis and single-modal sentiment analysis, making full use of the complementarity of information between modalities, thereby improving the model effect. Among them, multimodality refers to multiple single modalities.

[0067] The present invention provides a video emotion classification method based on gated fusion and multi-task learning. Figure 1 and 2 As shown, the method includes:

[0068] Input video, and obtain sentiment classification through sentiment classification model, where the sentiment classification model includes sequentially connected feature extraction module, GRU network, multimodal Transformer architecture based on gate mechanism, LSTM and third fully connected network, single modal label generation module, multi-task learning network and training module. The training method of sentiment classification model includes:

[0069] S1. The input video is transmitted to the feature extraction module, which separates the video to obtain multimodal data of the video, then extracts the features of each modal data and converts them into a multimodal initial feature vector. The multimodal feature vector contains multiple single-modal feature vectors.

[0070] S2. Input the multimodal initial feature vector into the GRU network, and encode the multimodal initial feature vector to obtain a plurality of first unimodal feature vectors with uniform dimensions. For example, the dimensions of the plurality of first unimodal feature vectors are all Z.

[0071] S3. Input multiple first single-modal feature vectors with unified dimensions into a multimodal Transformer architecture based on a gate mechanism to generate fusion codes, and concatenate the fusion codes to obtain a multimodal final vector representation.

[0072] For example, let the text vector be M t and the audio vector representation M a And the text vector representation M t and the image vector representation M v Perform fusion encoding and take the fusion vector output by the Transformer encoder and Concatenate to get the multimodal final vector representation F m .

[0073] S4. Input the multimodal initial feature vector basis into LSTM and the third fully connected network to generate a second unimodal feature vector with unified dimension for subsequent generation of unimodal labels.

[0074] For example, the initial feature vector representation F of each modality based on LSTM and the first fully connected network t 、F v and F a Further encoding obtains the second unimodal feature vector representation of each modality used to generate unimodal labels with the same dimension as the multimodal feature vector and

[0075] S5. Input the multimodal final vector representation, the second unimodal feature vector representation and the multimodal sentiment label into the unimodal label generation module, and the unimodal label generation module generates the unimodal sentiment label using a self-supervised strategy.

[0076] For example, the multimodal final vector F m , the second unimodal feature vector representation and multimodal sentiment labels y m Input the unimodal label generation module and use the self-supervised strategy to generate the unimodal sentiment label y i .

[0077] S6. Input the multimodal final vector representation and the second unimodal feature vector into the multi-task learning network, and output the multimodal sentiment prediction value and the unimodal sentiment prediction value.

[0078] S7. The multimodal emotion prediction value and the unimodal emotion prediction value, the known multimodal emotion label and the calculated unimodal emotion label are input into the training module for comparison. The training module adjusts all parameters of the emotion classification model through the back propagation algorithm and the like.

[0079] In step S1, the feature extraction module is used to extract a segment of the input video V i The separated multimodal data is used for feature extraction. The Bert model can be used to obtain the feature vector F of the word based on the features of the multimodal data. t , use Facet and Openface to obtain the feature vector F of the image v , use the COVAREP library to obtain the audio feature vector F a The text feature vector is represented by F ti Indicates that the image feature vector is represented by F vi Indicated by, the audio feature vector is represented by F ai Represents, where i represents the i-th segment of the input video data. These feature vector representations are collectively referred to as multimodal initial feature vectors.

[0080] In step S2, GRU (Gate Recurrent Unit) is a type of recurrent neural network (RNN). Like LSTM, it is also proposed to solve the problems of long-term memory and gradient in back propagation. GRU is a very effective variant of LSTM network. It has a simpler structure than LSTM network and has good effects. Therefore, it is also a very popular network at present.

[0081] Since the GRU network can effectively capture the long-distance semantic dependency information of the sequence, alleviate the gradient vanishing or gradient exploding problem, and its network structure is simpler and has fewer parameters than other sequence encoding networks such as LSTM, it can speed up the preprocessing before fusion. Therefore, in one embodiment, the GRU network is used to calculate the initial feature vector F of each modal input. t 、F v and F a The first unimodal feature vector M is obtained by encoding. t 、M v and M a The size of Z is 768, which is consistent with the text modality dimension after encoding by the pre-trained model Bert. Because the text modality has the greatest impact on sentiment, the first unimodal feature vector after encoding is used as the vector representation of each modality for subsequent fusion, which can facilitate the calculation of the attention mechanism in the gate mechanism Transformer. The usual method is to convert the initial feature vector F of each modality into t 、F v and F a Direct splicing will make it impossible to achieve the pairwise fusion of multimodal features, that is, the fusion of text modality and image modality, and the fusion of text modality and audio modality.

[0082] M i =ReLU(W i T (GRU(F i ))+b i )

[0083] Among them, i represents three modes, i∈{a,t,v}, ReLU represents the activation function, W i T and b i is a randomly initialized value that is automatically adjusted during training.

[0084] In step S3, a multimodal Transformer encoder architecture with a gate mechanism is used to fuse the multimodal feature vectors in pairs. The traditional Transformer encoder structure only has a self-attention mechanism, residual connection, and normalization calculation. For input vector A, vector B is obtained after the self-attention mechanism calculation. When the residual connection and normalization are performed for the first time, vector B is added to vector A element by element. The gate mechanism Transformer of the present invention adds a first mapping network, two memory gates, and a hybrid gate on the basis of the traditional Transformer encoder. Figure 3As shown. Among them, the multi-head attention mechanism of Transformer receives the first input vector and the second input vector and outputs a fusion vector. The first mapping network receives the first input vector and the second input vector, maps the first input vector and the second input vector to the same space, and then performs pooling and splicing. The spliced ​​vectors are respectively input into the first memory gate, the second memory gate and the mixed gate. The first memory gate outputs the proportional coefficient of the first input vector participating in the first residual connection and normalization calculation. The second memory gate outputs the proportional coefficient of the second input vector participating in the first residual connection and normalization calculation. The mixed gate outputs the proportional coefficient of the fused feature vector after multi-head attention participating in the first residual connection and normalization calculation. Then, the modal fusion vector is obtained through residual connection and normalization, which can better control the flow of information, so that the information of a single modality is retained while the modal fusion is performed, so that the information of the multi-modal joint representation vector output by the architecture is more complete, and the effect of the sentiment classification model is improved.

[0085] Both memory gates and the mixing gate include their own fully connected networks and ReLU activation functions, whose parameters are adjusted through the training module.

[0086] When fusing the first single-modal feature vectors pairwise, the Transformer architecture appears in pairs, with m1 and m2 used to represent the two modalities in each layer, and then the fusion is performed.

[0087] like Figure 4 As shown in the figure, the following takes the i-th layer Transformer encoder structure as an example to illustrate the fusion process. Specifically, it includes:

[0088] S31, the modal fusion vector with m1 as the main modality output from the previous layer and the modal fusion vector with m2 as the main mode Mapping is performed to obtain matrices Q, K, and V, and then the vector representation M is obtained by multi-head attention mechanism operation. i , the formula is as follows:

[0089]

[0090]

[0091] The Q matrix is ​​given by The K and V matrices are obtained by mapping the eigenvector of The feature vector mapping is obtained, MH ATT Represents multi-head attention mechanism.

[0092] Since the sources of Q, K and V matrices are not the same vector, this is not a strict self-attention mechanism. Represents the similarity of the sequence elements of the main mode and the auxiliary mode. The weight matrix on the left represents m 2 An element in the modal sequence is related to m 1 The greater the correlation of an element in the sequence of modalities, the greater the weight value. After Softmax normalization calculation, the attention score matrix is ​​calculated, and the value is between 0 and 1. The attention score matrix calculates the similarity between the two modalities. After Softmax calculation, it no longer contains the information of the modality itself. Therefore, the attention score matrix and the weight matrix V m1 Multiplying by and obtains the weighted result, which represents the mode m 2 From the mode m 1 The information of modal interaction, that is, the information after modal fusion, is learned from the vector; Q, K, and V are obtained by multiplying the input vector by the parameter matrix, which is adjusted during model training; d k is the dimension of matrix K, W 0 , are parameter matrices, K m1 The transposed matrix of , h represents the number of heads of the multi-head attention mechanism, and the value range of j is 1≤j≤4.

[0093] S32, the modal fusion vector output by the previous layer and Input the first mapping network to map it to the same vector space and concatenate it, that is, first input the modal fusion vector into the first fully connected network for mapping, and then perform average pooling calculation on its output, so as to map them to the same vector space to obtain the mapping vector representation and The formula is as follows:

[0094]

[0095]

[0096] in, and It is a randomly initialized value, which is automatically adjusted during training. The purpose of mapping it into the same vector space is to facilitate the calculation of subsequent gate mechanisms.

[0097] S33, the two mapping vectors are represented and The splicing input mixing gate and two memory gates are not part of the traditional Transformer encoder, but an additional part of the model of the present invention, which is used to control the proportion of the vector information after modal fusion and the vector information of a single modality that will continue to propagate on the network. The structure of the gate mechanism uses a fully connected layer plus a Relu activation function. The calculation formula is as follows:

[0098]

[0099]

[0100]

[0101] in, and is a randomly initialized value that is automatically adjusted during training. and Represents the output of the memory gate, which determines the proportion of the auxiliary modal and main modal information that will be subsequently propagated. The purpose of retaining single-modal information is to make the final modal representation information more complete. It represents the output of the mixing gate, which determines the proportion of the information after the two modes are mixed that will be subsequently propagated, thus better controlling the flow of information.

[0102] S34, the coefficient obtained in S33 and The input vectors of the previous layer (i.e., the i-1th layer of the Transformer) and the fusion vector obtained after multi-head attention are input into the residual connection and normalization module, in which the modal fusion vector of the i-th layer is calculated. The formula is as follows:

[0103]

[0104]

[0105] Among them, LN represents normalization, MLP represents a fully connected network, and Q represents the mapping from the input vector The matrix K represents the mapping from the input vector The matrix, M i represents the vector representation after modal fusion, so It indicates the proportion of the integrated information that will be subsequently disseminated. It indicates the proportion of the main modal information that will be subsequently disseminated. It indicates the proportion of auxiliary modal information that will be subsequently propagated. It is the output of the first Transformer architecture layer i with m2 as the main mode, and it also outputs the fusion vector of the second Transformer architecture layer i with m1 as the main mode in parallel. Next, these two fused vectors are input into the i+1th layer as the first input vector and the second input vector to continue fusion until the two outputs of the last layer are used as the final modal fusion vectors.

[0106] Initially, the first input vector is the first unimodal feature vector as the main modality, and the second input vector is the second unimodal feature vector as the auxiliary modality. After inputting the first Transformer architecture, the modal fusion vector is obtained. Then the first input vector is the second unimodal feature vector as the main modality, and the second input vector is the first unimodal feature vector as the auxiliary modality. After inputting into the second Transformer architecture, the modal fusion vector is obtained.

[0107] In the previous embodiment, four modal fusion vectors are obtained: and Then concatenate them to get the final multimodal vector representation F m The fusion process requires four gate mechanism Transformer encoders, each of which contains several layers.

[0108] In step S4, the multimodal initial feature vector F with dimension Z is t 、F v and F a , each eigenvector F t 、F v and F a Input into each LSTM for encoding to obtain the second unimodal feature vector and

[0109]

[0110] Among them, i represents three modes, i∈{a,t,v}, ReLU represents the activation function, and b i is a randomly initialized value that is automatically adjusted during training. and On the one hand, it is used to generate the sentiment prediction value of each single modality, and on the other hand, it participates in generating the sentiment label value of each single modality.

[0111] In step S5, according to the multimodal vector representation F m , the second unimodal eigenvector and multimodal sentiment labels y m Get each unimodal sentiment label y i The multimodal labels are already in the data set during model training. The purpose of calculating the unimodal labels is that multi-task learning requires simultaneous multimodal sentiment analysis and unimodal sentiment analysis, and unimodal sentiment analysis requires unimodal labels when calculating model loss. The method for calculating unimodal labels is proposed in the present invention. Figure 5 As shown, the calculation steps are as follows:

[0112] S51, calculate the central value of each single modal positive state and negative state center value

[0113]

[0114]

[0115] Among them, ∑ represents the sum, N represents the number of samples, that is, the total number of video clips in the dataset, and I pos,j represents the global representation of positive samples, that is, the feature vector representation of multimodal labels greater than 0 in each modality obtained in step S4, I neg,j represents the global representation of negative samples, that is, the feature vector representation of multimodal labels less than 0 in each modality obtained in step S4, It represents the vector representation after gated fusion, which is composed of the final multimodal feature vector F m After independent fully connected layer mapping, i represents multimodality and each single modality, i∈{m,a,t,v}. The fully fused vector can better represent the global information of multimodality, so the positive state center value and negative state center value generated are more accurate, which improves the accuracy of the subsequent single modality label;

[0116] S52, calculating the relative distance between the second unimodal feature vector and the two center values:

[0117]

[0118]

[0119] in, is the second unimodal eigenvector, d i yes The dimension of is the square value of the second norm, represents the central value of each modal positive state, represents the central value of each modal negative state, Represents the relative distance between each mode and the central value of the positive state, Indicates the relative distance of each mode from the central value of the negative state;

[0120] S53, calculation and The relative distance α i :

[0121]

[0122] Among them, ε is a constant with a value of 10-8 , which is used to prevent the divisor from approaching 0 and causing calculation anomalies;

[0123] S54, Multimodal Information α m The calculation process of α i Similarly, according to the multimodal label value y m , unimodal label value yi, relative distance α i and known multimodal information α m Is a proportional relationship, where the single-mode label value y is calculated i :

[0124]

[0125] In step S6, the multimodal final vector representation F m The first and second unimodal feature vectors are input into the multi-task learning network, and the multimodal sentiment classification and unimodal sentiment classification are output. The parameters of the multi-task learning network are adjusted by the training module.

[0126] In step S7, the multimodal sentiment classification and the unimodal sentiment classification are compared with the known multimodal sentiment labels and the calculated unimodal sentiment labels, and the parameters of the sentiment classification model are adjusted by a back propagation algorithm or the like.

[0127] Finally, it should be noted that the above embodiments are only used to describe the technical solution of the present invention rather than to limit the technical method. The present invention can be extended to other modifications, changes, applications and embodiments in application, and therefore it is believed that all such modifications, changes, applications and embodiments are within the spirit and teaching scope of the present invention.

Claims

1. A video emotion classification method based on gated fusion and multi-task learning, characterized in that, the method includes: inputting a video, and obtaining an emotion classification through an emotion classification model, where the emotion classification model includes a feature extraction module, a GRU network, a multi-modal Transformer architecture based on a gating mechanism, an LSTM and a third fully-connected network, a single-modal label generation module, a multi-task learning network, and a training module, the multi-modal Transformer architecture based on a gating mechanism includes a multi-head attention mechanism, a first mapping module, a residual connection and a normalization module, a first memory gate, a second memory gate, and a mixing gate, where, the multi-head attention mechanism receives a first input vector A1 and a second input vector A2, and outputs a fused vector M; the first mapping module receives the first input vector A1 and the second input vector A2, transforms A1 and A2 into the same vector space and performs pooling, and then inputs the obtained vector into the first memory gate, the second memory gate, and the mixing gate; the first memory gate, the second memory gate, and the mixing gate respectively obtain the coefficients for residual connection of the first input vector A1, the coefficients for residual connection of the second input vector A2, and the coefficients for residual connection of the fused vector M through their respective fully-connected networks and activation functions; the residual connection and normalization module is used for residual connection and normalization calculations of the fused vector M with the first input vector A1 and the second input vector A2, so as to obtain the modal fusion vectors of the first input vector A1 and the second input vector A2; the training method of the emotion classification model includes: S1. Transmit the video input into the emotion classification model to the feature extraction module. The feature extraction module separates the video to obtain multi-modal data, and then extracts and transforms the multi-modal data features into multi-modal initial feature vectors; S2. Input the multi-modal initial feature vectors into the GRU network to obtain a number of first single-modal feature vectors with a dimension of Z; S3. Input the first single-modal feature vectors into the multi-modal Transformer architecture based on a gating mechanism to generate modal fusion vectors. After the modal fusion vectors are concatenated, a multi-modal final vector representation is obtained; S4. Input the multi-modal initial feature vectors into the LSTM and the first fully-connected network to generate a second single-modal vector representation with a dimension of Z; S5. Input the multi-modal final vector representation, the second single-modal feature vector representation, and the multi-modal emotion label into the single-modal label generation module to generate single-modal emotion labels; S6. Input the multi-modal final vector representation and the second single-modal feature vector into the multi-task learning network, and output multi-modal emotion prediction values and single-modal emotion prediction values; S7. Input the multi-modal emotion prediction values, the single-modal emotion prediction values, the known multi-modal emotion labels, and the calculated single-modal emotion labels into the training module for comparison. The training module adjusts the parameters in the emotion classification model.

2. The method according to claim 1, characterized in that, the multi-modal is a text modality, a picture modality, and an audio modality.

3. The method according to claim 2, characterized in that, Obtain the feature vector representation F of words using the Bert model t , obtain the feature vector representation F of pictures using Facet and Openface v , obtain the feature vector representation F of audio using the COVAREP library a .

4. The method according to claim 2, characterized in that, In the said step S2, the dimension Z is the text modality dimension.

5. The method according to claim 1, wherein, in the said step S3, the method for generating the modality fusion vector includes: S31. Receive two first single-modality feature vectors, represented by m1 and m2 respectively; S32. The operations of the first layer of the first Transformer architecture and the second Transformer architecture include: Taking m1 as the main modality and m2 as the secondary modality as the first input vector and the second input vector, input them into the first layer of the first Transformer architecture to obtain a fused vector Taking m2 as the main modality and m1 as the auxiliary modality as the first input vector and the second input vector, input them into the first layer of the second Transformer architecture to obtain a fused vector S33. The operations after the first layer of the first Transformer architecture and the second Transformer architecture include: The fusion vector and are input as the first input vector and the second input vector into the first Transformer encoder of the i-th layer, and the fusion vector The fused vector and are input as the first input vector and the second input vector into the second Transformer encoder of the i-th layer, and the fused vector and Repeat the above steps until the modality fusion vector of the two first single-modality feature vectors is obtained.

6. The method according to claim 1, wherein, the said multi-head attention mechanism includes: The modality fusion vector with m1 as the main modality output by the (i-1)-th layer of the first Transformer and the modality fusion vector with m2 as the main modality output by the (i-1)-th layer of the second Transformer are mapped to obtain matrices Q, K, and V, and then the multi-head attention mechanism is used for operation to obtain the fused vector representation M i , and the formula is as follows: Among them, the Q matrix is mapped from the eigenvectors of , the K and V matrices are mapped from the eigenvectors of , MH ATT represents the multi-head attention mechanism, represents the similarity of the sequence elements of the main modality and the auxiliary modality. Q, K, and V are obtained by multiplying the input vectors by the parameter matrices, and the parameter matrices are adjusted during the training of the sentiment classification model. d k is the dimension of matrix K, W 0 , are both parameter matrices, represents the transpose matrix of K m1 , h represents the number of heads of the multi-head attention mechanism, and the value range of j is 1 ≤ j ≤ 4.

7. The method according to claim 1, wherein, In the residual connection and normalization module, calculate the modal fusion vector representation The formula is as follows: Among them, MLP represents a fully connected network, and Q represents a matrix mapped from the input vector ; K represents a matrix mapped from the input vector ; M i represents the fused vector, represents the output of the mixing gate, is the output of the first memory gate, is the output of the second memory gate, and LN represents normalization.

8. The method according to claim 7, wherein, the said multi-modalities are text modality, picture modality and audio modality, and in the said step S3, the text first single-modality feature vector and the audio first single-modality feature vector, and the text first single-modality feature vector and the picture first single-modality feature vector are fused and encoded.

9. The method according to claim 1, wherein, the calculation steps of the single-modality sentiment label in the said step S5 are as follows: S51. Calculate the central values of each single-modal positive state and the central values of the negative state Among them, ∑ represents summation, N represents the number of samples, and I pos,j represents the global representation of positive samples, and I neg,j represents the global representation of negative samples, represents the vector representation after gated fusion, where i represents each modality, and i ∈ {a, t, v}; S52. Calculate the relative distances between the second single-modality feature vector and two central values: Among them, is the second single-modal feature vector, and d i represents 's dimension, is the square value of taking the two-norm, represents the central value of the positive state of each modality, represents the central value of the negative state of each modality, represents the relative distance of each modality from the central value of the positive state, represents the relative distance of each modality from the central value of the negative state; S53. Calculate and the relative distance α i : Among them, represents the relative distance between each modality and the central value of the positive state, represents the relative distance between each modality and the central value of the negative state, and ε is a constant whose function is to prevent abnormal calculation caused by the divisor approaching 0; S54. According to the multi-modal tag value y m , the single-modal tag value y i , the relative distance α i and the known multi-modal information α m are in a proportional relationship, and the single-modal tag value y i is calculated as follows:

Citation Information

Patent Citations

  • Multi-modal fusion emotion recognition system and method based on multi-task learning and attention mechanism and experimental evaluation method

    CN113420807A

  • Method for estimating human emotions using deep psychological affect network and system therefor

    US20190347476A1