A Multimodal Sentiment Classification Method Based on Modal Space Assimilation and Contrastive Learning
Through TokenLearner module and supervision comparison learning, the problem of poor handling of modal balance and interaction relationships in the existing multimodal emotion analysis method is solved, and the assimilation of modal space and the accuracy of emotional classification is improved.
Patent Information
- Application Number
- CN202211139018.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-19
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-09-19
AI Technical Summary
The existing multimodal sentiment analysis methods are difficult to effectively balance the contribution of each modal to the final solution space, and ignore the interaction relationships and redundant information between modals, resulting in insufficient accuracy and generalization capabilities of sentiment analysis.
By introducing the TokenLearner module, the weight graph between modals is calculated using multi-head attention scores, and the information of the new vectors is complementary through orthogonal constraints, and the guide vector is constructed to assimilate the modal space. At the same time, supervised contrast learning is used to constrain multimodal representations, enhancing the model's expressive ability and distinctive ability.
Heterogeneous spatial assimilation between modals is realized, which avoids the problem of unbalanced contribution of each modal to the final solution space, and enhances the model's ability to explore multimodal emotional contexts and the accuracy of emotion classification.
Smart Images

Figure CN115310560B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of multi-modal emotion recognition in the cross-field of natural language processing, vision, and speech, and relates to a multi-modal emotion classification method based on modal space assimilation and contrast learning. Specifically, it is a method that uses a guiding vector to assimilate heterogeneous multi-modal spaces and constrains the obtained multi-modal representation through supervised contrast learning to judge the emotional state of the subject. Background Art
[0002] The field of emotion analysis usually includes data such as text, video, and audio. Previous studies have confirmed that these single-modal data usually contain discriminant information related to emotional states, and at the same time, it has been found that simply analyzing the data of a single modality often cannot obtain accurate emotion analysis. However, using the information of multiple modalities can ensure that the model can perform more accurate emotion analysis. By eliminating the singularity and uncertainty between modalities through the complementarity between modalities, the generalization ability and robustness of the model can be effectively enhanced, and the performance of the emotion analysis task can be improved.
[0003] Existing fusion models based on the attention mechanism construct a compact multi-modal representation by extracting information from each modality and perform emotion analysis based on this multi-modal representation. Therefore, it has received increasing attention from researchers. First, the attention coefficients between the information of the other two modalities (video and audio) and the text modality information are obtained through the attention mechanism, and then multi-modal fusion is performed according to the obtained attention coefficients. However, this ignores the interaction relationship existing between the information of multiple modalities. In addition, there are gaps between modalities and redundancies within each modality, both of which will increase the difficulty of learning the joint embedding space. However, existing multi-modal fusion methods rarely consider these two details, nor ensure that the information of the interacting multi-modalities is fine-grained, which has a certain impact on the final task performance.
[0004] Existing multi-modal fusion models based on transformation networks have great advantages in modeling time dependence, and the self-attention mechanism included can effectively solve the misalignment problem between multi-modal data, so it has attracted wide attention. This multi-modal fusion model obtains a cross-modal common subspace by transforming the distribution of the source modality into the distribution of the target modality and uses this as the multi-modal fusion information. In addition, a solution space is obtained during the process of transforming the source modality into another modality, which makes the solution space overly dependent on the contribution of the target modality. And when a certain modality data is missing, the solution space will lack the contribution from this modality data, which leads to the inability to effectively balance the contributions of each modality to the final solution space. On the other hand, existing transformation models usually only consider the transformation from text to audio and from text to video, and do not consider the possibility of other modality transformations, which has a certain impact on the final task performance.
[0005] Chinese Patent CN114722202A publicly proposes to use a bidirectional double-layer attention LSTM network to achieve multi-modal sentiment classification. The bidirectional attention LSTM network can explore more comprehensive temporal dependencies. Chinese Patent CN113064968A provides a sentiment analysis method based on a tensor fusion network, which uses a tensor network to model the interaction between modalities. However, it is difficult for the above two networks to effectively explore multi-modal sentiment contexts from long sequences, which may limit the expressive ability of the learning model. Chinese Patent CN114973062A discloses a multi-modal sentiment analysis method based on Transformer. This method uses a paired cross-modal attention mechanism to capture the interaction between multi-modal sequences across different time steps, thereby potentially mapping sequences from one modality to another. However, it ignores the redundant messages of the auxiliary modality, which increases the difficulty of effective reasoning about multi-modal messages. More importantly, the attention-based framework mainly focuses on static or implicit interactions between multi-modalities, which results in a relatively coarse-grained formation of multi-modal sentiment contexts. Summary of the Invention
[0006] The first object of the present invention is to address the deficiencies of the prior art and propose a multi-modal sentiment classification method based on modal space assimilation and contrast learning. A TokenLearner module is proposed to construct a guiding vector composed of complementary information between modalities. First, based on the multi-head attention scores of each modality, a weight map is calculated for each modality. Then, each modality is mapped to a new vector according to the obtained weight map, and orthogonal constraints are used to ensure that the information contained in these new vectors is complementary. Finally, the weighted average of the vectors is calculated to obtain the guiding vector. The learned guiding vector guides each modality to approach the solution space in parallel, which can make the heterogeneous spaces of the three modalities isomorphic. This strategy does not have the problem of unbalanced contributions of each modality to the final solution space and is suitable for effectively exploring more complex multi-modal sentiment backgrounds. To significantly improve the model's ability to distinguish various emotions, supervised contrast learning is used as an additional constraint when fine-tuning the model. With the help of label information, the model can capture more comprehensive multi-modal sentiment contexts.
[0007] The technical solution adopted by the present invention is as follows:
[0008] A fusion method based on modal space assimilation and contrast learning, comprising the following steps:
[0009] Step (1), obtaining multi-modal data:
[0010] Preprocess the multi-modal feature information, and extract the primary representations H t 、H a 、Hv ;
[0011] Step (2), construct a TokenLearner module to obtain a guiding vector:
[0012] Each modality m ∈ {t, a, v} is equipped with a TokenLearner module, where t, a, v represent the text, audio, and video modalities respectively; and these TokenLearner modules are reused in each guidance; the TokenLearner module calculates a weight map through the multi-head attention scores of the modality, and then obtains a new vector Z according to this weight map m :
[0013]
[0014]
[0015]
[0016] Z m = α m (MultiHead(H m , H m ))H m Equation (4)
[0017] where α m is a one-dimensional convolution layer followed by a softmax function, and are the weights of Q and K respectively, d k represents the dimension of H m , n represents the number of multi-heads; MultiHead(Q, K) represents the multi-head attention score; head i represents the i-th head attention score; Attention(Q, K) is a function to calculate the attention score; the superscript T represents transposing the matrix; Q, K are the two inputs of the function, which are the representations H m , H m .
[0018] To ensure that the information in Z m represents the complementary information of its corresponding modality, an orthogonality constraint is added to train each modality's TokenLeamer module, reduce redundant latent representations, and encourage the TokenLeamer module to encode different aspects of the multi-modal;
[0019] The orthogonality constraint is defined as:
[0020]
[0021] where represents the squared Frobenius norm;
[0022] By calculating the weighted average of Z m to obtain the guiding vector Z, which can be formulated as follows:
[0023]
[0024] where w m is the weight;
[0025] Step (3), guiding the modes close to the solution space:
[0026] According to the guiding vector Z obtained in step (2), parallelly guide the spaces where the three modes are located towards the solution space; during each guiding process, the guiding vector Z will be updated in real time according to the states of the spaces where the current three modes are located; more specifically, for the l-th guiding, the matrix representation after guiding each mode is as follows:
[0027]
[0028] where θ m represents the model parameters of the Transformer module, represents the concatenation of l and Z, and the guiding of the guiding vector Z for each mode is completed by the Transformer;
[0029] After expanding formula (7), it is specifically shown as:
[0030]
[0031] where MSA represents the multi-head self-attention module, LN represents the layer normalization module, and MLP represents the multi-layer perceptron;
[0032] Extract the last row data from the three-mode guided matrices obtained after L times of guiding and concatenate them into a multi-modal representation vector H final ; L represents the maximum number of guiding times;
[0033] Step (4), constraining the multi-modal representation vector H through supervised contrastive learning final :
[0034] Copy the hidden state of the multi-modal representation vector H final to form an augmented representation and remove its gradient; based on the above mechanism, after expanding N samples, there are 2N samples; it is expressed as follows:
[0035]
[0036]
[0037]
[0038] where represents the loss function of supervised contrastive learning, is the index of any sample in the multi-view batch, τ ∈ R + represents an adjustable coefficient for controlling class separation, P(i) is the set of samples that are different from z but have the same class, and A(i) represents all indices except i; SIM() is a function for calculating the similarity between samples.
[0039] Step (5), obtaining the classification result:
[0040] The multi-modal representation H final obtains the final prediction through a fully connected layer to implement multi-modal sentiment classification.
[0041] During training, the mean squared error loss is used to estimate the prediction quality during training:
[0042]
[0043] where y represents the true label;
[0044] The overall loss is composed of and weighted sum of, and is expressed as follows:
[0045]
[0046] where and respectively represent the loss function of the sentiment classification task, the orthogonal constraint loss function, and the loss function of supervised contrastive learning, and α, β, γ are respectively and weights.
[0047] The second object of the present invention is to provide an electronic device, which is characterized in that it includes a processor and a memory, the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the method.
[0048] The third object of the present invention is to provide a machine-readable storage medium, which is characterized in that the machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are called and executed by a processor, the machine-executable instructions cause the processor to implement the method.
[0049] The beneficial effects of the present invention are as follows:
[0050] The present invention introduces the concept of assimilation. A guiding vector is used to guide the spaces of each modality to approach the solution space simultaneously, enabling the assimilation of the heterogeneous spaces between modalities. This strategy does not have the problem of unbalanced contributions of each modality to the final solution space and is applicable to effectively exploring more complex multi-modal emotional contexts. At the same time, the steering vector guiding a single modality is composed of complementary information between multiple modalities, which can make the model pay more attention to emotional features, thereby naturally removing the intra-modal redundancy that increases the difficulty of obtaining multi-modal representations.
[0051] Combined with the dual learning mechanism and the self-attention mechanism, during the process of converting one modality to another, cross-modal fusion information with directional long-term interactions between modality pairs is mined. At the same time, the dual learning technology can enhance the robustness of the model, so it can well handle the inherent problem in multi-modal learning - the problem of missing modality data. Subsequently, a hierarchical fusion framework is constructed on this basis. The cross-modal fusion information with the same source modality is concatenated together, and a one-dimensional convolutional layer is further used for high-level multi-modal fusion, which is an effective supplement to the multi-modal fusion framework in the current field of emotion recognition. In addition, supervised contrast learning is introduced to help the model distinguish the differences between different categories, thereby achieving the purpose of improving the model's ability to distinguish different emotions. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is a flowchart of the present invention;
[0053] Figure 2 is an overall schematic diagram of step 3 of the present invention;
[0054] Figure 3 is a schematic diagram of the fusion framework of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0055] The method of the present invention will be described in detail below with reference to the accompanying drawings.
[0056] The method of the present invention is a multi-modal emotion classification method based on modality space assimilation and contrast learning, as Figure 1 shown, and includes the following steps:
[0057] Step 1: Obtain multi-modal information data
[0058] Under the condition that the subject performs a specific emotion task, the multi-modal data of the subject is recorded; the multi-modal includes text modality, audio modality, and video modality.
[0059] Step 2: Preprocess the multi-modal information data
[0060] Extract the primary features for each modality through a specific network:
[0061] Use BERT for the text modality;
[0062] Use Transformer for the audio modality and the video modality;
[0063] H t = BERT(T)
[0064] H a = Transformer(A)
[0065] H v = Transformer(V) Equation (1)
[0066] Where is the primary representation of the m-th modality, m ∈ {t, a, v}; t, a, v are the text, audio, and video modalities respectively; T, A, V are the original data of the text, audio, and video modalities; T m is the size of the time domain dimension, d m is the length of the feature vector at each moment;
[0067] Step 3, Construct a guiding vector to guide the modality space.
[0068] In the proposed multi-modal fusion framework, the TokenLearner module is one of the core processing modules. During the multi-modal fusion process, this module is designed for each modality to extract complementary information between modalities, thereby constructing a guiding vector to simultaneously guide each modality space towards the solution space, which ensures that each modality makes the same contribution to the final solution space.
[0069] First, calculate the multi-head attention score matrix MultiHead(Q, K) for each modality according to the multi-modal data H m (m ∈ {t, a, v}), then use one-dimensional convolution on this matrix and add the softmax function after convolution to obtain the weight matrix, and the number of rows of the weight matrix is much smaller than the number of rows of H m (m ∈ {t, a, v}). Multiply the weight matrix by the multi-modal data H m (m ∈ {t, a, v}) to extract the information Z m (m ∈ {t, a, v}):
[0070]
[0071]
[0072]
[0073] Z m = A m H m = α m (MultiHead(H m ,H m ))H m Equation (5)
[0074] where Attention(Q, K) is a function to calculate the attention score; the superscript T represents transpose; d k represents the dimension of H m .
[0075] Z m (m ∈ {t, a, v}) containing complementary information between modalities is averaged with weights to construct the guiding vector Z in the current situation.
[0076]
[0077]
[0078] Step 3 will be repeated multiple times. Each time, a new guiding vector Z will be generated according to the current situation of each modality to guide the modality space closer to the final solution space. At the same time, to ensure that the information extracted by the TokenLearner module is complementary between modalities, we will use an orthogonal constraint to train the three TokenLearner modules at the end:
[0079]
[0080] Step 4. Continue pre-training:
[0081] Based on Step 3, after multiple guidances, the last elements of the multi-modal data H m (m ∈ {t, a, v}) are extracted and integrated into a compact multi-modal representation H final . To make the model more easily distinguish various emotions, supervised contrastive learning is introduced to constrain the multi-modal representation H final . This strategy introduces label information. Under the condition of making full use of label information, samples with the same emotion are made cohesive, and samples with different emotions are mutually exclusive. Finally, the final fused information is input into the linear classification layer, and the output information is compared with the emotion category label to obtain the final classification result.
[0082] The present invention is compared with some superior fusion methods on two publicly available multimodal sentiment databases, CMU-MOSI and CMU-MOSEI. The CMU-MOSI (Multimodal Opinion Sentiment Intensity) dataset consists of 2,199 video clips collected from 93 opinion videos downloaded from YouTube. It contains the views of 89 different narrators on certain topics, and each clip of the video is manually labeled with an emotion intensity from -3 (strongly negative) to 3 (strongly positive).
[0083] Table 1 shows the results of the mean absolute error (MAE), correlation coefficient (Corr), accuracy for the binary sentiment classification task (Acc-2), F1-score (F1-Score), and accuracy for the seven-class sentiment classification task (Acc-7). Although Self-MM is better than other existing methods, it can still be observed from Table 1 the advantages and effectiveness of the present invention. On the CMU-MOSI dataset, the present invention outperforms the state-of-the-art Self-MM in all metrics. In addition, on the CMU-MOSEI dataset, the present invention is better than Self-MM, achieving an improvement of approximately 0.8% in Acc2 and 0.9% in F1-Score. Therefore, the effectiveness of the method proposed by the present invention is demonstrated.
[0084] Table 1. Comparison Table of Results
[0085]
Claims
1. A multi-modal sentiment classification method based on modal space assimilation and contrastive learning, characterized in that Including the following steps: Step (1), obtaining multimodal data: Preprocess the multi-modal feature information and extract the primary representations H t 、H a 、H v ; Step (2), constructing a TokenLearner module to obtain a guiding vector: Each modality m ∈ {t, a, v} is equipped with a TokenLearner module, where t, a, and v represent the text, audio, and video modalities respectively; and these TokenLearner modules are reused in each guidance; the TokenLearner module calculates a weight map through the multi-head attention scores of the modality, and then obtains a new vector Z according to this weight map m : Z m = α m (MultiHead(H m , H m ))H m Equation (4) where α m is a one-dimensional convolution layer followed by a softmax function, and are the weights of Q and K respectively, and d k represents the dimension of H m ; n represents the number of heads; MultiHead(Q, K) represents the multi-head attention score. head i represents the attention score of the $i$-th head; Attention(Q, K) is a function for calculating the attention score; To ensure that the information in Z m represents complementary information for its corresponding modality, an orthogonality constraint is added to train the TokenLearner module for each modality, reducing redundant latent representations and encouraging the TokenLearner module to encode different aspects of the multimodality; The orthogonality constraint is defined as: wherein represents the squared Frobenius norm; By calculating the weighted average of Z m to obtain the guiding vector Z, which can be expressed by the following formula: where w m is the weight; Step (3), guiding the modality towards the solution space: According to the guiding vector Z obtained in step (2), the spaces where the three modalities are located are guided in parallel towards the solution space; during each guiding process, the guiding vector Z is updated in real time according to the states of the spaces where the three modalities are located at present; more specifically, for the l-th guiding, the matrix representation after guiding each modality is as follows: where θ m represents the model parameters of the Transformer module, denotes the concatenation of l and Z, and the guidance of the guidance vector Z for each modality is completed by the Transformer; Specifically shown after expanding formula (7): Where MSA represents the multi-head self-attention module, LN represents the layer normalization module, and MLP represents the multi-layer perceptron; Extract the last row data from the three modal guidance matrices obtained after L times of guidance, and concatenate them into a multi-modal representation vector H final ; L represents the maximum number of guidance times; Step (4), constraining the multi-modal representation vector H through supervised contrastive learning final : Copy the multimodal representation vector H final 's hidden state to form an augmented representation and remove its gradient; Based on the above mechanism, after expanding N samples, there will be 2N samples; It is expressed as follows: Among them represents the loss function of supervised contrastive learning is the index of any sample in the multi-view batch, τ ∈ R + represents an adjustable coefficient for controlling class separation, P(i) is the set of samples that are different from i but have the same class, and A(i) represents all indices except i; SIM() is a function for calculating the similarity between samples Step (5), obtaining the classification result: Multimodal representation H final Obtain the final prediction through the fully connected layer Achieve multimodal sentiment classification.
2. The method according to claim 1, wherein During the training process, the mean squared error loss is used to estimate the prediction quality during training: Where y represents the true label; Overall loss is composed of the weighted sum of and and is expressed as follows: Among them and represent the loss function of the sentiment classification task, the orthogonal constraint loss function, and the loss function of supervised contrastive learning respectively. α, β, and γ are respectively and weights.
3. The method according to claim 1, wherein In step (1), the BERT model is used for preprocessing the text modality.
4. The method according to claim 1, wherein In step (1), the Transformer model is used for preprocessing the audio modality and the video modality.
5. An electronic device, characterized in that, Including a processor and a memory, the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the method according to any one of claims 1-4.
6. A machine-readable storage medium, characterized in that, The machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are called and executed by the processor, the machine-executable instructions cause the processor to implement the method according to any one of claims 1-4.
Citation Information
Patent Citations
Social media sentiment analysis method and system based on tensor fusion network
CN113064968A
Multimodal sentiment classification method and system based on bidirectional double-layer attention LSTM (Long Short Term Memory) network
CN114722202A
Transform-based multi-modal sentiment analysis method
CN114973062A