Multi-modal Sentiment Analysis System and Method Based on Cross-modal Feature State Transfer
Through the cross-modal feature state transfer method, combined with the SIS model of cross-modal multi-head mutual attention and network propagation dynamics, the problem of insufficient modal feature interaction in multi-modal sentiment analysis is solved, and the accuracy of deep-level modal feature interaction and emotion analysis is improved.
Patent Information
- Application Number
- CN202310838017.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2043-07-10
AI Technical Summary
The existing multimodal sentiment analysis model is prone to introduce noise in the cross-modal feature interaction process, and lacks theoretical guidance, making it difficult to achieve deep-level modal feature interaction, affecting the model performance.
The cross-modal feature state transfer method is adopted, and the cross-modal multi-head mutual attention mechanism is used to conduct infection interaction and recovery interaction, and the state transfer process is constructed with the SIS model in network propagation dynamics, and the feature interaction is optimized through the state recording matrix and error constraints, and emotional judgment is made in combination with adaptive weights.
It realizes stable and effective emotional correlation acquisition between modes, enhances the feature expression and accuracy of multimodal emotion analysis, and improves the performance of multimodal emotion analysis.
Smart Images

Figure CN116881843B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology and relates to a multi-modal sentiment analysis system and method based on cross-modal feature state transfer. Background Art
[0002] In multi-modal sentiment analysis, how to effectively learn the sentiment correlation between different modal data, combine the sentiment information of different modal data, and improve the performance of the multi-modal sentiment analysis model has always been a major problem. To solve this problem, many scholars at home and abroad have made many attempts. Some of them use the trained deep learning models to extract text features, visual features, and acoustic features respectively, and then directly fuse them into a multi-layer perceptron for classification to obtain the sentiment tendency. Although this method can perform sentiment analysis by combining the sentiment information of different modalities, the way of directly fusing different modal features is likely to introduce noise and it is difficult to effectively obtain the sentiment correlation between different modal data. More scholars, in order to effectively establish the sentiment correlation between different modal data, design a special interaction module before fusing different modal features to perform cross-modal interaction. Although this method effectively improves the performance of the multi-modal sentiment analysis model, their methods lack theoretical guidance and it is difficult to perform deeper modal feature interaction. In fact, the process of cross-modal interaction can be regarded as the process of information propagation between different modal features. From this point of view, many ideas and models in network propagation dynamics can provide solid theoretical support, and combined with deep learning, it can perform deep modal interaction and improve the model performance.
[0003] However, in addition to the main technologies such as word embedding, Convolutional Neural Networks (CNN), Recurrent Neural Network (RNN), and attention mechanism in deep learning, there are also many derivative models including BERT, Transformer, etc. There are also many models such as SI, SIS, SIR, etc. in network propagation dynamics. The research process and methods of network propagation dynamics are different from those of deep learning. These objective conditions make it very difficult to use the theoretical guidance of network propagation dynamics models for the modal interaction process in deep learning models. Therefore, designing a deep learning model combined with network propagation dynamics theory to more effectively learn the sentiment correlation between different modalities is a very important task. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a multi-modal sentiment analysis system and method based on cross-modal feature state transfer, so as to solve the technical problem of how to effectively learn the sentiment correlation between different modal data, combine the sentiment information of different modal data, and improve the performance of the multi-modal sentiment analysis model.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A multi-modal sentiment analysis method based on cross-modal feature state transfer, the method comprising the following steps:
[0007] S1: Use the pre-trained language representation model BERT to extract text features from text data, use the encoding part of Transformer to extract visual features and acoustic features from the pre-processed visual data and acoustic data, and calculate the mean of the extracted visual features and acoustic features to obtain visual initial state features and acoustic initial state features;
[0008] S2: Use the cross-modal multi-head mutual attention mechanism to perform infection interactions between text features and visual features, and between text features and acoustic features respectively to obtain infection interaction features, and change the input of the cross-modal multi-head mutual attention mechanism to perform recovery interactions to obtain recovery interaction features;
[0009] S3: Use the earth mover's distance and mean absolute error to constrain the infection interaction and recovery interaction processes respectively, calculate the mean of the infection interaction features to obtain the text-visual infection state features or text-acoustic infection state features after infection interaction respectively, and calculate the mean of the recovery interaction features to obtain the text-visual recovery state features or text-acoustic recovery state features after recovery interaction respectively;
[0010] S4: Use the state recording matrix to record the information contained in the initial state features, infection state features and recovery state features in the equilibrium state features, construct a state transfer process similar to the SIS model through the change of the state recording matrix, and use the mean absolute error to constrain the state transfer;
[0011] S5: Use weighted connection to fuse the mean of text features and equilibrium state features and perform sentiment judgment, and perform constraint through mean square error. Add the earth mover's distance, mean absolute error and mean square error constraints of all processes to train the deep learning model of multi-modal sentiment analysis, optimize the trained model to obtain the multi-modal sentiment analysis model, and given multi-modal data including text, video and voice, input the multi-modal sentiment analysis model to obtain the sentiment prediction result of the model.
[0012] Further, the S1 is specifically:
[0013] For the extraction of text feature T, first directly send the given original text into the pre-trained model of BERT to obtain the word vector matrix, and use the fully connected layer to map the text features to the common feature space to extract the text feature T;
[0014] For the extraction of visual feature V and acoustic feature A, first use the MMSA-FET open-source toolkit to preprocess the original visual data and acoustic data, send the preprocessed data into the Encoder part of the Transformer for feature extraction, and use a fully connected layer to map the extracted features to the common feature space to obtain visual feature V and acoustic feature A;
[0015] Calculate the mean of visual feature V and acoustic feature A to obtain the initial visual state feature and the initial acoustic state feature Calculate the mean of text feature T to obtain the initial text state feature
[0016] The initial visual state feature and the initial acoustic state feature are respectively expressed as:
[0017]
[0018]
[0019] where mean(·) dim=0 represents the function of taking the average in the 0th dimension.
[0020] Furthermore, the S2 is specifically:
[0021] Adopt a cross-modal multi-head mutual attention mechanism to achieve infection interaction and recovery interaction. During the infection interaction process, the query input of the cross-modal multi-head mutual attention mechanism is feature T, and the key and value inputs are both feature V and feature A;
[0022] During the recovery interaction process, the query input of the cross-modal multi-head mutual attention mechanism is feature V or A, and the key and value inputs are both feature T; then:
[0023] J T = MultiAtt(T, J, J), J ∈ {V, A} (3)
[0024]
[0025] where MultiAtt(query, key, value) represents the cross-modal multi-head mutual attention mechanism, query, key, and value represent the three inputs of the attention mechanism, and J T represents the infection interaction feature V T or A T , represents the recovery interaction feature or
[0026] Further, the S3 is specifically as follows:
[0027] Utilize the feature V for minimizing the distance of the bulldozer T and T or the feature A T to constrain the infection interaction process by the distance between T, and utilize the feature for minimizing the mean absolute error and V or the feature to constrain the recovery interaction process by the error between A, that is:
[0028]
[0029]
[0030] Wherein, and respectively represent the constraint functions of the infection interaction and the recovery interaction, wass(·) is the bulldozer distance calculation function, and MAE(·) is the mean absolute error calculation function;
[0031] For the infection interaction feature V T or A T calculate the mean value to obtain the text-visual infection state feature or the text-audio infection state feature
[0032] For the recovery interaction feature or calculate the mean value, and optimize the information distribution through the fully connected layer to extract the text-visual recovery state feature or the text-audio recovery state feature That is:
[0033]
[0034]
[0035] Wherein, mean(·) dim=0 represents the function for calculating the average in the 0th dimension, FC(·) is the fully connected layer, is the training parameter of the fully connected layer.
[0036] Further, the S4 is specifically as follows:
[0037] Determine the initial state feature through the contact rate or the vulnerable part in, introduce the infection rate and the random infection matrix to determine the initial state feature in the equilibrium state feature or information and the infection state feature or information, and record it through the state recording matrix, that is:
[0038]
[0039]
[0040] where θ is the infection rate, M rand is the random infection matrix, χ is the contact rate, O is the all-ones matrix, represents the characteristic of the infected state or the state record matrix of information, represents the characteristic of the initial state or the state record matrix of information, abs(·) is used to take the modulus of each element of the tensor, is the initial state characteristic of the text;
[0041] During the recovery process, the recovery rate and the recovery random matrix are introduced to determine the recovered state characteristic in the equilibrium state characteristic or the information part, which is recorded by the state record matrix, and the state record matrix of the information of the infected state characteristic is updated is i.e.:
[0042]
[0043] where ρ is the recovery rate, is the random recovery matrix, represents the characteristic of the recovered state or the state record matrix of information;
[0044] The equilibrium state characteristic V state or A state is extracted through the state information recorded in the state record matrix, and and are subjected to the Hadamard product and added to the corresponding state characteristics, i.e.:
[0045]
[0046] where ⊙ represents the Hadamard product, J state represents the equilibrium state characteristic V state or A state ;
[0047] The characteristics V state and A state are mapped to the emotion judgment domain, and are constrained and updated by the mean absolute error with the emotion label, and then the entire state transition process is constrained, then:
[0048]
[0049] L state = MAE(Y state , Y) (14)
[0050] Among them, and are the training parameters of the fully connected layer, Y state represents the tensor obtained by mapping V state and A state to the sentiment judgment domain, Y represents the sentiment label tensor, L state is the loss function used for constraint, and MAE() represents calculating the mean absolute error.
[0051] Furthermore, the S5 is specifically:
[0052] Adopt adaptive weights that can be updated through training to fully fuse the features V state , A state and , then:
[0053]
[0054] Among them, F represents the fused features, and are the adaptive weights, and concat(·) is the concatenation function;
[0055] Use a multi-layer perceptron to map the feature F to the sentiment judgment domain to achieve sentiment prediction, and optimize the loss function of the sentiment prediction of the model by minimizing the error between the prediction result of the model and the true sentiment label Y, then:
[0056]
[0057] Among them, L c is the loss function of the sentiment prediction of the optimized model, and MSE(·) represents the mean square error function;
[0058] Use the backpropagation algorithm to train the model, and the loss function L of the overall training of the model can be expressed as:
[0059]
[0060] Optimize and train the model through the AdamW optimizer to obtain a multi-modal sentiment analysis model, save it locally, and given multi-modal data including text, video, and speech, after inputting it into the multi-modal sentiment analysis model, the sentiment prediction result of the model can be obtained.
[0061] Multimodal sentiment analysis system based on cross-modal feature state transfer, which is successively provided with a feature extraction module, a cross-modal interaction module, a cross-modal feature state transfer module, and a sentiment prediction module.
[0062] The beneficial effects of the present invention are as follows:
[0063] First, the present invention uses a cross-modal multi-head mutual attention mechanism to achieve cross-modal interaction and initially obtain the correlation information between modalities.
[0064] Second, the present invention refers to the SIS model theory in network propagation dynamics, introduces a state recording matrix to construct the state transfer process between different modal features, thereby realizing the deep interaction between different modalities, stably and effectively obtaining the sentiment correlation of multimodal data, enhancing the feature expression, and enabling the present invention to more effectively realize multimodal sentiment analysis.
[0065] Third, the present invention can jointly analyze the sentiment of multiple modal information and classify the sentiment jointly expressed by multimodal representations, realizing the accurate extraction of multimodal features and the enhancement of multimodal information interaction, and having strong multimodal sentiment analysis capabilities.
[0066] Fourth, the present invention provides ideas and references for the research on the interpretability of deep learning.
[0067] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0069] Figure 1 is a process diagram of multimodal sentiment analysis based on cross-modal feature state transfer;
[0070] Figure 2 is a model diagram of multimodal sentiment analysis based on cross-modal feature state transfer. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0071] The following specific examples are used to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the drawings provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0072] Among them, the drawings are only for illustrative purposes, showing only schematic diagrams rather than actual diagrams, and should not be construed as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the drawings will be omitted, enlarged or reduced, which does not represent the size of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the drawings may be omitted.
[0073] In the drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0074] Please refer to Figures 1 to 2 , a multi-modal sentiment analysis system and method based on cross-modal feature state transfer.
[0075] The implementation scenario is to use the mentioned method to train a deep learning model for multi-modal sentiment analysis on a data set containing text, video, audio and corresponding emotions, so that the model can perform sentiment analysis by combining multiple modal data. The specific implementation steps are as follows:
[0076] Step 1: Extract text features, visual features and acoustic features, and calculate the mean of the visual features and acoustic features to obtain the initial state features.
[0077] First, download the BERT pre-trained model, directly send the original text data S into the BERT pre-trained model to obtain the word vector matrix. For convenient subsequent interaction processing, use a fully connected layer to map the text features to a common feature space to extract the text feature T. For the visual feature V and the acoustic feature A, first use the MMSA-FET open-source toolkit to preprocess the original visual data and acoustic data, then send them into the Encoder part of the Transformer, and then send them into a fully connected layer to map to the common feature space to obtain. In addition, obtain the initial visual state feature and the initial acoustic state feature Take the mean of the text feature T to obtain the feature
[0078] Step 2: Use the cross-modal multi-head mutual attention mechanism to perform infection interactions between the text feature and the visual feature, and between the text feature and the acoustic feature respectively, and change the input of the cross-modal multi-head mutual attention mechanism to perform recovery interactions.
[0079] Specifically, use a cross-modal multi-head mutual attention mechanism with the input query as the feature T, and the input key and value both as the feature V or A to achieve infection interaction to obtain the infection interaction feature V T or A T ; At the same time, use a cross-modal multi-head mutual attention mechanism with the input query as the feature V or A, and the input key and value both as the feature T to achieve recovery interaction to obtain the recovery interaction feature or Then:[[]]
[0080] J T = MultiAtt(T, J, J), J ∈ {V, A} (1)
[0081]
[0082] Where MultiAtt(query, key, value) represents the cross-modal multi-head mutual attention mechanism, query, key and value represent the three inputs of the attention mechanism, and J T represents the infection interaction feature V T or A T , represents the recovery interaction feature or
[0083] Step 3: Use the earth mover's distance and the mean absolute error to constrain the infection interaction and recovery interaction processes respectively, and calculate the mean to obtain the text-visual infection state feature and text-acoustic infection state feature after infection interaction, and the text-visual recovery state feature and text-acoustic recovery state feature after recovery interaction.
[0084] For feature V T or A T Calculate the mean value to obtain the text-vision infection state feature or the text-audio infection state feature And minimize the distance between feature V T and T or feature A T The distance between and T to constrain the infection interaction process. At the same time, for the feature or Calculate the mean value, and optimize the information distribution through the fully connected layer to extract the text-vision recovery state feature or the text-audio recovery state feature This recovery interaction process is constrained by minimizing the mean absolute error between the feature and V or the feature and A. The constraint function of the cross-modal feature interaction process is as follows:
[0085]
[0086]
[0087] Among them, and respectively represent the constraint functions of the infection interaction and the recovery interaction. wass(·) is the earth mover's distance calculation function, and MAE(·) is the mean absolute error calculation function.
[0088] Step 4: Use the state recording matrix to record the information contained in the initial state feature, the infection state feature, and the recovery state feature in the equilibrium state feature. Construct a state transition process similar to the SIS model through the change of the state recording matrix, and use the mean absolute error to constrain this process.
[0089] This step mainly draws on the SIS model in network propagation dynamics to construct the cross-modal feature state transition process. Different from that, in the infection process, first determine the part of the initial state feature or that is easily infected, and then introduce the infection rate and the random infection matrix to determine the initial state feature or information and the infection state feature or information in the equilibrium state feature, and record it through the state recording matrix, that is:
[0090]
[0091]
[0092] where θ is the infection rate, M randis a random infection matrix, χ is the contact rate, and O is a matrix of all 1s. represents the infected state characteristics or the state recording matrix of information represents the initial state characteristics or the state recording matrix of information, and abs(·) is used to take the modulus of each element of the tensor.
[0093] During the recovery process, the recovery rate and the recovery random matrix are introduced to determine the recovered state characteristics in the equilibrium state characteristics or the information part, and the state recording matrix is also used to record and update the state recording matrix of the infected state characteristic information is That is:
[0094]
[0095] where ρ is the recovery rate is the random recovery matrix represents the recovered state characteristics or the state recording matrix of information. In this way, a state transition process similar to the SIS is established through the state recording matrix for in-depth cross-modal interaction.
[0096] Since and are both 0-1 matrices, to extract the equilibrium state characteristics V state or A state , it is necessary to and perform the Hadamard product and addition with the corresponding state characteristics, that is:
[0097]
[0098] where ⊙ represents the Hadamard product, and J state represents the equilibrium state characteristics V state or A state . It is worth mentioning that the characteristics V state or A state are not the characteristics obtained in the equilibrium state defined in the network propagation dynamics, but are the characteristics composed of the state characteristic information saved through the state recording matrix after each state transition process in order to simplify the overall optimization process of the model. Therefore, in order to make the information of the equilibrium state characteristics V state or A state closer to the information distribution in the equilibrium state of the SIS model, it is necessary to state and A stateMap to the emotional judgment domain, and perform constraint update through the mean absolute error with emotional labels, thereby constraining the entire state transition process. Then:
[0099]
[0100] L state = MAE(Y state , Y) (10)
[0101] Among them, and are the training parameters of the fully connected layer. Y state represents the tensor obtained by mapping V state and A state to the emotional judgment domain, Y represents the emotional label tensor, and L state is the loss function used for constraint.
[0102] Step 5: Use weighted connection to fuse the mean of text features and equilibrium state features and perform emotional judgment. Constrain through the mean square error, and add the earth mover's distance, mean absolute error, and mean square error constraints of all processes to train the entire model.
[0103] Adopt adaptive weights that can be updated through training to fully fuse the features V state , A state and as follows:
[0104]
[0105] Among them, F represents the fused feature, and are the adaptive weights, and concat(·) is the concatenation function. Then use a multi-layer perceptron to map the feature F to the emotional judgment domain to achieve emotional prediction. To optimize the emotional prediction ability of the model, it is necessary to minimize the error between the predicted result of the model and the true emotional label Y, then:
[0106]
[0107] Among them, L c is the loss function for optimizing the emotional prediction of the model, and MSE(·) represents the mean square error function. Finally, use the backpropagation algorithm to train the model. Since there are loss functions for constraint in each part of the model, the loss function L for the overall training of the model can be expressed as:
[0108]
[0109] And optimize and train the model through the AdamW optimizer.
[0110] After the above steps, a multi-modal sentiment analysis model will be obtained and saved locally. Given multi-modal data containing text, video, and speech, after inputting it into the model, the sentiment prediction result of the model can be obtained.
[0111] Figure 2 This is the system model diagram of the present invention, which will be described below in conjunction with the accompanying drawings and includes the following several modules:
[0112] Module 1: For text data, for the visual data and acoustic data preprocessed by the MMSA-FET open-source toolkit, the BERT pre-trained model and Transformer are respectively used to extract text features, visual features, and acoustic features, and the mean values of the visual features and acoustic features are calculated to obtain the initial state features. Calculating the mean value of the text features is convenient for subsequent processing;
[0113] Module 2: The cross-modal multi-head mutual attention mechanism is used to respectively perform infection interactions between text features and visual features, and between text features and acoustic features, and the input of the cross-modal multi-head mutual attention mechanism is changed to perform recovery interactions. The earth mover's distance and mean absolute error are respectively used to constrain the infection interaction and recovery interaction processes, and the infection state features and recovery state features are respectively obtained through mean value calculation;
[0114] Module 3: The state recording matrix is used to record the information contained in the initial state features, infection state features, and recovery state features in the equilibrium state features. The state transition process similar to the SIS model is constructed through the change of the state recording matrix, and this process is constrained by the mean absolute error;
[0115] Module 4: The adaptive weight is used to fully fuse the text features and equilibrium state features, and the fused features are mapped to the sentiment judgment space through a multi-layer perceptron to further predict the sentiment of the multi-modal data.
[0116] Optionally, Module 1 specifically includes:
[0117] Feature extraction module. For the extraction of text feature T, first directly send the given original text into the pre-trained model of BERT to obtain the word vector matrix. For the convenience of subsequent cross-modal interaction and state transition processing, then send the word vector matrix into a fully connected layer to map it to the common feature space. For visual feature V and acoustic feature A, to avoid loss of important information and insufficient subsequent feature extraction, first use the MMSA-FET open-source toolkit to preprocess the original video data and audio data, then use an encoder containing two layers of Transformer Encoder Layer to further extract features, and then use a fully connected layer to map the features to the common feature space. In this way, text feature T, visual feature V, and acoustic feature A are all in the same feature space, which can avoid introducing noise or redundancy due to the heterogeneity of different modal features in subsequent processing. In addition, for the convenience of the implementation of the subsequent state transition process, take the mean of visual feature V and acoustic feature A to obtain the visual initial state feature and acoustic initial state feature Take the mean of text feature T to obtain the feature
[0118] Optionally, Module 2 specifically includes:
[0119] Cross-modal interaction module. Since the main structures of BERT and Transformer are both based on the attention mechanism, to avoid insufficient interaction between different modalities due to the differences in hidden layer features extracted by different deep learning techniques based on different structures, a cross-modal multi-head mutual attention mechanism is used to achieve infection interaction and recovery interaction. In the infection interaction process, the query input of the cross-modal multi-head mutual attention mechanism is feature T, and the key and value inputs are both feature V or A. In the recovery interaction process, the query input of the cross-modal multi-head mutual attention mechanism is feature V or A, and the key and value inputs are both feature T. Then:
[0120] J T = MultiAtt(T, J, J), J ∈ {V, A} (14)
[0121]
[0122] where MultiAtt(query, key, value) represents the cross-modal multi-head mutual attention mechanism, query, key, and value represent the three inputs of the attention mechanism, and J T represents the infection interaction feature V T or A T , represents the recovery interaction feature or
[0123] The feature V extracted through the above two types of interactions T , A T , and Actually, an emotional association between different modalities is initially established. However, due to the lack of constraints in this interaction process, this emotional association information is prone to instability, which is not conducive to establishing a deep emotional association during the subsequent state transition process.
[0124] To obtain stable emotional association information, the Earth Mover's Distance is used to minimize the distance between feature V T and T or feature A T and T to constrain the infection interaction process, and the Mean Absolute Error is used to minimize the error between feature and V or feature and A to constrain the recovery interaction process, that is:
[0125]
[0126]
[0127] where and represent the constraint functions for the infection interaction and the recovery interaction respectively. wass(·) is the Earth Mover's Distance calculation function, and MAE(·) is the Mean Absolute Error calculation function. In addition, for the smooth implementation of the subsequent state transition process, it is necessary to calculate the mean of feature V T or A T to obtain the text-vision infection state feature or the text-audio infection state feature To obtain the text-vision recovery state feature or the text-audio recovery state feature it is necessary to first calculate the mean of feature or To make the recovery state feature or closer to the initial state feature or it is also necessary to perform further processing using a fully connected layer. Then:
[0128]
[0129]
[0130] where mean(·) dim=0 represents the function of calculating the mean in the 0th dimension, FC(·) is the fully connected layer, are the training parameters of the fully connected layer.
[0131] Optionally, Module 3 specifically includes:
[0132] Cross-modal feature state transition module. As Figure 2 shown, this module mainly constructs the cross-modal feature state transition process by referring to the SIS model in network propagation dynamics. Different from it, during the infection process, the susceptible part of the initial state features or in the or information and the infected state features or information are first determined by the contact rate, and then the infection rate and the random infection matrix are introduced to determine the initial state features in the equilibrium state features
[0133]
[0134]
[0135] where θ is the infection rate, M rand is the random infection matrix, χ is the contact rate, O is the all-1 matrix, represents the state recording matrix of the infected state features or information, represents the state recording matrix of the initial state features or information, and abs(·) is used to take the modulus of each element of the tensor.
[0136] During the recovery process, the recovery rate and the recovery random matrix are introduced to determine the recovered state features or information part in the equilibrium state features, and the state recording matrix is also used to record and update the state recording matrix of the infected state feature information as i.e.:
[0137]
[0138] where ρ is the recovery rate, is the random recovery matrix, represents the state recording matrix of the recovered state features or information. In this way, a state transition process similar to SIS is established through the state recording matrix for in-depth cross-modal interaction.
[0139] Since and are both 0-1 matrices, to extract the equilibrium state features V state or A state from the state information recorded in these state recording matrices, it is necessary to combine and Perform a Hadamard product and addition with the corresponding state features, i.e.:
[0140]
[0141] where ⊙ represents the Hadamard product, and J state represents the equilibrium state feature V state or A state . It is worth noting that the feature V state or A state is not the feature obtained in the equilibrium state defined in the network propagation dynamics. Instead, to simplify the overall optimization process of the model, it is a feature composed of the state feature information saved through the state recording matrix after each state transition. Therefore, in order to make the information of the equilibrium state feature V state or A state closer to the information distribution in the equilibrium state of the SIS model, it is necessary to map the features V state and A state to the sentiment judgment domain, and perform constraint update through the mean absolute error with the sentiment labels, thereby constraining the entire state transition process. Then:
[0142]
[0143] L state = MAE(Y state , Y) (25)
[0144] where and are the training parameters of the fully connected layer, Y state represents the tensor obtained by mapping V state and A state to the sentiment judgment domain, Y represents the sentiment label tensor, and L state is the loss function used for constraint.
[0145] Optionally, Module Five specifically includes:
[0146] Sentiment prediction module. To avoid loss of important information and enhance the expression of effective sentiment information, it is necessary to use adaptive weights that can be updated through training to fully fuse the features V state , A state , and as follows:
[0147]
[0148] where F represents the fused feature, and are the adaptive weights, and concat(·) is the concatenation function.
[0149] Afterwards, a multi-layer perceptron is used to map the feature F to the sentiment judgment domain to achieve sentiment prediction. To optimize the sentiment prediction ability of the model, it is necessary to minimize the prediction result of the model through the mean square error The error between the true sentiment label Y, then:
[0150]
[0151] where L c is the loss function for optimizing the sentiment prediction of the model, and MSE(·) represents the mean square error function.
[0152] The model is trained using the backpropagation algorithm, and the model is optimized by minimizing the loss function. However, since there are loss functions for constraints in each part of the model, the loss function L for the overall training of the model can be expressed as:
[0153]
[0154] And the model is optimized and trained using the AdamW optimizer.
[0155] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A multi-modal sentiment analysis method based on cross-modal feature state transition, characterized in that: The method includes the following steps: S1: Use the pre-trained language representation model BERT to extract text features from text data, use the encoding part of the Transformer model to extract visual features and acoustic features from the pre-processed visual data and acoustic data, and calculate the mean of the extracted visual features and acoustic features to obtain the initial visual state features and the initial acoustic state features; S2: Use the cross-modal multi-head mutual attention mechanism to perform infection interaction between text features and visual features, and between text features and acoustic features to obtain infection interaction features, and change the input of the cross-modal multi-head mutual attention mechanism to perform recovery interaction to obtain recovery interaction features; S3: Use the earth mover's distance and the mean absolute error to respectively constrain the infection interaction and the recovery interaction processes, calculate the mean of the infection interaction features to obtain the text-visual infection state features or the text-acoustic infection state features after infection interaction respectively, and calculate the mean of the recovery interaction features to obtain the text-visual recovery state features or the text-acoustic recovery state features after recovery interaction respectively; S4: Use the state recording matrix to record the information contained in the initial state features, infection state features, and recovery state features in the equilibrium state features, construct a state transition process similar to the SIS model through the change of the state recording matrix, and use the mean absolute error to constrain the state transition; Specifically, the initial state characteristics are determined by the contact rate or the susceptible part in [[]], and the initial state characteristics in the equilibrium state characteristics are determined by introducing the infection rate and the random infection matrix or information and the infected state characteristics or information, and recorded through the state record matrix, that is: where θ is the infection rate, M rand is the random infection matrix, χ is the contact rate, O is the all-ones matrix, represents the characteristic of the infected state or the state record matrix of information, represents the characteristic of the initial state or the state record matrix of information, abs(·) is used to take the modulus of each element of the tensor, is the initial state characteristic of the text; During the recovery process, the recovery rate and the recovery random matrix are introduced to determine the characteristics of the recovery state among the equilibrium state characteristics or For the information part, it is recorded using the state record matrix, and the state record matrix for updating the information of the infected state characteristics is That is: where ρ is the recovery rate, is a random recovery matrix, represents the feature of the recovered state or the state record matrix of information; Extract the equilibrium state feature V from the state information recorded in the state record matrix state or A state , and and perform Hadamard product and addition with the corresponding state features, i.e.: where, ⊙ represents the Hadamard product, and J state represents the equilibrium state feature V state or A state ; Map the feature V state and A state to the emotion judgment domain, and perform constraint update through the mean absolute error with the emotion label, and then constrain the entire state transition process, then: L state = MAE(Y state , Y) Among them, and are the training parameters of the fully connected layer, Y state represents V state and A state the tensor mapped to the sentiment judgment domain, Y represents the sentiment label tensor, L state is the loss function for constraint, and MAE() represents the calculation of the mean absolute error; S5: Use weighted connection to fuse the mean of text features and the equilibrium state features and perform sentiment judgment, and perform constraint through the mean square error. Add up the earth mover's distance, mean absolute error, and mean square error constraints of all processes to train the deep learning model for multi-modal sentiment analysis, optimize the trained model to obtain the multi-modal sentiment analysis model, and given multi-modal data including text, video, and speech, input the multi-modal sentiment analysis model to obtain the sentiment prediction result of the model.
2. The multimodal sentiment analysis method based on cross-modal feature state transfer according to claim 1, wherein: The specific content of S1 is as follows: For the extraction of text feature T, first directly send the given original text into the pre-trained model of BERT to obtain the word vector matrix, and use the fully connected layer to map the text features to the common feature space to extract the text feature T; For the extraction of visual feature V and acoustic feature A, first use the MMSA-FET open source toolkit to pre-process the original visual data and acoustic data, send the pre-processed data into the encoding part of the Transformer for feature extraction, and use the fully connected layer to map the extracted features to the common feature space to obtain the visual feature V and the acoustic feature A; Calculate the mean of the visual feature V and the acoustic feature A to obtain the initial visual state feature and the initial acoustic state feature Calculate the mean of the text feature T to obtain the initial text state feature Visual initial state features Harmonic initial state features Are respectively expressed as: Among them, mean(·) dim=0 represents the function for calculating the average in the 0th dimension.
3. The multimodal sentiment analysis method based on cross-modal feature state transfer according to claim 2, characterized in that: The specific content of S2 is as follows: Use the cross-modal multi-head mutual attention mechanism to achieve infection interaction and recovery interaction. In the infection interaction process, the query input of the cross-modal multi-head mutual attention mechanism is feature T, and the key and value inputs are both feature V and feature A; In the recovery interaction process, the query input of the cross-modal multi-head mutual attention mechanism is feature V or A, and the key and value inputs are both feature T; then: J T = MultiAtt(T, J, J), J ∈ {V, A} Among them, MultiAtt(query, key, value) represents the cross-modal multi-head mutual attention mechanism, where query, key, and value represent the three inputs of the attention mechanism, and J T represents the infection interaction feature V T or A T , represents the recovery interaction feature or 4. The multimodal sentiment analysis method based on cross-modal feature state transfer according to claim 3, wherein: The specific content of S3 is as follows: Using the minimum distance of the bulldozer for feature V T With T or feature A T The distance between it and T is used to constrain the infection interaction process, and the feature for minimizing the mean absolute error With V or feature The error between it and A is used to constrain the recovery interaction process, that is: Among them, and respectively represent the constraint functions for infection interaction and recovery interaction, wass(·) is the earth mover's distance calculation function, and MAE(·) is the mean absolute error calculation function; For the infection interaction feature V T or A T Calculate the mean value to obtain the text-visual infection state feature or the text-audio infection state feature For restoring interaction features or calculate the mean value, and extract text-visual restoration state features by optimizing the information distribution through a fully connected layer or text-audio restoration state features That is: Among them, mean(·) dim=0 represents the function of taking the average in the 0th dimension, and FC(·) is the fully connected layer, which are the training parameters of the fully connected layer.
5. The multimodal sentiment analysis method based on cross-modal feature state transfer according to claim 4, wherein: The specific content of S5 is as follows: Adopt an adaptive weight that can be updated through training for feature V state , A state and to fully fuse them, then: Among them, F represents the fused feature, and ω T are adaptive weights, and concat(·) is a concatenation function; The feature F is mapped to the sentiment judgment domain by a multi-layer perceptron to achieve sentiment prediction, and the error between the prediction result of the model and the true sentiment label Y is minimized by the mean square error to obtain the loss function for sentiment prediction of the optimized model, then: The error between the prediction result of the model and the true sentiment label Y is minimized by the mean square error to obtain the loss function for sentiment prediction of the optimized model, then: Among them, L c is the loss function for optimizing the sentiment prediction of the model, and MSE(·) represents the mean squared error function; Use the backpropagation algorithm to train the model. The loss function L for the overall training of the model can be expressed as: The model is optimized and trained by the AdamW optimizer to obtain a multi-modal sentiment analysis model, which is saved locally. Given multi-modal data containing text, video, and speech, the sentiment prediction result of the model can be obtained after inputting it into the multi-modal sentiment analysis model.
6. A multimodal sentiment analysis system based on cross-modal feature state transfer, characterized in that: This system is used to execute the method described in any one of claims 1 to 5.
Citation Information
Patent Citations
Deep microblog sentiment analysis method using social context features
CN110188200A
Cross-modal BERT sentiment analysis method based on visual, audio and text fusion
CN115510224A