A multi-modal sentiment classification model training method and a multi-modal sentiment classification method
By using the method of learning masks and causal effect calculation in multimodal sentiment analysis, the problem of false association inside and outside the modal is solved, and the accuracy and robustness of multimodal sentiment classification are improved.
Patent Information
- Application Number
- CN202411601665.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2044-11-11
AI Technical Summary
There are false association problems in the existing multimodal sentiment analysis methods, which lead to degradation of model performance and insufficient generalization ability and robustness.
The multimodal features are filtered using a learning mask, extract multimodal causal features, and scheduling weights are calculated and jointly optimized through causal effects to build a multimodal emotion classification model to alleviate the problem of false association.
The performance of multimodal emotion classification is improved, the causal learning ability of the model is enhanced, and the accuracy and robustness of emotion classification is improved.
Smart Images

Figure CN119557785B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method for training a multi-modal sentiment classification model and a multi-modal sentiment classification method. Background Art
[0002] Goal-oriented multi-modal sentiment classification is a hot research direction in the fields of natural language processing and computer vision in recent years. It aims to automatically identify the emotional state of a target from multi-modal data (such as text, images, audio, etc.), and is widely applied to scenarios such as emotion computing, social media analysis, virtual assistants, etc. Different from traditional single-modal sentiment analysis, goal-oriented multi-modal sentiment classification utilizes data from different modalities, fuses multi-source information, and can more comprehensively understand and infer the emotional state of the target, improving the accuracy and robustness of the analysis.
[0003] In recent years, with the development of technologies such as deep learning, multi-modal sentiment analysis technology has made remarkable progress. By using deep learning models such as convolutional neural networks and long short-term memory networks, researchers can automatically learn multi-modal features from a large amount of data, significantly improving the robustness and accuracy of sentiment analysis.
[0004] However, existing multi-modal sentiment analysis methods still face a key challenge, namely the problem of spurious correlations within and between different modalities. In sentiment analysis, the problem of spurious correlations refers to the situation where the model wrongly establishes associations between irrelevant or noisy features and the target sentiment labels during the learning process, resulting in a decline in the performance of the model. Especially in multi-modal sentiment analysis, this problem is more prominent. The model will be affected by spurious correlations both between and within modalities, leading to a decline in performance. Among them, the problem of spurious correlations between modalities includes the problem of non-correspondence between different modalities, and the problem of spurious correlations within modalities includes the problem that some content contained in the modality does not match the specific expressed emotion. These problems of spurious correlations will reduce the generalization ability and robustness of the model. Summary of the Invention
[0005] In view of this, the present invention provides a method for training a multi-modal sentiment classification model and a multi-modal sentiment classification method to solve the problem of spurious correlations existing in existing multi-modal sentiment analysis.
[0006] In a first aspect, the present invention provides a method for training a multi-modal sentiment classification model, the method comprising: obtaining training samples containing multi-modal samples; extracting multi-modal features of the multi-modal samples using a preset neural network; filtering the multi-modal features using a learnable mask to obtain multi-modal causal features; calculating a causal effect using the sentiment classification loss of the multi-modal features and the sentiment classification loss of the multi-modal causal features, the sentiment classification loss being determined based on the sentiment classification results of the preset classifier for the multi-modal features and the multi-modal causal features; determining a joint loss using a scheduling weight determined by the causal effect and learnable parameters and the sentiment classification loss; performing parameter optimization on the parameters in the preset neural network, the preset classifier, and the joint loss using joint optimization, and constructing a multi-modal sentiment classification model based on the preset neural network, the learnable mask, and the preset classifier after parameter optimization.
[0007] In the present invention, during the model training process, a learnable mask is used to filter the multi-modal features in the samples to obtain multi-modal causal features, that is, multi-modal causal features are selected through simple and effective causal interventions, alleviating the problem of spurious associations. At the same time, the multi-modal learning process is scheduled by evaluating the causal effect of the features, and the optimal scheduling parameters are adaptively determined in combination with joint optimization. Thus, the trained multi-modal sentiment classification model realizes multi-modal feature learning based on causal relationships and improves the performance of multi-modal sentiment classification.
[0008] In an alternative embodiment, the multi-modal samples include text samples and image samples, and extracting multi-modal features of the multi-modal samples using a preset neural network includes: extracting text features of the text samples using a BERT model; extracting image features of the image samples using a CLIP model.
[0009] In the present invention, text features of text samples are extracted using a BERT model, and image features of image samples are extracted using a CLIP model, realizing accurate extraction of multi-modal features.
[0010] In an alternative embodiment, calculating a causal effect using the sentiment classification loss of the multi-modal features and the sentiment classification loss of the multi-modal causal features includes: determining the sentiment classification results of each modal feature and each modal causal feature using a preset classifier; calculating the sentiment classification loss corresponding to each modal feature and each modal causal feature based on the loss function corresponding to each modality, the sentiment classification results, and the corresponding sentiment labels; calculating a causal effect using the sentiment classification loss corresponding to each modal feature and each modal causal feature.
[0011] In the present invention, the causal effect is calculated through the above embodiment, and thus the influence of the corresponding causal features on the sentiment classification results can be quantified by this causal effect.
[0012] In an alternative embodiment, the causal effect is represented by the following formula:
[0013]
[0014]
[0015] Wherein, Δ∈ represents the causal effect, and respectively represent the counterfactual text causal feature and the counterfactual image causal feature in the multi-modal causal feature, y t represents the sentiment label of the text; y v represents the sentiment label of the image, respectively represent the sentiment classification results of the text feature and the image feature in the multi-modal feature determined by the preset classifier, respectively represent the sentiment classification results of the counterfactual text causal feature and the counterfactual image causal feature determined by the preset classifier, L ce represents the loss function corresponding to the text feature, L bce represents the loss function corresponding to the image feature.
[0016] In the present invention, by calculating the causal effect using the above formula, the calculated causal effect comprehensively takes into account the influences of the multi-modal feature and the multi-modal causal feature.
[0017] In an alternative embodiment, joint optimization is used to optimize the parameters in the preset neural network, the preset classifier, and the joint loss, including: using the preset learnable parameters to optimize the first parameters, where the first parameters include the parameters of the preset neural network, the learnable mask, and the parameters in the preset classifier; optimizing the learnable parameters based on the optimized first parameters; repeating the above steps until the obtained joint loss meets the preset requirements.
[0018] In the present invention, when optimizing the parameters, a two-level optimization method is adopted, that is, the parameters are divided into two levels. First, the first parameters are optimized by fixing the learnable parameters, and then the learnable parameters are optimized by the first parameters. Repeating this process enables the final joint loss to meet the preset requirements, thereby obtaining the optimal parameters.
[0019] Second aspect, the present invention provides a multi-modal sentiment classification method, which is applied to the multi-modal sentiment classification model trained by the multi-modal sentiment classification model training method described in the first aspect and any item of the first aspect of the present invention. The method includes: obtaining data to be classified, where the data to be classified is multi-modal data; extracting features of the data to be classified using a preset neural network with optimized parameters; filtering the features using a learnable mask with optimized parameters to obtain causal features; and predicting the causal features using a preset classifier with optimized parameters to determine the sentiment classification result of the data to be classified.
[0020] In the present invention, since the learnable mask is optimized in terms of parameters during the above model training process, thereby, the causal features obtained by filtering alleviate the problem of spurious associations in the features. Then, the classifier with optimized parameters is used to predict and classify the causal features, so as to obtain a more accurate sentiment classification effect.
[0021] Third aspect, the present invention provides a multi-modal sentiment classification model training device, which includes: a sample acquisition module for obtaining training samples including multi-modal samples; a feature extraction module for extracting multi-modal features of the multi-modal samples using a preset neural network; a filtering module for filtering the multi-modal features using a learnable mask to obtain multi-modal causal features; a causal effect calculation module for calculating the causal effect using the sentiment classification loss of the multi-modal features and the sentiment classification loss of the multi-modal causal features, where the sentiment classification loss is determined based on the sentiment classification results of the multi-modal features and the multi-modal causal features by a preset classifier; a joint loss determination module for determining the joint loss using the scheduling weight determined by the causal effect and learnable parameters and the sentiment classification loss; and a parameter optimization and model construction module for optimizing the parameters of the preset neural network, the preset classifier, and the parameters in the joint loss using joint optimization, and constructing a multi-modal sentiment classification model based on the preset neural network, the learnable mask, and the preset classifier with optimized parameters.
[0022] Fourth aspect, the present invention provides a multi-modal sentiment classification device, which is applied to the multi-modal sentiment classification model trained by the multi-modal sentiment classification model training method described in the first aspect and any item of the first aspect of the present invention. The device includes: a data acquisition module for obtaining data to be classified, where the data to be classified is multi-modal data; a data feature extraction module for extracting features of the data to be classified using a preset neural network with optimized parameters; a causal feature determination module for filtering the features using a learnable mask with optimized parameters to obtain causal features; and a prediction module for predicting the causal features using a preset classifier with optimized parameters to determine the sentiment classification result of the data to be classified.
[0023] Fifth aspect, the present invention provides a computer device, including: a memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the multi-modal sentiment classification model training method according to the first aspect or any corresponding embodiment thereof.
[0024] Sixth aspect, the present invention provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the multi-modal sentiment classification model training method according to the first aspect or any corresponding embodiment thereof.
[0025] Seventh aspect, the present invention provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the multi-modal sentiment classification model training method according to the first aspect or any corresponding embodiment thereof. Description of the Drawings
[0026] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0027] Figure 1 is a schematic flowchart of a multi-modal sentiment classification model training method according to an embodiment of the present invention;
[0028] Figure 2 is a schematic flowchart of a multi-modal sentiment classification method according to an embodiment of the present invention;
[0029] Figure 3 is a structural block diagram of a multi-modal sentiment classification model training device according to an embodiment of the present invention;
[0030] Figure 4 is a structural block diagram of another multi-modal sentiment classification model training device according to an embodiment of the present invention;
[0031] Figure 5 is a structural block diagram of a multi-modal sentiment classification device according to an embodiment of the present invention;
[0032] Figure 6 is a schematic hardware structure diagram of a computer device according to an embodiment of the present invention. Detailed Embodiments
[0033] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0034] According to an embodiment of the present invention, an embodiment of a multi-modal sentiment classification model training method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0035] In this embodiment, a multi-modal sentiment classification model training method is provided, which can be used in electronic devices such as computers, mobile phones, and tablet computers. Figure 1 It is a flowchart of the multi-modal sentiment classification model training method according to an embodiment of the present invention, as Figure 1 shown, and the process includes the following steps:
[0036] Step S101, obtain training samples including multi-modal samples. Specifically, the training samples may include multiple groups of multi-modal samples. Here, multi-modal refers to data from multiple different sources or types, such as text data, image data, audio data, or video data, etc. Among them, each group of multi-modal samples includes corresponding multi-modal data, and each group of multi-modal samples includes at least two types of samples. Here, the correspondence means the correspondence between each modality in each group of multi-modal samples. For example, each group of multi-modal samples may include a text sample and a corresponding image sample.
[0037] In practical applications, a group of multi-modal samples may be the text information and image information included in the dynamic posted by a certain user on a certain social platform. The text information and image information are the text sample and image sample in the multi-modal samples. In addition, for the text sample, it may include multiple text targets For example, if the following text information is included in the dynamic posted by a certain user on a certain social platform: "Pat celebrating her 90th birthday with Emily Roux in Chez Roux at the Newmarket Guineas Festival", then it can be determined that there are three targets Pat, Emily Roux, and Chez Roux in this text information. The multi-modal sentiment classification model is based on the text-image pair X = (X t, X v ) to perform sentiment classification on it.
[0038] Step S102: Extract multi-modal features of the multi-modal samples using a preset neural network. Specifically, when performing feature extraction, a neural network in related technologies can be used. At the same time, since the multi-modal samples include samples of different modalities, for samples of different modalities, a suitable neural network can be selected according to the actual situation to implement feature extraction for the corresponding modality, so that the extracted features are more accurate.
[0039] Step S103: Filter the multi-modal features using a learnable mask to obtain multi-modal causal features. Specifically, the element values of the learnable mask belong to (0, 1), that is, the learnable mask can give each multi-modal feature a weight from 0 to 1, realizing direct causal intervention on the features, or it can also be understood as screening or filtering of the multi-modal features, so as to obtain multi-modal causal features. That is, the learnable mask is used to filter out the noise features or background features that may mislead the sentiment classification in the multi-modal features. Therefore, the obtained multi-modal causal features can be understood as features that can better reflect the sentiment. At the same time, since the extracted multi-modal features include features of different modalities, when filtering the multi-modal features, the features of different modalities can be filtered separately, so that the obtained multi-modal causal features also include causal features of different modalities.
[0040] Step S104: Calculate the causal effect using the sentiment classification loss of the multi-modal features and the sentiment classification loss of the multi-modal causal features, and the sentiment classification loss is determined based on the sentiment classification results of the multi-modal features and the multi-modal causal features by a preset classifier.
[0041] Among them, for the obtained multi-modal causal features, the influence of the causal features of different modalities included in them on model training is dynamically changing. Therefore, when training the model, it is also necessary to comprehensively consider the joint causal influence of different modalities in the multi-modal causal features on the model. Thus, in order to quantify the influence of the multi-modal causal features on sentiment classification, the causal effect of the multi-modal causal features can be calculated through counterfactual reasoning. Specifically, when calculating the causal effect, the multi-modal features and the multi-modal causal features can be predicted respectively by a preset classifier to obtain the corresponding sentiment classification results, and then the causal effect is determined based on the sentiment classification loss calculated from the sentiment classification results and the corresponding sentiment labels.
[0042] Step S105: Determine the joint loss using the scheduling weights determined by the causal effect and learnable parameters and the sentiment classification loss. Specifically, according to the meaning of the above causal effect, the larger the calculated causal effect, the smaller the loss the corresponding causal feature can bring to the model, indicating that the sample corresponding to the causal feature is easier to learn. Therefore, the larger the causal effect, the greater the weight can be assigned to it during training. Thus, the scheduling weight corresponding to the causal feature can be determined based on the calculated causal effect of the causal feature. And based on this scheduling weight and the corresponding sentiment classification loss, the loss function of the model is determined. Since this loss function is determined based on the scheduling weights and sentiment classification losses corresponding to causal features of different modalities, this loss function can also be called the joint loss.
[0043] Step S106: Use joint optimization to optimize the parameters in the preset neural network, preset classifier, and joint loss, and construct a multi-modal sentiment classification model based on the optimized preset neural network, learnable mask, and preset classifier. Specifically, during the model training process, it is necessary to optimize the parameters involved in the processing procedures in the above steps S102 to S105, such as the parameters in the preset neural network, preset classifier, and joint loss, to reduce the joint loss. Among them, during training, the processing procedures in the above steps S102 to S105 can be first executed based on the initial parameters, and the corresponding joint loss is determined. Then, based on this joint loss, the parameters are adjusted using joint optimization, and the above steps are executed again based on the adjusted parameters. Repeat the above process until the obtained joint loss meets the requirements. Thus, a multi-modal sentiment classification model can be constructed based on the preset neural network, learnable mask, and preset classifier with optimized parameters for prediction.
[0044] The multi-modal sentiment classification model training method provided by the embodiments of the present invention, during the model training process, uses a learnable mask to filter the multi-modal features in the sample to obtain multi-modal causal features, that is, selects multi-modal causal features through simple and effective causal intervention, alleviating the problem of spurious associations. At the same time, the multi-modal learning process is scheduled by evaluating the causal effect of the features, and combined with joint optimization to adaptively determine the optimal scheduling parameters. Thus, the trained multi-modal sentiment classification model realizes multi-modal feature learning based on causal relationships and improves the performance of multi-modal sentiment classification.
[0045] In this embodiment, a multi-modal sentiment classification model training method is provided, and the process includes the following steps:
[0046] Step S201: Obtain training samples including multi-modal samples. Specifically, each group of multi-modal samples obtained in this embodiment includes a text sentence X (i.e., text sample) containing n words and the corresponding image Xv (i.e., the image sample).
[0047] Step S202: Extract the multi-modal features of the multi-modal sample using a preset neural network. Specifically, in this embodiment, the open-source BERT (Bidirectional Encoder Representations from Transformers) model and the CLIP (Contrastive Language-Image Pre-Training) model are used as the basic preset neural network models for feature extraction.
[0048] Specifically, the above step S202 includes:
[0049] Step S2021: Extract the text features of the text sample using the BERT model; specifically, when performing text feature extraction, for the text sample, i.e., the given text sentence X t , use BERT t = Φ t (X t ) to obtain the [CLS] feature of the sentence as the sentence feature, where Φ t is the encoder of BERT, and m×d is the dimension of the text feature.
[0050] Step S2022: Extract the image features of the image sample using the CLIP model. Specifically, when performing image feature extraction, given the image sample where c×h×w is the dimension of the image, use the image encoder of CLIP to obtain the image feature
[0051] Step S203: Filter the multi-modal features using a learnable mask to obtain multi-modal causal features. Specifically, for the extracted text features and image features, the corresponding multi-modal causal features can be calculated using the following formula:
[0052]
[0053]
[0054] In the formula, and respectively represent the counterfactual text causal feature and the counterfactual image causal feature in the multi-modal causal feature, and respectively represent learnable masks with the same dimension as the text feature and the image feature, and MLP (Multilayer Perceptron) represents a multi-layer perceptron.
[0055] Step S204: Calculate the causal effect by using the sentiment classification loss of multi-modal features and the sentiment classification loss of multi-modal causal features. The sentiment classification loss is determined based on the sentiment classification results of the multi-modal features and the multi-modal causal features by a preset classifier.
[0056] Specifically, the above step S204 includes:
[0057] Step S2041: Use a preset classifier to determine the sentiment classification results of each modal feature and each modal causal feature. Specifically, the preset classifier can use a multi-layer perceptron. Among them, the sentiment classification results can be calculated by the following formula:
[0058]
[0059]
[0060]
[0061]
[0062] In the formula, respectively represent the sentiment classification results of the text feature and the image feature in the multi-modal features determined by using the preset classifier. respectively represent the sentiment classification results of the counterfactual text causal feature and the counterfactual image causal feature determined by using the preset classifier. f t represents the classifier of the text, and f v represents the classifier of the image. In addition, in this embodiment, the sentiment is divided into three types: negative, neutral, and positive. The sentiment classification results include any one or more of negative, neutral, or positive.
[0063] Step S2042: Calculate the sentiment classification loss corresponding to each modal feature and each modal causal feature based on the loss function corresponding to each modality, the sentiment classification result, and the corresponding sentiment label. Specifically, in this embodiment, the cross-entropy loss is used as the loss function corresponding to the text modality (i.e., the text loss L ce ), and the binary cross-entropy loss is used as the loss function corresponding to the image modality (i.e., the image loss L bce ). Thus, the sentiment classification loss can be expressed as: and where y t represents the sentiment label of the text; y v represents the sentiment label of the image. For the sentiment label, it can be marked with a triple (negative, neutral, positive). For example, if the text associated with the image contains both "neutral" and "positive" emotions at the same time, its label is defined as (0, 1, 1).
[0064] Step S2043, calculate the causal effect using the sentiment classification loss corresponding to each modal feature and each modal causal feature. Specifically, the causal effect can be calculated using the following formula:
[0065]
[0066]
[0067] In the formula, Δ∈ represents the causal effect.
[0068] Step S205, determine the joint loss using the scheduling weights determined by the causal effect and learnable parameters and the sentiment classification loss; specifically, by determining the scheduling weights through the causal effect and learnable parameters, the model can assign larger learning weights to samples with larger causal effects and smaller learning weights to samples with smaller causal effects. Thus, the determined scheduling weights can also be called learnable scheduling weights. The scheduling weights can be expressed using the following formulas respectively:
[0069]
[0070]
[0071] In the formula, and represent the scheduling weights corresponding to the text modality and the image modality respectively. σ represents the softplus function, α and β represent learnable parameters, and b represents the batch index. Among them, when training with training samples, the training samples will be injected into the model in batches by setting the batch size batchsize for training. For example, if batchsize is 6, it means that 6 samples are input into the model for training at one time. And the batch index represents which batch of samples are input into the model for training. For the samples input into the model in the first batch, b = 1.
[0072] The joint loss determined based on the scheduling weights and the sentiment classification loss can be calculated using the following formula:
[0073]
[0074] where Dtrain represents the training set.
[0075] Step S206, perform parameter optimization on the parameters of the preset neural network, the preset classifier, and the joint loss using joint optimization, and construct a multi-modal sentiment classification model based on the preset neural network, the learnable mask, and the preset classifier after parameter optimization. Specifically, when performing parameter optimization using joint optimization, it can be optimized by dividing it into the following two-level optimization problem shown in the following formula:
[0076]
[0077]
[0078] Among them, represents the joint loss on the validation set D dev φ = {α, β} represents the learnable parameters, and θ represents the parameters of the preset neural network, the learnable mask, and the parameters in the preset classifier. When performing two-level optimization, the above formula can be referred to. Given θ, find the value of φ that minimizes the loss of the selected validation set. Given φ, optimize the parameters to obtain the value of θ that minimizes the loss of the training set.
[0079] Specifically, the above step S206 includes:
[0080] Step S2061, optimize the first parameter using the preset learnable parameters. The first parameter includes the parameters of the preset neural network, the learnable mask, and the parameters in the preset classifier. Specifically, when updating the first parameter, the weighted gradient sum of samples in different modalities can be used, and it can be specifically implemented using the following formula:
[0081]
[0082] In the formula, represents the derivative of the first parameter θ.
[0083] Step S2062, optimize the learnable parameters based on the optimized first parameter. Specifically, when optimizing the learnable parameters, it is necessary to calculate the gradient of the loss function L dev (θ * (φ)) with respect to the learnable parameter φ. Since the dependence of L dev (θ * (φ)) on φ is indirect and can be achieved through the parameter θ, implicit differentiation can be used to obtain its implicit gradient. Based on the Cauchy implicit function theorem, the gradient of the loss function L dev (θ * (φ)) with respect to the learnable parameter φ is systematically derived using the chain rule, and this process can be calculated using the following formula:
[0084]
[0085] However, for deep neural network models, due to their huge parameter scale and complex solution, directly calculating the inverse of the Hessian matrix in the above formula is usually computationally infeasible. To solve this problem, this embodiment uses the K-th truncated Neumann series to approximate the inverse of the Hessian matrix. By using this approximation method to approximate the inverse of the Hessian matrix, the implicit gradient can be effectively estimated. The process is as shown in the following formula:
[0086]
[0087]
[0088] Step S2063: Repeat the above steps until the obtained joint loss meets the preset requirements. Specifically, repeat the processes of optimizing the parameters in the above steps S2062 and S2063, and after each optimization, execute the above steps S202 to S205 until the obtained joint loss meets the preset requirements. This preset requirement can be determined according to the actual situation. For example, it can be to meet the preset iteration accuracy.
[0089] In addition, the above-mentioned trained multi-modal sentiment classification model and the models used for multi-modal sentiment classification in related technologies can be compared based on two datasets, Twitter2015 and Twitter2017. By comparing their accuracy rates and F1 scores, the accuracy rate and F1 score of the multi-modal sentiment classification model trained in the embodiments of the present invention are superior to those of the models used for multi-modal sentiment classification in related technologies.
[0090] In this embodiment, a multi-modal sentiment classification method is also provided, which is applied to the multi-modal sentiment classification model trained by the multi-modal sentiment classification model training method in the above embodiment. As Figure 2 shown, the method includes the following steps:
[0091] Step S301: Obtain the data to be classified, and the data to be classified is multi-modal data.
[0092] Step S302: Use the preset neural network with optimized parameters to extract the features of the data to be classified.
[0093] Step S303: Filter the features using the learnable mask with optimized parameters to obtain causal features.
[0094] Step S304: Use the preset classifier with optimized parameters to predict the causal features and determine the sentiment classification result of the data to be classified.
[0095] Specifically, the data to be classified may include text data and corresponding image data in multimodal data. When extracting features from the image to be classified, the above-mentioned BERT model and CLIP model with optimized parameters can be used to extract features from the text data and image data respectively, and then the learnable mask with optimized parameters is used to filter the extracted text features and image features to obtain counterfactual text causal features and counterfactual image causal features. Since the learnable mask is optimized in terms of parameters during the above model training process, thus, the obtained counterfactual text causal features and counterfactual image causal features alleviate the problem of spurious associations in the text features and image features. Then, the classifier with optimized parameters is used to predict and classify the counterfactual text causal features and counterfactual image causal features, thereby obtaining a more accurate sentiment classification effect.
[0096] In this embodiment, a multimodal sentiment classification model training device is further provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0097] This embodiment provides a multimodal sentiment classification model training device, as Figure 3 shown, including:
[0098] A sample acquisition module 31, configured to acquire training samples including multimodal samples;
[0099] A feature extraction module 32, configured to extract multimodal features of multimodal samples by using a preset neural network;
[0100] A filtering module 33, configured to filter the multimodal features by using a learnable mask to obtain multimodal causal features;
[0101] A causal effect calculation module 34, configured to calculate a causal effect by using the sentiment classification loss of multimodal features and the sentiment classification loss of multimodal causal features, where the sentiment classification loss is determined based on the sentiment classification results of the multimodal features and the multimodal causal features by a preset classifier;
[0102] A joint loss determination module 35, configured to determine a joint loss by using a scheduling weight determined by the causal effect and learnable parameters and the sentiment classification loss;
[0103] A parameter optimization and model construction module 36, configured to perform parameter optimization on the parameters in the preset neural network, the preset classifier, and the joint loss by using joint optimization, and construct a multimodal sentiment classification model based on the preset neural network, the learnable mask, and the preset classifier with optimized parameters.
[0104] As a specific application example of the embodiment of the present invention, as Figure 4 shown, the multi-modal sentiment classification model training device includes a multi-modal causal feature selection module, a cross-modal causal scheduler module, and a parameter solving module based on two-level optimization.
[0105] Among them, the multi-modal causal feature selection module includes two sub-modules: task initialization and multi-modal causal feature selection, which are used to learn the causal features of different modalities. Specifically, in the task initialization sub-module, the BERT model and the CLIP model are respectively used to extract features from the text samples and image samples in the training samples, and the corresponding text features E t and image features E v are obtained. In the multi-modal causal feature selection sub-module, learnable masks are used to perform causal intervention on the text features and image features to obtain the corresponding causal features and Then, a text classifier and an image classifier are respectively used to predict and classify the extracted features and the corresponding causal features to obtain the corresponding sentiment classification results After that, the cross-entropy loss is used as the text loss, and the binary cross-entropy loss is used as the image loss to calculate the losses L ce and L bce between the corresponding sentiment classification results and the sentiment labels.
[0106] The cross-modal causal scheduler module adaptively adjusts the weights of different losses according to the joint causal influence of different modalities on the model, including two sub-modules: the causal effect of features and the joint scheduling target. Specifically, in the causal effect of features sub-module, the causal effect of the selected features is calculated through counterfactual reasoning. That is, by given the features E t 、E v of the text and image and the counterfactual causal features after causal intervention the corresponding causal effects are calculated (the causal effects corresponding to the text modality and the image modality shown in the figure are 0.6 and 0.7 respectively). In the joint scheduling sub-module, learnable scheduling weights and are determined through the causal effect and learnable parameters. Then, through the scheduling weights and the losses L ce and L bce calculated by the multi-modal causal feature selection sub-module, the cross-modal joint scheduling target loss L (i.e., the joint loss) is determined.
[0107] The parameter solving module based on two-level optimization aims to solve the optimal modal scheduling weights, including two sub-modules: inner optimization and outer optimization. Specifically, in the inner optimization sub-module, a fixed parameter φ is used to update the parameter θ of the multi-modal causal feature selection module. The update of θ is carried out by using the weighted gradient sum of samples in different modalities. In the outer optimization sub-module, it is necessary to calculate the gradient of the loss function L dev (θ * (φ)) with respect to the hyperparameter φ. Since the dependence of L dev (θ * (φ)) on φ is indirect and can be achieved through the parameter θ, implicit differentiation is used to obtain its implicit gradient. In Figure 4 , the trained parameters represent the parameters that need to be updated in backpropagation. The frozen parameters represent the parameters that do not change in backpropagation. θ0 represents the parameters before the first update. H th represents that after H times of inner optimization, an outer optimization is performed once. η1 and η2 are two hyperparameters, representing the weights in front of the entire gradient.
[0108] In this embodiment, a multi-modal sentiment classification device is also provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated here. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0109] This embodiment provides a multi-modal sentiment classification device, which is applied to the multi-modal sentiment classification model trained by the multi-modal sentiment classification model training method of the above-mentioned embodiment. As Figure 5 shown, the device includes:
[0110] A data acquisition module 51, configured to acquire data to be classified, and the data to be classified is multi-modal data;
[0111] A data feature extraction module 52, which extracts the features of the data to be classified by using a preset neural network with optimized parameters;
[0112] A causal feature determination module 53, configured to filter the features by using a learnable mask with optimized parameters to obtain causal features;
[0113] A prediction module 54, configured to predict the causal features by using a preset classifier with optimized parameters to determine the sentiment classification result of the data to be classified.
[0114] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding above-mentioned embodiments, and will not be repeated here.
[0115] An embodiment of the present invention further provides a computer device having the above-mentioned Figure 3 multi-modal emotion classification model training device shown or Figure 5 the multi-modal emotion classification device shown.
[0116] Please refer to Figure 6 , Figure 6 which is a schematic structural diagram of a computer device provided by an alternative embodiment of the present invention. As Figure 6 shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Each component communicates with each other using different buses and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some alternative embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a set of blade servers, or a multi-processor system). Figure 6 In
[0117] , a single processor 10 is taken as an example.
[0118] The processor 10 can be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 can further include a hardware chip. The above-mentioned hardware chip can be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device can be a complex programmable logic device, a field programmable gate array, a general array logic, or any combination thereof.
[0119] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.
[0120] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, a hard disk, or a solid state drive; the memory 20 may further include a combination of the above types of memory.
[0121] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0122] An embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented by downloading through a network the original computer code stored in a remote storage medium or a non-transitory machine-readable storage medium and to be stored in a local storage medium, so that the method described herein can be stored in such software processes on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Wherein, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid state drive, etc.; further, the storage medium may further include a combination of the above types of memory. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiment is implemented.
[0123] A part of the present invention can be applied as a computer program product, such as computer program instructions, which when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should be able to understand that the forms of existence of computer program instructions in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways for computer program instructions to be executed by a computer include, but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible by the computer.
[0124] Although the embodiments of the present invention are described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present invention, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for training a multi-modal sentiment classification model, characterized in that The method includes: Obtaining training samples containing multimodal samples; Using a preset neural network to extract multimodal features of the multimodal samples; Filtering the multimodal features using a learnable mask to obtain multimodal causal features; Calculating the causal effect using the sentiment classification loss of the multimodal features and the sentiment classification loss of the multimodal causal features, where the sentiment classification loss is determined based on the sentiment classification results of the multimodal features and the multimodal causal features by a preset classifier; Determining the joint loss using the scheduling weight determined by the causal effect and learnable parameters and the sentiment classification loss; Performing parameter optimization on the parameters in the preset neural network, the preset classifier, and the joint loss using joint optimization, and constructing a multimodal sentiment classification model based on the preset neural network, the learnable mask, and the preset classifier after parameter optimization.
2. The method according to claim 1, wherein The multimodal samples include text samples and image samples. Using a preset neural network to extract multimodal features of the multimodal samples includes: Using a BERT model to extract text features of the text samples; Using a CLIP model to extract image features of the image samples.
3. The method according to claim 1, wherein Calculating the causal effect using the sentiment classification loss of the multimodal features and the sentiment classification loss of the multimodal causal features includes: Using a preset classifier to determine the sentiment classification results of each modal feature and each modal causal feature; Calculating the sentiment classification loss corresponding to each modal feature and each modal causal feature based on the loss function corresponding to each modality, the sentiment classification results, and the corresponding sentiment labels; Calculating the causal effect using the sentiment classification loss corresponding to each modal feature and each modal causal feature.
4. The method according to claim 1, wherein The causal effect is represented by the following formula: where, Δ ∈ represents the causal effect, and represent the counterfactual text causal feature and the counterfactual image causal feature in the multi-modal causal feature respectively, y t represents the sentiment label of the text; y v represents the sentiment label of the image, represent the sentiment classification results of the text feature and the image feature in the multi-modal feature determined by the preset classifier respectively, represent the sentiment classification results of the counterfactual text causal feature and the counterfactual image causal feature determined by the preset classifier respectively, L ce represents the loss function corresponding to the text feature, L bce represents the loss function corresponding to the image feature.
5. The method according to claim 1, characterized in that Performing parameter optimization on the parameters in the preset neural network, the preset classifier, and the joint loss using joint optimization includes: Using preset learnable parameters to optimize the first parameters, where the first parameters include the parameters of the preset neural network, the learnable mask, and the parameters in the preset classifier; Optimizing the learnable parameters based on the optimized first parameters; Repeating the above steps until the obtained joint loss meets the preset requirements.
6. A multi-modal sentiment classification method, characterized in that, Applying to the multimodal sentiment classification model trained by the multimodal sentiment classification model training method according to any one of claims 1-5, the method includes: Obtaining data to be classified, where the data to be classified is multimodal data; Using the preset neural network after parameter optimization to extract features of the data to be classified; Filtering the features using the learnable mask after parameter optimization to obtain causal features; Using the preset classifier after parameter optimization to predict the causal features and determining the sentiment classification result of the data to be classified.
7. A training device for a multi-modal sentiment classification model, characterized in that, The device includes: A sample acquisition module for obtaining training samples containing multimodal samples; A feature extraction module for using a preset neural network to extract multimodal features of the multimodal samples; A filtering module for filtering the multimodal features using a learnable mask to obtain multimodal causal features; A causal effect calculation module for calculating the causal effect using the sentiment classification loss of the multimodal features and the sentiment classification loss of the multimodal causal features, where the sentiment classification loss is determined based on the sentiment classification results of the multimodal features and the multimodal causal features by a preset classifier; A joint loss determination module, configured to determine a joint loss by using a scheduling weight determined by the causal effect and learnable parameters and a sentiment classification loss; A parameter optimization and model construction module, configured to perform parameter optimization on parameters in a preset neural network, a preset classifier, and a joint loss by using joint optimization, and construct a multi-modal sentiment classification model based on the preset neural network, a learnable mask, and the preset classifier after parameter optimization.
8. A multimodal sentiment classification device, characterized in that, A multi-modal sentiment classification model obtained by training using the multi-modal sentiment classification model training method according to any one of claims 1-5, the apparatus comprising: A data acquisition module, configured to acquire data to be classified, where the data to be classified is multi-modal data; A data feature extraction module, configured to extract features of the data to be classified by using the preset neural network after parameter optimization; A causal feature determination module, configured to filter the features by using the learnable mask after parameter optimization to obtain causal features; A prediction module, configured to predict the causal features by using the preset classifier after parameter optimization to determine a sentiment classification result of the data to be classified.
9. A computer device, characterized in that, Comprising: A memory and a processor, the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the multi-modal sentiment classification model training method according to any one of claims 1 to 5 or the multi-modal sentiment classification method according to claim 6.
10. A computer-readable storage medium, characterized in that, Computer instructions are stored on the computer-readable storage medium, and the computer instructions are used to cause a computer to execute the multi-modal sentiment classification model training method according to any one of claims 1 to 5 or the multi-modal sentiment classification method according to claim 6.
Citation Information
Patent Citations
Cross-modal emotion causal tracking method
CN118779725A
Multi-Modal Content Based Automated Feature Recognition
US20230068502A1