Multi-mode-based sentiment classification method and device, electronic equipment and medium
By employing a multimodal sentiment classification method, feature extraction and feature weight adjustment are performed on multimodal dialogue data, solving the problem of information imbalance in multimodal sentiment classification and achieving higher sentiment classification accuracy.
Patent Information
- Application Number
- CN202511011086.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-11-07
AI Technical Summary
Traditional sentiment classification methods fail to adequately consider the imbalance of sentiment semantic information among multimodal data, resulting in low sentiment classification accuracy.
By acquiring features from multimodal dialogue data, we assess the sentiment bias between pairs of modal features, and dynamically adjust the weights of each modal feature using multi-head attention processing and weighted aggregation to achieve cross-modal information aggregation.
It improves the accuracy of multimodal sentiment classification by identifying and quantifying the contribution of each modality to sentiment expression, thus solving the problem of other modal sentiment cues being ignored due to statistical advantages between modalities.
Smart Images

Figure CN120910264A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, and is applied to financial scenarios and medical scenarios, and in particular relates to a multi-modal sentiment classification method and device, an electronic device and a medium. BACKGROUND
[0002] Traditional sentiment classification methods usually extract sentiment tendencies (such as negative, positive and neutral) in multi-modal dialogues through natural language generation (NLG) technology. For example, in the financial business scenario, for the dialogue in the insurance product introduction video, the multi-modal sentiment tendency of the customer can be extracted as a neutral facial expression, a cheerful tone of voice, and a negative sentiment implied text through the NLG technology. However, this method may ignore the emotional clues in the audio due to the statistical advantage of the visual or text modalities, fail to fully consider the imbalance of emotional semantic information between different modalities, and cannot accurately fuse the sentiment tendencies of each modality, resulting in low accuracy of sentiment classification. Therefore, how to improve the accuracy of sentiment classification has become a problem to be solved. SUMMARY
[0003] The main purpose of the embodiments of the present application is to propose a multi-modal sentiment classification method and device, an electronic device and a medium, aiming to improve the efficiency of multi-modal sentiment classification.
[0004] To achieve the above-mentioned purpose, the first aspect of the embodiments of the present application proposes a multi-modal sentiment classification method, which comprises:
[0005] Obtaining original dialogue data, and extracting features of the original dialogue data through a pre-trained sentiment classification model to obtain multi-modal dialogue features; wherein the original dialogue data is multi-modal data; the multi-modal dialogue features include at least two of dialogue image features, dialogue text features and dialogue audio features;
[0006] According to the sentiment deviation evaluation between the dialogue image features, the dialogue text features and the dialogue audio features, a sentiment deviation score is obtained;
[0007] According to the sentiment deviation score, multi-head attention processing is performed on each modality feature in the multi-modal dialogue feature to obtain a multi-head attention weight of each modality;
[0008] According to the multi-head attention weight of each modality, the multi-modal dialogue feature is preliminarily aggregated to obtain a preliminarily aggregated modality feature;
[0009] The preliminarily aggregated modality feature and the multi-modal dialogue feature are weighted and aggregated to obtain a target aggregated modality feature;
[0010] According to the target poly-modal feature, the original dialogue data is classified according to the dialogue sentiment, and dialogue sentiment classification data is obtained.
[0011] In some embodiments, the sentiment bias evaluation between two modal features is performed according to the dialogue image feature, the dialogue text feature, and the dialogue audio feature, and a sentiment bias score is obtained.
[0012] According to the dialogue image feature, the sentiment contribution degree is calculated, and the image sentiment contribution degree is obtained.
[0013] According to the dialogue text feature, the sentiment contribution degree is calculated, and the text sentiment contribution degree is obtained.
[0014] According to the dialogue audio feature, the sentiment contribution degree is calculated, and the audio sentiment contribution degree is obtained.
[0015] The image sentiment contribution degree, the text sentiment contribution degree, and the audio sentiment contribution degree are calculated, and the modal contribution degree difference is obtained.
[0016] According to the modal contribution degree difference, the image sentiment contribution degree, the text sentiment contribution degree, and the audio sentiment contribution degree, the modal sentiment bias evaluation is performed, and the sentiment bias score is obtained.
[0017] In some embodiments, the modal contribution degree difference includes image-text contribution degree difference, image-audio contribution degree difference, and text-audio contribution degree difference.
[0018] The sentiment bias score includes image-text sentiment bias score, image-audio sentiment bias score, and audio-text sentiment bias score.
[0019] According to the modal contribution degree difference, the image sentiment contribution degree, the text sentiment contribution degree, and the audio sentiment contribution degree, the modal sentiment bias evaluation is performed, and the sentiment bias score is obtained.
[0020] The maximum sentiment contribution degree is selected from the image sentiment contribution degree, the text sentiment contribution degree, and the audio sentiment contribution degree as the target sentiment contribution degree.
[0021] The image-text contribution degree difference is multiplied by the target sentiment contribution degree to obtain the image-text sentiment bias score.
[0022] The image-audio contribution degree difference is multiplied by the target sentiment contribution degree to obtain the image-audio sentiment bias score.
[0023] The text-audio contribution degree difference is multiplied by the target sentiment contribution degree to obtain the audio-text sentiment bias score.
[0024] In some embodiments, the multi-head attention processing of each modality feature in the multi-modal dialogue feature according to the sentiment bias score comprises:
[0025] The multi-head attention allocation is performed on each modality feature in the multi-modal dialogue feature to obtain modality attention features;
[0026] A modality bias adjustment factor is obtained, and the multi-head modality attention constraint is performed on the modality attention features according to the modality bias adjustment factor and the sentiment bias score to obtain the multi-head attention weight of each modality.
[0027] In some embodiments, the weighted aggregation of the preliminary aggregated modality feature and the multi-modal dialogue feature to obtain the target aggregated modality feature comprises:
[0028] The dialogue image feature, the dialogue text feature, and the dialogue audio feature are spliced to obtain a multi-modal spliced feature;
[0029] A balance parameter of the multi-modal spliced feature is obtained, and the multi-modal spliced feature and the preliminary aggregated modality feature are weighted fused according to the balance parameter to obtain the target aggregated modality feature.
[0030] In some embodiments, before the feature extraction of the original dialogue data by the pre-trained sentiment classification model to obtain the multi-modal dialogue feature, the method further comprises:
[0031] An original sentiment classification model and a training multi-modal sample set are obtained, and a true sentiment classification label corresponding to each sample in the training multi-modal sample set is labeled;
[0032] The training multi-modal sample set is extracted by the original sentiment classification model to obtain training multi-modal features; the training multi-modal features comprise training image features, training text features, and training audio features;
[0033] The sentiment bias between two modalities is evaluated according to the training image features, the training text features, and the training audio features to obtain a training sentiment bias score;
[0034] The multi-head attention processing of each modality feature in the training multi-modal feature according to the training sentiment bias score comprises:
[0035] The training multi-modal features are preliminarily aggregated according to the training multi-head attention weight of each modality to obtain training preliminary aggregated modality features;
[0036] aggregate the training preliminary aggregated modality feature and the training multi-modal feature by weighted aggregation to obtain a training target aggregated modality feature;
[0037] perform dialogue sentiment classification prediction on the training multi-modal sample set according to the training target aggregated modality feature to obtain a predicted sentiment classification label;
[0038] calculate a target loss value of the training multi-head attention weight of each modality, the real sentiment classification label and the predicted sentiment classification label according to a preset loss function;
[0039] update model parameters of the original sentiment classification model based on the target loss value.
[0040] In some embodiments, the calculating a target loss value of the training multi-head attention weight of each modality, the real sentiment classification label and the predicted sentiment classification label according to a preset loss function comprises:
[0041] perform attention weight variance loss calculation on the training multi-head attention weight of each modality according to the loss function to obtain an attention weight loss value;
[0042] perform label loss calculation on the real sentiment classification label and the predicted sentiment classification label to obtain a sentiment loss value;
[0043] perform weighted summation according to the attention weight loss value and the sentiment loss value to obtain the target loss value.
[0044] To achieve the above object, a second aspect of the embodiment of the present application proposes a sentiment classification device based on multi-modalities, which comprises:
[0045] a multi-modal feature extraction module, configured to acquire original dialogue data, and perform feature extraction on the original dialogue data by a pre-trained sentiment classification model to obtain multi-modal dialogue features; wherein the original dialogue data is multi-modal data; and the multi-modal dialogue features comprise at least two of dialogue image features, dialogue text features and dialogue audio features;
[0046] a modality sentiment bias evaluation module, configured to perform sentiment bias evaluation between two modalities among the dialogue image features, the dialogue text features and the dialogue audio features to obtain a sentiment bias score;
[0047] a multi-head attention processing module, configured to perform multi-head attention processing on each modality feature in the multi-modal dialogue features according to the sentiment bias score to obtain a multi-head attention weight of each modality;
[0048] a preliminary aggregation module configured to preliminarily aggregate the multi-modal dialogue features according to the multi-modal multi-head attention weights to obtain preliminary aggregated modal features;
[0049] a weighted aggregation module configured to weighted aggregate the preliminary aggregated modal features and the multi-modal dialogue features to obtain target aggregated modal features;
[0050] a dialogue sentiment classification module configured to perform dialogue sentiment classification on the original dialogue data according to the target aggregated modal features to obtain dialogue sentiment classification data.
[0051] To achieve the above object, a third aspect of the embodiment of the present application provides an electronic device, which comprises a memory and a processor, the memory stores a computer program, and the processor implements the method of the first aspect when executing the computer program.
[0052] To achieve the above object, a fourth aspect of the embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the method of the first aspect.
[0053] The multi-modal based sentiment classification method and device, electronic device and medium provided by the present application firstly acquire original dialogue data and perform feature extraction on the original dialogue data, which can realize extraction of different modal features; secondly, sentiment bias evaluation is performed between two modal features according to dialogue image features, dialogue text features and dialogue audio features, which can effectively identify and quantify the contribution degree of each modal in sentiment expression, fully considers the sentiment semantic information imbalance problem between different modalities, and performs multi-head attention processing on each modal feature in the multi-modal dialogue feature according to the sentiment bias score, which realizes dynamic adjustment of the weight of each modal feature and helps to solve the problem of ignoring other modal sentiment clues due to the statistical advantage between modalities; further, the multi-modal dialogue features are preliminarily aggregated according to the multi-modal multi-head attention weights, and the preliminary aggregated modal features and the multi-modal dialogue features are weighted aggregated, which can retain the original modal features while realizing cross-modal information aggregation to accurately fuse the sentiment tendency of each modality; finally, the original dialogue data is classified according to the target aggregated modal features, which realizes the contribution balance of each modal feature in sentiment classification, thereby improving the sentiment classification accuracy based on multi-modal. BRIEF DESCRIPTION OF DRAWINGS
[0054] Figure 1 is a flowchart of the multi-modal based sentiment classification method provided by the embodiment of the present application;
[0055] Figure 2 is another flowchart of the multi-modal based sentiment classification method provided by the embodiment of the present application;
[0056] Figure 3 yes Figure 2 The flowchart of step S208 in the text;
[0057] Figure 4 yes Figure 1 The flowchart of step S102 in the document;
[0058] Figure 5 yes Figure 4 The flowchart of step S405 in the document;
[0059] Figure 6 yes Figure 1 The flowchart of step S103 in the process;
[0060] Figure 7 yes Figure 1 The flowchart of step S105 in the process;
[0061] Figure 8 This is a schematic diagram of the structure of the multimodal emotion classification device provided in the embodiments of this application;
[0062] Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0063] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0064] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0065] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0066] First, let's analyze some of the terms used in this application:
[0067] Artificial intelligence (AI): is a new technical science of studying, developing theories, methods, technologies and application systems for simulating, extending and expanding human intelligence; artificial intelligence is a branch of computer science, artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. The research in this field includes robots, language recognition, image recognition, natural language processing and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence is also the theory, method, technology and application system of using digital computer or digital computer controlled machine to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain the best results.
[0068] The embodiment of the present application provides a multi-modal based emotion classification method and device, electronic equipment and medium, aiming to improve the efficiency of multi-modal based emotion classification.
[0069] The multi-modal based emotion classification method and device, electronic equipment and medium provided by the embodiment of the present application are specifically explained by the following embodiment, first, the multi-modal based emotion classification method in the embodiment of the present application is described.
[0070] The embodiment of the present application can acquire and process related data based on artificial intelligence technology. Wherein, artificial intelligence (AI) is a theory, method, technology and application system of using digital computer or digital computer controlled machine to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain the best results.
[0071] The basic technology of artificial intelligence generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. The software technology of artificial intelligence mainly includes computer vision technology, robot technology, biometric technology, speech processing technology, natural language processing technology and machine learning / deep learning, etc.
[0072] The emotion classification method based on multiple modalities provided by the embodiments of the present application relates to the technical field of artificial intelligence. The emotion classification method based on multiple modalities provided by the embodiments of the present application can be applied to a terminal, can also be applied to a server end, and can also be software running in the terminal or the server end. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.; the server end can be configured as a separate physical server, can also be configured as a server cluster or a distributed system formed by multiple physical servers, and can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDNs, and big data and artificial intelligence platforms; and the software can be an application that implements the emotion classification method based on multiple modalities, etc., but is not limited to the above forms.
[0073] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment in which tasks are performed by remote processing devices connected by a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0074] Figure 1 is an optional flowchart of the emotion classification method based on multiple modalities provided by the embodiments of the present application, Figure 1 The method in the above embodiment can include but is not limited to including steps S101 to S106.
[0075] Step S101, obtaining original dialogue data, and performing feature extraction on the original dialogue data through a pre-trained emotion classification model to obtain multi-modal dialogue features; wherein the original dialogue data is multi-modal data; the multi-modal dialogue features include at least two of dialogue image features, dialogue text features, and dialogue audio features.
[0076] Step S102, performing emotion bias evaluation between two modalities according to the dialogue image features, the dialogue text features, and the dialogue audio features to obtain an emotion bias score.
[0077] In step S103, the multi-head attention processing is performed on each modality feature in the multi-modal dialogue feature according to the sentiment bias score, to obtain a multi-head attention weight of each modality.
[0078] In step S104, the multi-modal dialogue feature is preliminarily aggregated according to the multi-head attention weight of each modality, to obtain a preliminarily aggregated modality feature.
[0079] In step S105, the preliminarily aggregated modality feature and the multi-modal dialogue feature are weighted and aggregated, to obtain a target aggregated modality feature.
[0080] In step S106, the dialogue sentiment classification is performed on the original dialogue data according to the target aggregated modality feature, to obtain dialogue sentiment classification data.
[0081] The steps S101 to S106 shown in the embodiments of the present application first acquire the original dialogue data, and perform feature extraction on the original dialogue data, so that the extraction of different modality features can be realized. Secondly, the sentiment bias evaluation is performed between two modality features according to the dialogue image feature, the dialogue text feature and the dialogue audio feature, so that the contribution degree of each modality in the sentiment expression can be effectively identified and quantified. The sentiment semantic information imbalance problem between different modalities is fully considered, and the multi-head attention processing is performed on each modality feature in the multi-modal dialogue feature according to the sentiment bias score, so that the dynamic adjustment of the feature weight of each modality is realized, which helps to solve the problem of ignoring other modality sentiment clues due to the statistical advantage between modalities. Further, the multi-modal dialogue feature is preliminarily aggregated according to the multi-head attention weight of each modality, and the weighted aggregation is performed on the preliminarily aggregated modality feature and the multi-modal dialogue feature, so that the original modality feature can be retained while realizing the cross-modal information aggregation, to accurately fuse the sentiment tendency of each modality. Finally, the dialogue sentiment classification is performed on the original dialogue data according to the target aggregated modality feature, so that the contribution balance of each modality feature in the sentiment classification is realized, thereby improving the sentiment classification accuracy based on the multi-modal.
[0082] Please refer to Figure 2 In some embodiments, before the feature extraction is performed on the original dialogue data by the pre-trained sentiment classification model to obtain the multi-modal dialogue feature, the sentiment classification based on the multi-modal further includes but is not limited to steps S201 to S209:
[0083] In step S201, an original sentiment classification model and a training multi-modal sample set are acquired, and a real sentiment classification label corresponding to each sample in the training multi-modal sample set is labeled.
[0084] In step S202, the feature extraction is performed on the training multi-modal sample set by the original sentiment classification model, to obtain a training multi-modal feature. The training multi-modal feature includes a training image feature, a training text feature and a training audio feature.
[0085] In step S203, sentiment bias evaluation between two modalities is performed according to the training image features, the training text features and the training audio features, to obtain training sentiment bias scores.
[0086] In step S204, multi-head attention processing is performed on each modality feature in the training multi-modal features according to the training sentiment bias scores, to obtain training multi-head attention weights of each modality.
[0087] In step S205, preliminary aggregation is performed on the training multi-modal features according to the training multi-head attention weights of each modality, to obtain preliminary aggregated modality features.
[0088] In step S206, weighted aggregation is performed on the preliminary aggregated modality features and the training multi-modal features, to obtain training target aggregated modality features.
[0089] In step S207, dialog sentiment classification prediction is performed on the training multi-modal sample set according to the training target aggregated modality features, to obtain predicted sentiment classification labels.
[0090] In step S208, a target loss value of the training multi-head attention weights of each modality, the real sentiment classification labels and the predicted sentiment classification labels is calculated according to a preset loss function.
[0091] In step S209, the model parameters of the original sentiment classification model are updated based on the target loss value.
[0092] In step S201 of some embodiments, specifically, the original sentiment classification model includes a modality feature extraction layer, a modality sentiment bias analysis layer, a multi-head cross-attention layer, a modality feature aggregation layer, a loss function and an output layer. The modality feature extraction layer is used to extract multi-modal dialog features and linearly project the multi-modal dialog features. The modality sentiment bias analysis layer is used to perform sentiment bias evaluation between two modalities of dialog image features, dialog text features and dialog audio features. The multi-head cross-attention layer is used to calculate multi-head attention weights of each modality and preliminarily aggregate the multi-modal dialog features according to the multi-head attention weights of each modality. The modality feature aggregation layer is used to perform weighted aggregation on the preliminary aggregated modality features and the multi-modal dialog features. The loss function is used to train the original sentiment classification model. The output layer is used to output the sentiment category of the original dialog data.
[0093] For example, in a financial service scenario, the training multi-modal sample set can be customer facial expression images in an insurance product introduction video, dialog text between an insurance agent and a customer, and voice content of the customer, and the real sentiment classification label can be a positive label. In a medical service scenario, the training multi-modal sample set can be patient expression images in a remote consultation video, text describing the patient's condition, and voice between the doctor and the patient, and the real sentiment classification label can be a negative label.
[0094] In step S202 of some embodiments, specifically, the training multi-modal features can be extracted by the modal feature extraction layer.
[0095] For example, in the financial service scenario, for the insurance product introduction video, the training multi-modal features can be extracted by the modal feature extraction layer, including the facial expression features of the customer from the image modal, the dialogue text features from the training dialogue text modal, and the dialogue audio features from the training audio modal; in the medical service scenario, for the remote medical consultation video, the training multi-modal features can be extracted by the modal feature extraction layer, including the patient expression features from the patient expression image modal, the consultation dialogue text features from the text content describing the illness, and the consultation dialogue audio features from the audio modal of the doctor and the patient.
[0096] In step S203 of some embodiments, specifically, the training image features, the training text features, and the training audio features can be evaluated for the emotional bias between two modalities by the modal emotional bias analysis layer.
[0097] For example, in the financial service scenario, the modal emotional bias analysis layer can compare the emotional difference between the facial expression of the customer and the dialogue text content, and if the facial expression of the customer is neutral but the dialogue text content expresses positive emotion, the emotional bias evaluation can identify this difference and give the corresponding training emotional bias score; in the medical scenario, the modal emotional bias analysis layer can compare the emotional difference between the expression of the patient and the consultation dialogue text, and if the expression of the patient is neutral but the consultation dialogue text expresses negative emotion, the emotional bias evaluation can also identify this difference and give the corresponding training emotional bias score.
[0098] In this embodiment, the emotional bias evaluation between two modalities based on the training image features, the training text features, and the training audio features can effectively identify and quantify the contribution of each modality in emotional expression, fully consider the imbalance of emotional semantic information between different modalities, and avoid ignoring the emotional clues of other modalities due to the statistical advantage of a certain modality.
[0099] In step S204 of some embodiments, specifically, the training multi-modal features can be calculated by the multi-head cross-attention layer to obtain the training multi-modal multi-head attention weights.
[0100] For example, in the financial service scenario, the multi-head cross-attention layer can adjust the weights of the customer's facial expression, text content and voice tone according to the sentiment bias score. If the sentiment bias score shows that the customer's facial expression has a greater emotional contribution, the multi-head cross-attention layer will correspondingly increase the weight of the customer's facial expression. In the medical service scenario, the multi-head cross-attention layer can adjust the weights of the patient's expression, the text of the inquiry conversation and the voice communication according to the sentiment bias score. If the sentiment bias score shows that the text of the inquiry conversation has a greater emotional contribution, the multi-head cross-attention layer will correspondingly increase the weight of the text of the inquiry conversation.
[0101] In step S205 of some embodiments, specifically, the trained multi-modal features can be preliminarily aggregated by the multi-head cross-attention layer.
[0102] For example, in the financial service scenario, the preliminarily aggregated modal features can be the product conversation aggregation features obtained by weighted summation of the customer's facial expression, the text content of the insurance product introduction conversation and the audio features of the insurance product introduction conversation after the trained multi-modal multi-head attention weight adjustment. In the medical service scenario, the preliminarily aggregated modal features can be the inquiry conversation aggregation features obtained by weighted summation of the patient's expression, the text of the inquiry conversation and the audio features of the text conversation after the trained multi-modal multi-head attention weight adjustment.
[0103] In step S206 of some embodiments, specifically, the trained preliminarily aggregated modal features and the trained multi-modal features can be weighted and aggregated by the modal feature aggregation layer.
[0104] For example, in the financial service scenario, the weighted aggregated modal features can be the product conversation aggregation features further aggregated by the modal feature aggregation layer and the original multi-modal product conversation features. In the medical service scenario, the weighted aggregated modal features can be the inquiry conversation aggregation features further aggregated by the modal feature aggregation layer and the original multi-modal inquiry conversation features.
[0105] In step S207 of some embodiments, specifically, the output layer can output the predicted sentiment classification label of the training multi-modal sample set.
[0106] For example, in the financial service scenario, the training multi-modal sample set is predicted to be a positive sentiment classification label according to the training target aggregated modal features. In the medical service scenario, the sentiment state of the inquiry patient can be predicted to be a negative sentiment classification label according to the training target aggregated modal features.
[0107] Please refer to Figure 3 In some embodiments, step S208 includes but is not limited to steps S301 to S303:
[0108] Step S301, attention weight variance loss calculation is performed on the training of each modality multi-head attention weight according to the loss function, to obtain an attention weight loss value.
[0109] Step S302, label loss calculation is performed on the real emotion classification label and the predicted emotion classification label, to obtain an emotion loss value.
[0110] Step S303, weighted summation is performed according to the attention weight loss value and the emotion loss value, to obtain a target loss value.
[0111] In step S301 of some embodiments, specifically, the loss function includes an attention weight variance loss function and an emotion classification loss function.
[0112] Specifically, the attention weight loss value can be calculated by the following attention weight variance loss function:
[0113]
[0114] Wherein, L con represents the attention weight loss value; Var represents a variance function, used to calculate the dispersion degree of the attention weight; ξ represents a threshold parameter, which can be 0.2, and is not limited here; represents the modality attention weight set of the modality m of the kth attention head to all other modalities m ′ When the variance is lower than the threshold value (indicating that the attention distribution is too concentrated), the attention weight variance loss function will produce a penalty, so that the emotion classification model more evenly allocates attention.
[0115] In this embodiment, the attention weight variance loss calculation is performed on the training of each modality multi-head attention weight according to the loss function, which can effectively identify and quantify the imbalance problem between the feature weights of each modality, and ensure the balanced contribution of each modality feature in emotion classification.
[0116] In step S302 of some embodiments, for example, in a financial business scenario, if the emotion classification model predicts that the emotion classification label of the customer in the insurance product introduction conversation is positive, but the real emotion classification label is negative, the emotion loss value is lower; in a medical business scenario, if the emotion classification model predicts that the emotion classification label of the patient in the consultation conversation is neutral, and the real emotion classification label is also neutral, the emotion loss value is higher.
[0117] In step S303 of some embodiments, specifically, the target loss value can be calculated by the following formula:
[0118] L=L task +λL con
[0119] Wherein, L represents the target loss value; Ltask represents an emotion loss value; λ represents a balance coefficient for controlling the weight of the contrast loss; L con represents an attention weight loss value.
[0120] Through steps S301 to S303, the attention weight loss and the emotion classification task loss can be comprehensively considered through the weighted sum of the attention weight loss value and the emotion loss value, so as to ensure that the emotion classification model considers the balance of the modal feature weights in the optimization process, which helps to improve the accuracy of subsequent emotion classification.
[0121] In step S209 of some embodiments, specifically, the model parameters of the emotion classification model can be adjusted through a back propagation algorithm according to the calculated target loss value to optimize the performance of the emotion classification model.
[0122] Through steps S201 to S209, the model performance of the emotion classification model can be effectively optimized in combination with the loss function, so as to effectively consider the balance of the modal feature weights by the emotion classification model, so that the emotion classification model can better balance the contributions of each modality in the multi-modal emotion classification task, which helps to improve the accuracy of subsequent emotion classification.
[0123] In step S101 of some embodiments, specifically, the original dialogue data is multi-modal data, which can include at least two modal data of images, texts and audios.
[0124] For example, in a financial business scenario, for a video of an insurance product introduction, the original dialogue data can include customer facial expression images during the introduction of the insurance product, dialogue text content and dialogue voice, etc.; in a medical business scenario, for a video of remote medical consultation, the original dialogue data can include patient expression images during the consultation, consultation dialogue text content and voice of the doctor and the patient, etc.
[0125] Specifically, the dialogue image modal can be feature-extracted through a CNN (Convolutional Neural Network) of a modal feature extraction layer to determine dialogue image features, and the dialogue text modal can be processed through word embedding to extract dialogue text features, and further, the dialogue audio modal can be converted into dialogue audio text through an ASR (Automatic Speech Recognition) technology, and then the dialogue audio text can be processed through word embedding to extract dialogue audio features.
[0126] Further, after obtaining the multi-modal dialogue feature, feature projection needs to be performed on the multi-modal dialogue feature to map the dialogue image feature, dialogue text feature and dialogue audio feature to a latent space with the same dimension, and the projected multi-modal feature is taken as the final multi-modal dialogue feature.
[0127] Specifically, the feature projection can be realized by the following formula:
[0128] z m =ReLU(W m h m +b m )
[0129] wherein z m represents the projected dialogue feature vector of the modality m; ReLU represents a rectified linear unit activation function (defined as ReLU(x) = max(0, x)) for introducing nonlinearity; Wm m ∈R d represents the bias vector of the modality m; d m is the spatial vector dimension of the original dialogue feature h m before projection of the modality m; d represents the unified target spatial vector dimension after feature projection.
[0130] In this embodiment, by performing feature projection on the multi-modal dialogue feature, it can be ensured that the features of different modalities have consistent spatial vector dimensions, which is helpful for realizing subsequent cross-modal dialogue feature interaction.
[0131] Please refer to Figure 4 In some embodiments, step S102 includes but is not limited to steps S401 to S405:
[0132] Step S401: performing emotion contribution degree calculation according to the dialogue image feature to obtain an image emotion contribution degree.
[0133] Step S402: performing emotion contribution degree calculation according to the dialogue text feature to obtain a text emotion contribution degree.
[0134] Step S403: performing emotion contribution degree calculation according to the dialogue audio feature to obtain an audio emotion contribution degree.
[0135] Step S404: performing contribution degree difference calculation between two modalities on the image emotion contribution degree, text emotion contribution degree and audio emotion contribution degree to obtain a modality contribution degree difference.
[0136] Step S405: performing modality emotion bias evaluation according to the modality contribution degree difference, image emotion contribution degree, text emotion contribution degree and audio emotion contribution degree to obtain an emotion bias score.
[0137] In step S401 of some embodiments, specifically, the image sentiment contribution degree refers to the accuracy rate of sentiment classification prediction only through the dialogue image modality.
[0138] Specifically, the dialogue image features can be subjected to sentiment classification prediction by the single-modality sentiment classification network in the modal sentiment bias analysis layer to obtain a predicted image sentiment category, and a dialogue image verification set is acquired, and the dialogue image features are subjected to sentiment classification accuracy rate calculation according to the dialogue image verification set and the predicted image sentiment category to obtain the image sentiment contribution degree.
[0139] For example, in a financial service scenario, the customer's facial expression (i.e., dialogue image features) during the insurance introduction period is subjected to sentiment classification prediction by the single-modality sentiment classification network to obtain a predicted image sentiment category (such as positive), and a dialogue image verification set is acquired, and the dialogue image features are subjected to sentiment classification accuracy rate calculation according to the dialogue image verification set and the predicted image sentiment category, if the dialogue image verification set shows that the customer's sentiment is positive, and the predicted image sentiment category is also positive, then the image sentiment contribution degree is high, indicating that the customer's facial expression plays a positive role in sentiment expression.
[0140] In this embodiment, the image sentiment contribution degree is obtained by calculating the sentiment contribution degree according to the dialogue image features, which can quantify the contribution degree of the dialogue image features modality in sentiment classification, and provide contribution data of the image modality in sentiment classification for subsequent sentiment bias evaluation.
[0141] In step S402 of some embodiments, specifically, the text sentiment contribution degree refers to the accuracy rate of sentiment classification prediction only through the dialogue text modality.
[0142] Specifically, the dialogue text features can also be subjected to sentiment classification prediction by the single-modality sentiment classification network in the modal sentiment bias analysis layer to obtain a predicted text sentiment category, and a dialogue text verification set is acquired, and the dialogue text features are subjected to sentiment classification accuracy rate calculation according to the dialogue text verification set and the predicted text sentiment category to obtain the text sentiment contribution degree.
[0143] In step S403 of some embodiments, specifically, the audio sentiment contribution degree refers to the accuracy rate of sentiment classification prediction only through the dialogue audio modality.
[0144] Specifically, the dialogue audio features can also be subjected to sentiment classification prediction by the single-modality sentiment classification network in the modal sentiment bias analysis layer to obtain a predicted audio sentiment category, and a dialogue audio verification set is acquired, and the dialogue audio features are subjected to sentiment classification accuracy rate calculation according to the dialogue audio verification set and the predicted audio sentiment category to obtain the audio sentiment contribution degree.
[0145] In step S404 of some embodiments, specifically, the modal contribution difference refers to the modal contribution difference obtained by pairwise modal difference operation between the image sentiment contribution, the text sentiment contribution and the audio sentiment contribution, and is used to reflect the difference between the sentiment contributions of the modalities.
[0146] Specifically, in the financial service scenario, the difference in sentiment between the customer's facial expression during product introduction (such as image sentiment contribution 0.8) and the product introduction dialogue text (such as text sentiment contribution 0.6) can be compared to determine the image text contribution difference as 0.2; in the medical scenario, the difference in sentiment between the patient's expression during remote consultation (such as image sentiment contribution 0.85) and the doctor-patient dialogue voice (such as audio sentiment contribution 0.7) can be compared to determine the image audio contribution difference as 0.15.
[0147] Please refer to Figure 5 In some embodiments, the modal contribution difference includes the image text contribution difference, the image audio contribution difference and the text audio contribution difference; the sentiment bias score includes the image text sentiment bias score, the image audio sentiment bias score and the text audio sentiment bias score, and step S405 includes but is not limited to steps S501 to S504:
[0148] Step S501: selecting the maximum sentiment contribution from the image sentiment contribution, the text sentiment contribution and the audio sentiment contribution as the target sentiment contribution.
[0149] Step S502: performing a multiplication operation on the image text contribution difference and the target sentiment contribution to obtain the image text sentiment bias score.
[0150] Step S503: performing a multiplication operation on the image audio contribution difference and the target sentiment contribution to obtain the image audio sentiment bias score.
[0151] Step S504: performing a multiplication operation on the text audio contribution difference and the target sentiment contribution to obtain the text audio sentiment bias score.
[0152] In step S501 of some embodiments, specifically, the target sentiment contribution refers to the maximum sentiment contribution selected from the image sentiment contribution, the text sentiment contribution and the audio sentiment contribution, and is used to reflect the most significant modal information.
[0153] For example, in the financial service scenario, if the image sentiment contribution of the customer's facial expression emotion is 0.8, the text sentiment contribution of the insurance product introduction dialogue text is 0.6, and the audio sentiment contribution of the insurance product introduction dialogue audio is 0.7, then the target sentiment contribution is the image sentiment contribution, which is 0.8.
[0154] In step S502 of some embodiments, specifically, the image-text sentiment bias score is used to quantify the difference between the image modality and the text modality in terms of sentiment expression, for subsequent adjustment of attention weights.
[0155] Specifically, the image-text sentiment bias score can be calculated by the following formula:
[0156]
[0157] wherein G m,m′ represents the sentiment bias score from modality m to modality m ′ , ranging between [-1, 1], wherein if the sentiment bias score is positive, it indicates that modality m is relatively more important, and if the sentiment bias score is negative, it indicates that modality m ′ is relatively more important; acc m represents the sentiment contribution of modality m; acc m′ represents the sentiment contribution of modality m ′ ; acc V represents the image sentiment contribution; acc a represents the audio sentiment contribution; acc t represents the text sentiment contribution; and max(acc v , acc a , acc t ) represents the target sentiment contribution.
[0158] In step S503 of some embodiments, specifically, the image-audio sentiment bias score is used to quantify the difference between the image modality and the audio modality in terms of sentiment expression, for subsequent adjustment of attention weights.
[0159] Specifically, the method of performing the division operation of the image-audio contribution difference value and the target sentiment contribution is consistent with the method of performing the division operation of the image-text contribution difference value and the target sentiment contribution, which will not be repeated here.
[0160] In step S504 of some embodiments, specifically, the audio-text sentiment bias score is used to quantify the difference between the text modality and the audio modality in terms of sentiment expression, for subsequent adjustment of attention weights.
[0161] Specifically, the method of performing the division operation of the text-audio contribution difference value and the target sentiment contribution is consistent with the method of performing the division operation of the image-text contribution difference value and the target sentiment contribution, which will not be repeated here.
[0162] By selecting the target contribution degree and calculating the quotient of the contribution degree difference between each modal pair and the target contribution degree, the emotional bias score is obtained through steps S501 to S504, which can effectively evaluate the emotional bias between each modal, ensure that the imbalance between each modal is fully considered in multi-modal emotion classification, that is, ensure that the emotional clues of each modal are fully considered in the emotion classification process, which helps to improve the accuracy of subsequent emotion classification.
[0163] By steps S401 to S405, the emotional bias between each modal is evaluated, which can comprehensively consider the emotional contribution of each modal and the emotional contribution difference between each modal, ensure that the emotional contribution imbalance between each modal is fully considered in the emotion classification process, provide data support for subsequent dynamic adjustment of attention weight and multi-modal feature fusion, and thus improve the accuracy of subsequent emotion classification.
[0164] Please refer to Figure 6 In some embodiments, step S103 includes but is not limited to steps S601 to S602:
[0165] Step S601, multi-head attention is allocated to each modal feature in the multi-modal dialogue feature to obtain each modal attention feature.
[0166] Step S602, a modal bias adjustment factor is obtained, and multi-head modal attention constraint is performed on each modal attention feature according to the modal bias adjustment factor and the emotional bias score to obtain each modal multi-head attention weight.
[0167] In step S601 of some embodiments, specifically, the query-key value generation of each modal feature in the multi-modal dialogue feature can be performed by a Transformer network of a multi-head cross attention layer to determine each modal attention feature.
[0168] Specifically, each modal attention feature can be determined by the following formula:
[0169] Q m =z m W Q ,{K m′ ,V m′}=z m′ W K ,z m′ W v (m ′ ≠m)
[0170] Wherein, Q m represents the query feature vector of modal m, z m represents the dialogue feature vector of modal m, z m′ represents the dialogue feature vector of modal m ′Dialogue feature vector, V m′ Represents mode m ′ The value of the eigenvector, This represents the learnable parameter matrix of the query feature vector. The learnable parameter matrix representing the key feature vectors. The learnable parameter matrix d represents the eigenvectors of the values. h =d / K represents the spatial vector dimension of a single head, K m′ Represents mode m ′ The key feature vector.
[0171] In step S602 of some embodiments, specifically,
[0172] Specifically, the multi-head attention weights for each modality can be determined using the following formula:
[0173]
[0174] in, This indicates that mode m in the k-th attention head is related to all other modes m. ′ Attention weights; Softmax represents the normalization function (which can be defined as...) e represents the natural base, x i Let x represent the feature vector of the i-th modality m dialogue. j (represents the m-th modality dialogue feature vector), used to convert the input dialogue feature vector into a probability distribution; Let m represent the mode of the k-th attention head. ′ The transpose of the key eigenvectors; γ represents the modal deviation adjustment factor, used to control the strength of the modal constraints; G m,m′ Indicates the transition from mode m to mode m ′ The emotional bias score.
[0175] Through steps S601 to S602, by using the modal constraints of multi-modal attention, the sentiment classification model can dynamically adjust the attention weights of each modality according to the sentiment deviation between each modality, ensuring that the sentiment cues of each modality are fully considered during the sentiment classification process, which helps to further improve the accuracy of sentiment classification.
[0176] In step S104 of some embodiments, specifically, the preliminary aggregated modal features refer to the fused feature vector that comprehensively considers the dialogue image features, dialogue text features, and dialogue audio features after multi-head attention weight adjustment.
[0177] Specifically, the preliminary aggregation modal characteristics can be determined using the following formula:
[0178]
[0179] f m =LayerNorm(z m +Concat(head1,…,head k ))
[0180] Among them, head k This represents the output of k attention heads; This indicates that mode m in the k-th attention head is related to all other modes m. ′ Attention weights; Let m represent the mode of the k-th attention head. ′ The eigenvector matrix of the value; LayerNorm represents the layer normalization operation; z m Let f represent the dialogue feature vector of modality m; Concat represents the feature vector concatenation operation to merge the outputs of k attention heads; m ∑ represents the initial aggregated modal features of modality m after cross-attention; ∑ represents summation.
[0181] In this embodiment, the multimodal dialogue features are initially aggregated according to the multi-head attention weights of each modality to obtain preliminary aggregated modal features. This realizes the integration of multimodal dialogue features according to their importance based on the dynamically adjusted multi-head attention weights, forming a comprehensive feature representation that can more comprehensively reflect the emotional information in multimodal dialogue and further improve the accuracy of subsequent emotion classification.
[0182] Please see Figure 7 In some embodiments, step S105 includes, but is not limited to, steps S701 to S702:
[0183] Step S701: The dialogue image features, dialogue text features, and dialogue audio features are concatenated to obtain multimodal concatenated features.
[0184] Step S702: Obtain the balance parameters of the multimodal splicing features, and perform weighted fusion of the multimodal splicing features and the preliminary aggregated modal features according to the balance parameters to obtain the target aggregated modal features.
[0185] In step S701 of some embodiments, specifically, dialogue image features, dialogue text features, and dialogue audio features can be sequentially connected to form multimodal splicing features.
[0186] For example, if the dialogue image feature is a feature vector of length 128, the dialogue text feature is a feature vector of length 256, and the dialogue audio feature is a feature vector of length 128, then the concatenated multimodal concatenated feature will be a feature vector of length 512.
[0187] In step S702 of some embodiments, specifically, the feature fusion representation of the target aggregated modal feature can be generated by residual connection of the modal feature aggregation layer and modal weighted aggregation.
[0188] Specifically, the target aggregated modal feature can be determined by the following formula:
[0189]
[0190] wherein F represents the target aggregated modal feature; w m is a weight coefficient representing the modal m; f m represents the preliminary aggregated modal feature of the modal m after cross attention; MLP represents a multi-layer perceptron for establishing high-order modal interaction; β represents a learnable balance parameter for controlling the proportion of direct weighting of multi-modal features and high-order modal interaction; f v represents the dialogue image feature; f a represents the dialogue audio feature; f t represents the dialogue text feature; Concat(f v ,f a ,f t ) represents the multi-modal splicing feature; · represents dot product.
[0191] Through steps S701 to S702, the original modal feature can be preserved while cross-modal information aggregation is realized through feature weighted fusion, so as to accurately fuse the sentiment tendency of each modal, further ensure that the sentiment clues of each modal can be reasonably applied in the sentiment classification process, and thus further improve the accuracy of subsequent sentiment classification.
[0192] In step S106 of some embodiments, specifically, the dialogue sentiment classification data refers to the sentiment classification result obtained according to the target aggregated modal feature, which is usually represented in the form of class label.
[0193] For example, in the financial scenario, the dialogue sentiment classification data can be a sentiment classification label of the customer having a “positive” willingness to purchase an insurance product; in the medical scenario, the dialogue sentiment classification data can be a sentiment classification label of the patient having a “positive” willingness to treat the disease.
[0194] Specifically, the output layer can be used to output the sentiment class probability corresponding to the target aggregated modal feature through the activation function, and the sentiment class with the highest probability can be taken as the dialogue sentiment classification data.
[0195] The embodiment of the application first acquires original dialogue data, and performs feature extraction on the original dialogue data, which can realize the extraction of different modal features; secondly, the emotional bias between two modal features is evaluated according to the dialogue image features, dialogue text features and dialogue audio features, which can effectively identify and quantify the contribution degree of each modal in emotional expression, fully considers the emotional semantic information imbalance problem between different modalities, and performs multi-head attention processing on each modal feature in the multi-modal dialogue feature according to the emotional bias score, realizes the dynamic adjustment of the weight of each modal feature, and helps to solve the problem of ignoring other modal emotional clues due to the statistical advantage between modalities; further, the multi-modal dialogue features are preliminarily aggregated according to the multi-head attention weights of each modal, and the preliminarily aggregated modal features and the multi-modal dialogue features are weighted and aggregated, which can retain the original modal features while realizing cross-modal information aggregation, so as to accurately fuse the emotional tendencies of each modal; finally, the original dialogue data is classified according to the target aggregated modal features, which realizes the contribution balance of each modal feature in emotional classification, thereby improving the emotional classification accuracy based on multi-modal.
[0196] Please refer to Figure 8 The embodiment of the application also provides a multi-modal based emotional classification device, which can realize the above-mentioned multi-modal based emotional classification method, and the device comprises:
[0197] A multi-modal feature extraction module is configured to acquire original dialogue data, and perform feature extraction on the original dialogue data through a pre-trained emotional classification model to obtain multi-modal dialogue features; wherein the original dialogue data is multi-modal data; and the multi-modal dialogue features comprise at least two of dialogue image features, dialogue text features and dialogue audio features.
[0198] A modal emotional bias evaluation module is configured to evaluate the emotional bias between two modal features according to the dialogue image features, dialogue text features and dialogue audio features to obtain an emotional bias score.
[0199] A multi-head attention processing module is configured to perform multi-head attention processing on each modal feature in the multi-modal dialogue feature according to the emotional bias score to obtain multi-head attention weights of each modal.
[0200] A preliminary aggregation module is configured to preliminarily aggregate the multi-modal dialogue features according to the multi-head attention weights of each modal to obtain preliminarily aggregated modal features.
[0201] A weighted aggregation module is configured to perform weighted aggregation on the preliminarily aggregated modal features and the multi-modal dialogue features to obtain target aggregated modal features.
[0202] An emotional classification module is configured to classify the dialogue emotion according to the target aggregated modal features to obtain dialogue emotional classification data.
[0203] The specific implementation of the multi-modal based sentiment classification device is basically the same as the specific implementation of the multi-modal based sentiment classification method described above, and will not be repeated here.
[0204] The embodiments of the present application also provide an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor implements the multi-modal based sentiment classification method when executing the computer program. The electronic device can be any intelligent terminal, such as a tablet computer or a vehicle-mounted computer.
[0205] Please refer to Figure 9 , Figure 9 The hardware structure of the electronic device of another embodiment is illustrated, which includes:
[0206] The processor 901 can be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits, and is used to execute related programs to implement the technical solutions provided by the embodiments of the present application.
[0207] The memory 902 can be implemented in the form of a ROM (Read Only Memory), a static storage device, a dynamic storage device, or a RAM (Random Access Memory). The memory 902 can store a processing system and other application programs. When the technical solutions provided by the embodiments of the present application are implemented by software or firmware, the related program codes are stored in the memory 902 and are called and executed by the processor 901 to implement the multi-modal based sentiment classification method of the embodiments of the present application.
[0208] The input / output interface 903 is used to realize information input and output.
[0209] The communication interface 904 is used to realize the communication interaction between the device and other devices. The communication can be realized by a wired manner (such as USB, network cable, etc.) or a wireless manner (such as mobile network, WIFI, Bluetooth, etc.).
[0210] The bus 905 transmits information between various components (such as the processor 901, the memory 902, the input / output interface 903, and the communication interface 904) of the device.
[0211] The processor 901, the memory 902, the input / output interface 903, and the communication interface 904 are connected to each other through the bus 905 for communication within the device.
[0212] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to realize the multi-modal based sentiment classification method.
[0213] The memory, as a non-transitory computer readable storage medium, can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory disposed remotely relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0214] The embodiments described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the evolution of technology and the appearance of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0215] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and can include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0216] The device embodiments described above are only schematic, and the units described as separate components can or can not be physically separate, that is, can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments of the present application.
[0217] Those skilled in the art can understand that all or some of the steps in the above disclosed method, the functional modules / units in the system and the device can be implemented as software, firmware, hardware and their appropriate combinations.
[0218] The terms "first", "second", "third", "fourth", and the like in the description of this application and in the claims hereof, if any, are used for distinguishing between similar elements and not necessarily for describing a particular sequential or chronological order. It is to be understood that the use of the terms so termed herein is solely for the convenience of the reader and does not limit the scope of the application. It is also to be understood that the description and examples in this application are intended to cover all possible combinations where any of the several elements can represent one or more elements.
[0219] It should be understood that, in this application, "at least one" means one or more, "multiple" means two or more. "And / or" is used to describe the relationship between associated objects, which means that there can be three relationships, for example, "A and / or B" can mean: only A, only B, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c, can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0220] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative, for example, the division of the above-mentioned units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. The coupling or direct coupling or communication connection between the displayed or discussed each other can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0221] The units described above as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on multiple network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiment of the present application.
[0222] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.
[0223] When the integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in part, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes multiple instructions used to cause a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods in the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various other media that can store programs.
[0224] The preferred embodiments of the embodiments of the present application are described above with reference to the accompanying drawings, and the scope of the embodiments of the present application is not limited thereto. Any modification, equivalent replacement and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the embodiments of the present application.
Claims
1. A multi-modal based emotion classification method, characterized in that, The method comprises: obtaining original dialogue data, and performing feature extraction on the original dialogue data by using a pre-trained sentiment classification model to obtain multi-modal dialogue features; wherein the original dialogue data is multi-modal data; and the multi-modal dialogue features comprise at least two of dialogue image features, dialogue text features and dialogue audio features; performing sentiment bias evaluation between two modal features according to the dialogue image features, the dialogue text features and the dialogue audio features to obtain a sentiment bias score; performing multi-head attention processing on each modal feature in the multi-modal dialogue features according to the sentiment bias score to obtain a multi-head attention weight of each modal; performing preliminary aggregation on the multi-modal dialogue features according to the multi-head attention weight of each modal to obtain preliminary aggregated modal features; performing weighted aggregation on the preliminary aggregated modal features and the multi-modal dialogue features to obtain target aggregated modal features; performing dialogue sentiment classification on the original dialogue data according to the target aggregated modal features to obtain dialogue sentiment classification data.
2. The method of claim 1, wherein, The sentiment bias evaluation between two modal features according to the dialogue image features, the dialogue text features and the dialogue audio features to obtain a sentiment bias score comprises: performing sentiment contribution degree calculation on the dialogue image features to obtain an image sentiment contribution degree; performing sentiment contribution degree calculation on the dialogue text features to obtain a text sentiment contribution degree; performing sentiment contribution degree calculation on the dialogue audio features to obtain an audio sentiment contribution degree; performing contribution degree difference calculation between two modalities on the image sentiment contribution degree, the text sentiment contribution degree and the audio sentiment contribution degree to obtain a modal contribution degree difference; performing modal sentiment bias evaluation according to the modal contribution degree difference, the image sentiment contribution degree, the text sentiment contribution degree and the audio sentiment contribution degree to obtain the sentiment bias score.
3. The method of claim 2, wherein, The modal contribution degree difference comprises an image-text contribution degree difference, an image-audio contribution degree difference and a text-audio contribution degree difference; The sentiment bias score comprises an image-text sentiment bias score, an image-audio sentiment bias score and an audio-text sentiment bias score; The modal sentiment bias evaluation according to the modal contribution degree difference, the image sentiment contribution degree, the text sentiment contribution degree and the audio sentiment contribution degree to obtain the sentiment bias score comprises: selecting the maximum sentiment contribution degree from the image sentiment contribution degree, the text sentiment contribution degree and the audio sentiment contribution degree as a target sentiment contribution degree; performing quotient operation on the image-text contribution degree difference and the target sentiment contribution degree to obtain the image-text sentiment bias score; performing quotient operation on the image-audio contribution degree difference and the target sentiment contribution degree to obtain the image-audio sentiment bias score; performing quotient operation on the text-audio contribution degree difference and the target sentiment contribution degree to obtain the audio-text sentiment bias score.
4. The method of claim 1, wherein, The multi-head attention processing on each modal feature in the multi-modal dialogue features according to the sentiment bias score to obtain a multi-head attention weight of each modal comprises: The multi-modal dialogue feature is subjected to multi-head attention allocation, and each modality attention feature is obtained. A modality bias adjustment factor is obtained, and multi-head modality attention constraint is performed on the each modality attention feature according to the modality bias adjustment factor and the sentiment bias score, so as to obtain the each modality multi-head attention weight.
5. The method of claim 1, wherein, The preliminary aggregated modality feature and the multi-modal dialogue feature are subjected to weighted aggregation, and a target aggregated modality feature is obtained. The dialogue image feature, the dialogue text feature and the dialogue audio feature are subjected to feature splicing, and a multi-modal spliced feature is obtained. A balance parameter of the multi-modal spliced feature is obtained, and the multi-modal spliced feature and the preliminary aggregated modality feature are subjected to weighted fusion according to the balance parameter, so as to obtain the target aggregated modality feature.
6. The method of claim 1, wherein, Before the multi-modal dialogue feature is obtained by performing feature extraction on the original dialogue data through the pre-trained sentiment classification model, the method further comprises: An original sentiment classification model and a training multi-modal sample set are obtained, and a true sentiment classification label corresponding to each sample in the training multi-modal sample set is labeled. The training multi-modal sample set is subjected to feature extraction through the original sentiment classification model, and a training multi-modal feature is obtained; the training multi-modal feature comprises a training image feature, a training text feature and a training audio feature. Sentiment bias evaluation between two modalities is performed according to the training image feature, the training text feature and the training audio feature, and a training sentiment bias score is obtained. Multi-head attention processing is performed on each modality feature in the training multi-modal feature according to the training sentiment bias score, and a training each modality multi-head attention weight is obtained. The training multi-modal feature is subjected to preliminary aggregation according to the training each modality multi-head attention weight, and a training preliminary aggregated modality feature is obtained. The training preliminary aggregated modality feature and the training multi-modal feature are subjected to weighted aggregation, and a training target aggregated modality feature is obtained. Dialogue sentiment classification prediction is performed on the training multi-modal sample set according to the training target aggregated modality feature, and a predicted sentiment classification label is obtained. A target loss value of the training each modality multi-head attention weight, the true sentiment classification label and the predicted sentiment classification label is calculated according to a preset loss function. The model parameters of the original sentiment classification model are updated based on the target loss value.
7. The method of claim 6, wherein, The target loss value of the training each modality multi-head attention weight, the true sentiment classification label and the predicted sentiment classification label is calculated according to a preset loss function, and comprises: Attention weight variance loss calculation is performed on the training each modality multi-head attention weight according to the loss function, and an attention weight loss value is obtained. Label loss calculation is performed on the true sentiment classification label and the predicted sentiment classification label, and a sentiment loss value is obtained. Weighted summation is performed according to the attention weight loss value and the sentiment loss value, and the target loss value is obtained. 8.A multi-modal based emotion classification apparatus, characterized in that, The device comprises: The multi-modal feature extraction module is configured to obtain original dialogue data, and perform feature extraction on the original dialogue data by using a pre-trained sentiment classification model to obtain multi-modal dialogue features; wherein the original dialogue data is multi-modal data; and the multi-modal dialogue features include at least two of dialogue image features, dialogue text features, and dialogue audio features. The modal sentiment bias evaluation module is configured to perform sentiment bias evaluation between two modal features according to the dialogue image features, the dialogue text features, and the dialogue audio features to obtain a sentiment bias score. The multi-head attention processing module is configured to perform multi-head attention processing on each modal feature in the multi-modal dialogue features according to the sentiment bias score to obtain a multi-head attention weight of each modal. The preliminary aggregation module is configured to perform preliminary aggregation on the multi-modal dialogue features according to the multi-head attention weight of each modal to obtain preliminary aggregated modal features. The weighted aggregation module is configured to perform weighted aggregation on the preliminary aggregated modal features and the multi-modal dialogue features to obtain target aggregated modal features. The dialogue sentiment classification module is configured to perform dialogue sentiment classification on the original dialogue data according to the target aggregated modal features to obtain dialogue sentiment classification data.
9. An electronic device, comprising: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the multi-modal based sentiment classification method of any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program is executed by the processor to implement the multi-modal based sentiment classification method of any one of claims 1 to 7.