Multi-task sentiment analysis method based on cross-modal text enhancement

Through the cross-modal text enhancement module and the single-modal label generation module, the problem of low modal fusion efficiency and insufficient generalization ability in multimodal sentiment analysis is solved, and more efficient sentiment analysis results are achieved, suitable for human-computer interaction and mental health monitoring.

CN120387071APending Publication Date: 2025-07-29HEFEI UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510477575.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing multimodal emotion analysis methods have problems such as low modal fusion efficiency, defects in explicit alignment, insufficient model generalization ability, strong label dependence and redundancy of attention mechanisms across modal noise interference, which is difficult to meet the needs of human-computer interaction and mental health.

Method used

The cross-modal text enhancement module and After-Bert module are used to process multimodal features, and the explicit alignment between modals is reduced through the softmaxone cross-attention mechanism, the single-modal label generation module is introduced to self-supervise the generation of tags, and the target loss function is used to optimize model parameters, simplifying the dominant role of text modals and improving the generalization ability of the model.

Benefits of technology

It improves the accuracy and generalization ability of sentiment analysis, reduces the cost of manual labeling, enhances the transplantability and robustness of the model, and provides more reliable technical support for human-computer interaction and mental health monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387071A_ABST
    Figure CN120387071A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-task sentiment analysis method based on cross-modal text enhancement. The method comprises the following steps: respectively obtaining a video feature vector, an audio feature vector and a text feature vector; respectively inputting each vector into a cross-modal text enhancement module and an Aft-Bert module for processing to obtain a multi-modal feature and a single-modal feature; a single-mode label generation module is introduced, a single-mode label is generated based on the multi-mode features and the single-mode features, and model parameters are optimized through a target loss function. According to the invention, based on the strong emotion guidance of the text mode, the influence of the non-text mode on the text mode is also considered, and the long-distance important information of the voice and video modes is respectively captured by adopting a cross attention mechanism; the accuracy, the cost efficiency, the generalization ability, the practical application range and the like of the sentiment analysis task are remarkably improved, and more reliable technical support is provided for scenes such as human-computer interaction and mental health monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of cross-modal multi-task sentiment analysis, and in particular to a multi-task sentiment analysis method based on cross-modal text enhancement. Background Art

[0002] With the rapid progress of natural language processing technology, sentiment recognition has gradually emerged and become one of the research focuses. The research community has begun to use various forms of data such as text, sound, and facial expressions to perform sentiment recognition. Traditional single-modal text sentiment analysis methods, which rely on a single data source of text, although considering words, phrases, and their semantic connections during the recognition process, still have their inherent limitations and sometimes it is difficult to capture the diversity and complexity of human emotions. To overcome these limitations, recent research hotspots have turned to sentiment recognition methods that fuse multiple data sources, such as combining text, sound, and video materials, aiming to obtain a more comprehensive and accurate understanding of emotions through this multi-modal approach. Sentiment recognition methods based on multi-modal data have currently become an important research school in this field, covering numerous technologies and methods including convolutional neural networks, recurrent neural networks, graph neural networks, etc. For example, the CLIP model aims to be trained through a large number of text-image pairs, thereby learning to understand image content and being able to match this content with corresponding natural language descriptions. Its core idea is to use the method of contrastive learning, which is an unsupervised or weakly supervised learning method that learns representations by minimizing the distance between positive samples while maximizing the distance between negative samples. Another example is the MulT model, which realizes the interaction and fusion between different modalities through paired cross-modal Transformer modules. Specifically, the MulT model processes the input multi-modal data as queries, keys, and values respectively, transforms the vector space through a weight matrix and converts the feature dimensions to keep the dimensions of the queries and keys consistent. This mechanism allows the model to focus on the interaction relationships between multi-modal sequences at different time steps and implicitly adapt to the alignment of the data.

[0003] Although existing multi-modal methods have significantly improved the effectiveness of sentiment analysis, there are still the following core problems: (1) Low efficiency of modal fusion. Traditional methods (such as MulT) regard text, speech, and vision as equally important, ignoring the dominant role of the text modality. For example, in ironic scenarios, there may be contradictions between text and non-text information, and it is necessary to dynamically adjust the influence of other modalities based on the text. Explicit alignment defects: Existing methods rely on explicit modal alignment (such as time synchronization), but in actual scenarios, multi-modal data often has asynchrony (such as speech delay or expression lag). (2) Insufficient model generalization ability and poor pre-training compatibility: Existing fusion modules are difficult to seamlessly integrate into text pre-training models such as BERT, resulting in high migration costs. Strong label dependence: Most methods require manual annotation of single-modal labels, but in actual applications, the annotation cost of non-text labels (such as speech emotion labels) is high and subjective. (3) Cross-modal noise interference. Redundant attention mechanism: When the traditional softmax function calculates cross-modal attention, it forces attention on irrelevant information (such as background noise or non-emotion-related speech features), resulting in a decrease in the robustness of the model.

[0004] In reality, sentiment analysis has expanded from single-modal sentiment analysis to multi-modal interactive sentiment analysis. In the future, with the evolution of more modal large model technologies, sentiment analysis will be closer to real human interaction scenarios and become a key technology in fields such as human-computer interaction and mental health. Innovative cross-modal multi-task sentiment analysis is needed. Summary of the Invention

[0005] Aiming at the above existing problems, the purpose of the present invention is to provide a multi-task sentiment analysis method based on cross-modal text enhancement, which expands from single-modal sentiment analysis to multi-modal interactive sentiment analysis.

[0006] An embodiment of the present invention provides a multi-task sentiment analysis method based on cross-modal text enhancement, including the steps:

[0007] S1. Obtain video feature vectors, audio feature vectors, and text feature vectors respectively;

[0008] S2. Input the video feature vectors, audio feature vectors, and text feature vectors into a cross-modal text enhancement module and an After-Bert module respectively for processing to obtain multi-modal features, and at the same time perform multiple perceptron processes on the video feature vectors and audio feature vectors respectively to obtain single-modal features;

[0009] S3. Introduce a single-modal label generation module, generate single-modal labels based on the multi-modal features and single-modal features, and optimize the model parameters through an objective loss function.

[0010] Further, the acquisition of the video feature vectors in S1 includes:

[0011] The OpenFace tool is used to divide the input image into several small regions and calculate the gradient direction distribution of each region, and a simplified feature map for extracting the basic structure of the feature face is obtained;

[0012] Subsequently, key point localization is performed on the detected human face, and the spatial coordinates of the facial feature points are predicted through iterative optimization to achieve accurate facial geometric alignment;

[0013] Finally, a deep convolutional neural network is used to vectorize the features of the aligned face. After vectorization, the video feature vector is expressed as:

[0014]

[0015] where, T v is the length of the visual modality, and d v is the dimension of the visual modality.

[0016] The acquisition of the audio feature vector includes: using the COVAREP tool to extract audio features, and the audio features are all standardized by the zero-mean and variance normalization method. The obtained audio feature vector is expressed as:

[0017]

[0018] where, T a is the length of the speech modality, and d a is the dimension of the speech modality.

[0019] The acquisition of the text feature vector includes: inserting [CLS] and [SEP] into the corresponding positions of the text sequence through a tokenizer, and mapping them to the corresponding text ids. The mapped text sequence is input into the Pre-Bert model, and the obtained text feature vector is

[0020] Furthermore, the cross-modal text enhancement module has a processing flow including: inputting the text feature vector, through the softmaxone cross-attention mechanism, the text feature vector performs cross-modal information interaction with the video feature vector and the audio feature vector respectively, and then inputs it into a fully connected layer to improve the fusion effect between modalities, obtains a non-linear displacement vector, and after passing through a residual connection block, performs iterative processing to update the text feature vector.

[0021] Furthermore, the softmaxone cross-attention mechanism is expressed by the formula:

[0022]

[0023] where, represents the output text modality feature vector, represents the feature vector of the non-text modality u, u ∈ {v, a}; Tt and T u represent the sample lengths of the text modality and the non - text modality respectively, and d k is the sample dimension constant.

[0024] Furthermore, the non - linear displacement vector is expressed by the formula:

[0025]

[0026] where represents the operation of dimension concatenation on two matrices, W H and b H are the weight matrix and the bias parameter of the linear transformation respectively, and σ represents the Relu activation function.

[0027] Furthermore, the processing of the residual connection block is expressed by the formula:

[0028]

[0029] where α is the scaling factor, and H i is the non - text displacement vector.

[0030] Furthermore, the processing of the After - Bert module is expressed by the formula:

[0031]

[0032] where G t is the multi - modal feature.

[0033] Furthermore, the single - modality label generation module ULGM is introduced, which is used to learn the similarity information of modalities from multi - modal features, learn the difference information between modalities from single - modal features, and supervise the generation of single - modality task labels.

[0034] Furthermore, generating the single - modality label in S3 includes the steps:

[0035] S31. Create the positive and negative sample centers of the multi - modal feature, the video - modality feature, and the audio - modality feature, which is expressed by the formula:

[0036]

[0037] where m ∈ {t, v, a}, and represent the positive sample center and the negative sample center of modality m respectively, N represents the total number of samples, G m,n is the modality feature, representing the feature vector of modality m for sample n, y m is the modality label, and I(y(x)) is the function indicator, which incorporates the sample labels that meet the condition y(x) into its subset elements.

[0038] S32. Calculate modal characteristics G m,n The cosine similarity with the center of positive and negative samples is:

[0039]

[0040] Among them, m∈{t,a,v}, Represents the modal feature G m,n Similarity with the center of the positive sample; Represents the modal feature G m,n Similarity with the center of the negative sample; ||·|| represents the L2 norm operation on the vector;

[0041] S33, introduced the relative difference μ m , the formula is:

[0042]

[0043] Among them, μ m Represents the modal feature G m,n and the positive sample center and negative sample centers relative gap.

[0044] S34. Calculate the unimodal label. The formula is:

[0045]

[0046] Among them, u∈{a,v}, y t represents the multimodal label, y u represents a unimodal label, μ u and μ t Represent the relative differences between unimodal tasks and multimodal tasks, respectively.

[0047] S35. Dynamically update the single-modal tag. The dynamic update strategy formula is:

[0048]

[0049] in, represents the unimodal generated label obtained after the dynamic update strategy, Represents the newly generated unimodal label in the e-th iteration.

[0050] Furthermore, the optimization is performed through the objective loss function, and the formula is expressed as:

[0051]

[0052] Where N represents the total number of training samples. is the predicted value of the cross-modal fusion model for the nth sample in the previous subsection, represents the multi-modal sample label, is the weight adjustment strategy.

[0053] Advantages of the present invention:

[0054] 1. The present invention uses the powerful emotion orientation of the text modality as the basis, and also takes into account the influence of the non-text modality on the text modality, avoiding the misguidance of the emotion analysis task caused by only using the text modality in contexts such as irony and sarcasm; in the interaction between the text modality and the non-text modality, the present invention adopts a cross-attention mechanism based on softmaxone to capture the long-distance important information of the speech and video modalities respectively, avoiding the explicit alignment work between modalities, and reducing the attention of the text modality to irrelevant information through the softmaxone function.

[0055] 2. The cross-modal text enhancement module adopted by the present invention can be simply inserted into various text pre-training models based on Transformer, such as BERT, XLNet, RoBERTa, etc., and has strong portability and versatility. The addition of the single-modal label generation module generates single-modal labels in a self-supervised manner, reducing the additional cost and time brought by manual annotation of single-modal labels. This module also converts the original multi-modal emotion prediction task into multiple sub-tasks, learning the consistency and specificity between modalities and improving the generalization ability of the model.

[0056] 3. The present invention brings significant improvements in aspects such as the accuracy, cost efficiency, generalization ability, and actual application scope of the emotion analysis task through technologies such as cross-modal fusion, self-supervised label generation, and multi-task learning, providing more reliable technical support for scenarios such as human-computer interaction and mental health monitoring. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 is a schematic flow chart of a multi-task emotion analysis method based on cross-modal text enhancement of the present invention;

[0058] Figure 2 is a structural flow chart of a multi-task emotion analysis method based on cross-modal text enhancement of the present invention;

[0059] Figure 3 is a structural flow chart of the cross-modal text enhancement module of the present invention;

[0060] Figure 4 is a structural flow chart of the cross-attention mechanism based on softmaxone of the present invention;

[0061] Figure 5 is a schematic structural diagram of an electronic device of the present invention. Detailed implementation manners

[0062] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar symbols represent the same or similar elements or elements with the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.

[0063] The existing sentiment analysis methods cannot meet the needs in the fields of human-computer interaction, mental health, etc.

[0064] In view of the above problems, the present invention provides a multi-task sentiment analysis method based on cross-modal text enhancement. Figure 1 As shown in the flowchart of the multi-task sentiment analysis method based on cross-modal text enhancement provided by the embodiment of the present invention, the method includes the steps:

[0065] S1. Obtain video feature vector U v and audio feature vector U a and text feature vector

[0066] As Figure 2 shown, in the feature extraction stage, the open-source tools OpenFace and COVAREP are used to extract features from the original video and speech sequences respectively, and the BERT tokenizer is used to preprocess the original text sequence.

[0067] For the original video signal, the Openface tool uses the histogram of oriented gradients algorithm for face detection. By dividing the input image into several small regions and calculating the gradient direction distribution of each region, a simplified feature map of the basic facial structure is extracted; subsequently, a method based on an ensemble regression tree is used to perform key point localization on the detected face, and the spatial coordinates of 68 facial feature points are predicted through iterative optimization to achieve accurate facial geometric alignment; finally, a deep convolutional neural network is used to vectorize the aligned face. After the vectorized video features pass through the BiLSTM network, the video feature vector is

[0068] For the audio signal, the COVAREP acoustic analysis framework is used to extract audio features. Some of the features include mel-frequency cepstral coefficients, pitch, volume, glottal source parameters, and other features related to emotion and intonation. All features are standardized by the zero-mean and variance normalization method. Finally, the obtained high-dimensional audio features pass through the BiLSTM network, and the audio feature vector is

[0069] For the feature extraction and feature encoding of text sequences, the Per-Bert pre-trained model is used as the text feature encoder, which can provide rich semantic information for the text modality. Assume that the original text data is a text sequence S = {w1, w2, …, w n}, it is necessary to first insert placeholders such as [CLS] and [SEP] into the corresponding positions of the sequence through a tokenizer, and map the words and placeholders to the corresponding ids; the mapped text id sequence is then input into the embedding layer of the Per-Bert model, and the finally obtained text feature vector is

[0070] S2. Input the video feature vector U v and the audio feature vector U a and the text feature vector into the cross-modal text enhancement module and the After-Bert module respectively for processing to obtain the multi-modal feature G t ; perform multiple perceptron processes on the video feature vector and the audio feature vector respectively to obtain the single-modal features G v and G a .

[0071] The obtained video feature vector is U v and the audio feature vector is U a and the text feature vector is are respectively input into the cross-modal text enhancement module.

[0072] As Figure 3 shown, the processing flow of the cross-modal text enhancement module includes: input the text feature vector through the softmaxone cross-attention mechanism, the text feature vector respectively performs cross-modal information interaction with the video feature vector U v and the audio feature vector U a , then input into the fully connected layer to improve the fusion effect between modalities, obtain the non-linear displacement vector, and after passing through the residual connection block, perform iterative processing to update the text feature vector.

[0073] For the text modality, the video feature vector is U v and the audio feature vector is U a are non-text modalities as secondary modalities and respectively perform cross-modal information interaction with the text modality through the softmaxone cross-attention mechanism. In addition to being able to reduce the dimension of the concatenation operation of the previous layer, the fully connected layer can also learn higher-level interaction features in the cross-modal attention matrix to improve the fusion effect between modalities. The residual connection block passes a scaling factor α to generate the non-text displacement vector H iControl within a suitable range, and normalize the enhanced text vector after residual connection through layer normalization operation to make up for the slight impact brought by the softmaxone function in the cross-attention mechanism.

[0074] The softmaxone cross-attention mechanism uses cross-modal attention to capture long-distance non-verbal emotional information and generates a text-based non-verbal displacement vector. To reduce the influence caused by irrelevant displacement components of non-text modalities, the softmaxone function is used to replace the traditional softmax function. Although the softmax function in the traditional attention mechanism can enable the model to better focus on key elements, it will also force the attention head to focus on unimportant information. The formula for the traditional softmax function is:

[0075]

[0076] where x is a vector of dimension K, and x i is the i-th element of the vector x. When all elements in the vector x are much smaller than 0, the result of performing the softmax operation on it is still a value greater than 0, that is

[0077]

[0078] To overcome this shortcoming, the denominator of the traditional softmax function is adjusted, and the modified function is named the softmaxone function.

[0079] Using the softmaxone function adjustment can make the output value of the softmaxone function tend to 0 when the values of unimportant elements are small enough, avoiding too much unimportant information from being added to the attention score matrix. Using the softmaxone function, the cross-attention mechanism is as Figure 4 shown, and the formula is expressed as:

[0080]

[0081] where represents the output state of the Pro-Bert model, represents the feature vector of modality u, u ∈ {v, a}; subsequently, the text modality is mapped to the Q matrix through the weight matrix t matrix, and the non-text modality U u is mapped to the K and matrices and the V u matrix and the V u matrix respectively through tand T u represent the sample lengths of the text modality and the non - text modality respectively, and d k is the sample dimension constant.

[0082] In addition to being able to reduce the dimension of the concatenation operation of the previous layer, the fully - connected layer in the displacement vector fusion block can also learn higher - level interaction features in the cross - modality attention matrix, improving the fusion effect between modalities.

[0083] The two cross - modality attention score matrices and obtained through the softmaxone cross - attention mechanism need to be concatenated together through dimension connection, and then their non - text offset vectors are obtained through the fully - connected layer and the activation function.

[0084] The residual connection block in the cross - modality text enhancement module controls the generated non - text displacement vector H i within a suitable range through a scaling factor α, and the non - linear displacement vector, which is expressed by the formula:

[0085]

[0086] where represents the dimension concatenation operation on two matrices, W H and b H are the weight matrix and bias parameter of the linear transformation respectively, and σ represents the Relu activation function.

[0087] After concatenation, contains the cross - modality attention scores of the text for vision and speech respectively, which is used to map the dimension of the concatenated matrix vector to the same length as the text modality, and its non - linear representation is obtained through the Relu activation function as the non - text offset vector H i .

[0088] In the residual connection block, the text output state of the previous layer is used to i perform weighted summation with the non - text offset vector H of the current layer to obtain a new text vector representation

[0089] The enhanced text vector after residual connection is normalized through layer normalization to make up for the slight influence brought by the softmaxone function in the cross - attention mechanism, and the process of the residual connection block is expressed by the formula:

[0090]

[0091] Among them, α is a scaling factor, and its function is to scale the non-text offset vector H i to control an appropriate range and avoid the non-text offset vector from having an excessive impact on the semantic space of the text vector. In particular, when α takes the value of 0, it means ignoring the non-text displacement vector. The text vector representation after weighted summation After being normalized through the layer normalization operation LN(·), H i is the non-text displacement vector.

[0092] is input into the subsequent Transformer encoder of the After-Bert model to obtain the multimodal feature G t , and the formula is expressed as:

[0093]

[0094] where θ Bert are the parameters of the After-Bert model.

[0095] S3. Introduce a unimodal label generation module, which generates unimodal labels based on multimodal features and unimodal features, and optimizes the model parameters through the objective loss function.

[0096] Introduce the unimodal label generation module ULGM module, which performs unimodal label supervision by effectively using multimodal labels and unimodal feature vectors, thereby generating effective unimodal label annotations and converting the original single multimodal task into several different subtasks.

[0097] The ULGM module adopts a self-supervised strategy to jointly learn unimodal and multimodal tasks, learns the similarity information of modalities from multimodal tasks, and learns the difference information between modalities from unimodal tasks. Although in the real world, there will be expressions with implicit meanings such as irony and exaggeration, resulting in a large difference between unimodal and multimodal emotional expressions, due to the continuity of emotional expressions, there is still a certain connection between unimodal labels and multimodal labels within the same sample. In addition, there is also a positive correlation between label differences and the distance differences between modal features and class centers. In this subsection, a mapping relationship between known quantities and unknown quantities will be established based on the above relationships.

[0098] The multimodal task and the unimodal subtask share modal features, which is a form of multi-task with hard parameter sharing. To account for the dimensional differences between non-text modalities and text modalities, U a and U v are projected into a new feature space through a fully connected layer to obtain the unimodal features G v and G a , and the formula is expressed as:

[0099] Gu = σ(W u U u + b u )

[0100] where u ∈ {a, v}, W u and b u represent the weight parameters and biases of the fully connected layer, σ represents the Relu activation function, and G u is the newly obtained non-text modality feature.

[0101] To measure the connection between the modality features and the sentiment polarity labels, a pair of positive and negative sample centers are created for the multi-modal features and the single-modal features. In the subsequent sample iteration, the positive and negative sample centers also change dynamically. At the beginning of training, since the single-modal sample labels y u , u ∈ {a, v} are not generated, each single-modal label needs to be initialized with the multi-modal label before the first iteration. The positive and negative sample centers of the modality feature G m are expressed by the formula:

[0102]

[0103] where m ∈ {t, v, a}, and represent the positive sample center and the negative sample center of modality m respectively, N represents the total number of samples, G m,n is the modality feature, representing the feature vector of modality m for sample n, y m is the modality label, and I(y(x)) is the function indicator that includes the sample labels that meet the condition y(x) into its subset elements.

[0104] Considering the large distance differences between high-dimensional vectors, the cosine similarity is used to transform the distance problem from the modality feature to the positive and negative sample centers into the similarity problem of the modality feature and the positive and negative sample centers in terms of direction. The cosine similarity calculation only focuses on the direction of the vectors and does not consider the length of the vectors. Therefore, it can capture the similarity between vectors well without being affected by the vector length. Specifically, the cosine similarity between the modality feature G m and the positive and negative sample centers is expressed by the formula:

[0105]

[0106] where m ∈ {t, a, v}, represents the similarity between modality m and the positive sample center. represents the similarity between modality m and the negative sample center. ||·|| represents the operation of taking the L2 norm of the vector. The final result of the cosine similarity calculation is between 0 and 2, and the closer it is to 0, the more similar it is to the sample center.

[0107] To better measure the similarity between the modal features and the centers of positive and negative samples, μ is introduced. m to represent the modal feature G m respectively with the center of positive samples and the center of negative samples The relative differences are calculated by the following formula:

[0108]

[0109] where μ m represents the relative differences between the modal feature G m respectively with the center of positive samples and the center of negative samples When the unimodal feature is closer to the center of positive samples while the multimodal feature is closer to the center of negative samples, the multimodal label should increase a positive offset to make the unimodal generated label closer to the positive sentiment.

[0110] The correspondence between the unimodal label and the multimodal label can be established through the above-known quantities, and the unimodal label y u is obtained by adding an offset to the multimodal label y t Considering that there is a proportional relationship between the ratio of the unimodal label to the multimodal label and the ratio of the relative distance of the unimodal task to the relative distance of the multimodal task, the following calculation formula for the unimodal label can be obtained, and the formula is:

[0111]

[0112] where u ∈ {a, v}, y t represents the multimodal label, y u represents the unimodal label, μ u and μ t respectively represent the relative differences of the unimodal task and the multimodal task.

[0113] The unimodal label is updated dynamically. In each round of sample iteration, the update of the unimodal label is carried out through a dynamic update strategy (Momentum-based Update Policy), which is an update method that combines the newly generated label with the historically generated label, alleviating the impact caused by the instability of the generated label at lower iteration rounds. The formula of the dynamic update strategy is:

[0114]

[0115] where represents the unimodal generated label obtained after the dynamic adjustment strategy, Denote the newly generated unimodal labels in the e-th iteration; when the iteration round e = 1, the multimodal labels are adopted to initialize all unimodal sample labels. are the unimodal labels generated in the past multiple iterations and the unimodal labels generated in the current iteration which determines that this update strategy makes the later generated labels have a greater impact than the earlier generated labels, and as the iteration progresses, the generated label values gradually tend to be stable.

[0116] In the design of the multi-task objective function, the model adopts the L loss function as the objective function for both unimodal and multimodal tasks. In the balance of each sub-task, the difference between the unimodal generated labels and the multimodal labels is used as the weight of the unimodal task loss function, so that the network model pays more attention to the sample data with larger differences, and gives less attention to the sample data with smaller differences, thus generating more credible unimodal labels. Specifically, the calculation formula of the loss function is as follows:

[0117]

[0118] where N represents the total number of training samples, is the predicted value of the cross-modal fusion model for the n-th sample in the previous subsection, represents the multimodal sample labels, is the weight adjustment strategy, enabling the model to adopt a more stringent learning strategy for samples with large differences between multimodal labels and unimodal generated labels during the training process.

[0119] After completing the multimodal sentiment prediction task and the unimodal label generation task by optimizing the objective loss function, the final obtained result is the multimodal prediction result which numerically approximates the multimodal label y m indicating that the predicted value of the sentiment analysis task is close to the true value, and this prediction result can be used as the text modality label y t , and the labels y a and y v of the two non-text modalities.

[0120] The present invention also provides an electronic device, Figure 5 is the structural schematic diagram of the electronic device provided by the embodiment of the present invention, as Figure 5 shown, the electronic device may include: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete mutual communication through the communication bus. The processor can call the logical instructions in the memory, for example, to execute the following method:

[0121] S1. Obtain video feature vectors, audio feature vectors, and text feature vectors respectively;

[0122] S2. Input the video feature vectors, audio feature vectors, and text feature vectors into a cross-modal text enhancement module and an After-Bert module respectively for processing to obtain multi-modal features. At the same time, perform multiple perceptron processes on the video feature vectors and audio feature vectors respectively to obtain single-modal features;

[0123] S3. Introduce a single-modal label generation module, generate single-modal labels based on the multi-modal features and single-modal features, and optimize the model parameters through an objective loss function.

[0124] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0125] The embodiment of the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the methods provided in the above-mentioned various embodiments, for example, including:

[0126] S1. Obtain video feature vectors, audio feature vectors, and text feature vectors respectively;

[0127] S2. Input the video feature vectors, audio feature vectors, and text feature vectors into a cross-modal text enhancement module and an After-Bert module respectively for processing to obtain multi-modal features. At the same time, perform multiple perceptron processes on the video feature vectors and audio feature vectors respectively to obtain single-modal features;

[0128] S3. Introduce a single-modal label generation module, generate single-modal labels based on the multi-modal features and single-modal features, and optimize the model parameters through an objective loss function.

[0129] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0130] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-task sentiment analysis method based on cross-modal text enhancement, characterized in that Including: S1. Obtain a video feature vector, an audio feature vector, and a text feature vector respectively; S2. Input the video feature vector, the audio feature vector, and the text feature vector into a cross-modal text enhancement module and an After-Bert module respectively for processing to obtain multi-modal features. At the same time, perform multiple perceptron processes on the video feature vector and the audio feature vector respectively to obtain single-modal features; S3. Introduce a single-modal label generation module. Based on the multi-modal features and the single-modal features, generate single-modal labels, and optimize the model parameters through an objective loss function.

2. The multi-task sentiment analysis method based on cross-modal text enhancement according to claim 1, wherein, The acquisition of the video feature vector in S1 includes: Use the OpenFace tool to divide the input image into several small regions and calculate the gradient direction distribution of each region, and extract a simplified feature map of the basic facial structure; Subsequently, perform key point positioning on the detected human face, and iteratively optimize the spatial coordinates of the predicted facial feature points to achieve accurate facial geometric alignment; Finally, use a deep convolutional neural network to vectorize the features of the aligned human face. The video feature vector after vectorization is expressed as: Among them, T v is the length of the visual modality, and d v is the dimension of the visual modality; The acquisition of the audio feature vector includes: using the COVAREP tool to extract audio features, and normalizing the audio features by the zero-mean and variance normalization method to obtain the audio feature vector expressed as: Among them, T a is the length of the speech modality, and d a is the dimension of the speech modality; The acquisition of text feature vectors includes: inserting [CLS] and [SEP] into the corresponding positions of the text sequence through a tokenizer, mapping them to the corresponding text IDs, and inputting the mapped text sequence into the Pre-Bert model. The obtained text feature vectors are 3. A multi-task sentiment analysis method based on cross-modal text enhancement according to claim 1, characterized in that The processing flow of the cross-modal text enhancement module includes: Input the text feature vector. Through the softmaxone cross-attention mechanism, the text feature vector performs cross-modal information interaction with the video feature vector and the audio feature vector respectively, and then inputs it into a fully connected layer to improve the fusion effect between modalities, obtains a non-linear displacement vector, and after passing through a residual connection block, performs iterative processing to update the text feature vector.

4. A multi-task sentiment analysis method based on cross-modal text enhancement according to claim 3, characterized in that, The softmaxone cross-attention mechanism is expressed by the formula: Among them, represents the output text modal feature vector, represents the feature vector of non-text modality u, where u ∈ {v, a}; T t and T u represent the sample lengths of the text modality and non-text modality respectively, d k is the sample dimension constant.

5. A multi-task sentiment analysis method based on cross-modal text enhancement according to claim 4, characterized in that, The non-linear displacement vector is expressed by the formula: Among them, represents the operation of dimension concatenation on two matrices, W H and b H are the weight matrix and bias parameter of the linear transformation respectively, and σ represents the Relu activation function.

6. The multi-task sentiment analysis method based on cross-modal text enhancement according to claim 5, characterized in that, The processing of the residual connection block is expressed by the formula: where α is a scaling factor, and H i is a non-text displacement vector.

7. A multi-task sentiment analysis method based on cross-modal text enhancement according to claim 6, characterized in that The processing of the After-Bert module is expressed by the formula: Among them, G t is a multimodal feature.

8. A multi-task sentiment analysis method based on cross-modal text enhancement according to claim 1, characterized in that The introduced single-modal label generation module, i.e., the ULGM module, is used to learn the similarity information between modalities from multi-modal features, learn the difference information between modalities from single-modal features, and supervise the generation of single-modal task labels.

9. The multi-task sentiment analysis method based on cross-modal text enhancement according to claim 8, wherein, The generation of single-modal labels in S3 includes the steps: S31. Create positive and negative sample centers for multi-modal features, video modal features, and audio modal features, which is expressed by the formula: where m ∈ {t, v, a}, and respectively represent the positive sample center and negative sample center of modality m, N represents the total number of samples, G m,n is the modality feature, representing the feature vector of modality m for sample n, y m is the modality label, and I(y(x)) is the function indicator that incorporates the sample labels that meet the condition y(x) into the elements of its subset; S32. Calculate the modal feature G m,n The cosine similarity with the centers of positive and negative samples, and the formula is as follows: where m ∈ {t, a, v}, denotes the similarity of the modality feature G m,n to the positive sample center; denotes the similarity of the modality feature G m,n to the negative sample center; ||·|| represents the L2-norm operation on vectors; S33. The relative difference μ is introduced m , and the formula is: Among them, μ m represents the relative gap between the modal feature G m,n and the positive sample center and the negative sample center respectively; S34. Calculate the single-modal label, and the formula is: Among them, u ∈ {a, v}, y t represents a multi-modal label, y u represents a single-modal label, μ u and μ t respectively represent the relative differences between the single-modal task and the multi-modal task; S35. Dynamically update the single-modal label, and the dynamic update strategy formula is: Among them, represents the single-modal generation label obtained after the dynamic update strategy, represents the newly generated single-modal label in the e-th iteration.

10. A multi-task sentiment analysis method based on cross-modal text enhancement according to claim 9, characterized in that, The optimization through the objective loss function is expressed by the formula: Among them, N represents the total number of training samples, is the predicted value of the cross-modal fusion model for the nth sample in the previous subsection, represents the multi-modal sample label, is the weight adjustment strategy.