Multimodal Sentiment Recognition Method for Dynamically Adjusting Word Representations Based on Misaligned Behavior Information

Through the multimodal emotion recognition method that dynamically adjusts word representations with unaligned behavior information, the cross-modal attention mechanism is used for long-term fusion, which solves the problem of large parameters and high alignment operation cost in the processing of multimodal data, and achieves more efficient multimodal emotion recognition performance.

CN114936552BActive Publication Date: 2025-06-13HANGZHOU DIANZI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210624963.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-02
Publication Date
2025-06-13
Estimated Expiration
2042-06-02

AI Technical Summary

Technical Problem

When processing multimodal data, existing multimodal emotion recognition models need to perform multiple dual-modal fusion operations, resulting in the model retaining a large number of original parameters and affecting performance. In addition, existing models often rely on manually aligned multimodal sequence data, which consumes a lot of manpower and material costs.

Method used

A multimodal emotion recognition method that dynamically adjusts word representations with unaligned behavior information is adopted, and a long-term fusion is performed within the entire discourse range scale through a cross-modal attention mechanism, dynamically transfers the position of words in the semantic space, and uses unaligned visual and auditory modal information to integrate into the text modality.

Benefits of technology

It effectively reduces the number of parameters of the model, improves performance, and reduces the cost of alignment operations, and can better mine long-term interaction information in multimodal data and improves the accuracy of emotion recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114936552B_ABST
    Figure CN114936552B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimodal sentiment recognition method for dynamically adjusting word representations based on misaligned behavior information. The present invention utilizes a cross-modal attention mechanism to mine behavior information related to the text modality (composed of visual and auditory modalities), and then uses the behavior information to dynamically modify the positions of words in the text modality in the semantic space, thereby obtaining word representations adjusted by multimodal information. At the same time, the cross-modal attention mechanism can focus on behavior information related to the text modality within a long distance range, so it can well solve the inherent problem in multimodal learning - the frequency mismatch between various modality information. Secondly, based on this, several multimodal Transformer layers are constructed, which can further mine the high-level feature information of the word representations adjusted by multimodal information in the context environment, and it is an effective supplement to the multimodal fusion framework in the current sentiment recognition field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of multi-modal emotion recognition in the cross-field of natural language processing, vision, and speech, and specifically relates to a multi-modal emotion recognition method that uses unaligned behavior information to dynamically adjust word representations. By using a fusion network technology based on a cross-modal attention mechanism, it performs long-term fusion on behavior information and text modal information composed of vision and audition within the entire discourse range scale, and dynamically transfers the position of words in the semantic space to judge the emotional state of the subject. Background Art

[0002] The field of emotion analysis usually includes data information such as text modality, video modality, and speech modality. In previous studies, it has been verified that these single-modal data contain discriminant information related to emotional states. At the same time, it has been found that the consistency and complementarity existing between these single-modal data can effectively explain the internal correlated representations of multi-modal data, and can further enhance the model's expressive ability and stability, and improve the performance of emotion task analysis.

[0003] Existing multi-modal fusion models based on adjusting word representations have attracted wide attention because they can effectively model fine-grained multi-modal information data, thus to a certain extent reducing the impact caused by using an averaging strategy that ignores the complex interaction information within local modalities. The specific operation is as follows: during the process of multi-modal fusion, first, the two modalities between vision and text are fused respectively, and the two modalities between audition and text are fused. Then, the two fused pieces of information are further fused to obtain the fused information containing all modalities. However, when the number of multi-modalities exceeds two, multiple bimodal fusion operations need to be performed to obtain the fused information containing all modalities. This two-way fusion strategy will cause the model to retain a large number of original parameters, greatly affecting the performance of the model. In addition, existing networks for adjusting word representations usually use manually aligned multi-modal sequence data to dynamically adjust the representation of words in the semantic space. Since the sampling rates of each modality are different, the collected multi-modal sequence data is usually unaligned. To adjust the word representation in the aligned behavior information, first, the behavior information needs to be aligned with the text modality to make the three modality information consistent in the time dimension. However, in deep learning tasks, this annotation operation requires a large amount of human and material costs. Therefore, using unaligned behavior information to dynamically adjust word representations compared to aligned behavior information is a method with practical significance. Summary of the Invention

[0004] The purpose of the present invention is to propose a multi-modal emotion classification method that dynamically adjusts word representations using unaligned behavior information in view of the deficiencies of the prior art.

[0005] In a first aspect, the present invention provides a multimodal sentiment recognition method for dynamically adjusting word representations based on unaligned behavior information, which includes the following steps:

[0006] Step 1: Data collection.

[0007] Obtain a multimodal dataset collected under different sentiment categories.

[0008] Step 2: Multimodal information data preprocessing.

[0009] Convert the text modality, visual modality, and auditory modality data into primary representations respectively, and perform a pre-fusion operation on the auditory and visual modality data to reduce the time-domain dimension size and the feature vector length of the auditory and visual modality data.

[0010] Step 3: Cross-supermodal fusion.

[0011] 3-1. Obtain supermodal information

[0012] Concatenate the primary representations of the visual and auditory modalities after the pre-fusion operation in the time-domain dimension to obtain supermodal information X β .

[0013] 3-2. Dynamically adjust word representations.

[0014] Pass the supermodal information through two linear transformation networks respectively to obtain the key matrix K β and the real-valued matrix V β ; pass the text modality information through a linear transformation network to obtain the corresponding query matrix Q l .

[0015] Based on the query matrix Q l and the key matrix K β calculate the attention factor matrix e of the behavior information in the text modality as follows:

[0016]

[0017] e = softmax(a) Formula (6)

[0018] where a is the unnormalized attention factor matrix; d k is the feature length of the query matrix Q l .

[0019] Extract the text-related information H in the supermodal information as follows:

[0020] H = eV β Formula (7)

[0021] Obtain the text information incorporating unaligned behavior information;

[0022] Dynamically adjust each word representation in the text modality by using the text-related information H in the above-obtained hyper-modal information as follows:

[0023]

[0024] Among them, represents the text information incorporated with hyper-modal information. X l represents the initial representation of the text modality; α is a proportionality coefficient; λ is a preset hyperparameter.

[0025] Use the text information to input into the sentiment recognition model for training.

[0026] Step Four, Sentiment Recognition Output

[0027] Collect the multi-modal data of the measured object and send it into the sentiment recognition model obtained in Step Three to identify the sentiment category of the measured object.

[0028] Preferably, the sentiment categories include positive emotions and negative emotions.

[0029] Preferably, in Step 2, the text information is converted into a primary representation in the form of word embeddings through text encoding by a pre-trained language model.

[0030] Preferably, in Step 2, use a long short-term memory network to extract the primary features of visual and auditory data as follows:

[0031]

[0032] Among them, F m is the primary feature of visual or auditory data, is the primary representation of modality m; v and a respectively represent the visual and auditory modalities; I m is the original data of modality m; is the weight matrix of modality m; T m is the size of the time domain dimension; d m is the length of the feature vector at each moment.

[0033] Preferably, in Step 2, the result X {m} of the pre-fusion of auditory or visual modality data is expressed as follows:

[0034]

[0035] Among them, {m} is the primary representation of modality m; T m is the size of the time domain dimension, d m is the length of the feature vector at each moment; k {m} is the size of the convolution kernel of modality m.

[0036] Preferably, the key matrix K β and the real-valued matrix V β are expressed as follows:

[0037]

[0038]

[0039]

[0040]

[0041] wherein are the weight matrices of the linear networks of the matrices K β , V β respectively; d β , d k , d v are the eigenvector lengths of the hyper-modal information, the key matrix, and the real-valued matrix respectively.

[0042] Preferably, the query matrix Q l is expressed as follows:

[0043]

[0044]

[0045] wherein X l is the text-modal information, is the weight matrix of the query matrix; d l and d k are the eigenvector lengths of the text modality and the query matrix respectively.

[0046] Preferably, the sentiment recognition model adopts the BERT model (Bidirectional Encoder Representation from Transformers).

[0047] In a second aspect, the present invention provides a sentiment recognition system, which includes a processor and a memory. The memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the foregoing multi-modal sentiment recognition method. The machine-executable instructions include a data acquisition module, a data preprocessing module, a cross-hyper-modal fusion, and a sentiment recognition output module.

[0048] In a third aspect, the present invention provides a machine-readable storage medium; the machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are called and executed by a processor, the machine-executable instructions prompt the processor to implement the aforementioned multimodal emotion recognition method.

[0049] The beneficial effects of the present invention are:

[0050] The present invention combines the cross-modal attention mechanism and uses the unaligned behavioral information to dynamically adjust the word representation in the text modality, and mines the modal fusion information of the long-term interaction between the non-text modality and the text modality. In addition, the cross-modal attention mechanism can model multiple modal information at the same time, so it can deal well with the inherent problem in multimodal learning-multiple modalities cannot interact at the same time. Next, a multimodal Transformer framework was constructed on this basis, and the word representation dynamically adjusted by the behavioral information was sent into it to further perform high-level multimodal fusion, which is an effective supplement to the current multimodal fusion framework in the field of emotion recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a flow chart of the present invention;

[0052] Figure 2 A schematic diagram of dynamically adjusting a word network in the present invention;

[0053] Figure 3 Schematic diagram of three-modal fusion. DETAILED DESCRIPTION

[0054] The method of the present invention is described in detail below in conjunction with the accompanying drawings.

[0055] like Figure 1 As shown, a multimodal emotion recognition method for dynamically adjusting word representation based on unaligned behavior information includes the following steps:

[0056] Step 1: Obtain multimodal information data

[0057] When the subjects perform specific emotional tasks, the text modality data, voice modality data, and video modality data of the subjects are recorded as a multimodal dataset. The specific emotional tasks include positive emotions and negative emotions, which can be specifically subdivided into very negative, negative, weakly negative, neutral, weakly positive, positive, and very positive.

[0058] Step 2: Multimodal information data preprocessing

[0059] Multimodal data performs multimodal fusion operations at the feature level; for the text modality, a pre-trained language model is adopted, and the original text information is transformed into a primary representation in the form of word embeddings through a text encoder.

[0060] For the auditory and visual modalities, a long short-term memory network is used to extract the primary feature representations of visual and auditory data;

[0061]

[0062] where, F m is the primary feature of visual or auditory data, is the primary representation of modality m; v and a represent the visual and auditory modalities respectively; I m is the original data of modality m; is the weight matrix of modality m; T m is the size of the time domain dimension; d m is the length of the feature vector at each moment; due to the different standards of modality sampling rates, the time domain dimension sizes of non-text modalities (visual and auditory modalities) are usually much larger than those of the text modality, which is not conducive to multimodal fusion operations. Therefore, pre-fusion operations are performed on the auditory and visual modalities to reduce their time domain dimension sizes and the lengths of the feature vectors;

[0063]

[0064] where, is the result of pre-fusion of modality m; T m is the size of the time domain dimension, d m is the length of the feature vector at each moment; k {m} is the size of the convolution kernel of modality m. Conv 2D(·) is two-dimensional convolution processing.

[0065] Step 3: Based on the cross-supermodal fusion method, utilize the unaligned visual and auditory modality information to dynamically adjust the representation of the text modality in the semantic space. This method includes two tasks: obtaining supermodal information and dynamically adjusting word representations;

[0066] 3-1. Obtaining supermodal information

[0067] In the learning process of obtaining supermodal information, the primary representations of the unaligned visual and auditory modalities after pre-fusion operations are concatenated together in the time domain dimension to obtain supermodal information. This supermodal information contains all the information that affects the text representation. The expression of the supermodal information containing visual and auditory modalities is as follows:

[0068]

[0069] Among them, X β represents the obtained cross-modal information, v represents visual modal information, a represents auditory modal information, represents the concatenation operation.

[0070] 3-2. Dynamically adjust word representations. In the learning process of dynamically adjusting word representations, for each word representation in the text modality, utilize the aforementioned obtained cross-modal information to dynamically adjust each word representation in the text modality within the entire discourse scale, and integrate the cross-modal information composed of visual and auditory modalities into the text representation, thereby completing multi-modal fusion. The specific process is as follows:

[0071] Pass the cross-modal information through two linear transformation networks respectively to obtain the corresponding key matrix K β and real-valued matrix V β , which are expressed as follows:

[0072]

[0073]

[0074]

[0075]

[0076] Among them, are the weight matrices of the linear networks of matrices K β , V β respectively; d β , d k , d v are the characteristic vector lengths of the cross-modal information, key matrix, and real-valued matrix respectively.

[0077] Pass the text modality information through a linear transformation network to obtain the corresponding query matrix Q l , which is expressed as follows:

[0078]

[0079]

[0080] Among them, X l is the text modality information, is the weight matrix of the query matrix Q l ; d l and d k are the characteristic vector lengths of the text modality and the query matrix respectively.

[0081] Using the cross-modal attention mechanism, the hyper-modal information is integrated into the text modality, and the behavioral information is used to dynamically adjust the representation of words in the semantic space, as follows:

[0082] For the cross-modal attention mechanism, based on the query matrix Q l and the key matrix K β The attention factor matrix e of the behavioral information in the text modality is calculated as follows:

[0083]

[0084]

[0085] where a is the unnormalized attention factor matrix; d k is the feature length of the query matrix Q l of.

[0086] According to the interaction between the attention factor matrix and the real-valued matrix, the long-term correlation in the time domain between the hyper-modal information and the text information is obtained;

[0087]

[0088] where H represents the information related to the text in the hyper-modal information.

[0089] Using the information H related to the text in the hyper-modal information obtained above to dynamically adjust each word representation in the text modality, which is expressed as follows:

[0090]

[0091] where X l represents the unadjusted text modality information, represents the text information integrated with the unaligned behavioral information. α is the proportionality coefficient; λ is the preset hyperparameter; ||·|| 2 is the two-norm operation.

[0092] The text information integrated with the unaligned behavioral information adds video and audio modality information, greatly supplementing the limitations of the expression ability of single text modality information. We add a special marker (CLS) in front of each text modality to be used as the label for multi-modal sentiment classification. After the above operations, the original text modality information will obtain a new text modality representation vector, and the converging multi-modal information will be sent into the Transformers layer of BERT for further training to obtain a sentiment recognition model for downstream sentiment classification tasks. The loss function for training is

[0093] Step 4: Extract the text modality, visual modality, and auditory modality information of the object under test simultaneously, and input them into the emotion recognition model to obtain the emotion category of the object under test.

[0094] Figure 2 It is a flowchart for dynamically adjusting word representations using unaligned multimodal information. Figure 3 It is a multimodal fusion flowchart for three modalities A, V, and T.

[0095] The proposed method and multiple existing multimodal fusion methods are used to perform emotion state discrimination tasks on two publicly available multimodal emotion databases CMU-MOSI and CMU-MOSEI simultaneously. Each dataset has data in two formats: aligned and unaligned. The results are shown in Tables 1 and 2. The results in the tables are the mean absolute error (MAE), correlation coefficient (Corr), accuracy for binary emotion classification task (Acc-2), F1-score (F1-Score), and accuracy for seven-class emotion classification task (Acc-7). It can be seen that compared with the existing multimodal fusion frameworks that show excellent performance, the five evaluation indicators of the proposed method are better than those of the existing fusion models, which proves the effectiveness of the proposed method.

[0096] Table 1. Comparison Table of Results

[0097]

[0098] Table 2. Comparison Table of Results

[0099]

Claims

1. A multimodal sentiment recognition method for dynamically adjusting word representations based on misaligned behavior information, characterized in that: It includes the following steps: Step 1, data collection; Obtain a multimodal dataset collected under different sentiment categories; Step 2, multimodal information data preprocessing; Convert text modality, visual modality, and auditory modality data into primary representations respectively, and perform pre-fusion operations on auditory and visual modality data to reduce the time-domain dimension size and feature vector length of auditory and visual modality data; Use a long short-term memory network to extract the primary features of visual and auditory data as follows: Among them, F m is the primary feature of visual or auditory data, is the primary representation of modality m; v and a respectively represent the visual and auditory modalities; I m is the original data of modality m; is the weight matrix of modality m; T m is the size of the time domain dimension; d m is the length of the feature vector at each moment; The result X of pre-fusion of auditory or visual modality data {m} The expression thereof is as follows: Among them, {m} is the primary representation of mode m; T m is the size in the time domain dimension, d m is the length of the feature vector at each moment; k {m} is the size of the convolutional kernel of mode m; Step 3, cross-supermodal fusion; 3-1. Obtain supermodal information The primary representations of the visual and auditory modalities that have undergone the pre-fusion operation are concatenated together in the time domain dimension to obtain the cross-modal information X β ; 3-2. Dynamically adjust word representations; The hyper-modal information is respectively passed through two linear transformation networks to obtain the key matrix K β and the real-valued matrix V β ; The text-modal information is passed through a linear transformation network to obtain the corresponding query matrix Q l ; Based on the query matrix Q l and the key matrix K β The attention factor matrix e of the behavior information in the text modality is calculated as follows: e = softmax(a) Formula (6) where a is an unnormalized attention factor matrix; d k is the query matrix Q l 's characteristic length; Extract the text-related information H in the supermodal information as follows: H = eV β Equation (7) Obtain text information incorporating misaligned behavior information; Dynamically adjust the representation of each word in the text modality using the text-related information H in the supermodal information obtained above as follows: Among them, represents the text information integrated with hyper-modal information; X l represents the initial representation of the text modality; α is the proportionality coefficient; λ is the preset hyperparameter; Take the text information Input it into the emotion recognition model for training; Step Four, sentiment recognition output Collect the multimodal data of the object to be measured and send it to the sentiment recognition model obtained in Step Three to identify the sentiment category of the object to be measured.

2. The multimodal sentiment recognition method for dynamically adjusting word representations based on misaligned behavior information according to claim 1, characterized in that: The sentiment categories include positive emotions and negative emotions.

3. The multimodal sentiment recognition method for dynamically adjusting word representations based on misaligned behavior information according to claim 1, characterized in that: In Step 2, the text information is converted into a primary representation in the form of word embeddings through text encoding by a pre-trained language model.

4. The multimodal sentiment recognition method for dynamically adjusting word representations based on misaligned behavior information according to claim 1, characterized in that: Key matrix K β and real-valued matrix V β are expressed as follows: Among them, are the weight matrices of the linear networks of matrix K β , V β respectively; d β , d k , d v are the eigenvector lengths of the hypermodal information, the key matrix, and the real-valued matrix respectively.

5. The multimodal sentiment recognition method for dynamically adjusting word representations based on misaligned behavior information according to claim 1, characterized in that: Query matrix Q l The expression is as follows: Among them, X l is the text modality information, is the weight matrix of the query matrix; d l and d k are the eigenvector lengths of the text modality and the query matrix, respectively.

6. The multimodal sentiment recognition method for dynamically adjusting word representations based on misaligned behavior information according to claim 1, characterized in that: The sentiment recognition model uses a BERT model.

7. A sentiment recognition system, including a processor and a memory; characterized in that: The memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the multimodal sentiment recognition method described in any one of claims 1-6; the machine-executable instructions include a data collection module, a data preprocessing module, cross-supermodal fusion, and a sentiment recognition output module.

8. A machine-readable storage medium; characterized in that: This machine-readable storage medium stores machine-executable instructions, and when the machine-executable instructions are called and executed by the processor, the machine-executable instructions prompt the processor to implement the multimodal sentiment recognition method described in any one of claims 1-6.