A Multimodal Sentiment Analysis and Interaction Adaptation Method and System Based on Large Models

Through a multimodal emotion analysis method based on large models, combining video frames, audio and text data to perform cross-attention fusion and emotion classification, the problem of insufficient user emotion recognition and adjustment in the intelligent dialogue system is solved, and the user experience is improved.

CN119475099BActive Publication Date: 2025-07-08SHANDONG INSPUR SCI RES INST CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510025730.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-07-08
Estimated Expiration
2045-01-08

AI Technical Summary

Technical Problem

The existing intelligent dialogue systems lack the ability to accurately identify and dynamically adjust users' multimodal information, especially when it involves multimodal data processing, it is prone to information loss or misjudgment, and cannot effectively respond to fluctuations in user emotions, resulting in poor user experience.

Method used

A multimodal sentiment analysis method based on large models is adopted, and the video frames, audio and text data are preprocessed, image, speech and text features are extracted, and cross-attention mechanism is used to modally fusion to generate multimodal feature representations. Combined with a multi-layer perceptron for emotion classification, and dialogue strategies are dynamically adjusted to adapt to user emotional changes.

Benefits of technology

It realizes accurate identification and dynamic adjustment of user emotions, and improves the naturalness and user experience of the intelligent dialogue system in complex human-computer interaction scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119475099B_ABST
    Figure CN119475099B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-modal sentiment analysis and interaction adaptation method and system based on large models, which relates to the technical fields of artificial intelligence and natural language processing; it includes: Step 1: Collect dialogue video data and preprocess the video data. Step 2: Extract image features, speech features, and text features respectively according to video frames, audio, and text. Step 3: Interact and fuse every two types of modal features through a cross-attention mechanism. After the interactively fused features are processed by average pooling, they are input into a multi-layer perceptron for sentiment classification, and the category corresponding to the highest probability is obtained as the sentiment label recognition result. Step 4: According to the sentiment label in the recognition result, specify the requirements for sentiment adaptation, generate corresponding dialogue strategies, conduct interactions according to the dialogue strategies, and dynamically adjust the dialogue strategies according to the changes in the sentiment label; the present invention is applicable to enhancing the sentiment understanding and response capabilities of intelligent dialogue systems and improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention discloses a multi-modal sentiment analysis and interaction adjustment method and system based on a large model, which relates to the technical fields of artificial intelligence and natural language processing. Background Art

[0002] With the development of artificial intelligence technology, intelligent dialogue systems have been widely used in fields such as customer service, medical consultation, and education. However, existing dialogue systems mainly rely on single-text input for semantic understanding and generating responses, lacking the ability to accurately identify the user's emotional state and dynamically adjust. This limitation restricts the user experience of intelligent dialogue systems, especially in complex interactive application scenarios that require understanding multi-modal information of users, such as language, expression, voice, etc., and conducting sentiment analysis and adjustment, such as intelligent healthcare, psychological counseling, etc., where it is particularly insufficient.

[0003] Existing sentiment analysis methods usually use single-modal data, which easily leads to information loss or misjudgment when processing multi-modal data. In addition, existing sentiment analysis technologies often lack an adjustment process and are directly applied to intelligent dialogue systems, unable to dynamically adjust according to the user's real-time emotional state. Especially in scenarios involving long-term interactions or emotional sensitivity, they cannot effectively handle the fluctuations of the user's emotions, resulting in a poor user experience. Summary of the Invention

[0004] Aiming at the problems of the existing technology, the present invention provides a multi-modal sentiment analysis and interaction adjustment method and system based on a large model, which is suitable for enhancing the emotional understanding and response ability of intelligent dialogue systems.

[0005] The specific solution proposed by the present invention is as follows:

[0006] The present invention provides a multi-modal sentiment analysis and interaction adjustment method based on a large model, including:

[0007] Step 1: Collect dialogue video data and preprocess the video data: separately extract and process video frames, audio, and text.

[0008] Step 2: Respectively perform image feature extraction, speech feature extraction, and text feature extraction according to video frames, audio, and text.

[0009] Step 3: For each modality X, where X represents any one of text feature T, speech feature A, and image feature I, generate a corresponding query vector , key vector , value vector , which is represented by the following formula:

[0010] ;

[0011] ;

[0012] ;

[0013] is a learnable weight matrix, is a bias term, and fusion between any two modalities is performed based on the cross-attention mechanism, where for the text feature T corresponding to ( , , ) and the speech feature A corresponding to ( , , ) cross-attention calculation between the two modalities is performed to obtain , and for the text feature T corresponding to ( , , ) and the image feature I corresponding to ( , , ) cross-attention calculation between the two modalities is performed to obtain ; for the speech feature A corresponding to ( , , ) and the text feature T corresponding to ( , , ) cross-attention calculation between the two modalities is performed to obtain , and for the speech feature A corresponding to ( , , ) and the image feature I corresponding to ( , , ) cross-attention calculation between the two modalities is performed to obtain ; for the image feature I corresponding to ( , , ) and the text feature T corresponding to ( , , ) cross-attention calculation between the two modalities is performed to obtain , and for the image feature I corresponding to ( , , ) and the speech feature A corresponding to ( , , ) cross-attention calculation between the two modalities is performed to obtain ,

[0014] The text feature Concatenate the cross-attention results with other modalities to obtain ,

[0015] Concatenate the cross-attention results of the speech feature A with other modalities to obtain ,

[0016] Concatenate the cross-attention results of the image feature I with other modalities to obtain ,

[0017] For , and Apply average pooling to calculate the fused multi-modal feature representation, and input the multi-modal feature representation into a multi-layer perceptron MLP:

[0018] ;

[0019] is the probability distribution of the output sentiment categories, and the category corresponding to the highest probability is obtained as the sentiment label recognition result;

[0020] Step 4: According to the sentiment label in the recognition result, specify the requirements for sentiment adjustment, generate corresponding dialogue strategies, interact according to the dialogue strategies, and dynamically adjust the dialogue strategies according to the change of the sentiment label.

[0021] Furthermore, in step 1 of the method for multi-modal sentiment analysis and interaction adjustment based on a large model, preprocess the video data:

[0022] Extract and process video frames: Extract a sequence of image frames from the video at a preset frame rate,

[0023] Process the extracted sequence of image frames, adjust the resolution to 224×224, and perform operations of denoising and enhancing the contrast on the sequence of image frames,

[0024] Extract and process audio: Use a multimedia processing tool to separate the audio track from the video, separate the audio data according to the audio track, and perform preprocessing operations of noise reduction and speech enhancement on the audio data,

[0025] Extract and process text: Convert the speech data in the audio data into text content, and perform operations of removing noise words and duplicate words on the text content.

[0026] Further, in step 2 of the method for multi-modal sentiment analysis and interaction adjustment based on a large model, image feature extraction is performed on the image frame sequence extracted from the video frames: the preprocessed image frame sequence is input into the CLIP model, the visual encoder of the CLIP model is used to extract image visual features, and average pooling is applied according to the time series to generate global image visual features.

[0027] Further, in step 2 of the method for multi-modal sentiment analysis and interaction adjustment based on a large model, speech feature extraction is performed based on the audio: first, the audio data is normalized, and the audio is sampled at a sampling rate of 16 kHz. The sampled data is input into the HuBERT model, the average value of the features at all time steps is calculated, and a time-invariant global speech feature is generated according to the average value of the features. The global speech feature includes the content information of the speech and the emotional features of the speech, and the emotional features of the speech include intonation, volume, and speech rate.

[0028] Further, in step 2 of the method for multi-modal sentiment analysis and interaction adjustment based on a large model, text feature extraction is performed based on the text: the preprocessed text is input into the Baichuan-7B model, the Baichuan-7B model is used to segment the text, and the sentences are tokenized to obtain the global feature vectors of each segment of the text. Then, average pooling is applied to obtain the text features of the entire text.

[0029] Further, in step 4 of the method for multi-modal sentiment analysis and interaction adjustment based on a large model, the emotion labels of the recognition results involve happiness, sadness, anger, neutrality, fear, disgust, and surprise. According to the emotion labels, the requirements for emotion adjustment are specified, corresponding dialogue strategies are generated, the dialogue strategies are embedded into the instruction prompt of the large model, and interaction is performed according to the instruction prompt. At the same time, the dialogue strategies are dynamically adjusted according to the changes in the emotion labels.

[0030] Further, in step 4 of the method for multi-modal sentiment analysis and interaction adjustment based on a large model, according to the changes in the emotion labels, the emotion adjustment requirements specified by the emotion labels are dynamically adjusted, corresponding dialogue strategies are dynamically generated, the dialogue strategies are embedded into the instruction prompt of the large model, and interaction is performed according to the instruction prompt.

[0031] The present invention also provides a multi-modal sentiment analysis and interaction adjustment system based on a large model, including a collection and processing module, a feature extraction module, an emotion recognition module, and an interaction adjustment module.

[0032] The collection and processing module collects dialogue video data and preprocesses the video data: extracts and processes video frames, audio, and text respectively.

[0033] The feature extraction module extracts image features, speech features and text features based on video frames, audio and text respectively.

[0034] The emotion recognition module generates a corresponding query vector for each modality X, where X represents any one of the text feature T, speech feature A, and image feature I. , the key vector , value vector , expressed by the following formula:

[0035] ;

[0036] ;

[0037] ;

[0038] is a learnable weight matrix, is a bias term, which is used to fuse any two modalities based on the cross-attention mechanism, where the text feature T corresponds to ( , , ) and the speech feature A corresponding to ( , , ) Perform cross-attention calculation between the two modalities to obtain , the text feature T corresponds to ( , , ) and image feature I corresponding to ( , , ) Perform cross-attention calculation between the two modalities to obtain ; The speech feature A corresponds to ( , , ) and the text feature T corresponding to ( , , ) Perform cross-attention calculation between the two modalities to obtain , the speech feature A corresponds to ( , , ) and image feature I corresponding to ( , , ) Perform cross-attention calculation between the two modalities to obtain ; Image feature I corresponds to ( , , ) and the text feature T corresponding to ( , , ), perform cross-attention calculation between the two modalities to obtain , the one corresponding to the image feature I ( , , ), and the one corresponding to the speech feature A ( , , ), perform cross-attention calculation between the two modalities to obtain ,

[0039] Concatenate the text feature with the cross-attention results of other modalities to obtain ,

[0040] Concatenate the speech feature A with the cross-attention results of other modalities to obtain ,

[0041] Concatenate the image feature I with the cross-attention results of other modalities to obtain ,

[0042] For , and , apply average pooling calculation to obtain the fused multi-modal feature representation, and input the multi-modal feature representation into the multi-layer perceptron MLP:

[0043] ;

[0044] is the output probability distribution of the sentiment category, and obtain the category corresponding to the highest probability as the sentiment label recognition result;

[0045] The interaction adaptation module specifies the requirements for sentiment adaptation according to the sentiment label in the recognition result, generates the corresponding dialogue strategy, conducts interactions according to the dialogue strategy, and dynamically adjusts the dialogue strategy according to the change of the sentiment label.

[0046] The advantages of the present invention are:

[0047] The present invention effectively combines multi-modal data, such as text, speech, images, etc. with a large language model, can construct an intelligent dialogue system with sentiment analysis and adaptation capabilities, effectively fuses data of different modalities, and conducts dynamic sentiment adaptation, so that the intelligent dialogue interaction process can adjust the interaction strategy in a timely manner according to the change of the user's sentiment, solve the problems of insufficient sentiment understanding and interaction adaptation, and thus improve the naturalness and user experience in complex human-computer interaction scenarios during the intelligent dialogue process. Description of the Drawings

[0048] Figure 1 It is a schematic diagram of the method flow of the present invention. Specific embodiments

[0049] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, so that those skilled in the art can better understand the present invention and be able to implement it, but the specific embodiments cited are not intended to limit the present invention.

[0050] Embodiment 1: The present invention provides a multi-modal sentiment analysis and interaction adjustment method based on a large model, including:

[0051] Step 1: Collect dialogue video data and preprocess the video data: extract and process video frames, audio, and text respectively.

[0052] The video data usually contains a video frame sequence and an audio track, which can capture the user's facial expressions and speech information. In Step 1, the video data is preprocessed as follows:

[0053] Extract and process video frames: Extract an image frame sequence from the video at a preset frame rate,

[0054] Process the extracted image frame sequence, adjust the resolution to 224×224, and perform operations such as denoising and enhancing contrast on the image frame sequence for subsequent feature extraction.

[0055] Extract and process audio: Use the multimedia processing tool FFmpeg to separate the audio track from the video, separate the audio data according to the audio track, and perform preprocessing operations such as noise reduction and speech enhancement on the audio data to improve the clarity and intelligibility of the speech.

[0056] Extract and process text: Convert the speech data in the audio data into text content, and perform operations such as removing noise words and duplicate words on the text content.

[0057] Perform time synchronization processing on the text, speech, and image sequences extracted from the video, and align these data to ensure that they reflect the user's overall emotional state at the same time point.

[0058] Step 2: Perform image feature extraction, speech feature extraction, and text feature extraction according to the video frames, audio, and text respectively.

[0059] Among them, step 21: Extract image features based on the image frame sequence extracted from the video frames: Input the preprocessed image frame sequence into the CLIP model, use the visual encoder of the CLIP model to extract image visual features, and apply average pooling according to the time series to generate global image visual features. CLIP (Contrastive Language-Image Pretraining) is a joint vision and language model. The CLIP model can effectively map images and texts to the same feature space, thereby realizing the semantic understanding of images.

[0060] Step 22: Extract speech features according to the audio: First, normalize the audio data, sample the audio at a sampling rate of 16 kHz, input the sample into the HuBERT model, calculate the average value of the features at all time steps, and generate time-invariant global speech features according to the average value of the features. The global speech features include the content information of the speech and the emotional features of the speech. The emotional features of the speech include intonation, volume, and speech rate. HuBERT (Hidden-Unit BERT) is a speech model based on the Transformer and BERT architectures. It generates hidden unit representations through self-supervised learning of speech signals, thereby extracting high-quality speech features.

[0061] Step 23: Extract text features according to the text: Input the preprocessed text into the Baichuan-7B model, use the Baichuan-7B model to segment the text, and perform word segmentation on the sentences to obtain the global feature vectors of each segment of text, and then apply average pooling to obtain the text features of the entire text. The Baichuan-7B model is a powerful pre-trained language model. Through training on a large-scale corpus, it can effectively capture the semantic and emotional information in the text, and has strong context understanding ability and emotional recognition ability.

[0062] Step 3: For each modality X, where X represents any one of the text feature T, speech feature A, and image feature I, generate the corresponding query vector , key vector , value vector , which is represented by the following formula:

[0063] ;

[0064] ;

[0065] ;

[0066] is a learnable weight matrix, is the bias term, which is used for fusion between any two modalities based on the cross-attention mechanism. Among them, for the text feature T corresponding to ( , , ), and the speech feature A corresponding to ( , , ), cross-attention calculation between the two modalities is performed. The formula is as follows, where d represents the vector dimension and the superscript T represents the transpose operation:

[0067] ,

[0068] to obtain ,

[0069] For the text feature T corresponding to ( , , ), and the image feature I corresponding to ( , , ), cross-attention calculation between the two modalities is performed. The formula is as follows, where d represents the vector dimension and the superscript T represents the transpose operation:

[0070] ,

[0071] to obtain ;

[0072] Similarly, for the speech feature A corresponding to ( , , ), and the text feature T corresponding to ( , , ), cross-attention calculation between the two modalities is performed. The formula is as follows, where d represents the vector dimension and the superscript T represents the transpose operation:

[0073] ,

[0074] to obtain ,

[0075] For the speech feature A corresponding to ( , , ), and the image feature I corresponding to ( , , ), cross-attention calculation between the two modalities is performed. The formula is as follows, where d represents the vector dimension and the superscript T represents the transpose operation:

[0076] ,

[0077] to obtain ;

[0078] The ( , , ) corresponding to the image feature I and the ( , , ) corresponding to the text feature T perform cross - attention calculation between the two modalities. The formula is as follows, where d represents the vector dimension and the superscript T represents the transpose operation:

[0079] ,

[0080] Obtain ,

[0081] The ( , , ) corresponding to the image feature I and the ( , , ) corresponding to the speech feature A perform cross - attention calculation between the two modalities. The formula is as follows, where d represents the vector dimension and the superscript T represents the transpose operation:

[0082] ,

[0083] Obtain ,

[0084] Concatenate the cross - attention results of the text feature with those of other modalities to obtain ,

[0085] Concatenate the cross - attention results of the speech feature A with those of other modalities to obtain ,

[0086] Concatenate the cross - attention results of the image feature I with those of other modalities to obtain ,

[0087] For , and Apply average pooling calculation to obtain the fused multi - modal feature representation, and input the multi - modal feature representation into the multi - layer perceptron MLP:

[0088] ;

[0089] For the probability distribution of the output sentiment categories, the category corresponding to the highest probability is obtained as the sentiment label recognition result. In the above process, every two modal features interact and fuse through the cross-attention mechanism. After the interaction and fusion, the features are processed by average pooling and then input into a multi-layer perceptron for sentiment classification. The category corresponding to the highest probability is obtained as the sentiment label recognition result;

[0090] Step 4: According to the sentiment label in the recognition result, specify the requirements for sentiment adjustment, generate corresponding dialogue strategies, interact according to the dialogue strategies, and dynamically adjust the dialogue strategies according to the change of the sentiment label.

[0091] Among them, the sentiment labels of the recognition results can involve happy, sad, angry, neutral, fear, disgust, surprise. According to the sentiment labels, specify the requirements for sentiment adjustment, generate corresponding dialogue strategies, embed the dialogue strategies into the instruction prompt of the large model, and interact according to the instruction prompt. For example, if the current sentiment label of the user is recognized as "sad", for the "sad" sentiment, content for comfort, understanding or empathy needs to be generated. Therefore, the requirements for sentiment adjustment are specified as generating instruction prompts such as "comfort the user", "express empathy", etc., and generating dialogue strategies related to the sentiment, and embedding the dialogue strategies into the instruction prompt. That is, according to the sentiment label "sad", add the instruction prompts "comfort the user" and "express empathy" to the prompt: "The user's emotional state is {sad}, please generate a reply to comfort the user and express empathy", so as to clearly require the large model to generate dialogue content with this sentiment adjustment for interaction.

[0092] Furthermore, in Step 4, according to the change of the sentiment label, dynamically adjust the requirements for sentiment adjustment specified by the sentiment label, dynamically generate corresponding dialogue strategies, embed the dialogue strategies into the instruction prompt of the large model, and interact according to the instruction prompt. Among them, in the multi-round dialogue, continuously track the user's emotional state and dynamically adjust the reply generated during the interaction according to the emotional change. Specifically, after each round of dialogue, re-recognize the user's sentiment label, and according to the new sentiment label, adjust the requirements for sentiment adjustment in the next round of dialogue. For example, when changing from "sad" to "angry", the dialogue strategy that needs to be generated is a more calm and rational reply, and it is embedded in the instruction, dynamically adjusting the instruction prompt: "The user's emotional state is {angry}, please generate a rational and calm reply" to optimize the interaction, so as to achieve emotional dynamic adjustment and reply generation.

[0093] In summary of the above process, the present invention can analyze the user's emotional state more comprehensively and accurately, and generate dialogue content that matches the user's emotion, thereby improving the user experience.

[0094] Embodiment 2: The present invention also provides a multi-modal sentiment analysis and interaction adjustment system based on a large model, including a collection and processing module, a feature extraction module, a sentiment recognition module, and an interaction adjustment module.

[0095] The collection and processing module collects dialogue video data and preprocesses the video data: extracts and processes video frames, audio, and text respectively.

[0096] The feature extraction module performs image feature extraction, speech feature extraction, and text feature extraction based on video frames, audio, and text respectively.

[0097] For each modality X, where X represents any one of text feature T, speech feature A, and image feature I, the sentiment recognition module generates a corresponding query vector , key vector , value vector , which is represented by the following formula:

[0098] ;

[0099] ;

[0100] ;

[0101] is a learnable weight matrix, is a bias term, and based on the cross-attention mechanism, fusion between any two modalities is performed, where the cross-attention calculation between the ( , , ) corresponding to the text feature T and the ( , , ) corresponding to the speech feature A is performed to obtain , and the cross-attention calculation between the ( , , ) corresponding to the text feature T and the ( , , ) corresponding to the image feature I is performed to obtain ; the cross-attention calculation between the ( , , ) corresponding to the speech feature A and the ( , , ) corresponding to the text feature T is performed to obtain , and the cross-attention calculation between the ( , , corresponding to the image feature I , , perform cross-attention calculation between the two modalities to obtain ; for the ( , , ) corresponding to the image feature I and the ( , , ) corresponding to the text feature T, perform cross-attention calculation between the two modalities to obtain , for the ( , , ) corresponding to the image feature I and the ( , , ) corresponding to the speech feature A, perform cross-attention calculation between the two modalities to obtain ,

[0102] Concatenate the cross-attention results of the text feature with those of other modalities to obtain ,

[0103] Concatenate the speech feature A with the cross-attention results of other modalities to obtain ,

[0104] Concatenate the image feature I with the cross-attention results of other modalities to obtain ,

[0105] Apply average pooling calculation to , and to obtain the fused multi-modal feature representation, and input the multi-modal feature representation into a multi-layer perceptron MLP:

[0106] ;

[0107] is the output probability distribution of emotion categories, and the category corresponding to the highest probability is taken as the emotion label recognition result;

[0108] The interaction adaptation module specifies the requirements for emotion adaptation according to the emotion label in the recognition result, generates corresponding dialogue strategies, conducts interactions according to the dialogue strategies, and dynamically adjusts the dialogue strategies according to the change of the emotion label.

[0109] For the information interaction and execution process among the modules in the above system, since they are based on the same concept as the method embodiments of the present invention, the specific content can be referred to the description in the method embodiments of the present invention, and will not be elaborated here.

[0110] Similarly, the system of the present invention effectively combines multi-modal data, such as text, voice, images, etc. with a large language model, and can construct an intelligent dialogue system with emotion analysis and adjustment capabilities, effectively fuse different modal data, and perform dynamic emotion adjustment, so that the intelligent dialogue interaction process can timely adjust the interaction strategy according to the change of the user's emotion, solve the problems of insufficient emotion understanding and interaction adjustment, and thus improve the naturalness and user experience in complex human-computer interaction scenarios during the intelligent dialogue process.

[0111] It should be noted that not all steps and modules in the above processes and system structures are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted according to needs. The system structures described in the above embodiments can be physical structures or logical structures, that is, some modules may be implemented by the same physical entity, or some modules may be implemented separately by multiple physical entities, or some components in multiple independent devices can be jointly implemented.

[0112] The above-described embodiments are only preferred embodiments given to fully illustrate the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or transformations made by those skilled in the art on the basis of the present invention are all within the protection scope of the present invention. The protection scope of the present invention is subject to the claims.

Claims

1. A multi-modal sentiment analysis and interaction adjustment method based on a large model, characterized in that Including: Step 1: Collect dialogue video data and preprocess the video data: extract and process video frames, audio, and text respectively; Step 2: Extract image features, speech features, and text features based on video frames, audio, and text respectively; Step 3: For each modality X, where X represents any one of the text feature T, the speech feature A, and the image feature I, generate the corresponding query vector , key vector , value vector , which are represented by the following formula: , , , is a learnable weight matrix, is a bias term, and fusion between any two modalities is performed based on the cross-attention mechanism, where the ( , , corresponding to the text feature T and the ( , , corresponding to the speech feature A are used to perform cross-attention calculation between the two modalities to obtain , and the ( , , corresponding to the text feature T and the ( , , corresponding to the image feature I are used to perform cross-attention calculation between the two modalities to obtain ; the ( , , corresponding to the speech feature A and the ( , , corresponding to the text feature T are used to perform cross-attention calculation between the two modalities to obtain , and the ( , , corresponding to the speech feature A and the ( , , corresponding to the image feature I are used to perform cross-attention calculation between the two modalities to obtain ; The ones corresponding to the image feature I ( , , ) and the ones corresponding to the text feature T ( , , ) are used to perform cross-attention calculation between the two modalities to obtain . The ones corresponding to the image feature I ( , , ) and the ones corresponding to the speech feature A ( , , ) are used to perform cross-attention calculation between the two modalities to obtain , Concatenate the text features with the cross-attention results of other modalities to obtain , Concatenate the speech feature A with the cross-attention results of other modalities to obtain , Concatenate the image feature I with the cross-attention results of other modalities to obtain , For , and Apply average pooling to calculate the fused multi-modal feature representation, and input the multi-modal feature representation into a multi-layer perceptron MLP: , For the probability distribution of the output sentiment categories, the category corresponding to the highest probability is obtained as the sentiment label recognition result; Step 4: According to the emotion labels in the recognition results, specify the requirements for emotion adjustment, generate corresponding dialogue strategies, interact according to the dialogue strategies, and dynamically adjust the dialogue strategies according to the changes in the emotion labels.

2. The multimodal sentiment analysis and interaction adaptation method based on a large model according to claim 1, characterized in that In Step 1, preprocess the video data: Extract and process video frames: Extract a sequence of image frames from the video at a preset frame rate, Process the extracted sequence of image frames, adjust the resolution to 224×224, and perform denoising and contrast enhancement operations on the sequence of image frames, Extract and process audio: Use multimedia processing tools to separate the audio track from the video, separate the audio data according to the audio track, and perform preprocessing operations of noise reduction and speech enhancement on the audio data, Extract and process text: Convert the speech data in the audio data into text content, and perform operations to remove noise words and duplicate words on the text content.

3. A multimodal sentiment analysis and interaction adjustment method based on a large model according to claim 1, characterized in that In Step 2, extract image features based on the sequence of image frames extracted from the video frames: Input the preprocessed sequence of image frames into the CLIP model, use the visual encoder of the CLIP model to extract image visual features, and apply average pooling according to the time series to generate global image visual features.

4. A multi-modal sentiment analysis and interaction adaptation method based on a large model according to claim 1, characterized in that In Step 2, extract speech features based on audio: First, normalize the audio data, sample the audio at a sampling rate of 16 kHz, input the sample into the HuBERT model, calculate the average value of the features at all time steps, and generate time-invariant global speech features according to the average value of the features. The global speech features include the content information of the speech and the emotional features of the speech. The emotional features of the speech include intonation, volume, and speech rate.

5. A multimodal sentiment analysis and interaction adaptation method based on a large model according to claim 1, characterized in that In Step 2, extract text features based on text: Input the preprocessed text into the Baichuan-7B model, use the Baichuan-7B model to segment the text, and perform word segmentation on the sentences to obtain the global feature vectors of each segment of text, and then apply average pooling to obtain the text features of the entire text.

6. A multi-modal emotion analysis and interaction adjustment method based on a large model according to claim 1, characterized in that the emotion labels of the recognition results in Step 4 involve happy, sad, angry, neutral, fear, disgust, surprise. According to the emotion labels, specify the requirements for emotion adjustment, generate corresponding dialogue strategies, embed the dialogue strategies into the instruction prompt of the large model, interact according to the instruction prompt, and dynamically adjust the dialogue strategies according to the changes in the emotion labels.

7. A multimodal sentiment analysis and interaction adaptation method based on a large model according to claim 6, characterized in that In Step 4, dynamically adjust the emotion adjustment requirements specified by the emotion labels according to the changes in the emotion labels, and dynamically generate corresponding dialogue strategies.

8. A multi-modal sentiment analysis and interaction adjustment system based on a large model, characterized in that Including a collection and processing module, a feature extraction module, an emotion recognition module, and an interaction adjustment module, The acquisition and processing module acquires the dialogue video data and preprocesses the video data: extracting and processing video frames, audio, and text respectively; The feature extraction module performs image feature extraction, speech feature extraction, and text feature extraction based on video frames, audio, and text respectively; For each modality X, where X represents any one of text feature T, speech feature A, and image feature I, the emotion recognition module generates a corresponding query vector , key vector , value vector . The emotion recognition module is represented by the following formula: , , , is a learnable weight matrix, is a bias term, and fusion between any two modalities is performed based on the cross-attention mechanism, where for the text feature T corresponding to ( , , ), and the voice feature A corresponding to ( , , ), cross-attention calculation between the two modalities is performed to obtain , and for the text feature T corresponding to ( , , ), and the image feature I corresponding to ( , , ), cross-attention calculation between the two modalities is performed to obtain ; The ( , , ) corresponding to the speech feature A and the ( , , ) corresponding to the text feature T are used to perform cross-attention calculation between the two modalities to obtain . The ( , , ) corresponding to the speech feature A and the ( , , ) corresponding to the image feature I are used to perform cross-attention calculation between the two modalities to obtain ; The ones corresponding to the image feature I ( , , ) and the ones corresponding to the text feature T ( , , ) are subjected to cross-attention calculation between the two modalities to obtain . The ones corresponding to the image feature I ( , , ) and the ones corresponding to the speech feature A ( , , ) are subjected to cross-attention calculation between the two modalities to obtain , Concatenate the text features with the cross-attention results of other modalities to obtain , Concatenate the speech feature A with the cross-attention results of other modalities to obtain , Concatenate the image feature I with the cross-attention results of other modalities to obtain , Pair , and Apply average pooling to calculate the fused multi-modal feature representation, and input the multi-modal feature representation into the multi-layer perceptron MLP: , For the probability distribution of the output sentiment categories, the category corresponding to the highest probability is taken as the sentiment label recognition result; The interaction adaptation module specifies the requirements for emotion adaptation according to the emotion labels in the recognition results, generates corresponding dialogue strategies, conducts interactions according to the dialogue strategies, and dynamically adjusts the dialogue strategies according to the changes in the emotion labels.

Citation Information

Patent Citations

  • Multi-modal sentiment analysis method combining pre-training model and self-attention block

    CN118898046A