Multi-modal sentiment analysis method and system based on adaptive difficulty instruction

By using adaptive difficulty instructions in multimodal sentiment analysis, we adaptively add multimodal information to the samples, solving the problem of failing to make full use of multimodal information in the existing methods, and significantly improving the performance of the analysis task.

CN120046005APending Publication Date: 2025-05-27KUNMING UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510209342.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing methods fail to make full use of multimodal information in multimodal sentiment analysis tasks, resulting in insufficient learning of multimodal information.

Method used

A multimodal sentiment analysis method based on adaptive difficulty instructions is adopted. By aligning the multimodal information with the text, and adaptively adding multimodal information to the samples, the accuracy, robustness and generalization ability of the multimodal sentiment analysis task is improved.

Benefits of technology

It significantly improves the performance of multimodal sentiment analysis tasks, solves the problems of information redundancy and insufficient learning, and improves the accuracy and robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120046005A_ABST
    Figure CN120046005A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-modal sentiment analysis method and system based on an adaptive difficulty instruction, and belongs to the field of natural language processing. An existing large language model emotion analysis method has defects in the aspect of integrating audios and videos, neglects that part of audio or video modalities have no effect on model emotion analysis, and even possibly has negative effects on model prediction. According to the modal self-adaptive sentiment analysis method based on the instruction following difficulty, firstly, audio and video modals are converted into natural language description, and a large language model executes multi-modal sentiment analysis through text prompt; then, according to the method, the instruction follows the difficulty function, and the purpose is to add multi-modal information to the dialogue text in a self-adaptive mode, so that the efficiency and the effect of instruction tuning are improved. Experimental results on multi-modal sentiment analysis data sets MELD and IEMOCAP show that the modal adaptive sentiment analysis method based on the instruction following difficulty provided by the invention is effective.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a multimodal sentiment analysis method and system based on adaptive difficulty instructions, belonging to the field of natural language processing. Background Art

[0002] Sentiment analysis plays a crucial role in applications such as human-computer interaction, educational assistance, and psychological counseling. Although unimodal methods, including facial expression recognition, text sentiment analysis, and audio sentiment recognition, have shown effectiveness, real-world sentiment data is usually multimodal, integrating text, audio, and images. Multimodal emotion analysis aims to combine information from multiple modalities, such as text, visual, and sound signals, to understand different human emotions.

[0003] In recent years, the emergence of large language models (LLMs) such as GPT and LLaMA has brought substantial improvements to sentiment analysis tasks. These models are pre-trained based on a large amount of text data, enabling the large language models to initially possess emergent capabilities and reasoning capabilities. After the pre-training stage, the large language models already possess rich world knowledge and basic capabilities. However, essentially, what they do is still "predict the next word". Therefore, after the pre-training stage, it is still necessary to fine-tune and align the large language models to better solve specific domain-specific tasks. However, sentiment analysis becomes more complex in multimodal data because multimodal datasets contain image information and speech information, and not all of these large language models possess multimodal capabilities. Since image information and speech information usually use different feature extraction methods, they may be mapped to different semantic spaces. Therefore, the main challenge in multimodal sentiment analysis lies in how to utilize the information of these different modalities. To solve this problem, researchers have proposed various methods, such as converting the information of different modalities into text information. Although this method has achieved certain results to some extent, it ignores that some texts already have rich context information, resulting in information redundancy after introducing multimodal information.

[0004] Therefore, in view of the deficiencies of the existing methods, the present invention proposes a modality-adaptive sentiment analysis model based on instruction-following difficulty. Summary of the Invention

[0005] The technical problem solved by the present invention is: The present invention provides a multimodal sentiment analysis method and system based on adaptive difficulty instructions to solve the problem that the existing methods fail to fully utilize multimodal information and have insufficient learning of multimodal information in multimodal sentiment analysis tasks. By combining instruction-following difficulty, the model of the present invention adaptively adds multimodal information to samples while aligning multimodal information with text, improving the accuracy, robustness, and generalization ability of multimodal sentiment analysis tasks.

[0006] The technical solution of the present invention is: a multi-modal sentiment analysis method based on adaptive difficulty instructions, and the method includes:

[0007] Step1. Preprocess the videos in the original dataset, and extract the peak frames in the videos; analyze the peak frames, and extract the background information in the pictures; process the audio segments to generate descriptions related to tone, intonation, etc.;

[0008] Step2. Construct two training datasets. One training dataset is to concatenate visual and audio information with the original dialogue text into dialogue text with multi-modal information descriptions, and the other training dataset only has dialogue text;

[0009] Step3. Randomly select a part of the samples from the training datasets constructed in Step2, and fine-tune the large language model respectively, so that the large language model has basic instruction-following capabilities;

[0010] Step4. Use the large language model obtained in Step3 to calculate the instruction-following difficulty IFD m of the dialogue text with multi-modal information descriptions t and the instruction-following difficulty IFD of the dialogue text only. Consider the samples with as those requiring multi-modal descriptions, and consider the samples with h as those not requiring multi-modal descriptions. Combine the two parts into a dataset D h and reconstruct the dataset D j into a dataset D for judging whether the sample requires multi-modal descriptions;

[0011] Step5. Fine-tune the dataset D obtained in Step4 h to obtain a model capable of performing multi-modal sentiment analysis. Fine-tune the dataset D obtained in Step4 j to obtain a large language model capable of adaptively judging whether multi-modal information is required for predicting the sentiment of the current dialogue text;

[0012] Step6. Let the trained large language model adaptively judge whether the test set requires multi-modal information, and perform sentiment analysis on the test set through the model capable of performing multi-modal sentiment analysis.

[0013] Furthermore, the Step1 includes:

[0014] Step1.1. Use the OpenFace toolkit to extract facial features from each video frame. This toolkit detects and scores action units to identify the frame with the maximum cumulative intensity, that is, the peak frame; the peak frame of the l-th sample is expressed as:

[0015]

[0016] Among them, l represents the lth sample, represents the strength of the i-th action unit in the k-th frame, Indicates return The index of the maximum value;

[0017] Step 1.2: Use the visual model to find the peak frame of the lth sample Analyze and extract context information C v,l ; Context information includes: character activities or scene environment; context information C v,l It is expressed as:

[0018]

[0019] Among them, prompt represents the prompt to guide the model to perform the specified task, and l represents the lth sample;

[0020] Step 1.3: Use the large audio model to process the audio clips and generate descriptions related to tone and intonation C a,l ;

[0021] C a,l =Audio_Model(Audio l ,prompt l );

[0022] Audio l Represents the audio segment corresponding to the lth sample.

[0023] Furthermore, the Step 2 includes:

[0024] Step 2.1, construct a dialogue text dataset D. Assuming that the dialogue text is given with a dialogue length of n, then U = [u 1 ,u 2 ,…,u l ,...u n ], which includes M speakers p 1 ,p 2 ,…,p K ...p M (M≥2), and by the corresponding speaker p K (u l ) Every word spoken l ; Function K(u l ) is used to build each utterance u l The corresponding speaker p K (u l ), the conversation text dataset D is represented as D = [U1 , U 2 , …, U z , where z is the number of samples;

[0025] Step2.2. Construct a dialogue text dataset D with multimodal information description m , concatenate visual and audio information with the original dialogue text into a dialogue text with multimodal information description, and the dialogue text dataset with multimodal information description is denoted as D m = [(C a,1 , C v,1 , U 1 ), (C a,2 , C v,2 , U 2 ), …, (C a,z , C v,z , U z )].

[0026] Furthermore, the said Step3 includes:

[0027] Step3.1. Use the k-means method to select data from the dialogue text dataset D, and then train the large language model for 3 epochs to make M D The large language model has basic instruction-following ability;

[0028] Step3.2. Use the k-means method to select data from the dialogue text dataset D with multimodal information description m , and then train the large language model for 3 epochs to make The large language model has basic instruction-following ability.

[0029] Furthermore, the said Step4 includes:

[0030] Step4.1. Use the M D large language model obtained in Step3.1 to calculate the instruction-following difficulty of each sample in the dialogue text dataset D

[0031] Step4.2. Use the large language model obtained in Step3.2 to calculate the instruction-following difficulty of each sample in the dialogue text training set D with multimodal information description m

[0032] Step4.3. Construct the dataset D h , and the construction process is as follows: traverse the datasets D and D m , and respectively take the samples in D and D mThe sample corresponding to D in Calculate the value of, when then it indicates that the large language model is more adaptable at this time then should be added to D h When then it indicates that the large language model is more adaptable at this time then should be added to D h Repeat the above steps until the traversal ends;

[0033] Step4.4. Since the dataset D h contains the hidden information of whether the sample needs multimodal description, reconstruct the dataset D h into D j D j represents querying whether the sample needs multimodal description, and the label is that if the sample is from D m , the label is that it needs multimodal description, and if the sample is from D, the label is that it does not need multimodal description.

[0034] Furthermore, in the said Step4, the instruction following difficulty is judged by comparing the loss of the large language model generating the corresponding answer after a given instruction and the loss of the large language model generating the corresponding answer without a given instruction, to determine whether the large language model can follow the instruction well;

[0035]

[0036] Among them, L θ (O|Q) The essence of this score is the objective function of the large language model LLM instruction fine-tuning, which characterizes the difficulty of the large language model LLM generating the corresponding answer when a given instruction is provided; θ represents the trainable parameters, N represents the length of the output sequence, Q represents the instruction, represents the j-th element of the output sequence, and P(;) represents the conditional probability;

[0037]

[0038] Among them, s θ (O) This score is essentially the objective function of pre-training, indicating the difficulty of the large language model directly generating an answer without a given instruction;

[0039] The instruction following difficulty is defined as:

[0040]

[0041] IFD θ(Q, O) This score is obtained by comparing L θ (O|Q) and s θ (O) to determine the importance of the instruction data to the current large language model.

[0042] Furthermore, the said Step5 includes:

[0043] Step5.1: Fine-tune the large language model using LoRA, with the input being the dataset D h , to obtain a model capable of performing multi-modal sentiment analysis This model can judge the emotion category of the input samples;

[0044] Step5.1: Fine-tune the large language model using LoRA, with the input being the dataset D j , to obtain a large language model capable of adaptively judging whether multi-modal information is required for the predicted sentiment of the current dialogue text This model can judge whether multi-modal description is required for the predicted sentiment of the input samples.

[0045] Furthermore, the said Step6 includes:

[0046] Step6.1: Use the model obtained in Step5 to judge whether multi-modal information is required for the predicted sentiment of the samples in the test set, and then adaptively add multi-modal information to the samples;

[0047] Step6.2: Reconstruct the test set into a style similar to the dataset D h , where some instructions are dialogue texts with multi-modal information descriptions, and some instructions are only dialogue texts;

[0048] Step6.3: Use to perform sentiment analysis on the test set.

[0049] The present invention also provides a multi-modal sentiment analysis system based on adaptive difficulty instructions, and the system includes: a module for executing the multi-modal sentiment analysis method based on adaptive difficulty instructions.

[0050] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the multi-modal sentiment analysis method based on adaptive difficulty instructions.

[0051] The beneficial effects of the present invention are:

[0052] 1. The present invention designs a method for converting multi-modal features into natural language descriptions, allowing the large language model to perform multi-modal sentiment analysis through text prompts;

[0053] 2. The present invention preliminarily determines whether a sample requires multimodal information by using an instruction-following difficulty index, and learns whether multimodal information is needed through a large language model.

[0054] 3. For the test set data, the present invention determines whether a sample requires multimodal information through a large language model, adaptively adds multimodal information to the sample, optimizes the situation where the model is not suitable for multimodal information for some samples, and further improves the model performance.

[0055] 4. The method of the present invention solves the problem that the existing methods fail to fully utilize multimodal information in the multimodal sentiment analysis task. The comparative experimental results show that the method of the present invention has achieved significant performance improvement in the multimodal sentiment analysis task. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1 is the framework of the modality adaptive sentiment analysis method based on instruction following difficulty in the present invention;

[0057] Figure 2 is the multimodal instruction construction process in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0058] Example 1: The following method proposed in this example was implemented on two multimodal sentiment analysis datasets (MELD, IEMOCAP);

[0059] The present invention conducts experiments using the publicly available multimodal sentiment analysis datasets MELD and IEMOCAP. IEMOCAP is a multimodal dataset widely used in emotion recognition research. The dataset consists of approximately 12 hours of audio-visual recordings of 10 actors (5 males and 5 females) performing scripted scenarios. The IEMOCAP dataset is annotated with 6 emotions: happy, sad, neutral, angry, excited, and frustrated, and contains more than 7433 utterances and 151 conversations. MELD is a multimodal sentiment dataset. The dataset contains the original dialogue text, audio, video, and emotion annotations from movies and TV shows. The dataset is from the TV series "Friends". The MELD dataset has collected 13,118 conversations, covering 7 different emotion categories (angry, disgusted, fearful, happy, neutral, sad, and surprised). Each conversation involves interactions between more than two characters, and there are emotional changes and complexities in the conversations. The statistical information of the datasets is shown in Tables 1 and 2:

[0060] Table 1 IEMOCAP Dataset Statistics

[0061]

[0062] Table 2 MELD Dataset Statistics

[0063]

[0064] As shown Figure 1 - Figure 2 below, the multi-modal sentiment analysis method based on adaptive difficulty instructions includes:

[0065] Step1. Preprocess the videos in the original dataset, extract the peak frames in the videos; analyze the peak frames, extract the background information in the pictures; process the audio segments to generate descriptions related to tone, intonation, etc.;

[0066] Furthermore, Step1 includes:

[0067] Step1.1. Use the OpenFace toolkit to extract facial features from each video frame. This toolkit detects and scores action units to identify the frame with the maximum cumulative intensity, i.e., the peak frame; the peak frame of the l-th sample is represented as:

[0068]

[0069] where l represents the l-th sample, represents the intensity of the i-th action unit of the k-th frame, represents returning the index of the maximum value;

[0070] Step1.2. Use a vision large model to analyze the peak frame of the l-th sample to extract the context information C v,l ; the context information includes: human activities or scene environments; the context information C v,l is represented as:

[0071]

[0072] where prompt represents the prompt to guide the model to perform a specified task, and l represents the l-th sample;

[0073] Step1.3. Use an audio large model to process the audio segments to generate descriptions related to tone and intonation C a,l ;

[0074] C a,l = Audio_Model(Audio l , prompt l );

[0075] where Audio l represents the audio segment corresponding to the l-th sample.

[0076] Step 2. Construct two training datasets. One training dataset is the dialogue text concatenated with visual and audio information into a dialogue text with multimodal information description, and the other training dataset only contains the dialogue text;

[0077] Further, the Step 2 includes:

[0078] Step 2.1. Construct a dialogue text dataset D. Assuming a dialogue text with a given dialogue length of n, then U = [u 1 , u 2 , …, u l ,... u n , which includes M speakers p 1 , p 2 , …, p K ... p M (M ≥ 2), and each utterance u K (u l ) spoken by the corresponding speaker p l ; The function K(u l ) is used to establish the mapping between each utterance u l and its corresponding speaker p K (u l ), then the dialogue text dataset D is represented as D = [U 1 , U 2 , …, U z , where z is the number of samples;

[0079] Step 2.2. Construct a dialogue text dataset D m with multimodal information description, concatenate visual and audio information with the original dialogue text into a dialogue text with multimodal information description, and the dialogue text dataset with multimodal information description is represented as D m = [(C a,1 , C v,1 , U 1 ), (C a,2 , C v,2 , U 2 ), …, (C a,z , C v,z , U z )].

[0080] Step 3. Randomly select a part of the samples from the training datasets constructed in Step 2 respectively, and fine-tune the large language model respectively, so that the large language model has the basic instruction following ability;

[0081] Further, the Step 3 includes:

[0082] Step3.1. Select data from the dialogue text dataset D using the k-means method, and then train the large language model for 3 epochs to enable M D The large language model has basic instruction-following capabilities;

[0083] Step3.2. Select data from the dialogue text dataset D with multimodal information descriptions using the k-means method, and then train the large language model for 3 epochs to enable m The large language model has basic instruction-following capabilities. The large language model has basic instruction-following capabilities.

[0084] Step4. Use the large language model obtained in Step3 to calculate the instruction-following difficulty IFD of the dialogue text with multimodal information descriptions m and the instruction-following difficulty IFD of the dialogue text only t . Consider the samples with as those requiring multimodal descriptions, and the samples with as those not requiring multimodal descriptions. Combine the two parts into a dataset D h . Reconstruct the dataset D h into a dataset D for judging whether a sample requires multimodal descriptions j ;

[0085] Furthermore, Step4 includes:

[0086] Step4.1. Use the M D large language model obtained in Step3.1 to calculate the instruction-following difficulty of each sample in the dialogue text dataset D The instruction-following difficulty is determined by comparing the loss of the large language model in generating a corresponding answer after a given instruction and the loss of the large language model in generating a corresponding answer without a given instruction, to judge whether the large language model can follow the instruction well;

[0087]

[0088] where L θ (O|Q) This score is essentially the objective function of the instruction fine-tuning of the large language model LLM, representing the difficulty of the large language model LLM in generating a corresponding answer when a given instruction is provided; θ represents the trainable parameters, N represents the length of the output sequence, Q represents the instruction, represents the j-th element of the output sequence, and P(;) represents the conditional probability;

[0089]

[0090] where s θ(O) This score is essentially the pre-training objective function, indicating the difficulty for the large language model to directly generate answers without given instructions;

[0091] The instruction following difficulty is defined as:

[0092]

[0093] IFD θ (Q,O) This score determines the importance of the instruction data pair for the current large language model by comparing the magnitudes of L θ (O|Q) and s θ (O).

[0094] Step4.2. Use the M obtained in Step3.2 Dm The large language model calculates the instruction following difficulty of each sample in the dialogue text training set D with multimodal information description m

[0095] Step4.3. Construct the dataset D h , and the construction process is as follows: traverse the datasets D and D m , respectively take the samples in D and the corresponding samples in D m in D Calculate the value of , when , it means that the large language model is more adaptable at this time then should be added to D h , when , it means that the large language model is more adaptable at this time then should be added to D h , and repeat the above steps until the traversal ends;

[0096] Step4.4. Since the dataset D h contains the hidden information of whether the sample requires multimodal description, reconstruct the dataset D h as D j , D j represents querying whether the sample requires multimodal description, and the label is that if the sample is from D m , the label is that it requires multimodal description, and if the sample is from D, the label is that it does not require multimodal description.

[0097] Step5. Fine-tune the dataset D obtained in Step4 h to obtain a model capable of multimodal sentiment analysis. Fine-tune the dataset D obtained in Step4 j ​, obtain a large language model that can adaptively determine whether multimodal information is needed for predicting the sentiment of the current dialogue text;

[0098] Further, the said Step5 includes:

[0099] Step5.1. Use LoRA to fine-tune the large language model with the input being the dataset D h , and obtain a model capable of performing multimodal sentiment analysis This model can judge the emotion category of the input samples;

[0100] Step5.1. Use LoRA to fine-tune the large language model with the input being the dataset D j , and obtain a large language model that can adaptively determine whether multimodal information is needed for predicting the sentiment of the current dialogue text This model can judge whether multimodal description is needed for the predicted sentiment of the input samples.

[0101] Step6. Let the trained large language model adaptively determine whether the test set needs multimodal information, and perform sentiment analysis on the test set through the model capable of performing multimodal sentiment analysis.

[0102] Further, the said Step6 includes:

[0103] Step6.1 Use the model obtained in Step5 to judge whether multimodal information is needed for the predicted sentiment of the samples in the test set, and then adaptively add multimodal information to the samples;

[0104] Step6.2 Reconstruct the test set into a style similar to the dataset D h , where some instructions are dialogue texts with multimodal information descriptions, and some instructions only have dialogue texts;

[0105] Step6.3 Use to perform sentiment analysis on the test set.

[0106] The present invention also provides a multimodal sentiment analysis system based on adaptive difficulty instructions, and the said system includes: a module for executing the multimodal sentiment analysis method based on adaptive difficulty instructions.

[0107] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, it implements the multimodal sentiment analysis method based on adaptive difficulty instructions.

[0108] Evaluation metric: The present invention uses the F1 score as the main evaluation metric. The F1 score combines the precision and recall of the model and can comprehensively measure the prediction ability of the model on different categories. The higher the F1 score, the better the model can balance correct classification and missed classification when processing samples, thus showing stronger overall performance.

[0109] To verify the effectiveness of the model proposed by the present invention, the method of the present invention is compared with the following baseline methods with different annotation frameworks, specifically as follows:

[0110] DialogueRNN uses three different GRUs to separately track the speaker, context, and emotional state in the dialogue. This model connects text, acoustic, and visual features to obtain a multi-modal discourse representation.

[0111] MMGCN constructs a dialogue graph based on all three modalities and designs a multi-modal fusion graph convolutional network to simulate the context dependencies between multiple modalities.

[0112] DialogueTRM uses a hierarchical transformer to manage the differential context preferences in each modality and designs a multi-granularity interactive fusion for learning the different contributions of different modalities to the dialogue utterance.

[0113] MM-DFN designs a graph-based dynamic fusion module to fuse multi-modal context features, which can reduce redundancy and enhance the complementarity between modalities.

[0114] InstructERC adopts a retrieval multi-task LLMs framework, which revolutionizes emotion recognition in conversations. In addition, it introduces two auxiliary tasks, including speaker recognition and emotion prediction, to promote emotional consistency.

[0115] UniMSE uses T5 to fuse acoustic and visual modality features with multi-level text features and performs cross-modal contrast learning to obtain discriminative multi-modal representations.

[0116] BiosERC classifies the emotional labels of each utterance by studying the features of the speaker in the dialogue, extracting the "biographical information" of the speaker in the dialogue, and injecting it into the model as supplementary knowledge.

[0117] The method of the present invention has been effectively compared with the baseline model on the IEMOCAP and MELD datasets. To verify the effectiveness of the proposed model of the present invention, the method of the present invention has been compared with the following several baseline models, and the specific results are shown in Table 3: (1) Overall performance: On both datasets ("IEMOCAP" and "MELD"), the method of the present invention shows the best F1 score, indicating that the method of the present invention has strong advantages in the multi-modal sentiment analysis task. (2) Comparison with feature-based baseline methods: Compared with traditional feature-based methods, such as DialogueTRM, the F1 score of the method of the present invention is 2.61% and 4.37% higher than that of the best-performing feature-based baseline model respectively, highlighting the promotion effect of the large language model on task performance. (3) Comparison with large language model-based methods: Both the fine-tuning and prompt tuning strategies of InstructERC have achieved remarkable results, demonstrating the superior adaptability of the large language model in processing sentiment analysis tasks. Although InstructERC promotes sentiment consistency through two auxiliary tasks, it does not effectively utilize multi-modal information. Advantages of the method of the present invention: Compared with models of other large language tasks, such as InstructERC, the method of the present invention achieves more significant performance improvement by converting video and audio into natural language descriptions and introducing instruction-following difficulty for fine-grained adjustment of modalities. The method of the present invention shows obvious advantages in the large language model-based setting, and the F1 score is increased by 1.15% and 0.98% compared with the baseline method respectively, showing the strong potential of the model in multi-modal sentiment analysis. Through the above analysis, it can be seen that the method of the present invention shows more excellent performance under the settings of different benchmark datasets, especially more advantageous than feature-based methods.

[0118] Table 3 shows the experimental results on the IEMOCAP and MELD datasets

[0119]

[0120] The specific implementation manners of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above implementation manners. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention.

Claims

1. A multimodal sentiment analysis method based on adaptive difficulty instructions, characterized by: The method comprises: Step 1: Preprocess the videos in the original dataset and extract the peak frames in the videos; analyze the peak frames and extract the background information in the pictures; process the audio clips and generate descriptions related to tone and intonation; Step 2: Construct two training datasets. One training dataset is a dialogue text with multimodal information description that is concatenated with visual and audio information and original dialogue text, and the other training dataset is only dialogue text. Step 3: Randomly select a portion of samples from the training data set constructed in Step 2, and fine-tune the large language model respectively, so that the large language model has basic instruction following capabilities; Step 4: Use the large language model obtained in Step 3 to calculate the instruction following difficulty IFD of the dialogue text with multimodal information description m IFD with only dialogue text t ,Will The samples of The samples of are considered as not requiring multimodal description, and the two parts are merged into one dataset D h , the data set D h Reconstructed as a dataset D to determine whether a sample needs multimodal description j ; Step 5: Fine-tune the dataset D obtained in Step 4 h , to obtain a model capable of multimodal sentiment analysis, and fine-tune the dataset D obtained in Step 4 j , obtain a large language model that can adaptively determine whether multimodal information is needed to predict the emotion of the current conversation text; Step 6: Use the trained large language model to adaptively determine whether the test set needs multimodal information, and perform sentiment analysis on the test set using a model that can perform multimodal sentiment analysis.

2. The multimodal sentiment analysis method based on adaptive difficulty instructions according to claim 1, characterized in that: The Step 1 includes: Step 1.1, extract facial features from each video frame using the OpenFace toolkit, which detects and scores action units to identify the frame with the largest cumulative intensity, i.e., the peak frame; the peak frame of the lth sample It is expressed as: Among them, l represents the lth sample, represents the strength of the i-th action unit in the k-th frame, Indicates return The index of the maximum value; Step 1.2: Use the visual model to find the peak frame of the lth sample Analyze and extract context information C v,l ; Context information includes: character activities or scene environment; context information C v,l It is expressed as: Among them, prompt represents the prompt to guide the model to perform the specified task, and l represents the lth sample; Step 1.3: Use the large audio model to process the audio clips and generate descriptions related to tone and intonation C a,l ; C a,l =Audio_Model(Audio l ,prompt l ); Audio l Represents the audio segment corresponding to the lth sample.

3. The multimodal sentiment analysis method based on adaptive difficulty instructions according to claim 1, characterized in that: The Step 2 includes: Step 2.1, construct a conversation text dataset D. Assuming that the conversation length is n, then U = [u1,u2,…,u l ,...u n ], which includes M speakers p1, p2, …, p K ...p M (M≥2), and by the corresponding speaker p K (u l ) Every word spoken l ; The function K(u1) is used to establish each utterance u l The corresponding speaker p K (u l ), the conversation text dataset D is represented as D = [U1,U2,…,U z ], where z is the sample size; Step 2.2: Construct a dialogue text dataset D with multimodal information description m , the visual and audio information are concatenated with the original dialogue text into a dialogue text with multimodal information description. The dialogue text dataset with multimodal information description is denoted as D m =[(C a,1 ,C v,1 ,U1),(C a,2 ,C v,2 ,U2),…,(C a,z ,C v,z ,U z )].

4. The multimodal sentiment analysis method based on adaptive difficulty instructions according to claim 1, characterized in that: The Step 3 includes: Step 3.

1. Use the k-means method to select data from the conversation text dataset D, and then train the large language model for 3 epochs so that M D The large language model has basic command-following capabilities; Step 3.2: Use the k-means method to extract the conversation text dataset D with multimodal information description. m Select data from the dataset and train the large language model for 3 epochs. The large language model has basic instruction following capabilities.

5. The multimodal sentiment analysis method based on adaptive difficulty instructions according to claim 4, characterized in that: The Step 4 includes: Step 4.

1. Use the M obtained in Step 3.1 D The large language model calculates the instruction following difficulty of each sample in the dialogue text dataset D Step 4.2, use the information obtained in Step 3.2 Large language model calculation with multimodal information description of the dialogue text training set D m The instruction following difficulty of each sample Step 4.

3. Build dataset D h , the construction process is as follows: traverse the data sets D and D m , take samples from D respectively and D m The sample corresponding to D calculate When , it means that the large language model is more suitable for Then it should be Add to D h ,when , it means that the large language model is more suitable for Then it should be Add to D h , repeat the above steps until the traversal is completed; Step 4.4, due to the data set D h contains the hidden information of whether the sample needs multimodal description. h Refactoring to D j , D j It is expressed as whether the query sample needs a multimodal description, and the label is if the sample is from D m If , the label is that multimodal description is required. If the sample is from D, the label is that multimodal description is not required.

6. The multimodal sentiment analysis method based on adaptive difficulty instructions according to claim 1, characterized in that: In Step 4, the instruction following difficulty is determined by comparing the loss of the large language model generating a corresponding answer after a given instruction and the loss of the large language model generating a corresponding answer without a given instruction, to determine whether the large language model can follow the instruction well; Among them, L θ The essence of this score (O|Q) is the objective function of fine-tuning the LLM instruction, which represents the difficulty of the LLM to generate the corresponding answer when the current LLM is given an instruction; θ represents the trainable parameter, N represents the length of the output sequence, and Q represents the instruction. represents the jth element of the output sequence, and P(;) represents the conditional probability; Among them, s θ (O) This score is essentially the objective function of pre-training, indicating how easy it is for a large language model to directly generate answers without given instructions; The difficulty of following instructions is defined as: IFD θ (Q,O) This score is compared with L θ (O|Q) and s θ The size of (O) is used to determine the importance of instruction data to the current large language model.

7. The multimodal sentiment analysis method based on adaptive difficulty instructions according to claim 5, characterized in that: The Step 5 includes: Step 5.

1. Fine-tune the large language model using LoRA, with the input being dataset D h , obtain a model capable of multimodal sentiment analysis The model can determine the emotion category of the input sample; Step 5.

1. Fine-tune the large language model using LoRA, with the input being dataset D j , obtain a large language model that can adaptively determine whether multimodal information is needed to predict the emotion of the current dialogue text The model can determine whether multimodal description is needed to predict the sentiment of the input sample.

8. The multimodal sentiment analysis method based on adaptive difficulty instructions according to claim 7, characterized in that: The Step 6 includes: Step 6.1 Use the model obtained in Step 5 Determine whether the samples in the test set need multimodal information to predict emotions, and then adaptively add multimodal information to the samples; Step 6.2 Reconstruct the test set into the same h Similar style, some commands are dialogue text with multimodal information description, and some commands are only dialogue text; Step 6.3 Use Perform sentiment analysis on the test set.

9. A multimodal sentiment analysis system based on adaptive difficulty instructions, characterized in that: The system comprises: a module for executing the multimodal sentiment analysis method based on adaptive difficulty instructions as described in any one of claims 1 to 8.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the multimodal sentiment analysis method based on adaptive difficulty instructions as described in any one of claims 1 to 8 is implemented.