Multi-modal large model emotion understanding enhancement method and device

By constructing an emotion difficulty scoring function and a multimodal attention mechanism, a subset of multimodal large language models suitable for emotion training is selected, which solves the problems of poor model generalization ability and high training cost, improves the accuracy and robustness of emotion understanding, and is applicable to mental health and social emotion computing.

CN121859986APending Publication Date: 2026-04-14WUHAN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Multimodal large language models suffer from poor generalization ability, ambiguous emotional expression, redundant training data, and high fine-tuning costs in emotion understanding tasks. Existing methods have failed to effectively characterize the dominant role of visual modalities in emotion recognition, resulting in poor adaptability in emotional scenarios.

Method used

By constructing an 'emotional difficulty scoring function' that combines model confidence and emotional sensitivity, samples with clear or ambiguous emotional information are identified. By combining multimodal attention mechanisms and perturbation sensitivity analysis, emotionally relevant regions are quantified, and a subset suitable for emotional training is selected for fine-tuning the multimodal large language model.

Benefits of technology

Without relying on additional training, the model selects the most inspiring and emotionally challenging samples, reducing training costs while improving the accuracy and robustness of emotion recognition. This enables interpretable modeling of emotional cognitive behavior, making it suitable for practical applications such as mental health and social emotion computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121859986A_ABST
    Figure CN121859986A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal large model emotion understanding enhancement method and device, and belongs to the technical field of large language models.The method comprises the steps that an initial sample set is obtained; inputting the first training sample into a first multi-modal large language model, and obtaining a first emotional disturbance graph corresponding to the first training sample according to a cross attention operation result of an image and a text of the first training sample; inputting the first training sample and the first emotional disturbance graph into a first multi-modal large language model, and calculating a first confidence score and a first emotional disturbance score; determining an emotional difficulty score of the first training sample based on the first confidence score and the first emotional disturbance score; and according to the emotion difficulty score of each training sample in the initial sample set, screening an emotion training subset from the initial sample set. According to the method, the subset suitable for performing emotion training on the multi-modal large model can be effectively screened out from the training set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of large language model technology, and in particular to a method and apparatus for enhancing multimodal large model emotion understanding. Background Technology

[0002] In recent years, multimodal large language models have developed rapidly, demonstrating outstanding performance in areas such as image-text joint modeling, image captioning generation, and multimodal dialogue due to their powerful image-text joint modeling capabilities. These models typically employ large-scale image and text pairs for pre-training, combining a visual encoder and a language generation module to achieve the understanding and generation of complex multimodal semantics. However, despite their good performance in general tasks, multimodal large language models still face significant bottlenecks in the highly subjective, nuanced, and context-dependent task of sentiment understanding.

[0003] Emotional cognition tasks require models to capture subtle emotional cues from visual, linguistic, and even interactive sources, such as facial expressions, postures, scene atmosphere, and tone of voice. However, in existing large-scale model pre-training corpora, these highly emotion-related semantic signals are often scarce or even ignored. Furthermore, emotions themselves are highly subjective, diverse in expression, and context-sensitive; different samples exhibit significant differences in emotional intensity, clarity of expression, and cultural background, further increasing the challenge of emotion understanding.

[0004] Researchers use "data pruning" or "sample selection" techniques to select representative training samples with high emotional density or high cognitive value from pre-training corpora in order to improve training efficiency and downstream performance.

[0005] Currently, some methods have attempted to select a subset of pre-training corpora by using information such as model confidence, loss value, or training dynamics. However, most of these methods ignore the dominant role of visual modalities in emotion recognition and fail to effectively characterize the important dimension of "emotional difficulty," thus exhibiting poor adaptability in emotional scenarios.

[0006] In summary, how to identify the most inspiring and emotionally challenging samples from sentiment datasets without relying on additional training, thereby reducing training costs while preserving diversity and information density, is a crucial problem that urgently needs to be solved in current multimodal large-scale model sentiment cognition tasks. Summary of the Invention

[0007] This invention provides a method and apparatus for enhancing sentiment understanding in multimodal large models, which can effectively select a subset from the training set suitable for sentiment training of multimodal large models. The technical solution includes at least the following: Firstly, a method for enhancing sentiment understanding in a multimodal large-scale model is provided, comprising: acquiring an initial sample set, wherein the initial sample set contains multiple training samples, each training sample containing an image, text, and a sentiment label; inputting a first training sample into a first multimodal large-scale language model, and obtaining a first sentiment perturbation map corresponding to the first training sample based on the cross-attention operation result of the image and text of the first training sample, wherein sentiment-independent regions in the sentiment perturbation map are set as masks; inputting the first training sample and the first sentiment perturbation map into the first multimodal large-scale language model to obtain a first training sample probability corresponding to the first training sample and a first sentiment perturbation probability corresponding to the first sentiment perturbation map; calculating a first confidence score based on the first training sample probability, and calculating a first sentiment perturbation score based on the first sentiment perturbation probability; determining the sentiment difficulty score of the first training sample based on the first confidence score and the first sentiment perturbation score; and selecting a sentiment training subset from the initial sample set according to the sentiment difficulty score of each training sample in the initial sample set, wherein the sentiment training subset is used to fine-tune the first multimodal large-scale language model.

[0008] Optionally, the step of inputting the first training sample into the first multimodal large language model and obtaining the first emotion perturbation map corresponding to the first training sample based on the cross-attention operation result of the image and text of the first training sample includes: obtaining a first visual token based on the image of the first training sample; obtaining a first text token based on the text of the first training sample; inputting the first visual token and the first text token into the first multimodal large language model for cross-attention operation to obtain a first attention matrix; calculating the emotional attention score of each visual position of the first visual token on the first attention matrix according to a preset set of emotion cue words, wherein the set of emotion cue words is used to indicate the set of emotion-related tokens in the text token; obtaining an emotion-independent region according to a preset emotion threshold and the emotional attention score of each visual position of the first visual token, and setting the emotion-independent region as a mask in the first visual token to obtain the first emotion perturbation map.

[0009] Optionally, calculating the first confidence score based on the probability of the first training sample includes:

[0010] in, The first confidence level score is given. For the first training sample, The probability of the first training sample is the first The probability of each category, It is a positive integer. The value range is 1 to , Total number of categories; The calculation of the first emotional disturbance score based on the first emotional disturbance probability includes:

[0011] in, The first emotional disturbance score is given. It represents the probability of the k-th category in the first emotional disturbance probability.

[0012] Optionally, determining the sentiment difficulty score of the first training sample based on the first confidence score and the first sentiment perturbation score includes:

[0013] in, The emotion difficulty score for the first training sample. These are the preset weighting coefficients.

[0014] Optionally, the step of selecting an emotional training subset from the initial sample set according to the emotional difficulty score of each training sample in the initial sample set includes: sorting the training samples in the initial sample set in descending order according to the emotional difficulty score of each training sample in the initial sample set to obtain a first sort; dividing the first sort into a difficult sample set, a medium sample set, and an easy sample set, wherein the difficult sample set is the set consisting of the top q% of training samples in the first sort, the easy sample set is the set consisting of the bottom q% of training samples in the first sort, and the medium sample set is the set consisting of training samples in the first sort excluding the difficult sample set and the easy sample set; and selecting the emotional training subset from the difficult sample set, the medium sample set, and the easy sample set respectively.

[0015] Optionally, the step of filtering from the difficult sample set, the medium sample set, and the easy sample set to obtain the emotion training subset includes: filtering from the difficult sample set with the goal of maximizing diversity. Training samples, selected from the simple sample set training samples, according to The proportion is randomly selected from the medium sample set, and the set of all selected training samples constitutes the emotion training subset; wherein, This is a preset proportional coefficient. , These are the training budget parameters set by the user. , It is the total number of training samples in the initial sample set.

[0016] Secondly, a multimodal large-scale model emotion understanding enhancement device is also provided, comprising: a sample acquisition module for acquiring an initial sample set, wherein the initial sample set contains multiple training samples, each training sample containing an image, text, and an emotion label; a perturbation map acquisition module for inputting a first training sample into a first multimodal large-scale language model, and obtaining a first emotion perturbation map corresponding to the first training sample based on the cross-attention operation result of the image and text of the first training sample, wherein emotion-independent regions in the emotion perturbation map are set as masks; and a probability calculation module for inputting the first training sample and the first emotion perturbation map into the first multimodal large-scale language model to obtain... The system includes: a first training sample probability corresponding to the first training sample and a first emotion perturbation probability corresponding to the first emotion perturbation map; a scoring calculation module for calculating a first confidence score based on the first training sample probability and a first emotion perturbation score based on the first emotion perturbation probability; an emotion difficulty score determination module for determining the emotion difficulty score of the first training sample based on the first confidence score and the first emotion perturbation score; and a filtering module for filtering an emotion training subset from the initial sample set according to the emotion difficulty score of each training sample in the initial sample set, wherein the emotion training subset is used to fine-tune the first multimodal large language model.

[0017] Optionally, the perturbation map acquisition module is further configured to: acquire a first visual token based on the image of the first training sample; acquire a first text token based on the text of the first training sample; input the first visual token and the first text token into the first multimodal large language model for cross-attention operation to obtain a first attention matrix; calculate the emotional attention score of each visual position of the first visual token on the first attention matrix according to a preset set of emotional cue words, wherein the set of emotional cue words is used to indicate the set of emotionally related tokens in the text token; and obtain an emotion-independent region according to a preset emotion threshold and the emotional attention score of each visual position of the first visual token, and set the emotion-independent region as a mask in the first visual token to obtain the first emotion perturbation map.

[0018] Optionally, in the scoring calculation module, calculating the first confidence score based on the probability of the first training sample includes:

[0019] in, The first confidence level score is given. For the first training sample, The probability of the first training sample is the first The probability of each category, It is a positive integer. The value range is 1 to , Total number of categories; The calculation of the first emotional disturbance score based on the first emotional disturbance probability includes:

[0020] in, The first emotional disturbance score is given. It represents the probability of the k-th category in the first emotional disturbance probability.

[0021] Optionally, the sentiment difficulty score determination module is further configured to determine the sentiment difficulty score of the first training sample using the following formula:

[0022] in, The emotion difficulty score for the first training sample. These are the preset weighting coefficients.

[0023] Optionally, the filtering module is further configured to sort the training samples in the initial sample set in descending order according to the sentiment difficulty score of each training sample in the initial sample set to obtain a first sort; divide the first sort into a difficult sample set, a medium sample set, and an easy sample set, wherein the difficult sample set is the set of the top q% of training samples in the first sort, the easy sample set is the set of the bottom q% of training samples in the first sort, and the medium sample set is the set of training samples in the first sort excluding the difficult sample set and the easy sample set; and filter from the difficult sample set, the medium sample set, and the easy sample set respectively to obtain the sentiment training subset.

[0024] Optionally, the screening module is further configured to screen samples from the difficult sample set with the goal of maximizing diversity. Training samples, selected from the simple sample set training samples, according to The proportion is randomly selected from the medium sample set, and the set of all selected training samples constitutes the emotion training subset; wherein, This is a preset proportional coefficient. , These are the training budget parameters set by the user. , It is the total number of training samples in the initial sample set.

[0025] Thirdly, a computer device is also provided, comprising: a memory and a processor, wherein the memory stores at least one computer program, the at least one computer program being loaded and executed by the processor to perform the multimodal large model emotion understanding enhancement method described in the above embodiments.

[0026] Fourthly, a computer-readable storage medium is also provided, wherein at least one computer program is stored in the computer-readable storage medium, the at least one computer program being loaded and executed by a processor to perform the multimodal large model sentiment understanding enhancement method described in the above embodiments.

[0027] Fifthly, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the method described in the first aspect.

[0028] The beneficial effects of the technical solution provided by this invention include at least the following: In this embodiment, by considering confidence scores and emotional perturbation scores in the calculation of emotional difficulty scores, the ambiguity of sample emotional expression and model uncertainty can be estimated in real time without model gradients or backpropagation information. The selected emotional training subset can contain the most inspiring and emotionally challenging samples for the model, thereby reducing training costs while preserving diversity and information density. By masking emotionally irrelevant regions in the emotional perturbation map, this invention can quantitatively analyze the model's dependence on emotional regions, achieving interpretable modeling of emotional cognitive behavior. The multimodal large-model emotional understanding enhancement method in this invention does not depend on a specific model structure, has good transferability, and can be directly integrated into existing mainstream MLLM fine-tuning processes such as LLaVA and InternVL, making it suitable for practical applications such as mental health and social emotion computing. Attached Figure Description

[0029] To more clearly illustrate the technical solutions in this embodiment, the accompanying drawings used in the description of the embodiment will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 A flowchart of a multimodal large-model emotion understanding enhancement method provided by an exemplary embodiment of the present invention is shown; Figure 2 This is a schematic diagram of the entire process of the multimodal large-model sentiment understanding enhancement algorithm; Figure 3A flowchart of a multimodal large language model training method provided by an exemplary embodiment of the present invention is shown; Figure 4 This is a flowchart of the multimodal large language model training algorithm; Figure 5 This diagram illustrates the structure of a multimodal large-model emotion understanding enhancement device provided in an exemplary embodiment of the present invention. Figure 6 This is a schematic diagram of the structure of a computer device provided in an exemplary embodiment of the present invention. Detailed Implementation

[0031] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains. The terms “first,” “second,” “third,” and similar terms used in this patent application specification and claims do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an” or “a” and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms “comprising” or “including” and similar terms mean that the elements or objects preceding “comprising” or “including” encompass the elements or objects listed following “comprising” or “including” and their equivalents, but do not exclude other elements or objects.

[0032] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0033] To address the problems of poor generalization ability, ambiguous emotion expression, redundant training data, and high fine-tuning costs faced by multimodal large language models (MLLMs) in emotion understanding tasks, this invention discloses an adaptive multimodal large model emotion understanding enhancement method (Heart Prune). This method can effectively evaluate the "emotional challenge" of emotion samples without model retraining and can be used for fine screening of training samples and improvement of emotion cognition performance.

[0034] The core idea of ​​this invention is to identify samples with clear or ambiguous emotional information by constructing an "emotional difficulty scoring function" that combines model confidence and emotional sensitivity; to quantify the emotion-related regions in the input image by combining multimodal attention mechanism and perturbation sensitivity analysis; and finally to select training subsets hierarchically based on the scoring values, thereby significantly improving the accuracy and robustness of emotion recognition under limited training budget.

[0035] The Heart Prune method proposed in this invention can be integrated into the fine-tuning process of any pre-trained multimodal large model.

[0036] Figure 1 A flowchart of a multimodal large model emotion understanding enhancement method provided by an exemplary embodiment of the present invention is shown, which can be executed by a computer device. Figure 2 This is a schematic diagram illustrating the entire process of a multimodal large-model sentiment understanding enhancement algorithm. (See also...) Figure 1-2 The method includes: In step 101, the initial sample set is obtained.

[0037] The initial sample set contains multiple training samples, each containing an image, text, and emotion label.

[0038] The method in this embodiment is used to filter existing multimodal sentiment datasets. The filtered training samples (sentiment training subset) can be used to fine-tune a pre-trained multimodal large language model, thereby improving the sentiment understanding performance of the multimodal sentiment dataset. The initial sample set here can be any multimodal sentiment dataset that needs to be filtered.

[0039] A training sample in a multimodal sentiment dataset typically contains two modalities (images and text), and each training sample also has a corresponding sentiment label, which is pre-labeled.

[0040] In a training sample, images and text exist in pairs. The text portion can be a title, description, or instruction-style question corresponding to the image. Emotion labels are discrete emotion categories (such as happy, angry, sad, etc.).

[0041] In step 102, the first training sample is input into the first multimodal large language model, and the first emotion perturbation map corresponding to the first training sample is obtained based on the cross-attention operation result of the image and text of the first training sample.

[0042] The emotion-independent regions in the emotion perturbation map are masked. The first training sample is a sample from the initial sample set.

[0043] Optionally, step 102 includes steps a to e as follows.

[0044] Step a: Obtain the first visual token based on the image of the first training sample.

[0045] Step b: Obtain the first text token based on the text of the first training sample.

[0046] When implementing step a, the image can be compressed using an image encoder. In step b, the visual token of dimension 1 can be obtained by decomposing the text of the first training sample into multiple tokens using a tokenizer and an embedding layer. The first text token. Here, It is the total number of visual tokens in the first visual token. It represents the total number of text tokens in the first text token.

[0047] In this embodiment, the first multimodal large language model can be BLIP, Flamingo, MiniGPT-4, etc. The basic structure of the first multimodal large language model includes an image encoder, a text encoder (such as a tokenizer and an embedding layer), and a cross-modal fusion module (used to perform cross-attention operations in step c).

[0048] Step c: Input the first visual token and the first text token into the first multimodal large language model to perform cross-attention operation and obtain the first attention matrix.

[0049] For example, using This represents the first training sample. It is the first sample. It is a first-view token. It is the first text token.

[0050] Optionally, when the first visual token and the first text token are input into the first multimodal large language model, an explicit emotion guide (such as "Focus on emotions") can be inserted into the first text token. The explicit emotion guide can prompt the model to focus its attention on areas in the image that have the potential to express emotions, such as facial features, body movements, and color atmosphere, thereby completing the emotion classification task more accurately and guiding the model to focus on key emotional cues.

[0051] For example, the first attention matrix is ​​calculated using formula (1).

[0052] (1) In formula (1), This is the first attention matrix. , It is the query representation of the first text token. It is the key representation of the first visual token. For feature dimension, This represents the transpose of a matrix.

[0053] In this first attention matrix, the i-th row represents the attention weight of the i-th text token of the first text token on each visual token of the first visual token, and the j-th column represents the attention weight of the j-th visual token of the first visual token on each text token of the first text token. The value of i ranges from 1 to... The value of j ranges from 1 to , where i and j are both positive integers. The position of the first attention matrix is ​​the i-th row and j-th column, representing the attention weight of the j-th visual token of the first visual token on the i-th text token of the first text token.

[0054] In other words, the first visual token for each location corresponds to... Attention weights for the first text token.

[0055] By using a cross-attention mechanism, the semantic alignment of the first visual token and the first text token is achieved, forming a joint semantic space representation for sentiment discrimination.

[0056] Step d: Based on the preset set of emotional cue words, calculate the emotional attention score for each visual position of the first visual token on the first attention matrix.

[0057] The set of emotion cue words is used to indicate the set of text tokens that are related to emotions.

[0058] In this embodiment, the set of emotion prompt words is a preset set of emotion-related tokens, such as "emotion".

[0059] For the j-th visual position (equivalent to the j-th visual token of the first visual token), its emotional attention score is calculated using the following formula (2).

[0060] (2) In formula (2), This represents the emotional attention score for the j-th visual position. It is equivalent to the emotional attention score for the j-th visual token of the first visual token. A set of emotion cue words, Let be the i-th row and j-th column of the first attention matrix. It should be noted that when using formula (2) to calculate the emotional attention score for the j-th visual position, only the j-th column of the first attention matrix needs to be used, and the other columns do not need to be used.

[0061] Since column j represents the attention weight of the jth visual token of the first visual token on each text token of the first text token, formula (2) means that if some text tokens in the first text token belong to the set of emotional cue words, then from the attention weights of each text token corresponding to the jth visual token, the attention weights of the text tokens belonging to the set of emotional cue words are summed to obtain the emotional attention score corresponding to the jth visual token.

[0062] For example, the j-th column of the first attention matrix for Furthermore, the first and fourth text tokens of the first text token belong to the set of sentiment cue words. Therefore, for... and Summing these values, we get 0.6 + 0.2 = 0.8. Therefore, the attention score for the j-th visual token is 0.8.

[0063] After applying formula (2) to each visual position, the emotional attention score for each visual position of the first visual token can be obtained.

[0064] Emotional attention score is used to represent the relative importance of a visual location to emotional understanding.

[0065] Step e: Based on the preset emotion threshold and the emotion attention score of each visual position of the first visual token, obtain the emotion-independent region, and set the emotion-independent region as a mask in the first visual token to obtain the first emotion perturbation map.

[0066] In one alternative implementation, the emotion threshold If the value is fixed, then the emotional attention score of the first visual token is less than the emotional threshold. The visual location (i.e., the emotion-independent region) is set as a mask, and the emotional attention score of the first visual token is greater than or equal to the emotional threshold. The visual position remains unchanged, and the first visual token with the mask is the first emotion perturbation graph. For easier visualization, the first emotional disturbance graph can also be used. Restored to image format.

[0067] In another alternative implementation, the emotion threshold This is the proportional threshold. At this point, the multiple visual positions of the first visual token are sorted from largest to smallest according to their emotional attention scores, and the remaining positions in this sort are then... Each visual location (i.e., an emotion-independent region) is set as a mask, and the top of the sorted areas are... With all visual positions remaining unchanged, the first visual token containing the mask is the first emotion perturbation graph. .

[0068] The first emotional disturbance graph With First Vision Token The first perturbation can be obtained by combining the elements. .

[0069] Step 102 above aims to guide the multimodal large model to actively focus on the regions in the image most closely related to emotion understanding in the emotion recognition task, thereby quantifying the model's dependence on emotional visual elements and providing a basis for subsequent perturbation sensitivity analysis.

[0070] In step 103, the first training sample and the first emotion perturbation map are input into the first multimodal large language model to obtain the first training sample probability corresponding to the first training sample and the first emotion perturbation probability corresponding to the first emotion perturbation map.

[0071] Inputting the first training sample into the first multimodal large language model is equivalent to... The first emotion perturbation map is input into the first multimodal large language model. When inputting the first emotion perturbation map into the first multimodal large language model, the first text token must also be input simultaneously; that is, the first perturbation map is actually input into the first multimodal large language model. The input is fed into the first multimodal large language model.

[0072] When implementing step 103, first respectively... and The input is fed into the first multimodal large language model to obtain its output logits vector. Then, the probability distribution of each emotion category is obtained through the softmax function, thereby obtaining the probability of the first training sample and the probability of the first emotion perturbation. This process is represented by formulas (3) to (4).

[0073] (3) (4) In formulas (3) and (4), As the first multimodal large language model, To use the first training sample The logits vector obtained after inputting into the first multimodal large language model To make the first disturbance The logits vector obtained after inputting into the first multimodal large language model. The probability of the first training sample is the first The probability of each category, is the probability of the k-th category in the first emotional disturbance probability. It is a positive integer. The value range is 1 to , This represents the total number of emotion categories.

[0074] In step 104, a first confidence score is calculated based on the probability of the first training sample, and a first emotion perturbation score is calculated based on the probability of the first emotion perturbation.

[0075] In this embodiment, negative entropy is used to measure the model's certainty towards the original input, i.e., the first confidence score. Based on this, the first confidence score is calculated using formula (5).

[0076] (5) In formula (5), The first confidence score is given. The meanings of the other parameters in formula (5) are the same as those in formulas (2) and (3), and are omitted here.

[0077] A higher first confidence score indicates that the model is more certain about the first training sample (the emotional signal is clearer); a lower first confidence score indicates that the model is more uncertain about the first training sample.

[0078] In this embodiment, the KL divergence is used to assess the degree of change in the distribution of the model output before and after the perturbation, and to measure whether the model depends on the emotion-related visual region, i.e., the first emotion perturbation score. Based on this, the first emotion perturbation score is calculated using formula (6).

[0079] (6) In formula (6), The first emotional disturbance score is given. The meanings of the other parameters in formula (6) are the same as those in formulas (2) and (3), and are omitted here.

[0080] The higher the value of the first emotion perturbation score, the greater the influence of perturbation on the model's prediction. In other words, the more the model relies on the emotional visual cues of the first training sample, indicating that the first training sample is more valuable for training.

[0081] In step 105, the emotional difficulty score of the first training sample is determined based on the first confidence score and the first emotional perturbation score.

[0082] Optionally, step 105 is expressed by formula (7).

[0083] (7) In formula (7), The sentiment difficulty score for the first training sample. The weighting coefficients are preset. The meanings of the other parameters in formula (7) are the same as those in formulas (5) and (6), and are omitted here.

[0084] Among them when At times, more attention is paid to the uncertainty of the model; when At this time, the importance of the visual area for emotions is emphasized. The default value is... In other words, the two are weighted equally. This emotional difficulty score can be regarded as a measure of the "emotional challenge" of the sample. High-scoring samples are considered to be more difficult, contain more information, and are more crucial to improving the model's emotional cognitive ability.

[0085] By performing steps 102 to 105 above on each training sample in the initial training set, the sentiment difficulty score of each training sample can be obtained, and then step 106 can be performed.

[0086] In step 106, an emotional training subset is selected from the initial sample set according to the emotional difficulty score of each training sample in the initial sample set.

[0087] The emotion training subset is used to fine-tune the first multimodal large language model.

[0088] Optionally, step 106 includes steps f to h as follows.

[0089] Step f: Sort the training samples in the initial sample set in descending order according to the sentiment difficulty score of each training sample in the initial sample set to obtain the first sort.

[0090] Step g: Divide the first sort into a difficult sample set, a medium sample set, and an easy sample set.

[0091] Difficult sample set The set of training samples that constitute the top q% in the first ranking, a simple sample set. The set of training samples that constitute the last q% of the first sorted samples, the medium sample set. It is the set of training samples in the first ranking, excluding the hard and easy sample sets.

[0092] The difficult sample set has the most complex emotional expression and the most uncertain model; the simple sample set has clear emotional expression and high model confidence; the medium sample set is between the difficult and simple sample sets, with a moderate degree of emotional complexity or some ambiguity.

[0093] For example, q is 20, meaning the first 20% is the difficult sample set, the last 20% is the easy sample set, and the remaining 60% is the intermediate sample set. However, this is not a limitation, and the value of q can be changed according to the needs in practical applications.

[0094] Step h involves filtering from the difficult sample set, the medium sample set, and the easy sample set to obtain the emotion training subset.

[0095] Step h is essentially a difficulty-aware hierarchical sampling mechanism. This hierarchical sampling mechanism aims to select a subset of sentiment training data from the initial training set according to needs, given a limited training budget (e.g., only 10%-30% of the original data can be used).

[0096] Different objectives require different selection methods when selecting a subset of emotion training data. For example, if the goal is to maximize diversity, then it is necessary to uniformly sample from the difficult sample set, the medium sample set, and the easy sample set to obtain the emotion training subset.

[0097] If it is necessary to enhance the understanding of difficult samples by the first multimodal large language model, then when selecting the sentiment training subset, more training samples can be sampled from the difficult sample set, and fewer samples (or even no samples) can be sampled from the medium and simple sample sets to obtain the corresponding sentiment training subset.

[0098] For example, when maximizing diversity is the objective, the selection method for the emotion training subset includes: selecting samples from the difficult sample set to maximize diversity. Training samples were selected from a simple sample set. training samples, according to The proportion of samples selected is randomly chosen from a medium sample set, and the set of all selected training samples constitutes the emotion training subset. in, This is a preset proportional coefficient. , These are the training budget parameters set by the user. , It is the total number of training samples in the initial sample set.

[0099] The stratified sampling process described above ensures a balance between "challenge, ease of learning, and typicality" in the emotion training subset, helping the model to comprehensively understand the diversity and boundary conditions of emotions. This stratified sampling strategy covers emotion samples of varying difficulty, effectively improving the model's training efficiency and generalization ability in low-resource or rapid deployment scenarios.

[0100] In this embodiment, by considering confidence scores and emotional perturbation scores in the calculation of emotional difficulty scores, the ambiguity of sample emotional expression and model uncertainty can be estimated in real time without model gradients or backpropagation information. The selected emotional training subset can contain the most inspiring and emotionally challenging samples for the model, thereby reducing training costs while preserving diversity and information density. By masking emotionally irrelevant regions in the emotional perturbation map, this invention can quantitatively analyze the model's dependence on emotional regions, achieving interpretable modeling of emotional cognitive behavior. The multimodal large-model emotional understanding enhancement method in this invention does not depend on a specific model structure, has good transferability, and can be directly integrated into existing mainstream MLLM fine-tuning processes such as LLaVA and InternVL, making it suitable for practical applications such as mental health and social emotion computing.

[0101] Figure 3 A flowchart illustrating a multimodal large language model training method provided by an exemplary embodiment of the present invention is shown. This method can be executed by a computer device. Figure 4 This is a flowchart of the multimodal large language model training algorithm.

[0102] See Figure 3-4 The method includes: In step 301, the emotion training subset is obtained.

[0103] The emotional training subset is obtained using the methods described in steps 101 to 106.

[0104] In step 302, the first multimodal large language model is fine-tuned using an emotion training subset.

[0105] In practical applications, there is no need to modify the main structure of the first multimodal large language model; simply insert parameter-efficient and trainable structures (such as the LoRA module) into specific modules.

[0106] For any sample in the emotion training subset The first multimodal large language model The output can be represented as .in, This represents the main parameters of the first multimodal large language model.

[0107] In the process of fine-tuning the first multimodal large language model using an emotion training subset, training only updates the parameters of the LoRA interpolation layer. Maintain the main parameters of the first multimodal large language model Remain unchanged, forming a fine-tuning model . This indicates that only a very small number of parameters are fine-tuned and optimized.

[0108] The training task is defined as a multi-class classification problem, whereby the model predicts a probability distribution based on the hidden layer representation generated from the input. , where K is the predefined number of sentiment categories. The goal is to minimize the prediction distribution. With true emotion tags The difference is analyzed using cross-entropy loss as the objective function. The calculation method for cross-entropy loss is widely available in related technologies and will not be detailed here.

[0109] The loss function is computed individually on each training sample and optimized only through backpropagation. .

[0110] The Adam optimizer and mini-batch iterations are used during training, and parameter adaptation can be completed by traversing the subset once in each training round.

[0111] In step 303, the performance of the fine-tuned first multimodal large language model is verified.

[0112] In step 303, all parameters are frozen, including the inserted LoRA module, and only the trained model is used to perform forward prediction on the test samples. After inputting the test samples, the model outputs the probability distribution of the sentiment category for each training sample, and selects the category corresponding to the maximum value as the final predicted sentiment category.

[0113] Because the training data undergoes the emotion difficulty modeling and hierarchical sampling proposed in this invention, the emotion training subset contains a rich set of samples with ambiguous emotions, those easily confused by the model, or those structurally challenging. During training, the model can gain a stronger ability to construct discriminative boundaries from these cognitively valuable samples, thereby improving its ability to understand and generate subtle, implicit, and combinatorial emotions.

[0114] For example, the deep learning framework used is PyTorch, with CUDA support version 11.8. The experimental environment is based on an NVIDIA A100 graphics card and a high-performance server, capable of training and inference large-scale multimodal models. Experiments were conducted on two real-world emotion recognition datasets: EmoSet and WebEmo. The EmoSet dataset contains eight emotion labels, covering typical categories such as happiness, sadness, anger, surprise, fear, disgust, neutrality, and warmth. Its diverse sample sources make it suitable for evaluating the understanding of subjective emotional expressions by large multimodal models. The WebEmo dataset contains seven main emotion categories, composed of real-world image-text pairs. Emotion annotations are based on images and their associated text information, making it suitable for testing the model's generalization ability to complex emotional contexts. In the experiments, 10,000 image-text samples were selected from each dataset as candidate inputs, and a training subset was constructed by sampling according to a set ratio. The text portion of the training samples was constructed using the prompt template proposed by the LLaVA model to ensure consistent emotion recognition input to the model. The final results use accuracy (ACC) as the primary evaluation metric.

[0115] This invention is validated on two current mainstream large-scale multimodal language models: LLaVA (7B parameters) and InternVL (8B parameters). Both use officially pre-trained weights as a base, and then incorporate the LoRA mechanism for lightweight fine-tuning. The specific experimental settings are as follows: the LoRA rank is set to 32; all experiments are trained for only one epoch; and the emotion threshold is used in the emotion difficulty scoring parameters. The proportional threshold and Weighting parameters for emotional difficulty score For the proportion of sample subsets The experiment was conducted at 10%, 20%, and 30% scales to compare the impact of different training set sizes on performance; all experiments were run in parallel on eight NVIDIA A100 GPUs (40GB VRAM).

[0116] During the training phase, a perturbation version is first constructed for each image-text sample by obscuring emotion-independent image regions. Then, the entropy and KL divergence of the predicted output are calculated to construct an emotion difficulty scoring function. Candidate samples are sorted according to their scores, and then sampled from high, medium, and low difficulty ranges at predetermined proportions (e.g., 10%, 20%, 30%) to construct a broad-coverage training subset. The model is fine-tuned using a cross-entropy loss function; only the LoRA module parameters are updated during training, while the main model parameters remain unchanged.

[0117] During the testing phase, no gradient updates are performed. The trained, fine-tuned model is used to predict the test set, the model output category is recorded, and the accuracy (ACC) metric is calculated.

[0118] To verify the effectiveness of this invention, it is compared with training-free data screening methods in related technologies. Existing methods mainly include: (1) Random: Randomly sample the same number of training samples from the candidate samples; (2) DLC: Jiang W, Liu Z, Xie Z, et al. 2025. Exploring LearningComplexity for Efficient Downstream Dataset Pruning. in ICLR. Table 1 shows the accuracy of this invention (Heart Prune) on the EmoSet dataset, compared to Random and DLC.

[0119] Table 1: Accuracy differences between Heart Prune, Random, and DLC on the EmoSet dataset.

[0120]

[0121] Table 2 shows the accuracy of this invention (Heart Prune) on the WebEmo dataset, compared to Random and DLC.

[0122] Table 2: Accuracy differences between Heart Prune, Random, and DLC on the WebEmo dataset.

[0123]

[0124] As shown in Tables 1 and 2, compared with training-free data selection methods in related technologies, the accuracy of this invention on the EmoSet and WebEmo datasets is significantly improved. Especially under conditions of smaller training samples, the method of this invention demonstrates higher training efficiency and better emotion recognition capabilities. Particularly on the WebEmo dataset, the HeartPrune method of this invention improves accuracy by approximately 5% compared to traditional methods, proving its effectiveness in complex emotional scenarios. These improvements are mainly attributed to the emotion difficulty scoring mechanism proposed in this invention, which effectively selects representative and challenging samples, avoiding the impact of invalid samples on the training process. Simultaneously, the lightweight fine-tuning scheme of the LoRA mechanism enables this invention to achieve efficient and high-quality training with lower computational resources. In summary, this invention not only improves the performance of emotion understanding tasks but also demonstrates its practical application potential in real-world scenarios.

[0125] The following are device embodiments of this application. For details not described in detail in the device embodiments, please refer to the above method embodiments.

[0126] Figure 5 A schematic diagram of the structure of a multimodal large-model emotion understanding enhancement device provided in an exemplary embodiment of the present invention is shown. See also Figure 5 The device includes: a sample acquisition module 501, a perturbation map acquisition module 502, a probability calculation module 503, a score calculation module 504, an emotion difficulty score determination module 505, and a screening module 506.

[0127] The sample acquisition module 501 is used to acquire an initial sample set, which contains multiple training samples, each of which contains an image, text, and emotion label; The perturbation map acquisition module 502 is used to input the first training sample into the first multimodal large language model, and obtain the first emotion perturbation map corresponding to the first training sample based on the cross-attention operation result of the image and text of the first training sample. The emotion-independent region in the emotion perturbation map is set as a mask. The probability calculation module 503 is used to input the first training sample and the first emotion perturbation map into the first multimodal large language model to obtain the first training sample probability corresponding to the first training sample and the first emotion perturbation probability corresponding to the first emotion perturbation map. The scoring calculation module 504 is used to calculate a first confidence score based on the probability of the first training sample and to calculate a first emotion perturbation score based on the probability of the first emotion perturbation. The sentiment difficulty score determination module 505 is used to determine the sentiment difficulty score of the first training sample based on the first confidence score and the first sentiment perturbation score; The filtering module 506 is used to select a subset of sentiment training data from the initial sample set according to the sentiment difficulty score of each training sample in the initial sample set. The sentiment training subset is used to fine-tune the first multimodal large language model.

[0128] Optionally, the perturbation map acquisition module 502 is further configured to: acquire a first visual token based on the image of the first training sample; acquire a first text token based on the text of the first training sample; input the first visual token and the first text token into a first multimodal large language model for cross-attention operation to obtain a first attention matrix; calculate the emotional attention score of each visual position of the first visual token on the first attention matrix according to a preset set of emotional cue words, wherein the set of emotional cue words is used to indicate the set of emotionally related tokens in the text token; and obtain an emotion-independent region according to a preset emotion threshold and the emotional attention score of each visual position of the first visual token, and set the emotion-independent region as a mask in the first visual token to obtain a first emotion perturbation map.

[0129] Optionally, in the scoring calculation module 504, a first confidence score is calculated based on the probability of the first training sample, including:

[0130] in, The first confidence level score is given. As the first training sample, The probability of the first training sample is the first The probability of each category, It is a positive integer. The value range is 1 to , Total number of categories; The first emotion perturbation score is calculated based on the probability of the first emotion perturbation, including:

[0131] in, The first emotional disturbance score is given. is the probability of the k-th category in the first emotional disturbance probability.

[0132] Optionally, the sentiment difficulty score determination module 505 is also used to determine the sentiment difficulty score of the first training sample using the following formula:

[0133] in, The sentiment difficulty score for the first training sample. These are the preset weighting coefficients.

[0134] Optionally, the filtering module 506 is further configured to sort the training samples in the initial sample set in descending order according to the sentiment difficulty score of each training sample in the initial sample set to obtain a first sort; divide the first sort into a difficult sample set, a medium sample set, and an easy sample set, wherein the difficult sample set is the set of the top q% of training samples in the first sort, the easy sample set is the set of the bottom q% of training samples in the first sort, and the medium sample set is the set of training samples in the first sort excluding the difficult and easy sample sets; and filter from the difficult sample set, the medium sample set, and the easy sample set respectively to obtain a sentiment training subset.

[0135] Optionally, the screening module 506 is also used to screen from the difficult sample set with the goal of maximizing diversity. Training samples were selected from a simple sample set. training samples, according to The proportion of samples selected is randomly chosen from a medium-sized sample set, and the set of all selected training samples constitutes the sentiment training subset; among them, This is a preset proportional coefficient. , These are the training budget parameters set by the user. , It is the total number of training samples in the initial sample set.

[0136] It should be noted that the multimodal large-model sentiment understanding enhancement device provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the multimodal large-model sentiment understanding enhancement device and the multimodal large-model sentiment understanding enhancement method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0137] The module division in this embodiment of the invention is illustrative and represents only one logical functional division. In actual implementation, other division methods are possible. Furthermore, the functional modules in each embodiment of the invention can be integrated into a single processor, exist as separate physical entities, or consist of two or more modules integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.

[0138] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a terminal device (which may be a personal computer, mobile phone, or communication device, etc.) or processor to execute all or part of the steps of the method of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0139] Figure 6 This is a schematic diagram of the structure of a computer device provided in an exemplary embodiment of the present invention. For example... Figure 6 As shown, the computer device 600 includes a processor 601 and a memory 602.

[0140] Processor 601 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 601 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 601 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 601 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 601 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0141] The memory 602 may include one or more computer-readable storage media, which may be non-transitory. The memory 602 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 602 are used to store at least one instruction, which is executed by the processor 601 to implement the multimodal large-model sentiment understanding enhancement method provided in this embodiment of the invention.

[0142] Those skilled in the art will understand that Figure 6 The structure shown does not constitute a limitation on the computer device 600, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0143] This invention also provides a non-transitory computer-readable storage medium, which, when the instructions in the storage medium are executed by the processor of a computer device, enables the computer device to execute the multimodal large model sentiment understanding enhancement method provided in this invention.

[0144] This invention also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the multimodal large model sentiment understanding enhancement method provided in this invention.

[0145] The above description is merely an optional embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for enhancing multimodal large-scale emotion understanding, characterized in that, The method includes: Obtain an initial sample set, which contains multiple training samples, each of which contains an image, text, and a sentiment label; The first training sample is input into the first multimodal large language model. The first emotion perturbation map corresponding to the first training sample is obtained based on the cross-attention operation result of the image and text of the first training sample. The emotion-independent region in the emotion perturbation map is set as a mask. The first training sample and the first emotion perturbation map are input into the first multimodal large language model to obtain the first training sample probability corresponding to the first training sample and the first emotion perturbation probability corresponding to the first emotion perturbation map. A first confidence score is calculated based on the probability of the first training sample, and a first emotion perturbation score is calculated based on the probability of the first emotion perturbation. The emotional difficulty score of the first training sample is determined based on the first confidence score and the first emotional perturbation score. Based on the sentiment difficulty score of each training sample in the initial sample set, a sentiment training subset is selected from the initial sample set. This sentiment training subset is used to fine-tune the first multimodal large language model.

2. The method according to claim 1, characterized in that, The step of inputting the first training sample into the first multimodal large language model and obtaining the first sentiment perturbation map corresponding to the first training sample based on the cross-attention operation result of the image and text of the first training sample includes: Based on the images of the first training samples, obtain the first visual token; Based on the text of the first training sample, obtain the first text token; The first visual token and the first text token are input into the first multimodal large language model for cross-attention operation to obtain the first attention matrix; Based on a preset set of emotion cue words, the emotional attention score of each visual position of the first visual token is calculated on the first attention matrix. The set of emotion cue words is used to indicate the set of emotionally related tokens in the text token. Based on the preset emotion threshold and the emotion attention score of each visual position of the first visual token, the emotion-independent region is obtained, and the emotion-independent region is set as a mask in the first visual token to obtain the first emotion perturbation map.

3. The method according to claim 1, characterized in that, The calculation of the first confidence score based on the probability of the first training sample includes: in, The first confidence level score is given. For the first training sample, The probability of the first training sample is the first The probability of each category, It is a positive integer. The value range is 1 to , Total number of categories; The calculation of the first emotional disturbance score based on the first emotional disturbance probability includes: in, The first emotional disturbance score is given. It represents the probability of the k-th category in the first emotional disturbance probability.

4. The method according to claim 3, characterized in that, The step of determining the sentiment difficulty score of the first training sample based on the first confidence score and the first sentiment perturbation score includes: in, The emotion difficulty score for the first training sample. These are the preset weighting coefficients.

5. The method according to any one of claims 1 to 4, characterized in that, The step of selecting an emotional training subset from the initial sample set according to the emotional difficulty score of each training sample in the initial sample set includes: Based on the sentiment difficulty score of each training sample in the initial sample set, the training samples in the initial sample set are sorted in descending order to obtain the first sort; The first sorting is divided into a difficult sample set, a medium sample set, and an easy sample set. The difficult sample set is the set of the top q% of training samples in the first sorting. The easy sample set is the set of the bottom q% of training samples in the first sorting. The medium sample set is the set of training samples in the first sorting other than the difficult sample set and the easy sample set. The emotion training subset is obtained by filtering from the difficult sample set, the medium sample set, and the simple sample set respectively.

6. The method according to claim 5, characterized in that, The process of selecting the emotion training subset from the difficult sample set, the medium sample set, and the easy sample set, respectively, includes: With the goal of maximizing diversity, before screening from the difficult sample set Training samples, selected from the simple sample set training samples, according to The proportion is randomly selected from the medium sample set, and the set of all selected training samples constitutes the emotion training subset; in, This is a preset proportional coefficient. , These are the training budget parameters set by the user. , It is the total number of training samples in the initial sample set.

7. A multimodal large-scale emotion understanding enhancement device, characterized in that, The device includes: The sample acquisition module is used to acquire an initial sample set, which contains multiple training samples, each of which contains an image, text, and emotion label; The perturbation map acquisition module is used to input the first training sample into the first multimodal large language model, and obtain the first emotion perturbation map corresponding to the first training sample based on the cross-attention operation result of the image and text of the first training sample. The emotion-independent regions in the emotion perturbation map are set as masks. The probability calculation module is used to input the first training sample and the first emotion perturbation map into the first multimodal large language model to obtain the first training sample probability corresponding to the first training sample and the first emotion perturbation probability corresponding to the first emotion perturbation map. The scoring calculation module is used to calculate a first confidence score based on the probability of the first training sample and to calculate a first emotion perturbation score based on the probability of the first emotion perturbation. The emotion difficulty score determination module is used to determine the emotion difficulty score of the first training sample based on the first confidence score and the first emotion perturbation score. The filtering module is used to select an emotional training subset from the initial sample set according to the emotional difficulty score of each training sample in the initial sample set. The emotional training subset is used to fine-tune the first multimodal large language model.

8. A computer device, characterized in that, The computer device includes a memory and a processor, wherein the memory stores at least one computer program, which is loaded and executed by the processor to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to implement the method according to any one of claims 1 to 6.

10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the method described in any one of claims 1 to 6.