Multi-modal sentiment reasoning method and system based on fast-slow thinking cooperation
Patent Information
- Application Number
- CN202610914903.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2046-06-24
AI Technical Summary
现有强化学习方法通常直接优化基于F1的单一奖励,将召回率和精确率压缩为一个标量目标,容易造成二者相互干扰;同时,现有训练目标也没有显式利用快思考与慢思考之间的互补关系,难以同时保留快思考的覆盖能力和慢思考的筛选能力
1.缓解慢思考过度保守导致的情绪漏检问题。现有推理式多模态情绪模型在进行慢思考时,虽然能够生成较完整的推理过程,但容易在审慎筛选中压低部分正确情绪的置信度,导致最终答案遗漏真实存在的情绪类别。本发明通过引入快思考的直觉式覆盖能力,使慢思考最终答案能够保留更多可能正确的情绪信息。
Smart Images

Figure CN122433924B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of affective computing technology, specifically to a multimodal affective reasoning method and system based on fast and slow thinking collaboration. Background Technology
[0002] With the development of multimodal large language models, sentiment computing tasks are gradually shifting from traditional fixed-label classification to open-vocabulary multimodal sentiment recognition and multimodal sentiment reasoning. Compared to traditional methods that only output sentiment categories, multimodal sentiment reasoning requires models to utilize visual, audio, and textual cues simultaneously and generate interpretable reasoning processes, thereby improving the understandability and credibility of sentiment judgments.
[0003] However, practical research has shown that explicit reasoning does not necessarily lead to higher accuracy in multimodal emotion recognition. For large-scale multimodal emotion models with reasoning chain output capabilities, slow thinking typically involves a lengthy, deliberate reasoning process before providing the final answer; fast thinking, on the other hand, directly triggers the answer without generating a complete reasoning chain. Experimental analysis indicates that fast thinking tends to provide broader and more confident emotion predictions, thus improving recall; slow thinking, however, tends to conservatively filter out incorrect categories, thus improving precision, but may also suppress the confidence of correct emotion categories, leading to the omission of valid emotions.
[0004] The above phenomenon forms the "thinking paradox" in multimodal emotion reasoning: the reasoning process improves interpretability but does not necessarily improve the final recognition performance. Existing reinforcement learning methods usually directly optimize a single reward based on F1, compressing recall and precision into a single scalar objective, which easily leads to mutual interference between the two; at the same time, existing training objectives do not explicitly utilize the complementary relationship between fast thinking and slow thinking, making it difficult to simultaneously retain the coverage ability of fast thinking and the filtering ability of slow thinking.
[0005] Therefore, there is an urgent need for a new multimodal emotion reasoning optimization method that combines the intuitive coverage advantage of fast thinking with the deliberate selection advantage of slow thinking, so that the model can improve the accuracy, stability and robustness of the final emotion recognition while maintaining the interpretability of reasoning. Summary of the Invention
[0006] To address the aforementioned technical problems, this invention provides a multimodal emotion reasoning method and system based on fast and slow thinking collaboration.
[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: In a first aspect, the present invention provides a multimodal emotion reasoning method based on fast and slow thinking collaboration, comprising: Acquire multimodal emotion samples including visual, audio, and textual data; Based on multimodal emotion samples, a strategy model is constructed using slow thinking and fast thinking modes. The slow thinking mode generates an inference chain and the final emotional answer, while the fast thinking mode only generates the final emotional answer. Based on the emotion wheel, emotion words are mapped to primary emotion categories, and category-level confidence is calculated; The GRPO reinforcement learning framework is adopted to decouple the reward of the emotion recognition task into recall reward and precision reward. After in-group normalization, the two rewards are added together to obtain the dual-objective decoupling advantage. Calculate the category-level confidence difference between slow thinking and fast thinking in the same emotion category, construct positive calibration rewards and negative calibration rewards for correct emotion categories and incorrect emotion categories respectively, and obtain the confidence calibration advantage after normalization; The advantages of dual-objective decoupling, confidence calibration, and format are combined to obtain the total advantage. This advantage is then substituted into the optimization objective of the GRPO reinforcement learning framework to optimize the parameters of the policy model. After training, for new multimodal emotion samples, the policy model directly outputs a slow-thinking result containing the reasoning chain and the final emotion answer.
[0008] In one embodiment, obtaining multimodal emotion samples comprising visual, audio, and textual data specifically includes: Obtain multimodal emotion samples containing visual, audio, and textual information. : ; in, Indicates visual flow, Represents an audio stream. Represents a text stream.
[0009] In one embodiment, the strategy model constructed based on multimodal emotion samples includes a slow-thinking mode and a fast-thinking mode. The slow-thinking mode generates an inference chain and a final emotional answer, while the fast-thinking mode only generates the final emotional answer. Specifically, this includes: Given a multimodal emotion sample The output generated by the strategy model Including the chain of reasoning and the final emotional response: ; in, Represents a chain of reasoning. Indicates the final emotional response; The slow thinking mode employs a format of reasoning before answering, meaning it generates responses simultaneously. and Quick Thinking Mode uses a direct answer format and does not generate reasoning chains. .
[0010] In one embodiment, the mapping of emotion words to primary emotion categories based on the emotion wheel and the calculation of category-level confidence specifically includes: For primary emotion categories ,set up This indicates that the emotion wheel is mapped to A collection of emotional words; for thinking patterns ,definition Category-level confidence : ; These represent slow thinking mode and fast thinking mode, respectively. Indicating in thinking patterns The next strategy model assigns sentiment words The probability of.
[0011] In one embodiment, the decoupling of the reward for the emotion recognition task into recall reward and precision reward, and the summing of these after in-group normalization to obtain the dual-objective decoupling advantage, specifically includes: In terms of rewards, the rewards for emotion recognition tasks are broken down into recall rewards and precision rewards: , ; in, Indicates a recall reward; Indicates precise reward; This is the set of primary sentiment categories predicted by the strategy model. A collection of real primary emotion categories; For a given multimodal emotion sample From the strategy model before the update Medium sampling The candidate output; the... The recall advantage and precision advantage of each candidate output are defined as follows: , ; in, They respectively represent those based on the same The mean and standard deviation of the recall reward for the obtained candidate outputs; They respectively represent those based on the same The mean and standard deviation of the exact reward of the candidate outputs; Adding the recall advantage and the precision advantage yields the dual-objective decoupling advantage. .
[0012] In one embodiment, the calculation of the category-level confidence difference between slow thinking and fast thinking in the same emotion category, constructing positive and negative calibration rewards for correct and incorrect emotion categories respectively, and obtaining the confidence calibration advantage after normalization, specifically includes: For strategy model The candidate outputs Construct the slow-thinking emotional answer generation distribution separately. Distribution of Quick Thinking Emotional Response Generation : ; ; in, for The reasoning chain and emotional answer, Emotional words in the answer indicating emotion This indicates the generation of emotion words. The context of the previous emotional response; Each emotion word in the emotion answer Mapped to primary emotion categories And calculate the category-level confidence difference between the slow thinking mode and the fast thinking mode: ; For slow thinking mode Category-level confidence, For fast thinking mode Category-level confidence; Based on whether the emotion words in the emotion responses correspond to a true emotion category, the emotion words in the responses are divided into a set of correct emotion categories. and collection of incorrect emotion categories : ; ; A collection of real primary emotion categories; Positive calibration reward for: ; Negative calibration reward for: ; To each and Perform within-group normalization to obtain the normalized positive calibration reward. and normalized negative calibration reward And form a confidence calibration advantage. : .
[0013] In one embodiment, the combined advantages of dual-objective decoupling, confidence calibration, and formatting to obtain the overall advantage specifically include: Overall advantages for: ; in, The advantage of decoupling the two objectives For the advantage of confidence calibration, For format advantages; and The weights for confidence calibration advantage and format advantage are respectively. ; Indicates the first The format reward corresponding to each candidate output; and represents the mean and standard deviation of the reward for the candidate output format in the same group, respectively.
[0014] In one embodiment, the optimization objective of the GRPO reinforcement learning framework is... for: ; For strategy model The parameters, The total number of candidate outputs. As the overall advantage, For the clipping function, This is the cutting factor. For KL regularization weights, Let KL divergence be the KL divergence. For reference model; in, Importance sampling ratio: ; For multimodal emotion samples, For the policy model One candidate output; This is the strategy model before the update.
[0015] In one embodiment, the multimodal emotion reasoning model is first supervised and fine-tuned before reinforcement learning training, so that the multimodal emotion reasoning model has basic output format and emotion reasoning ability.
[0016] In a second aspect, the present invention provides a computer system including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method of any embodiment of the first aspect.
[0017] Compared with existing multimodal emotion recognition and multimodal emotion reasoning methods, this invention improves the stability, reliability and practicality of the multimodal emotion reasoning process by co-optimizing slow thinking and fast thinking. This allows the model to retain the advantages of fast thinking in emotion coverage and slow thinking in erroneous emotion screening.
[0018] Specifically, the present invention has the following technical effects: 1. Mitigating the problem of missed emotion detection caused by excessive conservatism in slow thinking. Existing reasoning-based multimodal emotion models, while generating relatively complete reasoning processes during slow thinking, tend to suppress the confidence of some correct emotions during careful screening, leading to the final answer omitting truly existing emotion categories. This invention introduces the intuitive coverage capability of fast thinking, enabling the final answer of slow thinking to retain more potentially correct emotional information.
[0019] 2. Reduce the introduction of erroneous emotions from rapid thinking output. Rapid thinking can quickly capture significant emotional cues in multimodal input, but due to a lack of sufficient reasoning, it is prone to outputting additional emotions without supporting evidence. This invention uses a selective filtering mechanism based on slow thinking to suppress noisy emotions in rapid thinking, making the final emotion prediction more robust.
[0020] 3. Improve the coverage and filtering of sentiment prediction. This invention decouples recall-related objectives from precision-related objectives, constructing reward and advantage signals separately, thus avoiding the problem of mutual interference between recall and precision in a single F1 reward. Therefore, the model can cover more potentially correct sentiment categories while reducing the output of irrelevant or incorrect sentiments.
[0021] 4. Enhance the model's confidence expression of correct emotions. This invention uses fast and slow confidence calibration to ensure that the final answer from slow thinking absorbs the high confidence of fast thinking for the correct emotion category, while retaining the ability of slow thinking to suppress incorrect emotion categories. This mechanism helps to improve the model's expression strength of correct emotions and reduces the possibility of incorrect emotions being mistakenly included in the final answer.
[0022] 5. Improve the credibility and stability of multimodal emotion reasoning results. This invention not only optimizes the final emotion label, but also constrains the model's reasoning process, emotion category coverage, and category confidence distribution, enabling the model to generate more stable, interpretable, and evidence-compliant emotion reasoning results when faced with complex scenarios where visual, audio, and textual information are complementary or inconsistent.
[0023] In summary, this invention, through dual-objective decoupling and fast / slow confidence calibration, enables multimodal emotion reasoning models to more reasonably combine intuitive emotion coverage with deliberate emotion screening, thereby improving the accuracy, interpretability, and reliability of model output in complex emotional scenarios. Attached Figure Description
[0024] Figure 1 This is a flowchart of the method of the present invention; Figure 2 This is a diagram of the overall architecture of the present invention. Detailed Implementation
[0025] A preferred embodiment of the present invention will now be described in detail with reference to the accompanying drawings.
[0026] like Figure 1 As shown, this invention provides a multimodal emotion reasoning method based on fast and slow thinking collaboration, comprising the following steps: S1, acquire multimodal emotion samples including visual, audio and text; S2, based on multimodal emotion samples, constructs a strategy model with slow thinking mode and fast thinking mode. The slow thinking mode generates the reasoning chain and the final emotion answer, while the fast thinking mode only generates the final emotion answer. S3, based on the emotion wheel, maps emotion words to primary emotion categories and calculates category-level confidence; S4 uses the GRPO reinforcement learning framework to decouple the reward of the emotion recognition task into recall reward and precision reward. After normalizing within the group, the two rewards are added together to obtain the dual-objective decoupling advantage. S5, calculate the category-level confidence difference between slow thinking and fast thinking in the same emotion category, construct positive calibration rewards and negative calibration rewards for correct emotion categories and incorrect emotion categories respectively, and obtain the confidence calibration advantage after normalization; S6. Combine the advantages of dual-objective decoupling, confidence calibration, and format to obtain the total advantage. Substitute this advantage into the optimization objective of the GRPO reinforcement learning framework to optimize the parameters of the policy model. S7. After training is complete, for new multimodal emotion samples, the policy model directly outputs the slow thinking result, which includes the reasoning chain and the final emotion answer.
[0027] Figure 2 This invention demonstrates the basic motivation and evaluation results of the proposed fast-slow thinking collaborative method (MER-R1). Figure 2(a) presents multimodal inputs, including video, audio, and text, and compares three reasoning methods: slow thinking, fast thinking, and fast-slow collaboration. Slow thinking allows for stepwise and deliberate reasoning, but it can be overly conservative and miss some valid emotions; fast thinking can quickly trigger broader emotional responses, but it may introduce unsupported, noisy emotions; fast-slow collaboration utilizes fast thinking to provide candidate emotional coverage, which is then filtered by slow thinking, resulting in a more balanced and accurate answer. Figure 2 (b) shows the evaluation trends of different methods on multiple datasets (OVMERD+, MER2023, MER2024, MELD, IEMOCAP, CMUMOSEI, CMUMOSI, SIMS, SIMSv2), demonstrating that fast and slow thinking collaboration can alleviate the performance conflict between fast and slow thinking.
[0028] The present invention will be described in detail below in several parts.
[0029] 1. Construct a multimodal emotion reasoning task and fast / slow thinking output formats.
[0030] Obtain multimodal emotion samples containing visual, audio, and textual information, and represent the input as follows: ; in, Indicates visual flow, Represents an audio stream. Represents a text stream.
[0031] Given multimodal input Strategy Model Generate output The output consists of two parts: the reasoning chain and the final emotional response. ; in, This represents a chain-like reasoning process. This indicates the final emotional response.
[0032] Slow thinking employs a format of reasoning before answering, that is, generating results simultaneously. and Quick Thinking uses a direct answer format and does not generate a complete reasoning chain, which can be represented as: .
[0033] Since both the model output and the standard answer may contain multiple open-ended emotion words, this invention maps emotion words to first-level emotion categories in the emotion wheel. Let the set of first-level emotion categories predicted by the model be denoted as . The set of true primary emotion categories is Then, the ensemble-level recall, precision, and F1 score are defined as follows: ; This paper analyzes the predictive characteristics of fast thinking and slow thinking separately. Fast thinking typically provides broader emotional coverage and higher recall; slow thinking typically provides more focused and conservative predictions and higher precision. Based on this, the present invention transforms the complementary relationship between the two into an optimizable reinforcement learning objective.
[0034] 2. Construct a category-level confidence metric and a speed difference analysis.
[0035] For primary emotion categories ,set up This indicates that the emotion wheel is mapped to categories. A fine-grained set of emotion words. For thinking patterns. Define the categorical log-level confidence level as: ; in, Indicating in thinking patterns The model assigns emotion words The probability of.
[0036] remember This is a set of true primary emotion categories. For thinking patterns The first set of several difficult-to-bear categories. The average confidence scores for the correct and incorrect categories are as follows: .
[0037] The relative confidence interval between the two is defined as follows: ; This metric measures whether the model can effectively distinguish between the correct emotion category and the difficult emotion category.
[0038] Based on the above metrics, it can be seen that fast thinking generally has stronger confidence in correct emotion categories, while slow thinking is better at suppressing incorrect emotion categories. This invention requires the final model to simultaneously meet two objectives: first, at the prediction level, retaining the recall and coverage capabilities of fast thinking while maintaining the precision filtering capabilities of slow thinking; second, at the confidence level, retaining the high confidence of fast thinking in correct categories while maintaining the ability of slow thinking to suppress incorrect categories.
[0039] 3. Basic reinforcement learning training based on GRPO.
[0040] During the training phase, in a preferred embodiment, the multimodal emotion reasoning model can first be supervised and fine-tuned to acquire basic output format and emotion reasoning capabilities before entering the GRPO-style reinforcement learning phase; alternatively, it can directly enter the GRPO-style reinforcement learning phase. Given input... From the old strategy Medium sampling Candidate outputs: ; And calculate the reward for each candidate output. .
[0041] Normalize the rewards to a relative advantage: ; in, and represents the mean and standard deviation of the output rewards for candidates in the same group, respectively.
[0042] The basic optimization objective of GRPO is: ; in, Importance sampling ratio: ; This is the cutting factor. For KL regularization weights, This is a reference model.
[0043] Basic rewards typically consist of the Emotion Wheel F1 reward and format rewards: ; in, This represents the F1 reward based on matching the set of emotion wheels. This is used to encourage the model to generate inference chains and answer structures that meet the requirements. However, since F1 compresses recall and precision into a single value, the two objectives may interfere with each other during training. Therefore, this invention further introduces a dual-objective decoupling mechanism.
[0044] 4. Construct a dual-objective decoupling mechanism.
[0045] In terms of rewards, the rewards for emotion recognition tasks are broken down into recall rewards and precision rewards: ; in, This will enable the model to cover more real-world emotion categories. Suppress the model from generating irrelevant or incorrect sentiment categories.
[0046] At the dominance function level, instead of merging the recall reward and the precision reward first, they are normalized separately within each group. For the first... For each of the candidate outputs, the recall advantage and precision advantage are defined as follows: ; in, represents the mean and standard deviation of the recall reward for candidate outputs in the same group, respectively; represents the mean and standard deviation of the exact reward for each candidate output in the same group, respectively.
[0047] Adding the two normalization advantages, we obtain the dual-objective decoupling advantage: ; This design avoids the high variance target dominance problem caused by directly merging rewards and then normalizing, so that recall and precision can be clearly preserved during the optimization process.
[0048] To illustrate the necessity of this design, let the within-group mean and standard deviation of the recall reward and the precision reward be respectively... and And define the normalized within-group variance ratio: ; If we directly use the F1 advantage Its correlation with recall rewards and precision rewards will vary. Bias towards high variance targets: ; By adopting the advantages of dual-objective decoupling, a more balanced correlation can be obtained: .
[0049] 5. Construct a fast and slow confidence calibration mechanism.
[0050] For the candidate outputs Construct the slow-thinking answer generation distribution and the fast-thinking answer generation distribution respectively: ; in, Words that express emotion in the answer, This indicates the generation of emotion words. The context of the previous answer.
[0051] For each emotion word in the output answer Map it to the primary emotion category And calculate the category-level confidence difference between slow thinking and fast thinking: ; Based on whether the emotion words correspond to the true categories, the emotion words in the answers are divided into a correct set and an incorrect set: ; ; For the correct emotion category, slow thinking is encouraged to maintain or exceed the confidence level of fast thinking; for the incorrect emotion category, slow thinking is encouraged to maintain its conservative inhibitory capacity. The corresponding calibration reward is defined as follows: ; ; To each and Perform within-group normalization to obtain and This allows for faster and slower confidence level calibration advantages. ; Ultimately, the advantages of dual-objective decoupling, confidence calibration, and formatting are combined into a total advantage: ; in, and These are the weights for confidence calibration advantage and format advantage, respectively. ; Indicates the first The format reward corresponding to each candidate output; and These represent the mean and standard deviation of the format rewards for candidate outputs within the same group, respectively. The format rewards are used to encourage the policy model to generate inference chains and final emotional responses according to a preset output structure; a higher format reward is given when a candidate output contains both an inference chain and a final emotional response and meets the preset format requirements; otherwise, a lower format reward or zero reward is given.
[0052] Overall Advantage By substituting the GRPO optimization objective, the reinforcement learning training of this invention (MER-R1) can be completed.
[0053] After training, during the inference phase, only new multimodal emotion samples need to be input, and the model can output a result containing the inference chain and the final emotion answer. Because the training phase has explicitly coordinated the recall and coverage capabilities of fast thinking and the precision filtering capabilities of slow thinking, the final slow thinking answer can more fully and stably utilize multimodal emotion cues.
[0054] This invention proposes a fast-slow thinking collaborative reinforcement learning framework for multimodal sentiment reasoning, which enables the model to simultaneously absorb the recall advantage of fast thinking and the precision advantage of slow thinking.
[0055] This invention proposes a dual-objective decoupling mechanism, which uses recall and precision in emotion recognition tasks as optimization signals respectively, and maintains their independence at the reward level and the advantage function level, thus avoiding target interference caused by a single F1 reward.
[0056] This invention proposes a fast-slow confidence calibration mechanism, which compares the confidence differences between slow thinking and fast thinking in a category-level emotion space, and applies calibration constraints in opposite directions to correct and incorrect emotion categories, thereby enhancing the confidence of correct categories and suppressing false categories.
[0057] This invention utilizes an emotion wheel to map open-vocabulary emotion words to first-level emotion categories, supports set-level evaluation of free-form emotion outputs, and makes the method applicable to open-vocabulary multimodal emotion recognition tasks.
[0058] Experimental results show that this method achieves better performance in the comprehensive evaluation of multimodal emotion recognition and multimodal emotion reasoning, and can make the final slow-thinking answer truly superior to the fast-thinking answer, thereby alleviating the thinking paradox in multimodal emotion reasoning.
[0059] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0060] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple steps or stages, which are not necessarily completed at the same time, but may be executed at different times, and the execution order of these steps or stages is not necessarily sequential, but may be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0061] In one embodiment, the present invention provides a computer system, which may be a server. The computer system includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores data used in the methods described above. The network interface communicates with external terminals via a network connection. The computer program is executed by the processor to implement the methods described above.
[0062] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0063] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention, and no reference numerals in the claims should be construed as limiting the scope of the claims.
[0064] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A multimodal affective reasoning method based on fast and slow thinking collaboration, characterized in that, include: Acquire multimodal emotion samples including visual, audio, and textual data; Based on multimodal emotion samples, a strategy model is constructed using slow thinking and fast thinking modes. The slow thinking mode generates an inference chain and a final emotional answer, while the fast thinking mode only generates the final emotional answer. Based on the emotion wheel, emotion words are mapped to primary emotion categories, and category-level confidence is calculated; The GRPO reinforcement learning framework is adopted to decouple the reward of the emotion recognition task into recall reward and precision reward. After in-group normalization, the two rewards are added together to obtain the dual-objective decoupling advantage. Calculate the category-level confidence difference between slow thinking and fast thinking in the same emotion category, construct positive calibration rewards and negative calibration rewards for correct emotion categories and incorrect emotion categories respectively, and obtain the confidence calibration advantage after normalization; The advantages of dual-objective decoupling, confidence calibration, and format are combined to obtain the total advantage. This advantage is then substituted into the optimization objective of the GRPO reinforcement learning framework to optimize the parameters of the policy model. After training, for new multimodal emotion samples, the policy model directly outputs a slow-thinking result containing the reasoning chain and the final emotion answer.
2. The multimodal affective reasoning method based on fast and slow thinking collaboration according to claim 1, characterized in that, The acquisition of multimodal emotion samples, including visual, audio, and textual data, specifically includes: Obtain multimodal emotion samples containing visual, audio, and textual information. : ; in, Indicates visual flow, Represents an audio stream. Represents a text stream.
3. The multimodal affective reasoning method based on fast and slow thinking collaboration according to claim 1, characterized in that, The strategy model constructed based on multimodal emotion samples includes a slow-thinking mode and a fast-thinking mode. The slow-thinking mode generates the reasoning chain and the final emotional answer, while the fast-thinking mode only generates the final emotional answer. Specifically, it includes: Given a multimodal emotion sample The output generated by the strategy model Includes the chain of reasoning and the final emotional response: ; in, Represents a chain of reasoning. Indicates the final emotional response; The slow thinking mode employs a format of reasoning before answering, meaning it generates responses simultaneously. and Quick Thinking Mode uses a direct answer format and does not generate reasoning chains. 。 4. The multimodal affective reasoning method based on fast and slow thinking collaboration according to claim 1, characterized in that, The process of mapping emotion words to primary emotion categories based on an emotion wheel and calculating category-level confidence specifically includes: For primary emotion categories ,set up This indicates that the emotion wheel is mapped to A collection of emotional words; for thinking patterns ,definition Category-level confidence : ; These represent slow thinking mode and fast thinking mode, respectively. Indicating in thinking patterns The next strategy model assigns sentiment words The probability of.
5. A multimodal affective reasoning method based on fast and slow thinking collaboration as described in claim 1, characterized in that, The decoupling of the reward for the emotion recognition task into recall reward and precision reward, which are then summed after in-group normalization to obtain the dual-objective decoupling advantage, specifically includes: In terms of rewards, the rewards for emotion recognition tasks are broken down into recall rewards and precision rewards: , ; in, Indicates a recall reward; Indicates precise reward; This is the set of primary sentiment categories predicted by the strategy model. A collection of real primary emotion categories; For a given multimodal emotion sample From the strategy model before the update Medium sampling The nth candidate output; The recall advantage and precision advantage of each candidate output are defined as follows: , ; in, They respectively represent those based on the same The mean and standard deviation of the recall reward for the obtained candidate outputs; They respectively represent those based on the same The mean and standard deviation of the exact reward of the candidate outputs; Adding the recall advantage and the precision advantage yields the dual-objective decoupling advantage. .
6. The multimodal affective reasoning method based on fast and slow thinking collaboration according to claim 1, characterized in that, The calculation of the category-level confidence difference between slow thinking and fast thinking in the same emotion category involves constructing positive and negative calibration rewards for correct and incorrect emotion categories, respectively. After normalization, the confidence calibration advantage is obtained, specifically including: For strategy model The Candidate outputs Construct the slow-thinking emotional answer generation distribution separately. Distribution of Quick Thinking Emotional Response Generation : ; ; in, for The reasoning chain and emotional answer, Emotional words in the answer indicating emotion This indicates the generation of emotion words. The context of the previous emotional response; Each emotion word in the emotion answer Mapped to primary emotion category And calculate the category-level confidence difference between the slow thinking mode and the fast thinking mode: ; For slow thinking mode Category-level confidence, For fast thinking mode Category-level confidence; Based on whether the emotion words in the emotion responses correspond to a true emotion category, the emotion words in the responses are divided into a set of correct emotion categories. and collection of incorrect emotion categories : ; ; A collection of real primary emotion categories; Positive calibration reward for: ; Negative calibration reward for: ; To each and Perform within-group normalization to obtain the normalized positive calibration reward. and normalized negative calibration reward And form a confidence calibration advantage. : 。 7. A multimodal affective reasoning method based on fast and slow thinking collaboration as described in claim 1, characterized in that, The combined advantages of dual-objective decoupling, confidence calibration, and format advantages yield a total advantage, which specifically includes: Overall advantages for: ; in, The advantage of decoupling the two objectives For the advantage of confidence calibration, Advantages of the format; and The weights for confidence calibration advantage and format advantage are respectively. ; Indicates the first The format reward corresponding to each candidate output; and represents the mean and standard deviation of the reward for the candidate output format in the same group, respectively.
8. A multimodal affective reasoning method based on fast and slow thinking collaboration as described in claim 1, characterized in that, The optimization objective of the GRPO reinforcement learning framework for: ; For strategy model The parameters, The total number of candidate outputs. As the overall advantage, For the clipping function, This is the cutting factor. For KL regularization weights, Let KL divergence be the KL divergence. For reference model; in, Importance sampling ratio: ; For multimodal emotion samples, For the policy model One candidate output; This is the strategy model before the update.
9. A multimodal affective reasoning method based on fast and slow thinking collaboration according to claim 1, characterized in that, Before reinforcement learning training, the multimodal emotion reasoning model is first supervised and fine-tuned to enable it to have basic output format and emotion reasoning ability.
10. A computer system comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Intelligent agent decision-making method and device, computer readable storage medium and equipment
CN121303373A
Ai-based response system and method based on automated user feedback analysis and dynamic transition between fast-thinking LLM and slow-thinking llm
US20260141251A1