A multi-modal large language model reinforcement learning method capable of verifying emotional reasoning
Patent Information
- Application Number
- CN202610849951.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-08-28
AI Technical Summary
[0004]有鉴于此,为了解决现有基于大语言模型的情感推理方法中仅依赖场景线索,忽略了人为中心的细粒度证据,进而导致推理精度不高的技术问题,本发明提出一种可验证情感推理的多模态大语言模型强化学习方法,该方法包括以下步骤:
[0005]Based on the above scheme, this invention provides a multimodal large language model reinforcement learning method for verifiable sentiment reasoning. By using character-focused rewards and temporally ordered rewards, the model is guided to pay attention to fine-grained cues such as micro-expressions and body movements, constructing a temporally coherent sentiment reasoning chain, effectively alleviating the shortcut reasoning problem. By replacing independent absolute scores with an intra-group comparison scoring mechanism, the score clustering and calibration drift problems are effectively alleviated, significantly improving reward discrimination and training stability. Through a reasoning-evidence closed-loop verification framework, the reasoning chain is compressed into the smallest evidence package and a pixel-level evidence mask is generated, establishing an externally auditable verification interface to achieve traceability and falsification of reasoning results.
Smart Images

Figure CN122655902A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal large language model technology, and in particular to a reinforcement learning method for multimodal large language models that can verify sentiment reasoning. Background Technology
[0002] Large Language Models (LMMs) are natural language processing (NLP) models based on deep learning. They typically employ the Transformer architecture, are pre-trained on massive amounts of text data, and possess powerful language understanding and generation capabilities.
[0003] In recent years, multimodal large language models (MLLM) have made significant progress in tasks such as image understanding and video question answering. Meanwhile, reinforcement learning (RL), especially group relative policy optimization (GRPO), has further enhanced the models' capabilities in complex reasoning tasks. However, in human-centered emotion understanding and interpersonal intention reasoning tasks, existing models generally suffer from shortcut reasoning: relying solely on global scene or background cues while ignoring fine-grained human-centered evidence such as micro-expressions, body language, and interpersonal interaction dynamics. This results in low reasoning accuracy and unreliable evidence, making it difficult to trace and verify with external evidence. Summary of the Invention
[0004] In view of this, in order to address the technical problem that existing sentiment reasoning methods based on large language models rely solely on scene cues and ignore human-centered fine-grained evidence, thus leading to low reasoning accuracy, this invention proposes a multimodal large language model reinforcement learning method for verifiable sentiment reasoning. This method includes the following steps: Multimodal data, including video, audio, and text, is collected and a training sample set is constructed. Using a group-based relative policy optimization framework, a policy model is trained on this sample set to obtain a converged sentiment reasoning model. This technical solution achieves synergistic enhancement of model performance and credibility through three core units: a human-centered fine-grained reward and intra-group comparative scoring mechanism is used during the training phase to guide the model to focus on key cues such as micro-expressions and body language, and to improve the discriminative power of reward signals; the reasoning result and pixel evidence closed-loop verification unit serves as an external audit interface, compressing the thought chain into a structured summary and generating a pixel-level evidence mask after the model completes reasoning, enabling traceability and falsification of the reasoning results. Feeding the test data into the trained sentiment reasoning model outputs structured sentiment reasoning results.
[0005] Based on the above scheme, this invention provides a multimodal large language model reinforcement learning method for verifiable sentiment reasoning. By using character-focused rewards and temporally ordered rewards, the model is guided to pay attention to fine-grained cues such as micro-expressions and body movements, constructing a temporally coherent sentiment reasoning chain, effectively alleviating the shortcut reasoning problem. By replacing independent absolute scores with an intra-group comparison scoring mechanism, the score clustering and calibration drift problems are effectively alleviated, significantly improving reward discrimination and training stability. Through a reasoning-evidence closed-loop verification framework, the reasoning chain is compressed into the smallest evidence package and a pixel-level evidence mask is generated, establishing an externally auditable verification interface to achieve traceability and falsification of reasoning results. Attached Figure Description
[0006] Figure 1 This is a flowchart illustrating the steps of a multimodal large language model reinforcement learning method for verifiable affective reasoning according to the present invention. Figure 2 This is a schematic diagram comparing the present invention with existing methods in shortcut reasoning and human-centered reasoning; Figure 3 This is a comparison chart of the reward distribution between the intra-group comparative score and the independent absolute score of this invention; Figure 4 This is a schematic diagram illustrating how the present invention transforms free-form reasoning chains into structured evidence; Figure 5 This is a visualization of the reasoning verification framework of this invention on a specific video question-and-answer example; Figure 6 This is an overall framework diagram of the method of the present invention; Figure 7 This is a qualitative comparison diagram between the present invention and the prior art (HumanOmniV2) in the micro-expression understanding task; Figure 8 This is a qualitative comparison diagram between the present invention and the prior art (HumanOmniV2) in a temporal reasoning task; Figure 9 This is a schematic diagram illustrating the performance of the invention in a typical failure case 1 (pragmatic reasoning failure); Figure 10 This is a schematic diagram illustrating the performance of the invention in a typical failure case 2 (reasoning-selection disconnect). Detailed Implementation
[0007] In addition to the technical issues mentioned in the background, existing large language model inference methods also have the following problems: Traditional LLM-as-Judge uses independent absolute scoring, which is prone to score clustering and calibration drift: different quality inference paths receive similar rewards, resulting in insufficient reward differentiation, leading to unstable training and inefficient optimization signals; the inference chain is lengthy, loosely structured, and lacks verifiable interfaces: the free text inference chain output by the model is lengthy and has low information density, making it impossible to establish an auditable association with pixel-level evidence from videos, and making it difficult to implement in high-risk scenarios.
[0008] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0009] It should be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0010] It should be understood that the terms "system," "apparatus," "unit," and / or "module" used in this application are a method of distinguishing different components, elements, parts, sections, or assemblies at different levels. However, if other terms can achieve the same purpose, they may be replaced by other expressions.
[0011] Unless the context explicitly indicates an exception, words such as "a," "an," "a kind," and / or "the" do not specifically refer to the singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list; a method or apparatus may also include other steps or elements. An element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, product, or apparatus that includes the element.
[0012] In the description of the embodiments of this application, "a plurality of" refers to two or more. The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.
[0013] Furthermore, flowcharts are used in this application to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed precisely in sequence. Instead, the steps can be processed in reverse order or simultaneously. Additionally, other operations can be added to these processes, or one or more steps can be removed from them.
[0014] Reference Figure 1 This is a flowchart illustrating an optional example of the multimodal large language model reinforcement learning method for verifiable sentiment reasoning proposed in this invention. The method can be applied to computer devices, and this embodiment may include, but is not limited to, the following steps: Step S1: Obtain multimodal data and construct a training dataset; Step S2: Based on the training data and the group relative strategy optimization framework, train the strategy model to obtain the trained sentiment reasoning model; wherein, the training process includes three core units: human-centered fine-grained reward, intra-group comparison scoring, and reasoning-evidence closed-loop verification. Step S3: Input the data to be tested into the trained sentiment reasoning model to generate structured sentiment reasoning output.
[0015] This invention includes a human-centered fine-grained reward shaping mechanism, an intra-group comparison scoring mechanism, and a reasoning-evidence grounding framework. The fine-grained reward shaping includes person-focused rewards and temporal sequence rewards, guiding the model to pay attention to human-centered cues such as micro-expressions, body language, and interpersonal interactions, and organizing temporally consistent emotional evolution reasoning. The intra-group comparison scoring mechanism submits multiple candidate answers under the same input to the evaluation model for relative comparison and ranking, effectively alleviating the score clustering and calibration drift problems caused by independent absolute scoring and improving reward discrimination. The reasoning-evidence grounding framework compresses free-form reasoning chains into minimal evidence packages through a thought compressor, and generates pixel-level evidence masks on video frames using the SAM3 segmentation model, establishing a verifiable closed loop that can be externally audited.
[0016] In some feasible embodiments, the objective function of the Group Relative Policy Optimization (GRPO) framework is: In the formula, This represents the policy optimization objective function. This represents the expectation operation. This indicates the number of candidate inference results for each sample group. This represents the sequence of candidate inference results for the i-th candidate. This represents the length of the token for the i-th candidate result. This represents the loss term corresponding to the t-th token of the i-th candidate result.
[0017] token loss item The calculation formula is: In the formula, This indicates the minimum value operation. This represents the importance sampling ratio of the i-th candidate result and the t-th token. This represents the standardized odds value of the i-th candidate result. This represents the clipping function. This represents the hyperparameters for strategy pruning.
[0018] Importance sampling ratio The calculation formula is: In the formula, Indicates the current update policy distribution. Let represent the t-th token of the i-th candidate result, and x represent the input multimodal query. This represents the sequence preceding the t-th token in the i-th candidate result. This represents the old policy distribution before the update.
[0019] like Figure 2 As shown, existing methods suffer from the shortcut reasoning problem: models often rely on background cues rather than person-centric, fine-grained evidence for sentiment inference, leading to unreliable inference results. To address this issue, this invention designs person-focused rewards and temporally ordered rewards. In some feasible embodiments, the total reward for model training is... The calculation formula is: In the formula, This indicates a reward for compliance with formatting regulations. Represents a reward based on reasoning accuracy. This indicates the weighting coefficient of the reward for focusing on specific individuals. This indicates that the focus is on the reward. This represents the time-ordered reward weighting coefficient. This indicates a sequential reward.
[0020] Character Focus Rewards The calculation formula is: In the formula, This represents the scoring function of the large language model judge. Indicates sample context information, Represents inference text, Indicates key words for focusing on the assessment of individuals. This indicates the range of values for the person-focused reward, which is used to guide the model to focus on fine-grained emotional cues centered on the person.
[0021] Time-ordered rewards The calculation formula is: In the formula, This represents the scoring function of the large language model judge. Indicates sample context information, Represents inference text, Indicates a time-ordered evaluation prompt. This indicates the range of values for the character's focus reward.
[0022] To reduce API call overhead, a joint evaluation scheme is adopted. and Merge into a single prompt word It returns two scores in a single call; it also applies a causal mask to make... and It only applies to the token positions in the context and thought segments to prevent rewards from leaking into the answer segment.
[0023] like Figure 3 As shown, traditional independent absolute scoring methods result in over 73% of scores clustering in the 9-10 range, exhibiting a serious score clustering problem. To address this issue, this invention proposes an intra-group comparison scoring mechanism. In some feasible embodiments, the intra-group comparison score output is as follows: In the formula, This represents the relative score of the i-th candidate. This represents a comparative scoring function. Indicates the candidate inference result, Indicates a contrasting prompt.
[0024] Intra-group comparison scoring is used in the model training and batch candidate inference result scoring stages to replace the traditional single independent scoring mode, improve the scoring clustering defects, optimize reward discrimination, and further improve the quality of model inference generation.
[0025] The formula for the weighted sum of total rewards for a single candidate is: In the formula, This represents the total reward for the i-th candidate result. This represents the weight of the k-th reward component. It represents the k-th reward component of the i-th candidate result.
[0026] The formula for calculating the standardized advantage value is: In the formula, This represents the advantage value of the i-th candidate result. This represents the total reward for the i-th candidate result. This represents the average reward within the group. This represents the total reward for the j-th candidate result.
[0027] To further improve training stability, the standardized advantage value is standardized within groups, and the calculation formula is as follows: In the formula, Indicates the standard deviation of the group's rewards. This represents the numerical smoothing coefficient to avoid the denominator being zero.
[0028] The obtained advantage parameters are used for the aforementioned GRPO loss calculation and model parameter update to optimize the final inference output performance of the model.
[0029] like Figure 4 As shown, this invention transforms free-form reasoning chains into structured summaries using a thought compressor and extracts key entities to generate SAM3 segmentation instructions. In some feasible embodiments, the loss function of the reasoning summarizer is: In the formula, Indicates the summarization loss. Indicates the parameters of the digester. Let 'o' represent the t-th token in the structured summary, and 'o' represent the original reasoning chain. This represents a summary history token.
[0030] This loss function is used to train the inference summarizer, which compresses the thought segments output by the model to generate a structured summary that can be used by the segmentation model.
[0031] like Figure 5 As shown, the formula for generating the SAM3 pixel-level segmentation mask is: In the formula, This represents the cross-frame segmentation mask for the k-th target entity. This represents a pixel-level segmentation model. Indicates the input video frame sequence. This represents the k-th target entity to be located. It represents the total number of target entities; by establishing an auditable association between text reasoning and pixel-level visual regions through segmentation success rate, external verification of sentiment reasoning results is completed.
[0032] The SAM3 segmentation model generates pixel masks based on the structured instructions output by the summary, and uses the masks to achieve external verification of the inference content.
[0033] The overall data flow reference of this invention Figure 6 Furthermore, this embodiment employs a three-stage training pipeline: the first stage is supervised cold-start fine-tuning (SFT), with a learning rate of... The first stage involves training for one epoch for format stabilization; the second stage is GRPO training with only accuracy rewards, at a learning rate of [missing information]. The training phase lasts for 2 epochs; the third phase involves the complete reward suite. GRPO training, with a learning rate of Training is performed for 2 epochs. For each instance, G=4 candidate results are sampled, and hyperparameters are pruned. The maximum output length is 2048 tokens. The reward weight is set to... , All experiments were run on four NVIDIA A100 80GB GPUs, with DeepSpeed ZeRO-2 acceleration and memory optimization.
[0034] Based on the overall process described above, this invention also provides relevant data examples: This embodiment verifies the performance of Embodiment 1 on three multimodal benchmarks covering emotion understanding and interpersonal intention reasoning, including IntentBench, Daily-Omni, and WorldSense.
[0035] This embodiment is trained based on the Qwen2.5-Omni-7B model, and the training corpus includes 24K video-audio samples from Video-R1, Social-IQ 2.0 and EMER.
[0036] In this embodiment, the present invention is compared with existing closed-source commercial models (GPT-4o, GPT-o1, Gemini series, etc.) and open-source multimodal models (Qwen2.5-Omni, MiniCPM-o, HumanOmniV2, etc.) on three benchmarks: IntentBench, Daily-Omni, and WorldSense. The performance comparison results are shown in Tables 1, 2, and 3.
[0037] Table 1. Performance comparison of the present invention model and existing technologies on IntentBench (%) As shown in Table 1, the average accuracy of this invention on IntentBench reaches 71.89%, significantly exceeding the 66.98% of the strongest open-source baseline model, HumanOmniV2 (7B). In particular, on the emotion recognition task, this invention achieves 85.44%, an improvement of 4.66 percentage points over the baseline; on the time-sensitive task (When), this invention achieves 71.43%, an improvement of 14.29 percentage points over the baseline.
[0038] Table 2 Performance comparison of the present invention model and existing technologies on Daily-Omni (%) As shown in Table 2, the average accuracy of this invention on Daily-Omni reached 61.90%, exceeding the 58.47% of HumanOmniV2 (7B). In particular, on a 60-second long task, this invention achieved 58.55%, an improvement of 5.46 percentage points over the baseline.
[0039] Table 3 Performance comparison of the present invention model and existing technologies on WorldSense (%) As can be seen from Table 3, the average accuracy of this invention on WorldSense reached 48.80%, exceeding HumanOmniV2 (7B)'s 47.70%, with the music domain reaching 47.20%, an improvement of 3.9 percentage points from the baseline.
[0040] To verify the focus on rewards for the characters and time-ordered rewards To assess the effectiveness, an ablation experiment was conducted on IntentBench in this embodiment, and the results are shown in Table 4.
[0041] Table 4. Ablation experiment of reward components (IntentBench, %) As can be seen from Table 4, adding The average accuracy improved from 66.98% to 69.78%, and the emotion recognition task accuracy improved to 84.23%; (Added) The average accuracy improved to 69.62%; the full model further improved the average accuracy to 71.89%, with a 14.29 percentage point improvement in time-sensitive tasks (When), validating the complementary effectiveness of the two reward mechanisms.
[0042] To verify the effectiveness of the within-group comparison scoring mechanism, this embodiment compares two strategies: independent absolute scoring and within-group comparison scoring. The results are shown in Table 5.
[0043] Table 5 Comparison of scoring strategies As shown in Table 5, the within-group comparison scoring of this invention increased the overall coefficient of variation from 0.153 to 0.325, and the within-group coefficient of variation from 0.049 to 0.138, indicating that the scoring mechanism of this invention can effectively alleviate the problems of score clustering and calibration drift, and significantly improve the discrimination of rewards. Correspondingly, the average accuracy of IntentBench increased from 69.78% to 71.89%. Figure 3 As shown, independent absolute scoring results in more than 73% of the scores being concentrated in the 9-10 range, while the intra-group comparison scoring of this invention makes the reward distribution more balanced, with the low score range (1-2) accounting for 7.31%, which significantly improves the reward discrimination and training stability.
[0044] To systematically evaluate the quality of the reasoning chain, this embodiment designed a five-dimensional diagnostic framework: D1 Person Anchoring, D2 Interaction Modeling, D3 Fine-Grained Evidence, D4 Traceability, and D5 Temporal Organization. 200 reasoning chains were randomly sampled from IntentBench for evaluation, and the results are shown in Table 6.
[0045] Table 6. Results of Inference Chain Quality Assessment As shown in Table 6, this invention outperforms the baseline in all five dimensions, with the average score increasing from 1.928 to 2.500, a relative improvement of 29.7%. The D5 (temporal organization) gain is the largest, indicating temporally ordered rewards. It played a key role.
[0046] The statistical test results of expert scoring and automatic scoring are as follows: Paired t-test showed t(4) = 0.847, p = 0.444; Pearson correlation coefficient r = 0.994, ICC = 0.987. The above results indicate that the automatic assessment has high reliability.
[0047] To verify the effectiveness of the reasoning-evidence closed-loop framework proposed in this invention, three sets of diagnostic experiments were conducted on the Minimal Evidence Package (MEP) in this embodiment.
[0048] (1) Evidence extractability experiment The occurrence rate of evidence types was statistically analyzed from 100 randomly sampled inference chains, and the results are shown in Table 7.
[0049] Table 7. Occurrence Rate of Evidence Types As shown in Table 7, visual evidence appeared in 84% of the samples, temporal evidence in 72%, and behavioral evidence in 89%. 93% of the samples contained at least two types of evidence, indicating that the model's reasoning chain contained rich structured information, which provided a foundation for subsequent grounding verification.
[0050] (2) Evidence-verification consistency experiment The occurrence rate of grounding-related elements and the grounding success rate in the inference chain were evaluated, and the results are shown in Tables 8 and 9.
[0051] Table 8. Occurrence rate of grounding elements and grounding success rate Table 9 Grounding Compliance Rate by Problem Type As can be seen from Tables 8-9, the person mention rate is 98%, the visual feature description rate is 91%, and the locatable action rate is 94%; the basic grounding rate reaches 98%, and the high-quality grounding rate reaches 86%, which proves that the MEP of the present invention has sufficient executable details and can provide effective positioning instructions for downstream visual verifiers.
[0052] (3) External audit and falsifiability experiment The invited annotators judged the relevance of the annotation to the question objective based solely on the SAM3 segmentation results, and statistically analyzed the correlation with the correctness of the answer. The results are shown in Table 10.
[0053] Table 10. Association between segmentation relevance and answer correctness (n=100) Table 10 shows that when the segmentation was determined to be relevant, the conditional accuracy was 83.3%; when it was determined to be irrelevant, the conditional accuracy dropped to 35.3%, a difference of 48.0 percentage points. The Phi coefficient was 0.48, indicating a moderate positive correlation between evidence grounding and answer correctness. Figure 5 As shown, the pixel-level segmentation mask generated by SAM3 visualizes the evidence referenced by the model, allowing external reviewers to obtain external signals regarding the reliability of the answer without reading the reasoning text. When the segmentation becomes disconnected from the problem objective, the accuracy drops significantly, demonstrating that the visual evidence grounding of this invention provides a falsifiable external auditing interface for the reasoning results.
[0054] like Figure 7 As shown, in the micro-expression understanding task, the present invention can capture more fine-grained facial cues (such as furrowed eyebrows and slightly parted lips) and eliminate hypotheses, while the baseline model ignores these key details.
[0055] like Figure 8As shown, in temporal reasoning tasks, this invention can model the dynamic evolution of emotions over time. For the question "Is the man wearing glasses passionate when discussing the electoral system?", the baseline model makes a judgment based only on a single moment, while this invention identifies the increasing changes in emotional intensity and the gradual clarification of the argument, thus making a more accurate judgment. This demonstrates that the temporally ordered reward of this invention can effectively guide the model to capture the evolutionary trajectory of emotions over time.
[0056] Figure 9 and Figure 10 Two representative failure cases are presented. For example... Figure 9 As shown, in a 60-second stand-up comedy clip, the model failed to detect the "surface-intention" rhetorical contrast in irony, misinterpreting sarcastic remarks as serious statements, leading to incorrect answers. This reveals the limitations of this invention in pragmatic reasoning (irony, satire, humor). Figure 10 As shown, in the 58-second interview segment, although the model correctly understood the respectful relationship between the people in the video, it was distracted by the background narration when choosing an answer, selecting an option irrelevant to the semantics of the question, resulting in a disconnect between reasoning and selection. These cases point the way for further improvements to this invention.
[0057] A multimodal large language model reinforcement learning system for verifiable sentiment reasoning includes: The data acquisition module is used to execute step S1; The model training module is used to execute step S2; The model inference module is used to execute step S3.
[0058] The content of the above method embodiments is applicable to this system embodiment. The specific functions implemented in this system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0059] A multimodal large language model reinforcement learning device for verifiable sentiment reasoning: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements a multimodal large language model reinforcement learning method for verifiable sentiment reasoning as described above.
[0060] The content of the above method embodiments is applicable to the device embodiments. The specific functions implemented by the device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0061] A storage medium storing processor-executable instructions, which, when executed by a processor, are used to implement a multimodal large language model reinforcement learning method for verifiable sentiment reasoning as described above.
[0062] The content of the above method embodiments is applicable to this storage medium embodiment. The specific functions implemented in this storage medium embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those achieved in the above method embodiments.
[0063] The above is a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.
Claims
1. A reinforcement learning method for a multimodal large language model that can verify sentiment reasoning, characterized in that, Includes the following steps: Acquire multimodal data and construct a training dataset, wherein the multimodal data includes video frame sequences, audio signals, and text queries; The training dataset is input into the multimodal large language model, and the multimodal large language model is trained by reinforcement learning based on the group relative policy optimization framework to obtain the trained sentiment reasoning model; the reward function during the training phase includes person focus reward; The multimodal data to be tested is input into the trained sentiment reasoning model. After the multimodal feature fusion and encoding are completed within the model, a structured reasoning output containing a context segment, a thought segment, and an answer segment is generated. The thought segment contains a fine-grained sentiment reasoning chain centered on the human. The thought segment is compressed into a structured summary using an inference digester, and a pixel-level evidence mask is generated based on the structured summary using the SAM3 segmentation model. The pixel-level evidence mask is then used to perform external verification of the inference results.
2. The multimodal large language model reinforcement learning method for verifiable sentiment reasoning according to claim 1, characterized in that, The objective function of the relative policy optimization framework is: In the formula, This represents the policy optimization objective function. This represents the expectation operation. This indicates the number of candidate inference results for each sample group. Indicates the first A sequence of candidate inference results, Indicates the first The length of the token for each candidate result. Indicates the first The candidate result is the The loss item corresponding to each token This indicates the minimum value operation. Indicates the first The candidate result is the The importance sampling ratio of each token Indicates the first The standardized odds value of each candidate outcome. This represents the clipping function. This indicates the hyperparameters of the pruning strategy. Indicates the current update policy distribution. Indicates the first The candidate result is the One token, This indicates that the input is a multimodal query. Indicates the first The candidate result is the The sequence preceding each token This represents the old policy distribution before the update.
3. The multimodal large language model reinforcement learning method for verifiable sentiment reasoning according to claim 2, characterized in that, The expression for the reward function in the training process is as follows: In the formula, This indicates a reward for compliance with formatting regulations. Represents a reward based on reasoning accuracy. This indicates the weighting coefficient of the reward for focusing on specific individuals. This indicates that the focus is on the reward. This represents the time-ordered reward weighting coefficient. Indicates a sequential reward. This represents the scoring function of the large language model judge. Indicates sample context information, Represents inference text, Indicates key words for focusing on the assessment of individuals. Indicates a sequential evaluation prompt.
4. The multimodal large language model reinforcement learning method for verifiable sentiment reasoning according to claim 2, characterized in that, The reasoning results are scored using intra-group comparison scoring, and the output of the intra-group comparison scoring is as follows: In the formula, This represents the relative score of the i-th candidate. This represents a comparative scoring function. Indicates the candidate inference result, Indicates a contrasting prompt.
5. The multimodal large language model reinforcement learning method for verifiable sentiment reasoning according to claim 4, characterized in that, By integrating the various reward scores, the total reward for a single candidate result is obtained. The weighted summation formula for the total reward of a single candidate result is as follows: In the formula, Indicates the first Total reward for each candidate result Indicates the first Each reward component weight, Indicates the first The first candidate result Each reward component.
6. The multimodal large language model reinforcement learning method for verifiable sentiment reasoning according to claim 5, characterized in that, The calculation and standardization of intra-group advantage are based on the total reward, where: The formula for calculating the standardized advantage value is: In the formula, Indicates the first The advantage value of each candidate result Indicates the first Total reward for each candidate result This represents the average reward within the group. Indicates the first Total reward for each candidate result; The formula for calculating the within-group standardized odds value is: In the formula, Indicates the standard deviation of the group's rewards. This represents the numerical smoothing coefficient.
7. The multimodal large language model reinforcement learning method for verifiable sentiment reasoning according to claim 6, characterized in that, The loss function for the inference digester is: In the formula, Indicates the summarization loss. Indicates the parameters of the digester. The structured summary is represented by the first... One token, Represents the original chain of reasoning. This represents a summary history token.
8. The multimodal large language model reinforcement learning method for verifiable sentiment reasoning according to claim 6, characterized in that, The formula for generating the SAM3 pixel-level segmentation mask is: In the formula, Indicates the first Cross-frame segmentation mask for each target entity This represents a pixel-level segmentation model. Indicates the input video frame sequence. Indicates the first A target entity to be located. Indicates the total number of target entities.
9. A multimodal large language model reinforcement learning system for verifiable sentiment reasoning, characterized in that, include: The data acquisition module is used to acquire multimodal data and build a training dataset; The model training module trains a strategy model based on the training data and the group relative strategy optimization framework to obtain a trained sentiment reasoning model; wherein, the reward function in the training process includes a person-focused reward. The model inference module is used to input the test data into the trained sentiment inference model and generate structured sentiment inference output.
10. A multimodal large language model reinforcement learning device for verifiable sentiment reasoning, characterized in that, include: At least one processor; At least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements a multimodal large language model reinforcement learning method for verifiable sentiment reasoning as described in any one of claims 1-8.