A dual-objective reinforcement learning method for multi-modal agent search
Patent Information
- Application Number
- CN202610865463.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-16
- Publication Date
- 2026-09-18
AI Technical Summary
[0004]然而,这些现有的基于强化学习的智能体搜索方法在有效性和效率方面仍然存在严重的不足
1:本发明提供一种面向多模态智能体搜索的双目标强化学习方法,该方法包括思考、搜索、信息获取、反思、总结和回答的多阶段推理流程,并在奖励建模阶段分别构建用于直接监督搜索质量的检索奖励,以及用于协同优化问答性能的答案正确性奖励、组内效率奖励和格式奖励,通过动态权重退火策略融合基于检索奖励的搜索目标函数与基于回答、格式与效率奖励的答案目标函数,解决了现有强化学习搜索方法忽视中间检索质量以及无法精细化抑制冗余搜索调用的技术问题,达到了在显著提高视觉问答准确率的同时大幅提升中间检索召回率、降低无效检索开销的技术效果。
Smart Images

Figure CN122779121A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal large model technology, and in particular to a dual-objective reinforcement learning method for multimodal agent search. Background Technology
[0002] Large language models and multimodal large language models have demonstrated significant reasoning and understanding capabilities in various visual question-answering tasks. To further enhance the ability of these large models to handle knowledge-intensive multimodal tasks, mainstream approaches in current technology are integrating multimodal large language models with external retrieval tools or web search engines. By iteratively and dynamically invoking retrieval tools during the reasoning process to acquire real-time external knowledge, large models can effectively overcome the inherent limitations of internal knowledge lag or forgetting caused by pre-training. This agent-based search architecture typically follows a multi-stage workflow: thinking, retrieving, and answering. The large model first analyzes and decomposes the input query and decides whether external knowledge retrieval is necessary; subsequently, the model invokes retrieval tools to obtain relevant documents or web page content, and can perform multiple iterations; finally, the large model integrates the retrieved contextual knowledge with its own parameterized knowledge to generate the final question-answering result.
[0003] Currently, to perform more efficient post-training of such search-capable agent systems, existing techniques (such as MMSearch-R1 and Search-R1) typically employ policy gradient-based reinforcement learning frameworks, such as Group Relative Policy Optimization (GRPO). While Search-R1 uses the correctness of the final answer as a feedback signal to guide the retrieval invocation behavior of the large model, MMSearch-R1 reduces unnecessary search invocations by introducing a fixed tool use penalty term into the reward function.
[0004] However, these existing reinforcement learning-based agent search methods still suffer from serious shortcomings in terms of effectiveness and efficiency. Regarding effectiveness, the reward modeling of the existing Search-R1 method only considers whether the final generated answer is correct, completely ignoring the intermediate retrieval strategies and the quality of the retrieval results. This leads to suboptimal search behavior, and this design results in the inability to provide direct retrieval supervision signals. For example, when a large model retrieves completely correct document information in an iteration, but then makes an incorrect answer in the subsequent final answer inference stage due to other interfering factors, the existing method will penalize the entire inference path, including the correct retrieval. This prevents the model from effectively absorbing positive retrieval behavior signals, and the learning of retrieval strategies is extremely unstable. Regarding efficiency, the existing MMSearch-R1 method cannot fully achieve fine-grained modeling of search efficiency. Although some methods suppress tool calls through simple search penalties, this approach is a coarse-grained, indiscriminate penalty, which often leads the model to hesitate to call searches for fear of penalty when facing complex, long-tailed problems requiring multi-step deep retrieval. Furthermore, if multiple inference paths that can all produce the correct answer exist within the same sampling group, but their search call counts differ, existing frameworks often give them the exact same positive reward. This makes it impossible to establish an explicit preference for concise and efficient retrieval behavior within the group, resulting in a large number of redundant and noisy invalid search calls during the inference phase, which significantly increases the computational overhead and token consumption for large model inference.
[0005] Therefore, how to improve the accuracy of visual question answering without reducing intermediate search recall and invalid retrieval overhead is an urgent problem to be solved. Summary of the Invention
[0006] In view of this, the present invention provides a dual-objective reinforcement learning method for multimodal agent search, which can significantly improve the accuracy of visual question answering while greatly improving the intermediate retrieval recall rate and reducing the cost of invalid retrieval.
[0007] To achieve the above objectives, the present invention provides a dual-objective reinforcement learning method for multimodal agent search, the technical solution of which includes the following steps:
[0008] Step 1: Construct a multimodal intelligent agent model. The multimodal intelligent agent model executes a multi-stage reasoning process. The multi-stage reasoning process consists of the thinking stage, the search stage, the information acquisition stage, the reflection stage, the summary stage, and the response stage.
[0009] Step 2: Design four types of reinforcement learning reward signals for the above multi-stage reasoning process, namely retrieval quality reward, answer correctness reward, search efficiency reward, and format reward.
[0010] Step 3: Construct a dual-objective reinforcement learning optimization system, decouple the optimization objectives of search behavior and response behavior, and construct local objective functions and global objective functions respectively; use dynamic weighted annealing strategy to fuse the local objective functions and global objective functions to obtain the joint total objective function.
[0011] Step 4: Iteratively train the constructed multimodal agent model, updating the parameters of the multimodal agent model iteratively according to the joint overall objective function until the preset training termination condition is met, thus obtaining the trained multimodal agent.
[0012] Step 5: Deploy the trained multimodal agent for multimodal visual question answering tasks.
[0013] Further, the thinking phase: receiving multimodal input samples consisting of input images and text questions, the multimodal large language model performs semantic understanding and knowledge requirement analysis on the questions, and enters the search phase when it determines that its own parameter knowledge cannot support the answer.
[0014] Search phase: The multimodal large language model generates search instructions containing search keywords and image usage tags. The image usage tags are used to control the multimodal retrieval system to perform plain text search or image-text combined search. The multimodal retrieval system matches the preset knowledge base according to the search instructions and returns a set of relevant knowledge fragments.
[0015] Information acquisition stage: The relevant knowledge fragments obtained in the search stage are integrated into the reasoning chain to support subsequent reasoning work as external knowledge.
[0016] Reflection phase: Evaluate the relevance and information completeness of the search results; if the search results are irrelevant to the question, reconstruct the search instructions and return to the search phase; if the search results are relevant but lack sufficient information, retain the existing evidence and continue the search; if the search results meet the answer requirements, proceed to the summary phase.
[0017] Summary phase: Integrate all external knowledge obtained from multiple rounds of retrieval.
[0018] Answering phase: The integrated knowledge is combined to generate and output the final question and answer.
[0019] Furthermore, the specific execution method of the search phase in step one is as follows: the retrieval instruction includes the retrieval query and the image usage tag with_image; when the image usage tag is enabled, the multimodal retrieval system performs image-text joint retrieval by combining the input image and the text query; when the image usage tag is disabled, the multimodal retrieval system performs retrieval based solely on the text query; after completing the knowledge base similarity matching, the multimodal retrieval system returns a set of Top-k relevant knowledge fragments, and encapsulates all knowledge fragments into a preset information module before sending them back to the multimodal large language model.
[0020] Furthermore, the reflection phase in step one supports multiple rounds of iterative operation, and the complete reasoning trajectory is as follows: thinking module, search module, information module, reflection module, summary module, and answer module. Among them, the search module, information module, and reflection module perform multiple rounds of cyclical iteration according to the complexity of the question. When the reflection phase determines that the search results are irrelevant to the question, the search query is reconstructed and the process jumps to the search phase. When the search results are determined to be relevant but the information is insufficient, the existing search evidence is retained and the search operation is performed again. When the search results are determined to support the answer, the iteration terminates and the process enters the summary phase.
[0021] Furthermore, the four types of reward signals in step two are as follows: The retrieval quality reward is determined based on whether the retrieval results contain the target knowledge document, and is used to supervise the quality of retrieval behavior; the answer correctness reward is calculated based on the degree of matching between the model-generated answer and the standard answer, and is used to ensure the accuracy of question-and-answer results; the intra-group search efficiency reward is calculated for multiple correct reasoning paths corresponding to the same question, combined with the number of search calls for each path and the intra-group normalization method, and is used to guide the agent to reduce redundant search behavior; the format reward is calculated based on whether the reasoning trajectory conforms to the preset module format, and is used to constrain the standardization of the reasoning process.
[0022] The specific calculation rules are as follows: 1) The retrieval quality reward is determined by an indicator function. When the retrieval result set contains the target knowledge document set, a positive reward is output; otherwise, the reward value is 0.
[0023] 2) The correctness reward adopts the Exact Match string matching method to measure the degree of matching between the model-generated answer and the standard answer and assign a value.
[0024] 3) Search efficiency reward: First, calculate the efficiency score based on the number of search calls for a single correct reasoning trajectory, and then obtain the final reward through group index normalization.
[0025] 4) Format rewards are calculated based on the ratio of "number of compliant modules / preset total number of modules" and are used to verify whether the reasoning trajectory includes complete thinking, searching, information, reflection, summarizing, and answering modules.
[0026] Furthermore, in step three, the local objective function and the global objective function are constructed as follows: for a sample group consisting of multiple reasoning trajectories corresponding to the same question, the mean and standard deviation of the retrieval quality reward within the group are calculated respectively, and the search advantage function is obtained after normalization; the mean and standard deviation of the global reward within the group are calculated respectively, and the answer advantage function is obtained after normalization; the two types of advantage functions are used to quantify the quality of a single reasoning trajectory within the group.
[0027] A local advantage function is constructed based on the retrieval quality reward, and a local objective function is generated. The local objective function only applies to the tokens corresponding to the search phase in the reasoning process, and is used to optimize the agent's search behavior. A global advantage function is constructed based on the answer correctness reward, search efficiency reward, and format reward, and a global objective function is generated. The global objective function applies to all tokens in the entire reasoning trajectory, and is used to optimize the agent's information integration and answer generation capabilities.
[0028] Further, in step four, perform iterative training of the model: sample multimodal training samples from the training dataset, and have the multimodal agent generate multiple inference trajectories for a single sample; calculate the four types of reward signals, local advantage function and global advantage function corresponding to each inference trajectory in sequence; iteratively update the parameters of the multimodal large language model according to the joint total objective function until the preset training termination condition is reached, and obtain the trained multimodal agent.
[0029] Furthermore, the model iterative training process in step four specifically includes: 1) Sample multimodal input samples from the training dataset.
[0030] 2) Control the multimodal agent to generate multiple complete inference trajectories in parallel for the same sample.
[0031] 3) Calculate the retrieval quality reward, answer correctness reward, format reward, and group search efficiency reward for each reasoning trajectory in sequence.
[0032] 4) Construct search advantage functions and answer advantage functions based on the four types of rewards respectively.
[0033] 5) The local objective function and the global objective function are fused using a dynamic weighted annealing strategy to obtain a joint total objective function, and the parameters of the multimodal large language model are updated based on the joint total objective function.
[0034] 6) Determine whether the preset training termination condition has been met. If not, return to the sampling step to continue iterating. If it has been met, terminate the training.
[0035] Furthermore, the joint overall objective function is obtained by multiplying the answer objective function and the search objective function by their respective weight coefficients and then summing them. During the iterative training of the model, in the early stage of training, the weight coefficient corresponding to the search objective function is increased to prioritize training the agent's retrieval ability. As the training rounds increase, the weight of the search objective function is gradually reduced and the weight of the answer objective function is increased, ultimately achieving synergistic optimization of the agent's retrieval ability and question answering ability.
[0036] Beneficial effects: 1. This invention provides a dual-objective reinforcement learning method for multimodal agent search. The method includes a multi-stage reasoning process encompassing thinking, searching, information acquisition, reflection, summarizing, and answering. In the reward modeling stage, it constructs retrieval rewards for directly supervising search quality, and answer correctness rewards, intra-group efficiency rewards, and format rewards for collaboratively optimizing question-answering performance. A dynamic weighted annealing strategy is used to fuse the search objective function based on retrieval rewards with the answer objective function based on answer, format, and efficiency rewards. This solves the technical problems of existing reinforcement learning search methods neglecting intermediate retrieval quality and failing to finely suppress redundant search calls. It achieves the technical effect of significantly improving visual question-answering accuracy while greatly increasing intermediate retrieval recall and reducing invalid retrieval overhead.
[0037] 2: This invention provides a dual-objective reinforcement learning method for multimodal agent search. By introducing a retrieval quality reward, it achieves direct supervision of search behavior, enabling the model to learn more accurate knowledge retrieval strategies.
[0038] 3: This invention provides a dual-objective reinforcement learning method for multimodal agent search. By introducing an in-group search efficiency reward, it achieves explicit optimization of the number of searches, reducing inference costs and tool call overhead.
[0039] 4. This invention provides a dual-objective reinforcement learning method for multimodal agent search. By introducing a retrieval reflection mechanism, it improves the utilization rate of retrieval results and reduces the interference of irrelevant knowledge on the reasoning process.
[0040] 5: This invention provides a dual-objective reinforcement learning method for multimodal agent search. By constructing a dual-objective reinforcement learning framework consisting of a search objective and a response objective, it achieves synergistic optimization of search capability and response capability.
[0041] 6: This invention provides a dual-objective reinforcement learning method for multimodal agent search, which achieves higher answer accuracy and better search efficiency compared to existing multimodal search enhancement methods on multiple knowledge-intensive visual question answering datasets. Attached Figure Description
[0042] Figure 1 This is a schematic diagram of a dual-objective reinforcement learning method architecture for multimodal agent search provided in an embodiment of the present invention; Figure 2 A multimodal intelligent agent reasoning flowchart provided in an embodiment of the present invention; Figure 3 This is a flowchart of a model training method provided in an embodiment of the present invention. Detailed Implementation
[0043] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.
[0044] Example 1: A dual-objective reinforcement learning method for multimodal agent search includes the following steps: Step 1: Construct a multimodal intelligent agent reasoning framework and set up a multi-stage reasoning process; the multi-stage reasoning process consists of the thinking stage, the search stage, the information acquisition stage, the reflection stage, the summary stage, and the response stage. Thinking phase: Receive multimodal input samples consisting of input images and text questions. The multimodal large language model performs semantic understanding and knowledge requirement analysis on the questions. If it determines that its own parameter knowledge cannot support the answer, it triggers a search operation. Search Phase: The multimodal large language model generates search instructions containing search keywords and image usage tags. Image usage tags control the multimodal retrieval system to perform plain text retrieval or combined text-image retrieval. The multimodal retrieval system matches the pre-defined knowledge base according to the search instructions and returns a set of relevant knowledge fragments. The specific execution method of the search phase is as follows: the search instructions include the search query and the image usage tag with_image. When the image usage tag is enabled, the multimodal retrieval system performs combined text-image retrieval by combining the input image and the text query. When the image usage tag is disabled, the multimodal retrieval system performs retrieval based only on the text query. After completing the knowledge base similarity matching, the multimodal retrieval system returns a set of the top-k relevant knowledge fragments, encapsulates all knowledge fragments into a pre-defined information module, and then sends it back to the multimodal large language model.
[0045] Information acquisition stage: The retrieved knowledge fragments are integrated into the reasoning chain to serve as external knowledge to support subsequent reasoning work; Reflection Phase: The relevance and information completeness of the search results are assessed. If the search results are irrelevant to the question, the search instructions are reconstructed and the search phase is returned. If the search results are relevant but lack sufficient information, existing evidence is retained and the search continues. If the search results meet the answer requirements, the summary phase begins. The reflection phase supports multiple iterations, with the complete reasoning trajectory consisting of the thinking module, search module, information module, reflection module, summary module, and answer module. The search module, information module, and reflection module perform multiple iterations based on the question's complexity. When the reflection phase determines that the search results are irrelevant to the question, the search query is reconstructed and the search phase resumes. When the search results are determined to be relevant but lack sufficient information, existing search evidence is retained and the search operation is performed again. When the search results are determined to support the answer, the iteration terminates and the summary phase begins.
[0046] Summary phase: Integrate all external knowledge obtained from multiple rounds of retrieval; Answer phase: Based on the integrated knowledge, generate and output the final question and answer; Based on the design at each stage above Figure 2 The multi-stage reasoning process is illustrated, with the specific steps as follows: First, in step S101, the system receives the input image. and question text To form multimodal input samples Then, step S102 is executed, where the multimodal large language model analyzes the input content and proceeds to... <think>< / think> In the first phase, the model performs reasoning and planning for the current problem. When the model determines that the existing parameter knowledge is insufficient to complete the answer, it generates a search request and enters the search phase.
[0047] Next, step S103 is executed to generate the model. <search>< / search> Request. A search request includes a retrieval query. query and image call tags with_image .in, query Used to describe the knowledge content to be searched; with_image Used to indicate whether the input image is used simultaneously during retrieval. with_image When the answer is "Yes", both images and text are used for retrieval; when... with_image When the value is "No", only text is used for retrieval. Then, step S104 is executed, where the retrieval module retrieves Top-k relevant knowledge fragments from the knowledge base based on the search request. The obtained search results are represented as follows:
[0048] in Indicates the first Each search segment is used, and the search results are encapsulated in... <information>< / information> The module returns the model.
[0049] Next, step S105 is executed, and the model enters... <reflect>< / reflect> In this stage, the search results are evaluated. When the search results are irrelevant to the problem, the model reconstructs the query and returns to step S103; when the search results are relevant but lack sufficient information, the existing evidence is retained and the search continues; when the search results can support the solution to the problem, the conclusion generation stage begins.
[0050] Then, step S106 is executed, and the model enters... <conclude>< / conclude> In this stage, the information obtained from multiple rounds of retrieval is integrated and summarized.
[0051] Finally, step S107 is executed, and the model enters... <answer>< / answer> In this stage, the final answer is generated and output based on the reasoning results.
[0052] Therefore, this embodiment forms the following reasoning trajectory: <think> → <search> → <information> → <reflect> → <conclude> → <answer>< / answer> < / conclude> < / reflect> < / information> < / search> < / think> in <search> 、 <information>< / information> < / search> and <reflect>< / reflect> The module can be executed in multiple iterations depending on the complexity of the task, until sufficient evidence is obtained to support the final answer.
[0053] Step 2: Design four types of reinforcement learning reward signals for the above multi-stage reasoning process: retrieval quality reward, answer correctness reward, intra-group search efficiency reward, and format reward. The retrieval quality reward is determined based on whether the retrieval results contain the target knowledge document, used to supervise the quality of retrieval behavior. The answer correctness reward is calculated based on the degree of matching between the model-generated answer and the standard answer, used to ensure the accuracy of question-answering results. The intra-group search efficiency reward is calculated for multiple correct reasoning paths corresponding to the same question, combined with the number of search calls for each path and the intra-group normalization method, used to guide the agent to reduce redundant search behavior. The format reward is calculated based on whether the reasoning trajectory conforms to the preset module format, used to constrain the standardization of the reasoning process. To simultaneously improve search and answer capabilities, this embodiment constructs four types of reward signals: retrieval quality reward, answer correctness reward, search efficiency reward, and format reward.
[0054] The retrieval quality reward is used to measure whether search results cover the target knowledge document. (For inference trajectory) ,set up This represents the corresponding set of search results. Let the target knowledge document set be represented, then the retrieval quality reward is defined as:
[0055] in, Indicates the reward coefficient. This represents the indicator function. A positive reward is given when the search results contain a document containing the target knowledge; otherwise, the reward is zero. This reward is used to encourage the model to learn more effective search strategies and improve the retrieval hit rate of the target knowledge.
[0056] The correctness reward is used to measure the degree of consistency between the final answer and the standard answer. Let... This indicates that the model generates the answer. If the standard answer is represented, then the reward for correct answer is defined as:
[0057] in, This invention uses the Exact Match string matching method to measure the degree of matching between the model's answer and the standard answer, representing the answer evaluation function. The model receives a higher reward when it generates a correct answer, and a lower reward or zero reward otherwise.
[0058] Search efficiency rewards are used to measure the efficiency of the search process. This applies to multiple correct reasoning paths corresponding to the same problem. First, count the number of searches corresponding to each trajectory. And construct an efficiency score:
[0059] in, This represents the scaling factor. The more searches performed, the lower the efficiency score. A search efficiency bonus is then calculated using within-group normalization.
[0060] in, This represents the set of correct reasoning trajectories corresponding to the same question. Through the normalization calculation method described above, correct trajectories with fewer search iterations receive higher rewards, thereby guiding the model to reduce redundant search behavior and improve search efficiency.
[0061] Format rewards are used to constrain the model to output inference in a predefined format, including <think>< / think> , <search>< / search> , < information>、 <reflect> 、 <conclude>< / conclude> < / reflect> as well as <answer>< / answer> Modules such as... (Settings are missing from the original text.) This indicates the number of modules that meet the format requirements. Given the total number of modules, the format reward is defined as follows:
[0062] A higher reward is given when the model output meets the predefined format requirements; otherwise, the reward is reduced accordingly. For example, in information acquisition... <information>< / information> Post-missing <reflect>< / reflect> The issue stems from a re-evaluation of tags, or the generation of queries with incorrect formats, such as missing tags. with_image "The tags all failed the format check. Format rewards can ensure the structure of the inference trajectory is standardized and improve the stability of the training process."
[0063] Based on the above reward mechanism, this embodiment supervises the model behavior from four aspects: retrieval quality, answer correctness, search efficiency, and inference format, providing training signals for the subsequent dual-objective reinforcement learning optimization process.
[0064] Step 3: Construct a dual-objective reinforcement learning optimization system based on the inter-group relative strategy optimization framework, decoupling the optimization objectives of search behavior and response behavior, and constructing search objective functions and response objective functions respectively; construct a search advantage function and generate a search objective function based on retrieval quality reward and intra-group search efficiency reward. The search objective function only applies to the tokens corresponding to the search stage in the inference process and is used to optimize the agent's retrieval decision behavior; construct a response advantage function and generate a response objective function based on answer correctness reward, retrieval quality reward, intra-group search efficiency reward, and format reward. The response objective function applies to all tokens in the entire inference trajectory and is used to optimize the agent's information integration and answer generation capabilities. Step 4: Employ a dynamic weighted annealing strategy to fuse the search objective function and the response objective function to obtain a joint overall objective function. In the early stages of training, increase the weight of the search objective function to prioritize training the agent's retrieval capabilities. As the training progresses, gradually increase the weight of the response objective function to collaboratively optimize the agent's knowledge integration and answer generation capabilities.
[0065] Traditional reinforcement learning methods typically optimize the entire reasoning process using a uniform reward, with the optimization objective primarily focused on the correctness of the final answer. However, in multimodal agent search tasks, search and response behaviors have different optimization objectives. Search behavior mainly focuses on retrieval quality, while response behavior focuses on answer correctness and reasoning quality. Using a uniform reward for optimization makes it difficult to effectively supervise the search strategy.
[0066] Therefore, this embodiment constructs a dual-objective reinforcement learning optimization mechanism to model and optimize search behavior and response behavior respectively.
[0067] First, based on the retrieval quality reward constructed in Example 3, a local reward is defined:
[0068] in, This indicates an input problem. Indicates the first A line of reasoning, This represents the corresponding retrieval quality reward. The search reward only reflects the retrieval performance of the current trajectory and is used to guide the model in learning more effective search strategies.
[0069] Furthermore, a global reward is constructed based on reward for correct answer, reward for format, and reward for search efficiency:
[0070] in, Rewards are given for correct answers. Indicates a formatted reward. This represents the search efficiency reward. The answer reward is used to measure the overall quality of the final answer generated by the model.
[0071] After obtaining a local reward, the inference trajectories corresponding to the same problem are normalized within the group to construct a local dominance function:
[0072] in, and These represent the mean and standard deviation of the rewards within the group, respectively.
[0073] Similarly, the global reward is normalized within the group to construct a global advantage function:
[0074] The above methods can be used to measure the quality of search behavior and response behavior relative to the same group of samples.
[0075] In the local dominant function and global advantage function Based on this, local objective functions are constructed using the GRPO optimization objective. and global objective function First, we define A function to tag a token. When When, it indicates the output sequence The Middle The local objective function is: If the target token is a local (search) token, then the target token is a local (search) token; otherwise, the target token is a response token. Therefore, the local objective function is:
[0076] in, Indicates the output sequence The set of locations of all local (search) tokens; The KL divergence constraint coefficients in the local objective; This represents the KL divergence between the current strategy and the reference strategy; This indicates that the current strategy and the old strategy are in the [number]th [phase]. The probability ratio at each token: This indicates a truncation operation.
[0077] Correspondingly, the global objective function is:
[0078] use As an optimization signal corresponding to the search behavior, it is used to guide the model to learn more effective local search strategies. use As an optimization signal corresponding to global behavior, it is used to improve the model's information integration and answer generation capabilities.
[0079] Finally, the two optimization objectives are jointly modeled to obtain the overall optimization objective:
[0080] in, Indicates model parameters, Indicates the current training phase. and These represent the weight coefficients corresponding to the answer target and the search target, respectively.
[0081] During training, the weight parameters are dynamically adjusted according to different training stages. In the early stages of training, the weight of the search target is increased, allowing the model to prioritize learning effective search strategies. As training progresses, the weight of the answer target is gradually increased, enabling the model to further improve its knowledge integration and answer generation capabilities. Through this dual-objective collaborative optimization mechanism, the joint improvement of search and answering abilities is achieved.
[0082] Step 5: Perform iterative training of the model: Sample multimodal training samples from the training dataset, and have the multimodal agent generate multiple inference trajectories for a single sample; calculate the four types of reward signals, search advantage function and answer advantage function corresponding to each inference trajectory in turn; iteratively update the parameters of the multimodal large language model according to the joint overall objective function until the preset training termination condition is reached, and obtain the trained multimodal agent.
[0083] Model training method reference Figure 3 As shown, this embodiment discloses a model training method based on dual-objective reinforcement learning.
[0084] First, step S201 is executed, sampling training questions from the training dataset and constructing corresponding multimodal input samples. Then, step S202 is executed, where the multimodal large language model generates multiple candidate reasoning trajectories for the same question. Each reasoning trajectory includes processes such as thinking, searching, information acquisition, reflection, summarizing, and answering. The model completes the search and reflection processes in parallel along each reasoning path. During the reasoning process, the model autonomously decides whether to call the external knowledge retrieval module based on the current question and performs reflective analysis based on the retrieval results, thereby deciding whether to continue searching or enter the answer generation stage.
[0085] After obtaining the reasoning results, steps S203 to S206 are executed respectively to construct reward signals. Specifically, step S203 calculates a retrieval quality reward based on the matching between the retrieval results and the target knowledge document; step S204 calculates an answer correctness reward based on the consistency between the model answer and the standard answer; step S205 calculates a format reward based on whether the output results meet a predefined reasoning format; and step S206 calculates a search efficiency reward based on the number of searches corresponding to the reasoning trajectory. After receiving the aforementioned rewards, step S207 is executed to construct a search advantage function, and step S208 is executed to construct a response advantage function. The search advantage function is used to evaluate the quality of the search behavior, and the response advantage function is used to evaluate the quality of the response behavior.
[0086] Then, step S209 is executed, in which the search objective function and the response objective function are constructed based on the search advantage function and the response advantage function, respectively, and the joint optimization objective is calculated based on the dual-objective reinforcement learning optimization mechanism.
[0087] Next, step S210 is executed to update the model parameters according to the joint optimization objective. Then, it is determined whether the current training has reached the preset termination condition; if the termination condition has not been reached, the process returns to step S201 to continue the next round of training; if the termination condition has been reached, the training ends and the final model is output.
[0088] Through the training process described above, the model can gradually learn when to perform a search, how to construct a search request, how to use the retrieval results to complete inference, and when to terminate the search. Upon completion of training, a multimodal intelligent agent model with autonomous search capabilities is obtained.
[0089] During the inference phase, the model can autonomously decide whether to perform a search, when to continue the search, and when to terminate the search based on the complexity of the question without additional supervision. It also combines the external knowledge obtained from the retrieval to complete the final answer, thereby achieving efficient solutions for complex knowledge-intensive visual question answering tasks.
[0090] Step Six: Deploy the trained multimodal agent to the multimodal visual question answering task. The agent will autonomously complete the entire process of knowledge demand judgment, retrieval decision, retrieval result evaluation, information integration and answer generation.
[0091] Example 2: Reference Figure 1 As shown, this embodiment proposes a multimodal agent search optimization method based on dual-objective reinforcement learning. The system mainly includes a multimodal large language model module, a multimodal knowledge retrieval module, a retrieval result reflection module, a reward calculation module, and a dual-objective reinforcement learning optimization module.
[0092] The multimodal large language model is used to complete problem understanding, search decision-making, reasoning planning, and answer generation; the multimodal knowledge retrieval module is used to obtain relevant knowledge from external knowledge bases based on the search requests generated by the model; the retrieval result reflection module is used to analyze and evaluate the retrieval content and decide whether to continue the search; the reward calculation module is used to construct reinforcement learning reward signals based on the search results and answer results; and the dual-objective reinforcement learning optimization module is used to update the model parameters based on the reward signals.
[0093] During the training phase, the model continuously interacts with the retrieval module through reinforcement learning, learning when to search, how to search, and how to use search results to complete inference. During the inference phase, the trained model can autonomously decide whether to invoke external knowledge based on the complexity of the question, and obtain supporting evidence through multiple rounds of searching and reflection, ultimately generating the answer.
[0094] This embodiment combines search capability optimization with answer capability optimization in joint modeling, enabling the model to simultaneously achieve high retrieval quality, answer accuracy, and search efficiency.
[0095] In summary, the above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A dual-objective reinforcement learning method for multi-modal agent search, characterized in that, Includes the following steps: Step 1: Construct a multimodal intelligent agent model, which executes a multi-stage reasoning process; the multi-stage reasoning process consists of the thinking stage, the search stage, the information acquisition stage, the reflection stage, the summary stage, and the response stage. Step 2: Design four types of reinforcement learning reward signals for the above multi-stage reasoning process, namely retrieval quality reward, answer correctness reward, search efficiency reward, and format reward; Step 3: Construct a dual-objective reinforcement learning optimization system, decouple the optimization objectives of search behavior and response behavior, and construct local objective functions and global objective functions respectively; use dynamic weighted annealing strategy to fuse the local objective functions and global objective functions to obtain the joint overall objective function; Step 4: Iteratively train the constructed multimodal agent model, updating the parameters of the multimodal agent model iteratively according to the joint overall objective function until the preset training termination condition is met, thus obtaining the trained multimodal agent. Step 5: Deploy the trained multimodal agent for multimodal visual question answering tasks.
2. The dual-objective reinforcement learning method for multi-modal agent search according to claim 1, wherein, The thinking phase involves receiving a multimodal input sample consisting of an input image and a text question. The multimodal large language model performs semantic understanding and knowledge requirement analysis on the question. If it determines that its own parameter knowledge cannot support the answer, it enters the search phase. The search phase involves the multimodal large language model generating search instructions containing search keywords and image usage tags. The image usage tags are used to control the multimodal retrieval system to perform plain text retrieval or combined text and image retrieval. The multimodal retrieval system matches the search instructions with a preset knowledge base and returns a set of relevant knowledge fragments. The information acquisition stage involves integrating relevant knowledge fragments obtained in the search stage into the reasoning chain, serving as external knowledge to support subsequent reasoning work. The reflection phase involves: evaluating the relevance and information completeness of the search results; if the search results are irrelevant to the question, reconstructing the search instructions and returning to the search phase; if the search results are relevant but lack sufficient information, retaining the existing evidence and continuing the search; and if the search results meet the answer requirements, proceeding to the summary phase. The summary phase involves integrating all external knowledge acquired through multiple rounds of retrieval. The answering stage: combining the integrated knowledge to generate and output the final question and answer answers.
3. The dual-objective reinforcement learning method for multi-modal agent search of claim 2, wherein, The specific execution method of the search phase in step one is as follows: the search instruction includes a search query and an image usage tag with_image; when the image usage tag is enabled, the multimodal retrieval system performs image-text joint retrieval by combining the input image and the text query; when the image usage tag is disabled, the multimodal retrieval system performs retrieval based solely on the text query; after completing the knowledge base similarity matching, the multimodal retrieval system returns a set of Top-k related knowledge fragments, and encapsulates all knowledge fragments into a preset information module before sending them back to the multimodal large language model.
4. The dual-objective reinforcement learning method for multi-modal agent search of claim 3, wherein, The reflection phase in step one supports multiple rounds of iterative operation, and the complete reasoning trajectory is as follows: thinking module, search module, information module, reflection module, summary module, and answer module. Among them, the search module, information module, and reflection module perform multiple rounds of iterative loops according to the complexity of the question. When the reflection phase determines that the search results are irrelevant to the question, the search query is reconstructed and the process jumps to the search phase. When the search results are determined to be relevant but the information is insufficient, the existing search evidence is retained and the search operation is performed again. When the search results are determined to support the answer, the iteration terminates and the process enters the summary phase.
5. The dual-objective reinforcement learning method for multi-modal agent search of claim 4, wherein, The four types of reward signals in step two are as follows: The retrieval quality reward is determined based on whether the retrieval results contain the target knowledge document, and is used to supervise the quality of retrieval behavior; the answer correctness reward is calculated based on the degree of matching between the model-generated answer and the standard answer, and is used to ensure the accuracy of question-and-answer results; the intra-group search efficiency reward is calculated for multiple correct reasoning paths corresponding to the same question, combined with the number of search calls for each path and the intra-group normalization method, and is used to guide the agent to reduce redundant search behavior; the format reward is calculated based on whether the reasoning trajectory conforms to the preset module format, and is used to constrain the standardization of the reasoning process; The specific calculation rules are as follows: 1) The retrieval quality reward is determined by an indicator function. When the retrieval result set contains the target knowledge document set, a positive reward is output; otherwise, the reward value is 0. 2) The correctness reward uses the Exact Match string matching method to measure and assign values to the degree of matching between the model-generated answer and the standard answer; 3) The search efficiency reward is first calculated based on the number of search calls for a single correct reasoning trajectory, and then the final reward is obtained through intra-group index normalization. 4) The format reward is calculated according to the ratio of "number of compliant modules / preset total number of modules" and is used to verify whether the reasoning trajectory contains complete thinking, searching, information, reflection, summarizing and answering modules.
6. The dual-objective reinforcement learning method for multimodal agent search as described in claim 5, characterized in that, In step three, the local objective function and the global objective function are constructed as follows: For a sample group consisting of multiple reasoning trajectories corresponding to the same question, the mean and standard deviation of the retrieval quality reward within the group are calculated and normalized to obtain the search advantage function; the mean and standard deviation of the global reward within the group are calculated and normalized to obtain the answer advantage function; the two types of advantage functions are used to quantify the quality of a single reasoning trajectory within the group. A local advantage function is constructed based on the retrieval quality reward, and a local objective function is generated. The local objective function only applies to the tokens corresponding to the search phase in the reasoning process, and is used to optimize the agent's search behavior. A global advantage function is constructed based on the answer correctness reward, search efficiency reward, and format reward, and a global objective function is generated. The global objective function applies to all tokens along the entire reasoning trajectory, and is used to optimize the agent's information integration and answer generation capabilities.
7. The dual-objective reinforcement learning method for multimodal agent search as described in claim 6, characterized in that, Step four, performing iterative model training: sample multimodal training samples from the training dataset, and have the multimodal agent generate multiple inference trajectories for a single sample; calculate the four types of reward signals, local dominance function, and global dominance function corresponding to each inference trajectory in sequence; iteratively update the parameters of the multimodal large language model according to the joint overall objective function until the preset training termination condition is reached, and obtain the trained multimodal agent.
8. The dual-objective reinforcement learning method for multimodal agent search as described in claim 7, characterized in that, The model iterative training process in step four specifically includes: 1) Sample multimodal input samples from the training dataset; 2) Control the multimodal agent to generate multiple complete inference trajectories in parallel for the same sample; 3) Calculate the retrieval quality reward, answer correctness reward, format reward, and intra-group search efficiency reward for each reasoning trajectory in sequence; 4) Construct search advantage functions and answer advantage functions based on the four types of rewards respectively; 5) A dynamic weighted annealing strategy is adopted to fuse the local objective function and the global objective function to obtain the joint total objective function, and the parameters of the multimodal large language model are updated according to the joint total objective function; 6) Determine whether the preset training termination condition has been met. If not, return to the sampling step to continue iterating. If it has been met, terminate the training.
9. The dual-objective reinforcement learning method for multimodal agent search as described in claim 8, characterized in that, The joint overall objective function is obtained by multiplying the answer objective function and the search objective function by their respective weight coefficients and then summing them. During the model iterative training process, in the early stage of training, the weight coefficient corresponding to the search objective function is increased to prioritize training the agent's retrieval ability. As the training rounds increase, the weight of the search objective function is gradually reduced and the weight of the answer objective function is increased, ultimately achieving synergistic optimization of the agent's retrieval ability and question answering ability.