Multi-agent cooperative strategy adjustment method and device, electronic equipment and medium

By employing a multi-round negotiation mechanism and a dynamic reward mechanism, the problem of low collaboration efficiency in multi-agent collaboration is solved, enabling agents to autonomously optimize and continuously collaborate in multi-round interactions, thereby improving task processing quality and overall system efficiency.

CN122433779APending Publication Date: 2026-07-21TENCENT TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TENCENT TECH (BEIJING) CO LTD
Filing Date
2026-03-30
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing multi-agent collaboration methods suffer from low collaboration efficiency, low task success rate, and a lack of dynamic adaptability and long-term collaborative capabilities.

Method used

Through a multi-round negotiation mechanism, the answer quality assessment value and the collaboration impact assessment value are calculated to generate correction rewards and collaboration contribution rewards, dynamically adjust the agent's collaboration strategy, and introduce a multi-dimensional dynamic reward mechanism and a collaboration contribution quantification system.

Benefits of technology

It enables agents to autonomously revise and continuously optimize their collaboration strategies during multiple rounds of negotiation, enhancing the collaboration capabilities among multi-agent systems, forming an adaptive division of labor mechanism, and significantly improving task processing quality and robustness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122433779A_ABST
    Figure CN122433779A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, and in particular to a multi-agent cooperation strategy adjustment method and device, electronic equipment and a medium, which are used to improve the cooperation capability among multiple agents. The method comprises the following steps: obtaining initial answers independently generated by multiple agents according to input target tasks; triggering the multiple agents to optimize the initial answers through multiple rounds of negotiation; and performing the following strategy adjustment operation according to the target answers generated by the multiple agents in each round of negotiation: for each agent, calculating an answer quality evaluation value and a cooperation influence evaluation value based on the target answer of the agent in the current round after each round of negotiation, and generating a corresponding correction reward when the target answer is corrected; generating the cooperation contribution reward of the agent by cumulatively processing the answer quality evaluation values and the cooperation influence evaluation values in each round; and adjusting the cooperation strategy of the agent according to the obtained cooperation contribution rewards and correction rewards.
Need to check novelty before this filing date? Find Prior Art