This invention discloses a multi-agent collaborative
reasoning system and method. The
system includes: a dynamic policy coordination module for analyzing the original problem input by the user and determining initial policy parameters; a multi-round reasoning module for sending the original problem and initial policy parameters to a
pool of agents, obtaining reasoning data from each agent, performing multiple rounds of interaction until round control conditions are met, and determining collaborative reasoning data based on the latest reasoning data; obtaining updated reasoning data from each agent during each round of interaction; a parameter optimization module for calculating trajectory scores on the reasoning data from the previous round for each agent, determining the current
interaction mode and current interaction parameters, and updating the trajectory
score weights; and a communication optimization module for
processing the reasoning data from each agent and sending it to other agents. This invention enables multi-agent
collaboration that organically integrates efficient interaction, deep reasoning evaluation, and optimal result selection.