Implementation method and system based on multi-agent strategy diversification

Through the adaptive convergence detection and historical strategy repository architecture, combined with MAPPO and SVGD algorithm, the multi-agent strategy is dynamically optimized, which solves the problems of exploring capacity loss and local optimization in traditional methods, and improves the overall performance and computing efficiency of the multi-agent system.

CN120258035APending Publication Date: 2025-07-04NANKAI UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510337217.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

After convergence, traditional multi-agent reinforcement learning algorithms lose their exploration capabilities, low exploration efficiency, and are prone to falling into local optimality, making it difficult to adapt to complex tasks and dynamic environments, resulting in a degradation of system performance.

Method used

Adaptive convergence detection, historical strategy repository architecture and a framework integrating MAPPO and SVGD, convergence through sliding window detection strategy, dynamically switch to the diversity-driven optimization mode, and combined with Stein's variable gradient descent algorithm-driven strategy to unexplored spatial optimization.

Benefits of technology

It significantly improves the diversity, adaptability and computing efficiency of multi-agent strategies, avoids local optimal traps, achieves efficient exploration and utilization balance, and reduces training costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258035A_ABST
    Figure CN120258035A_ABST
Patent Text Reader

Abstract

The invention relates to an implementation method and system based on multi-agent strategy diversification, and belongs to the technical field of multi-agent reinforcement learning. The method comprises the following steps of: firstly, performing strategy updating by adopting a multi-agent near-end strategy optimization framework in an exploration stage, and realizing global value evaluation and local action generation through a centralized training and decentralized execution framework; meanwhile, calculating a variable coefficient of a reward sequence through a sliding window, dynamically detecting a strategy convergence condition, and if the strategy is converged, collecting a current exploration strategy into a historical strategy library; and finally, based on a Stein variational gradient descent algorithm, a strategy particle set is extracted from a historical strategy storage library, the strategy is driven to be optimized to an unexplored behavior space, and local optimum is broken through. According to the method, the diversity, the adaptability and the calculation efficiency of a multi-agent strategy are remarkably improved through self-adaptive convergence detection, a historical strategy storage library framework and a complete framework integrating MAPPO and SVGD.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of multi-agent reinforcement learning, and particularly relates to a method and system for realizing multi-agent policy diversification. Background Art

[0002] Multi-Agent Reinforcement Learning (MARL) is the core technology of distributed intelligent systems and is widely used in collaborative tasks, competitive tasks, and mixed task scenarios. However, traditional multi-agent reinforcement learning algorithms (such as MAPPO) have the following problems in practical applications: 1. Loss of exploration ability after convergence: After traditional algorithms converge, the behavior patterns of agents tend to be single, resulting in their inability to adapt to changes in the dynamic environment. Due to the lack of an effective exploration mechanism, it is difficult for agents to quickly adjust their strategies when the environment changes, leading to a decline in system performance. 2. Low exploration efficiency: Existing methods usually rely on manually designed reward functions or complex evolutionary strategies to guide agent exploration. However, it is difficult for manually designed reward functions to meet the requirements of complex tasks, and the design cost is relatively high. In addition, although evolutionary algorithms can achieve policy diversification by eliminating old policies, they consume a large amount of computing resources, have low exploration efficiency, and are difficult to balance the relationship between exploration and exploitation, resulting in low data utilization. 3. Local optimum trap:

[0003] Traditional algorithms are prone to falling into a non-Pareto optimal Nash equilibrium after convergence, resulting in limited overall system performance. Due to the serious homogenization of agent policies, it is difficult for the system to break through the local optimum and reach the global optimum solution, thus limiting the overall performance of the multi-agent system.

[0004] To solve the above problems, the prior art usually adopts the following methods: 1. Reward design: Guide agent exploration by manually designing a reward function, but this method cannot automatically adapt to complex tasks, has a high design cost, and is difficult to cope with changes in the dynamic environment. 2. Evolutionary algorithm: Regard the agents in multi-agent as the population in the evolutionary algorithm, and achieve policy diversification by continuously eliminating old policies. However, this method consumes a large amount of computing resources, has low exploration efficiency, and is difficult to achieve efficient policy optimization in complex tasks.

[0005] Although the prior art can improve the diversity of agent policies to a certain extent, there are still the following deficiencies: First, it cannot be automatically adapted: relying on manually designed reward functions or complex evolutionary strategies, it is difficult to meet the requirements of complex tasks. Second, it consumes a large amount of computing resources: traditional methods such as evolutionary algorithms require a large amount of computing resources and have low exploration efficiency. Finally, it is difficult to balance exploration and exploitation: existing methods are difficult to find a balance between exploring new policies and exploiting existing policies, resulting in low data utilization.

[0006] Therefore, how to improve the exploration efficiency of multi-agent strategies, avoid local optimal traps, and achieve policy diversification is a technical problem that urgently needs to be solved in the field of multi-agent reinforcement learning. The present invention aims to provide a method and system for realizing multi-agent policy diversification to solve the above problems and improve the overall performance of multi-agent systems. Summary of the Invention

[0007] The purpose of the present invention is to provide a method and system for realizing multi-agent policy diversification. Through an adaptive convergence detection, a historical policy repository architecture, and a complete framework integrating MAPPO and SVGD, the diversity, adaptability, and computational efficiency of multi-agent policies are significantly improved, which has important theoretical value and practical application significance.

[0008] To achieve the above object, the present invention adopts the following technical solutions: A method for realizing multi-agent policy diversification, the method comprising the following steps:

[0009] Step 1: Exploration phase (CV≥a):

[0010] Use a standard multi-agent proximal policy optimization (MAPPO) framework for policy update, and realize global value evaluation and local action generation through a centralized training and decentralized execution (CTDE) architecture;

[0011] Constrain the policy update amplitude through a proximal policy optimization (PPO) objective function to ensure update stability. The proximal policy optimization (PPO) objective function is:

[0012] L CLIP (θ) = E t [min(r t (θ)A t , clip(r t (θ), 1 - ∈, 1 + ∈)A t )]

[0013] The centralized Critic network evaluates the global value function, and the decentralized Actor network generates local actions;

[0014] Step 2: Adaptive convergence detection (CV < a):

[0015] Use a sliding window to calculate the coefficient of variation CV of the reward sequence. The size of the sliding window W is W = 20, and the calculation formula is:

[0016]

[0017] When the coefficient of variation CV is lower than a preset threshold a, it is determined that the policy converges, archive the current policy parameters and gradients to the historical policy repository, and dynamically switch to the diversity-driven optimization mode;

[0018] Step 3: Diversity-driven optimization:

[0019] Extract a set of policy particles from the historical policy repository, calculate the policy update direction based on the Stein Variational Gradient Descent (SVGD) algorithm, and the update formula is:

[0020]

[0021] Drive the policy to optimize towards the unexplored behavior space to break through the local optimum.

[0022] The size W of the sliding window described in the present invention can be dynamically adjusted according to task requirements to achieve adaptive convergence detection; the value of the preset threshold a is dynamically adjusted according to task complexity and convergence requirements.

[0023] The historical policy repository described in the present invention is used to archive historical policy parameters and gradients, serving as a reference particle set for the SVGD algorithm to support diversity-driven policy optimization.

[0024] The method described in the present invention adopts a Centralized Training and Decentralized Execution (CTDE) architecture, integrating the MAPPO and SVGD algorithms to achieve efficient convergence and diversity balance.

[0025] The structures of the centralized Critic network and the decentralized Actor network described in the present invention can be designed according to specific task requirements to support the global value evaluation and local action generation of multi-agent systems.

[0026] The Stein Variational Gradient Descent (SVGD) algorithm described in the present invention calculates the similarity between policy particles π i , π j ), and drives the policy to update in other directions to increase policy diversity. i , π j by means of the kernel function k(π

[0027] The method described in the present invention is applicable to collaborative tasks, competitive tasks, and mixed task scenarios in multi-agent reinforcement learning.

[0028] Another object of the present invention is to provide a system for realizing multi-agent policy diversification, including: a policy update module for performing policy updates in the exploration stage, adopting the MAPPO framework to implement Centralized Training and Decentralized Execution (CTDE); an adaptive convergence detection module for dynamically detecting policy convergence based on the coefficient of variation (CV) of the sliding window and switching to the diversity-driven optimization mode; a historical policy storage module: for archiving historical policy parameters and gradients as a reference particle set for the SVGD algorithm; a diversity optimization module: for driving the policy to optimize towards the unexplored behavior space based on the SVGD algorithm to break through the local optimum.

[0029] The present invention also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

[0030] The present invention also provides an electronic device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

[0031] The design idea of the present invention is as follows: First, the initial exploration of the agent strategy is realized through the MAPPO algorithm. After obtaining the basic strategy of the agent exploration, the coefficient of variation of the agent strategy is calculated to detect the convergence of the agent strategy, and then the convergence of the agent strategy is put into the historical strategy library. SVGD is introduced as the driving force for the agent strategy exploration to drive the agent to explore other strategies. Through the present invention, it is possible to achieve multi-agent strategy breakthrough in strategy homogenization, SVGD drives the strategy to optimize the uncovered behavior space to avoid local optimum; adaptive exploration: the convergence detection mechanism dynamically switches the update mode without manual intervention; efficient calculation: combining the sampling efficiency of MAPPO and the multi-modal optimization ability of SVGD to reduce the training cost; wide applicability: it can be extended to multi-agent scenarios such as UAV formations and traffic control.

[0032] The basic idea of this method is: First, the initial exploration of the agent strategy is realized through the MAPPO algorithm. After obtaining the basic strategy of the agent exploration, the coefficient of variation of the agent strategy is calculated to detect the convergence of the agent strategy, and then the convergence of the agent strategy is put into the historical strategy library. SVGD is introduced as the driving force for the agent strategy exploration to drive the agent to explore other strategies. Through the present invention, it is possible to achieve multi-agent strategy breakthrough in strategy homogenization, SVGD drives the strategy to optimize the uncovered behavior space to avoid local optimum; adaptive exploration: the convergence detection mechanism dynamically switches the update mode without manual intervention; efficient calculation: combining the sampling efficiency of MAPPO and the multi-modal optimization ability of SVGD to reduce the training cost; wide applicability: it can be extended to multi-agent scenarios such as UAV formations and traffic control.

[0033] The beneficial effects of adopting the technical solution of the present invention are as follows: 1. By introducing the Stein variational gradient descent (SVGD) algorithm and the historical policy storage mechanism, the present invention dynamically quantifies the policy differences, drives the agent to explore the uncovered behavior space, significantly improves the policy diversity, and avoids the problem of homogeneous agent policies; by storing the historical policy parameters and gradients in the historical policy storage inventory as the reference particle set of the SVGD algorithm, the present invention dynamically quantifies the policy differences, guides the agent to explore the uncovered behavior space, and further improves the policy diversity. 2. As the driving force for policy exploration, the SVGD algorithm of the present invention can guide the agent to break through the local optimal trap and explore a better policy space, thereby improving the overall performance of the system; under the centralized training and decentralized execution (CTDE) architecture, the MAPPO and SVGD algorithms are integrated to achieve efficient convergence and diversity balance, and solve the problem that it is difficult for traditional methods to balance exploration and exploitation. 3. The present invention dynamically detects the policy convergence situation based on the coefficient of variation (CV) of the sliding window and automatically switches to the diversity-driven optimization mode without manual intervention, realizing the adaptive exploration of the agent policy. 4. By combining the sampling efficiency of the MAPPO algorithm and the multi-modal optimization ability of the SVGD algorithm, the present invention significantly reduces the training cost, improves the data utilization rate and the calculation efficiency. The present invention can be extended to various multi-agent application scenarios, such as unmanned aerial vehicle formation, traffic control, robot cooperation, etc., and has wide applicability and practicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0036] Embodiment 1

[0037] A method for realizing multi-agent policy diversification, the method comprising the following steps:

[0038] Step 1: Exploration stage (CV≥a):

[0039] 1. Initialization:

[0040] Initialize the centralized Critic network and the decentralized Actor network of the multi-agent system.

[0041] Set the sliding window size W = 20 and the preset threshold a = 0.1 (which can be adjusted according to task requirements).

[0042] 2. Policy update:

[0043] The standard Multi-Agent Proximal Policy Optimization (MAPPO) framework is adopted for policy update.

[0044] The magnitude of policy update is constrained by the Proximal Policy Optimization (PPO) objective function to ensure update stability. The PPO objective function is:

[0045] L CLIP (θ) = E t [min(r t (θ)A t , clip(r t (θ), 1 - ∈, 1 + ∈)A t )]

[0046] The centralized Critic network evaluates the global value function, and the decentralized Actor network generates local actions.

[0047] 3. Reward calculation: At each time step t, the reward sequence Ri of the agent is recorded, and the average reward μR and standard deviation σR within the sliding window are calculated:

[0048]

[0049]

[0050] The coefficient of variation (CV) is calculated:

[0051]

[0052] Step 2: Adaptive convergence detection (CV < a):

[0053] 1. Convergence judgment:

[0054] When the coefficient of variation (CV) is lower than the preset threshold a, the policy is determined to converge.

[0055] The current policy parameters and gradients are archived to the historical policy repository.

[0056] 2. Mode switching:

[0057] Dynamically switch to the diversity-driven optimization mode to prepare for policy diversification exploration.

[0058] Step 3: Diversity-driven optimization:

[0059] Policy particle extraction:

[0060] Randomly extract a set of policy particles {π j} from the historical policy repository as the reference particle set for the SVGD algorithm.

[0061] Policy update direction calculation:

[0062] Calculate the policy update direction based on the Stein Variational Gradient Descent (SVGD) algorithm, and the update formula is:

[0063]

[0064] where k(π i , π j ) is the kernel function used to calculate the similarity between policy particles.

[0065] 3. Policy optimization:

[0066] Drive the policy to optimize in the unexplored behavior space and break through the local optimum.

[0067] Regarding the parameter adjustment and task adaptation in the method of this embodiment

[0068] 1. Adjustment of the sliding window size: Dynamically adjust the sliding window size W according to the task requirements. For example, in complex tasks, W can be increased to improve the detection accuracy.

[0069] 2. Dynamic adjustment of the threshold: Dynamically adjust the preset threshold a according to the task complexity and convergence requirements. For example, in a dynamic environment, a can be reduced to improve the sensitivity.

[0070] 3. Network structure design: Design the structures of the centralized Critic network and the decentralized Actor network according to the specific task requirements to support the global value evaluation and local action generation of the multi-agent system.

[0071] Application scenarios of the method of this embodiment:

[0072] 1. UAV formation: Apply the present invention to the UAV formation task, and through adaptive convergence detection and diversity-driven optimization, achieve efficient cooperation and dynamic adaptation of the UAV formation.

[0073] 2. Traffic control: Apply the present invention to the traffic control task, and drive policy optimization through the SVGD algorithm to improve the overall efficiency of traffic flow.

[0074] 3. Robot cooperation: Apply the present invention to the robot cooperation task, and through the historical policy repository and the SVGD algorithm, achieve diverse exploration and efficient cooperation of robot policies.

[0075] Embodiment 2

[0076] A multi-agent policy diversification implementation system, including:

[0077] Policy update module: Implement centralized training and decentralized execution (CTDE) using the MAPPO framework to complete the policy update in the exploration stage;

[0078] Adaptive Convergence Detection Module: Dynamically detect the convergence of the mutation coefficient (CV) based on a sliding window and switch to the diversity-driven optimization mode;

[0079] Historical Policy Storage Module: Archive historical policy parameters and gradients as the reference particle set for the SVGD algorithm;

[0080] Diversity Optimization Module: Drive the policy to optimize the unexplored behavior space based on the SVGD algorithm to break through the local optimum.

[0081] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in Embodiment 1 is implemented.

[0082] An electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method in Embodiment 1 is implemented.

[0083] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: the specific implementation manners of the present invention can still be modified or equivalently replaced, and any modification or equivalent replacement without departing from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. A method for implementing multi-agent strategy diversification, characterized in that, The method includes the following steps: Step 1: Exploration phase: The standard multi-agent proximal policy optimization framework is adopted for policy update, and global value evaluation and local action generation are realized through a centralized training and decentralized execution architecture; The magnitude of policy update is constrained by the proximal policy optimization objective function to ensure update stability. The proximal policy optimization objective function is: L CLIP (θ) = E t [min(r t (θ)A t , clip(r t (θ), 1 - ∈, 1 + ∈)A t )] The centralized Critic network evaluates the global value function, and the decentralized Actor network generates local actions; Step 2: Adaptive convergence detection: The coefficient of variation CV of the reward sequence is calculated using a sliding window. The size of the sliding window W is W = 20, and the calculation formula is: When the coefficient of variation CV is lower than the preset threshold a, it is determined that the policy converges. The current policy parameters and gradients are archived in the historical policy repository, and the mode is dynamically switched to diversity-driven optimization; Step 3: Diversity-driven optimization: A set of policy particles is extracted from the historical policy repository, and the policy update direction is calculated based on the Stein variational gradient descent (SVGD) algorithm. The update formula is: Drive the policy to optimize in the unexplored behavior space to break through the local optimum.

2. The implementation method based on multi-agent strategy diversification according to claim 1, characterized in that The size W of the sliding window can be dynamically adjusted according to task requirements to achieve adaptive convergence detection; the value of the preset threshold a is dynamically adjusted according to task complexity and convergence requirements.

3. A method for implementing a multi-agent strategy diversification according to claim 1, characterized in that, The historical policy repository is used to archive historical policy parameters and gradients as the reference particle set for the SVGD algorithm, supporting diversity-driven policy optimization.

4. A method for implementing a multi-agent strategy diversification according to any one of claims 1-3, characterized in that The method adopts a centralized training and decentralized execution architecture, integrating the MAPPO and SVGD algorithms to achieve efficient convergence and diversity balance.

5. A method for implementing multi-agent policy diversification according to any one of claims 1-3, characterized in that The structures of the centralized Critic network and the decentralized Actor network can be designed according to specific task requirements to support global value evaluation and local action generation in multi-agent systems.

6. A method for implementing a multi-agent strategy diversification according to any one of claims 1-3, characterized in that The Stein variational gradient descent (SVGD) algorithm calculates the similarity between policy particles π i , π j through the kernel function k(π i , π j ), and drives the policy to update in other directions to increase policy diversity.

7. A method for implementing a multi-agent strategy diversification according to any one of claims 1-3, characterized in that, The method is applicable to collaborative tasks, competitive tasks, and mixed task scenarios in multi-agent reinforcement learning.

8. An implementation system based on multi-agent strategy diversification, characterized in that, Including: A policy update module for performing policy update in the exploration phase and implementing centralized training and decentralized execution (CTDE) using the MAPPO framework; An adaptive convergence detection module for dynamically detecting policy convergence based on the coefficient of variation CV of the sliding window and switching to the diversity-driven optimization mode; A historical policy storage module for archiving historical policy parameters and gradients as the reference particle set for the SVGD algorithm; A diversity optimization module for driving the policy to optimize in the unexplored behavior space based on the SVGD algorithm to break through the local optimum.

9. A computer-readable storage medium, characterized in that, Stored with a computer program, when the computer program is executed by a processor, it implements the method described in any one of claims 1 to 7.

10. An electronic device, characterized in that, Including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the method described in any one of claims 1 to 7.

Citation Information

Cited By

  • Self-adaptive strategy pruning method and device for balancing production efficiency

    CN121030762A

  • Intelligent factory environment adaptive access control method and device based on deep reinforcement learning

    CN121030762B