Distributed Multi-Agent Reinforcement Learning With Skewed JS-Divergence
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Multi-agent reinforcement learning algorithms face challenges in real-world applications due to low learning speed and non-stationarity issues in Decentralized Training with Decentralized Execution (DTDE) environments, particularly in scenarios where agents do not share observation or training information, leading to increased parameters and unpredictable action changes.
Innovation Solution
Quantify non-stationarity using skewed Jensen-Shannon (JS)-divergence to adjust the policy update span of each agent, employing a method that involves initializing policy parameters, exploring skew parameters, and performing stochastic policy training to set a target policy through interpolation, thereby mitigating non-stationarity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If Decentralized Training with Decentralized Execution (DTDE) method is used to maintain agent security and privacy, then agent information sharing is improved, but learning speed deteriorates due to increased parameters and lack of information sharing
Solution Approach 1:
The patent changes the parameter update mechanism by introducing skew parameters that control the degree of policy updates. Instead of uniform updates, agents adjust their policy parameters based on calculated skew parameters that limit the span of updates, thereby maintaining security while improving learning efficiency through controlled parameter evolution.
Solution Approach 2:
The patent introduces dynamic adjustment of update spans through skew parameters. The update span is not fixed but dynamically controlled based on the calculated Jensen-Shannon divergence and predefined maximum values, allowing the system to adaptively balance between maintaining security and improving learning speed.
2Stability of the object's composition
If the span of policy update is reduced to alleviate non-stationarity problem, then policy stability is improved, but training speed deteriorates further
Solution Approach 1:
The patent introduces skew parameters as control variables that adjust the span of policy updates. By changing the parameter update mechanism from fixed to controlled variable updates, the system can maintain policy stability while preventing excessive reduction of training speed through optimized parameter selection.
Solution Approach 2:
The patent implements feedback control by calculating Jensen-Shannon divergence between current and previous policies, comparing it against predefined maximum values, and adjusting skew parameters accordingly. This feedback mechanism ensures policy stability is maintained while avoiding unnecessary constraints that would slow down training.
3Productivity
If quantitative adjustment of update span is implemented to balance training speed and non-stationarity, then training efficiency is improved, but system complexity increases due to additional calculations
Solution Approach 1:
The patent manages system complexity by introducing skew parameters as additional control variables. While this does increase complexity, the structured approach of calculating Jensen-Shannon divergence and using predefined maximum values provides a systematic framework that balances the added complexity with improved training efficiency through quantitative control.
Data Source
AI summary
Disclosed herein is an apparatus and method for distributed multi-agent reinforcement learning. The method may include exploring a skew parameter at which a skewed Jensen-Shannon-(JS-)divergence, which is a change in a policy, becomes equal to or greater than a preassigned maximum value of a skewed JS-divergence stationarity and performing training based on a target policy set using the skew parameter.


