Multi-Agent Reinforcement Learning Agent Grouping for Stable Convergence

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In large-scale manufacturing facilities, multi-agent reinforcement learning (MARL) systems face challenges in achieving convergence and stability due to non-stationary environments and variable data distributions, leading to unpredictable agent behavior and reduced operational efficiency.

Innovation Solution

A method involving offline reinforcement learning training with least-squared temporal difference, followed by policy constraint updates using heuristic rule-based policies, and combining converged and non-converged agents to stabilize agent behavior and ensure convergence.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple RL-based scheduling techniques are utilized through multi-agent RL (MARL) based schedulers, then the ability to handle complex production increases, but coordination between agents becomes difficult and convergence is hard to achieve

Engineering Contradiction:
Improveability to handle complex productionVSAvoidconvergence and stability of agents
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent segments the MARL system into different agent groups (first group, second group, third group) with distinct roles. Converging agents are separated from non-converging agents, allowing independent training and policy updates. This segmentation enables complex production handling while maintaining stability by managing agent coordination in controlled groups rather than as a monolithic system.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary identification of converging agents before full deployment. By pre-training agents offline using historic data and identifying which agents achieve convergence, the system prepares stable agents in advance. This preliminary action ensures that only proven stable agents are deployed together, preventing coordination failures and achieving both versatility and reliability.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If human-based scheduling approach is used, then expertise can be applied, but it becomes difficult to coordinate as facility size grows and relies on years of training

Engineering Contradiction:
Improvescheduling qualityVSAvoidcoordination complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements self-service through automated agent identification and grouping. The system automatically identifies converging agents, separates them into appropriate groups, and manages their training without human intervention. This eliminates the need for years of human training while handling coordination complexity automatically, maintaining scheduling quality through algorithmic decision-making rather than human expertise.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical human-based coordination system with an automated computational system. Instead of relying on human supervisors to manually coordinate agents, the system uses automated algorithms to identify converging agents, manage groupings, and update policies. This substitution handles coordination complexity through computation rather than human cognitive effort, scaling effectively with facility size.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Reliability

If offline reinforcement learning training is performed, then agent stability improves, but training time and computational resources increase

Engineering Contradiction:
Improveagent stabilityVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies partial action by performing offline reinforcement learning training only for the first group of agents initially, rather than training all agents simultaneously. After identifying converging agents, subsequent groups are trained with reduced scope using the already-stabilized first group as a foundation. This partial training approach achieves necessary agent stability while reducing total training time and computational resources compared to comprehensive full-system training.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20250284969A1Systems and methods for stabilization of multi-agent reinforcement learning
Publication Date: 2025.09.11 SAMSUNG DISPLAY CO LTD
  • US20250284969A1 patent drawing
  • US20250284969A1 patent drawing
  • US20250284969A1 patent drawing

AI summary

A manufacturing system may include a processor and a memory storing instructions executed by the processor to cause the processor to identify one or more converging agents from a first group of agents, in response to the identification of the one or more converging agents, perform training of one or more non-converging agents of the first group of agents using historic data to form a second group of agents, identify one or more converging agents from the second group of agents, in response to the identification of the one or more converging agents from the second group of agents, update policy of the one or more non-converging agents of the second group of agents based on historic data collected by heuristic rule-based policy to form a third group of agents, and deploy, a policy based on the first, second, and third group of agents.