Multi-Agent Facility Control Learning for Faster Convergence
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional machine learning approaches in distributed control systems for optimizing manipulated devices in facilities face long convergence times and often fail to converge, making it impossible to calculate recommended control conditions effectively.
Innovation Solution
The system employs a plurality of agents that acquire state and control condition data to perform learning processing using kernel dynamic policy programming, outputting recommended control conditions that increase a reward value beyond a reference, thereby optimizing device control and reducing computational processes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional machine learning is performed to optimize control conditions of manipulated devices, then control optimization is achieved, but the learning convergence time becomes unrealistically long and learning may not converge
Solution Approach 1:
The patent segments the facility into multiple regions, with each region having its own agent that performs learning processing independently. This divides the large-scale learning problem into smaller, manageable sub-problems that converge faster. Each agent learns control conditions for devices in its specific region rather than learning all facility devices simultaneously, which resolves the contradiction between achieving comprehensive control optimization and maintaining reasonable convergence time.
Solution Approach 2:
Each agent performs learning processing for only a subset of devices in its region rather than all devices in the facility. This partial action approach allows learning to converge within realistic timeframes while still achieving effective control optimization for the covered devices. The patent accepts that not all devices are optimized by every agent, trading complete coverage for practical convergence.
2Productivity
If multiple agents perform learning processing for different devices, then distributed processing efficiency is improved, but redundant processing occurs when multiple agents learn the same device
Solution Approach 1:
The patent assigns specific regions to specific agents, creating a territorial division where each agent has exclusive learning responsibility for devices in its region. This local quality approach ensures that each device is learned by only one agent, eliminating redundant processing while maintaining distributed processing efficiency. The region assignment creates clear boundaries that prevent overlap and wasted computational resources.
Data Source
AI summary
Provided is an apparatus including a plurality of agents that each set some devices among a plurality of devices provided in a facility to be target devices, wherein each of the plurality of agents includes a state acquiring section that acquires state data indicating a state of the facility; a control condition acquiring section that acquires control condition data indicating a control condition of each target device; and a learning processing section that uses learning data including the state data and the control condition data to perform learning processing of a model that outputs recommended control condition data indicating a control condition recommended for each target device in response to input of the state data.


