Multi-Agent Facility Control Learning for Faster Convergence

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional machine learning approaches in distributed control systems for optimizing manipulated devices in facilities face long convergence times and often fail to converge, making it impossible to calculate recommended control conditions effectively.

Innovation Solution

The system employs a plurality of agents that acquire state and control condition data to perform learning processing using kernel dynamic policy programming, outputting recommended control conditions that increase a reward value beyond a reference, thereby optimizing device control and reducing computational processes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional machine learning is performed to optimize control conditions of manipulated devices, then control optimization is achieved, but the learning convergence time becomes unrealistically long and learning may not converge

Engineering Contradiction:
Improvelearning convergenceVSAvoidlearning convergence time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments the facility into multiple regions, with each region having its own agent that performs learning processing independently. This divides the large-scale learning problem into smaller, manageable sub-problems that converge faster. Each agent learns control conditions for devices in its specific region rather than learning all facility devices simultaneously, which resolves the contradiction between achieving comprehensive control optimization and maintaining reasonable convergence time.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Each agent performs learning processing for only a subset of devices in its region rather than all devices in the facility. This partial action approach allows learning to converge within realistic timeframes while still achieving effective control optimization for the covered devices. The patent accepts that not all devices are optimized by every agent, trading complete coverage for practical convergence.

Inventive Principle:
Principle #16Partial or excessive action

2Productivity

If multiple agents perform learning processing for different devices, then distributed processing efficiency is improved, but redundant processing occurs when multiple agents learn the same device

Engineering Contradiction:
Improvedistributed processing efficiencyVSAvoidredundant processing
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent assigns specific regions to specific agents, creating a territorial division where each agent has exclusive learning responsibility for devices in its region. This local quality approach ensures that each device is learned by only one agent, eliminating redundant processing while maintaining distributed processing efficiency. The region assignment creates clear boundaries that prevent overlap and wasted computational resources.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11119451B2Apparatus, method, program, and recording medium
Publication Date: 2021.09.14 YOKOGAWA ELECTRIC CORP
  • US11119451B2 patent drawing
  • US11119451B2 patent drawing
  • US11119451B2 patent drawing

AI summary

Provided is an apparatus including a plurality of agents that each set some devices among a plurality of devices provided in a facility to be target devices, wherein each of the plurality of agents includes a state acquiring section that acquires state data indicating a state of the facility; a control condition acquiring section that acquires control condition data indicating a control condition of each target device; and a learning processing section that uses learning data including the state data and the control condition data to perform learning processing of a model that outputs recommended control condition data indicating a control condition recommended for each target device in response to input of the state data.