Distributed Agent Learning for Convergent Facility Control Conditions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional machine learning approaches in distributed control systems for optimizing manipulated devices in facilities face long convergence times and often fail to converge, making it impossible to calculate recommended control conditions effectively.

Innovation Solution

The system employs a plurality of agents that acquire state and control condition data to perform learning processing using a model, with each agent focusing on a subset of target devices, utilizing kernel dynamic policy programming and a reward function to output recommended control conditions that increase a preset reward value, thereby optimizing control conditions for each target device.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional machine learning is performed to optimize control conditions for manipulated devices, then control optimization is achieved, but the learning convergence time becomes excessively long and learning may not converge

Engineering Contradiction:
Improvelearning convergenceVSAvoidlearning convergence time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent divides the facility devices into multiple groups and assigns each agent to learn control conditions for a specific group of target devices. This segmentation reduces the number of parameters each agent must learn simultaneously, enabling convergence while maintaining comprehensive control optimization across all devices.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Each agent performs learning processing only for its assigned target devices rather than all devices in the facility. This partial action approach reduces computational complexity and allows learning to converge within reasonable time frames while still achieving overall system optimization through coordinated multi-agent operation.

Inventive Principle:
Principle #16Partial or excessive action

2Productivity

If multiple agents perform learning processing for different target devices, then computational processes are reduced and learning convergence is achieved, but system complexity increases

Engineering Contradiction:
Improvelearning efficiencyVSAvoidmulti-agent system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

Each agent in the multi-agent system is designed with universal functionality to acquire state data, acquire control condition data, perform learning processing, and output recommended control conditions. This standardized multi-functional design reduces overall system complexity despite having multiple agents, as each agent follows the same operational framework.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Manufacturing precision

If all control conditions for all devices are learned simultaneously, then comprehensive optimization is achieved, but redundant learning occurs and resources are inefficiently utilized

Engineering Contradiction:
Improvecontrol condition optimizationVSAvoidcomputational resource efficiency
Core Design Contradiction:
Manufacturing precisionVSLoss of energy

Solution Approach 1:

The patent segments the learning task by device groups, assigning specific target devices to each agent. This eliminates redundant learning of the same device parameters by multiple agents simultaneously, optimizing computational resource efficiency while maintaining comprehensive control optimization through the collective output of all agents.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP3620868B1Apparatus, method, program, and recording medium to output recommended control conditions
Publication Date: 2021.03.10 YOKOGAWA ELECTRIC CORP
  • EP3620868B1 patent drawingFigure 1
  • EP3620868B1 patent drawingFigure 2
  • EP3620868B1 patent drawingFigure 3

AI summary

When simply performing machine learning, there are a huge number of parameters to be learned, the time needed until the learning converges is unrealistically long, and there are cases where the learning does not converge, and therefore it is impossible to calculate the control conditions recommended for the manipulated devices. Provided is an apparatus including a plurality of agents that each set some devices among a plurality of devices provided in a facility to be target devices, wherein each of the plurality of agents includes a state acquiring section that acquires state data indicating a state of the facility; a control condition acquiring section that acquires control condition data indicating a control condition of each target device; and a learning processing section that uses learning data including the state data and the control condition data to perform learning processing of a model that outputs recommended control condition data indicating a control condition recommended for each target device in response to input of the state data.