Reinforcement Learning Initialization for Delayed Process Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current reinforcement learning methods for process control, such as temperature and liquid level adjustment, require extensive learning time and may not converge to optimal control performance due to random action choices and inappropriate action widths, especially in N-order delay systems like temperature control with long response times.

Innovation Solution

The learning device initializes a machine learning model through preliminary learning before reinforcement learning, using prior knowledge from PID control or manual operations to select actions based on defined options, thereby reducing learning time and improving model accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If reinforcement learning is applied to process control without preliminary learning, then the system can learn optimal control policies, but the learning time becomes excessively long and convergence to optimal performance is not achieved

Engineering Contradiction:
Improvecontrol performanceVSAvoidlearning time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies preliminary learning before reinforcement learning to initialize the machine learning model with prior knowledge from PID control or manual operations. This preliminary action establishes a reasonable starting point for the reinforcement learning process, avoiding random initialization and enabling faster convergence to optimal control performance without excessive learning time.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If random action choices are used in reinforcement learning, then the system can explore the action space, but the learning efficiency decreases and convergence is delayed

Engineering Contradiction:
Improveaction explorationVSAvoidlearning efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent uses preliminary learning to pre-initialize the action selection strategy with prior knowledge before reinforcement learning begins. This allows the system to maintain adaptability and explore the action space effectively while avoiding the inefficiency of purely random action choices, thereby improving learning efficiency and convergence speed.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If inappropriate action widths are used in reinforcement learning, then the system can cover the action space, but the control precision deteriorates and optimal performance is not reached

Engineering Contradiction:
Improveaction coverageVSAvoidcontrol precision
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent applies preliminary learning to initialize the action width parameters with values derived from prior knowledge (PID control or manual operations) before reinforcement learning. This preliminary action ensures that the action width is appropriate from the start, maintaining both adequate action space coverage and high control precision, thereby preventing the deterioration of optimal performance.

Inventive Principle:
Principle #10Preliminary action

4Productivity

If prior knowledge from PID control or manual operations is introduced through preliminary learning, then the learning time is shortened and model accuracy is improved, but the system complexity increases

Engineering Contradiction:
Improvelearning efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent implements a preliminary learning phase that integrates prior knowledge from PID control or manual operations into the machine learning model initialization. While this adds a preliminary learning step, it significantly reduces the overall reinforcement learning time and improves model accuracy, making the increased initial complexity worthwhile for achieving faster convergence and better performance.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP4138005B1Learning device, learning method, learning program, and control
Publication Date: 2024.10.16 YOKOGAWA ELECTRIC CORP
  • EP4138005B1 patent drawingFigure 1
  • EP4138005B1 patent drawingFigure 2
  • EP4138005B1 patent drawingFigure 3

AI summary

Provided is a learning device including: a data acquisition unit configured to acquire, before control of a control target provided in equipment by a machine learning model that outputs an action corresponding to a state of the equipment, initialization data including state data indicating the state of the equipment and action data indicating an action on the control target; and a preliminary learning unit configured to initialize the machine learning model by performing preliminary learning on the basis of the initialization data before start of reinforcement learning corresponding to the control of the control target by the machine learning model. [Selected drawing] Fig. 1