Reinforcement Learning Initialization for Delayed Process Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current reinforcement learning methods for process control, such as temperature and liquid level adjustment, require extensive learning time and may not converge to optimal control performance due to random action choices and inappropriate action widths, especially in N-order delay systems like temperature control with long response times.
Innovation Solution
The learning device initializes a machine learning model through preliminary learning before reinforcement learning, using prior knowledge from PID control or manual operations to select actions based on defined options, thereby reducing learning time and improving model accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If reinforcement learning is applied to process control without preliminary learning, then the system can learn optimal control policies, but the learning time becomes excessively long and convergence to optimal performance is not achieved
Solution Approach 1:
The patent applies preliminary learning before reinforcement learning to initialize the machine learning model with prior knowledge from PID control or manual operations. This preliminary action establishes a reasonable starting point for the reinforcement learning process, avoiding random initialization and enabling faster convergence to optimal control performance without excessive learning time.
2Adaptability or versatility
If random action choices are used in reinforcement learning, then the system can explore the action space, but the learning efficiency decreases and convergence is delayed
Solution Approach 1:
The patent uses preliminary learning to pre-initialize the action selection strategy with prior knowledge before reinforcement learning begins. This allows the system to maintain adaptability and explore the action space effectively while avoiding the inefficiency of purely random action choices, thereby improving learning efficiency and convergence speed.
3Adaptability or versatility
If inappropriate action widths are used in reinforcement learning, then the system can cover the action space, but the control precision deteriorates and optimal performance is not reached
Solution Approach 1:
The patent applies preliminary learning to initialize the action width parameters with values derived from prior knowledge (PID control or manual operations) before reinforcement learning. This preliminary action ensures that the action width is appropriate from the start, maintaining both adequate action space coverage and high control precision, thereby preventing the deterioration of optimal performance.
4Productivity
If prior knowledge from PID control or manual operations is introduced through preliminary learning, then the learning time is shortened and model accuracy is improved, but the system complexity increases
Solution Approach 1:
The patent implements a preliminary learning phase that integrates prior knowledge from PID control or manual operations into the machine learning model initialization. While this adds a preliminary learning step, it significantly reduces the overall reinforcement learning time and improves model accuracy, making the increased initial complexity worthwhile for achieving faster convergence and better performance.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Provided is a learning device including: a data acquisition unit configured to acquire, before control of a control target provided in equipment by a machine learning model that outputs an action corresponding to a state of the equipment, initialization data including state data indicating the state of the equipment and action data indicating an action on the control target; and a preliminary learning unit configured to initialize the machine learning model by performing preliminary learning on the basis of the initialization data before start of reinforcement learning corresponding to the control of the control target by the machine learning model. [Selected drawing] Fig. 1