Reinforcement Learning Model Selection for Nonlinear Equipment Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing control systems for equipment like distillation apparatuses face challenges in achieving improved controllability, particularly in nonlinear operations, relying heavily on operator experience and manual adjustments for valve control, which limits efficiency and effectiveness in processes like quality assurance, energy saving, and yield improvement.
Innovation Solution
A model selection apparatus and method that utilizes reinforcement learning to generate and select operation models based on evaluation indicators, allowing for autonomous AI-controlled equipment operation by evaluating state data and selecting the most effective operation model for controlling equipment, such as valves, through a system comprising an evaluation model, operation model management, and model selection units.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If multiple candidate models are generated by reinforcement learning to improve control performance, then controllability and operational efficiency are improved, but device complexity and computational resources increase
Solution Approach 1:
The control system is segmented into multiple candidate models (first, second, third models) each trained for specific operational conditions. This segmentation allows the system to divide complex control tasks into manageable segments, improving operational efficiency while keeping each individual model relatively simple
Solution Approach 2:
The system dynamically selects and switches between different candidate models based on real-time operational conditions and evaluation results. This dynamic adaptation allows the system to optimize control performance for varying conditions without requiring a single overly complex model to handle all scenarios
2Ease of operation
If manual adjustments and operator experience are used for valve control, then ease of operation is maintained, but productivity and precision in nonlinear operations deteriorate
Solution Approach 1:
The control system performs self-optimization through automated model selection and switching based on evaluation results. The system serves itself by automatically adapting to operational conditions without requiring manual intervention, thereby improving productivity while maintaining ease of operation through the automated decision-making process
3Device complexity
If a single operation model is used to simplify the control system, then device complexity is reduced, but adaptability to different operational conditions deteriorates
Solution Approach 1:
Multiple candidate models are trained to handle different operational conditions and scenarios. Each model serves as a specialized component that can be universally applied when its specific conditions are met, providing adaptability across diverse operations while keeping the overall system structure manageable through standardized model architecture
4Manufacturing precision
If reinforcement learning is used to generate operation models, then manufacturing precision and control accuracy are improved, but loss of time and computational resources increase
Solution Approach 1:
Multiple candidate models are pre-trained offline using reinforcement learning before deployment. This preliminary action allows the system to perform computationally intensive training work in advance, reducing real-time computational requirements and enabling fast model selection during actual operation without significant time loss
Data Source
AI summary
Provided is a model selection apparatus including a candidate model storage unit configured to store a plurality of candidate models each of which is generated by reinforcement learning using, as at least a part of a reward, an output of an evaluation model configured to output an indicator obtained by evaluating a state of equipment and is capable of output an action, a state data acquisition unit configured to acquire a plurality of pieces of state data representing the state of the equipment in a case where each of manipulated variables based on outputs of the plurality of candidate models is applied to a controlled object in the equipment, an indicator acquisition unit configured to acquire a plurality of indicators, a model selection unit configured to select an object model for controlling the controlled object, and an object model output unit configured to output the object model.


