Reinforcement Learning Model Selection for Nonlinear Equipment Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing control systems for equipment like distillation apparatuses face challenges in achieving improved controllability, particularly in nonlinear operations, relying heavily on operator experience and manual adjustments for valve control, which limits efficiency and effectiveness in processes like quality assurance, energy saving, and yield improvement.

Innovation Solution

A model selection apparatus and method that utilizes reinforcement learning to generate and select operation models based on evaluation indicators, allowing for autonomous AI-controlled equipment operation by evaluating state data and selecting the most effective operation model for controlling equipment, such as valves, through a system comprising an evaluation model, operation model management, and model selection units.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If multiple candidate models are generated by reinforcement learning to improve control performance, then controllability and operational efficiency are improved, but device complexity and computational resources increase

Engineering Contradiction:
Improveoperational efficiencyVSAvoiddevice complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The control system is segmented into multiple candidate models (first, second, third models) each trained for specific operational conditions. This segmentation allows the system to divide complex control tasks into manageable segments, improving operational efficiency while keeping each individual model relatively simple

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically selects and switches between different candidate models based on real-time operational conditions and evaluation results. This dynamic adaptation allows the system to optimize control performance for varying conditions without requiring a single overly complex model to handle all scenarios

Inventive Principle:
Principle #15Dynamics

2Ease of operation

If manual adjustments and operator experience are used for valve control, then ease of operation is maintained, but productivity and precision in nonlinear operations deteriorate

Engineering Contradiction:
Improveease of operationVSAvoidproductivity
Core Design Contradiction:
Ease of operationVSProductivity

Solution Approach 1:

The control system performs self-optimization through automated model selection and switching based on evaluation results. The system serves itself by automatically adapting to operational conditions without requiring manual intervention, thereby improving productivity while maintaining ease of operation through the automated decision-making process

Inventive Principle:
Principle #25Self-service

3Device complexity

If a single operation model is used to simplify the control system, then device complexity is reduced, but adaptability to different operational conditions deteriorates

Engineering Contradiction:
Improvedevice complexityVSAvoidadaptability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

Multiple candidate models are trained to handle different operational conditions and scenarios. Each model serves as a specialized component that can be universally applied when its specific conditions are met, providing adaptability across diverse operations while keeping the overall system structure manageable through standardized model architecture

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Manufacturing precision

If reinforcement learning is used to generate operation models, then manufacturing precision and control accuracy are improved, but loss of time and computational resources increase

Engineering Contradiction:
Improvecontrol accuracyVSAvoidtraining time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

Multiple candidate models are pre-trained offline using reinforcement learning before deployment. This preliminary action allows the system to perform computationally intensive training work in advance, reducing real-time computational requirements and enabling fast model selection during actual operation without significant time loss

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20230384742A1Model selection apparatus, model selection method, and non-transitory computer readable medium
Publication Date: 2023.11.30 YOKOGAWA ELECTRIC CORP
  • US20230384742A1 patent drawing
  • US20230384742A1 patent drawing
  • US20230384742A1 patent drawing

AI summary

Provided is a model selection apparatus including a candidate model storage unit configured to store a plurality of candidate models each of which is generated by reinforcement learning using, as at least a part of a reward, an output of an evaluation model configured to output an indicator obtained by evaluating a state of equipment and is capable of output an action, a state data acquisition unit configured to acquire a plurality of pieces of state data representing the state of the equipment in a case where each of manipulated variables based on outputs of the plurality of candidate models is applied to a controlled object in the equipment, an indicator acquisition unit configured to acquire a plurality of indicators, a model selection unit configured to select an object model for controlling the controlled object, and an object model output unit configured to output the object model.