Gas Turbine RL Control With Randomized Setpoint Deviation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Complex systems like gas turbines face challenges in determining optimal control policies due to large or partially known state spaces, leading to inefficient control and difficulty in distinguishing between system state changes and control actions, especially when switching between finite control policies for changing objectives such as emissions and combustion dynamics.
Innovation Solution
A method using Reinforcement Learning with randomized deviations from set points to explore the state space for a control policy that maximizes expected total reward, allowing for efficient training and adaptation to changing objectives, and enabling better differentiation between control actions and system state effects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If Reinforcement Learning is used to learn control policy from training data, then control policy can be determined for complex systems, but training becomes unstable and data inefficient when control objectives change
Solution Approach 1:
The patent applies parameter changes by introducing a randomized offset to the control objective parameter. Instead of using fixed control objectives, the system adds random variations to the objective values during training, allowing the same training data to be effective across different control objectives. This resolves the contradiction by making the training process adaptable to changing objectives while maintaining stability through the consistent application of randomization.
2Measurement precision
If manual sweeps or subsampling are used to address high correlation between system state and control actions, then training data can be collected, but the process is time consuming and complex
Solution Approach 1:
The patent applies self-service by using the system's own operational data without requiring external manual intervention. Instead of manual sweeps or expert selection, the system automatically processes existing operational data with randomization, allowing it to self-train on real-world data. This resolves the contradiction by achieving precise distinction between control actions and system state effects without time-consuming manual processes.
3Adaptability or versatility
If multiple finite control policies are provided for gradations between generic control objectives, then control flexibility is improved, but device complexity increases
Solution Approach 1:
The patent applies universality by creating a single control policy that can handle multiple control objectives through randomization. Instead of maintaining separate finite control policies for different objectives, the system uses one universal policy trained with randomized objectives that can adapt to various control goals. This resolves the contradiction by providing control flexibility without increasing device complexity through multiple policies.
Data Source
Figure 1~3
Figure 4
Figure 5~6
AI summary
A method (10a; 10b) of controlling a complex system (50) and a gas turbine (50) being controlled by the method (10a; 10b) are provided. The method (10a; 10b) comprises providing (11) training data (40), which training data (40) represents at least a portion of a state space (S) of the system (50); setting (12) a generic control objective (32) for the system (50) and a corresponding set point (33); and exploring (13) the state space (S), using Reinforcement Learning, for a control policy for the system (50) which maximizes an expected total reward. The expected total reward depends on a randomized deviation (31) of the generic control objective (32) from the corresponding set point (33).