Gas Turbine RL Control With Randomized Setpoint Deviation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Complex systems like gas turbines face challenges in determining optimal control policies due to large or partially known state spaces, leading to inefficient control and difficulty in distinguishing between system state changes and control actions, especially when switching between finite control policies for changing objectives such as emissions and combustion dynamics.

Innovation Solution

A method using Reinforcement Learning with randomized deviations from set points to explore the state space for a control policy that maximizes expected total reward, allowing for efficient training and adaptation to changing objectives, and enabling better differentiation between control actions and system state effects.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If Reinforcement Learning is used to learn control policy from training data, then control policy can be determined for complex systems, but training becomes unstable and data inefficient when control objectives change

Engineering Contradiction:
Improveadaptability to changing control objectivesVSAvoidtraining stability
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent applies parameter changes by introducing a randomized offset to the control objective parameter. Instead of using fixed control objectives, the system adds random variations to the objective values during training, allowing the same training data to be effective across different control objectives. This resolves the contradiction by making the training process adaptable to changing objectives while maintaining stability through the consistent application of randomization.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If manual sweeps or subsampling are used to address high correlation between system state and control actions, then training data can be collected, but the process is time consuming and complex

Engineering Contradiction:
Improveability to distinguish control actions from system state effectsVSAvoidtime for data collection
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies self-service by using the system's own operational data without requiring external manual intervention. Instead of manual sweeps or expert selection, the system automatically processes existing operational data with randomization, allowing it to self-train on real-world data. This resolves the contradiction by achieving precise distinction between control actions and system state effects without time-consuming manual processes.

Inventive Principle:
Principle #25Self-service

3Adaptability or versatility

If multiple finite control policies are provided for gradations between generic control objectives, then control flexibility is improved, but device complexity increases

Engineering Contradiction:
Improvecontrol flexibilityVSAvoidnumber of control policies
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies universality by creating a single control policy that can handle multiple control objectives through randomization. Instead of maintaining separate finite control policies for different objectives, the system uses one universal policy trained with randomized objectives that can adapt to various control goals. This resolves the contradiction by providing control flexibility without increasing device complexity through multiple policies.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3682301B1Randomized reinforcement learning for control of complex systems
Publication Date: 2021.09.15 SIEMENS AG
  • EP3682301B1 patent drawingFigure 1~3
  • EP3682301B1 patent drawingFigure 4
  • EP3682301B1 patent drawingFigure 5~6

AI summary

A method (10a; 10b) of controlling a complex system (50) and a gas turbine (50) being controlled by the method (10a; 10b) are provided. The method (10a; 10b) comprises providing (11) training data (40), which training data (40) represents at least a portion of a state space (S) of the system (50); setting (12) a generic control objective (32) for the system (50) and a corresponding set point (33); and exploring (13) the state space (S), using Reinforcement Learning, for a control policy for the system (50) which maximizes an expected total reward. The expected total reward depends on a randomized deviation (31) of the generic control objective (32) from the corresponding set point (33).