Weighted Aggregated Control Policy for Turbine Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Complex dynamical systems like gas turbines and wind turbines require significant operational data for optimal control strategy development, leading to sub-optimal performance during changes or high change rates, as conventional machine learning methods are data-intensive and time-consuming.
Innovation Solution
A method utilizing a pool of control policies weighted by a processor to create a weighted aggregated control policy, with weights adjusted based on performance data, allowing for rapid learning and adaptation using neural networks, especially through 'transfer learning' from similar source systems, reducing the need for extensive data collection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional machine learning methods are used to develop control strategies, then optimization performance is improved, but data collection time and system change response time increase significantly
Solution Approach 1:
The patent applies preliminary action by pre-training multiple control policies on diverse source systems before deployment. These pre-trained policies are stored in a pool and can be quickly adapted to the target system through transfer learning, eliminating the need to collect extensive operational data from scratch when the system changes.
Solution Approach 2:
The patent segments the control strategy into multiple independent control policies that are trained separately on different source systems. Each policy captures specific operational patterns, and their weighted combination provides a comprehensive control strategy that adapts quickly to system changes without requiring complete retraining.
2Productivity
If a single control policy is trained on one data set, then training convergence is achieved, but adaptability to system changes and robustness decrease
Solution Approach 1:
The patent merges multiple control policies from different source systems into a single aggregated control policy through weighted combination. This ensemble approach maintains the fast convergence of individually trained policies while achieving superior adaptability to system changes and robustness against operational variations.
Solution Approach 2:
The patent creates a universal control policy framework that can handle multiple different operational conditions and system configurations. By training policies on diverse source systems and combining them, the aggregated policy achieves multi-functionality and can adapt to various target system scenarios without requiring separate training for each case.
3Adaptability or versatility
If multiple control policies are aggregated to improve robustness, then adaptability increases, but computing effort and complexity increase
Solution Approach 1:
The patent extracts the essential adaptive capability from complex multi-policy systems by isolating the weight adjustment mechanism. Instead of managing the full complexity of multiple policies, the system only needs to adjust weights of pre-trained policies, significantly reducing computational burden while maintaining adaptability.
Solution Approach 2:
The patent changes the control approach from modifying policy parameters to adjusting weights. This parameter transformation simplifies the adaptation process, as weight adjustment requires fewer computational resources and less complexity than retraining or fine-tuning multiple policy parameters, especially in reinforcement learning contexts.
Data Source
AI summary
For controlling a target system, e.g. a gas or wind turbine or another technical system, a pool of control policies is provided. The pool of control policies comprising a plurality of control policies and weights for weighting each of the plurality of control policies are received. The plurality of control policies is weighted by the weights to provide a weighted aggregated control policy. With that, the target system is controlled using the weighted aggregated control policy, and performance data relating to a performance of the controlled target system are received. Furthermore, the weights are adjusted on the basis of the received performance data to improve the performance of the controlled target system. With that, the plurality of control policies is reweighted by the adjusted weights to adjust the weighted aggregated control policy.

