Three-Module Machine Learning Control Device for Convergence and Repeatability
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning methods for controlling complex technical systems face convergence problems and repeatability issues due to incomplete consideration of the system's state space, noisy sensor data, and time delays in control actions, which impair learning success.
Innovation Solution
A control device and method utilizing three machine learning modules to reproduce system behavior without and with control actions, followed by optimizing control action performance based on deviations and specific behavioral signals, allowing efficient training without implicitly learning system behavior.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional machine learning methods are used to train control devices, then the control device can learn system behavior, but convergence problems and repeatability issues occur due to incomplete state space consideration, noisy sensor data, and time delays
Solution Approach 1:
The patent segments the control device into multiple specialized modules: a state space generator that creates comprehensive state representations, a training data generator that synthesizes clean training data, and a control policy trainer that learns optimal control actions. This segmentation allows each module to specialize in overcoming specific training problems, thereby improving both convergence and repeatability
Solution Approach 2:
The patent applies preliminary action by pre-generating comprehensive state space representations and synthetic training data before the actual control policy training. The state space generator creates a complete representation of all possible system states, and the training data generator prepares clean, noise-free training examples in advance, eliminating convergence issues during the actual training phase
2Reliability
If comprehensive state space is considered to improve training quality, then training repeatability improves, but computing time and resources increase
Solution Approach 1:
The patent uses copying by creating synthetic copies of training data through the training data generator. Instead of requiring extensive real-world data collection that would be time-consuming, the system generates synthetic training examples that replicate real system behavior, achieving comprehensive state space coverage without proportional time investment
Solution Approach 2:
The comprehensive state space is generated in advance by the state space generator module, allowing the control policy trainer to work with pre-organized state representations. This preliminary preparation of the state space eliminates the need for repeated state space exploration during training, reducing overall training time while maintaining completeness
3Measurement precision
If noise filtering is applied to sensor data to improve training accuracy, then measurement precision improves, but processing time and computational resources increase
Solution Approach 1:
The training data generator creates clean, noise-free synthetic copies of training data that replicate real system behavior without actual sensor noise. This eliminates the need for noise filtering processing while maintaining high measurement precision, as the synthetic data is generated in a clean state from the beginning
Solution Approach 2:
Instead of taking real noisy sensor data and filtering it to remove noise, the patent inverts the approach by directly generating clean training data synthetically. This reverse approach achieves high measurement precision without the computational overhead of noise filtering operations
Data Source
Figure 1~2
Figure 3
Figure 4
AI summary
According to the invention, an operating signal (BS) of the technical system (TS) is fed into a first machine learning module (NN1), which is trained to reproduce a specific behavior signal of the technical system that arises without any current application of a control action and to output the reproduced behavior signal (VSR1) as the first output signal. The first output signal (VSR1) is fed into a second machine learning module (NN2), which is trained to reproduce a resulting behavior signal of the technical system based on a control action signal (AS) and to output the reproduced behavior signal (VSR2) as the second output signal. Furthermore, an operating signal (BS) of the technical system is fed into a third machine learning module (NN3), and a third output signal (AS) from the third machine learning module (NN3) is fed into the trained second machine learning module (NN2).The second output signal (VSR2) is used to determine a control action performance (Q). This is then used to train the third machine learning module (NN3) to optimize the control action performance (Q). The training of the third machine learning module (NN3) then configures the control unit (CTL) to control the technical system.