ML Control Architecture for Stable Training Under Time Delays
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Convergence and repeatability issues arise in training machine learning models for controlling complex technical systems due to partial state space consideration, noisy sensor data, and time delays in control actions, impairing learning success in industrial environments.
Innovation Solution
A control device configuration using three machine learning modules, where the first module reproduces behavior signals without control actions, the second module reproduces behavior signals with control actions, and the third module optimizes control actions by learning from the outputs of the first two modules, allowing for efficient training and improved convergence and repeatability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a single machine learning model is used to control complex technical systems, then the system can learn from complete state space, but convergence problems and repeatability issues arise due to noisy sensor data and time delays
Solution Approach 1:
The patent divides the control system into three separate machine learning modules: a first module that processes operating signals, a second module that processes behavior signals, and a third module that generates control actions. This segmentation allows each module to specialize in specific aspects of the control task, improving convergence and repeatability while maintaining the ability to handle complex state spaces through modular architecture.
Solution Approach 2:
The patent introduces intermediate processed signals between the input operating signals and the final control actions. The first machine learning module produces intermediate representations from operating signals, which are then processed by the second module to produce behavior signals, finally processed by the third module to generate control actions. These intermediaries help decouple the complex relationships and reduce the impact of noisy sensor data and time delays.
2Productivity
If only a small part of the state space is considered, then training becomes faster and requires fewer resources, but convergence problems arise and learning success is impaired
Solution Approach 1:
By segmenting the state space processing across three modules, each module can focus on specific aspects of the state space. The first module handles operating signal processing, the second handles behavior signal processing, and the third handles control action generation. This allows the system to consider a comprehensive state space while maintaining training efficiency through specialized sub-processing in each module.
Solution Approach 2:
The patent applies partial action by having each module process only the specific portion of the state space relevant to its function, rather than processing the entire state space in one monolithic model. The first module processes operating signals partially, the second processes behavior signals partially, and the third generates control actions partially, with the combination of these partial actions covering the complete state space effectively.
3Loss of information
If noisy sensor data is processed directly, then all available information is used, but convergence problems and repeatability issues occur
Solution Approach 1:
The patent introduces intermediate processed signals that act as mediators between the noisy sensor data and the final control decisions. The first machine learning module processes the raw operating signals into intermediate representations, filtering and structuring the information. The second module then processes these intermediates into behavior signals, further refining the information. This multi-stage intermediary processing reduces the impact of noise while preserving essential information.
Solution Approach 2:
By segmenting the information processing into three distinct modules, each module can apply specialized processing techniques to handle noisy data in different ways. The first module can focus on feature extraction from operating signals, the second on pattern recognition in behavior signals, and the third on decision-making for control actions, with each segment contributing to noise reduction while maintaining information utility.
4Adaptability or versatility
If control actions with different time delays are handled by a single model, then the complete control scenario is captured, but repeatability problems arise
Solution Approach 1:
The patent makes the system dynamic by having the third machine learning module receive not only the current operating signals but also historical processed signals from the first and second modules. This allows the system to adapt to different time delays in control actions by considering the temporal evolution of the processed signals, improving repeatability while maintaining versatility in handling different control scenarios.
Solution Approach 2:
The first and second machine learning modules perform preliminary processing of the signals before they reach the third module that generates control actions. This preliminary action includes filtering, feature extraction, and pattern recognition that prepare the data in advance, reducing the impact of time delays and improving the repeatability of control actions across different scenarios.
Data Source
AI summary
An operating signal is fed to a first machine learning module to reproduce a behavior signal of a technical system, the behavior signal occurring specifically without the current use of a control action and output the reproduced behavior signal as a first output signal. The first output signal is fed to a second machine learning module to reproduce a resulting behavior signal using a control action signal and output the reproduced behavior signal as a second output signal. Furthermore, an operating signal is fed to a third machine learning module, and a third output signal is fed to the trained second machine learning module. A control action performance ascertained using the second output signal, and the control action performance is used to train the third machine learning module to optimize the control action performance. By training the third machine learning module, a control device controls the technical system.


