Learning-Based Machine Control With Dynamic Rule Weighting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing learning-based control systems for complex technical systems face risks due to untested output signals and difficulty in assessing their consequences, especially under conditions not covered by training data, necessitating a safer and more reliable control method.
Innovation Solution
A control method using a learning-based control device that incorporates a weighting of performance against a deviation from a specific action selection rule, allowing dynamic adjustment of objectives to prioritize either performance or adherence to a predefined rule based on the control device's prediction accuracy and historical data analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If a learning-based control device is used to optimize machine performance, then productivity and efficiency are improved, but reliability deteriorates due to untested output signals and difficulty in assessing consequences
Solution Approach 1:
The patent introduces a reference policy as an intermediary between the learning-based control device and the machine. The reference policy acts as a safety mediator that verifies control signals before execution, blocking signals that deviate too much from established safe operating procedures while allowing signals that conform to safe patterns, thus resolving the contradiction between performance optimization and reliability
Solution Approach 2:
The patent implements feedback by continuously monitoring control signals against the reference policy and adjusting the weighting between performance optimization and policy adherence based on observed deviations and their consequences. This feedback mechanism allows the system to learn from past experiences and improve reliability while maintaining productivity
2Productivity
If the control device prioritizes performance optimization, then productivity increases, but deviation from safe operating rules increases
Solution Approach 1:
The patent applies dynamics by making the weighting between performance optimization and reference policy adherence variable rather than fixed. The weighting is dynamically adjusted based on the current operating state, the confidence of the learning-based controller, and the severity of potential deviations, allowing the system to prioritize performance when safe and policy adherence when risks are detected
Solution Approach 2:
The patent changes the parameter of weighting between performance and policy adherence based on the operating conditions and confidence levels. When the learning-based controller is confident and operating within safe parameters, higher weight is given to performance optimization. When uncertainty increases or safe operating boundaries are approached, the weight shifts toward policy adherence, thus resolving the contradiction
3Reliability
If the control device follows predefined action selection rules, then reliability improves, but productivity decreases due to constrained optimization
Solution Approach 1:
The patent applies partial action by having the reference policy block only those control signals that exhibit excessive deviation from safe operating procedures, while allowing signals that fall within acceptable ranges to pass through for performance optimization. This partial blocking approach maintains reliability by preventing dangerous deviations while preserving productivity by allowing beneficial optimizations within safe boundaries
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The invention relates to a computer-implemented method for controlling a machine (M) by means of a trained learning-based control device (CTL), wherein the training of the control device (CTL) is carried out using an objective function which includes a weighting of a performance of the machine (M) against a deviation from a specific action selection rule. To control the machine (M), the previously trained control device (CTL) determines operating control signals (AO) based on the operating state signals (SO) of the machine (M) and values for the weighting provided to it and outputs them. These values for the weighting are determined by the control device (CTL) by calculating an error variable relating to a prediction of state signals by the control device (CTL) compared to operating state signals (SO) of the machine (M).Furthermore, the invention relates to a corresponding control device (CTL), a computer program, a computer-readable data carrier, and a transmission signal.