Learning-Based Machine Control with User-Guided Rule Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing learning-based control systems for complex technical systems face risks due to untested output signals and difficulty in assessing their consequences, especially under conditions not covered by training data, necessitating a safer and more reliable control method.
Innovation Solution
A control method using a learning-based control device that integrates a weighting of performance against a deviation from predefined action selection rules, allowing the control device to select the most appropriate rule based on user preferences and adapt the control strategy dynamically.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If learning-based control is used to optimize system performance, then productivity and efficiency are improved, but reliability deteriorates due to untested output signals and difficulty in assessing consequences
Solution Approach 1:
The patent introduces a reference policy as an intermediary between the learning-based control system and the technical system. This reference policy acts as a mediator that provides verified, safe control signals when the learning-based system produces unreliable outputs, thereby maintaining reliability while allowing performance optimization to proceed
Solution Approach 2:
The patent dynamically adjusts the weighting parameter α in the objective function based on the reliability assessment of the learning-based control signals. When signals are deemed reliable, the system prioritizes performance optimization; when unreliable, it shifts toward following the reference policy, thus adapting the balance between productivity and reliability in real-time
2Productivity
If learning-based control is used to achieve performance optimization, then efficiency is improved, but safety worsens due to untested operating control signals
Solution Approach 1:
The patent prepares a reference policy in advance that embodies safe, verified control strategies. This reference policy serves as a pre-prepared safety cushion that can be activated when the learning-based system produces potentially harmful untested signals, thereby cushioning against safety risks before they materialize
Solution Approach 2:
The patent implements a feedback mechanism where the reliability of learning-based control signals is continuously assessed based on training data coverage and operating conditions. This feedback loop allows the system to detect when safety risks arise and switch to the reference policy, thereby preventing harmful outcomes
3Adaptability or versatility
If multiple action selection rules are integrated into the control device, then adaptability is improved, but device complexity increases
Solution Approach 1:
The patent makes the selection between different action selection rules dynamic rather than static. The control device automatically selects which rule to apply based on real-time assessment of operating conditions and training data coverage, allowing adaptability without requiring complex manual configuration or switching mechanisms
Solution Approach 2:
The patent segments the control space into different regions based on the type of operating conditions and training data coverage. Each segment is associated with a specific action selection rule, allowing the system to handle different scenarios with appropriate rules while keeping the overall device complexity manageable through structured organization
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The invention relates to a computer-implemented method for controlling a machine (M) by means of a trained learning-based control device (CTL). The training of the control device (CTL) was carried out using an objective function which includes a weighting of two variables, namely a performance of the machine (M) versus a deviation from a specific action selection rule from a plurality of action selection rules. During control, the previously trained control device (CTL) determines and outputs operating control signals (AO) based on the operating state signals (SO) of the machine (M) provided to it, values to be used for the weighting (W), and information (c) regarding the action selection rule to be used from the plurality of action selection rules.The action selection rule to be used is selected by the control device (CTL) from the plurality of action selection rules by calculating, for each of the plurality of action selection rules, a deviation value relating to a deviation between control signals specified by a user based on operating state signals (SO) and operating control signals (AO) determined by the control device (CTL) based on the same operating state signals (SO).