ML Control Safety Module for Admissible Action Signals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Machine learning-based control systems for complex technical systems often fail to ensure that control actions adhere to predefined safety constraints, leading to potential safety issues in safety-critical applications.
Innovation Solution
A control device and method that integrate a safety module with a machine learning module, where safety information is used to validate and convert output signals into admissible control actions, ensuring both safety and optimization through techniques like backpropagation and gradient-based methods, and by considering expert knowledge and training data volume.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If machine learning methods are used to train a control policy for optimizing technical system performance, then the productivity and performance optimization is improved, but the reliability and safety compliance deteriorates because there is no guarantee that control actions will observe predefined limit values or technical constraints
Solution Approach 1:
The patent introduces a safety module as an intermediary component between the machine learning policy and the technical system. This safety module validates control actions generated by the policy against predefined safety constraints and limit values, blocking unsafe actions while allowing safe actions to pass through. This mediator resolves the contradiction by enabling performance optimization through machine learning while ensuring safety compliance through the intermediary validation layer.
2Reliability
If safety validation is applied to control actions output by a trained policy, then the reliability and safety compliance is improved, but the productivity and performance optimization deteriorates because the policy does not act in an optimum fashion
Solution Approach 1:
The patent implements a feedback mechanism where the safety module's validation results are used to train the machine learning policy. Control actions that are blocked by the safety module provide feedback signals that guide the policy to learn safe behavior patterns. This feedback loop enables the policy to progressively optimize its performance while naturally adhering to safety constraints, resolving the contradiction between safety compliance and performance optimization.
Solution Approach 2:
The patent applies preliminary action by pre-defining safety constraints and limit values before the machine learning policy is deployed. The safety module uses these predefined constraints to validate control actions in advance, preventing unsafe actions from being executed. This preliminary establishment of safety boundaries enables the policy to learn within safe operating regions, maintaining both safety compliance and performance optimization.
3Reliability
If the machine learning module is constrained to output only validated control actions, then the reliability is improved, but the adaptability and optimization capability deteriorates
Solution Approach 1:
The patent applies dynamics by making the control policy adaptive through continuous training with feedback from the safety module. The policy dynamically adjusts its output based on learned patterns from validated control actions, rather than being statically constrained. This dynamic adaptation enables the policy to maintain optimization capability while reliably producing safe control actions through the feedback-driven learning process.
Data Source
AI summary
A control device for a technical system, state-specific safety information about an admissibility of a control action signal is read in by a safety module is provided. Furthermore, a state signal indicating a state of the technical system is supplied to a machine learning module and to the safety module. In addition, an output signal of the machine learning module is supplied to the safety module. The output signal is converted into an admissible control action signal by the safety module on the basis of the safety information depending on the state signal. Furthermore, a performance for control of the technical system by the admissible control action signal is ascertained, and the machine learning module is trained to optimize the performance. The control device is then configured by the trained machine learning module.


