Learning-Based Machine Control With Dynamic Rule Weighting

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing learning-based control systems for complex technical systems face risks due to untested output signals and difficulty in assessing their consequences, especially under conditions not covered by training data, necessitating a safer and more reliable control method.

Innovation Solution

A control method using a learning-based control device that incorporates a weighting of performance against a deviation from a specific action selection rule, allowing dynamic adjustment of objectives to prioritize either performance or adherence to a predefined rule based on the control device's prediction accuracy and historical data analysis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If a learning-based control device is used to optimize machine performance, then productivity and efficiency are improved, but reliability deteriorates due to untested output signals and difficulty in assessing consequences

Engineering Contradiction:
Improvemachine performanceVSAvoidcontrol signal reliability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent introduces a reference policy as an intermediary between the learning-based control device and the machine. The reference policy acts as a safety mediator that verifies control signals before execution, blocking signals that deviate too much from established safe operating procedures while allowing signals that conform to safe patterns, thus resolving the contradiction between performance optimization and reliability

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements feedback by continuously monitoring control signals against the reference policy and adjusting the weighting between performance optimization and policy adherence based on observed deviations and their consequences. This feedback mechanism allows the system to learn from past experiences and improve reliability while maintaining productivity

Inventive Principle:
Principle #23Feedback

2Productivity

If the control device prioritizes performance optimization, then productivity increases, but deviation from safe operating rules increases

Engineering Contradiction:
Improvemachine performanceVSAvoiddeviation from safe operating rules
Core Design Contradiction:
ProductivityVSObject-affected harmful factors

Solution Approach 1:

The patent applies dynamics by making the weighting between performance optimization and reference policy adherence variable rather than fixed. The weighting is dynamically adjusted based on the current operating state, the confidence of the learning-based controller, and the severity of potential deviations, allowing the system to prioritize performance when safe and policy adherence when risks are detected

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter of weighting between performance and policy adherence based on the operating conditions and confidence levels. When the learning-based controller is confident and operating within safe parameters, higher weight is given to performance optimization. When uncertainty increases or safe operating boundaries are approached, the weight shifts toward policy adherence, thus resolving the contradiction

Inventive Principle:
Principle #35Parameter changes

3Reliability

If the control device follows predefined action selection rules, then reliability improves, but productivity decreases due to constrained optimization

Engineering Contradiction:
Improvecontrol signal reliabilityVSAvoidmachine performance
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies partial action by having the reference policy block only those control signals that exhibit excessive deviation from safe operating procedures, while allowing signals that fall within acceptable ranges to pass through for performance optimization. This partial blocking approach maintains reliability by preventing dangerous deviations while preserving productivity by allowing beneficial optimizations within safe boundaries

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentEP4600765A1Controlling a machine with learning-based control device
Publication Date: 2025.08.13 SIEMENS AG
  • EP4600765A1 patent drawingFigure 1
  • EP4600765A1 patent drawingFigure 2
  • EP4600765A1 patent drawingFigure 3

AI summary

The invention relates to a computer-implemented method for controlling a machine (M) by means of a trained learning-based control device (CTL), wherein the training of the control device (CTL) is carried out using an objective function which includes a weighting of a performance of the machine (M) against a deviation from a specific action selection rule. To control the machine (M), the previously trained control device (CTL) determines operating control signals (AO) based on the operating state signals (SO) of the machine (M) and values for the weighting provided to it and outputs them. These values for the weighting are determined by the control device (CTL) by calculating an error variable relating to a prediction of state signals by the control device (CTL) compared to operating state signals (SO) of the machine (M).Furthermore, the invention relates to a corresponding control device (CTL), a computer program, a computer-readable data carrier, and a transmission signal.