Industrial Automation Component Configuration via Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for predicting real-time behavior in industrial automation components are time-consuming and inefficient due to the need for iterative testing of numerous hardware and software permutations, making it impossible to make flexible and accurate predictions about cycle times and jitter.

Innovation Solution

A method using reinforcement learning to create a model that predicts real-time behavior by recording and optimizing properties of various hardware and software combinations, allowing for precise configuration and programming of industrial automation components.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If iterative testing of numerous hardware and software permutations is performed to predict real-time behavior, then prediction accuracy is improved, but time consumption and efficiency deteriorate

Engineering Contradiction:
Improveprediction accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies preliminary action by pre-collecting runtime property data from multiple hardware and software combinations and storing it in a database before actual prediction is needed. This pre-prepared data foundation enables the reinforcement learning model to make accurate predictions without requiring time-consuming iterative testing at the time of configuration, thus resolving the contradiction between prediction accuracy and time consumption

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by creating a reinforcement learning model that learns from copied data patterns in the database rather than physically testing each configuration. The model captures the essential relationships between hardware/software combinations and their runtime properties, enabling accurate predictions without replicating the actual testing process, thereby reducing time consumption while maintaining prediction accuracy

Inventive Principle:
Principle #26Copying

2Reliability

If hardware is generously dimensioned to guarantee real-time behavior for untested combinations, then reliability is improved, but resource utilization efficiency deteriorates

Engineering Contradiction:
Improvereal-time behavior guaranteeVSAvoidresource utilization efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies feedback by using the reinforcement learning model to predict runtime properties for specific hardware/software combinations before deployment. This feedback mechanism provides confidence levels and predicted performance metrics that allow engineers to make informed decisions about hardware sizing, avoiding both over-provisioning and under-provisioning, thus resolving the contradiction between reliability and resource utilization efficiency

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent uses parameter changes by adjusting hardware dimensions and configurations based on predicted runtime properties from the reinforcement learning model. Instead of using fixed generous dimensions for all cases, the system optimizes hardware parameters according to the specific requirements and predictions for each configuration, achieving reliable real-time behavior while improving resource utilization efficiency

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4441569B1Method, assembly and program for configuring and/or programming an industrial automation component
Publication Date: 2025.12.03 SIEMENS AG
  • EP4441569B1 patent drawingFigure 1
  • EP4441569B1 patent drawingFigure 2

AI summary

The invention relates to a method and an assembly for configuring and/or programming an industrial automation component. For a plurality of possible combinations of the hardware, each combination having a different hardware version, and/or the operating system, each combination having a different operating system version, and/or the application program, each combination having a different application program version, the respective properties at runtime are detected and stored in a database. A model for predicting the properties at runtime is generated from the data of the database and/or is optimized using a reinforcement learning process, wherein a reinforcement learning reward function which is used during the learning or optimization process of the model is aimed at an accurate prediction of the properties. The properties at runtime are then predicted for a number of intended or possible combinations using the model, and the predictions are then compared with a specified requirement. Finally, a suitable combination is ascertained using the comparison, and the industrial automation component is configured or programmed according to the selected combination. Thus, the real-time behavior can be very precisely predicted such that the automation component can be optimally configured or programmed, in particular the available computing power can be optimally used without needing to keep excessively large resources available and without the risk of violating real-time requirements.