Constrained Reinforcement Learning for Safe Manufacturing Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Manufacturing processes with complex, nonlinear relationships between input actions and output responses are difficult to model effectively using closed-form mathematical models, leading to inefficiencies and safety concerns in control algorithms, particularly in applications like robotics and chemical processing.

Innovation Solution

The implementation of constrained reinforcement learning, which uses a machine-learned network to adaptively control manufacturing processes by setting manipulated variables based on input states, incorporating safety constraints and device limitations to ensure safe and practical operation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If closed-form mathematical models are used for manufacturing process control, then control accuracy may be improved, but development time and engineering labor increase significantly

Engineering Contradiction:
Improvecontrol accuracyVSAvoiddevelopment time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces traditional mathematical modeling approaches with machine learning models that automatically learn process dynamics from data. This substitution eliminates the need for manual derivation of differential-algebraic equations and empirical model fitting, reducing development time from months to days while maintaining control accuracy through data-driven pattern recognition.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The machine learning model automatically learns process characteristics and control strategies from historical process data without requiring manual engineering intervention for model derivation. The system self-trains on available data, automatically identifying complex nonlinear relationships that would be difficult to capture with traditional mathematical modeling approaches.

Inventive Principle:
Principle #25Self-service

2Loss of time

If model-free reinforcement learning is used to reduce engineering cost and time, then development time decreases, but safety and practical operation may be compromised

Engineering Contradiction:
Improvedevelopment timeVSAvoidsafety
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The patent incorporates safety constraints and operational limits as feedback mechanisms in the reinforcement learning framework. The model is trained with reward functions that penalize unsafe or impractical actions, ensuring that learned policies respect device capabilities and safety requirements while still achieving rapid adaptation and reduced development time.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent pre-encodes safety constraints, device limitations, and operational boundaries into the reinforcement learning training process before deployment. By incorporating these constraints during the training phase rather than adding them afterward, the system learns safe operating behaviors from the outset, preventing unsafe actions before they occur in practice.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

3Adaptability or versatility

If traditional reinforcement learning is applied without constraints, then adaptability improves, but unsafe or impractical operations may occur

Engineering Contradiction:
ImproveadaptabilityVSAvoidunsafe operation
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent implements dynamic constraint satisfaction where the reinforcement learning model adapts to changing process conditions while continuously respecting safety boundaries. The constraints are integrated into the decision-making process, allowing the system to dynamically adjust its actions within safe operating envelopes, maintaining adaptability without compromising safety.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11914350B2Manufacturing process control using constrained reinforcement machine learning
Publication Date: 2024.02.27 SIEMENS AG
  • US11914350B2 patent drawing
  • US11914350B2 patent drawing

AI summary

For manufacturing process control, closed-loop control is provided (18) based on a constrained reinforcement learned network (32). The reinforcement is constrained (22) to account for the manufacturing application. The constraints may be for an amount of change, limits, or other factors reflecting capabilities of the controlled device and/or safety.