Reinforcement Learning Agents for Adaptive Manufacturing Scheduling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing manufacturing processes struggle to efficiently react to unforeseen events such as machine failures or high-priority product demands due to increasing complexity and flexibility, often relying on human experience that introduces inefficiencies and biases.

Innovation Solution

A computer-implemented method using reinforcement learning to train agents for each product, allowing them to determine optimal manufacturing steps based on product-specific state observations, with a neural network-based actor and critic system that considers overall manufacturing environment efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If human experience-based decision making is used to handle unforeseen events, then flexibility and adaptability are maintained, but efficiency and objectivity deteriorate due to human biases and inefficiencies

Engineering Contradiction:
Improveadaptability to unforeseen eventsVSAvoidmanufacturing efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system enables self-service through autonomous agents that independently make scheduling decisions without human intervention. Each agent represents a manufacturing resource and autonomously determines optimal scheduling actions based on learned policies, eliminating the need for human operators to manually handle unforeseen events while maintaining both adaptability and efficiency

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical system of human decision-making with an intelligent software-based agent system. The agents use reinforcement learning to substitute human cognitive processes with automated algorithms that can process information faster and without bias, thereby improving manufacturing efficiency while maintaining adaptability through learned behavioral policies

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Stability of the object's composition

If traditional fixed manufacturing schedules are used, then schedule stability is maintained, but responsiveness to real-time problems deteriorates

Engineering Contradiction:
Improveschedule stabilityVSAvoidresponsiveness to real-time events
Core Design Contradiction:
Stability of the object's compositionVSAdaptability or versatility

Solution Approach 1:

The patent applies dynamics by transitioning from static fixed schedules to dynamic adaptive scheduling. The agent-based system continuously monitors the manufacturing environment and adjusts schedules in real-time based on current conditions, allowing the schedule to be both stable in its structured approach and responsive to unforeseen events through automated reconfiguration

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements feedback mechanisms where agents continuously receive information about the current manufacturing state and adjust their scheduling decisions accordingly. This closed-loop control enables the schedule to respond to real-time problems while maintaining overall stability through systematic decision-making based on learned policies and current observations

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If reinforcement learning trained agents are deployed for each product, then responsiveness and adaptability to unforeseen events improve, but system complexity increases

Engineering Contradiction:
Improveresponsiveness to unforeseen eventsVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the complex scheduling problem into smaller sub-problems, with each agent responsible for a specific manufacturing resource or product. This modular approach allows the system to handle complexity through distributed decision-making, where each agent independently manages its local context while contributing to the global scheduling objective

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system uses universal agent templates that can be applied across different manufacturing resources and products. The agents follow common reinforcement learning algorithms and interaction protocols, allowing the same architectural pattern to handle diverse scheduling scenarios. This universality reduces system complexity by reusing proven components rather than creating custom solutions for each resource

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Speed

If agents make selfish decisions for their own product, then individual product throughput may improve, but overall manufacturing efficiency deteriorates

Engineering Contradiction:
Improveindividual product processing speedVSAvoidoverall manufacturing efficiency
Core Design Contradiction:
SpeedVSProductivity

Solution Approach 1:

The patent merges individual agent objectives with the global manufacturing goal through coordinated multi-agent reinforcement learning. Agents are designed to consider the impact of their decisions on other resources and products, combining local optimization with global coordination. This merging ensures that individual product throughput improvements do not come at the expense of overall manufacturing efficiency

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system implements feedback mechanisms where agents receive information about the global manufacturing state and the impact of their actions on other products. This feedback allows agents to adjust their decisions to avoid selfish behavior that would harm overall efficiency, aligning individual product throughput optimization with global manufacturing objectives through learned cooperative policies

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP4268026B1Controlling a manufacturing schedule by a reinforcement learning trained agent
Publication Date: 2025.07.30 SIEMENS AG
  • EP4268026B1 patent drawingFigure 1
  • EP4268026B1 patent drawingFigure 2
  • EP4268026B1 patent drawingFigure 3

AI summary

In conclusion the invention concerns a method for controlling a manufacturing schedule to manufacture a product (1, n) in an industrial manufacturing environment (100), a method for training an agent (AG1, AGn) by reinforcement learning to control manufacturing schedule for manufacturing a product (1, n) and an industrial controller (1000). To improve the capability to react to unforeseen events in an industrial manufacturing process, while producing products with a high efficiency it is proposed to: - provide a product specific state observation (s1, sn) corresponding to the product (1, n) to a trained agent (AG1, AGn) as an input, - determining a subsequent manufacturing step (MS+1) for the product (1, n) as an action (a1, an), and - performing the subsequent manufacturing step (MS+1) of the action (a1, an).