Warehouse Agent Control Using Digital Twin Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional algorithms for order picking and replenishment in warehouses are inflexible, requiring significant customization and struggle to adapt to changing warehouse operations and order conditions, leading to inefficiencies and suboptimal performance.

Innovation Solution

A system utilizing deep reinforcement learning and multi-agent reinforcement learning to control agents in warehouses, enabling continuous adaptation and optimization of order picking and replenishment processes, incorporating real-time data and simulation to improve efficiency and flexibility.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional heuristic-based algorithms are used to control order picking systems, then the system can address specific customer requirements and warehouse conditions, but the algorithms require significant effort to design, test, implement, optimize, program, and verify, and are not easily transferable to other warehouses

Engineering Contradiction:
Improveadaptability to customer requirementsVSAvoidalgorithm implementation complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent uses digital twins to create virtual copies of warehouse environments, agents, and operations. These digital twins are used to train reinforcement learning algorithms, allowing the system to learn optimal control strategies in a simulated environment before deployment in real warehouses, eliminating the need for complex manual algorithm design and verification

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces conventional heuristic-based control algorithms with reinforcement learning algorithms that learn optimal policies through interaction with digital twin environments. This substitution transforms the control approach from manually programmed heuristics to autonomously learned policies, reducing implementation complexity while maintaining adaptability

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If conventional algorithms are used for order picking control, then the system can operate with existing infrastructure, but the algorithms do not adjust well to changing warehouse operations and order conditions

Engineering Contradiction:
Improveoperational stabilityVSAvoidadaptability to changing conditions
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic adaptability through reinforcement learning algorithms that continuously learn from interactions with the environment. The algorithms can adjust their policies in response to changing warehouse operations, order conditions, and operational constraints, while the digital twin framework allows for continuous training and adaptation to new scenarios

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system incorporates feedback loops where operational data from real warehouses is used to update and retrain the reinforcement learning algorithms. The digital twin environment provides a feedback mechanism for testing and validating algorithm performance before deployment, ensuring reliable adaptation to changing conditions

Inventive Principle:
Principle #23Feedback

3Productivity

If multiple objectives are optimized simultaneously (lead time, energy consumption, distance travelled, labor cost), then the system can achieve holistic optimization, but the control problem becomes increasingly complex with varying warehouse sizes, geometries, and configurations

Engineering Contradiction:
Improveholistic optimizationVSAvoidcontrol system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent creates a universal reinforcement learning framework that can handle multiple optimization objectives simultaneously across different warehouse configurations. The digital twin environment allows the algorithm to learn policies that generalize across various warehouse sizes, geometries, and operational conditions, achieving holistic optimization without proportionally increasing control complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system uses reinforcement learning to dynamically adjust control parameters based on multiple objectives. The algorithm learns to balance lead time, energy consumption, distance travelled, and labor cost by adjusting agent behaviors and operational parameters in response to changing conditions and priority settings

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20260054929A1Artificial intelligence control and optimization of agent tasks in a warehouse
Publication Date: 2026.02.26 STILL GMBH
  • US20260054929A1 patent drawing
  • US20260054929A1 patent drawing
  • US20260054929A1 patent drawing

AI summary

A control system for a warehouse includes a controller for communicating commands for execution by item carrying vehicles, robotic pickers, and human workers. A warehouse simulation performs simulated runs of order picking and replenishment activities. The simulated results and experience data are recorded and stored in storage. The stored data includes operational data including live results and experience data that was recorded while the workers were performing according to the executable commands from the controller. A training module receives the simulation results, the simulated experience data, and the recorded operational data from the storage. The training module trains an algorithm using the simulated data and the operational data. The training module generates an updated algorithm for the controller. Using the updated algorithm, the controller communicates executable commands to the workers.