Pallet Routing Control Using Reinforcement Learning and Digital Twins

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional neural network routines fail to efficiently manage bottlenecks in manufacturing environments, leading to unsatisfied time and production constraints when controlling pallets in these settings.

Innovation Solution

A reinforcement learning system that uses sensor data from multiple routing control locations to calculate difference values, transient, and steady-state production values, generating state vectors and defining routes for pallets based on a digital twin of the environment, with the ability to perform corrective actions by adjusting reinforcement parameters.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional neural network routines are used to control pallets, then pallet routing decisions can be made, but bottlenecks cannot be efficiently inhibited and production constraints cannot be satisfied

Engineering Contradiction:
Improveproduction efficiencyVSAvoidconstraint satisfaction
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent transforms the control approach by changing the parameter representation from conventional neural network outputs to state vectors that explicitly encode production constraints, bottleneck indicators, and routing decisions. This parameter transformation enables the system to directly satisfy production constraints while maintaining high productivity through reinforcement learning optimization.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements a feedback mechanism where the system continuously monitors production constraints, bottleneck conditions, and pallet positions, then uses this feedback to dynamically adjust routing decisions through the reinforcement learning agent. This closed-loop feedback ensures that production constraints are satisfied while efficiently inhibiting bottlenecks.

Inventive Principle:
Principle #23Feedback

2Productivity

If reinforcement learning is used to dynamically control pallet routing, then bottlenecks can be inhibited and production constraints satisfied, but system complexity increases

Engineering Contradiction:
Improveproduction optimizationVSAvoidcontrol system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent creates a digital twin (virtual copy) of the manufacturing environment that mirrors the physical system's state, constraints, and dynamics. This digital twin allows the reinforcement learning agent to learn and optimize routing policies in a virtual environment before deploying them to the physical system, reducing the complexity burden on the actual control system while maintaining production optimization capabilities.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent segments the control problem into distinct components: state vector generation from sensor data, reinforcement learning policy selection, and routing decision execution. This segmentation allows each component to be independently optimized and managed, reducing overall system complexity while enabling sophisticated production optimization through the coordinated interaction of these modular elements.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20230196217A1Systems and methods for controlling pallets in a manufacturing environment using reinforcement learning
Publication Date: 2023.06.22 FORD GLOBAL TECH LLC
  • US20230196217A1 patent drawing
  • US20230196217A1 patent drawing
  • US20230196217A1 patent drawing

AI summary

A method includes obtaining sensor data from a plurality of sensors disposed at a plurality of routing control locations of an environment, where the sensor data is indicative of a number of a plurality of pallets at the plurality of routing control locations. The method includes calculating a plurality of difference values based on the sensor data, calculating a transient production value based on the sensor data and a transient objective function, and calculating a steady state production value based on the sensor data and a steady state objective function. The method includes generating a state vector based on the plurality of difference values, the transient production value, and the steady state production value, and defining a set of routes for a set of pallets from among the plurality of pallets based on the state vector and a digital twin of the environment.