Manufacturing Scheduling Policy Robustness Under Changing Production Goals
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Manufacturing facilities face challenges in efficiently coordinating the operation and utilization of hundreds of machines producing diverse products, especially when production priorities change, requiring robust and adaptive scheduling policies.
Innovation Solution
A manufacturing system that computes differences in evaluation metrics between scheduling policies, determines if these differences are within certain thresholds, and adjusts scheduling policies accordingly, combining meta learning and robust adversarial reinforcement learning techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If human-based scheduling approach is used, then scheduling can be adjusted based on expertise, but it becomes difficult to coordinate as facility size grows and requires many years of training
Solution Approach 1:
The patent replaces human-based scheduling with reinforcement learning-based automated scheduling systems. The RL agent learns optimal scheduling policies through interaction with the manufacturing environment, substituting human expertise with machine learning algorithms that can handle complex coordination of hundreds of machines and products without requiring years of training.
Solution Approach 2:
The patent employs meta-learning techniques that enable the scheduling system to adapt to changes in production targets and goals by learning from multiple tasks and scenarios. The system adjusts its scheduling policies dynamically based on changing parameters such as production demands, machine availability, and product priorities, maintaining robustness across different operating conditions.
2Extent of automation
If reinforcement learning-based scheduling is used, then automated scheduling can be achieved, but the policy may not remain optimal when production targets change
Solution Approach 1:
The patent applies meta-learning where the RL agent is pre-trained on multiple simulated tasks and scenarios before deployment. This preliminary training on diverse production scenarios enables the agent to quickly adapt to new production targets and goals without requiring complete retraining, maintaining optimal performance across changing conditions.
Solution Approach 2:
The system continuously monitors scheduling performance and uses this feedback to refine and update scheduling policies. The RL agent learns from actual production outcomes and adjusts its policies iteratively, ensuring that the automated scheduling remains robust and optimal even when production targets change.
3Productivity
If scheduling policies are frequently adjusted to meet changing production goals, then optimal performance can be maintained, but computational resources and time are consumed
Solution Approach 1:
The meta-learning approach pre-trains the RL agent on a wide variety of production scenarios and goals before actual deployment. This preliminary action creates a robust foundation that allows the agent to handle new production targets with minimal additional training time, reducing the computational resources and time needed for policy adjustments when production goals change.
Data Source
AI summary
According to one or more embodiments of the present disclosure, a manufacturing system may include a processor and a memory storing instructions executed by the processor to cause the processor to compute a difference between a first evaluation metric corresponding to a first scheduling policy and a second evaluation metric. The processor may determine that the computed difference is less than a first threshold, and in response, calculate a second scheduling policy, and determine that the second scheduling policy is greater than a second threshold, and in response, calculate a third scheduling policy such that the third scheduling policy is less than the second threshold.


