Substrate Processing Schedule Generation Using Adaptive RL Rewards
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing schedule creation methods for substrate processing apparatuses require developers to create separate processing flows for each unit type, leading to inefficiencies due to differing apparatus configurations, and do not effectively optimize time schedules for substrate processing.
Innovation Solution
A reinforcement learning-based method for generating a schedule creation program that adjusts a reward function to optimize the time schedule by changing the gradient of the reward function in specific sections, allowing for the efficient arrangement of planning factors in a timetable to minimize processing time across multiple components.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If developers create separate processing flows for each unit type, then the time schedule can reflect specific apparatus configuration rules, but the development complexity and time increase significantly
Solution Approach 1:
The patent implements a universal schedule creation program that can handle multiple unit types through a common processing flow. The system uses a reward function that adapts to different apparatus configurations without requiring separate development for each unit type, making the schedule creation mechanism multi-functional and universally applicable.
Solution Approach 2:
The patent changes the approach from creating separate flows for each unit type to using parameter-based adaptation. By modifying the reward function parameters to reflect different apparatus configurations, the system achieves unit-type-specific scheduling behavior through a single universal program, reducing development complexity while maintaining accuracy.
2Productivity
If the reward function gradient is increased in the partial section, then the optimization speed for time schedule improves, but the convergence stability may be affected
Solution Approach 1:
The patent implements a dynamic reward function where the gradient in the partial section can be adjusted during the reinforcement learning process. This allows the system to increase optimization speed when needed while maintaining convergence stability through adaptive adjustment, rather than using a fixed high gradient that would compromise stability.
Solution Approach 2:
The system employs periodic adjustment of the reward function gradient, alternating between phases of higher gradient (faster optimization) and lower gradient (stability maintenance). This periodic modulation allows the system to achieve both fast optimization and stable convergence over the learning process.
Data Source
AI summary
A schedule creation program generating method generates, through reinforcement learning, a schedule creation program for creating a time schedule for a plurality of components included in a substrate processing apparatus. The method includes: increasing a value of a reward by repeating experiencing through the reinforcement learning, the experiencing including arranging and reward determining; and changing a reward function that defines a relationship between an amount of time taken according to the time schedule and a value of a corresponding reward. The arranging includes sequentially arranging a plurality of planning factors in a timetable, the plurality of planning factors being given in advance to each of the plurality of substrates. The reward determining includes determining the value of the corresponding reward based on the reward function. The changing includes changing the reward function whose gradient in a partial section is larger than a gradient in the partial section of the reward function before the changing.


