Substrate Processing Schedule Generation Using Adaptive RL Rewards

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing schedule creation methods for substrate processing apparatuses require developers to create separate processing flows for each unit type, leading to inefficiencies due to differing apparatus configurations, and do not effectively optimize time schedules for substrate processing.

Innovation Solution

A reinforcement learning-based method for generating a schedule creation program that adjusts a reward function to optimize the time schedule by changing the gradient of the reward function in specific sections, allowing for the efficient arrangement of planning factors in a timetable to minimize processing time across multiple components.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If developers create separate processing flows for each unit type, then the time schedule can reflect specific apparatus configuration rules, but the development complexity and time increase significantly

Engineering Contradiction:
Improveschedule accuracyVSAvoiddevelopment complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a universal schedule creation program that can handle multiple unit types through a common processing flow. The system uses a reward function that adapts to different apparatus configurations without requiring separate development for each unit type, making the schedule creation mechanism multi-functional and universally applicable.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent changes the approach from creating separate flows for each unit type to using parameter-based adaptation. By modifying the reward function parameters to reflect different apparatus configurations, the system achieves unit-type-specific scheduling behavior through a single universal program, reducing development complexity while maintaining accuracy.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If the reward function gradient is increased in the partial section, then the optimization speed for time schedule improves, but the convergence stability may be affected

Engineering Contradiction:
Improveoptimization speedVSAvoidconvergence stability
Core Design Contradiction:
ProductivityVSStability of the object's composition

Solution Approach 1:

The patent implements a dynamic reward function where the gradient in the partial section can be adjusted during the reinforcement learning process. This allows the system to increase optimization speed when needed while maintaining convergence stability through adaptive adjustment, rather than using a fixed high gradient that would compromise stability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system employs periodic adjustment of the reward function gradient, alternating between phases of higher gradient (faster optimization) and lower gradient (stability maintenance). This periodic modulation allows the system to achieve both fast optimization and stable convergence over the learning process.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS20240329624A1Schedule creation program generating method, schedule creation program generating apparatus, schedule creating apparatus, recording medium, substrate processing apparatus, and substrate processing system
Publication Date: 2024.10.03 SCREEN HOLDINGS CO LTD
  • US20240329624A1 patent drawing
  • US20240329624A1 patent drawing
  • US20240329624A1 patent drawing

AI summary

A schedule creation program generating method generates, through reinforcement learning, a schedule creation program for creating a time schedule for a plurality of components included in a substrate processing apparatus. The method includes: increasing a value of a reward by repeating experiencing through the reinforcement learning, the experiencing including arranging and reward determining; and changing a reward function that defines a relationship between an amount of time taken according to the time schedule and a value of a corresponding reward. The arranging includes sequentially arranging a plurality of planning factors in a timetable, the plurality of planning factors being given in advance to each of the plurality of substrates. The reward determining includes determining the value of the corresponding reward based on the reward function. The changing includes changing the reward function whose gradient in a partial section is larger than a gradient in the partial section of the reward function before the changing.