Wafer scheduling method based on reinforcement learning

By converting wafer transport tracks into graph networks and establishing reinforcement learning models, combining Hungarian algorithms and Q-Learning algorithms, the problems of congestion and insufficient dynamic performance during wafer transport in semiconductor production are solved, and efficient and real-time wafer scheduling is achieved.

CN120218367APending Publication Date: 2025-06-27SHENYANG SIASUN ROBOT & AUTOMATION
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311816945.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In semiconductor manufacturing systems, with the expansion of production scale and the increase in the number of handling vehicles, severe congestion often occurs during wafer transportation, and traditional scheduling algorithms cannot adapt to complex and changeable working conditions, resulting in performance degradation and real-time difficulty.

Method used

Using a wafer scheduling method based on reinforcement learning, a reinforcement learning model is established by converting the wafer transportation track into a graph network, a Hungarian algorithm is used to realize transportation task assignment, and wafer transportation path planning is carried out through the Q-Learning algorithm.

Benefits of technology

This method can effectively alleviate congestion, has good dynamic performance, adapt to complex and changeable production environments, reduce algorithm complexity, improve real-timeness, and improve the practicality of the system through offline training and online training stages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218367A_ABST
    Figure CN120218367A_ABST
Patent Text Reader

Abstract

The invention relates to a wafer scheduling method based on reinforcement learning. Firstly, wafer scheduling is divided into a transportation task assignment stage and a transportation path planning stage. In a transportation task assignment stage, task assignment is realized by using a Hungary algorithm based on real-time traffic information. In the transportation path planning stage, firstly, a reinforcement learning model is established, then offline training is carried out to obtain a static Q value table, finally online training is carried out, Q values are continuously updated according to actual working conditions, and therefore dynamic scheduling is achieved. The method has good dynamic performance and can adapt to complex and changeable production environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of material transportation in semiconductor manufacturing, and specifically relates to a wafer scheduling method based on reinforcement learning. Background Art

[0002] In semiconductor production and manufacturing systems, wafer transportation runs through the whole process. However, with the expansion of production scale and the increase in the number of transfer carts, serious congestion often occurs during wafer transportation. And due to the flexibility of semiconductor production, requirements are also put forward for the dynamic performance of scheduling algorithms.

[0003] Currently, common scheduling methods include rule-based heuristic scheduling algorithms, time-window-based priority scheduling strategies, etc. These existing scheduling algorithms can provide good scheduling effects at the beginning, but with the change of working conditions during the production process, their performance gradually declines and they cannot adapt to complex and changeable working conditions. In addition, the complexity of traditional scheduling algorithms is exponentially related to the number of transfer carts. With the increase in the number of transfer carts, it is difficult to ensure real-time performance.

[0004] Therefore, a wafer scheduling strategy that can relieve congestion and has good dynamic performance is needed. Summary of the Invention

[0005] The purpose of the present invention is to provide a wafer scheduling method based on reinforcement learning to overcome the above defects.

[0006] The wafer scheduling strategy based on reinforcement learning is a feasible method to solve this problem. Since the reinforcement learning algorithm has good dynamic performance and can adapt to different production situations, this method can well relieve congestion and has dynamic performance.

[0007] The technical solution adopted by the present invention to achieve the above purpose is: a wafer scheduling method based on reinforcement learning, including the following steps:

[0008] Convert the wafer transportation track into a graph network, and establish a reinforcement learning model according to the graph network;

[0009] Use the Hungarian algorithm to implement transportation task assignment according to the graph network;

[0010] Perform wafer transportation path planning through the reinforcement learning model to achieve wafer scheduling.

[0011] The step of converting the wafer transportation track into a graph network and establishing a reinforcement learning model according to the graph network includes the following steps:

[0012] Converting the wafer transportation track into a graph network is to take the wafer loading and unloading positions of each machine tool and the intersections of each track in the wafer transportation track as the nodes of the graph network, and connect any two nodes with a transportation track as the edges of the graph network.

[0013] Build a reinforcement learning model: regard the nodes in the graph network as agents, set the reward as the estimated travel time from the current node to the target node, the combined data of the current node position and the target node position is the state, all the optional next nodes of the current node are the actions, and adopt the state-action value function as the Q-value, that is, each action in the current state corresponds to a Q-value, use the Q-Learning algorithm to update the Q-value, and use the Boltzmann function as the exploration strategy to build a reinforcement learning model.

[0014] The implementation of the transportation task assignment according to the graph network using the Hungarian algorithm includes the following steps:

[0015] First, record the estimated travel time of each track corresponding to each edge in the graph network, regard the estimated travel time as the cost of each edge, and use the shortest path algorithm to find the cost matrix of the wafers to be transported and the idle transportation workshops.

[0016] Then transform the cost matrix into a square matrix, and use the Hungarian algorithm to obtain the optimal assignment with the minimum total cost, that is, the wafers to be transported and the idle transportation vehicle with the minimum total cost.

[0017] The estimated travel time is the travel time spent by the nearest passing transport vehicle on the current transport track.

[0018] The wafer transportation path planning through the reinforcement learning model includes the following steps:

[0019] Offline training stage: regard the ideal transport time of each transport track as the reward of the reinforcement learning model, use the Q-value update formula of the reinforcement learning model, and continuously update until the Q-value converges to obtain the initial Q-value;

[0020] Online training stage: The transport vehicle starts to transport the wafer to the target machine tool. At each intersection, determine the driving direction of the transport vehicle through the Q-value of the next node representing the action, and update the estimated travel time of the transport track; according to this estimated travel time, plus the distance reduction amount from the target machine tool as the reward, use the reinforcement learning model to update the Q-value to make the transport vehicle move towards the target machine tool.

[0021] The ideal transport time is the track length divided by the maximum driving speed of the transport vehicle.

[0022] The reinforcement learning model adopts the Q-Learning algorithm.

[0023] A wafer scheduling system based on reinforcement learning includes:

[0024] A reinforcement learning model construction module, which is used to transform the wafer transportation track into a graph network and build a reinforcement learning model according to the graph network;

[0025] A transportation task assignment module, which is used to implement transportation task assignment according to the graph network using the Hungarian algorithm;

[0026] A transportation path planning module, which is used to plan the wafer transportation path through a reinforcement learning model to achieve wafer scheduling.

[0027] The reinforcement learning model construction module includes an agent, which represents the nodes in the graph network and is used as a platform for the implementation of the transportation path planning module.

[0028] A wafer scheduling device based on reinforcement learning includes a memory and a processor; the memory is used to store computer programs; the processor is used to implement the wafer scheduling method based on reinforcement learning when executing the computer programs.

[0029] The present invention has the following beneficial effects and advantages:

[0030] 1. It has good dynamic performance and can adapt to complex and changeable production environments.

[0031] 2. Using track intersections rather than handling carts as agents makes the algorithm complexity independent of the number of carts and only related to the track layout, reducing the algorithm complexity and improving the real-time performance.

[0032] 3. Compared with traditional planning algorithms, using the reinforcement learning algorithm only needs to maintain the Q-value table, and there is no need to re-plan whenever a new delivery task appears, and the real-time performance is better.

[0033] 4. Using two stages of offline training and online training enables the application of the reinforcement learning algorithm on the system without having to first explore and collect data for training. Description of the Drawings

[0034] Figure 1 The method flow chart of the present invention. Detailed Implementation Manner

[0035] The following further describes the present invention in detail with reference to the drawings and embodiments.

[0036] As Figure 1 shown, the present invention proposes a wafer scheduling strategy based on reinforcement learning, and the specific content is as follows:

[0037] 1. First, convert the transportation track into a graph network, and then establish a reinforcement learning model;

[0038] Converting the wafer transportation track into a graph network means regarding the wafer loading and unloading positions of each machine tool in the original transportation track and each track intersection as the nodes of the graph network. Correspondingly, the transportation tracks connecting each node are regarded as the edges of the graph network. On this basis, a reinforcement learning model can be established. In the present invention, the nodes in the graph network model are regarded as agents, the reward is set as the estimated travel time from the current node to the target node, the combined data (current node position, target node position) is the state, all the optional next nodes of the current node are the actions, and the action-value function is used as the Q value, that is, each action corresponds to a Q value in the current state. The Q-Learning algorithm is used to update the Q value, and the Boltzmann function is used as the exploration strategy, and finally a reinforcement learning model is established.

[0039] 2. Then use the Hungarian algorithm to implement the transportation task assignment;

[0040] The wafer scheduling task is divided into two stages: transportation task assignment and transportation path planning. For the transportation task assignment stage, the Hungarian algorithm is used to achieve the optimal assignment. The application of the Hungarian algorithm is based on the graph network, and the specific steps are as follows. First, record the estimated travel time of the track corresponding to each edge of the graph network. The estimated travel time is obtained from the travel time of the nearest transportation cart passing through the track. Regarding the estimated travel time as the cost of each edge, use the shortest path algorithm to find the cost matrix between the wafers to be transported and the idle transportation carts. Then, convert the cost matrix into a square matrix by adding 0 rows or 0 columns. Finally, through the traditional Hungarian algorithm, the optimal assignment with the minimum total cost can be obtained.

[0041] 3. Then use the reinforcement learning algorithm to implement the transportation path planning

[0042] For the transportation path planning stage, the Q-Learning algorithm is used to achieve dynamic path decision-making. The key to reinforcement learning lies in training the Q value. The training of the present invention is divided into offline training and online training. The offline training stage occurs when the wafer transportation has not started yet. The training result is only related to the track layout. Regarding the ideal transportation time of each track (track length divided by the maximum travel speed of the transportation cart) as the reward of reinforcement learning, use the Q value update formula of the Q-Learning algorithm and continuously update until the Q value converges. The purpose of this stage is to obtain a reasonable initial Q value. In the online training stage, the cart starts to transport the wafer to the target machine tool. At each intersection, the Q value of the next node (that is, the action) is used to determine the driving direction of the cart, and the estimated travel time of the track is updated. According to this estimated travel time, plus the distance reduction amount to the target machine tool as the reward, use the Q-Learning algorithm to update the Q value. The Q value in the online training stage cannot converge, and it will change dynamically with the working conditions, so it has good dynamic performance.

[0043] The meanings of the technical terms in this application are as follows:

[0044] Q value: A basic concept in reinforcement learning, which can be understood as a metric used to measure the "goodness or badness" of taking a certain action in a state. Since this value is only related to the state and the action, it is also called the state-action value function.

[0045] OHT: OverHead Hoist Transport, an overhead hoist transport vehicle, a handling trolley running on an overhead track, commonly used in semiconductor factories.

[0046] AGV: Automated Guided Vehicle, a handling trolley running on the ground, running according to the ground wire or tape.

[0047] Q-Learning: A Q-value update method in reinforcement learning and also the most popular update method. This method uses offline update, and the actions used for update are usually obtained by the greedy strategy, which has nothing to do with the actually selected actions.

[0048] The present invention proposes a wafer scheduling strategy based on reinforcement learning. The application scenario is rail transportation and wafer scheduling. The handling trolley can be an OHT or an AGV. The following describes the specific implementation manners of the present invention step by step. A signal receiver is installed on each handling trolley, and a signal generator is installed at each wafer loading / unloading port and rail intersection. Centralized control is adopted, and each trolley can communicate with the control terminal.

[0049] 1. Offline training of Q value: Before the transportation task is assigned, the Q value is updated through the Q-Learning algorithm until the Q value converges.

[0050] 2. Release the transportation task

[0051] 3. Transportation task assignment: Whenever the released transportation task starts, based on the ideal transportation time, a cost matrix is established using the shortest path algorithm, and the transportation task is assigned according to the cost matrix to obtain the optimal assignment.

[0052] 4. The transportation trolley starts to transport, drives to the location where the wafer is located, loads the wafer, and then drives to the wafer destination.

[0053] 5. During the driving process of the transportation trolley, it continuously receives the signals sent by the track, records the actual driving time on the track, and sends it to the control terminal.

[0054] 6. When the transportation trolley arrives at the intersection, it decides the driving direction based on the current Q value and updates the Q value at the same time.

[0055] 7. Repeat steps 5 and 6, continuously update the Q value until the transportation trolley arrives at the destination.

[0056] 8. When the transport cart arrives at the wafer destination, return to step 3, that is, reassign tasks and assign new transport tasks to the idle carts.

Claims

1. A wafer scheduling method based on reinforcement learning, characterized in that, Including the following steps: Convert the wafer transportation track into a graph network, and establish a reinforcement learning model based on the graph network; Use the Hungarian algorithm to implement transportation task assignment according to the graph network; Perform wafer transportation path planning through the reinforcement learning model to achieve wafer scheduling.

2. The wafer scheduling method based on reinforcement learning according to claim 1, wherein The step of converting the wafer transportation track into a graph network and establishing a reinforcement learning model based on the graph network includes the following steps: Converting the wafer transportation track into a graph network means taking the wafer loading and unloading positions of each machine tool and each track intersection in the wafer transportation track as the nodes of the graph network, and connecting any two nodes with existing transportation tracks as the edges of the graph network. Establish a reinforcement learning model: regard the nodes in the graph network as agents, set the reward as the estimated travel time from the current node to the target node, the combined data of the current node position and the target node position is the state, all the optional next nodes of the current node are the actions, and adopt the state-action value function as the Q value, that is, each action corresponds to a Q value in the current state, use the Q-Learning algorithm to update the Q value, and use the Boltzmann function as the exploration strategy to establish the reinforcement learning model.

3. The wafer scheduling method based on reinforcement learning according to claim 1, characterized in that, The step of using the Hungarian algorithm to implement transportation task assignment according to the graph network includes the following steps: First, record the estimated travel time of the track corresponding to each edge in the graph network, regard the estimated travel time as the cost of each edge, and find the cost matrix of the wafers to be transported and the idle transfer vehicles through the shortest path algorithm; Then convert the cost matrix into a square matrix, and obtain the optimal assignment with the minimum total cost through the Hungarian algorithm, that is, the wafers to be transported and the idle transfer vehicle with the minimum total cost.

4. A wafer scheduling method based on reinforcement learning according to claim 3, wherein The estimated travel time is the travel time spent by the nearest passing transport vehicle on the current transportation track.

5. A wafer scheduling method based on reinforcement learning according to claim 1, characterized in that The step of performing wafer transportation path planning through the reinforcement learning model includes the following steps: Offline training stage: regard the ideal transportation time of each transportation track as the reward of the reinforcement learning model, use the Q value update formula of the reinforcement learning model, and continuously update until the Q value converges to obtain the initial Q value; Online training stage: The transport vehicle starts to transport the wafer to the target machine tool. At each intersection, determine the driving direction of the transport vehicle through the Q value of the next node representing the action, and update the estimated travel time of the transportation track; according to this estimated travel time, add the distance reduction amount from the target machine tool as the reward, and use the reinforcement learning model to update the Q value to make the transport vehicle move towards the target machine tool.

6. The wafer scheduling method based on reinforcement learning according to claim 5, wherein, The ideal transportation time is the track length divided by the maximum driving speed of the transport vehicle.

7. A wafer scheduling method based on reinforcement learning according to claim 5, characterized in that, The reinforcement learning model adopts the Q-Learning algorithm.

8. A wafer scheduling system based on reinforcement learning, characterized in that, Including: A reinforcement learning model construction module, used to convert the wafer transportation track into a graph network and establish a reinforcement learning model based on the graph network; A transportation task assignment module, used to implement transportation task assignment using the Hungarian algorithm according to the graph network; A transportation path planning module, used to perform wafer transportation path planning through the reinforcement learning model to achieve wafer scheduling.

9. A wafer scheduling system based on reinforcement learning according to claim 8, characterized in that, The reinforcement learning model construction module includes an agent, which represents the nodes in the graph network and is used as the platform for the implementation of the transportation path planning module.

10. A wafer scheduling device based on reinforcement learning, characterized in that, It includes a memory and a processor; the memory is used to store a computer program; the processor is used to implement a wafer scheduling method based on reinforcement learning as described in any one of claims 1-7 when executing the computer program.

Citation Information

Cited By

  • Ash carrying control method and system for unmanned slag carrying vehicle of thermal power plant

    CN120891807A