Double-arm robot combined equipment scheduling method in wafer processing
By using distribution algorithms and deep reinforcement learning methods in wafer processing equipment, the scheduling of multifunctional two-arm robots is optimized, and the equipment scheduling problem is solved, and efficient and high-quality processing of wafers is achieved, especially after photoresist is applied in time to avoid the influence of solvent diffusion.
Patent Information
- Application Number
- CN202510320192.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-07-25
AI Technical Summary
In existing wafer processing equipment, multifunctional double-arm robot combination equipment is difficult to efficiently schedule, resulting in difficult to ensure wafer processing efficiency and quality, especially when the photoresist is not dry in time after it is coated, it affects the dimensional accuracy.
The allocation algorithm is used to determine the distribution scheme and wafer stream of the multi-functional processing module, combined with the deep reinforcement learning method, and optimize the scheduling scheme of the two-arm robot through the interaction between the agent and the environment, and select the best scheduling scheme to ensure production efficiency and quality.
It effectively solves the scheduling problem of multi-functional double-arm robot combination equipment, ensures high-quality and efficient processing of wafers, reduces post-processing residence time, and improves production efficiency.
Smart Images

Figure CN120363176A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of semiconductor manufacturing technology, and particularly to a scheduling method for a dual-arm robot combined equipment in wafer processing. Background Art
[0002] A wafer is a silicon wafer carrier for manufacturing semiconductor chips. In the process of semiconductor chip manufacturing, combined equipment is widely used to perform most wafer processing processes. The combined equipment usually consists of 4-6 single-function processing modules (SPMs), a robot responsible for transporting wafers between processing modules, and two vacuum locks (LLs) responsible for storing wafers. According to the configuration of the robot, the combined equipment can be divided into two types: one is a single-arm robot combined equipment, which is equipped with one robotic arm and can only grasp one wafer at a time; the other is a dual-arm robot combined equipment, which is equipped with two robotic arms, and each robotic arm can grasp one wafer at the same time.
[0003] In order to improve the manufacturing efficiency of the combined equipment, multi-function processing modules (MPMs) have gradually been applied. Compared with SPM, MPM can continuously perform multiple processing steps, thereby reducing the time for the robotic arm to grasp and move. However, MPM can also choose to perform only one of the processing steps. Therefore, the allocation problem of MPMs in the combined equipment becomes crucial, and different allocation schemes will lead to different scheduling efficiencies. In actual production, the allocation of MPMs is usually operated by on-site engineers. Due to the huge and complex decision-making space, it is extremely challenging to select a suitable MPM allocation scheme and perform efficient scheduling. The existing scheduling scheme for the SPM dual-arm robot combined equipment cannot be directly applied to the MPMs dual-arm robot combined equipment system. In addition, after the photoresist coating step in the wafer processing process, a drying operation is required. If the wafer is not taken out and dried in time after coating, the solvent in the photoresist will diffuse, which will affect the dimensional accuracy. Therefore, the wafer should be taken out by the robot as soon as possible after coating to reduce the residence time after processing, which further increases the complexity of robot scheduling. Summary of the Invention
[0004] The object of the present invention is to overcome the problems existing in the prior art, and provide a scheduling method for a dual-arm robot combined equipment in wafer processing. The present invention can effectively solve the problem that it is difficult to schedule the multi-function dual-arm robot combined equipment, and ensure the high-quality and high-efficiency processing of wafers.
[0005] To achieve the above object, the present invention provides a scheduling method for a dual-arm robot combined equipment in wafer processing, the method comprising: Given the first parameters of the dual-arm robot combined equipment, and determining the allocation scheme of the multi-function processing modules in the dual-arm robot combined equipment and the wafer flow corresponding to each allocation scheme according to the allocation algorithm; A system for simulating the dual-arm robot combination equipment based on the wafer flow and the first parameter; Taking the system as the environment of deep reinforcement learning, using an agent to interact with the environment, and obtaining the scheduling scheme of the dual-arm robot corresponding to each allocation scheme; Taking the completion time and the residence time after wafer processing as performance indicators, selecting the allocation scheme corresponding to the best scheduling scheme from all the scheduling schemes, and deploying the best scheduling scheme to the actual wafer production, so as to complete the scheduling of the dual-arm robot combination equipment.
[0006] Further, the dual-arm robot combination equipment includes a multi-functional processing module, a single-functional processing module, and a dual-arm robot. The multi-functional processing module is used to complete the processing steps of the wafer j and steps j +1; the single-functional processing module is responsible for executing the remaining processing steps; the dual-arm robot is used to transport the wafers.
[0007] Further, the dual-arm robot combination equipment further includes a vacuum lock, and the vacuum lock is used to store wafers.
[0008] Further, the allocation algorithm needs to meet the following constraint conditions: m j + m j+1 + m j&j+1 = k If m j > 0 then m j+1 > 0 If m j+1 > 0 then m j > 0 Wherein, k is the total number of multi-functional processing modules in the combination equipment, j and j +1 are the processing steps completed by the multi-functional processing module, that is, the multi-functional processing module can respectively execute step j、 step j +1 and the merging step j&j +1, m j 、 m j+1 and m j&j+1 are respectively the steps of executing j、 step j+1 and the merging step j&j The number of multi-functional processing modules corresponding to +1, and the constraint conditions ensure the continuous use of the multi-functional processing modules and avoid idleness.
[0009] Furthermore, the input of the allocation algorithm is k and m i , and the output is a legal allocation scheme and the wafer flow corresponding to each allocation party. Specifically, the legal allocation schemes that meet the constraint conditions include: Case 1: m j&j+1 = k and m j = m j+1 = 0; Case 2: m j&j+1 > 0, m j > 0 and m j+1 > 0; Case 3: m j&j+1 = 0, m j > 0 and m j+1 > 0; Among them, the wafer flow corresponding to Case 1 is ([[]] m 1,..., m j-1 , m j&j+1 , m j+2 ,…, m n ), the wafer flow corresponding to Case 3 is ([[]] m 1,…, m j-1 , m j , m j+1 , m j+2 ,…, m n ), and the wafer flow corresponding to Case 2 is ([[]] m 1,…, m j-1 , m j&j+1 , m j+2 ,…, m n ) and ([[]] m 1,…,m j-1 , m j , m j+1 , m j+2 ,…, m n )。
[0010] Further, the first parameter includes: the process processing time of the processing module, the time for the robot to perform wafer unloading, the time for the robot to perform wafer loading, the time for the robot to move between processing modules, or the time for the robot to move between the processing module and the vacuum lock.
[0011] Further, taking the system as the environment of deep reinforcement learning, and using an agent to interact with the environment to obtain the scheduling scheme of the dual-arm robot corresponding to each allocation scheme specifically includes: (1). Interacting the system of the dual-arm robot combination device with the agent as the environment of deep reinforcement learning; (2). The Mask mechanism gives reasonable actions of the dual-arm robot in the current state based on the real-time state of the system s, = Mask( A valid = Mask( s ); (3). The agent selects the optimal robot action from the reasonable actions according to ϵ- the greedy policy A valid ; a ; (4). Using the environment to simulate the state of the system after executing the selected optimal robot action a ; s '; (5). Calculating the reward r and weight p i according to the set reward and weight scheme, and then storing them in the prioritized experience replay pool in the form of ( s , a , r , s ', done, p i ); D ; (6). Sampling from the prioritized experience replay pool D , and the online network calculates the Q value according to the real-time state s t and the action , denoted as , the target network calculates the future predicted target value according to the state s ' and records it as y t ; (7). Calculate the error between the predicted target value of the target network and the Q value calculated by the online network, and record it as δ ; (8). According to the error δ to backpropagate and update the network parameters, and finally output the scheduling scheme.
[0012] Further, the online network calculates the Q value according to the real-time state s t and the action , and the specific calculation method is as follows:
[0013] Among them, is the output of the value stream, is the output of the advantage stream for the action a t , and |A| is the size of the action space.
[0014] Further, the target network calculates the future predicted target value according to the state s ', and the calculation formula is as follows:
[0015] Among them, is the discount factor.
[0016] Further, calculate the error between the predicted target value of the target network and the Q value calculated by the online network, and the calculation formula is as follows:
[0017] Among them, is the error, is the reward at the next moment.
[0018] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention determines the allocation scheme of the multi-functional processing module in the dual-arm robot combination device and the corresponding wafer flow for each allocation scheme through the allocation algorithm, and then simulates the system of the dual-arm robot combination device according to the wafer flow and the first parameter. Taking the system as the environment of deep reinforcement learning, the dual-arm robots of the combined device system under different allocation schemes are scheduled through the deep reinforcement learning algorithm, effectively solving the problem that it is difficult to schedule the multi-functional dual-arm robot combination device; and according to the completion time of each scheduling scheme and the maximum post-processing residence time of the wafers during the processing, a scheme that ensures both production efficiency and production quality is selected, ensuring the high-quality and high-efficiency processing of the wafers. Brief Description of the Drawings
[0019] Figure 1 is a flowchart of a scheduling method for a dual-arm robot combined equipment in wafer processing according to Embodiment 1 of the present invention; Figure 2 is a schematic structural diagram of the combined equipment corresponding to the wafer flow output by the allocation algorithm according to Embodiment 1 of the present invention; Figure 3 is a schematic diagram of the MD3QN network architecture according to Embodiment 1 of the present invention; Figure 4 is a structural diagram of robot scheduling for a multi-functional combined equipment system based on MD3QN according to Embodiment 2 of the present invention. Detailed Embodiments
[0020] The following further describes in detail the specific embodiments of the present invention in conjunction with the drawings and embodiments. The following embodiments are used to illustrate the present invention but are not intended to limit the scope of the present invention.
[0021] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "longitudinal", "lateral", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as limiting the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0022] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0023] In addition, in the description of the present invention, unless otherwise stated, the meaning of "a plurality of" is two or more.
[0024] Embodiment 1 As Figure 1 shown, a scheduling method for a dual-arm robot combined equipment in wafer processing according to a preferred embodiment of an embodiment of the present invention includes: Step S1: Determine the allocation scheme of the multi-functional processing modules in the dual-arm robot combination device and the corresponding wafer flow for each allocation scheme. In some embodiments, the dual-arm robot combination device includes multi-functional processing modules (MPMs), single-functional processing modules (SPMs), a dual-arm robot (R), and vacuum locks (LLs). The multi-functional processing modules are used to complete the wafer processing steps j and steps j +1; the single-functional processing modules are responsible for performing the remaining processing steps; the dual-arm robot is used to transport wafers, and the vacuum locks are used to store wafers.
[0025] Problem description: Let N q + = {1, 2, …, q} represent the set of finite natural numbers; let n represent the total number of wafer processing steps.
[0026] In the dual-arm combination device, wafer processing usually involves n steps ( n > 2) in sequential order. Among these steps, the j th step and the j +1th step (where j < n ) can be executed as a combined step by a single multi-functional processing module (MPM), and at the same time, they can also be processed separately by two MPMs, with each MPM acting as a single-functional processing module (SPM). It should be noted that the j th step is a specified step with specific processing requirements, and the j th step and the j +1th step can only be executed by the MPM, while all other steps in the sequence are executed by the SPM.
[0027] According to the above settings, if the j th step and the j +1th step are executed by a single MPM, they will be processed as a combined step and represented as step j & j+ 1. On the contrary, if the j th step and the j +1th step are processed separately by different MPMs, they are regarded as independent steps. Processing these two steps independently can provide higher scheduling flexibility and achieve more efficient resource utilization under certain operating conditions. We use m i to represent the number of MPMs for processing step i ( i ∈ N n + ), thenm j&j+1 Indicates the number of MPMs assigned to the processing combination step j & j +1.
[0028] For a dual-arm robot combination device with k MPMs ( k > 2), when assigning MPMs to the j th step, the j +1th step, and the combination step j & j +1, the following constraint conditions should be satisfied according to the assignment algorithm: m j + m j+1 + m j&j+1 = k If m j > 0, then m j+1 > 0 If m j+1 > 0, then m j > 0 Among them, k is the total number of multi-functional processing modules in the combination device, j and j +1 are the processing steps completed by the multi-functional processing module, that is, the multi-functional processing module can execute step j、 step j +1 and the combination step j&j +1. m j , m j+1 and m j&j+1 are the numbers of multi-functional processing modules corresponding to executing step j、 step j +1 and the combination step j&j +1 respectively. All MPMs must be fully utilized without any idle situation. Therefore, m j + m j+1 + m j&j+1 = k . In addition, m j > 0 indicates that at least one MPM is assigned to the jStep. To ensure the continuity of the process, at least one MPM must be assigned to step j +1; otherwise, after completing step j , there will be no MPM to execute step j +1, thus violating the operation requirements. Similarly, if m j+1 >0, then m j must also be greater than zero; otherwise, the MPM assigned to step j +1 will become idle.
[0029] Assignment patterns that do not conform to the above criteria are considered invalid. We use < m j&j+1 , m j , m j+1 > to represent a valid MPM assignment pattern, where each element represents the number of MPMs assigned to combined step j & j +1, step j , and step j +1, respectively. The goal is to identify all valid assignment patterns and select one that minimizes the completion time for processing a batch of wafers.
[0030] Furthermore, Algorithm 1 is designed to determine all valid assignments of MPMs. The input of Algorithm 1 is k , m i . For all legal assignment schemes < m j&j+1 , m j , m j+1 >, according to m j&j+1 , m j , m j+1 's different values, we divide these schemes into three cases: Case 1: m j&j+1 = k and m j = m j+1 = 0; Case 2: m j&j+1 >0, m j >0 and m j+1 >0; Case 2:m j&j+1 = 0, m j > 0 and m j+1 > 0; Among them, for Case 1 and Case 3, the wafer processing flow is: ( m 1,..., m j-1 , m j&j+1 , m j+2 ,…, m n ) and ( m 1,…, m j-1 , m j , m j+1 , m j+2 ,…, m n ), which means there is only one processing step route for the wafer. For Case 2, the corresponding WFPs are ( m 1,…, m j-1 , m j&j+1 , m j+2 ,…, m n ) and ( m 1,…, m j-1 , m j , m j+1 , m j+2 ,…, m n ), at this time there are two processing step routes for the wafer, which makes the scheduling more complex and challenging. There is currently no patent discussing the scheduling method of DACT in this complex mode. Algorithm 1 outputs the WFPs corresponding to all solutions.
[0031] Step S2: Simulate the system of the dual-arm robot combined device based on the wafer flow and the first parameter; In some embodiments, generate the WFPs corresponding to the solutions, and then create a dual-arm robot multi-functional processing module combined device in combination with the first parameter, so as to facilitate the interaction between the environment and the agent, where the type, quantity, and process processing time of the processing modules for completing each step in the allocation solution, as well as the time for the dual-arm robot to execute wafer unloading and wafer loading, are allocated.
[0032] Step S3: Use the system as the environment for deep reinforcement learning, and let an agent interact with the environment to obtain the scheduling scheme of the dual-arm robot corresponding to each allocation scheme. In some embodiments, Step S3 specifically includes: Step S301: Use the system of the dual-arm robot combination device as the environment for deep reinforcement learning to interact with the agent. Step S302: The Mask mechanism gives reasonable actions of the dual-arm robot in the current state s, based on the real-time state of the system A valid = Mask( s ); Step S303: The agent selects the optimal robot action from the reasonable actions ϵ- according to the A valid greedy policy. a ; Step S304: Use the environment to simulate the state a after the system executes the selected optimal robot action s '; Step S305: Calculate the reward r and weight p i according to the set reward and weight scheme, and then store them in the prioritized experience replay pool in the form of ( s , a , r , s ', done, p i ). D ; Step S306: Sample from the prioritized experience replay pool D . The online network calculates the Q value according to the real-time state s t and the action , denoted as . The target network calculates the future predicted target value according to the state s ', denoted as y t ; Step S307: Calculate the error between the predicted target value of the target network and the Q value calculated by the online network, denoted as ; Step S308: Backpropagate and update the network parameters according to the error to finally output the scheduling scheme.
[0033] Further, the online network calculates the Q value based on the real-time status s t and actions , and the specific calculation method is as follows:
[0034] where is the value stream output, is the output of the advantage stream for action a t , and |A| is the size of the action space.
[0035] Further, the target network calculates the future predicted target value based on the status s ', and the calculation formula is as follows:
[0036] where is the discount factor.
[0037] Further, calculate the error between the predicted target value of the target network and the Q value calculated by the online network. The calculation formula is as follows:
[0038] where is the error, is the reward at the next moment.
[0039] Design an algorithm 2 for solving the robotic arm scheduling scheme: the MD3QN algorithm. The input of algorithm 2 is: Q θ , Q θ - , α , β , lr , γ , ϵ start ,ϵ decay ,B, M, M min ,τ ; The output is the trained online network.
[0040] where α represents the priority coefficient of the samples in the experience replay pool. A larger α value means that during training, those experiences with larger TD errors will be sampled preferentially to accelerate learning. Related to this is the importance sampling correction coefficient β , which is used to correct the sample bias caused by priority sampling. lris the learning rate, which controls the step size of each weight update of the algorithm model, and the discount factor γ is used to measure the impact of future rewards on the current decision, ϵ is ϵ- the exploration factor in the greedy policy, which controls the balance between exploration and exploitation of the agent during training, ϵ start sets the initial exploration value, ϵ decay determines the decay speed, so that the exploration behavior can be gradually reduced and the exploitation behavior can be increased. B , M and M min represent the batch size of each sampling, the maximum capacity of the experience pool, and the minimum training capacity respectively, all of which contribute to balancing the storage efficiency and training effect. Finally, τ is the soft update factor, which controls the update speed of the target network parameters.
[0041] In Algorithm 2, first, the environment corresponding to the WFPs simulation according to the output of Algorithm 1 is carried out, and the online network Q θ and the target network Q θ - are initialized. At the same time, a priority-based experience replay mechanism (PERB) and a masking mechanism are introduced to improve the training efficiency. In each training episode, the algorithm dynamically adjusts the exploration rate ϵTo balance exploration and exploitation, actions are selected from the set of effective robot actions. Specifically, it is decided whether the robotic arm unloads or loads wafers at a certain processing module, and interacts with the environment to obtain rewards and state transitions, that is, to judge whether the actions of this robot have caused damage to the wafers or reduced the overall efficiency. Subsequently, the algorithm calculates the priorities based on the TD error, stores the data in the experience replay pool, and removes the oldest samples when the pool size exceeds the set threshold. Through prioritized sampling, the algorithm selects a small batch of data from the experience pool, updates the gradients of the network combined with importance weights, and updates the sample priorities in real time. Every certain number of steps, the target network synchronizes with the online network to ensure the stability and convergence of the algorithm. MD3QN organically combines the double deep Q-network (D3QN), prioritized experience replay (PERB), and action masking strategy, significantly improving the solution efficiency and effect in complex environments. After continuous exploration and learning by the agent, a scheduling plan based on different MPM allocation modes can be obtained. This plan determines all the actions of the robot and the specific execution times after a batch of wafers are placed in the LL, ensuring that all wafers are processed correctly and then returned to the LL. At the same time, the maximum residence time of the wafers after processing in different PMs is considered during the scheduling process, aiming to ensure production efficiency while ensuring the quality of the wafers by controlling the residence time after processing.
[0042] Step S4: Use the completion time and the residence time of the wafers after processing as performance indicators, select the allocation plan corresponding to the best scheduling plan from all the scheduling plans, and deploy the best scheduling plan to the actual wafer production, thus completing the scheduling of the dual-arm robot combined equipment.
[0043] In this embodiment, the allocation plan of the multi-functional processing module in the dual-arm robot combined equipment and the wafer flow corresponding to each allocation plan are determined through an allocation algorithm. Then, the system of the dual-arm robot combined equipment is simulated according to the wafer flow and the first parameter. The system is used as the environment of deep reinforcement learning, and the dual-arm robots in the combined equipment system under different allocation plans are scheduled through the deep reinforcement learning algorithm, effectively solving the problem of difficult scheduling of the multi-functional dual-arm robot combined equipment; and according to the completion time of each scheduling plan and the maximum residence time of the wafers during processing, a plan that ensures both production efficiency and production quality is selected, ensuring the high-quality and high-efficiency processing of the wafers.
[0044] Embodiment 2 This embodiment is the specific implementation process of Embodiment 1. In the multi-functional processing module dual-arm robot combined equipment system of this embodiment, there are k = 3, m 3 = m4 = 2. The processing of the wafer requires four steps. Steps 1 and 2 are completed by MPM, and steps 3 and 4 are completed by SPM. The first step in scheduling the combined equipment is to obtain all legal MPM allocations. For the above example, the process of applying Algorithm 1 to solve for the legal allocation and give the specific WFP is as follows:
[0045] Applying Algorithm 1 will generate the corresponding Patterns and WFPs , Figure 2 which gives the specific result diagram for the corresponding combined equipment.
[0046] Next, steps S2 to S3 in the first embodiment will be further described in detail as follows: In the problem of scheduling a multi-functional module combined equipment based on reinforcement learning, abstracting it into a Markov decision process (MDP) is a key step. The MDP is defined by the five-tuple ⟨ S , A , P , R ⟩, where S represents the state space, which contains all possible states of the robots, processing modules, and vacuum stations in the system; A represents the robot action space, which defines the set of robot actions triggered in each state; P ( s ′ | s , a ) is the state transition probability, which describes the probability distribution of transferring to the next state s after the robot executes the action a in the state s′ ; R ( s , a ) is the reward function, which is used to evaluate the immediate reward obtained by executing the action s in the state a .
[0047] In the MDP, the current state s of the combined equipment system is sufficient to fully characterize the dynamics of the system, which makes the future system state depend only on the current state and the actions taken, and is independent of the past states. The goal of reinforcement learning is to determine an optimal policy π ∗ , which maximizes the expected cumulative discounted reward, expressed as: , where γ is the discount factor, which is used to balance the importance of short-term and long-term rewards; For a deep reinforcement learning environment, the WFPs generated by Algorithm 1 are used to configure the wafer processing steps associated with the PMs. The environment includes robots, MPMs, SPMs, and LLs. The wafers are moved by the robotic arm of the robot to complete the processing of all wafers on the corresponding PMs in sequence and finally returned to the LLs. In the scheduling of a dual-arm robot combined device with multi-functional processing modules, the decision-making process must be based on the current state of each PM and its associated wafers in order to select the most appropriate robot action. Define actions: Swap operation i Refers to the wafer swapping operation between the PMs responsible for step i The swapping operation of the robot involves unloading the processed wafer from the PM responsible for step i Then rotate the two robotic arms and load another wafer into the PM for processing. Specifically, when the robot performs the PM swapping operation for step 1, it must move to the vacuum lock (LLs) to unload the raw wafer, and then move to the PM responsible for step 1 to load the wafer into it. This ensures that the wafer can be quickly loaded into the PM after the unloading operation is completed, thus minimizing the time cost. The specific process is determined by the operation rules, which determine the actions of the robot. First, one robotic arm unloads the processed wafer from the PM, and then the other robotic arm rotates and loads the new wafer into the same PM. This wafer is the one unloaded in the previous Swap i -1 step. At this point, the entire swapping operation is completed. There are the following rules to guide the actions of the robot: 1) If a robotic arm is holding a wafer, the unloading operation cannot be performed; conversely, if a robotic arm is not holding a wafer, the loading operation cannot be performed.
[0048] 2) Each PM can only process one wafer at a time.
[0049] 3) The wafers must be processed strictly in the order of the predetermined processing steps.
[0050] Through these rules, the effective cooperation between the PMs, LLs, and the robot provides a solid foundation for optimizing the scheduling of the deep reinforcement learning algorithm.
[0051] Define states: W LL : The number of wafers to be processed in the LL.
[0052] S stg : The operating phase of the system (initial state, steady state, shutdown state).
[0053] P Arm: Position of the robotic arm.
[0054] N PM : Number of processed wafers.
[0055] T RP : Remaining processing time of the wafer being processed in PM.
[0056] T RR : Remaining residence time of the wafer in PM.
[0057] T Idle : Idle time of PM.
[0058] The dimension of the state space directly affects the efficiency and effect of reinforcement learning. Although a higher dimension can provide more detailed information and enable the agent to better understand the environment, the increase in dimension will significantly increase the complexity of the problem, thus slowing down the learning speed; Define the reward: If the multi-functional processing module combination equipment successfully completes the processing of a batch of wafers, a reward of 1000 is given. On the contrary, violating the basic operation rules will be severely punished with a penalty value of -1000. When the robot successfully executes an action, a small reward of 30 is given. When parallel PM is effectively utilized, a reward of 5 is given; otherwise, if PM is underutilized, a penalty of -5 is given. Since the goal is to minimize the completion time, the execution time of each action will be penalized; Design algorithm 2 for solving the scheduling robot action sequence corresponding to the environment. The input of algorithm 2 is: Q θ , Q θ - , α , β , lr , γ , ϵ start ,ϵ decay ,B, M, M min ,τ The output of algorithm 2 is: Trained online network Q θ Among them, α represents the priority coefficient of the samples in the experience replay pool. A larger α value means that those experiences with larger TD errors will be preferentially sampled during training to accelerate learning. Related to this is the importance sampling correction coefficient β , which is used to correct the sample bias caused by priority sampling.lr is the learning rate, which controls the step size of each weight update of the algorithm model, and the discount factor γ is used to measure the impact of future rewards on the current decision, ϵ is ϵ- the exploration factor in the greedy policy, which controls the balance between exploration and exploitation of the agent during training, ϵ start sets the initial exploration value, ϵ decay determines the decay speed, which can gradually reduce exploration behavior and increase exploitation behavior. B , M and M min respectively represent the batch size of each sampling, the maximum capacity of the experience pool, and the minimum training capacity, all of which help to balance storage efficiency and training effect. Finally, τ is the soft update factor, which controls the update speed of the target network parameters. As Figure 3 stated, the architecture of the MD3QN network ensures that the agent can explore efficiently.
[0059] In this embodiment, the designed Algorithm 2 is as follows:
[0060] As Figure 4 stated, this embodiment provides a scheduling method for a multi-functional processing module dual-arm robot combination device based on deep reinforcement learning. By applying the MD3QN algorithm, the learning and exploration of the multi-functional processing module dual-arm robot combination device system are completed; then, engineers can directly view and analyze the performance of each MPMs allocation scheme from the operation results to make further decisions; As shown in Table 1, scheduling is performed for all four MPMs allocation cases, and the scheduling scheme for each case is obtained. Thus, the case corresponding to WFP (3, 2, 2) can be selected as the optimal scheme; Table 1
[0061] In this example, when the combined device contains a multi-functional processing module, first determine all possible allocation schemes and generate the WFPs corresponding to the schemes. Then, create a combined device of a dual-arm robot multi-functional processing module in combination with other processing parameters to facilitate interaction as the environment and the agent. Then, use the deep reinforcement learning algorithm to obtain the scheduling scheme for the whole process of the system from the initial transient state to the shutdown transient state based on the action exploration and learning process of the agent. According to the completion time of each scheduling scheme and the maximum post-processing residence time of the wafers during the processing, select the scheme that ensures both production efficiency and production quality. This method effectively solves the problem of difficult scheduling of the dual-arm robot in the combined device with a multi-functional processing module and ensures high-quality and high-efficiency processing of the wafers.
[0062] In summary, the embodiment of the present invention provides a scheduling method for a combined device of a dual-arm robot in wafer processing. It determines the allocation scheme of the multi-functional processing module in the combined device of the dual-arm robot and the wafer flow corresponding to each allocation scheme through an allocation algorithm. Then, simulate the system of the combined device of the dual-arm robot according to the wafer flow and the first parameter, use the system as the environment of deep reinforcement learning, and schedule the dual-arm robot of the combined device system under different allocation schemes through the deep reinforcement learning algorithm, effectively solving the problem of difficult scheduling of the multi-functional dual-arm robot combined device. And according to the completion time of each scheduling scheme and the maximum post-processing residence time of the wafers during the processing, select the scheme that ensures both production efficiency and production quality, ensuring high-quality and high-efficiency processing of the wafers.
[0063] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and substitutions can be made, and these improvements and substitutions should also be regarded as the protection scope of the present invention.
Claims
1. A scheduling method for a dual-arm robot combination device in wafer processing, characterized in that, The method includes the following steps: Given the first parameters of the dual-arm robot combination device, and determining the allocation scheme of the multi-functional processing module in the dual-arm robot combination device and the corresponding wafer flow for each allocation scheme according to the allocation algorithm; Simulating the system of the dual-arm robot combination device based on the wafer flow and the first parameters; Taking the system as the environment of deep reinforcement learning, and using an agent to interact with the environment to obtain the scheduling scheme of the dual-arm robot corresponding to each allocation scheme; Taking the completion time and the residence time after wafer processing as performance indicators, selecting the allocation scheme corresponding to the best scheduling scheme from all the scheduling schemes, and deploying the best scheduling scheme to the actual wafer production, thereby completing the scheduling of the dual-arm robot combination device.
2. A scheduling method for a dual-arm robot combination device in wafer processing according to claim 1, characterized in that, The dual-arm robot combination device includes a multi-functional processing module, a single-functional processing module, and a dual-arm robot. The multi-functional processing module is used to complete the processing steps of the wafer. j and steps j +1; the single-functional processing module is responsible for performing the remaining processing steps; the dual-arm robot is used to transport the wafer.
3. A scheduling method for a dual-arm robot combination device in wafer processing according to claim 2, characterized in that, The dual-arm robot combination device further includes a vacuum lock for storing wafers.
4. A scheduling method for a dual-arm robot combination device in wafer processing according to claim 1, characterized in that, The allocation algorithm needs to satisfy the following constraint conditions: m j + m j+1 + m j&j+1 = k If m j > 0 then m j+1 > 0 If m j+1 > 0 then m j > 0 Among them, k is the total number of multi-functional processing modules in the combined device, j and j +1 are the processing steps completed by the multi-functional processing module, that is, the multi-functional processing module can respectively execute step j、 Step j +1 and the merging step j&j +1, m j 、 m j+1 and m j&j+1 are respectively the numbers of the multi-functional processing modules corresponding to the execution of step j、 Step j +1 and the merging step j&j +1. The constraint conditions ensure the continuous use of the multi-functional processing module and avoid idleness.
5. A scheduling method for a dual-arm robot combination device in wafer processing according to claim 4, characterized in that, The input of the allocation algorithm is k and m i , and the output is a legal allocation scheme and the wafer flow corresponding to each allocator. Specifically, the legal allocation schemes that meet the constraint conditions include: Case 1: m j&j+1 = k and m j = m j+1 = 0; Case 2: m j&j+1 > 0, m j > 0 and m j+1 > 0; Case 3: m j&j+1 = 0, m j > 0 and m j+1 > 0; Among them, the wafer flow corresponding to Case 1 is ( m 1,..., m j-1 , m j&j+1 , m j+2 ,…, m n ), the wafer flow corresponding to Case 3 is ( m 1,…, m j-1 , m j , m j+1 , m j+2 ,…, m n ), and the wafer flow corresponding to Case 2 is ( m 1,…, m j-1 , m j&j+1 , m j+2 ,…, m n ) and ( m 1,…, m j-1 , m j , m j+1 , m j+2 ,…, m n ).
6. A scheduling method for a dual-arm robot combination device in wafer processing according to claim 1, characterized in that, The first parameters include: the process processing time of the processing module, the time for the robot to unload the wafer, the time for the robot to load the wafer, the time for the robot to move between the processing modules, or the time for the robot to move between the processing module and the vacuum lock.
7. A scheduling method for a dual-arm robot combination device in wafer processing according to claim 1, characterized in that, The step of taking the system as the environment of deep reinforcement learning, using an agent to interact with the environment, and obtaining the scheduling scheme of the dual-arm robot corresponding to each allocation scheme specifically includes: (1) Interacting with the agent by taking the system of the dual-arm robot combination device as the environment of deep reinforcement learning; (2). The Mask mechanism is based on the real-time state of the system s, to give reasonable actions of the dual-arm robot in the current state A valid = Mask( s ); (3). The agent selects the optimal robot action from the reasonable actions according to ϵ- the greedy strategy A valid ; a ; (4). Use the environment to simulate the system to execute the selected optimal robot actions a The state after s '; (5) Calculate the rewards according to the set reward and weight scheme r and weights p i , and then store them in the prioritized experience replay pool in the form of ( s , a , r , s ', done, p i ) D ; (6). Sample from the prioritized experience replay pool D The online network calculates the Q value, denoted as s t according to the real-time state and the action The target network calculates the future predicted target value according to the state s ', denoted as y t ; (7). Calculate the error between the predicted target value of the target network and the Q value calculated by the online network, denoted as δ ; (8) According to the error δ to backpropagate and update network parameters, and finally output a scheduling scheme.
8. A scheduling method for a dual-arm robot combination device in wafer processing according to claim 7, characterized in that, The online network, based on the real-time status s t and actions , calculates the Q value, and the specific calculation method is as follows: Among them, is the value stream output, is the output of the advantage stream for the action a t , and |A| is the size of the action space.
9. A method for scheduling a combined device of a dual-arm robot in wafer processing according to claim 7, characterized in that The target network calculates the future predicted target value based on the state s 'The calculation formula is as follows: Among them, is the discount factor.
10. A scheduling method for a dual-arm robot combination device in wafer processing according to claim 9, characterized in that, The error between the predicted target value of the target network and the Q value calculated by the online network is calculated, and the calculation formula is as follows: where, is the error, is the reward at the next moment.