A wafer wafer transmission path optimization method, system, device and medium
Patent Information
- Application Number
- CN202610807124.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-06-05
AI Technical Summary
[0006]本申请提供一种晶圆传片路径优化方法、系统、设备及介质,以解决现有技术中低附着场景下晶圆传片路径参数依赖人工调试,且防碰撞约束、防滑片约束与传片效率难以在统一评价框架下协同优化的问题
[0023] The above technical solution also has the following advantages: by determining the maximum allowable acceleration threshold based on the static friction coefficient and equivalent normal action, the evaluation of anti-slip plates can have a clear physical constraint basis; by expanding the obstacle area according to the obstacle safety margin, a collision safety buffer can be reserved in the two-dimensional plane simulation environment; by characterizing the path shape parameters with the offset distance and offset direction of the curved node, the complex path planning problem can be transformed into a finite-dimensional parameter search problem, which is convenient for the reinforcement learning optimization model to express the state and adjust the action; by determining the allowable linear velocity based on the path curvature and the maximum allowable acceleration threshold and further determining the theoretical running time, the plate transfer efficiency evaluation can be matched with the low-adhesion anti-slip plate constraint; by performing a hierarchical bounding box collision verification on the three-dimensional motion trajectory corresponding to the final two-dimensional plate transfer path before the control command is generated, the three-dimensional spatial interference risk that may be missed in the two-dimensional simplified modeling can be supplemented and verified, thereby further improving the executability and reliability of the final plate transfer path in the real working chamber.
Smart Images

Figure CN122396267B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of wafer transfer control technology, and in particular to a wafer transfer path optimization method, system, device and medium. Background Technology
[0002] Wafer transfer equipment is typically used to transfer wafers between loading areas, waiting positions, process chambers, or other workstations. Due to the compact internal space of semiconductor manufacturing equipment and the complex distribution of process chambers, chuck interfaces, sensor bases, and other mechanical structures, wafer robotic arms must not only avoid chamber boundaries and static obstacles when performing wafer transfer actions, but also ensure that the transfer path has good smoothness to reduce impact, vibration, and positioning errors.
[0003] In low-attachment wafer transfer scenarios, the wafer and end effector typically maintain their relative position through static friction, vacuum suction, or a combination of both. When the wafer robotic arm accelerates excessively, especially at high speeds along paths with significant curvature, the wafer may slip relative to the end effector, leading to issues such as wafer slippage, collisions, contamination, or equipment downtime. Therefore, wafer transfer path planning must consider not only obstacle avoidance but also anti-slip constraints determined by static friction.
[0004] Existing wafer transfer path planning schemes often generate transfer trajectories by pre-setting geometric curves, manually setting control points, manually setting safety distances, and combining speed planning methods. While this approach can achieve basic obstacle avoidance and path smoothing in engineering, it still has the following shortcomings: First, parameters such as path control points, safety margins, and speed limits are heavily reliant on engineer experience, resulting in long debugging cycles and difficulty in consistently obtaining optimal paths across different chamber layouts. Second, directly performing complex collision detection and dynamic constraint solving in the 3D model involves a large computational load and complex engineering deployment. Third, existing schemes often mix collision avoidance constraints, speed constraints, acceleration constraints, and joint space constraints, failing to establish a clear and calculable path evaluation mechanism around the static friction safety boundary in low-adhesion wafer transfer. Fourth, the balance between path efficiency and anti-slip plate safety mainly relies on manual parameter tuning, making automatic optimization of path parameters difficult.
[0005] Therefore, a wafer transfer path optimization scheme is needed that can reduce modeling and computational complexity, focus on the core physical constraints corresponding to wafer slip risk, and automatically adjust path parameters. Summary of the Invention
[0006] This application provides a wafer transfer path optimization method, system, device, and medium to solve the problems in the prior art where wafer transfer path parameters in low-attachment scenarios rely on manual adjustment, and where anti-collision constraints, anti-slip constraints, and transfer efficiency are difficult to optimize in a unified evaluation framework.
[0007] This application provides a wafer transfer path optimization method in a first aspect, comprising: constructing a two-dimensional planar simulation environment based on the horizontal projection of a wafer robotic arm and a working chamber, the two-dimensional planar simulation environment including chamber boundaries, obstacle areas, and robotic arm working areas, and determining a maximum allowable acceleration threshold for limiting wafer slippage based on the static friction between the wafer and the end effector; constructing a parameterized path generator, the parameterized path generator generating candidate transfer paths with path parameters as input, the path parameters including path shape parameters and obstacle safety margins; sampling and evaluating the candidate transfer paths to obtain the theoretical running time, collision determination results, and centripetal acceleration determination results of the candidate transfer paths; using the current path parameters as a reinforcement learning state, using the path parameter adjustment amount as a reinforcement learning action, and constructing a reward function based on the theoretical running time, the collision determination results, and the centripetal acceleration determination results to train a reinforcement learning optimization model to obtain optimal path parameters; inputting the optimal path parameters into the parameterized path generator to generate a final two-dimensional transfer path, and converting the final two-dimensional transfer path into control commands for controlling the wafer robotic arm to perform transfer actions.
[0008] Furthermore, the maximum allowable acceleration threshold is determined based on the static friction coefficient between the wafer and the end effector and the equivalent normal force provided by the end effector to the wafer, wherein the equivalent normal force includes the equivalent force obtained by converting the wafer's gravity force and / or vacuum adsorption force.
[0009] Furthermore, it also includes: expanding the boundary of the obstacle area outward according to the obstacle safety margin to form a prohibited area that the candidate wafer transfer path cannot enter; the obstacle area includes the projection area of the process cavity, chuck interface, sensor base and / or other static structures in the working chamber onto the horizontal plane.
[0010] Furthermore, the path shape parameters include the offset distance and offset direction of at least one curved node set between the starting point and the ending point of the transfer; the parameterized path generator generates a continuous and smooth candidate transfer path based on the starting point of the transfer, the offset curved node, the ending point of the transfer, and the obstacle safety margin.
[0011] Furthermore, the sampling evaluation of the candidate transfer path includes: discretely sampling the candidate transfer path to obtain multiple path sampling points; calculating the distance between each path sampling point and the prohibited area, and determining whether the candidate transfer path meets the anti-collision constraint based on the distance; calculating the path curvature and centripetal acceleration at each path sampling point, and determining whether the candidate transfer path meets the anti-slip constraint based on the comparison result of the centripetal acceleration and the maximum allowable acceleration threshold.
[0012] Furthermore, the allowable linear velocity corresponding to each of the path sampling points is determined based on the path curvature at each of the path sampling points and the maximum allowable acceleration threshold, and a velocity profile of the candidate transfer path is formed based on the allowable linear velocity. The theoretical running time of the candidate transfer path is determined based on the velocity profile.
[0013] Furthermore, the reward function includes a time gain term that is negatively correlated with the theoretical running time and a constraint penalty term; the constraint penalty term is activated when the candidate transfer path does not satisfy the anti-collision constraint or the anti-slip constraint.
[0014] Furthermore, training the reinforcement learning optimization model includes: in the two-dimensional planar simulation environment, enabling the reinforcement learning optimization model to iteratively interact with the parameterized path generator; outputting path parameter adjustment amounts based on the current path parameters, and having the parameterized path generator generate new candidate transfer paths based on the adjusted path parameters; updating the policy parameters of the reinforcement learning optimization model based on the reward function feedback, until a policy for outputting optimal path parameters is obtained.
[0015] Furthermore, the final two-dimensional transfer path corresponding to the optimal path parameters satisfies the anti-collision constraint and the anti-slip constraint, and has a smaller theoretical running time among the candidate transfer paths that satisfy the anti-collision constraint and the anti-slip constraint.
[0016] Furthermore, before converting the final two-dimensional wafer transfer path into control commands, the final two-dimensional wafer transfer path is mapped to the three-dimensional motion trajectory of the wafer robotic arm in the working chamber, and a hierarchical bounding box algorithm is used to perform collision verification on the robotic arm links, wafers, and obstacles in the working chamber corresponding to the three-dimensional motion trajectory; when the collision verification fails, the obstacle safety margin and / or the path shape parameters are adjusted, and a candidate wafer transfer path is regenerated.
[0017] This application provides a wafer transfer path optimization system in a second aspect, comprising: an environment construction unit for constructing a two-dimensional planar simulation environment based on the horizontal projection of a wafer robotic arm and a working chamber, the two-dimensional planar simulation environment including chamber boundaries, obstacle areas, and robotic arm working areas, and determining a maximum allowable acceleration threshold for limiting wafer slippage relative to the end effector based on the static friction between the wafer and the end effector; a path generation unit for generating candidate wafer transfer paths through a parameterized path generator, using path parameters as input, the path parameters including path shape parameters and obstacle safety margins; and a path evaluation unit for sampling the candidate wafer transfer paths. The evaluation process obtains the theoretical running time, collision determination result, and centripetal acceleration determination result of the candidate wafer transfer path. A reinforcement learning optimization unit is used to take the current path parameters as the reinforcement learning state, the path parameter adjustment amount as the reinforcement learning action, and construct a reward function based on the theoretical running time, the collision determination result, and the centripetal acceleration determination result to train the reinforcement learning optimization model and obtain the optimal path parameters. A control command generation unit is used to input the optimal path parameters into the parameterized path generator to generate the final two-dimensional wafer transfer path and convert the final two-dimensional wafer transfer path into control commands for controlling the wafer transfer action of the wafer robotic arm.
[0018] Furthermore, the path evaluation unit is used to discretely sample the candidate transmission path, calculate the distance between the path sampling point and the prohibited area formed by the obstacle safety margin, the path curvature at the path sampling point, and the centripetal acceleration at the path sampling point; the reinforcement learning optimization unit is used to determine the constraint penalty term based on the distance, the centripetal acceleration, and the maximum allowable acceleration threshold, and construct a reward function based on the constraint penalty term and the time benefit term negatively correlated with the theoretical running time.
[0019] Furthermore, it also includes: a collision verification unit; the collision verification unit is used to map the final two-dimensional wafer transfer path to the three-dimensional motion trajectory of the wafer robot arm in the working chamber before the control command generation unit converts the final two-dimensional wafer transfer path into control commands, and to perform collision verification on the robot arm links, wafers and obstacles in the working chamber corresponding to the three-dimensional motion trajectory using a hierarchical bounding box algorithm; the path generation unit is also used to adjust the obstacle safety margin and / or the path shape parameters when the collision verification fails, and to regenerate candidate wafer transfer paths.
[0020] This application provides a computer device in a third aspect, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor causes the processor to perform a wafer transfer path optimization method as described in any of the technical solutions in the first aspect.
[0021] In a fourth aspect, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform a wafer transfer path optimization method as described in any of the technical solutions in the first aspect.
[0022] Compared with existing technologies, the above technical solution has at least the following advantages: By constructing a two-dimensional planar simulation environment from the horizontal projection of the wafer robotic arm and the working chamber, and converting the static friction between the wafer and the end effector into a maximum allowable acceleration threshold to limit wafer slippage, a parameterized path generator containing path shape parameters and obstacle safety margins is used to generate candidate wafer transfer paths. Then, a reinforcement learning optimization model iteratively adjusts the path parameters based on theoretical running time, collision detection results, and centripetal acceleration detection results. This unifies path generation, obstacle avoidance evaluation, anti-slip plate evaluation, and efficiency optimization in low-attachment wafer transfer into a single parameter optimization process. Therefore, the reliance on manual parameter tuning for wafer transfer paths is reduced. While ensuring that candidate paths do not enter obstacle-prohibited areas and suppressing wafer slippage risks, a final two-dimensional wafer transfer path with a shorter theoretical running time is obtained, thereby improving the safety, stability, and efficiency of wafer transfer path generation.
[0023] The above technical solution also has the following advantages: by determining the maximum allowable acceleration threshold based on the static friction coefficient and equivalent normal action, the evaluation of anti-slip plates can have a clear physical constraint basis; by expanding the obstacle area according to the obstacle safety margin, a collision safety buffer can be reserved in the two-dimensional plane simulation environment; by characterizing the path shape parameters with the offset distance and offset direction of the curved node, the complex path planning problem can be transformed into a finite-dimensional parameter search problem, which is convenient for the reinforcement learning optimization model to express the state and adjust the action; by determining the allowable linear velocity based on the path curvature and the maximum allowable acceleration threshold and further determining the theoretical running time, the plate transfer efficiency evaluation can be matched with the low-adhesion anti-slip plate constraint; by performing a hierarchical bounding box collision verification on the three-dimensional motion trajectory corresponding to the final two-dimensional plate transfer path before the control command is generated, the three-dimensional spatial interference risk that may be missed in the two-dimensional simplified modeling can be supplemented and verified, thereby further improving the executability and reliability of the final plate transfer path in the real working chamber. Attached Figure Description
[0024] Figure 1 A flowchart illustrating a wafer transfer path optimization method provided in this application; Figure 2 A flowchart illustrating another wafer transfer path optimization method provided in this application; Figure 3 A schematic diagram of a wafer transfer path optimization system provided in this application; Figure 4 A schematic diagram of another wafer transfer path optimization system provided in this application; Figure 5 A schematic diagram of the two-dimensional planar simulation environment provided in this application; Figure 6 A schematic diagram of the parameterized path generation provided in this application; Figure 7 A schematic diagram of the candidate path sampling and evaluation provided in this application; Figure 8 The reinforcement learning interaction diagram provided in this application; Figure 9 A schematic diagram of three-dimensional collision verification provided for this application; Figure 10 A schematic diagram of the computer device structure provided in this application. Detailed Implementation
[0025] The technical solution of this application will be clearly and completely described below with reference to the accompanying drawings and embodiments. It should be understood that the following embodiments are only used to illustrate this application and are not intended to limit the scope of protection of this application. In the absence of conflict, the technical features in the following embodiments can be combined with each other. The ordinal numbers such as "first" and "second" involved in this application are only used to distinguish the same or similar objects and do not indicate the order, quantity limitation or importance; expressions such as "including" and "having" indicate non-exclusive inclusion; "multiple" can mean two or more; "at least one" can mean one, two or more; "and / or" means any one or any combination of the listed objects.
[0026] In this application, "wafer transfer" can be understood as the process of moving wafers between loading areas, waiting positions, process chambers, or other workstations using a wafer robotic arm. "Low adhesion scenario" can be understood as a wafer transfer scenario where the wafer and the end effector maintain their relative position primarily through static friction, vacuum suction, or a combination of both, and where there is a risk of slippage relative to the end effector when the wafer's acceleration is too high. "End effector" can be understood as the component at the end of the wafer robotic arm used to carry, support, or hold the wafer, which may include a fork, a carrying end, a suction end, or other end structures capable of holding the wafer.
[0027] In this application, "Reinforcement Learning (RL)" can be understood as an agent interacting with the environment and updating its policy based on reward feedback to obtain action selection methods that satisfy the objective. "State" in this application can represent the current path parameters. "Action" in this application can represent the adjustment amount of the path parameters. "Reward Function" can be understood as a function used to evaluate the comprehensive performance of candidate transfer paths in terms of theoretical running time, collision avoidance constraints, and anti-slip constraints. "Proximal Policy Optimization (PPO)" and "Soft Actor-Critic (SAC)" can both be used as optional algorithms for training reinforcement learning optimization models. "Computer-Aided Design (CAD)" can be used to obtain chamber contours and obstacle geometry information. "Denavit-Hartenberg Parameters (DH parameters)" can be used to describe the kinematic relationships between the robotic arm links and joints. The "hierarchical bounding box algorithm" can be understood as representing the three-dimensional geometric model of the robotic arm links, wafers, and obstacles as a multi-layered bounding box structure, and performing collision verification through the intersection relationship between the bounding boxes.
[0028] Figure 1 This is a flowchart illustrating a wafer transfer path optimization method provided in this application. Figure 1 As shown, the wafer transfer path optimization method may include steps S100 to S500. This method abstracts the horizontal projection of the wafer robotic arm and the working chamber into a two-dimensional planar simulation environment, transforms the risk of wafer slippage relative to the end effector in low-attachment scenarios into a maximum allowable acceleration threshold constraint, and uses a reinforcement learning optimization model to self-tune the path parameters. This generates control commands for controlling the wafer robotic arm to perform the wafer transfer action while satisfying anti-collision and anti-slip constraints.
[0029] S100 constructs a two-dimensional planar simulation environment based on the horizontal projection of the wafer robotic arm and the working chamber, and determines the maximum allowable acceleration threshold for limiting wafer slippage based on the static friction between the wafer and the end effector.
[0030] In this embodiment, the wafer transfer process mainly occurs within an approximately horizontal plane. Therefore, the wafer robotic arm, working chamber, process chamber, chuck interface, sensor base, and other static structures that may affect the transfer path can be projected onto the horizontal plane to form a two-dimensional planar simulation environment. The two-dimensional planar simulation environment can include chamber boundaries, obstacle areas, and robotic arm operating areas. Chamber boundaries define the external space range within which the wafer robotic arm can perform transfer actions; obstacle areas characterize the inaccessible areas of the process chamber, chuck interface, sensor base, and other structures within the working chamber on the horizontal plane; and the robotic arm operating area characterizes the reachable or movable areas of the wafer robotic arm end effector and the wafer during the transfer process.
[0031] In low-adhesion scenarios, the wafer and end effector maintain their relative position primarily through static friction, vacuum adhesion, or a combination of both. When the force corresponding to the acceleration experienced by the wafer during wafer transfer is too large, the wafer may slip relative to the end effector. Therefore, this embodiment determines a maximum permissible acceleration threshold based on the static friction between the wafer and the end effector. This maximum permissible acceleration threshold serves as a safety boundary for the anti-slip plate during subsequent path evaluation. For example, the maximum static friction can be determined based on the static friction coefficient between the wafer and the end effector and the equivalent normal force provided by the end effector to the wafer, and the maximum permissible acceleration threshold can be determined based on the relationship between the maximum static friction and the wafer mass. If the equivalent normal force is converted into equivalent gravitational acceleration, the maximum permissible acceleration threshold can also be determined by the static friction coefficient and the equivalent gravitational acceleration.
[0032] S200, construct a parameterized path generator to generate candidate transport paths with path parameters as input. The path parameters include path shape parameters and obstacle safety margins.
[0033] In this embodiment, a parameterized path generator is used to transform the complex wafer transfer path planning problem into a finite-dimensional path parameter search problem. Path parameters may include path shape parameters and obstacle safety margins. Path shape parameters control the geometry of candidate wafer transfer paths, and may include, for example, the offset distance and direction of at least one curved node set between the wafer transfer start and end points. Obstacle safety margins are used to extend the boundary of the obstacle region to reserve a safe distance between the candidate wafer transfer path and actual obstacles.
[0034] Specifically, the path skeleton can be determined first based on the starting and ending points of the transport, and then one or more curved nodes can be set on the path skeleton. Each curved node can be adjusted in position according to its corresponding offset distance and direction to obtain the offset curved node. The parametric path generator generates a continuous and smooth candidate transport path based on the starting point of the transport, the offset curved node, the ending point of the transport, and the obstacle safety margin. The candidate transport path can be generated using curve interpolation, which can include Bézier curves, spline curves, or polynomial curves, as long as the generated path meets the requirements of continuity and smoothness.
[0035] In this embodiment, the boundary of the obstacle region can be expanded outward based on the obstacle safety margin to form a prohibited area that the candidate transfer path must not enter. The prohibited area is used to determine whether there is a collision risk in the candidate transfer path during subsequent sampling and evaluation. The obstacle safety margin can be used as part of the path parameters in reinforcement learning optimization, enabling the reinforcement learning optimization model to not only adjust the path shape, but also adaptively adjust the safety distance between the path and the obstacle according to different chamber layouts and safety requirements.
[0036] S300 samples and evaluates the candidate transmission paths to obtain the theoretical running time, collision determination results, and centripetal acceleration determination results of the candidate transmission paths.
[0037] In this embodiment, candidate transmission paths can be discretely sampled to obtain multiple path sampling points. For each path sampling point, the distance between the sampling point and the prohibited area can be calculated, and the candidate transmission path can be judged based on this distance to determine whether it meets the anti-collision constraints. When any path sampling point enters the prohibited area, or when the distance between the path sampling point and the prohibited area is less than a preset safety distance, it can be determined that the candidate transmission path does not meet the anti-collision constraints.
[0038] Simultaneously, the path curvature and centripetal acceleration at each sampling point can be calculated, and the comparison between the centripetal acceleration and the maximum allowable acceleration threshold can be used to determine whether the candidate wafer transfer path meets the anti-slip constraint. Specifically, the greater the path curvature, the sharper the bend of the candidate wafer transfer path at that location; with a constant local linear velocity, the greater the path curvature, the greater the corresponding centripetal acceleration. If the centripetal acceleration at a certain sampling point exceeds the maximum allowable acceleration threshold, it indicates that there is a risk of wafer slippage relative to the end effector at that location, and it can be determined that the candidate wafer transfer path does not meet the anti-slip constraint.
[0039] When determining the theoretical running time, the allowable linear velocity corresponding to each path sampling point can be determined based on the path curvature and the maximum allowable acceleration threshold. A velocity profile of the candidate transfer path is then formed based on this allowable linear velocity, and the theoretical running time of the candidate transfer path is determined based on this velocity profile. The velocity profile can be a piecewise velocity sequence obtained by interpolating the allowable linear velocities corresponding to each path sampling point, or it can be generated under the conditions of satisfying the maximum allowable acceleration threshold, the system's maximum linear velocity, and velocity continuity requirements. Therefore, the theoretical running time is not simply determined by the path length, but is simultaneously affected by the path curvature, the maximum allowable acceleration threshold, and the anti-slip plate constraint. This approach enables the path evaluation process to simultaneously reflect transfer efficiency and low-adhesion safety.
[0040] S400 uses the current path parameters as the reinforcement learning state and the path parameter adjustment amount as the reinforcement learning action. It constructs a reward function based on the theoretical running time, collision determination results, and centripetal acceleration determination results, and trains the reinforcement learning optimization model to obtain the optimal path parameters.
[0041] In this embodiment, the reinforcement learning optimization model, the parameterized path generator, the two-dimensional plane simulation environment, and the path evaluation process form an iterative interactive relationship. The current path parameters can be used as the reinforcement learning state, and the path parameter adjustment amount can be used as the reinforcement learning action. The reinforcement learning optimization model outputs the path parameter adjustment amount based on the current path parameters, the parameterized path generator generates new candidate transfer paths based on the adjusted path parameters, and then samples and evaluates the new candidate transfer paths in the two-dimensional plane simulation environment, and feeds back a reward to the reinforcement learning optimization model based on the evaluation results.
[0042] The reward function can include a time reward term negatively correlated with the theoretical running time and a constraint penalty term. The time reward term encourages the reinforcement learning optimization model to search for transport paths with shorter theoretical running times; the constraint penalty term reduces the reward when a candidate transport path does not meet the anti-collision constraint or the anti-slip constraint. For example, a collision penalty can be triggered when a candidate transport path enters a prohibited area or the distance between it and the prohibited area does not meet the safety requirements; an anti-slip penalty can be triggered when the centripetal acceleration at any sampling point in the candidate transport path exceeds the maximum allowable acceleration threshold. Through the above reward function, the reinforcement learning optimization model can search and trade off between shortening the theoretical running time and meeting the safety constraints, thereby obtaining the optimal path parameters.
[0043] During training, the reinforcement learning optimization model can iteratively update its policy parameters multiple times, gradually learning the path parameter adjustment strategies to be adopted under different path parameter states. After training, the reinforcement learning optimization model can output the optimal path parameters. The final two-dimensional slice transfer path corresponding to the optimal path parameters satisfies the anti-collision constraint and the anti-slip constraint, and has a smaller theoretical running time among the candidate slice transfer paths that satisfy the anti-collision constraint and the anti-slip constraint.
[0044] S500 inputs the optimal path parameters into the parameterized path generator to generate the final two-dimensional wafer transfer path, and converts the final two-dimensional wafer transfer path into control commands for controlling the wafer robotic arm to perform wafer transfer actions.
[0045] In this embodiment, after obtaining the optimal path parameters, these parameters can be input into a parameterized path generator to generate the final two-dimensional wafer transfer path. The final two-dimensional wafer transfer path represents the path the wafer takes from the starting point to the ending point in the horizontal projection plane. Subsequently, based on the kinematic model of the wafer robotic arm, the final two-dimensional wafer transfer path can be converted into control commands for the wafer robotic arm. These control commands can include the position, velocity, or acceleration information of the end effector at different times, and can be further converted into the pose, velocity, or drive control quantities of each joint, thereby controlling the wafer robotic arm to perform the wafer transfer action.
[0046] In one alternative implementation, a kinematic model can be established based on the DH parameters of the wafer manipulator, and the end effector pose sequence in the final two-dimensional wafer transfer path can be converted into a joint space pose sequence according to the kinematic model. For scenarios involving speed control requirements, the Cartesian space velocity of the end effector can also be mapped to the joint velocity using the Jacobian matrix to generate joint control commands for execution by the wafer manipulator controller.
[0047] Figure 2 This is a flowchart illustrating another wafer transfer path optimization method provided in this application. Figure 1 Compared to the process shown, Figure 2 The process shown further includes step S600 after step S400, which is used to perform collision verification on the three-dimensional motion trajectory before generating the final control command.
[0048] S600, perform a 3D collision check on the transfer path obtained based on the optimal path parameters; if the 3D collision check passes, proceed to step S500; if the 3D collision check fails, adjust the path shape parameters and / or obstacle safety margin, and return to step S200 to regenerate the candidate transfer path.
[0049] In this embodiment, a two-dimensional planar simulation environment is used to reduce the computational complexity of path optimization and can effectively represent the wafer transfer path constraints that mainly occur in the horizontal plane. However, before actual equipment deployment, the geometric relationships of the robotic arm links, wafers, end effectors, and obstacles in the working chamber in three-dimensional space can still be further considered. To this end, the two-dimensional wafer transfer path generated based on the optimal path parameters can be mapped to the three-dimensional motion trajectory of the wafer robotic arm in the working chamber, and multiple complex poses can be selected along the three-dimensional motion trajectory.
[0050] At each verification pose, a hierarchical bounding box algorithm can be used to perform collision verification on the robotic arm links, wafers, and obstacles in the working chamber. The hierarchical bounding box algorithm first uses the outer bounding box for coarse intersection judgment, and then uses the inner sub-bounding boxes for refined judgment in areas where collisions may occur, thereby reducing the computational load of 3D collision detection while ensuring verification accuracy. If the 3D collision verification passes, it means that the 3D motion trajectory corresponding to the final 2D wafer transfer path meets the 3D spatial safety requirements, and the process can proceed to step S500 to generate control commands. If the 3D collision verification fails, the obstacle safety margin and / or path shape parameters can be adjusted based on the collision verification results, and the process returns to step S200 to regenerate candidate wafer transfer paths to obtain new path optimization results.
[0051] pass Figure 2 As illustrated, this application enables efficient path parameter optimization in a two-dimensional plane and supplements the verification of potential spatial collision risks missed in the simplified two-dimensional model using three-dimensional collision verification before the generation of control commands. This reduces the computational complexity of wafer transfer path optimization in low-attachment scenarios and improves the safety and reliability of the final execution path in the actual working chamber.
[0052] Figure 3 This is a schematic diagram of a wafer transfer path optimization system provided in this application. Figure 3 As shown, the wafer transfer path optimization system may include an environment construction unit 11, a path generation unit 12, a path evaluation unit 13, a reinforcement learning optimization unit 14, and a control command generation unit 15. The environment construction unit 11, path generation unit 12, path evaluation unit 13, reinforcement learning optimization unit 14, and control command generation unit 15 can be implemented through software modules, hardware modules, or a combination of software and hardware, and can be deployed in the controller, host computer, edge computing device, or other computer equipment of the wafer transfer device.
[0053] The environment construction unit 11 is used to construct a two-dimensional planar simulation environment based on the horizontal projection of the wafer robotic arm and the working chamber, and to determine the maximum allowable acceleration threshold for limiting wafer slippage based on the static friction between the wafer and the end effector. The two-dimensional planar simulation environment may include the chamber boundary, obstacle area, and robotic arm working area. The environment construction unit 11 can also form a constraint basis for the two-dimensional planar simulation environment based on the obstacle area, robotic arm working area, maximum allowable acceleration threshold, and other safety constraint information. This constraint basis for the two-dimensional planar simulation environment can be sent to the path generation unit 12 for the generation of candidate wafer transfer paths; it can also be sent to the reinforcement learning optimization unit 14 for the reinforcement learning optimization model to identify the constraint environment of the current optimization problem during the path parameter search process.
[0054] The path generation unit 12 is used to construct a parameterized path generator and generate candidate transport paths using path parameters as input. Path parameters may include path shape parameters and obstacle safety margins. Path shape parameters may include the offset distance and offset direction of curved nodes, and obstacle safety margins are used to create prohibited areas outside obstacle regions to reserve a safe distance between the candidate transport path and obstacles. The path generation unit 12 can generate candidate transport paths based on the two-dimensional planar simulation environment constraints provided by the environment construction unit 11, combined with the current path parameters, and output the candidate transport paths to the path evaluation unit 13.
[0055] The path evaluation unit 13 is used to sample and evaluate the candidate transfer paths generated by the path generation unit 12, and obtain evaluation results. The evaluation results may include the theoretical running time of the candidate transfer path, collision determination results, and centripetal acceleration determination results. Specifically, the path evaluation unit 13 can perform discrete sampling on the candidate transfer paths, calculate the distance between the path sampling points and the prohibited area, and obtain the collision determination result based on this distance; the path evaluation unit 13 can also calculate the path curvature and centripetal acceleration at the path sampling points, and obtain the centripetal acceleration determination result based on the comparison between the centripetal acceleration and the maximum allowable acceleration threshold. The path evaluation unit 13 can feed the evaluation results back to the reinforcement learning optimization unit 14, so that the reinforcement learning optimization unit 14 can construct the reward function and update the policy parameters.
[0056] The reinforcement learning optimization unit 14 uses the current path parameters as the reinforcement learning state and the path parameter adjustment amount as the reinforcement learning action. Based on the evaluation results from the path evaluation unit 13, it constructs a reward function and trains the reinforcement learning optimization model to obtain the optimal path parameters. Specifically, the path generation unit 12 can send the current path parameters or state information represented by the current path parameters to the reinforcement learning optimization unit 14; the path evaluation unit 13 can send the theoretical running time, collision determination results, and centripetal acceleration determination results to the reinforcement learning optimization unit 14. The reinforcement learning optimization unit 14 determines the time reward term and constraint penalty term based on the above information and updates the policy parameters of the reinforcement learning optimization model based on the reward function feedback. After training, the reinforcement learning optimization unit 14 outputs the optimal path parameters and sends them to the path generation unit 12.
[0057] The path generation unit 12 is also used to generate a final two-dimensional wafer transfer path based on the optimal path parameters output by the reinforcement learning optimization unit 14. The final two-dimensional wafer transfer path can be a wafer transfer path that satisfies anti-collision constraints and anti-slip constraints and has a relatively short theoretical running time. The path generation unit 12 or the path evaluation unit 13 can output the final two-dimensional wafer transfer path to the control command generation unit 15. The control command generation unit 15 is used to convert the final two-dimensional wafer transfer path into control commands for controlling the wafer transfer action of the wafer robotic arm. These control commands may include the pose sequence, velocity information, and acceleration information of the end effector, or further converted robotic arm joint control information.
[0058] pass Figure 3 The system shown comprises an environment construction unit 11 providing the basic constraints of a two-dimensional planar simulation environment, a path generation unit 12 generating candidate wafer transfer paths based on path parameters, a path evaluation unit 13 evaluating the candidate wafer transfer paths, a reinforcement learning optimization unit 14 updating the path parameters based on the evaluation results, and a control command generation unit 15 converting the final two-dimensional wafer transfer path into control commands. Thus, the system forms a closed-loop optimization structure of "path generation—path evaluation—reinforcement learning optimization—final path execution," thereby reducing the reliance on manual debugging of wafer transfer path parameters in low-attachment scenarios and improving the safety and operational efficiency of the wafer transfer path.
[0059] Figure 4 This is a schematic diagram of another wafer transfer path optimization system provided in this application. Figure 4 As shown, this wafer transfer path optimization system... Figure 3 The system shown further includes a collision verification unit 16. The collision verification unit 16 is located between the path evaluation unit 13 and the control command generation unit 15, and is used to perform collision verification on the three-dimensional motion trajectory corresponding to the final two-dimensional transfer path before the control command is generated.
[0060] In this embodiment, the path evaluation unit 13 can output the final two-dimensional wafer transfer path to the collision verification unit 16. The collision verification unit 16 is used to map the final two-dimensional wafer transfer path to the three-dimensional motion trajectory of the wafer robotic arm in the working chamber, and to perform collision verification on the robotic arm links, wafers, and obstacles in the working chamber corresponding to the three-dimensional motion trajectory using a hierarchical bounding box algorithm. The hierarchical bounding box algorithm may include establishing bounding boxes for robotic arm links, wafers, and obstacles, and determining whether there is a collision risk in the corresponding three-dimensional motion trajectory through the intersection relationship between the bounding boxes.
[0061] If the collision verification unit 16 determines that the 3D collision verification is successful, it can output the successful final 2D wafer transfer path or its corresponding 3D motion trajectory to the control command generation unit 15. The control command generation unit 15 generates control commands for controlling the wafer robotic arm to perform wafer transfer actions based on the successful path. If the collision verification unit 16 determines that the 3D collision verification is unsuccessful, it feeds back parameter adjustment information to the path generation unit 12. The parameter adjustment information may include suggestions for adjusting obstacle safety margins, suggestions for adjusting path shape parameters, or a combination of both. The path generation unit 12 regenerates candidate wafer transfer paths based on the parameter adjustment information, and these paths are evaluated and verified again by the path evaluation unit 13, the reinforcement learning optimization unit 14, and the collision verification unit 16.
[0062] pass Figure 4 The system shown, based on efficient path generation and reinforcement learning optimization in a two-dimensional planar simulation environment, further introduces a three-dimensional collision verification mechanism. This mechanism can supplement and verify three-dimensional spatial collision risks that may not be fully reflected in the simplified two-dimensional modeling before the control command is generated. If the verification fails, it returns to the path generation stage to regenerate candidate transfer paths by adjusting obstacle safety margins and / or path shape parameters. Thus, it retains the computational efficiency of two-dimensional planar path optimization while improving the safety and reliability of the final control commands when executed in the actual working chamber.
[0063] Figure 5 This is a schematic diagram of the two-dimensional planar simulation environment provided in this application. Figure 5As shown, in an exemplary wafer transfer scenario, the horizontal projection of the working chamber can be defined by the chamber boundary, within which are set loading area S0, process chamber C0, process chamber C1, process chamber C2, process chamber C3, process chamber C4, and process chamber C5. Loading area S0 can be considered the starting point of wafer transfer, and process chamber C3 can be considered the ending point of wafer transfer. The wafer robotic arm is located in the middle of the chamber, and its reachable or movable range can be represented as the robotic arm's working area. Process chambers C0 to C5, loading area S0, and other static structures within the chamber can all be considered as obstacle or constraint areas in a two-dimensional planar simulation environment.
[0064] In this embodiment, the two-dimensional planar simulation environment is not a simple illustration of the actual working chamber, but rather a computational environment used for path generation, path evaluation, and reinforcement learning optimization. Specifically, the process chamber, chuck interface, sensor base, mechanical mounting base, and other structures in the actual working chamber that may interfere with the wafer or end effector can be projected onto a horizontal plane to form corresponding two-dimensional obstacle regions. Since wafer transfer motion mainly occurs within an approximately horizontal plane, using a two-dimensional planar simulation environment can reduce the complexity of path search and constraint evaluation while preserving the main spatial constraints.
[0065] like Figure 5 As shown, the boundaries of each obstacle region can be expanded outwards based on the obstacle safety margin to form a prohibited area where candidate transport paths are not allowed to enter. The obstacle safety margin corresponds to the interval between the solid and dashed boundaries of the obstacle, for example... Figure 5 The area between the solid frame and the outer dashed frame of the central process cavity C3 represents the obstacle safety margin. By setting the obstacle safety margin, a buffer distance can be reserved between the candidate transfer path and the actual obstacle, thereby reducing the risk of collisions caused by modeling errors, control errors, or mechanical assembly errors during actual execution.
[0066] exist Figure 5 In the illustrated embodiment, the wafer transfer starting point is located near the loading area S0, and the wafer transfer ending point is located near the process cavity C3. The candidate wafer transfer path needs to start from the starting point, pass through the robotic arm's operating area, and finally reach the ending point. During path generation and evaluation, the candidate wafer transfer path should avoid prohibited areas formed by the expansion of various obstacle areas, and should also ensure that the wafer satisfies the anti-slip constraint determined by static friction during the wafer transfer process. Thus, the two-dimensional planar simulation environment simultaneously provides the spatial basis for anti-collision constraints, anti-slip constraints, and wafer transfer efficiency evaluation.
[0067] In this embodiment, the static friction force between the wafer and the end effector can be used to determine the maximum permissible acceleration threshold. For example, if... The coefficient of static friction between the wafer and the end effector is expressed as... The maximum allowable acceleration threshold represents the equivalent gravitational acceleration calculated from the vacuum adsorption force. It can be represented as:
[0068] in, This indicates the maximum permissible acceleration threshold that a wafer can withstand without slippage. This represents the coefficient of static friction between the wafer and the end effector. This represents the equivalent gravitational acceleration calculated from the vacuum adsorption force. This maximum permissible acceleration threshold can be used as the safety boundary for anti-slip plates in the subsequent evaluation of candidate transfer paths. When the centripetal acceleration corresponding to any sampling point on the candidate transfer path exceeds this maximum permissible acceleration threshold, it can be determined that the candidate transfer path has a risk of slippage.
[0069] Figure 6 A schematic diagram illustrating the generation of the parameterized path provided in this application. For example... Figure 6 As shown, a path skeleton can be determined first between the starting and ending points of the film transfer process. This path skeleton can be a straight line segment connecting the starting and ending points. At least one curved node is set on the path skeleton, for example... Figure 6 The curved node shown in the figure and bending nodes Bending node and The initial position is located on the path skeleton. During path optimization, the curved node can be offset along a preset offset direction to obtain the offset curved node. and .
[0070] In this embodiment, the offset of each curved node can be determined by both the offset distance and the offset direction. Figure 6 The curved node in For example, its offset distance can be expressed as The offset direction can be expressed as With curved nodes For example, its offset distance can be expressed as The offset direction can be expressed as The offset direction can be represented by discrete values, such as "+1" and "-1" representing offsets in two opposite directions along the path skeleton normal, respectively; the offset distance can be represented by continuous values to control the degree to which the curved node deviates from the path skeleton. By adjusting the offset direction and offset distance of the curved node, the candidate transfer path can be made to bypass obstacles and improve the path curvature distribution.
[0071] For those with A parametric path generator for each curved node can combine the offset distance, offset direction, and obstacle safety margin of each curved node to form path parameters. These path parameters can be defined as follows: A dimensional adjustable hyperparameter vector can be represented as:
[0072] in, Indicates path parameters, Indicates the first The offset distance of each curved node. Indicates the first The offset direction of each curved node, Indicates the safety margin of obstacles. Used to expand the boundaries of obstacles outwards, thus creating a restricted area. Given path parameters. Then, the parameterized path generator can generate candidate transfer paths based on the transfer start point, the offset curved node, and the transfer end point.
[0073] Furthermore, if we take Indicates the starting point of the video transmission, with Indicates the end point of the video transmission, with , ... Representing the offset curved node, the parameterized path generator can generate candidate transfer paths using curve interpolation. A candidate transfer path can be represented as... .
[0074] In one alternative implementation, a candidate transfer path can be represented in the following form:
[0075] in, This represents the curve interpolation function. This function can be implemented using Bezier curves, spline curves, or polynomial curves, etc. Candidate transfer paths generated in this way exhibit good continuity in position and in their first and second derivatives, thus avoiding sharp inflection points and facilitating subsequent satisfaction of the maximum permissible acceleration threshold constraint.
[0076] exist Figure 6 In the illustrated embodiment, the dashed outline of the obstacle's outer side represents the obstacle's safety margin. The resulting forbidden region. Candidate transfer paths should geometrically avoid this forbidden region. During reinforcement learning optimization, the offset distance... , Offset direction , and obstacle safety margin All of these can be used as adjustable path parameters in the optimization process. Through these adjustment methods, the reinforcement learning optimization model can not only adjust the curvature of the path, but also dynamically change the safe buffer distance between the candidate transport path and obstacles based on obstacle distribution and path evaluation results.
[0077] Figure 7 This is a schematic diagram illustrating the sampling and evaluation of candidate paths provided in this application. Figure 7 As shown, candidate transmission paths can be discretely sampled to obtain multiple path sampling points. Figure 7 China-Israel path sampling points Taking this as an example, we will demonstrate the anti-collision constraints and anti-slip constraints of the candidate transfer path. The following section will combine... Figure 7 Explanation is provided regarding path sampling points. It can calculate its distance to the prohibited area. This distance It can be used to determine whether candidate slice paths meet anti-collision constraints. Typically, path sampling points... Distance to the prohibited area The larger the value, the safer the candidate transfer path is near the sampling point; if the path sampling point Entering prohibited areas or path sampling points If the distance to the prohibited area does not meet the preset safety requirements, it can be determined that there is a collision risk in the candidate transmission path.
[0078] In this embodiment, path sampling points It can also be used for local curvature analysis. Specifically, it can be used at path sampling points. Local radius of curvature in the vicinity of determining candidate transfer paths and center of curvature radius of curvature Used to reflect the candidate transmission path at the path sampling point The degree of curvature at that point. Radius of curvature. The larger the radius of curvature, the gentler the candidate slice path at that location; The smaller the value, the sharper the bend in the candidate wafer transfer path at that location. In wafer transfer scenarios, the sharper the bend, the greater the centripetal acceleration generated at the same local speed, and the higher the risk of wafer slippage relative to the end effector.
[0079] like Figure 7 As shown, path sampling points It also has local linear velocity. The centripetal acceleration at this point can be calculated using the local velocity and radius of curvature, and its form is as follows:
[0080] in, Indicates path sampling points Centripetal acceleration at the point, Indicates path sampling points Local linear velocity at that point Indicates path sampling points The radius of curvature at that point. Due to the curvature With radius of curvature The centripetal acceleration can also be expressed as the reciprocals of each other:
[0081] in, Indicates path sampling points The path curvature at that point. This is achieved through centripetal acceleration. With the maximum allowable acceleration threshold By comparing the results, it can be determined whether the candidate transfer path satisfies the anti-slip plate constraint. If any path sampling point exists... Make If so, it can be determined that there is a risk of slippage in the candidate transmission path, and the anti-slippage penalty will be triggered in the reward function.
[0082] In one alternative implementation, when determining the constraint of the anti-slip pad, the maximum permissible acceleration threshold can be used. and path sampling points Path curvature at Determine the allowable linear velocity at the sampling point of this path. And a velocity profile is formed based on the allowable linear velocity. Allowable linear velocity The calculation formula is: ,in For the square root function, when At that time, it can be Take as the system's maximum linear velocity Based on the local velocities corresponding to the sampling points along each candidate transmission path, a velocity profile of the candidate transmission path can be formed, and the theoretical running time required to run along the candidate transmission path can be calculated through numerical integration. The theoretical runtime obtained in this way takes into account path length, path curvature, and maximum allowable acceleration threshold, thus more accurately reflecting the actual executable efficiency of candidate transfer paths in low-attachment scenarios.
[0083] In this embodiment, the sampling evaluation results of the candidate transmission path may include theoretical running time, collision determination results, and centripetal acceleration determination results. The collision determination result can be determined by the distance from the path sampling point to the prohibited area; the centripetal acceleration determination result can be determined by comparing the centripetal acceleration at the path sampling point with the maximum allowable acceleration threshold. These evaluation results can be fed back to the reinforcement learning optimization model to construct the reward function and update the policy parameters.
[0084] This application provides a reward function, which can be expressed as:
[0085] in, Represents path parameters The corresponding reward value, Indicates the path along the candidate transfer path The theoretical minimum time required for operation This indicates the penalty. The shorter the theoretical minimum time, the better. The larger the value, the higher the corresponding time gain. Penalty item. It is a non-negative value and is activated when a candidate transmission path violates hard constraints.
[0086] For example, penalty items This can be determined according to the following logic, and the pseudocode form of this logic is as follows: “ if (path) (Distance from any point on the path to the expanded obstacle < 0) OR (path) curvature at any point on The corresponding centripetal acceleration : For a very large positive value else: " in, This indicates a pre-defined large positive value penalty term; "path" "The distance from any point on the path to the expanded obstacle is less than 0" indicates that the candidate transmission path has entered the prohibited area, triggering a collision penalty; "path curvature at any point on The corresponding centripetal acceleration "This indicates that the centripetal acceleration at a certain location on the candidate transmission path exceeds the maximum allowable acceleration threshold." This triggers a slippage penalty. Through the aforementioned reward function, the reinforcement learning optimization model can avoid selecting paths into prohibited areas, and also avoid selecting paths where excessive curvature or speed would cause centripetal acceleration to exceed the maximum permissible acceleration threshold. Therefore, Figure 7 The corresponding sampling and evaluation process unifies anti-collision safety, anti-slip plate safety, and plate transfer efficiency into the same evaluation framework, providing calculable feedback for the automatic optimization of path parameters.
[0087] Figure 8 A schematic diagram of the reinforcement learning interaction provided in this application. Figure 8 As shown, the overall interaction chain of reinforcement learning can include: inputting the current path parameters or current state into the reinforcement learning optimization model, which outputs path parameter adjustments or actions; inputting these adjustments into a parameterized path generator, which generates candidate transfer paths and evaluates them in a two-dimensional simulation environment; subsequently, the path evaluation / reward module generates reward feedback and an updated state based on the theoretical running time, collision determination results, and centripetal acceleration determination results output from the two-dimensional simulation environment, which are then fed back to the reinforcement learning optimization model for policy updates. Through this closed-loop interaction, the reinforcement learning optimization model can gradually converge to the path parameter configuration that maximizes the reward function.
[0088] In this embodiment, the state of the reinforcement learning optimization model can be defined as the current path parameter vector. An action can be defined as an action on the current path parameter vector. Incremental adjustment vector The updated state after an action is performed can be represented as:
[0089] in, This represents the current path parameter vector. This represents the adjustment amount of the path parameters in the reinforcement learning optimization model output. This represents the updated path parameter vector obtained after performing the action. In this application, It can be used as the state input for a reinforcement learning optimization model. It can be used as the action output for optimizing reinforcement learning models.
[0090] In one specific implementation, combined with Figures 5 to 8 The following example parameters can be used for training and verification. Specifically, the process cavities C0 to C5 and their corresponding chuck interfaces can be modeled as 300mm × 200mm bounding rectangles, the loading area S0 can be modeled as a circular region with a diameter of 250mm, and all obstacle boundaries can be expanded outward with an initial safety margin. Set the static friction coefficient between the wafer and the robotic arm fork. The equivalent gravitational acceleration is set by the vacuum adsorption force. Then the maximum permissible acceleration threshold At this point, two curved nodes can be set on the path skeleton connecting the center points of S0 and C3. and The PPO algorithm was used as the reinforcement learning agent for training. After a total of 25,000 simulation interactions, the reinforcement learning optimization model converged to obtain the optimal combination of path parameters. In the example provided in this embodiment, the optimal path parameter combination... It can be represented as: After inputting the optimal path parameter combination into the parameterized path generator, the final two-dimensional slice transfer path can be obtained; the minimum radius of curvature of this final two-dimensional slice transfer path is approximately 180 mm, and the peak centripetal acceleration throughout the path is approximately... lower than The actual measured transfer time of the physical device is approximately 1.68 seconds, which is about 8.7% more efficient than the baseline path of 1.84 seconds manually adjusted by engineers. Furthermore, no slippage or collision events occurred in 10,000 consecutive repeated tests.
[0091] Figure 9 A schematic diagram of the three-dimensional collision verification provided for this application. (See attached diagram.) Figure 9 As shown, after obtaining the final two-dimensional slice path, the sequence of path points in the final two-dimensional slice path can be, for example... , , , This is mapped to the three-dimensional motion trajectory of the robotic arm in the three-dimensional working chamber. Figure 9 The diagram illustrates a 3D working chamber, a wafer transfer robot arm, a robot arm link bounding box, a wafer bounding box, an obstacle bounding box, a hierarchical bounding box, and a 3D motion trajectory. Specifically, the robot arm links, wafers, and obstacles in the working chamber can all construct corresponding outer bounding boxes and inner sub-bounding boxes, thus forming a hierarchical bounding box structure. Collision verification can be completed by traversing multiple discrete poses along the 3D motion trajectory and detecting the intersection relationships between the robot arm link bounding box, the wafer bounding box, and the obstacle bounding box.
[0092] In this embodiment, collision verification can be considered as an optional safety enhancement step after 2D path optimization and before control command generation. Specifically, the final 2D wafer transfer path can be mapped to a 3D motion trajectory first. Then, bounding box models are established for the robotic arm links, wafers, and obstacles in the 3D working chamber, and a hierarchical bounding box algorithm is used for coarse-to-fine collision judgment. When there is no intersection between the outer bounding boxes, it can be directly determined that there is no collision between the corresponding objects; when there is an intersection between the outer bounding boxes, the inner sub-bounding boxes are used for further refinement verification to improve the accuracy of collision judgment and reduce the overall computational complexity.
[0093] exist Figure 9 In the illustrated implementation, the collision verification result can be categorized into two cases: "verification passed" and "verification failed." If the collision verification passes, it indicates that the three-dimensional motion trajectory mapped from the final two-dimensional wafer transfer path meets the three-dimensional spatial safety requirements. At this point, the control command generation stage can begin, where the control module generates control commands for controlling the wafer robotic arm to perform the wafer transfer action based on the verified path. If the collision verification fails, it indicates that although the path in the two-dimensional plane meets the anti-collision constraints and anti-slip constraints, there may still be a risk of spatial interference in the real three-dimensional space. In this case, it can be handled according to... Figure 9 As shown, the verification results are fed back to the path planning stage to adjust the obstacle safety margin and / or path shape parameters. Candidate transfer paths are then regenerated, and path evaluation and reinforcement learning optimization are performed again until a final path that satisfies both two-dimensional safety constraints and three-dimensional collision verification requirements is obtained. Therefore, Figure 9 The implementation shown can further improve the security and reliability of the final execution path in real devices while maintaining the high efficiency of two-dimensional path optimization.
[0094] Figure 10 A schematic diagram of the computer device structure provided in this application. (For example...) Figure 10 As shown, the computer device may include a processor, a system bus, internal memory, a non-volatile storage medium, a network interface, a display screen, and an input device. The non-volatile storage medium may store the operating system and computer programs; the processor communicates with the internal memory, non-volatile storage medium, network interface, display screen, and input device via the system bus. This computer device can be a host computer, industrial controller, edge computing device, embedded control device, or other computing platform capable of executing the wafer transfer path optimization method described in this application.
[0095] In this embodiment, when the computer program stored in the non-volatile storage medium is executed by the processor, the processor can perform the steps in the aforementioned method embodiments. Specifically, the processor can perform steps such as constructing a two-dimensional planar simulation environment, determining the maximum allowable acceleration threshold based on static friction, constructing a parameterized path generator, generating candidate transport paths, performing path sampling evaluation, calculating theoretical running time, collision determination results and centripetal acceleration determination results, constructing a reward function, training a reinforcement learning optimization model, generating the final two-dimensional transport path, performing three-dimensional collision verification, and converting the final two-dimensional transport path into control commands. The network interface can be used for data interaction with external devices, such as receiving chamber CAD layout information, receiving process equipment configuration data, and uploading path optimization results or control logs; the display screen can be used to display the current optimization process, path planning results, or alarm information; the input device can be used to receive manually configured parameters, running instructions, or maintenance instructions.
[0096] In one alternative implementation, Figure 10 The computer device shown can also communicate with the motion controller of the wafer transfer equipment to send control commands generated according to the method of this application to the physical wafer robotic arm for execution. For scenarios employing offline training and online deployment, the training process of the reinforcement learning optimization model can be completed on a training device with strong computing power. Figure 10 The computer equipment shown can primarily undertake online deployment tasks such as path generation, path evaluation, collision verification, and control command generation. For scenarios employing online adaptive optimization, Figure 10 The computer device shown can also perform the functions of policy updates and control outputs simultaneously.
[0097] Furthermore, this application can also correspond to a computer-readable storage medium. This computer-readable storage medium stores a computer program, which, when executed by a processor, causes the processor to perform the steps in the aforementioned method embodiments. The computer-readable storage medium can be a disk, optical disk, solid-state drive, flash memory, read-only memory, random access memory, or other tangible storage medium capable of storing program code. By employing the above-described computer device and computer-readable storage medium implementation methods, the wafer transfer path optimization method of this application can be implemented not only as a method and system, but also as a software product, program deployment package, or embedded control program, thereby improving the engineering applicability and deployment flexibility of the technical solution of this application.
[0098] In summary, combining Figures 5 to 10This application constructs a two-dimensional planar simulation environment to abstract the main constraints of wafer transfer in low-attachment scenarios into a two-dimensional plane. A parameterized path generator transforms the wafer transfer path planning problem into a finite-dimensional path parameter search problem. A reinforcement learning optimization model then automatically explores and updates strategies within the path parameter space. Reward feedback is constructed by combining theoretical running time, collision detection results, and centripetal acceleration detection results, thereby obtaining a final two-dimensional wafer transfer path that satisfies anti-collision and anti-slip constraints while achieving superior transfer efficiency. If further deployment security is required, a three-dimensional collision verification step can be used to supplement and verify the three-dimensional motion trajectory corresponding to the final two-dimensional wafer transfer path. Finally, using computer equipment and computer-readable storage media, the above path optimization process is implemented as an executable control program and control instructions, thereby achieving safe and efficient wafer transfer for the wafer robotic arm in low-attachment scenarios.
[0099] The above embodiments are merely illustrative of the technical concept and features of this application, intended to enable those skilled in the art to understand the content of this application and implement it accordingly, and should not be construed as limiting the scope of protection of this application. It is obvious to those skilled in the art that this application is not limited to the details of the above exemplary embodiments, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application. Therefore, the embodiments should be considered exemplary and non-limiting in all respects. The scope of this application is defined by the appended claims rather than the foregoing description, and thus all variations falling within the meaning and scope of the equivalents of the claims are intended to be included within this application.
Claims
1. A wafer transfer path optimization method, characterized in that, include: A two-dimensional planar simulation environment is constructed based on the horizontal projection of the wafer robotic arm and the working chamber. The two-dimensional planar simulation environment includes the chamber boundary, the obstacle area and the robotic arm working area. The maximum allowable acceleration threshold for limiting wafer slippage is determined based on the static friction between the wafer and the end effector. Construct a parameterized path generator, which generates candidate transport paths with path parameters as input, including path shape parameters and obstacle safety margins; The candidate transmission paths are sampled and evaluated to obtain the theoretical running time, collision determination results, and centripetal acceleration determination results of the candidate transmission paths; The current path parameters are used as the reinforcement learning state, the path parameter adjustment amount is used as the reinforcement learning action, and a reward function is constructed based on the theoretical running time, the collision determination result, and the centripetal acceleration determination result. The reinforcement learning optimization model is then trained to obtain the optimal path parameters. The optimal path parameters are input into the parameterized path generator to generate the final two-dimensional wafer transfer path, and the final two-dimensional wafer transfer path is converted into control commands for controlling the wafer robotic arm to perform wafer transfer actions.
2. The wafer transfer path optimization method according to claim 1, characterized in that, The maximum allowable acceleration threshold is determined based on the static friction coefficient between the wafer and the end effector and the equivalent normal force provided by the end effector to the wafer. The equivalent normal force includes the equivalent force obtained by converting the wafer's gravity force and / or vacuum adsorption force.
3. The wafer transfer path optimization method according to claim 1, characterized in that, Also includes: Based on the obstacle safety margin, the boundary of the obstacle area is expanded outward to form a prohibited area that the candidate transfer path is not allowed to enter; The obstacle area includes the projection area of the process chamber, chuck interface, sensor base and / or other static structures in the working chamber onto the horizontal plane.
4. The wafer transfer path optimization method according to claim 1, characterized in that, The path shape parameters include the offset distance and offset direction of at least one curved node set between the starting point and the ending point of the transfer; The parameterized path generator generates a continuous and smooth candidate transport path based on the transport start point, the offset bending node, the transport end point, and the obstacle safety margin.
5. The wafer transfer path optimization method according to claim 3, characterized in that, The sampling and evaluation of the candidate transmission paths includes: Discrete sampling is performed on the candidate transmission paths to obtain multiple path sampling points; Calculate the distance between each of the path sampling points and the prohibited area, and determine whether the candidate transmission path satisfies the anti-collision constraint based on the distance; Calculate the path curvature and centripetal acceleration at each of the path sampling points, and determine whether the candidate transfer path satisfies the anti-slip plate constraint based on the comparison result between the centripetal acceleration and the maximum allowable acceleration threshold.
6. The wafer transfer path optimization method according to claim 5, characterized in that, The allowable linear velocity corresponding to each of the path sampling points is determined based on the path curvature and the maximum allowable acceleration threshold, and a velocity profile of the candidate transmission path is formed based on the allowable linear velocity. The theoretical running time of the candidate transmission path is determined based on the velocity profile.
7. The wafer transfer path optimization method according to claim 5, characterized in that, The reward function includes a time benefit term that is negatively correlated with the theoretical running time and a constraint penalty term; The constraint penalty term is activated when the candidate transfer path does not satisfy the anti-collision constraint or the anti-slip constraint.
8. The wafer transfer path optimization method according to claim 1, characterized in that, Training the reinforcement learning optimization model includes: In the two-dimensional planar simulation environment, the reinforcement learning optimization model interacts iteratively with the parameterized path generator. The path parameter adjustment amount is output based on the current path parameters, and the parameterized path generator generates a new candidate transmission path based on the adjusted path parameters. The policy parameters of the reinforcement learning optimization model are updated based on the feedback from the reward function until a policy is obtained for outputting the optimal path parameters.
9. The wafer transfer path optimization method according to claim 5, characterized in that, The final two-dimensional transfer path corresponding to the optimal path parameters satisfies the anti-collision constraint and the anti-slip constraint, and has a smaller theoretical running time among the candidate transfer paths that satisfy the anti-collision constraint and the anti-slip constraint.
10. The wafer transfer path optimization method according to claim 1, characterized in that, Before converting the final two-dimensional wafer transfer path into control commands, the final two-dimensional wafer transfer path is mapped to the three-dimensional motion trajectory of the wafer robotic arm in the working chamber, and a hierarchical bounding box algorithm is used to perform collision verification on the robotic arm links, wafers and obstacles in the working chamber corresponding to the three-dimensional motion trajectory. If the collision verification fails, adjust the obstacle safety margin and / or the path shape parameters, and regenerate the candidate transfer path.
11. A wafer transfer path optimization system, characterized in that, include: An environment construction unit is used to construct a two-dimensional planar simulation environment based on the horizontal projection of the wafer robotic arm and the working chamber. The two-dimensional planar simulation environment includes the chamber boundary, obstacle area and robotic arm working area, and determines the maximum allowable acceleration threshold for limiting the slippage of the wafer relative to the end effector based on the static friction between the wafer and the end effector. The path generation unit is used to generate candidate transport paths by a parameterized path generator, taking path parameters as input. The path parameters include path shape parameters and obstacle safety margin. The path evaluation unit is used to sample and evaluate the candidate slice path to obtain the theoretical running time, collision determination result and centripetal acceleration determination result of the candidate slice path. The reinforcement learning optimization unit is used to take the current path parameters as the reinforcement learning state, the path parameter adjustment amount as the reinforcement learning action, and construct a reward function based on the theoretical running time, the collision determination result and the centripetal acceleration determination result to train the reinforcement learning optimization model to obtain the optimal path parameters. The control instruction generation unit is used to input the optimal path parameters into the parameterized path generator to generate the final two-dimensional wafer transfer path, and convert the final two-dimensional wafer transfer path into control instructions for controlling the wafer robotic arm to perform wafer transfer actions.
12. The wafer transfer path optimization system according to claim 11, characterized in that, The path evaluation unit is used to perform discrete sampling on the candidate transmission path, and calculate the distance between the path sampling point and the prohibited area formed by the obstacle safety margin, the path curvature at the path sampling point, and the centripetal acceleration at the path sampling point. The reinforcement learning optimization unit is used to determine a constraint penalty term based on the distance, the centripetal acceleration, and the maximum allowable acceleration threshold, and to construct a reward function based on the constraint penalty term and a time benefit term that is negatively correlated with the theoretical running time.
13. The wafer transfer path optimization system according to claim 11 or 12, characterized in that, Also includes: Collision verification unit; The collision verification unit is used to map the final two-dimensional wafer transfer path into the three-dimensional motion trajectory of the wafer robot arm in the working chamber before the control command generation unit converts the final two-dimensional wafer transfer path into control commands, and to use a hierarchical bounding box algorithm to perform collision verification on the robot arm links, wafers and obstacles in the working chamber corresponding to the three-dimensional motion trajectory. The path generation unit is also used to adjust the obstacle safety margin and / or the path shape parameters when the collision verification fails, and to regenerate the candidate transfer path.
14. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the computer program is executed by the processor, the processor performs the wafer transfer path optimization method as described in any one of claims 1 to 10.
15. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the processor performs the wafer transfer path optimization method as described in any one of claims 1 to 10.
Citation Information
Patent Citations
Glue uniformizing equipment capable of shortening wafer transmission time
CN119987139A
Multi-mode real-time obstacle avoidance method and system for wafer transmission robot
CN121008573A