Moire fringe automatic leveling method and device

The Moiré fringe images are trained through the Deep Q-Network network, and the problems of reduced alignment accuracy and insufficient multi-process fusion processing capabilities in the Moiré fringe alignment method are solved, real-time monitoring and rapid response of Moiré fringe images are realized, and lithography production efficiency is improved.

CN120469178APending Publication Date: 2025-08-12INST OF OPTICS & ELECTRONICS CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510802615.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing moiré stripe alignment methods lack the fusion processing capability of lithography multi-process, and the inclination of moiré stripes leads to a decrease in alignment accuracy, failing to monitor image changes in real time and respond quickly.

Method used

The Deep Q-Network network is used to train moiré stripes, and the images are recorded through the CCD camera. Moiré stripes are generated using the angle or displacement of the mask and the silicon wafer grating. The network is trained to output the silicon wafer action and iterate over the Q value to achieve automatic leveling of the mask and the silicon wafer.

Benefits of technology

Real-time monitoring and rapid response of moiré stripe images are realized, the fusion processing capability of multi-process lithography is improved, the alignment accuracy is reduced, the alignment time is shortened, and the production efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120469178A_ABST
    Figure CN120469178A_ABST
Patent Text Reader

Abstract

The invention discloses an automatic moire fringe leveling method and device, and belongs to the technical field of optical measurement, and the method comprises the steps: generating moire fringes through a preset included angle or relative displacement between a mask and a grating on a silicon wafer, and recording a generated moire fringe image; taking a plurality of continuous moire fringe images with different inclination angles as a Deep Q-Network network environment input value; through Deep Q-Network network training, outputting a silicon wafer action and acquiring feedback information; according to the feedback information, the output silicon wafer action of the Deep Q-Network network interacts with the input value, and an optimal silicon wafer action behavior Q value is output; and continuously iterating the Q value according to the state of the moire fringe image to realize automatic leveling of the mask and the silicon wafer. According to the method, the alignment precision and the image quality of the moire fringes are remarkably improved, the alignment time is effectively shortened, and the production efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of optical measurement, and in particular relates to a moiré fringe automatic leveling method and device. Background Art

[0002] In the field of optical measurement and optical technology, photolithography is one of the core technologies of modern semiconductor manufacturing, used to transfer circuit patterns on masks to silicon wafers. With the development of integrated circuit technology, the requirements for the accuracy and range of mask-to-silicon wafer alignment in the photolithography process are becoming increasingly higher. Typically, the alignment accuracy needs to reach 1 / 7-1 / 10 of the minimum feature size, and the alignment accuracy of the next generation of photolithography tools needs to reach the nanometer level. Moiré fringes are interference fringes generated when two gratings are superimposed, which have the effect of amplifying displacement. Moiré fringe technology is a high-precision measurement technology that can be used to measure physical quantities such as displacement, angle, and vibration, and is widely used in the new generation of micro / nanolithography systems. However, existing moiré fringe alignment methods have some limitations, mainly due to the lack of fusion processing capabilities for multiple photolithography processes and the problem of reduced alignment accuracy caused by the tilt of the moiré fringes. There is also the problem of insufficient real-time monitoring of moiré fringe image changes and insufficient rapid response. Summary of the Invention

[0003] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0004] A moiré fringe automatic leveling method, comprising:

[0005] Step 1: Moiré fringes are generated by using a preset angle or relative displacement between the mask and the grating on the silicon wafer, and a CCD camera is used to record the generated moiré fringe image;

[0006] Step 2: Several consecutive moiré fringe images at different tilt angles are used as input values for the Deep Q-Network environment. Through Deep Q-Network training, the silicon chip motion is output and feedback information is obtained.

[0007] Step 3: Based on the feedback information output by the Deep Q-Network, the silicon chip action output by the Deep Q-Network interacts with the Deep Q-Network environment input value to output the optimal silicon chip action behavior Q value;

[0008] Step 4: Based on the state of the moiré fringe image, the Deep Q-Network continuously iterates the Q value to achieve automatic leveling of the mask and the silicon wafer.

[0009] A moiré fringe automatic leveling device, comprising:

[0010] The image acquisition module uses the preset angle or relative displacement between the mask and the grating on the silicon wafer to generate moiré fringes, and uses a CCD camera to record the generated moiré fringe image;

[0011] The training module uses a series of moiré fringe images at different tilt angles as input to the Deep Q-Network environment. Through Deep Q-Network training, it outputs silicon chip motion and obtains feedback information.

[0012] The interaction module, based on the feedback information output by the Deep Q-Network network, interacts with the silicon chip action output by the Deep Q-Network network and the input value of the Deep Q-Network network environment to output the optimal silicon chip action behavior Q value;

[0013] Leveling module: Based on the state of the moiré fringe image, the Deep Q-Network continuously iterates the Q value to achieve automatic leveling of the mask and silicon wafer.

[0014] An electronic device comprises a memory, a processor and a computer program stored in the memory and runnable on the processor. When the processor executes the program, the steps of the moire fringe automatic leveling method are implemented.

[0015] A non-transitory computer-readable storage medium stores a computer program, which implements the steps of the moire fringe automatic leveling method when executed by a processor.

[0016] The present invention has the following beneficial effects:

[0017] By collecting moiré fringe image samples and rapidly generating an optimized adjustment strategy for silicon wafer leveling, the present invention can monitor changes in the moiré fringe image in real time, respond quickly, and adjust the position of the mask and the silicon wafer, thereby greatly shortening the alignment time and improving production efficiency. At the same time, it effectively solves the problem of traditional methods of untimely monitoring of moiré fringe image changes and delayed adjustments, and improves the fusion processing capability of multiple lithography processes.

[0018] Based on the Deep Q-Network network, the present invention introduces dynamic optimization of the parameters of specific image features (moiré fringe images) during the training process, which can adapt to different environmental conditions and interference factors, making the alignment system more stable and reliable, and reducing the problem of decreased alignment accuracy caused by the tilt of the moiré fringes. At the same time, it has strong learning and adaptability, and can self-adjust and optimize according to different types of moiré fringe images and alignment requirements, making it suitable for various complex alignment scenarios and tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1Flowchart of the Deep Q-Network-based moiré fringe automatic leveling method of the present invention;

[0020] Figure 2 This is a curve diagram of the convergence of Deep Q-Network training based on the present invention;

[0021] Figure 3 This is the moiré fringe image before adjustment based on the Deep Q-Network of the present invention;

[0022] Figure 4 This is the moiré fringe image adjusted based on the Deep Q-Network network of the present invention. DETAILED DESCRIPTION

[0023] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only intended to illustrate the present invention and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below may be combined with each other as long as they do not conflict with each other.

[0024] In recent years, Deep Q-Network (DQN), an emerging deep reinforcement learning architecture, has combined the powerful representational capabilities of deep learning with the decision-making capabilities of reinforcement learning. DQN, a pioneering work in deep reinforcement learning, successfully integrates deep learning with Q-learning, overcoming the limitations of traditional reinforcement learning methods in dealing with high-dimensional, continuous, or complex state-space problems. This approach also has implications for optical metrology and optical technology. Using DQN for adaptive alignment reduces human interference and makes the photolithography process more stable and reliable. Automatic alignment systems using DQN for adaptive alignment can operate continuously and stably, adapting to diverse working environments and conditions, and reducing production interruptions and scrap rates caused by human error.

[0025] like Figure 1 As shown, the moire fringe automatic leveling method of the present invention comprises the following steps:

[0026] Step 1: Moiré fringes are generated by using a preset angle or relative displacement between the mask and the grating on the silicon wafer, and a CCD camera (digital camera with a charge-coupled device image sensor) is used to record the generated moiré fringe image.

[0027] In step 2, multiple consecutive (e.g., four) moiré fringe images at different tilt angles are used as input values for the Deep Q-Network environment. Through Deep Q-Network training, the silicon chip motion is output and feedback information is obtained.

[0028] Step 3: Based on the feedback information output by the Deep Q-Network network, the silicon chip action output by the Deep Q-Network network interacts with the input value of the Deep Q-Network network environment to output the optimal silicon chip action behavior Q value (a core concept used to represent the total reward expected from taking a specific action in a given state); that is, the silicon chip action with the highest expected total reward value is output.

[0029] In step 4, based on the state of the moiré fringe image, the Deep Q-Network continuously iterates the Q value to achieve automatic leveling of the mask and the silicon wafer.

[0030] Furthermore, step 2 includes using the moiré fringe images at different tilt angles as input values for the Deep Q-Network environment, and Deep Q-Network training outputs the silicon chip execution action. The specific steps include:

[0031] Step 21: Build a Deep Q-Network architecture, whose input is a series of (e.g., 4) moiré fringe images with different tilt angles, and the output is the Q value of each possible silicon chip action.

[0032] Step 22: Under the current state, select a silicon action based on the epsilon-greedy (ε-greedy) strategy;

[0033] Based on formula (1), in the Deep Q-Network network, each decision moment , randomly select an action to explore with probability ε (decay rate), and select the best known action to exploit with probability 1-ε; represents the Q value of the optimal silicon chip action behavior, s represents the input value of the Deep Q-Network network environment, a represents the action, represents random action;

[0034] (1)

[0035] In step 23, the selected silicon chip action is executed in the Deep Q-Network environment to obtain feedback information generated by the interaction, such as the next state, reward signal, and whether the terminal state is reached.

[0036] Step 24, storing the feedback information generated by each interaction in the experience replay buffer for subsequent training; wherein, the feedback information generated by each interaction includes: the current state of the moiré fringe image (the current feedback value calculated based on the degree of alignment between the mask and the silicon wafer), the current action of the silicon wafer, the reward signal (the reward can be based on the characteristics of the moiré fringe image (the degree of inclination of the fringe) to measure the contribution of the current action to the alignment effect), the next state (the new state of the moiré fringe image after the current silicon wafer action is executed), and whether the termination state is reached (determining whether the target moiré fringe state preset by the network has been reached after the current silicon wafer action is executed, that is, whether the alignment of the moiré fringe has been completed).

[0037] Furthermore, in step 3, the chip action output by the Deep Q-Network interacts with the Deep Q-Network environment input value, and the chip action output by the Deep Q-Network replaces the Deep Q-Network environment input value of step 2. Step 3 specifically includes:

[0038] Step 31: randomly extract a small batch of experience samples from the experience replay buffer.

[0039] Step 32, calculating the target Q value based on the output of the Deep Q-Network target network and the current network;

[0040] Based on formula (2), in the Deep Q-Network, the calculation method of the target Q value is the key part for updating the neural network weights. The target moiré fringe state pre-set by the network is already in a leveled state and is used as the target Q value to train the Q network so that the Q network can predict the state in a given state. Take different actions The expected cumulative reward is the Q value. By continuously making the network's predicted Q value close to the target Q value, the network can learn more accurate Q value estimates, thus providing a basis for the intelligent agent to choose the optimal silicon action. represents the target Q value, which is used to measure the optimal reward expected under a specific moiré fringe state (leveling state) and silicon chip action; Is in state Take action The immediate reward after the action indicates the direct benefit of the current action; is a discount factor, ranging from 0 to 1, which is used to balance the weights of immediate rewards and future rewards. Its function is to enable the agent to consider both short-term and long-term interests when making decisions; For the next state Take every possible action A set of Q values, which are used to calculate the possible rewards in the next state; Indicates the next state Take the best action The maximum Q value that can be obtained is the maximum Q value obtained by the target network (the weight parameter of the target network is , Periodically read the weight parameters of the current network Synchronous update) calculated.

[0041] (2)

[0042] in, is the time step, for the next time step.

[0043] Step 33: Compare the Q value predicted by the current network and target Q value The difference between , the loss function is calculated, and the network weights are updated using backpropagation and optimization algorithms.

[0044] Based on formula (3), in Deep Q-Network, mean square error is used as the loss function , which represents the target Q value and the Q value predicted by the current network The difference between them. N is the number of samples and is the size of the mini-batch data; Indicates the action predicted by the network in the current state. The Q value predicted by the current network is calculated by the current network; is the weight parameter of the current network.

[0045] (3)

[0046] in, is the time step, is the previous time step.

[0047] Based on formula (4), back propagation and optimization algorithms are used to calculate the gradient of the loss function to the network weights and update the network weights. is the updated network weight, is the learning rate, which controls the step size of each update, is the gradient of the loss function with respect to the network weights, is the gradient symbol. The target network (the weight parameter of the network is ) Usually every certain number of time steps from the current network (the weight parameters of the network are ) instead of updating in real time, which can stabilize the training process.

[0048] (4)

[0049] Step 34: Update the weight parameters of the current network at fixed time intervals. Copy to the target network and output the optimal silicon action behavior Q value.

[0050] The step 4 comprises:

[0051] Step 41, according to the state of the moiré fringe image during training, adjust the hyperparameters to optimize the learning effect; the hyperparameters include: learning rate, decay rate ε, discount factor , batch size (the number of samples extracted from the experience replay buffer during each training), target network update frequency (determines how often the weights of the current network are copied to the target network), network structure related parameters (such as the number of layers of the neural network, the number of neurons in each layer, etc.) and optimizer related parameters (used to accelerate the gradient descent process), etc.

[0052] Step 42: Regularly evaluate the performance of the network in the test Deep Q-Network environment and monitor the learning progress (judging the learning progress through indicators such as the change in reward value, loss function value, and alignment accuracy during training). If the training effect is poor or divergence occurs, readjust the network structure or training parameters.

[0053] Step 43 , repeating the training process until the network converges, that is, being able to stably output the optimal silicon wafer motion behavior Q value in most cases, thereby achieving the automatic leveling task between the mask and the silicon wafer.

[0054] The present invention further provides a moiré fringe automatic leveling device, comprising:

[0055] The image acquisition module uses the preset angle or relative displacement between the mask and the grating on the silicon wafer to generate moiré fringes, and uses a CCD camera to record the generated moiré fringe image;

[0056] The training module uses a series of moiré fringe images at different tilt angles as input to the Deep Q-Network environment. Through Deep Q-Network training, it outputs silicon chip motion and obtains feedback information.

[0057] The interaction module, based on the feedback information output by the Deep Q-Network network, interacts with the silicon chip action output by the Deep Q-Network network and the input value of the Deep Q-Network network environment to output the optimal silicon chip action behavior Q value;

[0058] In the leveling module, the Deep Q-Network continuously iterates the Q value based on the state of the moiré fringe image to achieve automatic leveling of the mask and the silicon wafer.

[0059] The present invention further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the moire fringe automatic leveling method when executing the program.

[0060] The present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which implements the steps of the moiré fringe automatic leveling method when executed by a processor.

[0061] Figure 2 This is a curve diagram based on the convergence of Deep Q-Network network training. The current network Q value is continuously iterated through Deep Q-Network network training and finally reaches the convergence curve.

[0062] Figure 3 and Figure 4 The figures are respectively the moiré fringe images before and after adjustment based on the Deep Q-Network network. By comparing the moiré fringe images before and after adjustment, it can be intuitively seen that the moiré fringe automatic leveling method based on Deep Q-Network of the present invention significantly improves the alignment accuracy and image quality of the moiré fringes, effectively shortens the alignment time, and improves production efficiency.

[0063] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention may be implemented using various computer languages.

[0064] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0065] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0066] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0067] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0068] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

[0069] The above descriptions are merely embodiments of the present invention and are not intended to limit the scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied to other related system fields, are also included in the scope of protection of the present invention.

[0070] The contents not described in detail in the specification of the present invention belong to the prior art known to those skilled in the art.

Claims

1. A moiré fringe automatic leveling method, characterized in that: include: Step 1: Moiré fringes are generated by using a preset angle or relative displacement between the mask and the grating on the silicon wafer, and a CCD camera is used to record the generated moiré fringe image; Step 2: Several consecutive moiré fringe images at different tilt angles are used as input values for the Deep Q-Network environment. Through Deep Q-Network training, the silicon chip motion is output and feedback information is obtained. Step 3: Based on the feedback information output by the Deep Q-Network, the silicon chip action output by the Deep Q-Network interacts with the Deep Q-Network environment input value to output the optimal silicon chip action behavior Q value; In step 4, based on the state of the moiré fringe image, the Deep Q-Network continuously iterates the Q value to achieve automatic leveling of the mask and the silicon wafer.

2. The moiré fringe automatic leveling method according to claim 1, characterized in that: Step 2 includes using the moiré fringe images at different tilt angles as input values for the Deep Q-Network environment, and Deep Q-Network training outputs the silicon chip execution action, specifically: Step 21: Build a Deep Q-Network architecture, whose input is a series of moiré fringe images with different tilt angles, and output is the Q value of each possible silicon chip action; Step 22: Under the current state, select a silicon chip action based on the ε-greedy strategy; Based on formula (1), in the Deep Q-Network network, each decision moment , randomly select an action to explore with probability ε, and select the best known action to exploit with probability 1-ε; is the Q value of the optimal silicon chip action behavior, s is the input value of the Deep Q-Network network environment, a represents the action, represents random action; (1) Step 23: Execute the selected silicon chip action in the Deep Q-Network environment to obtain feedback information generated by the next interaction; Step 24: Store the feedback information generated by each interaction into the experience replay buffer for subsequent training.

3. The moiré fringe automatic leveling method according to claim 2, characterized in that: In step 24, the feedback information generated by each interaction includes: the current state of the moiré fringe image, the current action of the silicon chip, the reward signal, the next state, and whether the terminal state is reached.

4. The moiré fringe automatic leveling method according to claim 2 or 3, characterized in that: In step 3, the silicon chip action output by the DeepQ-Network network interacts with the DeepQ-Network network environment input value, and the DeepQ-Network network environment input value of step 2 is replaced by the silicon chip action output by the DeepQ-Network network, specifically including: Step 31, randomly extract a small batch of experience samples from the experience replay buffer; Step 32, based on formula (2), calculate the target Q value according to the output of the Deep Q-Network target network and the current network; is the target Q value, which is used to measure the optimal reward expected under a specific moiré fringe state and silicon chip action; In state The immediate reward after taking action a represents the direct benefit brought by the current action; The discount factor ranges from 0 to 1 and is used to balance the weight of immediate rewards and future rewards; For the next state Next, take the set of Q values of each possible action a to calculate the possible reward in the next state; For the next state The maximum Q value that can be obtained by taking the optimal action a is calculated by the target network, and the weight parameter of the target network is , Periodically read the weight parameters of the current network Synchronous updates; (2) in, is the time step, For the next time step; Step 33: Compare the Q value predicted by the current network and target Q value The difference between them is used to calculate the loss function and update the network weights using backpropagation and optimization algorithms; Step 34: Update the weight parameters of the current network at fixed time intervals. Copy to the target network and output the optimal silicon action behavior Q value.

5. The moiré fringe automatic leveling method according to claim 4, characterized in that: Step 33 includes: loss function Indicates the target Q value and the Q value predicted by the current network The differences between: (3) in, is the number of samples; Indicates the action predicted by the network in the current state. The Q value predicted by the current network is calculated by the current network; is the weight parameter of the current network; is the previous time step; Based on formula (4), the loss function is calculated using back propagation and optimization algorithms. Gradients of network weights and update network weights: (4) in, is the updated network weight, is the learning rate, which controls the step size of each update, is the gradient of the loss function with respect to the network weights, is the gradient symbol.

6. The moiré fringe automatic leveling method according to claim 1, characterized in that: The step 4 comprises: Step 41, adjusting hyperparameters to optimize learning effects according to the state of the moiré fringe image during training; Step 42: Regularly evaluate the network performance in the test Deep Q-Network environment to monitor the learning progress; if the training effect is poor or divergence occurs, readjust the network structure or training parameters; Step 43, repeat the training process until the network converges, and the automatic leveling task of the mask and the silicon wafer is achieved.

7. The moiré fringe automatic leveling method according to claim 6, characterized in that: In step 41, the hyperparameters include: learning rate, decay rate ε, discount factor , batch size, target network update frequency, network structure related parameters and optimizer related parameters.

8. A moiré fringe automatic leveling device, characterized in that: include: The image acquisition module uses the preset angle or relative displacement between the mask and the grating on the silicon wafer to generate moiré fringes, and uses a CCD camera to record the generated moiré fringe image; The training module uses a series of moiré fringe images at different tilt angles as input to the Deep Q-Network environment. Through Deep Q-Network training, it outputs silicon chip motion and obtains feedback information. The interaction module, based on the feedback information output by the Deep Q-Network network, interacts with the silicon chip action output by the Deep Q-Network network and the input value of the Deep Q-Network network environment to output the optimal silicon chip action behavior Q value; In the leveling module, the Deep Q-Network continuously iterates the Q value based on the state of the moiré fringe image to achieve automatic leveling of the mask and the silicon wafer.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the moire fringe automatic leveling method according to any one of claims 1 to 7 are implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the moire fringe automatic leveling method according to any one of claims 1 to 7 are implemented.