Photoetching machine mask alignment method based on Transform model
Through the Transformer model-based lithography mask alignment method, global feature regression is performed by combining image and physical property data, which achieves precise alignment of the lithography machine in complex environments, solves the problem of insufficient alignment accuracy of traditional lithography machines, and improves production efficiency and quality.
Patent Information
- Application Number
- CN202511106929.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-08
AI Technical Summary
Traditional lithography machines, when faced with complex environmental factors and the physical properties of the mask, find it difficult to achieve real-time and accurate alignment of the mask and wafer, resulting in errors and quality fluctuations in the production process.
A lithography mask alignment method based on the Transformer model is adopted. The image data is collected for preprocessing and data enhancement, and the dynamic coefficient is calculated based on the physical property data. The Transformer model is used for global feature regression, and the distance and angle deviation are output. The stepper motor is used for real-time compensation.
It significantly improves the alignment accuracy and robustness of lithography machines under complex working conditions, improves production efficiency and product quality, and meets the real-time and low-power industrial deployment requirements.
Smart Images

Figure CN120595546A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of photolithography machines, and more specifically, to a photolithography mask alignment method based on a Transformer model. Background Art
[0002] Photolithography is widely used in semiconductor manufacturing, particularly in integrated circuit (IC) fabrication. Reticle alignment accuracy directly impacts the quality of the lithographic pattern and chip performance. In traditional lithography processes, the alignment accuracy of the reticle and wafer typically relies on mechanical control and image matching techniques. However, as IC sizes continue to shrink, the demand for alignment accuracy is becoming increasingly stringent, and traditional technologies struggle to meet the dual demands of precision and efficiency. Furthermore, environmental factors (such as temperature and humidity fluctuations, variations in the air's refractive index), as well as the physical properties of the reticle (such as stress distribution and uneven photoresist thickness), can also affect alignment accuracy. Existing lithography machine alignment methods struggle to achieve real-time, precise adjustments in the face of these complex factors, leading to errors and quality fluctuations during production. Therefore, a new alignment method is urgently needed that can combine image data with environmental and physical property information to accurately compensate for alignment in real time, thereby improving lithography quality and production efficiency.
[0003] In view of the above problems, the present invention proposes a solution. Summary of the Invention
[0004] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a method for aligning a mask of a lithography machine based on a Transformer model to solve the problems raised in the above-mentioned background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions: In a preferred embodiment, it comprises: Step 1: Collect image data of the mask and wafer combination and perform denoising, normalization, and data enhancement to generate normalized distance and angle labels; Step 2: Divide the image into multiple sub-regions, calculate the dynamic coefficients and extract features based on the physical property data, input them into the Transformer model for global feature regression, and output the distance deviation and angle deviation; Step 3: Adjust the mask position and angle based on the predicted distance deviation and angle deviation, compensate through stepper motors, and implement low-latency real-time inference on the FPGA.
[0006] In a preferred embodiment, in step 1, image data of different mask and wafer combinations are collected, image data noise is removed, and the image data is normalized.
[0007] In a preferred embodiment, in step 1, anti-interference data enhancement processing is further performed on the denoised and normalized image data, and the image and the annotated key point coordinates are synchronously translated, rotated and scaled using an affine transformation matrix. The enhanced image is ensured to be consistent with the annotated coordinates through a key point synchronization mapping interface, and local nonlinear perturbations are applied to the image in combination with the elastic deformation enhancement strategy.
[0008] In a preferred embodiment, in step 2, the mask and wafer area are evenly divided into a plurality of sub-areas; Analyze the internal stress distribution of the mask glass in each sub-region to generate a stress distribution matrix; analyze the thickness and curing uniformity of the photoresist layer in each sub-region to generate a thickness matrix and a curing uniformity matrix, which are recorded as physical property data; collect temperature and humidity data in each sub-region, monitor the fluctuation of the air refractive index, calculate the refractive index change matrix, and record it as the air refractive index fluctuation data; Calculate the dynamic coefficient D(x,y) of each sub-region and compare it with the dynamic coefficient threshold Yd. When the dynamic coefficient D(x,y) of the sub-region is greater than or equal to the dynamic coefficient threshold Yd, it indicates that the sub-region belongs to the high dynamic region; when the dynamic coefficient D(x,y) of the sub-region is less than the dynamic coefficient threshold Yd, it indicates that the sub-region belongs to the low dynamic region.
[0009] In a preferred embodiment, in step 2, feature data of the image data in each sub-region is extracted and converted into a feature vector Xi of a fixed size, and the feature vector Xi of each sub-region and the dynamic coefficient D(x, y) are combined to form sequence data Fi; Perform a linear transformation on the sequence data Fi of each sub-region; calculate the correlation matrix between the sequence data of each sub-region; perform weighted fusion of the sequence data of each sub-region through the correlation matrix to determine the global feature information of each sub-region, and map the different relationship information in the global feature information between each sub-region extracted by multiple attention heads to determine the feature mapping value of each sub-region, and fuse the feature mapping values of each sub-region to determine the alignment deviation prediction value Y of the entire mask and wafer area.
[0010] In a preferred embodiment, in step 2, after linear transformation and weighted fusion of the sequence data, the sequence data is input into a regression model based on the VisionTransformer architecture, the input image is divided into image blocks and linear mapping is performed to obtain feature vectors, and a learnable position encoding is added to preserve spatial information; The regression model based on the VisionTransformer architecture includes twelve layers of encoders, each of which contains multi-head self-attention and feedforward neural networks.
[0011] In a preferred embodiment, in step 2, the input data is divided into image blocks and global features are extracted, and the distance deviation and angle deviation are output. The angle is obtained by calculating the inverse tangent after sine-cosine encoding, and a hierarchical learning rate is used for different layers in combination with the weighted mean square error during the training process.
[0012] In a preferred embodiment, in step 3, the mask image offset is calculated, and a compensation deviation scheme is designed for the high dynamic area and the low dynamic area.
[0013] In a preferred embodiment, in step 3, when performing alignment compensation, the model is quantized and pruned, and the calculation process is parallel pipelined.
[0014] In a preferred embodiment, the method for obtaining the dynamic coefficient threshold is: The mean and standard deviation of the physical property data and air refractive index fluctuation data in each sub-region are calculated, and K-means clustering is used to analyze the mean and standard deviation of the physical property data and air refractive index fluctuation data in each sub-region to determine the dynamic coefficient threshold Yd.
[0015] The technical effects and advantages of the present invention's reticle alignment method for a photolithography machine based on a Transformer model are as follows: The present invention introduces global context modeling and data enhancement mechanisms into the traditional lithography machine alignment process, significantly improving the alignment capability under complex working conditions. By adding affine synchronous enhancement and elastic deformation in the image preprocessing stage, the data diversity and annotation consistency are improved, enabling the model to adapt to the interference caused by noise, illumination and deformation. In the feature extraction and modeling stage, the VisionTransformer structure is used to regress the global features, and combined with the hierarchical learning rate, freezing strategy and weighted loss function, the convergence stability of the model training and the high accuracy of the prediction are guaranteed. In the execution compensation stage, by designing stepper motor adjustment schemes for high dynamic and low dynamic areas respectively, accurate compensation of translation and rotation deviations is achieved. Finally, through FPGA quantization, pruning and pipeline scheduling, the model still has real-time reasoning capabilities in a low-power hardware environment. Overall, the present invention has high precision, strong robustness and industrial deployment feasibility, and can significantly improve the production efficiency and product quality of lithography machines. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is an operational flow chart of a photolithography mask alignment method based on a Transformer model of the present invention.
[0017] Figure 2 The present invention is a flowchart of a method for aligning a mask of a lithography machine based on a Transformer model. DETAILED DESCRIPTION
[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0019] Example The present invention discloses a method for aligning a mask of a photolithography machine based on a Transformer model, such as Figure 1 Shown, including: Step 1: Collect image data of different mask and wafer combinations and pre-process the image data; A high-resolution CMOS camera is used to control the movement and positioning of the mask and wafer combination through an automated mechanical device. The high-resolution CMOS camera is kept at a fixed shooting angle parallel to the mask and wafer. The full-field scanning method is used to capture images of different mask and wafer combinations to obtain image data of different mask and wafer combinations. Gaussian filtering algorithm is used to remove image data noise and normalize the image data. The specific formula is: ; Where, is the normalized value; After image acquisition and basic preprocessing, the acquired image data is further enhanced to combat interference. Specifically, affine transformation matrices are used to synchronously translate, rotate, and scale the image and annotation coordinates. The KeypointParams interface of the Albumentations library is then used to achieve real-time mapping between keypoint coordinates and the image. This ensures that each image enhancement strictly corresponds to the accurate annotation coordinates, thereby enhancing the consistency of the image and annotation data.
[0020] At the same time, the elastic deformation enhancement strategy is used to simulate the local thermal deformation that may occur in the mask during actual processing and operation, and the random elastic distortion algorithm is used to introduce tiny nonlinear perturbations in the data enhancement process to improve the model's adaptability to local deformation of the mask under different working conditions.
[0021] In the label generation stage, the normalized distance and angle are calculated based on the synchronized key point coordinates. Specifically, the distance label is obtained by calculating the Euclidean distance between the centers of the two markers and dividing it by the length of the image diagonal. The angle label is obtained by calculating the angle between the line connecting the two points and the horizontal axis and mapping it to a range of zero to 180 degrees. This achieves a unified scale expression of distance and angle, enhancing the robustness of subsequent model training.
[0022] The specific mathematical implementation is as follows: ; Where M is the transformation matrix, (x, y) is the original coordinate, and (x′, y′) is the transformed coordinate.
[0023] Label design: Distance: Calculate the Euclidean distance between the centers of two markers and normalize it: ; Where Ldiagonal is the length of the image diagonal.
[0024] Angle: Calculates the angle between the line connecting two points and the horizontal axis, and maps it to the range of 0-180°: ; Step 2: Establish the alignment relationship between the mask and the wafer through Transformer model training; A grid-based region partitioning method is used to evenly divide the mask and wafer area in the entire image data into multiple sub-regions; The internal stress distribution of the mask glass in each sub-region is analyzed using photoelastic measurement technology to generate a stress distribution matrix S(x,y). The thickness and curing uniformity of the photoresist layer in each sub-region are analyzed using ellipsometer to generate a thickness matrix T(x,y) and a curing uniformity matrix C(x,y), which are recorded as physical property data. A temperature and humidity sensor array is used to monitor the temperature and humidity fluctuations in each sub-area. The temperature and humidity data in each sub-area are obtained to generate the temperature matrix H1(x, y) and the humidity matrix H2(x, y). Interferometry is then used to monitor the air refractive index fluctuations based on the temperature and humidity data. The refractive index change matrix R(x, y) is calculated and recorded as the air refractive index fluctuation data. Based on the physical property data and air refractive index fluctuation data in each sub-region, a multi-factor fusion algorithm is used to calculate the dynamic coefficient D(x,y) of each sub-region. The specific formula is: D(x,y)=W1×S(x,y)+W2×T(x,y)+W3×C(x,y)+W4×R(x,y), where W1 represents the weight coefficient of stress distribution, W2 represents the weight coefficient of photoresist thickness, W3 represents the weight coefficient of curing uniformity, and W4 represents the weight coefficient of air refractive index. The mean and standard deviation of the refractive index fluctuation data are analyzed using K-means clustering on the physical property data and the air refractive index fluctuation data in each sub-region to determine the dynamic coefficient threshold Yd. The dynamic coefficient D(x,y) of each sub-region is compared with the dynamic coefficient threshold Yd. When the dynamic coefficient D(x,y) of the sub-region is greater than or equal to the dynamic coefficient threshold Yd, it indicates that the sub-region belongs to the high dynamic region; when the dynamic coefficient D(x,y) of the sub-region is less than the dynamic coefficient threshold Yd, it indicates that the sub-region belongs to the low dynamic region. CNN is used to extract the feature data of the image data in each sub-region, and the feature map of each sub-region is output to the Transforme model; The CNN structure is designed as follows: The first convolution layer uses a 3×3 convolution kernel and a ReLU activation function to extract edge, texture, curve, and boundary features of the image data in each sub-region. The second pooling layer uses a 2×2 convolution kernel and maximum pooling to enhance translation invariance. The third convolutional layer uses a 5×5 convolution kernel and a ReLU activation function to extract the pattern and shape features of the image data in each sub-region; The fifth global connection layer: uses global average pooling to convert the feature data of the image data in each sub-region into a fixed-size feature vector Xi; It should be noted that the Transformer model requires the input data format to be serialized data. Therefore, the feature vector Xi and dynamic coefficient D(x, y) of each sub-region are formed into a sequence: Fi=[Xi, Di], where Di represents the dynamic coefficient D(x, y) of each sub-region, and the sequence data Fi of each sub-region is transmitted to the Transformer model; Next, perform a linear transformation on the sequence data Fi of each sub-region and calculate the Q, K, and V matrices: Q = FiWq, K = FiWk, and V = FiWv, where Wq, Wk, and Wv represent trainable parameters, Q represents the query matrix, K represents the key matrix, and V represents the value matrix. Calculate the correlation matrix between the sequence data of each sub-region : ; in, represents the dimension of K, represents the dot product of Q and K; Through the correlation matrix Perform weighted fusion on the sequence data of each sub-region: ; in, Represents the global feature information of each sub-region after fusion; Use multiple attention heads to extract different relationship information from the global feature information of each sub-region: ; Where h represents the number of attention heads and Wo represents the final projection matrix; The two-layer fully connected network MLP in the feedforward neural network FFN is used to map the different relationship information in the global feature information between the sub-regions output by the Transformer model and convert it into numerical predictions. The specific formula is: ; in, represents the feature map value of the i-th subregion, Wf1 and Wf2 represent weight matrices, and bf1 and bf2 represent bias terms; The feature map values of each sub-region are integrated to output the alignment deviation prediction value Y of the entire mask and wafer area. Specific formula: ; in, represents the learnable weight, and N represents the total number of sub-regions; Finally, the mean square error (MSE) is used as the loss function to optimize the model parameters; Among them, the loss function design is as follows: Periodic angle loss: Sine-cosine encoding is used to avoid periodic angle jumps: ; Weighted total loss: balancing distance and angle errors: ; Furthermore, based on the above feature extraction and serialization processing, the VisionTransformer architecture is further used to perform regression modeling on the serialized data. Specifically, the input image is divided into 16-by-16 pixel blocks, and a linear mapping is performed on each block to obtain a 768-dimensional vector feature. At the same time, a learnable positional encoding is added to preserve the spatial position information of each image block.
[0025] Subsequently, a ViT structure consisting of twelve layers of encoders was adopted. Each layer of encoders consists of multi-head self-attention and feedforward neural networks, and global context modeling is used to characterize the long-distance dependency relationship between the mask and wafer images.
[0026] Building on the aforementioned VisionTransformer architecture, dynamic sparsification is performed to further reduce computational complexity and improve feature expression. Specifically, during the multi-head self-attention calculation process, the contribution of each attention head is calculated in real time based on the average attention score of each head during training. A threshold screening mechanism is then used to retain only those heads with contributions above the set threshold for subsequent matrix product calculations. Redundant attention heads are dynamically removed, reducing the overall computational effort by approximately 43% while maintaining prediction accuracy.
[0027] The specific ViT regression architecture improvements are as follows: Model structure: Input layer: The image is segmented into 16×16 pixel blocks and linearly mapped into 768-dimensional vectors.
[0028] Position Encoding: Add learnable position embedding to preserve spatial information.
[0029] Transformer Encoder: It uses a 12-layer encoder, each layer of which contains a multi-head self-attention (Multi-HeadAttention) and MLP module.
[0030] Regression head: Input the CLS token features of ViT into the fully connected layer and output a two-dimensional vector (d^,θ^).
[0031] Fine-tuning strategy: The encoder parameters of the first 6 layers are frozen, and only the last 6 layers and the regression head are trained, with the learning rate set to 1×10−4.
[0032] The AdamW optimizer is used with a weight decay coefficient of λ = 0.01.
[0033] At the same time, in order to enhance the feature interaction between different coding layers, the output features of the previous layer and the input features of the current layer are spliced according to the channel dimension, integrated into a unified dimension through linear transformation, and then sent to the subsequent multi-head self-attention structure to achieve collaborative regression of local and global information.
[0034] In terms of training strategy, we freeze the first six encoder layers of the pre-trained model and only fine-tune the last six encoder layers and the regression head. We use the AdamW optimizer with the weight decay coefficient set to 0.01 and combine it with layer-wise learning rate scheduling, as follows: The learning rate of the pre-training layer is set to one times ten to the power of -4, and the learning rate of the regression head is set to one times ten to the power of -3. During the overall training process, the cosine annealing scheduler is used to decay the learning rate to one times ten to the power of -5 within the initial training cycles, thereby accelerating the convergence of new tasks while maintaining the stability of the pre-training features.
[0035] Furthermore, in the regression head design, the CLS token features of ViT are fed into the dual-output fully connected layer, with the output corresponding to the distance deviation and angle deviation, respectively. The angle deviation is encoded using a dual sine and cosine channel, and the output is ultimately converted back to an angle value using the arctan2 function, effectively avoiding the periodic jump problem that often occurs in angle prediction.
[0036] To prevent overflow in the output space or violations of the sine-cosine function constraints, a dual-channel tanh activation function is used to normalize and constrain the network's output predictions. Specifically, the network outputs ŷ_cos and ŷ_sin are each activated using tanh, ensuring their values are always constrained to the interval [−1, 1]. Constraint checks are performed on the output values during training and inference to ensure that ŷ_cos² + ŷ_sin² ≤ 1, thus ensuring the stability and physical plausibility of the angle restoration calculation.
[0037] It should be noted that in the design of the loss function, it is necessary to use the mean square error weighted by the distance error and the angle error to balance the optimization process of distance prediction and angle prediction, and ultimately achieve high-precision regression of global features.
[0038] Step 3: Calculate the mask image offset and adjust the position and angle of the mask in real time; Based on the alignment deviation prediction value Y output by the Transformer model in step 2 and the image data collected in step 1, calculate the mask image offset: ;in, and Indicates the mask image offset, Represents the image data collected in step 1; For high-dynamic areas, an adaptive stepper motor control strategy is used to dynamically adjust the stepper motor's step size according to the mask image offset to compensate for translation errors. A proportional-integral-differential controller is used to close the loop to control the stepper motor's rotation angle and compensate for rotation errors. For low dynamic areas, a fixed step control strategy is used to adjust the step size of the stepper motor to compensate for translation errors; a simple proportional control is used to compensate for rotation errors. To meet the real-time and low-power requirements of semiconductor production lines, the Vision Transformer regression model was deployed on an FPGA platform for hardware acceleration. Specifically, the model was layered and quantized using the VitisAI toolchain, with the attention layer using INT8 fixed-point representation, the fully connected layer using BF16 precision, and the output layer using INT16. This significantly reduced the model's computational complexity and power consumption.
[0039] At the same time, during the model compression process, the attention heads whose average attention scores in ViT are lower than the preset threshold are removed through the pruning strategy, which reduces the number of model parameters and effectively alleviates the storage and computing pressure of FPGA.
[0040] In terms of operation scheduling, multi-head attention calculation, patch division and matrix multiplication operations are divided into parallel sub-modules, and pipeline scheduling is used to improve the clock cycle utilization of each computing module, thereby achieving real-time inference latency reduction of single-frame images on the Xilinx DPU platform, and ultimately meeting the real-time control requirements of lithography mask alignment.
[0041] like Figure 2 As shown in the figure, the implementation process of the mask alignment method described in the present invention forms a complete execution closed loop according to the above steps. First, the combined image of the mask and wafer is acquired through image acquisition and denoising, normalization and data enhancement operations are performed to form a standardized input. Then, the image is divided into several sub-regions and physical property data such as stress, thickness, curing, temperature and humidity are extracted. The dynamic coefficient D(x,y) is fused to construct the dynamic coefficient, and the high-dynamic and low-dynamic regions are divided. Subsequently, the sequence Fi is constructed by combining the image features and the dynamic coefficient and input into the model based on the VisionTransformer structure to complete the global feature modeling and alignment deviation prediction. Finally, the differentiated compensation strategy is executed by region according to the deviation results, and a low-latency hardware inference and control closed loop is achieved through model pruning, quantization and parallel scheduling on the FPGA platform. Figure 2 The flowchart comprehensively reflects the full-chain alignment mechanism of the present invention from perception, modeling, prediction to control, ensuring the accuracy, stability and real-time performance of the method in an industrial environment.
[0042] The FPGA deployment optimization Dynamic quantization: The model is quantized to INT8 using the Vitis AI toolchain, and the calibration set contains 1,000 representative images.
[0043] Model pruning: Remove attention heads whose attention scores in ViT are lower than the threshold (α < 0.1), compressing the model size by 40%.
[0044] Hardware acceleration: Xilinx DPU is configured with parallel computing units to optimize matrix multiplication and softmax operations, with inference latency less than 10ms.
[0045] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.
[0046] The above embodiments may be implemented in whole or in part through software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product.
[0047] Those skilled in the art will appreciate that the modules and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application of the technical solution and the invention constraints. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0048] In addition, each functional module in each embodiment of the present application may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0049] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0050] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A photolithography mask alignment method based on the Transformer model, It is characterized by including: Step 1: Collect image data of the mask and wafer combination and perform denoising, normalization, and data enhancement to generate normalized distance and angle labels; Step 2: Divide the image into multiple sub-regions, calculate the dynamic coefficients and extract features based on the physical property data, input them into the Transformer model for global feature regression, and output the distance deviation and angle deviation; Step 3: Adjust the mask position and angle based on the predicted distance deviation and angle deviation, compensate through stepper motors, and implement low-latency real-time inference on the FPGA.
2. The method for aligning a mask of a lithography machine based on a Transformer model according to claim 1, wherein: In step 1, image data of different mask and wafer combinations are collected, image data noise is removed, and the image data is normalized.
3. The method for aligning a mask of a lithography machine based on a Transformer model according to claim 2, wherein: In step 1, the denoised and normalized image data is further enhanced with anti-interference data. The image and the annotated key point coordinates are synchronously translated, rotated, and scaled using an affine transformation matrix. The enhanced image is consistent with the annotated coordinates through a key point synchronization mapping interface. The elastic deformation enhancement strategy is combined to apply local nonlinear perturbations to the image.
4. The method for aligning a mask of a lithography machine based on a Transformer model according to claim 3, wherein: In step 2, the mask and wafer area are evenly divided into multiple sub-areas; Analyze the internal stress distribution of the mask glass in each sub-region to generate a stress distribution matrix; analyze the thickness and curing uniformity of the photoresist layer in each sub-region to generate a thickness matrix and a curing uniformity matrix, which are recorded as physical property data; collect temperature and humidity data in each sub-region, monitor the fluctuation of the air refractive index, calculate the refractive index change matrix, and record it as the air refractive index fluctuation data; Calculate the dynamic coefficient D(x,y) of each sub-region and compare it with the dynamic coefficient threshold Yd. When the dynamic coefficient D(x,y) of the sub-region is greater than or equal to the dynamic coefficient threshold Yd, it indicates that the sub-region belongs to the high dynamic region; when the dynamic coefficient D(x,y) of the sub-region is less than the dynamic coefficient threshold Yd, it indicates that the sub-region belongs to the low dynamic region.
5. The method for aligning a mask of a lithography machine based on a Transformer model according to claim 4, characterized in that: In step 2, the feature data of the image data in each sub-region is extracted and converted into a feature vector Xi of a fixed size, and the feature vector Xi of each sub-region and the dynamic coefficient D (x, y) are combined to form sequence data Fi; Perform a linear transformation on the sequence data Fi of each sub-region; calculate the correlation matrix between the sequence data of each sub-region; perform weighted fusion of the sequence data of each sub-region through the correlation matrix to determine the global feature information of each sub-region, and map the different relationship information in the global feature information between each sub-region extracted by multiple attention heads to determine the feature mapping value of each sub-region, and fuse the feature mapping values of each sub-region to determine the alignment deviation prediction value Y of the entire mask and wafer area.
6. The method for aligning a mask of a lithography machine based on a Transformer model according to claim 5, wherein: In step 2, after linear transformation and weighted fusion of the sequence data, the sequence data is input into a regression model based on the VisionTransformer architecture, which divides the input image into image blocks and performs linear mapping to obtain feature vectors. At the same time, a learnable position encoding is added to preserve spatial information. The regression model based on the VisionTransformer architecture includes twelve layers of encoders, each of which contains multi-head self-attention and feedforward neural networks.
7. The method for aligning a mask of a lithography machine based on a Transformer model according to claim 6, wherein: In step 2, the input data is divided into image blocks and global features are extracted. The distance deviation and angle deviation are output. The angle is obtained by calculating the inverse tangent after sine-cosine encoding. During the training process, a hierarchical learning rate is used for different layers and combined with the weighted mean square error.
8. The method for aligning a mask of a lithography machine based on a Transformer model according to claim 7, wherein: In step 3, the mask image offset is calculated and compensation deviation schemes are designed for high dynamic area and low dynamic area.
9. The method for aligning a mask of a lithography machine based on a Transformer model according to claim 8, wherein: In step 3, when implementing alignment compensation, the model is quantized and pruned, and the computation process is parallel pipelined.
10. The method for aligning a mask of a lithography machine based on a Transformer model according to claim 9, characterized in that: How to obtain the dynamic coefficient threshold: The mean and standard deviation of the physical property data and air refractive index fluctuation data in each sub-region are calculated, and K-means clustering is used to analyze the mean and standard deviation of the physical property data and air refractive index fluctuation data in each sub-region to determine the dynamic coefficient threshold Yd.
Citation Information
Patent Citations
Photoetching process window detection method based on contrast learning
CN116482943A
Asymmetric compensation method and device for overlay mark of multi-modal deep learning network
CN120010190A
Semiconductor photoetching system based on intelligent industrial robot
CN120255292A
Virtual gauging system for use in lithographic processing
US6633050B1