Wafer sub-pixel alignment method based on image super-resolution

By adopting a wafer subpixel alignment method based on image super resolution in the scriber, the cutting defects caused by the alignment error of the scriber are solved, and the alignment accuracy and product quality are improved.

CN120047528APending Publication Date: 2025-05-27ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510093667.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The cutting defects caused by alignment errors during the cutting process of existing scribers affect the overall quality and yield of the product.

Method used

The wafer subpixel alignment method based on image super resolution is adopted, and enhanced spatial attention is introduced through the cascade residual channel attention group, and the wafer features are fully retained in combination with the hierarchical feature fusion structure, and super-resolution reconstruction is carried out. Finally, the phase correlation and Gaussian functions are used for subpixel alignment.

Benefits of technology

The alignment accuracy of the scriber is improved, cutting defects are reduced, the overall quality and yield of the product are improved, and the specific accuracy is increased by about 0.05pixel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047528A_ABST
    Figure CN120047528A_ABST
Patent Text Reader

Abstract

The invention discloses a wafer sub-pixel alignment method based on image super-resolution, and relates to the technical field of scribing machine image alignment. The invention provides a wafer sub-pixel alignment method based on image super-resolution. The wafer sub-pixel alignment method comprises the following steps: S100, operating a scribing machine to collect wafer image pairs; s200, on the basis of the cascade residual channel attention group, introducing high-frequency information implied in an enhanced spatial attention capture spatial domain, fully retaining wafer features extracted by the residual group in combination with a hierarchical feature fusion structure, and performing super-resolution reconstruction on a wafer image pair; and S300, performing sub-pixel alignment based on phase correlation and a Gaussian function by using the wafer image reconstructed in the step S200. According to the wafer sub-pixel alignment method based on the image super-resolution, cutting defects caused by alignment errors of a scribing machine can be reduced, and the overall quality and yield of products are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of dicing machine image alignment, and particularly to a wafer sub-pixel alignment method based on image super-resolution. Background Art

[0002] The cutting of semiconductor chips is an important process in integrated circuit processing, and the dicing machine needs to achieve high-precision positioning and alignment. The dicing machine is a special equipment for precision cutting and is widely used in high-tech industries such as semiconductors, electronic ceramics, optics, and LED packaging. As integrated circuits transition from large-scale to ultra-large-scale, the integration density and patterns of chips are more precise and complex, and the dicing lanes are getting narrower and narrower. To ensure the precise separation of chips during the cutting process by the dicing machine, it is required to further improve the alignment accuracy of the dicing machine. Summary of the Invention

[0003] The purpose of the embodiments of the present invention is to provide a wafer sub-pixel alignment method based on image super-resolution, which can reduce the cutting defects caused by alignment errors of the dicing machine and improve the overall quality and yield of products.

[0004] To solve the above technical problems, the technical solution adopted by the present invention is:

[0005] The embodiments of the present invention provide a wafer sub-pixel alignment method based on image super-resolution, including the following steps:

[0006] S100, operate the dicing machine to collect a pair of wafer images;

[0007] S200, on the basis of the cascaded residual channel attention group, introduce enhanced spatial attention to capture the high-frequency information hidden in the spatial domain, and combine the hierarchical feature fusion structure to fully retain the wafer features extracted by the residual group, and perform super-resolution reconstruction on the pair of wafer images;

[0008] S300, use the pair of wafer images reconstructed in step S200 for sub-pixel alignment based on phase correlation and Gaussian function.

[0009] In some embodiments, in step S100, it includes the following steps: use a lens group with a magnification function in combination with an image sensor to collect a pair of wafer images multiple times, and combine an algorithm to extract the features of the wafer in the image, so as to analyze the deviation state of the wafer. According to the deviation state, control the motion platform to generate corresponding motion actions to complete the alignment operation of the wafer.

[0010] In some embodiments, in step S200, it includes a model, and the model is composed of three parts: shallow feature extraction, deep feature extraction, and reconstruction, including the following steps:

[0011] S210, use I LR and ISR which respectively represent the input and output images. First, shallow feature extraction is performed on the low-resolution image:

[0012] F SF = H SFE (I LR ) (1)

[0013] where H SFE (·) represents the convolutional operation of shallow feature extraction;

[0014] S220. Subsequently, F SF is fed into the network for deep feature extraction to obtain a more accurate feature map:

[0015]

[0016] where H DFF (·) represents the deep feature extraction operation, and W ESA (·) represents enhancing the spatial attention weight, represents the output of the nth residual group;

[0017] S230. After deep feature extraction, the obtained feature map and the feature map obtained by shallow feature extraction are added through a global skip connection to fully retain the shallow features; in addition, for stable training, both are also added to the feature map output by the last residual group;

[0018] S240. In the image reconstruction stage, the shallow features and deep features are respectively upsampled to obtain the super-resolution image of the required size:

[0019] I SR = H up (F SF ) + H up (F DF ) (3)

[0020] where H up (·) represents the upsampling module;

[0021] S250. Given a training set which contains N LR images and the corresponding HR images, the loss function can be expressed as:

[0022]

[0023] where θ represents all the parameters in the network, and H ESARHFN (·) represents the super-resolution image obtained through the model.

[0024] In some embodiments, in step S200, an enhanced spatial attention step is further included, which specifically includes the following steps:

[0025] S261, perform spatial domain weight assignment on the feature map by obtaining the attention map to retain precise spatial details; given the input ESA first extracts features as follows:

[0026]

[0027] wherein, is the weight of the 1×1 convolutional layer for reducing the embedding dimension;

[0028] S262, ESA further extracts features as follows:

[0029]

[0030] wherein, is the weight of the 3×3 convolution, with a stride of 2, H pool (·) is the max pooling operation, H up (·) is the upsampling operation implemented by bilinear interpolation, H g is a convolution group composed of three convolutions;

[0031] S263, the output of the ESA module can be calculated as:

[0032]

[0033] wherein, is the weight of the 1×1 convolutional layer for restoring the embedding dimension, H sigmoid (·) is the sigmoid function, and the symbol × is the pointwise multiplication operation.

[0034] In some embodiments, in step S200, a hierarchical feature fusion step is further included, specifically including the following steps:

[0035] S271, connect adjacent RG blocks;

[0036] S272, use 1×1 convolution to remove redundant information from adjacent blocks;

[0037] S273, repeat this process for all RG blocks and the resulting blocks generated by this mechanism until all blocks are integrated into a single RG block, and this block is convolved by 1×1 to generate output features;

[0038] S274, add this output to the output F of the shallow features SF in an element-wise manner.

[0039] The cutting of semiconductor chips is an important process in integrated circuit processing, and the dicing saw needs to achieve high-precision positioning and alignment. The dicing saw is a special equipment for precision cutting and is widely used in high-tech industries such as semiconductors, electronic ceramics, optics, and LED packaging. As integrated circuits transition from large-scale to ultra-large-scale, the integration density and patterns of chips are more precise and complex, and the dicing lanes are getting narrower. To ensure the precise separation of chips during the cutting process by the dicing saw, it is required to further improve the alignment accuracy of the dicing saw.

[0040] Super-resolution technology can break through the diffraction limit of traditional optical systems and reconstruct image details with higher resolution than the theoretical resolution of the system. It can effectively improve the positioning accuracy in the wafer alignment task. With the development of deep learning, compared with traditional methods, convolutional neural networks have achieved great success in the field of super-resolution. Networks such as EDSR and RDN have achieved good reconstruction effects by continuously increasing the network depth and using dense connections, but they ignore the transmission of underlying features. MSRN uses convolutional kernels of different sizes for feature extraction to increase the receptive field of the model, but treats different types of information equally. RCAN allows the network to focus on channels with more information, but does not consider the feature information in the spatial domain. To address the above problems, the present invention proposes an enhanced spatial attention residual hierarchical fusion network, which can fully extract the rich and complex texture information of the wafer, contributing to an approximately 0.05-pixel improvement in the accuracy of subsequent alignment experiments.

[0041] In recent decades, numerous experts and scholars have conducted in-depth exploration and research on image alignment. Among them, the accuracy of the interpolation method depends on the quality of the interpolation algorithm, and the prerequisite for the gradient method is that the image brightness remains unchanged. The principle of the optimization method can be attributed to the problem of minimizing the cost function with respect to the registration parameters, which can be applied to the registration of multi-spectral images and images with brightness changes, but this method has a high complexity. The phase correlation method is simple and easy to implement, with a small amount of computation, but the accuracy can only reach the integer-pixel level. To meet the high-precision alignment requirements of dicing saw cutting, the present invention adopts a method combining phase correlation and Gaussian curve fitting to obtain sub-pixel-level alignment accuracy.

[0042] Attention-based network

[0043] The features input in the network contain different types of information. Most networks uniformly process the input features and treat various types of information equally. The purpose of the attention mechanism is to effectively distinguish this information and give different responses to different types of information. HU et al. proposed SENet to learn the correlation between channels, and adaptively recalibrate the feature response intensity between channels through the global loss function of the network. This network structure has achieved significant performance improvement in image classification compared with traditional methods. Non-local attention was introduced into SISR in RNAN, but since pixel-by-pixel calculations are required, it consumes a large amount of computational resources and cannot be reused multiple times. HAN proposed a hierarchical attention to focus on the correlation between each residual block, and can assign different weights to different levels. Such a mechanism allows the network to focus on more information features and enhance the discrimination learning ability.

[0044] Feature fusion

[0045] Early super-resolution reconstruction models first perform upsampling on the LR image and then perform reconstruction, which reduces the learning difficulty but increases the time and space costs. A new framework was proposed to first extract features from the LR image and then insert the upsampling operation at the end of the network, which greatly reduces the computational amount while ensuring the reconstruction quality. This framework has also become one of the most popular frameworks in SR in recent years. EDSR uses this framework to improve ResNet by deleting some unnecessary modules, proving that deep networks are beneficial for reconstruction. RDN introduced the method of dense connection into the SR field, adding residual connections to increase the flow of information and gradients, which helps to build a very deep network. SRFBN uses a feedback mechanism to perform multiple backpropagations, which helps to increase the network depth without adding additional parameters, thereby improving the super-resolution reconstruction effect.

[0046] Sub-pixel alignment

[0047] Traditional image alignment methods are pixel-level matching and cannot meet the needs of visual positioning, medical image analysis, etc. that rely on high-precision image registration technology. Therefore, sub-pixel research has received extensive attention. Zhao Yang proposed improvements based on traditional surface fitting algorithms. Aiming at the previous binary cubic fitting with whole pixels as nodes, it adopted bicubic interpolation and cubic spline interpolation, and the fitting nodes became sub-pixel level, which has obvious advantages compared with traditional surface fitting algorithms. Dong Yi proposed a sub-pixel algorithm that combines gray interpolation and gradient methods, and its overall performance is better than that of bicubic interpolation, surface fitting, and gradient methods in terms of search speed and computational accuracy.

[0048] A wafer sub-pixel alignment method based on image super-resolution provided by the present invention shows through experiments that the alignment accuracy of the method proposed by the present invention can reach 0.05 pixel, meeting the requirements of the dicing machine for registration accuracy.

[0049] A wafer sub-pixel alignment method based on image super-resolution provided by the present invention can generally achieve the following technical effects:

[0050] (1) The introduced hierarchical feature fusion technology is improved to better fuse the extracted information to help the feature information at different levels smoothly transmit in the deep network.

[0051] (2) An enhanced spatial attention mechanism and global, long, and short skip connections are introduced. The former learns the weights at different positions in the spatial domain to improve the reconstruction of high-frequency details, and the latter utilizes hierarchical information as the network depth increases to simplify the information flow.

[0052] (3) The proposed model is trained using a self-made wafer dataset, the phase correlation is estimated using the super-resolution image reconstructed in step S200, and the Gaussian function fitting is adopted in the peak neighborhood to extend the estimation accuracy to the sub-pixel level. Description of the Drawings

[0053] To more clearly illustrate the technical solutions in the present disclosure, the drawings required for some embodiments of the present disclosure will be briefly introduced below. Obviously, the drawings in the following description are only the drawings of some embodiments of the present disclosure, and those of ordinary skill in the art can also obtain other drawings based on these drawings. In addition, the drawings in the following description can be regarded as schematic diagrams and do not limit the actual sizes of the products and the actual processes of the methods involved in the embodiments of the present disclosure.

[0054] Figure 1 It is a wafer alignment schematic diagram according to some embodiments of the present disclosure;

[0055] Figure 2 It is a schematic diagram of an enhanced spatial attention residual hierarchical fusion network according to some embodiments of the present disclosure;

[0056] Figure 3 It is a schematic diagram of enhanced spatial attention according to some embodiments of the present disclosure;

[0057] Figure 4 It is a comparison schematic diagram of different hierarchical feature fusion methods according to some embodiments of the present disclosure. Detailed Embodiments

[0058] Next, the technical solutions in some embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments provided by the present disclosure fall within the protection scope of the present disclosure.

[0059] Unless otherwise required by the context, throughout the specification and claims, the term "comprising" is interpreted in an open, inclusive sense, that is, "including, but not limited to". In the description of the specification, the terms "one embodiment", "some embodiments", "exemplary embodiments", "examples" or "some examples", etc., are intended to indicate that the specific features, structures, materials or characteristics related to the embodiment or example are included in at least one embodiment or example of the present disclosure. The schematic representations of the above terms are not necessarily referring to the same embodiment or example. In addition, the specific features, structures, materials or characteristics may be included in any one or more embodiments or examples in any appropriate manner.

[0060] An embodiment of the present invention provides a wafer sub-pixel alignment method based on image super-resolution, including steps: S100 to S300.

[0061] S100, operate a dicing machine to collect pairs of wafer images.

[0062] Before cutting the wafer, the dicing machine must perform an alignment operation. The alignment operation is to use a lens group with an amplification function in combination with an image sensor to collect pairs of wafer images multiple times, and combine corresponding algorithms to extract the geometric features of the wafer in the images, so as to analyze the deviation state of the wafer. According to the deviation state, control the motion platform to generate corresponding motion actions to complete the alignment operation of the wafer, as Figure 1 shown in the schematic diagram of wafer alignment.

[0063] Exemplarily, algorithms such as edge detection algorithm, Hough transform, template matching, and deep learning can be used to extract the geometric features of the wafer.

[0064] Specifically, in the automatic recognition and alignment system of the dicing machine, the cutting tool moves vertically along the Z axis, the workbench moves in the X and Y directions and rotates around the Z axis, and the wafer to be cut is fixed on the workbench by a vacuum chuck. When the system works, first move the alignment system workbench in the X and Y axis directions to make the wafer enter the field of view; then move the Z axis until the wafer imaging is clear; then the imaging system grabs the wafer image and inputs it into the computer, and the algorithm program analyzes the image to calculate the deviation value of the wafer; according to the calculation result, send an instruction to the motion control card, and the position of the workbench moves and rotates accordingly to adjust the wafer to the target position for cutting.

[0065] S200. Based on the cascaded residual channel attention group, an enhanced spatial attention is introduced to capture the high-frequency information hidden in the spatial domain, and the wafer features extracted by the residual group are fully retained by combining with the hierarchical feature fusion structure, and super-resolution reconstruction is performed on the wafer image pair.

[0066] S300. Use the reconstructed wafer image pair in step S200 for sub-pixel alignment based on phase correlation and Gaussian function.

[0067] Use the reconstructed wafer image pair to perform phase correlation whole-pixel displacement estimation, and obtain higher-precision sub-pixel matching points through Gaussian function fitting.

[0068] A wafer sub-pixel alignment method based on image super-resolution provided by the present invention can generally achieve the following technical effects:

[0069] (1). The introduced hierarchical feature fusion technology is improved, and the extracted information is better fused to help the feature information at different levels smoothly transmit in the deep network.

[0070] (2). An enhanced spatial attention mechanism and global, long, and short skip connections are introduced. The former learns the weights at different positions in the spatial domain to improve the reconstruction of high-frequency details, and the latter uses hierarchical information as the network depth increases to simplify the information flow.

[0071] (3). The proposed model is trained using a self-made wafer dataset, phase correlation estimation is performed using the super-resolution image reconstructed in step S200, and Gaussian function fitting is adopted in the peak neighborhood to extend the estimation accuracy to the sub-pixel level.

[0072] A super-resolution image is an image that improves its resolution by expanding and filling the pixel points of a low-resolution image through certain algorithm processing.

[0073] A low-resolution image is a type of image with fewer pixel points and unable to clearly display image details.

[0074] In step S200, it includes a model, which consists of three parts: shallow feature extraction, deep feature extraction, and reconstruction. The overall network structure is as Figure 2 shown. The shallow feature extraction consists of a 3×3 convolution. The deep feature extraction includes an enhanced spatial attention (ESA), a hierarchical feature fusion (HFF), and multiple cascaded residual groups (RG). In step S200, it includes steps: S210 to S250.

[0075] S210. Use I LR and I SR to represent the input and output images respectively. First, perform shallow feature extraction on the low-resolution image:

[0076] F SF = H SFE (I LR ) (1)

[0077] Among them, H SFE (·) represents the convolutional operation for shallow feature extraction.

[0078] S220, then send F SF into the network for deep feature extraction to obtain a more accurate feature map:

[0079]

[0080] Among them, H DFF (·) represents the deep feature extraction operation, and W ESA (·) represents enhancing the spatial attention weight, represents the output of the nth residual group.

[0081] S230, the feature map obtained after deep feature extraction and the feature map obtained by shallow feature extraction are added through a global skip connection to fully retain the shallow features; in addition, for stable training, both are also added to the feature map output by the last residual group.

[0082] S240, in the image reconstruction stage, the shallow features and deep features are respectively upsampled to obtain a super-resolution image of the required size:

[0083] I SR = H up (F SF ) + H up (F DF ) (3)

[0084] Among them, H up (·) represents the upsampling module.

[0085] There are many loss functions in SISR, such as L1 loss, L2 loss, perceptual loss, and adversarial loss. To optimize the network of the present invention, the L1 loss function is adopted.

[0086] S250, given a training set which contains N LR images and the corresponding HR images, the loss function can be expressed as:

[0087]

[0088] Among them, θ represents all the parameters in the network, and H ESARHFN (·) represents the super-resolution image obtained through the model.

[0089] LR is a low-resolution image, HR is a high-resolution image, and SR is a super-resolution image.

[0090] After the channel attention in the residual block adaptively learns the features of each channel, it is necessary to capture in the spatial domain to share the information within the features. The present invention introduces an enhanced spatial attention (ESA) module. Compared with the ordinary spatial attention (SA) module, ESA can obtain a larger receptive field and better adaptively reallocate features according to the spatial context content.

[0091] In step S200, it further includes: S260, the enhanced spatial attention step, which specifically includes steps: S261 to S263.

[0092] S261, as Figure 3 shown, weight distribution in the spatial domain of the feature map is performed by obtaining the attention map, and precise spatial details are retained. Given the input ESA first extracts features as follows:

[0093]

[0094] Among them, is the weight of the 1×1 convolutional layer used to reduce the embedding dimension.

[0095] S262, ESA further extracts features as follows:

[0096]

[0097] Among them, is the weight of the 3×3 convolution, the stride is 2, H pool (·) is the max pooling operation, H up (·) is the upsampling operation implemented by bilinear interpolation, H g is a convolution group composed of three convolutions.

[0098] The spatial dimension is reduced through the cross-row convolutional layer and the max pooling layer, and then the spatial dimension is restored through the upsampling layer.

[0099] S263, the output of the ESA module can be calculated as:

[0100]

[0101] Among them, is the weight of the 1×1 convolutional layer used to restore the embedding dimension, H sigmoid (·) is the sigmoid function, and the symbol × is the pointwise multiplication operation.

[0102] Calibrate the response of features in a local range using enhanced spatial attention, interact with channel attention, and improve the reconstruction ability for high-frequency details.

[0103] In the depth feature extraction, as the network depth increases, information loss is inevitable. As Figure 4 shown, the HFF structure proposed in MSRN directly concatenates all outputs to the last layer, but it cannot smoothly transform the low-level, middle-level, and high-level features, resulting in improper utilization of mid-low level features. Therefore, the present invention proposes a smoother feature fusion technology to layer-by-layer fuse the rich features extracted by the residual groups, better retaining useful information for reconstruction.

[0104] In step S200, it further includes: S270, a hierarchical feature fusion step, specifically including steps: S271 to S274.

[0105] S271, connect adjacent RG blocks.

[0106] S272, use 1×1 convolution to remove redundant information from adjacent blocks.

[0107] S273, repeat this process for all RG blocks and the resulting blocks generated by this mechanism until all blocks are integrated into a single RG block, and this block is convolved through 1×1 to generate output features.

[0108] S274, add this output to the output F of the shallow features in an element-wise manner SF in.

[0109] It includes the following formula:

[0110] F i+1 = H RG (F i ) (8)

[0111] Among them, H RG (·) represents the features extracted by a single RG block, and F i represents the i th feature map extracted.

[0112] After all features are extracted by the RG blocks, these RG blocks can be utilized in the HFF architecture.

[0113] M j = H 1×1 [F i+1 , F i+2 (9)

[0114]

[0115] Here, the outputs of two adjacent RG blocks are concatenated by channel and then passed to a 1x1 convolutional layer to avoid redundant information. Therefore, four RG blocks will generate two additional blocks, such as F i+1 = M j 、F i+2 = M j+1 , M j and M j+1 will be processed in a similar manner as two RG blocks. This process is repeated until all the RGs and the resulting blocks are integrated into an output, which is further used as the input for the reconstruction step.

[0116] In step S300, the sub-pixel alignment algorithm based on phase correlation and Gaussian function specifically includes the following steps:

[0117] The phase correlation method is based on the Fourier shift property to establish the correspondence between the phase difference in the frequency domain and the translation in the spatial domain. Assume f 1 (x,y) is the reference image, and f 2 (x,y) is the image obtained by translating f 1 (x,y) by (x 0 ,y 0 ). The two images satisfy the following relationship:

[0118] f 2 (x,y) = f 1 (x - x 0 , y - y 0 ) (11)

[0119] Taking the Fourier transform of both sides of equation (11), according to the Fourier shift property, we can obtain:

[0120] F 2 (u,v) = F 1 (u,v) exp(-j2π(ux 0 + vy 0 )) (12)

[0121] The normalized cross-power spectrum P(u,v) between two images can be expressed as:

[0122]

[0123] where * represents the complex conjugate. Taking the inverse Fourier transform of both sides of equation (13), we can obtain the phase correlation function as follows:

[0124] p(x,y) = F -1 {exp(-j2π(ux 0 + vy 0 ))} = δ(x - x0 , y - y 0 ) (14)

[0125] where δ(x - x 0 , y - y 0 ) is a typical Dirac function, also known as the impulse function. This function is non - zero at the center point (x 0 , y 0 ) and zero at other positions. According to the peak distribution of the function, the image translation parameters can be obtained as follows:

[0126]

[0127] The above - mentioned method can only produce displacement estimates with integer precision. To achieve sub - pixel precision, a parabola function or a Gaussian function can be fitted in the neighborhood of the phase - correlation peak. Using formula (15), the prototype function is fitted to the triple:

[0128] {p(x m - 1, y m ), p(x m , y m ), p(x m + 1, y m )}(16)

[0129]

[0130] The Gaussian function is horizontally fitted to the data triple of formula (16):

[0131]

[0132] The vertical component △y can be fitted to formula (17) in the same way to obtain it. (△x, △y) is the desired sub - pixel displacement.

[0133] Experimental setup:

[0134] The images used in the experiment were collected by an 8230 full - automatic double - axis dicing machine from ADT Company in Israel. The self - made wafer dataset contains 600 high - quality images with different sizes, processes, environments, and position scenarios. For the super - resolution experiment, the self - made dataset was used to train the proposed model, and the test set Wafer100 was used to verify the generalization of the model. At the same time, the peak signal - to - noise ratio (PSNR) and the structural similarity (SSIM) were used to measure the quality of the SR images. In addition, the resolution of the alignment experimental images was 512×512 pixels, and the images to be registered were taken at random pixel steps in the horizontal or vertical direction of the reference image.

[0135] The network was trained through the minimum loss function, and the ADAM optimizer was adopted, where β 1 = 0.9, β2 = 0.999, ε = 10 -8 。The initial learning rate is set to 10 -4 , and it is halved every 2×10 5 iterations. For training and testing, the Pytorch framework is used to build the network on an NVIDIA GeForce RTX 3090 GPU. For the wafer alignment experiment, sub-pixel measurements in the horizontal and vertical directions are carried out on MATLAB R2014a.

[0136] Experimental analysis:

[0137] As shown in Table 1, under the alignment algorithm based on phase correlation and Gaussian function fitting, comparative experiments on the measured displacement are carried out using the super-resolution reconstructed images of different models. The maximum alignment errors of the MSRN, RDN, and ESARHFN models are 0.0523, 0.04, and 0.0273 pixel respectively. The reason for the reduction in error is that the RDN ignores the contribution of different frequency information to the super-resolution reconstruction result, and the MSRN does not fuse the extracted features smoothly. In contrast, the model in the present invention introduces an ESA (Enhanced Spatial Attention) module, which improves the attention to high-frequency information. At the same time, an improved HFF (Hierarchical Feature Fusion) module is also introduced, which can process the features extracted from different residual groups more smoothly. Since the reconstructed images of the model in the present invention have obtained better results, the average accuracy has been improved by about 0.02 and 0.01 pixel compared with the MSRN and RDN respectively. It is proved that the super-resolution algorithm can increase the pixel density of the image, thereby providing more texture detail information, which helps to more accurately identify and match image features during the alignment experiment and improve the alignment accuracy.

[0138] Table 1 Comparison of measured displacements of reconstructed images of different models

[0139]

[0140]

[0141] As shown in Table 2, the comparative data of sub-pixel displacement measurement using the source image and the super-resolution image are shown. The method used is the phase correlation method based on Gaussian function fitting. The maximum error in the alignment experiment using the source image is 0.0587 pixel, and the minimum error is 0.0237 pixel. The maximum error in the alignment experiment using the super-resolution image is 0.0356 pixel, and the minimum error is 0.012 pixel. It can be seen from the fifth group that, compared with the source image, the error in the alignment experiment using the super-resolution image can be reduced by nearly 0.05 pixel.

[0142] Table 2 Comparison of measured displacements of the source image and the super-resolution image by the same method

[0143]

[0144] Table 3 shows the comparative experimental results of different fitting methods under the reconstructed images of the ESARHFN model. The results of the method used in the present invention are indicated in bold. Since the translation parameters of the known image are available, the mean and standard values of the translation value estimation error can be used to evaluate the registration accuracy of the method. The phase correlation method can only achieve pixel-level registration. By using a fitting function, the accuracy can be extended to the sub-pixel level, and this method is called the extended phase correlation method. The error of the extended phase correlation experiment mainly depends on the matching degree between the shape of the fitted function and the data distribution. The experimental results show that the Gaussian function fitting has a relatively small average error, and its curve has the characteristics of symmetry and rapid decay, which is more suitable for describing the phase correlation peak distribution of the wafer image. Specifically, compared with the other three algorithms, the average accuracy is improved by approximately 0.4, 0.3, and 0.02 pixels respectively. The accuracy of the parabola function is the second and similar, because the descending speed on both sides of its peak is not as fast as that of the Gaussian function, but it can still capture the center position of the peak well. The accuracy of the Taylor expansion depends on the choice of the expansion point and the order of expansion. The Lorentz function has a relatively slow decay and is suitable for describing a distribution with a wider tail. In addition, from the analysis of the average running time and standard deviation in Table 4, it can be seen that the Gaussian function fitting has low computational efficiency but excellent stability. Considering the specific industrial requirements of the dicing machine, the present invention selects the phase correlation method with Gaussian function fitting. This algorithm can accurately describe the distribution law of feature points under the periodic lattice structure of the wafer, so as to better achieve precise matching in the image alignment process.

[0145] Table 3 Error Comparison of Different Fitting Methods

[0146]

[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Any other modifications or equivalent replacements made by those of ordinary skill in the art to the technical solutions of the present invention should be covered within the scope of the claims of the present invention as long as they do not depart from the spirit and scope of the technical solutions of the present invention.

Claims

1. A wafer sub-pixel alignment method based on image super-resolution, characterized in that: The following steps are involved: S100, operating a dicing machine to collect wafer image pairs; S200, based on the cascaded residual channel attention group, introduces enhanced spatial attention to capture the high-frequency information implicit in the spatial domain, and combines the hierarchical feature fusion structure to fully retain the wafer features extracted by the residual group, and perform super-resolution reconstruction of the wafer image pair; S300 , performing sub-pixel alignment based on phase correlation and Gaussian function using the wafer image reconstructed in step S200 .

2. A wafer sub-pixel alignment method based on image super-resolution as claimed in claim 1, characterized in that: In step S100, the following steps are included: A lens group with a magnifying function is used in conjunction with an image sensor to collect wafer image pairs multiple times, and an algorithm is used to extract the features of the wafer in the image, thereby analyzing the wafer deviation state. Based on the deviation state, the motion platform is controlled to generate corresponding motion actions to complete the wafer alignment operation.

3. A wafer sub-pixel alignment method based on image super-resolution as claimed in claim 2, characterized in that: In step S200, a model is included, and the model consists of three parts: shallow feature extraction, deep feature extraction and reconstruction, including the following steps: S210, with I LR and I SR Represent the input and output images respectively. First, shallow feature extraction is performed on the low-resolution image: F SF =H SFE (I LR ) (1) Among them, H SFE (·) represents the convolution operation for shallow feature extraction; S220, then F SF Send it into the network for deep feature extraction to get a more accurate feature map: Among them, H DFF (·) represents the deep feature extraction operation, W ESA (·) represents the enhanced spatial attention weight, represents the output of the nth residual group; S230, the feature map obtained after deep feature extraction and the feature map extracted from shallow features are added through a global skip link to fully retain the shallow features; in addition, in order to stabilize the training, the two are also added to the feature map output by the last residual group; S240, in the image reconstruction stage, upsampling operations are performed on the shallow features and the deep features to obtain a super-resolution image of the required size: I SR =H up (F SF )+H up (F DF ) (3) Among them, H up (·) indicates upsampling module; S250, given a training set It contains N LR images and corresponding HR images, and the loss function can be expressed as: Among them, θ represents all the parameters in the network, H ESARHFN (·) represents the super-resolution image obtained by the model.

4. The wafer sub-pixel alignment method based on image super-resolution as claimed in claim 3, characterized in that: In step S200, an enhanced spatial attention step is also included, which specifically includes the following steps: S261, by obtaining the attention map, the feature map is weighted in the spatial domain to retain the precise spatial details; given the input ESA first extracts features as follows: in, is the weight of the 1×1 convolutional layer used to reduce the embedding dimensionality; S262, ESA further extracts features as follows: in, is the weight of 3×3 convolution, with a stride of 2, H pool (·) is the maximum pooling operation, H up (·) is the upsampling operation achieved by bilinear interpolation, H g is a convolution group consisting of three convolutions; S263, output of ESA module It can be calculated as: in, is the weight of the 1×1 convolutional layer used to recover the embedding dimension, H sigmoid (·) is the sigmoid function, and the symbol × is a point-by-point multiplication operation.

5. The wafer sub-pixel alignment method based on image super-resolution as claimed in claim 4, characterized in that: In step S200, a hierarchical feature fusion step is also included, which specifically includes the following steps: S271, connect adjacent RG blocks; S272, uses 1×1 convolution to remove redundant information from adjacent blocks; S273, repeat this process for all RG blocks and result blocks produced by this mechanism until all blocks are integrated into a single RG block, which is convolved by 1×1 to produce output features; S274, add this output to the output F of the shallow feature in an element-wise manner SF middle.