Photoetching alignment image enhancement method based on super-resolution generative adversarial network

By constructing a super-resolution generation adversarial network and adaptive histogram equalization processing of lithographic alignment images based on three-dimensional DenseNet fusion SE blocks, the problem of low positioning accuracy in lithographic alignment is solved, and efficient lithographic alignment and overturning alignment is achieved.

CN120259123AActive Publication Date: 2025-07-04CHONGQING UNIV OF POSTS & TELECOMM +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510343450.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-04
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

In the existing lithographic alignment technology, the low contrast of the alignment mark image after the wafer is coated and baked, the structure is blurred, and nonlinear distortion leads to low positioning accuracy, affecting the incision accuracy.

Method used

A lithographically aligned image super-resolution generation adversarial network (AlignGAN) based on three-dimensional DenseNet fusion SE blocks is constructed, combined with adaptive histogram equalization processing, super-resolution reconstruction and noise reduction of the alignment marks are carried out, and precise engraving is achieved using OpenCV template matching.

Benefits of technology

It significantly improves the accuracy and efficiency of lithographic alignment, reduces interference in image processing, and achieves high-precision intercalation alignment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259123A_ABST
    Figure CN120259123A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of photoetching alignment and image processing, and discloses a photoetching alignment image enhancement method based on a super-resolution generative adversarial network, which comprises the steps of establishment and preprocessing of an AlignNet-SR data set, overlay image super-resolution reconstruction and photoetching alignment template feature matching and center positioning calculation. And performing overlay alignment of multiple exposures of the photoetching machine. According to the method, a photoetching alignment image super-resolution generative adversarial network (AlignGAN) is constructed, adversarial training of a generator and a discriminator is used, and adaptive histogram equalization processing of a photoetching alignment visual image is combined, so that the visual quality and the resolution of the photoetching alignment image are remarkably improved. After center positioning of the alignment mark is obtained through template feature matching, rapid and accurate alignment is achieved through the inner contour mark and the outer contour mark of the photoetching image, the method can adapt to various marks, the alignment accuracy is effectively improved, different areas of the image can be processed more meticulously, and therefore the situation of excessive enhancement or detail loss is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image super-resolution, and specifically provides a method for enhancing lithography alignment images based on a super-resolution generative adversarial network. Background Art

[0002] With the development of integrated circuits, the demand for more advanced and mature process capabilities is also increasing day by day. Integrated circuits require mature and advanced processing solutions. Taking the field of autonomous driving as an example, for actuator control, it is necessary to achieve mature signal-guided control requirements, while for central computing chips, high-speed processing capabilities are required to meet the needs of service-oriented operations. Therefore, integrated circuits are a key demand and a field of concentrated development in all walks of life. In such a market environment, lithography machines, as key equipment for integrated circuit manufacturing, have huge application demands. With the continuous development of integrated circuits and the growth of demand, lithography machines will continue to play an important role in the future industry, meeting the industry's demand for higher-level process capabilities. In the entire integrated circuit manufacturing process, lithography is the most core and complex process step. Lithography technology is to expose through a photoresist under the irradiation of a specific light source, and then transfer the pattern on the mask to the silicon wafer through steps such as development and etching. In the lithography process, each layer of circuit pattern needs to be exposed once, and a special mask plate is used. To ensure the precise alignment of each layer of pattern, the position of the mask plate must be exactly the same as the pattern obtained from the previous exposure. Since the design and materials of each layer of circuit are different, different mask plates are required to achieve the unique patterns of each layer. This approach ensures that the patterns formed on each layer can be accurately superimposed together, avoiding offset or misalignment, thus guaranteeing the performance and quality of the final chip. The overlay accuracy between the mask and the silicon wafer is one of the core performance indicators of the lithography machine, and the overlay accuracy needs to be achieved through lithography alignment. Therefore, the lithography alignment technology, as one of the three core technologies of lithography, plays a crucial role in lithography production. Improving the accuracy of lithography alignment can directly improve the overlay accuracy, thereby improving product quality. At the same time, improving the speed and efficiency of lithography alignment can also improve the productivity of products.

[0003] The overlay alignment technology is an alignment technology applied to multiple exposures of proximity contact lithography machines, and widely uses a video image alignment method based on geometric patterns. With the development of semiconductors, various types of overlay marks have emerged. In some lithography processes, special alignment marks are no longer designed, but the patterns of lithography or the contours of the silicon wafer are directly used as alignment marks for alignment. The overlay alignment algorithm determines the key factor of the overlay accuracy of a fully automatic exposure machine. However, after the wafer is coated with glue and baked, interference such as low contrast, blurred structure, and non-linear distortion of the alignment mark image directly leads to low positioning accuracy.

[0004] Based on this, the present invention provides a method for enhancing lithography alignment images based on a super-resolution generative adversarial network. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the present invention provides a lithography alignment image enhancement method based on a super-resolution generative adversarial network, which has an alignment image denoising and blind image super-resolution reconstruction method based on deep learning. By constructing a lithography alignment image super-resolution generative adversarial network (AlignGAN) based on the fusion of a three-dimensional DenseNet and a SE block (Squeeze-and-Excitation), which consists of a generator network and a discriminator network, combined with an adaptive histogram equalization process for lithography alignment visual images, it effectively reduces various interferences during the overlay alignment process, realizes the advantages of super-resolution and high-definition enhancement of alignment marks, and solves the problem of low positioning accuracy caused by low contrast, blurred structure, non-linear distortion and other interferences in the alignment mark images of wafers after spin coating and baking in the background technology.

[0006] The present invention provides the following technical solutions: A lithography alignment image enhancement method based on a super-resolution generative adversarial network, comprising the following steps:

[0007] S1: Establishment of the AlignNet-SR dataset. The dataset sources are high-resolution overlay mark images covering various overlay mark designs and wafer process conditions and corresponding low-resolution simulated images;

[0008] S2: Preprocessing of the dataset, performing a series of operations such as cropping and standardization, data cleaning and enhancement, and dataset division on the dataset;

[0009] S3: Acquisition of the template image and the overlay image. The image acquisition system acquires the template image before exposure and the overlay image after each exposure;

[0010] S4: Apply adaptive histogram equalization to process the acquired overlay pattern area. This technology can effectively eliminate the square overlay marks on the wafer after exposure by adjusting the brightness and contrast of the local area, significantly improving the quality of the overlay image, thereby making the subsequent multiple exposure process more accurate and reliable;

[0011] S5: After S4, construct a lithography alignment image super-resolution generative adversarial network (AlignGAN) based on the fusion of a three-dimensional DenseNet and a SE block (Squeeze-and-Excitation). This network can perform alignment image denoising and blind image super-resolution reconstruction based on deep learning. The introduced compression excitation SE block can complete the adaptive learning of each channel of the feature map, which can not only improve the utilization rate of effective features, but also reduce the influence of network redundancy;

[0012] S6: Template feature matching. After denoising and enhancement of the overlay marks based on the deep adversarial network in S5, use the cv.matchTemplate() function of OpenCV for template matching. After obtaining the best matching position, calculate its center positioning data. Finally, the motion control system realizes precise overlay alignment through the center positioning data of the marks.

[0013] Preferably, the overlay mark image in S1 includes an outer contour mark and an inner contour mark. Among them, the outer contour mark is set on the mask plate, and the inner contour mark is set on the wafer, and a low-resolution simulated image is generated by simulating lithography blur such as downsampling, Gaussian blur, and adding noise to the high-resolution image.

[0014] Preferably, the preprocessing operation of the dataset in S2 specifically includes: First, crop the region of interest (ROI) containing the overlay marks from the original image pair, unify the size of the image patches, and normalize the pixel values of the images. Then, after data cleaning, use data augmentation techniques such as geometric transformation, noise addition, brightness and contrast adjustment to expand the dataset for the obtained image pair. Finally, divide the preprocessed dataset into a training set, a validation set, and a test set according to a certain proportion.

[0015] Preferably, the adaptive histogram equalization technique is applied in S4 to weaken or eliminate the square overlay marks printed on the wafer image after exposure, specifically including: first convert the color image to a grayscale image; then apply Gaussian blur to denoise the grayscale image; after dividing the grayscale image into multiple non-overlapping and equal-sized sub-image patches, calculate the histogram of each sub-image patch to obtain the frequency of gray levels, and then calculate the cumulative distribution function to calculate the equalized pixel values; finally, perform bilinear interpolation on each sub-image patch to eliminate block effects, and then recombine all the enhanced sub-image patches into a complete image.

[0016] Preferably, in S5, a lithography alignment image super-resolution generative adversarial network (AlignGAN) based on three-dimensional DenseNet fused with SE blocks (Squeeze-and-Excitation) is introduced. Its training data is sourced from the AlignNet-SR dataset, including processed high-resolution registration mark images and their corresponding low-resolution simulated images. After initializing the parameters of the generator G and the discriminator D, defining the loss function and the optimizer, the generator G and the discriminator D are alternately trained to obtain the reconstructed super-resolution image. The specific training steps are as follows: The generator network takes the generated low-resolution simulated image as input, performs preliminary feature extraction on it through the initial three-dimensional convolutional layer, and uses the Leaky ReLU activation function to increase its non-linear expression ability. When the image is input into the Dense-SENet unit composed of three-dimensional DenseNet and SE blocks (Squeeze-and-Excitation) (the number of which is set to 16), first, the bottleneck layer 1x1 convolution in the dense block reduces the dimension and the number of channels of the feature map, then three-dimensional convolution is used to extract spatial features, and an SE block is added after the dense block for channel calibration, enhancing important features through weighting and suppressing unimportant features. The transition layer generates global features through pooling operations. Each Dense-SENet unit contains multiple three-dimensional convolutional layers, and the feature maps of all previous layers can be accessed through dense connections. After stacking multiple units, the dimension is further reduced through the bottleneck layer, and finally, a super-resolution image is generated through two upsampling layers and a three-dimensional convolutional layer. The discriminator network consists of basic units composed of stacked three-dimensional convolutional layers, Leaky ReLU activation functions, normalization layers (Normal Block), and Dense layers. First, the real high-resolution template image and the generated high-resolution image are input, and the images are converted to grayscale images to simplify the calculation. Then, preliminary and advanced features of the images are extracted through four 3D convolutional layers, and the Leaky ReLU activation function is used to increase the non-linear expression ability and avoid the problem of neuron death. After the convolutional layer, a batch normalization layer (Batch Normalization) is applied to standardize the features, reduce the covariance shift, and stabilize and accelerate the training process. Then, the fully connected layer Dense(1024) is used to flatten the extracted features and map them to a feature space with a fixed dimension, and a scalar is output through the final fully connected layer. This scalar is converted into a probability value between 0 and 1 through the Sigmoid activation function, representing the possibility that the image is a real image. Finally, the discriminant loss is calculated through the cross-entropy loss function, and the parameters of the discriminator are updated using the backpropagation algorithm to improve its ability to distinguish real images and generated images.

[0017] Preferably, in S6, the template matching algorithm (cv.matchTemplate()) of OpenCV is used to match the super-resolution image obtained in S5 with the average template generated from the marker image obtained by mask overlay. The matching is performed within the region of interest (ROI) of the target image, the normalized correlation coefficient between the template and the target image is calculated, the position with the highest matching score is found, the matching points are extracted, the RANSAC algorithm is used to estimate the geometric transformation matrix, the center positioning of the marker is obtained, and the position of the target image is corrected to ensure that the mask and the alignment markers on the wafer are accurately aligned. The effect of template matching and correction can be verified by visualizing the binary alignment marker contour, so as to achieve high-precision lithography alignment. The system sends a motion signal to the motion system according to the alignment motion data, so that the outer contour marker and the inner contour marker coincide, and then precise overlay is achieved.

[0018] Compared with the prior art, the present invention has the following beneficial effects:

[0019] (1) In the present invention, Adaptive Histogram Equalization (AHE) is introduced in the processing of overlay markers, which significantly improves the contrast of the image and the equalization of detail display, equalizes the image brightness and contrast, weakens or eliminates the square overlay markers on the wafer image after exposure, provides a necessary prerequisite for the realization of multiple exposures, and compared with the traditional Histogram Equalization (HE), it can process different regions of the image more carefully, thus avoiding the situation of over-enhancement or detail loss;

[0020] (2) The present invention not only constructs a high-quality AlignNet-SR dataset suitable for overlay markers in the lithography process, but also after adversarial training, the deep generative adversarial network can enable the generator to efficiently generate realistic images that are almost indistinguishable from real images, overcoming the dependence of traditional deep learning models on a large number of training samples and having the characteristics of unsupervised learning. Therefore, it has significant real-time and efficiency advantages in generating realistic and diverse images;

[0021] (3) Compared with the traditional super-resolution generative adversarial network, the present invention improves the three-dimensional Dense Net structure, adds a transition layer to reduce the dimension of the feature maps in the dense block to reduce the training difficulty of the network, and further adds a bottleneck layer to perform feature map dimension reduction processing before the three-dimensional feature maps are input into the dense block, significantly reducing the dimension of the feature maps, improving the training efficiency, and alleviating the problem of computational consumption;

[0022] (4) To enable the network to fully learn image features, a squeeze-and-excitation (SE) block is introduced into the 3D DenseNet. Considering the characteristics of the SE block, we insert the SE block in the transition layer of the 3D-DenseNet, which can filter out some unimportant features, squeeze and excite the effective features in the transition layer. Therefore, introducing the SE block into the 3D-DenseNet can not only improve the utilization rate of effective features but also reduce the impact of network redundancy. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 FIG. is a flowchart of the lithography alignment method based on the super-resolution generative adversarial network of the present invention;

[0024] Figure 2 FIG. is a schematic diagram of the lithography alignment image super-resolution generative adversarial network (AlignGAN) of the present invention;

[0025] Figure 3 FIG. is a schematic diagram of the improved three-dimensional dense block DenseNet structure of the present invention;

[0026] Figure 4 FIG. is a schematic diagram of the fused squeeze-and-excitation (SE) block structure of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0028] Please refer to Figures 1-4 , a lithography alignment image enhancement method based on the super-resolution generative adversarial network, comprising the following steps:

[0029] S1: Establishment of the AlignNet-SR dataset. The dataset is sourced from high-resolution overlay mark images covering various overlay mark designs and wafer process conditions and corresponding low-resolution simulated images;

[0030] S2: Preprocessing of the dataset, including a series of operations such as cropping and normalizing the dataset, data cleaning and enhancement, and dataset partitioning;

[0031] S3: Acquisition of the template image and the overlay image. The image acquisition system acquires the template image before exposure and the overlay image after each exposure;

[0032] S4: Apply adaptive histogram equalization to process the collected overlay pattern area. This technology can effectively eliminate the square overlay marks on the wafer after exposure by adjusting the brightness and contrast of the local area, significantly improving the quality of the overlay image, thus making the subsequent multiple exposure process more accurate and reliable.

[0033] S5: After S4, construct a lithography alignment image super-resolution generative adversarial network (AlignGAN) based on a three-dimensional DenseNet fused with SE blocks (Squeeze-and-Excitation). This network can perform alignment image denoising and blind image super-resolution reconstruction based on deep learning. The introduced squeeze-and-excitation (SE) blocks can complete the adaptive learning of each channel of the feature map, which can not only improve the utilization rate of effective features but also reduce the impact of network redundancy.

[0034] S6: Template feature matching. After S5 performs denoising and enhancement on the overlay marks based on the deep adversarial network, use the cv.matchTemplate() function of OpenCV for template matching. After obtaining the best matching position, calculate its center positioning data. Finally, the motion control system realizes precise overlay alignment through the center positioning data of the mark.

[0035] Among them; the overlay mark image in S1 includes an outer contour mark and an inner contour mark. The outer contour mark is set on the mask plate, and the inner contour mark is set on the wafer. A low-resolution simulated image is generated by simulating lithography blur such as downsampling, Gaussian blur, and adding noise to the high-resolution image.

[0036] Among them; the preprocessing operations of the dataset in S2 specifically include: First, crop the region of interest (ROI) containing the overlay marks from the original image pair, unify the size of the image patches, and normalize the pixel values of the images. Then, after data cleaning, use data augmentation techniques such as geometric transformation, noise addition, brightness and contrast adjustment to expand the dataset for the obtained image pair. Finally, divide the preprocessed dataset into a training set, a validation set, and a test set according to a certain ratio.

[0037] Among them; applying the adaptive histogram equalization technology in S4 to weaken or eliminate the square overlay marks printed on the wafer image by the mask after exposure specifically includes: first converting the color image to a grayscale image; then applying Gaussian blur to denoise the grayscale image; after dividing the grayscale image into multiple non-overlapping and equal-sized sub-image patches, calculate the histogram of each sub-image patch to obtain the frequency of gray levels, and then calculate the cumulative distribution function to calculate the equalized pixel values; finally, perform bilinear interpolation on each sub-image patch to eliminate block effects, and then recombine all the enhanced sub-image patches into a complete image.

[0038] Among them, in S5, a lithography alignment image super-resolution generative adversarial network (AlignGAN) based on three-dimensional DenseNet fused with SE blocks (Squeeze-and-Excitation) is introduced. Its training data is sourced from the AlignNet-SR dataset, including processed high-resolution registration mark images and their corresponding low-resolution simulated images. After initializing the parameters of the generator G and discriminator D, defining the loss function and optimizer, the generator G and discriminator D are alternately trained to obtain the reconstructed super-resolution image. The specific training steps are as follows: The generator network takes the generated low-resolution simulated image as input, performs preliminary feature extraction on it through an initial three-dimensional convolutional layer, and uses the Leaky ReLU activation function to enhance its non-linear expression ability. When the image is input into the Dense-SENet unit composed of three-dimensional DenseNet and SE blocks (Squeeze-and-Excitation) (the number of which is set to 16), first, the dimensionality and number of channels of the feature map are reduced through the bottleneck layer 1x1 convolution in the dense block, then three-dimensional convolution is used to extract spatial features, and an SE block is added after the dense block for channel calibration, enhancing important features through weighting and suppressing unimportant features. The transition layer generates global features through pooling operations. Each Dense-SENet unit contains multiple three-dimensional convolutional layers and can access the feature maps of all previous layers through dense connections. After stacking multiple units, the dimensionality is further reduced through the bottleneck layer, and finally, a super-resolution image is generated through two upsampling layers and a three-dimensional convolutional layer. The discriminator network consists of basic units composed of stacked three-dimensional convolutional layers, Leaky ReLU activation functions, normalization layers (Normal Block), and Dense layers. First, the real high-resolution template image and the generated high-resolution image are input, and the images are converted to grayscale images to simplify the calculation. Then, preliminary and advanced features of the images are extracted through four 3D convolutional layers, and the non-linear expression ability is enhanced through the Leaky ReLU activation function to avoid the problem of neuron death. After the convolutional layer, a batch normalization layer (Batch Normalization) is applied to standardize the features, reduce covariance shift, and stabilize and accelerate the training process. Then, the extracted features are flattened and mapped to a feature space of a fixed dimension using the fully connected layer Dense(1024), and a scalar is output through the final fully connected layer. This scalar is converted into a probability value between 0 and 1 through the Sigmoid activation function, representing the likelihood that the image is a real image. Finally, the discriminant loss is calculated through the cross-entropy loss function, and the parameters of the discriminator are updated using the backpropagation algorithm to improve its ability to distinguish real images and generated images.

[0039] Among them; in S6, the template matching algorithm (cv.matchTemplate()) of OpenCV is used to match the super-resolution image obtained in S5 with the average template generated from the marker image obtained by mask overlay. The matching is performed within the region of interest (ROI) of the target image, the normalized correlation coefficient between the template and the target image is calculated, the position with the highest matching score is found, the matching points are extracted, the RANSAC algorithm is used to estimate the geometric transformation matrix, the center positioning of the marker is obtained, and the position of the target image is corrected to ensure that the mask and the alignment markers on the wafer are accurately aligned. The effect of template matching and correction can be verified by visualizing the binary alignment marker contour, so as to achieve high-precision lithography alignment. The system sends a motion signal to the motion system according to the alignment motion data, so that the outer contour marker and the inner contour marker coincide, and then precise overlay is achieved.

[0040] Please refer to Figure 1 , a lithography alignment method based on a super-resolution generative adversarial network, which generally includes three parts: establishment and preprocessing of the dataset, reconstruction of the overlay marker image, extraction of alignment marker features and template matching. The specific steps are as follows:

[0041] Step 1: Establishment of the AlignNet-SR dataset. Use a high-precision microscopic imaging device to collect the original overlay marker images without defocusing treatment to ensure that they have complete edge details and texture information. To meet the requirements of GAN training, at least thousands of high-resolution images are collected, covering a variety of overlay marker designs and wafer process conditions. Low-resolution images are generated from the high-resolution images by means of downsampling, Gaussian blur, and adding noise, so as to simulate as much as possible the blurring phenomenon caused by factors such as optical system imaging distortion, equipment vibration, and optical diffraction after exposure. And for training the generative adversarial network, the collected high-resolution images and the generated low-resolution images must be strictly paired;

[0042] Step 2: Perform preprocessing steps on the dataset, such as image cropping and normalization, data cleaning and enhancement, and dataset partitioning;

[0043] Specifically, first, the region of interest (ROI) containing the registration marks is cropped from the original image pair, uniformly adjusted to image blocks of a fixed size, and the image pixel values are normalized to the interval [0, 1] to enhance the stability of model training. Then, to ensure the quality of the training data, low-quality samples such as excessive blurring, mark missing, or severe background interference during the imaging process are removed; and data augmentation techniques such as geometric transformation, random noise addition, and brightness and contrast adjustment are used to expand the dataset, increase the diversity of samples, and improve the generalization ability of the model. During the augmentation process, it is necessary to ensure that the enhanced image pair still maintains the registration relationship. Finally, the preprocessed dataset is divided into a training set, a validation set, and a test set according to a ratio. 70% is used as the training set for model optimization; 15% is used as the validation set for parameter tuning; 15% is used as the test set for the final performance evaluation;

[0044] Step 3: Use a high-precision microscopic imaging device to collect the template image and the registration image. The template image refers to the image before exposure, while the registration image is the image after exposure. The collected images cover the inner contour marks on the wafer and the outer contour marks on the mask.

[0045] Step 4: Perform adaptive histogram equalization on the collected registration mark image to enhance the contrast of the image exposure area, specifically as follows:

[0046] (1) Convert the color image to a grayscale image:

[0047] I gray (x, y) = 0.299×R(x, y) + 0.587×G(x, y) + 0.114×B(x, y)

[0048] In the formula, R(x, y), G(x, y), and B(x, y) are the red, green, and blue components of the original image at the position (x, y), respectively, and I gray (x, y) is the grayscale image;

[0049] (2) Apply Gaussian blur to the grayscale image to reduce noise and weaken its high-contrast regions. Gaussian blur is a linear smoothing filter that can be implemented through convolution operations. For the grayscale image I gray (x, y), the Gaussian blur operation can be expressed as:

[0050]

[0051] where (i, j) traverses the set of neighborhood pixels centered on (x, y). The size of the Gaussian kernel (i.e., the size of the neighborhood) determines the range of smoothing. By experimenting with different kernel sizes and evaluating the results, the most suitable kernel size is determined to balance the requirements of noise reduction and detail preservation. The standard deviation σ controls the degree of blurring. A larger σ will result in a stronger blurring effect. Through these steps, the noise in the image can be effectively reduced and the high-contrast regions can be weakened, making the image smoother and more uniform;

[0052] (3) After dividing the grayscale image into multiple non-overlapping sub-image blocks, calculate the histogram H(i) = ∑ x,y δ(I gray (x,y) - i) for each sub-image block to obtain the frequency H(i) of the gray level, and then calculate the cumulative distribution function Finally, the equalized pixel value can be calculated where M and N are the width and height of the sub-image block, and L is the number of gray levels, which is set to 256 here. CDF min is the minimum non-zero value of the CDF. Before applying histogram equalization, the histogram needs to be clipped to limit the contrast, which is achieved by setting the maximum frequency threshold TTT for each gray level in the histogram. For gray levels exceeding the threshold TTT, their frequencies are distributed to other gray levels. Finally, bilinear interpolation is performed on each sub-image block to eliminate the block effect. Then, all enhanced sub-image blocks are recombined into a complete image;

[0053] Step Five: Perform deep learning-based alignment image denoising and blind image super-resolution reconstruction. After performing lithography alignment visual image adaptive histogram equalization processing in Step Four, construct a lithography alignment image super-resolution generative adversarial network (AlignGAN) based on a three-dimensional DenseNet fused with SE blocks (Squeeze-and-Excitation) for image super-resolution reconstruction:

[0054] (1) The generator network uses Dense-SENet, which is a three-dimensional DenseNet fused with SE blocks, as the basic unit. By adding a transition layer and fusing it with the SE block after the three-dimensional convolutional layer in the dense block. In the three-dimensional dense block, due to the use of dense connections, any layer can access the feature maps of the previous layers 0, 1, …, l - 1 within its dense block. The number of input feature maps of the l-th layer can be expressed as Among them, k0 is the number of channels in the network input layer, k is the growth rate of the input feature map, and the growth rate k significantly affects the number of input feature maps. Each layer within the dense block passes the feature maps generated by its own k convolutional kernels to the next layer, and the number of feature maps will increase as the number of layers in the dense block increases;

[0055] To increase the network training efficiency, it is necessary to reasonably limit the growth rate k to control the number of feature maps passed to the next layer. Therefore, a transition layer is added to reduce the dimension of the feature maps in the dense block to reduce the training difficulty of the network. The transition layer in the dense block mainly reduces the feature dimension through pooling operations. In addition, to alleviate the problem of computational consumption, before the three-dimensional feature maps are fed into the dense block, a bottleneck layer is further added for feature map dimensionality reduction processing, significantly reducing the feature map dimension and improving the training efficiency. The improved three-dimensional DenseNet structure is as Figure 3 shown;

[0056] To enable the network to fully learn image features, a squeeze-and-excitation (SE) block is introduced in the three-dimensional DenseNet. The structure of the SE layer is as Figure 4 shown. The working principle of SENet is as follows: First, the spatial information of the feature map is compressed into a channel information descriptor through global average pooling operation; then, a gate mechanism composed of two fully connected layers is used to obtain the channel relationship; finally, the information in the channels is weighted to achieve recalibration of the original features in the channel dimension. The convolutional output u is compressed through the spatial dimension H×W to generate the channel descriptor Z, and the C-th element Z C is expressed as:

[0057]

[0058] where u C is the feature map output by the c convolutional kernels. To obtain channel dependencies, the gating mechanism formed by two fully connected layers can be expressed as:

[0059] s = F ex (z, W) = σ(g(z, W)) = σ(W2δ(W1z))

[0060] where δ is the ReLU function, r is the reduction rate of the feature dimension. The output of SENet is weighted to the original features and expressed as:

[0061]

[0062] Through the above analysis, SENet first compresses the input, then excites it, maps the feature map to the global real numbers of the channel unit, and finally multiplies the real numbers with the corresponding input to complete the adaptive learning of each channel of the feature map;

[0063] In the embodiments of the present invention, since a large number of feature maps are generated by the dense blocks and transition layers in the network, the SE block after the dense block is fused with the non-linear transformation function of the 3D DenseNet. Considering the characteristics of the SE block, the SE block is inserted into the transmission layer of the 3D-DenseNet, which can filter out some unimportant features, squeeze and stimulate the effective features of the transmission layer. Therefore, introducing the SE block into the 3D-DenseNet can not only improve the utilization rate of effective features, but also reduce the influence of network redundancy. The global average pooling is used as the squeezing operation, and then two fully connected layers are applied to establish the correlation between channels and output weights with the same number as the input features. First, the feature dimension is reduced to 1 / 16 of the input, then it is activated by ReLU, and then it returns to the original dimension through another FC layer. Using two FC layers can not only have more non-linearity to better fit the complex correlation between channels, but also greatly reduce the number of parameters and the amount of calculation. Then, the normalized weights between are obtained through the sigmoid gate, and finally the normalized weights are weighted to the features of each channel through scaling;

[0064] (2) The discriminator network structure is as Figure 2 shown. The discriminator network is mainly composed of basic units stacked by three-dimensional convolution, Dense layer, normalization processing (NB) and LeakyReLU. The project constructs a relative discriminator to predict the probability that a real image is relatively more real than a fake image. The relative discriminator can be described as where x r and x f are the real image and the fake image in the training data respectively. Ex r [·] takes the average of all fake images in the mini-batch. C(x) is the output of the discriminator, and σ() is the Sigmoid function.

[0065] The discriminator loss function is:

[0066] The adversarial loss function is:

[0067] where, x f = G(x i ), x i is the input low-resolution (LR) image. It can be seen that the generator adversarial loss function contains x r and x f data, and more detailed features of the generated alignment mark image and the real alignment module image can be obtained;

[0068] ​Step 6: After the noise reduction and enhancement based on the deep adversarial network introduced above, start the template feature matching. First, determine a region of interest in the target image and perform template matching only within this region to improve the matching efficiency and accuracy. Calculate the normalized correlation coefficient between the template and the target image, find the position with the highest matching score, extract the matching feature point pairs in the template image and the target image, use the RANSAC algorithm to estimate the geometric transformation matrix, obtain the center positioning of the mark and correct the position of the target image to ensure that the alignment marks on the mask and the wafer are accurately aligned. The effect of template matching and correction can be verified by visualizing the binary alignment mark contour, thus achieving high-precision lithography alignment. The motion control system calculates the alignment motion data based on the center positioning data of the mark, and the system sends a motion signal to the motion system according to the alignment motion data to make the outer contour mark and the inner contour mark coincide, and then achieve precise overlay.

[0069] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.

[0070] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A lithography alignment image enhancement method based on a super-resolution generative adversarial network, characterized in that It includes the following steps: S1: Establishment of the AlignNet-SR dataset. The dataset sources are high-resolution overlay mark images covering various overlay mark designs and wafer process conditions, as well as corresponding low-resolution simulated images; S2: Preprocessing of the dataset, including a series of operations such as cropping and normalizing the dataset, data cleaning and augmentation, and dataset partitioning; S3: Acquisition of the template image and the overlay image. The image acquisition system acquires the template image before exposure and the overlay image after each exposure; S4: Establish an adaptive histogram equalization method to process the acquired overlay pattern area. This technology can effectively eliminate the square overlay marks on the wafer after exposure by adjusting the local area brightness and contrast, significantly improving the quality of the overlay image, thus making the subsequent multiple exposure process more accurate and reliable; S5: After S4, construct a lithography alignment image super-resolution generative adversarial network (AlignGAN) based on a three-dimensional DenseNet fused with a SE block (Squeeze-and-Excitation). This network can perform deep learning-based alignment image denoising and blind image super-resolution reconstruction. The introduced compression excitation SE block can complete the adaptive learning of each channel of the feature map, which can not only improve the utilization rate of effective features but also reduce the impact of network redundancy; S6: Template feature matching. After denoising and enhancement of the overlay marks based on the deep adversarial network in S5, use the cv.matchTemplate() function of OpenCV for template matching. After obtaining the best matching position, calculate its center positioning data. Finally, the motion control system realizes precise overlay alignment through the center positioning data of the mark.

2. The lithography alignment image enhancement method based on a super-resolution generative adversarial network according to claim 1, characterized in that: In S1, the overlay mark images include outer contour marks and inner contour marks. Among them, the outer contour marks are set on the mask plate, and the inner contour marks are set on the wafer. The low-resolution simulated images are generated by applying lithography blurring such as downsampling, Gaussian blurring, and adding noise to the high-resolution images.

3. The lithography alignment image enhancement method based on a super-resolution generative adversarial network according to claim 1, characterized in that: The preprocessing operations of the dataset in S2 specifically include: First, crop the region of interest (ROI) containing the overlay marks from the original image pairs and unify the size of the image blocks and normalize the pixel values of the images. Then, after data cleaning, use data augmentation techniques such as geometric transformation, noise addition, brightness and contrast adjustment to expand the dataset for the obtained image pairs. Finally, divide the preprocessed dataset into a training set, a validation set, and a test set according to a certain proportion.

4. A lithography alignment image enhancement method based on a super-resolution generative adversarial network according to claim 1, characterized in that: In S4, the adaptive histogram equalization technique is applied to weaken or eliminate the square alignment marks printed on the wafer image after exposure. Specifically, it includes: first converting the color image into a grayscale image; then applying Gaussian blur to denoise the grayscale image; after dividing the grayscale image into multiple non-overlapping and equal-sized sub-image blocks, calculating the histogram of each sub-image block to obtain the frequency of gray levels, and then calculating the cumulative distribution function to calculate the equalized pixel values; finally, performing bilinear interpolation on each sub-image block to eliminate block effects, and then recombining all the enhanced sub-image blocks into a complete image.

5. A lithography alignment image enhancement method based on a super-resolution generative adversarial network according to claim 1, characterized in that: In S5, a lithography alignment image super-resolution generative adversarial network (AlignGAN) based on the fusion of three-dimensional DenseNet and SE block (Squeeze-and-Excitation) is introduced. Its training data is sourced from the AlignNet-SR dataset, which includes processed high-resolution registration mark images and their corresponding low-resolution simulated images. After initializing the parameters of the generator G and discriminator D, defining the loss function and optimizer, the generator G and discriminator D are alternately trained to obtain the reconstructed super-resolution image. The specific training steps are as follows: The generator network takes the generated low-resolution simulated image as input, performs preliminary feature extraction on it through the initial three-dimensional convolutional layer, and uses the Leaky ReLU activation function to enhance its non-linear expression ability. When the image is input into the Dense-SENet unit composed of three-dimensional DenseNet and SE block (Squeeze-and-Excitation) (the number of which is set to 16), first, the dimensionality and number of channels of the feature map are reduced through the bottleneck layer 1x1 convolution in the dense block, then the spatial features are extracted using three-dimensional convolution, and the SE block is added after the dense block for channel calibration, enhancing important features through weighting and suppressing unimportant features. The transition layer generates global features through pooling operations. Each Dense-SENet unit contains multiple three-dimensional convolutional layers and can access the feature maps of all previous layers through dense connections. After stacking multiple units, the dimensionality is further reduced through the bottleneck layer, and finally, a super-resolution image is generated through two upsampling layers and a three-dimensional convolutional layer. The discriminator network consists of basic units composed of stacked three-dimensional convolutional layers, Leaky ReLU activation functions, normalization layers (Normal Block), and Dense layers. First, the real high-resolution template image and the generated high-resolution image are input, and the images are converted to grayscale images to simplify the calculation. Then, the preliminary and high-level features of the images are extracted through four 3D convolutional layers, and the non-linear expression ability is enhanced through the Leaky ReLU activation function to avoid the problem of neuron death. After the convolutional layer, the batch normalization layer (Batch Normalization) is applied to standardize the features, reduce the covariance shift, and stabilize and accelerate the training process. Next, the fully connected layer Dense(1024) is used to flatten the extracted features and map them to a feature space with a fixed dimension, and a scalar is output through the final fully connected layer. This scalar is converted into a probability value between 0 and 1 through the Sigmoid activation function, representing the likelihood that the image is a real image. Finally, the discriminant loss is calculated through the cross-entropy loss function, and the parameters of the discriminator are updated using the backpropagation algorithm to improve its ability to distinguish real images and generated images.

6. The lithography alignment image enhancement method based on a super-resolution generative adversarial network according to claim 1, characterized in that: In S6, the template matching algorithm (cv.matchTemplate()) of OpenCV is used to match the super-resolution image obtained in S5 with the average template generated from the marker image obtained by mask overlay. The matching is performed within the region of interest (ROI) of the target image. The normalized correlation coefficient between the template and the target image is calculated to find the position with the highest matching score. By extracting the matching points, the RANSAC algorithm is used to estimate the geometric transformation matrix, obtaining the center positioning of the marker and correcting the position of the target image to ensure that the mask and the alignment markers on the wafer are accurately aligned. The effect of template matching and correction can be verified by visualizing the binary alignment marker contours, thus achieving high-precision lithography alignment. The system sends motion signals to the motion system according to the alignment motion data, making the outer contour marker and the inner contour marker coincide, and then achieving precise overlay.

Citation Information

Patent Citations

  • Image demosaicing method based on super-resolution generative adversarial network SRGAN

    CN114972073A

  • Liver and liver tumor segmentation method based on parallel residual attention

    CN116883429A

  • Underwater image super-resolution reconstruction method based on generative adversarial network

    CN117853342A

  • Space image retrieval system and method based on multi-source feature fusion

    CN118708751A