A wafer feature module detection method based on attention mechanism
By introducing an attention mechanism-based detection method in wafer feature module detection, a multi-deep expansion convolution module and a dynamic dimension fusion module are used, combined with multi-level information aggregation and attention mechanism, the problem of scarcity of sample numbers and the influence of complex morphology is solved, and a high-precision and robust feature module detection is achieved.
Patent Information
- Application Number
- CN202510142955.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-10
AI Technical Summary
In the detection of wafer feature modules, the model training effect is poor due to the scarce sample number and poor sample quality in the detection of wafer feature modules, the feature modules have complex and diverse shapes, and it is difficult to identify stains, defects, occlusions or light source changes.
The wafer feature module detection method based on attention mechanism is adopted. By adding a multi-deep expansion convolution module and a dynamic dimension fusion module in the feature extraction process, combining multi-level information aggregation and attention mechanism, the accuracy and robustness of the detection are improved.
In the case of scarce sample size, the influence of insufficient sample size on model accuracy is overcome through meta-learning, the accuracy and robustness of feature module detection are improved, and the adaptability to complex scenarios is enhanced.
Smart Images

Figure CN119600308B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image recognition technology, and more particularly to a method for detecting wafer feature modules based on an attention mechanism. Background Art
[0002] Wafers are the core material of the semiconductor industry and the basis for manufacturing integrated circuits and other microelectronic devices. Wafers are usually made of high-purity silicon. After cutting, polishing and a series of complex process processing, they become the key carrier for chip manufacturing. The quality of wafers directly affects chip performance. The application of wafers has long been deeply integrated into life. For example, in modern electronic devices, from smartphones, computers to self-driving cars, almost all core computing and storage functions rely on semiconductor chips, and the starting point of these chips is wafers. It can be said that the quality of wafers determines the performance, stability and reliability of terminal electronic products.
[0003] The conventional wafer feature module detection and recognition process includes identifying the feature module, detecting specific patterns in the feature module, and measuring the boundary length information and area. Identifying the feature module is to identify the part information containing the feature module from the wafer image under the microscope.
[0004] The introduction of deep learning methods meets the urgent needs of modern semiconductor manufacturing for high precision and high efficiency. However, due to the small number of feature modules in the wafer, the number of samples available for learning is small, and the model training effect is not good. In addition, the complex and diverse shapes of feature modules, as well as the impact of stains, defects, occlusion or light source changes on the wafer, have caused a certain degree of deformation in the color and scale of the feature modules in the wafer image, resulting in unclear edges, difficulty in identification, and even affecting the subsequent boundary detection work. Summary of the invention
[0005] In view of the shortcomings of the prior art, the present invention proposes a wafer feature module detection method based on an attention mechanism. By adding a multi-depth dilated convolution module and a dynamic dimension fusion module in the feature extraction process, the model complexity and positioning accuracy are balanced, and multi-level information aggregation is combined with the attention mechanism to improve the accuracy and robustness of detection. It has advantages in dealing with the problems of scarce sample quantity and poor sample quality.
[0006] A wafer feature module detection method based on an attention mechanism specifically includes the following steps:
[0007] Step 1: Data preparation
[0008] Collect wafer images taken under different light sources as training samples, annotate the wafer feature modules in the images, save the location and category information of the feature modules as training labels, and form a training set. The training set is randomly divided into a support set and a query set. The support set has the same category as the training samples in the query set.
[0009] Step 2: Feature extraction
[0010] First, the backbone network is used to extract features from the support set and query set, respectively, to obtain feature maps X and Y. A set of convolution kernels are generated using feature map X, and deep convolution is performed on feature map Y. Then, the prediction box of the query set is generated through the RPN network, and finally, the training labels of the support set are used to discriminate the generated prediction boxes.
[0011] The backbone network first performs universal convolution to obtain a primary feature map, then performs multiple consecutive multi-depth dilated convolutions and maximum pooling operations to obtain features at different levels, fuses features at different levels in dimension, and finally performs multi-depth dilated convolutions and universal convolutions on the fused features to output a feature map.
[0012] Step 3: Model training
[0013] Constructing the loss function The network parameters in step 2. The loss function Includes the matching degree between the annotation box of the support set and the prediction box of the query set , and the regression loss of the predicted box of the query set :
[0014]
[0015] in, represents the loss weight.
[0016] Step 4: Model testing
[0017] The wafer image whose feature module position is to be detected is used as the query set and input into the model trained in step 3 together with the support set in step 1 to obtain the prediction result of whether the query set wafer image has the feature module and the feature module position.
[0018] The present invention has the following beneficial effects:
[0019] In the case of scarce feature module samples, meta-learning is used to overcome the impact of insufficient sample quantity on model accuracy. Through multi-depth dilated convolution and dynamic dimension fusion in the backbone network, the extracted features contain more sample details, greatly improving the detection accuracy of subsequent feature modules. And MARPN greatly improves the ability to predict feature modules in query sets. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 This is a flow chart of the wafer feature module detection method based on the attention mechanism;
[0021] Figure 2 A schematic diagram of the backbone network structure in the embodiment;
[0022] Figure 3 Schematic diagram of multi-depth dilated convolution in the embodiment;
[0023] Figure 4 This is a schematic diagram of dynamic dimension fusion in the embodiment;
[0024] Figure 5 This is a schematic diagram of the MARPN module structure in the embodiment;
[0025] Figure 6 Schematic diagram of DAMA attention mechanism in the embodiment;
[0026] Figure 7 Schematic diagram of ROI pooling in the embodiment;
[0027] Figure 8 Schematic diagram of a global relationship classifier GRC module in an embodiment;
[0028] Fig. 9 Detection results of the wafer image feature module in different scenarios in the embodiments. DETAILED DESCRIPTION
[0029] The present invention will be further explained below with reference to the accompanying drawings;
[0030] A wafer feature module detection method based on an attention mechanism specifically comprises the following steps:
[0031] Step 1: Data preparation
[0032] Collect wafer images taken under different light sources, mark the wafer feature modules in the image with rectangular frames, and save the coordinate information of the rectangular frames, including the horizontal and vertical coordinates of the upper left corner ( ) and the horizontal and vertical coordinates of the lower right corner ( ). According to the type of chip, 60 categories of wafer images are randomly selected as training sets, and the remaining 12 categories of images are used as test sets. For the training set data, it is randomly divided into a query set and a support set, and the query set and the support set both include 60 categories of wafer images.
[0033] Step 2: Feature extraction
[0034] like Figure 1As shown in the figure, firstly, the backbone network is used to extract features from the images of the support set and the query set, respectively, to obtain feature maps X and Y. At the same time, the convolution kernel generated by the feature map X is used to perform deep convolution on the feature map Y, and then the prediction box of the query set is generated by the RPN network. Finally, the rectangular box annotation information of the wafer feature module in the support set image is used to discriminate the prediction box of the query set.
[0035] s2.1. Use two feature extraction networks with the same structure as the backbone network to extract features from images in the support set and query set respectively.
[0036] like Figure 2 As shown in the figure, for the input image, the backbone network first obtains the primary feature map through a general convolution module DBL, and then obtains the low-dimensional feature map through three consecutive multi-depth expansion convolution modules MDCM and maximum pooling operations. , middle layer features and high-dimensional features , and then , and As the input of the dynamic dimension fusion module DDFM, feature fusion is performed to retain more feature information, and then a multi-feature high-dimensional feature map is obtained through the multi-depth dilated convolution module MDCM and maximum pooling. Finally, it is compressed through the general convolution module to obtain the feature map X for the image of the support set and the feature map Y for the image of the query set.
[0037] Dilated convolution can increase the receptive field with the minimum number of parameters, so it is widely used in target detection related tasks. Figure 3 In the multi-depth dilated convolution module MDCM shown in FIG. 1 , multiple depth-separable convolution layers are used to capture spatial features of various receptive field sizes at different expansion rates, thereby modeling the difference between the wafer feature module and the background in more detail and enhancing the backbone network's ability to distinguish small objects. ,The multi-depth dilated convolution module first divides it into four different heads along the channel to generate , and then perform depth-separable dilated convolution on each head with different dilation rates to obtain :
[0038]
[0039] Where i=1,2,3,4, H, W, C represent the height, width and number of channels of the input feature map of the multi-depth dilated convolution module respectively. DDWConv() represents depthwise separable dilated convolution. represents ReLU activation, Represents a normalization layer.
[0040] Then Perform channel segmentation and reorganization, staggered arrangement to form a variety of feature maps :
[0041]
[0042] in, express The jth channel of , j = 1, 2, ..., C / 4. Then the feature map Perform point-by-point convolution and splicing, and then perform one more point convolution to output the feature map :
[0043]
[0044] in, and Represents the point-wise convolution weight matrix.
[0045] Since the original samples of wafer feature modules are scarce and the area ratio of feature modules in the entire wafer image may be small, as the network depth continues to deepen, after multiple downsampling stages, high-dimensional features may lose information about small targets, making it impossible for low-dimensional features to provide sufficient background information. Therefore, in the feature extraction process, it is necessary to continuously fuse features at different levels to try to include feature module information in high-dimensional features. Figure 4 The dynamic dimension fusion module DDFM shown in the figure first transforms the high-dimensional features into , low-dimensional features With the middle layer features Perform preliminary alignment, and then transform , and Divided into 4 equal segments , and , use the Sigmoid function to calculate the intermediate features heavy ,according to Adjustment and The ratio of is merged to obtain Finally, the 4 After splicing, convolution is performed to obtain fusion features :
[0046]
[0047]
[0048]
[0049] Among them, k=1,2,3,4, represents convolution, Represents the Sigmoid function.
[0050] s2.2. The existing target detection task generates candidate boxes for all possible targets through the RPN module, and then performs classification and regression. However, RPN will activate all high-scoring objects, including corresponding categories that do not belong to the support set, without any supporting image information, which adds a lot of classification workload to the detector. Therefore, in the candidate box selection stage, it is not only hoped to filter out background boxes, but also to filter out candidate boxes that do not belong to the categories included in the support set. That is, the image features corresponding to the candidate boxes generated based on the query set image should be as similar as possible to the image data features in the support set, so as to focus more on generating specific types of candidate boxes.
[0051] Based on the RPN network, this application proposes the following Figure 5 The MARPN module shown in the figure first performs deep convolution on the feature map X of the support set image, and then generates a set of convolution kernels through the SCA module and global average pooling GAP. The generated convolution kernels are used to perform deep convolution on the feature map Y of the query set image, and then a convolution operation is performed to generate the feature map FP. FP is sent to the RPN network to obtain the prediction box of the query set image.
[0052] The SCA module first passes the input features through the deep convolution DWConv residual block to achieve parameter sharing and enhance the learning of local features. Then it performs normalization, processes in parallel through the two attention mechanisms of DAMA and CA, and then performs normalization after fusion. Finally, the processing results are output through the MLP layer:
[0053]
[0054]
[0055]
[0056] in, represents the output characteristics of the SCA module, , It is an intermediate feature of the SCA module. represents depthwise separable convolution, LN( ) represents layer normalization, MLP( ) represents multilayer perceptron, , They represent the CA attention mechanism and DAMA attention mechanism respectively.
[0057] The DAMA attention mechanism module is based on the self-attention mechanism and introduces multi-layer activation functions and matrix interaction operations, so that the module can capture more complex context dependencies and feature relationships. It also alleviates the gradient vanishing problem in deep networks through residual connections, while retaining the low-order information of input features. Combined with the CA module, it can capture accurate position features and information between channels, such as Figure 6 As shown:
[0058]
[0059]
[0060]
[0061]
[0062]
[0063] Among them, Q, K, and V represent the query matrix, key matrix, and value matrix. represents linear transformation, FC( ) represents full connection, represents the activation function, represents the scaling factor, represents matrix multiplication, Represents the features enhanced by the DAMA attention mechanism module.
[0064] s2.3. According to the size ratio of the input image in the training set and the feature map extracted by the backbone network, the position information of the predicted box of the query set image generated by MARPN and the real rectangular annotation box of the support set image are Map to , thereby extracting the information of the corresponding position in the feature map and obtaining the region of interest R:
[0065]
[0066]
[0067]
[0068]
[0069]
[0070]
[0071]
[0072]
[0073]
[0074] in, , denote the width and height of the support set image, respectively. , Respectively represent the width and height of the backbone network output feature map X, Y, Represents the features of the feature map in the candidate region, represents the boundary coordinates of the grid cell, , Represents the row index and column index of the grid cell, such as Figure 7 shown.
[0075] Since the sizes of the regions of interest vary and cannot meet the network requirements, RoIPooling is needed to resize the regions of interest. The size of the region of interest is adjusted. Divide into Grid , then for each grid Perform maximum pooling to obtain a feature map of fixed size.
[0076] s2.4, judging the pre-selected box is the last step of the entire network. The global relation classifier GRC module is used to evaluate and adjust whether the pre-selected box generated by the MARPN module meets the requirements. Figure 8 As shown,
[0077] The global relation classifier GRC module first concatenates the features with a channel number C obtained by inputting X and Y after RoI Pooling into a feature with a channel number The feature vector of , and then the concatenated features are averagely pooled into size, and then use two layers of fully connected layers with ReLU activation for processing. Finally, the matching probability between the predicted box of the query set and the labeled box of the support set, as well as the predicted box coordinate information of the query set are obtained.
[0078] Step 3: Model training
[0079] Through the loss function Optimize the network model parameters constructed in step 2, the loss function Includes the matching degree between the annotation box of the support set and the prediction box of the query set , and the regression loss of the predicted box of the query set :
[0080]
[0081] in, Represents the loss weight, which is set to 1 in this embodiment. Using binary cross entropy loss function, The smooth L1 loss function is used:
[0082]
[0083]
[0084]
[0085]
[0086] Among them, N is the total number of prediction boxes, is the matching probability between the predicted box of the z-th query set and the labeled box of the support set. , They represent the predicted box coordinate information of the query set and the labeled box coordinate information of the support set respectively.
[0087] Step 4: Model testing
[0088] The wafer images with known feature module positions are taken as the support set, and the wafer images with feature module positions to be detected are taken as the query set. They are input into the model trained in step 3 to obtain the prediction results of whether the query set wafer images have feature modules and the feature module positions.
[0089] Fig. 9 The results of wafer feature module detection in different scenarios by this method are shown. As can be seen from the figure, under normal circumstances, that is, when there are no defects in the wafer feature module or the background is blurred, this method can detect the position information of each new type of feature module in the wafer image while ensuring accuracy. In the case of differences in background color of wafer images caused by different exposure conditions, this method can accurately detect the position of the feature module, and after the overall transformation of the lines caused by shooting with different microscopes, the feature module can still be accurately identified and marked, indicating that the model trained by this method has abandoned the background noise to a certain extent, and can complete accurate identification under the conditions of different exposure scenes and microscopes of different brands. For complex scenes, including defective occlusion of the wafer feature module, or serious external influences on the light source scene where the microscope is located, resulting in abnormal halos, this application can still effectively locate the position of the feature module.
[0090] In addition, AP, AP50 and AP75 are used as indicators to evaluate the performance of this method and other mainstream object detection algorithms. The results are shown in Table 1:
[0091] Table 1
[0092] Method Name AP AP50 AP75 Faster R-CNN 67.4 74.4 66.1 Meta R-CNN 72.3 79.2 68.2 MSPR 70.6 77.2 67.9 TFA 69.4 78.4 68.7 FSCE 74.4 82.1 72.7 This method 76.2 83.4 73.9
[0093] AP (Average Precision) is the most important comprehensive evaluation indicator in the target detection task, which can fully reflect the performance of the model under different IoU thresholds. At the same time, combined with AP50 (IOU threshold is 0.5) and AP75 (IOU threshold is 0.75), the performance of the model in two levels, coarse-grained target detection and fine-grained target positioning, can be analyzed. Through these indicators, the applicability of the model in different tasks can be better evaluated.
[0094] Experimental data show that this method shows excellent performance in the wafer feature module detection task, with an AP of 76.2, which is significantly improved compared with existing methods. At the same time, this method has a strong advantage in coarse-grained target detection, with an AP75 of 73.9 at a high IoU threshold, reflecting its superiority in precise target positioning. This shows that this method has better robustness and generalization capabilities while taking into account the comprehensiveness and refinement of target detection.
Claims
1. A wafer feature module detection method based on an attention mechanism, which collects wafer images taken under different light sources as training samples, annotates the information of the wafer feature modules in the images as training labels; based on a meta-learning method, randomly divides the training set into a support set and a query set; the characteristics are: Step 1: Through the backbone network, feature maps X and Y are extracted from the images of the support set and query set respectively; The backbone network first obtains the primary feature map of the image through a general convolution module DBL, and then obtains the low-dimensional feature F through three consecutive multi-depth dilated convolution modules MDCM and maximum pooling operations. l , intermediate layer features F u and high-dimensional features F h Then, the dynamic dimension fusion module DDFM is used to transform F l 、F u and F h Perform feature fusion, and then pass the fused features through the multi-depth dilated convolution module MDCM and maximum pooling to obtain a multi-feature high-dimensional feature map, and finally compress it through the general convolution module to output the feature map; Step 2: Perform deep convolution on the feature map X, and then process it through the parallel self-attention mechanism and coordinated attention mechanism to generate a set of convolution kernels for deep convolution on the feature map Y, and then generate the prediction box of the query set through the RPN network; Step 3: Use the training labels to identify the generated prediction boxes and calculate the loss function value Complete the optimization of model parameters; Step 4: The wafer image whose feature module position is to be detected is used as a query set, and is input into the model trained in step 3 together with the support set to obtain the prediction result of whether the wafer image in the query set has a feature module and the position of the feature module.
2. A wafer feature module detection method based on an attention mechanism as claimed in claim 1, characterized in that: The multi-depth dilated convolution module MDCM is used for the input feature map First, divide it into four different heads along the channel: a1, a2, a3, Then, each head is subjected to depth-separable dilated convolution at different dilation rates to obtain a′ i =δ(B(DDWConv(a i ))); Where i = 1, 2, 3, 4, H, W, C represent the input feature map F respectively. a The height, width, and number of channels of ; DDWConv() represents depthwise separable dilated convolution, δ() represents ReLU activation, and B() represents normalization layer; Then for a′ i Perform channel segmentation and reorganization, staggered arrangement to form a diverse feature map h j : h j =W inner ([a′ 1j ,a′ 2j ,a′ 3j ,a′ 4j ]); in, Represents a′ i The jth channel of j Perform point-by-point convolution and splicing, and then perform one more point-wise convolution to output the feature map F o : Among them, W inner and W outer Represents the point-wise convolution weight matrix.
3. A wafer feature module detection method based on an attention mechanism as claimed in claim 1, characterized in that: The dynamic dimension fusion module DDFM first transforms the high-dimensional feature F h , low-dimensional features F l With the middle layer feature F u Perform preliminary alignment, and then transform F l 、F u and F h Divide into 4 equal segments l k 、u k and h k , and then use the Sigmoid function to calculate the intermediate feature u k The weight α k : α k =Sigmoid(u k ); According to α k Adjust k and h k The ratio of is integrated to obtain u′ k : u′ k =a k l k +(1-a k )h k ; Finally, the 4 u′ k After splicing, convolution is performed to obtain fusion features Among them, k = 1, 2, 3, 4, Conv( ) represents convolution, and Sigmoid( ) represents Sigmoid function.
4. A wafer feature module detection method based on an attention mechanism as claimed in claim 1, characterized in that: Perform deep convolution on the feature map X of the support set image, and then generate a set of convolution kernels through the SCA module and global average pooling GAP. Use the generated convolution kernels to perform deep convolution on the feature map Y of the query set image, and then perform another convolution operation to generate the feature map FP. Send FP to the RPN network to obtain the prediction box of the query set image; The SCA module first passes the input feature X through the deep convolution DWConv residual block to obtain the feature map X1: X1=X+DDWConv(X); After normalizing the feature map X1, it is fused through the parallel self-attention mechanism and coordinated attention mechanism to obtain the feature map X2: X2=CA(LN(X1))+DAMA(LN(X1)); After normalizing the feature map, it is input into the MLP layer and concatenated with the feature map X1 to output the feature map X′: X′=MLP(LN(X2))+X1; Among them, DDWConv( ) represents depthwise separable convolution, LN( ) represents layer normalization, MLP( ) represents multilayer perceptron, CA( ) and DAMA( ) represent coordinated attention mechanism and self-attention mechanism respectively.
5. A wafer feature module detection method based on an attention mechanism as claimed in claim 4, characterized in that: The self-attention mechanism DAMA ( ) generates the query matrix, key matrix and value matrix Q, K, V of the input feature LN (X1) through the fully connected layer, and transforms them into Q', K', V' through linear transformation. Through a series of activation and matrix interaction operations, the output feature map Among them, FC () represents full connection, SiLU () represents activation function, d represents scaling factor, Represents matrix multiplication.
6. A wafer feature module detection method based on an attention mechanism as claimed in claim 1, characterized in that: The information of the wafer feature module in the training label is mapped to the feature map X and the feature map output by the RPN network through ROI pooling to extract the region of interest.
7. A wafer feature module detection method based on an attention mechanism as claimed in claim 1, characterized in that: The loss function Includes the matching degree between the annotation box of the support set and the prediction box of the query set And the regression loss of the predicted box of the query set Among them, λ represents the loss weight.
8. A wafer feature module detection method based on an attention mechanism as claimed in claim 7, characterized in that: Set the loss weight λ = 1, Using binary cross entropy loss function, The smooth L1 loss function is used.
9. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Target detection method based on meta-learning combination attention mechanism network model
CN117576379A
Person re-identification method combining reverse attention and multi-scale deep supervision
US20210232813A1