Wafer body defect detection method based on vision transfer model

By applying the detection method based on the Vision Transformer model in wafer defect detection, the absolute difference method and L-ViT structure, combined with the improved multi-head self-attention mechanism, the problems of false detection and missed detection in the existing technology are solved, and defect detection with higher accuracy and reliability are achieved.

CN119991571AActive Publication Date: 2025-05-13SHENZHEN ZHIXIAN FUTURE IND SOFTWARE CO LTD

Patent Information

Application Number
CN202411970942.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-13
Estimated Expiration
2044-12-30

AI Technical Summary

Technical Problem

The prior art has problems of mis-detection and missed detection in wafer defect detection, especially when dealing with complex and diverse defect forms, it is difficult to meet high-precision requirements.

Method used

The wafer defect detection method based on the Vision Transformer (ViT) model is adopted to collect the top view of normal view angles and defective view angles, and the defect position is positioned using the absolute difference method, and the defect feature extraction and model detection capabilities are enhanced through the L-ViT structure and improved multi-head self-attention mechanism (p-MHA).

Benefits of technology

It improves the accuracy and reliability of wafer defect detection, reduces missed detection rates and false detection rates, can more effectively capture complex defect characteristics, and provides a solid data foundation for subsequent defect analysis and processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991571A_ABST
    Figure CN119991571A_ABST
Patent Text Reader

Abstract

The invention provides a wafer body defect detection method based on a vision transfer model, and belongs to the crossing field of semiconductor manufacturing and computer vision technologies. According to the method, a top view of a normal visual angle and a top view of a defective visual angle of a wafer body are collected, and a defect position is positioned through an absolute difference method. And information fusion is performed on the positioning area, and the defect features are amplified, so that the model learning is facilitated. An L-Vit structure is provided, and the structure converts data subjected to absolute difference processing into a vector sequence, and the vector sequence is combined with a vector sequence with a defective view angle to more highlight defect features. According to the method, a new multi-head self-attention mechanism p-MHA is provided, q and v associated formulas are added into the multi-head self-attention mechanism, and the main purpose is to enable an absolute difference sequence vector and a vector sequence with a defective view angle to be connected more closely, meanwhile, the data size of the multi-head attention mechanism is reduced, and the data size of the model is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the intersection field of semiconductor manufacturing and computer vision technology, and in particular relates to a wafer defect detection method based on a vision transformer model. Background Art

[0002] In the semiconductor manufacturing process, wafers are the basic carriers for the production of integrated circuits (ICs), and their quality and integrity are crucial to the performance and stability of subsequent devices. Since the wafer manufacturing process is complex and susceptible to external environmental influences, various defects such as cracks, scratches, and contamination are prone to appear on the wafer surface. These defects can cause circuit failure or performance degradation in subsequent processes, so it is crucial to detect and identify surface defects during wafer manufacturing and processing. However, with the increase in wafer size and the diversification of defect types, wafer defect detection is facing higher precision requirements and more complex challenges.

[0003] Traditional wafer defect detection methods mostly rely on algorithms based on image processing and manual feature extraction. These methods usually rely on a lot of experience and manually designed features, so they are prone to false detection and missed detection when faced with complex and diverse defect forms. In recent years, the rapid development of deep learning in the field of computer vision has brought new opportunities for wafer defect detection. Among them, convolutional neural networks (CNNs) have performed well in image classification and object detection, and have gradually been applied to wafer body defect detection tasks. However, convolutional networks have certain limitations in processing global dependencies and subtle features of images, especially in the detection of high-resolution, large-size wafer images, it is difficult to fully capture complex defect features.

[0004] The Vision Transformer (ViT) model is an emerging deep learning model that was originally developed based on the Transformer model in natural language processing. It uses the self-attention mechanism to effectively model the global features of the image. The ViT model divides the input image into a series of small patches and captures the relationship between the small patches through the self-attention mechanism, thereby learning the deep information of the image on a global scale. This feature makes ViT particularly suitable for image processing tasks with high resolution and complex scenes, and shows higher accuracy and flexibility in processing irregular and subtle defect features. Therefore, applying the ViT model to wafer defect detection can improve the detection accuracy of the model and reduce the missed detection rate and false detection rate.

[0005] The invention not only makes up for the deficiencies of the prior art, but also provides an effective solution for quality control in the semiconductor manufacturing process. Summary of the invention

[0006] The purpose of the present invention is to overcome the above-mentioned defects existing in the background technology, and to provide a wafer body defect detection method based on the vision transformer model, which collects top views of the wafer from a normal perspective and a defective perspective, and locates the defect position using the absolute difference method. The positioning area is upsampled to amplify the defect features to facilitate model learning. The proposed L-ViT structure converts the absolute difference data into a vector sequence, which is combined with the defective perspective vector sequence to highlight the defect features. The newly proposed p-MHA multi-head self-attention mechanism adds an association formula between Q and V to enhance the connection between the absolute difference sequence and the defect sequence, and reduce the amount of data and model complexity. The present invention adopts the following technical solutions to solve the above-mentioned technical problems:

[0007] A wafer defect detection method based on a vision transformer model comprises the following steps:

[0008] Step S1, collecting a top view of the wafer at a normal viewing angle and a top view of the wafer at a defective viewing angle, and converting the two into the same image format and size, using a bilinear interpolation algorithm, taking into account the information of four pixels around the target pixel to ensure that each pixel can correspond one to one;

[0009] Step S2, compare the pixels at the same coordinate position in the two images obtained in step S1, determine the threshold d comprehensively according to the wafer image characteristics, noise and minimum defect size, traverse the pixels of the image, obtain its RGB value and calculate the absolute difference in each channel with the pixels in the surrounding nine-square grid neighborhood. Once the absolute difference of at least one channel of a pixel exceeds the threshold, it is marked as a suspected defect point. Then, according to the expected defect size under 5x5 pixels, determine the expansion area with the point as the center, and mark the pixels in the area as defect positions through binary image marking;

[0010] Step S3, processing the pixel points of the defect position, and extracting the target position and target information of the defect position by using a method based on centroid matching; collecting a normal viewing angle picture of the wafer body, performing data preprocessing on the picture, and finding out the position where the defect position should be merged under the normal viewing angle through matrix similarity calculation, and performing feature fusion with the position of the defective feature;

[0011] Step S4, constructing an L-Vit structure based on the vision transformer, first extracting the position information of defects in the wafer body image, encoding it into a vector sequence, the sequence contains key data such as the defect center coordinates and bounding box parameters, and then performing preprocessing, using a normalization operation to adjust the value range, using a smoothing algorithm to reduce noise interference, and finally integrating the processed vector sequence into the vision transformer architecture;

[0012] Step S5: Use the improved multi-head self-attention mechanism p-MHA to calculate the defect features. First, the correlation between the current position and the actual position is obtained through interactive calculation of the fused sequence. Then, p-MHA preprocesses q and v before attention calculation, filters out the interference of noise and irrelevant information through interaction at key positions, and finally determines the position of q.

[0013] Further, in step S1, the bilinear interpolation algorithm is used. Given four adjacent pixel values I x1,y1 , I x1,y2 , I x2,y1 , I x2,y2 in the original image, where x1, x2 and y1, y2 are the positions of four pixels in the original image, and x1 < x2 and y1 < y2, and the target position is (x, y), which is located between the four known pixel points. The formula is:

[0014]

[0015]

[0016] Then, through vertical interpolation, interpolation is performed in the y-axis direction:

[0017]

[0018] Finally, the obtained I x,y is the interpolation result of the target position (x, y).

[0019] Further, in step S2, for the pixel points at the same coordinate position in two pictures, the absolute differences in each color channel are calculated respectively; the defect positions are located by collecting the top views of the wafer body from the normal perspective and the defective perspective and using the absolute difference method:

[0020] (R def (x, y), G def (x, y), B def (x, y))

[0021] (R ref (x, y), G ref (x, y), B ref (x, y))

[0022] d R (x, y) = |R def (x, y) - R ref (x, y)|

[0023] where R def (x, y) is the red pixel point of the top view of the wafer body collected from the normal perspective, G def(x, y) is the green pixel point of the top view of the wafer body at normal viewing angle, B def (x, y) is the red pixel and blue pixel of the top view of the normal viewing angle of the wafer, R ref (x, y) is the red pixel of the top view of the defective wafer, G def (x, y) is the red pixel point of the top view of the defective wafer, B def (x, y) is the blue pixel point of the top view of the defective wafer, d R (x, y) is the absolute difference between the pixel points corresponding to the top view of the wafer at a normal viewing angle and the top view of the wafer at a defective viewing angle. The defect position judgment standard is:

[0024]

[0025] Furthermore, in step S3, the target position and target information of the defect position are obtained by using a centroid-based matching method, and the specific steps are as follows:

[0026] Step S3-1, calculate the centroid coordinates of the defect position by weighted average method; the centroid coordinates are calculated by calculating the weighted average of the pixels in the x direction and the y direction respectively, with the weight being the gray value of each pixel, and the calculation formula is:

[0027]

[0028]

[0029] Where I(x, y) is the grayscale value of the pixel with coordinates (x, y), and ∑ represents the sum of the pixels within the entire defect location area;

[0030] Step S3-2, obtain the defect-free control image information, traverse each point in the defect-free control image or traverse each area through a sliding window, and calculate the Euclidean distance d between them and the centroid coordinates of the defect position; the calculation formula of the Euclidean distance is:

[0031]

[0032] Where (x1, y1) is the coordinate of the centroid of the defect area, and (x2, y2) is the coordinate of the point or area to be compared in the reference photo);

[0033] Step S3-3, using the particle swarm algorithm to find the point or area with the smallest distance to determine the corresponding position of the defect position in the defect-free control image, the objective function is the mean square error f, the formula is:

[0034]

[0035] Where P = (p x , p y ) is the particle position to be optimized, T(i, j) and D(P+(i, j)) are the pixel values ​​of the defect-free image and the defective image at position (i, j), respectively. The point with the shortest distance is found according to the minimum value of f to determine the defect position and the corresponding position in the defect-free control image.

[0036] Furthermore, in step S4, the L-Vit structure based on the vision transformer adds a vector sequence of defect positions on the basis of Vit, and expands the defect features by adding the preprocessed sequence. Let P = (x p ,y p ) represents the coordinates of the defect location, which are the center coordinates or other representative position coordinates. ij Represents the feature vector at position (i, j) in the image feature map, W pos is the position guidance weight matrix, the position guidance score S ij The calculation is as follows:

[0037]

[0038] Among them, ||(i, j)-P|| 2 Calculate the square of the Euclidean distance between the image feature position and the defect position. γ is a hyperparameter that controls the degree of influence of the distance. According to this position guidance score, the features are weighted and summed to obtain the feature vector F after position guidance. guided :

[0039]

[0040] By the position feature vector F guided Make features close to the defect location receive more attention and emphasis.

[0041] Furthermore, the improved multi-head self-attention mechanism p-MHA in step S5 first processes q and v on the fused sequence before the three parameters q, k, and v are input, and the position of q is more accurately determined by the interactive calculation of the current position q and the actual position v. p-MHA pre-processes q and v before attention calculation, and interacts only at key positions to accurately determine the position of q.

[0042] Furthermore, the steps of improving the multi-head self-attention mechanism p-MHA in step S5 are as follows:

[0043] Step S5-1: For the input query (q), (k) and value (v), the multi-head attention mechanism will first perform a linear transformation on them to obtain different head q, k, v versions. For each head i, there will be a corresponding linear transformation matrix Wqi , W ki , W vi So that:

[0044] q i =q*W qi

[0045] k i =k*W ki

[0046] v i =v*W vi

[0047] Where, i = 1, 2, 3, …, h;

[0048] Step S5-2: Calculate the attention score of each head i, where the attention score represents the query q i With key k j The degree of association between them; the attention score calculation formula is:

[0049]

[0050] where d k is the dimension of key k, and the softmax function is used to normalize the scores to a probability distribution form so that the sum of all possible association scores is 1;

[0051] Step S5-3: Calculate the output o of each head i according to the attention score i , that is, the attention score and the corresponding value v are weighted and summed, the formula is:

[0052]

[0053] Compared with the prior art, the present invention adopts the above technical solution and has the following beneficial effects:

[0054] (1) The present invention provides a wafer defect detection method based on a vision transformer model, which collects a top view of a wafer from a normal perspective and a top view from a defective perspective and locates the defect position by using an absolute difference method. By upsampling the location area, the defect features are amplified, which is more conducive to model learning.

[0055] (2) The present invention provides a wafer defect detection method based on the vision transformer model and proposes an L-Vit structure, which converts the data after absolute difference processing into a vector sequence and combines it with the vector sequence of the defective perspective to further highlight the characteristics of the defect.

[0056] (3) The present invention provides a wafer defect detection method based on the vision transformer model and proposes a new multi-head self-attention mechanism p-MHA, which reduces the data volume of the multi-head attention mechanism and reduces the data volume of the model.

[0057] (4) The present invention provides a wafer body defect detection method based on the vision transformer model and proposes a new multi-head self-attention mechanism p-MHA, so that when calculating the defect features, the model can focus on the key information closely related to the defects, greatly improving the accuracy and reliability of wafer body defect detection and providing a solid data foundation for subsequent defect analysis and processing.

[0058] (5) The present invention provides a wafer defect detection method based on a vision transformer model, which adopts a bilinear interpolation algorithm. When scaling the wafer image, it can make the color and brightness transition of the image more natural and smooth. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 A detection process block diagram of a wafer defect detection method based on a vision transformer model according to the present invention;

[0060] Figure 2 The new L-Vit structure based on vision transformer (vit) proposed by the present invention; DETAILED DESCRIPTION

[0061] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings of the specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in the art without creative work are within the scope of protection of the present invention.

[0062] Example:

[0063] like Figure 1 As shown, a wafer defect detection method based on a vision transformer model comprises the following steps:

[0064] Step S1, collecting a top view of the wafer at a normal viewing angle and a top view of the wafer at a defective viewing angle, and converting the two into the same image format and size, using a bilinear interpolation algorithm, taking into account the information of four pixels around the target pixel to ensure that each pixel can correspond one to one;

[0065] Step S2: Compare the pixel points at the same coordinate positions in the two images obtained in Step S1, comprehensively determine the threshold d1 based on the characteristics of the wafer image, noise, and the minimum defect size, traverse the pixels of the image, obtain their RGB values, and calculate the absolute differences of the pixels in the surrounding nine-grid neighborhood in each channel. Once the absolute difference of at least one channel of a certain pixel exceeds the threshold, it is marked as a suspected defect point. Then, according to the expected detection of the defect size under 5x5 pixels, determine the expansion area centered on this point, and mark the pixels in the area as the defect positions through binary image marking;

[0066] Step S3: Process the pixel points at the defect positions, and use the method based on centroid matching to extract the target positions and target information of the defect positions; collect the images of the normal view of the wafer body, perform data preprocessing on the images, and find the positions where the defect positions should be in the normal view through matrix similarity calculation for fusion, and perform feature fusion with the positions with defective features;

[0067] Step S4: Construct the L-Vit structure based on vision transformer. First, extract the position information of the defects in the wafer body image, encode it into a vector sequence, and the sequence contains key data such as the defect center coordinates and bounding box parameters. Then, perform preprocessing, adjust the numerical range using normalization operation, and reduce noise interference using the smoothing algorithm. Finally, integrate the processed vector sequence into the vision transformer architecture;

[0068] Step S5: Use the improved multi-head self-attention mechanism p-MHA to calculate the defect features. First, the fusion sequence undergoes interactive calculation to obtain the correlation between the current position and the actual position. Then, p-MHA preprocesses q and v before attention calculation, filters out noise and irrelevant information interference at key positions, and finally determines the position of q.

[0069] Further, in Step S1, the bilinear interpolation algorithm is used. Given four adjacent pixel values I x1,y1 ,I x1,y2 ,I x2,y1 ,I x2,y2 in the original image, where x1, x2 and y1, y2 are the positions of the four pixels in the original image, and x1 < x2 and y1 < y2, the target position is (x, y), and this position is between the four known pixel points. The formula is:

[0070]

[0071]

[0072] Then, through vertical interpolation, perform interpolation in the y-axis direction:

[0073]

[0074] The final I x,y It is the interpolation result of the target position (x, y).

[0075] Furthermore, in step S2, for the pixel points at the same coordinate position in the two images, the absolute difference in each color channel is calculated respectively; the defect position is located by collecting the top view of the wafer body at a normal viewing angle and the top view of the wafer body at a defective viewing angle and using the absolute difference method:

[0076] (R def (x,y),G def (x,y),B def (x, y)

[0077] (R ref (x,y),G ref (x,y),B ref (x, y)

[0078] d R (x, y) = |R def (x, y)-R ref (x, y)|

[0079] Where R def (x, y) is the red pixel point of the top view of the normal viewing angle of the wafer body, G def (x, y) is the green pixel point of the top view of the wafer body at normal viewing angle, B def (x, y) is the red pixel and blue pixel of the top view of the normal viewing angle of the wafer, R ref (x, y) is the red pixel of the top view of the defective wafer, G def (x, y) is the red pixel point of the top view of the defective wafer, B def (x, y) is the blue pixel point of the top view of the defective wafer, d R (x, y) is the absolute difference between the pixel points corresponding to the top view of the wafer at a normal viewing angle and the top view of the wafer at a defective viewing angle. The defect position judgment standard is:

[0080]

[0081] Furthermore, in step S3, the target position and target information of the defect position are obtained by using a centroid-based matching method, and the specific steps are as follows:

[0082] Step S3-1, calculate the centroid coordinates of the defect position by weighted average method; the centroid coordinates are calculated by calculating the weighted average of the pixels in the x direction and the y direction respectively, with the weight being the gray value of each pixel, and the calculation formula is:

[0083]

[0084]

[0085] Where I(x, y) is the grayscale value of the pixel with coordinates (x, y), and ∑ represents the sum of the pixels within the entire defect location area;

[0086] Step S3-2, obtain the defect-free control image information, traverse each point in the defect-free control image or traverse each area through a sliding window, and calculate the Euclidean distance d between them and the centroid coordinates of the defect position; the calculation formula of the Euclidean distance is:

[0087]

[0088] Where (x1, y1) is the coordinate of the centroid of the defect area, and (x2, y2) is the coordinate of the point or area to be compared in the reference photo);

[0089] Step S3-3, using the particle swarm algorithm to find the point or area with the smallest distance to determine the corresponding position of the defect position in the defect-free control image, the objective function is the mean square error f, the formula is:

[0090]

[0091] Where P = (p x , p y ) is the particle position to be optimized, T(i, j) and D(P+(i, j)) are the pixel values ​​of the defect-free image and the defective image at position (i, j), respectively. The point with the shortest distance is found according to the minimum value of f to determine the defect position and the corresponding position in the defect-free control image.

[0092] Furthermore, in step S4, the L-Vit structure based on the vision transformer adds a vector sequence of defect positions on the basis of Vit, and expands the defect features by adding the preprocessed sequence. Let P = (x p ,y p ) represents the coordinates of the defect location, which are the center coordinates or other representative position coordinates. ij Represents the feature vector at position (i, j) in the image feature map, W pos is the position guidance weight matrix, the position guidance score S ij The calculation is as follows:

[0093]

[0094] Where, ‖(i,j)-P‖ 2Calculate the square of the Euclidean distance between the image feature position and the defect position. γ is a hyperparameter that controls the degree of influence of the distance. According to this position guidance score, the features are weighted and summed to obtain the feature vector F after position guidance. guided :

[0095]

[0096] By the position feature vector F guided Make features close to the defect location receive more attention and emphasis.

[0097] Furthermore, the improved multi-head self-attention mechanism p-MHA in step S5 first processes q and v on the fused sequence before the three parameters q, k, and v are input, and the position of q is more accurately determined by the interactive calculation of the current position q and the actual position v. p-MHA pre-processes q and v before attention calculation, and interacts only at key positions to accurately determine the position of q.

[0098] Furthermore, the steps of improving the multi-head self-attention mechanism p-MHA in step S5 are as follows:

[0099] Step S5-1: For the input query (q), (k) and value (v), the multi-head attention mechanism will first perform a linear transformation on them to obtain different head q, k, v versions. For each head i, there will be a corresponding linear transformation matrix W qi , W ki , W vi So that:

[0100] q i =q*W qi

[0101] k i =k*W ki

[0102] v i =v*W vi

[0103] Where, i = 1, 2, 3, …, h;

[0104] Step S5-2: Calculate the attention score of each head i, where the attention score represents the query q i With key k j The degree of association between them; the attention score calculation formula is:

[0105]

[0106] where d kis the dimension of key k, and the softmax function is used to normalize the scores to a probability distribution form so that the sum of all possible association scores is 1;

[0107] Step S5-3: Calculate the output o of each head i according to the attention score i , that is, the attention score and the corresponding value v are weighted and summed, the formula is:

[0108]

Claims

1. A wafer defect detection method based on a vision transformer model, comprising the following steps: Step S1, collecting a top view of the wafer at a normal viewing angle and a top view of the wafer at a defective viewing angle, and converting the two into the same image format and size, using a bilinear interpolation algorithm, taking into account the information of four pixels around the target pixel to ensure that each pixel can correspond one to one; Step S2, compare the pixels at the same coordinate position in the two images obtained in step S1, determine the threshold d comprehensively according to the wafer image characteristics, noise and minimum defect size, traverse the pixels of the image, obtain its RGB value and calculate the absolute difference in each channel with the pixels in the surrounding nine-square grid neighborhood. Once the absolute difference of at least one channel of a pixel exceeds the threshold, it is marked as a suspected defect point. Then, according to the expected defect size under 5x5 pixels, determine the expansion area with the point as the center, and mark the pixels in the area as defect positions through binary image marking; Step S3, processing the pixel points of the defect position, and extracting the target position and target information of the defect position by using a method based on centroid matching; collecting a normal viewing angle picture of the wafer body, performing data preprocessing on the picture, and finding out the position where the defect position should be merged under the normal viewing angle through matrix similarity calculation, and performing feature fusion with the position of the defective feature; Step S4, constructing an L-Vit structure based on the vision transformer, first extracting the position information of defects in the wafer body image, encoding it into a vector sequence, the sequence contains key data such as the defect center coordinates and bounding box parameters, and then performing preprocessing, using a normalization operation to adjust the value range, using a smoothing algorithm to reduce noise interference, and finally integrating the processed vector sequence into the vision transformer architecture; Step S5: Use the improved multi-head self-attention mechanism p-MHA to calculate the defect features. First, the fusion sequence is interactively calculated to obtain the correlation between the current position and the actual position. Then p-MHA preprocesses q and v before attention calculation, and interactively filters the interference of noise and irrelevant information at key positions, and finally determines the position of q.

2. According to claim 1, a wafer defect detection method based on a vision transformer model is characterized in that: In step S1, the bilinear interpolation algorithm is adopted. Given four adjacent pixel values I x1,y1 , I x1,y2 , I x2,y1 , I x2,y2 in the original image, where x1, x2 and y1, y2 are the positions of four pixels in the original image, and x1 < x2 and y1 < y2, and the target position is (x, y), which is located between the four known pixel points. The formula is: Then, we interpolate vertically, using the following interpolation method: The final I x,y It is the interpolation result of the target position (x, y).

3. The wafer defect detection method based on the vision transformer model according to claim 1, characterized in that: Step S2: For the pixel points at the same coordinate position in the two images, the absolute difference in each color channel is calculated respectively; the defect position is located by collecting the top view of the normal angle of view of the wafer and the top view of the defect angle and using the absolute difference method: (R def (x,y),G def (x,y),B def (x,y)) (R ref (x,y),G ref (x,y),B ref (x,y)) d R (x,y)=|R def (x,y)-R ref (x,y)| Where R def (x, y) is the red pixel point of the top view of the normal viewing angle of the wafer body, G def (x, y) is the green pixel point of the top view of the wafer body at normal viewing angle, B def (x, y) is the red pixel and blue pixel of the top view of the normal viewing angle of the wafer, R ref (x, y) is the red pixel of the top view of the defective wafer, G def (x, y) is the red pixel point of the top view of the defective wafer, B def (x, y) is the blue pixel point of the top view of the defective wafer, d R (x, y) is the absolute difference between the pixel points corresponding to the top view of the wafer at a normal viewing angle and the top view of the wafer at a defective viewing angle. The defect position judgment standard is:

4. The wafer defect detection method based on the vision transformer model according to claim 1, characterized in that: In step S3, the target position and target information of the defect position are obtained by using a centroid-based matching method, and the specific steps are as follows: Step S3-1, calculate the centroid coordinates of the defect position by weighted average method; the centroid coordinates are calculated by calculating the weighted average of the pixels in the x direction and the y direction respectively, with the weight being the gray value of each pixel, and the calculation formula is: Where I(x, y) is the grayscale value of the pixel with coordinates (x, y), and ∑ represents the sum of the pixels within the entire defect location area; Step S3-2, obtain the defect-free control image information, traverse each point in the defect-free control image or traverse each area through a sliding window, and calculate the Euclidean distance d between them and the centroid coordinates of the defect position; the calculation formula of the Euclidean distance is: Where (x1, y1) is the coordinate of the centroid of the defect area, and (x2, y2) is the coordinate of the point or area to be compared in the reference photo); Step S3-3, using the particle swarm algorithm to find the point or area with the smallest distance to determine the corresponding position of the defect position in the defect-free control image, the objective function is the mean square error f, the formula is: Where P = (p x , p y ) is the particle position to be optimized, T(i, j) and D(P+(i, j)) are the pixel values ​​of the defect-free image and the defective image at position (i, j), respectively. The point with the shortest distance is found according to the minimum value of f to determine the defect position and the corresponding position in the defect-free control image.

5. The wafer defect detection method based on the vision transformer model according to claim 1, characterized in that: In step S4, the L-Vit structure based on the vision transformer is used. This structure adds a vector sequence of defect positions on the basis of Vit, and expands the defect features by adding the preprocessed sequence. Let P = (x p ,y p ) represents the coordinates of the defect location, which are the center coordinates or other representative position coordinates. ij Represents the feature vector at position (i, j) in the image feature map, W pos is the position guidance weight matrix, the position guidance score S ij The calculation is as follows: Among them, ||(i, j)-P|| 2 Calculate the square of the Euclidean distance between the image feature position and the defect position. γ is a hyperparameter that controls the degree of influence of the distance. According to this position guidance score, the features are weighted and summed to obtain the feature vector F after position guidance. guided : By the position feature vector F guided Make features close to the defect location receive more attention and emphasis.

6. The wafer defect detection method based on the vision transformer model according to claim 1, characterized in that: The improved multi-head self-attention mechanism p-MHA in step S5 first processes q and v on the fused sequence before the three parameters q, k, and v are input, and the position of q is more accurately determined by the interactive calculation of the current position q and the actual position v. p-MHA pre-processes q and v before attention calculation, and interacts only at key positions to accurately determine the position of q.

7. The wafer defect detection method based on the vision transformer model according to claim 1, characterized in that: The steps of improving the multi-head self-attention mechanism p-MHA in step S5 are as follows: Step S5-1: For the input query (q), (k) and value (v), the multi-head attention mechanism will first perform a linear transformation on them to obtain different head q, k, v versions. For each head i, there will be a corresponding linear transformation matrix W qi , W ki , W vi So that: q i =q*W qi k i =k*W ki v i =v*W vi Where, i = 1, 2, 3, …, h; Step S5-2: Calculate the attention score of each head i, where the attention score represents the query q i With key k j The degree of association between them; the attention score calculation formula is: where d k is the dimension of key k, and the softmax function is used to normalize the scores to a probability distribution form so that the sum of all possible association scores is 1; Step S5-3: Calculate the output o of each head i according to the attention score i , that is, the attention score and the corresponding value v are weighted and summed, the formula is:

Citation Information

Patent Citations

  • Transform-based wafer defect detection method and system

    CN115311203A

  • Transform-based flexible circuit board defect detection method

    CN118887209A

  • Method and apparatus for reviewing defect of subject to be inspected

    US20060104500A1

Cited By

  • Remote sensing image intelligent quality detection method and system based on image segmentation

    CN121616592A

  • Remote sensing image intelligent quality detection method and system based on image segmentation

    CN121616592B