Detection method and system
By designing a lightweight segmentation network, combining a global enhancement module with data enhancement and attention mechanism and a local refinement module, the problems of inaccurate traditional segmentation algorithms and low efficiency of deep learning models in semiconductor manufacturing are solved, and efficient and accurate detection of unqualified groove areas is achieved.
Patent Information
- Application Number
- CN202510797523.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-23
AI Technical Summary
In semiconductor manufacturing, existing technologies such as traditional image segmentation algorithms are inaccurate, deep learning models are inefficient and occupy a large amount of video memory, and cannot meet the real-time detection needs of unqualified grooves on semiconductor production lines.
A lightweight segmentation network is designed, which combines data enhancement and lightweight encoding modules, introduces a global enhancement module and a local refinement module, and improves the segmentation accuracy through the attention mechanism. It is suitable for semiconductor trench segmentation scenarios.
Under the premise of ensuring real-time requirements, the detection accuracy and efficiency of unqualified groove areas are improved to meet the detection needs of semiconductor production lines.
Smart Images

Figure CN120689313A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of semiconductors, and more specifically, to a detection method and system. Background Art
[0002] In semiconductor manufacturing, trenches are a common and necessary structure, precisely manufactured through design and process control to meet the functional requirements of semiconductor devices. Trenches typically appear on the surface or within wafers and chips, and have strict geometric specifications and electrical performance requirements. The normal trench process in semiconductors is a key step in achieving high-performance device design and is widely used in trench isolation, deep trench capacitors, trench gates, and other fields. With the continuous advancement of process technology, the requirements for trench processing precision and complexity are becoming increasingly higher, posing greater challenges to etching technology, filling technology, and defect control.
[0003] However, during semiconductor manufacturing, process defects or design issues often cause trench structures to deviate from expected requirements, impacting device performance, yield, or reliability. These defects can occur during trench formation, filling, or subsequent processing. Efficient and accurate detection of unqualified trench areas is crucial to improving chip yield and reliability.
[0004] Most existing technologies use traditional segmentation methods to segment and analyze the groove area. However, traditional segmentation algorithms are affected by the grayscale of other areas and are often unable to accurately segment the groove area, which affects the detection results. Existing deep learning-based segmentation methods often have relatively large models, require more computing resources, and cannot meet the real-time requirements. However, the detection of unqualified grooves on the production line requires real-time response, otherwise it will affect production capacity; existing real-time segmentation models often compress and quantize the model at the expense of segmentation accuracy, but directly applying existing lightweight segmentation models will not have sufficient segmentation sensitivity on the groove structure and the accuracy will not meet the requirements. In addition, deep learning models require a large amount of labeled data, but in actual application scenarios, the amount of unqualified groove sample data is insufficient. Directly training an unqualified groove segmentation model will lead to model overfitting. Summary of the Invention
[0005] The embodiments of the present application provide a detection method and system to facilitate efficient and accurate positioning of unqualified groove areas.
[0006] According to an embodiment of the present application, a detection method is provided, comprising the following steps:
[0007] S100: Preprocessing the collected groove image;
[0008] S200: performing feature processing on the pre-processed groove image to obtain fused features;
[0009] S300: Segment the fused features to obtain a groove segmentation image;
[0010] S400: Statistically analyzing the grayscale values of the groove area of the groove segmentation image to detect unqualified groove areas.
[0011] Furthermore, step S100 specifically includes:
[0012] S101: performing data enhancement on the collected groove image;
[0013] S102: Annotate the groove image after data enhancement, generate annotated image and perform preprocessing.
[0014] Furthermore, step S200 specifically includes:
[0015] S201: Extract feature maps of different scales from the preprocessed annotated image;
[0016] S202: Perform global enhancement on feature maps of different scales to obtain globally enhanced features;
[0017] S203: Refine the feature maps of different scales under the globally enhanced features to obtain refined features;
[0018] S204: Fusing the globally enhanced features and the refined features to obtain fused features.
[0019] Furthermore, in step S101 , data enhancement is to generate diversified training data by simulating the morphological changes of the trenches under different semiconductor manufacturing process conditions;
[0020] In step S102, Labelme software is used to label the groove image, generate a binary segmentation map of the same size as the groove image, and perform a normalization operation.
[0021] Furthermore, in step S201, the pre-processed annotated image is input into a lightweight encoding module to extract feature maps of different scales;
[0022] In step S202, the encoder features with the lowest resolution among the different scale features are input into the global enhancement module to obtain globally enhanced features;
[0023] In step S203, the low-level features among the features of different scales are input into the local refinement module, and are refined under the gating of the features after global enhancement to obtain refined features;
[0024] In step S204, the globally enhanced features and the refined features are input into the feature fusion module, which combines the refined encoder features and the enhanced decoder features by element-wise addition, and then inputs the results into the convolution, BN and ReLU layers for learning to obtain the fused features.
[0025] Furthermore, in step S300, the fused feature map is input into the segmentation head module. The segmentation head module uses a sigmoid layer to normalize all outputs to between 0 and 1, and obtains a groove segmentation binary image according to a threshold.
[0026] Furthermore, in step S400 , the ratio of the mean to the standard deviation is used to distinguish between normal grooves and unqualified groove areas, and the groove defect areas are detected and located.
[0027] According to another embodiment of the present application, a detection system is provided, including:
[0028] A preprocessing unit, used for preprocessing the collected groove image;
[0029] A feature processing unit, used for performing feature processing on the pre-processed groove image to obtain fused features;
[0030] The segmentation unit is used to segment the fused features to obtain a groove segmentation image;
[0031] The detection unit is used to perform statistical analysis on the grayscale values of the groove area of the groove segmentation image and detect unqualified groove areas.
[0032] Furthermore, the pre-processing unit includes:
[0033] A data enhancement unit, used for performing data enhancement on the collected groove image;
[0034] The annotation unit is used to annotate the groove image after data enhancement, generate annotated images and perform preprocessing.
[0035] Furthermore, the feature processing unit includes:
[0036] An extraction unit, used to extract feature maps of different scales from the preprocessed annotated image;
[0037] The global enhancement unit is used to globally enhance feature maps of different scales to obtain globally enhanced features;
[0038] The refinement unit is used to refine the feature maps of different scales under the global enhanced features to obtain refined features;
[0039] The fusion unit is used to fuse the globally enhanced features and the refined features to obtain fused features.
[0040] A storage medium stores a program file capable of implementing any one of the above detection methods.
[0041] A processor is used to run a program, wherein any one of the above detection methods is executed when the program is running.
[0042] The detection method and system in the embodiments of the present application solve the problems of inaccurate segmentation using traditional image segmentation algorithms, slow efficiency of commonly used deep learning models, and large memory usage. The lightweight segmentation network designed for the semiconductor groove segmentation scenario adopts a lightweight segmentation skeleton that is more suitable for segmentation tasks. It innovatively introduces a global enhancement module and a local refinement module based on the attention mechanism, further improving the segmentation accuracy while ensuring the lightweight real-time requirements, and efficiently and accurately completing the detection task of unqualified groove areas. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0044] Figure 1 This is a flow chart of the detection method of this application;
[0045] Figure 2 This is a preferred flow chart of the detection method of this application;
[0046] Figure 3 This is a preferred flow chart of the detection method of this application;
[0047] Figure 4 This is a module diagram of the detection system for this application;
[0048] Figure 5 This is the preferred module diagram of the detection system of this application;
[0049] Figure 6 This is the preferred module diagram of the detection system of this application;
[0050] Figure 7 This is a schematic diagram of the system modules in the detection system of this application. DETAILED DESCRIPTION
[0051] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0052] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0053] Example 1
[0054] According to an embodiment of the present application, a detection method is provided. Figure 1 , including the following steps:
[0055] S100: Preprocessing the collected groove image;
[0056] S200: performing feature processing on the pre-processed groove image to obtain fused features;
[0057] S300: Segment the fused features to obtain a groove segmentation image;
[0058] S400: Statistically analyzing the grayscale values of the groove area of the groove segmentation image to detect unqualified groove areas.
[0059] The detection method in the embodiment of the present application solves the problems of inaccurate segmentation using traditional image segmentation algorithms, slow efficiency of commonly used deep learning models, and large memory usage. The lightweight segmentation network designed for the semiconductor groove segmentation scenario adopts a lightweight segmentation skeleton that is more suitable for segmentation tasks. It innovatively introduces a global enhancement module and a local refinement module based on the attention mechanism, while ensuring the lightweight real-time requirements, further improving the segmentation accuracy and efficiently and accurately completing the detection task of unqualified groove areas.
[0060] Among them, see Figure 2, step S100 specifically includes:
[0061] S101: performing data enhancement on the collected groove image;
[0062] S102: Annotate the groove image after data enhancement, generate annotated image and perform preprocessing.
[0063] Among them, see Figure 3 , step S200 specifically includes:
[0064] S201: Extract feature maps of different scales from the preprocessed annotated image;
[0065] S202: Perform global enhancement on feature maps of different scales to obtain globally enhanced features;
[0066] S203: Refine the feature maps of different scales under the globally enhanced features to obtain refined features;
[0067] S204: Fusing the globally enhanced features and the refined features to obtain fused features.
[0068] The following describes the detection method of the present application in detail with reference to a preferred embodiment:
[0069] This application proposes a lightweight real-time detection algorithm. When the data of unqualified grooves is limited, a solution is designed to first segment the groove area and then screen the unqualified groove areas through adaptive thresholds. The lightweight and high-precision groove area segmentation model designed according to the characteristics of the groove area can reduce the computing and storage requirements while ensuring the detection accuracy, improve the processing speed, meet the real-time requirements, and conveniently and efficiently locate the unqualified groove areas accurately.
[0070] A detection method proposed in this application comprises the following steps:
[0071] Step 1: Perform data enhancement on the collected groove image;
[0072] Step 2: Label the enhanced groove image and perform preprocessing;
[0073] Step 3: Input the groove image obtained in step 2 and the corresponding annotated image into a lightweight encoding module to extract feature maps of different scales;
[0074] Step 4: Input the encoder features with the lowest resolution obtained in step 3 into the global enhancement module to obtain the globally enhanced features;
[0075] Step 5: Input the low-level features obtained in step 3 into the local refinement module, and refine them under the gating of the globally enhanced features obtained in step 4 to obtain refined features;
[0076] Step 6: Input the refined features in step 5 and the globally enhanced features in step 4 into the feature fusion module to obtain the fused features;
[0077] Step 7: Input the fused feature map in step 6 into the segmentation head module to obtain the groove segmentation binary image;
[0078] Step 8: Perform statistical analysis on the grayscale values of the segmented groove area to detect unqualified groove areas.
[0079] In step 1, data augmentation primarily involves combining semiconductor manufacturing processes to simulate trench morphology variations under different process conditions to generate diverse training data. Factors such as exposure dose and focus during the lithography process can cause lithographic pattern dimensions to differ from the designed values. This dimensional deviation is simulated by scaling the original image. Based on the dimensional deviation ratio in the actual process, the image is randomly enlarged or reduced within this range with a certain probability. When the lithography equipment has limited resolution, the trench edges become blurred. A filtering algorithm, such as Gaussian blur, is used to simulate this blurring effect on the original image. The degree of blurring is controlled by adjusting the size and standard deviation of the Gaussian kernel to simulate lithography results at different resolutions. Ideally, trench sidewalls should be vertical, but in practice, sidewall tilt can occur during etching. This is simulated by applying a perspective transformation to the image. Based on the range of tilt angles that can occur in the actual process, the image is randomly transformed at a selected angle within this range to alter the trench shape. Multiple process simulation operations are randomly selected and combined based on the probability of problems occurring in the actual process to generate more diverse training data and improve the model's adaptability to various complex process conditions.
[0080] In step 2, Labelme software is used to label the groove image in step 1, and a binary segmentation map of the same size as the groove image is generated; and the input data is normalized.
[0081] In step 3, the lightweight encoding module mainly includes 5 stages. The first two stages are composed of convolution, BN and ReLU, and the last three stages are composed of STDC modules.
[0082] In step 4, the global enhancement module mainly includes a semantic aggregation module and a semantic distribution module. The semantic aggregation module mainly generates an attention map based on the feature map of the lowest resolution of the encoder, and the semantic distribution module mainly performs global enhancement on the decoder feature map that is continuously upsampled from the encoder feature map of the lowest resolution based on the semantic descriptor.
[0083] In step 5, the local refinement module mainly refines the encoder features from two dimensions: channel resampling and spatial gating. Channel resampling obtains the channel attention map, increases the weight of channels that are beneficial to segmentation, and reduces the weight of feature channels that are not beneficial or even harmful to segmentation. Spatial gating mainly uses the globally enhanced features as spatial gating, focusing on valuable details in the region of interest and discarding useless or even harmful textures outside the target area to obtain the encoder features after local refinement.
[0084] In step 6, the feature fusion module mainly combines the refined encoder features and the enhanced decoder features through element-wise addition, and then inputs the results into the convolution, BN and ReLU layers to obtain the fused features.
[0085] In step 7, the segmentation head uses a sigmoid layer to normalize all outputs to between 0 and 1, and obtains a binary segmentation map based on the threshold.
[0086] In step 8, the ratio of the mean to the standard deviation can be used to distinguish normal grooves from unqualified groove areas, thereby completing the detection and positioning of the groove defect area.
[0087] Compared with the prior art, this application has the following beneficial effects:
[0088] This application discloses a detection method that solves the problems of inaccurate segmentation using traditional image segmentation algorithms, slow efficiency of commonly used deep learning models, and large memory usage. A lightweight segmentation network designed for semiconductor groove segmentation scenarios adopts a lightweight segmentation skeleton that is more suitable for segmentation tasks. It innovatively introduces a global enhancement module and a local refinement module based on an attention mechanism. While ensuring lightweight real-time requirements, it further improves segmentation accuracy and efficiently and accurately completes the detection task of unqualified groove areas.
[0089] according to Figure 7 The system shown in this application also includes the following specific steps:
[0090] 1. Network model construction
[0091] 1) The lightweight encoding module mainly consists of five stages, which downsamples the image to 1 / 32 of the input image. The first two stages are composed of convolution, batch normalization, and ReLU, and the last three stages are composed of STDC modules. The STDC module is a lightweight skeleton network designed for segmentation tasks. Each STDC module consists of multiple blocks. Assuming that the number of output channels of the STDC module is M, except that the number of convolution kernels in the convolution layer of the last block is the same as the number of convolution kernels in the convolution layer of the previous block, the number of convolution kernels in the i-th block can be expressed as In classification tasks, higher layers typically have more channels. However, in segmentation tasks, variable receptive field sizes and multi-scale information are more important. Lower layers require sufficient channels to encode finer-grained information with smaller receptive fields, while higher layers, with larger receptive fields, focus on higher-level semantic information. Setting the same number of channels for higher layers as for lower layers can lead to information redundancy. To enrich feature information, features from different blocks are concatenated using skip connections as the output of the STDC module.
[0092] 2) Global enhancement module: A large amount of semantic information is usually contained in the high-level output of the encoder. Currently, during the decoding process, upsampling is usually performed through interpolation or transposed convolution. Due to the limited convolution kernel size, it is almost impossible to encode global context information, and the ability to accurately restore the pixel level is limited. In semiconductor scenarios, photoresist residues, material reflectivity differences, interference from metal reflections, etc., all increase the difficulty of segmentation, resulting in mis-segmentation or missed segmentation. The global enhancement module summarizes the global semantic context obtained from high-level features and adaptively distributes it to the upsampled feature map. In this way, global semantic information can be aggregated and passed to various stages of the decoder to make up for the lack of global semantic information in the upsampling process, capture the spatial correlation between grooves, identify the global layout pattern of the wafer, and generate a globally enhanced decoder feature map.
[0093] First, a semantic aggregation module is introduced to aggregate global features. CxHxW Represents the output feature map of the last encoding layer, with a size of (C, H, W), where C represents the number of channels of the feature, H represents the height of the feature, and W represents the width of the feature. Global attention pooling is applied on F to selectively aggregate the visually important semantic features in the entire feature space to generate a semantic descriptor D∈R CxN , the size is (C,N). ,in and F atten =softmax(θ(F;W θ ))∈R NxHW are two different 1x1 convolutional layers of the semantic aggregation module ( and θ), the sizes are (C, HW) and (N, HW), respectively, and the corresponding learnable parameters are represented as W θ 、 Transform the feature F, where It is a convolutional layer The learnable parameters, W θ is the learnable parameter of the convolutional layer θ, T represents the atten Perform matrix transposition, N and θ represent the same number of output channels of the 1x1 convolution, and two different 1x1 convolution layers are represented by θ, The Softmax function transforms the spatial dimension to obtain the attention map F atten . × is the matrix multiplication operation.
[0094] In order to make up for the lack of global semantic information in repeated upsampling operations, a semantic distribution module is introduced. The decoder feature map that is continuously upsampled from the lowest resolution feature map F is represented as F a The semantic distribution module firstly calculates F a Apply 1x1 convolution and softmax on the channel dimension to convert the representations from all positions into attention vectors, and then adaptively integrate the semantic descriptor D into each position according to the attention vector to generate the semantic descriptor map F s =D×F a , where F a =softmax(ζ(F;W ζ )), ζ represents the 1x1 convolution of the semantic distribution module.
[0095] Next, the semantic descriptor map is added to the original decoder feature F by an addition operation a Fusion, and then through the convolution, BN and RELU layers, generate the globally enhanced decoder feature S = Ψ (F a +αF s ;W Ψ ), where Ψ represents the convolution, BN, and RELU layers of the global enhancement module.
[0096] For situations where there may be metal layer obstruction or photoresist residue on the chip surface, the global enhancement module can use global context information to infer the groove direction of the obscured area through the periodic pattern of adjacent grooves, thereby improving the groove segmentation capability under the influence of interference factors.
[0097] 3) Local Refinement Module: This module refines the encoder features primarily through channel resampling and spatial gating. Through dual channel-spatial attention, it dynamically enhances groove features in low-contrast areas. For minor disturbances such as small scratches, the gating mechanism selectively filters out noise to prevent false activation. Furthermore, the local refinement module refines only high-confidence regions of the global enhancement module's output, such as near groove boundaries, reducing computational redundancy and meeting the real-time requirements of semiconductor in-line inspection.
[0098] In the low-level features of the encoder, a considerable number of channels do not provide useful information for the final segmentation. Therefore, the channel attention mechanism is used to resample the encoder feature map in the channel dimension, so that the features with strong correlation with the decoder features in the encoder features are given a higher weight, thereby improving the representation of specific semantic features. Channel resampling is mainly achieved through F ch =Fl ×softmax(F l ×S T ), where softmax is mainly used to obtain the channel attention map, F l represents the encoder feature map, S T represents the transpose of the decoder feature matrix after global enhancement.
[0099] Since low-level features encode local details and textures, excessive texture representation outside the target area may interfere with the segmentation of the target. To this end, the semantic description map is used as a spatial gate to focus on the valuable details in the region of interest and discard useless or even harmful textures outside the target area. Here, the semantic descriptor map F calculated by the global enhancement module is s Execute sigmoid to obtain the spatial gating graph Sigmoid (F s ), and then compared with the resampled encoder feature F ch Multiply to get the locally refined encoder features: F sp =F ch Sigmoid(F s ), where · represents element-wise multiplication. The information from the global enhancement module can serve as prior knowledge for the local refinement module, guiding it to focus on error-prone areas, such as groove intersections.
[0100] 4) Feature fusion module: The refined encoder features and enhanced decoder features are combined by element-wise addition, and the result is then fed into the convolution, BN, and ReLU layers to obtain the fused features F o =Φ(F sp +βS;W Φ ), Φ represents the convolution, BN and ReLU layers of the feature fusion module, W Φ represents its learnable parameters, F sp represents the refined encoder features, S represents the enhanced decoder features, and β represents the learnable scalar weight parameter.
[0101] 5) Segmentation head module: The segmentation head uses a sigmoid layer to normalize all outputs to between 0 and 1, and obtains a binary segmentation map based on the threshold.
[0102] 6) Unqualified groove detection module: After the segmented groove image, its average grayscale value and grayscale standard deviation are counted, and a contrast index c = mean / std is designed, where mean represents the average grayscale value of the groove area, and std represents the grayscale standard deviation of the groove area. Through the adaptive threshold, the groove image that does not meet the c index requirement is identified as an unqualified groove area.
[0103] 2. Model Training
[0104] In the segmentation model of the groove area, the objective function of dice loss and bce loss weighted is adopted, which is specifically expressed as L = L bce +λL dice ,in where y and Represent the true label and predicted value of each pixel segmentation respectively. The segmentation network is trained and optimized by SGD method, using batch size of 8 and learning rate update using the "poly" learning rate strategy, i.e. The power is set to 0.9 and lr0 is set to 0.005. Where L represents the total loss function, L bce represents the bce loss function, L dice Represents the dice loss function, iter is the current number of iterations, max_iter is the maximum number of iterations, the power exponent is used to control the shape of the curve, lr0 is the baseline learning rate, and λ represents the weighted weight corresponding to the dice loss function.
[0105] Example 2
[0106] According to another embodiment of the present application, a detection system is provided. Figure 4 ,include:
[0107] A pre-processing unit 100 is used to pre-process the collected groove image;
[0108] A feature processing unit 200 is used to perform feature processing on the pre-processed groove image to obtain fused features;
[0109] A segmentation unit 300 is used to segment the fused features to obtain a groove segmentation image;
[0110] The detection unit 400 is used to perform statistical analysis on the grayscale values of the groove area of the groove segmentation image to detect unqualified groove areas.
[0111] The detection system in the embodiment of the present application solves the problems of inaccurate segmentation using traditional image segmentation algorithms, slow efficiency of commonly used deep learning models, and large memory usage. The lightweight segmentation network designed for semiconductor groove segmentation scenarios adopts a lightweight segmentation skeleton that is more suitable for segmentation tasks. It innovatively introduces a global enhancement module and a local refinement module based on the attention mechanism, further improving the segmentation accuracy while ensuring lightweight real-time requirements, and efficiently and accurately completing the detection task of unqualified groove areas.
[0112] Among them, see Figure 5 , the pre-processing unit 100 includes:
[0113] A data enhancement unit 101 is used to perform data enhancement on the collected groove image;
[0114] The labeling unit 102 is used to label the groove image after data enhancement, generate a labeled image and perform preprocessing.
[0115] Among them, see Figure 6 , the feature processing unit 200 includes:
[0116] Extraction unit 201, used to extract feature maps of different scales from the preprocessed annotated image;
[0117] A global enhancement unit 202 is configured to perform global enhancement on feature maps of different scales to obtain globally enhanced features;
[0118] A refinement unit 203 is used to refine the feature maps of different scales based on the globally enhanced features to obtain refined features;
[0119] The fusion unit 204 is used to fuse the globally enhanced features and the refined features to obtain fused features.
[0120] The following describes the detection system of the present application in detail with reference to specific embodiments:
[0121] This application proposes a lightweight real-time detection algorithm. When the data of unqualified grooves is limited, a solution is designed to first segment the groove area and then screen the unqualified groove areas through adaptive thresholds. The lightweight and high-precision groove area segmentation model designed according to the characteristics of the groove area can reduce the computing and storage requirements while ensuring the detection accuracy, improve the processing speed, meet the real-time requirements, and conveniently and efficiently locate the unqualified groove areas accurately.
[0122] A detection system proposed in this application comprises the following steps:
[0123] Step 1: Innovative data enhancement of the acquired groove images;
[0124] Step 2: Label the enhanced groove image and perform preprocessing;
[0125] Step 3: Input the groove image obtained in step 2 and the corresponding annotated image into a lightweight encoding module to extract feature maps of different scales;
[0126] Step 4: Input the encoder features with the lowest resolution obtained in step 3 into the global enhancement module to obtain the globally enhanced features;
[0127] Step 5: Input the low-level features obtained in step 3 into the local refinement module, and refine them under the gating of the globally enhanced features obtained in step 4 to obtain refined features;
[0128] Step 6: Input the refined features in step 5 and the globally enhanced features in step 4 into the feature fusion module to obtain the fused features;
[0129] Step 7: Input the fused feature map in step 6 into the segmentation head module to obtain the groove segmentation binary image;
[0130] Step 8: Perform statistical analysis on the grayscale values of the segmented groove area to detect unqualified groove areas.
[0131] In step 1, data augmentation primarily involves combining semiconductor manufacturing processes to simulate trench morphology variations under different process conditions to generate diverse training data. Factors such as exposure dose and focus during the lithography process can cause lithographic pattern dimensions to differ from the designed values. This dimensional deviation is simulated by scaling the original image. Based on the dimensional deviation ratio in the actual process, the image is randomly enlarged or reduced within this range with a certain probability. When the lithography equipment has limited resolution, the trench edges become blurred. A filtering algorithm, such as Gaussian blur, is used to simulate this blurring effect on the original image. The degree of blur is controlled by adjusting the size and standard deviation of the Gaussian kernel to simulate lithography results at different resolutions. Ideally, trench sidewalls should be vertical, but in practice, sidewalls may tilt during etching. This tilt is simulated by applying a perspective transformation to the image. Based on the range of tilt angles that may occur in the actual process, the image is randomly transformed at a selected angle within this range to alter the trench shape. Multiple process simulation operations are randomly selected and combined based on the probability of problems encountered in the actual process to generate more diverse training data and improve the model's adaptability to various complex process conditions.
[0132] In step 2, Labelme software is used to label the groove image in step 1, and a binary segmentation map of the same size as the groove image is generated; and the input data is normalized.
[0133] In step 3, the lightweight encoding module mainly includes 5 stages. The first two stages are composed of convolution, BN and ReLU, and the last three stages are composed of STDC modules.
[0134] In step 4, the global enhancement module mainly includes a semantic aggregation module and a semantic distribution module. The semantic aggregation module mainly generates an attention map based on the feature map of the lowest resolution of the encoder, and the semantic distribution module mainly performs global enhancement on the decoder feature map that is continuously upsampled from the encoder feature map of the lowest resolution based on the semantic descriptor.
[0135] In step 5, the local refinement module mainly refines the encoder features from two dimensions: channel resampling and spatial gating. Channel resampling obtains the channel attention map, increases the weight of channels that are beneficial to segmentation, and reduces the weight of feature channels that are not beneficial or even harmful to segmentation. Spatial gating mainly uses the globally enhanced features as spatial gating, focusing on valuable details in the region of interest and discarding useless or even harmful textures outside the target area to obtain the encoder features after local refinement.
[0136] In step 6, the feature fusion module mainly combines the refined encoder features and the enhanced decoder features through element-wise addition, and then inputs the results into the convolution, BN and ReLU layers to obtain the fused features.
[0137] In step 7, the segmentation head uses a sigmoid layer to normalize all outputs to between 0 and 1, and obtains a binary segmentation map based on the threshold.
[0138] In step 8, the ratio of the mean to the standard deviation can be used to distinguish normal grooves from unqualified groove areas, thereby completing the detection and positioning of the groove defect area.
[0139] Compared with the prior art, this application has the following beneficial effects:
[0140] This application discloses a detection system that solves the problems of inaccurate segmentation using traditional image segmentation algorithms, slow efficiency of commonly used deep learning models, and large video memory usage. A lightweight segmentation network designed for semiconductor groove segmentation scenarios adopts a lightweight segmentation skeleton that is more suitable for segmentation tasks. It innovatively introduces a global enhancement module and a local refinement module based on an attention mechanism. While ensuring lightweight real-time requirements, it further improves segmentation accuracy and efficiently and accurately completes the detection task of unqualified groove areas.
[0141] according to Figure 7 The detection system shown in the figure, the specific steps of this application are as follows:
[0142] 1. Network model construction
[0143] 1) The lightweight encoding module mainly consists of five stages, which downsamples the image to 1 / 32 of the input image. The first two stages are composed of convolution, batch normalization, and ReLU, and the last three stages are composed of STDC modules. The STDC module is a lightweight skeleton network designed for segmentation tasks. Each STDC module consists of multiple blocks. Assuming that the number of output channels of the STDC module is M, except that the number of convolution kernels in the convolution layer of the last block is the same as the number of convolution kernels in the convolution layer of the previous block, the number of convolution kernels in the i-th block can be expressed as In classification tasks, higher layers typically have more channels. However, in segmentation tasks, variable receptive field sizes and multi-scale information are more important. Lower layers require sufficient channels to encode finer-grained information with smaller receptive fields, while higher layers, with larger receptive fields, focus on higher-level semantic information. Setting the same number of channels for higher layers as for lower layers can lead to information redundancy. To enrich feature information, features from different blocks are concatenated using skip connections as the output of the STDC module.
[0144] 2) Global enhancement module: A large amount of semantic information is usually contained in the high-level output of the encoder. Currently, during the decoding process, upsampling is usually performed through interpolation or transposed convolution. Due to the limited convolution kernel size, it is almost impossible to encode global context information, and the ability to accurately restore the pixel level is limited. In semiconductor scenarios, photoresist residues, material reflectivity differences, interference from metal reflections, etc., all increase the difficulty of segmentation, resulting in mis-segmentation or missed segmentation. The global enhancement module summarizes the global semantic context obtained from high-level features and adaptively distributes it to the upsampled feature map. In this way, global semantic information can be aggregated and passed to various stages of the decoder to make up for the lack of global semantic information in the upsampling process, capture the spatial correlation between grooves, identify the global layout pattern of the wafer, and generate a globally enhanced decoder feature map.
[0145] First, a semantic aggregation module is introduced to aggregate global features. CxhxW Represents the output feature map of the last encoding layer, with a size of (C, H, W), where C represents the number of channels of the feature, H represents the height of the feature, and W represents the width of the feature. Global attention pooling is applied on F to selectively aggregate the visually important semantic features in the entire feature space to generate a semantic descriptor D∈R CxN , the size is (C,N). ,in and F atten =softmax(θ(F;W θ ))∈R NxHW are two different 1x1 convolutional layers of the semantic aggregation module ( and θ), the sizes are (C, HW) and (N, HW), respectively, and the corresponding learnable parameters are represented as W θ 、 Transform the feature F, where It is a convolutional layer The learnable parameters, W θ is the learnable parameter of the convolutional layer θ, T represents the atten Perform matrix transposition, N and θ represent the same number of output channels of the 1x1 convolution, and two different 1x1 convolution layers are represented by θ, The Softmax function transforms the spatial dimension to obtain the attention map F atten . × is the matrix multiplication operation.
[0146] In order to make up for the lack of global semantic information in repeated upsampling operations, a semantic distribution module is introduced. The decoder feature map that is continuously upsampled from the lowest resolution feature map F is represented as F a The semantic distribution module firstly calculates F a Apply 1x1 convolution and softmax on the channel dimension to convert the representations from all positions into attention vectors, and then adaptively integrate the semantic descriptor D into each position according to the attention vector to generate the semantic descriptor map F s =D×F a , where F a =softmax(ζ(F;W ζ )), ζ represents the 1x1 convolution of the semantic distribution module.
[0147] Next, the semantic descriptor map is added to the original decoder feature F by an addition operation a Fusion, and then through the convolution, BN and RELU layers, generate the globally enhanced decoder feature S = Ψ (F a +αF s ;W Ψ ), where Ψ represents the convolution, BN, and RELU layers of the global enhancement module.
[0148] For situations where there may be metal layer obstruction or photoresist residue on the chip surface, the global enhancement module can use global context information to infer the groove direction of the obscured area through the periodic pattern of adjacent grooves, thereby improving the groove segmentation capability under the influence of interference factors.
[0149] 3) Local refinement module: The local refinement module mainly refines the encoder features from two dimensions: channel resampling and spatial gating.
[0150] In the low-level features of the encoder, a considerable number of channels do not provide useful information for the final segmentation. Therefore, the channel attention mechanism is used to resample the encoder feature map in the channel dimension, so that the features with strong correlation with the decoder features in the encoder features are given a higher weight, thereby improving the representation of specific semantic features. Channel resampling is mainly achieved through F ch =F l ×softmax(F l ×S T ), where softmax is mainly used to obtain the channel attention map, F l represents the encoder feature map, S T represents the transpose of the decoder feature matrix after global enhancement.
[0151] Since low-level features encode local details and textures, excessive texture representation outside the target area may interfere with the segmentation of the target. To this end, the semantic description map is used as a spatial gate to focus on the valuable details in the region of interest and discard useless or even harmful textures outside the target area. Here, the semantic descriptor map F calculated by the global enhancement module is s Execute sigmoid to obtain the spatial gating graph Sigmoid (F s ), and then compared with the resampled encoder feature F ch Multiply to get the locally refined encoder features: F sp =F ch ·Sigmoid(F s ), where · represents element-wise multiplication. The information from the global enhancement module can serve as prior knowledge for the local refinement module, guiding it to focus on error-prone areas, such as groove intersections.
[0152] The local refinement module dynamically enhances groove features in low-contrast areas through channel-spatial dual attention. A gating mechanism selectively filters out minor interference, such as scratches, to prevent false activation. Furthermore, the local refinement module refines only high-confidence areas output by the global enhancement module, such as those near groove boundaries, reducing computational redundancy and meeting the real-time requirements of semiconductor in-line inspection.
[0153] 4) Feature fusion module: The refined encoder features and enhanced decoder features are combined by element-wise addition, and the result is then fed into the convolution, BN, and ReLU layers to obtain the fused features F o =Φ(F sp +βS;W Φ ), Φ represents the convolution, BN and ReLU layers of the feature fusion module, W Φ represents its learnable parameters, F sp represents the refined encoder features, S represents the enhanced decoder features, and β represents the learnable scalar weight parameter.
[0154] 5) Segmentation head module: The segmentation head uses a sigmoid layer to normalize all outputs to between 0 and 1, and obtains a binary segmentation map based on the threshold.
[0155] 6) Unqualified groove detection module: After the segmented groove image, its average grayscale value and grayscale standard deviation are counted, and a contrast index c = mean / std is designed, where mean represents the average grayscale value of the groove area, and std represents the grayscale standard deviation of the groove area. Through the adaptive threshold, the groove image that does not meet the c index requirement is identified as an unqualified groove area.
[0156] 2. Model Training
[0157] In the segmentation model of the groove area, the objective function of dice loss and bce loss weighted is adopted, which is specifically expressed as L = L bce +λL dice ,in where y and Represent the true label and predicted value of each pixel segmentation respectively. The segmentation network is trained and optimized by SGD method, using batch size of 8 and learning rate update using the "poly" learning rate strategy, i.e. The power is set to 0.9 and lr0 is set to 0.005. Where L represents the total loss function, L bce represents the bce loss function, L dice Represents the dice loss function, iter is the current number of iterations, max_iter is the maximum number of iterations, the power exponent is used to control the shape of the curve, lr0 is the baseline learning rate, and λ represents the weighted weight corresponding to the dice loss function.
[0158] Example 3
[0159] A storage medium stores a program file capable of implementing any one of the above detection methods.
[0160] In an exemplary embodiment, a computer-readable storage medium is further provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the detection step is implemented. The computer storage medium can be any available medium or data storage device that can be accessed by a computer, including but not limited to magnetic storage (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MOs), etc.), optical storage (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor storage (such as ROMs, EPROMs, EEPROMs, non-volatile memories (NAND FLASH), solid-state drives (SSDs), etc.).
[0161] Example 4
[0162] A processor is used to run a program, wherein any one of the above detection methods is executed when the program is running.
[0163] In an exemplary embodiment, a computer device is further provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the detecting step is implemented when the processor executes the computer program. The processor may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0164] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0165] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0166] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the system embodiments described above are only exemplary. For example, the division of units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0167] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected to achieve the purpose of the present embodiment according to actual needs.
[0168] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0169] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program code.
[0170] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A detection method, characterized in that: The following steps are involved: S100: Preprocessing the collected groove image; S200: performing feature processing on the pre-processed groove image to obtain fused features; S300: Segment the fused features to obtain a groove segmentation image; S400: Statistically analyzing the grayscale values of the groove area of the groove segmentation image to detect unqualified groove areas.
2. The detection method according to claim 1, characterized in that Step S100 specifically includes: S101: performing data enhancement on the collected groove image; S102: Annotate the groove image after data enhancement, generate annotated image and perform preprocessing.
3. The detection method according to claim 2, characterized in that Step S200 specifically includes: S201: Extract feature maps of different scales from the preprocessed annotated image; S202: Perform global enhancement on feature maps of different scales to obtain globally enhanced features; S203: Refine the feature maps of different scales under the globally enhanced features to obtain refined features; S204: Fusing the globally enhanced features and the refined features to obtain fused features.
4. The detection method according to claim 2, characterized in that In step S101 , data enhancement is to generate diversified training data by simulating the morphological changes of the trenches under different semiconductor manufacturing process conditions; In step S102, Labelme software is used to label the groove image, generate a binary segmentation map of the same size as the groove image, and perform a normalization operation.
5. The detection method according to claim 3, characterized in that In step S201, the pre-processed annotated image is input into a lightweight encoding module to extract feature maps of different scales; In step S202, the encoder features with the lowest resolution among the different scale features are input into the global enhancement module to obtain globally enhanced features; In step S203, the low-level features among the features of different scales are input into the local refinement module, and are refined under the gating of the features after global enhancement to obtain refined features; In step S204, the globally enhanced features and the refined features are input into the feature fusion module, which combines the refined encoder features and the enhanced decoder features by element-wise addition, and then inputs the results into the convolution, BN and ReLU layers for learning to obtain the fused features.
6. The detection method according to claim 1, characterized in that In step S300, the fused feature map is input to the segmentation head module. The segmentation head module uses a sigmoid layer to normalize all outputs to between 0 and 1, and obtains a groove segmentation binary image according to a threshold.
7. The detection method according to claim 1, characterized in that In step S400 , the ratio of the mean to the standard deviation is used to distinguish normal trenches from unqualified trench areas, and the trench defect areas are detected and located.
8. A detection system, characterized in that: include: A preprocessing unit, used for preprocessing the collected groove image; A feature processing unit, used for performing feature processing on the pre-processed groove image to obtain fused features; The segmentation unit is used to segment the fused features to obtain a groove segmentation image; The detection unit is used to perform statistical analysis on the grayscale values of the groove area of the groove segmentation image and detect unqualified groove areas.
9. The detection system according to claim 8, characterized in that: The pre-processing unit includes: A data enhancement unit, used for performing data enhancement on the collected groove image; The annotation unit is used to annotate the groove image after data enhancement, generate annotated images and perform preprocessing.
10. The detection system according to claim 9, characterized in that: The feature processing unit includes: An extraction unit, used to extract feature maps of different scales from the preprocessed annotated image; The global enhancement unit is used to globally enhance feature maps of different scales to obtain globally enhanced features; The refinement unit is used to refine the feature maps of different scales under the global enhanced features to obtain refined features; The fusion unit is used to fuse the globally enhanced features and the refined features to obtain fused features.