Wafer defect detection method and related device
Patent Information
- Application Number
- CN202611000504.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-06
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]相关技术中,工程师通常需要逐一比对缺陷形态,并手动查阅历史缺陷记录或设备日志,分析周期长达数小时甚至数天,导致缺陷信息难以及时反馈至前道工艺环节以进行工艺校正
[0061]借由上述技术方案,本申请通过从待测图像中提取比例尺区域,并分别获得文字区域图像和线段区域图像,能够降低晶圆图案、缺陷纹理、背景噪声以及文字与线段之间的相互干扰。基于串联设置的文本框定位模型和文字识别模型对比例尺文字进行提取和识别,能够提高物理标定长度的识别准确性;利用自适应阈值分割模型和形态学细化模型对比例尺线段进行像素级提取,并通过亚像素级线段端点确定比例尺对应的像素长度,能够进一步减小图像噪声、线段畸变及整数像素定位误差对测量结果的影响。由此,根据物理标定长度和像素长度建立单像素对应的物理尺寸映射关系,可将缺陷目标的像素尺寸准确转换为实际物理尺寸,从而提高晶圆缺陷检测及尺寸分析的准确性和效率,使缺陷目标的尺寸信息能够更及时地用于溯源和工艺校正,进而有助于提升产线良率和生产效率。
Smart Images

Figure CN122820618A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of semiconductor inspection technology, and in particular to a wafer defect detection method and related apparatus. Background Technology
[0002] In semiconductor manufacturing, to ensure smooth operation and improve product quality and reliability, monitoring stations are typically established after critical process steps to perform precise defect detection on wafers. These stations are equipped with advanced inspection equipment capable of high-resolution scanning and imaging of the wafer surface. When a defect is detected, the equipment automatically captures and saves an image of the defect for subsequent analysis and processing. To analyze the type of wafer surface defect and its underlying process cause, quality control engineers usually need to review the defect images captured by the inspection equipment and combine this information with their own experience to make an analysis and judgment, thereby tracing and assessing the source of the defect.
[0003] In related technologies, engineers typically need to compare defect morphologies one by one and manually review historical defect records or equipment logs, with analysis cycles lasting several hours or even days. This makes it difficult to promptly relay defect information to upstream processes for correction. Furthermore, different engineers may have varying interpretations of similar defect images, affecting the accuracy of defect tracing results and impacting overall production line yield and efficiency. Summary of the Invention
[0004] In view of the above problems, this application provides a wafer defect detection method and related apparatus to achieve rapid and accurate measurement of the physical dimensions of defects in wafer defect images, improve the efficiency of defect analysis, and thus provide timely feedback of process correction information, thereby improving production line yield and production efficiency. The specific solution is as follows:
[0005] The first aspect of this application provides a wafer defect detection method, including:
[0006] Determine the scale area based on the image to be tested and extract the image of the scale area;
[0007] Based on the scale area image, determine the text area image and line segment area image corresponding to the scale respectively;
[0008] Based on the text box positioning model and the character recognition model set in series, the text region image is used to extract and recognize the characters, and the physical calibration length corresponding to the scale is determined.
[0009] The pixel-level extraction of line segments in the image of the line segment region is performed using a preset adaptive threshold segmentation model and a morphological thinning model. The pixel length corresponding to the scale is determined by reverse derivation of the sub-pixel-level line segment endpoints.
[0010] The physical size mapping relationship of a single pixel in the current image under test is calculated based on the physical calibration length and the pixel length, so as to determine the physical size of the defect target in the image under test based on the physical size mapping relationship.
[0011] Optionally, the step of determining the scale region based on the image to be tested and extracting the scale region image includes:
[0012] The image to be tested is subjected to a first image preprocessing to obtain a first target image;
[0013] The first target image is input into the first detection model to obtain at least one scale candidate detection box and the confidence level corresponding to each scale candidate detection box.
[0014] A target scale detection box is determined from at least one of the scale candidate detection boxes based on the confidence level corresponding to each scale candidate detection box;
[0015] Based on the target scale detection box, a scale region image is obtained by cropping from the image to be tested.
[0016] Optionally, determining the text region image and line segment region image corresponding to the scale based on the scale region image includes:
[0017] The scaled area image is subjected to a second image preprocessing to obtain a second target image;
[0018] The second target image is input into the second detection model to obtain candidate detection boxes for text regions and candidate detection boxes for line segment regions.
[0019] Deduplication is performed on the candidate detection boxes for the text region and the candidate detection boxes for the line segment region, respectively, so as to determine the target text detection box in the candidate detection box for the text region and the target line segment detection box in the candidate detection box for the line segment region.
[0020] Based on the target text detection box and the target line segment detection box, text region images and line segment region images are respectively cropped from the scale region image.
[0021] Optionally, the text box positioning model and character recognition model based on the concatenated configuration extract and recognize characters from the text region image to determine the physical calibration length corresponding to the scale bar, including:
[0022] The text region image is subjected to third image preprocessing to obtain a third target image;
[0023] The third target image is input into the text box positioning model to locate the text region, and the target text sub-image is cropped from the third target image based on the located text region.
[0024] The target text sub-image is input into the character recognition model for character recognition to obtain the scale string;
[0025] The scale string is parsed to obtain numerical and unit information;
[0026] Based on the numerical information and the unit information, the physical calibration length corresponding to the scale is determined.
[0027] Optionally, the step of using a preset adaptive threshold segmentation model and morphological thinning model to extract line segments at the pixel level from the line segment region image, and determining the pixel length corresponding to the scale bar by reverse derivation of sub-pixel level line segment endpoints, includes:
[0028] The image of the line segment region is subjected to a fourth image preprocessing to obtain a fourth target image;
[0029] The fourth target image is input into the adaptive threshold segmentation model to perform adaptive threshold segmentation on the fourth target image based on each pixel in the fourth target image to obtain the fifth target image. The fifth target image is then input into the morphological thinning model to perform morphological optimization and skeleton extraction on the foreground region to obtain a single-pixel skeleton of a line segment.
[0030] The two endpoint regions of the line segment are determined based on the single-pixel skeleton of the line segment, and the coordinates of the two sub-pixel endpoints are derived in reverse based on the image information of the two endpoint regions. The pixel length corresponding to the scale is determined based on the coordinates of the two sub-pixel endpoints.
[0031] Optionally, the step of inputting the fourth target image into the adaptive thresholding model to perform adaptive thresholding on the fourth target image based on each pixel in the fourth target image to obtain a fifth target image, and inputting the fifth target image into the morphological thinning model to perform morphological optimization and skeleton extraction on the foreground region to obtain a single-pixel skeleton of a line segment, includes:
[0032] The fourth target image is input into the adaptive threshold segmentation model, so that the adaptive threshold segmentation model determines the local segmentation threshold corresponding to each pixel based on the neighborhood grayscale information of each pixel in the fourth target image.
[0033] For each pixel in the fourth target image, the adaptive threshold segmentation model determines the current pixel as a foreground pixel or a background pixel based on the comparison result between the gray value of the current pixel and the local segmentation threshold corresponding to the current pixel, thus obtaining a fifth target image containing a foreground region and a background region.
[0034] The fifth target image is input into the morphological thinning model, and the morphological thinning model is used to perform erosion and dilation processing on the fifth target image to remove isolated noise pixels and perform contour smoothing processing on the foreground region of the fifth target image to obtain the sixth target image.
[0035] The sixth target image is input into the morphological thinning model, and the foreground region in the sixth target image is skeletonized by the morphological thinning model to extract the central axis of the foreground region and obtain a single-pixel skeleton of a line segment.
[0036] Optionally, the step of determining the two endpoint regions of the line segment based on the single-pixel skeleton of the line segment, and deriving the coordinates of the two sub-pixel endpoints based on the image information of the two endpoint regions, and determining the pixel length corresponding to the scale bar based on the coordinates of the two sub-pixel endpoints, includes:
[0037] Perform connected component analysis on the single-pixel skeleton of the line segment to determine the target skeleton point set from the single-pixel skeleton of the line segment;
[0038] The target skeleton point set is fitted with a straight line to obtain the target straight line model;
[0039] Determine the two endpoint regions of the line segment based on the target straight line model;
[0040] Based on the grayscale distribution information of multiple pixels within each endpoint region, a sub-pixel position estimation method is used to locate the sub-pixel endpoints of each endpoint region and determine the coordinates of the two sub-pixel endpoints of the line segment; wherein, the sub-pixel position estimation method includes at least one of quadratic curve interpolation, Taylor series fitting, edge grayscale gradient distribution, or grayscale moment calculation.
[0041] The coordinate distance between the two sub-pixel endpoints is calculated based on the coordinates of the two sub-pixel endpoints, and the coordinate distance is used to determine the pixel length corresponding to the scale.
[0042] Optionally, the method further includes:
[0043] Obtain at least one of the following: a first confidence level corresponding to the target scale detection box, a second confidence level corresponding to the target text detection box, a third confidence level corresponding to the target line segment detection box, a fourth confidence level corresponding to the physical calibration length, and a fifth confidence level corresponding to the pixel length;
[0044] Based on the obtained confidence levels and the preset weights corresponding to each confidence level, the mapping confidence level of the physical size mapping relationship corresponding to the single pixel is determined;
[0045] If the mapping confidence level is less than the preset confidence threshold, an abnormal mapping message will be output.
[0046] A second aspect of this application provides a wafer defect detection device, comprising:
[0047] The scale region extraction unit is used to determine the scale region based on the image to be measured and extract the scale region image.
[0048] The scale information separation unit is used to determine the text area image and line segment area image corresponding to the scale based on the scale area image.
[0049] The physical calibration length determination unit is used to extract and recognize text in the text region image based on the serially set text box positioning model and text recognition model, and determine the physical calibration length corresponding to the scale.
[0050] The pixel length determination unit is used to extract the line segments at the pixel level from the line segment region image using a preset adaptive threshold segmentation model and morphological thinning model, and to determine the pixel length corresponding to the scale by reverse derivation of the sub-pixel level line segment endpoints.
[0051] The defect size determination unit is used to calculate the physical size mapping relationship of a single pixel in the current image under test based on the physical calibration length and the pixel length, so as to determine the physical size of the defect target in the image under test based on the physical size mapping relationship.
[0052] Optionally, the pixel length determination unit includes:
[0053] An image preprocessing subunit is used to perform a fourth image preprocessing on the line segment region image to obtain a fourth target image;
[0054] A pixel-level extraction subunit is used to input the fourth target image into the adaptive threshold segmentation model to extract the foreground region composed of foreground pixels from the fourth target image to obtain the fifth target image, and input the fifth target image into the morphological thinning model to perform morphological optimization and skeleton extraction on the foreground region to obtain a single-pixel skeleton of a line segment.
[0055] The pixel length calculation subunit is used to determine the two endpoint regions of the line segment based on the single-pixel skeleton of the line segment, and to deduce the coordinates of the two sub-pixel endpoints based on the image information of the two endpoint regions, and to determine the pixel length corresponding to the scale based on the coordinates of the two sub-pixel endpoints.
[0056] A third aspect of this application provides an electronic device, comprising at least one processor and a memory connected to the processor, wherein:
[0057] The memory is used to store computer programs;
[0058] The processor is used to execute the computer program so that the electronic device can implement the wafer defect detection method of the first aspect or any implementation thereof.
[0059] A fourth aspect of this application provides a computer storage medium carrying one or more computer programs, which, when executed by an electronic device, enable the electronic device to perform a wafer defect detection method as described in the first aspect or any implementation thereof.
[0060] The fifth aspect of this application provides a computer program product including computer-readable instructions that, when executed on an electronic device, cause the electronic device to implement the wafer defect detection method of the first aspect or any implementation thereof.
[0061] By employing the aforementioned technical solution, this application extracts the scale region from the image under test and obtains text region images and line segment region images respectively, thereby reducing wafer patterns, defect textures, background noise, and mutual interference between text and line segments. The extraction and recognition of scale text based on a concatenated text box positioning model and text recognition model improves the accuracy of physical calibration length recognition. The pixel-level extraction of scale line segments using an adaptive threshold segmentation model and a morphological thinning model, and the determination of the corresponding pixel length of the scale through sub-pixel-level line segment endpoints, further reduces the impact of image noise, line segment distortion, and integer pixel positioning errors on the measurement results. Therefore, by establishing a physical size mapping relationship between a single pixel and the pixel length based on the physical calibration length and pixel length, the pixel size of the defect target can be accurately converted into the actual physical size, thereby improving the accuracy and efficiency of wafer defect detection and size analysis. This allows the size information of the defect target to be used more promptly for traceability and process correction, ultimately contributing to improved production line yield and production efficiency. Attached Figure Description
[0062] Figure 1 This is a schematic flowchart of a wafer defect detection method provided in an embodiment of this application;
[0063] Figure 2 A schematic diagram illustrating the pixel-to-physical-size mapping process of a wafer defect image provided in this application embodiment;
[0064] Figure 3 This is a schematic diagram of a wafer defect detection device provided in an embodiment of this application;
[0065] Figure 4 This is a schematic diagram of a computer device structure provided in an embodiment of this application. Detailed Implementation
[0066] The terminology used in the implementation section of this application is only for explaining specific embodiments of this application and is not intended to limit this application. Those skilled in the art will recognize that, with technological advancements and the emergence of new scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0067] The terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such terms are interchangeable where appropriate; this is merely a way of distinguishing objects with the same attributes in the embodiments of this application. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, so that a process, method, system, product, or apparatus that comprises a series of elements is not necessarily limited to those elements, but may include other elements not explicitly listed or inherent to those processes, methods, products, or apparatuses.
[0068] The wafer defect detection method of this application embodiment will be described in detail below with reference to the accompanying drawings. (Refer to...) Figure 1 , Figure 1 This is a flowchart illustrating a wafer defect detection method provided in an embodiment of this application, as shown below. Figure 1 As shown in the embodiment of this application, a wafer defect detection method may include steps S110 to S150, which are described in detail below.
[0069] S110. Determine the scale area based on the image to be tested and extract the scale area image.
[0070] The image under test refers to the image acquired by wafer inspection equipment for defects on the wafer surface. The image under test may include defect targets, wafer surface patterns, image annotation information, and a scale bar representing the image size. Defect targets refer to abnormal areas in the image under test that require dimensional measurement, such as scratches, pits, particle contamination, or other abnormal morphological areas on the wafer surface. The pixel dimensions of the defect targets may include their length, width, diameter, or outline dimensions, representing the image size of the defect targets. The scale bar region refers to the image area in the image under test containing scale bar information, which represents the correspondence between pixel dimensions in the image and actual physical dimensions. By performing image analysis on the image under test, identifying the location of the scale bar in the image, and extracting the scale bar region image based on this location, the scale bar information can be separated from the entire image under test, reducing the influence of defect targets, wafer background textures, and other irrelevant image content on scale bar recognition.
[0071] S120. Determine the text area image and line segment area image corresponding to the scale based on the scale area image.
[0072] Scale area images typically include text representing actual lengths and corresponding line segments. The text represents the physical calibration length of the scale, while the line segments represent the pixel lengths corresponding to those calibration lengths in the image. To extract these two types of information separately, the scale area image needs further region segmentation to identify text and line segment regions, and then extract the text region images and line segment region images respectively. The text region image contains the scale value and units, while the line segment region image contains the scale line segments. By extracting the text and line segment region images separately, the mutual interference between text and line segments can be reduced in subsequent text recognition and line segment measurement processes, improving the stability of the physical calibration length and pixel length determination processes.
[0073] S130. Based on the text box positioning model and text recognition model set in series, the text region image is extracted and recognized to determine the physical calibration length corresponding to the scale.
[0074] The physical calibration length refers to the actual spatial length represented by the scale text, and the unit can be nanometers, micrometers, millimeters, or other length units. In one feasible implementation, the text region image can first be input into a text box positioning model to locate and extract the target text region containing the scale text from the text region image; then, the target text region is input into a character recognition model set in series with the text box positioning model to recognize the characters in the target text region and obtain the scale string; by parsing the numerical and unit information in the scale string, the physical calibration length corresponding to the scale can be determined. For example, when the scale text in the text region image is represented as 100 μm, the physical calibration length corresponding to this scale can be determined to be 100 micrometers.
[0075] S140. Using a preset adaptive threshold segmentation model and morphological thinning model, pixel-level extraction of line segments is performed on the line segment region image. The pixel length corresponding to the scale is determined by reverse derivation of the sub-pixel level line segment endpoints.
[0076] Pixel length refers to the length of the scale line segment in the image being measured, measured in pixels. In one feasible implementation, the image of the line segment region can be input into a preset adaptive threshold segmentation model to distinguish the foreground and background of the line segment based on the local grayscale information of the image, achieving pixel-level extraction of the scale line segment. The segmented image is then input into a morphological thinning model to optimize the foreground and extract the skeleton of the line segment, obtaining a skeleton representing the position and extension direction of the scale line segment. Based on this, the two endpoint regions of the scale line segment are determined according to the skeleton, and the corresponding sub-pixel endpoint coordinates are derived from the image information of each endpoint region. Finally, the pixel length corresponding to the scale is determined based on the distance between the two sub-pixel endpoint coordinates. Sub-pixel-level endpoint localization reduces the influence of integer pixel coordinate limitations and factors such as line segment noise and edge distortion on the pixel length measurement results, improving the accuracy of determining the scale pixel length. Here, pixel length can be understood as the distance between the two ends of the scale line segment in the image coordinate system, representing the actual length corresponding to a certain number of pixels in the current image.
[0077] S150. Calculate the physical size mapping relationship of a single pixel in the current image under test based on the physical calibration length and pixel length, so as to determine the physical size of the defect target in the image under test based on the physical size mapping relationship.
[0078] The physical size mapping relationship refers to the correspondence between pixel dimensions and actual physical dimensions in the image under test. The physical size corresponding to a single pixel can be represented by λ, which characterizes the actual physical length corresponding to the side length of a pixel in the image. Specifically, λ can be calculated based on the ratio between the physical calibration length corresponding to the scale and the pixel length corresponding to the scale; that is, λ equals the physical calibration length divided by the pixel length. For example, if the physical calibration length is in nanometers and the pixel length is in pixels, then the unit of λ can be nanometers per pixel. After obtaining λ, the pixel dimensions of the defective target in the image under test can be converted into physical dimensions based on λ. For example, if the length of the defective target in the image is several pixels, the pixel length can be multiplied by λ to obtain the actual physical length corresponding to the defective target; if it is necessary to determine the width or other linear dimensions of the defective target, the same conversion method can be used.
[0079] By employing the aforementioned technical solution, this application extracts the scale region from the image under test and obtains text region images and line segment region images respectively, thereby reducing wafer patterns, defect textures, background noise, and mutual interference between text and line segments. The extraction and recognition of scale text based on a concatenated text box positioning model and text recognition model improves the accuracy of physical calibration length recognition. The pixel-level extraction of scale line segments using an adaptive threshold segmentation model and a morphological thinning model, and the determination of the corresponding pixel length of the scale through sub-pixel-level line segment endpoints, further reduces the impact of image noise, line segment distortion, and integer pixel positioning errors on the measurement results. Therefore, by establishing a physical size mapping relationship between a single pixel and the pixel length based on the physical calibration length and pixel length, the pixel size of the defect target can be accurately converted into the actual physical size, thereby improving the accuracy and efficiency of wafer defect detection and size analysis. This allows the size information of the defect target to be used more promptly for traceability and process correction, ultimately contributing to improved production line yield and production efficiency.
[0080] In one optional implementation, the aforementioned wafer defect detection method can be deployed on the software processing end of a wafer defect detection system and implemented in a modular manner using the C++ language. After acquiring a single image of a defect on the wafer under test, the software processing end sequentially performs processes such as scale region extraction, determination of text and line segment images, physical calibration length recognition, pixel length measurement, and physical size mapping relationship calculation to output the physical size of the defect target. Since this implementation primarily uses software algorithms to complete image scale recognition and size conversion, no modifications are required to the optical imaging structure, image acquisition hardware, or production line hardware interface of existing automated optical inspection equipment. This allows for the conversion from pixel size to nanometer-level physical size based on wafer defect images acquired by existing equipment. In practical applications, the processing time for a single image of a defect on the wafer under test can be less than 30 milliseconds, thus meeting the requirements for online or near real-time defect analysis and reducing the uncertainty of size mapping results while ensuring measurement accuracy.
[0081] In high-resolution wafer defect images, the scale bar is typically located in a localized area, while the entire image may also contain wafer circuit patterns, defect scratches, background textures, and device character annotations. If global image processing or edge-feature-based localization methods are used, the scale bar location results may be affected by irrelevant patterns or characters in complex backgrounds, leading to subsequent parsing failures or decreased accuracy. To address this issue, this application employs a first detection model for coarse scale bar localization across the entire defect image. Leveraging the deep learning target detection capabilities of this first detection model, it automatically learns and identifies the visual features of the scale bar region even in environments with significant interference, outputting the scale bar's bounding box. This reduces false detections caused by wafer patterns, characters, or defect scratches. Compared to traditional localization methods based on fixed templates or manual features, this approach offers greater stability and adaptability while avoiding the computational overhead of complex image segmentation across the entire image.
[0082] Specifically, in an optional implementation, step S110, which involves determining the scale region and extracting the scale region image based on the image to be measured, may include the following steps S111 to S114:
[0083] S111. Perform first image preprocessing on the image to be tested to obtain the first target image.
[0084] The first image preprocessing is used to make the image under test meet the input requirements of the first detection model, and may include image scaling, normalization, and padding. Specifically, image scaling adjusts the image under test to the input size required by the first detection model, for example, scaling it to a fixed input size of 640×640 pixels; normalization converts the image pixel values to a preset numerical range to improve the stability of the model processing, for example, mapping pixel values to between 0 and 1; padding maintains the aspect ratio of the image under test when adjusting the image size, avoiding geometric distortion of lines or text in the scale area caused by directly stretching the image. The resulting first target image satisfies the input format requirements of the first detection model while preserving as much of the morphological characteristics of the scale area in the original image as possible.
[0085] S112. Input the first target image into the first detection model to obtain at least one scale candidate detection box and the confidence level corresponding to each scale candidate detection box.
[0086] The first detection model is used to detect regions in the image to be tested that may contain a scale bar. In one implementation, the first detection model can be a first-level YOLOv5 model. This first detection model extracts multi-scale features of the first target image through a convolutional backbone network and predicts multiple candidate bounding boxes that may contain scale bar regions on the feature map. Each bounding box is accompanied by a confidence score, which characterizes the likelihood that the corresponding candidate bounding box contains a scale bar region. Since the image to be tested may contain line segments, characters, or textures similar in shape to a scale bar, the first detection model can output multiple candidate bounding boxes, and each candidate bounding box can record its own coordinate position and confidence score.
[0087] S113. Determine the target scale detection box from at least one scale candidate detection box based on the confidence level corresponding to each scale candidate detection box.
[0088] Specifically, duplicate detection boxes can be filtered out from multiple candidate detection boxes of different scales to remove duplicate detection boxes that overlap in position or point to the same scale region. The detection box with the highest confidence among the remaining candidate detection boxes is then selected as the target scale detection box. The preset confidence threshold can be understood as a threshold parameter used to determine whether the detection result meets the reliability requirements. Optionally, if the confidence of the target scale detection box is lower than the preset confidence threshold (e.g., 0.8), it indicates that the reliability of the current scale region detection result is insufficient, and an abnormal alarm message can be output to indicate that the current image may have missing scales, blurred scales, or detection anomalies; if the confidence of the target scale detection box is greater than or equal to the preset confidence threshold, the target scale detection box can be confirmed as valid.
[0089] S114. Based on the target scale detection box, crop the scale region image from the image to be tested.
[0090] Specifically, based on the coordinates of the target scale detection box, the image can be cropped from the original image to be tested to obtain the scale region image. Since the cropping is done from the original image to be tested, rather than directly from the first target image after scaling, normalization, or padding, the scale text edges, line segment edges, and local details in the original high-resolution wafer defect image can be preserved as much as possible, facilitating subsequent text recognition and line segment measurement.
[0091] It is understandable that in a scale area image, text content and line segment content are usually placed adjacently. The scale area image obtained in step S110 has eliminated most of the interference from wafer patterns, defects, scratches, and background textures. However, the scale area image may still contain text, units, line segments, and edge noise simultaneously. If text recognition or line segment measurement is performed directly on the entire scale area image, the text recognition process may be affected by the line segments, and the line segment measurement process may also be affected by the text outline. For example, the vertical line structure in numbers, letters, or unit symbols may interfere with the determination of the line segment region, and the scale line segment may also be misidentified as part of the text content. Based on this, the embodiments of this application further detect the scale area image using a second detection model to determine the text region and line segment region respectively, and extract the text region image and line segment region image, thereby providing a less interference-prone input image for subsequent text recognition and pixel length measurement.
[0092] Specifically, in an optional implementation, step S120, which determines the text region image and line segment region image corresponding to the scale based on the scale region image, specifically includes steps S121 to S124:
[0093] S121. Perform a second image preprocessing on the scale area image to obtain the second target image.
[0094] The second image preprocessing is used to ensure that the scale region image meets the input requirements of the second detection model, while preserving as much text outline and line texture information as possible. It's understood that since the scale region image is usually a local image cropped from the image to be tested, the size of this local image may not match the input resolution required by the second detection model. Therefore, the scale region image can be normalized, padded, or resampled to obtain the second target image. Resampling here can be understood as adjusting the scale region image to the target resolution by adding, deleting, or interpolating pixels; the target resolution refers to the resolution required by the network layer of the second detection model for the input image size. Through the second image preprocessing, unnecessary geometric deformation of text and line segments can be avoided as much as possible while meeting the model's input requirements.
[0095] S122. Input the second target image into the second detection model to obtain candidate detection boxes for text regions and candidate detection boxes for line segment regions.
[0096] In this embodiment, the second detection model is used to perform fine-grained target detection within the scale region image to determine the positions of text regions and line segment regions, respectively. In an optional implementation, the second detection model can be a second-level YOLOv5 model. The model architecture of the second detection model can be the same as the first detection model used for scale region detection, but due to differences in training samples, training targets, or model parameters, the second detection model focuses more on extracting text contour features and line segment texture features within the scale. After extracting features from the second target image, the second detection model outputs candidate detection boxes for text regions and candidate detection boxes for line segment regions, each with a confidence score. The candidate detection boxes for text regions represent regions that may contain scale text, and the candidate detection boxes for line segment regions represent regions that may contain scale line segments. Since text and line segments in the scale are usually adjacent but do not overlap, even if they are slightly connected, the second detection model can still output corresponding candidate detection boxes based on the differences in their shape, texture, and spatial position.
[0097] S123. Perform deduplication on the candidate detection boxes for the text region and the candidate detection boxes for the line segment region, respectively, so as to determine the target text detection box in the candidate detection box for the text region and the target line segment detection box in the candidate detection box for the line segment region.
[0098] Since the second detection model may output multiple candidate detection boxes that are close in position or have a high degree of overlap for the same text region or the same line segment region, it is necessary to perform detection box deduplication for the candidate detection boxes of the text region and the candidate detection boxes of the line segment region separately. Specifically, based on the degree of positional overlap between candidate detection boxes of the same category and the confidence level of the candidate detection boxes, detection boxes that can represent the target region can be retained from multiple candidate detection boxes, and duplicate detection boxes can be removed.
[0099] S124. Based on the target text detection box and the target line segment detection box, the text region image and the line segment region image are respectively cropped from the scale region image.
[0100] Specifically, based on the position information of the target text detection box and the target line segment detection box in the second target image, the corresponding positions of the target text detection box and the target line segment detection box in the scale area image can be determined, and the text area image and the line segment area image can be cropped from the scale area image based on the corresponding positions. Here, cropping can be performed from the original scale area image to preserve as much text edge, line segment edge, and pixel detail information as possible in the original scale area image.
[0101] Please refer to Figure 2 , Figure 2 This is a schematic diagram illustrating the pixel-to-physical-size mapping process for a wafer defect image, provided as an embodiment of this application. Figure 2 In the processing flow shown, the image to be tested contains a wafer pattern, defect area, equipment markings, and scale information. The scale information includes scale text and line segments corresponding to the scale text. Figure 2 The example used is a test image with a scale of 87.5 μm for the text, where 87.5 μm represents the actual physical length corresponding to the line segment on the scale. Since the text region image is a local image further cropped from the scale region image, it may have inconsistent orientations, blurred character edges, insufficient contrast, or excessive blank space at the edges. Directly performing character recognition can easily lead to numerical misidentification, unit misidentification, or character omission, thus affecting the accuracy of subsequently determining the physical size of the defect target. Therefore, this embodiment of the application... Figure 2 In the process shown, after completing scale detection, YOLOv5 detection (that is, the step in the aforementioned embodiment of using the first-level YOLOv5 model to coarsely locate the scale region of the image to be tested and output the target scale detection box), and region of interest (ROI) extraction (that is, the step in the aforementioned embodiment of cropping the scale region image from the image to be tested based on the target scale detection box), the extracted scale region is further processed by text recognition to convert the scale text in the image into a physical calibration length that can be used for calculation.
[0102] Specifically, in an optional implementation, step S130 above extracts and recognizes text in the text region image based on a concatenated text box positioning model and a text recognition model, and determines the physical calibration length corresponding to the scale bar, specifically including the following steps S131~S135:
[0103] S131. Perform third image preprocessing on the text region image to obtain the third target image.
[0104] Combination Figure 2 As shown, the extracted scale area contains green line segments and text content such as 87.5μm. Before recognizing this text content, a third image preprocessing can be performed on the text region image to improve the recognizability of the text content. In this embodiment, the third image preprocessing may include orientation correction, grayscale conversion, noise reduction, contrast enhancement, edge sharpening, binarization, and size normalization. The cropped text region image may be somewhat skewed due to the detection box angle, image acquisition angle, or cropping position. Therefore, orientation correction processing can be performed on the text region image to adjust the text to a preset reference direction, such as horizontal arrangement, thereby facilitating the subsequent character recognition model to read the text order and recognize the characters. Grayscale conversion reduces the impact of color information on character recognition, allowing subsequent processing to primarily rely on the grayscale differences between character strokes and the background. Denoising reduces the impact of image acquisition noise, background noise, or local noise on character outlines. Contrast enhancement improves uneven brightness, low contrast, or blurred characters in text region images, highlighting the differences between character strokes and the background. Edge sharpening enhances character edge outlines, making it easier for text region localization and character recognition to capture character boundaries. Binarization distinguishes the foreground and background of characters in the text region image, reducing the interference of background texture on recognition results. Size normalization ensures the image size meets the input requirements for subsequent text region localization and character recognition. Through these preprocessing operations, the third target image can be presented more clearly. Figure 2 The numerical and unit information of medium-scale text is used to improve the stability and accuracy of subsequent text region positioning and character recognition, without changing the content expressed by the text itself.
[0105] S132. Input the third target image into the text box positioning model to locate the text region, and crop the target text sub-image from the third target image based on the located text region.
[0106] Combination Figure 2 As shown in the process, after extracting the ROI region, there may still be green line segments, blank text edges, or a small amount of background information in that region. Therefore, it is necessary to further locate the region that actually contains valid text content in the text region image. Figure 2 As shown in the text box under "Character Recognition," the text box localization steps allow input of the third target image into the text box localization model for text region localization. The text box localization model can employ text detection models within the PaddleOCR framework, such as detectors based on Differentiable Binarization Network (DBNet). These detectors select the smallest bounding rectangle containing the text, excluding blank areas at image edges, isolated noise points, or irrelevant symbols. The output is the smallest bounding rectangle containing only valid characters and units; this smallest bounding rectangle is the target text region.
[0107] Based on the target text region defined by the smallest bounding rectangle, a target text sub-image is cropped from the third target image. The target text sub-image may contain only numeric characters and unit characters, for example... Figure 2 The 87.5μm shown eliminates blank areas, isolated noise, scale line segments, or other irrelevant symbols at the edges of the third target image, thereby improving the recognition accuracy of scale strings.
[0108] S133. Input the target text sub-image into the character recognition model for character recognition to obtain the scale string.
[0109] The scale string can include numeric characters, decimal points, and unit characters. For example, Figure 2 The character recognition result corresponding to the target text region can be 87.5μm.
[0110] like Figure 2As shown in the Chinese character recognition steps, in one feasible implementation, a Convolutional Recurrent Neural Network (CRNN) can be used as the character recognition model. After inputting the target text sub-image into the CRNN, the visual features of the target text sub-image are first extracted through a Convolutional Neural Network (CNN) to obtain a feature sequence formed according to the character arrangement direction. Then, a Recurrent Neural Network (RNN) performs temporal modeling on these feature sequences to learn the contextual relationships between adjacent characters. Finally, a Connectionist Temporal Classification (CTC) layer decodes the character probability sequence output by the RNN to obtain the scale string. This entire recognition process does not require individual character segmentation and can adapt to situations where numerical and unit characters are continuously arranged, character spacing is uneven, or characters are slightly adhered in the scale text, thus directly outputting a scale string containing both numerical and unit information. In another optional implementation, other Optical Character Recognition (OCR) models can also be used to perform character recognition on the target text sub-image to convert the image-based scale text into a computer-processable scale string.
[0111] S134. Parse the scale string to obtain numerical and unit information.
[0112] like Figure 2 As shown in the string parsing steps, since the identified string consists of consecutive numerical and unit parts, such as 87.5 and μm in 87.5μm, it needs to be split according to preset rules. Specifically, regular expression matching or parsing methods based on unit keywords can be used to extract the purely numerical parts (including integers or decimals) from the string as numerical information, and extract the remaining letter or symbol parts as unit information (such as μm, nm, mm, um, etc.). If there are spaces or special characters in the string, the parsing module will first perform standardization processing (removing spaces, unifying capitalization, unifying the micrometer symbol representation, and recognizing decimal points, etc.) to avoid ambiguity in subsequent conversions caused by different character formats or unit writing.
[0113] S135. Based on the numerical and unit information, determine the physical calibration length corresponding to the scale.
[0114] Combination Figure 2The process shown, after obtaining the numerical information 87.5 and the unit information μm, determines the physical calibration length corresponding to the scale bar according to preset unit conversion rules. The physical calibration length refers to the actual physical length represented by the scale bar text, used to establish a correspondence with the pixel length obtained by measuring line segments. For example, if the preset length unit is nanometers, then 87.5μm can be converted to 87500nm; if the subsequent mapping relationship is expressed in micrometers per pixel, then it can also be retained as 87.5μm. Figure 2 The document also shows the processing result of a line segment with a length of 50. Based on this, the physical calibration length obtained from text recognition and the pixel length obtained from line segment recognition can be correlated to calculate the physical size mapping relationship of a single pixel. For example... Figure 2 The returned results show that the corresponding pixel-to-physical-size mapping results can be output.
[0115] In the image of the line segment region corresponding to the scale, the line segment usually has a certain width and may be affected by factors such as image noise, background grayscale changes, edge jaggedness, or local blurring. If the pixel length is directly calculated based on the width of the bounding rectangle of the line segment region image or the edge of the original line segment, it is easy to include the line segment width, background noise, or irregular edge parts in the pixel length, resulting in an inaccurate pixel length corresponding to the scale. Based on this, the embodiments of this application perform preprocessing, local thresholding, morphological optimization, and skeletonization processing on the line segment region image, transforming the line segment region from a strip-shaped region with a certain width into a single-pixel central axis that can represent the extension direction of the line segment. Then, the pixel length corresponding to the scale is determined based on the single-pixel skeleton of the line segment, thereby improving the accuracy and stability of the line segment measurement results.
[0116] Specifically, in an optional implementation, step S140 above uses a preset adaptive threshold segmentation model and morphological thinning model to extract line segments at the pixel level from the line segment region image, and determines the pixel length corresponding to the scale bar by reverse derivation of the sub-pixel level line segment endpoints, specifically including the following steps S141~S143:
[0117] S141. Perform fourth image preprocessing on the line segment region image to obtain the fourth target image.
[0118] The line segment region image can be a grayscale image or a color image. In one optional implementation, the fourth image preprocessing may include grayscale conversion and noise reduction. Grayscale conversion converts the color image into a single-channel grayscale image to reduce computational load and weaken the impact of color information on line segment segmentation, allowing the algorithm to primarily focus on brightness differences, edge gradient differences, and structural features between the line segment and the background. For scaled line segments, truly useful information is usually concentrated on the grayscale changes at the line segment edges and the main body of the line segment; therefore, a single-channel grayscale image is usually sufficient for subsequent processing needs.
[0119] Because wafer images may contain acquisition noise, background particles, compression noise, edge burrs, or local bright spots, direct thresholding may missegment these noises as foreground elements, resulting in numerous isolated points, holes, or broken areas. Therefore, median filtering can be used to remove salt-and-pepper noise, or Gaussian filtering can be used to suppress random Gaussian noise. In practical applications, a weak smoothing and edge-preserving processing strategy is typically adopted, which means preserving the gradient of line segments as much as possible while removing noise, avoiding blurring of line segment boundaries and unclear endpoint positions due to excessive smoothing.
[0120] S142. Input the fourth target image into the adaptive threshold segmentation model to perform adaptive threshold segmentation on the fourth target image based on each pixel in the fourth target image to obtain the fifth target image. Then, input the fifth target image into the morphological thinning model to perform morphological optimization and skeleton extraction on the foreground region to obtain the single-pixel skeleton of the line segment.
[0121] Specifically, the adaptive thresholding segmentation model divides the pixels in the fourth target image into foreground pixels or background pixels based on the gray-level distribution within the local neighborhood of each pixel, resulting in a binarized fifth target image. The foreground region, composed of foreground pixels, represents the main body of the scale line segment.
[0122] Since the foreground region of the fifth target image may still contain isolated noise, holes, or edge burrs, the fifth target image can be input into a morphological thinning model for morphological optimization to improve the integrity and contour continuity of the foreground region. Subsequently, the optimized foreground region is skeletonized, transforming a strip-shaped region of a certain width into a central axis with a width of a single pixel, resulting in a single-pixel line segment skeleton. The single-pixel line segment skeleton can represent the extension direction and pixel-level position of scale line segments, facilitating endpoint localization and pixel length calculation.
[0123] S143. Determine the two endpoint regions of the line segment based on the single-pixel skeleton of the line segment, and deduce the coordinates of the two sub-pixel endpoints based on the image information of the two endpoint regions. Determine the pixel length corresponding to the scale based on the coordinates of the two sub-pixel endpoints.
[0124] Specifically, when determining the two endpoint regions of a scale line segment based on its single-pixel skeleton, the approximate positions of the two ends of the scale line segment in the image can be determined based on the skeleton points in the single-pixel skeleton, and a local neighborhood range can be defined at each end position as the endpoint region. Since the single-pixel skeleton may have local offsets, discrete noise points, or slight bends due to segmentation errors or noise, directly using the original skeleton points to locate the ends may not be accurate enough. Therefore, based on the degree of matching between each skeleton branch in the single-pixel skeleton and the shape of the scale line segment, a target skeleton point set that reflects the extension direction of the main body of the line segment can be selected from the single-pixel skeleton. Then, a straight line fitting is performed on the target skeleton point set to determine the principal axis direction of the line segment, and the two endpoint regions are defined along this principal axis extension direction. Further, in one embodiment, RANSAC straight line fitting can be used to distinguish between valid points that conform to a linear distribution and outliers with large deviations from the skeleton points, and the approximate position of the line segment end in the image can be determined based on the fitted target straight line, thereby defining the two endpoint regions.
[0125] For each endpoint region, sub-pixel position estimation is performed using the grayscale distribution information of multiple pixels within it, to deduce the corresponding sub-pixel endpoint coordinates in reverse. For example, quadratic curve interpolation, Taylor series fitting, edge grayscale gradient distribution, or methods based on grayscale moments can be used to perform continuous domain fitting on the grayscale changes within the endpoint region, thereby determining the endpoint positions between adjacent integer pixels. Finally, the coordinate distance between the two sub-pixel endpoints is determined based on their coordinates, and this coordinate distance is used as the pixel length corresponding to the scale. By further locating sub-pixel endpoints on top of the coarse skeleton localization, the limitations of integer pixel coordinates and the influence of local line segment distortion and edge blurring on the pixel length measurement results can be reduced, improving the accuracy of determining the pixel length of the scale.
[0126] Specifically, in an optional implementation, step S142 above inputs the fourth target image into an adaptive threshold segmentation model to perform adaptive threshold segmentation on the fourth target image based on each pixel in the fourth target image to obtain a fifth target image, and inputs the fifth target image into a morphological thinning model to perform morphological optimization and skeleton extraction on the foreground region to obtain a single-pixel skeleton of a line segment, specifically including the following steps S1421~S1424:
[0127] S1421. Input the fourth target image into the adaptive threshold segmentation model, so that the local segmentation threshold corresponding to each pixel can be determined by the adaptive threshold segmentation model based on the neighborhood grayscale information of each pixel in the fourth target image.
[0128] After preprocessing, an adaptive thresholding model is used to separate the foreground and background of the line segment. Unlike fixed thresholding, the local segmentation threshold in this application is calculated based on the local neighborhood of each pixel, for example, 3×3 or 5×5 pixels. Specifically, the local segmentation threshold for a pixel can be obtained by subtracting an offset constant from the average gray value of several pixels surrounding it (i.e., neighborhood gray value information) or the local Gaussian weighted average. Then, the gray value of the pixel is compared with its local threshold to determine whether the pixel belongs to the foreground or background of the line segment. In this way, even if the line segment region has uneven brightness, slow changes in background gray value, or local shadows, the line segment region can be extracted relatively stably, ensuring the continuity and integrity of the line segment region and reducing breaks and large-area oversegmentation.
[0129] S1422. For each pixel in the fourth target image, the adaptive threshold segmentation model determines the current pixel as a foreground pixel or a background pixel based on the comparison result between the gray value of the current pixel and the local segmentation threshold corresponding to the current pixel, thus obtaining the fifth target image containing the foreground region and the background region.
[0130] Specifically, for each pixel in the fourth target image, its grayscale value can be compared with its corresponding local segmentation threshold. If the grayscale value is greater than or equal to the threshold, the pixel is marked as foreground, for example, mapped to white, representing a line segment region; if the grayscale value is less than the threshold, it is marked as background, for example, mapped to black, representing a non-line segment region. A region composed of multiple foreground pixels can be used as a foreground region, and a region composed of multiple background pixels can be used as a background region. The foreground region is used to represent the scale line segment region segmented from the fourth target image, and the background region is used to represent the region other than the scale line segments.
[0131] S1423. Input the fifth target image into the morphological thinning model, and perform erosion and dilation processing on the fifth target image through the morphological thinning model to remove isolated noise pixels in the fifth target image and perform contour smoothing processing on the foreground region in the fifth target image to obtain the sixth target image.
[0132] The binary image after adaptive thresholding, i.e., the fifth target image, may still contain isolated noise pixels, edge burrs, or uneven foreground contours. To eliminate these interferences, the fifth target image can be input into a morphological thinning model for morphological processing: first, an erosion operation is performed, where structuring elements (such as squares, rectangles, circles, or rhombuses, or elongated structuring elements matching the direction of line segments to better preserve long, strip-shaped targets like scale lines) are slid across the image. Only when the structuring element completely covers the foreground region are the center pixels retained, thus removing isolated noise and small burrs; then, a dilation operation is performed, which restores the size of the foreground region after erosion and improves local discontinuities and uneven contours, making the line segment contours smoother and more continuous. Through erosion and dilation, the impact of isolated noise generated after binarization on subsequent skeletonization processing can be reduced, and the contours of the foreground region can be made smoother and more continuous, resulting in the sixth target image.
[0133] S1424. Input the sixth target image into the morphological thinning model, and perform skeletonization processing on the foreground region in the sixth target image through the morphological thinning model to extract the central axis of the foreground region and obtain the single-pixel skeleton of the line segment.
[0134] Scale line segments in an image typically have a certain width. If the length is directly calculated based on the outer contour of the line segment region, the line width, edge blurring, and threshold selection will all affect the measurement results. In an optional implementation, a morphological thinning model integrating image thinning algorithms can be used to skeletonize the foreground region in the sixth target image. Specifically, the image thinning algorithm may include the Zhang-Suen thinning algorithm or the Guo-Hall thinning algorithm. In each iteration, boundary pixels that meet preset deletion conditions are identified and deleted. This process is repeated until no boundary pixels that meet the preset deletion conditions exist, thereby thinning the foreground region with a certain width into a connected central axis with a width of a single pixel, resulting in a single-pixel skeleton of the line segment. The preset deletion conditions may include at least one of the following: deleting the current boundary pixel does not change the connectivity of the foreground region, does not damage the end structure of the line segment, and does not affect the extension direction of the central axis.
[0135] When determining the pixel length corresponding to the scale based on the single-pixel skeleton of a line segment, the single-pixel skeleton may contain extra branches, discrete points, or local bends caused by image noise, scratches, or printing noise. If the pixel length is directly calculated based on the coordinates of the skeleton endpoints, endpoint positioning errors and abnormal skeleton points may affect the accuracy of the pixel length. Based on this, embodiments of this application further perform connected component analysis, line fitting, and sub-pixel endpoint positioning on the single-pixel skeleton of the line segment to determine the pixel length corresponding to the scale.
[0136] Specifically, in an optional implementation, step S143 above determines the two endpoint regions of the line segment based on the single-pixel skeleton of the line segment, and derives the coordinates of the two sub-pixel endpoints based on the image information of the two endpoint regions, and determines the pixel length corresponding to the scale based on the coordinates of the two sub-pixel endpoints, specifically including the following steps S1461~S1465:
[0137] S1461. Perform connected component analysis on the single-pixel skeleton of the line segment to determine the target skeleton point set from the single-pixel skeleton of the line segment.
[0138] Because the single-pixel skeleton obtained after thinning may contain redundant connected components (such as short lines, isolated points, or branches) caused by scratches or stains, it affects the accurate determination of the target skeleton point set. A connected component labeling algorithm is used to label all interconnected pixel regions in the skeleton image as independent connected components. Then, based on preset filtering conditions (such as the total pixel area of the connected components, the aspect ratio of the contour, the minimum bounding rectangle size, etc.), connected components belonging to the main line segments are identified and retained, while stray connected components with too small an area or abnormal shape are removed. This results in a skeleton point set containing only the target line segments, which serves as the target skeleton point set.
[0139] S1462. Perform linear fitting on the target skeleton point set to obtain the target linear model.
[0140] Since the target skeleton point set may still contain slight curvature, local offset, or discrete noise caused by image distortion, segmentation errors, or edge wear, these outliers may affect the fitting result if the ordinary least squares method is used to fit the line. Therefore, this application embodiment uses the Random Sample Consensus (RANSAC) algorithm for line fitting. Specifically, two points are randomly selected from the target skeleton point set as samples to fit an initial straight line model; a distance threshold is set, and the distance from all points to the line is calculated. Points with a distance less than the threshold are marked as inliers (valid points that conform to the straight line distribution), and the rest are outliers; random sampling is repeated multiple times (e.g., hundreds to thousands of times), and the straight line model with the most inliers and the smallest mean square error of the inliers is selected as the optimal fitted straight line, which is the target straight line model. This reduces the influence of discrete noise, curvature distortion, and other factors on the fitting result, and obtains the straight line that best represents the main direction of the line segment.
[0141] S1463. Determine the two endpoint regions of the line segment based on the target straight line model.
[0142] Specifically, all points in the target skeleton point set are projected onto the target straight line to obtain a one-dimensional coordinate sequence. On this one-dimensional coordinate sequence, the projection points corresponding to the minimum and maximum coordinate values are identified; these two projection points represent the approximate positions of the two endpoints of the line segment along the straight line. Centered on these two projection points, two local neighborhood windows (e.g., 5×5 or 7×7 pixel areas) are delineated in the sixth target image (or the fourth target image) as regions of interest for subsequent sub-pixel endpoint localization.
[0143] S1464. Based on the grayscale distribution information of multiple pixels in each endpoint region, a sub-pixel position estimation method is used to locate the sub-pixel endpoints in each endpoint region and determine the coordinates of the two sub-pixel endpoints of the line segment.
[0144] Traditional pixel-level endpoint localization typically represents endpoint positions using integer pixel coordinates, and its accuracy is limited by the pixel grid. To improve endpoint localization accuracy, based on the two endpoint regions determined by the single-pixel skeleton of the line segment, local image regions corresponding to each endpoint region can be determined in the fourth target image, and grayscale distribution information of multiple pixels can be extracted from these local image regions. Since the fourth target image retains the grayscale variation information of the scale line segment ends and their neighborhoods, the integer pixel-level endpoint positions can be further corrected based on the grayscale distribution information to obtain sub-pixel endpoint coordinates. In some implementations, endpoint localization accuracy can reach 0.1 to 0.5 pixels.
[0145] Specifically, grayscale values of multiple adjacent pixels within each endpoint region can be extracted from the fourth target image along the normal direction of the endpoint edge or along the extension direction of the scale line segment, and a grayscale change sequence or grayscale gradient sequence can be formed based on the grayscale values. Then, at least one sub-pixel position estimation method among quadratic curve interpolation, Taylor series fitting, edge grayscale gradient distribution, or grayscale moment calculation is used to fit or analyze the grayscale change sequence or grayscale gradient sequence to determine the continuous coordinate position of the endpoint edge between adjacent integer pixels.
[0146] Specific possible methods include: using quadratic curve interpolation based on the gray values within the endpoint region, i.e., fitting a quadratic function to the gray-level changes near the endpoint, and determining the endpoint position by finding the extreme points based on the fitting results; performing Taylor expansion on the gray-level gradient to solve for the sub-pixel offset; or using a gray-level moment-based method, analyzing the gray-level change pattern along the edge normal direction to deduce the edge position in the continuous domain, thereby obtaining the endpoint coordinates with sub-pixel precision. For example, in the vertical edge direction of the endpoint region, extracting the gray values of multiple adjacent pixels, fitting a gray-level change curve, and the position corresponding to the inflection point of the curve is the sub-pixel endpoint coordinate.
[0147] S1465. Calculate the coordinate distance between the two sub-pixel endpoints based on their coordinates, and determine the pixel length corresponding to the scale based on the coordinate distance.
[0148] Taking Euclidean distance as the coordinate distance as an example, after obtaining the sub-pixel coordinates of the left and right endpoints, the straight-line distance between the two points is calculated using the Euclidean distance formula: , where (x1, y1) and (x2, y2) are the coordinates of the two sub-pixel endpoints, respectively. This distance is the pixel length occupied by the scale line segment in the image.
[0149] Through the above processing, the embodiments of this application can determine the target skeleton point set used to characterize the main body of the scale line segment from the single-pixel skeleton of the line segment, and further determine the precise coordinates of the two ends of the line segment through line fitting and sub-pixel endpoint positioning. Compared with directly calculating the pixel length based on the width of the line segment bounding box or integer pixel endpoints, this method can reduce the influence of noise, local offset, line segment width and pixel discrete characteristics on the measurement results, and improve the accuracy and stability of the scale pixel length determination.
[0150] When determining the physical size mapping relationship of the image under test based on scale information, the mapping result depends on multiple intermediate results, such as scale region determination, text region determination, line segment region determination, physical calibration length determination, and pixel length determination. Deviation in any of these intermediate results may affect the accuracy of the final physical size mapping relationship and the determination of the physical size of the defect target. Especially when the scale text is small, the line segment edges are blurred, the image background is complex, or the image contains noise, simply outputting the final physical size value is insufficient to reflect the reliability of that value, hindering subsequent verification, screening, or traceability of the measurement results. Therefore, in an optional implementation, the wafer defect detection method of this application embodiment further includes the following steps S160~S162:
[0151] S160, Obtain at least one of the following: a first confidence level corresponding to the target scale detection box, a second confidence level corresponding to the target text detection box, a third confidence level corresponding to the target line segment detection box, a fourth confidence level corresponding to the physical calibration length, and a fifth confidence level corresponding to the pixel length.
[0152] The first confidence level is derived from the highest confidence score output by the first-level YOLOv5 model when detecting the scale region, reflecting the reliability of the coarse scale localization. The second and third confidence levels are derived from the confidence scores of the corresponding detection boxes output by the second-level YOLOv5 model when detecting text and line segment regions, respectively, reflecting the accuracy of text and line segment separation. The fourth confidence level can be the average recognition confidence level or CTC decoding probability of the PaddleOCR-CRNN model during character recognition, reflecting the reliability of the text recognition results. The fifth confidence level comes from indicators such as the proportion of interior points in the straight line fitting or the endpoint localization residual during sub-pixel line segment measurement, reflecting the credibility of the pixel length measurement results. In practical applications, all or some of the above confidence levels can be selected according to system requirements.
[0153] S161. Based on the obtained confidence levels and the preset weights corresponding to each confidence level, determine the mapping confidence level of the physical size mapping relationship corresponding to a single pixel.
[0154] Since different processing stages have varying degrees of influence on the final mapping result, a weighting coefficient needs to be pre-assigned to each confidence level. This coefficient can be determined through experimental calibration or based on the error sensitivity of each stage. For example, the weight of the confidence level for scale detection can be set to 0.2, the weights for text and line segment detection can each be 0.15, the weight for character recognition can be 0.3, and the weight for line segment measurement can be 0.2, with the sum of the weights being 1. Each obtained confidence score is multiplied by its corresponding weight, and the products are then summed to obtain the overall mapping confidence level. This mapping confidence level ranges from 0 to 1; a higher value indicates a higher reliability of the final physical size mapping result.
[0155] S162. If the mapping confidence is less than the preset confidence threshold, output a mapping error message.
[0156] The preset confidence threshold can be set according to the reliability requirements of the actual application, for example, 0.8. When the calculated mapping confidence is lower than this threshold, the system determines that the automatic mapping result of the current image is unreliable, which may be due to poor image quality, failure of scale area detection, text recognition errors, or abnormal line segment measurement. At this time, the system outputs mapping error prompt information (such as displaying a warning sign on the interface or generating a log record) to remind the quality inspection engineer to manually review the defective image or re-acquire the image. If the mapping confidence is higher than or equal to the threshold, the measurement result is accepted and used for subsequent defect analysis and yield evaluation.
[0157] The above describes a wafer defect detection method provided by an embodiment of this application. The following describes the apparatus for performing the above-described wafer defect detection method. Please refer to... Figure 3 , Figure 3This is a schematic diagram of a wafer defect detection device provided in an embodiment of this application. Figure 3 As shown, the wafer defect detection device includes:
[0158] The scale region extraction unit 301 is used to determine the scale region based on the image to be measured and extract the scale region image.
[0159] The scale information separation unit 302 is used to determine the text area image and line segment area image corresponding to the scale based on the scale area image.
[0160] The physical calibration length determination unit 303 is used to extract and recognize text in a text region image based on a text box positioning model and a text recognition model set in series, and to determine the physical calibration length corresponding to the scale.
[0161] The pixel length determination unit 304 is used to extract the pixel-level line segments of the line segment region image using a preset adaptive threshold segmentation model and morphological thinning model, and to determine the pixel length corresponding to the scale by reverse derivation of the sub-pixel-level line segment endpoints.
[0162] The defect size determination unit 305 is used to calculate the physical size mapping relationship of a single pixel in the current image under test based on the physical calibration length and pixel length, so as to determine the physical size of the defect target in the image under test based on the physical size mapping relationship.
[0163] Optionally, the scale region extraction unit 301 is specifically used for: performing a first image preprocessing on the image to be tested to obtain a first target image; inputting the first target image into a first detection model to obtain at least one scale candidate detection box and the confidence level corresponding to each scale candidate detection box; determining a target scale detection box from the at least one scale candidate detection box based on the confidence level corresponding to each scale candidate detection box; and cropping a scale region image from the image to be tested based on the target scale detection box.
[0164] Optionally, the scale information separation unit 302 is specifically used for: performing a second image preprocessing on the scale region image to obtain a second target image; inputting the second target image into a second detection model to obtain candidate detection boxes for text regions and candidate detection boxes for line segments; performing detection box deduplication on the candidate detection boxes for text regions and candidate detection boxes for line segments respectively, so as to determine the target text detection box in the candidate detection box for text regions and the target line segment detection box in the candidate detection box for line segments; and cropping the text region image and the line segment region image from the scale region image based on the target text detection box and the target line segment detection box respectively.
[0165] Optionally, the physical calibration length determination unit 303 is specifically used for: performing third image preprocessing on the text region image to obtain a third target image; inputting the third target image into a text box positioning model to locate the text region, and cropping a target text sub-image from the third target image based on the located text region; inputting the target text sub-image into a character recognition model to recognize characters and obtain a scale string; parsing the scale string to obtain numerical information and unit information; and determining the physical calibration length corresponding to the scale based on the numerical information and unit information.
[0166] Optionally, the pixel length determination unit 304 specifically includes:
[0167] The image preprocessing subunit is used to perform fourth image preprocessing on the line segment region image to obtain the fourth target image;
[0168] The pixel-level extraction subunit is used to input the fourth target image into the adaptive threshold segmentation model to perform adaptive threshold segmentation on the fourth target image based on each pixel in the fourth target image to obtain the fifth target image. The fifth target image is then input into the morphological thinning model to perform morphological optimization and skeleton extraction on the foreground region to obtain the single-pixel skeleton of the line segment.
[0169] The pixel length calculation subunit is used to determine the two endpoint regions of the line segment based on the single-pixel skeleton of the line segment, and to deduce the coordinates of the two sub-pixel endpoints based on the image information of the two endpoint regions, and to determine the pixel length corresponding to the scale based on the coordinates of the two sub-pixel endpoints.
[0170] Optionally, the pixel-level extraction subunit is specifically used for: inputting the fourth target image into an adaptive threshold segmentation model, so that the adaptive threshold segmentation model determines the local segmentation threshold corresponding to each pixel based on the neighborhood grayscale information of each pixel in the fourth target image; for each pixel in the fourth target image, the adaptive threshold segmentation model determines the current pixel as a foreground pixel or a background pixel based on the comparison result between the grayscale value of the current pixel and the local segmentation threshold corresponding to the current pixel, thereby obtaining a fifth target image containing foreground and background regions; inputting the fifth target image into a morphological thinning model, so that the morphological thinning model performs erosion and dilation processing on the fifth target image to remove isolated noise pixels in the fifth target image and performs contour smoothing processing on the foreground region in the fifth target image, thereby obtaining a sixth target image; inputting the sixth target image into a morphological thinning model, so that the morphological thinning model performs skeletonization processing on the foreground region in the sixth target image to extract the central axis of the foreground region, thereby obtaining a single-pixel skeleton of a line segment.
[0171] Optionally, the pixel length calculation subunit is specifically used for: performing connected component analysis on the single-pixel skeleton of the line segment to determine the target skeleton point set from the single-pixel skeleton of the line segment; performing line fitting on the target skeleton point set to obtain the target line model; determining the two endpoint regions of the line segment based on the target line model; using a sub-pixel position estimation method to locate the sub-pixel endpoints of each endpoint region based on the gray-level distribution information of multiple pixels in each endpoint region, and determining the coordinates of the two sub-pixel endpoints of the line segment; wherein, the sub-pixel position estimation method includes at least one of quadratic curve interpolation, Taylor series fitting, edge gray-level gradient distribution, or gray-level moment calculation; calculating the coordinate distance between the two sub-pixel endpoints based on the coordinates of the two sub-pixel endpoints, and determining the pixel length corresponding to the scale based on the coordinate distance.
[0172] Optionally, the device further includes a mapping confidence evaluation unit; the mapping confidence evaluation unit is specifically used to: obtain at least one of the following: a first confidence level corresponding to the target scale detection box, a second confidence level corresponding to the target text detection box, a third confidence level corresponding to the target line segment detection box, a fourth confidence level corresponding to the physical calibration length, and a fifth confidence level corresponding to the pixel length; determine the mapping confidence level of the physical size mapping relationship corresponding to a single pixel based on the obtained confidence levels and the preset weights corresponding to each confidence level; if the mapping confidence level is less than a preset confidence threshold, output a mapping abnormality prompt message.
[0173] This application also provides an electronic device in its embodiments. (See reference...) Figure 4 The diagram illustrates a structural schematic of an electronic device suitable for implementing the wafer defect detection method in the embodiments of this application. The electronic device in the embodiments of this application may include, but is not limited to, fixed terminals such as mobile phones, laptops, personal digital assistants (PDAs), tablet computers (PADs), desktop computers, etc. Figure 4 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0174] like Figure 4As shown, the electronic device may include a processing unit (e.g., a central processing unit (CPU), graphics processing unit (GPU), etc.) 401, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 402 or a program loaded from storage device 408 into random access memory (RAM) 403. When the electronic device is powered on, RAM 403 also stores various programs and data required for the operation of the electronic device. The processing unit 401, ROM 402, and RAM 403 are interconnected via bus 404. Input / output (I / O) interface 405 is also connected to bus 404.
[0175] Typically, the following devices can be connected to I / O interface 405: input devices 406 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 407 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 408 including, for example, memory cards, hard drives, etc.; and communication devices 409. Communication device 409 allows electronic devices to exchange data via wireless or wired communication with other devices. Although Figure 4 Electronic devices with various devices are shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0176] This application also provides a computer program product including computer-readable instructions, which, when executed on an electronic device, cause the electronic device to implement any of the wafer defect detection methods provided in this application.
[0177] This application also provides a computer-readable storage medium carrying one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement any of the wafer defect detection methods provided in this application.
[0178] It should also be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. In addition, in the accompanying drawings of the device embodiments provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be implemented as one or more communication buses or signal lines.
[0179] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware, or it can be implemented by special-purpose hardware including application-specific integrated circuits, special-purpose CPUs, special-purpose memory, special-purpose components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can be diverse, such as analog circuits, digital circuits, or special-purpose circuits. However, for this application, software program implementation is more often the preferred implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a computer floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disk, etc., and includes several instructions to cause an electronic device (which may be a personal computer, training equipment, or network device, etc.) to execute the methods of the various embodiments of this application.
[0180] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0181] A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disk, hard disk, magnetic tape), optical media (e.g., digital versatile disc (DVD)), or semiconductor media (e.g., solid state disk (SSD)).
Claims
1. A method for detecting wafer defects, characterized in that, include: Determine the scale area based on the image to be tested and extract the image of the scale area; Based on the scale area image, determine the text area image and line segment area image corresponding to the scale respectively; Based on the text box positioning model and the character recognition model set in series, the text region image is used to extract and recognize the characters, and the physical calibration length corresponding to the scale is determined. The pixel-level extraction of line segments in the image of the line segment region is performed using a preset adaptive threshold segmentation model and a morphological thinning model. The pixel length corresponding to the scale is determined by reverse derivation of the sub-pixel-level line segment endpoints. The physical size mapping relationship of a single pixel in the current image under test is calculated based on the physical calibration length and the pixel length, so as to determine the physical size of the defect target in the image under test based on the physical size mapping relationship.
2. The method according to claim 1, characterized in that, The step of determining the scale region based on the image to be tested and extracting the scale region image includes: The image to be tested is subjected to a first image preprocessing to obtain a first target image; The first target image is input into the first detection model to obtain at least one scale candidate detection box and the confidence level corresponding to each scale candidate detection box. A target scale detection box is determined from at least one of the scale candidate detection boxes based on the confidence level corresponding to each scale candidate detection box; Based on the target scale detection box, a scale region image is obtained by cropping from the image to be tested.
3. The method according to claim 1, characterized in that, The step of determining the text region image and line segment region image corresponding to the scale based on the scale region image includes: The scaled area image is subjected to a second image preprocessing to obtain a second target image; The second target image is input into the second detection model to obtain candidate detection boxes for text regions and candidate detection boxes for line segment regions. Deduplication is performed on the candidate detection boxes for the text region and the candidate detection boxes for the line segment region, respectively, so as to determine the target text detection box in the candidate detection box for the text region and the target line segment detection box in the candidate detection box for the line segment region. Based on the target text detection box and the target line segment detection box, text region images and line segment region images are respectively cropped from the scale region image.
4. The method according to claim 1, characterized in that, The text box positioning model and character recognition model based on the concatenated configuration extract and recognize characters from the text region image, and determine the physical calibration length corresponding to the scale bar, including: The text region image is subjected to third image preprocessing to obtain a third target image; The third target image is input into the text box positioning model to locate the text region, and the target text sub-image is cropped from the third target image based on the located text region. The target text sub-image is input into the character recognition model for character recognition to obtain the scale string; The scale string is parsed to obtain numerical and unit information; Based on the numerical information and the unit information, the physical calibration length corresponding to the scale is determined.
5. The method according to claim 1, characterized in that, The step of extracting line segments at the pixel level from the line segment region image using a preset adaptive threshold segmentation model and a morphological thinning model, and determining the pixel length corresponding to the scale bar by reverse derivation of sub-pixel level line segment endpoints, includes: The image of the line segment region is subjected to a fourth image preprocessing to obtain a fourth target image; The fourth target image is input into the adaptive threshold segmentation model to perform adaptive threshold segmentation on the fourth target image based on each pixel in the fourth target image to obtain the fifth target image. The fifth target image is then input into the morphological thinning model to perform morphological optimization and skeleton extraction on the foreground region to obtain a single-pixel skeleton of a line segment. The two endpoint regions of the line segment are determined based on the single-pixel skeleton of the line segment, and the coordinates of the two sub-pixel endpoints are derived in reverse based on the image information of the two endpoint regions. The pixel length corresponding to the scale is determined based on the coordinates of the two sub-pixel endpoints.
6. The method according to claim 5, characterized in that, The step involves inputting the fourth target image into the adaptive threshold segmentation model to perform adaptive threshold segmentation on the fourth target image based on each pixel in the fourth target image to obtain a fifth target image, and then inputting the fifth target image into the morphological thinning model to perform morphological optimization and skeleton extraction on the foreground region to obtain a single-pixel skeleton of a line segment, including: The fourth target image is input into the adaptive threshold segmentation model, so that the adaptive threshold segmentation model determines the local segmentation threshold corresponding to each pixel based on the neighborhood grayscale information of each pixel in the fourth target image. For each pixel in the fourth target image, the adaptive threshold segmentation model determines the current pixel as a foreground pixel or a background pixel based on the comparison result between the gray value of the current pixel and the local segmentation threshold corresponding to the current pixel, thus obtaining a fifth target image containing a foreground region and a background region. The fifth target image is input into the morphological thinning model, and the morphological thinning model is used to perform erosion and dilation processing on the fifth target image to remove isolated noise pixels and perform contour smoothing processing on the foreground region of the fifth target image to obtain the sixth target image. The sixth target image is input into the morphological thinning model, and the foreground region in the sixth target image is skeletonized by the morphological thinning model to extract the central axis of the foreground region and obtain a single-pixel skeleton of a line segment.
7. The method according to claim 5, characterized in that, The step of determining the two endpoint regions of the line segment based on the single-pixel skeleton of the line segment, and deriving the coordinates of the two sub-pixel endpoints based on the image information of the two endpoint regions, and determining the pixel length corresponding to the scale bar based on the coordinates of the two sub-pixel endpoints, includes: Perform connected component analysis on the single-pixel skeleton of the line segment to determine the target skeleton point set from the single-pixel skeleton of the line segment; A target straight line model is obtained by fitting the target skeleton point set with a straight line. Determine the two endpoint regions of the line segment based on the target straight line model; Based on the grayscale distribution information of multiple pixels within each endpoint region, a sub-pixel position estimation method is used to locate the sub-pixel endpoints of each endpoint region and determine the coordinates of the two sub-pixel endpoints of the line segment; wherein, the sub-pixel position estimation method includes at least one of quadratic curve interpolation, Taylor series fitting, edge grayscale gradient distribution, or grayscale moment calculation. The coordinate distance between the two sub-pixel endpoints is calculated based on the coordinates of the two sub-pixel endpoints, and the coordinate distance is used to determine the pixel length corresponding to the scale.
8. The method according to claim 1, characterized in that, The method further includes: Obtain at least one of the following: a first confidence level corresponding to the target scale detection box, a second confidence level corresponding to the target text detection box, a third confidence level corresponding to the target line segment detection box, a fourth confidence level corresponding to the physical calibration length, and a fifth confidence level corresponding to the pixel length; Based on the obtained confidence levels and the preset weights corresponding to each confidence level, the mapping confidence level of the physical size mapping relationship corresponding to the single pixel is determined; If the mapping confidence level is less than the preset confidence threshold, an abnormal mapping message will be output.
9. A wafer defect detection device, characterized in that, include: The scale region extraction unit is used to determine the scale region based on the image to be measured and extract the scale region image. The scale information separation unit is used to determine the text area image and line segment area image corresponding to the scale based on the scale area image. The physical calibration length determination unit is used to extract and recognize text in the text region image based on the serially set text box positioning model and text recognition model, and determine the physical calibration length corresponding to the scale. The pixel length determination unit is used to extract the line segments at the pixel level from the line segment region image using a preset adaptive threshold segmentation model and morphological thinning model, and to determine the pixel length corresponding to the scale by reverse derivation of the sub-pixel level line segment endpoints. The defect size determination unit is used to calculate the physical size mapping relationship of a single pixel in the current image under test based on the physical calibration length and the pixel length, so as to determine the physical size of the defect target in the image under test based on the physical size mapping relationship.
10. The wafer defect detection device according to claim 9, characterized in that, The pixel length determination unit includes: An image preprocessing subunit is used to perform a fourth image preprocessing on the line segment region image to obtain a fourth target image; A pixel-level extraction subunit is used to input the fourth target image into the adaptive threshold segmentation model to extract the foreground region composed of foreground pixels from the fourth target image to obtain the fifth target image, and input the fifth target image into the morphological thinning model to perform morphological optimization and skeleton extraction on the foreground region to obtain a single-pixel skeleton of a line segment. The pixel length calculation subunit is used to determine the two endpoint regions of the line segment based on the single-pixel skeleton of the line segment, and to deduce the coordinates of the two sub-pixel endpoints based on the image information of the two endpoint regions, and to determine the pixel length corresponding to the scale based on the coordinates of the two sub-pixel endpoints.
11. A computer storage medium, characterized in that, The storage medium carries one or more computer programs that, when executed by an electronic device, enable the electronic device to implement the wafer defect detection method as described in any one of claims 1 to 8.