Microchip Appearance Defect Detection Method Based on Convolutional Neural Network
By adopting adaptive brightness equalization, layered annotation and dual-branch lightweight convolutional neural network model in the appearance defect detection of microchips, combined with the defect sensitivity loss function, the problems of high detection complexity, low efficiency and insufficient model performance in the prior art are solved, and defect detection effects with high accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510200022.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-02-24
AI Technical Summary
The prior art has problems in the detection of appearance defects of microchips with high complexity defect characteristics and high-resolution microscopic image data are difficult to process, manual detection efficiency is low and susceptible to subjective factors, detection models based on ordinary machine learning methods lack the ability to detect small-size defects, and unbalanced sample number.
The microchip appearance defect detection method based on convolutional neural network is adopted, including adaptive brightness equalization processing of microscopic images, layered annotation strategy to train sample annotation, build and train a dual-branch lightweight convolutional neural network model, and optimize the model training using defect sensitivity loss function.
It improves the accuracy of appearance defect detection of microchips, significantly improves the ability to identify complex defect features, enhances the ability to detect small-size defects and different defect types, and solves the problem of unbalanced sample size.
Smart Images

Figure CN119693363B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image data processing, and in particular to a method for detecting appearance defects of a microchip based on a convolutional neural network. Background Art
[0002] In the prior art, the appearance defect detection of microchips usually relies on traditional image processing methods or manual detection methods. Traditional image processing methods mainly use static threshold segmentation, edge detection and other technologies to identify defects in microscopic images, while manual detection relies on experienced inspectors to observe defect characteristics with the naked eye. With the continuous improvement of chip manufacturing technology, the structural complexity and precision requirements of microchips have been significantly improved. The existing detection methods have gradually expanded to computer vision technology, and some methods have begun to introduce machine learning models for defect classification and detection.
[0003] However, when the existing technologies are applied to microchip appearance defect detection, the following problems exist: traditional image processing methods are difficult to cope with the highly complex defect features and the amount of data of high-resolution microscopic images; manual detection is inefficient and easily affected by subjective factors, and it is difficult to meet the real-time detection needs of large-scale production; the detection model based on ordinary machine learning methods has insufficient detection capabilities for small-sized defects and is difficult to effectively handle the problem of unbalanced number of samples of different defect types. In addition, these methods still have limitations in terms of edge feature positioning accuracy and defect classification robustness.
[0004] Therefore, it is necessary to develop a new method for detecting appearance defects of microchips. Summary of the invention
[0005] The present application provides a method for detecting appearance defects of microchips based on convolutional neural networks to improve the accuracy of appearance defect detection of microchips.
[0006] The present application provides a method for detecting appearance defects of microchips based on a convolutional neural network, comprising:
[0007] Acquire a high-resolution microscopic image of the microchip to be detected, and perform adaptive brightness equalization processing on the microscopic image to generate a pre-processed microscopic image, wherein the adaptive brightness equalization processing dynamically adjusts the brightness value of each pixel by calculating the brightness mean and variance of a local area of the image;
[0008] The pre-processed microscopic image is subjected to a layered annotation strategy to perform defect feature annotation of the training sample, wherein the annotation includes coarse-grained classification annotation based on the pre-processed microscopic image to classify the defects into surface deformation class and material abnormality class, and fine-grained positioning annotation combined with image features to generate annotated sample data containing specific defect types and their position coordinates;
[0009] Based on the labeled sample data, a dual-branch lightweight convolutional neural network model is constructed and trained, wherein the model includes a first branch and a second branch, wherein the first branch uses a depthwise separable convolution to extract global semantic features in the labeled sample, and the second branch uses a hole convolution to extract local detail features; wherein the dual-branch lightweight convolutional neural network model also includes an adaptive fusion gating unit for dynamically adjusting the fusion weights of the two branch features according to the scale features of different types of defects;
[0010] During the training process of the dual-branch lightweight convolutional neural network model, a defect sensitivity loss function is applied to optimize the training of the convolutional neural network model, wherein the defect sensitivity loss function includes: a loss term based on a defect area attention mechanism to improve the model's detection capability for small-size defect areas, a loss term based on edge feature retention to enhance the defect boundary positioning accuracy, and a loss term based on category balance to alleviate the impact of an imbalance in the number of samples of different defect types on the detection performance;
[0011] The preprocessed microscopic image to be inspected is input into the trained dual-branch convolutional neural network model to generate defect detection results for the microchip to be inspected, wherein the defect detection results include defect type, defect location coordinates and confidence score.
[0012] The beneficial effects of the technical solution provided by this application include:
[0013] (1) The contrast of the microscopic image is effectively optimized through adaptive brightness equalization processing. Combined with the global semantic feature extraction and local detail feature extraction capabilities of the dual-branch lightweight convolutional neural network model, the appearance defect detection results of microchips are more accurate, and the recognition ability of complex defect features is significantly improved. (2) The defect area attention mechanism in the defect sensitivity loss function highlights the focus on small-size defect areas, improves the model's ability to capture subtle defects, and solves the problem that traditional methods are difficult to detect small-size defects. (3) Through the edge preservation mechanism in the defect sensitivity loss function, while optimizing model training, the recognition and positioning capabilities of defect boundaries are enhanced, which helps to more accurately determine the scope and shape of defects. (4) The category balance mechanism in the defect sensitivity loss function alleviates the negative impact of the uneven distribution of the number of samples of different defect types on the detection performance, thereby improving the robustness and applicability of the model in various defect detection tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 This is a flowchart of a method for detecting appearance defects of microchips based on a convolutional neural network provided in the first embodiment of the present application. DETAILED DESCRIPTION
[0015] Many specific details are described in the following description to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of the present application. Therefore, the present application is not limited to the specific implementation disclosed below.
[0016] The first embodiment of the present application provides a method for detecting microchip appearance defects based on a convolutional neural network. Figure 1 , which is a schematic diagram of the first embodiment of the present application. Figure 1 A first embodiment of the present application provides a method for detecting appearance defects of microchips based on a convolutional neural network and is described in detail.
[0017] Step S101: obtaining a high-resolution microscopic image of the microchip to be detected, and performing adaptive brightness equalization processing on the microscopic image to generate a pre-processed microscopic image, wherein the adaptive brightness equalization processing dynamically adjusts the brightness value of each pixel by calculating the brightness mean and variance of the local area of the image.
[0018] In step S101, a high-resolution microscopic image of a microchip to be inspected is acquired and adaptive brightness equalization processing is performed on the image to generate a pre-processed microscopic image.
[0019] First, the microchip to be inspected is placed under a microscope for image acquisition. The microscope should be equipped with an imaging device capable of outputting high-resolution images. It is recommended that the resolution be at least 10 microns per pixel to meet the needs of detecting tiny defects. The focus distance and lighting intensity of the microscope need to be adjusted during image acquisition to ensure that the surface texture and edge details of the chip are clearly visible. It is preferred to use a ring-shaped LED light source with uniform lighting and resistance to light interference, while avoiding environmental conditions that produce excessive reflections or shadows.
[0020] After the image acquisition is completed, the original microscopic image is subjected to adaptive brightness equalization processing. This processing dynamically adjusts the brightness value of the pixel by analyzing the brightness characteristics of the local area of the image. Specifically, the image is divided into several local areas (such as small blocks of 10×10 pixels), and the mean and variance of the pixel brightness in each area are calculated to determine the brightness distribution characteristics of the current area. For areas with lower brightness, the pixel brightness value is appropriately increased; for areas with too high brightness, its brightness value is reduced through brightness compression.
[0021] In addition, in the process of brightness equalization, in order to prevent the introduction of image artifacts or noise due to excessive adjustment, the upper and lower thresholds of brightness adjustment can be set (for example, the brightness adjustment ratio is limited to between 0.5 and 1.5), and the image can be smoothed with a bilateral filter to retain edge details. After processing, the output pre-processed microscopic image not only has uniform brightness distribution, but also has a higher visual contrast, which can more effectively highlight the defect features on the chip surface, laying the foundation for subsequent defect feature annotation and detection.
[0022] Through the above steps, after the microscopic image is optimized, not only the overall image quality is improved, but also higher quality input data is provided for the feature extraction of the model, which helps to improve the accuracy and robustness of defect detection.
[0023] Furthermore, the performing adaptive brightness equalization processing on the microscopic image to generate a pre-processed microscopic image includes:
[0024] The following formula 1 is used to dynamically adjust the brightness value of each pixel of the microscopic image to generate a preprocessed microscopic image:
[0025]
[0026] in, To adjust the pixels in the preprocessed microscopic image The brightness value of Represents the horizontal coordinate of the pixel point, The vertical coordinate of the pixel is shown;
[0027] is the pixel point in the original microscopic image The brightness value of
[0028] Represents the original microscopic image in pixels The average brightness of the local area centered is used to measure the average brightness level of the pixel in its neighborhood. When calculating, usually select a Window (e.g. or ), calculate the average brightness of all pixels in the window:
[0029]
[0030] The local mean is used to adjust the local contrast to prevent detail loss caused by uneven brightness.
[0031] Represents the original microscopic image in pixels The standard deviation of brightness in the local area centered is used to describe the fluctuation of brightness distribution. The calculation formula is:
[0032]
[0033] The standard deviation can enhance the contrast of local details. If the standard deviation is large, it means that the brightness distribution in the area is uneven, and the local contrast needs to be enhanced during adjustment.
[0034] To prevent the denominator from being zero, the positive smoothing factor is usually taken as arrive .
[0035] is the global brightness mean of the entire image in the original microscopic image, which is used to adjust the balance of the overall brightness of the image; its calculation formula is:
[0036]
[0037] in and are the width and height (in pixels) of the image respectively. The global mean can correct the brightness distribution of the entire image and prevent the overall brightness from shifting due to local adjustments.
[0038] It is a global brightness control factor, which is used to control the amplitude of brightness adjustment. The recommended value range is 0.5 to 2.0, and it can be set dynamically according to the image contrast requirements.
[0039] It is the global balance coefficient, which is used to adjust the impact of global brightness contrast; the recommended value is 0.1 to 1.0.
[0040] The brightness smoothing factor controls the dynamic range of global brightness adjustment. The recommended value is 5 to 50.
[0041] The formula adjusts pixel brightness through two parts: the first part adjusts the local contrast based on the local mean and standard deviation to enhance image details; the second part balances the overall brightness distribution through the global brightness mean to avoid the problem of uneven brightness. The combination of the two parts can generate pre-processed microscopic images with more uniform brightness distribution and higher contrast, thereby significantly improving the accuracy and robustness of subsequent defect detection. Through parameter adjustment, the formula is applicable to different types of microscopic images and has a wide range of application value.
[0042] Step S102: A hierarchical annotation strategy is used to annotate the defect features of the training samples of the preprocessed microscopic images. The annotation includes coarse-grained classification annotation based on the preprocessed microscopic images to classify the defects into surface deformation and material anomaly categories, and fine-grained positioning annotation combined with image features to generate annotated sample data containing specific defect types and their location coordinates.
[0043] In step S102, a layered annotation strategy is used to annotate the defect features of the training samples for the preprocessed microscopic images. The layered annotation strategy includes two main stages: coarse-grained classification annotation and fine-grained positioning annotation, so as to generate high-quality annotated sample data and provide reliable data support for subsequent model training.
[0044] First, coarse-grained classification and annotation are performed based on the preprocessed microscopic images. The goal of this stage is to preliminarily classify the defects and clarify the major categories to which they belong. Specifically, the microscopic images are loaded into the annotation tool, and professional inspectors or semi-automatic annotation algorithms are used to classify the defects based on their macroscopic features (such as shape, texture, color, etc.). Common defect categories include surface deformation (such as scratches and dents) and material anomalies (such as impurities and holes). During the annotation process, polygons or rectangular frames can be drawn through interactive tools to encircle the defect area, and corresponding category labels can be assigned to each annotated area.
[0045] After completing the coarse-grained classification, the fine-grained positioning and annotation stage begins. In this stage, the coarse-grained annotation results are refined in combination with the local detail features of the image. The specific operation includes further determining the type and location of each labeled defect area. Fine-grained annotation requires accurate recording of the boundaries and geometric shapes of the defects, which can be achieved by drawing high-resolution polygonal contours or pixel-level segmentation masks. At the same time, fine-grained annotation also requires the extraction of specific feature parameters of the defects, such as area, aspect ratio, edge gradient value, etc., in order to fully describe the detailed features of the defects in the annotation data.
[0046] During the annotation process, attention should be paid to data consistency and high-quality control. To ensure the accuracy of annotation, it is recommended to introduce a multi-round review mechanism, where different annotators independently annotate the same image, and then unify the annotation results through comparison and correction. In addition, to avoid imbalance in the categories of sample data, diversified samples should be generated during the annotation process by adjusting the data collection ratio or using data enhancement techniques (such as rotation, scaling, mirroring, etc.).
[0047] The final output of the labeled sample data includes the category label of each defect area, the boundary position coordinates, the specific defect feature description, etc. The data is saved in a standard format (such as COCO, Pascal VOC or custom JSON format) for the subsequent construction and training of the dual-branch lightweight convolutional neural network model. The above-mentioned hierarchical labeling strategy can significantly improve the quality of the training sample data, laying a solid foundation for the performance optimization of the subsequent defect detection model.
[0048] Step S103: Based on the labeled sample data, construct and train a dual-branch lightweight convolutional neural network model, wherein the model includes a first branch and a second branch, wherein the first branch uses a deep separable convolution to extract global semantic features in the labeled samples, and the second branch uses a hole convolution to extract local detail features; wherein the dual-branch lightweight convolutional neural network model also includes an adaptive fusion gating unit for dynamically adjusting the fusion weights of the two branch features according to the scale characteristics of different types of defects.
[0049] In step S103, a dual-branch lightweight convolutional neural network model is constructed and trained based on the labeled sample data generated in step S102. Specifically, the model design includes two functional branches, namely, a first branch for extracting global semantic features and a second branch for extracting local detail features, and is equipped with an adaptive fusion gating unit to dynamically adjust the fusion weights of the two branch features, thereby achieving accurate detection of different types of defects.
[0050] First, in the process of model construction, the first branch extracts global semantic features through depthwise separable convolution operation. Depthwise separable convolution is a technology for optimizing convolution calculations. It effectively reduces the number of parameters and computational complexity by decomposing the standard convolution into two steps: depthwise convolution and pointwise convolution. In the specific operation, the depthwise convolution is performed independently on each channel, while the pointwise convolution fuses the information between channels through a 1×1 convolution kernel. This branch focuses on the extraction of macroscopic features in microchip images and is suitable for detecting defects involving larger areas.
[0051] At the same time, the second branch extracts local detail features through dilated convolution. Dilated convolution expands the receptive field by inserting holes (i.e. skipping a certain number of pixels) between convolution kernels while avoiding loss of resolution. This design enables the model to capture multi-scale information while retaining detailed features, making it highly sensitive to small-scale defects or complex textures.
[0052] In order to effectively combine the output features of the first branch and the second branch, the dual-branch lightweight convolutional neural network model introduces an adaptive fusion gating unit. This unit dynamically adjusts the fusion weights of the two branch features according to the scale characteristics of the defect. The specific implementation method includes designing a weighting mechanism that takes the output features of each branch as input and optimizes the information interaction between the two branches by learning dynamic weight coefficients. The fused features not only retain the advantages of global and local features, but also can achieve feature enhancement for specific defect types.
[0053] During the model training process, the input labeled sample data includes defect categories, location coordinates, and boundary features. These data guide the model to gradually optimize parameters through standard supervised learning. The training process uses batch processing to adapt to the large size and high resolution of microscopic images. In each training iteration, the training data is processed through the feature extraction path of the first branch and the second branch, and the generated feature map is integrated through the adaptive fusion unit, and then compared with the labeled data for error to adjust the model parameters.
[0054] Through this step, the dual-branch lightweight convolutional neural network model can make full use of the complementary advantages of global and local features, improve the detection accuracy and adaptability of microchip appearance defects, and provide a strong technical foundation for subsequent defect detection tasks.
[0055] Furthermore, the dual-branch lightweight convolutional neural network model includes an input layer, a first branch, a second branch, an adaptive fusion gating unit, and a classification and regression unit;
[0056] The input layer is used to receive the preprocessed microscopic image and standardize the received image data into a standardized microscopic image tensor suitable for model processing;
[0057] The first branch is used to receive the standardized microscopic image tensor provided by the input layer, and adopts a depth-separable convolution module to extract global semantic features in the microscopic image by separating the calculation of spatial convolution and channel convolution to obtain a global semantic feature map;
[0058] The second branch is used to receive the standardized microscopic image tensor provided by the input layer, and adopts a dilated convolution module to extract local detail features of the microscopic image by expanding the receptive field without increasing the number of parameters, thereby obtaining a local detail feature map;
[0059] The adaptive fusion gating unit is used to receive the global semantic feature map and the local detail feature map, and dynamically adjust the feature fusion weights of the first branch and the second branch according to the defect type and scale characteristics through the gating unit to generate a comprehensive feature map;
[0060] The classification and regression unit is used to receive the comprehensive feature map from the feature fusion module, and classify the defect type through the fully connected layer and the Softmax activation function; predict the specific location coordinates and confidence scores of the defects through the regression module; and generate defect detection results, which include defect types, defect location coordinates and confidence scores.
[0061] First, the input layer is used to receive the preprocessed microscopic image. The preprocessed microscopic image is normalized to adjust the pixel values to a uniform scale range (for example, normalized to [0, 1] or standardized to zero mean and unit variance) to meet the input requirements of the convolutional neural network. The output of the input layer is a standardized microscopic image tensor, whose shape is determined by the size and number of channels of the input image (for example, for a grayscale image it may be ,in and represents the height and width of the image respectively).
[0062] The first branch receives the normalized microscopic image tensor from the input layer and uses a depthwise separable convolution module to extract global semantic features. The depthwise separable convolution module splits the traditional convolution operation into two independent computational stages: depthwise convolution and pointwise convolution. The depthwise convolution performs spatial convolution operations separately in each channel to extract spatial features, while the pointwise convolution is fused between all channels to use Convolution achieves the integration of cross-channel features. This design significantly reduces the computational complexity and the number of parameters while retaining the integrity of the global semantic features. The output of the first branch is a global semantic feature map that captures large-scale defect information in microscopic images, such as macroscopic surface deformation.
[0063] The second branch also receives the standardized microscopic image tensor provided by the input layer, but uses a dilated convolution module to extract local detail features. Dilated convolution expands the receptive field by introducing a dilated rate parameter in the convolution kernel, so that a wider range of contextual information can be captured without increasing the number of convolution kernel parameters. Multiple dilated rate settings (such as 1, 2, and 4) can form multi-scale perception capabilities, thereby better capturing the tiny detail features of the defect area, such as cracks or point defects. The output of the second branch is a local detail feature map, which is mainly used to describe the fine structure information in the microscopic image.
[0064] The adaptive fusion gating unit receives the global semantic feature map from the first branch and the local detail feature map from the second branch, and uses the gating mechanism to dynamically adjust the feature fusion weight. The gating unit calculates the weight ratio of different features according to the content of the input feature map (such as defect type, size and shape) to ensure that the fused comprehensive feature map can retain both global and local information. This adaptive mechanism is usually implemented using an attention mechanism, which accurately controls the contribution of each feature to the final comprehensive feature by calculating channel weights and spatial weights. The fused comprehensive feature map has multi-scale perception capabilities, which can support subsequent modules to complete classification and positioning tasks more accurately.
[0065] The classification and regression unit receives the comprehensive feature map from the adaptive fusion gating unit. The classification part maps the comprehensive features through the fully connected layer and normalizes the prediction results to the category probability distribution through the Softmax activation function to achieve the classification of defect types. The regression part predicts the specific location coordinates of the defect (such as the coordinates of the bounding box or the center point) and the confidence score through another branch to quantify the credibility of the model for the defect detection results. The final output includes the defect type, defect location coordinates and confidence score, which can provide comprehensive results for microchip appearance defect detection.
[0066] Through the collaborative work of the above modules, the dual-branch lightweight convolutional neural network model realizes efficient and accurate detection of microchip appearance defects. It has high computational efficiency and detection accuracy and can adapt to a variety of complex defect detection scenarios.
[0067] Furthermore, the depthwise separable convolution module in the first branch includes an alternating stack of multiple depthwise convolutional layers and pointwise convolutional layers, wherein the depthwise convolutional layers are used to extract spatial features for each channel respectively, and the pointwise convolutional layers are used to perform feature fusion between all channels; by adjusting the number and parameter size of the depthwise convolutional layers and the pointwise convolutional layers, the extraction of global semantic features is ensured while reducing the amount of computation and retaining significant semantic information.
[0068] The deep convolution layer is one of the core components of this module, which is used to extract the feature information of the microscopic image in the spatial dimension. In the deep convolution operation, each input channel performs a convolution operation separately, and the convolution kernel only acts on the pixels of the current channel without interacting with other channels. This method significantly reduces the amount of computation required for the convolution operation while retaining the spatial feature information of each channel. The convolution kernel size of the deep convolution can be adjusted according to the resolution and feature distribution of the input image, and is usually set to or , to strike a balance between computational efficiency and feature extraction capability.
[0069] The function of the point convolution layer is to fuse the features output by the deep convolution layer between channels, thereby generating comprehensive features with global semantic information. The convolution kernel linearly combines all channels of each pixel to achieve information interaction and feature integration between channels. This method not only retains the spatial information extracted by deep convolution, but also significantly reduces the number of parameters and computational complexity. The output of point convolution is nonlinearly mapped through an activation function (such as ReLU) to enhance the expressiveness of the model.
[0070] The deep convolutional layers and point convolutional layers are arranged in an alternating stacking manner to form multiple stages, each of which can extract global semantic features of microscopic images at different levels. Through this alternating stacking design, the model can gradually enhance the perception of macro features while maintaining low computational overhead. The number of deep convolutional layers and point convolutional layers in each stage can be adjusted according to specific task requirements. For example, when the input resolution is high, the number of deep convolutional layers can be increased to extract more spatial features; when stronger channel fusion capabilities are required, the number of point convolutional layers can be increased.
[0071] In order to ensure that the extraction process of global semantic features is both efficient and accurate, the parameter sizes of the deep convolution layer and the point convolution layer also need to be set reasonably. For example, the number of channels of the deep convolution is usually kept the same as the input, while the number of output channels of the point convolution can be adjusted according to the model objectives and hardware limitations. In addition, regularization techniques such as batch normalization and dropout can be combined to further improve the robustness and generalization ability of the model.
[0072] Through the above design and optimization, the depthwise separable convolution module in the first branch can efficiently extract global semantic features while significantly reducing the amount of computation and parameters, providing strong support for subsequent feature fusion and classification tasks.
[0073] Furthermore, the deep convolution layer in the first branch introduces a mechanism of dynamic convolution kernel size to adaptively adjust the receptive field size of the convolution kernel according to the resolution of the input feature map, wherein the dynamic adjustment is determined in real time based on the resolution and complexity of the feature map. At low resolution, a smaller convolution kernel is used to reduce the amount of calculation, and at high resolution, a larger convolution kernel is used to enhance the semantic feature extraction capability, thereby further optimizing the computational efficiency and the accuracy of global feature extraction.
[0074] The core function of the deep convolutional layer is to perform convolution operations independently on each channel of the input feature map to extract spatial features. The traditional fixed convolution kernel size may have limitations when processing images of different resolutions or complexities. For example, a fixed-size convolution kernel may not be able to effectively capture global features when processing high-resolution images, and may introduce unnecessary computational burden when processing low-resolution images. To overcome this problem, the dynamic convolution kernel size mechanism adjusts the size of the convolution kernel in real time to adapt it to the characteristics of the input feature map.
[0075] In low-resolution feature maps, the dynamic mechanism selects smaller convolution kernels, such as 3×3 or 5×5, to reduce the amount of computation while extracting relatively compact spatial features. Smaller convolution kernels can quickly focus on local features with a small receptive field, thereby improving computational efficiency. In high-resolution feature maps, the dynamic mechanism selects larger convolution kernels, such as 7×7 or 9×9, to expand the receptive field and capture more global feature information. This adjustment method can effectively balance the computational requirements and feature extraction accuracy in high-resolution image processing.
[0076] The basis for dynamic adjustment includes the resolution and complexity of the feature map. During the execution of the model, the characteristics of the input image can be determined by simple statistical methods (such as calculating the width and height of the feature map) or complex analytical methods (such as calculating the texture complexity or entropy value of the feature map). The dynamic mechanism selects the appropriate convolution kernel size based on these characteristic parameters, allowing the network to flexibly adapt to different types of input data.
[0077] Furthermore, the atrous convolution module in the second branch includes multiple convolution layers with different atrous rates, and convolution layers with atrous rates of 1, 2 and 4 are set in parallel to form a multi-scale feature extraction capability, wherein the convolution results of different atrous rates are fused by pixel-by-pixel summation to further enhance the integrity of local detail features.
[0078] Dilated convolution is a convolution operation that expands the receptive field by inserting holes (discontinuous sampling points) in the convolution kernel without increasing the number of parameters or the amount of computation. In the dilated convolution module, convolution layers with dilation rates of 1, 2, and 4 are set. The convolution layer with a dilation rate of 1 is equivalent to the standard convolution and is used to capture the detail information of the defect. The convolution layer with a dilation rate of 2 samples at intervals of one pixel in the convolution kernel, which expands the receptive field and is suitable for extracting contextual features of smaller areas. The convolution layer with a dilation rate of 4 covers a wider area with a larger sampling interval and is used to extract feature information of larger scales. With this design, the dilated convolution module can fully describe the multi-scale characteristics of defects.
[0079] The convolutional layers with the above-mentioned multiple dilation rates are set in parallel, which means that the input microscopic image feature map is simultaneously processed by convolutions with different dilation rates to generate corresponding feature maps respectively. This parallel processing method can improve the feature extraction efficiency of the model and avoid the feature information that may be lost due to a single convolution scale.
[0080] To integrate the features extracted by the convolutional layers with different dilation rates, the module fuses these feature maps by pixel-wise summation. Pixel-wise summation means that at each pixel position, the output feature values of all convolutional layers are superimposed. This fusion method can maintain the spatial consistency of the features and effectively integrate multi-scale information. In addition, the fused feature map can more completely reflect the edge and texture detail information of the defect area, improving the integrity and expression ability of local detail features.
[0081] This dilated convolution module design is particularly suitable for scenarios in microchip defect detection. Defects in microscopic images often have significant scale diversity, ranging from micron-sized cracks to deformations in local surface areas. Through the processing of the multi-dilation rate convolution module, the details and context features of these defects can be captured simultaneously, ensuring the accuracy and robustness of the detection results. In addition, since dilated convolution does not require an increase in computational complexity, the module design meets the lightweight computational requirements while ensuring the detection performance, making it very suitable for deployment on resource-limited embedded devices.
[0082] Furthermore, the dilated convolution module in the second branch introduces a dynamic adjustment mechanism for the dilation rate. According to the size distribution of the defect area in the microscopic image, it adaptively selects an appropriate combination of dilation rates. The dynamic adjustment of the dilation rate is completed by an image pre-analysis module, which generates a dilation rate parameter configuration suitable for the current image characteristics through the boundary information and morphological features of the defect, achieving the ability to optimize the extraction of multi-scale defect areas and reducing redundant features at the same time.
[0083] The dynamic adjustment of the dilation rate is achieved through an image pre-analysis module. This module first performs preprocessing analysis on the input microscopic image, including detecting the boundary information and morphological features of the defect. The extraction of boundary information can be achieved through edge detection algorithms (such as Sobel operator, Canny operator) to identify the regions with obvious brightness changes in the microscopic image and preliminarily determine the contour and distribution range of the defect. The analysis of morphological features includes evaluating the size, shape, and distribution density of the defect area. These information provide a direct basis for the dynamic adjustment of the dilation rate.
[0084] When determining the void ratio, the image pre-analysis module generates a void ratio combination configuration suitable for the current image based on the size distribution of the defect area. For example, for defect areas with smaller sizes, the module will give priority to smaller void ratios (such as 1 or 2) to ensure that the receptive field covers a more compact spatial area, thereby more accurately capturing the detailed features of the defect. For defect areas with larger sizes, the module will select a larger void ratio (such as 4 or 6) to expand the receptive field range and more comprehensively extract the global information of the defect area. Through this adaptive adjustment, the model can take into account the feature extraction requirements of both small and large defects.
[0085] The dynamic adjustment of the dilation rate not only optimizes the range of feature extraction, but also effectively reduces the generation of redundant features. Traditional dilation convolution with a fixed dilation rate may introduce irrelevant or repeated feature information in some areas, increasing the computational burden and model complexity. Through the dynamic adjustment mechanism, dilation convolution only generates feature maps required for the current image characteristics, significantly improving the efficiency and accuracy of feature extraction.
[0086] In addition, this dynamic adjustment mechanism can learn the void ratio configuration strategy suitable for different defect scenarios during the training phase of the convolutional neural network. During the inference phase, the image pre-analysis module quickly analyzes and configures the input image according to the trained adjustment rules, thereby ensuring the real-time and robustness of the adjustment mechanism. The void convolution module combined with the dynamic adjustment mechanism is particularly suitable for processing complex and diverse defect features in microscopic images, and can further improve the accuracy and efficiency of microchip appearance defect detection without increasing the demand for computing resources.
[0087] Furthermore, the adaptive fusion gating unit jointly calculates the fusion weight based on the spatial information and channel information of the feature map, specifically including: calculating the weight of each channel through the channel attention mechanism, calculating the weights of different spatial positions through the spatial attention mechanism, and dynamically adjusting the fusion ratio of the first branch and the second branch features according to the channel weight and the spatial weight, so as to adapt to the scale characteristics of different types of defects.
[0088] In the adaptive fusion gating unit, the channel attention mechanism is used to assign weights to each channel of the feature map. This mechanism dynamically determines the contribution of each channel to the overall detection task by analyzing the importance of the features contained in each channel. The implementation usually includes global average pooling and global maximum pooling for each channel of the input feature map, extracting global statistics at the channel level, and then generating channel weights through a lightweight fully connected layer and activation function (such as Sigmoid). The allocation of channel weights ensures that the model can pay more attention to channels with significant features, such as semantic channels containing key defect information.
[0089] The spatial attention mechanism focuses on evaluating the importance of different spatial positions in the feature map, thereby capturing important areas of local features. The implementation method includes performing a global pooling operation on the feature map in the channel direction (i.e. calculating the average and maximum value of each pixel in all channels) to form a spatial attention map. The spatial attention map is further processed by convolution operations to obtain spatial weights at different positions. In this way, the model can focus more accurately on spatial areas where defects may appear, such as small cracks or edge discontinuities.
[0090] During the fusion process, the adaptive fusion gating unit combines the channel weight and the spatial weight, and dynamically adjusts the fusion ratio of the first branch and the second branch features through weighted operations. The calculation process of the fusion ratio can be regarded as a dynamic optimization process to ensure that the global semantic features and local detail features are reasonably expressed in the final comprehensive features. Feature fusion can be achieved not only through simple weighted summation, but also by combining dot multiplication operations to increase the interaction between features and enhance feature expression capabilities.
[0091] This design combines information from both channel and space dimensions, rather than relying solely on feature selection in a single dimension. The introduction of this dual attention mechanism makes the model more adaptable to different types of defects. For example, for large-scale surface deformations, the weight of global semantic features may be greater, while for tiny point defects or edge cracks, local detail features need to be given higher weights.
[0092] In addition, the implementation of the adaptive fusion gating unit can be combined with embedded hardware optimization strategies, such as using deep separable convolution or low computational complexity attention modules to reduce the computational overhead in practical applications. This design not only improves the accuracy of feature fusion, but also takes into account the optimization of computational efficiency, making this method very suitable for deployment in resource-sensitive devices.
[0093] Through this precise design, the adaptive fusion gating unit achieves efficient fusion of global and local features, significantly improving the detection performance of the model when dealing with complex and diverse defects, and ensuring efficient application in microscopic image environments. The model can dynamically adapt to different defect scenarios, providing more powerful detection capabilities and robustness, and providing a practical technical solution for microchip appearance defect detection.
[0094] Furthermore, the classification and regression unit jointly optimizes the classification and regression tasks by introducing a multi-task learning strategy, wherein the loss function of the classification task is a cross-entropy loss based on category balance, and the loss function of the regression task is a weighted square error loss. The classification loss and the regression loss are combined into a total loss according to preset weights, ensuring that the model is simultaneously optimized in terms of the accuracy of defect type classification and location prediction.
[0095] In the classification task, the goal of the model is to classify the detected defects into specific types, such as surface deformation, material anomalies, or other types of defects. The loss function for the classification task adopts the cross-entropy loss based on class balance. Since different types of defects may have significant imbalances in the data distribution, such as fewer samples of certain types of defects and more samples of other types, the traditional cross-entropy loss may cause the model to prefer large-class samples. To this end, the cross-entropy loss based on class balance adjusts the calculation of the loss value by introducing a weight coefficient for each defect category, so that the model can focus on minority class samples and improve the classification performance of unbalanced data. These weights can be calculated based on the frequency of each class of samples, such as the weight is proportional to the inverse of the number of class samples, thereby balancing the impact of different classes on the total loss.
[0096] In the regression task, the goal of the model is to accurately predict the location of the defect and the corresponding confidence score. The loss function of the regression task is the weighted square error loss, which is used to measure the deviation between the model's predicted location and the actual location. The introduction of weight coefficients in this loss function allows the prediction error of key locations (such as the defect boundary area) to receive higher attention. For example, weights can be dynamically assigned based on the importance of the location. For example, the weight of the defect center area can be set higher, while the weight of the area outside the boundary is relatively low, thereby improving the accuracy of location prediction.
[0097] The losses of the classification task and the regression task are combined into a total loss function through preset weights to guide the optimization of the entire model. The weight setting can be adjusted according to the specific task requirements. For example, in the defect detection task, the classification task and the regression task can be given weights of 0.6 and 0.4 respectively to ensure that the model strikes a balance between identifying the defect type and locating the defect location. This weight setting can also be automatically optimized during the training phase through hyperparameter search to further improve model performance.
[0098] In the multi-task learning strategy, the classification task and the regression task share the same feature extraction network, but each task is completed through independent branches. The classification branch classifies the defect type through a fully connected layer and a Softmax activation function, while the regression branch outputs the coordinate value and confidence score of the defect location through a continuous fully connected layer and a linear activation function. This shared feature design not only reduces the number of parameters in the model, but also enables the classification and regression tasks to promote each other. For example, the results of the classification branch can provide guidance for the regression branch, allowing it to pay more attention to specific defect areas when predicting locations.
[0099] The following is the reference implementation code of the dual-branch lightweight convolutional neural network model:
[0100] import torch
[0101] import torch.nn as nn
[0102] import torch.nn.functional as F
[0103] # Define a depth-wise separable convolution module
[0104] class DepthwiseSeparableConv(nn.Module):
[0105] def __init__(self, in_channels, out_channels, kernel_size, stride=1,padding=0):
[0106] super(DepthwiseSeparableConv, self).__init__()
[0107] # Deep convolutional layer: operate independently on each channel
[0108] self.depthwise = nn.Conv2d(in_channels, in_channels, kernel_size=kernel_size,
[0109] stride=stride, padding=padding, groups=in_channels)
[0110] # Point convolution layer: fuse all channels
[0111] self.pointwise = nn.Conv2d(in_channels, out_channels, kernel_size=1)
[0112] def forward(self, x):
[0113] # Perform depth convolution first, then point convolution
[0114] x = self.depthwise(x)
[0115] x = F.relu(x) # activation function
[0116] x = self.pointwise(x)
[0117] return x
[0118] # Dynamically adjust the convolution kernel size
[0119] class DynamicDepthwiseConv(nn.Module):
[0120] def __init__(self, in_channels, out_channels, base_kernel_size):
[0121] super(DynamicDepthwiseConv, self).__init__()
[0122] self.base_kernel_size = base_kernel_size
[0123] self.dynamic_convs = nn.ModuleList([
[0124] DepthwiseSeparableConv(in_channels, out_channels, kernel_size=k,padding=k / / 2)
[0125] for k in [base_kernel_size, base_kernel_size + 2, base_kernel_size +4] ])
[0127] def forward(self, x):
[0128] # Dynamically select the convolution kernel size based on the input resolution
[0129] res = x.size(2) # Input feature map resolution
[0130] if res < 64:
[0131] return self.dynamic_convs[0](x)
[0132] elif res < 128:
[0133] return self.dynamic_convs[1](x)
[0134] else:
[0135] return self.dynamic_convs[2](x)
[0136] # Define the dilated convolution module
[0137] class AtrousConvBlock(nn.Module):
[0138] def __init__(self, in_channels, out_channels, dilations):
[0139] super(AtrousConvBlock, self).__init__()
[0140] # Convolution with different dilation rates
[0141] self.convs = nn.ModuleList([
[0142] nn.Conv2d(in_channels, out_channels, kernel_size=3, padding=d,dilation=d)
[0143] for d in dilations ])
[0145] def forward(self, x):
[0146] # Perform multiple hole-rate convolutions in parallel and sum them pixel by pixel
[0147] features = [conv(x) for conv in self.convs]
[0148] return sum(features) # Merge feature maps
[0149] # Adaptive Fusion Gating Unit
[0150] class AdaptiveFusionGate(nn.Module):
[0151] def __init__(self, in_channels):
[0152] super(AdaptiveFusionGate, self).__init__()
[0153] # Channel Attention Mechanism
[0154] self.channel_attention = nn.Sequential(
[0155] nn.AdaptiveAvgPool2d(1),
[0156] nn.Conv2d(in_channels, in_channels / / 8, kernel_size=1),
[0157] nn.ReLU(),
[0158] nn.Conv2d(in_channels / / 8, in_channels, kernel_size=1),
[0159] nn.Sigmoid() )
[0161] # Spatial Attention Mechanism
[0162] self.spatial_attention = nn.Sequential(
[0163] nn.Conv2d(2, 1, kernel_size=7, padding=3),
[0164] nn.Sigmoid() )
[0166] def forward(self, global_features, local_features):
[0167] # Channel weights
[0168] ca = self.channel_attention(global_features + local_features)
[0169] # Spatial weights
[0170] max_pool = torch.max(global_features + local_features, dim=1, keepdim=True)[0]
[0171] avg_pool = torch.mean(global_features + local_features, dim=1,keepdim=True)
[0172] sa_input = torch.cat([max_pool, avg_pool], dim=1)
[0173] sa = self.spatial_attention(sa_input)
[0174] # Feature Fusion
[0175] fused_features = ca * sa * (global_features + local_features)
[0176] return fused_features
[0177] # Classification and regression unit
[0178] class ClassificationAndRegressionHead(nn.Module):
[0179] def __init__(self, in_channels, num_classes):
[0180] super(ClassificationAndRegressionHead, self).__init__()
[0181] # Classification header
[0182] self.classifier = nn.Sequential(
[0183] nn.Conv2d(in_channels, in_channels / / 2, kernel_size=3, padding=1),
[0184] nn.ReLU(),
[0185] nn.Conv2d(in_channels / / 2, num_classes, kernel_size=1) )
[0187] # Return header
[0188] self.regressor = nn.Sequential(
[0189] nn.Conv2d(in_channels, in_channels / / 2, kernel_size=3, padding=1),
[0190] nn.ReLU(),
[0191] nn.Conv2d(in_channels / / 2, 4, kernel_size=1) # Assume that 4 regression values are output: x, y, w, h )
[0193] def forward(self, x):
[0194] # Output classification probability and regression results
[0195] classification = self.classifier(x)
[0196] regression = self.regressor(x)
[0197] return classification, regression
[0198] # Main model implementation
[0199] class DualBranchCNN(nn.Module):
[0200] def __init__(self, in_channels, num_classes):
[0201] super(DualBranchCNN, self).__init__()
[0202] # Input layer
[0203] self.input_layer = nn.Conv2d(in_channels, 64, kernel_size=3, padding=1)
[0204] # First branch (global semantics)
[0205] self.global_branch = nn.Sequential(
[0206] DynamicDepthwiseConv(64, 128, base_kernel_size=3),
[0207] DepthwiseSeparableConv(128, 256, kernel_size=3, padding=1) )
[0209] # Second branch (local details)
[0210] self.local_branch = AtrousConvBlock(64, 256, dilations=[1, 2, 4])
[0211] # Adaptive Fusion
[0212] self.fusion_gate = AdaptiveFusionGate(256)
[0213] # Classification and regression heads
[0214] self.head = ClassificationAndRegressionHead(256, num_classes)
[0215] def forward(self, x):
[0216] # Input layer processing
[0217] x = self.input_layer(x)
[0218] # Branch feature extraction
[0219] global_features = self.global_branch(x)
[0220] local_features = self.local_branch(x)
[0221] # Adaptive feature fusion
[0222] fused_features = self.fusion_gate(global_features, local_features)
[0223] # Classification and regression output
[0224] classification, regression = self.head(fused_features)
[0225] return classification, regression
[0226] The steps for training a dual-branch lightweight convolutional neural network model are as follows: First, prepare a preprocessed microscopic image dataset, including training data with defect type, location, and size annotated, and divide it into a training set and a validation set according to a certain ratio. Define the model structure, initialize the model parameters, and select a suitable optimizer (such as Adam) and learning rate. Define a cross-entropy loss function based on class balance for the classification task and a weighted square error loss function for the regression task, and combine the two into a total loss function according to preset weights. In each training iteration, pass the input images to the model in batches, calculate the classification and regression losses separately, and then calculate the total loss and optimize the model parameters through back propagation. Use the validation set to monitor the model performance and adjust the hyperparameters (such as learning rate and loss weight) until the training converges. After training, use the test set to evaluate the classification and localization accuracy of the model.
[0227] Step S104: During the training process of the dual-branch lightweight convolutional neural network model, a defect sensitivity loss function is applied to optimize the training of the convolutional neural network model, wherein the defect sensitivity loss function includes: a loss term based on the defect area attention mechanism to improve the model's detection capability for small-size defect areas, a loss term based on edge feature retention to enhance the defect boundary positioning accuracy, and a loss term based on category balance to alleviate the impact of an imbalance in the number of samples of different defect types on the detection performance.
[0228] In step S104, during the training process of the dual-branch lightweight convolutional neural network model, in order to optimize the model performance, the defect sensitivity loss function is used to optimize the model. Specifically, the defect sensitivity loss function is composed of multiple loss items designed for specific detection difficulties, and these loss items work together to improve the model's ability to detect microchip appearance defects.
[0229] First, in order to enhance the model's ability to detect small-sized defect areas, a loss term based on the defect area attention mechanism is designed. This loss term guides the model to pay more attention to these areas during training by assigning higher weights to defect areas in the image. In specific implementations, a pixel-level weighting mechanism can be used to separate pixel areas marked as defects in preprocessed microscopic images from background areas, and give additional amplification factors to the prediction errors of these areas. This factor is usually set to several times the background weight, such as 3 to 5 times, to ensure that the model has a higher sensitivity to small and difficult-to-identify defects.
[0230] Secondly, in order to improve the positioning accuracy of the model for defect boundaries, a loss term based on edge feature preservation is designed. This loss term encourages the model to generate more accurate prediction results at the boundary by calculating the difference between the defect area boundary predicted by the model and the actual boundary. Specifically, a boundary-aware loss function can be used, such as a gradient calculation method based on the Sobel operator or the Laplacian operator, to extract the boundary features of the predicted results and the actual annotations, and calculate the similarity error between the two in the gradient space. This loss term can effectively reduce detection errors caused by blurred or over-smoothed boundaries.
[0231] In addition, in order to solve the adverse effect of uneven distribution of samples of different defect types on model training, a loss term based on class balance is introduced. This loss term re-weights the class loss during training to ensure that defect classes with fewer samples will not affect the model's learning ability due to their lower weights. In specific implementation, a weight allocation mechanism based on class frequency can be adopted, that is, weights are allocated according to the inverse of the number of samples. For example, if a certain class of defect samples only accounts for 5% of the training data, the weight of this class can be set to 20 times that of other classes, thereby improving the model's learning effect on a small number of sample classes.
[0232] During the training process, the weighted combination of these loss terms constitutes the final defect sensitivity loss function, which guides the optimization of model parameters. The optimization of the loss function is achieved through the back propagation algorithm, which calculates the gradient and updates the model parameters in each iteration so that the loss function value gradually converges. In actual operation, common optimization algorithms such as Adam or SGD can be used, and the learning rate and regularization parameters can be adjusted according to the characteristics of the training data to ensure the stability and efficiency of the training.
[0233] Through the above optimization steps, the model can accurately identify and locate the appearance defects of microchips, including small-size defects, complex boundary defects, and rare defect categories in the case of sample imbalance, thereby improving the overall detection performance and robustness, and providing a reliable model foundation for subsequent defect detection tasks.
[0234] Furthermore, in the training process of the dual-branch lightweight convolutional neural network model, a defect sensitivity loss function is applied to optimize the training of the convolutional neural network model, including:
[0235] The defect sensitivity loss function provided by the following formula 2 is used to optimize the convolutional neural network model:
[0236]
[0237] in, Represents the overall loss value of the defect sensitivity loss function, which is used to guide the optimization of the convolutional neural network model;
[0238] , and is the weight coefficient, which is used to balance the contribution of the three sub-loss items to the total loss. The recommended value is [0.1, 1.0]. For example, it can be set according to the data characteristics. , , .
[0239] is the loss term based on the defect area attention mechanism; is the loss term based on edge feature preservation; is the loss term based on category balance; among them, the loss term based on the defect area attention mechanism This is achieved using the following formula 3:
[0240]
[0241] in, is the set of pixels in the defect area, determined by the real annotation data;
[0242] Pixel The predicted probability value of belonging to the specified defect category is generated by the output layer of the model and ranges from [0,1].
[0243] Pixel The true label value of is in binary form ( Indicates that the pixel belongs to the defective category, the opposite is true).
[0244] is the weight coefficient, calculated according to Formula 4;
[0245] For the adjustment factor, a value of 1.5 to 2.0 is recommended.
[0246] It is an adjustment coefficient used to control the influence weight of the second-order gradient term, and the recommended value is 0.1 to 1.0.
[0247] Pixel The second-order gradient of the predicted probability value is used to capture the changing trend of the predicted value. The calculation formula is as follows:
[0248]
[0249] Indicates that the pixel is The square of the first-order gradient of the direction; Indicates that the pixel is The square of the first-order gradient in the direction; the calculation formula is as follows:
[0250]
[0251] Pixel The distance from the defect center point; the defect center point is a reference point for defining the defect area, which is used to calculate the distance from each pixel to the center point. The method for determining it can be selected based on the annotation information of the defect area: If the defect area is clearly annotated, the horizontal and vertical coordinates of all pixels in the defect area can be averaged, and the position of this average value is the defect center point. If the shape of the defect area is relatively regular, such as close to a rectangle, the boundary of the defect area (the outermost pixels on the top, bottom, left and right) can be found, and then the center point of the boundary can be taken as the defect center point. If each pixel in the defect area has different weights, such as brightness value or detection confidence, the part with a larger weight can be used as a reference for the center point, and the position closer to these parts can be used as the defect center point.
[0252] is the distance smoothing factor, with a recommended value of 5 to 10.
[0253] Weight coefficient The following formula 4 is used for calculation:
[0254]
[0255] in, For tuning parameters, the recommended value is 0.5 to 2.0.
[0256] is the maximum diameter of the defect area;
[0257] Pixel Distance from the defect center;
[0258] Loss term based on edge feature preservation The calculation is performed using the following formula 5:
[0259]
[0260] in, It is a set of pixels at the edge of the defect, representing the pixel points related to the defect boundary in the microscopic image. It is obtained through real annotated data, usually extracted by segmentation or edge detection algorithms (such as Canny edge detection). For example, after the real defect area is annotated, the edge pixels of the annotated area are extracted by the gradient operator.
[0261] Pixel The true annotated edge gradient is extracted from the annotated data using classic methods such as the Sobel operator.
[0262] represents the L2 norm;
[0263] Loss term based on class balance , calculated using the following formula 6:
[0264]
[0265] in, is the set of all defect categories; The pixel belongs to the defective category The predicted probability value of For pixels with respect to defect class The true label value of
[0266] Is the defect category The weight coefficient is calculated using the following formula 7:
[0267]
[0268] in, Is the defect category The frequency of samples in the training data; is the smoothing factor, the recommended value is arrive .
[0269] Step S105: input the pre-processed microscopic image to be inspected into the trained dual-branch convolutional neural network model to generate a defect detection result for the microchip to be inspected, wherein the defect detection result includes a defect type, a defect location coordinate, and a confidence score.
[0270] In step S105, the pre-processed microscopic image to be inspected is input into the trained dual-branch convolutional neural network model to generate defect detection results for the microchip to be inspected. This process involves the input of microscopic images, feature extraction, fusion, and final defect classification and location, all of which are based on the model structure and parameters optimized in the previous steps.
[0271] First, the pre-processed microscopic image to be inspected enters the dual-branch convolutional neural network model through the input layer. The image has been processed by brightness equalization in step S101 and has good contrast and detail information, which can ensure that the model can efficiently extract defect features. The resolution of the input image should be consistent with that in the training process to ensure that the network's convolution kernel can correctly perceive the image features.
[0272] After entering the network, the image data is passed to the first branch and the second branch at the same time. The first branch uses depthwise separable convolution to extract global semantic features in the image. These features usually describe large-scale defect characteristics in the image, such as material anomalies in large areas or obvious surface deformations. The second branch extracts local detail features of the image through dilated convolution. These features can capture more subtle defect information, such as small-sized flaws or edge discontinuities. The two branches generate corresponding feature maps, reflecting the feature performance of the input image at different scales and levels.
[0273] After feature extraction, the two sets of feature maps are dynamically fused through the adaptive fusion gating unit. The working mechanism of the adaptive fusion gating unit is to automatically adjust the fusion weights of the two branch features according to the scale information and defect type in the feature map. For example, when the network detects a small-sized defect in the input image, it will give priority to increasing the weight of the detail features of the second branch; when facing a large-scale surface defect, it will rely more on the global features of the first branch. This dynamic adjustment can ensure that the network maintains a high accuracy in different detection scenarios.
[0274] The fused feature map is passed through the network's classification and positioning modules to generate the final defect detection results. The classification module uses the fully connected layer and the Softmax activation function to classify the defect type and output the confidence score for each category. The positioning module predicts the specific location coordinates of the defect through regression analysis, including the center point coordinates of the defect area, the size of the bounding box and other geometric information. For small-sized defects, the module can also output a more fine-grained boundary description to accurately calibrate the defect range.
[0275] The final inspection results include three main parts, namely defect type, defect location coordinates, and confidence score. These results are saved in a standard format (such as JSON or CSV) to facilitate subsequent defect analysis and quality control. In addition, in order to improve the interpretability of the results, the inspection results can also be visualized and superimposed on the original image, such as marking the defect area and its category label by box selection, so that the operator can check it intuitively.
[0276] Through the above process, the model can quickly and accurately output the defect detection results of microchips, effectively meeting the needs of high-efficiency and high-precision defect detection in actual production environments.
[0277] A second embodiment of the present application provides an electronic device, the electronic device comprising:
[0278] processor;
[0279] The memory is used to store a program, and when the program is read and executed by the processor, it executes a microchip appearance defect detection method based on a convolutional neural network provided in the first embodiment of the present application.
[0280] The third embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, a microchip appearance defect detection method based on a convolutional neural network provided in the first embodiment of the present application is executed.
[0281] Although the present application is disclosed as above in the form of a preferred embodiment, it is not intended to limit the present application. Any technical personnel in this field may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.
Claims
1. A method for detecting appearance defects of microchips based on convolutional neural networks, characterized in that: include: Acquire a high-resolution microscopic image of the microchip to be detected, and perform adaptive brightness equalization processing on the microscopic image to generate a pre-processed microscopic image, wherein the adaptive brightness equalization processing dynamically adjusts the brightness value of each pixel by calculating the brightness mean and variance of a local area of the image; The pre-processed microscopic image is subjected to a layered annotation strategy to perform defect feature annotation of the training sample, wherein the annotation includes coarse-grained classification annotation based on the pre-processed microscopic image to classify the defects into surface deformation class and material abnormality class, and fine-grained positioning annotation combined with image features to generate annotated sample data containing specific defect types and their position coordinates; Based on the labeled sample data, a dual-branch lightweight convolutional neural network model is constructed and trained; During the training process of the dual-branch lightweight convolutional neural network model, a defect sensitivity loss function is applied to optimize the training of the convolutional neural network model, wherein the defect sensitivity loss function includes: a loss term based on a defect area attention mechanism to improve the model's detection capability for small-size defect areas, a loss term based on edge feature retention to enhance the defect boundary positioning accuracy, and a loss term based on category balance to alleviate the impact of an imbalance in the number of samples of different defect types on the detection performance; Inputting the preprocessed microscopic image to be inspected into the trained dual-branch convolutional neural network model to generate defect detection results for the microchip to be inspected, wherein the defect detection results include defect type, defect location coordinates, and confidence score; The dual-branch lightweight convolutional neural network model includes an input layer, a first branch, a second branch, an adaptive fusion gating unit, and a classification and regression unit; The input layer is used to receive the preprocessed microscopic image and standardize the received image data into a standardized microscopic image tensor suitable for model processing; The first branch is used to receive the standardized microscopic image tensor provided by the input layer, and adopts a depth-separable convolution module to extract global semantic features in the microscopic image by separating the calculation of spatial convolution and channel convolution to obtain a global semantic feature map; The second branch is used to receive the standardized microscopic image tensor provided by the input layer, and adopts a dilated convolution module to extract local detail features of the microscopic image by expanding the receptive field without increasing the number of parameters, thereby obtaining a local detail feature map; The adaptive fusion gating unit is used to receive the global semantic feature map and the local detail feature map, and dynamically adjust the feature fusion weights of the first branch and the second branch according to the defect type and scale characteristics through the gating unit to generate a comprehensive feature map; The classification and regression unit is used to receive the comprehensive feature map from the feature fusion module, and classify the defect type through the fully connected layer and the Softmax activation function; predict the specific location coordinates and confidence scores of the defects through the regression module; and generate defect detection results, which include defect types, defect location coordinates and confidence scores.
2. The microchip appearance defect detection method based on convolutional neural network according to claim 1 is characterized in that: The depthwise separable convolution module in the first branch includes an alternating stack of multiple depthwise convolutional layers and pointwise convolutional layers, wherein the depthwise convolutional layers are used to extract spatial features for each channel respectively, and the pointwise convolutional layers are used to perform feature fusion between all channels; by adjusting the number and parameter size of the depthwise convolutional layers and the pointwise convolutional layers, the extraction of global semantic features is ensured while reducing the amount of computation and retaining significant semantic information.
3. The microchip appearance defect detection method based on convolutional neural network according to claim 1 is characterized in that: The atrous convolution module in the second branch includes multiple convolution layers with different atrous rates, and multi-scale feature extraction capability is formed by setting convolution layers with atrous rates of 1, 2 and 4 in parallel, wherein the convolution results with different atrous rates are fused by pixel-by-pixel summation to further enhance the integrity of local detail features.
4. The microchip appearance defect detection method based on convolutional neural network according to claim 1 is characterized in that: The adaptive fusion gating unit jointly calculates the fusion weight based on the spatial information and channel information of the feature map, specifically including: calculating the weight of each channel through the channel attention mechanism, calculating the weights of different spatial positions through the spatial attention mechanism, and dynamically adjusting the fusion ratio of the first branch and the second branch features according to the channel weight and the spatial weight, so as to adapt to the scale characteristics of different types of defects.
5. The microchip appearance defect detection method based on convolutional neural network according to claim 1 is characterized in that: The classification and regression unit jointly optimizes the classification and regression tasks by introducing a multi-task learning strategy, wherein the loss function of the classification task is a cross entropy loss based on category balance, and the loss function of the regression task is a weighted square error loss. The classification loss and the regression loss are combined into a total loss according to preset weights, ensuring that the model is simultaneously optimized in terms of the accuracy of defect type classification and location prediction.
6. The microchip appearance defect detection method based on convolutional neural network according to claim 2 is characterized in that: The deep convolution layer in the first branch introduces a mechanism of dynamic convolution kernel size to adaptively adjust the receptive field size of the convolution kernel according to the resolution of the input feature map, wherein the dynamic adjustment is determined in real time based on the resolution and complexity of the feature map. At low resolution, a smaller convolution kernel is used to reduce the amount of calculation, and at high resolution, a larger convolution kernel is used to enhance the semantic feature extraction capability, thereby further optimizing the computational efficiency and the accuracy of global feature extraction.
7. The microchip appearance defect detection method based on convolutional neural network according to claim 3 is characterized in that: The void convolution module in the second branch introduces a dynamic adjustment mechanism of the void ratio and adaptively selects a suitable void ratio combination according to the size distribution of the defect area in the microscopic image. The dynamic adjustment of the void ratio is completed by the image pre-analysis module. The void ratio parameter configuration adapted to the current image characteristics is generated through the boundary information and morphological characteristics of the defects, thereby achieving optimized extraction capabilities for multi-scale defect areas and reducing redundant features.
8. The microchip appearance defect detection method based on convolutional neural network according to claim 1 is characterized in that: The step of performing adaptive brightness equalization processing on the microscopic image to generate a pre-processed microscopic image includes: The following formula 1 is used to dynamically adjust the brightness value of each pixel of the microscopic image to generate a preprocessed microscopic image: in, To adjust the pixels in the preprocessed microscopic image The brightness value of Represents the horizontal coordinate of the pixel point, The vertical coordinate of the pixel is shown; is the pixel point in the original microscopic image The brightness value of Represents the original microscopic image in pixels The average brightness in the local area centered at ; Represents the original microscopic image in pixels The standard deviation of brightness in the local area centered on A positive smoothing factor to prevent the denominator from being zero; It is the global brightness mean of the entire image in the original microscopic image, which is used to adjust the balance of the overall brightness of the image; is the global brightness control factor, which is used to control the amplitude of brightness adjustment; is the global balance coefficient, which is used to adjust the impact of global brightness contrast; It is the brightness smoothing factor, which controls the dynamic range of global brightness adjustment.
9. The microchip appearance defect detection method based on convolutional neural network according to claim 1, characterized in that: In the training process of the dual-branch lightweight convolutional neural network model, applying a defect sensitivity loss function to optimize the convolutional neural network model includes: The defect sensitivity loss function provided by the following formula 2 is used to optimize the convolutional neural network model: in, Represents the overall loss value of the defect sensitivity loss function, which is used to guide the optimization of the convolutional neural network model; and is the weight coefficient; is the loss term based on the defect area attention mechanism; is the loss term based on edge feature preservation; is the loss term based on category balance; among them, the loss term based on the defect area attention mechanism This is achieved using the following formula 3: in, is the set of pixels in the defect area, determined by the real annotation data; Pixel The predicted probability value of belonging to the specified defect category; Pixel The true label value of is the weight coefficient; is the adjustment factor; is the adjustment factor; Pixel The second-order gradient of the predicted probability value is used to capture the changing trend of the predicted value; Indicates that the pixel is The square of the first-order gradient of the direction; Indicates that the pixel is The square of the first-order gradient of the direction; Pixel Distance from the defect center; is the distance smoothing factor; weight coefficient The following formula 4 is used for calculation: in, To adjust the parameters; is the maximum diameter of the defect area; Pixel Distance from the defect center; Loss term based on edge feature preservation The calculation is performed using the following formula 5: in, is the pixel set of the defect edge; Pixel The edge gradient of the true annotation; represents the L2 norm; Loss term based on class balance , calculated using the following formula 6: in, is the set of all defect categories; The pixel belongs to the defective category The predicted probability value of For pixels with respect to defect class The true label value of Is the defect category The weight coefficient is calculated using the following formula 7: in, Is the defect category The frequency of samples in the training data; is the smoothing factor.
Citation Information
Patent Citations
Two-stage mainboard image defect detecting and positioning method based on machine vision
CN114972213A
Defect detection method based on joint optimization and mixed attention feature fusion
CN115294038A