Defect detection method, device, equipment and computer program product
By adding upsampling, splicing and C3K2 modules to the Yolov11 network, combined with an improved model of shallow feature detection heads, the problem of small target defect detection on small screens is solved, and the accuracy and stability of detection is improved.
Patent Information
- Application Number
- CN202510448150.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-11
AI Technical Summary
Existing defect detection technologies are difficult to effectively detect small target defects on small screens, especially point defects and line defects, and deep learning methods are not effective in small target detection.
Using an improved defect detection model, the detection capability of small target defects is enhanced by adding upsampling, splicing and C3K2 modules to the feature fusion layer of the Yolov11 network, combined with a shallow feature detection head.
It improves the accuracy and stability of detection of small target defects, and can effectively identify common screen defects such as points, lines, and bubbles.
Smart Images

Figure CN120298385A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision technology, and particularly to a defect detection method, device, equipment, and computer program product. Background Art
[0002] Defect detection of small screens is a crucial link in the manufacturing industry, especially in the fields of semiconductor, display panel, and electronic product manufacturing. However, this task faces many challenges. First, there are various types of defects on small screens, including point defects, line defects, surface defects, etc., and the sizes of these defects are often very small. For example, the size of a point defect may be only about 4-5 pixels, and the lateral resolution of a line defect may be as low as 2-3 pixels. These tiny defects make detection extremely difficult.
[0003] Existing detection technologies have unsatisfactory detection effects for small targets. Traditional image processing methods, such as threshold segmentation, edge detection, etc., although can identify defects to a certain extent, have limited ability to distinguish different types of defects, and are easily affected by noise and interference, resulting in unstable detection. In addition, these methods often rely on manually designed features and rules, and have poor adaptability to complex and variable defect morphologies and background environments.
[0004] In recent years, deep learning methods have made remarkable progress in the field of computer vision and have performed well in object detection tasks. However, existing deep learning object detection networks are generally designed for large targets. As the number of network layers deepens and convolutional operations are carried out, the feature information of small targets gradually gets lost during the process of layer-by-layer transmission, resulting in poor detection effects for small targets. Therefore, existing deep learning methods also have limitations in the field of small target defect detection.
[0005] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention
[0006] The main purpose of this application is to provide a defect detection method, device, equipment, and computer program product, aiming to solve the technical problem of difficult detection of small target defects during defect detection.
[0007] To achieve the above objective, this application proposes a defect detection method, and the method includes:
[0008] Obtain an image to be detected, and perform preprocessing of image adjustment and normalization on the image to be detected;
[0009] Input the preprocessed image to be detected into a pre-constructed defect detection model, and perform defect detection on the image to be detected through the defect detection model to obtain the defect detection result of the image to be detected. The defect detection model includes a feature fusion layer, and the feature fusion layer includes a shallow feature detection head.
[0010] In one embodiment, before the step of inputting the preprocessed image to be detected into a pre-constructed defect detection model, and performing defect detection on the image to be detected through the defect detection model to obtain the defect detection result of the image to be detected, the following steps are further included:
[0011] Obtain several groups of initial training images;
[0012] Use an annotation tool to annotate the target defects in each of the initial training images to generate annotated training images;
[0013] Perform preprocessing of adjustment and normalization on each of the annotated training images, and divide the preprocessed annotated training images into a training image set and a validation image set;
[0014] Input the training image set into the initial defect detection model for iterative training, and input the validation image set into the initial defect detection model after iterative training for performance evaluation to obtain the defect detection model.
[0015] In one embodiment, the defect detection model further includes a feature extraction layer and a localization prediction layer. The step of inputting the preprocessed image to be detected into a pre-constructed defect detection model, and performing defect detection on the image to be detected through the defect detection model to obtain the defect detection result of the image to be detected includes:
[0016] Extract features from the preprocessed image to be detected through the feature extraction layer to obtain a first data feature map;
[0017] Input the first data feature map into the feature fusion layer for feature fusion and enhancement to obtain a second data feature map after the first data feature map is fused and enhanced;
[0018] Input the second data feature map into the localization prediction layer for target localization and classification to obtain the defect detection result of the image to be detected.
[0019] In one embodiment, the step of extracting features from the preprocessed image to be detected through the feature extraction layer to obtain a first data feature map includes:
[0020] Input the preprocessed image to be detected into the first extraction module of the feature extraction layer for convolution operation to obtain a first image feature map;
[0021] Input the first image feature map into the second extraction module of the feature extraction layer for segmentation and splicing operations to obtain a second image feature map;
[0022] Input the second image feature map into the third extraction module of the feature extraction layer for feature enhancement operations to obtain a third image feature map;
[0023] Input the third image feature map into the fourth extraction module of the feature extraction layer for feature extraction and fusion operations to obtain the first data feature of the image to be detected.
[0024] In one embodiment, the step of inputting the first data feature map into the feature fusion layer for feature fusion and enhancement to obtain a second data feature map after fusion and enhancement of the first data feature map includes:
[0025] Perform upsampling on the first data feature map through the feature fusion layer;
[0026] Perform splicing operations on the upsampled first data feature map to obtain a first spliced feature map of the first data feature map;
[0027] Perform feature extraction and feature fusion operations on the first spliced feature map to obtain a second data feature map after fusion and enhancement of the first data feature map.
[0028] In one embodiment, the step of inputting the first data feature map into the feature fusion layer for feature fusion and enhancement to obtain a second data feature map after fusion and enhancement of the first data feature map further includes:
[0029] Perform regression prediction of small target defects on the first data feature map through the shallow feature detection head to obtain defect location information and class prediction information of the small target defects.
[0030] In one embodiment, the step of inputting the second data feature map into the location prediction layer for target location classification to obtain a defect detection result of the image to be detected includes:
[0031] Fuse the second data feature map through the location prediction layer to obtain a second fused feature map of the second data feature;
[0032] Perform boundary box prediction of target defects on the second fused feature map to obtain a boundary box prediction map of target defects on the image to be detected;
[0033] Perform class prediction and coordinate location on the boundary boxes on the boundary box prediction map to obtain the class and coordinate positions corresponding to the boundary boxes;
[0034] Based on the category and coordinate position corresponding to the bounding box, obtain the defect detection result of the target defect on the image to be detected.
[0035] In addition, to achieve the above object, the present application also proposes a defect detection device, which includes:
[0036] An image acquisition module, configured to acquire an image to be detected and perform preprocessing of image adjustment and normalization on the image to be detected;
[0037] A defect detection module, configured to input the preprocessed image to be detected into a pre-constructed defect detection model, and perform defect detection on the image to be detected through the defect detection model to obtain the defect detection result of the image to be detected. The defect detection model includes a feature fusion layer, and the feature fusion layer includes a shallow feature detection head.
[0038] In addition, to achieve the above object, the present application also proposes a defect detection device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. The computer program is configured to implement the steps of the defect detection method as described above.
[0039] In addition, to achieve the above object, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the defect detection method as described above.
[0040] In addition, to achieve the above object, the present application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the defect detection method as described above.
[0041] One or more technical solutions proposed by the present application have at least the following technical effects:
[0042] The present application discloses a defect detection method, device, device, and computer program product, including: acquiring an image to be detected, and performing preprocessing of image adjustment and normalization on the image to be detected; inputting the preprocessed image to be detected into a pre-constructed defect detection model, and performing defect detection on the image to be detected through the defect detection model to obtain the defect detection result of the image to be detected. The defect detection model includes a feature fusion layer, and the feature fusion layer includes a shallow feature detection head. The present application performs defect detection on the image to be detected through a pre-constructed defect detection model, and can detect small target defects on the image through a shallow feature detection head, improving the accuracy of detecting small target defects. Description of the Drawings
[0043] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application and, together with the specification, are used to explain the principles of this application.
[0044] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0045] Figure 1 It is a schematic flowchart provided for the first embodiment of the defect detection method of this application;
[0046] Figure 2 It is a schematic structural diagram of the Yolov11 network provided for the embodiments of the defect detection method of this application;
[0047] Figure 3 It is a schematic structural diagram of the defect detection model provided for the embodiments of the defect detection method of this application;
[0048] Figure 4 It is a schematic module structure diagram of the defect detection device for the embodiments of this application;
[0049] Figure 5 It is a schematic device structure diagram of the hardware operating environment involved in the defect detection method for the embodiments of this application.
[0050] The implementation, functional features, and advantages of the objectives of this application will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments
[0051] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of this application and are not used to limit this application.
[0052] To better understand the technical solutions of this application, the following will be described in detail in combination with the specification drawings and specific implementation manners.
[0053] The main solution of the embodiments of this application is: obtaining an image to be detected, and performing preprocessing of image adjustment and normalization on the image to be detected; inputting the preprocessed image to be detected into a pre-constructed defect detection model, and performing defect detection on the image to be detected through the defect detection model to obtain the defect detection result of the image to be detected, where the defect detection model includes a feature fusion layer, and the feature fusion layer includes a shallow feature detection head.
[0054] In this embodiment, for the convenience of description, the following will be described with the recognition defect detection device as the execution subject.
[0055] Small screen defect detection is a crucial process in the manufacturing industry, especially in the fields of semiconductor, display panel, and electronic product manufacturing. However, this task faces numerous challenges. Firstly, there are various types of defects on small screens, including point defects, line defects, surface defects, etc., and the sizes of these defects are often extremely small. For example, the size of a point defect may be only about 4 - 5 pixels, and the lateral resolution of a line defect may be as low as 2 - 3 pixels. These tiny defects make detection extremely difficult.
[0056] Existing detection technologies do not perform well in detecting small targets. Traditional image processing methods, such as threshold segmentation, edge detection, etc., although can identify defects to a certain extent, have limited ability to distinguish different types of defects and are easily affected by noise and interference, resulting in unstable detection. In addition, these methods often rely on manually designed features and rules and have poor adaptability to complex and variable defect morphologies and background environments.
[0057] In recent years, deep learning methods have made remarkable progress in the field of computer vision and have performed well in object detection tasks. However, existing deep learning object detection networks are generally designed for large targets. As the number of network layers deepens and convolutional operations are carried out, the feature information of small targets gradually gets lost during the process of layer - by - layer transmission, resulting in poor detection effects for small targets. Therefore, existing deep learning methods also have limitations in the field of small - target defect detection.
[0058] This application provides a solution. The defect detection model constructed in advance is used to detect defects in the image to be detected. Through the shallow - feature detection head, small - target defects on the image can be detected, improving the accuracy of small - target defect detection.
[0059] It should be noted that the execution entity of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a defect detection device, etc. that can implement the above functions. Hereinafter, taking the defect detection device as an example, this embodiment and the following embodiments will be described.
[0060] Based on this, the embodiments of this application provide a defect detection method, referring to Figure 1 , Figure 1 which is the flowchart of the first embodiment of the defect detection method of this application.
[0061] In this embodiment, the defect detection method includes steps S10 - S40:
[0062] Step S10, obtain the image to be detected and perform pre - processing of image adjustment and normalization on the image to be detected.
[0063] It should be noted that the image to be detected refers to the image for which defect detection is required.
[0064] Specifically, obtain the image to be detected that needs to be defect-detected, and adjust the size of the image to be detected to a size suitable for detection by the defect detection model. In an embodiment of the present application, the image to be detected is adjusted to 640x640 pixels, and 640x640 pixels is a commonly used image detection size. Normalize the image to be detected so that its pixel values are distributed within a specific range, usually by scaling the pixel values to between 0 and 1.
[0065] Step S20: Input the preprocessed image to be detected into a pre-constructed defect detection model, and perform defect detection on the image to be detected through the defect detection model to obtain the defect detection result of the image to be detected. The defect detection model includes a feature fusion layer, and the feature fusion layer includes a shallow feature detection head.
[0066] It should be noted that the defect detection model is an object detection model improved based on the Yolov11 network. By adding an upsampling, splicing, and C3K2 module to the feature fusion layer (Neck part) of the Yolov11 network, the feature fusion of the shallow network is completed, so as to achieve the improvement of the detection of small target defects through a lightweight detection head. The original Yolov11 network adds a shallow feature fusion and a small target detection head to increase the feature information of small targets during shallow feature extraction. The added small target detection layer mainly includes the upsampling module, Concat splicing module, C3K2 module, and detection head module of the Yolov11 network.
[0067] More specifically, the upsampling module is an interpolation method based on pixel positions, used to estimate the values of missing pixels. In the upsampling operation, first expand the image size according to a specified ratio, and insert virtual pixel points at blank positions. The gray value of this pixel point is obtained by weighting the pixel values of 16 surrounding points, thus completing the upsampling operation of the feature layer.
[0068] The Concat splicing module performs the operation of splicing two or more tensors along a certain dimension, used to connect multiple feature maps or vectors, so as to facilitate subsequent calculations and let the network learn how to fuse features. No information is lost during this process.
[0069] The C3K2 module is an innovative module newly proposed in the deep learning object detection Yolov11 network, which is transformed from the C2F module of Yolov8. Through the concatenation operation of splitting and then Concat on multiple feature maps, the extraction of feature information is completed.
[0070] The detection head module uses depthwise separable convolutions to reduce the number of parameters. Two depthwise separable convolutions are added to the classification detection head in the decoupled head. By convolving each channel of the single-channel input, an output feature map with the same number of channels as the input feature map is obtained. While effectively extracting feature information, the number of parameters is greatly reduced.
[0071] After the feature map is upsampled by the upsampling module, its size doubles. After extracting features through the C3K2 module, the number of channels becomes one-third of the original. Therefore, the C3K2 module can effectively extract features while greatly reducing the number of parameters of the model, effectively improving the speed.
[0072] In addition, it should be noted that the feature fusion layer refers to the Neck part of the feature enhancement network of the Yolov11 network. The feature enhancement network further processes the feature map output by the backbone network to enhance the feature expression ability. Its main function is to enhance and fuse the features extracted by the Backbone to improve the model's detection ability for targets of different scales. Through feature fusion, Neck can generate richer and more robust feature representations, which helps to improve the accuracy and robustness of object detection. The feature fusion layer includes an upsampling module, a splicing module, a segmentation and splicing module, and a convolution module. The feature fusion layer of this application also includes a shallow feature detection head.
[0073] In addition, it should be noted that the shallow feature detection head is a new detection module added to the feature fusion layer after improving the Yolov11 network, which can effectively improve the detection performance of small targets. By introducing a shallow feature fusion mechanism, during the low-level (shallow) feature extraction stage of the backbone network Backbone, a decoupled detection head composed of an upsampling module, a Concat splicing module, a C3K2 module, and depthwise separable convolutions is integrated. The shallow feature detection head utilizes the high-resolution characteristics of the shallow feature map (retaining more spatial details), combines upsampling to enhance feature expression, cross-layer feature splicing to fuse information, and the efficient feature compression ability of the C3K2 module. Finally, a lightweight detection head is used to locate and classify small target defects (such as 4-5 pixel screen defects). Its role is to alleviate the problem of small target information loss in deep convolutions and significantly improve the sensitivity and detection accuracy of the model for tiny targets.
[0074] For a better understanding of the defect detection model of this application, please refer to Figure 2 and Figure 3 where Figure 2 is a schematic diagram of the structure of the Yolov11 network, Figure 3This is a schematic diagram of the structure of the defect detection model improved based on the Yolov11 network. The digital labels 0, 1, 2... in the figure represent the layers where the modules are located in the Yolov11 network. Based on the Yolov11 network, the defect detection model of the present application adds a shallow feature detection head to the feature fusion layer, integrating an upsampling module, a Concat splicing module, a C3K2 module, and a detection head composed of depthwise separable convolutions. First, an upsampling module is added to the feature fusion layer to increase the size of the feature map, that is, the upsampling module of the 17th layer is added. Then, the segmentation splicing module of the second layer and the upsampling module of the 17th layer are spliced by the splicing module of the 18th layer to increase the number of channels of the feature map. Finally, through the feature extraction of the segmentation splicing module, the number of channels is reduced to one-third, and a detection head for small targets is added from the segmentation splicing module of the 19th layer to perform the localization and class prediction of small target defects.
[0075] More specifically, in Figure 2 In the schematic diagram of the structure of the Yolov11 network shown, the original Yolov11 network only has three detection heads, and these three detection heads are all performing class prediction and localization prediction on the deep network, which will cause serious loss of small target information during the convolution process of the backbone feature extraction network. Therefore, in the detection of small screen defects, point defects and line defects with only 4-5 pixels are very difficult to be extracted and detected by the network. In the defect detection model of the present application, a feature extraction module is introduced into the shallow backbone feature extraction network, and a detection head for small target defects is added, which can effectively detect small target defects.
[0076] In an example of the present application, assuming an input image to be detected of 640*640, four detection heads with shapes of 128*160*160, 64*80*80, 128*40*40, and 256*20*20 are set to perform predictions on the image to be detected respectively, realizing multi-scale target prediction. Without significantly changing the robustness of the detection accuracy of large targets, the stability and accuracy of small target detection are effectively improved.
[0077] Specifically, the feature extraction layer extracts features from the preprocessed image to be detected to obtain a first data feature map; the first data feature map is input into the feature fusion layer for feature fusion and enhancement to obtain a second data feature map after the first data feature map is fused and enhanced; the second data feature map is input into the localization prediction layer for target localization and classification to obtain the defect detection result of the image to be detected.
[0078] In this embodiment, through the above solution, the image to be detected is obtained, and preprocessing of image adjustment and normalization is performed on the image to be detected; the preprocessed image to be detected is input into a pre-constructed defect detection model, and the defect detection model is used to perform defect detection on the image to be detected to obtain the defect detection result of the image to be detected. The defect detection model includes a feature fusion layer, and the feature fusion layer includes a shallow feature detection head. This application performs defect detection on the image to be detected through the defect detection model, enhances the recognition ability of small target defects, can detect small target defects on the image, effectively solves the problem of unstable defect detection commonly found in screen detection such as points, lines, and bubbles, and improves the stability and accuracy of small target defect detection.
[0079] Based on the above implementation scheme, in a feasible implementation manner, before the step of inputting the preprocessed image to be detected into a pre-constructed defect detection model, and using the defect detection model to perform defect detection on the image to be detected to obtain the defect detection result of the image to be detected, steps S31 to S34 are further included:
[0080] Step S31, obtain several groups of initial training images.
[0081] It should be noted that the initial training images refer to images with small target defects, which are used for iterative training of the initial defect detection model.
[0082] Specifically, pictures containing small target defects are collected. Public datasets can be used, such as datasets provided by platforms like Kaggle or Roboflow, or self-made datasets can be created by taking or collecting relevant pictures.
[0083] Step S32, use a labeling tool to label the target defects in each of the initial training images to generate labeled training images.
[0084] It should be noted that the labeled training images refer to the images after labeling the target defects on the initial training images.
[0085] Specifically, use a labeling tool to label the initial training images, label the target defects in the images, and labeling tools such as Labelimg, etc.
[0086] Step S33, perform preprocessing of adjustment and normalization on each of the labeled training images, and divide the preprocessed labeled training images into a training image set and a validation image set.
[0087] Specifically, adjust the size of the labeled training images to a size suitable for detection by the defect detection model. In an embodiment of the present application, the labeled training images are adjusted to 640x640 pixels, which is a commonly used image detection size. Normalize the labeled training images so that their pixel values are distributed within a specific range, usually by scaling the pixel values to between 0 and 1. Divide the preprocessed labeled training images into a training image set and a validation image set. The training image set is used to input into the initial defect detection model for iterative training, and the validation image set is used to input into the trained initial defect detection model for performance evaluation, calculating and comparing the performance of different metrics, such as precision, recall, mAP, etc.
[0088] Step S34, input the training image set into the initial defect detection model for iterative training, and input the validation image set into the initial defect detection model after iterative training for performance evaluation to obtain the defect detection model.
[0089] Specifically, input the training image set into the initial defect detection model for iterative training. During the training process, the model will learn how to identify defects in the images; input the validation image set into the initial defect detection model after iterative training for performance evaluation. Calculate metrics such as the accuracy, recall, and F1 score of the model on the validation set to evaluate the performance of the model; optimize the model according to the evaluation results. Different network structures, hyperparameter settings, and training strategies, etc. can be tried to improve the performance of the model; after the model training is completed and passes the performance evaluation, the defect detection model is obtained.
[0090] Based on the above implementation solutions, in a feasible implementation manner, the defect detection model further includes a feature extraction layer and a localization prediction layer. The step of inputting the preprocessed image to be detected into the pre-constructed defect detection model and performing defect detection on the image to be detected by the defect detection model to obtain the defect detection result of the image to be detected includes steps S21 - S23:
[0091] Step S21, perform feature extraction on the preprocessed image to be detected through the feature extraction layer to obtain a first data feature map.
[0092] It should be noted that the feature extraction layer refers to the Backbone part of the Yolov11 network, and the Backbone is responsible for extracting image features. The feature extraction layer includes a convolution module, a segmentation and splicing module, a pooling and fusion module, and a channel and spatial attention module. The Backbone of Yolov11 extracts image features through multiple convolutional layers and residual connections. The feature extraction process includes convolutional operations, pooling operations (or downsampling), and possible residual connections. The main function of the Backbone is to extract deep features in the image, which are crucial for subsequent object detection. It extracts image information layer by layer through convolutional layers, and gradually reduces the resolution of the feature map through downsampling while increasing the degree of abstraction of the features.
[0093] In addition, it should be noted that the first data feature map refers to the feature map obtained by processing the image to be detected through the Backbone part of the defect detection model.
[0094] Specifically, the preprocessed image to be detected is input into the first extraction module of the feature extraction layer for convolutional operation to obtain a first image feature map; the first image feature map is input into the second extraction module of the feature extraction layer for segmentation and splicing operation to obtain a second image feature map; the second image feature map is input into the third extraction module of the feature extraction layer for feature enhancement operation to obtain a third image feature map; the third image feature map is input into the fourth extraction module of the feature extraction layer for feature extraction and fusion operation to obtain the first data feature of the image to be detected.
[0095] Step S22: Input the first data feature map into the feature fusion layer for feature fusion and enhancement to obtain a second data feature map after the first data feature map is fused and enhanced.
[0096] It should be noted that the second data feature refers to the feature map obtained by processing the first data feature map through the Neck part of the defect detection model.
[0097] Specifically, the first data feature map is upsampled through the feature fusion layer; a splicing operation is performed on the upsampled first data feature map to obtain a first spliced feature map of the first data feature map; feature extraction and feature fusion operations are performed on the first spliced feature map to obtain a second data feature map after the first data feature map is fused and enhanced.
[0098] Step S23: Input the second data feature map into the localization and prediction layer for object localization and classification to obtain the defect detection result of the image to be detected.
[0099] It should be noted that the positioning prediction layer refers to the detection head (Head) part of the Yolov11 network. The main function of the Head is to predict the location, category, and confidence of target defects. It uses the fused feature map output by the Neck part for prediction and outputs the final target detection result. The positioning prediction layer mainly includes the detection head.
[0100] In addition, it should be noted that the defect detection result of the image to be detected refers to the detection result obtained after positioning and prediction by the Head part of the defect detection model, including the location, category, and confidence of the predicted target defect.
[0101] Specifically, the second data feature map is fused through the positioning prediction layer to obtain the second fused feature map of the second data feature; the bounding box of the target defect is predicted for the second fused feature map to obtain the bounding box prediction map of the target defect on the image to be detected; the category prediction and coordinate positioning are performed on the bounding boxes on the bounding box prediction map to obtain the category and coordinate positions corresponding to the bounding boxes; based on the category and coordinate positions corresponding to the bounding boxes, the defect detection result of the target defect on the image to be detected is obtained.
[0102] Based on the above implementation scheme, in a feasible implementation manner, the step of extracting features from the preprocessed image to be detected through the feature extraction layer to obtain the first data feature map includes S211 to S214:
[0103] Step S211, input the preprocessed image to be detected into the first extraction module of the feature extraction layer for convolution operation to obtain the first image feature map.
[0104] It should be noted that the first extraction module of the feature extraction layer refers to the convolution module (Conv module), which is the basic component of the convolutional neural network (CNN). In Yolov11, the convolution module is responsible for performing convolution operations to extract the features of the input image. The convolution operation slides the convolution kernel (also called the filter) on the input image and calculates the dot product of the convolution kernel and the local area of the image to obtain the feature map.
[0105] Specifically, the preprocessed image to be detected is input into the first extraction module of the feature extraction layer for convolution operation. The image is convolved through a convolution kernel (such as a 3x3 convolution kernel), and then the activation function (such as ReLU) and batch normalization (BN) layer are applied to perform preliminary feature extraction on the image to be detected, extracting low-level features, usually with fewer channels and higher spatial resolution, to obtain the first image feature.
[0106] Step S212, input the first image feature map into the second extraction module of the feature extraction layer for segmentation and splicing operations to obtain a second image feature map.
[0107] It should be noted that the second extraction module of the feature extraction layer refers to the segmentation and splicing module (C3k2 module), which is an important feature extraction component in Yolov11. It is an improvement based on the traditional C3 module. The C3k2 module provides more powerful feature extraction capabilities by combining variable convolution kernels (such as 3x3, 5x5, etc.) and channel separation strategies. Specifically, the C3k2 module usually divides the input features into two parts: one part is directly passed through ordinary convolution operations, and the other part undergoes deep feature extraction through multiple C3k (when the c3k parameter is set to True) or Bottleneck structures. Finally, the two parts of the features are spliced and fused through 1x1 convolution. This structure can maintain light weight while effectively extracting deep features.
[0108] Specifically, the first image feature map is bisected and divided into two parts for separate processing. Apply the C3K structure (i.e., a variant of CSPNet) to one part of the feature map for complex feature extraction. The C3K structure usually includes multiple convolutional layers and Bottleneck layers to extract richer features; splice the processed feature map with the other part of the feature map to obtain the feature map corresponding to the second image feature.
[0109] Step S213, input the second image feature map into the third extraction module of the feature extraction layer for feature enhancement operations to obtain a third image feature map.
[0110] It should be noted that the third extraction module of the feature extraction layer refers to the pooling and fusion module (SPPF module). The Spatial Pyramid Pooling Fast (SPPF) module is a spatial pyramid pooling module that is used in the Yolov series of networks to extract multi-scale spatial features. The SPPF module captures context information at different scales by performing pooling operations on the feature map at different spatial scales. This multi-scale feature extraction method helps to improve the robustness of the model to changes in the size and position of the target.
[0111] Specifically, use pooling kernels of different sizes (such as 5x5, 9x9, 13x13, etc.) to perform pooling operations on the second image feature map, and then splice the pooled feature maps. The SPPF module improves the robustness and expressive power of the features by fusing features at different scales.
[0112] Step S214, input the third image feature map into the fourth extraction module of the feature extraction layer for feature extraction and fusion operations to obtain the first data feature of the image to be detected.
[0113] It should be noted that the fourth extraction module of the feature extraction layer refers to the channel and spatial attention module (C2PSA module), which combines the PSA (Pointwise Spatial Attention) block. By introducing the PSA block into the standard C2F module, a more powerful attention mechanism is achieved. This attention mechanism can selectively focus on the important parts of the input features, suppress unimportant information, and thus improve the model's ability to capture important features.
[0114] Specifically, perform further convolutional operations on the third image feature map to extract features, apply attention mechanisms (such as channel attention, spatial attention, etc.), and perform weighted processing on the feature map. The attention mechanism can enhance the expression of key features and suppress irrelevant features; fuse the weighted feature map with the original feature map (or the feature map after certain processing) to obtain a new feature map, that is, the feature map corresponding to the first data feature of the image to be detected; the feature map processed by the C2PSA module has stronger feature expression ability and higher robustness.
[0115] Based on the above implementation, in a feasible implementation, the step of inputting the first data feature map into the feature fusion layer for feature fusion and enhancement to obtain the second data feature map after fusion and enhancement of the first data feature map includes S221 - S223:
[0116] Step S221, perform upsampling on the first data feature map through the feature fusion layer.
[0117] Specifically, perform upsampling on the first data feature map output by the feature extraction layer through the upsampling module (Upsample module) of the feature fusion layer, improve the resolution of the first data feature map, and transfer the upsampled feature map to the Concat module.
[0118] Step S222, perform a splicing operation on the first data feature map after upsampling to obtain the first spliced feature map of the first data feature map.
[0119] Specifically, input the upsampled feature map into the splicing module (Concat module) of the feature fusion layer, and perform a splicing operation on the upsampled feature map and the corresponding scale feature map from the feature extraction layer through the splicing module to obtain the first spliced feature map, and transfer the first spliced feature map to the C3k2 module.
[0120] Step S223: Perform feature extraction and feature fusion operations on the first spliced feature map to obtain a second data feature map with enhanced fusion of the first data feature map.
[0121] Specifically, transfer the feature map after the splicing operation to the segmentation and splicing module (C3k2 module). The input feature map is segmented and evenly divided by the segmentation and splicing module and then convolutional operations are performed, including residual connections to enhance features. Then, the feature maps after the convolutional operations are spliced to obtain the second data feature map.
[0122] Furthermore, input the feature map processed by the segmentation and splicing module into the convolutional module (Conv module). Convolutional operations are performed by the convolutional module for further feature extraction. Then, transfer the feature map after feature extraction to the next segmentation and splicing module for feature extraction and feature fusion, and transfer the feature map after the operation to the localization prediction layer.
[0123] Based on the above implementation, in a feasible implementation, the step of inputting the first data feature map into the feature fusion layer for feature fusion and enhancement to obtain a second data feature map with enhanced fusion of the first data feature map further includes S224:
[0124] Step S224: Use the shallow feature detection head to perform regression prediction on small target defects in the first data feature map to obtain defect localization information and class prediction information of the small target defects.
[0125] It should be noted that small target defects refer to defects with a small area, such as point defects with a size of about 4-5 pixels and line defects with a horizontal resolution as low as 2-3 pixels. The specific defect size limit is determined by the specific actual situation.
[0126] In addition, it should be noted that the defect localization information refers to the bounding box coordinates of the small target defects; the class prediction information refers to the defect type of the small target defects.
[0127] Specifically, the shallow feature map extracted from the backbone network, that is, the first data feature map, is upsampled by the upsampling module. The feature map is enlarged according to the specified upsampling ratio to obtain the upsampled feature map; the upsampled feature map and the shallow original feature map are spliced along the channel dimension by the splicing module to obtain a spliced feature map with an increased number of channels; the spliced feature map is subjected to feature equalization and concatenation operations by the C3K2 module to extract key information, and the number of channels is compressed. Finally, the lightweight detection head is used to perform localization prediction and classification prediction on small target defects to obtain the localization information and class prediction information of the small target defects.
[0128] Based on the above implementation solutions, in a feasible implementation manner, the step of inputting the second data feature map into the location prediction layer for object location classification to obtain the defect detection result of the image to be detected includes S231 to S234:
[0129] Step S231: Through the location prediction layer, fuse the second data feature map to obtain the second fused feature map of the second data feature.
[0130] Specifically, through the location prediction layer, the second data feature map is fused through a series of convolutional layers, upsampling layers, splicing layers, etc. to obtain the second fused feature map.
[0131] Step S232: Predict the bounding box of the target defect for the second fused feature map to obtain the bounding box prediction map of the target defect on the image to be detected.
[0132] Specifically, there is a series of anchor boxes or grid cells in the location prediction layer, and each anchor box or grid cell is responsible for predicting one or more bounding boxes. The coordinates and confidence scores of these bounding boxes are predicted through convolutional layers. The bounding box of the target defect in the second fused feature map is predicted through the anchor box or grid cell to obtain the bounding box prediction map.
[0133] Step S233: Perform class prediction and coordinate location for the bounding boxes on the bounding box prediction map to obtain the class and coordinate positions corresponding to the bounding boxes.
[0134] Specifically, for each bounding box prediction result, the location prediction layer also predicts the defect class it contains through another convolutional layer; the classification of the bounding box prediction result is completed through a classifier (such as a softmax layer or a sigmoid layer) to obtain the class label of the bounding box prediction. For each bounding box prediction result, obtain the bounding box coordinates of the target defect.
[0135] Step S234: Based on the class and coordinate positions corresponding to the bounding boxes, obtain the defect detection result of the target defect on the image to be detected.
[0136] Specifically, calculate non-maximum suppression based on the class and coordinate positions corresponding to the bounding boxes. Non-maximum suppression is a method for removing overlapping bounding boxes. It selects the optimal bounding box and removes redundant bounding boxes according to the confidence scores and overlapping degrees of the bounding boxes to obtain the final defect detection result.
[0137] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the defect detection method of this application. Based on this technical concept, more forms of simple transformations are within the protection scope of this application.
[0138] This application also provides a defect detection device. Please refer to Figure 4 , the defect detection device includes:
[0139] An image acquisition module 401, configured to acquire an image to be detected, and perform preprocessing of image adjustment and normalization on the image to be detected;
[0140] A defect detection module 402, configured to input the preprocessed image to be detected into a pre-constructed defect detection model, and perform defect detection on the image to be detected through the defect detection model to obtain a defect detection result of the image to be detected. The defect detection model includes a feature fusion layer, and the feature fusion layer includes a shallow feature detection head.
[0141] The defect detection device provided by this application adopts the defect detection method in the above embodiment, and can solve the technical problem of difficult detection of small target defects during defect detection. Compared with the prior art, the beneficial effects of the defect detection device provided by this application are the same as those of the defect detection method provided by the above embodiment, and other technical features in the defect detection device are the same as those disclosed in the method of the above embodiment, and will not be elaborated here.
[0142] This application provides a defect detection device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the defect detection method in the first embodiment above.
[0143] Next, refer to Figure 5 , which shows a schematic structural diagram of a defect detection device suitable for implementing the embodiments of this application. The defect detection device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The defect detection device shown is only an example, and should not bring any limitation to the functions and usage scopes of the embodiments of this application.
[0144] As Figure 5As shown, the defect detection device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the defect detection device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the defect detection device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a defect detection device having various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be implemented or had alternatively.
[0145] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.
[0146] The defect detection device provided by the present application adopts the defect detection method in the above-mentioned embodiment, and can solve the technical problem of difficult detection of small target defects during defect detection. Compared with the prior art, the beneficial effects of the defect detection device provided by the present application are the same as those of the defect detection method provided by the above-mentioned embodiment, and other technical features in the defect detection device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.
[0147] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0148] As described above, the above are only the specific embodiments of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0149] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the defect detection method in the above embodiments.
[0150] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems, or devices, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM: Random Access Memory), read-only memory (ROM: Read Only Memory), erasable programmable read-only memory (EPROM: Erasable Programmable Read Only Memory or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM: CD-Read Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0151] The above computer-readable storage medium can be included in the defect detection device; it can also exist separately without being assembled into the defect detection device.
[0152] The above computer-readable storage medium stores one or more programs, which, when executed by a defect detection device, cause the defect detection device to: obtain an image to be detected, and perform preprocessing of image adjustment and normalization on the image to be detected; input the preprocessed image to be detected into a pre-constructed defect detection model, and perform defect detection on the image to be detected through the defect detection model to obtain a defect detection result of the image to be detected. The defect detection model includes a feature fusion layer, and the feature fusion layer includes a shallow feature detection head.
[0153] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0154] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, which contains one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0155] The modules described in the embodiments of the present application may be implemented in software or in hardware. In some cases, the name of the module does not constitute a limitation on the unit itself.
[0156] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned defect detection method, which can solve the technical problem of difficult detection of small target defects during defect detection. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the defect detection method provided by the above embodiments, and will not be elaborated here.
[0157] This application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the defect detection method as described above are implemented.
[0158] The computer program product provided by this application can solve the technical problem of difficult detection of small target defects during defect detection. Compared with the prior art, the beneficial effects of the computer program product provided by this application are the same as those of the defect detection method provided by the above embodiments, and will not be elaborated here.
[0159] The above are only partial embodiments of this application, and do not limit the patent scope of this application accordingly. Any equivalent structural transformation made by using the content of the specification and drawings of this application under the technical concept of this application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of this application.
Claims
1. A defect detection method, characterized in that, The method includes: Obtain the image to be detected, and perform preprocessing of image adjustment and normalization on the image to be detected; Input the preprocessed image to be detected into a pre-constructed defect detection model, and perform defect detection on the image to be detected through the defect detection model to obtain the defect detection result of the image to be detected. The defect detection model includes a feature fusion layer, and the feature fusion layer includes a shallow feature detection head.
2. The method according to claim 1, wherein Before the step of inputting the preprocessed image to be detected into a pre-constructed defect detection model and performing defect detection on the image to be detected through the defect detection model to obtain the defect detection result of the image to be detected, it further includes: Obtain a number of groups of initial training images; Label the target defects in each of the initial training images through a labeling tool to generate labeled training images; Perform preprocessing of adjustment and normalization on each of the labeled training images, and divide the preprocessed labeled training images into a training image set and a validation image set; Input the training image set into the initial defect detection model for iterative training, and input the validation image set into the initial defect detection model after iterative training for performance evaluation to obtain the defect detection model.
3. The method according to claim 1, wherein The defect detection model further includes a feature extraction layer and a localization prediction layer. The step of inputting the preprocessed image to be detected into a pre-constructed defect detection model and performing defect detection on the image to be detected through the defect detection model to obtain the defect detection result of the image to be detected includes: Extract features from the preprocessed image to be detected through the feature extraction layer to obtain a first data feature map; Input the first data feature map into the feature fusion layer for feature fusion and enhancement to obtain a second data feature map after the first data feature map is fused and enhanced; Input the second data feature map into the localization prediction layer for target localization and classification to obtain the defect detection result of the image to be detected.
4. The method according to claim 3, wherein The step of extracting features from the preprocessed image to be detected through the feature extraction layer to obtain a first data feature map includes: Input the preprocessed image to be detected into the first extraction module of the feature extraction layer for convolution operation to obtain a first image feature map; Input the first image feature map into the second extraction module of the feature extraction layer for segmentation and splicing operation to obtain a second image feature map; Input the second image feature map into the third extraction module of the feature extraction layer for feature enhancement operation to obtain a third image feature map; Input the third image feature map into the fourth extraction module of the feature extraction layer for feature extraction and fusion operation to obtain the first data feature of the image to be detected.
5. The method according to claim 3, wherein The step of inputting the first data feature map into the feature fusion layer for feature fusion and enhancement to obtain a second data feature map after the first data feature map is fused and enhanced includes: Perform upsampling on the first data feature map through the feature fusion layer; Perform a splicing operation on the first data feature map after upsampling is completed to obtain a first spliced feature map of the first data feature map; Feature extraction and feature fusion operations are performed on the first spliced feature map to obtain a second data feature map with enhanced fusion of the first data feature map.
6. The method according to claim 5, characterized in that, The step of inputting the first data feature map into the feature fusion layer for enhanced feature fusion to obtain a second data feature map with enhanced fusion of the first data feature map further includes: The shallow feature detection head is used to perform regression prediction of small target defects on the first data feature map to obtain defect location information and category prediction information of the small target defects.
7. The method according to claim 3, characterized in that The step of inputting the second data feature map into the location prediction layer for target location classification to obtain the defect detection result of the image to be detected includes: The second data feature map is fused through the location prediction layer to obtain a second fused feature map of the second data feature; Boundary box prediction of target defects is performed on the second fused feature map to obtain a boundary box prediction map of target defects on the image to be detected; Category prediction and coordinate location are performed on the boundary boxes on the boundary box prediction map to obtain the category and coordinate positions corresponding to the boundary boxes; Based on the category and coordinate positions corresponding to the boundary boxes, the defect detection result of the target defects on the image to be detected is obtained.
8. A defect detection device, characterized in that, The device includes: An image acquisition module, configured to acquire an image to be detected and perform preprocessing of image adjustment and normalization on the image to be detected; A defect detection module, configured to input the preprocessed image to be detected into a pre-constructed defect detection model, and perform defect detection on the image to be detected through the defect detection model to obtain the defect detection result of the image to be detected. The defect detection model includes a feature fusion layer, and the feature fusion layer includes a shallow feature detection head.
9. A defect detection device, characterized in that, The device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. The computer program is configured to implement the steps of the defect detection method according to any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps of the defect detection method according to any one of claims 1 to 7.