A defect detection method and storage medium based on high-resolution images

Through the composite feature extraction and multi-scale high-frequency feature extraction of the dual-branch network, combined with the feature fusion module, the problem of poor detection of small target defects in high-resolution images is solved, and efficient defect detection and real-time detection are achieved.

CN116596866BActive Publication Date: 2025-07-22SHENZHEN HUAHAN WEIYE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310498557.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-05
Publication Date
2025-07-22
Estimated Expiration
2043-05-05

AI Technical Summary

Technical Problem

The existing defect detection methods have poor detection of small target defects in high-resolution images, and due to the large image size, the memory limit and the detection time are too long, which cannot meet the industrial real-time requirements.

Method used

A dual-branch network is adopted, including a composite feature extraction module and a multi-scale high-frequency feature extraction module, which extracts low-frequency and high-frequency features of high-resolution images respectively, and features are fused through a multi-scale feature fusion module, and finally defect detection is used to use the detection module.

Benefits of technology

It improves the accuracy of small-target defect detection, solves the memory limitation problem, and realizes the whole image training and real-time detection of high-resolution images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116596866B_ABST
    Figure CN116596866B_ABST
Patent Text Reader

Abstract

A defect detection method and storage medium based on high-resolution images. The method includes: obtaining a high-resolution image of an object to be measured; inputting the high-resolution image into a trained defect detection model to obtain a defect detection result of the high-resolution image, including: using a composite feature extraction module to extract features from the high-resolution image to obtain a composite feature map of the high-resolution image; using a multi-scale high-frequency feature extraction module to extract high-frequency features from the high-resolution image to obtain multiple high-frequency feature maps of different scales of the high-resolution image; using a multi-scale feature fusion module to perform feature fusion on the composite feature map and the high-frequency feature maps to obtain a fused feature map; using a detection module to process the fused feature map to obtain a defect detection result of the high-resolution image. By fusing low-frequency features and multi-scale high-frequency features, the detailed information of the high-resolution image is retained as much as possible, thereby improving the accuracy of small target defect detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of defect detection, and in particular to a defect detection method and a storage medium based on high-resolution images. Background Art

[0002] At present, the premise of surface inspection of various products in industrial production is to image the above products. With the continuous improvement of the quality requirements for the above various products, low-resolution cameras cannot image the fine defects of the above products, and the inability to image the fine defects of the above products means that machine vision technology cannot be used to automatically detect the defects of industrial production products, thus making visual inspection meaningless. And using manual methods to inspect the above products on-site is time-consuming, laborious, and costly, which further causes production enterprises to be unable to bear such manual inspection methods. In actual industrial scenarios, in order to accurately control the quality of products, manufacturers generally use high-resolution cameras to perform refined imaging on the surface of products to obtain high-resolution images, so as to maximize the defects in the products through the above high-resolution images, and then achieve accurate detection of products. Since high-resolution cameras can accurately image the above products with more pixels to better present defects of various scales, more and more production enterprises use high-resolution cameras to collect product images, which has led to a strong demand for defect detection of high-resolution images. However, due to the huge size of high-resolution images, it poses a major challenge to existing defect detection methods / systems. The above major challenges are mainly reflected in: the poor detection effect on small target defects. Since the number of defect pixels in the above high-resolution images is small, the area of the defects is small, and the ratio of defect pixels to background pixels is huge, it is difficult to detect the defects in the above images; in addition, since the proportion of pixels of small target defects in the above high-resolution images to the total pixels is very small, and most of them present the characteristics of "fine and weak", it further increases the difficulty of detecting the above defects. Therefore, it is necessary to improve the existing technology in view of its deficiencies. Summary of the Invention

[0003] This application provides a defect detection method based on high-resolution images to solve the technical defect that the existing defect detection methods have poor detection effects on small target defects.

[0004] According to a first aspect, in one embodiment, a defect detection method based on a high-resolution image is provided. The above-mentioned defect detection method includes: obtaining a high-resolution image of an object to be measured; inputting the high-resolution image into a trained defect detection model to obtain a defect detection result of the high-resolution image, where the defect detection result includes a defect area and / or a defect category in the high-resolution image; wherein the defect detection model includes a composite feature extraction module, a multi-scale high-frequency feature extraction module, a multi-scale feature fusion module, and a detection module; the step of inputting the high-resolution image into the trained defect detection model to obtain the defect detection result of the high-resolution image includes: using the composite feature extraction module to extract features from the high-resolution image to obtain a composite feature map of the high-resolution image, where the composite feature map is used to represent the low-frequency and high-frequency features of the high-resolution image; using the multi-scale high-frequency feature extraction module to extract high-frequency features from the high-resolution image to obtain multiple high-frequency feature maps with different scales of the high-resolution image, where the high-frequency feature maps are used to represent the high-frequency features of the high-resolution image; using the multi-scale feature fusion module to perform feature fusion on the composite feature map and the high-frequency feature maps to obtain a fused feature map; using the detection module to process the fused feature map to obtain the defect detection result of the high-resolution image.

[0005] In some embodiments, the composite feature extraction module includes a convolutional layer, a max pooling layer, and multiple cascaded residual modules, and each of the residual modules includes two convolutional sub-layers; the input feature map of the max pooling layer is the output feature map of the convolutional layer, and the output of each residual module is added to the output of the previous-level residual module as the input of the next-level residual module, where the input feature map of the first residual module is the output feature map of the max pooling layer, and the input feature map of the second residual module is the feature map obtained by adding the output of the first residual module and the output feature map of the max pooling layer.

[0006] In some embodiments, the step of using the multi-scale high-frequency feature extraction module to extract high-frequency features from the high-resolution image to obtain multiple high-frequency feature maps with different scales of the high-resolution image includes: extracting the high-frequency component of the high-resolution image to obtain a high-frequency component image of the high-resolution image; inputting the high-frequency component image into the multi-scale high-frequency feature extraction module for feature extraction to obtain the multiple high-frequency feature maps with different scales.

[0007] In some embodiments, extracting the high-frequency component of the high-resolution image to obtain the high-frequency component image of the high-resolution image includes: performing low-pass filtering on the high-resolution image to obtain a first filtered image, subtracting the first filtered image from the high-resolution image to obtain a first high-frequency residual image; performing low-pass filtering on the first high-frequency residual image to obtain a second filtered image, subtracting the second filtered image from the first high-frequency residual image to obtain a second high-frequency residual image; performing channel splicing on the first high-frequency residual image and the second high-frequency residual image to obtain the high-frequency component image, and the resolution of the high-frequency component image is the same as that of the high-resolution image.

[0008] In some embodiments, the high-frequency feature map includes a first high-frequency feature map and a second high-frequency feature map, and the multi-scale high-frequency feature extraction module includes a first standard convolutional layer, a max-pooling layer, a first high-frequency feature extraction module, and a second high-frequency feature extraction module, where the standard convolutional layer includes a convolutional layer, a batch normalization layer, and an activation layer connected in sequence; inputting the high-frequency component image into the multi-scale high-frequency feature extraction module for feature extraction to obtain the multiple high-frequency feature maps with different scales includes: inputting the high-frequency component image into the first standard convolutional layer, and inputting the output feature map of the first standard convolutional layer into the max-pooling layer to obtain a first pooled feature map; inputting the first pooled feature map into the first high-frequency feature extraction module for feature extraction to obtain the first high-frequency feature map; inputting the first high-frequency feature map into the second high-frequency feature extraction module for further feature extraction to obtain the second high-frequency feature map, and the resolution of the first high-frequency feature map is greater than that of the second high-frequency feature map.

[0009] In some embodiments, the first high-frequency feature extraction module includes a second standard convolutional layer, a third standard convolutional layer, a fourth standard convolutional layer, an attention sub-module, and a fifth standard convolutional layer, and the convolutional kernel sizes of the second standard convolutional layer and the third standard convolutional layer are different; the step of inputting the first pooled feature map into the first high-frequency feature extraction module for feature extraction to obtain the first high-frequency feature map includes: performing a channel separation operation on the first pooled feature map to obtain a first separated feature map and a second separated feature map, where the number of channels of the first separated feature map and the second separated feature map are equal; inputting the first separated feature map into the second standard convolutional layer to obtain a second convolutional feature map, and inputting the second separated feature map into the third standard convolutional layer to obtain a third convolutional feature map; after concatenating the second convolutional feature map and the third convolutional feature map in channels, inputting them into the fourth standard convolutional layer to obtain a fourth convolutional feature map; inputting the fourth convolutional feature map into the attention sub-module to obtain a feature attention map; adding the feature attention map and the first pooled feature map, and then inputting the result into the fifth standard convolutional layer to obtain the first high-frequency feature map; the second high-frequency feature extraction module has the same structure as the first high-frequency feature extraction module, where the input feature map of the second standard convolutional layer in the second high-frequency feature extraction module is a third separated feature map, and the input feature map of the third standard convolutional layer is a fourth separated feature map, and the third separated feature map and the fourth separated feature map are obtained by performing a channel separation operation on the first high-frequency feature map, and the number of channels of the third separated feature map and the fourth separated feature map are equal; the output of the fifth standard convolutional layer in the second high-frequency feature extraction module is the second high-frequency feature map.

[0010] In some embodiments, the attention sub-module includes an average pooling layer, a sixth standard convolutional layer, and a Sigmoid function layer; the step of inputting the fourth convolutional feature map into the attention sub-module to obtain a feature attention map includes: inputting the fourth convolutional feature map into the average pooling layer for pooling and then inputting it into the sixth standard convolutional layer to obtain a sixth convolutional feature map; inputting the sixth convolutional feature map into the Sigmoid function layer to obtain the channel weights of the fourth convolutional feature map; multiplying the fourth convolutional feature map by the channel weights to obtain the feature attention map.

[0011] In some embodiments, the multi-scale feature fusion module includes multiple seventh standard convolutional layers with different dilation rates and an eighth standard convolutional layer; the method of using the multi-scale feature fusion module to perform feature fusion on the composite feature map and the high-frequency feature map to obtain a fused feature map includes: inputting the composite feature map into the multiple seventh standard convolutional layers with different dilation rates respectively to obtain multiple seventh convolutional feature maps with different scale features; concatenating all the seventh convolutional feature maps in channels and inputting the result into the eighth standard convolutional layer to obtain a first intermediate fused feature map, the resolution of the first intermediate fused feature map being smaller than that of the second high-frequency feature map; upsampling the first intermediate fused feature map to the same resolution as the second high-frequency feature map and concatenating it with the second high-frequency feature map in channels to obtain a second intermediate fused feature map; upsampling the second intermediate fused feature map to the same resolution as the first high-frequency feature map and concatenating it with the first high-frequency feature map in channels to obtain the fused feature map.

[0012] In some embodiments, the detection module includes three ninth standard convolutional layers; the defect detection model is trained in the following manner: obtaining a training sample image and corresponding annotation data; inputting the training sample image into the defect detection model to obtain a first intermediate fused feature map, a second intermediate fused feature map, and a fused feature map of the training sample image; inputting the first intermediate fused feature map, the second intermediate fused feature map, and the fused feature map into different ninth standard convolutional layers respectively to obtain ninth convolutional feature maps of the first intermediate fused feature map, the second intermediate fused feature map, and the fused feature map respectively; upsampling the ninth convolutional feature maps of the first intermediate fused feature map, the second intermediate fused feature map, and the fused feature map to the same resolution as the training sample image to obtain a first defect prediction map, a second defect prediction map, and a third defect prediction map; wherein, the elements of the defect prediction map represent the probabilities of the pixel points at the corresponding positions belonging to each classification category, where the classification categories include background pixels and each defect category; determining a total loss function L according to a first loss function L1 determined by the first defect prediction map and the annotation data, a second loss function L2 determined by the second defect prediction map and the annotation data, and a third loss function L3 determined by the third defect prediction map and the annotation data total , and training the defect detection model according to the total loss function.

[0013] In some embodiments, the expressions of the first loss function, the second loss function, and the third loss function are:

[0014] L = αL cross + βL dice ,

[0015] where \(L\) represents any one of the first loss function, the second loss function, and the third loss function, \(\alpha\) and \(\beta\) are weight coefficients, and \(L\) cross is the cross - entropy loss function, and \(L\) dice is the intersection - over - union loss function, and

[0016]

[0017] where \(i\) represents the pixel point serial number, \(n\) represents the total number of pixel points of the first defect prediction map, the second defect prediction map, or the third defect prediction map, and \(LP\) i represents the cross - entropy loss value of the \(i\) - th pixel point. The cross - entropy loss value \(LP\) of each pixel point is:[[]]

[0018]

[0019] where \(c\) is the defect category serial number, \(C\) is the number of defect categories, and \(p\) c is the probability that the pixel point belongs to the \(c\) - th defect category, and \(g\) c is the corresponding category annotation value;

[0020]

[0021] where \(X\) is the defect region segmented from the first defect prediction map, the second defect prediction map, or the third defect prediction map, and \(Y\) is the corresponding defect annotation region.

[0022] In some embodiments, the expression of the total loss function is:

[0023] \(L\) total = \(L1 + L2+L3\)

[0024] In some embodiments, the detection module includes a ninth standard convolutional layer; obtaining the defect detection result of the high - resolution image from the fused feature map includes: inputting the fused feature map into the ninth standard convolutional layer to obtain a ninth convolutional feature map, upsampling the ninth convolutional feature map to the same resolution as the high - resolution image to obtain a defect prediction map, where the elements of the defect prediction map represent the probabilities that the pixel points at the corresponding positions belong to each classification category, and the classification categories include background pixels and each defect category; obtaining a defect segmentation map of the high - resolution image from the defect prediction map, where the defect segmentation map is used to display the defect regions and the corresponding defect categories in the high - resolution image.

[0025] In some embodiments, the multi-scale high-frequency feature extraction module is a lightweight module compared to the composite feature extraction module; the step of using the composite feature extraction module to extract features from the high-resolution image to obtain the composite feature map of the high-resolution image includes: performing downsampling processing on the high-resolution image, and inputting the downsampled high-resolution image into the composite feature extraction module to obtain the composite feature map of the high-resolution image.

[0026] According to a second aspect, in one embodiment, a computer-readable storage medium is provided. A program is stored on the medium, and the program can be executed by a processor to implement the defect detection method as described in any embodiment of the present application.

[0027] Aiming at the technical defect that the existing defect detection methods have poor detection effect on small target defects, the defect detection model of the present application adopts a dual-branch network (i.e., the composite feature extraction module and the multi-scale high-frequency feature extraction module in the present application); the composite feature extraction module is responsible for extracting features from the high-resolution image to obtain the composite feature map of the high-resolution image, and the composite feature map is used to represent the low-frequency and high-frequency features of the high-resolution image; while the multi-scale high-frequency feature extraction module extracts high-frequency features from the high-resolution image to obtain multiple high-frequency feature maps with different scales of the high-resolution image, and the high-frequency feature maps are used to represent the high-frequency features of the high-resolution image; the above high-frequency feature maps can retain the position information and the feature information of small targets in the high-resolution image; then, the multi-scale feature fusion module of the defect detection model is used to fuse the composite feature map and the high-frequency feature maps to obtain a fused feature map; finally, the detection module of the defect detection model is used to process the fused feature map to obtain the defect detection result of the high-resolution image. Since the low-frequency features and multi-scale high-frequency features are fused, it is beneficial to retain the detail information of the high-resolution image as much as possible, thereby greatly improving the accuracy of small target defect detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 It is a flowchart of a defect detection method based on a high-resolution image in an embodiment of the present application;

[0029] Figure 2 It is a schematic diagram of the overall network structure of a defect detection model in an embodiment of the present application;

[0030] Figure 3 It is a schematic diagram of the network structure of a composite feature extraction module in an embodiment of the present application;

[0031] Figure 4 It is a schematic diagram of the network structure of a multi-scale high-frequency feature extraction module in an embodiment of the present application;

[0032] Figure 5 Schematic diagram of the network structure of the first high-frequency feature extraction module in an embodiment of the present application;

[0033] Figure 6 Schematic diagram of the network structure of the attention sub-module in an embodiment of the present application;

[0034] Figure 7 Schematic diagram of the network structure of the multi-scale feature fusion module in an embodiment of the present application;

[0035] Figure 8 Schematic diagram of the overall network structure of the defect detection model in an embodiment of the present application;

[0036] Figure 9 Schematic diagram of the input and output of the defect detection model in an embodiment of the present application; wherein, Figure 9 Figure a in it represents the high-resolution image input into the defect detection model; Figure 9 Figure b in it represents the defect segmentation map output by the defect detection model;

[0037] Figure 10 Schematic diagram of the overall network structure of the defect detection model in another embodiment of the present application;

[0038] Figure 11 Flow chart of the training of the defect detection model in an embodiment of the present application. Detailed implementation manners

[0039] The present application will be further described in detail below in conjunction with the accompanying drawings through specific implementation manners. Similar elements in different implementation manners are labeled with related similar element numbers. In the following implementation manners, many detailed descriptions are provided to enable a better understanding of the present application. However, those skilled in the art can easily recognize that some of the features can be omitted in different situations, or can be replaced by other elements, materials, or methods. In some cases, some operations related to the present application are not shown or described in the specification to avoid the core part of the present application being overwhelmed by excessive descriptions. For those skilled in the art, it is not necessary to describe these related operations in detail, and they can fully understand the related operations based on the descriptions in the specification and the general technical knowledge in the art.

[0040] In addition, the features, operations, or characteristics described in the specification can be combined in any appropriate manner to form various implementation manners. At the same time, the steps or actions in the method description can also be reordered or adjusted in an obvious manner for those skilled in the art. Therefore, the various sequences in the specification and the drawings are only for clearly describing a certain embodiment and do not mean that they are the necessary sequences, unless it is stated that a certain sequence must be followed.

[0041] The serial numbers assigned to components in this text, such as "first", "second", etc., are only used to distinguish the described objects and do not have any sequential or technical meaning. The terms "connection" and "coupling" as used in this application, unless otherwise specified, both include direct and indirect connection (coupling).

[0042] Due to the huge size of the high-resolution image of the object to be measured, significant challenges are posed to existing defect detection methods. The above-mentioned significant challenges are mainly reflected in: poor detection effect for small target defects. Since the number of defective pixels in the above high-resolution image is small, the area of the defect is small, and the ratio of defective pixels to background pixels is extremely different, it is difficult to detect the defects in the above image; in addition, since the proportion of pixels of small target defects in the above high-resolution image to the total pixels is very small and mostly presents the characteristics of "fine, small, and weak", it further increases the difficulty of detecting the above defects.

[0043] The above-mentioned significant challenges are also reflected in: memory limitation. For high-resolution images, due to their huge size, they require a high amount of computing memory. Limited by hardware resources, existing defect detection methods cannot process the entire high-resolution image at once. To address the above technical deficiencies, the common solution adopted by existing defect detection methods is to downsample the high-resolution image or crop it into blocks and then process them separately. However, the process of downsampling the high-resolution image will result in the loss of many detailed features in the high-resolution image, thereby leading to poor detection effects (especially poor detection effects for small target defects); while the method of cropping the high-resolution image into blocks for processing will result in an incomplete context semantic structure of the image, destroying the original semantic information structure of the image, and making the subsequent inference process take too long, thus unable to meet the real-time requirements of the industrial community for object detection.

[0044] In summary, the main characteristics of detecting the product (i.e., the object to be measured) through the high-resolution image of the object to be measured are the large size of the original image (i.e., the high-resolution image) and the poor detection effect for small target defects. Therefore, it is necessary to improve the existing technology to enhance the detection effect. Among them, the above high resolution generally refers to 2K resolution or higher (including 2K), such as a resolution of 2048*1080.

[0045] The present application provides a defect detection method based on high-resolution images to solve some or all of the above problems. Among them, aiming at the technical defect that the existing defect detection methods have poor detection effect on small target defects, the present application designs a dual-branch network (that is, the composite feature extraction module and the multi-scale high-frequency feature extraction module in the present application). Among them, the composite feature extraction module is responsible for extracting features from the high-resolution image to obtain a composite feature map of the high-resolution image, and the composite feature map is used to represent the low-frequency and high-frequency features of the high-resolution image; while the multi-scale high-frequency feature extraction module extracts the high-frequency features in the high-resolution image to obtain multiple high-frequency feature maps with different scales of the high-resolution image, and the high-frequency feature map is used to represent the high-frequency features of the high-resolution image; the above high-frequency feature maps can retain the position information and the feature information of small targets in the high-resolution image; then, the multi-scale feature fusion module is used to fuse the composite feature map and the high-frequency feature map to obtain a fused feature map; then, the detection module is used to process the fused feature map to obtain the defect detection result of the high-resolution image and complete the task of defect segmentation of the high-resolution image, ultimately greatly improving the accuracy of small target defect detection and enabling whole-image training and inference of high-resolution images, etc.

[0046] The technical solution of the present application will be specifically described below in conjunction with embodiments.

[0047] Some embodiments of the present application disclose a defect detection method based on high-resolution images. Please refer to Figure 1 , the above defect detection method includes:

[0048] Step S100: Obtain a high-resolution image of the object to be measured;

[0049] Step S200: Input the high-resolution image into the trained defect detection model to obtain the defect detection result of the high-resolution image, and the defect detection result includes the defect area and / or defect category in the high-resolution image.

[0050] It should be noted that the above-mentioned object to be measured can be products on an industrial assembly line, mechanical parts in an object box, tools on an operating table, etc. There is no specific limitation on the object to be measured here. A high-resolution vision sensor such as a high-resolution camera can be used to image the object to be measured to obtain a high-resolution image of the object to be measured. This image is the graphic information on the surface of the object to be measured, reflecting a part of the appearance shape of the surface of the object to be measured. The high-resolution camera used to photograph the object to be measured can be a grayscale or color camera, that is, the obtained high-resolution image can be a grayscale image or a color image. When the high-resolution image is a color image, it can be converted into a grayscale image for processing. If there are defects such as cracks, flaws, dirt, etc. on the surface of the object to be measured, then in the high-resolution image of the object to be measured obtained by photographing, these defects will be shown or presented on the corresponding abnormal continuous background pattern.

[0051] The detection result of the above high-resolution image can be the defect type and / or defect area in the high-resolution image. Those skilled in the art can determine / define the specific types of the above defect types according to the specific application scenarios. The defect area is used to show the part of the object to be measured where there are defects. Based on the defect area, the detector can locate the position of the defect on the object to be measured. It can be understood that the above defect detection result can be presented in the form of a defect segmentation image, and there is no limitation on the specific presentation form of the defect area here. In the defect segmentation image, the defect area and the background area are displayed in different colors. For example, the defect area is displayed in white and the background area is displayed in black; in some embodiments, the defect segmentation image can also display the defect type corresponding to the defect area. Specifically, different colors can be used to display the defect area according to the defect type, and each color corresponds to a defect type. In this way, the detector can know the part where there are defects on the object to be measured and the corresponding defect type according to the defect segmentation image.

[0052] Please refer to Figure 2 , the defect detection model provided by this application includes a composite feature extraction module 100, a multi-scale high-frequency feature extraction module 200, a multi-scale feature fusion module 300, and a detection module 400; please refer to Figure 1 , based on this defect detection model, the above step S200 includes:

[0053] Step S210: Use the composite feature extraction module 100 to extract features from the high-resolution image to obtain a composite feature map of the high-resolution image. The composite feature map is used to represent the low-frequency and high-frequency features of the high-resolution image;

[0054] Step S220: Use the multi-scale high-frequency feature extraction module 200 to extract high-frequency features from the high-resolution image, so as to obtain multiple high-frequency feature maps with different scales of the high-resolution image, and the high-frequency feature maps are used to characterize the high-frequency features of the high-resolution image;

[0055] Step S230: Use the multi-scale feature fusion module 300 to perform feature fusion on the composite feature map and the high-frequency feature map to obtain a fused feature map;

[0056] Step S240: Use the detection module 400 to process the fused feature map to obtain the defect detection result of the high-resolution image.

[0057] Among them, the low-frequency feature is the area where the image intensity (such as brightness or gray value, etc.) in the image changes slowly, that is, the feature of the large flat (i.e., slow-changing) area in the image, such as the area where large color blocks are located in the image. The high-frequency feature is the feature of some parts of the image where the image intensity (such as brightness or gray value, etc.) changes violently, such as the edge (outline) of the image or the feature of the detail part or noise (i.e., noise points). It can be seen that this application fuses the low-frequency feature and the multi-scale high-frequency feature for defect detection, and retains the detail information of the high-resolution image as much as possible, thereby improving the accuracy of small target defect detection.

[0058] The following will specifically describe each module in the defect detection model and the above steps S210-S240 with reference to the accompanying drawings.

[0059] The composite feature extraction module 100 is mainly used to extract features from the complete or downsampled high-resolution image, and obtain the low-frequency features and high-frequency features therein. The composite feature extraction module 100 can extract features by filtering. For example, the high-resolution image is respectively subjected to low-pass and high-pass filtering, and the obtained results are fused to obtain a composite feature map. Or the composite feature extraction module 100 can be a deep learning-based feature extractor, such as a convolutional neural network; input the complete or downsampled high-resolution image into the feature extractor for feature extraction, and the obtained feature map includes both low-frequency features and high-frequency features, that is, a composite feature map is obtained. The feature extractor can adopt Res2Net, VGG, etc.

[0060] In some embodiments, please refer to Figure 3, the composite feature extraction module 100 includes a convolutional layer 101, a max pooling layer 102, and multiple cascaded residual modules 103. Each residual module 103 includes two convolutional sub-layers; the input feature map of the max pooling layer 102 is the output feature map of the convolutional layer 101. The output of each residual module 103 is added to the output of the previous-level residual module 103 as the input of the next-level residual module 103. Among them, the input feature map of the first residual module 103 is the output feature map of the max pooling layer 102, and the input feature map of the second residual module 103 is the feature map obtained by adding the output feature map of the first residual module 103 and the output feature map of the max pooling layer 102.

[0061] Those skilled in the art can select the specific structural parameters of the convolutional layer and the max pooling layer 102 of the composite feature extraction module 100 according to the actual application scenario (such as the stride of the convolutional kernel, the number of output channels, etc.). For example, the size of the convolutional kernel of the convolutional layer 101 of the composite feature extraction module 100 can be 7×7, the stride can be 2, and the number of output channels can be 64. The pooling window size of the max pooling layer 102 of the composite feature extraction module 100 can be 3×3, and the stride is 2. The number of cascaded residual modules 103 can be 8. Those skilled in the art can also determine the number of cascaded residual modules 103 in the composite feature extraction module according to the actual application scenario requirements. Figure 3 Taking 4 as an example, but not limited thereto. As Figure 3 shown, the input feature map of the first residual module 103 is the output feature map of the max pooling layer 102, and the input feature map of the second residual module 103 is the feature map obtained by adding the output feature map of the first residual module 103 and the output feature map of the max pooling layer 102. The output of each residual module 103 after the first residual module 103 (i.e., the second-level residual module 103 and each subsequent level of residual module 103) is added to the output of the previous-level residual module 103 as the input of the next-level residual module 103. For example, the output feature map of the second residual module 103 is added to the output feature map of the first residual module 103 as the input feature map of the third-level residual module 103, and so on for the residual modules after the second residual module. Each residual module 103 includes two convolutional sub-layers. Among them, the two convolutional sub-layers can both be 3×3 convolutional layers (that is, the size of the convolutional kernel in both convolutional sub-layers is 3×3), and the number of output channels of the two convolutional sub-layers in the residual module 103 is the same. That is to say, the feature map input to the residual module 103 is first subjected to convolutional processing by the previous convolutional sub-layer within the residual module 103; the feature result obtained by the convolutional processing of the previous convolutional sub-layer is input to the subsequent convolutional sub-layer within the residual module 103 for further convolutional processing; the output result of the subsequent convolutional sub-layer within the residual module 103 is used as the output of the residual module 103.

[0062] In some embodiments, step S220 includes step S221 and step S222, which will be described separately below.

[0063] Step S221: Extract the high-frequency components of the high-resolution image to obtain the high-frequency component image of the high-resolution image.

[0064] For the extraction of the high-frequency components of the high-resolution image, methods such as high-frequency filtering can be used. In addition, in an embodiment of the present application, another method for extracting the high-frequency components of the high-resolution image is provided. Specifically, step S221 includes: performing low-pass filtering on the high-resolution image to obtain a first filtered image, subtracting the first filtered image from the high-resolution image to obtain a first high-frequency residual image; performing low-pass filtering on the first high-frequency residual image to obtain a second filtered image, subtracting the second filtered image from the first high-frequency residual image to obtain a second high-frequency residual image; performing channel splicing on the first high-frequency residual image and the second high-frequency residual image to obtain a high-frequency component image, where the high-frequency component image has the same resolution as the high-resolution image. Here, the first filtered image is the result of performing low-pass filtering on the high-resolution image, and the second filtered image is the result of performing low-pass filtering on the first high-frequency residual image.

[0065] The low-pass filtering can use Gaussian filtering. Specifically, a Gaussian kernel K can be preset, and the high-resolution image is processed by Gaussian filtering using the Gaussian kernel K to obtain the first filtered image, and then the first high-frequency residual image is obtained by subtracting the first filtered image from the high-resolution image. Then, the first high-frequency residual image is continuously processed by Gaussian filtering using the Gaussian kernel K to obtain the second filtered image, and the second high-frequency residual image is obtained by subtracting the second filtered image from the first high-frequency residual image. In one embodiment, the Gaussian kernel K can be:

[0066]

[0067] Those skilled in the art can also select other suitable Gaussian kernels K according to the actual scenario requirements.

[0068] It should be noted that the above high-frequency component image has the same resolution as the high-resolution image, and the purpose is to retain high-frequency features such as position information and details of small target defects in the high-resolution image. In the specific implementation process, the first high-frequency residual image, the second high-frequency residual image, and the high-frequency component image can all have the same size as the input image (i.e., the high-resolution image). Therefore, appropriate padding operations (such as mirror padding) need to be performed on the high-resolution image or the first high-frequency residual image before performing the low-pass filtering process.

[0069] In addition, in order to increase the receptive field and obtain better detection effects, two Gaussian filtering and difference processing operations are performed in this embodiment. That is, on the basis of the first obtained first high-frequency residual map, Gaussian filtering is continued to be used, and then the first high-frequency residual map is subtracted from the result of the second filtering (i.e., the second filtered map) to obtain a second high-frequency residual map; the first high-frequency residual map and the second high-frequency residual map are concatenated on the channel to obtain a high-frequency component image.

[0070] Step S222: Input the high-frequency component image into the multi-scale high-frequency feature extraction module 200 for feature extraction to obtain multiple high-frequency feature maps with different scales.

[0071] Since the above high-frequency component image retains the position information in the original image (i.e., the high-resolution image) and the detailed information of small target defects, the high-frequency component image can be input into the multi-scale high-frequency feature extraction module 200 for feature extraction to obtain high-frequency feature maps representing the high-frequency features in the original image, and multiple high-frequency feature maps with different scales containing high-frequency features can be further obtained, which is beneficial to the subsequent accurate detection and segmentation of small target defects in the high-resolution image.

[0072] The multi-scale high-frequency feature extraction module 200 can be implemented by using existing technologies. For example, a feature extractor based on deep learning can be used. Those skilled in the art can, according to needs, set the multi-scale high-frequency feature extraction module 200 to directly output multiple high-frequency feature maps with different scales, or set the multi-scale high-frequency feature extraction module 200 to first perform feature extraction on the high-frequency component image to obtain a high-frequency feature map, and then perform further feature extraction and downsampling on the high-frequency feature map to obtain another high-frequency feature map of a different scale.

[0073] Please refer to Figure 4 , in some embodiments, the above high-frequency feature maps include a first high-frequency feature map f3 and a second high-frequency feature map f2, and the multi-scale high-frequency feature extraction module 200 includes a first standard convolutional layer 210, a max pooling layer 220, a first high-frequency feature extraction module 230, and a second high-frequency feature extraction module 240, where the standard convolutional layer includes a convolutional layer, a batch normalization layer, and an activation layer connected in sequence; the above step S222 includes:

[0074] Step S222a: Input the high-frequency component image into the first standard convolutional layer 210, and input the output feature map of the first standard convolutional layer 210 into the max pooling layer 220 to obtain a first pooled feature map;

[0075] Step S222b: Input the first pooled feature map into the first high-frequency feature extraction module 230 for feature extraction to obtain the first high-frequency feature map f3;

[0076] Step S222c: Input the first high-frequency feature map f3 into the second high-frequency feature extraction module 240 for further feature extraction to obtain the second high-frequency feature map f2, where the resolution of the first high-frequency feature map f3 is greater than that of the second high-frequency feature map f2. For example, the resolution of the second high-frequency feature map f2 is half of that of the first high-frequency feature map f3.

[0077] In some embodiments, the convolutional layer of the first standard convolutional layer 210 in the multi-scale high-frequency feature extraction module 200 may adopt a 3×3 convolutional layer (i.e., a convolutional layer with a convolutional kernel size of 3×3), and the activation layer may adopt a ReLU activation function.

[0078] Preferably, the high-frequency feature maps output by the multi-scale high-frequency feature extraction module 200 only include the first high-frequency feature map f3 and the second high-frequency feature map f2. The reason is as follows: If the multi-scale high-frequency feature extraction module 200 only outputs one type of high-frequency feature map, the information extraction completed by the multi-scale high-frequency feature extraction module 200 is still not sufficient, and the receptive field of the uniquely output high-frequency feature map is also small, so its limitations are large; If the multi-scale high-frequency feature extraction module 200 outputs two high-frequency feature maps of different scales (i.e., the first high-frequency feature map and the second high-frequency feature map) through two consecutive steps of information extraction, then the above two high-frequency feature maps can meet the comprehensiveness and effectiveness of feature extraction, and will not cause the resolution of the feature map to be greatly reduced and affect the final defect detection; If the multi-scale high-frequency feature extraction module 200 outputs more than two high-frequency feature maps, since the resolution of the subsequent high-frequency feature map is lower than that of the previous high-frequency feature map, it will result in the resolution of some high-frequency feature maps being very low and lacking effective feature information, making these high-frequency feature maps less useful in subsequent defect detection.

[0079] The function of the multi-scale high-frequency feature extraction module 200 is to extract / retain high-frequency features such as position information and small target defect feature information in the high-resolution image as much as possible; Among them, the first high-frequency feature extraction module 230 and the second high-frequency feature extraction module 240 in the multi-scale high-frequency feature extraction module 200 are used to extract high-frequency features such as position information and small target defect feature information in the above high-resolution image to obtain high-frequency feature maps. By designing this dedicated branch module of the multi-scale high-frequency feature extraction module 200, the detailed information (i.e., high-frequency features) of the high-resolution image is extracted as much as possible, enabling the defect detection model to make full use of the defect features of small targets, so as to be able to accurately predict, greatly improving the defect detection ability of the defect detection model for small target defects and achieving accurate positioning of object defects.

[0080] Please refer to Figure 5, in some embodiments, the first high-frequency feature extraction module 230 includes a second standard convolutional layer 231, a third standard convolutional layer 232, a fourth standard convolutional layer 233, an attention sub-module 234, and a fifth standard convolutional layer 235. The convolutional kernels of the second standard convolutional layer 231 and the third standard convolutional layer 232 have different sizes. The above step S222b includes: performing a channel separation operation on the first pooled feature map to obtain a first separated feature map and a second separated feature map, where the number of channels of the first separated feature map and the second separated feature map is equal; inputting the first separated feature map into the second standard convolutional layer 231 to obtain a second convolutional feature map, and inputting the second separated feature map into the third standard convolutional layer 232 to obtain a third convolutional feature map; after concatenating the second convolutional feature map and the third convolutional feature map in channels, inputting them into the fourth standard convolutional layer 233 to obtain a fourth convolutional feature map; inputting the fourth convolutional feature map into the attention sub-module 234 to obtain a feature attention map; adding the feature attention map to the first pooled feature map and then inputting it into the fifth standard convolutional layer 235 to obtain the first high-frequency feature map f3.

[0081] , in some embodiments, the convolutional layer of the second standard convolutional layer 231 in the first high-frequency feature extraction module 230 may adopt a 1×1 convolutional layer, and the activation layer may adopt a ReLU activation function; the convolutional layer of the third standard convolutional layer 232 may adopt a 3×3 convolutional layer, and the activation layer may adopt a ReLU activation function; the convolutional layer of the fourth standard convolutional layer 233 may adopt a 1×1 convolutional layer, and the activation layer may adopt a ReLU activation function; the convolutional layer of the fifth standard convolutional layer 235 may adopt a 3×3 convolutional layer, and the activation layer may adopt a ReLU activation function, and the stride of the convolutional kernel in this 3×3 convolutional layer is 2, so that the resolution of the feature map output by the first high-frequency feature extraction module 230 is reduced to half of the resolution of the feature map input to the first high-frequency feature extraction module 230.

[0082] , in some embodiments, the structure of the second high-frequency feature extraction module 240 is the same as that of the first high-frequency feature extraction module 230, except that the number of channels of the input / output feature maps of the second high-frequency feature extraction module 240 and the first high-frequency feature extraction module 230 is different, and the input feature map of the second standard convolutional layer 231 in the second high-frequency feature extraction module 240 is the third separated feature map, and the input feature map of the third standard convolutional layer 232 is the fourth separated feature map. The fifth standard convolutional layer 235 in the second high-frequency feature extraction module 240 outputs the second high-frequency feature map f2; where the third separated feature map and the fourth separated feature map are obtained by performing a channel separation operation on the first high-frequency feature map f3, and the number of channels of the third separated feature map and the fourth separated feature map is equal.

[0083] In this embodiment, the feature maps input to the first high-frequency feature extraction module 230 and the second high-frequency feature extraction module 240 are first subjected to a channel separation operation, and then the two separated feature maps are respectively subjected to convolution processing with different convolutional kernel sizes. On the one hand, since the receptive fields obtained by performing convolution processing with different sizes of convolutional kernels are different, the defect detection model can adapt to the detection of defects of different sizes, which can enhance the generalization ability of the defect detection model; on the other hand, compared with performing convolution operations using the complete feature map (such as the first pooled feature map), separating the feature map into two parts (such as the first separated feature map and the second separated feature map) and inputting these two parts into the second standard convolutional layer 231 and the third standard convolutional layer 232 respectively for convolution operations can significantly reduce the computational amount; in addition, if the convolutional layer of the second standard convolutional layer 231 is set as a 1×1 convolutional layer and the convolutional layer of the third standard convolutional layer 232 is set as a 3×3 convolutional layer, the second standard convolutional layer 231 focuses on the extraction of information on the feature channels, and the third standard convolutional layer 232 focuses on the extraction of information on the feature space.

[0084] After respectively performing convolution operations on the two separated feature maps obtained by channel separation, the results of the convolution operations (such as the second convolutional feature map and the third convolutional feature map) are subjected to channel concatenation to achieve information combination, thereby effectively improving the information representation ability of the feature map obtained by channel concatenation. In addition, the attention sub-module 234 is beneficial to realizing noise reduction of noise information and enhancement of effective information. In the first high-frequency feature extraction module 230, the feature attention map output by the attention sub-module 234 is added to the input feature map of the first high-frequency feature extraction module 230 (i.e., the first pooled feature map). In the second high-frequency feature extraction module 240, the feature attention map output by the attention sub-module 234 is added to the input feature map of the second high-frequency feature extraction module 240 (i.e., the first high-frequency feature map f3), which can further enrich the features learned by the defect detection model, reduce the possibility of gradient disappearance, and prevent feature degradation.

[0085] In some embodiments, the attention sub-module 234 is a lightweight channel attention module with a relatively simple structure. Please refer to Figure 6 , in some embodiments, the attention sub-module 234 includes an average pooling layer 234a, a sixth standard convolutional layer 234b, and a Sigmoid function layer 234c; the above-mentioned input of the fourth convolutional feature map into the attention sub-module 234 to obtain a feature attention map includes: inputting the fourth convolutional feature map into the average pooling layer 234a for pooling and then inputting it into the sixth standard convolutional layer 234b to obtain a sixth convolutional feature map; inputting the sixth convolutional feature map into the Sigmoid function layer 234c to obtain the channel weight of the fourth convolutional feature map; multiplying the fourth convolutional feature map by the channel weight to obtain a feature attention map.

[0086] The output of the Sigmoid function layer 234c is a vector. The elements in the vector correspond one by one to the channels of the fourth convolutional feature map, and each element is the weight of its corresponding channel. Multiplying the fourth convolutional feature map by the channel weights means multiplying each channel of the fourth convolutional feature map by its corresponding weight.

[0087] In some embodiments, the convolutional layer of the sixth standard convolutional layer 234b in the attention sub-module 234 may adopt a 1×1 convolutional layer, and the activation layer may adopt a ReLU activation function.

[0088] The multi-scale feature fusion module 300 is mainly used to fuse the extracted composite feature map and high-frequency feature maps of multiple scales, that is, to fuse low-frequency features and high-frequency features of multiple scales. Please refer to Figure 7 , in some embodiments, the multi-scale feature fusion module 300 includes multiple seventh standard convolutional layers 301 with different dilation rates and an eighth standard convolutional layer 302. When the high-frequency feature maps output by the multi-scale high-frequency feature extraction module 200 include a first high-frequency feature map f3 and a second high-frequency feature map f2, step S230 includes: inputting the composite feature map f1 into multiple seventh standard convolutional layers 301 with different dilation rates respectively to obtain multiple seventh convolutional feature maps with different scale features; concatenating all the seventh convolutional feature maps along the channel dimension and inputting them into the eighth standard convolutional layer 302 to obtain a first intermediate fusion feature map g1, where the resolution of the first intermediate fusion feature map g1 is less than that of the second high-frequency feature map f2; upsampling the first intermediate fusion feature map g1 to the same resolution as the second high-frequency feature map f2 and concatenating them along the channel dimension to obtain a second intermediate fusion feature map g2; upsampling the second intermediate fusion feature map to the same resolution as the first high-frequency feature map f3 and concatenating them along the channel dimension to obtain a fusion feature map g3.

[0089] In this embodiment, the multi-scale feature fusion module 300 increases the receptive field by using dilated convolution. And since multiple seventh standard convolutional layers 301 in the multi-scale feature fusion module 300 have different dilation rates, dilated convolutions with different dilation rates can be realized, enabling the multi-scale feature fusion module 300 to perform feature extraction on the input feature map by using multiple dilated convolutions with different dilation rates, thereby effectively extracting features of different scales of the input feature map, which is beneficial to improving the segmentation accuracy of defects in high-resolution images in the subsequent process.

[0090] In some embodiments, the convolutional layer of the seventh standard convolutional layer 301 in the multi-scale feature fusion module 300 may adopt a 3×3 convolutional layer, and the activation layer may adopt a ReLU activation function. Please refer to Figure 7, in some embodiments, the multi-scale feature fusion module 300 may be provided with four seventh standard convolutional layers 301, and the dilation rates of the four seventh standard convolutional layers 301 may be 1, 6, 12, and 18 respectively. Of course, those skilled in the art can also flexibly adjust parameters such as the dilation rate and the number of the seventh standard convolutional layers 301 in the multi-scale feature fusion module 300 according to actual scenario requirements. The convolutional layer of the eighth standard convolutional layer 302 in the multi-scale feature fusion module 300 may adopt a 1×1 convolutional layer, and the activation layer may adopt a ReLU activation function.

[0091] The detection module 400 is mainly used to implement the judgment and positioning of defects. Please refer to Figure 8 , in some embodiments, the detection module 400 includes a ninth standard convolutional layer 401; the foregoing step S240 includes: inputting the fused feature map into the ninth standard convolutional layer 401 to obtain a ninth convolutional feature map, upsampling the ninth convolutional feature map to the same resolution as the high-resolution image to obtain a defect prediction map, where the elements of the defect prediction map represent the probabilities that the pixel points at the corresponding positions belong to each classification category, and the classification categories include background pixels and each defect category; obtaining a defect segmentation map of the high-resolution image according to the defect prediction map, where the defect segmentation map is used to display the defect regions and the corresponding defect categories in the high-resolution image.

[0092] In some embodiments, the element at each pixel point position in the defect segmentation map can be regarded as a vector, and the elements in the vector respectively represent the probabilities that the pixel point belongs to each classification category. In order to make the values of the elements in the above vector conform to the characteristics of probabilities, the values of the elements in the above vector can be processed by softmax so that they are in the interval [0,1].

[0093] A defect segmentation map of the high-resolution image can be obtained according to the defect prediction map. Specifically, for each pixel point in the defect prediction map, the classification category with the highest probability is selected as its category. If the probability of belonging to the background pixel is the highest, the pixel point is a background pixel. If the probability of belonging to a certain defect category is the highest, the pixel point is a defect pixel, and its defect category is the defect category with the highest probability. Then, different colors are used to display the background pixels and the defect pixels of different categories to obtain the defect segmentation map. For example, the background of the defect segmentation map can be represented by black, and the defect pixels can be displayed in colors such as white, yellow, and red. Those skilled in the art can use different colors to display the above defect pixels according to different defect categories. Figure 9 For the defect detection result of an embodiment, please refer to Figure 9 , Figure 9 Figure a in Figure 9Figure b in [reference] shows a defect segmentation map obtained by processing with a defect detection model, where the black part represents the background and the white part represents the defect. Based on the defect segmentation map, the position and corresponding defect type of the defect on the object to be measured can be obtained.

[0094] In some embodiments, through reasonable design, the multi-scale high-frequency feature extraction module 200 can be a lightweight module compared to the composite feature extraction module 100, that is, the weight parameters of the multi-scale high-frequency feature extraction module 200 are less than those of the composite feature extraction module 100; step S210 may include: performing downsampling on the high-resolution image, and inputting the downsampled high-resolution image into the composite feature extraction module 100 to obtain a composite feature map of the high-resolution image.

[0095] In this embodiment, in order to enable the network model to process the high-resolution image as a whole under the current hardware resources, first, the resolution of the original image is reduced through a downsampling operation. The downsampled high-resolution image is input into the composite feature extraction module 100 to obtain complete semantic information. However, the reduction of the image resolution will cause a lot of detailed information to be lost. Therefore, this application designs a multi-scale high-frequency feature extraction module 200 to obtain high-frequency feature maps of multiple scales of the high-resolution image to make up for the lost detailed information. Moreover, the multi-scale high-frequency feature extraction module 200 adopts a lightweight design, which is also beneficial to adapting to the limitations of hardware resources. The image input into the multi-scale high-frequency feature extraction module 200 for processing can be an image with a relatively large resolution, but at the same time, the multi-scale high-frequency feature extraction module 200 also needs to meet the requirement of being able to effectively extract features.

[0096] The downsampled high-resolution image is relatively a small-scale image, and the composite feature extraction module 100 is a large network compared to the multi-scale high-frequency feature extraction module 200, which reflects the design idea of "small image, large network"; the image input into the multi-scale high-frequency feature extraction module 200 for processing is relatively large in scale (such as the above-mentioned high-frequency component image), and the multi-scale high-frequency feature extraction module 200 is a small network compared to the composite feature extraction module 100, which reflects the design idea of "large image, small network". In this way, it is beneficial to break through the memory limit, solve the problem of insufficient computing memory for high-resolution images, realize the whole-image training and inference of high-resolution images, and meet the detection requirements in industry.

[0097] The training process of the defect detection model will be described below. Please refer to Figure 10 , in some embodiments, the detection module 400 of the defect detection model includes three ninth standard convolutional layers 401, Figure 10 The dotted line in [reference] indicates that it is only executed in the training stage of the defect detection model, and the solid line indicates that it is executed in both the training and inference stages of the defect detection model; on this basis, please refer toFigure 11 , the training process of the defect detection model includes steps S310 to S350, which are specifically described below.

[0098] Step S310: Obtain training sample images and corresponding annotation data.

[0099] The training sample images can include normal sample images and defect sample images. The normal sample images are high-resolution images of defect-free normal products corresponding to the object to be measured, while the defect sample images are high-resolution images of defective products corresponding to the object to be measured. The normal sample images can be obtained by a high-resolution camera shooting the defect-free normal products corresponding to the object to be measured, and the defect sample images can be obtained by a high-resolution camera shooting the defective products corresponding to the object to be measured.

[0100] Since the acquisition of training sample images and corresponding annotation data is common knowledge in this technical field, the description of how to obtain training sample images and corresponding annotation data will not be elaborated here.

[0101] Step S320: Input the training sample images into the defect detection model to obtain the first intermediate fusion feature map g1, the second intermediate fusion feature map g2, and the fusion feature map g3 of the training sample images.

[0102] The process of obtaining the first intermediate fusion feature map g1, the second intermediate fusion feature map g2, and the fusion feature map g3 can refer to the relevant embodiments above and will not be elaborated here.

[0103] Step S330: Input the first intermediate fusion feature map, the second intermediate fusion feature map, and the fusion feature map into different ninth standard convolutional layers 401 respectively to obtain the ninth convolutional feature maps of the first intermediate fusion feature map g1, the second intermediate fusion feature map g2, and the fusion feature map g3 respectively.

[0104] Step S340: Upsample the ninth convolutional feature maps of the first intermediate fusion feature map g1, the second intermediate fusion feature map g2, and the fusion feature map g3 to the same resolution as the training sample images to obtain the first defect prediction map, the second defect prediction map, and the third defect prediction map.

[0105] Among them, the elements of the defect prediction map represent the probabilities of the pixel points at the corresponding positions belonging to each classification category, where the classification categories include background pixels and each defect category. For specific reference, please refer to the relevant description above. The defect regions can be segmented from the defect prediction map, that is, for each pixel point of the defect prediction map, it can be determined whether the pixel point belongs to the defect pixels according to the probabilities of the pixel point belonging to each classification category, and the region composed of the defect pixels is the defect region.

[0106] Step S350: Determine the total loss function L according to the first loss function L1 determined by the first defect prediction map and the annotation data, the second loss function L2 determined by the second defect prediction map and the annotation data, and the third loss function L3 determined by the third defect prediction map and the annotation data. total Train the defect detection model according to the total loss function. That is, the first defect prediction map, the second defect prediction map, and the third defect prediction map are respectively trained under the control of their respective loss functions with the corresponding annotation data.

[0107] The role of the loss function is to guide the update of the weight parameters of the defect detection model, so that the obtained first defect prediction map, second defect prediction map, and third defect prediction map are close to the corresponding annotation data. Those skilled in the art can design the above loss functions according to actual needs.

[0108] In some embodiments, the first loss function L1, the second loss function L2, and the third loss function L3 adopt the same form, and the specific expression is:

[0109] L = αL cross + βL dice ,

[0110] where L represents any one of the first loss function L1, the second loss function L2, and the third loss function L3, α and β are preset weight coefficients, L cross is the cross-entropy loss function, and L dice is the intersection over union loss function. Those skilled in the art can design appropriate cross-entropy loss functions and intersection over union loss functions according to specific scenarios. In some embodiments, the weight coefficients α and β can be 1.0 and 3.0 respectively. Since the intersection over union loss function can effectively deal with the situation of pixel category imbalance and is beneficial to the detection of small target defects, a larger weight value can be set.

[0111] In some embodiments, the annotation data includes the defect annotation regions of the first defect prediction map, the second defect prediction map, and the third defect prediction map, as well as the actual probability values of each pixel belonging to each classification category, that is, the category annotation values. Based on this, the present application provides a cross-entropy loss function and an intersection over union loss function, where the cross-entropy loss function is

[0112]

[0113] where i represents the pixel point serial number, n represents the total number of pixel points of the first defect prediction map or the second defect prediction map or the third defect prediction map, and LP i represents the cross-entropy loss value of the i-th pixel point, and the cross-entropy loss value LP of each pixel point is:

[0114]

[0115] where c is the defect category serial number, C is the number of defect categories, and p c is the probability that the pixel belongs to the c-th defect category, and g c is the corresponding category annotation value. In some embodiments, the category annotation value g c can be obtained through One-hot encoding processing, and the probability p c can be obtained through Softmax function processing.

[0116] The intersection over union loss function is

[0117]

[0118] where X is the defect region segmented from the first defect prediction map, the second defect prediction map, or the third defect prediction map, and Y is the corresponding defect annotation region.

[0119] For the total loss function L total , it can be obtained by weighted summation of the first loss function L1, the second loss function L2, and the third loss function L3. In some embodiments, the total loss function L total is

[0120] L total = L1 + L2 + L3.

[0121] Those skilled in the art can understand that, according to the defect detection method based on high-resolution images in the above embodiments, by constructing a dual-branch network (i.e., a composite feature extraction module and a multi-scale high-frequency feature extraction module), the composite feature map of the high-resolution image is extracted by using the composite feature extraction module, and multiple high-frequency feature maps with different scales of the high-resolution image are extracted by using the multi-scale high-frequency feature extraction module. Then, the composite feature map and the multi-scale high-frequency feature maps are fused for defect detection, which is beneficial to retaining as much detail information of the high-resolution image as possible, thereby greatly improving the accuracy of small target defect detection.

[0122] In some embodiments of the present application, in view of the large size of high-resolution images and the poor detection effect of small target defects, a decoupled processing technical solution is proposed. The specific idea is as follows: Since the size of the original image (i.e., the high-resolution image) is large, the present application reduces the resolution of the original image by using a downsampling operation, and complete semantic information can be obtained by extracting features from the downsampled high-resolution image; for the loss of detailed information caused by scaling the image and the difficulty of detecting small target defects, the present application designs a dual-branch network, uses a composite feature extraction module to extract features from the downsampled high-resolution image to obtain a composite feature map of the high-resolution image, and uses a multi-scale high-frequency feature extraction module to extract high-frequency features from the high-resolution image to obtain multiple high-frequency feature maps of different scales of the high-resolution image. The high-frequency feature maps can retain the position information and small target feature information in the original image, making up for the lost detailed information of the image. Finally, the composite feature map and the high-frequency feature map are fused to complete the defect detection of the high-resolution image; among them, the composite feature extraction module and the multi-scale high-frequency feature extraction module adopt the design idea of "small image with large network, large image with small network". By adopting the above technical solution, not only can the accuracy of small target defect detection be improved, the accurate positioning of object defects be realized, but also the hardware resource limitation can be broken through, the problem of insufficient computing memory of high-resolution images can be solved, the whole-image training and real-time inference of high-resolution images can be realized, and the detection requirements in industry can be met.

[0123] In summary, the beneficial effects of the defect detection method provided by the present application are mainly as follows:

[0124] (1) Through the design of the dual-branch network, the overall semantic information, rich position information and detailed information of the high-resolution image can be extracted, and these information are fused for defect detection, greatly improving the detection ability of small target defects in high-resolution images and realizing accurate detection;

[0125] (2) End-to-end processing can be realized. The defect detection model proposed by the present application can automatically extract and fuse the corresponding features of the image to complete the task of semantic segmentation without complex post-processing, thereby making the defect detection method simple to use and convenient to deploy;

[0126] (3) In some embodiments, through the decoupled processing technical solution, the hardware resource limitation existing in the existing defect detection methods is solved, and the original image can be directly processed. Under the condition of ensuring the detection accuracy, the whole-image training and inference of high-resolution images are realized;

[0127] (4) In some embodiments, the defect detection model proposed in this application adopts a lightweight design, using the ingenious design of "small images with large networks and large images with small networks", which is characterized by being concise, efficient and having a fast inference speed, thus being able to meet the real-time performance requirements of production enterprises for defect detection.

[0128] Those skilled in the art can understand that all or part of the functions of the above methods can be implemented in a hardware manner or in a computer program manner. When all or part of the functions in the above embodiments are implemented in a computer program manner, the program can be stored in a computer-readable storage medium, and the storage medium can include: read-only memory, random access memory, magnetic disk, optical disk, hard disk, etc. The above functions can be realized by a computer executing this program. For example, storing the program in the memory of the device, when the processor executes the program in the memory, the above all or part of the functions can be realized. In addition, when all or part of the functions in the above embodiments are implemented in a computer program manner, the program can also be stored in storage media such as a server, another computer, magnetic disk, optical disk, flash drive or mobile hard disk, and saved to the memory of the local device by downloading or copying, or the system of the local device is updated with a version. When the processor executes the program in the memory, all or part of the functions in the above embodiments can be realized.

[0129] The above uses specific examples to elaborate on this application, which is only used to help understand the technical solution of this application and is not intended to limit this application. For those skilled in the art in the technical field, based on the idea of this application, several simple deductions, deformations or substitutions can also be made.

Claims

1. A defect detection method based on high-resolution images, characterized in that, Including: Obtaining a high-resolution image of the object to be measured; Inputting the high-resolution image into a trained defect detection model to obtain a defect detection result of the high-resolution image, where the defect detection result includes a defect area and / or a defect category in the high-resolution image; Wherein the defect detection model includes a composite feature extraction module, a multi-scale high-frequency feature extraction module, a multi-scale feature fusion module, and a detection module; The step of inputting the high-resolution image into the trained defect detection model to obtain the defect detection result of the high-resolution image includes: Using the composite feature extraction module to extract features from the high-resolution image to obtain a composite feature map of the high-resolution image, where the composite feature map is used to represent the low-frequency and high-frequency features of the high-resolution image; wherein, the composite feature extraction module includes a convolutional layer, a max-pooling layer, and multiple cascaded residual modules, each of the residual modules includes two convolutional sub-layers, the input feature map of the max-pooling layer is the output feature map of the convolutional layer, the output of each residual module is added to the output of the previous-level residual module as the input of the next-level residual module, where the input feature map of the first residual module is the output feature map of the max-pooling layer, and the input feature map of the second residual module is the feature map obtained by adding the output of the first residual module and the output feature map of the max-pooling layer; Using the multi-scale high-frequency feature extraction module to extract high-frequency features from the high-resolution image to obtain multiple high-frequency feature maps of different scales of the high-resolution image, where the high-frequency feature maps are used to represent the high-frequency features of the high-resolution image; Using the multi-scale feature fusion module to perform feature fusion on the composite feature map and the high-frequency feature maps to obtain a fused feature map; Using the detection module to process the fused feature map to obtain the defect detection result of the high-resolution image.

2. The defect detection method according to claim 1, wherein, The step of using the multi-scale high-frequency feature extraction module to extract high-frequency features from the high-resolution image to obtain multiple high-frequency feature maps of different scales of the high-resolution image includes: Extracting the high-frequency component of the high-resolution image to obtain a high-frequency component image of the high-resolution image; Inputting the high-frequency component image into the multi-scale high-frequency feature extraction module for feature extraction to obtain the multiple high-frequency feature maps of different scales.

3. The defect detection method according to claim 2, wherein, The step of extracting the high-frequency component of the high-resolution image to obtain a high-frequency component image of the high-resolution image includes: Performing low-pass filtering on the high-resolution image to obtain a first filtered image, subtracting the first filtered image from the high-resolution image to obtain a first high-frequency residual image; Performing low-pass filtering on the first high-frequency residual image to obtain a second filtered image, subtracting the second filtered image from the first high-frequency residual image to obtain a second high-frequency residual image; Performing channel splicing on the first high-frequency residual image and the second high-frequency residual image to obtain the high-frequency component image, where the resolution of the high-frequency component image is the same as that of the high-resolution image.

4. The defect detection method according to claim 2, wherein The high-frequency feature maps include a first high-frequency feature map and a second high-frequency feature map. The multi-scale high-frequency feature extraction module includes a first standard convolutional layer, a max-pooling layer, a first high-frequency feature extraction module, and a second high-frequency feature extraction module. The standard convolutional layer includes a convolutional layer, a batch normalization layer, and an activation layer connected in sequence. The step of inputting the high-frequency component image into the multi-scale high-frequency feature extraction module for feature extraction to obtain the multiple high-frequency feature maps with different scales includes: Inputting the high-frequency component image into the first standard convolutional layer, and inputting the output feature map of the first standard convolutional layer into the max-pooling layer to obtain a first pooled feature map; Inputting the first pooled feature map into the first high-frequency feature extraction module for feature extraction to obtain the first high-frequency feature map; Inputting the first high-frequency feature map into the second high-frequency feature extraction module for further feature extraction to obtain the second high-frequency feature map, where the resolution of the first high-frequency feature map is greater than that of the second high-frequency feature map.

5. The defect detection method according to claim 4, wherein, The first high-frequency feature extraction module includes a second standard convolutional layer, a third standard convolutional layer, a fourth standard convolutional layer, an attention sub-module, and a fifth standard convolutional layer. The convolutional kernel sizes of the second standard convolutional layer and the third standard convolutional layer are different. The step of inputting the first pooled feature map into the first high-frequency feature extraction module for feature extraction to obtain the first high-frequency feature map includes: Performing a channel separation operation on the first pooled feature map to obtain a first separated feature map and a second separated feature map, where the number of channels of the first separated feature map and the second separated feature map are equal; Inputting the first separated feature map into the second standard convolutional layer to obtain a second convolutional feature map, and inputting the second separated feature map into the third standard convolutional layer to obtain a third convolutional feature map; After concatenating the second convolutional feature map and the third convolutional feature map in channels, inputting them into the fourth standard convolutional layer to obtain a fourth convolutional feature map; Inputting the fourth convolutional feature map into the attention sub-module to obtain a feature attention map; Adding the feature attention map to the first pooled feature map and then inputting it into the fifth standard convolutional layer to obtain the first high-frequency feature map; The second high-frequency feature extraction module has the same structure as the first high-frequency feature extraction module. Among them, the input feature map of the second standard convolutional layer in the second high-frequency feature extraction module is a third separated feature map, and the input feature map of the third standard convolutional layer is a fourth separated feature map. The third separated feature map and the fourth separated feature map are obtained by performing a channel separation operation on the first high-frequency feature map, and the number of channels of the third separated feature map and the fourth separated feature map are equal. The output of the fifth standard convolutional layer in the second high-frequency feature extraction module is the second high-frequency feature map.

6. The defect detection method according to claim 5, characterized in that, The attention sub-module includes an average pooling layer, a sixth standard convolutional layer, and a Sigmoid function layer. The step of inputting the fourth convolutional feature map into the attention sub-module to obtain a feature attention map includes: Input the fourth convolutional feature map into the average pooling layer for pooling and then input it into the sixth standard convolutional layer to obtain a sixth convolutional feature map; Input the sixth convolutional feature map into the Sigmoid function layer to obtain the channel weights of the fourth convolutional feature map; Multiply the fourth convolutional feature map by the channel weights to obtain the feature attention map.

7. The defect detection method according to claim 4, characterized in that, The multi-scale feature fusion module includes multiple seventh standard convolutional layers with different dilation rates and an eighth standard convolutional layer; the method for using the multi-scale feature fusion module to perform feature fusion on the composite feature map and the high-frequency feature map to obtain a fused feature map includes: Input the composite feature map into the multiple seventh standard convolutional layers with different dilation rates respectively to obtain multiple seventh convolutional feature maps with different scale features; Perform channel concatenation on all the seventh convolutional feature maps and then input them into the eighth standard convolutional layer to obtain a first intermediate fused feature map, and the resolution of the first intermediate fused feature map is smaller than that of the second high-frequency feature map; Upsample the first intermediate fused feature map to the same resolution as the second high-frequency feature map and perform channel concatenation with the second high-frequency feature map to obtain a second intermediate fused feature map; Upsample the second intermediate fused feature map to the same resolution as the first high-frequency feature map and perform channel concatenation with the first high-frequency feature map to obtain the fused feature map.

8. The defect detection method according to claim 7, characterized in that, The detection module includes three ninth standard convolutional layers; the defect detection model is trained in the following manner: Obtain training sample images and corresponding annotation data; Input the training sample images into the defect detection model to obtain a first intermediate fused feature map, a second intermediate fused feature map, and a fused feature map of the training sample images; Input the first intermediate fused feature map, the second intermediate fused feature map, and the fused feature map into different ninth standard convolutional layers respectively to obtain ninth convolutional feature maps of the first intermediate fused feature map, the second intermediate fused feature map, and the fused feature map respectively; Upsample the ninth convolutional feature maps of the first intermediate fused feature map, the second intermediate fused feature map, and the fused feature map to the same resolution as the training sample images to obtain a first defect prediction map, a second defect prediction map, and a third defect prediction map; wherein, the elements of the defect prediction map represent the probabilities of the pixel points at the corresponding positions belonging to each classification category, and the classification categories include background pixels and each defect category; According to the first loss function determined by the first defect prediction graph and the labeled data L 1. The second loss function determined by the second defect prediction graph and the labeled data L 2, and the third loss function determined by the third defect prediction graph and the labeled data L 3. Determine the total loss function L total , and train the defect detection model according to the total loss function.

9. The defect detection method according to claim 8, wherein The expressions of the first loss function, the second loss function, and the third loss function are: L = αL cross + βL dice , Among them L represents any one of the first loss function, the second loss function, and the third loss function, α and β is a weight coefficient, L cross is a cross-entropy loss function, L dice is an intersection over union loss function, and , Among them i represents the pixel point serial number n represents the total number of pixel points of the first defect prediction map or the second defect prediction map or the third defect prediction map LP i represents the i cross-entropy loss value of the nth pixel point, and the cross-entropy loss value of each pixel point LP is as follows , where c is the defect category serial number, C is the number of defect categories, p c is the probability that the pixel belongs to the c th defect category, g c is the corresponding category annotation value; , wherein X is a defect area segmented according to the first defect prediction map or the second defect prediction map or the third defect prediction map Y is the corresponding defect annotation area 10. The defect detection method according to claim 9, characterized in that, The expression of the total loss function is: L total = L 1+ L 2+ L 3。 11. The defect detection method according to claim 1, wherein The detection module includes a ninth standard convolutional layer; the method for using the detection module to process the fused feature map to obtain the defect detection result of the high-resolution image includes: Input the fused feature map into the ninth standard convolutional layer to obtain a ninth convolutional feature map, and upsample the ninth convolutional feature map to the same resolution as the high-resolution image to obtain a defect prediction map. The elements of the defect prediction map represent the probabilities that the pixel points at the corresponding positions belong to each classification category, where the classification categories include background pixels and each defect category. Obtain a defect segmentation map of the high-resolution image according to the defect prediction map, where the defect segmentation map is used to display the defect regions and corresponding defect categories in the high-resolution image.

12. The defect detection method according to claim 1, characterized in that, The multi-scale high-frequency feature extraction module is a lightweight module compared to the composite feature extraction module; the use of the composite feature extraction module to perform feature extraction on the high-resolution image to obtain a composite feature map of the high-resolution image includes: Perform downsampling processing on the high-resolution image, and input the downsampled high-resolution image into the composite feature extraction module to obtain a composite feature map of the high-resolution image.

13. A computer-readable storage medium, characterized in that, A program is stored on the medium, and the program can be executed by a processor to implement the defect detection method according to any one of claims 1-12.

Citation Information

Patent Citations

  • Surface defect detection method and equipment based on feature fusion

    CN114550021A

  • Defect detection method and system based on battery surface image and related equipment

    CN115272330A