A Few-Shot Defect Detection Method Combining Gradient Decoupling and Contrastive Learning

By combining gradient decoupling and contrast learning, the defect detection problem in small sample size and medium sample scenarios is solved, and efficient defect detection performance is achieved, suitable for short-term production models.

CN114943698BActive Publication Date: 2025-06-10NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210527181.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-16
Publication Date
2025-06-10
Estimated Expiration
2042-05-16

AI Technical Summary

Technical Problem

In industrial manufacturing scenarios, it is difficult for the prior art to achieve efficient defect detection in small sample scenarios, especially when product models are temporarily produced, it is difficult to collect defect samples, making it difficult to effectively apply deep learning models.

Method used

Combining the small sample defect detection method of gradient decoupling and contrast learning, the object detection model is pre-trained using a large-scale object detection data set, the gradient decoupling module and comparison branch are added, and the anchor points are generated using the adaptive anchor point algorithm, and the candidate box prediction network is adjusted to adapt to the number of defect categories.

Benefits of technology

This method significantly reduces the model training time, enhances the model's learning ability and target positioning ability, improves the ability to distinguish targets easily confused, and realizes efficient defect detection in small sample scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114943698B_ABST
    Figure CN114943698B_ABST
Patent Text Reader

Abstract

The present invention discloses a few-shot defect detection method combining gradient decoupling and contrastive learning. The method of the present invention mainly includes four stages: pre-training of the object detection dataset, generating anchors that conform to the characteristics of defect data, adjusting the object detection model and inputting the generated anchors for training, and detecting defects. In the pre-training stage, a large-scale object detection dataset is used to train the object detection model, which can greatly reduce the training time of the model and enable the model to have the ability of object localization. Among them, the gradient decoupling module is used to decouple the candidate box extraction network and the candidate box prediction network, which can obtain information more in line with the network characteristics during training and strengthen the model learning ability. The adjusted object detection model is trained with the defect dataset and anchors to generate a defect detection model. During training, the backbone network weights are frozen and a contrastive branch is added to the candidate box prediction network. The contrastive branch can make the feature differences of candidate boxes of different categories larger, strengthen the discrimination ability of the model, and have higher accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of machine vision, and particularly relates to a small-sample defect detection method combining gradient decoupling and contrast learning, which is used to detect defects on the surface of industrial products in a small number of manually labeled scenarios. Background Art

[0002] With the continuous development of current industrial production technology, the requirements for the quality of industrial products are also constantly increasing. In some highly automated industrial manufacturing scenarios, the requirement for the yield of products is particularly high. However, due to factors such as processing, design, and equipment failures, there are often some defects and damages on the surface of the produced products, resulting in an increase in production costs and waste of resources, and even causing harm to the personal safety of users in severe cases. Therefore, it is necessary to conduct defect detection in a timely manner after product production to confirm whether there are foreign objects, defects, and flaws on the surface of components or products.

[0003] To solve the problem of small-sample defects in industrial defect detection, there are usually two main methods, namely the engineering path and the algorithm path. Among them, there are two common methods in the engineering path. One is manual manufacturing defect detection based on real products, and the other is manual simulation defect detection based on real images. Manual manufacturing defect detection based on real products has a high cost, certain operation difficulties, and is irreversible if the product is damaged during detection. Manual simulation defect detection based on real images has a slow detection speed and high difficulty, and has high requirements for operators. To solve the small-sample problem from the algorithm path, there are two basic solutions. The first is to increase samples, and the second is to reduce the dependence of the algorithm on samples. Currently, the main defect detection methods are based on deep learning object detection algorithms.

[0004] The deep learning-based object detection algorithm builds a model based on a large number of defect samples. However, in reality, due to the incomplete defect samples, it is difficult to put the model into use. For example, in the scenario of multi-model small-batch production in the automotive industry, each model of product may only be produced for a short number of days. It is very likely that these models of products are no longer produced before the collection of defect samples is completed. Then it is very difficult to collect a large number of defect samples in such a scenario. In addition, since defect flaws are generated by non-controlled factors in the production process, and the forms of defects are diverse, it is also very difficult to completely collect samples of various forms, which limits the application of deep learning in the field of industrial detection. Therefore, it is crucial to design a defect detection method that can effectively obtain good detection performance in small-sample scenarios. Summary of the Invention

[0005] The content part of this application is used to introduce ideas in a brief form, and these ideas will be described in detail in the following detailed implementation part. The content part of this application is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0006] Aiming at the problems and deficiencies in the prior art, the purpose of the present invention is to provide a small-sample defect detection method combining gradient decoupling and contrast learning. From the perspective of transfer learning, using a large-scale dataset to pre-train the object detection model can significantly reduce the training time of the model, and at the same time enable the model to have a certain object localization ability. The gradient decoupling module can allow the candidate box prediction network and the candidate box extraction network to obtain information more in line with the network characteristics during training during forward propagation, and at the same time can optimize the candidate box prediction network and the candidate box extraction network to different degrees during backpropagation to enhance the learning ability of the model. Using the adaptive anchor point algorithm to generate anchors and input them into the candidate box extraction network can generate candidate boxes that more conform to the characteristics of the defect dataset. The added contrast branch calculates the contrast feature loss, making the difference between candidate box features of the same class smaller and the difference between candidate box features of different classes larger, enhancing the discrimination ability of the object detection model for easily confused objects. It is used to solve the problems raised in the above background technology.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] The present invention discloses a small-sample defect detection method combining gradient decoupling and contrast learning, mainly including the following steps:

[0009] Step 1, using a large-scale object detection dataset to obtain an object detection model through pre-training. The target defect detection model includes a backbone network, a feature pyramid network, a candidate box extraction network, a candidate box prediction network, and a gradient decoupling module;

[0010] Step 2, generating multiple anchors for the defect dataset through the adaptive anchor point algorithm;

[0011] Step 3, adjusting the candidate box prediction network in the object detection model, setting the number of outputs of the candidate box classification branch of the candidate box prediction network to the number of defect categories, and adding a contrast branch to the candidate box prediction network;

[0012] Step 4, inputting the defect dataset and the generated anchors into the adjusted object detection model for training to obtain a defect detection model;

[0013] Step 5, inputting the image to be detected into the defect detection model, and correspondingly outputting the position and category of the defect target.

[0014] Further, the large-scale object detection dataset described in step 1 is pre-trained to obtain an object detection model, which specifically includes the following steps:

[0015] Step 1.1, initialize the data loader used by the object detection model, and extract images and ground truth annotations;

[0016] Step 1.2, extract multi-layer feature maps of the image through the backbone network;

[0017] Step 1.3, use the Feature Pyramid Network to fuse the information of the multi-layer feature maps to obtain the feature maps of each layer of the pyramid;

[0018] Step 1.4, use the gradient decoupling module to perform two affine transformations on the pyramid feature maps to obtain two sets of affine feature maps;

[0019] Step 1.5, input the two sets of affine feature maps into the candidate box extraction network and the candidate box prediction network respectively. The candidate box extraction network extracts candidate boxes containing objects, and uses the candidate box prediction network to calculate the candidate box categories and candidate box regression parameters;

[0020] Step 1.6, calculate the classification loss, regression loss and gradient of the network according to the candidate box categories and candidate box regression parameters;

[0021] Step 1.7, use the gradient decoupling module to perform gradient decoupling operations on the candidate box extraction network and the candidate box prediction network;

[0022] Step 1.8, judge that after reaching the maximum number of training epochs, end and output the object detection model.

[0023] Further, in step 2, multiple anchor points are generated for the defect dataset through the adaptive anchor point algorithm, which specifically includes the following steps:

[0024] Step 2.1, randomly select k from the defect annotations of all samples in the defect dataset as the current clustering centers;

[0025] Step 2.2, judge whether there are defect annotations that have not been calculated. If so, calculate the clustering center closest to the current defect annotation, and add the current defect annotation to the list of the affiliated clustering center;

[0026] Step 2.3, judge whether all the defect annotations have been calculated. If so, calculate the mean value of the defect annotations in all the lists as the new clustering center;

[0027] Step 2.4, compare the current clustering center with the new clustering center. If they are the same, output the anchor points corresponding to the defective dataset of the current clustering center. If they are different, return to Step 2.2 and continue to execute.

[0028] Further, in Step 4, the defective dataset and the multiple anchor points are used to train a defect detection model, which specifically includes the following steps:

[0029] Step 4.1, initialize the data loader used by the defect detection model, and extract images and real annotations;

[0030] Step 4.2, extract multi-layer feature maps of the defective images through the backbone network;

[0031] Step 4.3, use the Feature Pyramid Network to fuse the information of the multi-layer defective feature maps to obtain defective pyramid feature maps;

[0032] Step 4.4, use the gradient decoupling module to perform two affine transformations on the defective pyramid feature maps to obtain two sets of defective affine feature maps;

[0033] Step 4.5, input the two sets of defective affine feature maps into the candidate box extraction network and the candidate box prediction network respectively. The candidate box extraction network extracts candidate boxes containing defects, and uses the candidate box prediction network to calculate the candidate box categories and candidate box regression parameters;

[0034] Step 4.6, calculate the classification loss, regression loss, contrast loss and gradient of the network according to the candidate box categories and candidate box regression parameters;

[0035] Step 4.7, use the gradient decoupling module to perform gradient decoupling operations on the candidate box extraction network and the candidate box prediction network;

[0036] Step 4.8, judge that after reaching the maximum number of training rounds, end and output the defect detection model.

[0037] Further, in Step 5, for the defective image to be detected, input it into the defective training model, and the specific detection steps are as follows:

[0038] Step 5.1, input the defective image to be detected, use the backbone network to extract multi-layer image features, and use the Feature Pyramid Network to fuse the multi-layer image features to obtain an image feature map;

[0039] Step 5.2, perform two affine transformations on the image feature map using the gradient decoupling module, and then output the affine image feature map, and input it into the candidate box extraction network and the candidate box prediction network respectively;

[0040] Step 5.3, use the target candidate box extraction network to extract the target candidate boxes containing defects, and use the candidate box prediction network to combine the target candidate boxes and the affine image feature map to calculate the category of each target candidate box and the target candidate box regression parameters;

[0041] Step 5.4, finally use the target candidate box regression parameters to adjust the position, width and height of the target candidate box, and then output the position of the target candidate box and the corresponding category.

[0042] Further, the input of the contrast branch added in the candidate box prediction network is the feature vector of the target candidate box. The feature vector of the target candidate box is transformed into a contrast feature vector through a fully connected layer, and then the ReLU activation function is used to process the contrast feature vector to obtain a contrast feature, and then the contrast loss of the model is calculated for the contrast feature.

[0043] Further, the calculation formulas of the classification loss and the regression loss are as follows:

[0044]

[0045]

[0046]

[0047] Among them, L cls represents the classification loss, L reg represents the regression loss, N represents the number of candidate boxes calculated in this round of training, p i represents the probability that the i-th candidate box is predicted as the target, t i represents the regression parameters of the i-th candidate box predicted, represents the regression parameters of the target candidate box matched by the i-th candidate box.

[0048] Further, the calculation formula of the contrast loss is as follows:

[0049]

[0050] Among them, N represents the number of candidate boxes calculated in the current training round, u i represents the IOU value between the current candidate box and the true annotation, f(·) is used to control the target candidate box to be calculated, N y represents the number of candidate boxes belonging to the category y, and z represents the contrast feature vector.

[0051] Further, in step 2.2, the 1-IOU value is used to calculate the distance between the current defect annotation and the nearest clustering center, and the specific calculation method is as follows:

[0052]

[0053] Among them, A represents the anchor point, i.e., the clustering center, G represents the defect annotation, and Area() represents calculating the area of the block diagram within the brackets.

[0054] Furthermore, the candidate box regression parameters shown are used to adjust the position and size of the target candidate box, and the calculation method for the adjustment is as follows:

[0055] x′ = t x ·w + x,

[0056] y′ = t y ·h + y,

[0057]

[0058]

[0059] Among them, x, y, w, and h represent the position of the upper left corner coordinate of the original candidate box and the width and height of the candidate box, x′, y′, w′, and h′ represent the position of the upper left corner coordinate of the adjusted candidate box and the adjusted width and height, and t x 、t y 、t w 、t h represent the candidate box regression parameters.

[0060] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention provides a small-sample defect detection method combining gradient decoupling and contrast learning. A target detection model is obtained through pre-training using a large-scale object detection dataset. Multiple anchor points are generated for the defect dataset through an adaptive anchor point algorithm, and the candidate box prediction network in the target detection model is adjusted. A contrast branch is added to the candidate box prediction network, and the defect dataset and the anchor points are input into the adjusted target detection model for training to obtain a defect detection model. For the defect image to be detected, the position and category of the defect target are correspondingly output when input into the defect detection model. The target detection model includes a backbone network, a feature pyramid network, a candidate box extraction network, a candidate box prediction network, and a gradient decoupling module. Using a large-scale dataset to pre-train the target detection model can significantly reduce the training time of the model and endow the model with a certain target localization ability. The gradient decoupling module can enable the candidate box prediction network and the candidate box extraction network to obtain information more in line with the network characteristics during training, strengthening the learning ability of the model. Using the defect dataset and the anchor points to train the adjusted target detection model to generate a defect detection model, freezing the backbone network weights during training and adding a contrast branch to the candidate box prediction network. The contrast branch can make the feature differences of candidate boxes of different categories larger, strengthening the model's ability to distinguish easily confused targets and having higher accuracy. Description of the Drawings

[0061] The accompanying drawings, which form a part of this application, are used to provide a further understanding of this application, making other features, objectives, and advantages of this application more apparent. The schematic embodiments and their descriptions of this application are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0062] Figure 1 : It is the main process structure diagram of the few-shot defect detection method that combines gradient decoupling and contrast learning implemented in the present invention;

[0063] Figure 2 : It is the main step flow diagram of the few-shot defect detection method that combines gradient decoupling and contrast learning implemented in the present invention;

[0064] Figure 3 : It is the step structure diagram of pre-training in the few-shot defect detection method that combines gradient decoupling and contrast learning implemented in the present invention;

[0065] Figure 4 : It is the structure diagram of the backbone network in the few-shot defect detection method that combines gradient decoupling and contrast learning implemented in the present invention;

[0066] Figure 5 : It is the structure diagram of the feature pyramid network in the few-shot defect detection method that combines gradient decoupling and contrast learning implemented in the present invention;

[0067] Figure 6 : It is the structure diagram of the candidate box extraction network in the few-shot defect detection method that combines gradient decoupling and contrast learning implemented in the present invention;

[0068] Figure 7 : It is the structure diagram of the candidate box prediction network in the few-shot defect detection method that combines gradient decoupling and contrast learning implemented in the present invention;

[0069] Figure 8 : It is the step structure diagram of defect detection model training in the few-shot defect detection method that combines gradient decoupling and contrast learning implemented in the present invention;

[0070] Figure 9 : It is the step flow diagram of the adaptive anchor point algorithm in the few-shot defect detection method that combines gradient decoupling and contrast learning implemented in the present invention;

[0071] Figure 10 : It is the structure diagram of the contrast branch in the few-shot defect detection method that combines gradient decoupling and contrast learning implemented in the present invention;

[0072] Figure 11 : It is the step flow diagram of the few-shot defect detection method that combines gradient decoupling and contrast learning implemented in the present invention for defect detection. Detailed implementation manners

[0073] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0074] In addition, it should be noted that for ease of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0075] The present invention discloses a few-shot defect detection method combining gradient decoupling and contrastive learning. The present disclosure will be described in detail below with reference to the accompanying drawings and in combination with embodiments.

[0076] Refer to Figures 1 to 2 as shown, and mainly includes the following steps:

[0077] Step 1, use a large-scale object detection dataset to obtain an object detection model through pre-training. The object detection model includes a backbone network, a feature pyramid network, a candidate box extraction network, a candidate box prediction network, and a gradient decoupling module;

[0078] Step 2, generate multiple anchors for the defect dataset through an adaptive anchor algorithm;

[0079] Step 3, adjust the candidate box prediction network in the object detection model, set the number of outputs of the candidate box classification branch of the candidate box prediction network to the number of defect categories, and add a contrast branch to the candidate box prediction network;

[0080] Step 4, input the defect dataset and the generated anchors into the adjusted object detection model for training to obtain a defect detection model;

[0081] Step 5, input the defect image to be detected into the defect detection model, and correspondingly output the position and category of the defect target.

[0082] Specifically, first, a pre-training process of the object detection model is carried out, that is, a large-scale object detection dataset is used to pre-train a model with good object detection and localization capabilities. Then, the candidate box prediction network is adjusted. The output quantity of the candidate box classification branch of the candidate box prediction network is set to the number of defect categories, and a contrast branch is added to the candidate box prediction network. The adaptive anchor algorithm is used to calculate all the anchors of the defect dataset, and then the calculated anchors are used to train the adjusted object detection model to obtain a defect detection model. Finally, the defect image to be detected is input into the defect detection model, and the defect detection model calculates and outputs the upper left coordinates, width and height of the defect target, as well as the position of the target candidate box and the corresponding category.

[0083] Furthermore, the candidate box prediction network includes a candidate box classification branch and a candidate box regression branch. The candidate box classification branch inputs the candidate box features into a fully connected layer network to calculate the category of the candidate box. When adjusting the output quantity of the candidate box classification branch, the weights of the newly generated candidate box classification branch satisfy a random distribution with a mean of 0 and a variance of 1, and then a contrast branch is added on the basis of this adjustment.

[0084] As Figure 3 shown, the specific process of pre-training a large-scale object detection dataset to obtain an object detection model in step 1 is as follows: Step 1.1, initialize the data loader used by the object detection model and extract images and ground truth annotations;

[0085] Step 1.2, extract multi-layer feature maps of the images in the large-scale dataset through the backbone network;

[0086] Step 1.3, use the feature pyramid network to fuse the information of the multi-layer feature maps to obtain the feature maps of each layer of the pyramid;

[0087] Step 1.4, use the gradient decoupling module to perform two affine transformations on the feature maps of each layer of the pyramid to obtain two groups of pyramid affine feature maps;

[0088] Step 1.5, input the two groups of affine feature maps into the candidate box extraction network and the candidate box prediction network respectively. The candidate box extraction network extracts the candidate boxes containing the objects, and the candidate box prediction network calculates the candidate box category and the candidate box regression parameters;

[0089] Step 1.6, calculate the classification loss, regression loss and gradient of the network according to the candidate box category and the candidate box regression parameters;

[0090] Step 1.7, use the gradient decoupling module to perform gradient decoupling operations on the candidate box extraction network and the candidate box prediction network;

[0091] Step 1.8, determine to end and output the object detection model after reaching the maximum number of training rounds.

[0092] Specifically, first prepare the Microsoft COCO dataset for model training. Using a large-scale dataset to pre-train the model can significantly reduce the model training time. Then initialize the data loader used by the object detection model. The data loader is used to select a batch of images and ground truth annotations as pre-training data each time. After preparing the data, it is necessary to determine whether the current round of training has reached the maximum number of training rounds before each round of training. If it has reached, output the model and end; if it has not reached the maximum number of training rounds, start the current round of training. During each round of training, first extract a batch of images and ground truth annotations from the pre-training data, use the backbone network to extract multi-level feature maps of the images, and use the Feature Pyramid Network (FPN) to fuse these multi-level feature maps to obtain more informative feature maps for each layer of the pyramid. Perform two affine transformations on the feature maps for each layer of the pyramid through the gradient decoupling module and input them into the region proposal network (RPN) and the region proposal regression network. Use the RPN to extract region proposals containing objects. The region proposal regression network uses the region proposals containing objects and the affine feature maps to calculate the class of each region proposal and the region proposal regression parameters. Calculate the classification loss, regression loss, and gradients of the network according to the region proposal classes and region proposal regression parameters. During the gradient descent process, multiply the gradients of the RPN and the region proposal regression network by different coefficients for gradient decoupling operations. Finally, return a judgment on whether the current round of training has reached the maximum number of training rounds to complete the current round of training.

[0093] The object detection model includes a backbone network, a Feature Pyramid Network (FPN), a region proposal network (RPN), a region proposal regression network, and a gradient decoupling module.

[0094] Such as Figure 4The structure of the backbone network is shown. The data in the second column of the figure respectively represent the size of the convolutional kernel, the number of convolutional kernels, the stride of the convolutional kernel, and the number of convolutional kernels in this convolutional block. Specifically, the backbone network consists of 5 convolutional blocks. Among them, convolutional block 1 contains a convolutional kernel with a size of 7x7 and a stride of 2. The stride in convolutional blocks 2 to 5 is 1. Convolutional block 2 contains 3 convolutional groups, and each convolutional group consists of a convolutional kernel with a size of 1x1 and 64 channels, a convolutional kernel with a size of 3x3 and 64 channels, and a convolutional kernel with a size of 1x1 and 256 channels. After each convolutional kernel, a batch normalization operation is performed. Convolutional block 3 contains 4 convolutional groups, and each convolutional group consists of a convolutional kernel with a size of 1x1 and 128 channels, a convolutional kernel with a size of 3x3 and 128 channels, and a convolutional kernel with a size of 1x1 and 512 channels. After each convolutional kernel, a batch normalization operation is performed. Convolutional block 4 contains 23 convolutional groups, and each convolutional group consists of a convolutional kernel with a size of 1x1 and 256 channels, a convolutional kernel with a size of 3x3 and 256 channels, and a convolutional kernel with a size of 1x1 and 1024 channels. After each convolutional kernel, a batch normalization operation is performed. Convolutional block 5 contains 3 convolutional groups, and each convolutional group consists of a convolutional kernel with a size of 1x1 and 512 channels, a convolutional kernel with a size of 3x3 and 512 channels, and a convolutional kernel with a size of 1x1 and 2048 channels. After each convolutional kernel, a batch normalization operation is performed. The output of each convolutional group is added to the input to obtain the multi-layer feature map of the final output of the convolutional group.

[0095] As Figure 5 shown is the structure of the Feature Pyramid Network, where the features Figures 1 to 4 correspond to the multi-layer feature maps output by convolutional blocks 2 to 5 of the backbone network. The higher the layer of the feature map, the fewer the pixel points and the more semantic information it contains. Among them, the feature Figure 1 is directly used as the first layer of the Feature Pyramid Network. The feature pyramid network of the 0th layer is obtained by performing max-pooling downsampling with a size of 3x3 and a stride of 2 on the pyramid feature map of the first layer of the Feature Pyramid Network. For the pyramid feature maps of the 2nd to 4th layers, first, the previous layer of the pyramid feature map is upsampled to obtain a pyramid feature map with doubled length and width, then a 1x1 convolutional operation is performed, and then it is added to the output feature map of the corresponding backbone network of this layer of the pyramid feature map to obtain the final feature map.

[0096] As Figure 6As shown, it is the candidate extraction network structure. The input of the network is the feature maps of each layer of the feature pyramid. First, a convolution operation is performed on the feature map with a convolution kernel size of 3x3 and a stride of 1. Then, two 1x1 convolution operations are performed respectively. After these two convolution operations, two feature maps are output. The first feature map represents the probability that each anchor point at each point on the feature map may contain a target, and the second feature map represents the regression parameters of each anchor point at each point on the feature map.

[0097] As Figure 7 shown, it is the candidate box prediction network structure. The input of the candidate box prediction network is the feature map corresponding to the candidate box output by the candidate box extraction network. First, a region of interest pooling operation with a size of 7x7 is performed to convert the feature map of each candidate box into a feature vector of a fixed size. Then, a fully connected layer is used to convert the feature vector of the candidate box into a 1024-dimensional feature vector. Finally, two fully connected layers are respectively passed through to calculate the probability that the candidate box corresponds to different category targets, and the regression parameters of the candidate box corresponding to different categories.

[0098] During the model training, the gradient decoupling module has the following functions: during the forward propagation of the model, the feature map is respectively input into the candidate box prediction network and the candidate box extraction network through two affine transformations. The parameters of the two affine transformations are continuously optimized through training, and an affine feature map that better conforms to the characteristics of the candidate box extraction network and the candidate box prediction network can be obtained. During the backpropagation of the model, the gradients input into the candidate box extraction network and the candidate box prediction network are multiplied by different coefficients, so that the two networks can optimize the model to different extents.

[0099] Furthermore, the candidate box regression parameters are used to adjust the position and size of the target candidate box, and the calculation method of the adjustment is as follows:

[0100] x′ = t x ·w + x,

[0101] y′ = t y ·h + y,

[0102]

[0103]

[0104] where x, y, w, and h represent the position of the upper left corner coordinate of the original candidate box and the width and height of the candidate box, and x′, y′, w′, and h′ represent the position of the upper left corner coordinate of the adjusted candidate box and the adjusted width and height, and t x , t y , t w , t h represent the candidate box regression parameters.

[0105] The calculation of the classification loss and the regression loss is as follows:

[0106]

[0107]

[0108]

[0109] Among them, L cls represents the classification loss, L reg represents the regression loss, N represents the number of candidate boxes calculated in this round of training, and p i represents the probability that the i-th candidate box is predicted to be a target. When the i-th candidate box is a positive sample, the value of is 1, otherwise the regression parameter of the i-th candidate box, represents the regression parameter of the true target candidate box matched by the i-th candidate box. The loss function used can control the magnitude of the gradient so that the training is not likely to go out of control.

[0110] Refer to Figure 9 as shown, the specific process steps of calculating the anchor points using the adaptive anchor point algorithm in step 2 are as follows:

[0111] Step 2.1, randomly select k from the defect annotations of all samples in the defect dataset as the current clustering centers;

[0112] Step 2.2, determine whether there are defect annotations that have not been calculated. If so, calculate the clustering center closest to the current defect annotation and add the current defect annotation to the list of the affiliated clustering center;

[0113] Step 2.3, determine whether all defect annotations have been calculated. If so, calculate the mean value of the defect annotations in all lists as the new clustering center;

[0114] Step 2.4, compare the current clustering center with the new clustering center. If they are the same, output the anchor points corresponding to the defect dataset of the current clustering center. If they are not the same, return to step 2.2 to continue execution.

[0115] Specifically, first, prepare the defect annotations for all samples in the defect dataset, and randomly select k defect annotations as the current clustering centers. Then, determine whether there are still uncalculated defect annotations. If so, calculate the clustering center closest to the current defect annotation, and add the current defect annotation to the list of the belonging clustering center. After calculating all the defect annotations, calculate the mean value of the defect annotations in each clustering center list, and use the mean value of each clustering center list as the new clustering center. If the new clustering center is not consistent with the current clustering center, repeat the above steps. Otherwise, output the anchor points corresponding to the defect dataset of the current clustering center.

[0116] In the above process, the 1-IOU value is used to calculate the distance between the current defect annotation and the closest clustering center. The specific calculation method of IOU is as follows:

[0117]

[0118] Among them, A represents the anchor point, that is, the clustering center, G represents the defect annotation, and Area() represents the area of the block diagram in the brackets.

[0119] As Figure 8 shown, the specific process of obtaining the defect detection model by training the defect dataset and the multiple anchor points in step 4 is as follows:

[0120] Step 4.1, initialize the data loader used by the defect detection model, and extract the images and ground truth annotations;

[0121] Step 4.2, extract multi-layer feature maps of the defect images through the backbone network;

[0122] Step 4.3, use the Feature Pyramid Network to fuse the information of the multi-layer defect feature maps to obtain the defect pyramid feature maps;

[0123] Step 4.4, use the gradient decoupling module to perform two affine transformations on the defect pyramid feature maps to obtain two sets of defect affine feature maps;

[0124] Step 4.5, input the two sets of defect affine feature maps into the candidate box extraction network and the candidate box prediction network respectively. The candidate box extraction network extracts the candidate boxes containing defects, and uses the candidate box prediction network to calculate the candidate box categories and candidate box regression parameters;

[0125] Step 4.6, calculate the classification loss, regression loss, contrast loss and gradient of the network according to the candidate box categories and candidate box regression parameters;

[0126] Step 4.7, use the gradient decoupling module to perform gradient decoupling operations on the candidate box extraction network and the candidate box prediction network;

[0127] Step 4.8, determine to end and output the defect detection model after reaching the maximum number of training rounds.

[0128] Refer to Figure 8 As shown, in Step 4, the defect dataset and the anchor points are input into the adjusted object detection model for training to obtain the defect detection model. The prepared defect dataset and the anchor points obtained in Step 2 are used for the training of the defect detection model. Its specific training process is similar to the pre-training process of the object detection model. The difference is that: the candidate box extraction network calculates the candidate boxes containing defects by combining the anchor points calculated in Step 2 and the affine feature map, and the candidate box prediction network calculates the candidate box feature vectors using the candidate boxes containing defects and the affine feature map. According to the calculated candidate box feature vectors, the category and regression parameters of each candidate box and the contrast feature of the candidate box are calculated respectively. Calculate the classification loss, regression loss, contrast loss and gradient of the network. During the process of gradient descent, the gradients of the candidate box extraction network and the candidate box prediction network are multiplied by different coefficients for gradient decoupling operation. Finally, it is judged whether the current round of training has reached the maximum number of training rounds to complete the current round of training.

[0129] Refer to Figure 10 As shown, the structure of the contrast branch is shown. The input of the contrast branch is the feature vector of the target candidate box calculated in each round of training. First, the 1024-dimensional target candidate box feature vector is transformed into a 128-dimensional target contrast feature vector through two fully connected layers. Then, the ReLU activation function is used to set the numbers less than 0 in the target contrast feature vector to 0 and output. Finally, the contrast loss of the defect detection model is calculated through all the target contrast feature vectors calculated in each round. The calculation method of the contrast loss is as follows:

[0130]

[0131] Among them, N represents the number of candidate boxes calculated in the current training round, u i represents the IOU value between the current candidate box and the true annotation, and f(·) is used to control the target candidate box to be calculated. In the present invention, it is set that when the value of u i is greater than 0.8, the value of f(·) is 1, otherwise it is 0. N y represents the number of candidate boxes belonging to the category y of the candidate box, z represents the contrast feature vector, and τ is set to 20.

[0132] Refer to Figure 11 As shown, for the defect image to be detected, it is input into the defect training model for defect detection. The specific steps are as follows:

[0133] Step 5.1, input the defect image to be detected, use the backbone network to extract multi-layer image features, and use the feature pyramid network to fuse the multi-layer image features to obtain the image feature map;

[0134] Step 5.2: After performing two affine transformations on the image feature map using the gradient decoupling module, the affine image feature map is output and input into the candidate box extraction network and the candidate box prediction network respectively;

[0135] Step 5.3: Use the candidate box extraction network to extract the target candidate boxes containing defects, and use the candidate box prediction network to combine the target candidate boxes and the affine image feature map to calculate the class of each target candidate box and the target candidate box regression parameters;

[0136] Step 5.4: Finally, use the target candidate box regression parameters to adjust the position, width, and height of the target candidate boxes, and then output the positions of the target candidate boxes and the corresponding classes.

[0137] Specifically, first, prepare the defect image to be detected and input it into the trained defect detection model. Each time of detection, use the backbone network to extract multiple layers of image features, use the feature pyramid network to fuse the multiple layers of image features to obtain the image features, and use the gradient decoupling module to perform two affine transformations on the image features and input them into the candidate box extraction network and the candidate box prediction network. Then use the candidate box extraction network to extract the candidate boxes containing the targets, and the candidate box prediction network combines and calculates the target candidate boxes and the image features output by the feature pyramid network to calculate the class of each target candidate box and the target candidate box regression parameters. Finally, use the target candidate box regression parameters to adjust the position, width, and height of the target candidate boxes, and then output the target candidate boxes and the corresponding classes.

[0138] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A few-shot defect detection method combining gradient decoupling and contrast learning, characterized in that, it mainly includes the following steps: Step 1, use a large-scale object detection dataset to obtain an object detection model through pre-training. The object detection model includes a backbone network, a feature pyramid network, a candidate box extraction network, a candidate box prediction network, and a gradient decoupling module; Step 2, generate multiple anchor points for the defect dataset through an adaptive anchor point algorithm; Step 3, adjust the candidate box prediction network in the object detection model, set the number of outputs of the candidate box classification branch of the candidate box prediction network to the number of defect categories, and add a contrast branch to the candidate box prediction network; Step 4, input the defect dataset and the generated anchor points into the adjusted object detection model for training to obtain a defect detection model; Step 5, input the defect image to be detected into the defect detection model, and correspondingly output the position and category of the defect target; Among them, the training steps of the defect detection model in Step 4 are as follows: Step 4.1, initialize the data loader used by the defect detection model, and extract images and ground truth annotations; Step 4.2, extract multi-layer defect feature maps of the defect image through the backbone network; Step 4.3, use the feature pyramid network to fuse the information of the multi-layer defect feature maps to obtain a defect pyramid feature map; Step 4.4, use the gradient decoupling module to perform two affine transformations on the defect pyramid feature map to obtain two groups of defect affine feature maps; Step 4.5, input the two groups of defect affine feature maps into the candidate box extraction network and the candidate box prediction network respectively. The candidate box extraction network extracts candidate boxes containing defects, and uses the candidate box prediction network to calculate the candidate box category and candidate box regression parameters; Step 4.6, calculate the classification loss, regression loss, contrast loss and gradient of the network according to the candidate box category and candidate box regression parameters; Step 4.7, use the gradient decoupling module to perform gradient decoupling operations on the candidate box extraction network and the candidate box prediction network; Step 4.8, judge that after reaching the maximum number of training rounds, end and output the defect detection model.

2. The few-shot defect detection method combining gradient decoupling and contrast learning according to claim 1, characterized in that, the step of obtaining the object detection model through pre-training using the large-scale object detection dataset in Step 1 specifically includes the following steps: Step 1.1, initialize the data loader used by the object detection model, and extract images and ground truth annotations; Step 1.2, extract multi-layer feature maps of the image through the backbone network; Step 1.3, use the feature pyramid network to fuse the information of the multi-layer feature maps to obtain a pyramid feature map; Step 1.4, use the gradient decoupling module to perform two affine transformations on the pyramid feature map to obtain two groups of affine feature maps; Step 1.5: Input the two groups of the affine feature maps into the candidate box extraction network and the candidate box prediction network respectively. The candidate box extraction network extracts the candidate boxes containing the target, and the candidate box prediction network calculates the candidate box categories and the candidate box regression parameters. Step 1.6: Calculate the classification loss, regression loss and gradients of the network according to the candidate box categories and the candidate box regression parameters. Step 1.7: Use the gradient decoupling module to perform gradient decoupling operations on the candidate box extraction network and the candidate box prediction network. Step 1.8: Judge whether the maximum number of training rounds is reached, and if so, end and output the object detection model.

3. A small sample defect detection method combining gradient decoupling and contrast learning according to claim 2, wherein, In step 2, generating multiple anchor points for the defect data set through the adaptive anchor point algorithm specifically includes the following steps: Step 2.1: Randomly select k from the defect annotations of all samples in the defect data set as the current clustering centers. Step 2.2: Judge whether there are defect annotations that have not been calculated. If so, calculate the clustering center closest to the current defect annotation, and add the current defect annotation to the list of the affiliated clustering center. Step 2.3: Judge whether all the defect annotations have been calculated. If so, calculate the mean value of the defect annotations in all the lists as the new clustering center. Step 2.4: Compare the current clustering center with the new clustering center. If they are the same, output the anchor points of the defect data set corresponding to the current clustering center. If they are different, return to step 2.2 and continue to execute.

4. A small sample defect detection method combining gradient decoupling and contrast learning according to claim 3, wherein, In step 5, for the defect image to be detected, input it into the defect training model, and the specific detection steps are as follows: Step 5.1: Input the image to be detected, use the backbone network to extract multi-layer image features, and use the feature pyramid network to fuse the multi-layer image features to obtain an image feature map. Step 5.2: Perform two affine transformations on the image feature map using the gradient decoupling module and then output an affine image feature map, and input it into the candidate box extraction network and the candidate box prediction network respectively. Step 5.3: Use the candidate box extraction network to extract the target candidate boxes containing defects, and use the candidate box prediction network to combine the target candidate boxes and the affine image feature map to calculate the category of each target candidate box and the target candidate box regression parameters. Step 5.4: Finally, use the target candidate box regression parameters to adjust the position, width and height of the target candidate box, and then output the position of the target candidate box and the corresponding category.

5. A small sample defect detection method combining gradient decoupling and contrast learning according to claim 4, wherein, The input of the contrast branch added to the candidate box prediction network is the feature vector of the target candidate box. The feature vector of the target candidate box is transformed into a contrast feature vector through a fully connected layer. Then, the ReLU activation function is used to process the contrast feature vector to obtain a contrast feature. Next, the contrast loss of the model is calculated for the contrast feature.

6. A small sample defect detection method combining gradient decoupling and contrast learning according to claim 5, wherein, the calculation formulas of the classification loss and the regression loss are as follows: Among them, L cls represents the classification loss, L reg represents the regression loss, N represents the number of candidate boxes calculated in this round of training, p i represents the probability that the i-th candidate box is predicted as the target, t i represents the regression parameter of the i-th candidate box predicted, represents the regression parameter of the target candidate box matched by the i-th candidate box.

7. A small sample defect detection method combining gradient decoupling and contrast learning according to claim 6, wherein, the calculation formula of the contrast loss is as follows: Among them, N represents the number of candidate boxes calculated in the current training round, and u i represents the IOU value between the current candidate box and the true annotation. f(·) is used to control the target candidate boxes to be calculated. N y represents the number belonging to the candidate box category y, and z represents the contrast feature vector.

8. A small sample defect detection method combining gradient decoupling and contrast learning according to claim 7, wherein, in step 2.2, the 1-IOU value is used to calculate the distance between the current defect annotation and the nearest clustering center. The specific calculation method is as follows: where A represents the anchor point, i.e., the clustering center, G represents the defect annotation, and Area() represents the area of the block diagram in the parentheses.

9. A small sample defect detection method combining gradient decoupling and contrast learning according to claim 8, wherein, the candidate box regression parameters are used to adjust the position and size of the target candidate box. The adjustment calculation method is as follows: x′ = t x ·w + x, y′ = t y ·h + y Among them, x, y, w, and h represent the position of the upper left corner coordinates of the original candidate box and the width and height of the candidate box, and x′, y′, w′, and h′ represent the position of the upper left corner coordinates of the adjusted candidate box and the adjusted width and height, and t x , t y , t w , t h represent the candidate box regression parameters.

Citation Information

Patent Citations

  • Image contour identification method and device, equipment and medium

    CN110222703A

  • Convolutional neural network cloth defect detection method based on extreme learning machine

    CN111260614A