Generalized deep learning defect detection method based on conditional token

By introducing conditional tokens and the self-attention mechanism of Transformer network in deep learning defect detection, the problem of insufficient generalization ability in the existing technology is solved, and higher defect detection accuracy and robustness are achieved.

CN120182279AActive Publication Date: 2025-06-20SHENZHEN DEEPVISION INNOVATION TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510662026.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-06-20
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

The existing deep learning defect detection methods have weak generalization capabilities when diversified and uncertainty defects, especially when training data is insufficient or sample distribution changes, the model performance is significantly reduced.

Method used

Deep learning defect detection method based on conditional tokens is adopted, and image block feature extraction and conditional token generation are selected by selecting qualified sample maps as reference images. Combined with the self-attention mechanism of the Transformer network and the cross-image attention fusion module, the generalization ability of the model is improved.

Benefits of technology

It improves the accuracy and robustness of defect detection, can more effectively detect defects in different scenarios, and maintains high detection accuracy and robustness in a variety of industrial applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182279A_ABST
    Figure CN120182279A_ABST
Patent Text Reader

Abstract

The invention discloses a generalized deep learning defect detection method based on a conditional token, and relates to the technical field of computer vision and deep learning defect detection, and the method comprises the steps: selecting # imgabs0 qualified sample images as reference images for the same batch of samples, dividing a to-be-detected sample image into to-be-detected area small images, and carrying out the detection of the to-be-detected area small images; obtaining a reference small image at the same position on the reference image; the method comprises the following steps of: amplifying a reference small image, taking the amplified reference small image and a to-be-detected area small image as inputs of a Transform network, acquiring key feature information between the reference small image and the to-be-detected area small image through a multi-layer self-attention mechanism, carrying out feature fusion, and carrying out defect detection by utilizing a decision network to obtain a position mask and a category of a defect. Therefore, by adopting the generalized deep learning defect detection method based on the conditional token, various environmental noises and product process changes in the industrial detection field can be handled, and the accuracy and robustness of defect detection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and deep learning defect detection, and in particular to a generalizable deep learning defect detection method based on conditional tokens. Background Art

[0002] With the rapid development of the manufacturing industry towards intelligence and automation, the importance of product quality inspection in industrial production has been increasing day by day. Traditional defect detection methods mainly rely on manual visual inspection or rule-based algorithms. Such methods have disadvantages such as high labor intensity, low efficiency, and low accuracy, and are easily affected by human factors. To solve these problems, deep learning-based defect detection technologies have been widely applied in the industrial field, especially showing great potential in improving detection accuracy and reducing manual intervention.

[0003] Convolutional neural networks (CNNs) in deep learning have achieved certain results in defect detection, but there are certain limitations in capturing global image information and processing long-range dependencies in this convolutional neural network. Transformer was initially a model for natural language processing tasks. Relying on its powerful self-attention mechanism, it can effectively capture global context information. As the Transformer architecture continues to penetrate into the field of computer vision, some research has applied it to the field of image processing, enhancing the performance of defect detection, especially showing advantages under complex structures and large-scale data.

[0004] However, when dealing with diverse and uncertain defects, the generalization ability of existing deep learning defect detection methods is weak. Especially when the training data is insufficient or the sample distribution changes, the performance of the model often drops significantly. For this reason, a defect detection method based on conditional Transformer tokens is proposed. By introducing conditional information, the generalization ability of the model is improved, so as to achieve efficient detection of defects in different scenarios and maintain high detection accuracy and robustness in a variety of industrial applications. Summary of the Invention

[0005] The purpose of the present invention is to provide a generalizable deep learning defect detection method based on conditional tokens, which can solve the generalization problem of defect detection, learn deeper image features, and improve the defect detection performance of deep learning models.

[0006] To achieve the above purpose, the present invention provides a generalizable deep learning defect detection method based on conditional tokens, including the following steps: S1. For the same batch of samples, select Use a qualified sample image as a reference image, divide the sample image to be detected into small sub-images of the area to be detected. At the same time, obtain the reference small sub-images at the same positions on the reference image, and amplify the reference small sub-images to increase the diversity of the samples; S2. Use the amplified reference small sub-images and the small sub-images of the area to be detected as the input of the Transformer network. Use the shared encoder to extract features from each reference small sub-image, then extract the corresponding global semantic representation, aggregate it into conditional tokens. At the same time, extract the image patch features of the small sub-images of the area to be detected, integrate the conditional tokens and the image patch features of the small sub-images of the area to be detected as the combined input, and obtain the first feature through the self-attention mechanism; S3. Through the cross-image attention fusion module, extract the image patch features of the small sub-images of the area to be detected and the reference small sub-images respectively. Use the image patch features of the small sub-images of the area to be detected to generate query vectors, and use the image patch features of the reference small sub-images to generate key vectors and value vectors. Obtain the second feature through the cross-attention mechanism, and perform weighted fusion on the first feature and the second feature to obtain the fused feature; S4. Based on the fused feature, construct the total loss function by combining the embedding distance loss and the cross-entropy loss, and use the decision network to perform defect detection to obtain the trained deep learning defect detection model; S5. For the samples to be detected in the same batch, use the trained deep learning defect detection model to extract the key features between the small sub-images of the area to be detected and the reference small sub-images, and then obtain the softmax score map through the decision network, and obtain the category mask corresponding to the maximum probability, so as to obtain the defect location and category information of the samples to be detected.

[0007] Preferably, in step S1, it includes selecting a reference image as the standard image, and performing calibration and registration processing on other reference images and the sample image to be detected with the standard image.

[0008] Preferably, in step S2, the amplification of the reference small sub-images includes processing the images by random translation and noise.

[0009] Preferably, in step S2, using the decision network to perform defect detection includes updating and optimizing the network parameters based on the total loss function through the stochastic gradient descent and backpropagation algorithms.

[0010] Preferably, the expression of the total loss function is as follows: ; Wherein, ; ; ; In the formula, is the embedding distance loss, is the cross-entropy loss, , is the corresponding weight coefficient, is the image patch feature of the th small image in the area to be detected, is the conditional token, represents the matching situation between the th image patch of the small image in the area to be detected and the conditional token, is the interval boundary, is the corresponding ground truth, is the input of the softmax layer, is the output of the softmax layer, represents the number of classes, represents the Softmax cross-entropy loss.

[0011] Therefore, the present invention adopts the above-mentioned generalizable deep learning defect detection method based on conditional tokens, and has the following technical effects: (1) Based on the comparison learning method of Transformer conditional tokens, the same area of qualified products is used as the reference image for comparative analysis to obtain deeper image features and improve the accuracy of defect detection.

[0012] (2) The pre-trained defect detection model of the present invention has stronger robustness, can cope with various environmental noises and product process changes faced in the industrial detection field, and has a wider application range.

[0013] Next, through the drawings and embodiments, the technical solutions of the present invention will be further described in detail. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is a schematic diagram of a small image to be detected and a reference small image in an embodiment of the generalizable deep learning defect detection method based on conditional tokens; Figure 2 is a schematic diagram of the structure of a deep learning defect detection model in an embodiment of the generalizable deep learning defect detection method based on conditional tokens. DETAILED DESCRIPTION OF THE INVENTION

[0015] The present invention can be more specifically explained through the following embodiments. The purpose of disclosing the present invention is to protect all changes and improvements within the scope of the present invention. The present invention is not limited to the following embodiments.

[0016] Deep learning defect detection methods usually use image data recognition and analysis to detect defects on the surface of objects, such as cracks, scratches, stains, etc. during the manufacturing process. They are widely used in the automated quality inspection and control of industrial manufacturing and production lines, helping to automatically detect product defects and improve production efficiency. Generally, deep learning models usually need a large amount of training data to automatically learn the patterns and features for detecting specific defects. However, the biggest problem faced by defect detection is the generalization problem. When a model trained on a batch of pictures is switched to test on a new batch of pictures, the performance will drop significantly. Based on this, the present invention provides a generalizable deep learning defect detection method based on conditional tokens, which uses a set of reference images to generate conditional tokens (Condition Token) representing the "normal sample distribution", guiding the model to focus on the difference regions when comparing the reference images and the images to be tested, thereby realizing unsupervised or weakly supervised contrast learning, including the following steps: S1. For samples in the same batch, select qualified sample images as reference images, that is, reference OK sample images. To ensure that the differences between the reference OK sample images are minimized overall, before defect detection, the reference OK sample images need to be corrected and registered. The specific operation is to use one reference OK sample image as a benchmark to correct and register all other reference OK samples. Similarly, perform correction and registration operations on all sample images to be tested.

[0017] Please refer to Figure 1 . Given a sample image to be inspected, cut it into small sub-images of the area to be inspected , for example or in size, and intercept the area at the same position on the reference OK sample image as the reference sub-image .

[0018] To reduce the influence of various environmental noises and product process variations on the samples, in this embodiment, the reference sub-images are randomly augmented, such as random translation and noise operations, to increase the diversity of the images.

[0019] S2. Please refer to Figure 2 . First, extract the features of each augmented reference sub-image through the shared encoder of the Transformer model , and extract the global semantic representation (such as the CLS token) from each feature . Then, aggregate the global semantic representations of the extracted reference sub-images into a conditional token , which is used to represent the average feature of the reference images and can be regarded as the semantic "prototype" of qualified products.

[0020] Then, the same shared encoder of the Transformer model is used to extract the image patch features of , and combine it with as the joint input, that is , and extract the first feature through the self-attention mechanism to achieve attention contrast learning under the condition.

[0021] S3. To better capture the differences between the reference image and the image to be tested, this embodiment also designs a cross-image attention fusion module to perform feature comparison between different images, and find key information such as similar features and different features between the two through a multi-layer attention mechanism to obtain the fusion feature. The specific operation is as follows: Both the image to be tested and the reference small image are input into the shared encoder to extract the corresponding image patch features. The image patch features of the image to be tested are used to generate the query vector , and the image patch features of the reference small image are used to generate the key vector and the value vector , and obtain the second feature through the cross-attention mechanism . Then, is weighted and fused with to obtain the fused feature , as follows: ; In the formula, can take a constant value, which is 0.5 in this embodiment. In another embodiment, can be obtained through learning by the attention mechanism.

[0022] S4. Based on the fused feature , by defining the embedding distance loss between positive and negative sample pairs, combined with the cross-entropy loss of defect detection, a total loss function is constructed, and then a decision network is used for defect detection to output the position mask and category of X defects.

[0023] The expression of the total loss function is: ; In the formula, is the embedding distance loss, is the cross-entropy loss, , are the corresponding weight coefficients.

[0024] Among them, the expression of the embedding distance loss is: ; In the formula, Indicates a figure Matches the conditional token (normal), Indicates non - matching (abnormal), Is the interval boundary.

[0025] The cross - entropy loss is used to measure the error between the output result and the defect mask annotation information as follows: ; ; In the formula, is the corresponding true value, is the input of the softmax layer, is the output of the softmax layer, represents the number of classes, represents the Softmax cross - entropy loss.

[0026] Based on the above total loss function, the network parameters are updated and optimized using the stochastic gradient descent and backpropagation algorithms to obtain the defect class mask closest to realize the pre - training process of the deep - learning defect detection model.

[0027] S5. For the same batch of samples to be tested, the trained deep - learning defect detection model is used to extract the key features between the image to be tested and the reference image, then through the decision network to obtain the softmax score map, and obtain the class mask corresponding to the maximum probability, so as to obtain the defect location and class information of the samples to be tested, which can greatly improve the accuracy and robustness of defect detection.

[0028] Therefore, the present invention adopts the above - mentioned generalizable deep - learning defect detection method based on conditional tokens, introduces the qualified sample image as a condition, and obtains deeper image features through contrast learning, which can cope with various environmental noises and product process changes faced in the industrial detection field and has stronger robustness.

[0029] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that: they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A deep learning defect detection method with generalizability based on conditional tokens, characterized in that, It includes the following steps: S1. For the same batch of samples, select qualified sample images as reference images, divide the sample images to be detected into small sub-images of the areas to be detected. At the same time, obtain the reference small sub-images at the same positions on the reference images, and amplify the reference small sub-images to increase the diversity of the samples; S2. Use the amplified reference small image and the small image of the area to be detected as the input of the Transformer network. Use the shared encoder to extract features from each reference small image, then extract the corresponding global semantic representation, aggregate it into conditional tokens, and at the same time extract the patch features of the small image of the area to be detected. Integrate the conditional tokens and the patch features of the small image of the area to be detected as the combined input, and obtain the first feature through the self-attention mechanism. S3. Through the cross-image attention fusion module, extract the patch features of the small image of the area to be detected and the reference small image respectively. Use the patch features of the small image of the area to be detected to generate query vectors, and use the patch features of the reference small image to generate key vectors and value vectors. Obtain the second feature through the cross-attention mechanism, and perform weighted fusion on the first feature and the second feature to obtain the fused feature. S4. Based on the fused feature, construct the total loss function by combining the embedding distance loss and the cross-entropy loss, and use the decision network for defect detection to obtain the trained deep learning defect detection model. S5. For the same batch of samples to be tested, use the trained deep learning defect detection model to extract the key features between the small image of the area to be detected and the reference small image, then obtain the softmax score map through the decision network, and obtain the category mask corresponding to the maximum probability, so as to obtain the defect location and category information of the samples to be tested.

2. The deep learning defect detection method with generalizability based on conditional tokens according to claim 1, characterized in that, In step S1, it includes selecting a reference image as the standard image, and performing calibration and registration processing on other reference images and the sample images to be tested with the standard image.

3. The deep learning defect detection method with generalizability based on conditional tokens according to claim 1, characterized in that, In step S2, the amplification of the reference small image includes processing the image by random translation and noise.

4. The deep learning defect detection method with generalizability based on conditional tokens according to claim 1, characterized in that, In step S2, using the decision network for defect detection includes updating and optimizing the network parameters based on the total loss function through the stochastic gradient descent and backpropagation algorithms.

5. The deep learning defect detection method with generalizability based on conditional tokens according to claim 1, characterized in that, The expression of the total loss function is as follows: ; Where, ; ; ; Wherein, is the embedding distance loss, is the cross-entropy loss, , are the corresponding weight coefficients, is the image patch feature of the th small image of the area to be detected, is the conditional token, represents the matching situation between the th image patch of the small image of the area to be detected and the conditional token, is the interval boundary, is 's corresponding ground truth, is the input of the softmax layer, is the output of the softmax layer, represents the number of classes, represents the Softmax cross-entropy loss.

Citation Information

Patent Citations

  • Industrial product defect quality inspection method and device based on deep learning and storage medium

    CN113643268A

  • Chip inductor surface defect detection method and system based on token fusion

    CN116309451A

  • Steel surface defect detection method based on MASK R-CNN

    CN117788359A

  • Deep learning-based jet printing color two-dimensional code defect detection method

    CN118115782A

  • Defect detection method based on deep contrast learning

    CN118967690A