Progressive industrial product surface defect detection method based on decoupling characterization
By combining convolutional neural networks and Transformer networks in surface defect detection in industrial products, a progressive detection method with decoupled representation is designed, which solves the problem that existing models are difficult to detect local and structural abnormalities at the same time, and achieves more accurate defect detection and positioning effects.
Patent Information
- Application Number
- CN202510218584.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-30
AI Technical Summary
The existing neural network models used for surface defect detection of industrial products are mainly convolutional neural networks and Transformer. It is difficult to effectively combine the two and fuse their respective advantages, resulting in insufficient performance when detecting local and structural abnormalities.
By designing a progressive defect detection method based on decoupled representation, combining convolutional neural networks and Transformer networks, using structural repair subnets and apparent reconstructed subnets, using pre-trained visual Transformer for feature extraction, and using zero-initialization residual tandem strategy and cross-attention calculation, shape-appearance progressive repair and defect positioning are achieved.
It achieves more accurate image repair and defect detection effects, can show good detection performance on small local defects and global structural defects, and improves the detection and positioning capabilities of surface defects of industrial products.
Smart Images

Figure CN120070401A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a progressive industrial product surface defect detection method based on decoupling characterization, and belongs to the technical field of computer vision. Background Art
[0002] Automatically detecting whether there are defects on the product surface and accurately locating the defective area is an important part of the intelligent manufacturing system, which plays a very important role in product quality control, technical equipment performance evaluation, production parameter optimization, etc. The defect detection method based on computer vision obtains the product surface image through optical sensors and locates and segments the defective area.
[0003] From the perspective of representation, the existing mainstream methods of surface defect detection can be divided into two categories: convolutional neural network-based methods and Transformer-based methods. The former extracts local features with strong representation capabilities through convolutional neural networks, focusing on capturing changes in texture and color, and can accurately locate small-scale local anomalies. However, due to the limited receptive field and inductive bias of the convolution kernel, the convolutional neural network lacks an overall understanding of the image, making it unsuitable for processing structural changes such as shape and spatial relationships. Transformer-based methods can capture long-distance dependencies, so they can better extract global information and contextual relationships. However, these models often have more parameters and require a lot of computing resources. Due to the lack of abnormal data in reality, this will weaken the performance of Transformers, causing them to learn suboptimal long-distance dependencies and inaccurate local details.
[0004] The convolutional neural network-based method is not suitable for processing structural anomalies and large-scale abnormal areas, while the Transformer lacks accuracy in abnormal locations and often mistakenly identifies normal areas as abnormal. How to make the model accurately locate local anomalies and detect structural and logical anomalies still needs further development. Summary of the invention
[0005] Technical issues:
[0006] The existing neural networks used for defect detection tasks are mainly convolutional neural networks and Transformers, but there are few methods that can effectively combine these two models and integrate the advantages of each model. The present invention organically integrates these two common models to perform shape-appearance progressive repair and defect location to achieve more accurate image repair and defect detection effects.
[0007] Technical solution:
[0008] A progressive industrial product surface defect detection method based on decoupled characterization comprises the following steps:
[0009] Step 1. Take an image of the surface of a normal industrial product using an industrial camera to obtain a normal image.
[0010] Step 2. Since the number of real defect samples is small, generate simulated defect images by random mask template sampling and linear weighting as the input during model training. Specifically, randomly generate a mask template according to Perlin noise, sample the abnormal area from other natural scene image datasets according to the mask template, and fuse it with the original normal image by linear weighting to generate a defect image with local abnormalities. The linear weighting parameter can control the intensity of the defect. The mask template serves as the ground truth of the defect location.
[0011] Step 3. Use the pre-trained Segment Anything Model (SAM) to obtain the two-dimensional shape of the product in the normal image as the supervision information for the structure repair sub-network. Specifically, given a fixed anchor box for each category as a rough hint of the target location, input the normal image and the hint information into the SAM model, and the model will output the target shape image.
[0012] Step 4. Use the two-dimensional shape image obtained in Step 3 as the supervision information to train the structure repair sub-network to predict the target shape information of the defect image. Specifically, the structure repair sub-network is based on the Transformer architecture. The input is the artificially simulated defect image and several randomly sampled normal images, and the output is the repaired shape information. The shape information of the corresponding normal image obtained by SAM is used as the supervision information to train this sub-network. Specifically, use a pre-trained feature extractor to extract the features of the image. During model training, the parameters of the feature extractor are frozen. The feature extractor uses the Vision Transformer (ViT), which has been pre-trained on ImageNet-1K through an image reconstruction task. To make full use of the structural information of the normal image, after the normal image is extracted by this pre-trained ViT, it is stored in the memory bank, and 30% of the features are sampled by the kernel set sampling method to reduce the storage cost. After the input image passes through ViT to obtain the feature F, the 4 most similar normal features F′ are queried in the memory bank through cosine similarity. Taking the input feature F as the query, after the input feature and the normal feature are concatenated and mapped through a linear layer as the key and value, cross-attention calculation between abnormal and normal is performed to obtain the fused feature. The fused feature is input into 8 stacked self-attention layers for further feature aggregation, and smoothed through a two-layer convolutional neural network to output the repaired target shape image. The output is compared with the two-dimensional shape image of the corresponding normal image to do l 2Calculate the loss and the gradient similarity loss, and train the structure repair sub-network. Through the above method, the structure repair module can decouple and repair the target shape of the input image, providing guidance for subsequent appearance information reconstruction.
[0013] Step 5. Adaptively fuse the repaired target shape information as an increment into the appearance reconstruction sub-network, and precisely reconstruct the image through global structure guidance to obtain the reconstructed image. Specifically, the appearance reconstruction sub-network based on the convolutional neural network aims to reconstruct the input image into a normal image and adopts an encoder-decoder structure. The repaired shape information is used as incremental information, multiplied by a learnable weight initialized to zero, and then concatenated to each layer of the encoder of the appearance reconstruction sub-network to guide the reconstruction of the image. Calculate the l 2 loss and the structure consistency loss between the output of the reconstruction sub-network and the real normal image.
[0014] Step 6. Input the simulated defective image, the reconstructed image, and the repaired target shape information into the localization sub-network to detect defects using the appearance reconstruction error and the structure information. Specifically, the localization sub-network aims to detect the defective regions in the image. Perform the following processing on the input image I, the reconstructed image I r and the repaired two-dimensional shape image S r :
[0015]
[0016] I d = I c + α(I C · S r )
[0017] where α is a learnable weight initialized to zero. denotes concatenation of the input image I and the reconstructed image I r along the channel dimension, and "·" is the dot product operation.
[0018] Step 7. Calculate the loss between the output of the appearance reconstruction sub-network and the corresponding normal image, and calculate the loss between the output of the localization sub-network and the mask template in Step 2, and jointly train and optimize the appearance reconstruction sub-network and the localization sub-network. Specifically, use I d as the final input of the localization sub-network. The output of the localization sub-network is the defect prediction score map. After setting a threshold for binarization, calculate the focal loss with the mask template. Train the structure repair sub-network separately, and jointly train the appearance reconstruction sub-network and the localization sub-network. Finally, use the trained model to detect and locate the surface defects of industrial products.
[0019] Beneficial effects:
[0020] The present invention designs a progressive defect detection method based on decoupled representation, which has the following beneficial effects:
[0021] 1. This method organically integrates the respective advantages of convolutional neural network and Transformer network, enabling the model to exhibit good detection performance for both tiny local defects and global structural defects.
[0022] 2. This method proposes a structure repair sub-network based on Transformer. By decoupling the shape information from the image and performing independent repair, it avoids the interference of texture, color, etc. on the structural information.
[0023] 3. This method proposes a zero-initialization residual concatenation strategy, thereby adaptively integrating the repaired shape information into the multi-scale appearance reconstruction module, enhancing the accuracy of image reconstruction.
[0024] 4. This method proposes a sub-network for locating defects guided by the target shape information, improving the detection and localization capabilities for surface defects of various industrial products.
[0025] In summary, the progressive defect detection method based on decoupled representation improves the defect detection algorithm based on reconstruction, obtains more stable and accurate image reconstruction results under the guidance of shape information, and effectively improves the defect detection performance. Description of the Drawings
[0026] Figure 1 is the overall flowchart of the model of the present invention;
[0027] Figure 2 is the flowchart of the structure repair sub-network in the present invention;
[0028] Figure 3 is the flowchart of the anomaly-normal attention layer in the present invention. Detailed Embodiments
[0029] The following further clarifies the present invention in conjunction with the drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.
[0030] As shown in the figure, the progressive defect detection method based on decoupled representation proposed by the present invention includes the following steps:
[0031] Step 1: Take an image of the surface of a normal industrial product through an industrial camera to obtain a normal image.
[0032] Step 2: Since defect samples are usually scarce and difficult to obtain, the present invention adopts a synthesis method based on random masks to generate defect images.
[0033] 1) Randomly generate masked regions on the normal image, and these regions are marked as the parts to be replaced.
[0034] 2) Replace the masked regions with a linear combination of the defect-free image and any image from the external data source. Assume I M represents the masked region and is a binary image. I n represents the normal image, and I E represents any image from the external data source. Then the synthesized defective region can be expressed as:
[0035] I n =I O +γI M I E
[0036] where γ is the mixing weight used to adjust the ratio of the normal image to the external image.
[0037] 3) Combine the replaced masked regions with the rest of the original image to generate the final simulated defective image I.
[0038] The defective images generated in this way introduce defective features in the local regions and contain the global context information of the normal image. The strength of the generated defects can be controlled by adjusting the magnitude of the weighting coefficient. Training with the defective images generated by this method can effectively enhance the model's defect detection ability.
[0039] Step Three: Automatically generate the shape of the target as supervision information using the Segment Anything Model (SAM). For each type of product, set a rough bounding box as the prompt information for SAM and output an accurate binary shape image.
[0040] Step Four: To capture the global and local features of the input image, the present invention uses a pre-trained Vision Transformer for feature extraction.
[0041] 1) The input image I is divided into small patches of size p×p. Each small patch is flattened and mapped to a high-dimensional space to form an initial feature embedding. The embedded features are input into the Vision Transformer, and the global context information is captured through the self-attention mechanism to obtain multi-layer feature representations where N is the number of small patches, C is the dimension of each embedding, and L is the number of layers.
[0042] 2) For each normal image, extract its multi-level features and store them in the memory bank. Since the memory bank may occupy a large storage space due to the excessive number of samples, the kernel set sampling method is adopted to sample 30% of the features according to importance to reduce the scale of the memory bank and obtain the most representative features of the normal images.
[0043] Step 5: The defect may stem from changes in multiple factors such as structure, color, or texture. The present invention independently extracts and repairs the target shape information of the input image through a structure repair sub-network.
[0044] 1) Calculate the defect-normal cross-attention score: For the feature φ(I) of the input image, retrieve k nearest neighbor normal features from the memory bank. Take the feature of the input image as the query (Query), concatenate the feature of the input image and the k nearest neighbor normal features and then perform a linear transformation to obtain the key (Key) and value (Value), and then calculate the cross-attention score.
[0045] Q = Φ(I)
[0046]
[0047] where MLP is a multi-layer linear perceptron.
[0048] 2) Abnormal feature repair: Multiply the obtained cross-attention score by the value (Value) to obtain the feature repaired by fusing normal and abnormal information.
[0049] Attention(Q, K, V) = Softmax(Q(K) T )V
[0050] 3) Stack multiple self-attention modules, take the repaired feature as the input, and further improve the global perception ability of the feature. Two layers of 3×3 convolutional operations further aggregate information, map the fused feature to the structure space, and obtain a binary shape image with the same size as the input image.
[0051] 4) Use SAM to automatically generate the shape of the target as supervision information, and use the gradient magnitude similarity (GMS) and l 2 loss to jointly optimize the structure repair sub-network:
[0052]
[0053] where S(I n ) is the normal image target shape obtained by the SAM model, and S r is the shape image predicted by the model. N p is the total number of pixels in the image, and respectively represent the gradients of the two images I n and S r at point i, and λ is a balanced hyperparameter, which is set to 2 here.
[0054] Step 6: Use the repaired target shape information to guide the detailed reconstruction of the image.
[0055] 1) Adopt an encoder-decoder structure, whose input is the defective image I containing a random mask a , and the goal is to reconstruct the complete original image I.
[0056] 2) Take the repaired two-dimensional shape image as incremental information, multiply it by the learnable weight initialized to 0, and then splice it to each layer of the encoder of the reconstruction network:
[0057]
[0058] where X k represents the feature of the k-th layer of the encoder, β k is the learnable weight, S k represents the structural information of the repaired shape image downsampled to the same length and width as X k , represents channel-wise splicing. X ′ k replaces X k input to the k+1-th layer of the encoder. Through the above fusion, the structural information can adaptively guide the reconstruction of the image, especially in the restoration of large-scale structures and long-distance texture patterns.
[0059] 3) The optimization goal of the reconstruction module combines pixel-level loss and structural similarity loss:
[0060]
[0061] where I r is the reconstructed image, and N p is the total number of pixels in the image.
[0062] Step 7: Defect localization and detection.
[0063] 1) The input of the localization sub-network is the channel splicing result I r of the input image I and the reconstructed image I C . Use the shape information to guide the network to focus on the target area:
[0064] I d = I C + α(I C · S r )
[0065] where α is the learnable weight initialized to zero. S r is the shape information predicted by the structural repair sub-network, and "·" is the dot product operation. The two-dimensional shape image S rAfter its size is changed from 224×224×1 to 256×256×1 through bilinear interpolation, it is then dot-producted with I C for dot product operation.
[0066] 2) The localization module is optimized by the Focal Loss function. The reconstruction sub-network and the localization sub-network are jointly trained, and the total loss function is:
[0067]
[0068] where, represents the result predicted by the localization model, and M gt is the image of the defect area manually annotated.
[0069] This method is trained using PyTorch, and relevant parameters are set with reference to the engineering parameter setting experience. The batch size is set to 8, which means 8 images are loaded each time, with 4 normal images and 4 defective images. We use the Vision Transformer pre-trained on ImageNet-1k as the feature extractor. The parameters of the feature extractor are frozen during the training process. The structure repair module is trained separately. Then, under the guidance of the structure information, the appearance repair module and the defect localization module are jointly trained, and the epoch is set to 700. After the defect prediction map is smoothed, the maximum value is taken at the pixel level to obtain the image-level defect score.
[0070] Using the trained model, the surface defects of industrial products are detected and located. The input image passes through the structure repair, appearance reconstruction, and defect localization modules to generate a defect score map. By setting a threshold, the defect score Figure 2 is binarized to obtain the final defect detection result.
[0071] The above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any modification made on the basis of this technical solution belongs to the technical idea proposed by the present invention and falls within the protection scope of the claims of the present invention.
Claims
1. A progressive industrial product surface defect detection method based on decoupled characterization, characterized in that: The following steps are involved: Step 1. Use an industrial camera to capture a normal surface image of an industrial product to obtain a normal image; Step 2. Generate a simulated defect image by random mask template sampling and linear weighting as input for model training; Step 3. Use the segmentation model to obtain the two-dimensional shape image of the product in the normal image; Step 4. Use the two-dimensional shape image obtained in step 3 as supervision information to train the structure repair sub-network to predict the target shape information of the defect image; Step 5. Adaptively fuse the inpainted target shape information into the appearance reconstruction subnetwork as an increment, and accurately reconstruct the image through the global structure to obtain the reconstructed image. Step 6. Input the simulated defect image, reconstructed image, and repaired target shape information into the localization subnetwork, and detect defects using apparent reconstruction error and structural information; Step 7. Calculate the loss of the output of the appearance reconstruction subnetwork and the corresponding normal image, calculate the loss of the output of the positioning subnetwork and the mask template in step 2, and jointly train and optimize the appearance reconstruction subnetwork and the positioning subnetwork.
2. The progressive industrial product surface defect detection method based on decoupled characterization according to claim 1 is characterized in that: In the step 2, A mask template with a random shape is generated by random sampling of Perlin noise. The mask template is used to sample defect areas with different distributions from normal images from other natural scene image datasets. The defect areas are linearly weighted with the normal image to generate a simulated defect image. The binary mask template is used as the true value of the defect position, and the strength of the generated defect is controlled by the weighting coefficient.
3. The progressive industrial product surface defect detection method based on decoupled characterization according to claim 1 is characterized in that: In step 3, Use the pre-trained segmentation model SAM to automatically obtain the shape information corresponding to industrial products; for each category of industrial product surface image, only a rough product position anchor box prompt is given; the image and position prompt are input into the SAM model at the same time, and the output is a two-dimensional target shape image.
4. The progressive industrial product surface defect detection method based on decoupled characterization according to claim 1 is characterized in that: In step 4, The structural repair subnetwork is based on the Transformer architecture. Its input is an artificially simulated defect image and several randomly sampled normal images, and its output is the shape information of the repaired target. The pre-trained visual Transformer is used for feature extraction. The features of the normal image are sampled by 30% through the core set sampling method and stored in the memory bank. An abnormal-normal attention layer is designed. The feature F of the input image is queried for the four most similar features F′ in the memory bank through cosine similarity. F is used as the query. After concatenating F and F′, they are mapped through linear layers to obtain the key and value, and the normal-abnormal cross attention is calculated for feature screening and fusion. The fused features are input into 8 stacked self-attention layers and 2 convolutional layers, and the repaired two-dimensional target shape information is output.
5. The progressive industrial product surface defect detection method based on decoupled characterization according to claim 1 is characterized in that: In step 5, The appearance reconstruction subnetwork is based on a convolutional neural network and adopts an encoder-decoder structure. The target shape information repaired in step 4 is downsampled, multiplied by the zero-initialized learnable weights, and then spliced to each layer of the encoder in the appearance reconstruction subnetwork. The output of the appearance reconstruction subnetwork is compared with the corresponding normal image using the l2 loss and structural similarity loss, and the loss function is minimized for training.
6. The progressive industrial product surface defect detection method based on decoupled characterization according to claim 1 is characterized in that: In step 6, Input image I, reconstructed image I r And the restored target shape image S r Do the following: I d =I c +α(I C ·S r ) in represents concatenation along the channel dimension, "·" is the dot product operation, and α is the learnable weight initialized to zero; I d As the final input of the positioning sub-network.
7. The progressive industrial product surface defect detection method based on decoupled characterization according to claim 1 is characterized in that: In step 7, After setting a threshold for binarization on the output of the localization subnetwork, the focal loss is calculated with the mask template in step 2. The structure repair subnetwork is trained separately, and the reconstruction subnetwork and the localization subnetwork are jointly trained and optimized using the repaired target shape information as a guide.
Citation Information
Cited By
Industrial image change anomaly detection method and system based on artificial intelligence
CN120876941A
Industrial defect classification method based on comparative learning and feature decoupling
CN121305245A
Industrial defect visual detection method and system for decoupling defect features and imaging conditions
CN121544611A
Industrial Defect Visual Inspection Method and System for Decoupling Defect Features and Imaging Conditions
CN121544611B