Industrial defect intelligent detection method based on auxiliary supervision
The supervised industrial defect detection method addresses feature extraction and fusion challenges by using a variable convolutional alignment and auxiliary branch, enhancing accuracy and adaptability in complex industrial environments.
Patent Information
- Application Number
- CN202510419641.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-07-15
AI Technical Summary
The prior art has problems such as large complex background interference, insufficient feature extraction and fusion, and lack of generalization capabilities in industrial defect detection, resulting in poor detection results.
Using an intelligent industrial defect detection method based on auxiliary supervision, the large-core expansion of the large-core span gradient path aggregation network can be used to expand the effective receptive field, and the deformable convolution feature alignment module is aligned with adjacent features in the deep supervision architecture, combined with the auxiliary branch calculation objective function to update the main branch weight, and improve detection accuracy.
It improves the accuracy of industrial surface defect detection, enhances the generalization ability of the model in a variety of contexts, and improves the real-time performance and robustness of the detection.
Smart Images

Figure CN120318186A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing, and particularly relates to an intelligent industrial defect detection method based on auxiliary supervision. Background Art
[0002] In the current era of intelligent manufacturing, ensuring the quality and safety of industrial products is of utmost importance. In industrial production, surface defect detection has become an indispensable aspect, as it directly affects the performance, aesthetics, and even the overall service life of the finished products. However, due to the diverse types of defects on industrial products, ranging from minor point defects to severe damages, with various forms, combined with the huge interference brought by the complex background in the complex manufacturing process, these factors together pose great challenges to the detection work.
[0003] Early process detection mainly relied on manual visual inspection. Although this method is intuitive, it is limited by human judgment, sensitivity to environmental factors, and difficulty in maintaining consistency. With the progress of technology, machine learning methods gradually came into view. At this stage, predefined rules or learning models were used to identify features, such as support vector machines (SVMs) and decision trees. These methods achieved process automation to a certain extent, but still faced problems such as high complexity, over - reliance on specific features, and difficulty in adapting to new scenarios in defect recognition, especially in defect detection with variable and complex backgrounds. The limitations of traditional machine learning prompted people to start exploring deep learning, especially in the field of defect detection. Deep learning, especially detection models based on convolutional neural networks (CNNs), such as YOLO, RCNN, and SSD, significantly improved real - time performance, accuracy, and robustness by directly extracting features automatically from the image end - to - end without manual design.
[0004] A large number of research results have emerged in improving the defect detection accuracy in the presence of multi - scale defects and low - contrast backgrounds. However, these methods focus on using attention to enhance feature extraction and expression or using modified pyramid networks to optimize feature fusion, while ignoring the problem of the effective receptive field in the feature extraction process and the spatial misalignment problem caused by step - by - step sampling in the feature fusion process, making their methods only suitable for fixed application scenarios and lacking the generalization ability required for the variability of industrial surface defects. Summary of the Invention
[0005] To solve the above - mentioned problems of the prior art, the present invention adopts an intelligent industrial defect detection method based on auxiliary supervision, including: obtaining industrial product surface defect data, inputting the industrial product surface defect data into a trained industrial defect detection model to obtain a defect detection result; the industrial defect detection model includes: a main branch, a deformable convolution feature alignment module, an auxiliary branch, a feature fusion layer, and a detection head;
[0006] The training process of the industrial defect detection model includes:
[0007] S1. Obtain a dataset of industrial product surface defects; the dataset of industrial product surface defects includes multiple industrial product surface defect image data;
[0008] S2. Input the industrial product surface defect image data into the main branch to obtain feature maps at different stages;
[0009] S3. Input the feature maps at different stages into the deformable convolution feature alignment module to obtain aligned features at different stages;
[0010] S4. Input the aligned features at different stages and the industrial product surface defect image data into the auxiliary branch to obtain feature maps at different stages; input the feature maps at different stages into the detection head to obtain defect detection results; calculate the loss function value according to the defect detection results;
[0011] S5. Update the parameters of the main branch, the deformable convolution feature alignment module, and the auxiliary branch according to the loss function value; input the industrial product surface defect image data into the updated main branch to obtain updated feature maps at different stages;
[0012] S6. Input the updated feature maps at different stages into the feature fusion layer, and input the output of the feature fusion layer into the detection head to obtain defect detection results;
[0013] S7. Calculate the loss function value according to the defect detection results, update the parameters of the industrial defect detection model according to the loss function value, and when the loss function value is the smallest, obtain the trained industrial defect detection model.
[0014] Beneficial effects:
[0015] 1. The present invention uses the large-kernel dilated separable convolution module and the cross-gradient path aggregation module in the large-kernel cross-gradient path aggregation network to expand the effective receptive field and eliminate redundant gradient information, improving the accuracy of industrial surface defect detection; 2. The present invention uses an auxiliary supervision structure to additionally introduce an auxiliary branch to calculate the objective function, and obtains reliable gradient information according to the objective function to update the weights of the early network of the main branch, improving the accuracy of industrial surface defect detection; 3. The present invention aligns adjacent features in the deep supervision architecture through the deformable convolution feature alignment module to replace the traditional step-by-step sampling feature fusion operation, thereby improving the accuracy of industrial surface defect detection. Description of the Drawings
[0016] Figure 1 It is a flowchart of an industrial defect intelligent detection method based on auxiliary supervision provided by an embodiment of the present invention;
[0017] Figure 2Structural diagram of the large kernel cross-gradient path aggregation network module provided by the embodiment of the present invention;
[0018] Figure 3 Structural diagram of the large kernel dilated separable convolution module provided by the embodiment of the present invention;
[0019] Figure 4 Structural diagram of the cross-gradient path aggregation module provided by the embodiment of the present invention. Detailed implementation manners
[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0021] As Figure 1 shown, the present invention adopts an intelligent industrial defect detection method based on auxiliary supervision, including: obtaining industrial product surface defect data, inputting the industrial product surface defect data into a trained industrial defect detection model to obtain a detection result; the industrial defect detection model includes: a main branch, a deformable convolution feature alignment module, an auxiliary branch, a feature fusion layer, and a detection head;
[0022] The training process of the industrial defect detection model includes:
[0023] S1. Obtain an industrial product surface defect data set; the industrial product surface defect data set includes multiple industrial product surface defect image data;
[0024] The acquisition of the industrial surface defect data set includes: using a camera to collect industrial product surface defect image data, and calibrating the defect types for each image data with a labeling tool; wherein, the industrial product surface defect image can be an unlabeled image to be detected or a labeled image. If it is a labeled image, it can be used to train and optimize the deep network. If it is an unlabeled image, defect detection can be performed through the deep network.
[0025] S2. Input the industrial product surface defect image data into the main branch to obtain feature maps at different stages;
[0026] The main branch is a large kernel cross-gradient path aggregation network; the large kernel cross-gradient path aggregation network includes multiple cascaded stage modules;
[0027] The main branch processes the image data including:
[0028] S21. Input the image data into the first stage module to obtain the first stage feature map F1;
[0029] S22. Input the feature map F1 of the first stage into the second stage module to obtain the feature map F2 of the second stage;
[0030] S23. Input the feature map F of the previous stage into the current stage module to obtain the feature map F of the current stage; where l is the index of the stage; l-1 Input the feature map of the previous stage into the current stage module to obtain the feature map of the current stage; l ; where l is the index of the stage;
[0031] S24. Repeat step S23 until the feature map F output by the last stage module is obtained; where L is the number of stage modules. L ; where L is the number of stage modules.
[0032] As shown in Figure 2 , each stage module includes: a DDC module and a CGP module; where DDC is a large kernel dilated separable convolution and CGP is a cross-gradient path aggregation.
[0033] The processing of the feature map of the previous stage by the current stage module includes:
[0034] S231. Input the feature map F of the previous stage into the DDC module and perform batch normalization on the features output by the DDC module to obtain the feature F'; l-1 Input the feature map of the previous stage into the DDC module and perform batch normalization on the features output by the DDC module to obtain the feature F'; l-1 ;
[0035] As shown in Figure 3 , the processing process of the DDC module includes:
[0036]
[0037] Among them, C l , H l , W l are the number of channels, height, and width of the feature F', and β is batch normalization. l-1 ; where β is batch normalization.
[0038] The DDC module uses separable dilated convolution to design large kernel convolution blocks with different dilation rates and convolution kernel sizes. The specific steps include: performing convolution on the feature map F l-1 , performing separable dilated convolution on the convolved features to obtain M features Adding the M features and normalizing the added features. The specific formula is as follows:
[0039]
[0040] Among them, is the feature obtained by processing the feature F l-1 through Conv 1×1 convolution and separable dilated convolution ;k (m) is the size of the m-th convolution kernel of the separable dilated convolution, M is the number of convolution kernels of the separable dilated convolution, and M are added and batch-normalized to obtain the feature F l ' -1 .
[0041] Using the reparameterization technique, the separable dilated convolution can be transformed into a large-kernel convolution during the inference stage.
[0042] S232. The feature F′ l-1 is split in half along the channel dimension to obtain the feature and the feature
[0043]
[0044] S233. The feature is input into the CGP module, and the output of the CGP module is successively convolved, batch-normalized, and activated to obtain the feature The formula is as follows:
[0045]
[0046] where represents the SiLU activation function, and Conv 3×3 is a 3×3 convolution.
[0047] As Figure 4 shown, the CGP module processes the feature including:
[0048] The feature is successively convolved, batch-normalized, and activated to obtain the feature X l-1 ; the feature X l-1 is split in half along the channel dimension to obtain the feature and the feature
[0049]
[0050]
[0051] The features and are added, and the result of the addition is convolved and batch-normalized to obtain the feature The feature is processed to obtain the feature
[0052]
[0053] For the feature Feature and the feature perform processing to obtain the feature For the feature For the feature perform processing to obtain the feature
[0054]
[0055] For the feature Feature Feature perform processing to obtain Feature X l ; For Feature X l perform processing to obtain the feature
[0056]
[0057] Among them, represents the SiLU activation function, β is batch normalization, Conv 1×1 represents a 1×1 convolution, represents a separable dilated convolution, repConv 3×3 is a 3×3 reparameterized convolution block, Concat represents channel concatenation, and the output is obtained through a series of intermediate variables
[0058] S234. Concatenate the feature Feature and the feature F l ( 2 ) to perform channel concatenation to obtain the concatenated feature F l ″; For the concatenated feature F l ' perform convolution, batch normalization, and activation to obtain the feature map F l at the current stage; The formula is as follows:
[0059]
[0060] Among them, Conv 1×1 is a 1×1 convolution, Concat represents channel concatenation, and n is the number of input features, which is 3.
[0061] According to the image data resolution, customize the convolution kernel size of the large kernel separable dilated convolution module at each stage, and the overall network convolution kernel size is designed to increase stage by stage.
[0062] S3. Input the feature maps of different stages into the deformable convolution feature alignment module to obtain the aligned features of different stages;
[0063] There are multiple deformable convolution feature alignment modules, which correspond one by one to the stage modules of the main branch; each deformable convolution feature alignment module DFA l processes the feature maps F l 、F l+1 output by the corresponding stage module and the next stage module to obtain the aligned feature P l for each stage; each deformable convolution feature alignment module DFA l processes the feature maps F l 、F l+1 output by the corresponding stage module and the next stage module, including:
[0064] S31. Sample the feature maps F l 、F l+1 to the same size to obtain the sampled feature maps
[0065] S32. Fuse the sampled feature maps to obtain the fused feature A l ;
[0066] Fusing the sampled feature maps includes: concatenating the sampled feature maps and convolving the concatenated features to obtain the fused feature A l .
[0067] S33. Calculate the defect feature bias map Δ l according to the fused feature A l ;
[0068] Specifically, after splitting the channels of the fused feature A l in half, the features and the features are obtained. Convolve the feature to obtain the feature A l ', perform deformable convolution on the special A l ' to obtain the feature . Fuse the feature feature and the feature to obtain the defect feature bias map Δ l :
[0069]
[0070] Among them, DConv is deformable convolution, represents the SiLU activation function, β is batch normalization, Conv 1×11×1 convolution is denoted as, and Concat represents channel concatenation.
[0071] S34. Split the defect feature offset map Δ l by channel to obtain the split defect feature offset maps Δ′ l , Δ′ l+1 and Δ′ k ;
[0072] S35. Align the feature map l according to the split defect feature offset maps Δ′ l+1 to obtain the aligned feature map U , U l and U l+1 ;
[0073]
[0074] where U l , U l+1 is the aligned feature map, (h, w) are the feature points of the feature map, h' is the index of the height, w' is the index of the width, H l is the height of the feature map , and W l is the width of the feature map . Δ′ l (h) and Δ′ l (w) are the offsets of the feature map , and Δ′ l+1 (h) and Δ′ l+1 (w) are the offsets of the feature map , i.e., the offset transformation amounts.
[0075] S36. Weightedly combine the aligned feature map according to the defect feature offset map Δ′ k to obtain the aligned feature P l ; The specific formula is as follows:
[0076] λ = Sigmoid(Δ′ k )
[0077] P l = λU l + (1 - λ)U l+1
[0078] where Sigmoid is the Sigmoid function, λ is the learned weight, and P l represents the final aligned feature;
[0079] S4. Input the alignment features at different stages and the industrial product surface defect image data into the auxiliary branch to obtain feature maps at different stages; input the feature maps at different stages into the detection head to obtain the defect detection results; calculate the loss function value according to the defect detection results; the auxiliary branch has the same structure as the main branch;
[0080] The processing of the image data by the auxiliary branch includes:
[0081] S41. Input the image data and the alignment features of the first stage into the first stage module to obtain the first stage feature map;
[0082] S42. Input the first stage feature map and the alignment features of the second stage into the second stage module to obtain the second stage feature map;
[0083] S43. Input the feature map of the previous stage and the alignment features of the current stage into the current stage module to obtain the current stage feature map; where l is the index of the stage;
[0084] S44. Repeat step S43 until the feature map output by the last stage module is obtained.
[0085] S5. Update the parameters of the main branch, the deformable convolution feature alignment module, and the auxiliary branch according to the objective function value; input the industrial product surface defect image data into the updated main branch to obtain the updated feature maps at different stages;
[0086] S6. Input the updated feature maps at different stages into the feature fusion layer, and input the output of the feature fusion layer into the detection head to obtain the defect detection results;
[0087] The feature fusion layer includes four fusion modules; the processing of the updated feature maps at different stages by the feature fusion layer includes: inputting the updated feature maps at different stages into the first fusion module respectively, inputting the output of the first fusion module and the updated feature maps at different stages into the second fusion module, inputting the outputs of the first fusion module and the second fusion module into the third fusion module, and inputting the output of the third fusion module and the updated feature maps at different stages into the fourth fusion module to obtain the final fused features.
[0088] The detection head includes: a bounding box detection module and a class detection module.
[0089] S7. Calculate the loss function value according to the defect detection results, update the parameters of the industrial defect detection model according to the loss function value, and when the loss function value is the smallest, obtain the trained industrial defect detection model.
[0090] The loss function Loss formula is as follows:
[0091] Loss = λL B+(1 - λ)L C
[0092]
[0093] where L B is the bounding box loss function, L C is the classification loss function, λ is the weight coefficient, S is the grid size of the image, B is the number of bounding boxes predicted for each grid cell of the image, indicates whether the j-th bounding box in the i-th grid of the image predicts an object, (x ij , y ij ) are the coordinates of the center point of the j-th ground truth bounding box in the i-th grid of the image, are the coordinates of the center point of the j-th bounding box predicted for the i-th grid of the image, (w ij , h ij ) are the width and height of the j-th ground truth bounding box in the i-th grid of the image, are the width and height of the j-th bounding box predicted for the i-th grid of the image, y c is the ground truth label for class c, p c is the probability of predicting class c.
[0094] Furthermore, the trained model weight file is used to validate the validation set, and the performance of the model can be comprehensively evaluated by the average precision, recall rate, mean average precision, and FPS.
[0095] The above metrics are defined as follows:
[0096]
[0097] Where TP indicates that the model determines a positive sample as a positive sample; FP indicates that a negative sample is determined as a positive sample; FN indicates that a positive sample is determined as a negative sample; p(r) indicates the precision of the recall rate; n indicates the number of classes; AP i indicates the detection precision of the class.
[0098] Finally, the trained model weight file is used to test the test set to obtain the actual defect prediction results.
[0099] The above embodiments further elaborate on the purpose, technical solution, and advantages of the present invention. It should be understood that the above embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An intelligent industrial defect detection method based on auxiliary supervision, characterized in that, Including: Obtain the surface defect data of industrial products, input the surface defect data of industrial products into the trained industrial defect detection model, and obtain the detection result; The industrial defect detection model includes: a main branch, a deformable convolutional feature alignment module, an auxiliary branch, a feature fusion layer, and a detection head; The training process of the industrial defect detection model includes: S1. Obtain the surface defect dataset of industrial products; the surface defect dataset of industrial products includes multiple surface defect image data of industrial products; S2. Input the surface defect image data of industrial products into the main branch to obtain feature maps at different stages; S3. Input the feature maps at different stages into the deformable convolutional feature alignment module to obtain aligned features at different stages; S4. Input the aligned features at different stages and the surface defect image data of industrial products into the auxiliary branch to obtain feature maps at different stages; input the feature maps at different stages into the detection head to obtain the defect detection result; calculate the loss function value according to the defect detection result; S5. Update the parameters of the main branch, the deformable convolutional feature alignment module, and the auxiliary branch according to the loss function value; input the surface defect image data of industrial products into the updated main branch to obtain the updated feature maps at different stages; S6. Input the updated feature maps at different stages into the feature fusion layer, and input the output of the feature fusion layer into the detection head to obtain the defect detection result; S7. Calculate the loss function value according to the defect detection result, update the parameters of the industrial defect detection model according to the loss function value, and when the loss function value is the smallest, obtain the trained industrial defect detection model.
2. The intelligent industrial defect detection method based on auxiliary supervision according to claim 1, characterized in that, The main branch is a large kernel cross-gradient path aggregation network; the large kernel cross-gradient path aggregation network includes multiple cascaded stage modules; the processing of the surface defect image data of industrial products by the main branch includes: S21. Input the surface defect image data of industrial products into the first stage module to obtain the first stage feature map; S22. Input the first stage feature map into the second stage module to obtain the second stage feature map; S23. Input the feature map of the previous stage into the current stage module to obtain the feature map of the current stage; S24. Repeat step S23 until the feature map output by the last stage module is obtained.
3. An intelligent industrial defect detection method based on auxiliary supervision according to claim 2, characterized in that Each stage module includes: a DDC module and a CGP module; the processing of the feature map of the previous stage by the current stage module includes: S231. Input the feature map F of the previous stage into the DDC module, perform batch normalization on the output of the DDC module, and obtain the feature F'. l-1 l-1 ; S232. Split the feature F' l-1 into two halves after splitting the channel to obtain the features and the feature S233. Input the feature into the CGP module, and perform convolution, batch normalization, and activation on the output of the CGP module in sequence to obtain the feature F l ( 2 ); S234. Fuse feature feature and feature F l ( 2 ) to obtain the feature map F l ; Among them, DDC is a large kernel dilated separable convolution, and CGP is a cross-gradient path aggregation.
4. An intelligent industrial defect detection method based on auxiliary supervision according to claim 3, characterized in that, The DDC module processes the feature map F of the previous stage l-1 The processing includes: performing convolution on the feature map F l-1 Performing separable dilated convolution on the convolved features to obtain multiple features Adding multiple features And normalizing the added features 5. An intelligent industrial defect detection method based on auxiliary supervision according to claim 2, characterized in that, There are multiple deformable convolution feature alignment modules, which correspond one by one to the stage modules of the main branch; each deformable convolution feature alignment module DFA l processes the feature maps F l and F l+1 output by the corresponding stage module and the next stage module to obtain the aligned feature P l for each stage; each deformable convolution feature alignment module DFA l processing the feature maps F l and F l+1 output by the corresponding stage module and the next stage module includes: S31. Sample the feature maps F l and F l+1 to the same size to obtain the sampled feature maps S32. Fuse the sampled feature maps to obtain the fused feature A l ; S33. According to the fusion feature A l Calculate the defect feature offset map Δ l ; S34. Bias the defect feature map Δ l Split it by channel to obtain the split defect feature bias map Δ′ l , Δ′ l+1 and Δ′ k ; S35. According to the split defect feature offset maps Δ′ l , Δ′ l+1 align the feature map to obtain the aligned feature map; S36. According to the defect feature offset map Δ′ k Perform weighted combination on the aligned feature map to obtain the aligned feature P l .
6. The intelligent industrial defect detection method based on auxiliary supervision according to claim 5, wherein Calculating the defect feature bias map for each fused feature includes: for the fused feature A l After splitting the channels in half, the feature is obtained and the feature For the feature Perform convolution to obtain the feature A l ', for the special feature A l ' Perform deformable convolution to obtain the feature Combine the feature feature and the feature to perform fusion to obtain the defect feature bias map Δ l .
7. An intelligent industrial defect detection method based on auxiliary supervision according to claim 5, characterized in that Aligning the feature maps in the feature map group includes: Among them, U l , U l+1 is the aligned feature map, (h, w) are the feature points of the feature map, h' is the index of the height, w' is the index of the width, H l is the feature map height, W l is the feature map width, Δ′ l (h), Δ′ l (w) are the biases of the feature map Δ′ l+1 (h), Δ′ l+1 (w) are the biases of the feature map .
8. An intelligent industrial defect detection method based on auxiliary supervision according to claim 5, characterized in that Weighted combination of the aligned feature maps in the feature map group includes: λ = Sigmoid(Δ′ k ) P l = λU l +(1 - λ)U l+1 Among them, λ is the weight, and P l represents the alignment feature.
9. The intelligent industrial defect detection method based on auxiliary supervision according to claim 1, wherein, The feature fusion layer includes four fusion modules; The processing of the updated feature maps at different stages by the feature fusion layer includes: respectively inputting the updated feature maps at different stages into the first fusion module, inputting the output of the first fusion module and the updated feature maps at different stages into the second fusion module, inputting the outputs of the first fusion module and the second fusion module into the third fusion module, and inputting the output of the third fusion module and the updated feature maps at different stages into the fourth fusion module to obtain the final fused feature.
10. The intelligent industrial defect detection method based on auxiliary supervision according to claim 1, characterized in that, The loss function value Loss = λL B +(1 - λ)L C ; where, L B is the bounding box loss function, L C is the classification loss function, and λ is the weight coefficient.