A method for detecting surface defects of lightweight industrial products
By designing a lightweight hierarchical multi-scale feature extraction network and an improved object detection module, the problems of high computational complexity and insufficient defect recognition capabilities in the prior art are solved, and efficient and accurate surface defect detection of industrial products are achieved.
Patent Information
- Application Number
- CN202411847662.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-12-16
AI Technical Summary
The existing automated surface defect detection system has high computational complexity, which makes the detection process take a long time, making it difficult to achieve real-time or near-real-time detection, and is difficult to identify small-size defects, and insufficient detection accuracy, especially in products with complex backgrounds or various defect types.
A lightweight industrial product surface defect detection method is designed, and a lightweight hierarchical multi-scale feature extraction network is used as the backbone network. Combined with a cross-stage feature reshaping module and a scale feature adaptive aggregation unit, the capture capability of multi-scale features is improved, and an improved IoU loss function and dual attention mechanism are introduced to optimize the target detection performance.
It significantly reduces the computing complexity of the system, improves the identification ability and detection accuracy of small defects, and can realize real-time or near-real-time defect detection on high-speed production lines, improving product quality and production efficiency.
Smart Images

Figure CN119313659B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and artificial intelligence applications, and specifically relates to a method for detecting surface defects of lightweight industrial products. Background Art
[0002] In modern industrial production, automated surface defect detection technology is essential to ensure product quality and improve production efficiency. However, existing automated detection systems often have high computational complexity, which not only makes the detection process time-consuming, but also increases the demand for computing resources accordingly. In high-speed production lines, such high-computational complexity systems are difficult to achieve real-time or near-real-time defect detection. Secondly, the identification of small-size defects is a difficulty in existing technologies. Many automated detection systems are prone to missed detection when facing subtle surface defects due to resolution limitations or insufficient algorithm sensitivity. Although these small defects are not visually noticeable, they may have a significant impact on the final performance of the product, especially in industries with high precision requirements, such as semiconductors and aviation. Furthermore, existing detection systems often find it difficult to achieve satisfactory detection accuracy when dealing with products with complex backgrounds or diverse defect types. This may be due to the limitations of the algorithm for feature extraction, or because the data set used in model training is not comprehensive enough, resulting in insufficient model generalization ability. Such low-precision detection results not only lead to missed detection of unqualified products, but may also cause misjudgment of qualified products, thereby affecting production costs and product quality.
[0003] The existence of these problems not only limits the application effect of the detection system in industrial production, but also poses a challenge to improving production efficiency and ensuring product quality. Therefore, developing an automated detection technology that can reduce computational complexity, improve small defect recognition capabilities and detection accuracy is of great practical significance for promoting the development of industrial automation. Summary of the invention
[0004] In order to solve the shortcomings of the prior art, reduce the computational complexity of the system, and improve the recognition ability and detection accuracy in the field of industrial product surface defect detection, the present invention adopts the following technical solutions:
[0005] A method for detecting surface defects of lightweight industrial products comprises the following steps:
[0006] Step 1: Obtain a dataset of industrial product surface defect images;
[0007] Step 2: Construct an industrial product surface defect detection model, including a backbone network, a neck network, and a head network. The backbone network is a lightweight hierarchical multi-scale feature extraction network, and the neck network is a scale-sensitive neck. The features of different depths output by the lightweight hierarchical multi-scale feature extraction network are reshaped across stages, and features of different scales are aggregated. Local features are generated based on the reshaped and aggregated features for defect detection in the head network.
[0008] Step 3: Training the YOLO-IDL deep learning model using the industrial product surface defect image dataset;
[0009] Step 4: Deploy the trained YOLO-IDL deep learning model and detect surface defects of industrial products, and output the defect type and location.
[0010] Furthermore, in step 1, the acquired data set is preprocessed, including the following steps:
[0011] Step 1.1: Collect surface images of industrial products from actual industrial production lines, and obtain high-resolution images through industrial cameras, high-definition cameras and other equipment; combine public industrial product defect image datasets and other data resources related to the target scene; use simulation tools and defect generation algorithms to simulate and generate defect images similar to actual scenes to supplement the sample size;
[0012] Step 1.2: Use the labeling tool Label Studio to label the specific area of the defect and generate a VOC format file containing the defect type and location; classify the images by industrial product type and defect type to facilitate on-demand loading and data enhancement in model training.
[0013] Furthermore, the lightweight hierarchical multi-scale feature extraction network in step 2 includes a context layer, a lightweight hierarchical multi-scale feature extraction layer and a depth-separable convolution layer connected in sequence, and the subsequent lightweight hierarchical multi-scale feature extraction layer and the depth-separable convolution layer are alternately connected in sequence, and the feature extraction includes the following steps:
[0014] Step 2.1.1: The context layer is processed through multi-level convolution and padding, and then through maximum pooling, which gradually reduces the feature map resolution and reduces the amount of calculation while maintaining key features. Finally, the maximum pooling result is concatenated with the input feature after convolution, and the information at different levels is combined to obtain a feature map with lower resolution but richer channel information. This is used as the input of subsequent network layers, which enhances the multi-scale perception of small defects, thereby improving recognition ability and detection accuracy.
[0015] The formula for the choroid layer is as follows:
[0016] Fout1 = Concat(MaxPool(Conv(P(Conv(P(Conv(X)))))),Conv(X 1 ))
[0017] Among them, X 1 Represents the input features, with a dimension of H×W×C, Conv and P represent the convolution layer and padding operation, MaxPool represents the maximum pooling layer, Concat represents the concatenation operation, and F out1 Represents the output of the context layer.
[0018] Step 2.1.2: The lightweight hierarchical multi-scale feature extraction layer processes the network input by stacking multiple lightweight convolutions GhostConv. The lightweight convolution module reduces redundant calculations, thereby reducing computational complexity. Then, the output features of different lightweight convolutions are spliced to enhance the perception of small defects. The spliced features are then channel-weighted through the SE attention module to complete the extraction of key features to improve detection accuracy. Finally, the features processed by the attention module are residually connected with the original input features to prevent feature loss and further improve the recognition ability of small defects.
[0019] The formula of the lightweight hierarchical multi-scale feature extraction layer is as follows:
[0020] F out2 = Attention(Concat(LightConv 1 (X),LightConv 3 (X),...,LightConv 11 (X))) + X 2
[0021] Among them, X 2 represents the original input features of the lightweight hierarchical multi-scale feature extraction layer, LightConv represents lightweight convolution, Concat represents the concatenation operation of feature channels, Attention represents feature enhancement through the attention module, and F out2 Represents the output of a lightweight hierarchical multi-scale feature extraction layer.
[0022] Step 2.1.3: The depthwise separable convolution layer is used to extract features efficiently. The process includes deep convolution and point-by-point convolution. The deep convolution processes the spatial information of each channel and only performs simple channel-by-channel convolution to reduce the amount of calculation. Then, the point-by-point convolution is used to fuse the information between channels instead of directly performing standard convolution, which significantly reduces the parameters and computational complexity. At the same time, the deep convolution retains the detail information, and the point-by-point convolution integrates the global features, which enhances the detail perception and overall expression ability of small defects.
[0023] Furthermore, in step 2.1.3, the formula of the depthwise separable convolutional layer is as follows:
[0024]
[0025] Where c represents the input channel index, represents the depth convolution kernel, represents the point-by-point convolution kernel, σ is the nonlinear activation function, b represents the bias term, which prevents the model output from being completely dependent on the input features, and F sep Represents the output of a depthwise separable convolutional layer.
[0026] Furthermore, the scale-based consonant neck in step 2 includes a cross-stage feature reshaping module, a scale feature adaptive aggregation unit, an aggregation method cross-stage local network module VoVGSCSP and a lightweight convolution GSConv, and its execution process is as follows:
[0027] Step 2.2.1: The cross-stage feature reshaping module receives the feature maps of different depths output by the lightweight hierarchical multi-scale feature extraction network, aligns the number of channels through convolution operations on the feature maps of each stage, then concatenates the features along the depth dimension, extracts and optimizes the concatenated features, improves the network's expression ability and detection accuracy, and obtains the reshaped output features;
[0028] Step 2.2.2: Before feature aggregation, the scale feature adaptive aggregation module adjusts the features of different scales to keep the scale features consistent, and then splices them along the channel dimension to obtain the aggregated features;
[0029] Step 2.2.3: The aggregation method cross-stage local network module VoVGSCSP obtains aggregate features and reshapes features to generate local features. The aggregation method cross-stage local network module inputs a feature map, and first uses ordinary convolution to obtain the first feature map with adjusted channel number; then the input feature map is subjected to one ordinary convolution and two lightweight convolutions to obtain the other two feature maps; then the two feature maps are fused by element-by-element addition, and then concatenated with the first feature map; finally, the concatenated feature map is subjected to one ordinary convolution to obtain an output feature map, which helps to capture the subtle features of small defects and improve recognition ability and detection accuracy through multi-path fusion and feature retention mechanism;
[0030] The formula of the cross-stage local network module of the aggregation method is as follows:
[0031] F out6 = CBS(Concat(CBS(X 6 ),CBS(X 6 )+LightConv2(LightConv1(CBS(X 6)))))
[0032] Among them, X 6 represents the input feature map of the cross-stage local network module of the aggregation method, CBS(X 6 ) represents the input feature X 6 The result after being processed by the CBS convolution module. The CBS convolution module includes convolution, batch normalization and activation function. LightConv represents depthwise separable convolution, Concat represents concatenation operation, and F out6 represents the output of the cross-stage local network module of the aggregation method;
[0033] Step 2.2.4: Lightweight convolution GSConv, first use ordinary convolution to capture basic feature information of the local features to obtain the first set of feature maps; then reduce the number of parameters through depth-separable convolution operations, focus on important spatial features, and obtain the second set of feature maps; then splice the two sets of feature maps along the channel dimension; finally, in order to improve the expressiveness of the features and the computational efficiency of the model, perform channel shuffling on the spliced features. Rearranging the channels helps to reintegrate information, enhance the expressiveness of the features, and ensure information interaction between different channels, which not only retains the detailed features but also integrates multi-scale information. The final output features are used as the output of the scale-sensitive neck and input into the head network;
[0034] The formula of the lightweight convolution is as follows:
[0035] F out5 = Shuffle(Concat(CBS(X 5 ),DepthwiseConv(X 5 )))
[0036] Among them, X 5 Represents the input feature map of lightweight convolution, CBS(X 5 ) represents the input feature X 5 The result after being processed by the CBS convolution module. The CBS convolution module includes convolution, batch normalization and activation function. DepthwiseConv represents depthwise separable convolution, Concat represents concatenation operation, and Shuffle represents channel shuffling operation. By rearranging the channel order, the feature mixing effect is enhanced. out5 Represents the output of a lightweight convolution.
[0037] Furthermore, in the step 2, there are two cross-stage feature reshaping modules, four aggregation method cross-stage local network modules and two lightweight convolutions; the first cross-stage feature reshaping module receives the deep, middle and final output features of the lightweight hierarchical multi-scale feature extraction network, and the first local features are obtained after feature reshaping by the first aggregation method cross-stage local network module; the second cross-stage feature reshaping module receives the shallow, middle and first local features of the lightweight hierarchical multi-scale feature extraction network, and the reshaped features are together with the aggregation features, and the second local features are obtained by the second aggregation method cross-stage local network module. The second local features are then subjected to the first lightweight convolution and then spliced with the first local features after ordinary convolution, and then subjected to the third aggregation method cross-stage local network module to obtain the third local features. The third local features are then subjected to the second lightweight convolution and then spliced with the final output features of the lightweight hierarchical multi-scale feature extraction network after ordinary convolution to obtain the fourth local features. The second, third and fourth local features are used as the output of the scale-sensitive neck and input into the head network.
[0038] Furthermore, the cross-stage feature reshaping module in step 2.2.1 receives feature maps of three different stages, namely, deep, middle and shallow, output by the lightweight hierarchical multi-scale feature extraction network; an expansion convolution operation is performed on the shallow feature map to expand the receptive field while maintaining the resolution; a linear interpolation method is used to perform upsampling operations on the deep and middle feature maps to reduce the computational cost and match the shallow features to achieve multi-scale feature fusion; then, a dimension-upgrading operation is performed on the feature map to transform it from a three-dimensional feature with the shape of channel, height and width to a four-dimensional feature with the shape of channel, depth, height and width to enhance the recognition capability of small defects; then, the dimension-upgraded features are spliced along the depth dimension, and the spliced features are extracted and optimized through three-dimensional convolution, batch normalization and PReLU activation function to improve the network's expression capability and detection accuracy, and finally, the reshaped output features are obtained by using three-dimensional convolution;
[0039] The cross-stage feature reshaping module formula is as follows:
[0040] F out3 = Conv3D(PReLU(BN(Conv3D(Concat(UpAlign(F deep ,F mid ,F shallow ))))))
[0041] Among them, F deep 、F mid 、F shallow They represent the input features of the deep, middle and shallow layers respectively, UpAlign represents the alignment of feature dimensions, Conv3D represents three-dimensional convolution, BN represents batch normalization, PreLU represents the activation function, and Fout3 represents the output of the cross-stage feature reshaping module, and Concat represents the concatenation operation.
[0042] Furthermore, the scale feature adaptive aggregation module in step 2.2.2, for the large-scale feature map, first adjusts its channel number to be consistent with the mesoscale feature through the convolution module, and then uses adaptive average pooling to perform downsampling operation to make its spatial resolution match the mesoscale feature; for the small-scale feature map, first uses the convolution module to adjust its channel number to be consistent with the mesoscale feature, and then uses linear interpolation to perform upsampling to adjust its resolution to align with the mesoscale feature; for the mesoscale feature, uses different 3×3 convolution and 1×1 convolution mixed convolution structures for feature extraction, and then adds the results of the two, and obtains the enhanced feature through the Sigmoid activation function; then the large and small scale features with the same feature size are fused through the convolution operation; finally, the fused large and small scale features and the enhanced mesoscale feature are spliced along the channel dimension;
[0043] The formula of the scale feature adaptive aggregation module is as follows:
[0044] F out4 = Concat(CBS(F s ),UpSample(CBS(F d )),Sigmoid(MidConv(F m )))
[0045] Among them, CBS represents the convolution module, including convolution, batch normalization and activation function. s ) represents the large-scale feature map F s The result after processing by the CBS convolution module, UpSample(CBS(F d )) represents the small-scale feature map F d After being processed by the CBS convolution module and adjusted to the same resolution through UpSample upsampling, SigmoidMidConv (F m )) represents the mid-scale feature map F m After the MidConv multi-layer convolution, the Sigmoid activation function is applied to form the features. Concat represents the concatenation operation. out4 Represents the output of the feature adaptive aggregation module.
[0046] Furthermore, in step 3, the head network adopts the YOLO-IDL deep learning model, and the improved IoU loss function Unified-IoU optimized for high-quality target detection scenarios is introduced in its training. Through the dynamic scaling mechanism and weight adjustment, the optimization of high-quality anchor frames is enhanced, while the gradient of low-quality anchor frames is suppressed. The formula is as follows:
[0047]
[0048] in, represents the IoU loss function, FocusFactor represents the focus factor, and the weight is dynamically assigned according to the quality of the anchor box;
[0049] In view of the contradiction between the convergence speed of the YOLO-IDL deep learning model and high-quality detection, the hyperparameter ratio is used to dynamically shift the attention of the model. The formula is as follows:
[0050]
[0051] Among them, β represents the sensitivity of controlling weight adjustment, γ represents the threshold of the IoU loss function, and dynamically determines the quality range of anchor boxes that the model should prioritize to optimize. At the same time, a dual attention mechanism is designed for the bounding box regression loss, including high-quality anchor boxes and low-quality anchor boxes, to avoid the model's excessive attention to high-quality anchor boxes in the later stage of training and prevent overfitting. At the same time, the gradient gain is adjusted to ensure that the model continues to optimize the anchor boxes of ordinary quality and improve the overall regression performance.
[0052] Furthermore, step 4 comprises the following steps:
[0053] Step 4.1: Model deployment: Deploy the YOLO-IDL deep learning model on edge devices such as embedded devices and industrial cameras, and adapt the model's lightweight design to hardware with lower computing power.
[0054] Step 4.2: Data input: collect surface images of industrial products in real time and input them into the deployed YOLO-IDL deep learning model;
[0055] Step 4.3: Data processing; the YOLO-IDL deep learning model performs forward reasoning on the input image and outputs the type and location of each target; the detection results are then post-processed, including non-maximum suppression, removing overlapping detection boxes, retaining the boxes with the highest confidence, refining the detection boxes, ensuring boundary accuracy, and filtering low-confidence targets according to the confidence threshold;
[0056] Step 4.4: Output the results; the results include the location of the defect, the confidence level of the defect, and the type of defect.
[0057] The advantages and beneficial effects of the present invention are:
[0058] The present invention designs a lightweight hierarchical multi-scale feature extraction network as the backbone network. Compared with the existing YOLOv8 model, the designed backbone network significantly reduces the number of parameters and improves the detection frame rate. Secondly, a cross-stage feature reshaping module and a scale feature adaptive aggregation unit are designed, and the Slim-neck architecture is integrated to propose a neck network module named "Scale Lingxi Neck" to improve the detail detection and feature capture capabilities of multi-scale SAR images. Finally, an improved IoU loss function Unified-IoU optimized for high-quality target detection scenarios is introduced. Through a dynamic scaling mechanism and a hyperparameter ratio, the optimization of high-quality anchor frames is enhanced and the gradient of low-quality anchor frames is suppressed. At the same time, a dual attention mechanism is designed to further improve the overall regression performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 is a flow chart of a method in an embodiment of the present invention.
[0060] Figure 2 2 is a diagram of the structure of the venation layer in an embodiment of the present invention.
[0061] Figure 3 It is a structural diagram of a lightweight hierarchical multi-scale feature extraction layer in an embodiment of the present invention.
[0062] Figure 4 It is a structural diagram of a cross-stage feature reshaping module in an embodiment of the present invention.
[0063] Figure 5 1 is a structural diagram of a scale-feature adaptive polymerization unit in an embodiment of the present invention.
[0064] Figure 6 It is a structural diagram of the cross-stage local network module VoVGSCSP of the aggregation method in an embodiment of the present invention.
[0065] Figure 7 It is a structural diagram of the lightweight convolution GSConv in an embodiment of the present invention. DETAILED DESCRIPTION
[0066] The specific implementation of the present invention is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described here is only used to illustrate and explain the present invention, and is not used to limit the present invention.
[0067] like Figure 1 As shown, a method for detecting surface defects of lightweight industrial products of the present invention is based on YOLO-IDL and includes the following steps:
[0068] Step 1: Obtain the industrial product surface defect image dataset and preprocess the data; the construction process of the industrial product surface defect image dataset is as follows:
[0069] Step 1.1: Dataset source: Collect surface images of industrial products from actual industrial production lines, and obtain high-resolution images through industrial cameras, high-definition cameras and other equipment; combine public industrial product defect image datasets and other data resources related to the target scene; use simulation tools and defect generation algorithms to simulate and generate defect images similar to actual scenes to supplement the sample size;
[0070] Step 1.2: Dataset construction: Use the labeling tool Label Studio to label the specific areas of the defects and generate a VOC format file containing the defect type and location; classify the images by industrial product type and defect type to facilitate on-demand loading and data enhancement in model training.
[0071] Step 2: Build a YOLO-IDL industrial product surface defect detection model, design a lightweight hierarchical multi-scale feature extraction network as the backbone network, design a scaled Lingxi neck as the neck network, and use the YOLOv8 head network as the head network;
[0072] The lightweight hierarchical multi-scale feature extraction network consists of a context layer, a depth-separable convolutional layer, and a lightweight hierarchical multi-scale feature extraction layer. The execution process includes the following steps:
[0073] Step 2.1.1: Figure 2 As shown in the figure, the context layer gradually reduces the feature map resolution and reduces the amount of calculation through multi-level convolution and maximum pooling, while maintaining key features. Finally, feature splicing is performed to combine information from different levels and convert it into a feature map with lower resolution but richer channel information. As the input of subsequent network layers, the multi-scale perception of small defects is enhanced, thereby improving recognition ability and detection accuracy. The formula is as follows:
[0074] F out = Concat(MaxPool(Conv(P(Conv(P(Conv(X)))))),Conv(X)) (1)
[0075] Among them, X is the input feature, the dimension is H×W×C, Conv and P represent the convolution layer and padding operation, MaxPool represents the maximum pooling layer, and Concat represents concatenation;
[0076] Step 2.1.2: If Figure 3As shown in the figure, the lightweight hierarchical multi-scale feature extraction layer processes the network input by stacking multiple lightweight convolution modules GhostConv. The lightweight convolution module reduces redundant calculations, thereby reducing the computational complexity. Then, the output features of different convolution layers are spliced to enhance the perception of small defects. Then, the spliced features are channel-weighted through the SE attention module to complete the extraction of key features to improve the detection accuracy. Finally, the features processed by the attention module are residually connected with the original input features to prevent feature loss and further improve the recognition ability of small defects. The formula is as follows:
[0077] F out = Attention(Concat(LightConv 1 (X),LightConv 3 (X),...,LightConv 11 (X))) + X' (3)
[0078] Among them, X' is the original input feature, LightConv represents lightweight convolution, Concat represents the concatenation operation of feature channels, and Attention represents feature enhancement through the attention module.
[0079] Step 2.1.3: The depthwise separable convolution layer is used to extract features efficiently. The process includes deep convolution and point-by-point convolution. The spatial information of each channel is processed by deep convolution, and only simple channel-by-channel convolution is performed to reduce the amount of calculation. Then, point-by-point convolution is used to fuse the information between channels instead of directly performing standard convolution, which significantly reduces the parameters and computational complexity. At the same time, deep convolution retains the detail information, and point-by-point convolution integrates the global features, which enhances the detail perception and overall expression ability of small defects. The formula is as follows:
[0080] (2)
[0081] Where c represents the input channel index, is the depth convolution kernel, represents the point-wise convolution kernel, and σ is the nonlinear activation function.
[0082] The construction of the scale-based consonant neck is composed of a cross-stage feature reshaping module, a scale feature adaptive aggregation unit, a lightweight convolution GSConv, and an aggregation method cross-stage local network module VoVGSCSP. The execution process includes the following steps:
[0083] Step 2.2.1: Figure 4As shown in the figure, the cross-stage feature reshaping module receives the feature maps of the three different stages of deep, middle and shallow layers output by the proposed lightweight hierarchical multi-scale feature extraction network as input. First, the feature maps of each stage are convolved to align the number of channels to a uniform value; the shallow feature maps are dilated by convolution to expand the receptive field while maintaining the resolution; the deep and middle feature maps are upsampled by linear interpolation to reduce the computational cost and match the shallow features to achieve multi-scale feature fusion; then the feature maps are dimensionally upgraded to transform them from three-dimensional features with the shape of [channel, height, width] to four-dimensional features with the shape of [channel, depth, height, width] to enhance the recognition ability of small defects; then the four-dimensional features are spliced along the depth dimension, and the spliced features are extracted and optimized by three-dimensional convolution, batch normalization and PReLU activation function to improve the network's expression ability and detection accuracy. Finally, the reshaped output features are obtained by three-dimensional convolution, and the formula is as follows:
[0084] F out = Conv3D(PReLU(BN(Conv3D(Concat(UpAlign(F deep ,F mid ,F shallow ))))))(4)
[0085] Among them, F deep ,F mid ,F shallow They represent deep, middle and shallow input features respectively, UpAlign means aligning feature dimensions, Conv3D means three-dimensional convolution, BN means batch normalization, and PReLU is the activation function;
[0086] Step 2.2.2: If Figure 5 As shown in the figure, the implementation of the scale feature adaptive aggregation module is as follows: before feature aggregation, the features of different scales are adjusted. For the large-scale feature map, the number of channels is first adjusted to be consistent with the medium-scale features through the convolution module, and then the adaptive average pooling is used for downsampling to match its spatial resolution with the medium-scale features; for the small-scale feature map, the number of channels is first adjusted to be consistent with the medium-scale features through the convolution module, and then the linear interpolation is used for upsampling to adjust its resolution to align with the medium-scale features; for the medium-scale features, a hybrid structure of 3×3 convolution and 1×1 convolution is used for feature extraction, and then the results of the two are added, and the enhanced features are obtained through the Sigmoid activation function; then the large and small scale features with the same feature size are fused through the convolution operation; finally, the fused large and small scale features are spliced with the enhanced medium-scale features along the channel dimension, and the formula is as follows:
[0087] Fout = Concat(CBS(F s ),UpSample(CBS(F d )),Attention(MidConv(F m ))) (5)
[0088] Among them, CBS represents the convolution module, which is implemented as a combination of convolution, batch normalization and activation function. s ) represents the result of processing the large-scale feature map by the CBS convolution block, UpSample(CBS(F d )) indicates that the small-scale feature map is processed by CBS and adjusted to the same resolution by upsampling. Attention(MidConv(F m )) represents the features formed by applying Sigmoid activation function to the mid-scale feature map after multiple layers of convolution;
[0089] Step 2.2.3: If Figure 6 As shown in the figure, the implementation of the cross-stage local network module VoVGSCSP of the aggregation method is as follows: first, ordinary convolution is performed on the input feature map to obtain the first feature map with adjusted channel number; then, the feature map is subjected to one ordinary convolution and two lightweight convolutions to obtain the other two feature maps; then, the two feature maps are fused by element-by-element addition and then concatenated with the first feature map; finally, the concatenated feature map is subjected to one ordinary convolution again to output the feature map. Through multi-path fusion and feature retention mechanism, it is helpful to capture the subtle features of small defects and improve the recognition ability and detection accuracy. The formula is as follows:
[0090] F out = CBS(Concat(CBS(X),CBS(X)+LightConv2(LightConv1(CBS(X)))))(7)
[0091] Where X is the input feature map and LightConv stands for Depthwise Separable Convolution.
[0092] Step 2.2.4: Figure 7 As shown in the figure, the implementation of lightweight convolution GSConv includes the following steps: first, ordinary convolution is used to capture the basic feature information of the input feature map to obtain the first set of feature maps; then, the depth-separable convolution operation is used to reduce the number of parameters and focus on important spatial features to obtain the second set of feature maps; then the two sets of feature maps are spliced along the channel dimension; finally, in order to improve the expressiveness of the features and the computational efficiency of the model, the spliced features are shuffled through the channels. Rearranging the channels helps to reintegrate the information, enhance the expressiveness of the features, and ensure the information interaction between different channels, which not only retains the detailed features but also integrates the multi-scale information. The formula is as follows:
[0093] F out = Shuffle(Concat(CBS(X),DepthwiseConv(X))) (6)
[0094] Where X is the input feature map, DepthwiseConv represents the depthwise separable convolution, and Shuffle represents the channel shuffle operation, which enhances the feature mixing effect by rearranging the channel order.
[0095] Furthermore, at the end of the lightweight hierarchical multi-scale feature extraction network, the output of the last lightweight hierarchical multi-scale feature extraction layer is subjected to spatial pyramid pooling to obtain the final output, and the spatial pyramid pooling includes a first CBS module, three maximum pooling modules, a splicing module and a second CBS module connected in sequence, wherein the splicing module splices the outputs of the first CBS module and the three maximum pooling modules.
[0096] Step 3: Use the industrial product surface defect image dataset preprocessed in step 1 to train the YOLO-IDL model;
[0097] The training of the YOLO-IDL model introduces an improved IoU loss function Unified-IoU optimized for high-quality target detection scenarios. It enhances the optimization of high-quality anchor boxes through dynamic scaling mechanism and weight adjustment, while suppressing the gradient of low-quality anchor boxes. The formula is as follows:
[0098] (8)
[0099] in, is the standard IoU loss function, FocusFactor is the focus factor, and the weight is dynamically assigned according to the quality of the anchor box;
[0100] In order to solve the contradiction between model convergence speed and high-quality detection, the hyperparameter ratio is used to dynamically shift the model's attention formula as follows:
[0101] (9)
[0102] Among them, β represents the sensitivity of controlling weight adjustment, and γ is the threshold of IoU, which dynamically determines the quality range of anchor boxes that the model should prioritize to optimize. At the same time, a dual attention mechanism is designed for the bounding box regression loss, including high-quality anchor boxes and low-quality anchor boxes, to avoid the model's excessive attention to high-quality anchor boxes in the later stage of training and prevent overfitting. At the same time, the gradient gain is adjusted to ensure that the model continues to optimize the anchor boxes of ordinary quality and improve the overall regression performance.
[0103] Step 4: Deploy the trained YOLO-IDL model and detect surface defects of industrial products, output the defect type and location, including the following steps:
[0104] Step 4.1: Model deployment: Deploy the YOLO-IDL model on edge devices such as embedded devices and industrial cameras, and adapt the model's lightweight design to hardware with lower computing power.
[0105] Step 4.2: Data input: The surface image of the industrial product is collected in real time by the camera and input into the deployed YOLO-IDL model;
[0106] Step 4.3: Data processing; the YOLO-IDL model performs forward reasoning on the input image and outputs the type and location of each target; the detection results are then post-processed, including non-maximum suppression, removing overlapping detection boxes, retaining the boxes with the highest confidence, refining the detection boxes to ensure boundary accuracy, and filtering low-confidence targets according to the confidence threshold;
[0107] Step 4.4: Output the results; the results include the location of the defect, the confidence level of the defect, and the type of defect, such as scratches, cracks, contamination, etc.
[0108] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some or all of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting surface defects of lightweight industrial products, characterized in that The steps include: Step 1: Obtain a dataset of industrial product surface defect images; Step 2: Construct an industrial product surface defect detection model, including a backbone network, a neck network, and a head network. The backbone network is a lightweight hierarchical multi-scale feature extraction network, and the neck network is a scale-sensitive neck. The features of different depths output by the lightweight hierarchical multi-scale feature extraction network are reshaped across stages, and features of different scales are aggregated. Local features are generated based on the reshaped and aggregated features for defect detection in the head network. The lightweight hierarchical multi-scale feature extraction network includes a context layer, a lightweight hierarchical multi-scale feature extraction layer and a depth-separable convolution layer connected in sequence. The subsequent lightweight hierarchical multi-scale feature extraction layer and the depth-separable convolution layer are alternately connected in sequence. The feature extraction includes the following steps: Step 2.1.1: The context layer is processed through multi-level convolution and padding, and then through maximum pooling, to gradually reduce the feature map resolution while maintaining the key features. Finally, the maximum pooling result is concatenated with the input features after convolution, combining information from different levels to obtain the feature map of channel information. Step 2.1.2: The lightweight hierarchical multi-scale feature extraction layer processes the network input by stacking multiple lightweight convolutions, then concatenates the output features of different lightweight convolutions, and then performs channel weighting on the concatenated features through the attention module. Finally, the features processed by the attention module are residually connected with the original input features. Step 2.1.3: The depthwise separable convolution layer includes depthwise convolution and pointwise convolution. The depthwise convolution processes the spatial information of each channel, only performs simple channel-by-channel convolution, and then uses point-by-point convolution to fuse the information between channels. At the same time, the depthwise convolution retains the detail information, and the point-by-point convolution integrates the global features. The scale-based consonant neck includes a cross-stage feature reshaping module, a scale feature adaptive aggregation unit, an aggregation method cross-stage local network module, and a lightweight convolution. The execution process is as follows: Step 2.2.1: The cross-stage feature reshaping module receives the feature maps of different depths output by the lightweight hierarchical multi-scale feature extraction network, performs convolution operations on the feature maps of each stage to align the number of channels; then splices the features along the depth dimension, extracts and optimizes the spliced features, and obtains the reshaped output features; Step 2.2.2: Before feature aggregation, the scale feature adaptive aggregation module adjusts the features of different scales to keep the scale features consistent, and then splices them along the channel dimension to obtain the aggregated features; Step 2.2.3: The cross-stage local network module of the aggregation method obtains aggregate features and reshapes features to generate local features. The feature map is input to the cross-stage local network module of the aggregation method. First, ordinary convolution is used to obtain the first feature map with adjusted channel number; then the input feature map is subjected to one ordinary convolution and two lightweight convolutions to obtain the other two feature maps; then the two feature maps are fused by element-by-element addition and then concatenated with the first feature map; finally, ordinary convolution is performed on the concatenated feature map to obtain the output feature map; Step 2.2.4: Lightweight convolution first uses ordinary convolution to capture basic feature information of the local features to obtain the first set of feature maps; then, through the depth-separable convolution operation, focus on important spatial features to obtain the second set of feature maps; then, the two sets of feature maps are spliced along the channel dimension; finally, the spliced features are shuffled, and the final output features are used as the output of the scale-based Lingxi neck and input into the head network; Step 3: training a deep learning model using the industrial product surface defect image dataset; Step 4: Deploy the trained deep learning model and detect surface defects of industrial products.
2. A method for detecting surface defects of lightweight industrial products according to claim 1, characterized in that: In step 1, the acquired data set is preprocessed, including the following steps: Step 1.1: Collect high-resolution images of the surface of industrial products from actual industrial production lines; combine public industrial product defect image datasets and other data resources related to the target scene; use simulation tools and defect generation algorithms to simulate and generate defect images similar to actual scenes to supplement the sample size; Step 1.2: Mark the specific area of the defect and generate a file containing the defect type and location; classify the images by industrial product type and defect type.
3. A method for detecting surface defects of lightweight industrial products according to claim 1, characterized in that: In step 2.1.3, the formula of the depth-wise separable convolutional layer is as follows: Where c represents the input channel index, represents the depth convolution kernel, represents the point-by-point convolution kernel, σ is the nonlinear activation function, b represents the bias term, and F sep Represents the output of a depthwise separable convolutional layer.
4. A method for detecting surface defects of lightweight industrial products according to claim 1, characterized in that: In the step 2, there are two cross-stage feature reshaping modules, four aggregation method cross-stage local network modules and two lightweight convolutions; the first cross-stage feature reshaping module receives the deep layer, middle layer and final output features of the lightweight hierarchical multi-scale feature extraction network, and the first local features obtained by the first aggregation method cross-stage local network module after feature reshaping; The second cross-stage feature reshaping module receives the shallow and middle layer features of the lightweight hierarchical multi-scale feature extraction network and the first local feature. The reshaped features are together with the aggregation features, and are subjected to the second aggregation method cross-stage local network module to obtain the second local feature. The second local feature is subjected to the first lightweight convolution and then concatenated with the first local feature after ordinary convolution, and then subjected to the third aggregation method cross-stage local network module to obtain the third local feature. The third local feature is subjected to the second lightweight convolution and then concatenated with the final output feature of the lightweight hierarchical multi-scale feature extraction network after ordinary convolution to obtain the fourth local feature. The second, third and fourth local features are used as the output of the scale-sensitive neck and input into the head network.
5. The method for detecting surface defects of lightweight industrial products according to claim 1, characterized in that: The cross-stage feature reshaping module in the step 2.2.1 receives feature maps of three different stages, namely, deep, middle and shallow, output by a lightweight hierarchical multi-scale feature extraction network; an expansion convolution operation is performed on the shallow feature map; a linear interpolation method is used to upsample the deep and middle feature maps; then a dimension-upgrading operation is performed on the feature map, transforming it from a three-dimensional feature with the shape of channel, height and width to a four-dimensional feature with the shape of channel, depth, height and width; then the dimension-upgraded features are spliced along the depth dimension, and the spliced features are subjected to feature extraction and optimization through three-dimensional convolution, batch normalization and activation function, and finally the reshaped output features are obtained by using three-dimensional convolution.
6. A method for detecting surface defects of lightweight industrial products according to claim 1, characterized in that: The scale feature adaptive aggregation module in step 2.2.2, for large-scale feature maps, first adjusts its channel number to be consistent with the medium-scale features through the convolution module, and then uses adaptive average pooling to perform downsampling operations to make its spatial resolution match the medium-scale features; for small-scale feature maps, first uses the convolution module to adjust its channel number to be consistent with the medium-scale features, and then uses linear interpolation to perform upsampling to adjust its resolution to align with the medium-scale features; for medium-scale features, different hybrid convolution structures are used for feature extraction, and then the results are added to obtain enhanced features through activation functions; then the large and small scale features with the same feature size are fused through convolution operations; finally, the fused large and small scale features and the enhanced medium-scale features are spliced along the channel dimension.
7. A method for detecting surface defects of lightweight industrial products according to claim 1, characterized in that: In step 3, the head network adopts a deep learning model, and a loss function is introduced in its training. Through a dynamic scaling mechanism and weight adjustment, the optimization of high-quality anchor frames is enhanced, while the gradient of low-quality anchor frames is suppressed. The formula is as follows: Among them, LIoU represents the IoU loss function, FocusFactor represents the focus factor, and the weight is dynamically assigned according to the quality of the anchor box; In order to solve the contradiction between the convergence speed of deep learning models and high-quality detection, the hyperparameter ratio is used to dynamically shift the attention of the model. The formula is as follows: Among them, β represents the sensitivity of controlling weight adjustment, γ represents the threshold of the IoU loss function, and dynamically determines the quality range of anchor boxes that the model should prioritize to optimize. At the same time, a dual attention mechanism is designed for the bounding box regression loss, including high-quality anchor boxes and low-quality anchor boxes, and the gradient gain is adjusted at the same time.
8. A method for detecting surface defects of lightweight industrial products according to claim 1, characterized in that: The step 4 comprises the following steps: Step 4.1: Model deployment: Deploy the deep learning model in the edge device and adapt it to hardware with lower computing power through the lightweight design of the model; Step 4.2: Data input: collect surface images of industrial products in real time and input them into the deployed deep learning model; Step 4.3: Data processing; The deep learning model performs forward reasoning on the input image and outputs the type and location of each target. The detection results are then post-processed, including non-maximum suppression, removing overlapping detection frames, retaining the frames with the highest confidence, refining the detection frames, ensuring boundary accuracy, and filtering low-confidence targets based on the confidence threshold. Step 4.4: Output the results; the results include the location of the defect, the confidence level of the defect, and the type of defect.
Citation Information
Patent Citations
Wood board small target surface defect identification and detection method based on machine vision
CN118396968A
Fan surface defect detection method, device and equipment based on lightweight PC-EMA algorithm and storage medium
CN118691573A