Multi-scale metal surface defect detection method and system fusing dynamic convolution and attention

By integrating dynamic convolution and attention into a multi-scale detection method, the problems of insufficient accuracy and poor adaptability in metal surface defect detection are solved, and efficient identification and accurate detection of complex multi-scale defects are achieved.

CN120431027BActive Publication Date: 2025-12-05ZHUHAI COLLEGE OF JILIN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510454222.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-12-05
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

Existing methods for detecting defects on metal surfaces suffer from insufficient detection accuracy, weak feature extraction capabilities, poor adaptability, and multi-scale detection problems when faced with complex, multi-scale, and multi-morphological defects. They are also unable to effectively identify minute and irregular defects.

Method used

A multi-scale metal surface defect detection method integrating dynamic convolution and attention is adopted. The C3_DAD-CE module is used for deep feature extraction and the HSF-CA module is used for feature fusion. By combining dynamic deformation convolution and channel attention mechanism, the shape and weight of the convolution kernel are adaptively adjusted to improve the network's ability to identify multi-scale defects.

Benefits of technology

It significantly improves the accuracy and robustness of metal surface defect detection, better identifies small and irregular defects, enhances the efficiency and generalization ability of the detection system, and is suitable for complex production environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431027B_ABST
    Figure CN120431027B_ABST
Patent Text Reader

Abstract

The application discloses a multi-scale metal surface defect detection method fusing dynamic convolution and attention, relates to surface defect detection, and acquires a metal surface image to be detected; a C3_DAD-CE module formed by convolution stacking is used for extracting deep features of the metal surface image, so that a multi-level feature map capable of preliminarily characterizing a defect area of the metal surface is obtained; a channel attention mechanism is used for fusing the multi-level feature map, so that an enhanced feature map is obtained; and a classification network is used for predicting a defect category of the enhanced feature map, so that a detection result image is generated. The application discloses a multi-scale metal surface defect detection system fusing dynamic convolution and attention. The application aims to overcome the limitations of current methods in processing complex shapes and multi-scale defects, and comprehensively improve the precision and efficiency of detection.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to surface defect detection, more particularly, it relates to a multi-scale metal surface defect detection method and system fusing dynamic convolution and attention. BACKGROUND

[0002] With the accelerating industrialization process, especially in the fields of automobile manufacturing, aerospace, machining, etc., the quality control of metal materials becomes increasingly important. Metal surface defects not only affect the appearance and performance of products, but also can cause structural failures, and even endanger safety. Therefore, metal surface defect detection has become a crucial link in the production process, directly affecting the quality and production efficiency of products.

[0003] Traditional metal surface defect detection methods, such as manual visual inspection and automated detection methods based on traditional image processing, while in some cases can meet the basic detection needs, but with the diversification of metal surface defect types and the improvement of detection requirements, these methods have gradually shown many limitations. Manual detection is not only low in efficiency and strong in subjectivity, but also cannot meet the high-intensity detection needs in large-scale production; in addition, traditional image processing methods usually rely on handcrafted feature extraction, which is difficult to handle subtle defects in complex scenes, and there are certain deficiencies in processing speed and accuracy.

[0004] In recent years, deep learning technology has made significant progress in image recognition and defect detection, especially in the field of convolutional neural networks (CNN) and object detection. Deep learning models can automatically learn rich features from a large amount of image data, significantly improving the accuracy and efficiency of defect detection. Metal surface defect detection methods based on deep learning have become a research hotspot and have been applied in multiple industrial fields. These methods can effectively perform automated, efficient and accurate metal surface defect recognition by training complex neural network models, gradually replacing traditional detection methods.

[0005] In existing metal surface defect detection systems, common image processing techniques include traditional machine learning-based methods and deep learning methods. Traditional image processing methods, such as edge detection, template matching and morphological operations, rely on handcrafted feature extraction algorithms and are difficult to fully mine deep information in images. Especially when facing complex, subtle or morphologically diverse defects, the detection effect of traditional methods is often limited, and false positives or false negatives are prone to occur

[0006] In recent years, deep learning, especially Convolutional Neural Networks (CNN), has made significant progress in computer vision, widely used in image classification, object detection, semantic segmentation, and other tasks. In particular, deep learning methods have great potential in metal surface defect detection, which can automatically learn and extract features from images, enabling efficient and accurate defect recognition.

[0007] Yolo (You Only Look Once) is an efficient real-time object detection model that has been widely used in various real-time detection tasks due to its speed and accuracy. Yolo model converts object detection into a regression problem, directly predicting bounding boxes and classes from images, thus avoiding the complex candidate box generation and post-processing steps in traditional methods. Due to its excellent detection speed, Yolo is widely used in various industrial detection scenarios.

[0008] However, traditional object detection models such as Yolo have some problems when it comes to metal surface defect detection, which can be summarized as follows:

[0009] 1. Limited detection accuracy

[0010] Existing methods based on traditional image processing techniques, such as edge detection, template matching, and morphological operations, usually rely on manually designed feature extraction algorithms. These methods cannot effectively deal with complex, subtle, and diverse defect types on metal surfaces, especially when facing large morphological changes or complex surface textures, resulting in low detection accuracy and easy misclassification or missed detection.

[0011] 2. Insufficient feature extraction capability

[0012] Deep learning methods, especially those based on Convolutional Neural Networks (CNN) and Yolo object detection models, have improved the automation and accuracy of metal surface defect detection to some extent. However, existing models still have problems such as insufficient receptive field, neglect of local information, and ineffective use of contextual information when dealing with different sizes and shapes of metal surface defects. This leads to poor performance of the detection model in complex backgrounds or small size defects, especially when metal surface defects exhibit multi-scale and multi-morphology. Traditional convolutional neural networks tend to miss some important detailed features, affecting the overall detection performance.

[0013] 3. Poor adaptability

[0014] Existing deep learning methods are mostly trained for specific defect types, leading to inconsistent performance in different types of metal surfaces or variable production environments. Especially in cases where the morphology and scale of defects vary greatly, the adaptability of existing models is poor. Traditional Yolo models can quickly detect metal surface defects, but due to the static nature and fixed scale of convolution kernels, the model has difficulty effectively processing variable defect types and complex background information.

[0015] 4. Multi-scale detection problem

[0016] Metal surface defects have large size differences, which may include small cracks, large pits, or complex scratches. This requires a detection method that can effectively extract features at different scales to ensure accurate detection of defects of different sizes. However, existing Yolo models and other traditional convolutional neural networks often face performance bottlenecks when dealing with multi-scale problems. Although multi-scale training and data augmentation methods can alleviate this problem, there are still limitations in handling defects of different scales simultaneously. SUMMARY

[0017] The technical problem to be solved by the present application is to overcome the limitations of current methods in handling complex shapes and multi-scale defects, and to improve the accuracy and efficiency of detection.

[0018] The multi-scale metal surface defect detection method fusing dynamic convolution and attention described in the present application includes the following steps:

[0019] S1, obtaining a metal surface image to be detected;

[0020] S2, extracting deep features of the metal surface image through a C3_DAD-CE module formed by convolution stacking to obtain multi-level feature maps that can preliminarily characterize the defect regions of the metal surface;

[0021] S3, fusing the multi-level feature maps through a channel attention mechanism to obtain enhanced feature maps;

[0022] S4, predicting the defect categories of the enhanced feature maps through a classification network to generate a detection result image.

[0023] Preferably, in step S2, the deep features of the metal surface image are extracted through a C3_DAD-CE module formed by convolution stacking, specifically:

[0024] Surface defect features are initially extracted from the metal surface image to obtain an original feature map; a DAD-CE block is constructed, and the original feature map is input into a stacked block formed by multiple DAD-CE blocks connected in series for processing, and a multi-level feature map is output.

[0025] Preferably, the method for constructing the DAD-CE block is as follows:

[0026] S21. Generate the offset and mask of the convolution kernel through a dynamically deformable convolutional layer;

[0027] S22. Adjust the position of the convolution kernel according to the offset and assign weights according to the mask to complete the dynamic deformation convolution of the original feature map;

[0028] S23. Perform batch standardization on the original feature maps that have undergone dynamic deformation convolution;

[0029] S24. Perform a nonlinear transformation on the feature map output in step S23 using the SiLU activation function;

[0030] S25. Perform pooling operations at different scales on the feature map output in step S24 to obtain feature maps after pooling at different scales. Then, fuse the features at multiple scales of the feature maps after pooling at different scales through a convolutional layer.

[0031] Preferably, in step S21, the method for calculating the offset and the mask is as follows:

[0032] The output tensor is calculated using a convolutional layer:

[0033] conv_offset_mask=Conv 2D(x,W offset_mask b offset_mask );

[0034] Among them, W offset_mask b represents the kernel weights; offset_mask For bias terms; conv_offset_mask is the output tensor;

[0035] The output tensor is divided into offset and mask by a segmentation operation:

[0036] offset, mask=chunk(conv_offset_mask, 3, dim=1);

[0037] Where offset is the offset value; mask is the mask.

[0038] Preferably, in step S23, the batch normalization processing of the original feature map after dynamic deformation convolution is specifically as follows:

[0039] The mean and variance of all the original feature maps after the dynamic deformation convolution are calculated to obtain statistics required for normalization, and the original feature maps after the dynamic deformation convolution are normalized using the statistics; the normalized feature maps are linearly transformed by a learnable scaling factor γ and an offset term β to obtain the batch normalization output feature maps.

[0040] Preferably, the linear transformation is implemented by the following formula:

[0041]

[0042] wherein BN_output is the batch normalization output feature map; x i is the feature value input data of the i-th input feature map sample in the current batch; μ is the mean; σ 2 is the variance; ε is a small constant to prevent division by zero error.

[0043] Preferably, in step S3, the multi-level feature maps are fused by the channel attention mechanism, specifically:

[0044] S31, respectively performing adaptive average pooling and maximum pooling operations on the multi-level feature maps;

[0045] S32, respectively inputting the multi-level feature maps after the adaptive average pooling and maximum pooling operations into a convolution layer for dimension reduction operation, and activating the dimension-reduced feature maps by a ReLU activation function;

[0046] S33, restoring the two processing results output in step S32 to the original feature dimension consistent with the feature dimension of the multi-level feature maps by another convolution layer;

[0047] S34, adding the two output results output in step S33, and generating a channel attention map by Sigmoid activation of the addition result;

[0048] S35, element-wise multiplying the channel attention map and the multi-level feature maps to obtain enhanced feature maps.

[0049] Preferably, the surface defect features are preliminarily extracted from the metal surface image by 3x3 convolution.

[0050] Preferably, in step S4, the defect region marked on the detection result image, and the defect region is attached with a defect category and a confidence score.

[0051] A system for implementing the multi-scale metal surface defect detection method of fusing dynamic convolution and attention, comprising:

[0052] a metal surface defect picture input module configured to obtain a metal surface image to be detected;

[0053] a C3_DAD-CE module configured to extract deep features of the metal surface image through a convolution stacking operation to obtain a multi-level feature map capable of preliminarily characterizing a defect region of the metal surface;

[0054] an HSF-CA module configured to fuse the multi-level feature map through a channel attention mechanism to obtain an enhanced feature map;

[0055] a result output module configured to predict a defect category of the enhanced feature map through a classification network to generate a labeled detection result image.

[0056] Advantages

[0057] The present application has the advantages that:

[0058] 1. By introducing dynamic deformation convolution, the method can adaptively adjust the shape of the convolution kernel, effectively capturing fine-grained defect features; combined with the attention mechanism, the network's attention to key features is enhanced; the model uses multi-scale detection and feature fusion strategies to further optimize the network's context awareness and structural design. The present application not only can accurately detect various types of metal defects, especially small and irregular defects, but also significantly improves the robustness and generalization ability of the detection system, providing an efficient and reliable solution for metal surface quality detection.

[0059] 2. The C3_DAD-CE module of the present application combines dynamic deformation convolution and multiple stacking designs, further enhancing the depth and expressiveness of the network by stacking multiple DAD-CE blocks. Each DAD-CE block introduces a dynamic deformation convolution operation, allowing each sub-module to have adaptive convolution capabilities. By stacking multiple such modules, the network can learn more levels of features, thereby improving the recognition ability of complex patterns.

[0060] 3. The HSF-CA module of the present application greatly improves the network's perception of the importance of different channel features by combining adaptive pooling and maximum pooling operations in the channel attention mechanism. In traditional convolutional neural networks (CNN), each channel is usually considered equally important, but in reality, the contributions of different channel features to the task are different. HSF-CA dynamically adjusts the channel weights, solving this problem, so that the network can better focus on task-related key information while suppressing irrelevant features. BRIEF DESCRIPTION OF DRAWINGS

[0061] Figure 1Flow chart of the multi-scale metal surface defect detection method of the present application;

[0062] Figure 2 Structure diagram of the C3_DAD-CE module of the present application;

[0063] Figure 3 Implementation flow chart of the HSF-CA module of the present application;

[0064] Figure 4 System structure schematic diagram of the present application;

[0065] Figure 5 Actual recognition image of Tianchi aluminum dataset;

[0066] Figure 6 Image based on Figure 5 the image processed by the method of the present application;

[0067] Figure 7 Model overview comparison table of the Yolov5-DAD-CE-HSFPCA network optimized based on the method of the present application and the traditional Yolov5m network;

[0068] Figure 8 Model performance comparison table of the Yolov5-DAD-CE-HSFPCA network optimized based on the method of the present application and the traditional Yolov5m network;

[0069] Figure 9 Model speed comparison table of the Yolov5-DAD-CE-HSFPCA network optimized based on the method of the present application and the traditional Yolov5m network. DETAILED DESCRIPTION

[0070] The present application will be further described below in conjunction with examples, but does not constitute any limitation on the present application, and any limited number of modifications made by anyone within the scope of the claims of the present application is still within the scope of the claims of the present application.

[0071] Example 1

[0072] Referring to Figure 1 , the multi-scale metal surface defect detection method of the present application fusing dynamic convolution and attention, mainly based on the Yolov5m network, combines the DAD-CE block and the HSF-CA module, respectively optimizes the traditional convolutional neural network from multiple aspects such as feature selection, calculation efficiency and accuracy. Through these innovative modules, the present application significantly improves the performance of the metal surface defect detection task, especially in processing high-resolution images and complex defect patterns, showing stronger feature learning ability and calculation efficiency. The method includes the following steps:

[0073] S1, acquire a metal surface image to be detected. The metal surface image is input into the algorithm in a specified format (such as an RGB image or a grayscale image).

[0074] S2, extract deep features of the original feature map through a C3_DAD-CE module formed by convolution stacking to obtain a multi-level feature map capable of preliminarily characterizing the defect area of the metal surface. The convolution stacking operation is the stacking of multiple convolution blocks, which can further enhance the depth and feature extraction capability of the network.

[0075] Specifically, the convolution stacking operation specifically includes: preliminarily extracting surface defect features from the metal surface image to obtain an original feature map. Specifically, the surface defect features are preliminarily extracted from the metal surface image through 3x3 convolution. A DAD-CE block is constructed, and the original feature map is input into a C3_DAD-CE module formed by a plurality of DAD-CE blocks in series for processing to output a multi-level feature map. Each DAD-CE block uses dynamic deformation convolution, thereby introducing adaptive convolution operation in each sub-module, which enables the entire network to better adapt to variable input data. In the application of metal surface defect detection, the C3_DAD-CE module can effectively enhance the recognition ability of the model to defects of different scales and shapes by connecting a plurality of dynamic deformation convolution modules in series, thereby improving the detection precision and robustness.

[0076] As shown in the following table, the DAD-CE block includes a dynamic deformation convolution layer, a batch normalization layer, a ReLU activation function layer, and a pooling layer. Figure 2 The specific implementation process of the DAD-CE block is as follows:

[0077] S21, generate the offset and mask of the convolution kernel through the dynamic deformation convolution layer. Wherein, the original feature map x input into the module satisfies x∈R B×C×H×W , B is the batch size, C is the number of channels, H and W are the height and width of the original feature map respectively. The calculation method of the offset and the mask is as follows:

[0078] The output tensor is calculated by applying the convolution layer:

[0079] conv_offset_mask=Conv2D(x,W offset_mask ,b offet_mask );

[0080] Wherein, W offset_mask is the convolution kernel weight; b offset_mask is the bias term; conv_offset_mask is the output tensor. The output tensor is divided into offset and mask through the segmentation operation:

[0081] offset,mask=chunk(conv_offset_mask,3,dim=1);

[0082] wherein offset is the offset; and mask is the mask.

[0083] Specifically, the original feature map x is convolved by the convolution layer to generate an output tensor conv_offset_mask containing the offset and the mask. The output tensor is divided into three parts along the channel dimension through chunk operation, two of which are used as the offset and the other as the mask.

[0084] S22, adjusting the position of the convolution kernel according to the offset and assigning weights according to the mask to complete the dynamic deformation convolution of the original feature map.

[0085] After the offset and the mask are calculated, the deformation convolution operation is performed. The deformation convolution dynamically adjusts the position of the convolution kernel through the calculated offset and mask to better adapt to the local changes of the input feature map. Specifically, the original feature map x, the convolution kernel weight W, the offset offset and the mask mask are input into the deformation convolution operation. The deformation convolution dynamically adjusts the receptive field of the convolution kernel according to the offset and the mask of each position, and performs convolution calculation to obtain the output feature map. The specific operation is performed by torch.ops.torchvision.deform_conv2d, and the specific process is as follows:

[0086] output=DeformConv2D(x,W,offset,mask,b,stride,padding,dilation,groups).

[0087] wherein b is the bias term; stride, padding, dilation and groups are the step length, padding, dilation rate and grouping number in the convolution operation respectively. The operation dynamically adjusts the position of the convolution kernel, so that the convolution processing better adapts to the complex patterns of the metal surface defects.

[0088] In the DAD-CE block, first, the offset of the convolution kernel and the adaptive mask are generated by the dynamic deformation convolution layer. Specifically, by dynamically generating the offset of the convolution kernel, the spatial position of the convolution kernel is adjusted according to the changes of the local region of the original feature map, so as to better capture the detailed features in the image. Unlike the static offset of the traditional dynamic convolution, the offset of the present module is dynamically calculated according to the requirements of the specific task and the local information of the input image. Specifically, the input original feature map x generates three output tensors through convolution operation, two channels of which are used to generate the offset of deformation, and the other channel is used to generate the mask. The offset represents the displacement adjustment of the convolution kernel on the feature map, and the mask assigns an adaptive weight to each position to enhance the expression of local features. The generated offset and mask are then passed to the deformation convolution operation, which dynamically adjusts the position of the convolution kernel, so that the convolution operation can adapt to the diversity of metal surface defects.

[0089] S23, performing batch normalization processing on the original feature map completed by the dynamic deformation convolution.

[0090] After the dynamic deformation convolution operation in step S22 is completed, the batch normalization (Batch Normalization) process is introduced to prevent gradient explosion or disappearance and speed up the training process of the network. Specifically, after each convolution layer, the batch normalization operation is used to normalize the input feature map. First, the mean and variance of all original feature maps completed by the dynamic deformation convolution are calculated to obtain the statistical quantities required for normalization, and the original feature maps completed by the dynamic deformation convolution are normalized using the statistical quantities. The normalized feature map is linearly transformed by a learnable scaling factor γ and a bias term β to obtain the batch normalized output feature map. The specific calculation process is as follows:

[0091]

[0092] where N is the batch size, which is consistent with the batch size B in the above; x i is the feature value input of the i-th input feature map sample in the current batch; BN_output is the batch normalized output feature map; μ is the mean; σ 2 is the variance; ε is a small constant to prevent division by zero error.

[0093] S24, performing non-linear transformation on the feature map output by step S23 by SiLU activation function.

[0094] After performing the batch normalization as in step S23, the SiLU (Sigmoid Linear Unit) activation function is used to perform non-linear transformation on the output of step S23. The calculation formula of the SiLU activation function is:

[0095]

[0096] wherein x j is the output of step S23.

[0097] S25, performing a pooling operation of different scales on the feature map output by step S24 to obtain a feature map after different scale pooling, and fusing features of multiple scales of the feature map after different scale pooling through a convolution layer.

[0098] In order to further improve the fusion capability of multi-scale features, after step S24, a multi-scale adaptive pooling (MSAP) layer is introduced. The MSAP layer extracts context information at different scales by performing a pooling operation of different scales on the feature map, and fuses information at different scales through an adaptive convolution layer. This operation can effectively enhance the adaptability of the model to defects of different sizes, especially when processing metal surface defects, it can simultaneously identify large-area and small-range defect features. Specifically, the calculation formula of the MSAP layer is:

[0099] y = Conv (Concat (MaxPool (x k ) for k = 3, 7, 11)).

[0100] wherein x k is the feature map after different scale pooling.

[0101] The present application introduces a dynamic deformation convolution mechanism and a multi-scale adaptive pooling, so that the DAD-CE block provides stronger feature extraction and processing capability in metal surface defect detection. The dynamic deformation convolution can adaptively adjust according to the local features of different defects, and the MSAP layer improves the perception ability of the model to defects at multiple scales. The DAD-CE block can effectively identify and classify defects of different morphologies, sizes and complexities when processing metal surface defects, has higher robustness and accuracy, and is suitable for efficient detection of metal surface defects in the industrial field.

[0102] S3, fusing the multi-level feature maps through a channel attention mechanism to obtain an enhanced feature map. Specifically, this step is mainly realized through an HSF-CA module, which designs a channel attention mechanism combining adaptive pooling and max pooling operations to solve the problem that traditional convolutional neural networks cannot effectively capture the importance difference between channels.

[0103] As Figure 3 shown, the specific implementation process is as follows:

[0104] S31, respectively performing adaptive average pooling and max pooling operations on the multi-level feature maps. The formula of the adaptive average pooling operation is as follows:

[0105] avg_out = Conv2D(ReLU(Conv1(AvgPool(x i ))).

[0106] This operation is to pool the multi-level feature map x i into a feature map with a size of (B, C, 1, 1). Then, it is reduced to C / ratio through the convolution layer Conv1 and activated through the ReLU activation function. Then, the convolution layer Conv2 is restored to the original channel number. Similarly, MaxPool(x) is the max pooling operation, and the goal is also to pool the input feature map into a feature map with a size of (B, C, 1, 1). The convolution operation is similar to the average pooling operation, which is reduced through Conv1 and activated through ReLU, and finally restored to the original channel number through Conv2. The formula is as follows:

[0107] max_out = Conv2D(ReLU(Conv1(MaxPool(x i ))).

[0108] S32, respectively inputting the multi-level feature maps after the adaptive average pooling and max pooling operations into a convolution layer for dimension reduction operation, and activating the reduced feature maps through the ReLU activation function.

[0109] S33, restoring both processing results output in step S32 to the original feature dimension consistent with the feature dimension of the multi-level feature map through another convolution layer. After the first layer of convolution and ReLU activation, the channel number of the obtained feature map has been compressed (i.e. reduced), and then the channel number of the feature map is restored to the original C through the second layer of convolution operation, which is represented as follows:

[0110] avg_out, max_out = Conv2(ReLU(Conv1(·))).

[0111] Wherein, the convolution kernel size of the convolution layer Conv2 is 1x1, and the size of the feature map output in this step S33 is (B, C, 1, 1), which will be restored to the same channel number C as the multi-level feature map.

[0112] S34, adding the two output results output in step S33, and generating a channel attention map out through Sigmoid activation of the addition result. Wherein, the addition result can be represented as:

[0113] am_out = avg_out + max_out.

[0114] The calculation formula of sigmoid activation is:

[0115]

[0116] S35, element-wise multiply the channel attention map with the multi-level feature map to obtain an enhanced feature map.

[0117] In the HSF-CA module for fusing multi-level feature maps through a channel attention mechanism, the innovation lies in that it dynamically adjusts the channel weights of the feature map using global information, thereby enhancing the network's perception ability of key information. In traditional networks, each channel of the feature map is usually treated equally, ignoring the importance difference of different channels' features to the task. The HSF-CA module introduces global average pooling and max pooling operations to extract information from the global perspective and the maximum feature perspective, respectively, and then integrates these information through two convolutional layers to generate a channel attention map. This attention map is element-wise multiplied with the original feature map, thereby weighting the output features of each channel, enhancing the response of important channels, and suppressing the influence of irrelevant features.

[0118] In specific implementation, the HSF-CA module first applies adaptive average pooling and max pooling to the input feature map respectively to obtain two different global information representations (avg_out and max_out). Both of them are mapped to a lower dimension through a convolution operation (conv1), and then activated through a ReLU activation function, and then restored to the original feature dimension through a second convolutional layer (conv2). Through this design, the network can consider the feature information extracted by different channels under different pooling methods. Finally, the outputs of the two pooling methods are added and a channel attention map is generated through a Sigmoid function, which is element-wise multiplied with the original feature map, thereby enhancing the ability to capture key information.

[0119] S4, predict the defect category of the enhanced feature map through the classification network to generate a detection result image. Specifically, the defect area is marked on the detection result image, and the defect category and confidence score are attached to the defect area, so as to intuitively display the position and category of the metal surface defect, supporting subsequent analysis and processing.

[0120] Finally, the recognized metal surface defect image and its related data (such as defect category, position, etc.) are stored in the local storage of the computer for subsequent analysis, report generation, and production line monitoring.

[0121] As Figure 5 and Figure 6As shown, the actual recognition image of the existing Tianchi aluminum dataset is compared with the image recognized by the method of the application. As can be seen from the two graphs, in Figure 6 the confidence value corresponding to each type of defect is displayed after the annotation box, which intuitively shows the recognition result of the model for the defect. As shown in Figure 7-9 the model overview, performance and speed comparison table of the YoloV5-DAD-CE-HSFPCA network optimized based on the method of the application and the traditional YoloV5m network.

[0122] As shown in Figure 7 in terms of the number of parameters, YoloV5-DAD-CE-HSFPCA is 18,736,912, which is significantly less than 25,051,006 of YoloV5m. This indicates that the improved model reduces the complexity of the model and reduces the consumption of computing resources by optimizing the structure design. At the same time, its GFLOPs value is 59.2, which is lower than 64.0 of YoloV5m, indicating that the model has improved in computing efficiency and can complete more inference tasks in unit time. In addition, the weight size of YoloV5-DAD-CE-HSFPCA is 37.9MB, which is smaller than 50.5MB of YoloV5m, which is of great significance for the storage and deployment of the model on resource-limited devices (such as edge devices), which can effectively reduce the storage space requirement and transmission cost.

[0123] As shown in Figure 8 in terms of accuracy, for the overall classification

[0124] The mAP50 of YoloV5-DAD-CE-HSFPCA is 0.499, which is higher than 0.488 of YoloV5m. This indicates that under the standard of IoU threshold value of 0.5, the improved model can more accurately locate the target object, so that the overlap degree of the detection box and the real annotation box is higher, thereby reducing the false detection and missing detection, and the position and range of the target can be more accurately recognized.

[0125] In terms of mAP50-95, YoloV5-DAD-CE-HSFPCA reaches 0.301, while YoloV5m is 0.284. This improvement means that the improved model also has better detection performance under strict IoU threshold value. When the IoU threshold value is higher, the overlap requirement of the detection box and the real annotation box is more strict, and the improved model can still maintain high detection accuracy, which shows that it has stronger robustness and adaptability, which shows that in some industrial detection scenarios with very high requirements for target positioning accuracy, the improved model can better meet the needs.

[0126] As shown in Figure 9As shown, in terms of post-processing time, the Yolov5-DAD-CE-HSFPCA post-processing time is 0.5 ms, and the Yolov5m is 0.6 ms. The improved model has a shorter post-processing time, indicating that the improved model is more efficient in filtering, adjusting, and outputting the final detection results.

[0127] In general, the present application is designed and optimized for the difficulties of small scale, irregular shape and multi-scale characteristics in metal surface defect detection. The backbone network inherits the standard structure of Yolov5m and introduces the C3_DAD-CE module to improve the feature extraction capability. The C3 module of Yolov5m optimizes the feature expression by enhancing the cross-layer information transmission, while the C3_DAD-CE module introduces the offset and mask mechanism by using dynamic deformation convolution, realizing the adaptive adjustment of the convolution kernel shape, and significantly enhancing the sensitivity of the network to small and irregular defects.

[0128] The head network (Head) is improved on the basis of the detection head structure of the original Yolov5m, and the HSF-CA module is introduced to enhance feature fusion. This module combines adaptive average pooling and maximum pooling operations, and uses channel attention mechanism to dynamically adjust the importance of feature channels, thereby significantly improving the attention to key features, especially when detecting small or irregular defects. Moreover, the present application fuses multiple layers of features through Multiply and Add operations, making full use of the complementarity of features at different levels, and further improving the multi-scale feature expression capability of the model.

[0129] Example two

[0130] As Figure 4 shown, a system for implementing the above-mentioned multi-scale metal surface defect detection method combining dynamic convolution and attention, comprising:

[0131] A metal surface defect image input module, which is responsible for receiving metal surface images to be detected. These images may contain different types of defects, such as cracks, bubbles or other irregularities. The images can be obtained by industrial cameras or scanning devices, and the image size and format can be adjusted according to actual needs. The images need to be preprocessed, including steps such as grayscale, denoising, normalization, etc., to ensure that the input images meet the input requirements of the network model.

[0132] A C3_DAD-CE module is used to extract deep features from the original feature map through convolution stacking operation, to obtain multi-level feature maps that can preliminarily characterize the defect areas of the metal surface.

[0133] The traditional convolution operation often performs poorly in the face of object deformation, especially when dealing with metal surface defects, the morphology of the defects changes greatly, and the standard convolution kernel cannot flexibly adapt to these changes. In order to overcome this problem, the present application introduces a dynamic deformation convolution, which can dynamically adjust the position and shape of the convolution kernel according to the characteristics of the input image, thereby adapting to the morphological changes in different regions of the image. Specifically, the dynamic deformation convolution learns a spatial transformation mechanism to automatically adjust the convolution kernel, further improving the network's ability to handle deformation and irregular patterns. This mechanism is particularly suitable for metal surface defect detection tasks, because metal surface defects often have irregular and large deformation characteristics. Compared with traditional convolution operations, dynamic deformation convolution can better capture the subtle differences of metal defects, improving detection accuracy and robustness. The C3_DAD-CE module is a combination of dynamic deformation convolution and multiple stacking designs, which further enhances the depth and expressiveness of the network by stacking multiple DAD-CE blocks. Each DAD-CE block introduces dynamic deformation convolution operation, so that each sub-module has adaptive convolution ability. By stacking multiple such modules, the network can learn more levels of features, thereby improving the ability to recognize complex patterns. For example, in the metal surface defect detection task, defects of different scales and morphologies can be effectively captured through multi-level feature extraction. This design can significantly improve the detection accuracy and robustness of the network. In addition, the MSAP layer introduces a multi-scale pooling operation, effectively enhancing the ability of the metal surface defect detection model in different scale feature fusion and diverse defect recognition. By pooling the input feature map at different scales, the model can simultaneously extract global and local information, thereby improving the perception ability of large-scale and small-scale defects. This design makes the model more adaptable and robust when dealing with complex surface defects, while reducing information loss, improving feature expression ability, and enhancing model accuracy and flexibility through convolution fusion of multi-scale information.

[0134] The HSF-CA module is used to fuse multi-level feature maps through a channel attention mechanism to obtain enhanced feature maps.

[0135] The HSF-CA module greatly improves the network's ability to perceive the importance of different channel features by combining adaptive pooling and maximum pooling operations in a channel attention mechanism. In traditional convolutional neural networks (CNNs), each channel is usually considered equally important, but in reality, different channels have different contributions to the task. The HSF-CA module dynamically adjusts the channel weights, solving this problem, so that the network can better focus on task-related key information while suppressing irrelevant features.

[0136] The innovation of the module lies in that it extracts global information through two pooling methods (global average pooling and max pooling) and fuses the extracted features using convolution operation to generate a channel attention map. Specifically, the HSF-CA module first applies adaptive average pooling and max pooling operations to the input feature map. These two pooling methods extract information from global and local maximum feature perspectives respectively, capturing features at different levels. Adaptive average pooling helps to capture global information, while max pooling can retain the features of local maximum response.

[0137] Next, the two pooling results are respectively mapped to a lower dimension through a convolution operation (conv1) to reduce the amount of calculation and are activated through a ReLU activation function. Then, these low-dimensional features are restored to the same number of channels as the original feature map through a second convolution layer (conv2), thereby generating a channel attention map containing global and local information. Finally, the outputs of the two pooling methods are added and normalized by a Sigmoid function to generate the final channel attention map.

[0138] The channel attention map is element-wise multiplied with the original feature map, which can weight the output features of each channel. Specifically, the channel attention map can enhance the response to key information while suppressing unimportant channels. This design makes the network pay more attention to task-related features during training, improving the model's expression ability and robustness.

[0139] In summary, the HSF-CA module dynamically adjusts the channel weights of the feature map by introducing a channel-level attention mechanism and combining global and local information, enabling the network to more accurately capture useful feature information and improving the network's performance in visual tasks.

[0140] The result output module is used to predict the defect class of the enhanced feature map through a classification network and generate a labeled detection result image. Finally, the detection result image and its related data (such as defect class, position, etc.) are stored in the local storage of the computer for subsequent analysis, report generation, and production line monitoring.

[0141] The above only describes the preferred embodiments of the present application. It should be noted that for those skilled in the art, without departing from the structure of the present application, several modifications and improvements can be made, which will not affect the effect and practicality of the present application.

Claims

1. A multi-scale metal surface defect detection method fusing dynamic convolution and attention, characterized in that, The method comprises the following steps: S1, acquiring a metal surface image to be detected; S2, extracting deep features of the metal surface image through a C3_DAD-CE module formed by a convolution stack to obtain a multi-level feature map capable of preliminarily characterizing a defect area of the metal surface; specifically: preliminarily extracting surface defect features from the metal surface image to obtain an original feature map; constructing a DAD-CE block, inputting the original feature map into a stack block formed by a plurality of DAD-CE blocks in series for processing, and outputting a multi-level feature map; S3, fusing the multi-level feature map through a channel attention mechanism to obtain an enhanced feature map; S4, predicting a defect category of the enhanced feature map through a classification network to generate a detection result image; The construction method of the DAD-CE block is as follows: S21, generating an offset and a mask of a convolution kernel through a dynamic deformation convolution layer; S22, adjusting the position of the convolution kernel according to the offset and assigning weights according to the mask to complete dynamic deformation convolution of the original feature map; S23, performing batch normalization processing on the original feature map that has completed dynamic deformation convolution; S24, performing nonlinear transformation on the feature map output by step S23 through a SiLU activation function; S25, performing pooling operations of different scales on the feature map output by step S24 to obtain feature maps after different scale pooling, and fusing the features of multiple scales of the feature maps after different scale pooling through a convolution layer.

2. The method of claim 1, wherein the method of fusing dynamic convolution with attention for multi-scale metal surface defect detection is characterized by, In step S21, the calculation method of the offset and the mask is as follows: an output tensor is calculated by applying a convolution layer: ; wherein, is a convolution kernel weight; is a bias term; is an output tensor; the output tensor is segmented into an offset and a mask through a segmentation operation: ; wherein is an offset amount; is a mask.

3. The method of claim 1, wherein the method of fusing dynamic convolution with attention for multi-scale metal surface defect detection is characterized by, In step S23, the batch normalization processing on the original feature map that has completed dynamic deformation convolution is as follows: the mean and variance of all the original feature maps that have completed dynamic deformation convolution are calculated to obtain statistical quantities required for normalization, and the original feature maps that have completed dynamic deformation convolution are normalized using the statistical quantities; the normalized feature map is linearly transformed through a learnable scaling factor γ and an offset term β to obtain a batch-normalized output feature map.

4. The method of claim 3, wherein the method is a multi-scale metal surface defect detection method fusing dynamic convolution and attention, characterized in that, The linear transformation is realized by the following formula: ; wherein, is the batch normalized output feature map; x i is the input data of the feature value of the feature map sample of the i-th input in the current batch; μ is the mean; σ 2 is the variance; ε is a small constant to prevent division by zero error.

5. The method of claim 1, wherein the method of fusing dynamic convolution with attention for multi-scale metal surface defect detection is characterized by, In step S3, the multi-level feature map is fused through a channel attention mechanism, specifically: S31, performing adaptive average pooling and maximum pooling operations on the multi-level feature map, respectively; S32, inputting the multi-level feature maps after the adaptive average pooling and maximum pooling operations into a convolution layer for dimension reduction operation, and activating the dimension-reduced feature maps through a ReLU activation function; S33, restoring the two processing results output in step S32 to the original feature dimension consistent with the feature dimension of the multi-level feature map through another convolution layer; S34, adding the two output results output in step S33, and generating a channel attention map through Sigmoid activation of the addition result; S35, multiply the channel attention map with the multi-level feature map element by element to obtain an enhanced feature map.

6. The method of claim 1, wherein the method of fusing dynamic convolution with attention for multi-scale metal surface defect detection is characterized by, Preliminary surface defect features are extracted from the metal surface image through 3x3 convolution.

7. The method of claim 1, wherein the method of fusing dynamic convolution with attention for multi-scale metal surface defect detection is characterized by, In step S4, the defect area marked on the detection result image is labeled, and the defect area is attached with a defect category and a confidence score.

8. A system for implementing the multi-scale metal surface defect detection method fusing dynamic convolution and attention according to any one of claims 1-7, characterized in that, Comprise: a metal surface defect picture input module for obtaining a metal surface image to be detected; a C3_DAD-CE module for extracting deep features of the metal surface image through convolution stacking operation to obtain a multi-level feature map capable of preliminarily characterizing the defect area of the metal surface; an HSF-CA module for fusing the multi-level feature map through a channel attention mechanism to obtain an enhanced feature map; a result output module for predicting the defect category of the enhanced feature map through a classification network to generate a labeled detection result image.

Citation Information

Patent Citations

  • Underwater fish target detection method, device and equipment and storage medium

    CN118334703A

  • Shape memory polymer printing defect detection method, device and system and medium

    CN119722622A