YOLOv5-based accurate detection method for complex and diversified welding defects of steel pipe

By constructing the XCM-YOLOv5 model based on YOLOv5, combined with technical means such as Focus network and CBAM attention mechanism, the existing detection methods have low accuracy, slow speed and prone to human error in the detection of steel pipe weld defects, and high-precision and fast defect detection effects are achieved.

CN120198770APending Publication Date: 2025-06-24HEFEI UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510249561.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing manual detection and X-ray detection methods have problems such as high labor intensity, long time-consuming, easy to cause human errors, expensive and complex equipment in the detection of steel pipe weld defects, and it is difficult to effectively detect small defects.

Method used

The YOLOv5 model based on convolutional neural network is adopted to construct the XCM-YOLOv5 model, and high-precision detection of steel pipe weld defects through the Focus network, CBAM attention mechanism, CSP module, FPN+PAN structure and small object detection layer are achieved.

Benefits of technology

It improves the accuracy and speed of steel pipe weld defect detection, can effectively identify multiple complex defects, reduces the possibility of error detection, and improves the reliability and robustness of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198770A_ABST
    Figure CN120198770A_ABST
Patent Text Reader

Abstract

The invention discloses a YOLOv5-based accurate detection method for complex and diverse welding defects of a steel pipe, which comprises the following steps of: firstly, dividing acquired X-ray image data into a training set and a test set, preprocessing, and secondly, constructing a YOLOv5-based novel detection classification algorithm model XCM-YOLOv5 model which comprises a Backbone module, a Neck module and a Head module; the Backbone module comprises a Focus network, a CBAM module and a CSP module, the complexity and diversity of welding defects are solved through a CBAM attention mechanism by enhancing the capability of a model focusing on most relevant characteristics, and then multiple small target defects existing in a steel pipe welding seam are detected through the Neck module comprising an FPN + PAN module and a small target detection layer. The method has the characteristics of small calculation amount and high network identification precision, can be used for detecting various weld defects, and has the advantages of high detection speed and high accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision, image processing, and defect detection. Specifically, it is a method for improving the detection accuracy of steel pipe weld defects. Background Art

[0002] Steel pipes are widely used in high-risk and high-pressure scenarios such as oil, natural gas, and shale gas. During the production process, due to factors such as welding technology, defects such as pores, cracks, lack of fusion, and inclusions inevitably exist in the steel pipe welds. These defect areas are prone to stress concentration, weak mechanical strength, corrosion, etc., which can easily lead to sudden changes in the performance of steel pipes, fatigue damage, and seriously affect the usability and safety of steel pipes. The impacts of these defects on the mechanical properties and reliability of welding vary. Different types of defects have different impacts on the quality, performance, and service life of steel pipes, and there are also different national or industry standards. The detection and type identification of steel pipe defects are of great significance for adjusting the production process, improving the equipment status, and evaluating the performance and service life of steel pipes. Therefore, the development of advanced steel pipe defect detection and type identification technologies is crucial for ensuring the integrity and reliability of steel pipe systems in various industrial applications.

[0003] Currently, the main methods for detecting steel pipe weld defects - manual inspection and X-ray inspection - both have obvious defects and limitations. Manual inspection has a high labor intensity, takes a long time, requires skilled personnel to visually inspect the welds, is ineffective for detecting internal defects, is prone to human errors, and has inconsistent quality due to different skill levels of inspectors. On the other hand, the inspection that directly judges defects through X-ray intensity requires specialized equipment and professional knowledge, so it is both expensive and complex to implement. The interpretation of direct observation of X-ray images usually depends on the professional knowledge of inspectors, making it a subjective and time-consuming process. In addition, this method may overlook certain types of defects, especially small defects that are easily overlooked. These challenges highlight the need for more advanced, automated, and precise detection methods to improve the integrity and reliability of steel pipe systems.

[0004] Based on this, the present invention provides an algorithm for improving the detection accuracy and type identification rate of steel pipe weld defects. By using the convolutional neural network YOLOv5 model, an XCM-YOLOv5 model is proposed to improve the accuracy and speed of detecting steel pipe weld defects. Summary of the Invention

[0005] The present invention aims to provide an accurate detection method for complex and diverse welding defects of steel pipes based on YOLOv5, which can detect a variety of complex steel pipe weld defects and has the advantages of fast speed, high accuracy, simple model, and high applicability.

[0006] To achieve the above object, the present invention adopts the following technical solutions to be implemented:

[0007] An accurate detection method for complex and diverse welding defects of steel pipes based on YOLOv5, including:

[0008] Dividing the collected steel pipe weld defect dataset into a training set and a test set, and performing image preprocessing to obtain input images;

[0009] Construct a new detection and classification algorithm model XCM-YOLOv5 model based on YOLOv5, and input the preprocessed pictures into the XCM-YOLOv5 model for training;

[0010] The XCM-YOLOv5 model includes a Backbone module, a Neck module, and a Head module; the Backbone module includes a Focus network, a CBAM module, and a CSP module; the Neck module includes an FPN+PAN module and a small target detection layer;

[0011] Detect whether there are defects in different steel pipe welds under different environments;

[0012] For the input image, first use the Focus network structure to slice the picture. The spliced picture becomes 12 channels relative to the original RGB three-channel mode. Finally, the obtained new picture undergoes a convolution operation to finally obtain a two-fold downsampled feature map without information loss; then use the CBAM attention mechanism module to optimize the feature recognition and classification effect of different steel pipe weld defect types. The feature map passed through CBAM is downsampled by a 3×3 convolution, doubling the number of channels while halving the height and width; then the feature map is passed through the CSP module, directly combining the input and output of the backbone, and implementing different filter extractions for the features of different scales input from the previous layer; subsequently, 1×1 convolutions are added in parallel to constrain the number of channels and enhance the expression ability of the network;

[0013] In the Neck module, the FPN+PAN structure is adopted and used in the fusion of feature maps of different scales in FPN, changing a fixed convolution kernel into a convolution kernel that can adaptively change the attention according to the input; combined with the small target detection layer, adding a detection head for the 160x160 feature map to accurately identify various complex defects;

[0014] The small target detection layer upsamples the 80x80 feature map in the Neck to 160x160, and performs concat fusion with the 160x160 feature map of the second layer in the Backbone to obtain a 160x160 feature map, and then through downsampling, an 80x80 feature map is obtained again and sent to the Head module for detection;

[0015] Divide the running process of CBAM into CAM and SAM, and give the intermediate feature map F∈R C×H×W As the input, first, perform global max pooling and global average pooling on the input channels, input the one-dimensional vectors of the two poolings into the fully connected layer, and sum to generate the one-dimensional channel attention Mc∈R C×H×W , then multiply the channel attention by the input elements to obtain the channel attention adjusted feature map F'. Secondly, F is the global max and mean pool space, and the two two-dimensional vectors generated by the pooling are concatenated and convolved to finally generate the two-dimensional spatial attention Mc∈R C×H×W , the spatial attention is associated with the elements of F, and the attention process is described by the following equation:

[0016] F ′ =M c (F)

[0017]

[0018] where represents the corresponding element multiplication. Before the multiplication operation, it is necessary to broadcast the channel attention and spatial attention according to the spatial dimension and channel dimension respectively;

[0019] The channel attention equation is: M c (F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F)));

[0020] The spatial attention sequence is:

[0021] Further technology of the present invention:

[0022] Preferably, divide the data set into a training set and a test set according to a ratio of 8:2.

[0023] Preferably, the image preprocessing is specifically to perform random mirroring, flipping, cropping, scaling, adding Gaussian noise, adjusting brightness and adjusting contrast on the weld defect image for image enhancement; at the same time, combine the Mosaic data augmentation algorithm to combine multiple pictures into one picture.

[0024] Preferably, use the graphic image annotation tool LabelImg to annotate the defects in each preprocessed image.

[0025] Preferably, the preprocessed images are input into the XCM-YOLOv5 model for training. Specifically: First, the CBAM module is used to enhance the features of different types of steel pipe weld defects in different situations. Then, through multi-scale feature fusion, the MPD IOU is used as the loss function for the bounding boxes, while reducing the problem of imbalance between positive and negative samples within the prediction boxes. The overall model uses the SiLU activation function to generate non-linear relationships.

[0026] Preferably, the calculation method of the MPD IOU is as follows:

[0027] Preferably, for the SiLU activation function, its calculation formula is:

[0028] Compared with the prior art, the present invention has the following technical effects:

[0029] The present invention introduces a small target detection layer in the Neck part, which can better extract small defects, and at the same time, without changing the accuracy and precision of large target defects, it can better target the complexity and diversity of steel pipe weld defects;

[0030] The attention mechanism added by the present invention in the backbone network part adopts different size combinations for different types of steel pipe defects, and has better recognition ability and detection accuracy compared with the original attention mechanism or other attention mechanisms. The spatial attention component allows the model to focus on specific regions of the image, which are crucial for defect recognition, thereby enhancing the model's ability to accurately locate welding defects. The enhanced feature extraction and focusing provided by CBAM contribute to the detection of various types of defects in steel pipe welds, thus achieving more effective multi-class type recognition;

[0031] The MPD IoU in the present invention can more effectively distinguish real defects from background noise, reduce the possibility of false detection, and improve the overall accuracy. This overall evaluation of the model performance can provide a deeper understanding of the model's performance in terms of accuracy and localization, making it particularly suitable for the detection of steel pipe weld defects, and is equally effective for overlapping and non-overlapping bounding box regression scenarios, ensuring robust and reliable defect detection under various welding conditions;

[0032] The SiLU in the present invention provides a smooth non-linear feature, which can effectively capture complex feature relationships and enhance the model's ability to represent welding defects. Then, by combining linear and non-linear characteristics, SiLU enables the network to better select and retain important features, thereby improving the detection accuracy of weld defects. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 It is a schematic flow chart of the present invention.

[0034] Figure 2 This is a schematic diagram of the XCM-YOLOv5 model of the present invention.

[0035] Figure 3 It is a curve graph showing the change of spectral line intensity with the opening angle θ, and the data in the graph is normalized by the maximum value.

[0036] Figure 4 It is a calibration curve graph of spectral line self-absorption. (a) is without a cavity, and (b) is with a cavity.

[0037] Figure 5 It is a calibration curve graph for quantitative detection of Fe and Cr elements based on the internal standard method; circular marks represent data without a cavity, and triangular marks represent data with a cavity; (a) is for Fe element, and (b) is for Cr element. Detailed implementation manners

[0038] In order to make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the present invention will be further clarified below with reference to specific drawings.

[0039] The present invention discloses a precise detection method for complex and diverse welding defects of steel pipes based on YOLOv5, as Figure 1 shown, which specifically includes the following steps:

[0040] Step 1: First, divide the obtained data set into a training set and a test set, and perform preprocessing.

[0041] We first divide the data set into a training set and a test set according to a ratio of 8:2; perform image enhancement on the weld defect images by means of random mirroring, flipping, cropping, scaling, adding Gaussian noise, adjusting brightness and adjusting contrast; at the same time, combine the Mosaic data enhancement algorithm to combine multiple pictures into one picture according to a certain ratio, so that the model can identify targets in a smaller range; use the graphic image annotation tool LabelImg to annotate the defects in each preprocessed image.

[0042] Step 2: Construct an XCM-YOLOv5 model, and input the training set to train it.

[0043] The XCM-YOLOv5 model constructed by the present invention is as Figure 2 shown, and mainly includes a Backbone module, a Neck module and a Head module; the Backbone module includes a Focus network, a CBAM module and a CSP module; the Neck module includes an FPN+PAN module and a small target detection layer.

[0044] The training is specifically as follows: First, the CBAM module is used to enhance the features of different types of steel pipe weld defects in different situations. Then, through multi-scale feature fusion, the MPD IOU is used as the loss function for the bounding box, while reducing the problem of imbalance between positive and negative samples within the prediction box. The overall model uses the SiLU activation function to generate non-linear relationships.

[0045] The MPD IOU metric at the Head end provides a more accurate evaluation of the localization accuracy by considering the distance between the predicted bounding box and the ground truth box. By adopting an improved distance metric, MPD IOU can more effectively distinguish real defects from background noise, reducing the possibility of false detection and improving the overall accuracy. The calculation method is as follows:

[0046] The SiLU activation function used in the overall model (also known as the Sigmoid-weighted linear unit) is a new type of non-linear activation function, which is an improved version of Sigmoid and ReLU, with the characteristics of no upper and lower bounds, smoothness, and non-monotonicity. Its calculation formula is as follows:

[0047] Specifically for detection, for the input image, first, the Focus network structure is used to slice the picture. The spliced picture becomes 12 channels relative to the original RGB three-channel mode. Finally, the information of the obtained new picture is convolved to finally obtain a two-fold downsampled feature map without information loss. Then, the CBAM attention mechanism module is used to optimize the feature recognition and classification effects of different types of steel pipe weld defects. The feature map after passing through CBAM is downsampled by a 3*3 convolution, doubling the number of channels while halving the height and width. Then, the feature map passes through the CSP module. The CSP module is divided into two CSP structures. One CSP1_X structure is applied to the Backbone main network, and the other CSP2X structure is applied to the Neck. The input and output of the main backbone are directly combined to extract different filters for the features of different scales input from the previous layer. Subsequently, 1*1 convolutions are added in parallel to constrain the number of channels and enhance the expression ability of the network.

[0048] The FPN+PAN structure adopted in the Neck module is used for the fusion of feature maps of different scales in FPN, changing a fixed convolution kernel into a convolution kernel that can adaptively change the attention according to the input. Then, combined with small target detection layers such as Figure 3As shown in the figure, a detection head for the 160x160 feature map is added to achieve accurate identification and classification of various complex defects. The small target detection layer upsamples the 80x80 feature map in the Neck to 160x160, so that it can be concatenated and fused with the 160x160 feature map of the second layer in the Backbone to obtain a 160x160 feature map. After downsampling, an 80x80 feature map (FPN+TPN) is obtained and transmitted to the Head module for detection.

[0049] As Figure 4 The described CBAM module starts from two ranges of channels and spaces to implement a sequential attention structure from channels to spaces. Spatial attention allows the neural network to pay more attention to the pixel regions that play a decisive role in type recognition in the image and ignore irrelevant regions. At the same time, channel attention is used to process the allocation relationship of the feature map channels. At the same time, allocating attention to both dimensions enhances the impact of the attention mechanism on the model performance. The specific process is as follows: The operation process of CBAM is divided into CAM and SAM, and the intermediate feature map F∈R C×H×W is given as the input. First, global max pooling and global average pooling are performed on the input channels. The one-dimensional vectors of the two poolings are input into the fully connected layer, and the sum is used to generate a one-dimensional channel attention Mc∈R C×H×W . Then, the channel attention is multiplied by the input elements to obtain the channel attention adjusted feature map F'. Secondly, F is the global max and mean pool space, and the two two-dimensional vectors generated by the pool are concatenated and convolved, and finally a two-dimensional spatial attention Mc∈R C ×H×W is generated. The spatial attention is associated with the elements of F. The attention process can be described by the following equation:

[0050] F ′ =M c (F)

[0051]

[0052] where represents the corresponding element multiplication. Before the multiplication operation, channel attention and spatial attention need to be broadcast according to the spatial dimension and channel dimension respectively.

[0053] The channel attention equation is: M c (F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F)));

[0054] The spatial attention sequence is:

[0055] Next, specific experimental results are used to verify the technical effects of the present invention:

[0056] The embodiments of the present invention provide an accurate detection method for complex and diverse welding defects of steel pipes. Table

[0057] Ⅰ. Results of evaluation metrics

[0058]

[0059] As shown in Table 1 and Figure 5 shown, the model performs very well in most defect categories. For defects such as undercut, arc blowout, crack, overlap, slag inclusion, and lack of fusion, the precision is always between 0.917 and 1.00. Similarly, the recall rates of these defects are almost perfect, with values between 0.99 and 1.00. The F1 scores that balance precision and recall are also very high, between 0.957 and 1.00. This indicates that the model is well-calibrated and effective in identifying and correctly classifying these defect types. Although the detection precision of the air hole defect is relatively low, at 0.509, this does not mean that the model has poor performance. The effectiveness of air hole defect detection is affected by various factors. Small target size: Air hole defects are usually very small and may be densely packed inside a steel pipe, which makes them more challenging to identify in images and have ambiguity in sample data: The sample data of air holes may be ambiguous, resulting in difficulty in accurately labeling during the training and testing processes. Despite these challenges, the recall rate of the model for air hole defects reaches 0.830, which indicates that it can still detect a large part of these defects. This performance shows the potential of the model in dealing with complex and small defects. Although slag inclusion defects are not easily distinguishable from the background by the naked eye and often visually resemble concave edge defects, the model handles them well. This success is likely due to repeated training, which enables the model to learn subtle differences that may not be immediately obvious.

[0060] We used the same dataset to conduct experiments in Faster R-CNN, YOLOv5, and the optimized YOLOv5 model (XCM-YOLOv5) proposed in this paper respectively, and compared the results. As shown in Table 2, a comparison is made between YOLOv5, Faster R-CNN, and YOLOv5q.

[0061] Table 2. Comparison results of different model effects

[0062]

[0063] Generally speaking, the performance of deep learning-based defect detection algorithms is significantly better than traditional computer vision algorithms. Compared with Faster R-CNN and YOLOv5, the proposed XCM-YOLOv5 model in this paper performs excellently in various evaluation metrics such as mAP, F1 score, and accuracy. Specifically, the mAP value of XCM-YOLOv5 is 0.961, which is about 10% higher than 0.871 of YOLOv5 and 0.879 of Faster R-CNN. In terms of F1 score, YOLOv5q reaches 0.960, which is much higher than 0.880 of YOLOv5 and 0.860 of Faster R-CNN, indicating stronger stability and prediction accuracy in defect detection. In addition, the accuracy of YOLOv5q is 0.992, significantly exceeding 0.943 of Faster R-CNN and 0.899 of YOLOv5. In terms of the detection time of a single image, XCM-YOLOv5 also shows significant advantages. Although the detection time of XCM-YOLOv5 is 0.133 seconds, slightly longer than 0.120 seconds of YOLOv5, its speed is more than twice that of 0.437 seconds of Faster R-CNN.

[0064] The above shows and describes the basic principles, main features, and characteristics of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. The method for accurately detecting complex and diverse welding defects of steel pipes based on YOLOv5 is characterized by: include: The collected steel pipe weld defect data set is divided into a training set and a test set, and image preprocessing is performed to obtain input images; Build a new detection and classification algorithm model XCM-YOLOv5 based on YOLOv5, and input the preprocessed images into the XCM-YOLOv5 model for training; Detect whether there are defects in the welds of different steel pipes under different environments.

2. The method for accurately detecting complex and diverse welding defects of steel pipes based on YOLOv5 according to claim 1, characterized in that: The dataset is divided into training set and test set in a ratio of 8:

2.

3. The method for accurately detecting complex and diverse welding defects of steel pipes based on YOLOv5 according to claim 1, characterized in that: The image preprocessing specifically includes randomly mirroring, flipping, cropping, scaling, adding Gaussian noise, adjusting brightness and contrast for image enhancement of the weld defect image; and combining multiple images into one image using the Mosaic data enhancement algorithm.

4. The method for accurately detecting complex and diverse welding defects of steel pipes based on YOLOv5 according to claim 3 is characterized in that: Use the graphic image annotation tool Label lmg to annotate the defects in each preprocessed image.

5. The method for accurately detecting complex and diverse welding defects of steel pipes based on YOLOv5 according to claim 3 is characterized in that: The XCM-YOLOv5 model includes a Backbone module, a Neck module and a Head module; the Backbone module includes a Focus network, a CBAM module and a CSP module; the Neck module includes an FPN+PAN module and a small target detection layer.

6. The method for accurately detecting complex and diverse welding defects of steel pipes based on YOLOv5 according to claim 5, characterized in that: The preprocessed images are input into the XCM-YOLOv5 model for training. Specifically, the CBAM module is used to enhance the features of different types of steel pipe weld defects under different conditions. Then, multi-scale feature fusion is performed and MPDIOU is used as the loss function of the bounding box. At the same time, the imbalance of positive and negative samples in the prediction box is reduced. The overall model uses the SiLU activation function to generate nonlinear relationships.

7. The method for accurately detecting complex and diverse welding defects of steel pipes based on YOLOv5 according to claim 6, characterized in that: The MPDIOU calculation method is: θ), where B gt is the set of actual bounding boxes, B prd is the set of predicted bounding boxes and Θ is the parameters of the deep model used for regression.

8. The method for accurately detecting complex and diverse welding defects of steel pipes based on YOLOv5 according to claim 6, characterized in that: The SiLU activation function is calculated as follows:

9. The method for accurately detecting complex and diverse welding defects of steel pipes based on YOLOv5 according to claim 5, characterized in that: Detect whether there are defects in different steel pipe welds under different environments, specifically: For the input image, the Focus network structure is first used to slice the image. The spliced ​​image becomes 12 channels compared to the original RGB three-channel mode. Finally, the new image is subjected to convolution operation, and finally a double-downsampled feature map without information loss is obtained; then the CBAM attention mechanism module is used to optimize the feature recognition and classification effects of different types of steel pipe weld defects. The feature map after CBAM is downsampled through 3×3 convolution, so that its height and width are halved while the number of channels is doubled; then the feature map is passed through the CSP module to directly combine the input and output of the backbone, and different filters are extracted for the features of different scales of the previous layer input; then 1×1 convolution is added in parallel to constrain the number of channels and improve the network's expression ability; The FPN+PAN structure adopted in the Neck module is used to fuse feature maps of different scales in FPN, turning a fixed convolution kernel into a convolution kernel that can adaptively change attention according to the input; combined with the small target detection layer, a detection head with a 160x160 feature map is added to accurately identify a variety of complex defects; The small target detection layer upsamples the 80x80 feature map in Neck to 160x160, concats it with the 160x160 feature map of the second layer in Backbone to obtain a 160x160 feature map, and then downsamples it to obtain an 80x80 feature map, which is passed to the Head module for detection; The operation process of CBAM is divided into CAM and SAM, and the intermediate feature map F∈R is given C×H×W As input, first, perform global maximum pooling and global average pooling on the input channel, input the two pooled one-dimensional vectors into the fully connected layer, and sum them to generate a one-dimensional channel attention Mc∈R C×H×W , then multiply the channel attention by the input element to get the channel attention adjustment feature map F'. Secondly, F is the global maximum and mean pool space. The two two-dimensional vector pools generated by the pool are concatenated and finally generate the two-dimensional space attention Ms∈R C×H×W , spatial attention is associated with the elements of F, and the attention process is described by the following equation: F ′ =M c (F) in Represents the corresponding element-wise multiplication. Before the multiplication operation, channel attention and spatial attention need to be broadcasted according to the spatial dimension and channel dimension respectively; The channel attention equation is: c (F)=σ(MLP(AvgPool(F))+MLP(MaxPool(F))); The spatial attention sequence is:

Citation Information

Cited By

  • Weld defect color image detection method and system based on YOLO-World

    CN121810699A

  • A weld defect color image detection method and system based on YOLO-World

    CN121810699B