Weld DR image defect intelligent recognition method based on local-global efficient model improved YOLOv5

CN119091265BActive Publication Date: 2026-09-29CHINA NUCLEAR IND 23 CONSTR +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411092286.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2026-09-29
Estimated Expiration
2044-08-09

AI Technical Summary

Technical Problem

然而,对于小尺度缺陷的检测,却面临着漏检的问题

Benefits of technology

[0035](1)本发明对YOLOv5进行改进,在Backbone网络中,设计了一种高效特征提取模块iRMB_Conv,该模块兼顾动态全局建模和静态局部信息融合的优势,同时能够有效地增加模型的感受野。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119091265B_ABST
    Figure CN119091265B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on local-global efficient model building improvement YOLOv5's weld DR image defect intelligent identification method.The method designs an iRMB_Conv efficient feature extraction module, the module combines the technical advantages of dynamic global modeling and static local information fusion.In Backbone network, iRMB_Conv module can fully extract the multiscale defect features of weld DR image, help to improve the recognition accuracy of different scale defects.In addition, in Neck network, the application designs a context-aware module CSM, by using different expansion rates of dilated convolution, the features of different receptive fields are obtained, which effectively enhances the feature information of small defects in the weld DR image, and thus improves the recognition accuracy of the model for small defects.The intelligent identification model has high recognition accuracy and good robustness, and is suitable for industrial fields such as aviation, aerospace, nuclear industry and special equipment, promoting the application and deployment of weld DR image defect intelligent identification technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer vision and digital radiographic imaging technology. It designs an intelligent identification technology for defects in weld digital radiographic images, specifically an intelligent identification method for weld DR (Digital Radiography) images based on a local-global efficient modeling improvement of YOLOv5. Background Technology

[0002] Welding, as a fundamental method of workpiece joining, is widely used in manufacturing, aerospace, nuclear industry, and construction. Its quality has a direct and significant impact on the structural performance, making the detection of internal defects in welds crucial. Currently, commonly used non-destructive testing (NDT) methods for welds include radiographic testing (RT) and ultrasonic testing (UT). X-ray testing utilizes X-rays to penetrate the workpiece and form an image on a detector. By detecting the degree of X-ray absorption, it reveals the internal and external structure and defects of the inspected workpiece and is widely used in weld NDT. In traditional image interpretation methods, defect assessment of weld radiographic images mainly relies on subjective judgment by qualified personnel. This method is easily affected by subjective factors such as the assessor's eyesight and fatigue, leading to inconsistent assessment results. Furthermore, manual defect assessment is inefficient and cannot meet the ever-increasing inspection demands. Therefore, to reduce labor costs and improve inspection efficiency, it is essential to develop intelligent weld defect target recognition methods based on artificial intelligence.

[0003] With the rapid development of artificial intelligence technology, significant breakthroughs have been achieved in the field of defect detection, driving the development of a series of intelligent defect assessment methods based on machine learning. Traditional machine learning algorithms require manual design of feature vectors based on data distribution, making them applicable only to datasets with specific environments and targets. Furthermore, designing effective image features for classification models in complex welding environments remains a challenge. Deep learning is a machine learning technique based on artificial neural networks. Defect recognition methods powered by deep learning can automatically extract deep features of defects from raw images, eliminating the complexity and inaccuracies associated with manually designed feature extraction methods. Moreover, it can utilize large amounts of images for end-to-end training, thereby enhancing the generalization ability and robustness of the defect recognition model. In the field of weld defect recognition, this method has demonstrated excellent performance.

[0004] YOLOv5 is a lightweight network widely used in industrial object detection tasks. YOLOv5 mainly consists of a Backbone network, a Neck network, and a Head network. The Backbone network uses the CSPDarknet53 network structure to extract features from the input image. The Neck network includes PANet and FPN modules for further feature extraction and fusion. The Head network employs a lightweight, single-stage detector based on an anchor box design, responsible for generating predicted bounding boxes and class probabilities for object detection.

[0005] Defects in welds, such as porosity, inclusions, cracks, incomplete penetration, and lack of fusion, exhibit characteristics at multiple scales. The YOLOv5 target detection algorithm performs reasonably well in identifying larger-scale weld defects, meeting basic detection requirements. However, it faces the problem of missed detections for small-scale defects. Therefore, it is necessary to optimize the existing YOLOv5 network model to improve the accurate identification of multi-scale weld defects. Summary of the Invention

[0006] To address the challenges of multi-scale defect identification in existing weld seams, this invention provides an intelligent weld seam DR image defect identification method based on an improved YOLOv5 model using a local-global efficient modeling approach. An iRMB_Conv module is constructed within the Backbone network for dynamic global feature modeling and static local feature fusion. A Context-Aware Module (CSM) is built within the Neck network, which effectively enhances the feature information of small-scale defects, thereby improving the accuracy of multi-scale defect identification. This invention exhibits highly accurate defect identification capabilities and recall, providing strong support for the development of intelligent weld seam defect identification technology.

[0007] To address the aforementioned technical problems, this invention proposes an intelligent defect recognition method for weld seam DR images based on a local-global efficient modeling improvement of YOLOv5, comprising the following steps:

[0008] Step 1: Collect DR images of welds containing different types of weld defects and annotate the weld defects;

[0009] Step 2: Enhance the labeled data by rotation, mirroring, and cropping, and then randomly divide the enhanced labeled data into training set, validation set, and test set in an 8:1:1 ratio;

[0010] Step 3: On the PyTorch platform, construct a network model for intelligent identification of weld seam DR image defects based on the improved YOLOv5 model with efficient local-global model building.

[0011] Step 4: Input the data from the training set into the model constructed in Step 3 for training, and use the validation set to validate the model; after each epoch, calculate the loss and update the model parameters through backpropagation; training stops when the preset number of epochs is reached, and the trained network model is saved; and the generalization performance of the trained network model is evaluated using the test set.

[0012] Step 5: Input the DR image of the weld to be identified into the network model, and output the detection result through model inference to realize intelligent identification of weld defects.

[0013] Furthermore, in the intelligent defect recognition method for weld seam DR images described in this invention, wherein:

[0014] Step 1 includes the following:

[0015] 1-1) Conduct weld DR tests according to NBT 47013.11-2015 standard. The sensitivity, resolution, grayscale value and normalized signal-to-noise ratio of the weld DR image must meet at least the requirements of AB level detection technology.

[0016] 1-2) Use ImageJ image processing software to adjust the window width and window level of the weld DR image and export it as a JPG image;

[0017] 1-3) The labelme image annotation tool was used to annotate the five types of weld defects in the JPG format weld DR image and generate a JSON format annotation file. The five types of weld defects are porosity, inclusion, crack, incomplete penetration and lack of fusion. PO, SL, CRK, LP and LF are used to represent the weld defect categories of porosity, inclusion, crack, incomplete penetration and lack of fusion.

[0018] 1-4) Convert the JSON format annotation file into XML format.

[0019] In step 2, the labeled data is enhanced by rotation at angles of 90°, 180°, and 270°; mirroring includes horizontal and vertical mirroring; cropping enhancement is performed with the defect center as the center, and the cropping size is randomly selected within the range of 900-1200 pixels; the rotation and mirroring enhancement methods are applied randomly, while the cropping enhancement method is a mandatory application.

[0020] Step 3 includes the following:

[0021] 3-1) An iRMB_Conv module for feature extraction is constructed in the Backbone network. This module is implemented by concatenating M iRMB modules with a 3×3 convolutional layer. The core components of the iRMB module include an Expanded Window Multi-Head Self-Attention (EW-MHSA) module, a Depth-Wise Convolution (DW-Conv) module with integrated short connections, a 1×1 convolutional layer, and residual connections. In the EW-MHSA module, a prerequisite judgment step is first performed to determine whether to perform a self-attention operation. If it is determined that a self-attention operation should not be performed, the input features are directly subjected to a 1×1 convolution operation with channel expansion, and the result of this convolution is used as the output of the EW-MHSA module. Conversely, if a self-attention operation is required, the input features of the EW-MHSA module are first subjected to preliminary convolution processing. Subsequently, the convolutional features are split into matrices K and Q using the Split operator. Next, a dot product operation is performed on K and Q, and the Softmax function is applied to obtain the attention weights (AttentionMap). Finally, the attention weights are weighted with the original input features of the EW-MHSA module, and then subjected to a 1×1 convolution operation with channel expansion. The result is the output feature of the EW-MHSA module. The specific expression of the iRMB module is as follows:

[0022] Y1 = EW - MHSA(Y)

[0023] Y2=DW-Conv(Y1)+Y1

[0024]

[0025] Where Y represents the input feature of the iRMB module. The output features of the iRMB module are applied to Y2 by a 1×1 convolutional layer with channel shrinkage.

[0026] 3-2) Instantiate multiple iRMB_Conv modules by setting input channels, prerequisite judgment settings, number of multi-head attention, iRMB module depth M, output channels, and DW-Conv convolution kernels; the Backbone network can contain 4 to 5 instantiated iRMB_Conv modules for feature extraction.

[0027] Preferably, four iRMB_Conv modules are instantiated. The input channels of the four iRMB_Conv modules are 24, 32, 48, and 120 respectively. The prerequisite judgment settings are No, No, Yes, and Yes respectively. The number of multi-head attention is 16, 16, 20, and 20 respectively. The depth M of the iRMB modules is 3, 3, 9, and 3 respectively. The output channels of the four iRMB_Conv modules are 32, 48, 120, and 200 respectively. The kernel size of the depth-separating convolution (DW-Conv) is 3, 3, 5, and 5 respectively. Further, the output channels of the 3×3 convolutional layer in the iRMB_Conv module are the same as the output channels of the module, which are set to 32, 48, 120, and 200 respectively.

[0028] 3-3) Constructing a context-aware module (CSM) includes: extracting multi-scale features F1, F2, and F3 through three convolutional layers with different dilation rates; adjusting their dimensions using 1×1 convolutions to obtain F′1, F′2, and F′3, and concatenating F′1, F′2, and F′3 along the channel dimension to form a new feature; generating normalized weights from the concatenated feature through 1×1 convolutions and Softmax operations, and splitting it into three parts along the channel dimension, which serve as adaptive weights for the multi-scale features F1, F2, and F3 respectively; finally, multiplying the multi-scale features F1, F2, and F3 with their respective weights and summing them to obtain the final context fusion feature, which is the output feature of the context-aware module.

[0029] 3-4) Replace the C3 module in the 13th layer of the Neck part with the context-aware module CSM mentioned above, so as to provide fully fused feature information for the defect classification and bounding box regression tasks in the downstream Head part.

[0030] Step 4 includes the following:

[0031] 4-1) Set the hyperparameters, including: batch size of 4, optimizer of stochastic gradient descent (SGD), initial learning rate of 0.01, total training epochs of 200, and input image size of 640×640.

[0032] 4-2) During the model training phase, the training set data is input into the model, and the weights in the model are adjusted by the optimization algorithm through backpropagation and gradient descent, so that the model can learn useful features and patterns from the data.

[0033] 4-3) The validation set is used to evaluate the model's performance and tune hyperparameters during training. At the end of each epoch, the model is evaluated on the validation set to determine whether the model is overfitting or underfitting, thereby selecting the optimal hyperparameters.

[0034] Compared with the prior art, the beneficial effects of the present invention are:

[0035] (1) This invention improves YOLOv5 by designing an efficient feature extraction module iRMB_Conv in the Backbone network. This module combines the advantages of dynamic global modeling and static local information fusion, and can effectively increase the receptive field of the model.

[0036] (2) The Backbone network constructed by the iRMB_Conv module in this invention can make full use of the excellent processing capability of the convolutional neural network in the early stage of the visual task and the advantage of the Transformer in extracting long-distance features, fully extract the multi-scale defect features of the weld, which is beneficial to the intelligent recognition of multi-scale defects.

[0037] (3) In the Neck network, the present invention designs a context-aware module CSM, which obtains features of different receptive fields through dilated convolution with different dilation rates, obtains spatial adaptive weights through cascaded operations, and obtains context information features through weighted summation.

[0038] (4) This invention introduces a CSM module into the deep feature fusion part of the Neck network, which enhances the ability to capture the feature information of minute defects in the weld, thereby improving the accuracy of the intelligent recognition algorithm in recognizing minute defects in the weld. Therefore, this invention improves the overall recognition accuracy of multi-scale defects.

[0039] (5) This invention utilizes deep learning technology to achieve intelligent identification of weld defects. The designed identification model can effectively extract multi-scale defect features and enhance the features of minute defects, significantly reducing the false negative rate. Therefore, this model has high identification accuracy and strong robustness, providing reliable technical support for intelligent identification technology of weld defects in industries such as aviation, aerospace, nuclear industry, and special equipment, and promoting the application and implementation of related technologies. Attached Figure Description

[0040] Figure 1 This is a network structure diagram of the intelligent identification method for weld DR images based on local-global efficient model building and improved YOLOv5, which is the implementation method of the present invention.

[0041] Figure 2 This is a schematic diagram of the iRMB_Conv module used for feature extraction in this invention;

[0042] Figure 3 This is a schematic diagram of the Context-Aware Module (CSM) for feature enhancement in this invention.

[0043] Figures 4(a) to 4(i) This is a diagram showing the training results of the present invention:

[0044] Figure 5-1 and Figure 5-2 These are all images showing the defect identification results of the present invention, wherein:

[0045] Figure 5-1 In the diagram, (a) shows the defect identification results of YOLOv5, (b) shows the defect identification results of YOLOv7, (c) shows the defect identification results of YOLOv8, (d) shows the defect identification results of YOLOv9, and (e) shows the defect identification results of the improved YOLOv5 model of this invention.

[0046] Figure 5-2 In order to express clearly Figure 5-1 The relevant data, along Figure 5-1 The diagram shows a magnified view of the area captured by the dashed line. Detailed Implementation

[0047] The design concept of the intelligent defect recognition method for weld DR images based on efficient local-global modeling and improved YOLOv5 proposed in this invention is as follows: This method designs an efficient feature extraction module, iRMB_Conv, which combines the advantages of dynamic global modeling and static local information fusion. In the Backbone network, the iRMB_Conv module can fully extract multi-scale defect features from the weld DR image, helping to improve the recognition accuracy of defects at different scales. Furthermore, in the Neck network, this invention designs a context-aware module, CSM, which uses dilated convolutions with different dilation rates to obtain features from different receptive fields, effectively enhancing the feature information of minute defects in the weld DR image, thereby improving the model's recognition accuracy for minute defects. This intelligent recognition model has high recognition accuracy and good robustness, and is suitable for industrial fields such as aviation, aerospace, nuclear industry, and special equipment, promoting the application and deployment of intelligent defect recognition technology for weld DR images.

[0048] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but the following embodiments are by no means intended to limit the present invention.

[0049] Example

[0050] This invention proposes an intelligent defect recognition method for weld seam DR images based on local-global efficient model building and improved YOLOv5, comprising the following steps:

[0051] Step 1, Data Collection: Collect DR images of welds containing different types of defects and annotate the defects;

[0052] Step 2, Dataset Creation: The labeled data is augmented using rotation, mirroring, and cropping, and the augmented labeled data is randomly divided into training set, validation set, and test set in an 8:1:1 ratio;

[0053] Step 3, Model Building: On the open-source deep learning framework PyTorch platform, build a weld defect intelligent recognition network model based on local-global efficient model building improvement YOLOv5.

[0054] Step 4, Model Training: Input the training set data into the model built in Step 3 for training, and use the validation set to validate the model. After each epoch (when the model completes a full traversal of the entire training dataset), calculate the loss and update the model parameters through backpropagation. Training stops when the preset number of epochs is reached, and the best-performing network model is saved.

[0055] Step 5, Model Testing and Inference: Evaluate the generalization performance of the network model saved in Step 4 using the test set. The main purpose of model inference is to apply the network model to weld DR images in actual engineering for intelligent defect identification.

[0056] The steps described above are explained in detail below:

[0057] Step 1: The specific steps for data collection are as follows:

[0058] 1-1) Conduct weld DR tests in accordance with NBT 47013.11-2015 (Non-destructive testing of pressure equipment - Part 11: X-ray digital imaging testing) standard. The sensitivity, resolution, grayscale value and normalized signal-to-noise ratio of the weld DR image must at least meet the requirements of AB level testing technology.

[0059] 1-2) Use ImageJ (image processing) software to adjust the window width and window level of the weld DR images acquired in the experiment, and export the images in JPG format;

[0060] 1-3) Labelme (image annotation tool) is used to annotate the five types of defects in the JPG format weld DR image and generate a JSON file corresponding to the DR image. The five types of defects are porosity, inclusion, crack, incomplete penetration and lack of fusion. PO, SL, CRK, LP and LF are used to represent the defect categories of porosity, inclusion, crack, incomplete penetration and lack of fusion.

[0061] 1-4) Convert the JSON format annotation file generated by labelme into XML format to conform to the VOC dataset specification annotation file.

[0062] Step 2, the specific steps of creating the dataset, include:

[0063] 2-1) Rotation enhancement: Randomly select a rotation angle from a predetermined set of rotation angles for data enhancement, the set of rotation angles including 90°, 180° and 270°; Rotate the DR image to be enhanced according to the selected rotation angle, and save the rotated image as a new DR image; In the rotated new image, calculate the coordinate values ​​of the defects, and update the coordinate information in the XML annotation file corresponding to the DR image to be enhanced based on the coordinate values, so as to generate a new XML annotation file corresponding to the new image;

[0064] 2-2) Mirroring enhancement includes horizontal mirroring and vertical mirroring; a mirroring method is randomly selected to perform mirroring operation on the image to be enhanced, and the image is saved as a new DR image; in the new image after the mirroring operation, the coordinate values ​​of the defects are calculated, and the coordinate information in the XML annotation file corresponding to the DR image to be enhanced is updated based on the coordinate values ​​to generate a new XML annotation file corresponding to the new image;

[0065] 2-3) Cropping enhancement: The cropping size is randomly selected within the range of 900-1200 pixels, centered on the defect center, and the DR image to be enhanced is cropped. The cropped image is then saved as a new DR image. The coordinate values ​​of the defect are calculated in the new cropped image, and the coordinate information in the XML annotation file corresponding to the DR image to be enhanced is updated based on these coordinate values ​​to generate a new XML annotation file corresponding to the new image.

[0066] 2-4) Rotation enhancement and mirror enhancement methods are applied randomly, with a 60% probability of random application. Cropping enhancement is a mandatory application method. The final data enhancement methods include four combinations: cropping enhancement, random rotation enhancement combined with cropping enhancement, random mirror enhancement combined with cropping enhancement, and random rotation enhancement combined with random mirror enhancement and cropping enhancement. The original experimental DR images and their corresponding XML annotation files are processed by the above four data enhancement methods in sequence to generate the final enhanced data.

[0067] 2-5) The enhanced VOC format dataset is randomly divided into training, validation, and test sets in an 8:1:1 ratio, and the data files for each part are saved as train.txt, val.txt, and test.txt, respectively. The weld defect categories, labels, and number of defects in each category in the dataset are shown in Table 1.

[0068] Table 1

[0069]

[0070]

[0071] Step 3: Improve the YOLOv5 weld defect intelligent identification network model based on efficient local-global model building, such as... Figure 1 As shown, where N defect N represents the number of defect categories. defect =5, specifically including the following steps:

[0072] 3-1) Construct the iRMB_Conv module for feature extraction in the Backbone network, such as... Figure 2 As shown, this module is implemented by concatenating M iRMB modules with a 3×3 convolutional layer. The core components of the iRMB module include an Expanded Window Multi-Head Self-Attention (EW-MHSA) module, a Depth-Wise Convolution (DW-Conv) with integrated short connections, a 1×1 convolutional layer, and residual connections. In the EW-MHSA module, a prerequisite judgment step is first performed to determine whether to perform self-attention. If it is determined not to perform self-attention, the input features are directly subjected to a 1×1 convolution operation with channel expansion, and the result of this convolution is used as the output of the EW-MHSA module. Conversely, if self-attention is required, the input features of the EW-MHSA module are first subjected to preliminary convolution processing. Subsequently, the convolutional features are split into matrices K and Q using the Split operator. Then, a dot product operation is performed on K and Q, and the Softmax function is applied to obtain the attention weights (Attention Map). Finally, the attention weights are weighted with the original input features of the EW-MHSA module, and then subjected to a 1×1 convolution operation with channel expansion. The result is the output feature of the EW-MHSA module. The specific expression of the iRMB module is:

[0073] Y1 = EW - MHSA(Y)

[0074] Y2=DW-Conv(Y1)+Y1

[0075]

[0076] Where Y represents the input feature of the iRMB module. The output features of the iRMB module are applied to Y2 by a 1×1 convolutional layer with channel shrinkage.

[0077] 3-2) Multiple iRMB_Conv modules are instantiated by adjusting the input channels, prerequisite settings, number of multi-head attention points, iRMB module depth M, output channels, and DW-Conv convolutional kernels. Specifically, four iRMB_Conv modules are instantiated with input channels of 24, 32, 48, and 120 respectively; prerequisite settings of No, No, Yes, and Yes respectively; number of multi-head attention points of 16, 16, 20, and 20 respectively; iRMB module depth M of 3, 3, 9, and 3 respectively; output channels of 32, 48, 120, and 200 respectively; and DW-Conv convolutional kernel sizes of 3, 3, 5, and 5 respectively. Furthermore, the output channels of the 3×3 convolutional layer in the iRMB_Conv module are the same as the module's output channels, set to 32, 48, 120, and 200 respectively. The four instantiated iRMB_Conv modules are used for feature extraction in the Backbone network, such as... Figure 1 As shown;

[0078] 3-3) Construct a context-aware module (CSM), such as Figure 3 As shown, firstly, multi-scale features F1, F2, and F3 are extracted through three convolutional layers with different dilation rates. Then, their dimensions are adjusted using 1×1 convolutions to obtain F1′, F2′, and F3′. These F1′, F2′, and F3′ are then concatenated along the channel dimension to form a new feature, where the three dilation rates are set to 1, 3, and 5, respectively. Next, the concatenated feature is processed through 1×1 convolutions and Softmax operations to generate normalized weights. The Split operator is then used to split these normalized weights into W1, W2, and W3 along the channel dimension, which are then used as adaptive weights for the multi-scale features F1, F2, and F3, respectively. Finally, these features are multiplied by their respective weights and summed to obtain the final context fusion feature, which is W1·F1 + W2·F2 + W3·F3.

[0079] 3-4) Replace the C3 module in the 13th layer of the Neck part of the YOLOv5 model with the context-aware module CSM to fully mine the contextual information of defects in the weld DR image, and provide fully fused feature information for the defect classification and bounding box regression tasks in the downstream Head part.

[0080] Step 4, model training specifically includes:

[0081] 4-1) Model hyperparameter settings, including batch size of 4, optimizer of stochastic gradient descent (SGD), initial learning rate of 0.01, total training epochs of 200, and input image size of 640×640.

[0082] 4-2) During the model training phase, the training set data is input into the model, and the weights in the model are adjusted by the optimization algorithm through backpropagation and gradient descent, so that the model can learn useful features and patterns from the data.

[0083] 4-3) The validation set is used to evaluate the model's performance and tune hyperparameters during training. At the end of each epoch, the model is evaluated on the validation set to determine whether it is overfitting or underfitting and to help select the optimal hyperparameters.

[0084] In this embodiment, the weld defect intelligent recognition network model based on the local-global efficient model building improvement YOLOv5 is trained using the training set. The training results are as follows: Figures 4(a) to 4(i) As shown, Figure 4(a) is the bounding box regression loss curve for the 0-200 Epoch training set; Figure 4(b) is the target confidence loss curve for the 0-200 Epoch training set; Figure 4(c) is the classification loss curve for the 0-200 Epoch training set; Figure 4(d) is the bounding box regression loss curve for the 0-200 Epoch validation set; Figure 4(e) is the target confidence loss curve for the 0-200 Epoch validation set; Figure 4(f) is the classification loss curve for the 0-200 Epoch validation set; Figure 4(g) is the average precision (mAP@0.5) curve for the 0-200 Epoch model; Figure 4(h) is the precision curve for the 0-200 Epoch model; and Figure 4(i) is the recall curve for the 0-200 Epoch model.

[0085] In this embodiment, the improved YOLOv5 weld defect intelligent identification network model based on efficient local-global construction is trained using the training set.

[0086] YOLOv5, YOLOv7, YOLOv8 and YOLOv9 were trained using the training set, and the models were validated on the validation set. The experimental results of different models are shown in Table 2.

[0087] Table 2

[0088] YOLOv5 86.5 87.5 83.5 YOLOv7 54.5 56.4 58.7 YOLOv8 76.7 82.8 68.5 YOLOv9 81.3 82.8 74.1 This invention model 91.5 88.3 87.3

[0089] As shown in Table 2, the optimized YOLOv5 model of this invention achieves average precision (mAP@0.5), precision, and recall of 91.5%, 88.3%, and 87.3% for weld seam DR image defect recognition, respectively, which are significantly better than other models. In particular, compared with the original YOLOv5 model, the optimized model of this invention improves the average precision (mAP@0.5), precision, and recall by 5.0%, 0.8%, and 3.8%, respectively.

[0090] Step 5, model testing and inference, mainly includes:

[0091] 5-1) Evaluate the generalization performance of the network model saved in step 4 by using a test set;

[0092] 5-2) The main purpose of model inference is to apply the network model to weld DR images in actual engineering to perform intelligent defect identification;

[0093] 5-3) Input a DR image of the weld to be identified into YOLOv5, YOLOv7, YOLOv8, YOLOv9 and the model of this invention respectively, and output the detection results, such as... Figure 5-1 and Figure 5-2 As shown, the YOLOv5 and YOLOv7 models misclassified the porosity (PO) defect on the far left of the weld as an inclusion (SL) defect, the YOLOv8 model missed detecting the leftmost PO defect, and the YOLOv9 model misclassified the low-density inclusion (SL) defect located in the middle of the weld as a PO defect. In contrast, the YOLOv5 optimized model proposed in this invention can effectively identify defects in the weld of this DR image, with accurate defect classification.

[0094] Although the present invention has been described above in conjunction with the accompanying drawings, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many improvements and changes under the guidance of the present invention without departing from the spirit of the present invention, and these improvements and changes are all within the protection scope of the present invention.

Claims

1. A method for intelligent defect recognition of weld seam DR images based on efficient local-global model building to improve YOLOv5, comprising the following steps: Step 1: Collect DR images of welds containing different types of weld defects and annotate the weld defects; Step 2: Enhance the labeled data by rotation, mirroring, and cropping, and then randomly divide the enhanced labeled data into training set, validation set, and test set in an 8:1:1 ratio; Step 3: On the PyTorch platform, construct a network model for intelligent identification of weld seam DR image defects based on the improved YOLOv5 model with efficient local-global model building. Step 4: Input the data from the training set into the model constructed in Step 3 for training, and use the validation set to validate the model; after each epoch, calculate the loss and update the model parameters through backpropagation; training stops when the preset number of epochs is reached, and the trained network model is saved; and the generalization performance of the trained network model is evaluated using the test set. Step 5: Input the DR image of the weld to be identified into the network model, and output the detection result after model inference to realize intelligent identification of weld defects; Step 1 includes the following: S1-1) Conduct weld DR tests according to NBT 47013.11-2015 standard. The sensitivity, resolution, grayscale value and normalized signal-to-noise ratio of the weld DR image must meet at least the requirements of AB level detection technology. S1-2) Use ImageJ image processing software to adjust the window width and window level of the weld seam DR image and export the image in JPG format; S1-3) The labelme image annotation tool is used to annotate five types of weld defects in the JPG format weld DR image and generate a JSON format annotation file. The five types of weld defects are porosity, inclusions, cracks, incomplete penetration and lack of fusion. PO, SL, CRK, LP and LF are used to represent the weld defect categories of porosity, inclusions, cracks, incomplete penetration and lack of fusion. S1-4) Convert the JSON format annotation file into XML format; In step 2, during the enhancement of the labeled data, the rotation enhancement angles are 90°, 180°, and 270°; the mirror enhancement includes horizontal mirroring and vertical mirroring; the cropping enhancement is centered on the defect center, and the cropping size is randomly selected within the range of 900-1200 pixels; among them, the rotation enhancement and mirror enhancement methods are applied in a random manner, and the cropping enhancement method is applied in a specific manner. Step 3 includes the following: S3-1) In the Backbone network, an iRMB_Conv module for feature extraction is constructed. This module is implemented by concatenating M iRMB modules with a 3×3 convolutional layer. The core components of the iRMB module include an extended window multi-head self-attention EW-MHSA module, a depthwise segregating convolution DW-Conv with integrated short connections, a 1×1 convolutional layer, and residual connections. In the EW-MHSA module, a prerequisite judgment step is first performed to determine whether to perform the self-attention operation. If it is determined that Self-Attention is not to be performed, the input features are directly subjected to a 1×1 convolution operation with one channel expansion, and the result of this convolution is used as the output of the EW-MHSA module. Conversely, if Self-Attention is required, the input features of the EW-MHSA module are first subjected to preliminary convolution processing. Subsequently, the convolutional features are split into matrices using the Split operator. sum matrix Next, regarding and The dot product operation is performed, and the Softmax function is applied to obtain the attention weights. Finally, the attention weights are weighted with the original input features of the EW-MHSA module, and then subjected to a 1×1 convolution operation with channel expansion. The result is the output feature of the EW-MHSA module. The specific expression of the iRMB module is as follows: in, For the input characteristics of the iRMB module, The output characteristics of the iRMB module, acting on The layer above is a 1×1 convolutional layer with channel shrinkage; S3-2) Instantiate multiple iRMB_Conv modules by setting input channels, prerequisite judgment settings, number of multi-head attention, iRMB module depth M, output channels, and DW-Conv convolution kernels; the Backbone network contains 4 to 5 instantiated iRMB_Conv modules for feature extraction. S3-3) Construct a context-aware module (CSM), including extracting multi-scale features through three convolutional layers with different dilation rates. , and Adjust their dimensions using 1×1 convolutions to obtain... , and and along the channel dimension direction , and The concatenated features are then combined to form a new feature. The combined feature is then subjected to 1×1 convolution and Softmax operations to generate normalized weights, which are then split into three parts along the channel dimension, serving as multi-scale features. , and The adaptive weights; finally, the multi-scale features... , and After multiplying each feature by its respective weight and summing them, the final context fusion feature is the output feature of the context-aware module CSM. (S3-4) Replace the 13th layer C3 module in the Neck part with the aforementioned context-aware module CSM, thereby providing fully fused feature information for the defect classification and bounding box regression tasks in the downstream Head part.

2. The intelligent defect recognition method for weld seam DR images according to claim 1, characterized in that, In step S3-2), the Backbone network contains four instantiated iRMB_Conv modules; the input channels of the four iRMB_Conv modules are 24, 32, 48 and 120 respectively, the prerequisite judgment settings are No, No, Yes and Yes respectively, the number of multi-head attention is 16, 16, 20 and 20 respectively, the iRMB module depth M is 3, 3, 9 and 3 respectively, the output channels of the four iRMB_Conv modules are 32, 48, 120 and 200 respectively, and the depth separation convolution DW-Conv convolution kernel size is 3, 3, 5 and 5 respectively; The output channels of the 3×3 convolutional layer in the iRMB_Conv module are the same as the output channels of the module, and are set to 32, 48, 120 and 200 respectively.

3. The intelligent defect recognition method for weld seam DR images according to claim 1, characterized in that, Step 4 includes the following: S4-1) Set the hyperparameters, including: batch size of 4, optimizer of stochastic gradient descent (SGD), initial learning rate of 0.01, total training epochs of 200, and input image size of 640×640. S4-2) During the model training phase, the training set data is input into the model, and the weights in the model are adjusted by the optimization algorithm through backpropagation and gradient descent, so that the model can learn useful features and patterns from the data. S4-3) The validation set is used to evaluate the model's performance and tune hyperparameters during training; at the end of each epoch, the model is evaluated on the validation set to determine whether the model is overfitting or underfitting, thereby selecting the optimal hyperparameters.

Citation Information

Patent Citations

  • Improved lightweight weld defect recognition algorithm based on YOLOv5

    CN117095234A

  • Automatic driving small target detection model based on YOLOv8

    CN117437407A