Welding defect detection method based on lightweight RT-DETR model

By reducing the lightweight RT-DETR model and replacing Bottleneck with SMPConv-CGLU and FPN with DRFD, the high cost and low efficiency problems of traditional welding defect detection are solved, and real-time detection with high accuracy and low complexity is achieved.

CN120387991APending Publication Date: 2025-07-29SOUTHWEST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510466070.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

Traditional welding defect detection methods rely on high-cost professional equipment, high model complexity and slow inference speed, limited feature expression capabilities, and downsampling information loss affects detection accuracy.

Method used

The lightweight RT-DETR model is adopted, the Bottleneck that replaces the C2f module is SMPConv-CGLU composite structure, the downsampling module that replaces the FPN module is DRFD, and the data set is built using ordinary camera mechanisms and trained and verified.

Benefits of technology

It reduces equipment costs, improves detection accuracy and inference speed, meets the real-time inspection needs of industrial production lines, and is suitable for small and medium-sized welding and assembly workshops.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387991A_ABST
    Figure CN120387991A_ABST
Patent Text Reader

Abstract

The invention discloses a welding defect detection method based on a lightweight RT-DETR model, and relates to the technical field of computer vision and industrial detection, and the welding defect detection method comprises the following steps: constructing a data set by using welding defect images collected by a common camera; the Bottleneck of a C2f module in the RT-DETR model is replaced with an SMPConv-CGLU composite structure, and the SMPConv-CGLU composite structure is used as a model of the RT-DETR model; a down-sampling module of an FPN module in the RT-DETR model is replaced with a multi-modal feature down-sampling module DRFD, and a lightweight RT-DETR model is obtained; and training, testing and verifying the lightweight RT-DETR model by using the constructed data set, and applying the lightweight RT-DETR model to an actual scene. According to the invention, high-precision and low-complexity welding defect real-time detection is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and industrial inspection technology, and particularly to a welding defect detection method based on a lightweight RT-DETR model. Background Art

[0002] The traditional welding defect detection methods have the following problems:

[0003] They usually rely on professional detection equipment such as HDR images, radiographs or X-ray images. Although these devices are accurate, they are costly;

[0004] The model has a high complexity and a slow inference speed. It is necessary to improve the model to be lightweight, reduce the amount of calculation and improve the inference speed. The mainstream detection model (such as RT-DETR-R18) has 19.88M parameters and 57GFLOPs of computational volume, which is difficult to meet the deployment requirements of embedded devices;

[0005] The feature expression ability is limited: the traditional Bottleneck module uses fixed convolution kernel sampling points and cannot adapt to the multi-scale geometric features of welding defects;

[0006] Downsampling information loss: Conventional stride convolution or pooling operations result in the loss of tiny defect features, affecting the detection accuracy. Summary of the Invention

[0007] Aiming at the above deficiencies in the prior art, the welding defect detection method based on a lightweight RT-DETR model provided by the present invention realizes real-time detection of welding defects with high precision and low complexity.

[0008] In order to achieve the above invention purpose, the technical solution adopted by the present invention is: A welding defect detection method based on a lightweight RT-DETR model, comprising the following steps:

[0009] S1: Construct a data set by using the welding defect images collected by an ordinary camera;

[0010] S2: Replace the Bottleneck in the C2f module of the RT-DETR model with an SMPConv-CGLU composite structure;

[0011] S3: Replace the downsampling module in the FPN module of the RT-DETR model with a multi-modal feature downsampling module DRFD to obtain a lightweight RT-DETR model;

[0012] S4: Train, test and verify the lightweight RT-DETR model by using the constructed data set, and apply the lightweight RT-DETR model to the actual scenario.

[0013] Further, in S2, the SMPConv-CGLU composite structure includes a dynamic sampling point offset module SMPConv and a gated feature screening module CGLU connected in sequence.

[0014] Further, the continuous kernel function of the dynamic sampling point offset module SMPConv is:

[0015]

[0016] where SMP(·) is the kernel function, x is the query point, φ is the set of learnable parameters, φ = {(p i ), {w i}, {r i}}, N(·) is the set of neighbor points related to the query point x, g(·) is the distance function, p i is the coordinate of the self-moving point, r i is the radius of each self-moving point, w i is the weight of the self-moving point;

[0017]

[0018] where |·|1 represents the L1 distance;

[0019] Convolve the input function f in the continuous domain, and the formula is:

[0020]

[0021] where c is the channel index, N c is the total number of channels, R is the set of real numbers, f c (·) is the function corresponding to the channel index c, SMP c (·) is the kernel function corresponding to the channel index c, and τ is the integration variable.

[0022] Further, the gated feature screening module CGLU performs the following operations:

[0023] A1: Generate a feature map by convolving the input feature with a 1x1 convolution, and divide the feature map into two parts, v and x, in the channel dimension;

[0024] A2: Perform local feature extraction on x through depthwise separable convolution, and use the GLU gating mechanism to multiply x and v element-wise to obtain an activated feature map;

[0025] A3: Perform Dropout processing on the activated feature map, x passes through a 1x1 convolution to output the final feature map, and add the final feature map to the original input feature to form a residual connection.

[0026] Furthermore, the multi-modal feature downsampling module DRFD in S3 includes parallel Dcut branch, Dconv branch, and Dmax branch;

[0027] The Dcut branch realizes lightweight dimensionality reduction through channel pruning and average pooling;

[0028] The Dconv branch uses depthwise separable convolution to extract local detail features;

[0029] The Dmax branch retains significant features through 2×2 max pooling.

[0030] Furthermore, downsampling is performed using the Dcut branch, Dconv branch, and Dmax branch, and the formula is:

[0031] x1 = GELU(D conv (Y))

[0032] x2 = D cut (Y)

[0033] x3 = D max (Y)

[0034] where x1, x2, and x3 are the output feature maps obtained through the Dconv branch, Dcut branch, and Dmax branch respectively, GELU(·) is an activation function based on the Gaussian error function, D conv (·), D cut (·), and D max (·) represent the Dconv branch, Dcut branch, and Dmax branch respectively, and Y is the input feature map;

[0035] The three-way feature channels are concatenated through a 1×1 convolutional feature fusion layer, and the formula is:

[0036] X = fusion(x1, x2, x3)

[0037] where X is the concatenated feature map, and fusion(·) represents the feature fusion operation.

[0038] The beneficial effects of the present invention are as follows: By replacing the Bottleneck in the original C2f module with the SMPConv-CGLU composite structure (C2f_SGConv) and combining the dual-branch residual fusion downsampling module (DRFD), the present invention significantly reduces the model parameter quantity and computational complexity, and improves the deployment efficiency; SMPConv enhances the multi-scale feature expression ability for irregular defects such as weld cracks and pores by dynamically adjusting the sampling point positions of the convolution kernels, and the detection accuracy (mAP50) is improved; the DRFD module reduces the downsampling information loss through multi-branch feature fusion, and the inference speed reaches 64.9 FPS, meeting the real-time detection requirements of industrial production lines; it supports the input of ordinary cameras, reduces the equipment cost by 90%, solves the problem that traditional methods rely on high-cost professional equipment (such as HDR, X-rays), and at the same time overcomes the defects of large model parameter quantities and insufficient detection accuracy for tiny defects in existing models, and is applicable to small and medium-sized welding workshops. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 FIG. is a flowchart of a welding defect detection method based on a lightweight RT-DETR model according to the present invention.

[0040] Figure 2 FIG. is a structural diagram of a lightweight RT-DETR model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] The following further describes the present invention in conjunction with the drawings and specific embodiments.

[0042] As Figure 1 and Figure 2 shown, a welding defect detection method based on a lightweight RT-DETR model is characterized by including the following steps:

[0043] S1: Construct a data set using the welding defect images collected by an ordinary camera;

[0044] In this embodiment, the data set is intended to be used for training or testing a computer vision model to detect weld defects using traditional camera images (not HDR, radiographic, or X-ray images). The data set is roughly distributed as 85% for training, 10% for validation, and 5% for testing. The data set format is in the yolo format, and the welding defects are divided into five categories, including adjacent defects (spatter, local melting, or surface damage), integrity defects (slag inclusions, blisters, burn-through, etc.), geometric defects (lack of fusion, surface irregularities, etc.), post-processing defects (burrs, scratches, dents, etc.), and undelivered defects (lack of fusion).

[0045] By using traditional camera images to construct the dataset (replacing HDR, ray, or X-ray images), the cost of data acquisition equipment is significantly reduced (by more than 90%), while ensuring the universality of the model in industrial scenarios; the dataset is reasonably divided into 85% for training, 10% for validation, and 5% for testing, taking into account the sufficiency of model training and the reliability of generalization ability verification; in view of the diversity of welding defects, five types of defect labels, namely adjacent defects, integrity defects, geometric defects, post-processing defects, and lack of fusion, are defined to cover the core issues of welding quality inspection.

[0046] S2: Replace the Bottleneck of the C2f module in the RT-DETR model with the SMPConv-CGLU composite structure;

[0047] The reference model rtdetr-r18 of RT-DETR has 19,878,180 parameters, a computational complexity of 57 GFLOPs, an FPS of 57.8, and an mAP50 of 80.1%. In this study, in view of the problems of large number of parameters and high computational complexity existing in the deployment of the RT-DETR object detection model, a lightweight improvement scheme is proposed. By replacing the original Bottleneck of the C2f module with the SGConv (SMPConv + CGLU) composite structure (SMPConv realizes the adaptive offset of the sampling points of the convolution kernel, and CGLU filters the effective feature channels through the gating mechanism), the feature expression ability is improved while reducing the number of parameters; combined with the Backbone channel compression strategy, a lightweight detection network is constructed. Verified by the dataset, on the premise that the accuracy of the improved model is increased to 81.4% mAP, the computational complexity is reduced to 42.6 GFLOPs (a decrease of 25.3%), and the inference speed is increased by 10%.

[0048] The SMPConv-CGLU composite structure in S2 includes a dynamic sampling point offset module SMPConv and a gating feature screening module CGLU connected in sequence.

[0049] SMPConv allows the sampling points of the convolutional kernel to dynamically adjust their positions according to the input features, so as to adaptively weld the local morphology of defects. For example, for irregular defects such as cracks and pores, the sampling points can automatically focus on the defect edges or abnormal areas, enhancing the pertinence of feature expression. By adjusting the radius of the points, SMPConv can flexibly adapt to defects of different sizes. Small defects can achieve high-precision positioning by reducing the radius, while large defects can capture the global context by expanding the radius. Welding defects often manifest as tiny or local anomalies (such as welding slag and lack of fusion). By optimizing the positions of the sampling points, SMPConv makes the convolutional kernel more focused on the key areas, reduces the interference of background noise, and improves the feature discriminability. SMPConv replaces the dense parameters of the traditional large-kernel convolution with a small number of dynamic points, reducing model redundancy. For example, in the backbone network of RT-DETR, SMPConv can reduce the computational cost while maintaining support for a large receptive field.

[0050] The continuous kernel function of the dynamic sampling point offset module SMPConv is as follows:

[0051]

[0052] where SMP(·) is the kernel function, x is the query point, representing the data point to be processed, the relative coordinates of the area covered by the convolutional kernel, φ is a set of learnable parameters, φ = {(p i ), {w i}, {r i}}, N(·) is the set of neighbor points related to the query point x, and the neighbor points are determined according to the distance function g(x, p i , r i ), that is, only those points that are very close to x will be selected as neighbors, g(·) is the distance function, used to measure the relationship between the query point x and the self-moving point p i , specifically, g calculates the weighted distance between point x and point p i , and this distance determines the contribution of p i to the query point. p i is the coordinate of the self-moving point, representing the position of each point in the convolutional kernel, and the coordinates of these points are learnable. r i is the radius of each self-moving point, representing the influence range of the point, that is, the "distance" limit of its influence on the query point. w i is the weight of the self-moving point, representing the "importance" of each point, and it is a vector;

[0053] The formula can be understood as follows: SMP(x; φ) is a mapping function from the input coordinate space (image position) to the output kernel vector. It generates the corresponding output value at any query point x by dynamically adjusting the contributions of neighboring points. The final output is generated by taking the weighted average of a series of neighboring points around the query point x. The weights are controlled by the distance function g(x, p i , r i ), and the closer the point is, the greater its contribution to the result. The calculated result is a value between 0 and 1, representing the influence of point p i on the query point x. The closer the distance, the greater the g value and the greater the contribution to the query point. The distance formula is as follows:

[0054]

[0055] where |·|1 represents the L1 distance, that is, the Manhattan distance between the query point x and the self-moving point p i . The L1 distance is the sum of the absolute values of the coordinate differences.

[0056] Convolving the input function f over the continuous domain, the formula is:

[0057]

[0058] where c is the channel index, N c is the total number of channels, R is the set of real numbers, f c (·) is the function corresponding to the channel index c, SMP c (·) is the kernel function corresponding to the channel index c, and τ is the integration variable.

[0059] Each channel has its corresponding function f c and SMP c . The final summation represents the accumulation of the convolution results of all channels. That is, the convolution of each channel is calculated separately first, and then these results are added together to obtain the final result.

[0060] The core algorithm of the ConvolutionalGLU module is to combine the convolution operation with the GLU gating mechanism, efficiently extract local features using depthwise separable convolution, and dynamically adjust the activation of features through the gating mechanism. In addition, residual connections and Dropout regularization enhance the training stability and generalization ability of the model. First, the input data passes through the fc1 linear layer, is divided into two tensors x and v, x passes through the custom depth convolution layer DWConv, then through the activation function, multiplies the output by v, performs the Dropout operation, the result passes through the fc2 linear layer, and finally performs the Dropout operation again. However, due to processing 4D tensors instead of 2D tensors in the source code, some adjustments are made to the source code.

[0061] Replace the linear layer with a convolutional layer: The convolutional layer (nn.Conv2d) is used, which is suitable for input data with spatial information (image data in 4D tensors). The convolutional layer can preserve the spatial structure of the input data, such as the height and width information in an image, while the linear layer flattens this spatial information. Therefore, the modified code is more suitable for processing images or other input data with spatial dimensions.

[0062] Standard nn.Conv2d is used for depthwise convolution operations, and activation functions are added. This improves the readability and maintainability of the code.

[0063] Residual connections are used, adding the input x_shortcut to the processed output. Residual connections can help alleviate the vanishing gradient problem in deep neural networks, making the network easier to train and improving the performance and stability of the model.

[0064] Specifically, first, a 1x1 convolution is performed to generate a feature map of size hidden_features*2. Then, the feature map is split into two parts in the channel dimension: x and v. x undergoes local feature extraction through depthwise separable convolution dwconv. This convolution operation is performed channel by channel, so it can effectively capture local information for each channel. Using the GLU gating mechanism, x and v are multiplied element-wise to form the final activated feature map. Here, v serves as the gating signal, determining which parts of the features need to be retained. Dropout is applied to the feature map to enhance generalization ability and prevent overfitting. x then passes through a 1x1 convolution to output the final feature map. The final output is added to the original input feature to form a residual connection.

[0065] The gating feature screening module CGLU performs the following operations:

[0066] A1: Pass the input feature through a 1x1 convolution to generate a feature map, and divide the feature map into two parts, v and x, in the channel dimension;

[0067] A2: Perform local feature extraction on x through depthwise separable convolution and use the GLU gating mechanism to multiply x and v element-wise to obtain the activated feature map;

[0068] A3: Apply Dropout to the activated feature map, x then passes through a 1x1 convolution to output the final feature map, and the final feature map is added to the original input feature to form a residual connection.

[0069] Combine the above two algorithms and replace the Bottleneck module in the original C2f with the name C2f_SGConv. In the specific operation, first input the feature map x, create a shortcut to save a reference to the input x for residual connection. Process x through the SMPConv module, then process x through the CGLU module, and perform dropout through the drop_path module. Add the shortcut and the processed x and return the result.

[0070] Through the above improvements, the model is lightweighted, and the detection speed and efficiency are optimized. The local feature extraction ability is improved, the adaptability of the model to diverse data is enhanced, and the flexibility of information screening and transmission is improved.

[0071] S3: Replace the downsampling module of the FPN module in the RT-DETR model with the multi-modal feature downsampling module DRFD to obtain a lightweight RT-DETR model;

[0072] By introducing a multi-path parallel downsampling architecture (the Dcut branch realizes channel pruning and dimensionality reduction, the Dconv branch uses grouped convolution to extract local details, and the Dmax branch retains significant features), combined with a dynamic weight fusion mechanism (1×1 convolution adaptively integrates multi-source features), the feature expression ability of multi-scale targets is significantly improved. Compared with the RT-DETR model with C2f_SGConv added, the improved model has a 1.9% reduction in the number of parameters, a 2.0% increase in the inference speed, and a 2.6% increase in mAP50 (0.84 vs 0.814) while keeping the computational cost basically the same (42.1 GFLOPs vs the original 42.6 GFLOPs). Experiments show that the DRFD module, through the channel reduction strategy and the heterogeneous feature complementary mechanism, enhances the ability to retain small target features while reducing redundant calculations, providing a better accuracy-efficiency balance for real-time detection tasks.

[0073] The multi-modal feature downsampling module DRFD in S3 includes parallel Dcut branch, Dconv branch and Dmax branch;

[0074] The Dcut branch realizes lightweight dimensionality reduction through channel pruning (retaining the first 50% of channels) and average pooling;

[0075] The Dconv branch uses depthwise separable convolution (3×3 grouped convolution → 3×3 stride 2 convolution) to extract local detail features;

[0076] The Dmax branch retains significant features through 2×2 max pooling.

[0077] The input feature map Y is downsampled through Dconv, Dcut, and Dmax to obtain output feature maps x1, x2, and x3. x1 is obtained through Dconv, and at the same time, the Gaussian Error Linear Unit (GELUs) activation function is used, and the number of channels is increased from C to 2C. x2 is downsampled through Dcut, and the number of channels is increased from C to 2C. x3 is downsampled through Dmax, and the number of channels is increased from C to 2C.

[0078] The Dcut branch, Dconv branch, and Dmax branch are used for downsampling, and the formula is:

[0079] x1 = GELU(D conv (Y))

[0080] x2 = D cut (Y)

[0081] x3 = D max (Y)

[0082] Among them, x1, x2, and x3 are the output feature maps obtained through the Dconv branch, Dcut branch, and Dmax branch respectively. GELU(·) is the activation function based on the Gaussian error function. D conv (·), D cut (·), and D max (·) represent the Dconv branch, Dcut branch, and Dmax branch respectively. Y is the input feature map;

[0083] The three-way feature channels are concatenated through a 1×1 convolutional feature fusion layer, and the formula is:

[0084] X = fusion(x1, x2, x3)

[0085] Among them, X is the concatenated feature map, and fusion(·) represents the feature fusion operation, that is, the channels of multiple input feature maps are weighted and summed through 1×1 convolution, so as to fuse features from different sources or scales into a unified representation.

[0086] By concatenating x1, x2, and x3, the number of channels of the feature map is increased from C to 6C. Subsequently, a 1×1 convolutional fusion layer is designed. After concatenating the three-way feature channels, it is reduced from 6C to 2C.

[0087] By designing a cross-level feature interaction mechanism and a three-path dynamic fusion downsampling mechanism, the fusion of low-level details and high-level semantics is enhanced. Multi-path feature fusion enhances information retention, fuses the features of three branches, and combines the advantages of different downsampling methods.

[0088] S4: Train, test, and validate the lightweight RT-DETR model using the constructed dataset, and apply the lightweight RT-DETR model to the actual scenario.

[0089] The present invention provides a welding defect detection method based on a lightweight RT-DETR model. First, a dataset is constructed using images collected by an ordinary camera instead of expensive devices such as X-rays, which can reduce the cost of welding defect detection for enterprises. At the same time, the number of parameters of the model is reduced, and the inference speed of the model is improved to meet the real-time requirements of industrial detection. Most importantly, the detection accuracy of the model is improved.

[0090] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the principles of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the invention.

Claims

1. A welding defect detection method based on a lightweight RT-DETR model, characterized in that, It includes the following steps: S1: Construct a dataset using the welding defect images collected by an ordinary camera; S2: Replace the Bottleneck in the C2f module of the RT-DETR model with the SMPConv-CGLU composite structure; S3: Replace the downsampling module in the FPN module of the RT-DETR model with the multi-modal feature downsampling module DRFD to obtain a lightweight RT-DETR model; S4: Use the constructed dataset to train, test, and validate the lightweight RT-DETR model, and apply the lightweight RT-DETR model to the actual scenario.

2. The welding defect detection method based on the lightweight RT-DETR model according to claim 1, wherein, The SMPConv-CGLU composite structure in S2 includes a dynamic sampling point offset module SMPConv and a gated feature screening module CGLU connected in sequence.

3. The welding defect detection method based on the lightweight RT-DETR model according to claim 2, wherein, The continuous kernel function of the dynamic sampling point offset module SMPConv is: Among them, SMP(·) is the kernel function, x is the query point, φ is the set of learnable parameters, φ = {(p i ), {w i}, {r i}}, N(·) is the set of neighbor points related to the query point x, g(·) is the distance function, p i is the coordinate of the self-moving point, r i is the radius of each self-moving point, w i is the weight of the self-moving point; where, |·|1 represents the L1 distance; Convolve the input function f in the continuous domain, and the formula is: Among them, c is the channel index, N c is the total number of channels, R is the set of real numbers, f c (·) is the function corresponding to the channel index c, SMP c (·) is the kernel function corresponding to the channel index c, and τ is the integration variable.

4. The welding defect detection method based on the lightweight RT-DETR model according to claim 2, characterized in that, The gated feature screening module CGLU performs the following operations: A1: Generate a feature map by convolving the input feature with a 1x1 convolution, and divide the feature map into two parts, v and x, in the channel dimension; A2: Perform local feature extraction on x through depthwise separable convolution, and use the GLU gating mechanism to multiply x and v element by element to obtain an activation feature map; A3: Perform Dropout processing on the activation feature map, x is then convolved with a 1x1 convolution to output the final feature map, and the final feature map is added to the original input feature to form a residual connection.

5. The welding defect detection method based on the lightweight RT-DETR model according to claim 1, wherein, The multi-modal feature downsampling module DRFD in S3 includes parallel Dcut branch, Dconv branch, and Dmax branch; The Dcut branch realizes lightweight dimensionality reduction through channel cropping and average pooling; The Dconv branch uses depthwise separable convolution to extract local detail features; The Dmax branch retains significant features through 2×2 max pooling.

6. The welding defect detection method based on the lightweight RT-DETR model according to claim 5, characterized in that Use the Dcut branch, Dconv branch, and Dmax branch for downsampling, and the formula is: x1 = GELU(D conv (Y)) x2 = D cut (Y) x3 = D max (Y) Among them, x1, x2, and x3 are the output feature maps obtained through the Dconv branch, Dcut branch, and Dmax branch respectively, GELU(·) is an activation function based on the Gaussian error function, D conv (·), D cut (·), and D max (·) represent the Dconv branch, Dcut branch, and Dmax branch respectively, and Y is the input feature map; Concatenate the three-way feature channels through a 1×1 convolution feature fusion layer, and the formula is: X = fusion(x1, x2, x3) where, X is the concatenated feature map, and fusion(·) represents the feature fusion operation.