Welding defect detection method based on deep learning
By improving the YOLOv11 model, a multi-scale convolution attention module and a full-dimensional dynamic convolution module are introduced, combined with DBSCAN clustering and MPDIoU loss function, the problem of insufficient detection accuracy and speed in welding defect detection is solved, and the performance and applicability of welding defect detection is improved.
Patent Information
- Application Number
- CN202510207099.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-07-22
AI Technical Summary
The existing welding defect detection methods have low detection accuracy and poor robustness. Traditional image processing methods are difficult to effectively extract features in the face of complex welding images. Deep learning models are prone to lose edge information when detecting small targets. The existing deep learning algorithms lack detection speed and accuracy in welding defect detection.
The improved YOLOv11 model is adopted, and a multi-scale convolution attention module and a full-dimensional dynamic convolution module are introduced. The anchor box ratio is adjusted in combination with the DBSCAN clustering algorithm, and the model is optimized for training using MPDIoU loss function to improve the model's ability to detect welding defects with small sizes and relatively large lengths and widths.
It improves the accuracy and speed of welding defect detection, especially the detection performance of multi-scale defects in complex contexts, enhances the applicability and stability of the model, is suitable for the shape and size changes of weld defects, and has faster convergence speed and higher accuracy when dealing with overlapping or non-overlapping targets.
Smart Images

Figure CN120355645A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a welding defect detection method, specifically a welding defect detection method based on deep learning, belonging to the field of industrial inspection and automation technology. Background Art
[0002] Welding plays a key role in many industrial fields, and its quality is directly related to product performance and safety. Traditional welding defect detection methods, such as visual inspection, non-destructive testing (radiographic testing, ultrasonic testing, magnetic particle testing, etc.) and destructive testing, have many limitations. For example, visual inspection is greatly affected by human factors, and each non-destructive testing method has its own deficiencies. For example, radiographic testing has high costs and radiation risks, ultrasonic testing is complex to operate, and magnetic particle testing has a limited scope of application; destructive testing is highly destructive and cannot be fully inspected.
[0003] With the development of technology, detection methods based on image analysis have emerged. However, traditional image processing is unable to cope with complex welding images, with difficulties in feature extraction and poor detection accuracy and robustness. The emergence of deep learning technology provides a new solution for welding defect detection. Deep learning models have powerful feature learning capabilities and can automatically learn rich and discriminative feature representations from a large amount of welding image data without the need for manual design of complex feature extraction algorithms. By constructing a deep learning model, efficient and accurate detection of welding defects can be achieved, improving the automation level and reliability of welding quality detection.
[0004] In the prior art, 1) as disclosed in the patent with publication number CN112101091A, a deep learning-based automatic welding defect detection method and system classifies and locates welding defects by training a deep learning model. The system includes an image acquisition module, a preprocessing module, a deep learning model, and a defect recognition module. This detection method focuses more on the application of the CNN model; 2) as disclosed in the patent with publication number CN113034399A, a deep learning-based welding defect detection and classification method extracts features from welding images using pre-trained models (such as ResNet, VGG), and improves the generalization ability of the model through data augmentation technology. By introducing transfer learning and data augmentation technology, it is suitable for small sample datasets; 3) as disclosed in the patent with publication number US20210110674A1, a deep learning-based real-time welding defect detection system uses a lightweight deep learning model (such as MobileNet) to achieve real-time detection on embedded devices. The system includes an image acquisition, model inference, and defect alarm module. The highlight of this detection system lies in the real-time performance and the application of lightweight models, which is suitable for industrial field deployment, while the method in the present invention may focus more on offline detection.
[0005] More and more scholars and engineers have begun to introduce deep learning algorithms into the field of defect detection. In terms of implementation methods, object detection algorithms are mainly divided into two major schools:
[0006] One is the two-stage method, which divides object detection into two stages: generating candidate boxes and recognizing objects within the boxes. Represented by the RCNN series of networks, such as RCNN, Faster-RCNN, and Mask-RCNN. However, RCNN uses multi-stage training, consuming huge amounts of computing and storage resources. Although Faster RCNN improves accuracy with the help of the Region Proposal Network (RPNs), its speed is still not satisfactory.
[0007] The other is the one-stage method, which integrates the detection process and directly outputs the detection results. Algorithms such as SSD, YOLOv8, and YOLOv10 belong to this category. Due to their fast detection speed and high precision, these algorithms are widely used in real-time industrial detection. However, when facing small target defect detection, they are prone to losing edge information, resulting in poor overall prediction effects. Summary of the Invention
[0008] The purpose of the present invention is to provide a welding defect detection method based on deep learning to solve at least one of the above technical problems.
[0009] The present invention realizes the above purpose through the following technical solutions: A welding defect detection method based on deep learning, the welding defect detection method includes the following steps:
[0010] S1. Collect welding defect images;
[0011] S2. Perform image preprocessing on the collected welding defect images, and then classify and label the welding defects in the welding defect images to form a data set for model training;
[0012] S3. Divide the data set according to the ratio of training set: validation set: test set = 8:1:1;
[0013] S4. Build a weld defect detection model based on the improved YOLOv11;
[0014] S5. Use the training set to train the weld defect detection model, and use the validation set for verification. Calculate the gradient of the loss function and loss value of each training of the weld defect detection model with respect to the parameters of the weld defect detection model, and update the parameters of the weld defect detection model according to the gradient and hyperparameters such as the learning rate.
[0015] S6. Determine whether the performance of the weld defect detection model meets the expectation or reaches the maximum number of training times. If it does, proceed to step S7. If not, adjust the hyperparameters of the weld defect detection model and then repeat step S5 to train the weld defect detection model.
[0016] S7. Output the trained weld defect detection model and use the validation set to evaluate the performance of the trained weld defect detection model. The evaluation metrics are AP, mAP, F1–score, and processing speed (FPS).
[0017] As a further technical solution of the present invention: The specific steps for constructing a weld defect detection model based on the improved YOLOv11 include:
[0018] S41. Introduce a multi-scale convolutional attention module into the Backbone part of the YOLOv11 object detection model.
[0019] S42: Introduce a full-dimensional dynamic convolutional module into the Neck part of the YOLOv11 object detection model and perform this operation before each feature fusion to improve the multi-scale defect detection performance of the model in complex backgrounds.
[0020] S43: Modify the preset anchor box ratios of the YOLOv11 object detection model. The new anchor box ratios are generated by using the DBSCAN clustering algorithm on the dataset.
[0021] As a further technical solution of the present invention: The multi-scale convolutional attention module consists of three parts: depthwise separable convolution, multi-branch depthwise strip convolution, and 1×1 convolution.
[0022] Among them, the expression of the multi-scale convolutional attention module is as follows:
[0023]
[0024] In the formula: F represents the input feature; Att and Out represent the attention map and the output feature respectively; represents the element-wise matrix multiplication operation; DW-Conv represents the depthwise separable convolution, Scale i represents the i-th branch, i ∈ {0, 1, 2, 3}; Scale0 is the identity connection.
[0025] As a further technical solution of the present invention: The depthwise separable convolution separates the spatial and channel dimensions of the standard convolution.
[0026] As a further technical solution of the present invention: the multi-branch depthwise separable convolution includes multiple branches, and each branch uses depthwise separable convolutions of different sizes to extract multi-scale features; each branch includes two depthwise separable convolutions, which are respectively used for feature extraction in the horizontal and vertical directions.
[0027] As a further technical solution of the present invention: the 1×1 convolution is used to simulate the relationship between different channels and generate attention weights. The attention weights re-weight the input features through element-wise multiplication, thereby implementing the spatial attention mechanism, and the output of the 1×1 convolution is directly used as the attention weights.
[0028] As a further technical solution of the present invention: the full-dimensional dynamic convolution module includes GAP, FC, ReLU, Sigmoid, and Softmax, where GAP represents the global average pooling layer, FC represents the fully connected layer, and ReLU, Sigmoid, and Softmax are activation functions;
[0029] The expression of the full-dimensional dynamic convolution module is as follows:
[0030] y = (α w1 ⊙α f1 ⊙α c1 ⊙α s1 ⊙W1 + … + α wn ⊙α fn ⊙α cn ⊙α sn ⊙W n ) * x
[0031] In the formula: x is the input feature, y is the output feature; α wi represents the attention scalar of the convolution kernel, α fi represents the attention scalar of the output channel, α ci represents the attention scalar of the input channel, α si represents the attention scalar of the convolution parameter at each position, W i is the convolution kernel, ⊙ represents element-wise multiplication, and * represents the convolution operation;
[0032] The full-dimensional dynamic convolution module introduces a multi-dimensional attention mechanism in four dimensions: the spatial size of the convolution kernel, the input channel, the output channel, and the number of convolution kernels.
[0033] As a further technical solution of the present invention: the DBSCAN clustering algorithm automatically identifies clusters and noise points in the data, calculates the central ratio of each cluster, and uses these central ratios closer to the actual distribution of the data as the new anchor box ratios to improve the performance of the object detection model;
[0034] The steps of using the DBSCAN clustering algorithm to modify the anchor box ratios are as follows:
[0035] First, extract the aspect ratios of all true bounding boxes from the dataset, and represent these aspect ratios as data points (w i , h i ), where w i and h i are the width and height of the bounding box respectively;
[0036] Secondly, apply the DBSCAN algorithm. By setting the neighborhood radius ε and the minimum number of points MinPts parameters, identify the core objects, that is, the points that satisfy that there are at least MinPts points in their ε neighborhood;
[0037] Then, starting from these core objects, cluster the density-reachable points into clusters. That is, if point q is in the ε neighborhood of point p and p is a core object, then q is directly density-reachable from p; if there exists a point chain p1, p2, …, p n , where p1 = p, p n = q, and pi+1 is directly density-reachable from pi, then q is density-reachable from p;
[0038] The expression is:
[0039]
[0040] such that
[0041]
[0042] the above formula holds, then q is density-reachable from p; for each cluster, calculate the average aspect ratio of all its points, and these average aspect ratios are the generated anchor box ratios;
[0043] Finally, collect the average aspect ratios of all clusters to form new anchor box ratios.
[0044] As a further technical solution of the present invention: use the MPDIoU loss function to calculate the loss value. In welding defect detection, MPDIoU combines the traditional IoU and the penalty term of the minimum point distance.
[0045] As a further technical solution of the present invention: The calculation formula of MPDIOU is as follows:
[0046]
[0047] In the formula: is the traditional intersection over union, d1 and d2 are the Euclidean distances between the upper left corner points and the lower right corner points of the predicted box and the true box respectively, and w and h are the average values of the widths and heights of the predicted box and the true box.
[0048] The beneficial effects of the present invention are:
[0049] 1) Based on YOLOv11, the present invention introduces a multi-scale convolutional attention module into the backbone network, enhancing the model's detection ability for small-sized and large aspect ratio welding defects. A full-dimensional dynamic convolutional module is introduced in the Neck layer. This module introduces a multi-dimensional attention mechanism in the four dimensions of the convolutional kernel (spatial size, input channels, output channels, and the number of convolutional kernels). These attention mechanisms can dynamically adjust the weights of the convolutional kernel, thereby improving the feature extraction ability. The multi-scale convolutional attention module not only reduces the computational amount but also enables the model to more effectively capture the features of strip-shaped welding defects, enhancing the multi-scale defect detection performance of the model in complex backgrounds.
[0050] 2) The present invention adjusts the original anchor box ratios of the YOLOv11 object detection model through the DBSCAN clustering algorithm, making the model more suitable for welding defect detection. The new anchor box ratios are generated using the DBSCAN clustering algorithm. Using these new anchor box ratios can improve the performance of the object detection model.
[0051] 3) Regarding the situation where the shape and size of weld defects may vary greatly and there may be overlapping or non-overlapping between target boxes, MPDIoU can handle it better, providing more stable loss calculation and faster convergence speed, thereby improving the performance and accuracy of the detection model. Using MPDIoU to optimize the loss function improves the model's detection ability for small-sized welding defects and speeds up the model's convergence speed, avoiding falling into local optima. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 is the flowchart of the weld defect detection method based on deep learning of the present invention;
[0053] Figure 2 is the network structure diagram of the improved YOLOv11 of the present invention;
[0054] Figure 3 is the structure diagram of the multi-scale convolutional attention module of the present invention;
[0055] Figure 4 is the structure diagram of the full-dimensional dynamic convolutional module of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0057] Embodiment 1, as Figures 1 to 4As shown, this embodiment provides a welding defect detection method based on deep learning. The welding defect detection method includes the following steps:
[0058] S1. Collect welding defect images;
[0059] S2. Perform image preprocessing on the collected welding defect images, then classify and label the welding defects in the welding defect images, and form a dataset (for training the weld defect detection model based on the improved YOLOv11);
[0060] S3. Divide the dataset according to the ratio of training set: validation set: test set = 8:1:1;
[0061] S4. Build a weld defect detection model based on the improved YOLOv11;
[0062] S5. Use the training set to train the weld defect detection model; and use the validation set for verification. Calculate the gradient of the loss function and loss value (loss value) of the weld defect detection model for each training with the results of training and verification, and update the parameters of the weld defect detection model according to the gradient and hyperparameters such as the learning rate;
[0063] S6. Judge whether the performance of the weld defect detection model reaches the expectation or the maximum number of training times. If it reaches, perform step S7. If not, adjust the hyperparameters of the weld defect detection model and then repeat step S5 to train the weld defect detection model;
[0064] S7. Output the trained weld defect detection model and use the validation set to evaluate the performance of the trained welding defect detection model. The evaluation indicators are AP, mAP, F1–score, and processing speed (FPS).
[0065] Embodiment 2. In this embodiment, in addition to including all the technical features in Embodiment 1, it further includes:
[0066] The specific steps for building a weld defect detection model based on the improved YOLOv11 include:
[0067] S41. Introduce a multi-scale convolutional attention (MSCA) module in the Backbone part of the YOLOv11 object detection model;
[0068] Among them, the multi-scale convolutional attention module consists of three parts: depth-wise convolution, multi-branch depth-wise strip convolutions, and 1×1 convolution, which can improve the performance of the welding defect detection model in detecting welding defects with a large aspect ratio (such as cracks), such as Figure 3 as shown;
[0069] The expression of the multi-scale convolutional attention module is as follows:
[0070]
[0071] In the formula: F represents the input feature; Att and Out represent the attention map and the output feature respectively; represents the element-wise matrix multiplication operation; DW-Conv represents the depth-wise convolution, Scale i represents the i-th branch, i ∈ {0, 1, 2, 3}; Scale0 is the identity connection.
[0072] The depth-wise convolution reduces the computational complexity by separating the spatial and channel dimensions of the standard convolution, while maintaining the sensitivity to local features.
[0073] The multi-branch depth-wise strip convolution contains multiple branches. Each branch uses depth-wise strip convolutions of different sizes to extract multi-scale features. Each branch contains two depth-wise strip convolutions, which are used for feature extraction in the horizontal and vertical directions respectively. This design not only reduces the computational complexity, but also can effectively capture the features of strip-shaped welding defects.
[0074] In each branch, two depth-wise strip convolutions are used to approximate the standard depth convolution with a large kernel. The kernel sizes of each branch are set to 7, 11, and 21 respectively. There are two reasons for choosing the depth-wise strip convolution: one is that the strip convolution is lightweight. To simulate a standard 2D convolution with a kernel size of 7×7, only a pair of 7×1 and 1×7 convolutions are needed; on the other hand, it helps to extract strip-shaped features. This design not only reduces the computational complexity, but also can effectively capture the features of strip-shaped welding defects.
[0075] The 1×1 convolution is used to simulate the relationship between different channels and generate attention weights. The attention weights re-weight the input features through element-wise multiplication, thus realizing the spatial attention mechanism; the output of the 1×1 convolution is directly used as the attention weight, and the input features are re-weighted through element-wise multiplication, making the model pay more attention to important spatial regions, thereby improving the segmentation accuracy.
[0076] S42. Introduce the Omni-Dimensional Dynamic Convolution module (ODConv) in the Neck part of the YOLOv11 object detection model, and perform this operation before each feature fusion to improve the multi-scale defect detection performance of the weld defect detection model in complex backgrounds;
[0077] The Omni-Dimensional Dynamic Convolution module includes GAP, FC, ReLU, Sigmoid, and Softmax; among them, GAP represents the global average pooling layer, FC represents the fully connected layer, and ReLU, Sigmoid, and Softmax are activation functions. By introducing non-linear factors through the activation functions, the generalization ability and expression ability of the model can be improved, further enhancing the performance of the model;
[0078] The expression of the Omni-Dimensional Dynamic Convolution module is as follows:
[0079] y = (α w1 ⊙ α f1 ⊙ α c1 ⊙ α s1 ⊙ W1 + … + α wn ⊙ α fn ⊙ α cn ⊙ α sn ⊙ W n ) * x
[0080] In the formula: x is the input feature, y is the output feature; α wi represents the attention scalar of the convolution kernel, α fi represents the attention scalar of the output channel, α ci represents the attention scalar of the input channel, α si represents the attention scalar of the convolution parameter at each position, W i is the convolution kernel, ⊙ represents element-wise multiplication, and * represents the convolution operation.
[0081] The Omni-Dimensional Dynamic Convolution module introduces a multi-dimensional attention mechanism in four dimensions: the spatial size of the convolution kernel, the input channel, the output channel, and the number of convolution kernels. The attention mechanism can dynamically adjust the weights of the convolution kernel, thereby improving the feature extraction ability; the Omni-Dimensional Dynamic Convolution module is applied to the Neck part of the detection model, specifically by adding the Omni-Dimensional Dynamic Convolution module before each deep and shallow feature fusion operation; this can make full use of the convolution kernel space and channel information, capture rich context information, and thus improve the detection performance of the model for multi-scale weld defects and the detection ability in complex backgrounds.
[0082] S43. Modify the preset anchor box ratios of the YOLOv11 object detection model, and the new anchor box ratios are generated by using the DBSCAN clustering algorithm on the dataset.
[0083] The DBSCAN clustering algorithm automatically identifies clusters and noise points in the data, calculates the central ratio of each cluster, and uses these central ratios that are closer to the actual distribution of the data as the new anchor box ratios to improve the performance of the object detection model.
[0084] The steps to modify the anchor box ratio using the DBSCAN clustering algorithm are as follows:
[0085] First, extract the aspect ratios of all the true bounding boxes from the dataset. Represent these aspect ratios as data points (w i , h i ), where w i and h i are the width and height of the bounding box respectively;
[0086] Secondly, apply the DBSCAN clustering algorithm. By setting the neighborhood radius ε and the minimum number of points MinPts parameters, identify the core objects, that is, the points that satisfy having at least MinPts points in their ε neighborhood;
[0087] Then, starting from these core objects, cluster the density-reachable points into clusters. That is, if point q is in the ε neighborhood of point p and p is a core object, then q is directly density-reachable from p. If there exists a point chain p1, p2, …, p n , where p1 = p, p n = q, and pi+1 is directly density-reachable from pi, then q is density-reachable from p;
[0088] The expression is:
[0089]
[0090] Such that
[0091]
[0092] The above formula holds, then q is density-reachable from p; for each cluster, calculate the average aspect ratio of all its points, and these average aspect ratios are the generated anchor box ratios;
[0093] Finally, collect the average aspect ratios of all the clusters to form the new anchor box ratios, which are used in the weld defect detection model based on the improved YOLOv11 of the present invention, thereby improving the detection performance of the model for welding defects.
[0094] Example 3. In addition to all the technical features included in Example 1, this example further includes: using the MPDIoU loss function to calculate the loss value. In welding defect detection, MPDIoU comprehensively considers the overlapping area, non-overlapping area, center point distance, and deviations in width and height between the predicted bounding box and the ground truth bounding box by combining the traditional IoU and the penalty term of the minimum point distance, thereby providing more accurate loss calculation. It is particularly suitable for handling overlapping and non-overlapping objects, with simple calculation and fast convergence speed.
[0095] The calculation formula of MPDIoU is as follows:
[0096]
[0097] In the formula: is the traditional intersection over union (IoU), d1 and d2 are the Euclidean distances between the upper left and lower right points of the predicted bounding box and the ground truth bounding box respectively, and w and h are the average values of the widths and heights of the predicted bounding box and the ground truth bounding box.
[0098] In weld defect detection, MPDIoU has significant advantages. MPDIoU comprehensively considers the overlapping area, non-overlapping area, center point distance, and deviations in width and height between the predicted bounding box and the ground truth bounding box by combining the traditional IoU and the penalty term of the minimum point distance, thereby providing more accurate loss calculation. This method is particularly suitable for handling overlapping and non-overlapping objects, with simple calculation and fast convergence speed, and can effectively improve the performance and accuracy of the detection model. This loss function is more effective than the traditional IoU and its variants (such as GIoU, DIoU, CIoU, EIoU) when dealing with the situation where the predicted bounding box and the ground truth bounding box have the same aspect ratio but different width and height values.
[0099] In weld defect detection, this characteristic is particularly important because the shapes and sizes of weld defects may vary greatly, and there may be overlapping or non-overlapping situations between the target bounding boxes. MPDIoU can better handle these situations, provide more stable loss calculation and faster convergence speed, thereby improving the performance and accuracy of the detection model.
[0100] A multi-scale convolutional attention module is introduced into the Backbone part of the YOLOv11 object detection model. It consists of depthwise separable convolution, multi-branch depth strip convolution, and 1×1 convolution, which not only reduces the computational amount but also enables the model to more effectively capture the features of strip-shaped welding defects. A full-dimensional dynamic convolution module is introduced into the Neck part of the YOLOv11 object detection model to improve the multi-scale defect detection performance of the model in complex backgrounds. The preset anchor box ratios of the model are modified, and the new anchor box ratios are generated using the DBSCAN clustering algorithm. Using these new anchor box ratios can improve the performance of the object detection model because they are closer to the actual distribution of the data. The MPDIoU loss function is used to calculate the loss value to guide the optimization direction of the model. For the situation where the shape and size of the weld defects may vary greatly and there may be overlap or non-overlap between the target boxes, MPDIoU can handle it better, providing more stable loss calculation and faster convergence speed, thereby improving the performance and accuracy of the detection model.
[0101] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be included in the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
[0102] In addition, it should be understood that although this specification is described according to embodiments, not every embodiment only contains an independent technical solution. This narrative way of the specification is only for clarity. Those skilled in the art should regard the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A welding defect detection method based on deep learning, characterized in that, The welding defect detection method includes the following steps: S1. Collect welding defect images; S2. Perform image preprocessing on the collected welding defect images, classify and label the welding defects in the images, and form a data set; S3. Divide the data set according to the ratio of training set: validation set: test set = 8:1:1; S4. Build a weld defect detection model based on the improved YOLOv11; S5. Use the training set to train the weld defect detection model, use the validation set for validation, calculate the gradient of the loss function and loss value of each training of the weld defect detection model with respect to the parameters of the weld defect detection model based on the training and validation results, and update the parameters of the weld defect detection model according to the gradient and learning rate; S6. Determine whether the performance of the weld defect detection model reaches the expectation or the maximum number of training times. If it reaches, proceed to step S7. If not, adjust the parameters of the weld defect detection model and then repeat step S5 to train the weld defect detection model; S7. Output the trained weld defect detection model and use the validation set to evaluate the performance of the trained welding defect detection model. The evaluation indicators are AP, mAP, F1–score, and processing speed.
2. The welding defect detection method according to claim 1, characterized in that: In the above S4, the specific steps for building a weld defect detection model based on the improved YOLOv11 include: S41. Introduce a multi-scale convolutional attention module into the Backbone part of the YOLOv11 object detection model; S42. Introduce a full-dimensional dynamic convolution module into the Neck part of the YOLOv11 object detection model and perform this operation before each feature fusion; S43. Modify the preset anchor box ratios of the YOLOv11 object detection model, and the new anchor box ratios are generated by using the DBSCAN clustering algorithm on the data set.
3. The welding defect detection method according to claim 2, characterized in that: The multi-scale convolutional attention module consists of three parts: depthwise separable convolution, multi-branch depthwise strip convolution, and 1×1 convolution; Among them, the expression of the multi-scale convolutional attention module is as follows: Where: F represents the input feature; Att and Out represent the attention map and the output feature respectively; represents the element-wise matrix multiplication operation; DW-Conv represents the depthwise separable convolution, Scale i represents the i-th branch, i ∈ {0, 1, 2, 3}; Scale0 is the identity connection.
4. The welding defect detection method according to claim 3, characterized in that: The depthwise separable convolution separates the spatial and channel dimensions of the standard convolution.
5. The welding defect detection method according to claim 3, characterized in that: The multi-branch depthwise strip convolution contains multiple branches. Each branch uses depthwise strip convolutions of different sizes to extract multi-scale features, and each branch contains two depthwise strip convolutions for feature extraction in the horizontal and vertical directions respectively.
6. The welding defect detection method according to claim 3, characterized in that: The 1×1 convolution is used to simulate the relationship between different channels and generate attention weights. The attention weights reweight the input features through element-wise multiplication to achieve the spatial attention mechanism, and the output of the 1×1 convolution is directly used as the attention weights.
7. The welding defect detection method according to claim 2, wherein: In the above S42, the full-dimensional dynamic convolution module includes GAP, FC, ReLU, Sigmoid, and Softmax; where GAP represents the global average pooling layer, FC represents the fully connected layer, and ReLU, Sigmoid, and Softmax are activation functions; The expression of the full-dimensional dynamic convolution module is as follows: y = (α w1 ⊙ α f1 ⊙ α c1 ⊙ α s1 ⊙ W1 + … + α wn ⊙ α fn ⊙ α cn ⊙ α sn ⊙ W n ) * x where: x is the input feature, y is the output feature; α wi represents the attention scalar of the convolutional kernel, α fi represents the attention scalar of the output channel, α ci represents the attention scalar of the input channel, α si represents the attention scalar of the convolutional parameter at each position, W i is the convolutional kernel, ⊙ represents element-wise multiplication, and * represents the convolution operation; The full-dimensional dynamic convolution module introduces a multi-dimensional attention mechanism in four dimensions: the spatial size of the convolution kernel, the number of input channels, the number of output channels, and the number of convolution kernels.
8. The welding defect detection method according to claim 2, characterized in that: In S43, the DBSCAN clustering algorithm automatically identifies clusters and noise points in the data and calculates the central ratio of each cluster. The steps for modifying the anchor box ratio using the DBSCAN clustering algorithm are as follows: First, extract the aspect ratios of all the true bounding boxes from the dataset, and represent the aspect ratios as data points (w i , h i ), where w i and h i are the width and height of the bounding box, respectively; Secondly, using the DBSCAN clustering algorithm, by setting the neighborhood radius ε and the minimum number of points MinPts parameters, the core objects are identified, that is, the points that satisfy at least MinPts points in their ε neighborhood. Then, starting from the core object, cluster the density-reachable points into clusters, that is, if point q is within the ε-neighborhood of point p and p is a core object, then q is directly density-reachable from p. If there exists a point chain p1, p2, …, p n , where p1 = p, p n = q, and pi+1 is directly density-reachable from pi, then q is density-reachable from p; The expression is: Such that If the above equation holds, then q is density-reachable from p; for each cluster, calculate the average width-to-height ratio of all its points, and the average width-to-height ratio is the generated anchor box ratio. Finally, collect the average width-to-height ratios of all clusters to form a new anchor box ratio.
9. The welding defect detection method according to claim 1, characterized in that: In S5, the MPDIoU loss function is used to calculate the loss value. In welding defect detection, MPDIoU combines the traditional IoU and the penalty term of the minimum point distance.
10. The welding defect detection method according to claim 9, wherein: The calculation formula of the MPDIoU is as follows: Wherein: is the traditional intersection over union (IoU), d1 and d2 are the Euclidean distances between the upper left corner points and the lower right corner points of the predicted bounding box and the ground truth bounding box respectively, and w and h are the average values of the widths and heights of the predicted bounding box and the ground truth bounding box.
Citation Information
Patent Citations
Video classification method, electronic equipment and storage medium
CN112101091A
Binocular vision-based autonomous underwater robot recovery guidance pseudo light source removal method
CN113034399A
Gaming device having multiplier poker game
US20210110674A1
Cited By
PID (Proportion Integration Differentiation) drawing element intelligent identification and topology reconstruction method based on visual inspection
CN121281087A
PID paper element intelligent recognition and topological reconstruction method based on visual detection
CN121281087B
Multi-layer and multi-pass welding defect detection method based on multi-source sensor feature fusion
CN121980510A