Plate welding part detection method based on improved image stitching and deep learning network

CN119205654BActive Publication Date: 2026-09-08NANJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411240279.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-05
Publication Date
2026-09-08
Estimated Expiration
2044-09-05

AI Technical Summary

Technical Problem

然而,基于深度学习的板材类焊接件目标检测算法在实际应用中仍存在一些问题

Benefits of technology

[0043] First, the plate weldment detection method of this invention, based on improved image stitching and deep learning networks, shortens the feature point extraction time by using a rectangular feature region constraint function and the SURF algorithm to limit the feature point extraction range. It also shortens the feature point extraction time and improves the feature point pair matching accuracy and computation speed by using a kd-tree to accelerate nearest neighbor search and introducing an improved Manhattan distance. Furthermore, the introduction of the RANSAC algorithm can eliminate incorrectly matched feature point pairs, improving the stitching quality of weldment images and removing obvious stitching artifacts, thus facilitating the creation of weldment datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119205654B_ABST
    Figure CN119205654B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on improved image splicing and deep learning network's plate class welding piece detection method, comprising: the image of plate class welding piece that camera group is photographed at same time is fused by the fusion method of improved SURF algorithm and is fused by image splicing;YOLOv8 neural network is structurally improved, form the improved YOLOv8 neural network for plate class welding piece target detection identification;Comprehensively obtain NWD- Inner- IoU loss function;Splicing fusion image dataset is imported into improved YOLOv8 neural network, and NWD- Inner- IoU loss function is combined to train YOLOv8 neural network after structural improvement, and the model after training is used for the detection of plate class welding piece.The application can effectively improve the detection accuracy and detection precision of plate class welding piece, and the robustness of the model, strengthen the practicability, so that it can be deployed in edge device, and the performance requirement of required equipment can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of real-time detection technology for sheet metal welded parts, specifically relating to a method for detecting sheet metal welded parts based on improved image stitching and deep learning networks. Background Technology

[0002] Computer vision technology is rapidly developing in the industrial sector. Previously, parts sorting relied on manual labor; in recent years, semi-automated methods combined with manual labor have been used; and now, deep learning-based object detection methods are being employed for parts sorting. However, machine-based parts sorting places certain demands on operators and processes, and the complex factory environment can easily lead to errors in traditional machine sorting. Conversely, deep learning-based object detection methods can automatically extract features from large datasets and improve the accuracy of object detection. Therefore, deep learning-based object detection methods have become a hot research topic.

[0003] Intelligent manufacturing is becoming increasingly widespread, and the detection of mechanical parts, as the foundation of automated mechanical production, boasts significant advantages such as non-contact operation, rapid recognition, high accuracy, and strong anti-interference capabilities, greatly improving economic and social benefits. However, deep learning-based algorithms for detecting welded sheet metal parts still face some challenges in practical applications. On one hand, feature extraction and precise target localization are difficult in welded part detection tasks. On the other hand, many industrial devices on-site have limited computing power, and traditional deep learning models tend to generate a large number of computational parameters, requiring high-performance equipment for calculation. This makes traditional deep learning models unsuitable for some industrial equipment. Therefore, how to compress deep learning models to reduce their size and accelerate detection speed remains a challenge. Summary of the Invention

[0004] To address the problems existing in the prior art, this invention proposes a method for detecting welded sheet metal parts based on improved image stitching and deep learning networks. It uses the YOLOv8 network, which currently has good detection performance, as a foundation to effectively improve the detection accuracy and precision of welded sheet metal parts, as well as the robustness of the model, enhance its practicality, and enable it to be deployed in edge devices, while reducing the performance requirements of the required equipment.

[0005] Technical solution:

[0006] In a first aspect, the present invention discloses a method for detecting welded parts made of sheet metal based on improved image stitching and deep learning networks, the method comprising the following steps:

[0007] S1: Images of welded sheet metal parts captured simultaneously by a camera group are stitched together using an improved SURF algorithm. After annotation, a stitched and fused image dataset is constructed. Specifically, this includes:

[0008] By using a rectangular feature region constraint function, the SURF algorithm is used to limit the feature point extraction range to a rectangular area selected in the image of the welded sheet metal. A kd-tree is used to perform nearest neighbor search on the extracted feature points. During the nearest neighbor search, an improved Manhattan distance is used for distance measurement, and the measurement result is compared with a set threshold to obtain feature point pairs. The RANSAC algorithm is introduced to remove mismatched feature point pairs from the two images and solve for the homography matrix. The homography matrix is ​​then used to transform the two images to the same coordinate system for image fusion, resulting in a stitched fused image.

[0009] S2: Improve the structure of the YOLOv8 neural network to form an improved YOLOv8 neural network for target detection and recognition of welded sheet metal parts; specifically including:

[0010] An EMA attention mechanism is added after the 1*1 Conv convolutional layer in the FasterNet Block of the FasterNet network to perform a weighted average of historical data, forming a Faster-EMA improved C2f module. The Faster-EMA improved C2f module is used in layers 2, 4, 6, and 8 of the YOLOv8 neural network. Dense convolution SC and depthwise separable convolution DSC are mixed to obtain hybrid convolution GSConv, which is introduced in layers 16 and 19 of the YOLOv8 neural network.

[0011] S3: Based on the normalized Wasserstein distance, a loss function based on NWD is proposed. The NWD-Inner-IoU loss function is obtained by combining the NWD-based loss function and the Inner-IoU loss function. The stitched and fused image dataset is imported into the improved YOLOv8 neural network. The improved YOLOv8 neural network is trained by combining the NWD-Inner-IoU loss function. The trained model is used for the detection of plate welded parts.

[0012] Furthermore, the rectangular feature region restriction function is:

[0013]

[0014] Where (x,y) and (x',y') are the horizontal and vertical coordinates of the pixel and the coordinates of the top-left corner of the rectangular feature-restricted region, respectively; w is the width of the rectangular feature-restricted region; and h is the height of the rectangular feature-restricted region.

[0015] Furthermore, the nearest neighbor search is performed on the extracted feature points using a kd-tree. During the nearest neighbor search process, an improved Manhattan distance is used as a distance metric. The process of comparing the metric results with a set threshold to obtain feature point pairs includes the following steps:

[0016] Assuming k is a spatial dimension, a kd-tree is used to partition the data points in k-dimensional space. Feature points are detected using the Hessian matrix, and the feature points are described using Haar wavelet features to obtain a sample set. Then, a kd-tree nearest neighbor search is performed, where the nearest neighbor search is defined as follows:

[0017]

[0018] Where E is the sample set obtained after feature point detection and description, and d is the set to be queried; d is the sample point, d' is the nearest neighbor point to be queried, and d” is the other points in the set;

[0019] In nearest neighbor search, an improved Manhattan distance is used as a distance metric. Feature-matching point pairs are obtained by comparing their magnitudes with a set threshold. Two k-dimensional vectors A(x) are then used. 11 ,x 12 ,x 13 ,...,x 1k ) and B(x 21 ,x 22 ,x 23 ,...x 2k The improved Manhattan distance between them is:

[0020]

[0021] Where D is the improved Manhattan distance between the vectorizations of two feature point pairs, and w i It is the reciprocal of the standard deviation in the i-th dimension.

[0022] Furthermore, the RANSAC algorithm is introduced to remove mismatched feature point pairs from the two images and solve for the homography matrix. The homography matrix is ​​then used to transform the two images to the same coordinate system for image fusion, resulting in the stitched and fused image. The process includes the following steps:

[0023] Randomly select four pairs of feature points that have been matched using the improved Manhattan distance, and calculate the transformation matrix M. Use the transformation matrix M to transform the points in the remaining feature point pairs that belong to the first image, and calculate the distance between the transformed points and the corresponding points in the second image. If the distance is less than a set threshold, save the feature point pair. Repeat the above steps until the correct matching points are obtained.

[0024] The transformation matrix with the most matching points is selected as the homography matrix H, and the two images are transformed to the same coordinate system for stitching; the transformation formula is:

[0025] [X,Y]=H[x,y]

[0026] Where H is the homography matrix, x and y are the coordinates of the feature points to be matched, and X and Y are the coordinates of the matched feature points.

[0027] Furthermore, the historical data is weighted and averaged using the EMA attention mechanism, where the weight of each data point decreases as it gets closer to the current time point; the EMA calculation formula is:

[0028] EMA[t]=α*x[t]+(1-α)*EMA[t-1]

[0029] Where t represents the time step, x[t] represents the original data at the t-th time point; α is a smoothing factor, with a value between 0 and 1, representing the weight of the current sample; (1-α) represents the weight of the historical data; EMA[t-1] represents the EMA value of the previous time point.

[0030] Furthermore, the hybrid convolution GSConv takes the input of the previous neural network as the output, with the number of input channels being Q. First, it obtains feature M1 through dense convolution SC, with the number of channels being Q / 2; then, it performs pointwise convolution and channelwise convolution through depthwise separable convolution DSC to obtain feature M2, with the number of channels being Q / 2; after concatenating feature M1 and feature M2 into M3, a shuffle operation is performed to shuffle and recombine the concatenated information to obtain the final output feature M4, which is then fed into the next layer of the YOLOV8 neural network.

[0031] Furthermore, the NWD-Inner-IoU loss function is:

[0032]

[0033]

[0034] In the formula, m is the coefficient weight, which takes a value between 0 and 1; P is the predicted region, G is the labeled region, I represents the intersection area of ​​the predicted region and the labeled region, and U represents their union area. It is the bounding box A = (cx a ,cy a ,w a ,h a ), B = (cx b ,cy b ,w b ,h b The second-order Wasserstein distance of N;a N b The rectangle follows a Gaussian distribution, where cx is the x-coordinate of the center point, cy is the y-coordinate of the center point, w is the width of the rectangle, and h is the height of the rectangle.

[0035] Secondly, this invention discloses a detection device for welded sheet metal parts based on improved image stitching and deep learning networks, the detection device for welded sheet metal parts comprising:

[0036] The image stitching and fusion module is used to stitch together images of welded sheet metal parts captured simultaneously by a camera group using an improved SURF algorithm. After annotation, a stitched and fused image dataset is constructed. Specifically, it includes:

[0037] By using a rectangular feature region constraint function, the SURF algorithm is used to limit the feature point extraction range to a rectangular area selected in the image of the welded sheet metal. A kd-tree is used to perform nearest neighbor search on the extracted feature points. During the nearest neighbor search, an improved Manhattan distance is used for distance measurement, and the measurement result is compared with a set threshold to obtain feature point pairs. The RANSAC algorithm is introduced to remove mismatched feature point pairs from the two images and solve for the homography matrix. The homography matrix is ​​then used to transform the two images to the same coordinate system for image fusion, resulting in a stitched fused image.

[0038] The model improvement module is used to structurally improve the YOLOv8 neural network, forming an improved YOLOv8 neural network for target detection and recognition of sheet metal welded parts. Specifically, an EMA attention mechanism is added after the 1*1 Conv convolutional layer in the FasterNet Block of the FasterNet network to perform weighted averaging of historical data, forming a Faster-EMA improved C2f module. The Faster-EMA improved C2f module is used in layers 2, 4, 6, and 8 of the YOLOv8 neural network. Dense convolution SC and depthwise separable convolution DSC are mixed to obtain hybrid convolution GSConv, which is introduced into layers 16 and 19 of the YOLOv8 neural network.

[0039] The model training module is used to propose a loss function based on normalized Wasserstein distance (NWD), and to obtain the NWD-Inner-IoU loss function by combining the NWD-based loss function and the Inner-IoU loss function. The stitched and fused image dataset is imported into the improved YOLOv8 neural network, and the improved YOLOv8 neural network is trained by combining the NWD-Inner-IoU loss function. The trained model is then used for the detection of welded parts of sheet metal.

[0040] Thirdly, the present invention discloses a computer storage device containing a computer program, which enables the computer to execute the aforementioned method for detecting welded sheet metal parts.

[0041] Fourthly, the present invention discloses an electronic device, including: a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When the processor executes the program, it implements the aforementioned method for detecting welded sheet metal parts.

[0042] Beneficial effects:

[0043] First, the plate weldment detection method of this invention, based on improved image stitching and deep learning networks, shortens the feature point extraction time by using a rectangular feature region constraint function and the SURF algorithm to limit the feature point extraction range. It also shortens the feature point extraction time and improves the feature point pair matching accuracy and computation speed by using a kd-tree to accelerate nearest neighbor search and introducing an improved Manhattan distance. Furthermore, the introduction of the RANSAC algorithm can eliminate incorrectly matched feature point pairs, improving the stitching quality of weldment images and removing obvious stitching artifacts, thus facilitating the creation of weldment datasets.

[0044] Second, the plate welded component detection method of the present invention based on improved image stitching and deep learning network utilizes Faster-EMA structure to reduce redundant calculations, memory access and model sensitivity to local noise, strengthen feature fusion, reduce overfitting risk and improve the detection accuracy of plate welded components.

[0045] Third, the plate welding component detection method of the present invention based on improved image stitching and deep learning network utilizes GSConv, which not only retains key feature information and makes the network structure lightweight, but also reduces model parameters and computational load, accelerates feature fusion and model inference speed, and facilitates deployment to edge devices.

[0046] Fourth, the plate welding component detection method of the present invention based on improved image stitching and deep learning network proposes the NWD-Inner-IoU loss function to address the problem of poor similarity in judging non-overlapping or mutually inclusive bounding boxes, thereby improving the accuracy of measuring bounding box similarity and enhancing the robustness of the model. Attached Figure Description

[0047] Figure 1 This is a schematic diagram of the improved YOLOv8 neural network structure of this invention;

[0048] Figure 2 This is a schematic diagram of the improved Faster-EMA structure of the present invention;

[0049] Figure 3This is a schematic diagram of the improved GSConv structure of the present invention;

[0050] Figure 4 This is a schematic diagram of the images to be spliced ​​and welded according to the present invention; wherein, (a) and (b) are two images taken by different cameras at the same time;

[0051] Figure 5 This is a schematic diagram of feature point matching;

[0052] Figure 6 The images shown are a stitched image of the welded sheet metal parts of the present invention and the original image of the welded parts; wherein, (a) is the stitched image and (b) is the original image of the welded parts;

[0053] Figure 7 This is a schematic diagram of the detection results of the improved detection method of the present invention;

[0054] Figure 8 This is a flowchart of the detection method for welded sheet metal parts based on improved image stitching and deep learning networks according to the present invention. Detailed Implementation

[0055] The following embodiments are provided to enable those skilled in the art to more fully understand the present invention, but do not limit the invention in any way.

[0056] This invention discloses a method for detecting welded sheet metal parts based on improved image stitching and deep learning networks. See [link to relevant documentation]. Figure 8 It mainly includes the following steps:

[0057] S1: The sheet metal welded parts on the conveyor belt are photographed using a camera group (at least two cameras at different positions) on the production line. The images of the sheet metal welded parts captured simultaneously by the camera group are then stitched together using an improved SURF algorithm. After labeling, a stitched and fused image dataset is generated. In this embodiment, the label categories include nuts (M8, M10), studs (9R-2160, 160-0240), and wire clips. Specifically, this includes:

[0058] (1) By using a rectangular feature region constraint function, the SURF algorithm is used to limit the feature point extraction range, thereby shortening the feature point extraction time. In this part, the specific content of the rectangular feature region constraint function g(x,y) is as follows:

[0059]

[0060] Where (x,y) and (x',y') are the horizontal and vertical coordinates of the pixel and the coordinates of the upper left corner of the rectangular feature restriction area, respectively; w is the width of the rectangular feature restriction area; and h is the height of the rectangular feature restriction area. Formula (1) is used to limit the feature point extraction range to the rectangular area selected in the image of the plate welded part.

[0061] (2) Use kd-tree to accelerate nearest neighbor search and introduce improved Manhattan distance to shorten computation time and improve feature point matching accuracy.

[0062] Specifically, assuming k is the spatial dimension, the data points are partitioned in the k-dimensional space using a kd-tree, feature points are detected using the Hessian matrix, and a sample set is obtained by describing the feature points using Haar wavelet features. Then, a kd-tree nearest neighbor search is performed, where the nearest neighbor search is defined as follows:

[0063]

[0064] Where E is the sample set obtained after feature point detection and description, which is a set to be queried, d is the sample point, and d' is the number of points.

[0065] "d" represents the nearest neighbor to be queried, and "d" represents the other points in the set.

[0066] The specific details of the improved Manhattan distance are as follows:

[0067] In the feature point matching stage, a distance metric function is introduced. In nearest neighbor search, an improved Manhattan distance is used for distance measurement. Feature matching point pairs are obtained by comparing their magnitudes with a set threshold. Two k-dimensional vectors A(x) are then used. 11 ,x 12 ,x 13 ,...,x 1k ) and B(x 21 ,x 22 ,x 23 ,...x 2k The improved Manhattan distance between them is:

[0068]

[0069] Where D is the improved Manhattan distance between the vectorizations of two feature point pairs, and w i It is the reciprocal of the standard deviation in the i-th dimension.

[0070] (3) The RANSAC algorithm is introduced to remove the feature point pairs with incorrect feature matching on the two images and solve the homography matrix. The homography matrix is ​​then used to transform the two images to the same coordinate system for image fusion.

[0071] Specifically, four pairs of feature points that have been matched using the improved Manhattan distance are randomly selected, and the transformation matrix M is calculated. The remaining feature point pairs belonging to the first image are transformed using M, and the distance between the transformed points and their corresponding points in the second image is calculated. If the distance is less than a set threshold, the feature point pair is saved. This process is repeated several times to obtain the correct matching points. The transformation matrix with the most matching points is selected as the homography matrix H, and the two images are transformed to the same coordinate system for stitching.

[0072] The transformation formula is as follows:

[0073] [X,Y]=H[x,y] (4); where H is the homography matrix, x and y are the coordinates of the feature points to be matched, and X and Y are the coordinates of the matched feature points.

[0074] S2: An improved YOLOv8 neural network is developed based on the existing YOLOv8 neural network for target detection and recognition of welded sheet metal parts. See also Figure 1 and Figure 2 The improvements specifically include:

[0075] (1) Improve the C2f module with Faster-EMA in the 2nd, 4th, 6th and 8th layers of the existing YOLOv8 neural network. The EMA attention mechanism is connected after the 1*1 Conv convolutional layer in the FasterNetBlock of the FasterNet network to form Faster-EMA, so as to reduce redundant computation and memory access, strengthen feature fusion and reduce the risk of overfitting.

[0076] The EMA attention mechanism, based on a grouping structure, corrects the sequential processing method of CA, eliminating the need for dimensionality reduction and exponential moving average smoothing to perform weighted averaging on the sequence data. EMA assigns greater weight to more recent data points and less weight to earlier data points. The EMA calculation formula is:

[0077] EMA[t]=α*x[t]+(1-α)*EMA[t-1] (5);

[0078] Where t represents the time step, x[t] represents the original data at the t-th time point, α is the smoothing factor, which usually takes a value between 0 and 1, representing the weight of the current sample, (1-α) represents the weight of the historical data, and EMA[t-1] represents the EMA value of the previous time point.

[0079] The EMA attention mechanism is used to perform a weighted average of historical data, where the weight of each data point decreases as it gets closer to the current time point. This effectively smooths the time series data, making it more continuous and stable. The EMA attention mechanism is introduced to enhance the model's stability, interpretability, and attention to overall semantics, while also strengthening the fusion of dynamic features, reducing the model's sensitivity to local noise, and mitigating the risk of overfitting. To address the issue of poor detection accuracy for small targets on welded sheet metal, the C2f module is improved using Faster-EMA to enhance the detection accuracy of welded sheet metal components.

[0080] The entire feature extraction process is as follows: The stitched image is fed into the backbone part (layers 1-9) of the model for feature extraction. The C2f module in layers 2, 4, 6, and 8, which uses the Faster-EMA structure to improve the input features, ensures feature diversity, and achieves low latency. The image is downsampled by 2, 4, 8, 16, and 32 times to obtain five features, which are denoted as P1, P2, P3, P4, and P5, respectively.

[0081] (2) See Figure 3 In the neck part of the YOLOv8 network, a shuffling method is introduced to shuffle SC and DSC to obtain a hybrid convolution GSConv. Hybrid convolution GSConv is introduced into the 16th and 19th layers of the existing YOLOv8 neural network. Hybrid convolution GSConv concatenates the features obtained by dense convolution and depthwise separable convolution and performs a shuffle operation to shuffle and reorganize the concatenated information to obtain the final output features. The output features are then used as input to the next layer of the YOLOv8 neural network, making the network structure lightweight and reducing the model parameters and computational load.

[0082] The hybrid convolutional layer GSC takes the input of the previous neural network as its output, with Q input channels. First, it performs dense convolution SC to obtain feature M1 with Q / 2 channels. Next, it performs depthwise separable convolution DSC, which is further divided into pointwise convolution and channelwise convolution, to obtain feature M2 with Q / 2 channels. M1 and M2 are concatenated to form M3. Finally, a shuffle operation is performed to shuffle and recombine the concatenated information to obtain the final output feature M4, which is then fed into the next layer of the YOLOv8 neural network.

[0083] (3) To address the problem of poor similarity in judging non-overlapping or mutually containing bounding boxes, an NWD-Inner-IoU loss function is proposed. Inner-IoU is used to measure the similarity of overlapping bounding boxes, and NWD is further used to overcome the problem of inaccurate similarity measurement between non-overlapping or mutually containing bounding boxes. By using loss functions with different weight coefficients, the stability, accuracy and speed of small target detection are improved, and the performance gap between training and testing is eliminated.

[0084] In this embodiment, the feature P3 output by fusing the neck part is obtained by using NWD and Inner-IoU loss functions with different weight coefficients. out P4 out P5 out The data is fed into the head section as input, and the NWD-Inner-IoU loss function is used to predict the location of objects at three scales: 80×80×256, 40×40×512, and 20×20×512, and then the results are output.

[0085] The NWD loss function possesses scale invariance and smoothness against positional biases, and in object detection, it can also measure the similarity between non-overlapping or mutually containing bounding boxes. Assume there are bounding boxes A = (cx... a ,cy a ,w a ,h a ), B = (cx b ,cy b ,w b ,h b For two bounding boxes, their second-order Wasserstein distance can be defined as:

[0086]

[0087] Where, N a N b The rectangle follows a Gaussian distribution, where cx is the x-coordinate of the center point, cy is the y-coordinate of the center point, w is the width of the rectangle, and h is the height of the rectangle.

[0088] Since formula (7) is a distance metric and cannot be directly used for similarity, a normalized exponent (i.e., Wasserstein distance) is used here for measurement, and the formula is:

[0089]

[0090] The loss value based on NWD is:

[0091] L NWD =1-NWD(N aN b (8);

[0092] The Inner-IoU loss function is commonly used in image segmentation tasks to measure the degree of overlap between two regions. In object detection, it is mainly used to match the model's predictions with the ground truth annotations. The mathematical expression for the Inner-IoU loss function is:

[0093]

[0094] Where P is the predicted region, G is the labeled region, I represents the intersection area of ​​the predicted region and the labeled region, and U represents their union area.

[0095] To address the issue of poor similarity assessment between non-overlapping or mutually containing bounding boxes, the NWD-Inner-IoU loss function is proposed. This allows the model to learn target features more comprehensively. It overcomes the inaccuracy in similarity measurement between non-overlapping or mutually containing bounding boxes, ensures the model's predictions match the ground truth annotations, improves the stability, accuracy, and speed of small object detection, enhances model performance, reduces overfitting risk and strengthens the model's robustness, and eliminates the performance gap between training and testing, ultimately improving the accuracy and precision of object detection. The mathematical expression of the fused loss function is:

[0096]

[0097] Where m is the coefficient weight, and its value is between 0 and 1.

[0098] S3: Input the stitched and fused image dataset from step S1 into the improved YOLOv8 model for training, and use the trained model for the detection of plate welded parts.

[0099] The improved image stitching algorithm produces images of the parts to be stitched and welded, such as... Figure 4 As shown, (a) and (b) are two images taken by different cameras at the same time; the feature point matching results are as follows. Figure 5 As shown, the feature point pair matching accuracy is high and the matching effect is good; the stitched image and the original image of the welded part are compared as follows. Figure 6 As shown, (a) is the stitched image, and (b) is the original image of the welded part. The stitching effect is good, with no obvious traces. The improved inspection results for plate welded parts are as follows: Figure 7 As shown, it can accurately identify nuts, studs (blot_none, bolt_tiny), and thread clips (u_thread). The final improved algorithm, compared to the original YOLOv8 algorithm, improves the average accuracy (mAP) by 0.8 percentage points and reduces the number of parameters to 70% of the original.

[0100] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should be considered within the scope of protection of the present invention.

Claims

1. A method for detecting welded sheet metal parts based on improved image stitching and deep learning networks, characterized in that, The inspection method for welded sheet metal components includes the following steps: S1: Images of welded sheet metal parts captured simultaneously by a camera group are stitched together using an improved SURF algorithm. After annotation, a stitched and fused image dataset is constructed. Specifically, this includes: By using a rectangular feature region constraint function, the SURF algorithm is used to limit the feature point extraction range to a rectangular area selected in the image of the welded sheet metal. A kd-tree is used to perform nearest neighbor search on the extracted feature points. During the nearest neighbor search, an improved Manhattan distance is used for distance measurement, and the measurement result is compared with a set threshold to obtain feature point pairs. The RANSAC algorithm is introduced to remove mismatched feature point pairs from the two images and solve for the homography matrix. The homography matrix is ​​then used to transform the two images to the same coordinate system for image fusion, resulting in a stitched fused image. S2: Improve the structure of the YOLOv8 neural network to form an improved YOLOv8 neural network for target detection and recognition of welded sheet metal parts; specifically including: An EMA attention mechanism is added after the 1*1 Conv convolutional layer in the FasterNet Block of the FasterNet network to perform a weighted average of historical data, forming a Faster-EMA improved C2f module. The Faster-EMA improved C2f module is used in layers 2, 4, 6, and 8 of the YOLOv8 neural network. Dense convolution SC and depthwise separable convolution DSC are mixed to obtain hybrid convolution GSConv, which is introduced in layers 16 and 19 of the YOLOv8 neural network. S3: Based on the normalized Wasserstein distance, a loss function based on NWD is proposed. The NWD-Inner-IoU loss function is obtained by combining the NWD-based loss function and the Inner-IoU loss function. The stitched and fused image dataset is imported into the improved YOLOv8 neural network. The improved YOLOv8 neural network is trained by combining the NWD-Inner-IoU loss function. The trained model is used for the detection of plate welded parts.

2. The method for detecting welded sheet metal parts based on improved image stitching and deep learning networks according to claim 1, characterized in that, The rectangular feature region constraint function is: Where (x,y) and (x',y') are the horizontal and vertical coordinates of the pixel and the coordinates of the top-left corner of the rectangular feature-restricted region, respectively; w is the width of the rectangular feature-restricted region; and h is the height of the rectangular feature-restricted region.

3. The method for detecting welded sheet metal parts based on improved image stitching and deep learning networks according to claim 1, characterized in that, The process of using a kd-tree to perform nearest neighbor search on the extracted feature points, and using an improved Manhattan distance as a distance metric during the nearest neighbor search, and comparing the metric results with a set threshold to obtain feature point pairs, includes the following steps: Assuming k is a spatial dimension, a kd-tree is used to partition the data points in k-dimensional space. Feature points are detected using the Hessian matrix, and the feature points are described using Haar wavelet features to obtain a sample set. Then, a kd-tree nearest neighbor search is performed, where the nearest neighbor search is defined as follows: Where E is the sample set obtained after feature point detection and description, and d is the set to be queried; d is the sample point, d' is the nearest neighbor point to be queried, and d” is the other points in the set; In nearest neighbor search, an improved Manhattan distance is used as a distance metric. Feature-matching point pairs are obtained by comparing their magnitudes with a set threshold. Two k-dimensional vectors A(x) are then used. 11 ,x 12 ,x 13 ,...,x 1k ) and B(x 21 ,x 22 ,x 23 ,...x 2k The improved Manhattan distance between them is: Where D is the improved Manhattan distance between the vectorizations of two feature point pairs, and w i It is the reciprocal of the standard deviation in the i-th dimension.

4. The method for detecting welded sheet metal parts based on improved image stitching and deep learning networks according to claim 1, characterized in that, The RANSAC algorithm is introduced to remove mismatched feature point pairs from two images and solve for the homography matrix. The homography matrix is ​​then used to transform the two images to the same coordinate system for image fusion, resulting in the stitched fused image. The process includes the following steps: Randomly select four pairs of feature points that have been matched using the improved Manhattan distance, and calculate the transformation matrix M. Use the transformation matrix M to transform the points in the remaining feature point pairs that belong to the first image, and calculate the distance between the transformed points and the corresponding points in the second image. If the distance is less than a set threshold, save the feature point pair. Repeat the above steps until the correct matching points are obtained. The transformation matrix with the most matching points is selected as the homography matrix H, and the two images are transformed to the same coordinate system for stitching; the transformation formula is: [X,Y]=H[x,y] Where H is the homography matrix, x and y are the coordinates of the feature points to be matched, and X and Y are the coordinates of the matched feature points.

5. The method for detecting welded sheet metal parts based on improved image stitching and deep learning networks according to claim 1, characterized in that, The EMA attention mechanism is used to perform a weighted average of historical data, where the weight of each data point decreases as it becomes closer to the current time point; the EMA calculation formula is: EMA[t]=α*x[t]+(1-α)*EMA[t-1] Where t represents the time step, x[t] represents the original data at the t-th time point; α is a smoothing factor, with a value between 0 and 1, representing the weight of the current sample; (1-α) represents the weight of the historical data; EMA[t-1] represents the EMA value of the previous time point.

6. The method for detecting welded sheet metal parts based on improved image stitching and deep learning networks according to claim 1, characterized in that, The hybrid convolution GSConv takes the input of the previous neural network as the output, with Q input channels. First, it obtains feature M1 through dense convolution SC, with Q / 2 channels. Then, it performs pointwise and channelwise convolution through depthwise separable convolution DSC to obtain feature M2, with Q / 2 channels. After concatenating feature M1 and feature M2 into M3, a shuffle operation is performed to shuffle and recombine the concatenated information to obtain the final output feature M4, which is then fed into the next layer of the YOLOV8 neural network.

7. The method for detecting welded sheet metal parts based on improved image stitching and deep learning networks according to claim 1, characterized in that, The NWD-Inner-IoU loss function is: In the formula, m is the coefficient weight, which takes a value between 0 and 1; P is the predicted region, G is the labeled region, I represents the intersection area of ​​the predicted region and the labeled region, and U represents their union area. It is the bounding box A = (cx a ,cy a ,w a ,h a ), B = (cx b ,cy b ,w b ,h b The second-order Wasserstein distance of N; a N b The rectangle follows a Gaussian distribution, where cx is the x-coordinate of the center point, cy is the y-coordinate of the center point, w is the width of the rectangle, and h is the height of the rectangle.

8. A detection device for welded sheet metal parts based on improved image stitching and deep learning networks, characterized in that, The plate welded component testing device includes: The image stitching and fusion module is used to stitch together images of welded sheet metal parts captured simultaneously by a camera group using an improved SURF algorithm. After annotation, a stitched and fused image dataset is constructed. Specifically, it includes: By using a rectangular feature region constraint function, the SURF algorithm is used to limit the feature point extraction range to a rectangular area selected in the image of the welded sheet metal. A kd-tree is used to perform nearest neighbor search on the extracted feature points. During the nearest neighbor search, an improved Manhattan distance is used for distance measurement, and the measurement result is compared with a set threshold to obtain feature point pairs. The RANSAC algorithm is introduced to remove mismatched feature point pairs from the two images and solve for the homography matrix. The homography matrix is ​​then used to transform the two images to the same coordinate system for image fusion, resulting in a stitched fused image. The model improvement module is used to structurally improve the YOLOv8 neural network, forming an improved YOLOv8 neural network for target detection and recognition of sheet metal welded parts. Specifically, an EMA attention mechanism is added after the 1*1 Conv convolutional layer in the FasterNet Block of the FasterNet network to perform weighted averaging of historical data, forming a Faster-EMA improved C2f module. The Faster-EMA improved C2f module is used in layers 2, 4, 6, and 8 of the YOLOv8 neural network. Dense convolution SC and depthwise separable convolution DSC are mixed to obtain hybrid convolution GSConv, which is introduced into layers 16 and 19 of the YOLOv8 neural network. The model training module is used to propose a loss function based on normalized Wasserstein distance (NWD), and to obtain the NWD-Inner-IoU loss function by combining the NWD-based loss function and the Inner-IoU loss function. The stitched and fused image dataset is imported into the improved YOLOv8 neural network, and the improved YOLOv8 neural network is trained by combining the NWD-Inner-IoU loss function. The trained model is then used for the detection of welded parts of sheet metal.

9. A computer storage device, characterized in that, The computer contains a computer program, and the computer running the computer program can perform the inspection method for welded sheet metal parts as described in any one of claims 1-7.

10. An electronic device, characterized in that, include: The processor, the memory, and the computer program stored in the memory and capable of running on the processor, wherein the processor, when executing the program, implements the plate weldment inspection method as described in any one of claims 1-7.