Welding quality detection method based on three-dimensional convolutional neural network

Through the welding quality detection method based on three-dimensional convolutional neural network, the problems of low efficiency, unstable accuracy and poor real-time performance of traditional detection methods are solved, and real-time monitoring and efficient detection of welding quality are achieved.

CN120147223APending Publication Date: 2025-06-13TAIYUAN HEAVY IND
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510100204.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

Traditional welding quality detection methods rely on manual operation, low detection efficiency and unstable accuracy, making it difficult to meet the manufacturing needs of high precision and high efficiency, and have poor real-time performance, so it is impossible to detect defects in the welding process in a timely manner.

Method used

Welding quality detection method based on three-dimensional convolutional neural network is adopted, and welding video stream data is obtained, multi-frame welding image sequence is extracted, and a pre-trained three-dimensional convolutional neural network is used for real-time detection to realize real-time monitoring of welding quality.

Benefits of technology

Real-time monitoring of welding quality is achieved, and abnormal welding quality can be detected in a timely manner, which improves detection efficiency and accuracy, and avoids the shortcomings of manual inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147223A_ABST
    Figure CN120147223A_ABST
Patent Text Reader

Abstract

The invention discloses a welding quality detection method based on a three-dimensional convolutional neural network. The method comprises the following steps: acquiring a training data set; training a pre-constructed three-dimensional convolutional neural network by using the training data set; obtaining welding video stream data to be detected; extracting a to-be-detected multi-frame welding image sequence from to-be-detected welding video stream data, and inputting the to-be-detected multi-frame welding image sequence into the trained three-dimensional convolutional neural network to obtain a welding quality condition corresponding to the to-be-detected multi-frame welding image sequence output by the three-dimensional convolutional neural network and a probability value of the welding quality condition; and determining and outputting a final welding quality condition corresponding to the to-be-detected multi-frame welding image sequence according to the welding quality condition corresponding to the to-be-detected multi-frame welding image sequence output by the three-dimensional convolutional neural network, the probability value of the welding quality condition and a preset probability threshold value. According to the method, the welding quality condition can be monitored in real time, and welding quality abnormity can be found in time to make up for the abnormity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of welding quality detection, and particularly to a welding quality detection method based on a three-dimensional convolutional neural network. Background Art

[0002] In the current industrial manufacturing field, as an important metal processing equipment, the manufacturing quality and efficiency of a straightening machine directly affect the stability and cost-effectiveness of the entire production line. Welding, as a key link in the manufacturing process of the straightening machine, not only determines the performance and service life of the straightening machine, but also is related to the safety and reliability of the production line. Therefore, how to improve the intelligence and refinement level of straightening machine welding has become an important issue faced by the current manufacturing industry.

[0003] During the welding process of the straightening machine, welding quality detection is required. Traditional welding quality detection methods mainly rely on manual operation and empirical judgment, with low detection efficiency and unstable detection accuracy, making it difficult to meet the manufacturing requirements of high precision and high efficiency. Moreover, traditional weld quality inspection methods usually perform detection after welding is completed, with poor real-time performance, unable to obtain defects in the molten pool information during the welding process, and once welding defects occur, they cannot be discovered and remedied in time. Summary of the Invention

[0004] To solve some or all of the above-mentioned technical problems existing in the prior art, the present invention provides a welding quality detection method based on a three-dimensional convolutional neural network.

[0005] The technical solution of the present invention is as follows:

[0006] A welding quality detection method based on a three-dimensional convolutional neural network is provided, and the method includes:

[0007] Step 1, obtaining a training data set, where the training data includes a multi-frame welding image sequence extracted from welding video stream data and its corresponding welding quality situation, and the welding quality situation includes normal welding quality and abnormal welding quality;

[0008] Step 2, training a pre-constructed three-dimensional convolutional neural network using the training data set. The three-dimensional convolutional neural network includes a 3D residual block based on 3D convolution, a conditional bidirectional attention module based on an attention mechanism, and a classification network. The 3D residual block is used to extract three-dimensional features of the multi-frame welding image sequence, the conditional bidirectional attention module is used to perform feature extraction and feature fusion on the three-dimensional features output by the 3D residual block, and the classification network is used to classify according to the input features to obtain and output the welding quality situation and its probability value corresponding to the multi-frame welding image sequence;

[0009] Step 3, obtaining the welding video stream data to be detected;

[0010] Step 4: Extract a sequence of multiple welding images to be detected from the welding video stream data to be detected, and input the sequence of multiple welding images to be detected into the trained three-dimensional convolutional neural network to obtain the welding quality condition and its probability value corresponding to the sequence of multiple welding images to be detected output by the three-dimensional convolutional neural network;

[0011] Step 5: Determine and output the final welding quality condition corresponding to the sequence of multiple welding images to be detected according to the welding quality condition and its probability value corresponding to the sequence of multiple welding images to be detected output by the three-dimensional convolutional neural network and a preset probability threshold.

[0012] In some alternative embodiments, when obtaining training data, measure the geometric dimensions of the weld forming molten pool in each welding image in the sequence of multiple welding images to obtain the geometric dimensions of the weld forming molten pool in each welding image, and detect and judge the welding quality condition according to the change situation of the geometric dimensions of the weld forming molten pool in the sequence of multiple welding images. If the change of the geometric dimensions of the weld forming molten pool is normal, it is judged that the welding quality is normal; if the change of the geometric dimensions of the weld forming molten pool is abnormal, it is judged that the welding quality is abnormal.

[0013] In some alternative embodiments, the three-dimensional convolutional neural network includes: a three-dimensional convolutional unit, a three-dimensional batch normalization unit, an activation function unit, a three-dimensional max pooling unit, a first 3D residual block, a first conditional bidirectional attention module, a second 3D residual block, a second conditional bidirectional attention module, a third 3D residual block, a third conditional bidirectional attention module, a fourth 3D residual block, and a classification network, which are connected in sequence;

[0014] The three-dimensional convolutional unit is used to perform three-dimensional convolutional operations on the input to extract features;

[0015] The three-dimensional batch normalization unit is used to perform normalization processing on the input;

[0016] The activation function unit is used to perform non-linear transformation processing on the input by using an activation function;

[0017] The three-dimensional max pooling unit is used to perform max pooling operations on the input in the spatial and temporal dimensions;

[0018] The 3D residual block is used to extract features from the input through a combination of multiple convolutional operations and non-linear transformation processing;

[0019] The conditional bidirectional attention module is used to extract features and fuse the input in the channel and spatial dimensions based on the attention mechanism;

[0020] The classification network is used to classify the input into predefined categories to obtain and output the final classification result.

[0021] In some alternative embodiments, the 3D residual block includes: a first three-dimensional convolutional unit, a first three-dimensional batch normalization unit, a first ReLU activation function unit, a second three-dimensional convolutional unit, a second three-dimensional batch normalization unit, a second ReLU activation function unit, and a splicing unit connected in sequence. One input end of the splicing unit is connected to the output end of the second ReLU activation function unit, and the other input end of the splicing unit is jump-connected to the input end of the 3D residual block;

[0022] The three-dimensional convolutional unit is used to perform three-dimensional convolution operations on the input to extract features;

[0023] The three-dimensional batch normalization unit is used to perform normalization processing on the input;

[0024] The ReLU activation function unit is used to perform non-linear transformation processing on the input by using the ReLU activation function;

[0025] The splicing unit is used to perform splicing operations on the input features to obtain and output fused features.

[0026] In some alternative embodiments, the conditional bidirectional attention module includes: a channel attention module and a spatial attention module;

[0027] The input end of the channel attention module is connected to the input end of the conditional bidirectional attention module, the output end of the channel attention module is connected to the input end of the spatial attention module, and the output end of the spatial attention module is connected to the output end of the conditional bidirectional attention module;

[0028] The channel attention module is used to extract features from the input in the channel dimension;

[0029] The spatial attention module is used to extract features and perform feature fusion on the input in the spatial dimension to obtain and output attention features.

[0030] In some alternative embodiments, the channel attention module includes: a three-dimensional pooling unit, a third three-dimensional convolutional unit, a ReLU activation function unit, a fourth three-dimensional convolutional unit, and a Sigmoid activation function unit connected in sequence;

[0031] The three-dimensional pooling unit is used to perform pooling processing on the input;

[0032] The three-dimensional convolutional unit is used to perform three-dimensional convolution operations on the input to extract features;

[0033] The ReLU activation function unit is used to perform non-linear transformation processing on the input by using the ReLU activation function;

[0034] The Sigmoid activation function unit is used to perform non-linear transformation processing on the input using the Sigmoid activation function.

[0035] In some alternative embodiments, the spatial attention module includes: a temporal convolution unit, a spatial convolution unit, a first Softmax function unit, a second Softmax function unit, and a splicing unit;

[0036] The input ends of the temporal convolution unit and the spatial convolution unit are respectively connected to the output end of the channel attention module. The output end of the temporal convolution unit is connected to the input end of the first Softmax function unit. The output end of the spatial convolution unit is connected to the input end of the second Softmax function unit. The output ends of the first Softmax function unit and the second Softmax function unit are respectively connected to the two input ends of the splicing unit. The output end of the splicing unit is connected to the output end of the spatial attention module;

[0037] The temporal convolution unit is used to extract the features of the input in the temporal dimension;

[0038] The spatial convolution unit is used to extract the features of the input in the spatial dimension;

[0039] The Softmax function unit is used to convert the input into a probability distribution;

[0040] The splicing unit is used to perform a splicing operation on the input to obtain and output the attention features.

[0041] In some alternative embodiments, the classification network includes: a patch vector conversion unit, a Transformer layer, a first fully connected layer, and a second fully connected layer connected in sequence;

[0042] The patch vector conversion unit is used to split the input into smaller patches and assign an embedding vector to each patch;

[0043] The Transformer layer is used to capture the global dependencies in the input data using the self-attention mechanism;

[0044] The fully connected layer is used to perform non-linear transformation processing on the input features. The first fully connected layer is used to map the features output by the Transformer layer to a smaller dimensional space. The second fully connected layer is used to map the features to a dimensional space with the same number of predefined categories to obtain and output the final classification result.

[0045] In some alternative embodiments, the three-dimensional convolutional neural network is trained in the following manner:

[0046] Use a sequence of multiple frames of welding images in the training data set as the input of the three-dimensional convolutional neural network, and use the welding quality situation corresponding to the input sequence of multiple frames of welding images as the output to train the three-dimensional convolutional neural network.

[0047] In some alternative embodiments, the final welding quality situation corresponding to the sequence of multiple frames of welding images to be detected is determined by the following method:

[0048] Compare the probability value of welding quality abnormality corresponding to the sequence of multiple frames of welding images to be detected output by the three-dimensional convolutional neural network with a preset probability threshold;

[0049] If the probability value of welding quality abnormality is greater than the probability threshold, it is determined that the welding quality situation corresponding to the sequence of multiple frames of welding images to be detected is welding quality abnormality;

[0050] If the probability value of welding quality abnormality is not greater than the probability threshold, it is determined that the welding quality situation corresponding to the sequence of multiple frames of welding images to be detected is normal welding quality.

[0051] The main advantages of the technical solution of the present invention are as follows:

[0052] The welding quality detection method based on a three-dimensional convolutional neural network of the present invention applies deep learning technology and 3D vision technology to the welding quality detection process. By real-time acquiring a welding video stream during the welding process and using a pre-trained three-dimensional convolutional neural network for real-time detection, it can achieve real-time monitoring of the welding quality situation, and can timely detect and make up for it when welding quality abnormalities occur; moreover, there is no need for manual detection, the detection efficiency is relatively high, and the detection accuracy is relatively stable. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] The drawings described herein are used to provide a further understanding of the embodiments of the present invention, and constitute a part of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation of the present invention. In the drawings:

[0054] Figure 1 is a flowchart of a welding quality detection method based on a three-dimensional convolutional neural network provided by an embodiment of the present invention;

[0055] Figure 2 is a structural block diagram of a three-dimensional convolutional neural network provided by an embodiment of the present invention;

[0056] Figure 3 is Figure 2 a structural block diagram of a 3D residual block in the three-dimensional convolutional neural network shown;

[0057] Figure 4 is Figure 2Structural block diagram of the conditional bidirectional attention module in the three-dimensional convolutional neural network shown;

[0058] Figure 5 is Figure 2 Structural block diagram of the classification network in the three-dimensional convolutional neural network described. Specific implementation manners

[0059] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.

[0060] The technical solutions provided by the embodiments of the present invention will be described in detail below with reference to the drawings.

[0061] Refer to Figure 1-2 , the embodiments of the present invention provide a welding quality detection method based on a three-dimensional convolutional neural network, and the method includes the following steps 1-step 5:

[0062] Step 1, obtain a training data set;

[0063] In the embodiments of the present invention, the training data set includes a plurality of training data, and the training data includes a multi-frame welding image sequence extracted from the welding video stream data and its corresponding welding quality condition.

[0064] In the embodiments of the present invention, the welding quality condition includes normal welding quality and abnormal welding quality.

[0065] It should be noted that the number of training data included in the training data set is specifically set according to the actual situation.

[0066] Step 2, use the training data set to train a pre-constructed three-dimensional convolutional neural network;

[0067] In the embodiments of the present invention, the three-dimensional convolutional neural network includes a 3D residual block based on 3D convolution, a conditional bidirectional attention module based on the attention mechanism, and a classification network. The 3D residual block is used to extract the three-dimensional features of the multi-frame welding image sequence, the conditional bidirectional attention module is used to perform feature extraction and feature fusion on the three-dimensional features output by the 3D residual block, and the classification network is used to classify according to the input features to obtain and output the welding quality condition and its probability value corresponding to the multi-frame welding image sequence.

[0068] Step 3, obtain the welding video stream data to be detected;

[0069] In the embodiments of the present invention, according to the actual detection requirements, welding video stream data to be detected is obtained.

[0070] Step 4: Extract a sequence of multiple frames of welding images to be detected from the welding video stream data to be detected, and input the sequence of multiple frames of welding images to be detected into the trained three-dimensional convolutional neural network to obtain the welding quality situation and its probability value corresponding to the sequence of multiple frames of welding images to be detected output by the three-dimensional convolutional neural network.

[0071] In the embodiments of the present invention, when extracting a sequence of multiple frames of welding images to be detected from the welding video stream data to be detected, sequential extraction is performed according to the time sequence of the video stream data, and multiple groups of sequences of multiple frames of welding images to be detected are obtained in sequence. For example, if the sequence of multiple frames of welding images is an N-frame sequence of welding images, then the 1st to Nth frames of welding images in the welding video stream data to be detected are extracted as the first group of N-frame sequences of welding images, and the (N + 1)th to 2Nth frames of welding images in the welding video stream data to be detected are extracted as the second group of N-frame sequences of welding images, and extraction is performed in this way in sequence until the extraction process of the entire welding video stream data to be detected is completed.

[0072] Step 5: Determine and output the final welding quality situation corresponding to the sequence of multiple frames of welding images to be detected according to the welding quality situation and its probability value corresponding to the sequence of multiple frames of welding images to be detected output by the three-dimensional convolutional neural network and a preset probability threshold.

[0073] In the embodiments of the present invention, the probability threshold is specifically set according to the actual situation. The probability threshold is used to compare with the probability value of the welding quality situation output by the three-dimensional convolutional neural network to determine whether to adopt the prediction result of the three-dimensional convolutional neural network, and further determine the final welding quality situation.

[0074] The welding quality detection method based on a three-dimensional convolutional neural network provided by the embodiments of the present invention applies deep learning technology and 3D vision technology to the welding quality detection process. By obtaining the welding video stream in real time during the welding process and using the pre-trained three-dimensional convolutional neural network for real-time detection, it can realize the real-time monitoring of the welding quality situation, and can timely detect and make up for the abnormal welding quality; moreover, there is no need for manual detection, the detection efficiency is relatively high, and the detection accuracy is relatively stable.

[0075] Further, in the embodiments of the present invention, in Step 1, when obtaining the training data set, the welding video stream data is obtained by actually shooting the welding process or selected from existing video data.

[0076] In the embodiments of the present invention, in Step 1, multiple groups of sequences of multiple frames of welding images are sequentially extracted from the welding video stream data, and the welding quality situation corresponding to each group of sequences of multiple frames of welding images is detected and judged.

[0077] It should be noted that there is no overlapping part in the multi-frame image sequences of different groups extracted.

[0078] In an embodiment of the present invention, in step 1, one or more welding video stream data can be obtained, and multiple training data are obtained from each welding video stream data, so as to obtain a training data set.

[0079] In an embodiment of the present invention, in step 1, the welding quality situation corresponding to each group of multi-frame welding image sequences is detected and judged by means of manual detection and judgment. Among them, if there is already corresponding welding quality situation data, it can be directly used.

[0080] Furthermore, in an alternative embodiment of the embodiment of the present invention, in step 1, when obtaining training data, by measuring the geometric dimensions of the weld forming molten pool in each frame of the multi-frame welding image sequence, the geometric dimensions of the weld forming molten pool in each frame of the welding image are obtained, and the welding quality situation is detected and judged according to the change situation of the geometric dimensions of the weld forming molten pool in the multi-frame welding image sequence. If the change of the geometric dimensions of the weld forming molten pool is normal, it is judged that the welding quality is normal; if the change of the geometric dimensions of the weld forming molten pool is abnormal, it is judged that the welding quality is abnormal.

[0081] Reference Figure 2 , furthermore, in an alternative embodiment of the embodiment of the present invention, in order to implement the functions of the three-dimensional convolutional neural network defined above, the three-dimensional convolutional neural network includes: a three-dimensional convolutional unit, a three-dimensional batch normalization unit, an activation function unit, a three-dimensional max pooling unit, a first 3D residual block, a first conditional bidirectional attention module, a second 3D residual block, a second conditional bidirectional attention module, a third 3D residual block, a third conditional bidirectional attention module, a fourth 3D residual block and a classification network connected in sequence;

[0082] The three-dimensional convolutional unit is used to perform three-dimensional convolutional operations on the input to extract features;

[0083] The three-dimensional batch normalization unit is used to perform normalization processing on the input to accelerate the training speed, alleviate the problems of gradient disappearance and explosion, and improve the generalization ability of the neural network;

[0084] The activation function unit is used to perform non-linear transformation processing on the input by using an activation function to improve the non-linear expression ability of the neural network, enable the neural network to learn complex patterns, and avoid the problem of gradient disappearance;

[0085] The three-dimensional max pooling unit is used to perform max pooling operations on the input in the spatial and temporal dimensions to reduce the feature size while retaining important features;

[0086] The 3D residual block is used to extract features from the input through a combination of multiple convolutional operations and non-linear transformation processes, so as to improve the feature extraction ability and the stability of the training process while maintaining the network depth;

[0087] The conditional bidirectional attention module is used to extract and fuse features from the input in the channel and spatial dimensions based on the attention mechanism, so as to enhance the learning of important features, suppress irrelevant features, and improve the feature expression ability and performance of the neural network;

[0088] The classification network is used to classify the input into predefined categories and obtain and output the final classification result.

[0089] In the embodiment of the present invention, by adopting the three-dimensional convolutional neural network of the above structure type, features can be effectively extracted from the input welding image sequence and classified. By combining different processing units, 3D residual blocks and conditional bidirectional attention modules, rich spatio-temporal information and context relationships can be captured, thereby improving the classification performance.

[0090] In the embodiment of the present invention, the 3D residual blocks form a deep residual network, and the conditional bidirectional attention modules form a self-attention network.

[0091] In the embodiment of the present invention, based on the above-set welding quality conditions, the predefined categories include: normal welding quality and abnormal welding quality, and the final classification result output by the classification network is the welding quality condition and its probability value.

[0092] In the embodiment of the present invention, in the three-dimensional convolutional neural network, the convolution kernel of the convolution operation in the three-dimensional convolution unit is set to 7x7x7, and the activation function adopted in the activation function unit is the ReLU activation function.

[0093] It should be noted that in the appendix Figure 2 Conv3D represents a three-dimensional convolution unit, BatchNorm3d represents a three-dimensional batch normalization unit, ReLU represents an activation function unit, MaxPool3d represents a three-dimensional max pooling unit, Res3DBlock represents a 3D residual block, CBAM represents a conditional bidirectional attention module, and Classifier represents a classification network.

[0094] Reference Figure 3, Further, in an alternative embodiment of the present invention, in order to implement the functions of the 3D residual block defined above, the 3D residual block includes: a first 3D convolutional unit, a first 3D batch normalization unit, a first ReLU activation function unit, a second 3D convolutional unit, a second 3D batch normalization unit, a second ReLU activation function unit, and a splicing unit connected in sequence. One input end of the splicing unit is connected to the output end of the second ReLU activation function unit, and the other input end of the splicing unit is jump-connected to the input end of the 3D residual block;

[0095] The 3D convolutional unit is used to perform 3D convolution operations on the input to extract features;

[0096] The 3D batch normalization unit is used to perform normalization processing on the input to accelerate the training speed, alleviate the problems of gradient vanishing and explosion, and improve the generalization ability of the neural network;

[0097] The ReLU activation function unit is used to perform non-linear transformation processing on the input using the ReLU activation function to improve the non-linear expression ability of the neural network, enable the neural network to learn complex patterns, and avoid the problem of gradient vanishing;

[0098] The splicing unit is used to perform splicing operations on the input features to obtain and output fused features.

[0099] In the embodiment of the present invention, by adopting the 3D residual block of the above structure type, in the 3D residual block, the corresponding input is added back to the original input after two convolution operations and appropriate transformation operations to form a shortcut path. This design can allow the gradient to directly pass through the network, avoiding the problem of gradient vanishing in deep neural networks, effectively improving the training efficiency and stability of the neural network, and further improving the prediction and classification accuracy of the trained neural network.

[0100] In the embodiment of the present invention, in the 3D residual block, the convolution kernel of the convolution operation in the 3D convolutional unit is set to 3x3x3.

[0101] It should be noted that in the embodiment of the present invention, the structures and functions of the first 3D residual block to the fourth 3D residual block are the same.

[0102] It should be noted that in the appendix Figure 3 , Conv3D represents the 3D convolutional unit, BatchNorm3d represents the 3D batch normalization unit, ReLU represents the ReLU activation function unit, and + represents the splicing unit.

[0103] Reference Figure 4, Further, in an alternative embodiment of the present invention, in order to implement the functions of the conditional bidirectional attention module defined above, the conditional bidirectional attention module includes: a channel attention module and a spatial attention module;

[0104] The input end of the channel attention module is connected to the input end of the conditional bidirectional attention module, the output end of the channel attention module is connected to the input end of the spatial attention module, and the output end of the spatial attention module is connected to the output end of the conditional bidirectional attention module;

[0105] The channel attention module is used to extract features from the input in the channel dimension;

[0106] The spatial attention module is used to extract features and perform feature fusion on the input in the spatial dimension, and obtain and output attention features.

[0107] In the embodiments of the present invention, the channel attention module is based on the channel attention mechanism, aiming to focus on the importance of different channels in the input. Since different channels may represent different features or semantic information, the channel attention module adaptively assigns higher weights to important channels while suppressing irrelevant channels, which can help the neural network better focus on key features and improve the feature representation ability. The spatial attention module is based on the spatial attention mechanism, mainly focusing on the spatial positions of the feature maps. At different spatial positions, some regions may be more important for the current task. The spatial attention module assigns different weights to the pixel points at different spatial positions, enabling the neural network to pay more attention to the regions containing important information while ignoring the background or other irrelevant regions.

[0108] In the embodiments of the present invention, by extracting the attention features of channels and spaces and combining these features, the feature expression ability and performance of the neural network can be improved.

[0109] Reference Figure 4 , Further, in an alternative embodiment of the present invention, the channel attention module includes: a three-dimensional pooling unit, a third three-dimensional convolutional unit, a ReLU activation function unit, a fourth three-dimensional convolutional unit, and a Sigmoid activation function unit connected in sequence;

[0110] The three-dimensional pooling unit is used to perform pooling processing on the input to reduce the size of the input and achieve dimensionality reduction of the input while retaining important information;

[0111] The three-dimensional convolutional unit is used to perform three-dimensional convolutional operations on the input to extract features;

[0112] The ReLU activation function unit is used to perform non-linear transformation processing on the input using the ReLU activation function to improve the non-linear expression ability of the neural network;

[0113] The Sigmoid activation function unit is used to perform a non-linear transformation process on the input using the Sigmoid activation function.

[0114] Furthermore, in the embodiments of the present invention, the spatial attention module includes: a temporal convolution unit, a spatial convolution unit, a first Softmax function unit, a second Softmax function unit, and a splicing unit;

[0115] The input ends of the temporal convolution unit and the spatial convolution unit are respectively connected to the output end of the channel attention module. The output end of the temporal convolution unit is connected to the input end of the first Softmax function unit. The output end of the spatial convolution unit is connected to the input end of the second Softmax function unit. The output ends of the first Softmax function unit and the second Softmax function unit are respectively connected to the two input ends of the splicing unit. The output end of the splicing unit is connected to the output end of the spatial attention module;

[0116] The temporal convolution unit is used to extract the features of the input in the temporal dimension;

[0117] The spatial convolution unit is used to extract the features of the input in the spatial dimension;

[0118] The Softmax function unit is used to convert the input into a probability distribution;

[0119] The splicing unit is used to perform a splicing operation on the input to obtain and output the attention features.

[0120] In the embodiments of the present invention, by adopting the conditional bidirectional attention module with the above structure type, it is possible to effectively extract the attention features of the channel and the space, and combine these features to improve the feature expression ability and performance of the neural network, thereby improving the prediction classification accuracy of the trained neural network.

[0121] It should be noted that, in the embodiments of the present invention, the output end of the spatial attention module is the output end of the conditional bidirectional attention module.

[0122] It should be noted that, in the embodiments of the present invention, the structures and functions of the first conditional bidirectional attention module to the third conditional bidirectional attention module are the same.

[0123] It should be noted that in the appendix Figure 4Among them, Channel Attention represents the channel attention module, SpatioTemporal Attention represents the spatio-temporal attention module, Pool3D represents the 3D pooling unit, Conv3D represents the 3D convolutional unit, ReLU represents the ReLU activation function unit, Sigmoid represents the Sigmoid activation function unit, temporal_conv represents the temporal convolutional unit, spatial_conv represents the spatial convolutional unit, Softmax represents the Softmax function unit, and + represents the concatenation unit.

[0124] Reference Figure 5 Furthermore, in an alternative embodiment of the present invention, in order to implement the functions of the classification network defined above, the classification network includes: a patch vector conversion unit, a Transformer layer, a first fully connected layer, and a second fully connected layer connected in sequence;

[0125] The patch vector conversion unit is used to divide the input into smaller patches and assign an embedding vector to each patch for subsequent processing by the Transformer layer;

[0126] The Transformer layer is used to capture global dependencies in the input data using the self-attention mechanism;

[0127] The fully connected layers are used to perform non-linear transformation processing on the input features. The first fully connected layer is used to map the features output by the Transformer layer to a smaller dimensional space, and the second fully connected layer is used to map the features to a dimensional space with the same number of dimensions as the predefined number of classes to obtain and output the final classification result.

[0128] In the embodiment of the present invention, the Transformer is composed of multiple encoder layers, and each encoder layer includes a multi-head self-attention sub-layer and a feed-forward neural network sub-layer. The multi-head self-attention mechanism allows multiple attention heads to work in parallel, so as to be able to capture input information in different aspects.

[0129] In the embodiment of the present invention, by adopting the classification network with the above structure type, the prediction classification accuracy of the trained neural network can be effectively improved.

[0130] It should be noted that in the appendix Figure 5 Patch embedding represents the patch vector conversion unit, Transformer represents the Transformer layer, and Linear represents the fully connected layer.

[0131] Furthermore, in an alternative embodiment of the present invention, the 3D convolutional neural network is trained in the following manner:

[0132] Use a sequence of multiple frames of welding images in the training data set as the input of a three-dimensional convolutional neural network, and use the welding quality corresponding to the input sequence of multiple frames of welding images as the output to train the three-dimensional convolutional neural network.

[0133] Further, in an embodiment of the present invention, using a sequence of multiple frames of welding images in the training data set as the input of a three-dimensional convolutional neural network, and using the welding quality corresponding to the input sequence of multiple frames of welding images as the output to train the three-dimensional convolutional neural network further includes the following steps 201-step 203:

[0134] Step 201, input the sequences of multiple frames of welding images in all training data into the three-dimensional convolutional neural network in batches, and obtain the welding quality corresponding to each sequence of multiple frames of welding images output by the three-dimensional convolutional neural network;

[0135] In an embodiment of the present invention, the sequence of multiple frames of welding images in the training data is input from the input end of the three-dimensional convolutional neural network, sequentially processed through each layer in the three-dimensional convolutional neural network, and output from the output end of the three-dimensional convolutional neural network. The information output by the three-dimensional convolutional neural network is the welding quality corresponding to the input sequence of multiple frames of welding images and its probability value.

[0136] In an embodiment of the present invention, the parameters of the three-dimensional convolutional neural network are initialization parameters, and during the training process, the parameters of the three-dimensional convolutional neural network are continuously updated and learned.

[0137] Step 202, calculate a preset loss function according to the welding quality in the training data and the prediction result of the welding quality corresponding to the training data output by the three-dimensional convolutional neural network;

[0138] In an embodiment of the present invention, the loss function is specifically set according to the actual situation. For example, a cross-entropy loss function is used.

[0139] Step 203, determine whether the preset training stop condition is reached. If so, use the current three-dimensional convolutional neural network as the trained three-dimensional convolutional neural network. If not, update the parameters of the three-dimensional convolutional neural network using the loss function and return to step 201.

[0140] In an embodiment of the present invention, the training stop condition is specifically set according to the actual situation. For example, the training iteration number reaches the set iteration number or the optimization index reaches the set threshold. Among them, the loss function can be used as the optimization index.

[0141] Further, in an embodiment of the present invention, the random gradient descent method is used to train and update the parameters of the three-dimensional convolutional neural network.

[0142] Specifically, the parameters of the three-dimensional convolutional neural network are updated using the following formula:

[0143]

[0144] Among them, θ represents the set of parameters of the three-dimensional convolutional neural network, Δ[·] represents the optimizer, η represents the learning rate, and Loss represents the loss function. Among them, the optimizer is specifically set according to the actual situation, such as Adam, and the learning rate is preset to control the speed of parameter update of the three-dimensional convolutional neural network.

[0145] Furthermore, in the embodiment of the present invention, in step 5, according to the welding quality situation and its probability value corresponding to the multi-frame welding image sequence to be detected output by the three-dimensional convolutional neural network, and a preset probability threshold, the final welding quality situation corresponding to the multi-frame welding image sequence to be detected is determined in the following manner:

[0146] Compare the probability value of welding quality abnormality corresponding to the multi-frame welding image sequence to be detected output by the three-dimensional convolutional neural network with the preset probability threshold;

[0147] If the probability value of welding quality abnormality is greater than the probability threshold, it is determined that the welding quality situation corresponding to the multi-frame welding image sequence to be detected is welding quality abnormality;

[0148] If the probability value of welding quality abnormality is not greater than the probability threshold, it is determined that the welding quality situation corresponding to the multi-frame welding image sequence to be detected is normal welding quality.

[0149] In the embodiment of the present invention, by obtaining welding video stream data in real time during the welding process, extracting a multi-frame welding image sequence from the welding video stream data in real time and inputting it into the three-dimensional convolutional neural network, obtaining the welding quality situation and its probability value corresponding to the multi-frame welding image sequence output by the three-dimensional convolutional neural network, and determining whether there is a welding quality abnormality in the current welding process according to the welding quality situation and its probability value output by the three-dimensional convolutional neural network, real-time monitoring of the welding quality situation can be achieved, and timely discovery and compensation can be made when there is a welding quality abnormality.

[0150] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0151] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A welding quality detection method based on a three-dimensional convolutional neural network, characterized in that: include: Step 1, obtaining a training data set, the training data including a multi-frame welding image sequence extracted from welding video stream data and its corresponding welding quality conditions, wherein the welding quality conditions include normal welding quality and abnormal welding quality; Step 2, using the training data set to train a pre-built three-dimensional convolutional neural network, the three-dimensional convolutional neural network includes a 3D residual block based on 3D convolution, a conditional bidirectional attention module based on an attention mechanism, and a classification network, the 3D residual block is used to extract three-dimensional features of a multi-frame welding image sequence, the conditional bidirectional attention module is used to extract and fuse the three-dimensional features output by the 3D residual block, and the classification network is used to classify according to the input features, and obtain and output the welding quality situation and its probability value corresponding to the multi-frame welding image sequence; Step 3, obtaining welding video stream data to be detected; Step 4, extracting a multi-frame welding image sequence to be detected from the welding video stream data to be detected, inputting the multi-frame welding image sequence to be detected into the trained three-dimensional convolutional neural network, and obtaining the welding quality status and probability value corresponding to the multi-frame welding image sequence to be detected output by the three-dimensional convolutional neural network; Step 5: Determine and output the final welding quality status corresponding to the multi-frame welding image sequence to be detected according to the welding quality status and probability value corresponding to the multi-frame welding image sequence to be detected output by the three-dimensional convolutional neural network, and the preset probability threshold.

2. The welding quality detection method based on three-dimensional convolutional neural network according to claim 1 is characterized in that: When acquiring training data, the geometric dimensions of the weld forming molten pool in each welding image in a multi-frame welding image sequence are measured to obtain the geometric dimensions of the weld forming molten pool in each welding image frame, and the welding quality is detected and judged according to the changes in the geometric dimensions of the weld forming molten pool in the multi-frame welding image sequence. If the changes in the geometric dimensions of the weld forming molten pool are normal, the welding quality is judged to be normal; if the changes in the geometric dimensions of the weld forming molten pool are abnormal, the welding quality is judged to be abnormal.

3. The welding quality detection method based on three-dimensional convolutional neural network according to claim 1 is characterized in that: The three-dimensional convolutional neural network includes: a three-dimensional convolution unit, a three-dimensional batch normalization unit, an activation function unit, a three-dimensional maximum pooling unit, a first 3D residual block, a first conditional bidirectional attention module, a second 3D residual block, a second conditional bidirectional attention module, a third 3D residual block, a third conditional bidirectional attention module, a fourth 3D residual block and a classification network connected in sequence; The 3D convolution unit is used to perform a 3D convolution operation on the input to extract features; The three-dimensional batch normalization unit is used to normalize the input; The activation function unit is used to perform nonlinear transformation processing on the input using the activation function; The three-dimensional maximum pooling unit is used to perform maximum pooling operations on the input in spatial and temporal dimensions; The 3D residual block is used to extract features from the input through a combination of multiple convolution operations and nonlinear transformation processing; The conditional bidirectional attention module is used to extract and fuse input features in channel and spatial dimensions based on the attention mechanism; The classification network is used to classify the input into predefined categories and obtain and output the final classification result.

4. The welding quality detection method based on three-dimensional convolutional neural network according to claim 3 is characterized in that: The 3D residual block includes: a first three-dimensional convolution unit, a first three-dimensional batch normalization unit, a first ReLU activation function unit, a second three-dimensional convolution unit, a second three-dimensional batch normalization unit, a second ReLU activation function unit and a splicing unit connected in sequence, one input end of the splicing unit is connected to the output end of the second ReLU activation function unit, and the other input end of the splicing unit is jump-connected to the input end of the 3D residual block; The 3D convolution unit is used to perform a 3D convolution operation on the input to extract features; The three-dimensional batch normalization unit is used to normalize the input; The ReLU activation function unit is used to perform nonlinear transformation processing on the input using the ReLU activation function; The splicing unit is used to perform splicing operations on the input features to obtain and output fused features.

5. The welding quality detection method based on three-dimensional convolutional neural network according to claim 4 is characterized in that: The conditional bidirectional attention module includes: a channel attention module and a spatial attention module; The input end of the channel attention module is connected to the input end of the conditional bidirectional attention module, the output end of the channel attention module is connected to the input end of the spatial attention module, and the output end of the spatial attention module is connected to the output end of the conditional bidirectional attention module; The channel attention module is used to extract features from the input in the channel dimension; The spatial attention module is used to extract and fuse features of the input in the spatial dimension to obtain and output attention features.

6. The welding quality detection method based on three-dimensional convolutional neural network according to claim 5 is characterized in that: The channel attention module includes: a three-dimensional pooling unit, a third three-dimensional convolution unit, a ReLU activation function unit, a fourth three-dimensional convolution unit and a Sigmoid activation function unit connected in sequence; The three-dimensional pooling unit is used to perform pooling on the input; The 3D convolution unit is used to perform a 3D convolution operation on the input to extract features; The ReLU activation function unit is used to perform nonlinear transformation processing on the input using the ReLU activation function; The Sigmoid activation function unit is used to perform nonlinear transformation processing on the input using the Sigmoid activation function.

7. The welding quality detection method based on three-dimensional convolutional neural network according to claim 6 is characterized in that: The spatial attention module includes: a temporal convolution unit, a spatial convolution unit, a first Softmax function unit, a second Softmax function unit and a splicing unit; The input ends of the temporal convolution unit and the spatial convolution unit are respectively connected to the output end of the channel attention module, the output end of the temporal convolution unit is connected to the input end of the first Softmax function unit, the output end of the spatial convolution unit is connected to the input end of the second Softmax function unit, the output end of the first Softmax function unit and the output end of the second Softmax function unit are respectively connected to the two input ends of the splicing unit, and the output end of the splicing unit is connected to the output end of the spatial attention module; The temporal convolution unit is used to extract input features in the time dimension; The spatial convolution unit is used to extract input features in the spatial dimension; The Softmax function unit is used to convert the input into a probability distribution; The concatenation unit is used to concatenate the input to obtain and output the attention features.

8. The welding quality detection method based on three-dimensional convolutional neural network according to claim 7 is characterized in that: The classification network includes: a patch vector conversion unit, a Transformer layer, a first fully connected layer and a second fully connected layer connected in sequence; The patch vector conversion unit is used to split the input into smaller patches and assign an embedding vector to each patch; The Transformer layer is used to capture global dependencies in input data using a self-attention mechanism; The fully connected layer is used to perform nonlinear transformation on the input features, the first fully connected layer is used to map the features output by the Transformer layer to a smaller dimensional space, and the second fully connected layer is used to map the features to a dimensional space with the same number of predefined categories to obtain and output the final classification result.

9. The welding quality detection method based on three-dimensional convolutional neural network according to claim 1 is characterized in that: The three-dimensional convolutional neural network is trained in the following way: A multi-frame welding image sequence in the training data in the training data set is used as the input of the three-dimensional convolutional neural network, and the welding quality conditions corresponding to the input multi-frame welding image sequence are used as the output to train the three-dimensional convolutional neural network.

10. The welding quality detection method based on three-dimensional convolutional neural network according to claim 1 is characterized in that: The final welding quality corresponding to the multi-frame welding image sequence to be inspected is determined by the following method: Compare the probability value of welding quality abnormality corresponding to the multi-frame welding image sequence to be detected output by the three-dimensional convolutional neural network with a preset probability threshold; If the probability value of abnormal welding quality is greater than the probability threshold, the welding quality situation corresponding to the multi-frame welding image sequence to be detected is determined to be abnormal welding quality; If the probability value of abnormal welding quality is not greater than the probability threshold, the welding quality corresponding to the multi-frame welding image sequence to be detected is determined to be normal.