A high-precision and high-quality vision-guided robot automatic welding method based on deep learning
Through an improved semi-supervised segmentation network based on deep learning, weld location and molten pool centroid, and real-time adjustment of welding gun position and welding parameters, the problems of long welding assistance time and high technical requirements are solved, and an efficient and automated welding process is achieved.
Patent Information
- Application Number
- CN202510378119.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-03-28
AI Technical Summary
The existing welding robot technology has problems such as long welding assistance time, high labor intensity for workers, high technical requirements, and difficult to guarantee welding quality and real-time performance.
Adopting an improved semi-supervised segmentation network based on deep learning, the welding process is automated and intelligent by identifying the weld position, the center of mass of the molten pool and the number of splashes, and adjusting the welding gun position and welding parameters in real time.
It improves welding quality and efficiency, reduces workers' labor intensity, enhances the automation and intelligence of the welding process, and solves the problems of long welding assistance time and high technical requirements.
Smart Images

Figure CN119871466B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of robotic welding, and particularly relates to a high-precision and high-quality vision-guided robotic automatic welding method based on deep learning. Background Technique
[0002] Before using a welding robot for welding operations, technicians need to teach and program the welding path of the welding robot according to the weld path. The "teach-playback" mode and offline programming (OLP) are still the two main working modes of modern welding robots. When replacing welding workpieces, the "teach-playback" mode and offline programming (OLP) welding robots need to be re-taught and programmed. Both teaching and offline programming require professional technicians to spend a lot of time setting up the welding robot before welding. Although it can improve welding quality and reduce the labor intensity of workers to a certain extent, the welding auxiliary time is long and the requirements for technical workers are relatively high. Not only do technical workers need to have welding skill knowledge and rich experience, but also they need to have robot operation and programming skills. These weaknesses seriously hinder the wide application of welding robots. Therefore, in order to meet the development needs of modern industrial automated welding, it is necessary to study the real-time tracking technology of welding robot welds.
[0003] Prior art 1 proposed a method for robot trajectory tracking and deviation correction during welding CN116512270A, which uses a structured light scanning device to obtain the welding trajectory data on the workpiece in real time and the robot control end to synchronously obtain the real-time control data during the workpiece welding process to achieve the deviation correction of the welding trajectory. However, this method realizes the deviation correction of the welding trajectory on the basis of a preliminary planning of the welding trajectory, and the system flexibility is poor.
[0004] Prior art 2 is a vision control system for a welding robot based on deep learning CN116587288A. This method combines a deep learning algorithm with visual detection to identify welds, including the identification of weld types and weld positions, and adopts different welding processes for different types of welds. At the same time, during the welding process, the time series information and spatial series information of the welding position of the welding torch head are continuously learned to optimize the welding process of the welding robot. The method proposed in this patent includes modules such as a weld trajectory determination module, a movement parameter determination module, and a welding parameter determination module. The combination of multiple models has a high computational complexity and limited real-time performance.
[0005] Prior Art 3: A Gas Shielded Welding Process Parameter Optimization System and Method (CN116511652A). By processing the welding process images, the data of the molten pool area in the welding images is obtained, and the parameter optimization model is optimized. The welding process is monitored in real time, and at the same time, the current welding parameters are optimized and adjusted according to the parameter optimization model. This method has high requirements for the quality of training data, the accuracy of sensors, and the anti-interference ability, and is insufficient in dealing with abnormal working conditions.
[0006] Prior Art 4: An Automatic Welding Monitoring System, Method, Equipment and Readable Storage Medium (CN116586719A), including a welding control module, a welding quality detection module, and a welding adjustment module. The welding parameters are selected by the welding control module to weld the workpiece. A variety of sensor data is integrated in the welding quality detection module, including welding torch sensor data, visual sensor data, ultrasonic sensor data, and infrared thermal imager data. By fusing and processing the variety of sensor data, it is judged whether the welding data needs to be adjusted and how to adjust it. Finally, the welding parameters are adjusted in real time by the welding adjustment module. The method proposed in this patent has complex equipment, high cost, and high later maintenance and repair costs.
[0007] Since the visual sensor is installed at the end of the welding torch, the weld image it recognizes is the image of the previous moment. During the welding process, due to factors such as robot errors and workpiece deformation, the weld position recognized by the visual sensor is not accurate enough.
[0008] During the welding process, due to improper selection of welding parameters and poor weld processing accuracy, etc., the amount of spatter increases. The increase in spatter affects the surface quality after welding on the one hand, and seriously affects the environment and the health of the welding operator on the other hand. On the other hand, the spatter increases the stress concentration at the weld, affecting the service life of the welded part. Therefore, there is an urgent need for a high-precision and high-quality vision-guided robot automatic welding method based on deep learning. Summary of the Invention
[0009] To solve the above technical problems, the present invention proposes a high-precision and high-quality vision-guided robot automatic welding method based on deep learning, which can realize the automation and intelligence of the welding process, improve the welding quality, and reduce the labor intensity of workers.
[0010] The present invention provides a high-precision and high-quality vision-guided robot automatic welding method based on deep learning, including:
[0011] Obtain the image to be recognized;
[0012] Input the image to be recognized into the segmentation network model to identify the weld position, the centroid of the molten pool, and the number of spatter. Among them, the segmentation network model is an improved semi-supervised segmentation network, and the improved semi-supervised segmentation network is composed of the encoder and decoder of the optimized teacher-student model;
[0013] Send the weld position to the robot controller to adjust the position of the welding torch in real time and guide the robot to perform automatic welding.
[0014] Optionally, the encoder and decoder of the optimized teacher-student model include: improving the encoder and decoder by combining dilated convolution, multi-scale feature fusion, and residual network. In the Layer2, Layer3, and Layer4 of the encoder structure, dilated convolution is used instead of the traditional stride convolution. At the same time, the features output by the Layer1, Layer2, Layer3, and Layer4 are adjusted by the convolutional layer to adjust the number of channels and finally stitched and output. Improve the batch normalization layer in the encoder and decoder, improve the loss function, and replace the ResNet network structure in the decoder backbone network with the Xception network structure.
[0015] Optionally, the improved encoder structure includes: several convolutional layers, a max pooling layer, and several bottleneck structure dilated convolutional layers. The improved decoder structure includes: replacing the ResNet network structure in the decoder backbone network with the Xception network structure;
[0016] The convolutional layer is used to extract image features and obtain feature maps;
[0017] The max pooling layer is used to reduce the spatial size of the feature map;
[0018] The bottleneck structure dilated convolutional layer is used to expand the receptive field, maintain the resolution of the feature map, reduce parameters and computational complexity without increasing the number of model parameters.
[0019] Optionally, the improved decoder includes: atrous spatial pyramid pooling, a head structure, an upsampling layer, convolutional layers, a concatenation layer, and an output layer;
[0020] The atrous spatial pyramid pooling is used to introduce multiple parallel branches to capture context information at different scales and obtain multi-level features;
[0021] The head structure is used to prevent overfitting;
[0022] The upsampling is used to fuse high-level features with low-level features and optimize the accuracy of the segmentation boundary;
[0023] The convolutional layer is used to extract multi-scale features of the input image;
[0024] The splicing layer is used to splice the upsampling result and the convolution result;
[0025] The output layer is used to output the splicing result.
[0026] Optionally, the improved batch normalization layer includes: replacing the original batch normalization layer with a synchronous batch normalization method.
[0027] Optionally, the improved loss function includes: improving the original loss function by combining the affinity field loss and the weighted loss fusion method to obtain the improved loss function;
[0028] The improved loss function is:
[0029]
[0030] Wherein, is the improved loss function, is the supervised loss, is the unsupervised loss, is the contrastive loss, , and are the weight coefficients of the corresponding items, is the affinity field loss.
[0031] Optionally, sending the weld position to the robot controller and adjusting the welding torch position in real time includes:
[0032] Sending the weld position to the robot controller for welding. During welding, use the improved semi-supervised segmentation network to identify the number of spatter and the centroid of the molten pool in the image, compare the centroid position of the molten pool with the weld position, calculate the deviation, and use the deviation as the weld position adjustment amount for the next moment.
[0033] Optionally, sending the weld position to the robot controller for welding includes:
[0034] Converting the coordinates of the identified weld position according to the transformation matrix to obtain the coordinates in the robot coordinate system;
[0035] Welding according to the coordinates in the robot coordinate system.
[0036] Optionally, the transformation matrix is:
[0037]
[0038] Wherein, is the transformation matrix from the workpiece coordinate system to the base coordinate system of the six-axis robotic arm, is the transformation matrix from the end - effector coordinate system of the six - axis robotic arm to the base coordinate system of the six - axis robotic arm. is the transformation matrix from the workpiece coordinate system to the camera coordinate system. is the transformation matrix from the end - effector coordinate system of the six - axis robotic arm to the camera coordinate system.
[0039] Optionally, sending the weld position to the robot controller to adjust the position of the welding torch in real - time, and further includes: comparing the current spatter quantity with the spatter quantity at the previous moment, and adjusting the welding process parameters.
[0040] Comparing the current spatter quantity with the spatter quantity at the previous moment and adjusting the welding process parameters includes:
[0041] Calculating the spatter quantity of the i - th frame and the (i + 1) - th frame of the image. If the spatter quantity decreases, adjust the welding process parameters forward; otherwise, adjust the welding inclination angle and nozzle height backward.
[0042] Compared with the prior art, the present invention has the following advantages and technical effects:
[0043] 1. The encoder structure of the improved semi - supervised segmentation network adopted by the present invention introduces a multi - scale convolution structure, an affinity field loss function, and a residual network, and innovatively designs a new feature extraction module, which simultaneously segments welds, spatters, molten pools, distortion points, and lasers, reducing the complexity of the detection system.
[0044] 2. By identifying the weld position, the present invention guides the robot to perform automatic welding, avoiding manual teaching and improving efficiency.
[0045] 3. According to the recognized centroid of the welding molten pool, the present invention adjusts the deviation caused by the front - placed vision sensor and the deformation of the welded part during the welding process, improving the welding quality.
[0046] 4. By identifying the spatter quantity and the molten pool, the present invention adjusts the welding parameters in real - time to obtain a weld with better forming. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0048] Figure 1 is the flowchart of a high - precision and high - quality vision - guided robot automatic welding method based on deep learning according to an embodiment of the present invention;
[0049] Figure 2 is the schematic diagram of the system framework according to an embodiment of the present invention;
[0050] Figure 3It is a schematic diagram of weld seam tracking according to an embodiment of the present invention;
[0051] Figure 4 It is a schematic diagram of the working principle of the sensor according to an embodiment of the present invention;
[0052] Figure 5 It is a diagram of the welding deviation model according to an embodiment of the present invention;
[0053] Figure 6 It is a schematic diagram of the adjustment of welding process parameters according to an embodiment of the present invention;
[0054] Figure 7 It is a schematic diagram of the encoder network architecture according to an embodiment of the present invention;
[0055] Figure 8 It is a schematic diagram of the decoder network architecture according to an embodiment of the present invention. Detailed implementation manners
[0056] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0057] It should be noted that the steps shown in the flowchart of the drawings may be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0058] The professional knowledge related to the embodiments will be elaborated below:
[0059] The semi-supervised segmentation network model adopts the Momentum Teacher structure commonly used in the self-supervised technical route. It consists of two branch networks, the teacher and the student, with exactly the same structure. The teacher model is a model with fixed parameters during the training process. The prediction result of the teacher model is regarded as a kind of "soft label", corresponding to the hard label (i.e., the true label); the student model is the model to be trained, and it is trained by minimizing the difference from the prediction of the teacher model. During the training process, for the labeled samples, the true label is used to calculate the loss; for the unlabeled samples, the prediction result of the Teacher model is used to calculate the loss. The network weights of the teacher branch are corrected by exponential moving average (EMA) according to the network weights of the student branch, making full use of the unlabeled sample information, reducing model oscillation, forcing the prediction of the model on the unlabeled samples to be close to the prediction of the Teacher model, and reducing overfitting. Both network branches consist of an encoder and a decoder. The encoder adopts the ResNet101 model structure of the residual network deep residual network, and the decoder adopts the Deeplabv3+ model structure of deep learning image segmentation. The encoder is mainly responsible for feature extraction and dimension compression, removing redundancy and noise; the decoder is responsible for generating the final output sequence from the output of the encoder, including the category, position, and pixel information of the required recognition features. In the encoder architecture, by integrating multi-scale convolution, a comprehensive consideration of local details and global information is achieved, providing a more flexible and comprehensive feature representation method.
[0060] This embodiment proposes a high-precision and high-quality vision-guided robot automatic welding method based on deep learning, as Figure 1 shown, which specifically includes the following steps:
[0061] Obtain the image to be recognized;
[0062] Input the image to be recognized into the segmentation network model to identify the weld position, the centroid of the molten pool, and the number of spatter. Among them, the segmentation network model is an improved semi-supervised segmentation network, and the improved semi-supervised segmentation network is composed of the encoder and decoder of the optimized teacher-student model;
[0063] Send the weld position to the robot controller to adjust the position of the welding torch in real time and guide the robot to perform automatic welding.
[0064] Specifically, the recognition features of the semi-supervised segmentation network model are welds, spatter, molten pools, distortion points, and lasers. It mainly recognizes the distortion points to guide welding. However, sometimes the field of view is limited and the distortion points are not obvious, so the position of the distortion points is determined by recognizing the intersection of the laser and the weld to guide welding.
[0065] Furthermore, the encoder and decoder of the optimized teacher-student model include: improving the encoder and decoder by combining dilated convolution, multi-scale feature fusion, and residual network. The dilated convolution technique is adopted to expand the receptive field of the convolution kernel without increasing the parameter burden of the model, ensuring the accuracy of segmentation and maintaining the efficiency of the model. Dilated convolution is used to replace the traditional stride convolution in the Layer2, Layer3, and Layer4 of the encoder structure. At the same time, the features output by the Layer1, Layer2, Layer3, and Layer4 are adjusted in the number of channels through the convolutional layer and finally concatenated and output. The batch normalization layer in the encoder and decoder is improved, and the loss function is improved. The residual ResNet network structure in the decoder backbone network is replaced with the convolutional neural Xception network structure.
[0066] Specifically, as Figure 2 shown, it mainly consists of 100 welding power supply, 200 robot controller, 300 host computer, 400 structured light vision device, and 500 welding execution device. Among them, the welding execution device includes 501 six-axis robotic arm and 502 welding torch. The structured light vision device 400 is installed at the end of the welding torch 502 and moves with the welding torch.
[0067] Furthermore, the improved encoder and decoder architectures are respectively as Figure 7 and Figure 8 shown. The improved encoder structure includes: several convolutional layers, max pooling layer, and several bottleneck structure dilated convolutional layers;
[0068] Convolutional layers are used to extract image features. Shallow convolutional layers can reveal detailed features such as spatter and distortion points, while deep convolutional layers grasp more extensive context information, which helps to identify the overall contours of welds, laser stripes, and molten pools;
[0069] Max pooling layer is used to reduce the spatial size of the feature map, reduce the computational amount and the number of parameters of the subsequent layers, and increase the receptive field;
[0070] Bottleneck structure dilated convolutional layers are used to expand the receptive field, maintain the resolution of the feature map, reduce parameters and computational amount without increasing the number of model parameters.
[0071] Specifically, in the encoder structure, the input layer receives the image information of the dataset; the basic building unit of the bottleneck structure module is the residual structure. By adding cross-stage connections, the network learns the residual mapping between the input and output data, avoiding the problems of gradient disappearance and gradient explosion; the convolutional layer is used to extract the features of the input data, such as the color and texture of splashes, lasers, and molten pools, and passes the output data to the next residual structure, forming a feature extraction process that goes deeper layer by layer; the pooling layer uses max pooling to reduce the dimensionality of the feature map, reducing the computational amount while retaining the key information of the features, suppressing noise, reducing information redundancy, and making the model more robust to changes in the feature position of the input. The entire ResNet101 network structure of the encoder is regarded as a feature extractor.
[0072] Furthermore, replacing the ResNet network structure in the decoder backbone network with the Xception network structure includes: replacing the max pooling layer with a depthwise separable convolution with a stride of 2, reducing the number of parameters and the computational amount, and adding batch normalization and ReLU activation functions after the depth convolution to improve the learning ability of the model.
[0073] The improved decoder includes: an atrous spatial pyramid pooling, a head structure, an upsampling layer, a convolutional layer, a concatenation layer, and an output layer;
[0074] The atrous spatial pyramid pooling is used to introduce multiple parallel branches to capture context information at different scales and obtain multi-level features;
[0075] The head structure is used to prevent overfitting and improve the generalization ability of the model;
[0076] Upsampling is used to fuse high-level features with low-level features to optimize the accuracy of the segmentation boundary;
[0077] The convolutional layer is used to extract multi-scale features of the input image;
[0078] The concatenation layer is used to concatenate the upsampling result and the convolutional result;
[0079] The output layer is used to output the concatenation result.
[0080] Specifically, the Deeplabv3+ model structure serves as the mask predictor in the decoder. Its Atrous Spatial Pyramid Pooling (ASPP) module captures context information at different scales by introducing multiple parallel branches. Each branch employs different atrous convolutions and pooling operations to obtain multi-level features. By using different dilation rates for atrous convolutions, the receptive field can be effectively expanded, which is equivalent to capturing the context information of the welding image dataset at multiple scales, thereby processing objects and structures of different scales and enhancing the network model's perception ability of weld seams, lasers, distortion points, spatter, and molten pools. The output result after concatenation reduces the number of channels to 256 through the Head layer, then performs upsampling and concatenates it with the output result of the original input data through the convolutional layer. Finally, the output class features are mapped to a 256-dimensional representation space through the representation layer.
[0081] Furthermore, the improved batch normalization layer includes: replacing the original batch normalization layer with synchronous batch normalization to ensure that the parameters of the normalization layer remain consistent across different devices, reducing communication overhead, and improving the training effect, convergence speed, and stability of the model.
[0082] Furthermore, the improved loss function includes: improving the original loss function by combining the affinity field loss and weighted loss fusion method to obtain the improved loss function;
[0083] The improved loss function is:
[0084]
[0085] Where, is the improved loss function, is the supervised loss, is the unsupervised loss, is the contrastive loss, , and are the weight coefficients of the corresponding terms, is the affinity field loss.
[0086] Specifically, the loss function of the original model combines supervised loss, unsupervised loss, and contrastive loss. The supervised loss is used to calculate the difference between the model prediction and the true label, and the formula is expressed as follows:
[0087]
[0088] Where, represents the number of samples in the labeled data batch, is the cross-entropy loss function; represents the prediction function of the model, where is usually the feature extraction network, Typically, it is a classification network, are model parameters, including all weights and biases that need to be learned through training; is a sample in the labeled data, is the corresponding true label.
[0089] Unsupervised loss is used to process reliable prediction results in unlabeled data. The model's own predictions are used to generate pseudo-labels, and these pseudo-labels are used to further train the model. The formula is as follows:
[0090]
[0091] Among them, represents the number of samples in the unlabeled data batch, is a sample in the unlabeled data, is the pseudo-label generated through the model's own predictions.
[0092] Contrastive loss is used to process unreliable prediction results in unlabeled data. By comparing positive and negative samples, the model is prompted to learn feature representations that distinguish different classes. The formula is as follows:
[0093]
[0094] Among them, and represent the total number of classes and the total number of pixels for each sample respectively; represents the model's prediction of the class for the th sample; is the temperature parameter that controls the smoothness of the softmax function in the contrastive loss. Each anchor pixel is followed by a positive sample and negative samples, which are represented as and respectively, is the cosine similarity between the features from two different pixels, and its range is restricted between -1 and 1.
[0095] The final loss function is expressed as:
[0096]
[0097] Among them, and are the weight coefficients for the corresponding terms, and during the training process, this coefficient is updated in real-time using a linear change strategy.
[0098] To further optimize the clarity of the weld edge and spatter boundary in the semantic segmentation of welding images by the semi-supervised segmentation model, an improved loss function that fuses the Affinity Field Loss is proposed. This loss function strengthens the recognition of the fine edges in the weld area and achieves more accurate pixel-level segmentation. The improved loss function consists of the following parts:
[0099]
[0100]
[0101] Among them, is the number of all pixel pairs in the image; is the set of all pixel pairs, where each pair of pixels are adjacent pixels; is a function used to calculate the affinity loss value of a given pixel pair, specifically:
[0102]
[0103] and are the true labels or pseudo-labels of pixels and ; is the probability that the model predicts pixels and to be in the same category.
[0104] The improved loss function enhances the performance of the model in complex welding scenarios by considering the spatial relationship and coherence between pixels. A constraint on the relationship between pixel pairs is introduced, which requires adjacent pixels to be semantically consistent, thus generating a more refined segmentation boundary, encouraging the model to learn more accurate edges and boundaries, and effectively improving the segmentation quality of the model for the weld edge, especially at the irregular boundaries generated by the molten pool and spatter during the welding process.
[0105] Furthermore, sending the weld position to the robot controller for real-time adjustment of the welding torch position includes:
[0106] Sending the weld position to the robot controller for welding. During welding, the centroid of the molten pool in the image is identified using the improved semi-supervised segmentation network, the centroid position of the molten pool is compared with the weld position, the deviation is calculated, and the deviation is used as the adjustment amount of the weld position at the next moment.
[0107] Furthermore, sending the weld position to the robot controller for welding includes:
[0108] Convert the coordinates of the recognized weld position according to the transformation matrix to obtain the coordinates in the robot coordinate system;
[0109] Perform welding according to the coordinates in the robot coordinate system.
[0110] Specifically, use the weights of the trained semi-supervised segmentation model to identify the weld position. For the recognized weld position, calculate the position of the weld in the robot polar coordinate system through the hand-eye calibration matrix obtained by calibration, so as to guide the six-axis robotic arm to drive the welding torch to weld the weld. The form of the hand-eye calibration matrix is as follows:
[0111]
[0112] Where: is the transformation matrix from the workpiece coordinate system to the base coordinate system of the six-axis robotic arm;
[0113] is the transformation matrix from the end coordinate system of the six-axis robotic arm to the base coordinate system of the six-axis robotic arm;
[0114] is the transformation matrix from the workpiece coordinate system to the camera coordinate system;
[0115] is the transformation matrix from the end coordinate system of the six-axis robotic arm to the camera coordinate system.
[0116] Furthermore, as Figure 6 shown, send the weld position to the robot controller to adjust the position of the welding torch in real time, and further include: compare the current spatter quantity with the spatter quantity at the previous moment, and adjust the welding process parameters;
[0117] Comparing the current spatter quantity with the spatter quantity at the previous moment and adjusting the welding process parameters includes:
[0118] Calculate the spatter quantity of the i-th frame and the (i + 1)-th frame of the image. If the spatter quantity decreases, adjust the welding process parameters forward; otherwise, adjust the welding process parameters backward. Among them, the process parameters include: welding inclination angle and nozzle height. In addition, the so-called forward adjustment in this embodiment means increasing the value, and the backward adjustment means decreasing the value;
[0119] Specifically, after the welding arc is struck, the image collected by the structured light vision device 400 not only contains the weld position information, but also includes splash information and molten pool information. Therefore, after the welding arc is struck, on the one hand, continuously identify the weld position to guide the six-axis robotic arm 501 to drive the welding torch 502 to weld the weld. At the same time, identify the number of splashes and the centroid of the molten pool in the image through the trained deep learning algorithm, and optimize and adjust the welding process parameters through the splash number information and the molten pool information, so as to achieve the purpose of reducing the number of splashes and improving the weld forming quality.
[0120] The following elaborates on this embodiment in conjunction with the accompanying drawings:
[0121] When using the welding equipment as Figure 2 shown for welding operations, first, take a photo of the welded part in the welding area through the line structured light device 400 installed on the welding torch 502, and transmit the obtained image to the host computer 300. In the host computer 300, use the trained semi-supervised segmentation model weights to identify the weld, laser, molten pool, distortion points, and splash features in the welding image.
[0122] Through the identification of the weld position, and through the coordinate system transformation matrix, convert the weld position coordinates identified in the image to the robot base coordinate system coordinates, and send the coordinates to the robot controller 200, thereby completing guiding the six-axis robotic arm 501 to move along the identified weld position.
[0123] After the welding arc is struck, the structured light vision device 400 continuously transmits the captured welding image to the host computer 300. In the host computer 300, input the welding image into the trained deep learning algorithm. Continuously identify the weld position through the deep learning algorithm, and continuously input the identified weld position into the robot controller 200 after coordinate transformation, driving the six-axis robotic arm 501 to drive the welding torch 502 for welding, realizing the structured light vision device 400 to guide the welding. When identifying the centroid of the molten pool in the welding image, as Figures 3 - 5 shown, the plane where the welding torch is located is perpendicular to the plane where the weld is located , and the welding torch rotates a certain angle around as the axis during the welding process, which is called the welding torch inclination angle. During the rotation of the welding torch, the perpendicular relationship between the plane where the welding torch is located and the plane where the weld is located remains unchanged. Due to the thermal deformation during the welding process, welding errors will occur. During the butt welding process in the plane, this error is considered to be the plane parallel to the weld plane where the end of the welding torch is located Deviations in the plane. Therefore, when adjusting the deviations, only the deviations in the x and y directions are considered, and the deviations in the z direction are not considered. The weld position recognized by the laser vision system is the position at the previous moment of the current welding position, and there is a certain tracking deviation in the movement of the robot. Due to various reasons, there is a certain deviation in the current position of the welding torch. The current position deviation and are combined with the weld position information recognized by the laser vision system to obtain the adjusted coordinate information and then transmitted to the robot to achieve high-precision welding operations.
[0124] Identify the number of spatters in the image. As Figure 6 shown, according to the number of spatters recognized during the welding process, adjust the welding process parameters to suppress the number of spatters during the welding process.
[0125] The above is only a preferred specific embodiment of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A high-precision and high-quality vision-guided robot automatic welding method based on deep learning, characterized in that: include: Obtain an image to be recognized; Inputting the image to be identified into a segmentation network model to identify the weld position, the centroid of the molten pool and the amount of spatter, wherein the segmentation network model is an improved semi-supervised segmentation network, and the improved semi-supervised segmentation network is composed of an encoder and a decoder of an optimized teacher-student model; The encoder and decoder of the optimized teacher-student model include: combining dilated convolution, multi-scale feature fusion and residual network to improve the encoder and decoder. In the Layer2, Layer3 and Layer4 layers of the encoder structure, dilated convolution is used instead of traditional stride convolution. At the same time, the features output by Layer1, Layer2, Layer3 and Layer4 are transmitted through The convolution layer adjusts the number of channels and finally splices the output, improves the batch normalization layer in the encoder and decoder, improves the loss function, and replaces the ResNet network structure in the decoder backbone network with the Xception network structure; The improved encoder structure includes: several convolutional layers, maximum pooling layers and several bottleneck structure hole convolutional layers. The improved decoder structure includes: replacing the ResNet network structure in the decoder backbone network with the Xception network structure; The convolution layer is used to extract image features and obtain feature maps; The maximum pooling layer is used to reduce the spatial size of the feature map; The bottleneck structure hole convolution layer is used to expand the receptive field, maintain the feature map resolution, and reduce parameters and calculation amount without increasing the number of model parameters; The improved decoder comprises: a dilated spatial convolutional pooling pyramid, a head structure, an upsampling layer, a convolutional layer, a concatenation layer and an output layer; The atrous spatial convolutional pooling pyramid is used to introduce multiple parallel branches to capture context information of different scales and obtain multi-level features; The head structure is used to prevent overfitting; The upsampling layer is used to fuse high-level features with low-level features to optimize the accuracy of segmentation boundaries; The convolution layer is used to extract multi-scale features of the input image; The splicing layer is used to splice the upsampling result and the convolution result; The output layer is used to output the splicing result; The improved batch normalization layer includes: replacing the original batch normalization layer with a synchronous batch normalization method; The improved loss function includes: improving the original loss function by combining affinity field loss and weighted loss fusion method to obtain an improved loss function; The improved loss function is: in, is the improved loss function, To monitor losses, is the unsupervised loss, is the contrast loss, , and is the weight coefficient of the corresponding item, Affinity Field Loss The weld position is sent to the robot controller, the welding gun position is adjusted in real time, and the robot is guided to weld automatically.
2. According to the high-precision and high-quality vision-guided robot automatic welding method based on deep learning in claim 1, it is characterized in that: Sending the weld position to the robot controller and adjusting the welding gun position in real time include: The weld position is sent to the robot controller for welding. During welding, the improved semi-supervised segmentation network is used to identify the centroid of the molten pool in the image, and the centroid position of the molten pool is compared with the weld position, the deviation is calculated, and the deviation is used as the weld position adjustment amount at the next moment.
3. According to claim 2, a high-precision and high-quality vision-guided robot automatic welding method based on deep learning is characterized in that: Sending the welding gun position to the robot controller for welding includes: The identified weld position coordinates are transformed according to a transformation matrix to obtain the coordinates in a robot coordinate system; Welding is performed according to the coordinates in the robot coordinate system.
4. According to claim 3, a high-precision and high-quality vision-guided robot automatic welding method based on deep learning is characterized in that: The transformation matrix is: in, is the transformation matrix from the workpiece coordinate system to the six-axis robot base coordinate system, is the transformation matrix from the six-axis robot end coordinate system to the six-axis robot base coordinate system, is the transformation matrix from the workpiece coordinate system to the camera coordinate system, It is the transformation matrix from the six-axis robot end coordinate system to the camera coordinate system.
5. The high-precision and high-quality vision-guided robot automatic welding method based on deep learning according to claim 2 is characterized in that: The weld position is sent to the robot controller to adjust the welding gun position in real time, and further includes: comparing the current spatter quantity with the spatter quantity at the previous moment to adjust the welding process parameters; Compare the current spatter quantity with the spatter quantity at the previous moment and adjust the welding process parameters including: The spatter amount of the i-th frame and the i+1-th frame image is calculated. If the current spatter amount decreases, the welding process parameters are adjusted positively; otherwise, the welding inclination angle and the nozzle height are adjusted negatively.
Citation Information
Patent Citations
Gas shielded welding process parameter optimization system and method
CN116511652A
Trajectory tracking and process deviation rectifying method based on welding robot
CN116512270A
Automatic welding monitoring system, method and equipment and readable storage medium
CN116586719A
Welding robot visual control system based on deep learning
CN116587288A
Multi-person posture recognition method and device
CN112101326A