A welding seam recognition and tracking method based on multi-modal information, an automatic welding device and a computer device

By combining multimodal information from RGB and depth images, and using a feature extraction and fusion module, the weld seam recognition method solves the problem of single visual images being affected by light and smoke, thereby improving the accuracy of weld seam recognition and welding quality.

CN119681385BActive Publication Date: 2025-12-05ZHENGZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510078466.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2025-01-10
Filing Date
2025-01-17
Publication Date
2025-12-05
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

In existing welding technologies, single visual images are easily affected by changes in lighting and smoke, resulting in low accuracy in weld seam recognition and affecting welding quality.

Method used

A weld seam recognition method based on multimodal information is adopted, which combines RGB images and depth images. Through feature extraction and feature fusion modules, RGB and depth features at different levels are extracted and fused. A convolutional neural network with ResNet architecture is used for weld seam recognition, and a laser weld seam tracking system is used for weld seam tracking.

Benefits of technology

It improves weld seam recognition accuracy, enhances welding quality, avoids the interference of light and smoke, and achieves higher precision welding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119681385B_ABST
    Figure CN119681385B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of weld seam identification and tracking, and particularly relates to a weld seam identification and tracking method based on multi-modal information, an automatic welding device and computer equipment. The method comprises the following steps: obtaining an RGB image and a depth image of a workpiece to be welded, inputting the obtained RGB image and depth image into a trained weld seam identification model to obtain an identification result of the weld seam; the weld seam identification model comprises a feature extraction module and a feature fusion module, the feature extraction module is used for extracting different levels of RGB features and different levels of depth features, the feature fusion module is used for fusing the same level of RGB features and depth features to obtain multi-modal fusion features of the level, and the multi-modal fusion features of different levels are spliced and output. The weld seam can be accurately identified, and the smoke and light interference in the welding process can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of weld seam recognition and tracking, and particularly relates to a weld seam recognition and tracking method based on multi-modal information guidance, an automatic welding device and computer equipment. BACKGROUND

[0002] Welding is a common metal connection method and is widely used in various fields of manufacturing industry. The welding quality directly affects the product quality and reliability. With the improvement of industrial automation level, welding robots gradually replace manual welding. Weld seam recognition plays a crucial role in automatic welding. If the weld seam recognition is not accurate, the welding quality is difficult to improve. The existing weld seam recognition method mainly depends on a visual system, captures images or videos in the welding process through a camera, and then analyzes and processes the images or videos by using image processing technology to obtain the weld seam recognition result. However, a single visual image is easily affected by changes in ambient light and smoke generated in the welding process, resulting in low weld seam recognition accuracy and low welding quality of the automatic welding device. SUMMARY

[0003] The application aims to provide a weld seam recognition and tracking method based on multi-modal information, an automatic welding device and computer equipment, so as to solve the technical problem that a single visual image is easily affected by changes in ambient light and smoke, resulting in low welding accuracy and affecting the welding quality.

[0004] To solve the above technical problem, the application provides a weld seam recognition method based on multi-modal information, which comprises the following steps:

[0005] An RGB image and a depth image of a workpiece to be welded are obtained, and the obtained RGB image and depth image are input into a trained weld seam recognition model to obtain a recognition result of the weld seam.

[0006] The weld seam recognition model comprises a feature extraction module and a feature fusion module. The feature extraction module is used to extract RGB features of different levels and depth features of different levels. The feature fusion module is used to fuse the RGB features and the depth features of the same level to obtain multi-modal fusion features of the level, and to splice and output the multi-modal fusion features of different levels.

[0007] Further, the process of fusing the RGB feature and the depth feature of the same level by the feature fusion module comprises: converting the RGB feature of the level into an RGB query feature, an RGB key feature and an RGB value feature, converting the depth feature of the level into a depth query feature, a depth key feature and a depth value feature; multiplying the RGB query feature of the level and the depth key feature element by element, and then obtaining an attention weight through a softmax operation, multiplying the attention weight and the depth value feature of the level, and obtaining a first fusion feature of the level through product and operation, connecting the first fusion feature of the level and the RGB feature of the level to obtain an enhanced RGB feature of the level; multiplying the depth query feature of the level and the RGB key feature of the level element by element, and then obtaining an attention weight of the product through a softmax operation, multiplying the attention weight and the value of the RGB feature of the level, and obtaining a second fusion feature of the level through product and operation, merging the second fusion feature of the level and the depth feature to obtain an enhanced depth feature of the level; and fusing the enhanced RGB feature and the enhanced depth feature to obtain a multi-modal fusion feature.

[0008] Further, the feature extraction module comprises two ResNet architecture convolutional neural networks, wherein each residual block connected in sequence in one of the ResNet architecture convolutional neural networks outputs different levels of RGB features, and each residual block connected in sequence in the other of the ResNet architecture convolutional neural networks outputs different levels of depth features.

[0009] Further, the multi-modal fusion features of different levels comprise: a deep multi-modal fusion feature, a middle multi-modal fusion feature and a shallow multi-modal fusion feature.

[0010] Correspondingly, the process of splicing the multi-modal fusion features of different levels comprises: splicing the deep multi-modal fusion feature and the middle multi-modal fusion feature, and splicing the splicing result and the shallow multi-modal fusion feature, and then outputting.

[0011] The present application is an improved invention, which has the beneficial effect that the present application inputs two visual images of RGB images and depth images into a weld joint recognition model to obtain multi-modal fusion features, and according to the recognition result of the weld joint obtained by the multi-modal features, the problem that the weld joint recognition accuracy is affected by the light change and smoke interference in the welding process is avoided, and the welding quality is improved.

[0012] To solve the above technical problems, the present application further provides a computer device comprising a processor, which implements the method steps of the weld joint recognition method based on multi-modal information of the present application when executing a computer program.

[0013] The present application is an improved invention, which has the same beneficial effect as the weld joint recognition method based on multi-modal information of the present application.

[0014] To solve the above technical problems, the present application further provides a weld joint tracking method based on multi-modal information, comprising the following steps:

[0015] Obtaining the RGB image and the depth image of the workpiece to be welded, and inputting the obtained RGB image and the depth image into the trained weld joint recognition model to obtain the recognition result of the weld joint;

[0016] The weld joint recognition model comprises a feature extraction module and a feature fusion module, the feature extraction module is used for extracting different levels of RGB features and different levels of depth features, and the feature fusion module is used for fusing the same level of RGB features and depth features to obtain the multi-modal fusion features of the level, so that the weld joint recognition model splices the multi-modal fusion features of different levels to obtain the recognition result of the weld joint;

[0017] According to the recognition result of the weld joint, the weld joint tracking is performed.

[0018] Further, the process of tracking the weld joint according to the recognition result of the weld joint comprises: if the recognition result of the weld joint is that there is a weld joint, controlling the welding gun to move to the starting position of the recognized weld joint, then inputting the image comprising the laser beam intersecting with the weld joint into the trained weld point recognition model to recognize the position of the weld point, and tracking the weld joint according to the position of the weld point.

[0019] Further, the weld point recognition model is a yolov5-face model.

[0020] Further, the starting position of the weld is the starting position of the weld in the mechanical arm coordinate system obtained according to the weld identification result and the hand-eye conversion matrix.

[0021] Further, the welding point position is the position of the welding point in the mechanical arm coordinate system.

[0022] Further, the hand-eye conversion matrix is obtained by a checkerboard calibration method.

[0023] The present application is an improved invention, and the beneficial effects thereof are the same as those of the weld identification method based on multi-modal information of the present application.

[0024] To solve the above technical problems, the present application also provides an automatic welding device, comprising a processor and a mechanical arm for controlling the movement of a welding gun, and further comprising an image detection module for acquiring an RGB image and a depth image of a workpiece to be welded, wherein the processor realizes the method steps of the weld tracking method based on multi-modal information of the present application when executing a computer program.

[0025] Further, it further comprises a laser weld tracking system, wherein the laser weld tracking system comprises a laser emitter for emitting a laser beam intersecting the weld.

[0026] The present application is an improved invention, and the beneficial effects thereof are the same as those of the weld identification method based on multi-modal information of the present application. BRIEF DESCRIPTION OF DRAWINGS

[0027] Figure 1a and Figure 1b are respectively an RGB image and a depth image of an automatic welding device embodiment of the present application;

[0028] Figure 2 is a schematic diagram of a weld tracking system of an automatic welding device embodiment of the present application;

[0029] Figure 3 is a structural schematic diagram of a DFM feature fusion module of an automatic welding device embodiment of the present application;

[0030] Figure 4 is a structural schematic diagram of a weld identification model of an automatic welding device embodiment of the present application;

[0031] Figure 5 is a weld prediction result diagram of an automatic welding device embodiment of the present application;

[0032] Figure 6 is a schematic diagram of a welding point collected by a laser weld tracking system of an automatic welding device embodiment of the present application;

[0033] Figure 7 is a schematic diagram of a welding point prediction of an automatic welding device embodiment of the present application. DETAILED DESCRIPTION

[0034] A welding seam recognition and tracking method based on multi-modal information, an automatic welding device and a computer device of the present application input two visual images of RGB images and depth images into a welding seam recognition model to obtain multi-modal fusion features, and according to the recognition result of the welding seam obtained by the multi-modal features, the problem that the welding seam recognition accuracy is affected by light changes and smoke interference in the welding process is avoided, and the welding quality is improved. Specifically, the welding seam recognition model includes a feature extraction module and a feature fusion module. The feature extraction module is used to extract different levels of RGB features and different levels of depth features. The feature fusion module is used to fuse the same level of RGB features and depth features to obtain the multi-modal fusion features of the level. The welding seam recognition model obtains the recognition result of the welding seam according to the multi-modal fusion features of different levels, improves the accuracy of the welding seam recognition, and thus improves the welding quality.

[0035] In order to make the purpose, technical scheme and advantages of the present application more clear and obvious, the present application will be further described in detail below in combination with the drawings and examples.

[0036] Automatic welding device embodiment:

[0037] The automatic welding device of the present application includes a mechanical arm, the end of the mechanical arm is fixed with a welding device and is installed with a camera (hereinafter also referred to as a 3D camera or a high-speed camera), the camera is used to obtain a depth image (as shown in Figure 1b ) and an RGB image (as shown in Figure 1a ) of a workpiece to be welded, and process these multi-modal data through a deep learning model to realize accurate recognition and segmentation of the welding seam; the end of the mechanical arm is also installed with a laser welding seam tracking system, the laser welding seam tracking system cooperates with the camera through a laser beam to accurately capture the image details of the welding point area and determine the welding point position, so as to adjust the position of the mechanical arm according to the position of the welding point and the recognition result of the welding seam, thereby improving the quality of automatic welding. The structure of the welding seam tracking system is shown in Figure 2 .

[0038] The automatic welding device of the present application further includes a processor, which realizes the following steps when executing a computer program:

[0039] I. Recognize the welding seam.

[0040] In order to improve the welding quality, it is necessary to improve the accuracy of the weld identification first, therefore, the application uses depth image and RGB image fusion to identify the weld, which avoids the problem that single visual information may be affected by environmental light changes, smoke interference and other factors, resulting in insufficient recognition accuracy. Specifically, the weld identification steps of the application are as follows:

[0041] 1.1, obtain the depth image and RGB image of the workpiece to be welded collected by the camera.

[0042] 1.2, input the depth image and RGB image of the workpiece to be welded into the weld identification model to obtain the identification result of the weld.

[0043] The weld identification model includes a feature extraction module and a feature fusion module, the feature extraction module is used to extract different levels of RGB features and different levels of depth features, and the feature fusion module is used to fuse the RGB features and the depth features of the same level to obtain the multi-modal fusion features of the level, so that the weld identification model splices the multi-modal fusion features of different levels to obtain the identification result of the weld.

[0044] The weld identification model is used for feature extraction of the obtained depth image data and RGB color information, and the obtained depth image data and RGB color information are input into the weld identification model for segmentation. The feature extraction module includes two ResNet architecture convolutional neural networks, and each residual block connected in sequence in one of the ResNet architecture convolutional neural networks outputs different levels of RGB features, and each residual block connected in sequence in the other ResNet architecture convolutional neural network outputs different levels of depth features.

[0045] The specific structure is: two ResNet architecture convolutional neural networks (CNN) are used as backbone networks, and the specific structure is as shown in Figure 4 The 3-channel RGB image and the 1-channel depth image are used as input, the depth image is copied to 3 channels for processing, and two networks with the same structure (i.e. ResNet50) are used to extract appearance features and depth features respectively.

[0046] ResNet-50 is a widely used deep learning model with a depth of 50 layers, which solves the degradation problem in training of deeper networks by using a residual learning framework. The model includes multiple residual blocks, each of which has a weight layer. In this embodiment, a skip connection is added between the input and output of the residual block, so that the input can skip some layers, allowing the signal to propagate directly through the skip connection during network training, thereby solving the gradient vanishing problem in deep network training and maintaining network performance.

[0047] In this embodiment, the feature maps of the intermediate convolutional layers of ResNet-50 are used as shared feature maps, and the outputs of the 10th, 22nd and 40th layers are specifically selected as shared feature maps, which contain rich color information, texture information and geometric information, and are crucial for fusing two modalities. The shared feature maps respectively share the rich color and texture information in the RGB image and the geometric information of the scene in the depth image. Through the shared feature maps, the model can learn the association between the two kinds of information, thereby improving the understanding of the scene.

[0048] The output of each residual block in the first three residual blocks (ResBlock-1, ResBlock-2, ResBlock-3) of ResNet50 is input to the depth fusion module DFM corresponding to the residual block, and the DFM module is used to enhance the perception ability of the model to important features through attention mechanism, and promote effective fusion of features between different modalities. The structure of the depth fusion module DFM is as shown in Figure 3 The figure, RGB-feature map is an RGB feature map, and Depth-feature map is a depth feature map. The definition of the DFM module is as follows:

[0049] F= S R || S D

[0050] S R =X_R ||(softmax(Q_R⊙K_D )⊕V_D)

[0051] S D =X_D||(softmax(Q_D⊙K_R )⊕V_R)

[0052] Where || represents the joint of the features, which is used to fuse the RGB feature and the depth feature, and improve the accuracy and robustness of feature fusion.

[0053] In the formula, S R is the enhanced RGB feature, X_R (X R in the figure) is the RGB feature, Q_R (Q R in the figure) represents the RGB query feature, K_D (K D in the figure) is the depth key feature, and V_D (V D in the figure) is the depth value feature; S D is the enhanced depth feature, X_D (X D in the figure) is the depth feature, Q_D (Q D in the figure) is the depth query feature, K_R (K R in the figure) is the RGB key feature, and V_R (V R() represents the RGB value feature.

[0054] The DFM module learns cross-modal information features from X_R and X_D. Prior to this, X_R is transformed into Q_R, K_R, and V_R, and X_D is transformed into Q_D, K_D, and V_D. Then, we generate complementary cross-modal features through the DFM module. Specifically, the DFM module calculates attention weights between the query and the key using a softmax operation, and then uses these weights to weight the features, achieving attention point localization. The weighted deep feature K_D is multiplied by the original RGB query feature Q_R using a Hadamard product (element-wise multiplication), and then the attention weights are obtained using a softmax function (not shown in the figure). The attention weights are multiplied by the vector of deep feature values ​​(V_D), and the results are summed. Figure 3 The "Product and Summation" operation (product and summation) of the DFM module yields the first fused feature. Then, a concatenation operation (also known as splicing or joining) merges the first fused feature with the RGB features to produce an enhanced RGB feature. The weighted RGB feature K_R is multiplied by the original deep query feature Q_D using a Hadamard product (element-wise multiplication). Attention weights are obtained through a softmax function and then applied to the RGB feature values ​​(V_R). The attention weights are multiplied by the value vector, and the results are summed to obtain the output vector, which is the second fused feature. This second fused feature is then merged with the deep features through a concatenation operation to produce an enhanced deep feature. Finally, the enhanced RGB features and the deep features are fused through a concatenation operation, outputting the deep fused features of both modalities.

[0055] Correspondingly, in this embodiment, the multimodal fusion features at different levels include: deep multimodal fusion features (output of the first DFM module), mid-level multimodal fusion features (output of the second DFM module), and shallow multimodal fusion features (output of the third DFM module).

[0056] The process of stitching together multimodal fusion features at different levels to obtain weld seam recognition results includes: stitching together deep multimodal fusion features and mid-level multimodal fusion features, and then stitching the stitched result together with shallow multimodal fusion features to obtain weld seam recognition results. Figure 4 (Predict). The output weld identification results are as follows: Figure 5 As shown.

[0057] II. Weld tracking.

[0058] Due to the welding process, due to the vibration of the workpiece to be welded or other influences, the position and shape of the weld seam will change slightly. In order to solve this problem, the automatic welding device of the embodiment adds a weld seam tracking system to detect the welding point and control the movement of the mechanical arm to make the welding point track the welding route. The laser emitter emits a laser beam that intersects the weld seam. The laser beam does not return at the weld seam, and the brightness is low, as shown in Figure 6 The white line in the figure is the image of the laser beam.

[0059] The specific weld seam tracking method is executed by the controller, including the following steps:

[0060] 2.1, construct a hand-eye conversion matrix to obtain the welding route in the mechanical arm coordinate system.

[0061] According to the weld seam recognition result, the starting point, the end point and the welding route of the weld seam are obtained. Since the final controlled object is the mechanical arm, in order to accurately control the mechanical arm, it is necessary to convert the weld seam recognition result from the camera coordinate system to the mechanical arm coordinate system. In this embodiment, a hand-eye conversion matrix of the camera coordinate system and the mechanical arm coordinate system is constructed, and then the hand-eye conversion matrix Rclb is used to convert the welding route from the camera coordinate system to the mechanical arm coordinate system. And according to the welding route, the starting point and the end point of the welding are determined. The hand-eye conversion matrix is obtained by hand-eye calibration. And when calibrating, the checkerboard calibration method is used for calibration.

[0062] 2.2, control the mechanical arm to move to the starting point of welding, and weld according to the welding route.

[0063] The control module controls the mechanical arm to move to the starting point of welding, and then controls the mechanical arm to gradually move along the welding path according to the calculated welding path to weld the workpiece to be welded. First, the control module controls the mechanical arm to move to the starting point of welding, and after moving to the starting point, the mechanical arm can move stably according to the predetermined welding route or the real-time calculated welding route to control the welding gun to move for welding work. In this embodiment, in order to improve the quality of welding, the control module also needs to accurately control the motion parameters of the mechanical arm, including speed, acceleration and trajectory smoothness. At the same time, the control module will also control the welding gun to adjust the welding parameters such as current, voltage and welding speed according to the welding process requirements. Through a series of accurate control measures, the system can ensure that the welding process is stable and reliable, and the expected welding effect is achieved.

[0064] 2.3, obtain the real-time position of the welding point.

[0065] The laser weld seam tracking system is used to obtain the welding point position. By detecting the position of the key welding point, it is determined whether the position of the welding point deviates. During the welding process, through the cooperation of the high-speed camera and the laser beam, the control module can quickly capture the slight changes in the position of the welding point and convert these changes into image data. In this embodiment, the laser image intersecting with the weld seam is input into the trained welding point recognition model to recognize the position of the welding point, and the welding point recognition result is as shown in Figure 7 In this embodiment, the yolov5-face model is used in the welding point recognition model, and the detection of five coordinates in the original model is modified to the detection of one coordinate. The model inputs the image collected by the laser weld seam tracking system L1, and returns the pixel coordinates of the welding point in the laser weld seam tracking system L1.

[0066] 2.4. Adjust the position of the welding gun according to the welding point position and the welding route to realize the movement of the welding point according to the welding route.

[0067] An eye-in-hand relationship matrix Rc2 b between the welding seam final system and the mechanical arm coordinates is constructed. In the real-time welding process, the system adjusts the position of the mechanical arm according to the welding point position data obtained by the laser weld seam tracking system L1 in real time and the eye-in-hand conversion matrix Rc2 b to adjust the position of the welding gun welding point, so that the welding point tracks the welding seam operation.

[0068] Welding seam recognition method embodiment:

[0069] The welding seam recognition method of the present application is as described in the "one, welding seam recognition" of the automatic welding device embodiment. The specific process, principles and beneficial effects of the method have been described in detail in the automatic welding device embodiment, and will not be repeated here.

[0070] Welding seam tracking method embodiment:

[0071] The welding seam tracking method of the present application is as described in the "two, welding seam tracking" of the automatic welding device embodiment. The specific process, principles and beneficial effects of the method have been described in detail in the automatic welding device embodiment, and will not be repeated here.

[0072] Computer device embodiment:

[0073] A computer device of the present application includes a processor that implements the method steps of the welding seam recognition method as described in the automatic welding device embodiment when executing a computer program. The specific process, principles and beneficial effects of the method have been described in detail in the automatic welding device embodiment, and will not be repeated here. The processor can be a microprocessor MCU, a programmable logic device FPGA, etc.

Claims

1. A weld seam recognition method based on multi-modal information, characterized by, The method comprises the steps of: The weld recognition model comprises a feature extraction module and a feature fusion module, the feature extraction module comprises two ResNet architecture convolutional neural networks, each residual block connected in sequence in one of the ResNet architecture convolutional neural networks outputs different levels of RGB features to a corresponding level of the feature fusion module, and each residual block connected in sequence in the other ResNet architecture convolutional neural network outputs different levels of depth features to a corresponding level of the feature fusion module; The feature fusion module is used for fusing the RGB features and the depth features of the same level to obtain multi-modal fusion features of the level, and outputs the multi-modal fusion features of different levels after splicing. The definition of the feature fusion module is as follows: The process of fusing the RGB features and the depth features of the same level by the feature fusion module comprises the following steps: the RGB features of the level are converted into RGB query features, RGB key features and RGB value features, the depth features of the level are converted into depth query features, depth key features and depth value features; the RGB query features of the level and the depth key features are multiplied element by element, and then attention weights are obtained through a softmax operation; the attention weights are multiplied with the depth value features of the level to obtain first fusion features of the level; the first fusion features of the level are connected with the RGB features of the level to obtain enhanced RGB features of the level; the depth query features of the level and the RGB key features of the level are multiplied element by element, and then attention weights of the product are obtained through a softmax operation; the attention weights are multiplied with the values of the RGB features of the level to obtain second fusion features of the level; the second fusion features of the level are combined with the depth features to obtain enhanced depth features of the level; and the enhanced RGB features and the enhanced depth features are fused to obtain multi-modal fusion features. F = S R ||S D S R = X_R || (softmax(Q_R ⊙ K_D) ⊕ V_D) S D = X_D || (softmax(Q_D ⊙ K_R) ⊕ V_R) where || denotes the combination of features, S R is the enhanced RGB feature, X_R is the RGB feature, Q_R is the RGB query feature, K_D is the depth key feature, and V_D is the depth value feature. D is the enhanced depth feature, X_D is the depth feature, Q_D is the depth query feature, K_R is the RGB key feature, and V_R is the RGB value feature.

2. The multi-modal information based weld seam recognition method according to claim 1, characterized in that The multi-modal fusion features of different levels comprise deep multi-modal fusion features, middle multi-modal fusion features and shallow multi-modal fusion features.

3. The multi-modal information based weld recognition method of claim 1, wherein, Correspondingly, the process of splicing the multi-modal fusion features of different levels comprises the following steps: the deep multi-modal fusion features and the middle multi-modal fusion features are spliced, and the spliced result is spliced with the shallow multi-modal fusion features and then output. The processor realizes the steps of the weld recognition method based on multi-modal information according to any one of claims 1-3 when executing the computer program.

4. A computer device comprising a processor, characterized in that, The method comprises the steps of:

5. A weld seam tracking method based on multi-modal information, characterized by The method comprises the steps of: ​ The weld seam recognition model comprises a feature extraction module and a feature fusion module, the feature extraction module comprises two ResNet architecture convolutional neural networks, each residual block connected in sequence in one of the ResNet architecture convolutional neural networks outputs different levels of RGB features to a corresponding level of the feature fusion module, and each residual block connected in sequence in the other ResNet architecture convolutional neural network outputs different levels of depth features to a corresponding level of the feature fusion module; The feature fusion module is configured to fuse the RGB features and the depth features of the same level to obtain multi-modal fusion features of the level, so that the weld seam recognition model splices the multi-modal fusion features of different levels to obtain a recognition result of the weld seam; The definition of the feature fusion module is as follows: F = S R ||S D S R = X_R || (softmax(Q_R ⊙ K_D) ⊕ V_D) S D = X_D || (softmax(Q_D Θ K_R) Θ V_R) where || denotes the combination of features, S R is the enhanced RGB feature, X_R is the RGB feature, Q_R is the RGB query feature, K_D is the depth key feature, and V_D is the depth value feature; S D is the enhanced depth feature, X_D is the depth feature, Q_D is the depth query feature, K_R is the RGB key feature, and V_R is the RGB value feature. Weld seam tracking is performed according to the recognition result of the weld seam.

6. The multi-modal information based seam tracking method of claim 5, wherein, The process of weld seam tracking according to the recognition result of the weld seam comprises: if the weld seam recognition result is a weld seam, the welding gun is controlled to move to the starting position of the recognized weld seam, then an image comprising a laser beam intersecting the weld seam is input into the trained weld spot recognition model to recognize the position of the weld spot, and weld seam tracking is performed according to the weld spot position.

7. The multi-modal information based seam tracking method of claim 6, wherein, The weld spot recognition model is a yolov5-face model.

8. The multi-modal information based seam tracking method of claim 6, wherein, The starting position of the weld seam is a starting position of the weld seam in the mechanical arm coordinate system obtained according to the weld seam recognition result and a hand-eye conversion matrix.

9. The multi-modal information based seam tracking method of claim 6, wherein, The weld spot position is the position of the weld spot in the mechanical arm coordinate system.

10. The multi-modal information based seam tracking method of claim 8, wherein, The hand-eye conversion matrix is obtained by a checkerboard calibration method.

11. An automated welding apparatus comprising a processor and a robotic arm for controlling movement of a welding torch, characterized in that, The image detection module is further configured to acquire the RGB image and the depth image of the workpiece to be welded, and the processor, when executing the computer program, implements the weld seam tracking method based on multi-modal information according to any one of claims 5-10.

12. The automatic welding apparatus according to claim 11, wherein The laser weld seam tracking system comprises a laser emitter configured to emit a laser beam intersecting the weld seam.

Citation Information

Patent Citations

  • Welding seam detection and segmentation method and device based on RGB-D feature layered fusion

    CN115393294A

  • RGB-D feature fusion method based on attention mechanism

    CN118918426A