Method and system for automatically rectifying deviation of welding robot body based on multi-modal sensing fusion
By combining a multimodal sensing fusion architecture and specific algorithms, the problem of insufficient perception by a single sensor in complex welding environments is solved, achieving high-precision and fast-response autonomous correction, thus improving welding quality and efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHENGZHOU KEHUI TECH CO LTD
- Filing Date
- 2026-04-09
- Publication Date
- 2026-05-15
AI Technical Summary
Existing automated welding technologies struggle to achieve high-precision, fully autonomous welding under complex working conditions. Single sensors suffer from incomplete sensing dimensions, poor anti-interference capabilities, and response delays, making it difficult to operate stably and reliably in dynamic welding environments.
Employing a multimodal sensing fusion architecture, combining wavelet transform, Kalman filtering, and DS evidence theory, a body-based intelligent correction system is constructed by using millimeter-wave radar, cameras, six-dimensional force/torque sensors, welding power source data acquisition modules, and acoustic emission sensors for data processing and decision-making.
It achieves high-precision and rapid-response autonomous correction in complex welding environments, improving system reliability and real-time performance, reducing the frequency of manual intervention, and improving welding quality and production efficiency.
Smart Images

Figure FT_1 
Figure QLYQS_4 
Figure QLYQS_5
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent control technology for industrial robots, specifically to a method and system for autonomous correction of welding robots based on multimodal sensor fusion. Background Technology
[0002] Welding is a critical process in high-end equipment manufacturing industries such as shipbuilding, pressure vessels, and heavy machinery. Welding quality directly determines the structural safety and service life of products. With the development of industrial automation, welding robots are widely used in these fields to replace traditional manual labor and improve production efficiency and consistency. However, achieving high-precision, fully autonomous welding under complex working conditions remains a major challenge in the field of intelligent control of industrial robots. Existing automated welding technologies rely on offline programming and single-sensor feedback, which have significant limitations.
[0003] Currently, there are two main types of automated welding solutions in the industry: one is pre-programmed welding based on offline programming, which is suitable for standardized production lines with fixed workpiece models and precise positioning, but cannot cope with trajectory deviations caused by workpiece assembly errors and thermal deformation in actual production; the other is adaptive welding based on online correction of a single sensor, which is based on sensing the welding environment in real time through a sensor and guiding the robot to correct the path, but this type of technology has many shortcomings.
[0004] Monocular vision-based correction techniques (such as in reference document CN112139684A) acquire weld seam images for tracking using a visual sensor. However, intense arc light, fumes, and metal spatter during welding severely interfere with image quality, leading to feature point extraction failures. Detection rates drop by approximately 30% under strong light conditions, and the system is prone to failure. Force-controlled sensor-based correction techniques (such as in reference document CN113305483B) use a six-dimensional force / torque sensor to sense the contact force between the welding torch and the workpiece bevel side. However, the response delay of force-controlled sensors is typically greater than 200ms, making it difficult to meet the real-time requirements of high-speed welding (such as GMAW welding speeds often exceeding 1m / min). Dynamic response is insufficient, prone to overshooting or oscillation, and unable to predict trends, only responding after passive contact. Arc-sensing-based correction techniques infer welding torch height or lateral deviation by detecting changes in welding current or voltage. While the response speed is fast, it only indirectly reflects the molten pool state and has limited ability to perceive bevel geometry deviations (such as misalignment and gap changes).
[0005] Existing technologies all rely on a single information source, which has problems such as incomplete perception dimensions, poor anti-interference ability, response delay or inability to fully perceive deviations. They are difficult to work stably and reliably in complex and dynamic real welding environments. There is an urgent need for an embodied intelligent deviation correction method that integrates multi-source information, has strong anti-interference ability and fast response speed. Summary of the Invention
[0006] To address the aforementioned issues, a method and system for autonomous correction of welding robots based on multimodal sensor fusion is presented. This method utilizes a multimodal sensor fusion architecture combined with wavelet transform, Kalman filtering, and DS evidence theory to ensure the accuracy of data processing and decision-making at each stage.
[0007] To achieve the above objectives, the first aspect of the present invention proposes a welding robot embodied autonomous correction system based on multimodal sensor fusion, comprising a three-level distributed architecture of a sensing layer, an edge layer, and a control layer.
[0008] The sensing layer includes a millimeter-wave radar, a camera, a six-dimensional force / torque sensor, a welding power data acquisition module, and an acoustic emission sensor. The millimeter-wave radar and camera are fixed to the front side of the welding torch by an adjustable tilt bracket, with the central axis of the lens forming a forward tilt angle with the central axis of the welding torch.
[0009] The six-dimensional force / torque sensor is integrated between the robot's sixth-axis flange and the welding torch holder;
[0010] The welding power data acquisition module is electrically connected to the welding power output terminal through an isolated sensor and a communication cable. The acoustic emission sensor is fixed to the welding torch body or the conductive nozzle by magnetic attraction or high-temperature adhesive.
[0011] All sensors are connected to the multi-channel synchronous acquisition card in the edge layer via shielded cables;
[0012] The edge layer includes an edge computing unit, which is installed inside the robot control cabinet or on the robot body. The edge computing unit is equipped with an Ethernet cable to communicate with the acquisition card of the sensing layer and the robot controller of the control layer, and is used to run the fusion algorithm model, receive and process all raw sensor data.
[0013] The control layer includes a robot controller, which receives real-time correction commands from the edge computing unit, drives the robot body to perform motion corrections, and sends parameter adjustment commands to the welding power source through a digital interface.
[0014] Furthermore, the high-speed infrared camera is equipped with an 850nm narrowband filter, and the millimeter-wave radar is a short-range detection radar;
[0015] The edge computing unit includes the NVIDIA Jetson AGX Orin developer kit;
[0016] The six-dimensional force / torque sensor includes the ATIOmega160 six-dimensional force / torque sensor;
[0017] The welding power data acquisition module includes the NIUSB-6366X series multi-functional DAQ module;
[0018] The acoustic emission sensor includes the Micro80D high-frequency acoustic emission sensor from Physical Acoustics.
[0019] The second aspect of this invention proposes a self-adjusting method for welding robots based on multimodal sensor fusion, applied to the system described in claim 1 or 2, comprising the following steps:
[0020] S1: Real-time acquisition and synchronization of multimodal data. During the welding process, the four sensors of the sensing layer work in parallel. All data are stamped with a unified high-precision timestamp through the IEEE1588 protocol and collected by the edge computing unit.
[0021] S2: Data preprocessing and feature extraction: Wavelet transform algorithm is used to denoise the images captured by the camera, and deep learning is used to process the radar data. Then the two data are fused and modeled.
[0022] The six-dimensional force / torque sensor data is decoupled, the contact force component perpendicular to the weld seam advance direction is separated, and the force guidance deviation value is calculated by combining the robot kinematic model;
[0023] Fast Fourier transform analysis was performed on the welding current / voltage signal to extract the characteristic frequency components reflecting the arc length and penetration state.
[0024] Time-frequency analysis of acoustic emission sensor signals is performed to identify characteristics of molten pool stability and spatter intensity.
[0025] S3: Multi-source information fusion and decision-making, which unifies the extracted feature values to the robot base coordinate system to achieve spatiotemporal alignment; uses Kalman filter to fuse force and electrical signal data to predict trajectory deviation trends; inputs the wavelet-denoised visual deviation, the Kalman-predicted deviation trend, and the molten pool state evidence provided by acoustic emission into the DS evidence theory model for decision-level fusion, and outputs the confidence value of the proposition "correction is needed";
[0026] When the confidence value exceeds the preset threshold of 0.85, a three-dimensional pose compensation vector (ΔX, ΔY, ΔZ, ΔRx, ΔRy, ΔRz) is generated.
[0027] S4: Real-time control and execution. The pose compensation vector is sent to the robot controller via high-speed Ethernet. The robot controller adds the compensation amount to the current motion command, drives the robot body to perform pose correction, and fine-tunes the welding parameters based on acoustic emission and electrical signal characteristics.
[0028] Furthermore, the DS evidence theory model fusion process in step S3 includes:
[0029] S31: Calculate the basic probability values based on visual bias, bias trend, and evidence of molten pool state.
[0030] S32: The Dempster combination rule is used to fuse the basic probability assignments to obtain the combined basic probability assignments;
[0031] S33: Calculate the confidence value of the proposition "needs correction" based on the combined basic probability assignment, and calculate the pose compensation vector by combining visual bias and bias trend.
[0032] Furthermore, in step S4, the overall system response delay is controlled within 50ms, and the correction accuracy is stabilized at ±0.25mm.
[0033] Furthermore, the deep learning processing of the radar acquisition data described in S2 includes:
[0034] PointNet++ was used as the backbone network for radar point cloud processing to carry out deep learning processing. The network has two SetAbstraction modules as encoding layers with sampling radii of r1=5mm and r2=10mm, and sampling points of N1=512 and N2=128, respectively. The dimensions of each MLP layer are [32, 32, 64] and [64, 64, 128], respectively. The feature propagation layer uses inverse distance weighted interpolation to upsample the features back to the original point cloud scale, generating a point-by-point feature vector Fradar.
[0035] The segmentation head then outputs a point-by-point classification score through 1×1 convolution to distinguish between "weld area" and "background noise";
[0036] The network training uses a joint loss algorithm of FocalLoss and ChamferDistance, with the loss function formula being L. total =α⋅Focal(p t ,γ)+(1−α)⋅CD(P pred ,P gt );
[0037] Where α is the joint loss weight, Focal(p) t ,γ) is the classification loss, p t The model predicts a weld seam, γ is the focusing coefficient, CD(Ppred,Pgt) is the geometric loss, and P is the constrained predicted point cloud. pred and real point cloud P gt To ensure spatial geometric consistency, during training, the weighting coefficient α=0.8, FocalLoss parameter αfocal=0.75, and focusing coefficient γ=2.0 were set.
[0038] After pre-training the geometric feature extractor on a public dataset, the segmentation head was fine-tuned for this scenario. Combined with data preprocessing for denoising and point cloud enhancement, the hyperparameters were set as follows: batch size 8-16, initial learning rate 1e-3, cosine annealing learning rate strategy, training epochs 100-200, AdamW optimizer, weight decay 1e-4, and a dropout layer with a dropout rate of 0.1 was added before the segmentation head to enhance generalization ability.
[0039] Furthermore, S2 also includes:
[0040] After the camera images are denoised by wavelet transform, the visual features related to the weld are extracted by Canny edge detection combined with Hough linear transform algorithm, including the visual pixel set of the weld center line and the visual contour point set of the bevel edge, and converted into three-dimensional coordinates in the robot base coordinate system.
[0041] After deep learning processing, radar data is used to extract key geometric features of the weld area, including the radar 3D point set of the weld centerline, the radar boundary point set of the bevel edge, and the curvature k of each point. r Normal vector n r Local features reflect the smoothness of the bevel surface.
[0042] Furthermore, S3 specifically includes:
[0043] The spatiotemporal alignment includes spatial coordinate alignment and feature alignment. The spatial coordinates are obtained through hand-eye calibration to obtain the pose matrix of the radar relative to the robot's base coordinate system. The pose matrix includes a rotation matrix Rr and a translation vector Tr, and is then aligned using a coordinate transformation algorithm formula. ; Complete the conversion, among which, These are the original three-dimensional coordinates of the radar. These are the coordinates in the robot's base coordinate system after transformation;
[0044] When transforming from the visual coordinate system to the robot's base coordinate system, the focal length and principal point coordinates are first calibrated using the camera's intrinsic parameters. This transforms the image pixel coordinates u,v of the high-speed infrared camera into the camera coordinate system Xc,Yc,Zc. The Zc is obtained from the weld depth information estimated with the assistance of radar data. Finally, the pose matrix of the camera relative to the robot's base coordinate system is obtained through hand-eye calibration to complete the final coordinate transformation. ;
[0045] Feature alignment includes a bidirectional matching algorithm based on distance and feature similarity. Taking the radar weld centerline as a reference, for each point in the visual centerline, the algorithm finds the radar point with the closest Euclidean distance in the robot base coordinate system, establishes an initial matching pair, and sets a 0.5mm distance threshold to filter invalid matches.
[0046] For the initial matching pairs, calculate the local feature similarity. The similarity algorithm is as follows:
[0047] ;
[0048] in, The curvature of the visual feature points is calculated from the image gradient. The direction angle of the radar point normal vector. The angle of the visual point edge. and As the weighting coefficient, a similarity threshold of 0.7 is set to retain high-similarity matching pairs and remove false matches;
[0049] The RANSAC (Random Sampling Consensus) algorithm is used to optimize the matching pairs, fit a uniform straight line or curve model of the weld centerline, remove external interference, and ensure that the radar and visual feature points are consistent in global trend.
[0050] The beneficial effects of the present invention through the above technical solution are as follows:
[0051] (1) This invention adopts a multimodal sensing fusion architecture, and innovatively integrates four heterogeneous sensors, namely vision, force, current / voltage and acoustic emission, to construct a complementary sensing system. This fundamentally overcomes the limitations of single sensor sensing dimension and poor environmental adaptability. The data of each sensor complement each other. When one sensor is interfered with, other sensors can still provide reliable data, thus improving the overall reliability of the system.
[0052] (2) This invention incorporates an intelligent control closed loop, proposing a complete real-time autonomous correction closed loop of "perception, fusion, decision-making to control", deeply embedding environmental perception into the robot motion control loop, enabling the robot to respond to environmental changes instinctively like a "body", which is different from traditional offline programming or single-loop feedback control, and achieves more intelligent and timely dynamic adjustment.
[0053] (3) This invention employs a specific combination of fusion algorithms, namely “wavelet transform + Kalman filtering + DS evidence theory”, for data processing and decision-making. Wavelet transform effectively solves the problem of visual image denoising, Kalman filtering enables bias trend prediction to compensate for delay, and DS evidence theory handles the uncertainty of multi-source information and makes high-confidence decisions. The synergistic effect of the three is the core of achieving high-precision and high-reliability fusion, ensuring the accuracy of data processing and decision-making in each stage. Attached Figure Description
[0054] Figure 1 This is a flowchart illustrating the steps of the self-autonomous deviation correction method and system for welding robots based on multimodal sensor fusion, as described in this invention. Detailed Implementation
[0055] The present invention will be further described below with reference to the accompanying drawings and specific embodiments:
[0056] Example 1
[0057] like Figure 1 As shown, a welding robot embodied autonomous correction system based on multimodal sensor fusion includes a three-level distributed architecture of sensing layer, edge layer and control layer.
[0058] The sensing layer includes a high-speed infrared camera, a millimeter-wave radar, a six-dimensional force / torque sensor, a welding power data acquisition module, and an acoustic emission sensor. The high-speed infrared camera and the millimeter-wave radar are fixed to the front side of the welding torch by an adjustable tilt bracket, and the central axis of the lens forms a forward tilt angle with the central axis of the welding torch.
[0059] The six-dimensional force / torque sensor is integrated between the robot's sixth-axis flange and the welding torch holder;
[0060] The welding power data acquisition module is electrically connected to the welding power output terminal through an isolated sensor and a communication cable. The acoustic emission sensor is fixed to the welding torch body or the conductive nozzle by magnetic attraction or high-temperature adhesive.
[0061] All sensors are connected to the multi-channel synchronous acquisition card in the edge layer via shielded cables;
[0062] The edge layer includes an edge computing unit, which is installed inside the robot control cabinet or on the robot body. The edge computing unit is equipped with an Ethernet cable to communicate with the acquisition card of the sensing layer and the robot controller of the control layer, and is used to run the fusion algorithm model, receive and process all raw sensor data.
[0063] The control layer includes a robot controller, which receives real-time correction commands from the edge computing unit, drives the robot body to perform motion corrections, and sends parameter adjustment commands to the welding power source through a digital interface.
[0064] In this embodiment, the industrial robot selected is the KUKAKR20R1810 model, which has a load capacity of 20kg, a working radius of 1810mm, and six-axis motion capability. It can meet the motion requirements of most welding scenarios and ensure the flexibility and accuracy of the welding torch movement during the welding process.
[0065] Edge computing unit: It adopts the NVIDIA Jetson AGX Orin developer kit, which has a built-in high-performance GPU with a computing power of up to 275 TOPS (trillion operations per second). It can efficiently run lightweight fusion algorithm models. At the same time, its integrated high-precision time synchronization module supports the IEEE1588 protocol and can achieve a time synchronization accuracy of ≤10μs, meeting the requirements of multimodal data synchronization processing and real-time computing.
[0066] Visual sensor: A FLIRBlackflySBFS-U3-51S5PC-C high-speed infrared camera is selected. This camera uses a global shutter, has 5 megapixels, supports infrared band imaging, and is equipped with an 850nm narrow-bandpass filter, which can effectively suppress welding arc interference and clearly capture weld seam images, providing high-quality image data for visual deviation calculation. Millimeter-wave radar is used in conjunction with visual data for fusion modeling; the millimeter-wave radar is a short-range detection radar.
[0067] Six-dimensional force / torque sensor: The ATIOmega160 six-dimensional force / torque sensor is selected. Its accuracy can reach ±0.1N. It can accurately measure the force and torque on the end of the welding torch. It is also small in size and suitable for integration between the sixth axis flange of the robot and the welding torch holder to directly obtain force feedback data.
[0068] Data acquisition card: The NIUSB-6366X series multi-functional DAQ module (data acquisition) is adopted. This module supports high-speed synchronous acquisition of multiple analog signals with a sampling rate of not less than 10kHz. It can meet the acquisition requirements of electrical signals such as welding current and voltage as well as acoustic emission signals, ensuring the accuracy and real-time performance of signal acquisition.
[0069] Acoustic emission sensor: The Micro80D high-frequency acoustic emission sensor from Physical Acoustics (PAC) is selected. Its frequency range is 100-1000kHz. It can effectively collect the acoustic wave signals generated in the molten pool area during welding. It can be easily installed near the welding torch body or the conductive nozzle by magnetic attraction or high-temperature adhesive fixing to ensure the signal acquisition effect.
[0070] Welding power source: A high-performance welding power source that matches the welding process is selected. It has a digital interface and can receive parameter adjustment commands sent by the robot controller to realize real-time fine adjustment of current and voltage, ensuring the stability of the welding process and the welding quality.
[0071] The aforementioned hardware is assembled according to the connection relationships in the system architecture. The high-speed infrared camera and radar are fixed to the front side of the welding torch using an adjustable tilt mounting bracket. The bracket is adjusted so that the central axis of the camera lens forms a forward tilt angle of 15°±5° with the central axis of the welding torch, and the distance between the lens and the weld surface is set to 30mm. The six-dimensional force / torque sensor is installed between the robot's sixth-axis flange and the welding torch holder, ensuring a tight connection between the sensor and both without any looseness. The welding power data acquisition module is connected to the welding power output terminal through an isolated sensor and communication cable, ensuring electrical safety and stable signal transmission. The acoustic emission sensor is fixed to the welding torch body or near the conductive nozzle using magnetic attraction or high-temperature adhesive, ensuring good contact between the sensor and the welding torch for effective acquisition of acoustic signals. All sensors are connected to a multi-channel synchronous acquisition card through shielded cables. The acquisition card is connected to the edge computing unit, which communicates with the acquisition card and the robot controller via Ethernet cables. The robot controller establishes connections with the robot body and the welding power supply, completing the hardware assembly of the entire system.
[0072] A Linux-based operating system was deployed on the edge computing unit (NVIDIA Jetson AGX Orin Developer Kit), and corresponding drivers were installed to ensure that all hardware devices could be correctly recognized and used. Deep learning frameworks (such as TensorFlow and PyTorch) and signal processing libraries (such as OpenCV and SciPy) were built to provide a software environment for the fusion algorithm. The fusion algorithms, including wavelet transform, Kalman filtering, and DS evidence theory models, were written as lightweight software modules and integrated into the edge computing unit's software system to ensure efficient algorithm operation. Simultaneously, a data acquisition program was written to achieve real-time acquisition and time-synchronized processing of data from various sensors. Corresponding control software was installed on the robot controller to enable communication with the edge computing unit, receive correction commands, and drive the robot's movement. A communication program with the welding power source was also written to adjust welding parameters.
[0073] A self-autonomous deviation correction method for welding robots based on multimodal sensor fusion, applied to the system described in claim 1 or 2, includes the following steps:
[0074] S1: Real-time acquisition and synchronization of multimodal data. During the welding process, the four sensors of the sensing layer work in parallel. All data are stamped with a unified high-precision timestamp through the IEEE1588 protocol and collected by the edge computing unit.
[0075] S2: Data preprocessing and feature extraction. Wavelet transform algorithm is used to denoise the images captured by the camera, and deep learning is used to process the radar data. Then, the two data are fused and modeled to extract the weld centerline and bevel edge features and calculate the visual deviation value.
[0076] The six-dimensional force / torque sensor data is decoupled, the contact force component perpendicular to the weld seam advance direction is separated, and the force guidance deviation value is calculated by combining the robot kinematic model;
[0077] Fast Fourier transform analysis was performed on the welding current / voltage signal to extract the characteristic frequency components reflecting the arc length and penetration state.
[0078] Time-frequency analysis of acoustic emission sensor signals is performed to identify characteristics of molten pool stability and spatter intensity.
[0079] S3: Multi-source information fusion and decision-making, which unifies the extracted feature values to the robot base coordinate system to achieve spatiotemporal alignment; uses Kalman filter to fuse force and electrical signal data to predict trajectory deviation trends; inputs the wavelet-denoised visual deviation, the Kalman-predicted deviation trend, and the molten pool state evidence provided by acoustic emission into the DS evidence theory model for decision-level fusion, and outputs the confidence value of the proposition "correction is needed";
[0080] When the confidence value exceeds the preset threshold of 0.85, a three-dimensional pose compensation vector (ΔX, ΔY, ΔZ, ΔRx, ΔRy, ΔRz) is generated.
[0081] S4: Real-time control and execution. The pose compensation vector is sent to the robot controller via high-speed Ethernet. The robot controller adds the compensation amount to the current motion command, drives the robot body to perform pose correction, and fine-tunes the welding parameters based on acoustic emission and electrical signal characteristics.
[0082] Step S3, the DS evidence theory model fusion process, includes:
[0083] S31: Calculate the basic probability values based on visual bias, bias trend, and evidence of molten pool state.
[0084] S32: The Dempster combination rule is used to fuse the basic probability assignments to obtain the combined basic probability assignments;
[0085] S33: Calculate the confidence value of the proposition "needs correction" based on the combined basic probability assignment, and calculate the pose compensation vector by combining visual bias and bias trend.
[0086] In step S4, the overall system response delay is controlled within 50ms, and the correction accuracy is stabilized at ±0.25mm.
[0087] The deep learning processing of the radar acquisition data mentioned in S2 includes:
[0088] PointNet++ was used as the backbone network for radar point cloud processing to carry out deep learning processing. The network has two SetAbstraction modules as encoding layers with sampling radii of r1=5mm and r2=10mm, and sampling points of N1=512 and N2=128, respectively. The dimensions of each MLP layer are [32, 32, 64] and [64, 64, 128], respectively. The feature propagation layer uses inverse distance weighted interpolation to upsample the features back to the original point cloud scale, generating a point-by-point feature vector Fradar.
[0089] The segmentation head then outputs a point-by-point classification score through 1×1 convolution to distinguish between "weld area" and "background noise";
[0090] The network training uses a joint loss algorithm of FocalLoss and ChamferDistance, with the loss function formula being L. total =α⋅Focal(p t ,γ)+(1−α)⋅CD(P pred ,P gt );
[0091] Where α is the joint loss weight, Focal(p) t ,γ) is the classification loss, p t The model predicts a weld seam, γ is the focusing coefficient, CD(Ppred,Pgt) is the geometric loss, and P is the constrained predicted point cloud. pred and real point cloud P gt To ensure spatial geometric consistency, during training, the weighting coefficient α=0.8, FocalLoss parameter αfocal=0.75, and focusing coefficient γ=2.0 were set.
[0092] After pre-training the geometric feature extractor on a public dataset, the segmentation head was fine-tuned for this scenario. Combined with data preprocessing for denoising and point cloud enhancement, the hyperparameters were set as follows: batch size 8-16, initial learning rate 1e-3, cosine annealing learning rate strategy, training epochs 100-200, AdamW optimizer, weight decay 1e-4, and a dropout layer with a dropout rate of 0.1 was added before the segmentation head to enhance generalization ability.
[0093] S2 further includes:
[0094] After the camera images are denoised by wavelet transform, the visual features related to the weld are extracted by Canny edge detection combined with Hough linear transform algorithm, including the visual pixel set of the weld center line and the visual contour point set of the bevel edge, and converted into three-dimensional coordinates in the robot base coordinate system.
[0095] After deep learning processing, radar data is used to extract key geometric features of the weld area, including the radar 3D point set of the weld centerline, the radar boundary point set of the bevel edge, and the curvature k of each point. r Normal vector n r Local features reflect the smoothness of the bevel surface.
[0096] S3 specifically includes:
[0097] The spatiotemporal alignment includes spatial coordinate alignment and feature alignment. The spatial coordinates are obtained through hand-eye calibration to obtain the pose matrix of the radar relative to the robot's base coordinate system. The pose matrix includes a rotation matrix Rr and a translation vector Tr, and is then aligned using a coordinate transformation algorithm formula. ; Complete the conversion, among which, These are the original three-dimensional coordinates of the radar. These are the coordinates in the robot's base coordinate system after transformation;
[0098] When transforming from the visual coordinate system to the robot's base coordinate system, the focal length and principal point coordinates are first calibrated using the camera's intrinsic parameters. This transforms the image pixel coordinates u,v of the high-speed infrared camera into the camera coordinate system Xc,Yc,Zc. The Zc is obtained from the weld depth information estimated with the assistance of radar data. Finally, the pose matrix of the camera relative to the robot's base coordinate system is obtained through hand-eye calibration to complete the final coordinate transformation. ;
[0099] Feature alignment includes a bidirectional matching algorithm based on distance and feature similarity. Taking the radar weld centerline as a reference, for each point in the visual centerline, the algorithm finds the radar point with the closest Euclidean distance in the robot base coordinate system, establishes an initial matching pair, and sets a 0.5mm distance threshold to filter invalid matches.
[0100] For the initial matching pairs, calculate the local feature similarity. The similarity algorithm is as follows:
[0101] ;
[0102] in, The curvature of the visual feature points is calculated from the image gradient. The direction angle of the radar point normal vector. The angle of the visual point edge. and As the weighting coefficient, a similarity threshold of 0.7 is set to retain high-similarity matching pairs and remove false matches;
[0103] The RANSAC (Random Sampling Consensus) algorithm is used to optimize the matching pairs, fit a uniform straight line or curve model of the weld centerline, remove external interference, and ensure that the radar and visual feature points are consistent in global trend.
[0104] Wavelet transform algorithm implementation: The db4 wavelet basis function is used to perform multi-scale decomposition on the weld seam images and data acquired by the high-speed infrared camera, and deep learning algorithms are combined to process the radar data. The decomposition yields approximation coefficients and detail coefficients. A soft thresholding denoising method is used on the detail coefficients to suppress noise interference. Then, the image is reconstructed through inverse wavelet transform to obtain a clear, denoised weld seam image. Finally, an edge detection algorithm (such as the Canny algorithm) is used to extract features such as the weld seam centerline and bevel edges, and visual deviation values are calculated based on these features.
[0105] Kalman filter algorithm implementation: Establish a state-space model of the force and electrical signals and the deviation of the weld trajectory. Use the force guidance deviation value and the characteristic frequency component of the electrical signal as observation values and input them into the Kalman filter. Through prediction, updating and other steps, predict the deviation of the weld trajectory to obtain the deviation trend in the short term, and provide predictive information for decision fusion.
[0106] The DS evidence theory model is implemented as follows: Two propositions, "correction is needed" and "correction is not needed," are defined as the recognition framework. Basic probability assignments (BPAs) for each information source to the two propositions are determined based on visual bias values, Kalman prediction bias trends, and acoustic emission characteristics. These assignment rules can be obtained through training with a large amount of experimental data or set empirically. The Dempster combination rule is used to fuse the BPAs of the three information sources, resulting in a combined BPA. The confidence value of the "correction is needed" proposition is calculated based on the combined BPA. When the confidence value exceeds 0.85, a pose compensation vector is generated. The calculation of the pose compensation vector combines visual bias values and Kalman prediction bias trends, establishing a mapping relationship between bias and compensation amount.
[0107] Example 2
[0108] The following experiments were conducted based on Example 1, demonstrating that the present invention has good performance.
[0109] Experimental conditions set:
[0110] Two common welding materials (Q235B low-carbon steel and 304 stainless steel) and two typical bevel shapes (V-groove and U-groove) were selected as experimental objects to simulate welding conditions in actual industrial scenarios such as shipbuilding and pressure vessel welding. Welding process parameters were set as follows: for Q235B low-carbon steel, GMAW (gas metal arc welding) was used, with a welding current of 180-220A, a welding voltage of 22-26V, and a welding speed of 1-1.2m / min; for 304 stainless steel, the same GMAW welding process was used, with a welding current of 160-200A, a welding voltage of 20-24V, and a welding speed of 0.8-1m / min. During the welding process, some interfering factors were artificially introduced, such as strong arc light, spatter, and certain assembly errors and thermal deformation of the workpiece, to simulate a complex and harsh welding environment and test the system's performance under such conditions.
[0111] Experimental Results and Analysis
[0112] Perception reliability testing: Under conditions of different welding materials and bevel forms, and with interference factors such as arc light and spatter, the operating status and data acquisition of each sensor in the system were recorded. The results show that when the high-speed infrared camera is interfered with by arc light and spatter, the radar, six-dimensional force / torque sensor, and acoustic emission sensor can still stably acquire reliable data. The system can continuously perform perception and correction operations, and there were no instances of the system stopping work due to the failure of a single sensor. This proves that the system's perception reliability and environmental adaptability have been significantly improved, and the multimodal sensing fusion architecture effectively overcomes the limitations of a single sensor.
[0113] Real-time performance testing: A high-precision timer was used to record the time interval between the sensor acquiring deviation data and the robot executing the correction action, i.e., the system response delay. Multiple tests were conducted at different welding speeds, and the results showed that the overall system response delay was stably controlled within 50ms, far lower than the 200ms or more delay of existing force control solutions. This effectively matches the real-time requirements of high-speed welding (1-1.2m / min), avoiding overshoot or oscillation problems caused by delay and ensuring the stability of the welding process.
[0114] Accuracy Testing: After welding, high-precision measuring instruments (such as coordinate measuring machines) are used to measure the deviation between the actual trajectory and the ideal trajectory of the weld, and the correction accuracy is calculated. Multiple measurements were performed on welds of different materials and with different bevel forms. The results show that the system's correction accuracy is stable within ±0.25mm, meeting the accuracy requirements of ISO5817 Class B welds (this standard requires weld deviation within ±0.5mm), and even exceeding the standard requirements. This effectively overcomes the trajectory deviation problem caused by workpiece assembly errors and thermal deformation, ensuring high-precision weld formation.
[0115] Automation and Quality Testing: The frequency of manual intervention during the welding process was recorded, and the first-pass yield rate was statistically analyzed. Experimental results show that, using the system of this invention, the frequency of manual intervention was reduced from 8 times / meter in the existing scheme to 0.3 times / meter, significantly reducing manual intervention and dependence on highly skilled welders; the first-pass yield rate was increased from approximately 90% in the existing scheme to 99.2%, significantly improving production efficiency. Simultaneously, through acoustic emission sensor monitoring and welding parameter fine-tuning, spatter during the welding process was effectively suppressed, resulting in good weld formation and a significant improvement in welding quality.
[0116] The above experiments demonstrate that the self-autonomous correction method and system for welding robots based on multimodal sensor fusion provided by this invention exhibits excellent performance in terms of perception reliability, real-time performance, accuracy, automation level, and welding quality. It can effectively solve the problems existing in the prior art, meet the needs of actual industrial welding scenarios, and has broad application prospects.
[0117] The embodiments described above are merely preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Therefore, all equivalent changes or modifications made to the structure, features and principles described in the claims of the present invention should be included within the scope of the present invention.
Claims
1. A self-contained autonomous deviation correction system for welding robots based on multimodal sensor fusion, characterized in that, A three-tiered distributed architecture comprising a sensing layer, an edge layer, and a control layer; The sensing layer includes a millimeter-wave radar, a camera, a six-dimensional force / torque sensor, a welding power data acquisition module, and an acoustic emission sensor. The millimeter-wave radar and camera are fixed to the front side of the welding torch by an adjustable tilt bracket, and the central axis of the lens forms a forward tilt angle with the central axis of the welding torch. The six-dimensional force / torque sensor is integrated between the robot's sixth-axis flange and the welding torch holder; The welding power data acquisition module is electrically connected to the welding power output terminal through an isolated sensor and a communication cable. The acoustic emission sensor is fixed to the welding torch body or the conductive nozzle by magnetic attraction or high-temperature adhesive. All sensors are connected to the multi-channel synchronous acquisition card in the edge layer via shielded cables; The edge layer includes an edge computing unit, which is installed inside the robot control cabinet or on the robot body. The edge computing unit is equipped with an Ethernet cable to communicate with the acquisition card of the sensing layer and the robot controller of the control layer, and is used to run the fusion algorithm model, receive and process all raw sensor data. The control layer includes a robot controller, which receives real-time correction commands from the edge computing unit, drives the robot body to perform motion corrections, and sends parameter adjustment commands to the welding power source through a digital interface.
2. The welding robot autonomous correction system based on multimodal sensor fusion according to claim 1, characterized in that, The high-speed infrared camera is equipped with an 850nm narrowband pass filter. The millimeter-wave radar is a short-range detection radar. The edge computing unit includes the NVIDIA Jetson AGX Orin developer kit; The six-dimensional force / torque sensor includes the ATIOmega160 six-dimensional force / torque sensor; The welding power data acquisition module includes the NIUSB-6366X series multi-functional DAQ module; The acoustic emission sensor includes the Micro80D high-frequency acoustic emission sensor from Physical Acoustics.
3. A method for autonomous deviation correction of a welding robot based on multimodal sensor fusion, applied to the system described in claim 1 or 2, characterized in that, Includes the following steps: S1: Real-time acquisition and synchronization of multimodal data. During the welding process, the four sensors of the sensing layer work in parallel. All data are stamped with a unified high-precision timestamp through the IEEE1588 protocol and collected by the edge computing unit. S2: Data preprocessing and feature extraction. Wavelet transform algorithm is used to denoise the images captured by the camera, and deep learning is used to process the radar data. Then, the two data are fused and modeled to extract the weld centerline and bevel edge features and calculate the visual deviation value. The six-dimensional force / torque sensor data is decoupled, the contact force component perpendicular to the weld seam advance direction is separated, and the force guidance deviation value is calculated by combining the robot kinematic model; Fast Fourier transform analysis was performed on the welding current / voltage signal to extract the characteristic frequency components reflecting the arc length and penetration state. Time-frequency analysis of acoustic emission sensor signals is performed to identify characteristics of molten pool stability and spatter intensity. S3: Multi-source information fusion and decision-making, which unifies the extracted feature values to the robot base coordinate system to achieve spatiotemporal alignment; uses Kalman filter to fuse force and electrical signal data to predict trajectory deviation trends; inputs the wavelet-denoised visual deviation, the Kalman-predicted deviation trend, and the molten pool state evidence provided by acoustic emission into the DS evidence theory model for decision-level fusion, and outputs the confidence value of the proposition "correction is needed"; When the confidence value exceeds the preset threshold of 0.85, a three-dimensional pose compensation vector (ΔX, ΔY, ΔZ, ΔRx, ΔRy, ΔRz) is generated. S4: Real-time control and execution. The pose compensation vector is sent to the robot controller via high-speed Ethernet. The robot controller adds the compensation amount to the current motion command, drives the robot body to perform pose correction, and fine-tunes the welding parameters based on acoustic emission and electrical signal characteristics.
4. The self-adjusting method for welding robots based on multimodal sensor fusion according to claim 3, characterized in that, Step S3, the DS evidence theory model fusion process, includes: S31: Calculate the basic probability values based on visual bias, bias trend, and evidence of molten pool state. S32: The Dempster combination rule is used to fuse the basic probability assignments to obtain the combined basic probability assignments; S33: Calculate the confidence value of the proposition "needs correction" based on the combined basic probability assignment, and calculate the pose compensation vector by combining visual bias and bias trend.
5. The self-adjusting autonomous deviation correction method for welding robots based on multimodal sensor fusion according to claim 3, characterized in that, In step S4, the overall system response delay is controlled within 50ms, and the correction accuracy is stabilized at ±0.25mm.
6. The self-adjusting autonomous deviation correction method for welding robots based on multimodal sensor fusion according to claim 3, characterized in that, The deep learning processing of the radar acquisition data mentioned in S2 includes: PointNet++ was used as the backbone network for radar point cloud processing to carry out deep learning processing. The network has two SetAbstraction modules as encoding layers with sampling radii of r1=5mm and r2=10mm, and sampling points of N1=512 and N2=128, respectively. The dimensions of each MLP layer are [32, 32, 64] and [64, 64, 128], respectively. The feature propagation layer uses inverse distance weighted interpolation to upsample the features back to the original point cloud scale, generating a point-by-point feature vector Fradar. The segmentation head then outputs a point-by-point classification score through 1×1 convolution to distinguish between "weld area" and "background noise"; The network training uses a joint loss algorithm of FocalLoss and ChamferDistance, with the loss function formula being L. total =α⋅Focal(p t ,γ)+(1−α)⋅CD(P pred ,P gt ); Where α is the joint loss weight, Focal(p) t ,γ) is the classification loss, p t The model predicts a weld seam, γ is the focusing coefficient, CD(Ppred,Pgt) is the geometric loss, and P is the constrained predicted point cloud. pred and real point cloud P gt To ensure spatial geometric consistency, during training, the weighting coefficient α=0.8, FocalLoss parameter αfocal=0.75, and focusing coefficient γ=2.0 were set. After pre-training the geometric feature extractor on a public dataset, the segmentation head was fine-tuned for this scenario. Combined with data preprocessing for denoising and point cloud enhancement, the hyperparameters were set as follows: batch size 8-16, initial learning rate 1e-3, cosine annealing learning rate strategy, training epochs 100-200, AdamW optimizer, weight decay 1e-4, and a dropout layer with a dropout rate of 0.1 was added before the segmentation head to enhance generalization ability.
7. The self-adjusting autonomous deviation correction method for welding robots based on multimodal sensor fusion according to claim 3, characterized in that, S2 further includes: After the camera images are denoised by wavelet transform, the visual features related to the weld are extracted by Canny edge detection combined with Hough linear transform algorithm, including the visual pixel set of the weld center line and the visual contour point set of the bevel edge, and converted into three-dimensional coordinates in the robot base coordinate system. After deep learning processing, radar data is used to extract key geometric features of the weld area, including the radar 3D point set of the weld centerline, the radar boundary point set of the bevel edge, and the curvature k of each point. r Normal vector n r Local features reflect the smoothness of the bevel surface.
8. The self-adjusting autonomous deviation correction method for welding robots based on multimodal sensor fusion according to claim 3, characterized in that, S3 specifically includes: The spatiotemporal alignment includes spatial coordinate alignment and feature alignment. The spatial coordinates are obtained through hand-eye calibration to obtain the pose matrix of the radar relative to the robot's base coordinate system. The pose matrix includes a rotation matrix Rr and a translation vector Tr, and is then aligned using a coordinate transformation algorithm formula. ; Complete the conversion, among which, These are the original three-dimensional coordinates of the radar. These are the coordinates in the robot's base coordinate system after transformation; When transforming from the visual coordinate system to the robot's base coordinate system, the focal length and principal point coordinates are first calibrated using the camera's intrinsic parameters. This transforms the image pixel coordinates u,v of the high-speed infrared camera into the camera coordinate system Xc,Yc,Zc. The Zc is obtained from the weld depth information estimated with the assistance of radar data. Finally, the pose matrix of the camera relative to the robot's base coordinate system is obtained through hand-eye calibration to complete the final coordinate transformation. ; Feature alignment includes a bidirectional matching algorithm based on distance and feature similarity. Taking the radar weld centerline as a reference, for each point in the visual centerline, the algorithm finds the radar point with the closest Euclidean distance in the robot base coordinate system, establishes an initial matching pair, and sets a 0.5mm distance threshold to filter invalid matches. For the initial matching pairs, calculate the local feature similarity. The similarity algorithm is as follows: ; in, The curvature of the visual feature points is calculated from the image gradient. The direction angle of the radar point normal vector. The angle of the visual point edge. and As the weighting coefficient, a similarity threshold of 0.7 is set to retain high-similarity matching pairs and remove false matches; The RANSAC (Random Sampling Consensus) algorithm is used to optimize the matching pairs, fit a uniform straight line or curve model of the weld centerline, remove external interference, and ensure that the radar and visual feature points are consistent in global trend.