Welding robot welding seam real-time sensing method and system based on multi-modal data fusion, electronic device and storage medium

By using multimodal data fusion technology, high robustness and high precision of weld seam real-time perception under complex working conditions are achieved, which solves the problem of insufficient robustness and accuracy of weld seam perception in existing technologies. The multimodal data fusion method is used to dynamically adjust the weights and generate weld seam tracking parameters, thereby improving the adaptive control capability of the welding robot.

CN121972765BActive Publication Date: 2026-06-16INNO CIRCUITS LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INNO CIRCUITS LTD
Filing Date
2026-04-07
Publication Date
2026-06-16

Smart Images

  • Figure CN121972765B_ABST
    Figure CN121972765B_ABST
Patent Text Reader

Abstract

The application provides a multi-modal data fusion welding robot weld seam real-time sensing method, system, electronic device and storage medium, and belongs to welding automation technology. The method comprises the following steps: synchronously collecting multiple types of sensing data in a welding process, adding a uniform timestamp to the collected sensing data, filtering, and obtaining initial multi-modal data; aligning the time axis based on the timestamp in the same coordinate system with the welding torch as the reference, extracting quality parameters and state features of each mode; a feature fusion network dynamically generates fusion weights corresponding to each mode according to the quality parameters of each mode, and dynamically weights the state features of each mode using the fusion weights, generates and outputs an optimized feature at the current time; the optimized feature is input into a pre-constructed decision model to calculate the weld seam tracking parameters required for real-time control of the welding robot. The scheme can more effectively resist complex working condition interference and accurately sense the weld seam information in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of welding automation and relates to welding robots and sensor information fusion processing technology, especially to a method, system, electronic device and storage medium for real-time sensing of weld seams in welding robots using multimodal data fusion. Background Technology

[0002] With the rapid development of intelligent manufacturing and industrial automation, intelligent welding robots are increasingly being used in pressure vessels, shipbuilding, aerospace, and steel structures. Accurate and real-time perception and identification of weld conditions are prerequisites for achieving autonomous and adaptive control of the welding process.

[0003] Current weld seam sensing technologies have the following drawbacks:

[0004] 1. Based on a single visual sensing solution, this type of solution typically uses a CCD camera, laser structured light, or depth camera to acquire two-dimensional or three-dimensional image information of the weld, and extracts the weld position and bevel features through image processing algorithms. However, under actual welding conditions, visual sensing is easily affected by strong arc radiation, welding fume dispersion, metal spatter interference, and oil or reflection on the workpiece surface, leading to decreased image quality and feature extraction failure; especially under complex conditions such as multi-layer multi-pass welding, narrow space welds, and highly reflective materials, reliability is even more difficult to guarantee.

[0005] 2. Relying on a single physiological sensing scheme, such as force control or arc sensing-based tracking, although these methods are not sensitive to optical interference, they are difficult to obtain accurate three-dimensional geometric information of the weld (such as groove width, depth, etc.), have low spatial resolution, and have limited response to changes in working conditions such as workpiece assembly errors and uneven groove processing.

[0006] 3. Simple Sensor Data Complementarity Schemes. Some existing studies attempt to improve perception capabilities by combining visual guidance with force-controlled contact. However, these methods often employ simple data switching logic or data fusion strategies based on fixed weights, failing to consider the confidence differences of different sensors under dynamically changing conditions. When visual sensors experience data failure due to strong light or smoke interference, and force sensors generate noise due to abrupt changes in contact state, traditional fixed fusion modes struggle to dynamically evaluate and adaptively adjust the confidence of multi-source information, leading to decreased fusion accuracy and insufficient system robustness.

[0007] Therefore, there is an urgent need for a new method that can dynamically assess the reliability and intelligently fuse multi-source heterogeneous information such as vision and force under strong interference and high noise conditions, so as to achieve high robustness and high precision real-time perception of weld seams. Summary of the Invention

[0008] To address the shortcomings of the aforementioned existing technologies, this application provides a method, system, electronic device, and storage medium for real-time weld seam perception in welding robots using multimodal data fusion, which can more effectively resist interference from complex working conditions and perceive weld seam information in real time and accurately.

[0009] To achieve the above objectives, the present invention employs the following techniques:

[0010] A real-time weld seam perception method for welding robots based on multimodal data fusion includes:

[0011] Step 1: Multimodal data acquisition and preprocessing:

[0012] Simultaneously acquire multiple types of sensor data during the welding process, including:

[0013] Three-dimensional point cloud data of the weld area in front of the welding torch;

[0014] Welding process data includes mechanical process data and electrical process data. Mechanical process data is obtained by measuring a torque sensor installed at the end of the welding torch. Electrical process data includes welding current signals collected from the output of the welding power source and arc voltage signals collected between the welding torch contact tip and the workpiece.

[0015] Temperature field distribution data of the weld pool and its surrounding predetermined area;

[0016] A unified timestamp is added to the collected sensor data, and filtering preprocessing is performed separately to obtain the initial multimodal data;

[0017] Step 2: Spatiotemporal registration and feature extraction

[0018] The initial multimodal data are unified into the same coordinate system with the welding gun as the reference, and the time axis is aligned based on the timestamp to ensure that the initial multimodal data collected at the same time correspond one-to-one, so as to complete the spatiotemporal registration, and then the quality parameters and state characteristics of each mode are extracted from it.

[0019] Step 3: Feature Fusion Based on Dynamic Weights

[0020] The quality parameters and state features of each modality at the same time are input into a pre-constructed feature fusion network. The feature fusion network dynamically generates fusion weights corresponding to each modality based on the quality parameters of each modality, and uses the fusion weights to dynamically weight the state features of each modality to generate and output a set of optimized features at the current time. Among them, the quality parameters are used to characterize the credibility of the state features of their respective modalities at the current time.

[0021] Step 4: Weld Parameter Calculation and Output

[0022] The quality parameters and state features of each modality at the same time are input into a pre-constructed feature fusion network. The feature fusion network dynamically generates fusion weights corresponding to each modality based on the quality parameters of each modality, and uses the fusion weights to dynamically weight the state features of each modality to generate and output a set of optimized features at the current time. Among them, the quality parameters are used to characterize the credibility of the state features of their respective modalities at the current time.

[0023] Furthermore, the mass parameters and state characteristics of each mode are extracted, including:

[0024] Three-dimensional point cloud geometric analysis and weld area feature analysis are performed on spatiotemporally registered three-dimensional point cloud data to extract point cloud quality parameters to characterize the quality of the point cloud data itself and initial visual state features to characterize the weld geometry.

[0025] Force signal time-domain analysis and contact state characteristic analysis are performed on the spatiotemporally registered mechanical process data to extract mechanical quality parameters to characterize the quality of the mechanical signal itself and initial mechanical state characteristics to characterize the contact state of the welding torch.

[0026] We perform time-domain statistics and arc stability analysis on the spatiotemporally registered electrical process data to extract electrical quality parameters to characterize the quality of the electrical signal itself and initial electrical state characteristics to characterize the arc state.

[0027] Temperature field distribution analysis and molten pool feature extraction are performed on the spatiotemporally registered temperature field distribution data. Thermal quality parameters used to characterize the quality of the temperature field data itself and initial thermal state features used to characterize the state of the molten pool are extracted.

[0028] Furthermore, the feature fusion network includes an input layer, a feature encoder, a confidence generator, a cross-modal attention interaction layer, a softmax normalization layer, and a weighted fusion layer;

[0029] There are multiple feature encoders connected in parallel, each feature encoder corresponding to the state features of a mode; there are multiple confidence generators connected in parallel, each confidence generator corresponding to the quality parameters of a mode;

[0030] The input layer is used to group and distinguish the quality parameters and state features of each modality at the same time of input. The initial visual state features, initial mechanical state features, initial electrical state features, and initial thermal state features are respectively fed into the corresponding feature encoders for dimensional transformation or nonlinear mapping to obtain geometric intermediate feature vectors, mechanical intermediate feature vectors, electrical intermediate feature vectors, and thermal intermediate feature vectors with unified dimensions. The point cloud quality parameters, mechanical quality parameters, electrical quality parameters, and thermal quality parameters are fed into the corresponding confidence generators for confidence calculation to obtain the initial geometric feature confidence, mechanical feature confidence, electrical feature confidence, and thermal feature confidence, respectively.

[0031] The cross-modal attention interaction layer is used to perform cross-modal association and adaptive correction on the initial geometric feature confidence, mechanical feature confidence, electrical feature confidence, and thermal feature confidence, and output the corrected geometric feature confidence, mechanical feature confidence, electrical feature confidence, and thermal feature confidence.

[0032] The Softmax normalization layer is used to normalize the corrected confidence scores of geometric features, mechanical features, electrical features, and thermal features, and outputs a set of weight coefficients that sum to 1 as fusion weights; the higher the confidence score, the larger the corresponding fusion weight.

[0033] The weighted fusion layer is used to multiply the geometric intermediate feature vector, mechanical intermediate feature vector, electrical intermediate feature vector, and thermal intermediate feature vector element by element with their respective fusion weights and sum them to generate a set of optimized features for the current time step.

[0034] The cross-modal attention interaction layer is used to establish the intermodal interaction coefficients based on the correlation strength between the initial geometric feature confidence, mechanical feature confidence, electrical feature confidence, and thermal feature confidence. Based on the interaction coefficients, the layer performs consistency verification on the confidence of each modality, suppresses confidence that contradicts the global operating condition trend, enhances confidence that is consistent with the global operating condition trend, and outputs the corrected confidence of each modality.

[0035] Furthermore, the confidence generator includes:

[0036] The normalization sublayer is used to normalize the mass parameters of the corresponding input modes to the 0-1 interval, thereby obtaining the corresponding normalized mass scores.

[0037] The confidence mapping sublayer is used to calculate the initial confidence of the corresponding mode based on the normalized quality score. For any mode, when the quality parameter of the mode contains only one item, the normalized quality score corresponding to the quality parameter is used as the initial confidence of the mode. When the quality parameter of the mode contains multiple items, the initial confidence of the mode is calculated by weighted averaging, taking the minimum value, or taking the product of the normalized quality scores of each item.

[0038] Furthermore, the decision model includes a shared feature layer, a task branch layer, and a parameter output layer connected in sequence. The task branch layer includes a first regression branch, a second regression branch, and a third classification branch set in parallel.

[0039] The shared feature layer is used to perform nonlinear transformations on the input optimization features to extract shared decision features;

[0040] The first regression branch receives the shared decision features and calculates the continuous three-dimensional spatial trajectory coordinates of the weld seam through at least one fully connected network layer.

[0041] The second regression branch receives shared decision features and calculates continuous weld geometry parameter predictions through at least one fully connected network layer. The weld geometry parameter predictions include weld lateral deviation predictions, height deviation predictions, and groove width and depth predictions.

[0042] The third classification branch receives shared decision features, which are activated by a Softmax layer after passing through at least one fully connected network to calculate the discrete welding torch tilt state, which includes the welding torch tilting to the left, centering, or right.

[0043] The parameter output layer is used to assemble weld geometry parameters and welding torch attitude state into weld tracking parameters in a preset format.

[0044] A real-time weld seam perception system for welding robots based on multimodal data fusion includes:

[0045] The multimodal data acquisition and preprocessing module is used to simultaneously acquire various types of sensor data during the welding process, including: three-dimensional point cloud data of the weld area in front of the welding torch acquired by a three-dimensional vision sensor; welding process data, including mechanical process data and electrical process data, with mechanical process data acquired by a torque sensor installed at the end of the welding torch; electrical process data including welding current signals acquired from the output of the welding power source and arc voltage signals acquired between the welding torch contact tip and the workpiece; temperature field distribution data of the weld pool and its surrounding predetermined area acquired by a temperature sensor or thermal imager; and adding a unified timestamp to the acquired sensor data and performing filtering preprocessing to obtain initial multimodal data.

[0046] The spatiotemporal registration and feature extraction module is used to unify the initial multimodal data into the same coordinate system based on the welding gun, and to align the time axis based on the timestamp to ensure that the initial multimodal data collected at the same time correspond one-to-one. This allows the three-dimensional point cloud data, welding process data and temperature field distribution data at the same time to form corresponding data groups to complete the spatiotemporal registration. Then, the quality parameters and state features of each mode are extracted from them.

[0047] The dynamic weight and feature fusion module is used to input the quality parameters and state features of each modality at the same time into a pre-constructed feature fusion network. The feature fusion network dynamically generates fusion weights corresponding to each modality based on the quality parameters of each modality, and uses the fusion weights to dynamically weight the state features of each modality to generate and output a set of optimized features at the current time. Among them, the quality parameters are used to characterize the credibility of the state features of their respective modalities at the current time.

[0048] The weld parameter calculation and output module is used to input optimized features into a pre-built decision model to calculate the weld tracking parameters required for real-time control of the welding robot. These parameters include the weld's three-dimensional spatial trajectory coordinates, predicted weld bevel width and depth, and predicted position and attitude deviation of the welding torch relative to the weld centerline. The weld tracking parameters are transmitted to the welding robot's motion controller, enabling the controller to guide the welding torch to perform adaptive welding based on these parameters.

[0049] An electronic device includes at least one processor and a memory; wherein the memory stores computer-executable instructions; the at least one processor executes the computer-executable instructions stored in the memory, causing the at least one processor to execute the multimodal data fusion welding robot weld seam real-time perception method.

[0050] A computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, controls the device where the storage medium is located to perform the multimodal data fusion welding robot weld seam real-time perception method.

[0051] The beneficial effects of this invention are as follows:

[0052] 1. This invention collects multimodal data for spatiotemporal registration and feature extraction, and then uses dynamic weights for weighted fusion to calculate weld seam tracking parameters. It has the advantages of strong anti-interference ability and high perception robustness. When a certain mode is disturbed, the contribution of that mode can be automatically reduced by dynamic weights and supplemented by other reliable modes, thereby ensuring the continuity and stability of weld seam perception.

[0053] 2. Different modal sensors have different measurement errors and noise characteristics. This invention uses dynamic weight fusion to complementarily correct multi-source information, effectively suppressing outliers in a single modality. This results in the accuracy of the final output parameters, such as the weld seam three-dimensional trajectory, groove size, and position deviation, being significantly higher than the original output of any single sensor.

[0054] 3. Compared with traditional fixed weights or simple rule-based switching methods, this invention dynamically generates fusion weights based on the quality parameters of each mode (such as point cloud noise ratio, force signal signal-to-noise ratio, temperature gradient change rate, etc.), realizing automatic adaptation to changes in operating conditions. It can respond more delicately to continuous changes in interference intensity and avoid control abrupt changes caused by hard switching. Attached Figure Description

[0055] Figure 1 This is a flowchart of the method steps in an embodiment of this application.

[0056] Figure 2 This is a structural diagram of the integrated design of feature fusion network and decision model in an embodiment of this application.

[0057] Figure 3 This is a system structure block diagram of an embodiment of this application. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the implementation methods of the present invention will be described in detail below with reference to the accompanying drawings. However, the embodiments described in this invention are only some embodiments of the present invention, and not all embodiments.

[0059] One aspect of this application provides a method for real-time weld seam perception in welding robots based on multimodal data fusion, such as... Figure 1 As shown, it includes the following steps:

[0060] Step 1: Multimodal data acquisition and preprocessing:

[0061] First, multiple types of sensor data are collected simultaneously during the welding process, including 3D point cloud data, welding process data, and temperature field distribution data.

[0062] Among them, the three-dimensional point cloud data of the weld area in front of the welding torch is acquired by a three-dimensional vision sensor. For example, it is installed at an angle of 15°-45° on the neck of the welding torch or on the side of the six-axis flange of the welding robot, or installed on the side of the torch body through a fixed or rotating bracket. The acquisition frequency can be set to 30 frames / second-60 frames / second, and the point cloud resolution is not less than 0.1mm.

[0063] The welding process data includes mechanical process data and electrical process data. The mechanical process data is obtained by measuring the torque sensor installed at the end of the welding torch, which can collect the three-dimensional force (Fx, Fy, Fz) and three-dimensional torque (Mx, My, Mz) at the end of the welding torch, and the sampling frequency can be set to 1000Hz. The electrical process data includes the welding current signal collected from the output end of the welding power source and the arc voltage signal collected from the contact tip of the welding torch and the workpiece, and the sampling frequency can be set to 5000Hz.

[0064] Temperature field distribution data of the weld pool and its surrounding predetermined area, acquired by a temperature sensor or thermal imager; the temperature sensor or thermal imager is mounted on the neck of the welding torch and moves with the torch, or mounted sideways to the flange of the welding robot. The acquisition frequency is 30 frames / second to 60 frames / second, and the temperature resolution is not less than 0.1℃.

[0065] Then, a uniform timestamp (format: yyyy-MM-dd HH:mm:ss.ffffff, accurate to microseconds) is added to the collected sensor data, and filtering preprocessing is performed to obtain the initial multimodal data.

[0066] The filtering preprocessing of 3D point cloud data includes at least one of the following: using pass-through filtering, setting the effective range of the X / Y / Z axes, such as 100mm in front of the welding torch, and removing invalid point clouds that exceed the range to coarsely screen outout points; using radius filtering, setting the core point radius, and removing isolated points whose neighboring points within the radius are smaller than the threshold to eliminate isolated noise points generated by spatter; using statistical filtering, calculating the neighborhood standard deviation of each point, and removing noise points that are greater than twice the standard deviation to eliminate arc light / smoke noise; and using voxel downsampling filtering, dividing the point cloud space into fixed-size voxels, retaining the center point of each voxel, reducing the amount of data, and reducing the dimensionality of the point cloud without losing features.

[0067] For mechanical process data, a Butterworth low-pass filter is used with a cutoff frequency of 200Hz to filter out high-frequency vibrations and mechanical noise. For electrical process data, a median filtering algorithm is used to remove electromagnetic interference, arc spikes and other pulse noise, while retaining the true waveform characteristics of droplet transition and arc state while removing noise.

[0068] For temperature field distribution data: Gaussian filtering algorithm is used to smooth the temperature distribution image and suppress random noise.

[0069] Step 2: Spatiotemporal registration and feature extraction

[0070] (1) Spatial registration of spatiotemporal registration

[0071] The initial multimodal data were unified into the same coordinate system with the welding torch as the reference, mainly involving three-dimensional point cloud data, mechanical process data, and temperature field distribution data.

[0072] This example uses a tool coordinate system with the welding torch tip as the origin (0,0,0) as the reference. The X-axis points in the direction of the welding torch axis, towards the direction of weld advancement; the Y-axis is perpendicular to the X-axis and points horizontally towards the side of the weld bevel; the Z-axis is perpendicular to the XY plane and points vertically upward, using the right-hand rule.

[0073] The mechanical process data itself is the tool coordinate system, so no transformation is needed. Specifically, the coordinate system of the 3D sensor / 3D camera for 3D point cloud data, the pixel coordinate system of the thermal imager for temperature field distribution data, or the coordinate system of the temperature sensor / thermal imager itself is transformed into the tool coordinate system.

[0074] Wherein, the coordinates ps of the 3D point cloud data are transformed by the transformation matrix T SW Transform to the tool coordinate system to get pw 3D ,pw 3D =ps·T SW , among which, T SW The solution is obtained through hand-eye calibration. First, a high-precision calibration board is prepared. Then, a robot drives a 3D sensor / 3D camera to take pictures of the calibration board in different poses and collect multiple sets of data to obtain the solution.

[0075] For temperature field distribution data, the filtered temperature field matrix (pixel coordinates + corresponding temperature) is input and transformed in the following three steps: 1) Calculate the thermal imager coordinate system from pixel coordinates: Distortion is eliminated using the thermal imager's own intrinsic parameter matrix, and normalized coordinates in the thermal imager coordinate system are calculated; 2) Solve for spatial coordinates using the thermal imager coordinate system: Spatial coordinates (x, y, z) are solved by combining the depth values ​​of the 3D point cloud data. t ,y t ,z t 3) Obtain pw by transforming the spatial coordinates to the tool coordinate system. 热 ,pw 热 =T TW ·[x t ,y t ,z t ,1] T superscript T The transpose symbol is used; the extrinsic matrix T TW It can be obtained through hand-eye calibration.

[0076] (2) Time alignment in spatiotemporal registration

[0077] Timeline alignment is performed based on timestamps to ensure a one-to-one correspondence between initial multimodal data acquired at the same time. This ensures that 3D point cloud data, welding process data, and temperature field distribution data at the same time form corresponding data sets, thus completing spatiotemporal registration. Optionally, linear interpolation is used to align low-frequency 3D point cloud data and temperature field distribution data with high-frequency mechanical and electrical process data, ensuring that various types of data at the same time form corresponding data sets. Specifically, for any time t, if the 3D point cloud data acquisition time is t_vis, then the 3D point cloud data at time t is obtained through linear interpolation.

[0078] (3) Feature extraction

[0079] The quality parameters and state features of each mode are extracted from the initial multimodal data after spatiotemporal registration is completed.

[0080] Specifically, 3D point cloud geometric analysis and weld area feature analysis are performed on the spatiotemporally registered 3D point cloud data to extract point cloud quality parameters characterizing the quality of the point cloud data itself and initial visual state features characterizing the weld geometry. The point cloud quality parameters include at least one of the following: bevel edge sharpness (edge ​​detection is performed on the 3D point cloud data of the weld area, and the average gradient magnitude of edge points is calculated as an edge sharpness index); noise point ratio (the proportion of outliers in the 3D point cloud data is calculated, with outliers defined as those exceeding a preset threshold (e.g., 2mm from their nearest neighbor); and effective point ratio (the proportion of effective points in the 3D point cloud data is calculated, with effective points defined as those with depth values ​​within a reasonable range and reflection intensity exceeding a threshold).

[0081] Initial visual characteristics include the initial value of the lateral deviation of the weld center, ∆x. in Initial value of welding torch height deviation ∆z in Initial value of bevel width W in With initial depth D in The initial value of the weld trajectory position is calculated. Specifically, the weld region is extracted by cropping the ROI (Region of Interest) from the 3D point cloud data; the left and right edge points of the bevel are extracted by fitting the bevel plane using the RANSAC (Random Sample Consensus) algorithm; the weld centerline is fitted based on the edge points, and the spatial trajectory of the centerline is calculated as the initial value of the weld trajectory position; the initial value of the lateral deviation of the weld center, ∆x, is calculated. in (i.e., the horizontal distance between the centerline of the welding torch and the centerline of the weld), initial value of the welding torch height ∆z in (i.e., the vertical distance between the welding torch tip and the weld surface), initial width value W in With initial depth D in .

[0082] Specifically, the spatiotemporally registered mechanical process data undergoes time-domain force signal analysis and contact state characteristic analysis to extract mechanical quality parameters characterizing the quality of the mechanical signal itself and initial mechanical state characteristics characterizing the welding torch contact state. The mechanical quality parameters include at least one of the following: force signal signal-to-noise ratio (SNR), extracted by calculating the ratio of the force signal's mean to its standard deviation; torque drift, extracted by calculating the amplitude of the linear trend term of the torque signal within a 1-second time window; impact anomalies, extracted by the number of pulses exceeding the mean ± 3σ in the statistical force signal, where σ is the standard deviation; and contact stability, extracted by the energy percentage of the calculated force signal in a predetermined frequency range (e.g., 10Hz-100Hz), with a higher percentage indicating more stable contact.

[0083] The initial mechanical state characteristics include: welding torch normal torque, lateral offset torque, force fluctuation amplitude, and weld contact deviation; wherein, the welding torch normal torque is obtained by extracting the torque component in the direction of the welding torch axis; the lateral offset torque is obtained by extracting the torque component in the direction perpendicular to the welding torch axis; the force fluctuation amplitude is obtained by calculating the peak value of the force signal; and the weld contact deviation is obtained by solving the contact position deviation between the welding torch and the bevel sidewall through the force / torque balance equation.

[0084] Specifically, time-domain statistics of electrical signals and arc stability analysis are performed on the spatiotemporally registered electrical process data to extract electrical quality parameters characterizing the quality of the electrical signals themselves and initial electrical state characteristics characterizing the arc state. Among them, the electrical quality parameters include at least one of the following: current-voltage fluctuation coefficient: calculated as the ratio of the standard deviation to the mean of the current / voltage signal; short-circuit frequency: obtained by statistically analyzing the number of short-circuit events per unit time; arcing ratio: obtained by calculating the proportion of arcing time to the total time; deviation from the set value: obtained by calculating the relative deviation between the measured current / voltage and the set value.

[0085] Initial electrical state characteristics include real-time welding current, instantaneous values ​​of arc voltage, droplet transfer characteristics, and arc position offset. Droplet transfer characteristics are extracted by identifying short-circuit transfer, droplet transfer, or spray transfer modes based on the current waveform. Arc position offset is extracted by estimating the lateral offset of the arc within the bevel through the relationship between arc voltage and welding torch position. For example, if the arc is in the middle of the bevel, the arc length and voltage are stable; if the arc is deviated towards the bevel sidewall, the arc is compressed or elongated, and the voltage shifts.

[0086] Temperature field distribution analysis and molten pool feature extraction are performed on the spatiotemporally registered temperature field distribution data. Thermal quality parameters characterizing the quality of the temperature field data itself and initial thermal state features characterizing the molten pool state are extracted. The thermal quality parameters include the regularity of the molten pool isotherms and / or the rate of change of the temperature gradient. The initial thermal state features include the temperature gradient, the width of the heat-affected zone, the center position of the molten pool, and the size of the molten pool. Among these:

[0087] Regularity of molten pool isotherms: Extract isotherms (e.g., the 1500℃ isotherm) from the molten pool region and calculate the circularity or elliptic fitting error of the isotherms. Higher regularity indicates a more stable molten pool morphology. Temperature gradient change rate: Extracted by calculating the time change rate of the temperature gradient at the edge of the molten pool. A smaller change rate indicates a more stable heat input. Temperature gradient: Extracted by calculating the spatial temperature gradient field of the molten pool region. Width of heat-affected zone: Obtained by extracting the boundary of the heat-affected zone and calculating its width. Center position of molten pool: Extracted by calculating the centroid coordinates of the molten pool temperature field. Dimensions of molten pool: Obtained by calculating the length and width of the molten pool.

[0088] Step 3: Feature Fusion Based on Dynamic Weights

[0089] The quality parameters and state features of each modality at the same time are input into a pre-constructed feature fusion network. The feature fusion network dynamically generates fusion weights corresponding to each modality based on the quality parameters of each modality, and uses the fusion weights to dynamically weight the state features of each modality to generate and output a set of optimized features at the current time. Among them, the quality parameters are used to characterize the credibility of the state features of their respective modalities at the current time.

[0090] Specifically, such as Figure 2 As shown, the feature fusion network includes an input layer, a feature encoder, a confidence generator, a cross-modal attention interaction layer, a softmax normalization layer, and a weighted fusion layer.

[0091] There are multiple feature encoders connected in parallel, each corresponding to the state features of a mode; there are multiple confidence generators connected in parallel, each corresponding to the quality parameters of a mode.

[0092] The input layer is used to group and distinguish the quality parameters and state features of each modality at the same time of input. The initial visual state features, initial mechanical state features, initial electrical state features, and initial thermal state features are respectively fed into the corresponding feature encoders for dimensional transformation or nonlinear mapping to obtain geometric intermediate feature vectors, mechanical intermediate feature vectors, electrical intermediate feature vectors, and thermal intermediate feature vectors with unified dimensions. The point cloud quality parameters, mechanical quality parameters, electrical quality parameters, and thermal quality parameters are fed into the corresponding confidence generators for confidence calculation to obtain the initial geometric feature confidence, mechanical feature confidence, electrical feature confidence, and thermal feature confidence, respectively.

[0093] Specifically, the confidence generator includes a normalization sublayer and a confidence mapping sublayer. The normalization sublayer is used to normalize the quality parameters of the corresponding modes of the input to the 0-1 interval, thereby obtaining the corresponding normalized quality scores.

[0094] For the visual modality, the noise point ratio in the point cloud quality parameters is complemented, the effective point ratio remains unchanged, and the edge sharpness is linearly scaled to obtain the noise point quality score, the effective point quality score, and the edge sharpness quality score, respectively.

[0095] For the mechanical modes, the upper limit truncation normalization, torque drift compensation, impact anomaly compensation, and contact stability are performed on the force signal signal-to-noise ratio in the mechanical mass parameters to obtain the signal-to-noise ratio mass score, drift mass score, impact mass score, and contact stability mass score, respectively.

[0096] For the electrical modes, the current and voltage fluctuation coefficients, short-circuit frequency, arcing ratio, and deviation from the set value are complemented in the electrical quality parameters to obtain fluctuation quality score, short-circuit quality score, arcing quality score, and deviation quality score, respectively; wherein, the short-circuit frequency is appropriately converted, specifically including: preset the ideal range of the short-circuit frequency […]. f _min, f [_max], for example, for short-circuit transition welding, f _min=50Hz, f _max=150Hz, when the measured short-circuit frequency f When the short circuit quality falls within this range, the short circuit quality is 1; when f Below f At _min, the short-circuit quality is divided into f / f _min; when f Higher than f When _max, the short-circuit quality is divided into f _max / f The process involves a moderate conversion of the arc ratio, specifically including: presetting an ideal value r_opt for the arc ratio, for example, for pulse welding, r_opt=0.7, and the arc quality is divided into 1−|r−r_opt| / max(r_opt,1−r_opt), where r is the measured arc ratio, || is the absolute value sign, and max() is to take the maximum value.

[0097] For the thermal mode, the regularity of the molten pool isotherm remains unchanged and the temperature gradient change rate is complemented in the thermal quality parameters to obtain the regularity mass score and the gradient change mass score, respectively.

[0098] The confidence mapping sublayer is used to calculate the initial confidence of the corresponding mode based on the normalized quality score. For any mode, when the quality parameter of the mode contains only one item, the normalized quality score corresponding to the quality parameter is used as the initial confidence of the mode. When the quality parameter of the mode contains multiple items, the initial confidence of the mode is calculated by weighted averaging, taking the minimum value, or taking the product of the normalized quality scores of each item.

[0099] The confidence generator in this example first unifies the quality parameters with different dimensions to the 0-1 range, making subsequent operations such as weighted averaging and minimum value taking mathematically valid, and also making the neural network mapping easier to converge. However, if the original quality parameters are directly weighted and averaged without normalization, the different dimensions will cause the weights to lose their meaning.

[0100] The cross-modal attention interaction layer is used to perform cross-modal correlation and adaptive correction on the initial confidence levels of geometric, mechanical, electrical, and thermal features, outputting the corrected confidence levels for these features. Specifically, the cross-modal attention interaction layer establishes intermodal mutual influence coefficients based on the correlation strength among the initial confidence levels of geometric, mechanical, electrical, and thermal features; and based on these mutual influence coefficients, performs consistency checks on the confidence levels of each modality, suppressing confidence levels that contradict the global operating condition trend and enhancing confidence levels that are consistent with the global operating condition trend, outputting the corrected confidence levels for each modality.

[0101] The Softmax normalization layer is used to normalize the corrected confidence scores of geometric features, mechanical features, electrical features, and thermal features, and outputs a set of weight coefficients that sum to 1 as fusion weights; the higher the confidence score, the larger the corresponding fusion weight.

[0102] The weighted fusion layer is used to multiply the geometric intermediate feature vector, mechanical intermediate feature vector, electrical intermediate feature vector, and thermal intermediate feature vector element by element with their respective fusion weights and sum them to generate a set of optimized features for the current time step.

[0103] For example, the confidence generator calculates the initial confidence of each modality based on the quality parameters. For instance, when encountering a splash, the confidence of geometric features drops to 0.3, while the confidence of mechanical features remains at 0.8. The cross-modal attention interaction layer calibrates the confidence, and the confidence of geometric features is mechanically calibrated to 0.5. The Softmax normalization layer generates fusion weights, where the visual weight is 0.2, the mechanical weight is 0.5, the electrical weight is 0.2, and the thermal weight is 0.1. The weighted fusion layer generates a 64-dimensional fusion feature vector.

[0104] This example's feature fusion network distinguishes between quality parameters and state features. Based on the confidence scores of the quality parameters for each modality, it achieves a learnable mapping from data quality to fusion weights, which is more flexible and accurate than fixed thresholds. Furthermore, it performs cross-modal attention interaction correction on the confidence scores to avoid prematurely discarding valid information due to transient interference. This is more consistent with actual physical processes than independently evaluating the confidence scores of each modality. Then, based on the normalized confidence scores as fusion weights, the contributions of each modality are directly comparable. Finally, dynamic weighting of state features is performed, achieving true soft fusion. This solves the problems of lacking explicit modeling of signal quality, poor interpretability, and difficulty in guaranteeing robustness when a certain modality completely fails, which are inherent in simple feature concatenation followed by fully connected layers.

[0105] Step 4: Weld Parameter Calculation and Output

[0106] The optimized features are input into a pre-built decision model to calculate the weld seam tracking parameters required for real-time control of the welding robot. These parameters include the weld seam's three-dimensional spatial trajectory coordinates and the predicted weld bevel width W. out With depth prediction value D out The predicted positional deviation of the welding torch relative to the weld centerline. This predicted positional deviation includes the predicted lateral deviation ∆x. out Predicted height deviation ∆z out The welding torch tilt angle is set to left tilt / center tilt / right tilt. Weld seam tracking parameters are transmitted to the welding robot's motion controller, enabling the motion controller to guide the welding torch to perform adaptive welding based on these parameters.

[0107] Optional, such as Figure 2 As shown, the decision model includes a shared feature layer, a task branch layer, and a parameter output layer connected in sequence. The task branch layer includes a first regression branch, a second regression branch, and a third classification branch set in parallel.

[0108] The shared feature layer is used to perform nonlinear transformations on the input optimization features to extract shared decision features;

[0109] The first regression branch receives the shared decision features and calculates the continuous three-dimensional spatial trajectory coordinates of the weld seam through at least one fully connected network layer.

[0110] The second regression branch receives shared decision features and calculates continuous weld geometry parameter predictions through at least one fully connected network layer. The weld geometry parameter predictions include weld lateral deviation predictions, height deviation predictions, and groove width and depth predictions.

[0111] The third classification branch receives shared decision features, which are activated by a Softmax layer after passing through at least one fully connected network to calculate the discrete welding torch tilt state, which includes the welding torch tilting to the left, centering, or right.

[0112] The parameter output layer is used to assemble the predicted values ​​of weld geometry parameters and the welding torch tilt state into weld tracking parameters in a preset format.

[0113] The decision model in this example uses a shared feature layer to allow the regression task (bias, size) and the classification task (welding torch tilt) to share underlying features. Multi-task learning enhances the generalization ability of the shared features while reducing computational cost. Decoupling the regression and classification branches avoids gradient conflicts between different tasks, resulting in better performance for each task. Furthermore, the discrete states (left tilt / center / right tilt) output by the classification branch can serve as auxiliary information for the control strategy; for example, prioritizing left-side compensation during left tilt improves control smoothness. If a single regression network is used to predict both continuous parameters and discrete states simultaneously, the discrete states need to be forcibly converted to continuous values ​​(e.g., left tilt = 0, center = 1, right tilt = 2), which introduces incorrect numerical ordering and misleads the network. This example avoids this problem by decoupling the network into two branches.

[0114] As a preferred option, such as Figure 2 As shown, the feature fusion network and the decision model are integrated into a unified end-to-end neural network. The weighted fusion layer is connected to the shared feature layer, and the entire network is trained jointly. The loss function is: L=L_reg+λL_cls, where L_reg is the regression loss, L_cls is the classification loss, and λ is the balance coefficient, which can be set to 0.1.

[0115] Specifically: L_reg = 1 / 4[(Δx) out -Δx true ) 2 +(ΔZ out -ΔZ true ) 2 +(W out -W true ) 2 +(D out -D true ) 2 ], where Δx true ΔZ true W true D true These are the actual values ​​of weld transverse deviation, weld height deviation, groove width, and groove depth obtained during offline measurement during training.

[0116] The classification loss L_cls can be calculated as follows:

[0117] 1) First determine the encoding of the true pose; for example, if the true pose is "left tilt", then the encoding is (1,0,0); if the true pose is "center", then the encoding is (0,1,0); if the true pose is "right tilt", then the encoding is (0,0,1).

[0118] 2) After the output of the third classification branch is processed by the Softmax layer function, three probability values ​​between 0 and 1 are obtained, which represent the probability that the model judges it as left-skewed, center-skewed, and right-skewed, respectively, and the sum of the three probabilities is 1; for example, the output may be (0.2, 0.7, 0.1), which means that the model thinks there is a 20% probability of left-skewed, a 70% probability of center-skewed, and a 10% probability of right-skewed.

[0119] 3) Calculate the loss value by multiplying the encoding of the true pose by the corresponding positions of the model's predicted probabilities. Only the predicted probabilities corresponding to the true class are retained. Then, take the negative logarithm of this probability. The other positions do not contribute because they are multiplied by 0. For example, if the encoding is (0,1,0) and the model's predicted probabilities are (0.2,0.7,0.1), then:

[0120] The classification loss L_cls = -(0·log(0.2) + 1·log(0.7) + 0·log(0.1)) = -log(0.7) = 0.3567;

[0121] If the model predicts accurately, the loss is small. For example, if the predicted probability is (0.05, 0.9, 0.05), L_cls = -log(0.9) ≈ 0.1054. If the model predicts incorrectly, the loss is large. For example, if the predicted probability is (0.7, 0.2, 0.1), L_cls = -log(0.2) ≈ 1.6094.

[0122] Compared to optimizing each module independently or in combination with decision-making separately, the global end-to-end optimization in this example avoids the trap of local optima. The integrated design enables the mechanism of dynamically generating fusion weights through quality parameters to form a closed loop with the final control objective, ensuring that the learning of the confidence generator is not carried out in isolation, but is directly linked to the final welding quality.

[0123] Through an integrated network, the parameter updates of the cross-modal attention interaction layer are supervised by the final weld seam tracking parameters. This means that the attention interaction strategy, which uses mechanical confidence to correct visual confidence, is optimized directly with the final accuracy as the goal, rather than just with confidence consistency. For example, when the visual confidence is moderate (0.5) and the mechanical confidence is high (0.9), the visual confidence is increased to 0.7 instead of simply averaging.

[0124] Another aspect of this application provides a real-time weld seam perception system for welding robots based on multimodal data fusion, such as... Figure 3As shown, it includes: a multimodal data acquisition and preprocessing module, a spatiotemporal registration and feature extraction module, a dynamic weighting and feature fusion module, and a weld parameter calculation and output module.

[0125] The multimodal data acquisition and preprocessing module is used to simultaneously acquire various types of sensor data during the welding process, including: three-dimensional point cloud data of the weld area in front of the welding torch; welding process data, including mechanical process data and electrical process data. The mechanical process data is obtained by measuring the torque sensor installed at the end of the welding torch, and the electrical process data includes the welding current signal acquired from the output end of the welding power source and the arc voltage signal acquired from the contact tip of the welding torch and the workpiece; and temperature field distribution data of the weld pool and its surrounding predetermined area. The module is used to add a unified timestamp to the acquired sensor data and perform filtering preprocessing to obtain the initial multimodal data.

[0126] The spatiotemporal registration and feature extraction module is used to unify the initial multimodal data into the same coordinate system based on the welding torch, and to align the time axis based on the timestamp to ensure that the initial multimodal data collected at the same time correspond one-to-one. This allows the three-dimensional point cloud data, welding process data and temperature field distribution data at the same time to form corresponding data groups to complete the spatiotemporal registration. Then, the quality parameters and state features of each mode are extracted from them.

[0127] The dynamic weight and feature fusion module is used to input the quality parameters and state features of each modality at the same time into a pre-constructed feature fusion network. The feature fusion network dynamically generates fusion weights corresponding to each modality based on the quality parameters of each modality, and uses the fusion weights to dynamically weight the state features of each modality to generate and output a set of optimized features at the current time. Among them, the quality parameters are used to characterize the credibility of the state features of their respective modalities at the current time.

[0128] The weld parameter calculation and output module is used to input optimized features into a pre-built decision model to calculate the weld tracking parameters required for real-time control of the welding robot. These parameters include the weld's three-dimensional spatial trajectory coordinates, predicted weld bevel width and depth, and predicted position and attitude deviation of the welding torch relative to the weld centerline. The weld tracking parameters are then transmitted to the welding robot's motion controller, enabling the controller to guide the welding torch to perform adaptive welding based on these parameters.

[0129] In another aspect of this application, an electronic device is provided, including at least one processor and a memory; wherein the memory stores computer execution instructions; the at least one processor executes the computer execution instructions stored in the memory, causing the at least one processor to perform the multimodal data fusion welding robot weld seam real-time perception method as described in the foregoing embodiments.

[0130] In another aspect of this application, a computer-readable storage medium is provided, on which a computer program is stored, wherein the computer program, when executed by a processor, controls the device where the storage medium is located to perform the real-time sensing method for weld seams of a welding robot using multimodal data fusion as described in the preceding embodiments.

[0131] The solution in this application addresses the challenges of sensing in complex environments such as smoke, splashes, reflections, and uneven workpiece surfaces. When a single sensor (such as vision) is subjected to strong interference, it can rely on other sensors (mechanical, electrical, and temperature) to maintain basic sensing capabilities and dynamically adjust the confidence level, significantly improving robustness under complex working conditions such as smoke, splashes, and reflections. By integrating three-dimensional geometric, mechanical, electrical processes, and thermodynamic state information, it can provide a more comprehensive and accurate description of weld features compared to a single sensor, providing more basis for adjusting high-quality welding processes.

[0132] The above description is only a preferred embodiment of this application and is not intended to limit this application. Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application.

Claims

1. A method for real-time weld seam perception in welding robots based on multimodal data fusion, characterized in that, include: Step 1: Multimodal data acquisition and preprocessing: Simultaneously acquire multiple types of sensor data during the welding process, including: Three-dimensional point cloud data of the weld seam area in front of the welding torch; Welding process data includes mechanical process data and electrical process data. Mechanical process data is obtained by measuring a torque sensor installed at the end of the welding torch. Electrical process data includes welding current signals collected from the output of the welding power source and arc voltage signals collected between the welding torch contact tip and the workpiece. Temperature field distribution data of the weld pool and its surrounding predetermined area; A unified timestamp is added to the collected sensor data, and filtering preprocessing is performed separately to obtain the initial multimodal data; Step 2: Spatiotemporal registration and feature extraction The initial multimodal data are unified into the same coordinate system with the welding gun as the reference, and the time axis is aligned based on the timestamp to ensure that the initial multimodal data collected at the same time correspond one-to-one, so as to complete the spatiotemporal registration, and then the quality parameters and state characteristics of each mode are extracted from it. Step 3: Feature Fusion Based on Dynamic Weights The quality parameters and state features of each modality at the same time are input into a pre-constructed feature fusion network. The feature fusion network dynamically generates fusion weights corresponding to each modality based on the quality parameters of each modality, and uses the fusion weights to dynamically weight the state features of each modality to generate and output a set of optimized features at the current time. Among them, the quality parameters are used to characterize the credibility of the state features of their respective modalities at the current time. Step 4: Weld Parameter Calculation and Output The optimized features are input into a pre-built decision model to calculate the weld tracking parameters required for real-time control of the welding robot. The weld tracking parameters include the three-dimensional spatial trajectory coordinates of the weld, the predicted values ​​of the weld groove width and depth, and the predicted values ​​of the position and attitude deviation of the welding torch relative to the weld centerline. The quality parameters and state characteristics of each mode are extracted, including: Three-dimensional point cloud geometric analysis and weld area feature analysis are performed on spatiotemporally registered three-dimensional point cloud data to extract point cloud quality parameters to characterize the quality of the point cloud data itself and initial visual state features to characterize the weld geometry. Force signal time-domain analysis and contact state characteristic analysis are performed on the spatiotemporally registered mechanical process data to extract mechanical quality parameters to characterize the quality of the mechanical signal itself and initial mechanical state characteristics to characterize the contact state of the welding torch. We perform time-domain statistics and arc stability analysis on the spatiotemporally registered electrical process data to extract electrical quality parameters to characterize the quality of the electrical signal itself and initial electrical state characteristics to characterize the arc state. Temperature field distribution analysis and molten pool feature extraction are performed on the spatiotemporally registered temperature field distribution data to extract thermal quality parameters for characterizing the quality of the temperature field data itself and initial thermal state features for characterizing the state of the molten pool. The feature fusion network includes an input layer, a feature encoder, a confidence generator, a cross-modal attention interaction layer, a softmax normalization layer, and a weighted fusion layer. There are multiple feature encoders connected in parallel, each feature encoder corresponding to the state features of a mode; there are multiple confidence generators connected in parallel, each confidence generator corresponding to the quality parameters of a mode; The input layer is used to group and distinguish the quality parameters and state features of each modality at the same time of input. The initial visual state features, initial mechanical state features, initial electrical state features, and initial thermal state features are respectively fed into the corresponding feature encoders for dimensional transformation or nonlinear mapping to obtain geometric intermediate feature vectors, mechanical intermediate feature vectors, electrical intermediate feature vectors, and thermal intermediate feature vectors with unified dimensions. The point cloud quality parameters, mechanical quality parameters, electrical quality parameters, and thermal quality parameters are fed into the corresponding confidence generators for confidence calculation to obtain the initial geometric feature confidence, mechanical feature confidence, electrical feature confidence, and thermal feature confidence, respectively. The cross-modal attention interaction layer is used to perform cross-modal association and adaptive correction on the initial geometric feature confidence, mechanical feature confidence, electrical feature confidence, and thermal feature confidence, and output the corrected geometric feature confidence, mechanical feature confidence, electrical feature confidence, and thermal feature confidence. The Softmax normalization layer is used to normalize the corrected confidence scores of geometric features, mechanical features, electrical features, and thermal features, and outputs a set of weight coefficients that sum to 1 as fusion weights; the higher the confidence score, the larger the corresponding fusion weight. The weighted fusion layer is used to multiply the geometric intermediate feature vector, mechanical intermediate feature vector, electrical intermediate feature vector, and thermal intermediate feature vector element by element with their respective fusion weights and sum them to generate a set of optimized features for the current time step.

2. The method for real-time weld seam perception of a welding robot based on multimodal data fusion according to claim 1, characterized in that, The cross-modal attention interaction layer is used to establish the intermodal interaction coefficients based on the correlation strength between the initial geometric feature confidence, mechanical feature confidence, electrical feature confidence, and thermal feature confidence. Based on the interaction coefficients, the layer performs consistency verification on the confidence of each modality, suppresses confidence that contradicts the global operating condition trend, enhances confidence that is consistent with the global operating condition trend, and outputs the corrected confidence of each modality.

3. The method for real-time weld seam perception of a welding robot based on multimodal data fusion according to claim 1, characterized in that, Point cloud quality parameters include at least one of the following: bevel edge sharpness, noise point ratio, and effective point ratio; initial visual state features include initial values ​​of weld center lateral deviation, welding torch height deviation, bevel width and depth, and weld trajectory position. Mechanical quality parameters include at least one of force signal signal-to-noise ratio, torque drift, impact anomaly, and contact stability. Initial mechanical state characteristics include welding torch normal torque, lateral offset torque, force fluctuation amplitude, and weld contact deviation. Electrical quality parameters include at least one of the following: current and voltage fluctuation coefficient, short-circuit frequency, arc ratio, and deviation from set value. Initial electrical state characteristics include real-time welding current, arc voltage, droplet transfer characteristics, and arc position offset. Thermal quality parameters include the regularity of the isotherms of the molten pool and / or the rate of change of the temperature gradient. Initial thermal state characteristics include temperature gradient, width of the heat-affected zone, center position of the molten pool, and size of the molten pool.

4. The method for real-time weld seam perception of a welding robot based on multimodal data fusion according to claim 1, characterized in that, Confidence generators include: The normalization sublayer is used to normalize the mass parameters of the corresponding input modes to the 0-1 interval, thereby obtaining the corresponding normalized mass scores. The confidence mapping sublayer is used to calculate the initial confidence of the corresponding mode based on the normalized quality score. For any mode, when the quality parameter of the mode contains only one item, the normalized quality score corresponding to the quality parameter is used as the initial confidence of the mode. When the quality parameter of the mode contains multiple items, the initial confidence of the mode is calculated by weighted averaging, taking the minimum value, or taking the product of the normalized quality scores of each item.

5. The real-time weld seam perception method for welding robots based on multimodal data fusion according to claim 4, characterized in that, The decision model consists of a shared feature layer, a task branch layer, and a parameter output layer connected in sequence. The task branch layer includes a first regression branch, a second regression branch, and a third classification branch set in parallel. The shared feature layer is used to perform nonlinear transformations on the input optimization features to extract shared decision features; The first regression branch receives the shared decision features and calculates the continuous three-dimensional spatial trajectory coordinates of the weld seam through at least one fully connected network layer. The second regression branch receives shared decision features and calculates continuous weld geometry parameter predictions through at least one fully connected network layer. The weld geometry parameter predictions include weld lateral deviation predictions, height deviation predictions, and groove width and depth predictions. The third classification branch receives shared decision features, which are activated by a Softmax layer after passing through at least one fully connected network to calculate the discrete welding torch tilt state, which includes the welding torch tilting to the left, centering, or right. The parameter output layer is used to assemble weld geometry parameters and welding torch attitude state into weld tracking parameters in a preset format.

6. A real-time weld seam perception system for welding robots based on multimodal data fusion, characterized in that, include: The multimodal data acquisition and preprocessing module is used to simultaneously acquire various types of sensor data during the welding process, including: three-dimensional point cloud data of the weld area in front of the welding torch; welding process data, including mechanical process data and electrical process data. The mechanical process data is obtained by measuring the torque sensor installed at the end of the welding torch, and the electrical process data includes the welding current signal acquired from the output end of the welding power source and the arc voltage signal acquired from the contact tip of the welding torch and the workpiece; and temperature field distribution data of the weld pool and its surrounding predetermined area. The module is used to add a unified timestamp to the acquired sensor data and perform filtering preprocessing to obtain the initial multimodal data. The spatiotemporal registration and feature extraction module unifies the initial multimodal data into the same coordinate system based on the welding torch, and aligns the time axis based on timestamps to ensure that the initial multimodal data collected at the same time correspond one-to-one, thus completing the spatiotemporal registration. Then, it extracts the quality parameters and state features of each modality. The extraction of quality parameters and state features includes: performing 3D point cloud geometric analysis and weld region feature analysis on the spatiotemporally registered 3D point cloud data to extract point cloud quality parameters characterizing the quality of the point cloud data itself and initial visual state features characterizing the weld geometry; and performing mechanical analysis on the spatiotemporal registration. The process data undergoes time-domain force signal analysis and contact state characteristic analysis to extract mechanical quality parameters characterizing the quality of the mechanical signal itself and initial mechanical state characteristics characterizing the welding torch contact state. The spatiotemporally registered electrical process data undergoes time-domain electrical signal statistics and arc stability analysis to extract electrical quality parameters characterizing the quality of the electrical signal itself and initial electrical state characteristics characterizing the arc state. The spatiotemporally registered temperature field distribution data undergoes temperature field distribution analysis and molten pool feature extraction to extract thermal quality parameters characterizing the quality of the temperature field data itself and initial thermal state characteristics characterizing the molten pool state. The dynamic weight and feature fusion module is used to input the quality parameters and state features of each modality at the same time into a pre-constructed feature fusion network. The feature fusion network dynamically generates fusion weights corresponding to each modality based on the quality parameters of each modality, and uses the fusion weights to dynamically weight the state features of each modality to generate and output a set of optimized features at the current time. Among them, the quality parameters are used to characterize the credibility of the state features of their respective modalities at the current time. The weld parameter calculation and output module is used to input the optimized features into the pre-built decision model and calculate the weld tracking parameters required for real-time control of the welding robot. The weld tracking parameters include the three-dimensional spatial trajectory coordinates of the weld, the predicted values ​​of the weld groove width and depth, and the predicted values ​​of the position and attitude deviation of the welding torch relative to the weld centerline. The feature fusion network includes an input layer, a feature encoder, a confidence generator, a cross-modal attention interaction layer, a softmax normalization layer, and a weighted fusion layer. There are multiple feature encoders connected in parallel, each feature encoder corresponding to the state features of a mode; there are multiple confidence generators connected in parallel, each confidence generator corresponding to the quality parameters of a mode; The input layer is used to group and distinguish the quality parameters and state features of each modality at the same time of input. The initial visual state features, initial mechanical state features, initial electrical state features, and initial thermal state features are respectively fed into the corresponding feature encoders for dimensional transformation or nonlinear mapping to obtain geometric intermediate feature vectors, mechanical intermediate feature vectors, electrical intermediate feature vectors, and thermal intermediate feature vectors with unified dimensions. The point cloud quality parameters, mechanical quality parameters, electrical quality parameters, and thermal quality parameters are fed into the corresponding confidence generators for confidence calculation to obtain the initial geometric feature confidence, mechanical feature confidence, electrical feature confidence, and thermal feature confidence, respectively. The cross-modal attention interaction layer is used to perform cross-modal association and adaptive correction on the initial geometric feature confidence, mechanical feature confidence, electrical feature confidence, and thermal feature confidence, and output the corrected geometric feature confidence, mechanical feature confidence, electrical feature confidence, and thermal feature confidence. The Softmax normalization layer is used to normalize the corrected confidence scores of geometric features, mechanical features, electrical features, and thermal features, and outputs a set of weight coefficients that sum to 1 as fusion weights; the higher the confidence score, the larger the corresponding fusion weight. The weighted fusion layer is used to multiply the geometric intermediate feature vector, mechanical intermediate feature vector, electrical intermediate feature vector, and thermal intermediate feature vector element by element with their respective fusion weights and sum them to generate a set of optimized features for the current time step.

7. An electronic device comprising at least one processor and a memory; wherein, The memory stores computer execution instructions; characterized in that, when the at least one processor executes the computer execution instructions stored in the memory, the at least one processor performs the multimodal data fusion welding robot weld seam real-time perception method as described in any one of claims 1-5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is run by the processor, it controls the device containing the storage medium to execute the real-time sensing method for weld seams of welding robots based on multimodal data fusion as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Welding robot system based on body perception and intelligent route generation method

    CN121572333A

  • Steel pipe welding seam regulation and control method and system, electronic equipment and storage medium

    CN121788445A