Training method of welding imitation learning model and intelligent welding method
By building a welding master-slave control system and using multimodal data to train welding imitation learning models, the problems of cumbersome operation and lack of flexibility in traditional welding robot systems are solved, efficient and precise control of welding robots is achieved, and welding quality and stability are improved.
Patent Information
- Application Number
- CN202510696750.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-28
AI Technical Summary
Traditional welding robot systems are cumbersome to operate and lack flexibility, making it difficult to adapt to changes in workpiece geometry, material characteristics and welding conditions, resulting in frequent welding defects and high difficulty in skill transfer.
The training method of welding imitation learning model is adopted. By constructing a welding master-slave control system, using optical motion capture cameras and molten pool cameras to collect multimodal data, train welding imitation learning model, and realize accurate control of the position of the welding gun at the end of the robot.
It improves the environmental adaptability and skill transfer capabilities of welding robots, reduces manual intervention, improves welding quality and stability, and reduces welding defects.
Smart Images

Figure CN120218115A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of robotic welding, and particularly to a training method for a welding imitation learning model and an intelligent welding method. Background Art
[0002] Traditional welding robot systems either require elaborate manual programming or a skilled technician to drag the robotic arm for trajectory teaching and other extensive manual operations. Both of these solutions require a significant amount of time and manpower. This process is both cumbersome and lacks flexibility, especially when dealing with complex geometries or frequent product changes. Existing welding robots are difficult to adapt to changes in workpiece geometry, material properties, and welding conditions, which may lead to welding defects and require manual intervention. Therefore, there are adaptability issues. The skill transfer of existing welding robot systems is difficult. For human welders with valuable expertise in welding technology and process control, the process of controlling welding robots using traditional methods (such as teach pendants, offline programming) is not intuitive; excellent human welders are often not proficient in programming, and the feel of dragging for teaching is very different from normal welding. In addition, it is difficult for traditional welding robot systems to capture and replicate the skills of professional welders. Therefore, it is difficult to transform the fine welding technology of humans into an efficient automated production process.
[0003] Therefore, an intelligent welding solution tailored specifically for robotic welding applications is needed to address the above limitations of existing methods and enable robots to learn welding skills more efficiently and effectively. Summary of the Invention
[0004] To this end, this application proposes a training method for a welding imitation learning model and an intelligent welding method to at least to some extent solve one of the technical problems in the related art.
[0005] The first aspect of the embodiments of this application proposes a training method for a welding imitation learning model, including the following steps: Construct a welding master-slave control system. The welding master-slave control system includes a master device, a slave device, and a working host. The master device includes a simulation welding torch and an optical motion capture camera. The positioning tool carried by the optical motion capture camera is fixed on the simulation welding torch. The slave device includes a robotic arm, a welding torch, and a molten pool camera. The welding torch is fixedly connected to the end of the robotic arm. The molten pool camera is installed on the side of the welding torch. The working host is connected to the industrial control computer of the optical motion capture camera, the molten pool camera, and the robotic arm. A data acquisition system is configured on the working host; During the welding process of operating the simulation welding torch, the working host controls the welding torch to run synchronously with the simulation welding torch, and the multi-modal data stream during the welding process is obtained through the data acquisition system. The multi-modal data stream includes the pose change amount of the positioning tool collected in real time by the optical motion capture camera, the molten pool image collected in real time by the molten pool camera, the process parameters uploaded by the welding machine of the welding torch and the industrial control computer of the robotic arm. Based on the multi-modal data stream, training sample data for training the welding imitation learning model is obtained. Based on the welding imitation learning model constructed by training with the training sample data, a trained welding imitation learning model is obtained.
[0006] An embodiment of the second aspect of the present application proposes an intelligent welding method based on a welding imitation learning model. The welding imitation learning model is obtained by the training method of the welding imitation learning model described in the first aspect. The welding method includes: Obtain the image sequence collected in real time by the molten pool camera and the corresponding process parameter sequence. Input the image sequence and the corresponding process parameter sequence into the welding imitation learning model to obtain the predicted pose change amount and process parameter adjustment amount. Convert the predicted pose change amount into the target pose in the robot base coordinate system, and then calculate the corresponding robot joint angle control instruction through the inverse kinematics solver. Send the robot joint angle control instruction to the robot controller of the robotic arm at a fixed period through the robot control interface to achieve intelligent welding.
[0007] The beneficial effects of the training method of the welding imitation learning model and the intelligent welding method provided by the present application are as follows: This solution trains the imitation learning model through the data collected by the motion capture camera and the molten pool camera, and combines the molten pool image collected by the molten pool camera to achieve precise control of the pose of the welding torch at the end of the robot, so as to solve the problems of high manual intervention, low environmental adaptability and difficult skill transfer existing in traditional welding robots.
[0008] Additional aspects and advantages of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application. Description of the Drawings
[0009] The above and / or additional aspects and advantages of the present application will become obvious and easy to understand from the following description of the embodiments in conjunction with the drawings, where: Figure 1 is a schematic flow chart of a training method of a welding imitation learning model provided by an embodiment of the present application; Figure 2 A schematic flow chart of an intelligent welding method based on a welding imitation learning model provided in an embodiment of the present application. DETAILED DESCRIPTION
[0010] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0011] The following describes the training method of the welding imitation learning model and the intelligent welding method of the embodiment of the present application with reference to the accompanying drawings.
[0012] Figure 1 A flow chart of a training method for a welding imitation learning model provided in an embodiment of the present application. Figure 1 As shown, the training method of the welding imitation learning model includes the following steps: Step S101, constructing a welding master-slave control system, which includes a master-end device, a slave-end device and a working host. The master-end device includes a simulated welding gun and an optical motion capture camera, and the positioning tool of the optical motion capture camera is fixed on the simulated welding gun; the slave-end device includes a robotic arm, a welding gun and a molten pool camera, the end of the robotic arm is fixedly connected to the welding gun, and the molten pool camera is installed on the side of the welding gun; the working host is connected to the optical motion capture camera, the molten pool camera and the industrial computer of the robotic arm, and the working host is configured with a data acquisition system.
[0013] This step constructs a welding master-slave control system so that the robot can learn welding skills by observing and imitating the welding operations of human welding experts, thereby realizing intelligent welding technology.
[0014] As an example, the dimensions, weight, and grip feel (surface texture, center of gravity position) height of the simulated welding torch are the same as those of the welding torch of the slave device, and the simulated welding torch of the master device can also be replaced by a real welding torch. The positioning tool of the optical motion capture camera includes 4 reflective balls, which form a rigid body coordinate system; the positioning tool is fixed in the middle section 20 cm away from the end of the simulated welding torch, and the positioning tool is fixed through a rigid bracket (material: aluminum alloy) to ensure the relative pose of the positioning tool and the simulated welding torch is fixed (calibration error < 0.3 mm). The optical motion capture camera is used to collect the pose change amount (Δpi = (x, y, z, r, p, y), where x / y / z is the translation amount (unit: mm), and r / p / y is the Euler angle (unit: °) around the X / Y / Z axis) of the positioning tool in real time with 6 degrees of freedom (6DoF). The robotic arm is a UR5 collaborative robotic arm (6 degrees of freedom), and the end flange of the robotic arm is fixed to the actual welding torch (such as a MIG welding torch) through a customized fixture. The molten pool camera is installed on the side of the welding torch at a 45° angle and 200 mm away from the molten pool (after optical simulation, this position can cover the entire field of view of the molten pool without occlusion), and is used to collect the molten pool image in real time to obtain the dynamic characteristics of the molten pool (molten width, molten depth, hump). The working host is the system control center, which is deployed in the master operation area and is connected to the industrial control computers of the optical motion capture camera, the molten pool camera, and the robotic arm through data lines (USB cable + network cable). It is used to display the molten pool image (30 Hz) in real time, assist the welder in observing the welding state, and synchronously collect the pose change amount Δpi output by the optical motion capture camera, the joint angles and the end pose of the robotic arm, the molten pool image Ii, and the process parameters Si (welding current, welding voltage, and the moving speed of the welding torch).
[0015] In this embodiment, the master device and the slave device are physically isolated by a black light-shielding board. The height of the black light-shielding board is 2 m, and the light-shielding rate > 99%, so as to ensure that the field of view of the optical motion capture camera of the master device is not interfered by the welding arc light, and at the same time protect the operator from arc light damage.
[0016] In this embodiment, after constructing the master-slave welding control system, in order to ensure the control accuracy, it is necessary to clearly define and calibrate the relevant coordinate systems. It is necessary to obtain the calibration relationship between the motion capture camera coordinate system and the world coordinate system, the calibration relationship between the positioning tool coordinate system and the motion capture camera coordinate system, the calibration relationship between the robotic arm base coordinate system and the world coordinate system, the calibration relationship between the welding torch tool coordinate system and the robotic arm end flange coordinate system, and the calibration relationship between the molten pool camera coordinate system and the robotic arm end flange coordinate system. Among them, the world coordinate system is the system reference benchmark, the motion capture camera coordinate system is the coordinate system of the optical motion capture system, the positioning tool coordinate system is the coordinate system of the positioning tool fixed on the simulated welding torch, the robotic arm base coordinate system is the robotic arm base coordinate system, the robotic arm end flange coordinate system is the robotic arm end interface coordinate system, and the welding torch tool coordinate system is the coordinate system of the real welding torch end point (TCP).
[0017] In step S102, during the process of operating the simulation welding torch for welding, the working host controls the welding torch to run synchronously with the simulation welding torch, that is, the master-slave teleoperation mode is realized; and the multi-modal data stream during the welding process is obtained through the data acquisition system. The multi-modal data stream includes the pose change amount of the positioning tool collected by the optical motion capture camera in real time, the molten pool image collected by the molten pool camera in real time, and the process parameters uploaded by the industrial control computers of the welding machine of the welding torch and the robotic arm.
[0018] As an implementation method, during the data acquisition process of this solution, a welder with good welding skills holds the simulation welding torch, starts the data acquisition system, and the welder observes the molten pool image through the screen of the working host, manually adjusts the pose of the simulation welding torch (simulating the actual welding action), and the optical motion capture camera collects the 6DoF pose change amount Δpi (the increment relative to the initial pose) of the positioning tool at a frequency of 100Hz. The working host converts the collected Δpi into the joint angle command of the robotic arm through inverse kinematics solution, and controls the welding torch at the slave end to move synchronously with the simulation welding torch until the welding is completed. During this process, the molten pool camera collects the molten pool image at 30Hz (224×224 resolution); the welding machine connected to the welding torch and the controller of the robotic arm report the process parameter Si at 100Hz.
[0019] Step S103, based on the multi-modal data stream, obtain the training sample data for training the welding imitation learning model.
[0020] As an implementation method, a method for obtaining the training sample data for training the welding imitation learning model includes: performing timestamp alignment on the pose change amount, molten pool image, welding voltage, welding current, and moving speed in the multi-modal data stream to generate a multi-modal data packet with time series alignment; based on the multi-modal data packet with time series alignment, obtain the training sample data.
[0021] As an example, the timestamp alignment of Δpi, Ii, and Si can be performed through the time synchronization module of the Robot Operating System (ROS) to generate a 30Hz multi-modal data packet with time series alignment (Ii, Δpi, Si), and store it in a preset path; then, the data in the multi-modal data packet is divided into training sample data and model verification data.
[0022] Step S104, based on the training sample data, train and construct the welding imitation learning model to obtain the trained welding imitation learning model.
[0023] The welding imitation learning model constructed in this embodiment includes an input layer, a feature extraction layer, a feature fusion layer, a time series modeling layer, and an output layer: among them, The input layer is used to obtain a sequence of molten pool images and the corresponding process parameter sequence from the training sample data by using a sliding window. The sequence of molten pool images includes n consecutive frames of molten pool images; The feature extraction layer includes a visual feature extraction module and a process parameter embedding module. The visual feature extraction module is used to extract the spatial features of n frames of molten pool images respectively through multiple convolutional neural networks and perform temporal splicing to obtain a temporal visual feature vector. The process parameter embedding module is used to form an n×3-dimensional temporal sequence from the process parameter sequence, and then map the n×3-dimensional temporal sequence to an n×64-dimensional space through a linear transformation to obtain a parameter feature vector; The feature fusion layer is used to reduce the dimension of the temporal visual feature vector through a linear transformation to obtain a reduced-dimensional temporal visual feature vector. Then, the reduced-dimensional temporal visual feature vector is used as the Key and Value, and the parameter feature vector is used as the Query to perform feature fusion through cross-modal attention to obtain a fused feature vector; The temporal modeling layer is used to capture the long-range temporal dependence in the fused feature vector through a multi-layer Transformer Encoder network structure to obtain a long-range feature vector. The multi-layer Transformer Encoder network structure includes a multi-dimensional hidden layer, multi-head attention, and a feed-forward network; The output layer is used to perform pooling processing on the long-range feature vector, and then map the pooled long-range feature vector to a predicted control instruction through a fully connected layer. The predicted control instruction includes a pose change amount, a current adjustment amount, a voltage adjustment amount, and a speed adjustment amount.
[0024] The model training objective of the welding imitation learning model in this embodiment is: input 5 consecutive frames of molten pool images (temporal window) and the corresponding process parameters, and output the 6DoF pose change amount Δpi and the process parameter adjustment amount ΔS at the current moment, so as to realize the end-to-end mapping of "molten pool state → torch movement".
[0025] As an example, the model input includes a sequence of molten pool images and corresponding process parameter sequences at the current moment and several (e.g., a total of 5 frames) past time steps. The visual feature extraction module inputs the 5-frame molten pool images (224×224×3×5) into a convolutional neural network (such as ResNet-34) respectively to extract spatial features. Each convolutional neural network outputs a 512-dimensional vector, and after temporal concatenation, a 5×512-dimensional temporal visual feature vector is obtained to capture the dynamic changes of the molten pool. The process parameter sequences at 5 moments form a 5×3-dimensional temporal sequence. The process parameter embedding module maps the 5×3-dimensional temporal sequence to a 5×64-dimensional space through a linear transformation (fully connected layer, weight matrix 3×64) to obtain a parameter feature vector (5×64) to retain the temporal correlation. The feature fusion layer adopts a cross-modal attention mechanism, taking the parameter feature vector as Query (abbreviated as Q), the temporal visual feature vector as Key (abbreviated as K) and Value (abbreviated as V). This setting allows the model to dynamically focus on the most relevant regions or features in the molten pool image according to the current process parameters, thereby enhancing the understanding of the correlation between the molten pool dynamics and the torch movement; among them, Key / Value is the temporal visual feature vector (5×512→5×64, dimension reduction through linear transformation); the dot product attention of the attention mechanism ( , where d k is the dimension of Key) outputs a 5×64-dimensional fusion feature to strengthen the dynamic focus of the process parameters on the molten pool features. For example, when the current increases, focus on the change in the weld width. The temporal modeling layer inputs the obtained fusion feature vector into a 6-layer Transformer Encoder with a hidden layer dimension of 64, multi-head attention (9 heads), and a feed-forward network (64→256→64) to capture long-range temporal dependencies. For example, the humps of the molten pool in the first 3 frames indicate that the torch angle needs to be adjusted currently. The output layer maps the output of the temporal modeling layer (e.g., 5x64 dimensions) to the predicted control instruction through a fully connected network (64→128→9), outputting the pose change amount and process parameter adjustment amount (3 dimensions: current adjustment amount, voltage adjustment amount, movement speed adjustment amount); the prediction target of the model is the pose change amount (Δppred) of the 6-DoF torch end at the next moment and the possible process parameter adjustment amount (ΔSpred). The output sequence of the Transformer can be aggregated using methods such as self-attention weighted average to generate the final single control instruction.
[0026] As an implementation, a method for training and constructing a welding imitation learning model based on training sample data to obtain a trained welding imitation learning model, including: performing data augmentation processing on the training sample data; using the molten pool image and process parameters in the training sample data as feature data, and using the pose change amount as label data, and training the welding imitation learning model with the loss function aiming to minimize the mean square error between the predicted pose change amount and the actual value of the pose change amount to obtain a trained welding imitation learning model.
[0027] Among them, the method for performing data augmentation processing on the training sample data includes: adding Gaussian noise and random brightness offset to the molten pool image; adding random perturbations to the process parameters. For example, adding Gaussian noise (σ = 0.01) and random brightness offset (±10%) to the molten pool image, and adding ±5% random perturbations to the process parameters (simulating the fluctuations in actual working conditions).
[0028] As an example, during model training, the training objective can be to minimize the difference between the predicted Δppred and the collected Δpi, and the prediction of process parameters is also processed similarly. For example, the loss function is as follows: L = α・Lpose + β・Lparam Among them, Lpose is the mean square error of the pose change amount (weight α = 0.8), and Lparam is the mean square error of the process parameters (weight β = 0.2).
[0029] During training, the optimizer Adam (learning rate 1e-4, weight decay 1e-5) can be used; the training platform is NVIDIA A100 GPU, the batch size is 32, and the training cycle is 50 rounds (about 48 hours).
[0030] In one embodiment, after obtaining the trained welding imitation learning model, it includes: performing lightweight processing on the trained welding imitation learning model to obtain a lightweight welding imitation learning model; deploying the lightweight welding imitation learning model to the ROS control node of the working host.
[0031] As an example, the trained model is deployed to the ROS control node of the working host, and model lightweighting can be performed according to the actual situation: the model volume is compressed to <15MB (inference latency <10ms) through pruning (removing redundant channels) and quantization (FP32→INT8); when the model is applied, the molten pool images collected by the molten pool camera are transmitted to the GPU for model inference. The model outputs the predicted pose change amount Δppred of the welding torch and the process parameter adjustment amount Δspred according to the real-time input molten pool image sequence and process parameter sequence. Among them, the pose change amount Δppred is converted into the target pose in the base coordinate system of the robot system of the robotic arm, and then the corresponding joint angle command of the robotic arm is calculated through the inverse kinematics solver. The joint angle command is sent to the robotic arm controller at a fixed period of 30Hz through the robotic arm control interface (such as UR5 SDK) to achieve closed-loop motion control.
[0032] The training method of the welding imitation learning model according to the embodiments of the present application constructs a master-slave control system to implement the master-slave remote control mode, collects the fine operation data of human welders through a motion capture device, and combines the real-time visual feedback of the molten pool camera to train the welding imitation learning model to achieve precise closed-loop control of the pose of the welding torch at the end of the robotic arm, so as to solve the problems of high manual intervention, low environmental adaptability and difficult skill transfer existing in traditional welding robots. By constructing a master-slave control system to collect data during the master-slave remote operation, the welder does not need to perform complex programming or directly teach in a harsh environment, which significantly reduces the labor cost and operation difficulty and improves the data collection efficiency. Combining the high-precision pose measurement of optical motion capture and the process information of the molten pool camera, the multi-modal data fusion provides a richer and more reliable data basis for model training. This solution collects real-time molten pool images through the molten pool camera, enabling the welder to perceive the welding state online and dynamically adjust the pose of the welding torch and process parameters, so as to better adapt to the geometric differences of workpieces, assembly errors and disturbances during the welding process, effectively improve the welding quality and stability, reduce the defect rate, and enhance the environmental and task adaptability. This solution can directly learn the fine control strategy and subconscious compensation actions of human experts in actual operations through imitation learning, breaking through the limitations of traditional methods in the fidelity of skill transfer, and realizing the digital and automatic inheritance of expert welding skills, thereby improving the reliability of the collected data. This solution effectively avoids the influence of welding strong light, high temperature and electromagnetic interference on the motion capture camera and operators through the master-slave physical isolation design. Based on any of the above embodiments, the embodiments of the present application further provide an intelligent welding method based on a welding imitation learning model, as Figure 2 shown. The intelligent welding method based on the welding imitation learning model includes the following steps: Step S201, obtain the image sequence and the corresponding process parameter sequence collected in real time by the molten pool camera.
[0033] Step S202: Input the image sequence and the corresponding process parameter sequence into the welding imitation learning model to obtain the predicted pose change amount and process parameter adjustment amount.
[0034] Step S203: Convert the predicted pose change amount into the target pose in the robot base coordinate system, and then calculate the corresponding robot joint angle control command through the inverse kinematics solver.
[0035] Step S204: Send the robot joint angle control command to the robot controller of the robotic arm at a fixed period through the robot control interface to achieve intelligent welding.
[0036] The intelligent welding method based on the welding imitation learning model in the embodiment of the present application is based on the fine operation data of human welders collected by the motion capture device, combined with the real-time visual feedback of the molten pool camera, and the trained welding imitation learning model to achieve precise closed-loop control of the pose of the welding torch at the end of the robotic arm, realizing intelligent welding with strong adaptability and less manual intervention.
[0037] In the description of the foregoing embodiments, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0038] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0039] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A training method for a welding imitation learning model, characterized in that, The following steps are involved: A welding master-slave control system is constructed, the welding master-slave control system includes a master-end device, a slave-end device and a working host, the master-end device includes a simulation welding gun and an optical motion capture camera, the positioning tool of the optical motion capture camera is fixed on the simulation welding gun; the slave-end device includes a mechanical arm, a welding gun and a molten pool camera, the end of the mechanical arm is fixedly connected to the welding gun, and the molten pool camera is installed on the side of the welding gun; the working host is connected to the optical motion capture camera, the molten pool camera and the industrial computer of the mechanical arm, and the working host is equipped with a data acquisition system; During the operation of the simulated welding gun for welding, the welding gun is controlled by the working host to run synchronously with the simulated welding gun, and the multimodal data stream in the welding process is obtained by the data acquisition system, wherein the multimodal data stream includes the posture change of the positioning tool acquired in real time by the optical motion capture camera, the molten pool image acquired in real time by the molten pool camera, and the process parameters uploaded by the welding machine of the welding gun and the industrial computer of the robotic arm; Based on the multimodal data stream, obtaining training sample data for training a welding imitation learning model; The welding imitation learning model constructed based on the training sample data is trained to obtain a trained welding imitation learning model.
2. The method according to claim 1, characterized in that The process parameters include welding voltage, welding current and moving speed of the welding gun.
3. The method according to claim 2, wherein The method of acquiring training sample data for training a welding imitation learning model based on the multimodal data stream comprises: Performing timestamp alignment on the posture change, the molten pool image, the welding voltage, the welding current and the moving speed in the multimodal data stream to generate a time-aligned multimodal data packet; Based on the time-aligned multimodal data packets, training sample data is acquired.
4. The method according to claim 3, characterized in that, The welding imitation learning model includes: An input layer, used for acquiring a molten pool image sequence and a corresponding process parameter sequence from the training sample data by using a sliding window, wherein the molten pool image sequence includes n consecutive frames of molten pool images; A feature extraction layer, the feature extraction layer includes a visual feature extraction module and a process parameter embedding module, the visual feature extraction module is used to extract the spatial features of the n frames of molten pool images respectively through multiple convolutional neural networks and perform time-series splicing to obtain a time-series visual feature vector; the process parameter embedding module is used to form the process parameter sequence into an n×3-dimensional time-series sequence, and then map the n×3-dimensional time-series sequence to an n×64-dimensional space through linear transformation to obtain a parameter feature vector; A feature fusion layer is used to reduce the dimension of the temporal visual feature vector by linear transformation to obtain a temporal visual feature vector after dimension reduction; then the temporal visual feature vector after dimension reduction is used as Key and Value, and the parameter feature vector is used as Query, and feature fusion is performed through cross-modal attention to obtain a fused feature vector; The temporal modeling layer is used to capture the long-range temporal dependencies in the fused feature vectors through a multi-layer Transformer Encoder network structure, obtaining long-range feature vectors; the multi-layer Transformer Encoder network structure includes a multi-dimensional hidden layer, multi-head attention, and a feed-forward network; The output layer is used to perform pooling processing on the long-range feature vectors, and then map the pooled long-range feature vectors to the predicted control instructions through a fully connected layer; the predicted control instructions include pose change amounts, current adjustment amounts, voltage adjustment amounts, and speed adjustment amounts.
5. The method according to claim 4, wherein The welding imitation learning model trained and constructed based on the training sample data to obtain a trained welding imitation learning model; includes: Perform data augmentation processing on the training sample data; Use the molten pool images and process parameters in the training sample data as feature data, and the pose change amounts as label data. The loss function trains the welding imitation learning model with the goal of minimizing the mean square error between the predicted pose change amounts and the actual values of the pose change amounts, obtaining a trained welding imitation learning model.
6. The method according to claim 5, wherein Performing data augmentation processing on the training sample data includes: Adding Gaussian noise and random brightness offsets to the molten pool images; Adding random perturbations to the process parameters.
7. The method according to claim 1, characterized in that, After obtaining the trained welding imitation learning model, includes: Perform lightweight processing on the trained welding imitation learning model to obtain a lightweight welding imitation learning model; Deploy the lightweight welding imitation learning model to the ROS control node of the working host.
8. The method according to claim 1, characterized in that, The robotic arm is fixedly connected to the welding torch through the end flange, and the method further includes: Obtain the calibration relationships between the motion capture camera coordinate system and the world coordinate system, the positioning tool coordinate system and the motion capture camera coordinate system, the robotic arm base coordinate system and the world coordinate system, the welding torch tool coordinate system and the robotic arm end flange coordinate system, and the molten pool camera coordinate system and the robotic arm end flange coordinate system.
9. The method according to claim 1, characterized in that, The master device and the slave device are physically isolated by a black light-shielding plate.
10. An intelligent welding method based on a welding imitation learning model, characterized in that, The welding imitation learning model is obtained by the training method of the welding imitation learning model according to any one of claims 1 to 9; the welding method includes: Obtain the image sequence and the corresponding process parameter sequence collected in real time by the molten pool camera; Input the image sequence and the corresponding process parameter sequence into the welding imitation learning model to obtain the predicted pose change amounts and process parameter adjustment amounts; Convert the predicted pose change amounts to the target poses in the robot base coordinate system, and then calculate the corresponding robot joint angle control instructions through an inverse kinematics solver; Send the robot joint angle control instructions to the robot controller of the robotic arm at a fixed period through the robot control interface to achieve intelligent welding.
Citation Information
Patent Citations
Prediction method for perforated plasma-arc weld pool based on deep learning algorithm
CN112894101A
Intelligent welding method based on causal learning and robot system
CN116618915A
Welding robot with parameter prediction function
CN116900582A
Method for robot AI training manual operation
CN118097790A
Mechanical arm control system calibration method based on deep reinforcement learning, storage device and electronic equipment
CN118123838A