Distributed reinforcement learning control method and system for multi-robot cooperative welding
Through the distributed reinforcement learning control method, combined with multi-agent reinforcement learning and global coordination modules, the problem of insufficient coupling between task allocation and process parameters in heterogeneous multi-robot collaborative welding is solved, and efficient and stable weld formation is achieved.
Patent Information
- Application Number
- CN202511082327.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-09-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the existing technology of heterogeneous multi-robot collaborative welding, the task allocation algorithm is difficult to cope with faults and environmental interference in the welding process, resulting in idle resources and insufficient collaborative accuracy. In addition, the multi-agent reinforcement learning model does not fully couple the welding process parameters, affecting the weld formation accuracy.
A distributed reinforcement learning control method is adopted to construct a heterogeneous multi-robot collaborative welding state space. By combining a multi-agent reinforcement learning model and a global coordination module, task allocation and robot decision-making are dynamically adjusted, welding state parameters are collected in real time, and the welding process is optimized.
It achieves rapid response to faults and parameter mutations during the welding process, avoids idle resources, improves robot collaboration accuracy and weld formation quality, and enhances the adaptability and stability of welding.
Smart Images

Figure CN120645228A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multi-robot collaborative welding, and in particular to a distributed reinforcement learning control method and system for multi-robot collaborative welding. Background Art
[0002] In modern industrial manufacturing, welding operations for large and complex components place stringent demands on efficiency, precision, and stability. A single robot, limited by its operating range, load capacity, and welding parameter adjustment range, struggles to meet the one-time formation requirements for complex welds. Heterogeneous multi-robot collaborative welding, leveraging the differentiated advantages of different robots in welding current adaptability, welding gun posture adjustment capabilities, and movement speed, has become an important technical approach for addressing these issues. With the development of distributed computing and intelligent decision-making technologies, autonomous collaborative decision-making among robots through multi-agent reinforcement learning models, combined with dynamic task allocation algorithms to optimize welding processes, has gradually become a key path to improving the intelligence level of collaborative welding and driving welding production towards flexibility and efficiency.
[0003] Existing technologies have two significant deficiencies in heterogeneous multi-robot collaborative welding control. On the one hand, traditional task allocation algorithms are mostly based on preset rules or static optimization models. When a robot fails, weld parameters suddenly change, or the external environment interferes during the welding process, it is difficult to quickly reconstruct the task allocation plan, resulting in subtask execution conflicts or idle resources, affecting the overall welding progress. On the other hand, the decision-making process of the multi-agent reinforcement learning model lacks deep coupling with the welding process parameters. The action instructions output by the local decision-making module do not fully consider the impact of changes in the molten pool temperature and the dynamic characteristics of weld formation on the operation of neighboring robots. The global coordination mechanism only relies on position information for simple corrections, resulting in insufficient accuracy in the coordinated actions between robots and unable to meet the continuous formation requirements of high-precision welds. Summary of the Invention
[0004] In order to overcome the shortcomings and deficiencies of the prior art, the present invention provides a distributed reinforcement learning control method and system for multi-robot collaborative welding.
[0005] The technical solution adopted by the present invention is a distributed reinforcement learning control method for multi-robot collaborative welding, comprising the following steps:
[0006] Step S1: Constructing a heterogeneous multi-robot collaborative welding state space including welding current, welding voltage, welding speed, welding gun posture, weld position, and welding environment obstacle distribution, and determining the initial feasible operating area of each robot in the state space based on the welding operating range, motion constraints, and welding capability parameters of each robot;
[0007] Step S2: Based on the state space and the initial feasible operation area of each robot, a heterogeneous multi-robot task allocation algorithm is used to decompose the overall welding task into several subtasks. Each subtask includes a preset weld segment welding requirement, a welding accuracy threshold, and a task execution time window, and at least one candidate robot is assigned to each subtask;
[0008] Step S3: Each robot is equipped with a local decision-making module of a multi-agent reinforcement learning model. The local decision-making module takes the robot's current welding state parameters, the welding state parameters of neighboring robots, and the subtask requirement parameters as input, and outputs a local action decision including the welding gun movement trajectory adjustment amount and the welding parameter correction value;
[0009] Step S4: Setting up a global coordination module of the multi-agent reinforcement learning model. This module receives the local action decisions of all robots and the completion progress parameters of the current overall welding task. Through the dynamic adjustment mechanism of the heterogeneous multi-robot task allocation algorithm, it generates a global coordination factor for correcting the local decisions of each robot.
[0010] Step S5: The local decision module of each robot updates the output local action decision according to the global coordination factor, and converts the updated local action decision into a specific robot drive instruction to drive the robot to perform the welding operation, while simultaneously collecting the molten pool temperature and weld formation size parameters in real time during the welding process;
[0011] Step S6: The collected molten pool temperature and weld formation size parameters are fed back to the multi-agent reinforcement learning model. The model adjusts the state space representation method and the decision weight of the local decision module according to the feedback parameters, and enters the decision cycle for the next round of welding operation.
[0012] Furthermore, in step S2, the heterogeneous multi-robot task allocation algorithm matches subtasks with candidate robots by constructing a task allocation optimization model, and the optimization model is:
[0013]
[0014] Among them, T ij represents the comprehensive fitness of the i-th subtask assigned to the j-th robot, α, β, γ are weight coefficients and α+β+γ=1, W i is the welding workload of the ith subtask, C j is the welding capability coefficient of the jth robot, D ij is the distance from the jth robot to the starting position of the ith subtask, V j is the moving speed of the jth robot, P j is the welding accuracy parameter of the jth robot, P maxIt has the largest welding accuracy parameter among all robots;
[0015] In step S3, the action value function of the local decision module is:
[0016]
[0017] Among them, Q j (s, a) is the action value of the jth robot performing action a in the global state s, ω1 and ω2 are weight parameters and ω1+ω2=1, For the jth robot based on its own state s j and its own action a j The local action value, N j is the set of neighboring robots of the j-th robot, For the kth neighboring robot based on its own state s k and its own action a k The local action value.
[0018] Furthermore, in step S3, the decision-making process of the local decision module includes: based on the current welding current I of the robot j , welding voltage U j , welding gun posture angle θ j Construct its own state vector s j ;
[0019] Receive the state vector s of other robots within the radius R k And the corresponding subtask requirements in the weld thickness h i , welding speed threshold v i,min 、v i,max ;
[0020] The state vector s j , the set of neighboring robot state vectors {s k} and subtask requirement parameters {h i , v i,min , v i,max The input is sent to the local decision network, which consists of three fully connected layers. The number of input neurons in the first layer is the sum of the state vector dimension and the number of neighboring robots. The number of neurons in the second layer is 64. The number of output neurons in the third layer is the dimension of the action space.
[0021] The local decision network outputs the adjustment amount Δx of the welding gun movement trajectory in the x, y, and z axis directions j , Δy j , Δz j and welding current correction value ΔI j , welding voltage correction value ΔU j , forming a local action decision aj .
[0022] Furthermore, in step S4, the process of generating the global coordination factor by the global coordination module includes:
[0023] Collect the local action decisions of all robots a j , extract the welding gun moving speed v in each decision j , welding current I j and the corresponding subtask number;
[0024] Based on the welding progress P of each subtask i and the overlapping area O where the robot performs subtasks jk , calculate the global task coordination index Among them, n is the total number of subtasks, S i is the total area of the welding region of the i-th subtask;
[0025] Construct the global coordination factor matrix Λ=[λ jk ], where λ jk =exp(-μ·D jk )·(1+v·G), μ and v are adjustment parameters, D jk is the real-time distance between the jth and kth robots;
[0026] The elements in the global coordination factor matrix Λ are used as correction coefficients for the local decisions of each robot, where λ jj Used to correct the robot's own decision weight, λ jk Used to modify the influence weight of neighboring robots on decision making.
[0027] Furthermore, in step S5, the robot uses a dynamic welding parameter adjustment model when performing a welding operation according to the updated local action decision:
[0028]
[0029] Among them, I j (t) is the welding current of the jth robot at time t, is the initial welding current, ΔI j is the welding current correction value output in step S3, φ is the current adjustment frequency coefficient, t is the welding time, ζ is the temperature correction coefficient, T j (t) is the molten pool temperature collected at time t, and T0 is the target molten pool temperature;
[0030] At the same time, the real-time adjustment of the welding gun posture angle satisfies:
[0031]
[0032] Among them, θ j (t) is the welding gun posture angle of the j-th robot at time t, is the initial welding gun posture angle, Δθ j is the welding gun posture adjustment amount in local action decision, and κ is the posture adjustment rate coefficient.
[0033] Furthermore, in step S6, the process of adjusting the decision weights of the multi-agent reinforcement learning model includes:
[0034] Melt pool temperature T based on feedback j (t) and weld size, H j (t) is the weld height, W j (t) is the weld width, and the weld quality evaluation value Q is calculated j =α T ·|T j (t)-T0|+α H ·|H j (t)-H0|+α W |W j (t)-W0|, where α T , α H , α W is the quality assessment weight, T0, H0, and W0 are the target molten pool temperature, target weld height, and target weld width, respectively;
[0035] The weld quality evaluation value Q j As the reward and punishment signals are input to the multi-agent reinforcement learning model, the model uses the following weight update formula:
[0036]
[0037] Among them, ω(t+1) and ω(t) are the decision weights of the j-th robot at time t+1 and t respectively, η is the learning rate, is the gradient of the weld quality assessment value to the weight, ξ is the neighborhood influence coefficient, Q k is the weld quality evaluation value of the kth neighboring robot.
[0038] Furthermore, configuring the local decision module of the multi-agent reinforcement learning model for each robot in step S3 includes the following sub-steps:
[0039] S3.1: Deploy an independent neural network computing unit for each robot. This unit consists of an input layer, a hidden layer, and an output layer. The number of neurons in the input layer is consistent with the dimension of the state parameters. The hidden layer uses the ReLU activation function, and the output layer uses the Tanh activation function to limit the range of action decisions.
[0040] S3.2: Embed the robot hardware constraint parameters into the neural network computing unit, including the maximum welding gun movement speed, welding current adjustment range, and welding gun posture angle rotation limit. By setting a threshold truncation mechanism in the output layer, the output local action decision is ensured to be within the hardware constraint range;
[0041] S3.3: Preload the state-action-feedback data pairs from historical welding tasks and use offline training to initialize the weight parameters of the neural network calculation unit. During the training process, the weld formation errors in the historical data are used as the correction basis for back propagation;
[0042] S3.4: Configure a real-time data acquisition interface for the local decision module. The interface is connected to the robot's sensor system, collects welding state parameters in a 10ms cycle and converts them into digital signals that can be recognized by the neural network.
[0043] Furthermore, setting up the global coordination module of the multi-agent reinforcement learning model in step S4 includes the following sub-steps:
[0044] S4.1: Build a global data fusion center, which is connected to the local decision modules of all robots via industrial Ethernet. It uses a timestamp synchronization mechanism to receive local action decisions and state parameters sent by each module. Data transmission uses the UDP protocol to reduce transmission delays.
[0045] S4.2: Deploy a task progress monitoring unit in the global data fusion center. This unit calculates the completed length percentage of each subtask in real time based on the welding position information reported by each robot and generates a task progress heat map. Different colors in the heat map represent different progress intervals.
[0046] S4.3: Configure the global coordination algorithm execution unit, which calls the core function of the heterogeneous multi-robot task allocation algorithm, takes the task progress heat map data, robot position distribution data, and subtask priority parameters as input, and calculates the coordination priority of each robot;
[0047] S4.4: Based on the coordination priority and the distance parameters between robots, a global coordination factor is generated and sent to the corresponding local decision module via broadcast. The sending frequency is consistent with the decision cycle of the local decision module.
[0048] Furthermore, in step S5, the local decision module of each robot updates the local action decision according to the global coordination factor and drives the robot to perform the welding operation, which includes the following sub-steps:
[0049] S5.1: After receiving the global coordination factor, the local decision module performs a matrix multiplication operation on the local action decision output by itself. During the operation, weighted corrections are made according to the welding current, voltage, and welding gun movement trajectory parameter categories. Different parameter categories correspond to different coordination factor components.
[0050] S5.2: Convert the weighted corrected action decision into the drive signal of each actuator of the robot. The welding gun movement trajectory adjustment is converted into the number of pulses of the servo motor, and the welding current and voltage correction values are converted into analog voltage signals of the power controller.
[0051] S5.3: The drive signal is processed by the robot's control system and sent to the corresponding actuator. The actuator changes its operating state according to the signal instruction.
[0052] S5.4: While performing the welding operation, the robot's feedback sensor group is activated. The temperature sensor is inserted 5 mm from the edge of the molten pool to collect temperature data. The vision sensor is installed 30 cm behind the welding gun to capture the weld formation image. The position sensor uses an encoder to record the three-dimensional coordinates of the welding gun in real time.
[0053] The distributed reinforcement learning control system for multi-robot collaborative welding includes the following units:
[0054] Heterogeneous robot state parameter acquisition and preprocessing unit, which is connected to the sensor array of each robot via a wired cable and is used to receive raw sensor signals and convert them into standardized digital signals;
[0055] The multi-source data fusion and state space construction unit has its input connected to the output of the heterogeneous robot state parameter acquisition and preprocessing unit. It generates state space data including all robot and environment information through data splicing and dimension expansion operations.
[0056] The distributed task decomposition and candidate allocation unit has its input connected to the output of the multi-source data fusion and state space construction unit. It runs a heterogeneous multi-robot task allocation algorithm internally and outputs a subtask list and candidate robot matching relationships.
[0057] A multi-agent local decision-making neural network cluster, in which each neural network node corresponds to a robot. The node input is connected to the output of the distributed task decomposition and candidate allocation unit and the output of the heterogeneous robot state parameter acquisition and preprocessing unit, and the output generates a local action decision.
[0058] The global coordination factor calculation and distribution unit has its input connected to the output of the multi-agent local decision-making neural network cluster. After generating the global coordination factor through calculation, it will be sent to the corresponding neural network nodes respectively.
[0059] A robot drive and welding execution unit cluster, where each execution unit corresponds to a robot. The input end is connected to the output end of the multi-agent local decision-making neural network cluster, and the output end is connected to the robot's drive motor and welding power supply to drive the robot to complete the welding operation.
[0060] Beneficial effects: The present invention proposes a distributed reinforcement learning control method and system for multi-robot collaborative welding. By combining a heterogeneous multi-robot task allocation algorithm with a multi-agent reinforcement learning model, real-time welding state parameters are incorporated into the task decomposition and allocation process, and the matching relationship between subtasks and robots is dynamically adjusted. When a robot failure or a sudden change in weld parameters occurs, the global coordination module can quickly generate a new allocation plan to avoid subtask conflicts and idle resources. To address the problem of insufficient coupling between multi-agent decision-making and welding process parameters, the local decision-making module directly uses parameters such as molten pool temperature and weld formation size as input. The generation of the global coordination factor comprehensively considers factors such as welding area overlap and task progress, so that the robot action decision is deeply associated with the welding process characteristics, thereby improving the collaborative accuracy between robots and meeting the requirements of high-precision weld continuous formation. At the same time, the decision-making model is continuously optimized through a real-time feedback mechanism, further enhancing the adaptability and stability of collaborative welding. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 is a flow chart of the method steps of the present invention;
[0062] Figure 2 It is a diagram of the system unit composition of the present invention. DETAILED DESCRIPTION
[0063] It should be noted that, unless there is a conflict, the embodiments in this application and the features described in the embodiments can be combined with each other. The application is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0064] like Figure 1 As shown in FIG, the distributed reinforcement learning control method for multi-robot collaborative welding includes the following steps:
[0065] Step S1: Constructing a heterogeneous multi-robot collaborative welding state space including welding current, welding voltage, welding speed, welding gun posture, weld position, and welding environment obstacle distribution, and determining the initial feasible operating area of each robot in the state space based on the welding operating range, motion constraints, and welding capability parameters of each robot;
[0066] Specifically, constructing the state space for heterogeneous multi-robot collaborative welding forms the foundation for all subsequent decision-making and control. This state space, encompassing parameters such as welding current, welding voltage, welding speed, welding gun posture, weld seam location, and obstacle distribution in the welding environment, directly reflects key characteristics and environmental constraints of the welding process. Welding current typically fluctuates between 50 and 500 A, and welding voltage varies between 10 and 40 V. These parameters determine the arc energy density, which in turn influences molten pool formation and weld quality. Welding speed is generally controlled between 5 and 50 mm / s. Gun posture includes factors such as the angle and height between the welding gun and the workpiece, and weld seam location is accurate to the millimeter level. Together, these parameters form the core basis for robot operation. Determining each robot's initial feasible operating area requires considering the robot's welding operating range (e.g., some robots may have a welding operating radius of 1 to 3 meters), motion constraints such as a maximum travel speed of 0.5 to 2 m / s, and welding capability parameters such as the maximum welding current it can withstand. This process ensures that the robot's initial operation does not exceed its capabilities or environmental limitations, providing a reliable spatial boundary for subsequent task allocation and decision-making.
[0067] During implementation, a sensor network is used to collect relevant data about the welding environment. This includes using a current sensor to collect welding current, a voltage sensor to collect welding voltage, an encoder to collect welding speed, a posture sensor to collect the pitch, roll, and yaw angles of the welding gun, a laser scanner to obtain the three-dimensional coordinates of the weld seam location, and an ultrasonic sensor to detect obstacles in the environment. The welding current is collected within a range of 50-500A at a frequency of 1kHz; the welding voltage is collected within a range of 10-40V at a frequency of 1kHz. The welding speed is collected with an accuracy of ±0.1mm / s, the welding gun posture angle is measured with an accuracy of ±0.5 degrees, and the weld seam location is measured with an error of no more than ±0.5mm. The obstacle detection range is 0.5-5 meters with an accuracy of ±5mm. Based on this collected data, a state space is constructed with each parameter as a dimension. Next, based on the parameters of each robot, for example, a certain robot model has a welding radius of 2 meters, a maximum movement speed of 1.5 m / s, a maximum welding current of 400 A, and a minimum welding current of 60 A, its initial feasible operating area in state space is determined. For weld position parameters, the robot's initial feasible operating area must include all weld segments within its operating radius; for welding current parameters, the feasible range is limited to between 60 and 400 A. This process is repeated until the initial feasible operating areas of all robots are determined.
[0068] Step S2: Based on the state space and the initial feasible operation area of each robot, a heterogeneous multi-robot task allocation algorithm is used to decompose the overall welding task into several subtasks. Each subtask includes a preset weld segment welding requirement, a welding accuracy threshold, and a task execution time window, and at least one candidate robot is assigned to each subtask;
[0069] Specifically, task decomposition and allocation based on the state space and each robot's initial feasible operating area is a key step in achieving efficient multi-robot collaboration. The decomposition of the overall welding task must take into account the continuity and complexity of the weld. After dividing it into several subtasks, each subtask has clear welding requirements for the weld segment. For example, the weld thickness may be required to be between 2 and 10 mm, the welding accuracy threshold may be a position error of no more than ±0.3 mm, and the task execution time window, such as a subtask must be completed within 10 to 30 minutes. A heterogeneous multi-robot task allocation algorithm is used to allocate tasks, fully leveraging the strengths of different robots. For example, some robots excel at high-precision welding and can be assigned to subtasks requiring high welding accuracy, while others have high mobility and can be assigned to subtasks requiring rapid transfer. Assigning at least one candidate robot to each subtask increases the flexibility and reliability of task execution. If a candidate robot is unable to perform a task, it can be replaced by another candidate robot to ensure the smooth progress of the welding task.
[0070] During implementation, a comprehensive analysis of the overall welding task is first conducted to determine the total length, distribution, and characteristics of each weld segment. For example, a 5-meter weld is broken down into eight subtasks based on its direction and thickness variation, with each subtask having a weld segment length between 0.5 and 0.8 meters. Specific parameters are set for each subtask, such as a 3mm weld thickness requirement for the first subtask, a position error of ±0.2mm welding accuracy threshold, and a 15-25 minute task execution window; a 5mm weld thickness requirement for the second subtask, a position error of ±0.3mm welding accuracy threshold, and a 20-35 minute task execution window. Then, a heterogeneous multi-robot task allocation algorithm is invoked, which first calculates each robot's suitability for each subtask. This suitability is calculated based on factors such as the robot's welding accuracy parameters, movement speed, and welding thickness range. For example, Robot A has a welding accuracy of ±0.1mm, a welding thickness range of 2-8mm, and a moving speed of 1.2m / s; Robot B has a welding accuracy of ±0.2mm, a welding thickness range of 3-10mm, and a moving speed of 0.8m / s. For the first subtask, the algorithm calculates that Robot A has a higher suitability than Robot B, and therefore selects Robot A as the primary candidate, while also listing Robot B as a candidate. Similarly, two or three candidate robots are assigned to each subtask, forming a list of correspondences between subtasks and candidate robots, completing the task assignment process.
[0071] Step S3: Each robot is equipped with a local decision-making module of a multi-agent reinforcement learning model. The local decision-making module takes the robot's current welding state parameters, the welding state parameters of neighboring robots, and the subtask requirement parameters as input, and outputs a local action decision including the welding gun movement trajectory adjustment amount and the welding parameter correction value;
[0072] Specifically, configuring each robot with a local decision-making module based on a multi-agent reinforcement learning model is the core of the robot's autonomous decision-making. The input to this local decision-making module includes various information. The robot's current welding state parameters, such as real-time welding current, voltage, and torch posture, reflect its current working status. The welding state parameters of neighboring robots enable the robot to understand the conditions of its companions, avoiding conflicts and duplication of work. Subtask requirement parameters provide goals and constraints for the robot's decision-making. The welding torch trajectory adjustment in the output local action decision ensures precise alignment of the welding torch with the weld seam, and welding parameter corrections optimize welding current, voltage, and other parameters based on actual conditions to ensure welding quality. This local decision-making model enables each robot to quickly respond to changes in the environment and tasks, improving the timeliness and flexibility of decision-making while also providing a foundation for global coordination.
[0073] During implementation, each robot was equipped with a separate local decision-making module hardware, utilizing a high-performance processor to ensure rapid data processing and decision-making. The module's input interface connects to the robot's various sensors, receiving real-time data on its current welding state parameters. These include welding current data within a range of 50-500A with an accuracy of ±1A; welding voltage data within a range of 10-40V with an accuracy of ±0.1V; and welding gun pitch and roll angles within a range of -30 to 30 degrees with an accuracy of ±0.1 degrees. Simultaneously, wireless communication modules receive these state parameters from neighboring robots within a 5-meter radius, with communication latency kept to less than 10ms. Subtask-required parameters, such as weld thickness, welding accuracy threshold, and time window, are transmitted to the local decision-making module via the task allocation system, with a data update frequency of 10Hz. The multi-agent reinforcement learning model within the local decision-making module is pre-trained, with the number of input layer nodes determined by the dimensionality of the input parameters: five state parameters for the robot itself, five state parameters for each of the two neighboring robots, and three required parameters for the subtask, for a total of 18 input nodes. The hidden layer adopts a multi-layer neural network structure. After training, the model can output reasonable local action decisions based on the input data. The adjustment range of the welding gun movement trajectory in the x, y, and z axis directions is -5mm to 5mm, with an accuracy of ±0.01mm; the welding current correction value range is -20A to 20A, and the welding voltage correction value range is -2V to 2V, ensuring that the decision is within a reasonable range.
[0074] Step S4: Setting up a global coordination module of the multi-agent reinforcement learning model. This module receives the local action decisions of all robots and the completion progress parameters of the current overall welding task. Through the dynamic adjustment mechanism of the heterogeneous multi-robot task allocation algorithm, it generates a global coordination factor for correcting the local decisions of each robot.
[0075] Specifically, establishing a global coordination module within the multi-agent reinforcement learning model is key to achieving globally optimal multi-robot collaboration. This module receives the local action decisions of all robots and can understand each robot's specific operational plan. The completion progress parameters of the current overall welding task, such as the completion percentage of each subtask, reflect the overall progress of the task. A global coordination factor is generated through the dynamic adjustment mechanism of the heterogeneous multi-robot task allocation algorithm. This factor modifies each robot's local decisions, aligning them with the global interest and avoiding situations where local optimality leads to global suboptimality. The introduction of this global coordination factor balances the actions of each robot, reduces conflicts and resource waste, improves the efficiency and quality of the overall welding task, and ensures more orderly and efficient collaboration among the robots.
[0076] In specific implementation, the global coordination module utilizes a distributed computing architecture, connecting to each robot's local decision-making module via a high-speed communication network with a communication rate of at least 100 Mbps, ensuring real-time reception of all robots' local action decisions. These decisions include information such as each robot's welding torch trajectory adjustment plan and welding parameter corrections, with a data sampling frequency of 10 Hz. Simultaneously, the module connects to the task management system to obtain completion progress parameters for the current overall welding task, such as the proportion of each subtask's welded length to the total length, with a data update cycle of 1 minute. The dynamic adjustment mechanism of the heterogeneous multi-robot task allocation algorithm operates continuously within the module. This mechanism calculates whether there are conflicts among the robots' local decisions (e.g., intersecting welding torch trajectories) and whether the completion progress of subtasks is balanced. For example, if it detects that the welding torch trajectories of robots A and B will intersect in 30 seconds, and if the completion progress of the subtask assigned to robot C lags behind plan by 20%, the algorithm begins calculating the global coordination factor. The coordination factor ranges from 0.8 to 1.2. Robots that need to slow down or adjust their trajectory are assigned a coordination factor less than 1, while robots that need to speed up are assigned a coordination factor greater than 1. After calculation, the global coordination module transmits the corresponding global coordination factor to each robot's local decision module via the communication network, with a transmission delay of less than 50ms.
[0077] Step S5: The local decision module of each robot updates the output local action decision according to the global coordination factor, and converts the updated local action decision into a specific robot drive instruction to drive the robot to perform the welding operation, while simultaneously collecting the molten pool temperature and weld formation size parameters in real time during the welding process;
[0078] Specifically, each robot's local decision-making module updates its local action decisions based on the global coordination factor and converts them into drive instructions to execute welding operations. This is a key step in putting decisions into practice. The application of the global coordination factor ensures the consistency of local decisions with the global goal, and the updated local action decisions are more reasonable. The updated decisions are converted into robot drive instructions, realizing the transition from decision-making to action, driving the robot's mechanical structure and welding equipment to operate according to the instructions. Real-time collection of molten pool temperature and weld formation size parameters during the welding process provides feedback information for subsequent model adjustments. These parameters directly reflect the welding quality. Through feedback, the model can continuously optimize decisions, form a closed-loop control, and improve the stability and reliability of welding.
[0079] During implementation, upon receiving the global coordination factor, each robot's local decision module immediately initiates an update program. This program, using a pre-set algorithm, calculates the global coordination factor and the parameters in the original local action decision. For example, if the original welding gun movement speed is 5 mm / s and the global coordination factor is 0.9, the updated movement speed is 4.5 mm / s. If the original welding current is 200 A and the global coordination factor is 1.1, the updated welding current is 220 A. The updated local action decision includes detailed parameters, such as a welding gun movement speed of 4.5 mm / s in the x-axis, 3 mm / s in the y-axis, and 0.5 mm / s in the z-axis; a welding current of 220 A; and a welding voltage of 25 V. The module then converts these digital decisions into analog drive commands via a digital-to-analog converter. The command for the servo motor controlling the welding gun movement is a 4-20 mA current signal, corresponding to a motor speed of 0-3000 rpm. The command for controlling the welding power supply is a 0-10 V voltage signal, corresponding to a welding current of 50-500 A and a welding voltage of 10-40 V. Drive commands are transmitted via cables to the robot's actuators. The drive motor drives the welding torch along the programmed trajectory, and the welding power supply outputs the corresponding current and voltage to perform welding. Simultaneously, a temperature sensor installed on the robot, inserted 5mm above the molten pool, measures the temperature within a range of 800-2000°C, with an accuracy of ±5°C and a sampling frequency of 10Hz. A visual sensor, mounted 10cm to the side of the welding torch, captures images of the weld seam. Image analysis reveals weld seam dimensions, such as width (2-15mm) with an accuracy of ±0.1mm and height (1-8mm) with an accuracy of ±0.05mm. These parameters are transmitted in real time to a data storage module.
[0080] Step S6: The collected molten pool temperature and weld formation size parameters are fed back to the multi-agent reinforcement learning model. The model adjusts the state space representation method and the decision weight of the local decision module according to the feedback parameters, and enters the decision cycle for the next round of welding operation.
[0081] Specifically, feeding back the collected molten pool temperature and weld size parameters to the multi-agent reinforcement learning model is the core link in achieving continuous model optimization. These feedback parameters directly reflect the actual effect of the welding operation. Excessively high or low molten pool temperature will affect the strength and formation of the weld, and the deviation of the weld size directly reflects whether the welding accuracy meets the requirements. The model adjusts the state space representation method based on these feedback parameters, which can make the state space more accurately reflect the actual welding situation; adjusting the decision weights of the local decision module can optimize the rationality and accuracy of the decision. This continuous feedback and adjustment mechanism enables the multi-agent reinforcement learning model to continuously adapt to changes in the welding environment and tasks, improve the robot's decision-making ability and welding quality, and form a continuously optimizing closed-loop system.
[0082] In specific implementation, feedback parameters are transmitted via a high-speed data bus. Molten pool temperature and weld size parameters are read from the data storage module and sent to the multi-agent reinforcement learning model at a frequency of 10Hz. The model first verifies the validity of the received parameters, eliminating outliers (e.g., data with temperatures outside the range of 800-2000°C) to ensure the reliability of the input data. Regarding the adjustment of the state space representation, if the molten pool temperature is detected to frequently exceed the normal range (e.g., 1500-1800°C) over a period of time, the model increases the weight of the molten pool temperature in the state space, making the decision process more focused on the temperature parameter. If the deviation in weld size is primarily due to the width, the width parameter is refined in the state space, for example, from 2mm intervals between 2-15mm to 1mm intervals. Regarding the adjustment of the decision weights of the local decision modules, the model calculates the adjustment amount based on the deviation of the weld size from the target value. When the width deviation of a certain weld segment is positive (greater than the target value) and continues to increase, the model reduces the weight of the decision related to welding current and increases the weight of the decision related to welding speed. During the adjustment process, the weight changes are kept within ±5% each time to avoid drastic fluctuations in the decision-making process. The adjusted state space representation and decision weights are written into the parameter storage area of the local decision module and take effect immediately in the next welding operation decision cycle. The entire feedback and adjustment process has a delay of less than 1 second, ensuring that the model can respond promptly to changes in the welding process.
[0083] Preferably, in step S2, the heterogeneous multi-robot task allocation algorithm matches subtasks with candidate robots by constructing a task allocation optimization model, and the optimization model is:
[0084]
[0085] Among them, T ij represents the comprehensive fitness of the i-th subtask assigned to the j-th robot, α, β, γ are weight coefficients and α+β+γ=1, W i is the welding workload of the ith subtask, C j is the welding capability coefficient of the jth robot, D ij is the distance from the jth robot to the starting position of the ith subtask, V j is the moving speed of the jth robot, P j is the welding accuracy parameter of the jth robot, P max It has the largest welding accuracy parameter among all robots;
[0086] In step S3, the action value function of the local decision module is:
[0087]
[0088] Among them, Q j (s, a) is the action value of the jth robot performing action a in the global state s, ω1 and ω2 are weight parameters and ω1+ω2=1, For the jth robot based on its own state s j and its own action a j The local action value, N j is the set of neighboring robots of the j-th robot, For the kth neighboring robot based on its own state s k and its own action a k The local action value.
[0089] Specifically, the heterogeneous multi-robot task allocation algorithm in step S2 matches subtasks with candidate robots through a specific optimization model that comprehensively considers multiple parameters, including subtask workload, robot welding capability, distance, movement speed, and welding accuracy. The selection of these parameters is directly related to the actual welding scenario, and the setting of weight coefficients balances the influence of each parameter in the matching process, ensuring that the overall fitness truly reflects the suitability of the robot to perform the subtask. In the local decision module in step S3, the action value function combines the local action value of the robot itself with the local action value of neighboring robots. The setting of weight parameters reflects the proportion of the robot's own decision and the decision of neighboring robots in the overall action value. This allows the robot to consider both its own situation and the influence of its surrounding robots when making decisions, thus avoiding decision conflicts. During implementation, the initial values of each weight coefficient are first determined based on historical welding data, and then the weights of the optimization model are adjusted through multiple experiments to ensure that the sub-task allocation results are reasonable; for the action value function, during the model training phase, the weight parameters are trained based on a large amount of welding scene data so that the output action decisions can adapt to different collaborative welding environments. The welding quality is continuously monitored during the training process, and the parameters are corrected based on quality indicators to ensure the effectiveness of the decision.
[0090] Preferably, in step S3, the decision-making process of the local decision module includes: based on the current welding current I of the robot j , welding voltage U j , welding gun posture angle θ j Construct its own state vector s j ;
[0091] Receive the state vector s of other robots within the radius R k (k∈N j ) and the corresponding subtask requirements in the weld thickness h i , welding speed threshold v i,min 、v i,max ;
[0092] The state vector s j , the set of neighboring robot state vectors {s k} and subtask requirement parameters {h i , v i,min , v i,max The input is sent to the local decision network, which consists of three fully connected layers. The number of input neurons in the first layer is the sum of the state vector dimension and the number of neighboring robots. The number of neurons in the second layer is 64. The number of output neurons in the third layer is the dimension of the action space.
[0093] The local decision network outputs the adjustment amount Δx of the welding gun movement trajectory in the x, y, and z axis directions j , Δy j, Δz j and welding current correction value ΔI j , welding voltage correction value ΔU j , forming a local action decision a j .
[0094] Specifically, the decision-making process of the local decision module in step S3 is closely centered around the robot's own and neighboring robot's state parameters and subtask parameters. The construction of the robot's own state vector selects core parameters such as welding current, voltage, and welding gun posture angle. These parameters directly determine the robot's current welding state and form the basis for decision-making. Receiving the state parameters of neighboring robots and subtask requirements such as weld thickness and welding speed thresholds allows the robot to understand its surroundings and task objectives, providing a reference for decision-making. The structural design of the local decision network meets the requirements of parameter processing and decision output. The number of neurons in the input layer matches the state parameter dimensions and the number of neighboring robots to ensure comprehensive information reception and processing. The configuration of the hidden and output layers ensures the depth of the decision calculation and the accuracy of the output. During implementation, the frequency and accuracy of state parameter collection are first determined to ensure the timeliness and reliability of input data. A local decision network that meets the structural requirements is then constructed. The network is trained using a large amount of labeled state-action data. During training, network parameters are continuously adjusted to ensure that the output welding gun trajectory adjustment and welding parameter correction values accurately meet welding requirements. After training, actual welding tests are conducted, and the network is further optimized based on the test results to improve decision-making accuracy.
[0095] Preferably, in step S4, the process of generating the global coordination factor by the global coordination module includes:
[0096] Collect the local action decisions of all robots a j , extract the welding gun moving speed v in each decision j , welding current I j and the corresponding subtask number;
[0097] Based on the welding progress P of each subtask i (ratio of welded length to total length) and the overlapping area O of the robot performing subtasks jk (the overlapping area of the welding area of the jth and kth robots), calculate the global task coordination index Among them, n is the total number of subtasks, S i is the total area of the welding region of the i-th subtask;
[0098] Construct the global coordination factor matrix Λ=[λ jk ], where λ jk =exp(-μ·D jk )·(1+ν·G), μ and v are adjustment parameters, Djk is the real-time distance between the jth and kth robots;
[0099] The elements in the global coordination factor matrix Λ are used as correction coefficients for the local decisions of each robot, where λ jj Used to correct the robot's own decision weight, λ jk (j≠k) is used to correct the influence weight of neighboring robots on the decision.
[0100] Specifically, the global coordination module in step S4 generates the global coordination factor by fully integrating key parameters in the robot's local motion decisions and overall task progress parameters. By collecting the welding gun movement speed, welding current, and subtask number from each robot's local motion decision, each robot's operation plan and task assignment can be determined. When calculating the global task coordination index, the subtask welding progress and the overlapping area between the robots performing the subtasks are important considerations. Progress reflects task progress, while the overlapping area reflects the degree of operational conflict. The index value comprehensively assesses the coordination status of the global task. The construction of the global coordination factor matrix takes into account the distance between robots and the global task coordination index. The distance parameter influences the strength of interaction between robots, while the coordination index determines the degree to which the overall coordination factors are modified. During implementation, a reasonable collection cycle should be set to obtain local action decisions and task progress parameters to ensure the real-time nature of the data; when calculating the global task coordination index, the calculation standards of each parameter should be clarified, such as the measurement method of the overlapping area and the determination method of the total area of the welding area; for the global coordination factor matrix, appropriate initial values of the adjustment parameters should be set according to the actual size of the welding scene and the movement range of the robot, and the parameters should be adjusted through the coordination effect in actual operation, so that the generated coordination factor can effectively correct the local decision of each robot and improve the global collaboration efficiency.
[0101] Preferably, in step S5, the robot uses a dynamic welding parameter adjustment model when performing the welding operation according to the updated local action decision:
[0102]
[0103] Among them, I j (t) is the welding current of the jth robot at time t, is the initial welding current, ΔI j is the welding current correction value output in step S3, φ is the current adjustment frequency coefficient, t is the welding time, ζ is the temperature correction coefficient, T j (t) is the molten pool temperature collected at time t, and T0 is the target molten pool temperature;
[0104] At the same time, the real-time adjustment of the welding gun posture angle satisfies:
[0105]
[0106] Among them, θ j (t) is the welding gun posture angle of the j-th robot at time t, is the initial welding gun posture angle, Δθ j is the welding gun posture adjustment amount in local action decision, and κ is the posture adjustment rate coefficient.
[0107] Specifically, the dynamic welding parameter adjustment model involved in step S5 performs real-time adjustments to two key welding parameters: welding current and gun attitude angle. Welding current adjustment integrates the initial current, correction value, time factor, and the deviation between the weld pool temperature and the target temperature. The introduction of the time factor results in a dynamic change pattern in the current adjustment, while the correction of the temperature deviation ensures that the weld pool temperature remains within the appropriate range, ensuring weld quality. Gun attitude angle adjustment is based on the initial attitude angle, adjustment amount, time, and rate coefficient. The rate coefficient determines the speed of attitude angle adjustment, ensuring that the attitude angle reaches the target value smoothly and accurately to accommodate different welding positions and weld requirements. During implementation, first determine the initial welding current and welding gun posture angle according to the welding material and weld requirements, then set the initial values of the current adjustment frequency coefficient and the temperature correction coefficient, and adjust these coefficients by observing the temperature changes of the molten pool and the weld formation through experiments; for the welding gun posture angle, determine the rotation limit and rate coefficient range according to the mechanical properties of the welding gun. In actual welding, monitor the welding gun posture angle and welding effect in real time, and continuously optimize the rate coefficient to ensure that the posture angle adjustment can meet the welding accuracy requirements without causing excessive wear and tear on the mechanical structure.
[0108] Preferably, in step S6, the process of adjusting the decision weights of the multi-agent reinforcement learning model includes:
[0109] Melt pool temperature T based on feedback j (t) and weld size H j (t)(weld height), W j (t) (weld width), calculate the weld quality assessment value Q j =α T ·|T j (t)-T0|+α H ·|H j (t)-H0|·+α W |W j (t)-W0|, where α T , α H , α W is the quality assessment weight, T0, H0, and W0 are the target molten pool temperature, target weld height, and target weld width, respectively;
[0110] The weld quality evaluation value Qj As the reward and punishment signals are input to the multi-agent reinforcement learning model, the model uses the following weight update formula:
[0111]
[0112] Among them, ω j (t+1),ω j (t) are the decision weights of the jth robot at time t+1 and time t, η is the learning rate, is the gradient of the weld quality assessment value to the weight, ξ is the neighborhood influence coefficient, Q k is the weld quality evaluation value of the kth neighboring robot.
[0113] Specifically, in step S6, the multi-agent reinforcement learning model adjusts its decision weights based on the weld quality assessment value. This assessment value integrates the deviations of the molten pool temperature, weld height, and weld width from their respective target values. The weight setting reflects the importance of each parameter in assessing weld quality, ensuring that the assessment results can comprehensively and accurately reflect welding quality. After the assessment value is input into the model as a reward and punishment signal, the weight update formula considers the current weight, learning rate, gradient, neighborhood influence coefficient, and the difference between the neighboring robot's assessment value and its own assessment value. The learning rate determines the step size of the weight update, and the neighborhood influence coefficient reflects the impact of the neighboring robot's welding quality on its own weight adjustment, allowing weight adjustment to be based on both its own welding quality and the situation of its surrounding companions. During implementation, the target parameter values and evaluation weights are determined according to the welding process standards. During the model training phase, the weld quality evaluation value is calculated using a large amount of welding data to train the model to master the weight update rules. In actual applications, appropriate learning rates and initial values of neighborhood influence coefficients are set, the weight update amount is calculated in real time, changes in welding quality are monitored, and these parameters are adjusted based on whether the quality has improved, ensuring that the model can continuously optimize decisions through weight updates and improve the stability of welding quality.
[0114] Preferably, configuring a local decision module of a multi-agent reinforcement learning model for each robot in step S3 includes the following sub-steps:
[0115] S3.1: Deploy an independent neural network computing unit for each robot. This unit consists of an input layer, a hidden layer, and an output layer. The number of neurons in the input layer is consistent with the dimension of the state parameters (welding current and voltage parameters each occupy one dimension, and the state parameters of adjacent robots increase in number). The hidden layer uses the ReLU activation function, and the output layer uses the Tanh activation function to limit the range of action decisions.
[0116] S3.2: Embed the robot hardware constraint parameters into the neural network computing unit, including the maximum welding gun movement speed, welding current adjustment range, and welding gun posture angle rotation limit. By setting a threshold truncation mechanism in the output layer, the output local action decision is ensured to be within the hardware constraint range;
[0117] S3.3: Preload the state-action-feedback data pairs from historical welding tasks and use offline training to initialize the weight parameters of the neural network calculation unit. During the training process, the weld formation errors in the historical data are used as the correction basis for back propagation;
[0118] S3.4: Configure a real-time data acquisition interface for the local decision-making module. The interface is connected to the robot's sensor system (temperature sensor, visual sensor, position sensor), and collects welding state parameters in a 10ms cycle and converts them into digital signals that can be recognized by the neural network.
[0119] Preferably, setting up the global coordination module of the multi-agent reinforcement learning model in step S4 includes the following sub-steps:
[0120] S4.1: Build a global data fusion center, which is connected to the local decision modules of all robots via industrial Ethernet. It uses a timestamp synchronization mechanism to receive local action decisions and state parameters sent by each module. Data transmission uses the UDP protocol to reduce transmission delays.
[0121] S4.2: Deploy a task progress monitoring unit in the global data fusion center. This unit calculates the completed length percentage of each subtask in real time based on the welding position information reported by each robot and generates a task progress heat map. Different colors in the heat map represent different progress intervals.
[0122] S4.3: Configure the global coordination algorithm execution unit, which calls the core function of the heterogeneous multi-robot task allocation algorithm, takes the task progress heat map data, robot position distribution data, and subtask priority parameters as input, and calculates the coordination priority of each robot;
[0123] S4.4: Based on the coordination priority and the distance parameters between robots, a global coordination factor is generated and sent to the corresponding local decision module via broadcast. The sending frequency is consistent with the decision cycle of the local decision module (that is, each time a local decision is generated, a global coordination factor is received synchronously).
[0124] Preferably, in step S5, the local decision module of each robot updates the local action decision according to the global coordination factor and drives the robot to perform the welding operation, which includes the following sub-steps:
[0125] S5.1: After receiving the global coordination factor, the local decision module performs a matrix multiplication operation on the local action decision output by itself. During the operation, weighted corrections are made according to the welding current, voltage, and welding gun movement trajectory parameter categories. Different parameter categories correspond to different coordination factor components.
[0126] S5.2: Convert the weighted corrected action decision into the drive signal of each actuator of the robot. Specifically, the welding gun movement trajectory adjustment is converted into the number of pulses of the servo motor, and the welding current and voltage correction values are converted into analog voltage signals (4-20mA standard signal) of the power controller.
[0127] S5.3: The drive signal is processed by the robot's control system (PLC controller) and sent to the corresponding actuator (the servo motor drives the welding gun to move, and the welding power supply adjusts the output current and voltage). The actuator changes its operating state according to the signal instruction;
[0128] S5.4: While performing the welding operation, the robot's feedback sensor group is activated. The temperature sensor is inserted 5 mm from the edge of the molten pool to collect temperature data. The vision sensor is installed 30 cm behind the welding gun to capture the weld formation image. The position sensor uses an encoder to record the three-dimensional coordinates of the welding gun in real time.
[0129] like Figure 2 As shown in the figure, the distributed reinforcement learning control system for multi-robot collaborative welding includes the following units:
[0130] Heterogeneous robot state parameter acquisition and preprocessing unit, which is connected to the sensor array (temperature, vision, position sensor) of each robot through a wired cable, and is used to receive the original sensor signal and convert it into a standardized digital signal;
[0131] The multi-source data fusion and state space construction unit has its input connected to the output of the heterogeneous robot state parameter acquisition and preprocessing unit. It generates state space data including all robot and environment information through data splicing and dimension expansion operations.
[0132] The distributed task decomposition and candidate allocation unit has its input connected to the output of the multi-source data fusion and state space construction unit. It runs a heterogeneous multi-robot task allocation algorithm internally and outputs a subtask list and candidate robot matching relationships.
[0133] A multi-agent local decision-making neural network cluster, in which each neural network node corresponds to a robot. The node input is connected to the output of the distributed task decomposition and candidate allocation unit and the output of the heterogeneous robot state parameter acquisition and preprocessing unit, and the output generates a local action decision.
[0134] The global coordination factor calculation and distribution unit has its input connected to the output of the multi-agent local decision-making neural network cluster. After generating the global coordination factor through calculation, it will be sent to the corresponding neural network nodes respectively.
[0135] A robot drive and welding execution unit cluster, in which each execution unit corresponds to a robot, has its input end connected to the output end of a multi-agent local decision-making neural network cluster (corrected by a global coordination factor), and its output end is connected to the robot's drive motor and welding power supply to drive the robot to complete the welding operation.
[0136] A distributed reinforcement learning control method and system for multi-robot collaborative welding offers significant advantages in task allocation, effectively overcoming the inflexibility of traditional algorithms. By combining a heterogeneous multi-robot task allocation algorithm with a dynamic adjustment mechanism, the overall welding task is decomposed into subtasks, fully incorporating the welding capability parameters and real-time status of each robot. During task execution, the allocation scheme is continuously optimized based on progress parameters. In the event of a robot failure or a sudden change in weld parameters, the global coordination module quickly reconfigures the matching relationship between subtasks and robots, avoiding idle resources and task conflicts caused by static allocation and improving the smoothness of task execution.
[0137] In terms of multi-agent decision-making collaboration, its advantage lies in achieving a deep coupling between decision-making and welding process parameters, resolving the lack of correlation between the two in existing technologies. The local decision-making module directly uses key process parameters such as molten pool temperature and weld formation dimensions, as well as the status of neighboring robots, as input. The output action decision includes welding gun trajectory adjustment and welding parameter correction, ensuring that the action instructions are consistent with actual welding requirements. The generation of the global coordination factor comprehensively considers task progress, robot spacing, and regional overlap, enabling each robot's local decision-making to form effective coordination at the global level, significantly improving the accuracy of multi-robot collaboration.
[0138] The system also possesses powerful adaptive optimization capabilities, further enhancing its adaptability to complex welding environments. By collecting welding process parameters in real time and feeding them back to the multi-agent reinforcement learning model, the model continuously adjusts the state space representation and decision weights, dynamically optimizing the robot's action decisions as welding conditions change. Furthermore, the hardware constraints embedded in the local decision module and the offline pre-training mechanism ensure the feasibility and initial accuracy of the decision output. Combined with online dynamic adjustment, the system maintains stable welding quality in the face of environmental interference, meeting the requirements of high-precision and high-stability collaborative welding.
[0139] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "disposed," "installed," "connected," "connected," and "fixed" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; they may refer to mechanical connections or electrical connections; they may refer to direct connections or indirect connections through an intermediate medium; and they may refer to internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0140] While embodiments of the present invention have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A distributed reinforcement learning control method for multi-robot collaborative welding, characterized in that: The following steps are involved: Step S1: Construct a heterogeneous multi-robot collaborative welding state space including welding current, welding voltage, welding speed, welding gun posture, weld position and welding environment obstacle distribution, and determine the initial feasible operation area of each robot in the state space based on the welding operation range, motion constraints and welding capability parameters of each robot; Step S2: Based on the state space and the initial feasible operation area of each robot, a heterogeneous multi-robot task allocation algorithm is used to decompose the overall welding task into several subtasks, each subtask includes a preset weld segment welding requirement, a welding accuracy threshold and a task execution time window, and at least one candidate robot is assigned to each subtask; Step S3: Configure a local decision module of a multi-agent reinforcement learning model for each robot, the local decision module takes the robot's own current welding state parameters, the welding state parameters of neighboring robots and the subtask requirement parameters as input, and outputs a welding gun movement trajectory adjustment local action decisions based on the overall amount and welding parameter correction value; step S4: setting a global coordination module of the multi-agent reinforcement learning model, which receives the local action decisions of all robots and the completion progress parameters of the current overall welding task, and generates a global coordination factor for correcting the local decisions of each robot through the dynamic adjustment mechanism of the heterogeneous multi-robot task allocation algorithm; step S5: the local decision module of each robot updates the output local action decision according to the global coordination factor, and converts the updated local action decision into a specific robot drive instruction to drive the robot to perform the welding operation, while collecting the molten pool temperature and weld formation size parameters in real time during the welding process; step S6: feeding back the collected molten pool temperature and weld formation size parameters to the multi-agent reinforcement learning model, and the model adjusts the state space representation method and the decision weight of the local decision module according to the feedback parameters, and enters the decision cycle for the next round of welding operation.
2. The method according to claim 1, characterized in that In step S2, the heterogeneous multi-robot task allocation algorithm matches subtasks with candidate robots by constructing a task allocation optimization model, and the optimization model is: Among them, T ij represents the comprehensive fitness of the i-th subtask assigned to the j-th robot, α, β, γ are weight coefficients and α+β+γ=1, W i is the welding workload of the ith subtask, C j is the welding capability coefficient of the jth robot, D ij is the distance from the jth robot to the starting position of the ith subtask, V j is the moving speed of the jth robot, P j is the welding accuracy parameter of the jth robot, P max It has the largest welding accuracy parameter among all robots; In step S3, the action value function of the local decision module is: Among them, Q j (s, a) is the action value of the jth robot performing action a in the global state s, ω1 and ω2 are weight parameters and ω1+ω2=1, For the jth robot based on its own state s j and its own action a j The local action value, N j is the set of neighboring robots of the j-th robot, For the kth neighboring robot based on its own state s k and its own action a k The local action value.
3. The method according to claim 1, characterized in that In step S3, the decision-making process of the local decision module includes: based on the current welding current I of the robot j , welding voltage U j , welding gun posture angle θ j Construct its own state vector s j ; Receive the state vector s of other robots within the radius R k And the corresponding subtask requirements in the weld thickness h i , welding speed threshold v i,min 、v i,max ; The state vector s j , the set of neighboring robot state vectors {s k } and subtask requirement parameters {h i , v i,min , v i,max The input is sent to the local decision network, which consists of three fully connected layers. The number of input neurons in the first layer is the sum of the state vector dimension and the number of neighboring robots. The number of neurons in the second layer is 64. The number of output neurons in the third layer is the dimension of the action space. The local decision network outputs the adjustment amount Δx of the welding gun movement trajectory in the x, y, and z axis directions j , Δy j , Δz j and welding current correction value ΔI j , welding voltage correction value ΔU j , forming a local action decision a j .
4. The method according to claim 1, wherein In step S4, the process of generating the global coordination factor by the global coordination module includes: Collect the local action decisions of all robots a j , extract the welding gun moving speed v in each decision j , welding current I j And the corresponding subtask number; Based on the welding progress P of each subtask i and the overlapping area O where the robot performs subtasks jk , calculate the global task coordination index Among them, n is the total number of subtasks, S i is the total area of the welding region of the i-th subtask; Construct the global coordination factor matrix Λ=[λ jk ], where λ jk =exp(-μ·D jk )·(1+v·G), μ and v are adjustment parameters, D jk is the real-time distance between the jth and kth robots; The elements in the global coordination factor matrix Λ are used as correction coefficients for the local decisions of each robot, where λ jj Used to correct the robot's own decision weight, λ jk Used to modify the influence weight of neighboring robots on decision making.
5. The method according to claim 1, characterized in that In step S5, the robot uses a dynamic welding parameter adjustment model when performing the welding operation according to the updated local action decision: Among them, I j (t) is the welding current of the jth robot at time t, is the initial welding current, ΔI j is the welding current correction value output in step S3, φ is the current adjustment frequency coefficient, t is the welding time, ζ is the temperature correction coefficient, T j (t) is the molten pool temperature collected at time t, and T0 is the target molten pool temperature; At the same time, the real-time adjustment of the welding gun posture angle satisfies: Among them, θ j (t) is the welding gun posture angle of the j-th robot at time t, is the initial welding gun posture angle, Δθ j is the welding gun posture adjustment amount in local action decision, and κ is the posture adjustment rate coefficient.
6. The method according to claim 1, characterized in that In step S6, the process of adjusting the decision weights of the multi-agent reinforcement learning model includes: Melt pool temperature T based on feedback j (t) and weld size, H j (t) is the weld height, W j (t) is the weld width, and the weld quality evaluation value Q is calculated j =α T ·|T j (t)-T0|+α H ·|H j (t)-H0|+α W |W j (t)-W0|, where α T , α H , α W is the quality assessment weight, T0, H0, and W0 are the target molten pool temperature, target weld height, and target weld width, respectively; The weld quality evaluation value Q j As the reward and punishment signals are input to the multi-agent reinforcement learning model, the model uses the following weight update formula: Among them, ω j (t+1),ω j (t) are the decision weights of the jth robot at time t+1 and time t, η is the learning rate, is the gradient of the weld quality assessment value to the weight, ξ is the neighborhood influence coefficient, Q k is the weld quality evaluation value of the kth neighboring robot.
7. The method according to claim 1, characterized in that Configuring the local decision-making module of the multi-agent reinforcement learning model for each robot in step S3 includes the following sub-steps: S3.1: Deploy an independent neural network computing unit for each robot. This unit consists of an input layer, a hidden layer, and an output layer. The number of neurons in the input layer is consistent with the dimension of the state parameters. The hidden layer uses the ReLU activation function, and the output layer uses the Tanh activation function to limit the range of action decisions. S3.2: Embed the robot hardware constraint parameters into the neural network computing unit, including the maximum welding gun movement speed, welding current adjustment range, and welding gun posture angle rotation limit. By setting a threshold truncation mechanism in the output layer, the output local action decision is ensured to be within the hardware constraint range; S3.3: Preload the state-action-feedback data pairs from historical welding tasks and use offline training to initialize the weight parameters of the neural network calculation unit. During the training process, the weld formation errors in the historical data are used as the correction basis for back propagation; S3.4: Configure a real-time data acquisition interface for the local decision module. The interface is connected to the robot's sensor system, collects welding state parameters in a 10ms cycle and converts them into digital signals that can be recognized by the neural network.
8. The method according to claim 1, characterized in that Setting up the global coordination module of the multi-agent reinforcement learning model in step S4 includes the following sub-steps: S4.1: Build a global data fusion center, which is connected to the local decision modules of all robots via industrial Ethernet. It uses a timestamp synchronization mechanism to receive local action decisions and state parameters sent by each module. Data transmission uses the UDP protocol to reduce transmission delays. S4.2: Deploy a task progress monitoring unit in the global data fusion center. This unit calculates the completed length percentage of each subtask in real time based on the welding position information reported by each robot and generates a task progress heat map. Different colors in the heat map represent different progress intervals. S4.3: Configure the global coordination algorithm execution unit, which calls the core function of the heterogeneous multi-robot task allocation algorithm, takes the task progress heat map data, robot position distribution data, and subtask priority parameters as input, and calculates the coordination priority of each robot; S4.4: Based on the coordination priority and the distance parameters between robots, a global coordination factor is generated and sent to the corresponding local decision module via broadcast. The sending frequency is consistent with the decision cycle of the local decision module.
9. The method according to claim 1, characterized in that In step S5, the local decision module of each robot updates the local action decision according to the global coordination factor and drives the robot to perform the welding operation, which includes the following sub-steps: S5.1: After receiving the global coordination factor, the local decision module performs a matrix multiplication operation on the local action decision output by itself. During the operation, weighted corrections are made according to the welding current, voltage, and welding gun movement trajectory parameter categories. Different parameter categories correspond to different coordination factor components. S5.2: Convert the weighted corrected action decision into the drive signal of each actuator of the robot. The welding gun movement trajectory adjustment is converted into the number of pulses of the servo motor, and the welding current and voltage correction values are converted into analog voltage signals of the power controller. S5.3: The drive signal is processed by the robot's control system and sent to the corresponding actuator. The actuator changes its operating state according to the signal instruction. S5.4: While performing the welding operation, the robot's feedback sensor group is activated. The temperature sensor is inserted 5 mm from the edge of the molten pool to collect temperature data. The vision sensor is installed 30 cm behind the welding gun to capture the weld formation image. The position sensor uses an encoder to record the three-dimensional coordinates of the welding gun in real time.
10. A distributed reinforcement learning control system for multi-robot collaborative welding, characterized by: The following units are included: Heterogeneous robot state parameter acquisition and preprocessing unit, which is connected to the sensor array of each robot via a wired cable and is used to receive raw sensor signals and convert them into standardized digital signals; The multi-source data fusion and state space construction unit has its input connected to the output of the heterogeneous robot state parameter acquisition and preprocessing unit. It generates state space data including all robot and environment information through data splicing and dimension expansion operations. The distributed task decomposition and candidate allocation unit has its input connected to the output of the multi-source data fusion and state space construction unit. It runs a heterogeneous multi-robot task allocation algorithm internally and outputs a subtask list and candidate robot matching relationships. A multi-agent local decision-making neural network cluster, in which each neural network node corresponds to a robot. The node input is connected to the output of the distributed task decomposition and candidate allocation unit and the output of the heterogeneous robot state parameter acquisition and preprocessing unit, and the output generates a local action decision. The global coordination factor calculation and distribution unit has its input connected to the output of the multi-agent local decision-making neural network cluster. After generating the global coordination factor through calculation, it will be sent to the corresponding neural network nodes respectively. A robot drive and welding execution unit cluster, where each execution unit corresponds to a robot. The input end is connected to the output end of the multi-agent local decision-making neural network cluster, and the output end is connected to the robot's drive motor and welding power supply to drive the robot to complete the welding operation.
Citation Information
Cited By
Multi-point intelligent welding system and method
CN121004398A
Multi-point intelligent welding system and method
CN121004398B
Multi-welding gun path cooperation welding system and method based on parameter dynamic adjustment
CN121083011A
Robot collaborative welding control method and system for water purifier barrel and end cover
CN121361085A
Robot collaborative welding control method and system for water purifier cylinder and end cover
CN121361085B