Cluster robot space on-orbit assembly sequence planning method based on reinforcement learning

By employing a reinforcement learning-based on-orbit assembly sequence planning method for swarm robots, the problems of poor flexibility and high energy consumption in existing technologies are solved, enabling efficient and robust multi-robot collaborative assembly, and improving assembly efficiency and the ability to cope with environmental interference.

CN121492020APending Publication Date: 2026-02-10HARBIN INST OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511661530.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-13
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

In existing technologies, on-orbit assembly methods in space suffer from problems such as poor flexibility, high energy consumption, and conflicts in multi-robot collaboration. Single-robot trajectory planning is difficult to meet the needs of complex and large-scale assembly tasks, and multi-robot collaboration lacks dynamic optimization capabilities.

Method used

A reinforcement learning-based on-orbit assembly sequence planning method for swarm robots is adopted. Through dynamic simulation and trajectory planning, reward functions and constraints are designed, and multi-agent reinforcement learning is combined for simulation training to optimize the assembly process.

Benefits of technology

It achieves autonomous optimization of the assembly sequence, reduces energy consumption, minimizes trajectory planning conflicts, improves assembly efficiency and robustness, and adapts to faults and environmental interference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121492020A_ABST
    Figure CN121492020A_ABST
Patent Text Reader

Abstract

The invention discloses a cluster robot space on-orbit assembly sequence planning method based on reinforcement learning, and belongs to the technical field of robot automation control. Comprising the following steps: S1, performing dynamic simulation and target trajectory planning on a cluster robot; s2, based on a robot working scene, modeling an assembly task into a Markov decision process, and providing an interaction environment for multi-agent reinforcement learning; s3, designing a reward function and a constraint condition based on a target trajectory planning result and a Markov decision process model; s4, performing simulation training on the assembly process of the cluster robot; and S5, performing sequence planning on a space on-orbit assembly task by using the trained agent model. By the adoption of the cluster robot space on-orbit assembly sequence planning method based on reinforcement learning, the assembly sequence can be autonomously optimized, energy consumption is reduced, and conflicts are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot automation control technology, and in particular to a method for planning spatial on-orbit assembly sequences of swarm robots based on reinforcement learning. Background Technology

[0002] In-orbit assembly is a key technology for future deep space exploration and space infrastructure construction. Traditional methods rely on preset commands or manual remote operation, which suffers from poor flexibility, high energy consumption, and conflicts in multi-robot collaboration. In existing technologies, single-robot trajectory planning is difficult to meet the needs of complex and large-scale assembly tasks, while multi-robot collaborative sequence planning lacks dynamic optimization capabilities. Therefore, current technologies suffer from the following problems: assembly sequences lack autonomous optimization capabilities, energy consumption is not ideal, and conflicts exist in the trajectory planning process. Summary of the Invention

[0003] The purpose of this invention is to provide a method for planning the spatial on-orbit assembly sequence of swarm robots based on reinforcement learning, thereby solving the above-mentioned technical problems.

[0004] To achieve the above objectives, this invention provides a reinforcement learning-based method for on-orbit assembly sequence planning of swarm robots, comprising the following steps: S1: Perform dynamic simulation and target trajectory planning for the swarm robots; S2: Based on the robot's working scenario, the assembly task is modeled as a Markov decision process, providing an interactive environment for multi-agent reinforcement learning; S3: Based on the target trajectory planning results and the Markov decision process model, design the reward function and constraints; S4: Simulate and train the assembly process of the cluster robot; S5: Use the trained agent model to perform sequence planning for on-orbit assembly tasks in space.

[0005] Preferably, S1 specifically includes: S11. Construct a dynamic model of the floating base and robotic arm, consisting of a floating base and a robotic arm model. Input the robot's working range parameters into the dynamic model of the floating base and robotic arm. The working range parameters include the coordinates of the working area boundary and the robot's movement radius. ; S12. Set the constraints and optimization objectives for the dynamic model of the floating base robotic arm. The constraint is that the end effector of the floating base robotic arm reaches the set assembly position. The optimization objective is to achieve energy conservation in motion using a trajectory planning algorithm. The trajectory planning must satisfy the following constraints: (1); (2); (3); in, Joint angle vector, The total time taken for the robotic arm to move from the initial position to the target position is given by formula (1), which indicates that the initial velocity and end velocity of the robotic arm joints are both 0. For the terminal velocity, Let be the Jacobian matrix of the robot system. Formula (2) represents the mapping relationship between the end-effector velocity and the joint angle vector; in Formula (3) and Indicates the physical limit range of the robotic arm; S13. Design an optimization function based on the optimization objective to calculate the total energy consumption of trajectory planning. The optimization function is shown below: (4); in, Disturbance to the position of the floating base For attitude disturbance of the floating base. For the position deviation of the robotic arm end effector, This is a penalty term for the joint angular acceleration of the robotic arm. , , , These are the weighting coefficients for the floating base position disturbance, the floating base attitude disturbance, the robotic arm end-effector position deviation, and the robotic arm joint angular acceleration penalty term, respectively.

[0006] Preferably, S2 specifically includes: S21: Model the Markov decision process for the space on-orbit assembly task. Input the modular design parameters of the target space structure and the robot's working range parameters in S11 into the Markov decision process model. The modular design parameters include the total number of modules, the size of a single module, and the spatial distribution coordinates of the modules. S22: Using a Markov decision process model, the functional units in the target spatial structure are decomposed into independent assembly modules. The spatial coordinates of each assembly module are recorded, and then the coordinates of each assembly module are mapped onto a planar mesh, while preserving the relative positional relationships between the assembly modules. Specifically, the working area where the robot is located in the Markov decision process model is the state space, and the working areas of the current assembly module and the next assembly module are the action space. The state space function is shown below: (5); in, This refers to the position parameters of the robot's current module area, used to describe the robot's real-time position status; The action space function is shown below: (6); in, This means to remain in place. These represent the robot's movement in the up, down, left, and right directions within the work area where the current assembly module is located.

[0007] Preferably, S3 specifically includes: S31: Total energy consumption for trajectory planning based on S12 Based on the state-space function and action-space function of S22, and considering assembly accuracy, design the reward function and set of constraints. The reward function is as follows: (7); in, This is the overall reward value for a single-step assembly action. The reward value for successfully assembling an unassembled module after selecting a movement direction for the currently assembled module. The energy consumed by the robotic arm during assembly. The energy consumed for the floating base to move For actions that do not violate the set of constraints Laziness punishment at the time For the set of constraints that are violated Punishment at that time; Constraint set This includes robot non-interference constraints and assembly physical constraints. Robot non-interference constraints mean that two robots do not enter the same work area of ​​the module to be assembled. Assembly physical constraints mean that when a new assembly module is assembled, it must have a physical connection with the already assembled assembly module, the assembly module must correspond to the assembly position, the assembly path of the assembly module must be unobstructed, and the robot's movement range must not exceed the set work area boundary. Here, unobstructed assembly path means that there is one or more unassembled areas around the module to be assembled. S32: Total energy consumption of the robotic arm Energy consumption of floating base movement The calculation formula is as follows: (8); in, This is the energy consumption penalty coefficient for the robotic arm. For the first The angular change of each robotic arm joint; (9); in, Energy consumption coefficient per unit of movement steps This represents the number of steps taken.

[0008] Preferably, S4 specifically includes: S41: Based on the Markov decision process model and combined with the MAPPO algorithm, the swarm robot is regarded as a multi-agent and the assembly process is simulated and trained. The multi-agent executes the state space perception and action space selection of S22 in the interactive environment of the Markov decision process model, and then obtains training feedback according to the reward function of S3. The decision logic is adjusted through the policy update mechanism of the MAPPO algorithm to continuously simulate the assembly process under different initial states. S42: Use a preset cumulative reward value as the pre-training performance metric, where the preset cumulative reward value is the reward value of the model in a set of assembly scenarios from the initial state to the completion of assembly of all modules. The cumulative sum, when the actual single-step reward value The cumulative sum is greater than or equal to the preset cumulative reward value, and no trigger is given. and When the multi-agent assembly planning model training reaches the target, training stops.

[0009] Preferably, S5 specifically includes: performing sequence planning for the actual assembly task, training a qualified multi-agent assembly planning model based on S42, and inputting actual assembly task parameters into this model. The actual assembly task parameters include the total number of modules in the target structure, the size and spatial distribution coordinates of a single module, the initial position coordinates of the robot, the assembly accuracy, the number of robots, and the robot's movement radius. .

[0010] Preferably, in the dynamic model of the floating base robotic arm based on S1, the positional deviation between the end effector of the robotic arm and the module assembly interface is less than or equal to... To match assembly precision.

[0011] Preferably, the methods for determining the successful assembly of the assembly module in S3 are as follows: Method 1: The module to be assembled is directly adjacent to the assembled module and the spatial distance between them is less than or equal to 1.1 times the external dimensions of the module to be assembled; Method 2: The module to be assembled is indirectly adjacent to the assembled module through an unassembled intermediate module, and the assembly path of the intermediate module is unobstructed.

[0012] Preferably, S4 employs a distributed training framework, which simultaneously simulates... The scenarios have different initial states. The initial state of each scenario includes the robot's initial position and the number of unassembled modules.

[0013] Preferably, in S5, when performing sequence planning for the actual assembly task, the multi-agent assembly planning model collects two types of data in real time: the actual position of the robot and the connection status of the assembled modules, through the sensors mounted on the robot. The two types of data are then fed back to the multi-agent assembly planning model. When an abnormal situation is detected, the multi-agent assembly planning model reassigns the assembly task.

[0014] Therefore, the present invention employs the above-mentioned reinforcement learning-based on-orbit assembly sequence planning method for swarm robots, which has the following beneficial effects: 1. By using dynamic simulation and energy-optimal trajectory planning of the floating base robotic arm, an energy consumption benchmark is provided for reinforcement learning; at the same time, a composite reward function is designed to integrate assembly efficiency, energy consumption and constraint penalties to guide the agent to converge quickly.

[0015] 2. By modular task modeling, the assembly task is abstracted into a state-action space, providing a foundation for multi-robot collaborative decision-making. Combined with multi-agent collaborative training, distributed learning is used to solve the conflict and collaboration problems between robots.

[0016] 3. Compared with single robot assembly, cluster robots improve assembly efficiency due to their numerical advantage, and are more capable of coping with faults and external environmental interference, with better robustness.

[0017] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the spatial on-orbit assembly sequence planning method for swarm robots based on reinforcement learning according to the present invention. Figure 2 This is a schematic diagram illustrating the logical relationship between the reinforcement learning agent and the environment according to the present invention; Figure 3 This is a schematic diagram of the assembly physical constraints of the present invention; Figure 4 This is a diagram illustrating the pre-training process of the reinforcement learning model of this invention. Figure 5 This describes the convergence process of the reward function in this invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages disclosed in the embodiments of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the embodiments of the present invention and are not intended to limit the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of this application. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout.

[0020] It should be noted that the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, such as a process, method, system, product, or server that includes a series of steps or units, not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or device.

[0021] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0022] like Figures 1-5 As shown, this invention provides a method for on-orbit assembly sequence planning of swarm robots based on reinforcement learning, comprising the following steps: S1: performing dynamic simulation and target trajectory planning for the swarm robots; S2: modeling the assembly task as a Markov decision process based on the robot's working scenario, providing an interactive environment for multi-agent reinforcement learning; S3: designing reward functions and constraints based on the target trajectory planning results and the Markov decision process model; S4: simulating and training the assembly process of the swarm robots; S5: using the trained agent model to perform sequence planning for the on-orbit assembly task. This invention is applicable to modular on-orbit autonomous assembly scenarios of large space structures, such as solar panels for space solar power stations, space station module segments, and large space trusses.

[0023] S1 specifically includes: S11, constructing a dynamic model of the floating base and robotic arm, consisting of a floating base and a robotic arm model. This model references the mass characteristics of the robotic arm, such as the mass characteristics of the UR5 robotic arm, and combines the mass and size parameters of the base to construct the model. The working range parameters of the robot are then input into the dynamic model of the floating base and robotic arm. These working range parameters include the coordinates of the work area boundaries, as in the space solar power station assembly scenario. Planar range The coordinate range, the robot's movement radius Can be taken as It can cover adjacent Each module area; S12, set the constraints and optimization objectives of the dynamic model of the floating base manipulator, with the end of the floating base manipulator reaching the set assembly position as the constraint, such as the positional deviation of the module assembly interface center. To meet the docking accuracy requirements of space station modules, and with the optimization goal of achieving maneuver energy conservation through trajectory planning algorithms, this invention calculates the minimum energy consumption as the means to achieve energy conservation. The trajectory of the robotic arm's end effector is planned, and the constraints that the trajectory planning must meet are as follows: (1); (2); (3); in, Joint angle vector, The total time taken for the robotic arm to move from its initial position to its target position is given by formula (1), which indicates that both the initial velocity and the end effector velocity of the robotic arm joints are... To avoid start-stop shock; For the terminal velocity, The Jacobian matrix of the robot system. Based on the DH parameters of the UR5 robotic arm, formula (2) represents the mapping relationship between the end-effector velocity and the joint angle vector; in formula (3) and S13. Indicate the minimum and maximum values ​​of the physical limit range of the robotic arm; S14. Design an optimization function based on the optimization objective to calculate the total energy consumption of trajectory planning. The optimization function is shown below: (4); in, Disturbance to the position of the floating base For attitude disturbance of the floating base. For the position deviation of the robotic arm end effector, This is a penalty term for the joint angular acceleration of the robotic arm. , , , These are the weighting coefficients for the floating base position disturbance, floating base attitude disturbance, robotic arm end-effector position deviation, and robotic arm joint angular acceleration penalty term, respectively. The range of values ​​for the weighting coefficients is as follows: .

[0024] Preferably, S2 specifically includes: S21: Model a Markov decision process for the on-orbit assembly task. Input the modular design parameters of the target space structure and the robot's working range parameters from S11 into the Markov decision process model. The modular design parameters include the total number of modules, the size of each module, and the spatial distribution coordinates of the modules. For example, the solar panels of a space solar power station have... Each module Work area by The module is divided into modules, and the size of a single module is... The rectangular prism structure, with the spatial distribution coordinates of the modules represented as follows, is based on the modules. For example, module The coordinates are , , S22: Using a Markov decision process model, functional units in the target spatial structure, such as the power generation unit of the solar panel and the support unit of the truss, are decomposed into independent assembly modules. The spatial coordinates of each assembly module are recorded, and then the coordinates of each assembly module are mapped onto a planar mesh. Specifically, the working area where the robot is located in the Markov decision process model is the state space, and the working areas of the current assembly module and the next assembly module are the action space. The state space function is shown below: (5); in, The position parameters of the robot's current module region are used to describe the robot's real-time position state; the motion space function is shown below: (6); in, This means to remain in place. These represent the robot's movement in the up, down, left, and right directions within the work area of ​​the current assembly module. For example, the assembly module can be mapped onto a 4m × 4m planar grid, with a grid cell size of... Consistent with module size, preserving the relative positional relationship between modules, such as modules being... The top of The right side is .

[0025] S3 specifically includes: S31: Total energy consumption for trajectory planning based on S12 Based on the state-space function and action-space function of S22, and considering assembly accuracy, design the reward function and set of constraints. The reward function is as follows: (7); in, This is the overall reward value for a single-step assembly action. The reward value for successfully assembling an unassembled module after selecting a movement direction for the currently assembled module. The energy consumed by the robotic arm during assembly. The energy consumed for the floating base to move For actions that do not violate the set of constraints Laziness punishment at the time For the set of constraints that are violated Time-based penalties; set of constraints This includes robot non-interference constraints and assembly physical constraints. Robot non-interference constraints mean that two robots cannot enter the same work area of ​​the module to be assembled. Assembly physical constraints mean that when a new assembly module is assembled, it must have a physical connection with an already assembled module, the assembly module must correspond to the assembly position, the assembly path of the assembly module must be unobstructed, and the robot's movement range must not exceed the set work area boundary. Here, "unobstructed assembly path" means that there is one or more unassembled areas around the module to be assembled; S32: Total energy consumption of the robotic arm. Energy consumption of floating base movement The calculation formula is as follows: (8); in, This is the energy consumption penalty coefficient for the robotic arm. For the first The angular change of each robotic arm joint; (9); in, Energy consumption coefficient per unit of movement steps The assembly efficiency is improved by using the reward function of S3 and the set of constraints to determine the number of steps to move.

[0026] S4 specifically includes: S41: Based on the Markov decision process model, combined with the MAPPO algorithm, which is a distributed policy optimization algorithm that supports... Parallel training of the robots is conducted, treating the swarm of robots as multiple agents. Assembly process simulation training is performed, with the multiple agents executing state-space perception and action-space selection (S22) within the interactive environment of a Markov decision process model. Training feedback is then obtained based on the reward function (S3). During the feedback acquisition process, if assembly is successful... Energy consumption Then deduct Final single-step reward value Then, the decision logic is adjusted through the policy update mechanism of the MAPPO algorithm to continuously simulate the assembly process under different initial states; S42: The preset cumulative reward value is used as the pre-training effect indicator, where the preset cumulative reward value is the reward value of the model in a set of assembly scenarios from the initial state to the completion of assembly of all modules. The cumulative sum, such as the total number of modules. At that time, When the actual cumulative reward value The condition that the cumulative sum is greater than or equal to the preset cumulative reward value is met, and no trigger is made. and At that time, the multi-agent assembly planning model training reached the target and training stopped. However, during the training process, the convergence trend of the reward function was as follows: Figure 5 As shown: Iteration The cumulative reward for this time period reaches iteration The convergence value stabilized above 144,000 at this time, and the convergence speed was improved by approximately [missing information] compared to the DDPG algorithm. Once the training is successful, the model can be directly used for actual assembly tasks of space solar power stations and space stations.

[0027] S5 specifically includes: sequence planning for actual assembly tasks, training a qualified multi-agent assembly planning model based on S42, and inputting actual assembly task parameters into this model. These parameters include the total number of modules in the target structure, the size and spatial distribution coordinates of individual modules, the robot's initial position coordinates, assembly accuracy, the number of robots, and the robot's movement radius. The multi-agent assembly planning model can output planning results after analyzing the assembly module sequence, movement direction, energy consumption estimate, and docking force control value of each robot.

[0028] Based on the dynamic model of the floating base robotic arm using S1, the positional deviation between the end effector of the robotic arm and the module assembly interface is less than or equal to... To match assembly precision; the methods for determining successful assembly of assembly modules in S3 are as follows: Method 1: The module to be assembled is directly adjacent to the already assembled module and the spatial distance between them is less than or equal to the external dimensions of the module to be assembled. Method 2: The module to be assembled passes... If an unassembled intermediate module is indirectly adjacent to an assembled module, and the assembly path of the intermediate module is unobstructed, the judgment results of both methods are considered as valid. The criteria for triggering the algorithm are determined, and the judgment process is time-efficient and does not affect real-time assembly. S4 employs a distributed training framework, which simultaneously simulates... Scenarios with different initial states of groups. The range of values ​​is Each group's initial state includes the robot's initial position and the number of unassembled modules. The robot's initial position covers the edge, center, and half-edge of the S11 working area, and the number of unassembled modules covers a percentage of the total number of modules. , , , Scene It needs to cover all combinations of "position × quantity" to ensure the model adapts to different assembly progress and robot deployment locations. Distributed training shortens model training time and improves generalization ability, allowing it to adapt to different cluster sizes of multiple robots without retraining. In S5, when performing sequence planning for actual assembly tasks, the multi-agent assembly planning model uses sensors on the robot, such as vision sensors and force sensors, to collect real-time data on the robot's actual position and the connection status of assembled modules. This data is then fed back to the multi-agent assembly planning model. When anomalies are detected, such as excessive robot position deviation, insufficient docking force, or robot base drive failure, the multi-agent assembly planning model reassigns assembly tasks, reducing assembly delay time.

[0029] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the technical solutions of the present invention, and these modifications or equivalent substitutions cannot cause the modified technical solutions to deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for planning on-orbit assembly sequences of swarm robots based on reinforcement learning, characterized in that: Includes the following steps: S1: Perform dynamic simulation and target trajectory planning for the swarm robots; S2: Based on the robot's working scenario, the assembly task is modeled as a Markov decision process, providing an interactive environment for multi-agent reinforcement learning; S3: Based on the target trajectory planning results and the Markov decision process model, design the reward function and constraints; S4: Simulate and train the assembly process of the cluster robot; S5: Use the trained agent model to perform sequence planning for on-orbit assembly tasks in space.

2. The method for planning the spatial on-orbit assembly sequence of swarm robots based on reinforcement learning according to claim 1, characterized in that: S1 specifically includes: S11. Construct a dynamic model of the floating base and robotic arm, consisting of a floating base and a robotic arm model. Input the robot's working range parameters into the dynamic model of the floating base and robotic arm. The working range parameters include the coordinates of the working area boundary and the robot's movement radius. ; S12. Set the constraints and optimization objectives for the dynamic model of the floating base robotic arm. The constraint is that the end effector of the floating base robotic arm reaches the set assembly position. The optimization objective is to achieve energy conservation in motion using a trajectory planning algorithm. The trajectory planning must satisfy the following constraints: (1); (2); (3); in, Joint angle vector, The total time taken for the robotic arm to move from the initial position to the target position is given by formula (1), which indicates that the initial velocity and end velocity of the robotic arm joints are both 0. For the terminal velocity, Let be the Jacobian matrix of the robot system. Formula (2) represents the mapping relationship between the end-effector velocity and the joint angle vector; in Formula (3) and Indicates the physical limit range of the robotic arm; S13. Design an optimization function based on the optimization objective to calculate the total energy consumption of trajectory planning. The optimization function is shown below: (4); in, Disturbance to the position of the floating base For attitude disturbance of the floating base. For the position deviation of the robotic arm end effector, This is a penalty term for the joint angular acceleration of the robotic arm. , , , These are the weighting coefficients for the floating base position disturbance, the floating base attitude disturbance, the robotic arm end-effector position deviation, and the robotic arm joint angular acceleration penalty term, respectively.

3. The method for planning the spatial on-orbit assembly sequence of swarm robots based on reinforcement learning according to claim 2, characterized in that: S2 specifically includes: S21: Model the Markov decision process for the space on-orbit assembly task. Input the modular design parameters of the target space structure and the robot's working range parameters in S11 into the Markov decision process model. The modular design parameters include the total number of modules, the size of a single module, and the spatial distribution coordinates of the modules. S22: Using a Markov decision process model, the functional units in the target spatial structure are decomposed into independent assembly modules. The spatial coordinates of each assembly module are recorded, and then the coordinates of each assembly module are mapped onto a planar mesh, while preserving the relative positional relationships between the assembly modules. Specifically, the working area where the robot is located in the Markov decision process model is the state space, and the working areas of the current assembly module and the next assembly module are the action space. The state space function is shown below: (5); in, This refers to the position parameters of the robot's current module area, used to describe the robot's real-time position status; The action space function is shown below: (6); in, This means to remain in place. These represent the robot's movement in the up, down, left, and right directions within the work area where the current assembly module is located.

4. The method for planning the spatial on-orbit assembly sequence of swarm robots based on reinforcement learning according to claim 3, characterized in that: S3 specifically includes: S31: Total energy consumption for trajectory planning based on S12 Based on the state-space function and action-space function of S22, and considering assembly accuracy, design the reward function and set of constraints. The reward function is as follows: (7); in, This is the overall reward value for a single-step assembly action. The reward value for successfully assembling an unassembled module after selecting a movement direction for the currently assembled module. The energy consumed by the robotic arm during assembly. The energy consumed for the floating base to move For actions that do not violate the set of constraints Laziness punishment at the time For the set of constraints that are violated Punishment at that time; Constraint set This includes robot non-interference constraints and assembly physical constraints. Robot non-interference constraints mean that two robots do not enter the same work area of ​​the module to be assembled. Assembly physical constraints mean that when a new assembly module is assembled, it must have a physical connection with the already assembled assembly module, the assembly module must correspond to the assembly position, the assembly path of the assembly module must be unobstructed, and the robot's movement range must not exceed the set work area boundary. Here, unobstructed assembly path means that there is one or more unassembled areas around the module to be assembled. S32: Total energy consumption of the robotic arm Energy consumption of floating base movement The calculation formula is as follows: (8); in, This is the energy consumption penalty coefficient for the robotic arm. For the first The angular change of each robotic arm joint; (9); in, Energy consumption coefficient per unit of movement steps This represents the number of steps taken.

5. The method for planning the spatial on-orbit assembly sequence of swarm robots based on reinforcement learning according to claim 4, characterized in that: S4 specifically includes: S41: Based on the Markov decision process model and combined with the MAPPO algorithm, the swarm robot is regarded as a multi-agent and the assembly process is simulated and trained. The multi-agent executes the state space perception and action space selection of S22 in the interactive environment of the Markov decision process model, and then obtains training feedback according to the reward function of S3. The decision logic is adjusted through the policy update mechanism of the MAPPO algorithm to continuously simulate the assembly process under different initial states. S42: Use a preset cumulative reward value as the pre-training performance metric, where the preset cumulative reward value is the reward value of the model in a set of assembly scenarios from the initial state to the completion of assembly of all modules. The cumulative sum, when the actual single-step reward value The cumulative sum is greater than or equal to the preset cumulative reward value, and no trigger is given. and When the multi-agent assembly planning model training reaches the target, training stops.

6. The method for planning the spatial on-orbit assembly sequence of swarm robots based on reinforcement learning according to claim 5, characterized in that: S5 specifically includes: sequence planning for actual assembly tasks, training a qualified multi-agent assembly planning model based on S42, and inputting actual assembly task parameters into this model. These parameters include the total number of modules in the target structure, the size and spatial distribution coordinates of individual modules, the robot's initial position coordinates, assembly accuracy, the number of robots, and the robot's movement radius. .

7. The method for planning the spatial on-orbit assembly sequence of swarm robots based on reinforcement learning according to claim 2, characterized in that: Based on the dynamic model of the floating base robotic arm using S1, the positional deviation between the end of the robotic arm and the module assembly interface is less than or equal to ±0.1mm to match the assembly accuracy.

8. The method for planning the spatial on-orbit assembly sequence of swarm robots based on reinforcement learning according to claim 4, characterized in that: The methods for determining whether the assembly module in S3 is successfully assembled are as follows: Method 1: The module to be assembled is directly adjacent to the assembled module and the spatial distance between them is less than or equal to 1.1 times the external dimensions of the module to be assembled; Method 2: The module to be assembled is indirectly adjacent to the assembled module through an unassembled intermediate module, and the assembly path of the intermediate module is unobstructed.

9. The method for planning the spatial on-orbit assembly sequence of swarm robots based on reinforcement learning according to claim 5, characterized in that: S4 employs a distributed training framework, which simultaneously simulates... The scenarios have different initial states. The initial state of each scenario includes the robot's initial position and the number of unassembled modules.

10. The method for planning the spatial on-orbit assembly sequence of swarm robots based on reinforcement learning according to claim 6, characterized in that: When performing sequence planning for actual assembly tasks in S5, the multi-agent assembly planning model collects two types of data in real time: the robot's actual position and the connection status of assembled modules, through the sensors mounted on the robot. These two types of data are then fed back to the multi-agent assembly planning model. When an abnormal situation is detected, the multi-agent assembly planning model reassigns the assembly tasks.