Multi-robot navigation and calculation unloading method based on MARL

Through a MARL-based approach, a MEC-assisted multi-robot navigation system supported by a satellite-ground network was designed. The DAOMAN algorithm was used to optimize mobile control and computational offloading, solving the problems of high computational latency and communication bottlenecks in the multi-robot system and achieving efficient task completion.

CN120620174APending Publication Date: 2025-09-12SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510594226.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

In the existing technology, multi-robot systems have problems such as high computing delay, large communication delay, low task allocation efficiency and data exchange bottlenecks between robots during computation offloading and navigation, resulting in low task completion efficiency.

Method used

A multi-robot navigation system supported by MEC-assisted satellite-ground network is designed using a multi-agent reinforcement learning (MARL) method. By decomposing the joint optimization problem into mobility control and computation offloading sub-problems, and using the DAOMAN algorithm optimization strategy, combined with the MATD3 and PPO algorithms, the joint optimization of robot task allocation, mobility control and computation offloading is achieved.

Benefits of technology

It effectively shortens the time it takes for robots to reach the target location and the time it takes to complete computing tasks, improves the task completion efficiency of the multi-robot system, and reduces communication delays and computing delays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120620174A_ABST
    Figure CN120620174A_ABST
Patent Text Reader

Abstract

The invention discloses an MARL-based multi-robot navigation and calculation unloading method, relates to the technical field of intersection of multi-robot navigation and edge calculation, and designs an MEC-assisted satellite-ground network supported multi-robot navigation system. Based on a multi-robot navigation system, a joint optimization problem of multi-robot movement control and calculation unloading is formalized; decomposing a joint optimization problem into a movement control sub-problem and a calculation unloading sub-problem, and designing three elements of reinforcement learning; on the basis of three elements of reinforcement learning, an optimal strategy for solving movement control and calculating unloading through a DAOMAN algorithm is provided. In the aspect of movement control, the acceleration of each robot is optimized, so that each robot can reach each target location within the time as short as possible. In the calculation unloading level, the edge calculation equipment around the robot is used for providing services for the robot, and the average completion time of the calculation task is shortened as much as possible by optimizing the unloading strategy of the calculation task of the robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of the intersection of multi-robot navigation and edge computing, and in particular to a multi-robot navigation and computation offloading method based on MARL. Background Art

[0002] When performing a task, an embodied intelligent robot typically uses its perception hardware to collect data from its environment and then processes this data to derive the next action instruction. Therefore, the smoothness of the robot's actions often depends on the efficiency of its data processing. However, due to cost constraints, robots are generally equipped with limited computing hardware and low computing power, resulting in high latency when processing data and generating the next action. In the field of mobile edge computing, computing tasks can be offloaded from devices with limited computing resources to external mobile edge computing servers that provide computing services. Through satellite-ground networks, the robot's computationally intensive tasks can be offloaded to the mobile edge computing servers. However, due to the robot's mobility, its positional relationship with the mobile edge server (MEC) is constantly changing, and the robot's computing offload environment is dynamically changing.

[0003] Furthermore, when solving complex real-world problems, the use of multiple robots is common. Robots often collaborate to perform specific tasks, which improves efficiency. The first step in collaborative tasks is task allocation, which forms the foundation of multi-robot coordinated control. Task allocation involves breaking down complex tasks into multiple subtasks and assigning them to different robots. An effective task allocation solution can maximize overall system efficiency and resource utilization.

[0004] Compared to traditional mobile computing devices, a notable feature of robot swarms is their collaborative nature. Robots often need to exchange large amounts of data with each other to complete tasks. In such collaborative tasks, cooperation among team members can significantly improve the overall performance of the robot team. However, data-intensive tasks often involve significant communication costs. As the size of the swarm continues to expand, the large amount of data exchange between robots can become a bottleneck for multi-robot applications, a problem that does not exist in traditional mobile computing devices (such as smartphones or the Internet of Things). In this case, simply offloading the robots' computing tasks to edge servers will result in significant communication delays, which is counterproductive.

[0005] At the same time, robots are subject to certain constraints when performing actions. For example, during movement, they must avoid collisions with obstacles and other robots. Furthermore, the robot's movement can influence the decision to offload its computational tasks. Specifically, in multi-robot navigation tasks, a robot's movement causes its position to change, while the positions of nearby edge devices that provide computing services are often fixed. This changes the conditions for offloading the robot's computational tasks. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to overcome the shortcomings of the existing technology and provide a multi-robot navigation and computation offloading method based on multi-agent reinforcement learning (MARL). The present invention aims to ensure the efficiency of multiple robots in completing a complex real-world task together while minimizing the average completion time of the robots' computational tasks.

[0007] The present invention adopts the following technical solutions to solve the above technical problems:

[0008] A multi-robot navigation and computation offloading method based on MARL proposed in the present invention includes:

[0009] Design a multi-robot navigation system supported by MEC-assisted satellite-ground network;

[0010] Based on a multi-robot navigation system, the joint optimization problem of multi-robot motion control and computation offloading is formalized.

[0011] Decompose the joint optimization problem into a mobility control subproblem and a computation offloading subproblem, and design the three elements of multi-agent reinforcement learning (MARL).

[0012] Based on MARL, a decentralized adaptive offloading and multi-agent navigation DAOMAN algorithm is proposed to solve the optimal strategy of mobile control and computation offloading.

[0013] As a further optimization scheme of the MARL-based multi-robot navigation and computation offloading method described in the present invention, a multi-robot navigation system supported by a MEC-assisted satellite-ground network is designed as follows:

[0014] The multi-robot navigation system includes N robots, N target locations that the robots need to reach, and N SA satellites and N BS base stations;

[0015] Assume that the robot's initial velocity and acceleration are both 0, the robot and the target location are located in different places, and no target location is assigned to each robot in advance. There are several obstacles in the environment, and the robot avoids collisions with obstacles or other robots during movement. The robot generates a computing task every time interval ΔT, which is processed by the robot's own computing device or offloaded to an edge device. Satellites and base stations have coverage areas, and the robot can only offload its computing tasks to the satellite or base station when it is within the coverage area of ​​the satellite or base station.

[0016] The multi-robot navigation system includes a target allocation model, a motion control model, and a computation offloading model.

[0017] As a further optimization scheme of the MARL-based multi-robot navigation and computation offloading method described in the present invention,

[0018] The target allocation model defines the following parameters: i is the number of the target location that the i-th robot wants to reach; T i is the number of time steps it takes for the i-th robot to reach the target location of the robot;

[0019] The robots cannot reach the same destination. There are the following restrictions:

[0020]

[0021] Among them, tar j is the number of the target location that the j-th robot wants to reach;

[0022] The i-th robot spends T to reach its target location i time steps, the distance between the robot and the target location satisfies the following relationship:

[0023]

[0024] Among them, pos i,t is the coordinate of the i-th robot at the t-th time step, is the coordinate of the target location that the i-th robot wants to reach, It represents the distance between the i-th robot and its target location at the t-th time step, ε is the judgment distance for the robot to reach the target location, and it is considered to have reached the target location if the distance is less than ε.

[0025] As a further optimization scheme of the MARL-based multi-robot navigation and computation offloading method described in the present invention,

[0026] The mobility control model defines the following parameters: (x i,t ,y i,t) represents the coordinates of the i-th robot at the t-th time step; v i,t represents the velocity of the i-th robot at the t-th time step, v i,t It is composed of the speed in the x-axis and y-axis directions respectively. is the velocity of the ith robot in the x-axis direction at the tth time step, is the velocity of the ith robot in the y-axis direction at the tth time step; a i,t represents the acceleration of the i-th robot at the t-th time step, a i,t It is composed of the acceleration in the x-axis and y-axis directions respectively. is the acceleration of the ith robot in the x-axis direction at the tth time step, is the acceleration of the ith robot in the y-axis direction at the tth time step; for the robot's velocity v i,t , v i,t The value range is [-V limit ,V limit ];clip(v i,t ,-V limit ,V limit ) means if v i,t Exceeding [-V limit ,V limit ], then it is clipped to [-V limit ,V limit ]Inside; T move It represents the average moving time of N robots each moving to a target location;

[0027]

[0028]

[0029] Among them, V limit is the maximum speed of the robot, is the velocity of the ith robot in the x-axis direction at the t+1th time step, is the velocity of the ith robot in the y-axis direction at the t+1th time step, x i,t+1 is the x-axis coordinate of the i-th robot at the t+1th time step, y i,t+1 is the y-axis coordinate of the i-th robot at the t+1-th time step, and ΔT is the time step.

[0030] As a further optimization scheme of the MARL-based multi-robot navigation and computation offloading method described in the present invention,

[0031] The computational unloading model defines the following parameters: D i,t Indicates the amount of data for the task; delayi,t represents the maximum delay that the task can tolerate; p represents the transmission power; h represents the channel gain; W represents the channel bandwidth; σ 2 represents the noise power generated by the complex white noise channel; r represents the wireless transmission rate; f is the computing power of the device; Indicates the transmission time of the task; is the computation time of the task; Represents the total time spent on the task locally, offloaded to the base station, and the satellite respectively; is the unloading time of the task; w is the equilibrium T i move and T i offload The weight factor of

[0032] Each variable is calculated using the following formula:

[0033]

[0034] Among them, T i move is the time it takes for the i-th robot to move to its target location; T i offload Calculate the total completion time of the task for the i-th robot during the movement process; decision i,t is the offloading decision of the computational task of the i-th robot at the t-th time step; when decision i,t =0, [decision i,t =0] is 1, otherwise it is 0; when decision i,t =1, [decision i,t =1] is 1, otherwise it is 0; when decision i,t =2, [decision i,t =2] is 1, otherwise it is 0; T mean It is the weighted average of all robot movement times and computational task completion times.

[0035] As a further optimization scheme of the MARL-based multi-robot navigation and computation offloading method described in the present invention,

[0036] Based on a multi-robot navigation system, the joint optimization problem of multi-robot motion control and computational task offloading is formalized. It is a multi-objective optimization problem based on POMDP, which aims to minimize the time it takes for the robots to reach their respective target locations while shortening the average completion time of the robots' own computational tasks. Specifically,

[0037]

[0038] Among them, A limit Indicates the acceleration limit of the robot, obst j Represents the coordinates of the jth obstacle, C4 restricts different robots from reaching the same target location; C6 ensures that the robot cannot collide with other robots; C7 ensures that the robot cannot collide with obstacles.

[0039] As a further optimization scheme of the MARL-based multi-robot navigation and computation offloading method described in the present invention,

[0040] Decompose the joint optimization problem into a mobility control subproblem and a computational task offloading subproblem, and design the three elements of reinforcement learning;

[0041] The three elements of reinforcement learning include the state space State(t), the action space Action(t), and the reward function Reward(t), where:

[0042] For the mobility control subproblem:

[0043] State space: State(t)={pos(t),v x (t),v y (t), dest(t), obst(t)}

[0044] Where, pos(t)={pos1(t),...,pos N (t)}, pos(t) represents the position set of N robots, pos i (t) represents the coordinates of the i-th robot at the t-th time step; v x (t) represents the set of velocities of N robots in the x-axis direction, is the velocity of the ith robot in the x-axis direction at the tth time step; v y (t) represents the set of velocities of N robots in the y-axis direction, is the velocity of the ith robot in the tth time step in the y-axis direction; dest(t) = {dest1(t),...,dest N (t)}, dest(t) represents the coordinate set of N target locations, dest i (t) is the coordinate of the i-th destination; N obst is the number of obstacles, obst(t) represents the coordinate set of all obstacles, obst i (t) represents the coordinates of the i-th obstacle;

[0045] Action space: Action(t) = {a x (t),ay (t)}

[0046] in, a x (t), a y (t) represents the acceleration of the robot in the x-axis and y-axis directions, is the acceleration of the ith robot in the x-axis direction at the tth time step, is the acceleration of the ith robot in the y-axis direction at the tth time step;

[0047] Reward function:

[0048] in, is the coordinate of the target location that the i-th robot wants to reach, represents the distance between the i-th robot and its target location; collision_penalty represents the penalty term for robot collision;

[0049] For the computation offloading subproblem:

[0050] State space: State(t) = {pos(t), D(t), f(t), sate(t), BS(t)}

[0051] Among them, pos(t) represents the position of N robots, D(t) represents the data size of the computing task at the current moment; f(t) represents the computing power of the robot itself; sate(t) represents the state information set of the satellite, sate i' (t) is the state of the i'th satellite at the tth time step, including coverage area, channel bandwidth, and computing power, where i' is in the range [1,N SA ]; BS(t) represents the status information of the base station. i” (t) is the state of the i'th base station at the tth time step, including coverage area, channel bandwidth, and computing power, where the range of i'' is [1,N BS ];

[0052] Action space: Action(t) = {0, 1, 2};

[0053] Wherein, Action(t)=0 means that the computing task is processed locally; Action(t)=1 means that the computing task is offloaded to the base station for processing; Action(t)=2 means that the computing task is offloaded to the satellite for processing;

[0054] Reward function:

[0055] As a further optimization scheme for the MARL-based multi-robot navigation and computation offloading method described in the present invention, a DAOMAN algorithm is proposed based on MARL to solve the optimal strategy for mobile control and computation offloading; it includes:

[0056] To address the joint optimization problem, we use a layered approach and, based on MARL, propose a decentralized adaptive offloading and multi-agent navigation DAOMAN algorithm to solve the optimal strategy for mobile control and computation offloading. The DAOMAN algorithm includes a multi-agent double-delayed deep deterministic policy gradient (MATD3) and a proximal policy optimization (PPO) algorithm.

[0057] First, for the mobility control subproblem, a multi-agent dual-delay deep deterministic policy gradient (MATD3) algorithm is used to solve the optimal mobility control policy. The algorithm consists of an actor network and two critic networks. The smallest Q value output by the two critic networks is used as the actual Q value. The actor network is updated using the policy gradient method, and the two critic networks are updated using the mean square error (MSE) loss function.

[0058] Then, for the computational offloading sub-problem, the proximal strategy is used to optimize the PPO algorithm to solve the optimal strategy for computational task offloading; in this sub-problem, each robot shares a PPO model.

[0059] A computer device includes a memory, a processor, and a computer program stored in the memory and capable of running on the processor, wherein the processor implements the steps of the above-mentioned MARL-based multi-robot navigation and computation offloading method when executing the computer program.

[0060] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the above-mentioned MARL-based multi-robot navigation and computation offloading method.

[0061] Compared with the prior art, the present invention adopts the above technical solution and has the following technical effects:

[0062] This paper formulates this problem as a multi-objective optimization problem by minimizing the time it takes for robots to reach their respective destinations while simultaneously shortening the time it takes for the robots to complete their own computational tasks. Based on MARL, a decentralized adaptive offloading and multi-agent navigation DAOMAN algorithm is proposed to jointly optimize goal allocation, multi-robot navigation, and computation offloading strategies. Through extensive comparative experiments, this paper demonstrates that compared to existing methods, this method is superior in minimizing the time it takes for robots to reach their respective destinations while simultaneously shortening the time it takes for the robots to complete their own computational tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 Schematic diagram of the MEC-assisted satellite-ground network multi-robot navigation system model;

[0064] Figure 2 This is the DAOMAN algorithm architecture diagram;

[0065] Figure 3 The reward curves for each round of training for the DAOMAN algorithm and the baseline algorithms MADDPG and MATD3;

[0066] Figure 4 Figure 2 shows the trajectory diagrams of the DAOMAN algorithm and the baseline algorithms MADDPG and MATD3 during the testing phase. (a) is the trajectory diagram of the MADDPG algorithm, (b) is the trajectory diagram of the MATD3 algorithm, and (c) is the trajectory diagram of the DAOMAN algorithm.

[0067] Figure 5 The calculation task completion time of the DAOMAN algorithm and the task local processing method under different data sizes. DETAILED DESCRIPTION

[0068] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0069] The MARL-based satellite-to-ground network multi-robot navigation and computation offloading method of the present invention comprises the following steps:

[0070] S1: Design a MEC-assisted satellite-ground network multi-robot navigation system.

[0071] like Figure 1 As shown, there are N robots in the system, N target locations that the robots need to reach, and N SA Satellites, N BS The system has the following settings: the robot's initial velocity and acceleration are both zero, the robots and the target locations are located in different locations, and no target location is assigned to each robot in advance; there are several obstacles in the environment, and the robot must avoid collisions with obstacles or other robots during movement; the robot generates a computational task every time interval ΔT, which can be processed by the robot's own computing equipment or offloaded to edge devices; satellites and base stations have coverage areas, and the robot can only offload its computing tasks to the satellite or base station when it is within the coverage area. The entire system model is divided into three parts: the target allocation model, the mobility control model, and the computation offloading model.

[0072] S2: Formalization of the optimization problem.

[0073] The problem is formulated as a multi-objective optimization problem based on POMDP, which aims to minimize the time it takes for the robots to reach their respective destinations while shortening the average completion time of the robots' own computational tasks. Specifically:

[0074]

[0075] Among them, A limit Indicates the acceleration limit of the robot, obst j Represents the coordinates of the jth obstacle, C4 restricts different robots from reaching the same target location; C6 ensures that the robot cannot collide with other robots; C7 ensures that the robot cannot collide with obstacles.

[0076] S3: Decompose the original problem into a mobility control subproblem and a computation offloading subproblem, and design the three elements of reinforcement learning, as follows:

[0077] For the mobility control subproblem:

[0078] State space: State(t)={pos(t),v x (t),v y (t), dest(t), obst(t)}

[0079] Where, pos(t)={pos1(t),...,pos N (t)}, pos(t) represents the position set of N robots, pos i (t) represents the coordinates of the i-th robot at the t-th time step; v x (t) represents the set of velocities of N robots in the x-axis direction, is the velocity of the ith robot in the x-axis direction at the tth time step; v y (t) represents the set of velocities of N robots in the y-axis direction, is the velocity of the ith robot in the tth time step in the y-axis direction; dest(t) = {dest1(t),...,dest N (t)}, dest(t) represents the coordinate set of N target locations, dest i (t) is the coordinate of the i-th destination; N obst is the number of obstacles, obst(t) represents the coordinate set of all obstacles, obst i (t) represents the coordinates of the i-th obstacle;

[0080] Action space: Action(t) = {a x (t),a y (t)}

[0081] in, a x (t), a y (t) represents the acceleration of the robot in the x-axis and y-axis directions, is the acceleration of the ith robot in the x-axis direction at the tth time step, is the acceleration of the ith robot in the y-axis direction at the tth time step;

[0082] Reward function:

[0083] in, is the coordinate of the target location that the i-th robot wants to reach, represents the distance between the i-th robot and its target location; collision_penalty represents the penalty term for robot collision;

[0084] For the computation offloading subproblem:

[0085] State space: State(t) = {pos(t), D(t), f(t), sate(t), BS(t)}

[0086] Among them, pos(t) represents the position of N robots, D(t) represents the data size of the computing task at the current moment; f(t) represents the computing power of the robot itself; sate(t) represents the state information set of the satellite, sate i' (t) is the state of the i'th satellite at the tth time step, including coverage area, channel bandwidth, and computing power, where i' is in the range [1,N SA ]; BS(t) represents the status information of the base station. i” (t) is the state of the i'th base station at the tth time step, including coverage area, channel bandwidth, and computing power, where the range of i'' is [1,N BS ];

[0087] Action space: Action(t) = {0, 1, 2}

[0088] Wherein, Action(t)=0 means that the computing task is processed locally; Action(t)=1 means that the computing task is offloaded to the base station for processing; Action(t)=2 means that the computing task is offloaded to the satellite for processing;

[0089] Reward function: Reward(t) = -T i,t offload .

[0090] S4: Use the DAOMAN algorithm to solve the optimal strategy.

[0091] like Figure 2 As shown in Figure 3, the DAOMAN algorithm is divided into two stages: the first stage is the target allocation stage, and the second stage is the mobility control and computation offloading stage.

[0092] In the first stage, a different target location is assigned to each robot using the Hungarian algorithm based on its straight-line distance to the target location.

[0093] In the second stage, for the above-mentioned mobile control subproblem, the MATD3 algorithm is used, which includes a weight of The Actor network and the two weights are and By using the smaller of the Q values ​​output by the two critic networks as the actual Q value, the problem of Q value overestimation in a single critic network is alleviated.

[0094] Update the Actor network through policy gradient, the formula is as follows:

[0095]

[0096] The two critic networks are updated using the mean square error (MSE) loss function, as follows:

[0097]

[0098] For the computational offloading subproblem mentioned above, the PPO algorithm is used, with the weights of the Actor and Critic networks being μ and θ, respectively. It should be noted that in this subproblem, all robots share a PPO model. The Clip operation is introduced to prevent excessive updates from causing training instability. The objective function for the policy update is:

[0099]

[0100] Figure 3 The reward curves for each round of the DAOMAN algorithm and the baseline algorithms MADDPG and MATD3 during training. It can be seen from the figure that the proposed DAOMAN can obtain higher rewards than the MADDPG and MATD3 algorithms.

[0101] Figure 4 This is the trajectory diagram of the DAOMAN algorithm and the baseline algorithms MADDPG and MATD3 during the testing phase; Figure 4 (a) and Figure 4 In (b), not all robots reach their respective target locations, and the robots collide. Figure 4 This situation does not occur in (c). This shows that the proposed DAOMAN algorithm is effective in the actual testing phase.

[0102] Figure 5 The computational task completion time of the DAOMAN algorithm and the task local processing method under different data sizes is shown in Figure 2. It can be seen that the proposed DAOMAN algorithm can achieve efficient computation offloading decisions, allowing the robot's computational tasks to be completed as quickly as possible.

[0103] An embodiment of the present invention also provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, the steps of the above-mentioned MARL-based multi-robot navigation and computation offloading method are implemented.

[0104] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the computer program implements the steps of the above-mentioned MARL-based multi-robot navigation and computation offloading method.

[0105] It will be understood by those skilled in the art that the embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present invention may be implemented in various computer languages, for example, the object-oriented programming language Java and the interpreted scripting language JavaScript.

[0106] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1A device that provides the functions specified in a block or multiple blocks.

[0107] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0108] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0109] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0110] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A multi-robot navigation and computation offloading method based on MARL, characterized in that: include: Design a multi-robot navigation system supported by MEC-assisted satellite-ground network; Based on a multi-robot navigation system, the joint optimization problem of multi-robot motion control and computation offloading is formalized. Decompose the joint optimization problem into a mobility control subproblem and a computation offloading subproblem, and design the three elements of multi-agent reinforcement learning (MARL). Based on MARL, a decentralized adaptive offloading and multi-agent navigation DAOMAN algorithm is proposed to solve the optimal strategy of mobile control and computation offloading.

2. A MARL-based multi-robot navigation and computation offloading method according to claim 1, characterized in that: Design a multi-robot navigation system supported by MEC-assisted satellite-ground network, as follows: The multi-robot navigation system includes N robots, N target locations that the robots need to reach, and N SA satellites and N BS base stations; Assume that the robot's initial velocity and acceleration are both 0, the robot and the target location are located in different places, and no target location is assigned to each robot in advance. There are several obstacles in the environment, and the robot avoids collisions with obstacles or other robots during movement. The robot generates a computation task every time interval ΔT, which is processed by the robot's own computing device or offloaded to an edge device for processing. Satellites and base stations have coverage areas. Only when the robot is within the coverage area of ​​a satellite or base station can it offload its computing tasks to the satellite or base station. The multi-robot navigation system includes a target allocation model, a motion control model, and a computation offloading model.

3. The MARL-based multi-robot navigation and computation offloading method according to claim 2, characterized in that: The target allocation model defines the following parameters: i is the number of the target location that the i-th robot wants to reach; T i is the number of time steps it takes for the i-th robot to reach the target location of the robot; The robots cannot reach the same destination. There are the following restrictions: Among them, tar j is the number of the target location that the j-th robot wants to reach; The i-th robot spends T to reach its target location i time steps, the distance between the robot and the target location satisfies the following relationship: Among them, pos i,t is the coordinate of the i-th robot at the t-th time step, is the coordinate of the target location that the i-th robot wants to reach, It represents the distance between the i-th robot and its target location at the t-th time step, ε is the judgment distance for the robot to reach the target location, and it is considered to have reached the target location if the distance is less than ε.

4. The MARL-based multi-robot navigation and computation offloading method according to claim 2, characterized in that: The mobility control model defines the following parameters: (x i,t ,y i,t ) represents the coordinates of the i-th robot at the t-th time step; v i,t represents the velocity of the i-th robot at the t-th time step, v i,t It is composed of the speed in the x-axis and y-axis directions respectively. is the velocity of the ith robot in the x-axis direction at the tth time step, is the velocity of the ith robot in the y-axis direction at the tth time step; a i,t represents the acceleration of the i-th robot at the t-th time step, a i,t It is composed of the acceleration in the x-axis and y-axis directions respectively. is the acceleration of the ith robot in the x-axis direction at the tth time step, is the acceleration of the ith robot in the y-axis direction at the tth time step; for the robot's velocity v i,t , v i,t The value range is [-V limit ,V limit ];clip(v i,t ,-V limit ,V limit ) means if v i,t Beyond [-V limit ,V limit ], then it is clipped to [-V limit ,V limit ]Inside; T move It represents the average moving time of N robots each moving to a target location; Among them, V limit is the maximum speed of the robot, is the velocity of the ith robot in the x-axis direction at the t+1th time step, is the velocity of the ith robot in the y-axis direction at the t+1th time step, x i,t+1 is the x-axis coordinate of the i-th robot at the t+1th time step, y i,t+1 is the y-axis coordinate of the i-th robot at the t+1-th time step, and ΔT is the time step.

5. The MARL-based multi-robot navigation and computation offloading method according to claim 2, characterized in that: The computational unloading model defines the following parameters: D i,t Indicates the amount of data for the task; delay i,t represents the maximum delay that the task can tolerate; p represents the transmission power; h represents the channel gain; W represents the channel bandwidth; σ 2 represents the noise power generated by the complex white noise channel; r represents the wireless transmission rate; f is the computing power of the device; Indicates the transmission time of the task; is the computation time of the task; Represents the total time spent on the task locally, offloaded to the base station, and the satellite respectively; is the unloading time of the task; w is the equilibrium T i move and T i offload The weight factor of Each variable is calculated using the following formula: Among them, T i move is the time it takes for the i-th robot to move to its target location; T i offload Calculate the total completion time of the task for the i-th robot during the movement process; decision i,t is the offloading decision of the computational task of the i-th robot at the t-th time step; when decision i,t =0, [decision i,t =0] is 1, otherwise it is 0; when decision i,t =1, [decision i,t =1] is 1, otherwise it is 0; when decision i,t =2, [decision i,t =2] is 1, otherwise it is 0; T mean It is the weighted average of all robot movement times and computational task completion times.

6. The MARL-based multi-robot navigation and computation offloading method according to claim 2, characterized in that: Based on a multi-robot navigation system, the joint optimization problem of multi-robot motion control and computational task offloading is formalized. It is a multi-objective optimization problem based on POMDP, which aims to minimize the time it takes for the robots to reach their respective target locations while shortening the average completion time of the robots' own computational tasks. Specifically, s.t.C1:tar i ∈{1,2,...,N} C2:decision i,t ∈{0,1,2} Among them, A limit Indicates the acceleration limit of the robot, obst j Represents the coordinates of the jth obstacle, C4 restricts different robots from reaching the same target location; C6 ensures that the robot cannot collide with other robots; C7 ensures that the robot cannot collide with obstacles.

7. The MARL-based multi-robot navigation and computation offloading method according to claim 6, characterized in that: Decompose the joint optimization problem into a mobility control subproblem and a computational task offloading subproblem, and design the three elements of reinforcement learning; The three elements of reinforcement learning include the state space State(t), the action space Action(t), and the reward function Reward(t), where: For the mobility control subproblem: State space: State(t)={pos(t),v x (t),v y (t), dest(t), obst(t)} Where, pos(t)={pos1(t),...,pos N (t)}, pos(t) represents the position set of N robots, pos i (t) represents the coordinates of the i-th robot at the t-th time step; v x (t) represents the set of velocities of N robots in the x-axis direction, is the velocity of the ith robot in the x-axis direction at the tth time step; v y (t) represents the set of velocities of N robots in the y-axis direction, is the velocity of the ith robot in the tth time step in the y-axis direction; dest(t) = {dest1(t),...,dest N (t)}, dest(t) represents the coordinate set of N target locations, dest i (t) is the coordinate of the i-th destination; N obst is the number of obstacles, obst(t) represents the coordinate set of all obstacles, obst i (t) represents the coordinates of the i-th obstacle; Action space: Action(t) = {a x (t),a y (t)} in, a x (t), a y (t) represents the acceleration of the robot in the x-axis and y-axis directions, is the acceleration of the ith robot in the x-axis direction at the tth time step, is the acceleration of the ith robot in the y-axis direction at the tth time step; Reward function: in, is the coordinate of the target location that the i-th robot wants to reach, represents the distance between the i-th robot and its target location; collision_penalty represents the penalty term for robot collision; For the computation offloading subproblem: State space: State(t) = {pos(t), D(t), f(t), sate(t), BS(t)} Among them, pos(t) represents the position of N robots, D(t) represents the data size of the computing task at the current moment; f(t) represents the computing power of the robot itself; sate(t) represents the state information set of the satellite, sate i' (t) is the state of the i'th satellite at the tth time step, including coverage area, channel bandwidth, and computing power, where i' is in the range [1,N SA ]; BS(t) represents the status information of the base station. i” (t) is the state of the i'th base station at the tth time step, including coverage area, channel bandwidth, and computing power, where the range of i'' is [1,N BS ]; Action space: Action(t) = {0, 1, 2}; Wherein, Action(t)=0 means that the computing task is processed locally; Action(t)=1 means that the computing task is offloaded to the base station for processing; Action(t)=2 means that the computing task is offloaded to the satellite for processing; Reward function:

8. The MARL-based multi-robot navigation and computation offloading method according to claim 7, characterized in that: Based on MARL, the DAOMAN algorithm is proposed to solve the optimal strategy for mobility control and computing task offloading; including: To address the joint optimization problem, we use a layered approach and, based on MARL, propose a decentralized adaptive offloading and multi-agent navigation DAOMAN algorithm to solve the optimal strategy for mobile control and computation offloading. The DAOMAN algorithm includes a multi-agent double-delayed deep deterministic policy gradient (MATD3) and a proximal policy optimization (PPO) algorithm. First, for the mobility control subproblem, a multi-agent dual-delay deep deterministic policy gradient (MATD3) algorithm is used to solve the optimal mobility control policy. The algorithm consists of an actor network and two critic networks. The smallest Q value output by the two critic networks is used as the actual Q value. The actor network is updated using the policy gradient method, and the two critic networks are updated using the mean square error (MSE) loss function. Then, for the computational offloading sub-problem, the proximal strategy is used to optimize the PPO algorithm to solve the optimal strategy for computational task offloading; in this sub-problem, each robot shares a PPO model.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, characterized in that: When the processor executes the computer program, the steps of the MARL-based multi-robot navigation and computation offloading method are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the MARL-based multi-robot navigation and computation offloading method are implemented.