Multi-Machine Cooperative Task Scheduling Method for Edge Computing Scenarios in Aerospace Networks

By adopting a multi-machine collaborative task scheduling method in the edge computing scenario of aerospace networks, and using the task collection model and task delay model, the problem of unbalanced load of drones is solved, and efficient task processing and excellent user experience are achieved.

CN119759587BActive Publication Date: 2025-06-24BEIJING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510259114.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-24
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

In the edge computing scenario of aerospace networks, the coverage and computing resources of a single drone are not sufficient to meet the needs of large-area or high-density users, resulting in load imbalance, wasting processing power and increasing data transmission delay.

Method used

The multi-machine collaborative task scheduling method is adopted to minimize the user-task completion delay model through pre-constructed task collection models and task delay models, match target drones in the multi-drone cluster, and schedule other drones or satellites to perform task collaborative calculations, so as to minimize the user-task completion delay.

Benefits of technology

It effectively solves the problem of uneven load distribution of drones, improves task processing efficiency and overall system performance, and ensures that a high level of service quality and user experience are maintained in complex dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119759587B_ABST
    Figure CN119759587B_ABST
Patent Text Reader

Abstract

The present application provides a multi-aircraft collaborative task scheduling method for the edge computing scenario of the space-air network, which relates to the technical field of edge computing of the space-air network. It includes that in response to the user tasks initiated by ground users, the unmanned aerial vehicle (UAV) collects user tasks based on a pre-constructed task collection model (modeled based on Markov decision process and multi-agent reinforcement learning) to maximize the task collection rate in the current time slot. When the user tasks are offloaded to the target UAV, the task scheduling is optimized based on the deep deterministic policy gradient algorithm, and scheduled to other UAVs in the multi-UAV cluster or the satellite communicating with the multi-UAV cluster, so that other UAVs and / or satellites in the multi-UAV cluster perform task collaborative computing with the target UAV, realizing the minimization of the completion delay of user tasks. The present application not only solves the problem of the difference in user distribution density, but also overcomes the problem of UAV resource limitation, improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of edge computing for space-air networks, and in particular, to a multi-aircraft collaborative task scheduling method for space-air network edge computing scenarios. Background Art

[0002] With the development of space-air network technology, edge computing, as a new computing mode, has been widely applied in various scenarios. As an important mobile node, unmanned aerial vehicles (UAVs) play an important role in various application scenarios, especially in edge computing scenarios. In space-air network edge computing scenarios, UAVs, as important computing nodes, undertake a large number of data processing tasks.

[0003] In the prior art, UAVs can collect and process data through their own sensors and communication modules, and at the same time provide computing services to other devices. However, the capabilities of a single UAV are limited, and its coverage and computing resources are insufficient to meet the needs of large-area or high-density users.

[0004] In addition, due to the uneven distribution of users, it may lead to overloading of UAVs in some areas, while UAVs in other areas are in a light-load state. For example, there may be a large number of users gathering in some areas, resulting in overloading of UAVs in these areas, while UAVs in other areas may be in a low-load state. This load imbalance will not only cause waste of the processing capabilities of some UAVs, but also increase the data transmission delay, thus affecting the overall network service quality and user experience. Summary of the Invention

[0005] The purpose of this application is to provide a multi-aircraft collaborative task scheduling method for space-air network edge computing scenarios to alleviate the above technical problems existing in the prior art.

[0006] In a first aspect, the present invention provides a multi-aircraft collaborative task scheduling method for space-air network edge computing scenarios, the method comprising:

[0007] In response to a user task initiated by a ground user, matching a target UAV in a multi-UAV cluster, the target UAV collecting the user task based on a pre-constructed task collection model; wherein, the pre-constructed task collection model is a model that maximizes the task collection rate based on Markov decision process and multi-agent reinforcement learning modeling;

[0008] When the user task is offloaded to the target UAV, other UAVs in the multi-UAV cluster or satellites communicating with the multi-UAV cluster are scheduled through a pre-constructed task delay model, so that other UAVs and / or satellites in the multi-UAV cluster perform task collaborative computing with the target UAV to minimize the completion delay of the user task; wherein, the task delay model is a model constructed by optimizing the task scheduling strategy based on the deep deterministic policy gradient algorithm.

[0009] In an alternative embodiment, the steps for constructing the task collection model include:

[0010] Construct an association matrix between the user tasks initiated by the ground user and the UAVs in the multi-UAV cluster; the association matrix is used to represent the association relationship between the user tasks initiated by the ground user and the UAVs in the current time slot.

[0011] Determine the task collection rate of the current time slot according to the association matrix, the first task data volume of the user task published, the second data volume of the UAV collecting the user task, the first location information of the ground user, and the second location information of the UAV.

[0012] During the flight process of the UAV for task collection, optimize the flight trajectory of the UAV based on the flight constraint conditions and the Markov decision process, and determine the task collection model of the UAV based on the maximum task collection rate of the UAV corresponding to the optimized flight trajectory.

[0013] In an alternative embodiment, the task collection rate of the current time slot is determined according to the association matrix, the first task data volume of the user task published, the second data volume of the UAV collecting the user task, the first location information of the ground user, and the second location information of the UAV, and is represented by the following formula:

[0014] =

[0015] wherein, is the task collection rate, N is the number of UAVs, M is the number of ground users, is the task volume generated by user i at time t, is the task volume generated by user j at time t, is the association relationship between UAV i and user j , and t is the time slot.

[0016] In an alternative embodiment, during the flight of the UAV for task collection, the flight trajectory of the UAV is optimized based on flight constraint conditions and the Markov decision process, and the task collection model of the UAV is determined based on the maximum task collection rate corresponding to the optimized flight trajectory, including:

[0017] An agent is deployed for each UAV. In each time slot, based on the flight constraint conditions, the agent interacts with the environment, takes a first agent action according to the first agent state, and the agent will obtain a first immediate reward after executing the first agent action; wherein, the first immediate reward is related to the task collection rate; the flight constraint conditions include the flight speed constraint condition of the UAV, the flight angle constraint condition, the constraint conditions of the moving ranges of the UAV and the ground users, and the collision distance constraint condition between UAVs.

[0018] When the agent interacts with the environment, a policy network and a value network are initialized for each agent, the environment is interacted with using a preset policy, and the corresponding agent state, agent action, and immediate reward are obtained; the advantage value of each time step is calculated through generalized advantage estimation; the objective function of the multi-agent reinforcement learning algorithm is maximized, and the value function of the multi-agent reinforcement learning algorithm is minimized, and this step is repeated until convergence to determine the task collection model of the UAV.

[0019] In an alternative embodiment, when the user task is offloaded to the target UAV, other UAVs in the multi-UAV cluster or a satellite communicating with the multi-UAV cluster are scheduled through a pre-constructed task delay model, so that other UAVs and / or the satellite in the multi-UAV cluster perform task collaborative computing with the target UAV, including:

[0020] When the user task is offloaded to the target UAV, an initial task delay model is constructed; wherein, the initial task delay model is a model related to communication delay, computing delay, and transmission delay.

[0021] The satellite side is determined as the central scheduling agent. The satellite side obtains the task information and computing queue status of all UAVs, and optimizes the task scheduling policy based on the deep deterministic policy gradient algorithm to obtain the target task delay model.

[0022] Other UAVs in the multi-UAV cluster or a satellite communicating with the multi-UAV cluster are scheduled through the target task delay model, so that other UAVs and / or the satellite in the multi-UAV cluster perform task collaborative computing with the target UAV.

[0023] In an alternative embodiment, the satellite side is determined as the central scheduling agent. The satellite side obtains the task information and the computing queue status of all UAVs, and optimizes the task scheduling policy based on the deep deterministic policy gradient algorithm to obtain a target task delay model, including:

[0024] Initialize the Actor network and the Critic network of the deep deterministic policy gradient algorithm, and initialize the experience replay pool;

[0025] Determine the satellite side as the central scheduling agent. For each user task, select the second agent action corresponding to the central scheduling agent according to the current policy. After executing the second agent action, obtain the second immediate reward and the next second agent state, and store them in the experience replay pool;

[0026] Randomly sample a batch of samples from the experience replay pool, calculate the loss through the loss function, and update the parameters of the policy network and the evaluation network;

[0027] Determine the continuous action policy corresponding to minimizing the task completion delay through the above steps to determine the target task delay model.

[0028] In an alternative embodiment, matching the target UAV in the multi-UAV cluster includes:

[0029] Match the UAVs within the flight coverage range corresponding to the location information in the multi-UAV cluster based on the location information of the ground user, and determine the UAVs as the target UAVs.

[0030] In an alternative embodiment, the method further includes:

[0031] When the user task corresponds to multiple UAVs that can collect user tasks, match the target UAVs in the multi-UAV cluster based on the load of the UAVs.

[0032] In a second aspect, the present invention provides a multi-aircraft collaborative task scheduling system for an airspace network edge computing scenario, which is used to implement the multi-aircraft collaborative task scheduling method for the airspace network edge computing scenario described in the foregoing embodiments. The system includes a satellite side, a multi-UAV cluster, and multiple ground users, where the multi-UAV cluster includes multiple UAVs.

[0033] In a third aspect, the present invention provides a UAV, which includes a processor and a memory. The memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the multi-aircraft collaborative task scheduling method for the airspace network edge computing scenario described in any one of the foregoing embodiments.

[0034] The beneficial effects of the multi - drone collaborative task scheduling method for the edge - computing scenario of the space - air network provided by this application are as follows:

[0035] By modeling the operation scenario as a Markov decision process (MDP) and using the model - based multi - agent proximal policy optimization (MAPPO) algorithm to perform the trajectory planning of the target drone, the maximization of task acquisition efficiency is ensured; the deep deterministic policy gradient (DDPG) algorithm is adopted to promote the cooperation between drones, which can make efficient scheduling decisions for user requests, effectively reduce the average task response time, and thus effectively improve the problem of uneven load distribution among drones. This method not only solves the challenges brought by the difference in user distribution density but also overcomes the problem of drone resource limitations, achieving a significant improvement in task - processing efficiency and the overall performance of the system. Furthermore, it ensures that a high level of service quality can be maintained even in complex dynamic environments, enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] To more clearly illustrate the specific implementation manners of this application or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific implementation manners or the prior art. Obviously, the drawings in the following description are some implementation manners of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0037] Figure 1 It is a flowchart of a multi - drone collaborative task scheduling method for the edge - computing scenario of the space - air network provided by an embodiment of this application;

[0038] Figure 2 It is a schematic diagram of the edge - computing scenario of the space - air network provided by an embodiment of this application;

[0039] Figure 3 It is a flowchart of another multi - drone collaborative task scheduling method for the edge - computing scenario of the space - air network provided by an embodiment of this application;

[0040] Figure 4 It is a structural logic diagram of a drone provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. Generally, the components of the embodiments of this application described and illustrated in the drawings here can be arranged and designed in various different configurations.

[0042] Accordingly, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the claimed present application, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.

[0043] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0044] The embodiments of the present application provide a multi-aircraft collaborative task scheduling method for the edge computing scenario of the space-air network. Refer to Figure 1 as shown, the method mainly includes the following steps:

[0045] Step S110, in response to a user task initiated by a ground user, match the target unmanned aerial vehicle (UAV) in the multi-UAV cluster, and the target UAV collects the user task based on a pre-constructed task collection model.

[0046] The above-mentioned pre-constructed task collection model is a model that maximizes the task collection rate based on the Markov decision process and multi-agent reinforcement learning. This model is a task collection model with the goal of maximizing the task collection rate. The embodiments of the present application represent the model optimization problem as a Markov decision process (MDP), including the definitions of state, action, and reward. In specific implementation, the multi-agent reinforcement learning algorithm MAPPO (Multi-Agent Proximal Policy Optimization) is used to optimize the action selection of the UAVs. The MAPPO algorithm includes steps such as the initialization of the policy network and value network, trajectory sampling, advantage estimation calculation, policy update, and value function update. Through this method, the algorithm can gradually optimize the action strategy of each UAV and maximize the task collection rate on the premise of ensuring safety.

[0047] In one implementation, when a ground user initiates a user task through a specific interface or device (such as a smartphone application, web interface, etc.), this request will first be received by the system. Next, the system parses the request to determine the specific requirements of the user task, which may include key information such as the target location of the task, the type of data to be collected, time limit, priority, etc. The system maintains a multi-UAV cluster, and the UAVs in the multi-UAV cluster are distributed at different locations and may also have different capabilities and characteristics. According to the requirements of the user task, the system needs to select a suitable UAV (i.e., the target UAV) from the cluster to perform the task collection.

[0048] The above-mentioned target UAV in the multi-UAV cluster can match the UAV within the flight coverage corresponding to the location information in the multi-UAV cluster based on the location information of the ground user, and determine the UAV as the target UAV. When the user task corresponds to multiple UAVs that can collect the user task, the target UAV in the multi-UAV cluster can also be matched based on the load of the UAV. For example, the UAV with less load is preferentially selected for collection.

[0049] Step S120, when the user task is offloaded to the target UAV, schedule other UAVs in the multi-UAV cluster or the satellite communicating with the multi-UAV cluster through a pre-constructed task delay model, so that other UAVs and / or satellites in the multi-UAV cluster perform task collaborative computing with the target UAV to minimize the completion delay of the user task.

[0050] Among them, the task delay model is a model constructed by optimizing the task scheduling strategy based on the deep deterministic policy gradient algorithm. The task delay model is a delay model including communication delay, computing delay and queuing delay, and this task delay model is determined with the goal of minimizing the average completion delay of the computing task. In one implementation, the Markov decision process (MDP) is used to model the task scheduling problem, and the satellite side is defined as the central scheduling agent to achieve unified decision-making. To solve this problem, the deep deterministic policy gradient (DDPG) algorithm is used to optimize the task scheduling strategy. The optimization steps of the DDPG algorithm include network initialization, exploration and experience collection, network update and target network update. Through this method, the load balancing problem in the collaborative computing of the UAV cluster and the satellite can be effectively solved.

[0051] The multi-UAV collaborative task scheduling method provided by the embodiment of the present application for the airspace network edge computing scenario models the operation scenario as a Markov decision process (MDP), and uses the model-based multi-agent proximal policy optimization (MAPPO) algorithm to perform the trajectory planning of the target UAV, thereby ensuring the maximization of the task acquisition rate; the deep deterministic policy gradient (DDPG) algorithm is used to promote the cooperation between UAVs, and can make efficient scheduling decisions for user requests, effectively reducing the average task response time, thereby effectively improving the problem of uneven UAV load distribution. This method not only solves the challenges brought by the difference in user distribution density, but also overcomes the problem of UAV resource limitations, realizes a significant improvement in task processing efficiency and the overall performance of the system, and further ensures that a high level of service quality can be maintained even in a complex dynamic environment, improving the user experience.

[0052] For ease of understanding, the multi-UAV collaborative task scheduling method provided by the embodiment of the present application for the airspace network edge computing scenario will be described in detail below.

[0053] First, this application is applied to the edge computing scenario of the space-air network. Refer to Figure 2 As shown, in this scenario, it includes a satellite node, a number of UAV nodes, and a number of user nodes. The total simulation duration of the entire system is T, which is evenly divided into K time slots. Task collection, UAV movement decision-making, and task scheduling are all carried out in the time slots. Each time slot, users generate tasks with a certain probability.

[0054] For the task initiated by the ground user m, it can be represented by a triple as follows, represents the size of the task data volume, represents the computing resources required by the task, represents the maximum tolerable delay of the task.

[0055] The ground user can only move within the area with length X and width Y. The position coordinates of the ground user are represented as . At the end of each time slot, the user moves randomly. The movement model is as follows:

[0056]

[0057] where is the maximum movement speed of the user, , represents the actual speed of the user, is the time slot size, is the movement angle of the user.

[0058] The UAV flies over this area, and its height remains constant at H. Its position is represented as . Considering that the coverage range of a single UAV is limited, the coverage radius of the UAV in this application is R. During the task collection process, the UAV can only collect the user tasks within its coverage range. For the case where a single user is within the coverage range of multiple UAVs, the UAV with less load is preferentially selected for collection.

[0059] In one implementation manner, the steps for constructing the task collection model may include the following steps 1-1 to step 1-3:

[0060] Step 1-1, construct an association matrix between the user tasks initiated by the ground users and the UAVs in the multi-UAV cluster; the association matrix is used to characterize the association relationship between the user tasks initiated by the ground users and the UAVs in the current time slot.

[0061] The association relationship between the user task m initiated by the ground user and the UAV n can be characterized by an association matrix as follows:

[0062] .

[0063] Explanation of each item in the association matrix: N represents the number of UAVs; M represents the number of user tasks initiated by ground users; t represents the time point; the element o n,m (t) represents the association relationship between UAV n and user task m initiated by ground users at time point t.

[0064] Step 1-2: Determine the task collection rate of the current time slot according to the association matrix, the first task data volume of the published user task, the second data volume collected by the UAV for the user task, the first location information of the ground user, and the second location information of the UAV. Considering the limited coverage range of a single UAV, multiple UAVs are combined for task collection and preliminary load balancing is achieved. According to the above association matrix the association relationship between the UAV and the user in the current time slot can be obtained, and the tasks collected by UAV n are Therefore, the modeling of the task collection rate in the current time slot can be expressed by the following formula:

[0065] =

[0066] where is the task collection rate, N is the number of UAVs, M is the number of ground users, is the task volume generated by user i at time t, is the task volume generated by user j at time t, is the association relationship between UAV i and user j and t is the time slot.

[0067] Based on the above model, an optimization problem can be established to maximize the task collection rate, expressed as:

[0068]

[0069] Step 1-3: During the flight of the UAV for task collection, optimize the flight trajectory of the UAV based on the flight constraint conditions and the Markov decision process, and determine the task collection model of the UAV based on the maximum task collection rate corresponding to the optimized flight trajectory. Specifically, it can further include the following steps 1-3.1 and step 1-3.2:

[0070] Step 1-3.1: Deploy corresponding agents for each UAV. In each time slot, based on the flight constraint conditions, enable the agent to interact with the environment, take the first agent action according to the first agent state, and the agent will obtain the first immediate reward after executing the first agent action; wherein, the first immediate reward is related to the task collection rate.

[0071] Model the Markov decision process (MDP) according to the optimization problem. In this application, each UAV deploys an agent to optimize the flight trajectory of the UAV and maximize the task collection rate. It mainly includes three parts: state, action, and reward. In each time slot, the agent interacts with the environment, takes action a according to its own state s, makes the environment enter a new state, and obtains an immediate reward r. Specifically as follows:

[0072] The state of each agent includes its own position information, the position information of all users, and the task data size of all users, which is defined as:

[0073]

[0074] Where is the position of agent i, is the set of positions of all users, is the set of task data sizes of all users.

[0075] The action of each agent includes the flight speed and flight angle of the UAV, which is defined as:

[0076]

[0077] Where is the flight speed parameter, is the flight angle, and the movement model of the UAV is as follows:

[0078]

[0079] After the agent executes a specific action, it will obtain environmental feedback, and this feedback is the reward. In this application, the reward of the agent is defined as:

[0080]

[0081] Where is the total task collection rate, represents the penalty set for the UAV flying beyond the motion range to ensure that the UAV flies within the motion range. represents the penalty set when the distance between UAVs is less than the minimum safe distance to ensure that the UAVs avoid collisions. 、 and is a weight coefficient used to balance the importance of different objectives.

[0082] The above flight constraint conditions include the flight speed constraint condition of the UAV, the flight angle constraint condition, the constraint conditions of the moving ranges of the UAV and the ground user, and the collision distance constraint condition between UAVs. Specifically,

[0083]

[0084] Among them C 1 and C 2 are the constraints on the flight speed and flight angle of the UAV, C 3~ C 6 are the constraints on the moving ranges of the UAV and the user, C 7 is the constraint on the collision distance between UAVs.

[0085] Step 1-3.2, when the agent conducts environmental interaction, initialize the policy network and value network for each agent, conduct environmental interaction using the preset policy, and obtain the corresponding agent state, agent action, and immediate reward; calculate the advantage value at each time step through generalized advantage estimation; maximize the objective function of the multi-agent reinforcement learning algorithm, and, minimize the value function of the multi-agent reinforcement learning algorithm, and repeat this step until convergence to determine the task collection model of the UAV.

[0086] To optimize the action selection of the UAV, this application adopts the multi-agent reinforcement learning algorithm MAPPO (Multi-Agent Proximal Policy Optimization). MAPPO is a multi-agent extended version of the PPO algorithm and can handle scenarios where multiple agents learn simultaneously.

[0087] The MAPPO algorithm optimizes the actions of the UAV through the following steps:

[0088] 1. Policy network and value network initialization: Initialize the policy network and value network .

[0089] 2. Sampling trajectory: Use the current policy to interact with the environment and collect data such as states, actions, and rewards.

[0090] 3. Calculate advantage estimation: Calculate the advantage value at each time step using generalized advantage estimation (GAE) .

[0091] 4. Policy update: Maximize the objective function , while ensuring that the difference between the new and old policies is not too large:

[0092]

[0093] wherein is the clipping parameter, is the old policy before parameter update, is the policy entropy, is the hyperparameter of the entropy.

[0094] 5. Value function update: Minimize the loss function of the value network:

[0095]

[0096] 6. Repeat steps 2 - 5 until convergence or reaching the predetermined number of iterations.

[0097] In this way, the MAPPO algorithm can gradually optimize the action policy of each UAV to maximize the task collection rate while ensuring safety. The algorithm will balance the flight speed, angle selection, and cooperation with other UAVs to achieve the optimization of the overall performance.

[0098] Furthermore, when the user task is offloaded to the target UAV, other UAVs in the multi - UAV cluster or satellites communicating with the multi - UAV cluster are scheduled through the pre - constructed task delay model to perform task collaborative computing with the target UAV. In specific implementation, it may include the following steps 2 - 1 to step 2 - 3:

[0099] Step 2 - 1, when the user task is offloaded to the target UAV, construct an initial task delay model; wherein, the initial task delay model is a model related to communication delay, computing delay, and transmission delay.

[0100] During the task completion process, it mainly includes communication delay, computing delay, and transmission delay.

[0101] I. Regarding the communication delay, it mainly includes offloading delay and scheduling delay:

[0102] 1) The offloading delay is the time required for task m to be offloaded to UAV n:

[0103]

[0104] wherein is the transmission rate between UAV n and user m, defined as:

[0105]

[0106] represents the UAV 's channel bandwidth, is the maximum transmission power of the user equipment, is the noise power, is the channel gain between the drone n and the user equipment m.

[0107] 2) The scheduling delay is the maximum transmission delay of task m during the task scheduling process.

[0108]

[0109] where j is the drone number to which task m is offloaded, is the offloading ratio of drone j scheduled to drone n, is the task offloading ratio of drone j scheduled to the satellite. and are the transmission rates between drone n and drone j and between drone j and the satellite respectively. is defined as:

[0110]

[0111] where, represents the channel bandwidth, is the maximum transmission power of the drone. represents the channel gain between drone n and drone j.

[0112] The transmission rate between the drone and the satellite is defined as follows:

[0113]

[0114] where, represents the channel bandwidth, is the path loss of the communication channel between the drone and the satellite.

[0115] Therefore, the communication delay is .

[0116] II. For task m, its computation delay is defined as follows

[0117]

[0118] where represents the number of cycles required for task m to compute one unit of task data, is the processing capacity of the drone, is the processing capacity of the satellite.

[0119] Meanwhile, considering the limited processing capacity of the drones, a computing queue is maintained at the drone side for task processing. After a task is scheduled to other drones according to the offloading ratio, it will be stored in the computing queue of the drone according to the arrival time. Since the computing and storage capacities of the satellite are much larger than those of the drones, no computing queue is maintained at the satellite side, and the computing tasks scheduled to the satellite can be computed immediately without waiting. The computing queuing delay of task m is the sum of the computing delays of the previous tasks, that is

[0120]

[0121] In summary, the total completion delay of a task = computing delay + queuing delay + communication delay, and the communication delay = offloading delay + scheduling delay. That is, the total completion delay for task m is:

[0122] 。

[0123] Step 2-2: Determine the satellite side as the central scheduling agent. The satellite side obtains the task information and computing queue status of all drones, and optimizes the task scheduling strategy based on the deep deterministic policy gradient algorithm to obtain the target task delay model.

[0124] Based on the task delay model constructed in the above steps, the total completion delay of the task initiated by user m at time slot t is 。This application aims to minimize the average completion delay of computing tasks and proposes an optimization scheme for task scheduling decisions. The optimization problem is modeled as follows:

[0125]

[0126]

[0127] Among them, C1 is the constraint on the task scheduling variable, ensuring that the sum of the task scheduling ratios is 1.

[0128] In the embodiments of the present application, for the task scheduling problem of the collaborative computing of the UAV cluster and the satellite, the Markov Decision Process (MDP) is used for modeling. Considering the problem that the inconsistent state dimensions may be caused by the load differences of the UAVs in the heterogeneous UAV cluster, preferably, the satellite side is defined as the central scheduling agent in the present application to achieve unified decision-making. The satellite side can obtain the task information and the computing queue status of all UAVs, so as to perform global optimization. In view of the fact that the scheduling decision needs to be made independently for each task, the Deep Deterministic Policy Gradient (DDPG) algorithm is used to optimize the task scheduling strategy. As a deep reinforcement learning method based on the Actor-Critic architecture, the DDPG algorithm can effectively handle the decision-making problem in the continuous action space and improve the accuracy of decision-making optimization in this application scenario.

[0129] The state of the agent includes the data size of the current task and the computing queue situation of all UAVs in the previous time slot, which is defined as follows:

[0130]

[0131] Among them, is the computing queue situation of UAV n in the (t - 1) time slot, and its update method is as follows:

[0132]

[0133] Considering that the task is divisible and can be scheduled to other UAVs and the satellite according to a certain ratio, therefore, the agent in the present application is defined as the scheduling ratio of task n, that is, it is defined as follows:

[0134]

[0135] After the agent outputs an action, it will interact with the environment to obtain the completion delay of the task, which is the reward - .

[0136] Furthermore, in the specific implementation, it may include the following specific steps 2 - 2.1 to step 2 - 2.4:

[0137] Step 2 - 2.1, initialize the Actor network and the Critic network of the Deep Deterministic Policy Gradient algorithm, and initialize the experience replay pool;

[0138] Step 2 - 2.2, determine the satellite side as the central scheduling agent, and for each user task , select the second agent action corresponding to the central scheduling agent according to the current policy (Add a certain amount of noise in the early stage of training to increase the exploration rate), and execute the second agent action After that, obtain the second immediate reward and the next second agent state , and store them in the experience replay pool;

[0139] Step 2-2.3, randomly sample a batch of samples from the experience replay pool, calculate the loss through the loss function, and update the parameters of the policy network and the evaluation network;

[0140] Step 2-2.4, determine the continuous action policy corresponding to minimizing the task completion delay through the above steps to determine the target task delay model.

[0141] Step 2-3, use the target task delay model to control other drones in the multi-drone cluster, or schedule satellites communicating with the multi-drone cluster to perform task collaborative computing with the target drone.

[0142] In this way, using the DDPG algorithm for decision optimization can gradually optimize the task scheduling strategy, learn a continuous action policy that can minimize the task completion delay, and thus enhance the decision-making of drone scheduling, effectively reducing the average task response time.

[0143] In summary, the embodiment of the present application models the operation scenario as a Markov decision process (MDP), and uses the model-based multi-agent proximal policy optimization (MAPPO) algorithm to perform the trajectory planning of the target drone, thereby ensuring the maximization of the task acquisition efficiency; adopting the deep deterministic policy gradient (DDPG) algorithm to promote the cooperation between drones, can make efficient scheduling decisions for user requests, effectively reduce the average task response time, and thus effectively improve the problem of uneven load distribution of drones.

[0144] In the actual task collection and task scheduling process, Figure 3 shows another task scheduling method for multi-aircraft cooperation in the airspace network edge computing scenario, including the following steps:

[0145] Step S1, for the ground users in any time slot, the drone cluster collects the user location information and data size information to determine its local state of the environment. The local state refers to the state observed by each drone agent. In this simulation system, the local state is [the location information of the current drone, the location information of all users, the task size of all users].

[0146] Step S2, the agent observes the local state and outputs a trajectory optimization decision.

[0147] Step S3, the UAV flies according to the actions output by the agent and collects tasks, and obtains the total task collection rate of the current time slot as the task reward of the agent.

[0148] Step S4, for the tasks unloaded to the UAV, the satellite side makes task scheduling decisions. For any task, obtain the task data size and the reinforcement of each UAV queue in the previous time slot as the input of the task scheduling agent, and the actor network outputs the scheduling decision of the task.

[0149] Step S5, schedule the tasks according to the scheduling decision, count the completion delay of the tasks, use it as the reward of the scheduling algorithm and update the environment, and enter the next state.

[0150] Step S6, the user moves randomly and enters the next time slot. The user generates tasks with a certain probability. In practical applications, the entire simulation duration is T. For the convenience of simulation, it is pre-divided into K time slots. At the beginning of each time slot, the user will generate tasks with a certain probability, and then the UAV adjusts its flight trajectory to collect tasks and complete the collaborative scheduling of tasks.

[0151] In summary, the agent in the above manner can dynamically adjust the flight path and task scheduling strategy of the UAV according to environmental changes and user behaviors, so as to better adapt to complex real-world scenarios; through the decision-making process of the agent, it ensures the optimization of the UAV's flight path and the rationalization of task allocation, and improves the resource utilization rate of the entire system.

[0152] Furthermore, the multi-UAV collaborative task scheduling system provided in the embodiment of the present application for the edge computing scenario of the space-air network is used to implement the multi-UAV collaborative task scheduling method for the edge computing scenario of the space-air network described in the foregoing implementation manner. Its implementation principle and the technical effects produced are the same as those of the foregoing method embodiment. For the sake of brief description, for the parts not mentioned in the embodiment of the multi-UAV collaborative task scheduling device for the edge computing scenario of the space-air network, reference may be made to the corresponding content in the foregoing embodiment of the multi-UAV collaborative task scheduling method for the edge computing scenario of the space-air network. The system includes a satellite side, a multi-UAV cluster, and multiple ground users, where the multi-UAV cluster includes multiple UAVs.

[0153] The embodiment of the present application also provides a UAV, as Figure 4 shown, which is a schematic structural diagram of the UAV. Among them, the electronic device 100 includes a processor 41 and a memory 40. The memory 40 stores computer-executable instructions that can be executed by the processor 41, and the processor 41 executes the computer-executable instructions to implement any one of the foregoing multi-UAV collaborative task scheduling methods for the edge computing scenario of the space-air network.

[0154] In Figure 4In the illustrated embodiment, the electronic device further includes a bus 42 and a communication interface 43. Among them, the processor 41, the communication interface 43, and the memory 40 are connected through the bus 42.

[0155] Among them, the memory 40 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. The communication connection between the system network element and at least one other network element is realized through at least one communication interface 43 (which can be wired or wireless), and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 42 may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, etc. The bus 42 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 only a single bidirectional arrow is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0156] The processor 41 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 41 or the instructions in the form of software. The above-mentioned processor 41 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor 41 reads the information in the memory and combines its hardware to complete the steps of the multi-machine collaborative task scheduling method for the space-air network edge computing scenario in the foregoing embodiments.

[0157] Unless otherwise specifically stated, the relative steps, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present application.

[0158] If the above function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium executable by a processor. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0159] In the description of the present application, it should be noted that the terms "first", "second", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0160] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A multi-machine collaborative task scheduling method for aerospace network edge computing scenarios, characterized in that: Applied to a scenario where users are unevenly distributed, the method includes: In response to a user task initiated by a randomly moving ground user, a target drone in a multi-drone cluster is matched, and the target drone collects the user task based on a pre-built task collection model; wherein the pre-built task collection model is a model for maximizing the task collection rate based on Markov decision process and multi-agent reinforcement learning modeling; The steps of constructing the task collection model include: Constructing an association matrix between the user tasks initiated by the ground user and the drones in the multi-drone cluster; the association matrix is ​​used to characterize the association relationship between the user tasks initiated by the ground user and the drones in the current time slot; Determine the task collection rate of the current time slot according to the association matrix, the first task data volume of publishing the user task, the second data volume of collecting the user task by the drone, the first location information of the ground user, and the second location information of the drone; During the flight of the UAV for task collection, the flight trajectory of the UAV is optimized based on flight constraints and a Markov decision process, and a task collection model of the UAV is determined based on a maximized task collection rate of the UAV corresponding to the optimized flight trajectory; When the user task is unloaded to the target drone, other drones in the multi-drone cluster are scheduled through a pre-built task delay model, or a satellite communicating with the multi-drone cluster is scheduled, so that other drones and / or satellites in the multi-drone cluster perform task collaborative calculation with the target drone, so as to minimize the completion delay of the user task collected under the maximum task collection rate; wherein the task delay model is a model constructed based on a deep deterministic policy gradient algorithm to optimize the task scheduling strategy; this step includes the following process: When the user task is unloaded to the target UAV, an initial task delay model is constructed; wherein the initial task delay model is a model related to communication delay, calculation delay and transmission delay; The satellite end is determined as the central scheduling agent. The satellite end obtains the task information and calculation queue status of all drones, optimizes the task scheduling strategy based on the deep deterministic policy gradient algorithm, and obtains the target task delay model. The target task delay model is used to schedule other drones in the multi-drone cluster, and / or schedule satellites communicating with the multi-drone cluster, so that other drones and / or satellites in the multi-drone cluster can perform task collaborative calculations with the target drone.

2. The multi-machine collaborative task scheduling method for aerospace network edge computing scenarios according to claim 1 is characterized in that: The task collection rate of the current time slot is determined according to the association matrix, the first task data volume of the user task released, the second data volume of the user task collected by the drone, the first location information of the ground user, and the second location information of the drone, and is expressed by the following formula: = in, is the task collection rate, N is the number of drones, M is the number of ground users, For user at time t i The amount of tasks generated, is the user at time t j The amount of tasks generated, For drones i and users j The association relationship between them, t is the time slot.

3. The multi-machine collaborative task scheduling method for aerospace network edge computing scenarios according to claim 1 is characterized in that: During the flight process of the UAV for task collection, the flight trajectory of the UAV is optimized based on flight constraints and a Markov decision process, and the task collection model of the UAV is determined based on the maximum task collection rate of the UAV corresponding to the optimized flight trajectory, including: A corresponding intelligent agent is deployed on each UAV. In each time slot, the intelligent agent interacts with the environment based on the flight constraints, takes a first intelligent agent action according to the first intelligent agent state, and obtains a first immediate reward after the intelligent agent performs the first intelligent agent action; wherein the first immediate reward is related to the task collection rate; the flight constraints include the flight speed constraints of the UAV, the flight angle constraints, the movement range constraints of the UAV and the ground user, and the collision distance constraints between the UAVs; When the agent interacts with the environment, a policy network and a value network are initialized for each agent, and the preset strategy is used to interact with the environment, and the corresponding agent state, agent action and immediate reward are obtained; the advantage value of each time step is calculated by generalized advantage estimation; the objective function of the multi-agent reinforcement learning algorithm is maximized, and the value function of the multi-agent reinforcement learning algorithm is minimized, and this step is repeated until convergence to determine the task collection model of the drone.

4. The multi-machine collaborative task scheduling method for aerospace network edge computing scenarios according to claim 1 is characterized in that: The satellite end is determined as the central scheduling agent. The satellite end obtains the task information and calculation queue status of all drones, optimizes the task scheduling strategy based on the deep deterministic policy gradient algorithm, and obtains the target task delay model, including: Initialize the Actor network and Critic network of the deep deterministic policy gradient algorithm, and initialize the experience replay pool; Determine the satellite end as the central scheduling agent, select the second agent action corresponding to the central scheduling agent according to the current strategy for each user task, obtain the second instant reward and the next second agent state after executing the second agent action, and store them in the experience replay pool; Randomly sampling a batch of samples from the experience replay pool, calculating the loss through the loss function, and updating the parameters of the policy network and the evaluation network; The above steps are used to determine the continuous action strategy corresponding to minimizing the task completion delay, so as to determine the target task delay model.

5. The multi-machine collaborative task scheduling method for aerospace network edge computing scenarios according to claim 1 is characterized in that: Match target drones in a multi-drone cluster, including: Based on the location information of the ground user, a drone in the multi-drone cluster within a flight coverage range corresponding to the location information is matched, and the drone is determined as a target drone.

6. The multi-machine collaborative task scheduling method for aerospace network edge computing scenarios according to claim 1 or 5 is characterized in that: The method further comprises: When the user task corresponds to multiple drones that can collect the user task, the target drone in the multi-drone cluster is matched based on the load of the drone.

7. A multi-machine collaborative task scheduling system for aerospace network edge computing scenarios, characterized in that: A task scheduling method for multi-machine collaboration for aerospace network edge computing scenarios according to claim 1 is used to implement the system, which includes a satellite terminal, a multi-UAV cluster and multiple ground users, wherein the multi-UAV cluster includes multiple UAVs.

8. A drone, characterized in that: It includes a processor and a memory, the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the multi-machine collaborative task scheduling method for aerospace network edge computing scenarios as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Time delay minimization calculation task unloading method and system in space-air-ground integrated network

    CN113346944A