Edge task scheduling method based on multi-agent near-end strategy optimization offline training
By adopting an offline training method optimized by multi-agent near-end strategy in post-disaster emergency networks, combining the temporary network and load control mechanism of drones and low-orbit satellites, the problems of low task scheduling efficiency and high computing load in the existing technology are solved, and more efficient task scheduling and cost reduction are achieved.
Patent Information
- Application Number
- CN202510578055.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-05-07
AI Technical Summary
When scheduling edge tasks in post-disaster emergency networks, existing methods only consider service deployment or load balancing, resulting in uneven node load and low overall system scheduling efficiency; at the same time, the use of online training reinforcement learning algorithms increases the computational load and cost, affecting real-time response capabilities.
The offline training method based on multi-agent near-end strategy optimization is adopted. By laying drones and low-orbit satellites in the ground-free infrastructure area to form temporary networks, adding service deployment restrictions and load control mechanisms, and offline simulation environment data is used to conduct offline training of multi-agent near-end strategy optimization algorithms, and obtaining an approximate optimal scheduling strategy model.
It effectively avoids the high computing load brought by online training, reduces the task scheduling cost, improves the overall task scheduling efficiency and real-time response capabilities of the system, and reduces the task scheduling cost by about 40.3% compared with the online training method.
Smart Images

Figure CN120085998A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of edge computing, and specifically, relates to an edge task scheduling method based on multi-agent proximal policy optimization offline training. Background Art
[0002] After natural disasters such as earthquakes, floods, and landslides occur, ground communication infrastructure is often severely damaged and traditional network services are interrupted, bringing great difficulties to information transmission and rescue coordination in the disaster area. In recent years, an air-ground-space collaborative network composed of unmanned aerial vehicles (UAVs) and low-earth orbit (LEO) satellites combined with mobile edge computing (MEC) technology has become an important means to build an emergency communication system. In the post-disaster emergency network, the help information and data processing tasks of ground users need to be responded to in a timely manner. Therefore, efficient computing task scheduling in the temporarily constructed emergency communication system becomes particularly important.
[0003] However, existing methods usually only consider either service deployment or load balancing alone for task scheduling, and rarely consider both at the same time. This single scheduling mechanism leads to uneven task allocation, node overload, and difficulty in adapting to the actual situation of the post-disaster network service distribution, reducing the overall scheduling efficiency of the system. For example, the applicant Hangzhou Dianzi University disclosed a task scheduling method based on service deployment update in its patent application document "Task Scheduling Method and System Based on Service Deployment Update" (application date: October 31, 2022, application number: 202211347782.7, application publication number: CN115696588 A, the content of this application can still be cited). The task scheduling method based on service deployment update has the following deficiencies: load balancing is not considered during the task scheduling process. When the computing tasks in the network surge, uneven task allocation will lead to node overload; and when an edge server deployed with the required service cannot be found within the communication range of the user, the task may not be offloaded. The patent application document "A Task Load Balancing Scheduling Method Based on Cloud-Edge-Terminal Collaboration" (application date: May 15, 2024, application number: 202410602801.9, application publication number: CN118394520 A, the content of this application can still be cited) disclosed a task scheduling method considering load balancing. The task scheduling method considering load balancing has the following deficiencies: it is assumed that all required services are available on the edge server during task scheduling, which does not conform to the actual situation of limited storage capacity of the edge server and reduces the overall scheduling efficiency of the system.
[0004] With the development of artificial intelligence technology, some researchers have begun to use reinforcement learning algorithms to solve task scheduling problems. However, when existing reinforcement learning algorithms are used to solve task scheduling problems, online training methods are mostly adopted, which brings high computational energy consumption and large computational latency. This not only increases the cost of task scheduling but also seriously affects the real-time response ability of emergency communication networks. For example, the applicant Jilin University disclosed a method for MEC task scheduling based on proximal policy optimization in its patent application document "MEC Task Scheduling Method Based on Proximal Policy Optimization in Vehicular Ad Hoc Networks" (application date: August 18, 2023, application number: 202311041910.X, application publication number: CN117014832 A, and the content of this application can still be cited). The method for MEC task scheduling based on proximal policy optimization has the following deficiencies: using online training to perform proximal policy optimization for the MEC task scheduling problem, the additional energy consumption and latency brought by online training increase the task scheduling cost; and the online training method occupies computing resources, bringing a high computational load to the MEC system.
[0005] In summary, when existing methods are used for edge task scheduling in post-disaster emergency networks, there are problems such as low overall system scheduling efficiency, and the online training of algorithms increases the computational load, resulting in a high task scheduling cost. Summary of the Invention
[0006] Object of the Invention: When existing methods are used for edge task scheduling in post-disaster emergency networks, there are problems that only considering service deployment or load control mechanisms leads to uneven node loads or difficulty in adapting to the post-disaster network service distribution, resulting in low overall system scheduling efficiency; and when using reinforcement learning algorithms to solve task scheduling problems, online training methods are adopted, which brings a high computational load, increases the cost of task scheduling, and seriously affects the real-time response ability of emergency communication networks. The present invention proposes an edge task scheduling method based on multi-agent proximal policy optimization offline training to make up for the above technical deficiencies.
[0007] To achieve the above object of the invention, the present invention adopts the following technical solutions: An edge task scheduling method based on multi-agent proximal policy optimization offline training, comprising the following steps: S1: Deploy unmanned aerial vehicles in areas without ground infrastructure, form a temporary network with low-earth orbit satellites, and incorporate service deployment restrictions and load control mechanisms. In the temporary network, only a part of the services are deployed on the edge servers carried by each unmanned aerial vehicle, and all services are deployed on the edge servers carried by the low-earth orbit satellites; S2: Define the task scheduling delay and the task scheduling energy consumption respectively, and define the task scheduling cost according to the task scheduling delay and the task scheduling energy consumption; the task scheduling delay including the task upload time , the task migration time and the task calculation time ; the energy consumption of task scheduling including the calculation energy consumption , the migration energy consumption and the hovering energy consumption ; S3: Model the task scheduling cost minimization problem as a discrete partially observable Markov decision process, and the discrete partially observable Markov decision process includes a local observation state, a global state, an action space, and a reward function; S4: Simulate the environmental data through an offline simulation, and use the multi-agent proximal policy optimization algorithm to perform offline training on the task scheduling problem in the network to obtain an approximate optimal scheduling policy model; the environmental data of the offline simulation is obtained through a digital twin simulation platform and includes state-action-reward tuples , as the data input for training the multi-agent proximal policy optimization algorithm; S5: Pre-deploy the model parameters obtained from the offline training in the UAV. When a computing task request is received, directly call the model for online inference to obtain the optimal scheduling policy.
[0008] Furthermore, the specific steps of step S1 are as follows: S1-1: Deploy UAVs in areas without ground infrastructure and form a temporary network with low-earth orbit satellites: Deploy UAVs in areas without ground infrastructure and form a temporary network with low-earth orbit satellites and ground users; the sets of the three types of nodes, namely UAVs , low-earth orbit satellites and users are respectively represented as , and ; among them, edge servers are equipped on both UAVs and low-earth orbit satellites to provide computing services for ground users
[0009] S1-2: Add service deployment restrictions and load control mechanisms: Let represent the number of services in the network, represent the set of services in the network, and only a part of the services are deployed on the edge server carried by each UAV, that is , where represents UAV A set of services deployed on the edge servers carried by high-orbit satellites; while all services are deployed on the edge servers carried by low-earth orbit satellites , the computing tasks generated by ground users can only be scheduled to the edge servers with the corresponding services for processing. For example, if the computing task generated by a ground user is an image recognition computing task, then this computing task needs to be scheduled to the edge server where the image recognition service is deployed for processing; Add a load control mechanism to the edge servers carried by drones. When the drone receives a computing task from a ground user after that, first store it in the buffer; if the drone has the service corresponding to processing this computing task, and the drone detects that the occupancy rate of its own buffer is lower than the load threshold , then it will directly process the computing task locally; otherwise, trigger the load control mechanism and forward the computing task to a drone with a load buffer lower than the load threshold and having the service corresponding to processing this computing task for processing. If no eligible drone can be found within the communication range of the drone , then forward the computing task to the low-earth orbit satellite for processing.
[0010] Furthermore, the specific steps of step S2 are as follows: S2-1: Define the task scheduling delay and the task scheduling energy consumption respectively: The task processing modes of drones are divided into three types: local processing, forwarding to other drones for processing, and forwarding to low-earth orbit satellites for processing; therefore, the task scheduling delay includes the task upload time , the task migration time and the task calculation time respectively: (1) The task scheduling energy consumption of the drone includes the computing energy consumption , the migration energy consumption and the hovering energy consumption respectively: (2) S2-2: Define the scheduling cost according to the task scheduling delay and the task scheduling energy consumption : The goal of the task scheduling problem is to find the optimal task scheduling strategy , minimizing the scheduling cost; the cost of task scheduling is defined as: (3) where and represent the weight factors of task scheduling delay and UAV task scheduling energy consumption respectively, and satisfy ; and represent the maximum task scheduling delay and the maximum UAV task scheduling energy consumption respectively.
[0011] Furthermore, the specific steps of step S3 are as follows: Regarding each UAV as an agent, the discrete partially observable Markov decision process includes the following elements: (1) Local observation state: At time slot , the agent collects environmental information within the communication range, and the local observation state is defined as: (4) where is the computing load of UAV ; is the service deployment set of UAV ; is the task queue state of UAV ; is the channel state of UAV ; is the remaining computing resources of UAV . (2) Global state: At time slot , the global state includes all UAV local observation states , the global unprocessed task queue and the global computing task request information , and the global state is defined as: (5) (3) Action space: At time slot , the action is used to represent the action space of UAV , and the action is defined as: (6) Constraints: and ; wherein, is the target node, and the target node is a drone or a low-earth orbit satellite; represents the target node has deployed the corresponding service, represents the target node the occupancy rate of the buffer is lower than the load threshold ; Reward function: The goal of task scheduling is to find the optimal task scheduling strategy , minimizing the scheduling cost. Therefore, at time slot , the reward obtained after the agent interacts with the environment is defined as: (7) wherein, is the resource overrun penalty coefficient, and its value ranges from 0 to 1; is the drone allocated to the computing task of the computing resource, is the drone available computing resources.
[0012] Furthermore, the specific steps of step S4 are as follows: S4-1: Simulate the environmental data through offline simulation: First, build a digital twin simulation platform. Based on the temporary network topology of the drone-low-earth orbit satellite-ground user defined in steps S1 and S2, clarify the simulation parameters, including the drone units, low-earth orbit satellite satellites, ground user individuals, the services deployed by each drone , load threshold , task generation rate ; During the simulation process, simulate the computing task flow dynamically generated by the ground user, and use the Poisson distribution model to generate computing tasks , each task carries service type requirements, and record the task upload time , migration time and computing time , in order to obtain the task scheduling delay data; Record the task computing energy consumption of the drone , migration energy consumption and hovering energy consumption , in order to obtain the task scheduling energy consumption of the drone data; meanwhile, in each time slot generate the local observation state , including the computing load of the UAV 、the service deployment set of the UAV 、the task queue state of the UAV 、the channel state of the UAV 、the remaining computing resources of the UAV ; constitute the state-action-reward tuple , as the data input for training the multi-agent proximal policy optimization algorithm; S4-2: Use the multi-agent proximal policy optimization algorithm to perform offline training on the task scheduling problem in the network to obtain an approximate optimal scheduling policy model: Using the state-action-reward tuple obtained in step S4-1 as the input, perform offline training on the task scheduling problem in the network. The loss function of the multi-agent proximal policy optimization algorithm is defined as: (8) where represents the action probability ratio, is the scheduling policy parameter, is the advantage function, is the clipping range parameter, ∈[0.1,0.3]; represents clipping the action probability ratio , and the clipping range is ; the advantage function is defined as: (9) where is the value function at time slot , is the value function at time slot , is the discount factor, and its value ranges from 0.9 to 0.99.
[0013] Furthermore, the specific steps of step S5 are as follows: S5-1: Pre-deploy the model parameters obtained from offline training to the UAV: Compress the approximate optimal scheduling policy model obtained from offline training in S4; and pre-deploy the compressed model parameters to the edge server of the UAV; S5-2: When receiving a computing task request, directly call the model for online inference to obtain the optimal scheduling strategy: When the UAV receives a computing task from a ground user , first obtain the local observation state , including the computing load of the UAV , the service deployment set of the UAV , the task queue state of the UAV , the channel state of the UAV , the remaining computing resources of the UAV ; then input the local observation state into the deployed policy model for forward inference, and output the optimal scheduling strategy .
[0014] The advantages and technical effects of the present invention are as follows: The present invention first performs task scheduling by simultaneously considering service deployment and adding a load control mechanism, evenly distributes tasks under the actual service deployment of the edge network, avoids overloading of node tasks, and improves the overall system scheduling efficiency; secondly, simulates environmental data offline through simulation, and uses the multi-agent proximal policy optimization algorithm to perform offline training on the task scheduling problem in the network to obtain an approximately optimal scheduling strategy model, avoiding the additional computational load, energy consumption, and latency brought by the online training process of the multi-agent proximal policy optimization algorithm, reducing the task scheduling cost, and enabling the system to quickly respond to task requests; finally, through simulation verification, when the UAV computing power is set to 0.6 (billion cycles per second), the method in the present invention reduces the task scheduling cost by about 40.3% compared with the online training multi-agent proximal policy optimization (MAPPO) method.
[0015] In summary, the present invention effectively avoids the high computational load of online training, reduces the task scheduling cost, and improves the overall task scheduling efficiency and real-time response ability of the system. Brief Description of the Drawings
[0016] Figure 1 is the overall flowchart of an embodiment of the present invention; Figure 2 is the network architecture diagram of an embodiment of the present invention; Figure 3Schematic diagram of the load control mechanism with an edge server mounted on a drone in an embodiment of the present invention; Figure 4 Schematic diagram of three task scheduling and processing modes of a drone in an embodiment of the present invention; Figure 5 Comparison chart of the task scheduling cost simulation results of the method provided by the present invention and the MAPPO method with online training in an embodiment of the present invention. Detailed implementation manners
[0017] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0018] This embodiment proposes an edge task scheduling method based on multi-agent proximal policy optimization offline training. Its overall flowchart is as Figure 1 shown, and includes the following steps: S1: Deploy drones in areas without ground infrastructure, form a temporary network with low-earth orbit satellites, and incorporate service deployment restrictions and a load control mechanism. In the temporary network, only a part of the services are deployed on the edge server carried by each drone, and all services are deployed on the edge server carried by the low-earth orbit satellite. The specific steps are as follows: S1-1: Deploy drones in areas without ground infrastructure and form a temporary network with low-earth orbit satellites: As Figure 2 shown, deploy drones in areas without ground infrastructure, and form a temporary network with low-earth orbit satellites and ground users. In this embodiment, = 8, = 3, = 12 are taken respectively; the sets of the three types of nodes, namely drones , low-earth orbit satellites and users are respectively represented as , and ; among them, edge servers are equipped on both drones and low-earth orbit satellites to provide computing services for ground users; S1-2: Incorporate service deployment restrictions and a load control mechanism: Let represent the number of services in the network. In this embodiment, = 10 is taken; represents the set of services in the network. Only a part of the services are deployed on the edge server carried by each drone, that is, , where represents drone The set of services deployed by the edge servers carried on it; while all services are deployed by the edge servers carried on the low-earth orbit satellites , the computing tasks generated by the ground users can only be scheduled to the edge servers with the corresponding services for processing. For example, if the computing task generated by the ground user is an image recognition computing task, then this computing task needs to be scheduled to the edge server deployed with the image recognition service for processing; Such as Figure 3 shown, a load control mechanism is added to the edge server carried on the unmanned aerial vehicle. When the unmanned aerial vehicle receives the computing task from the ground user , it first stores it in the buffer; if the unmanned aerial vehicle has the service corresponding to processing this computing task, and the unmanned aerial vehicle detects that the occupancy rate of its own buffer is lower than the load threshold , then it will directly process the computing task locally; otherwise, it triggers the load control mechanism and forwards the computing task to the unmanned aerial vehicle with a load buffer lower than the load threshold and having the service corresponding to processing this computing task for processing. If no eligible unmanned aerial vehicle can be found within the communication range of the unmanned aerial vehicle , then the computing task is forwarded to the low-earth orbit satellite for processing. In this embodiment, the load threshold is taken.
[0019] S2: Define the task scheduling delay and the task scheduling energy consumption respectively, and define the task scheduling cost and the task scheduling energy consumption based on the task scheduling delay ; the task scheduling delay includes the task upload time , the task migration time and the task calculation time ; the task scheduling energy consumption includes the computing energy consumption , the migration energy consumption and the hovering energy consumption , and the specific steps are as follows: S2-1: Define the task scheduling delay and the task scheduling energy consumption respectively: Such as Figure 4As shown in Figure 2, the UAV task processing modes are divided into three types: local processing, forwarding to other UAVs for processing, and forwarding to low-orbit satellites for processing; therefore, the task scheduling delay Including task upload time , task migration time and task calculation time : (1) Drones Task scheduling energy consumption Including calculation of energy consumption , Migration energy consumption and hovering energy consumption : (2) S2-2: Delay according to task scheduling and task scheduling energy consumption Defining Scheduling Costs : The goal of the task scheduling problem is to find the optimal task scheduling strategy , minimize the scheduling cost; the cost of scheduling tasks Defined as: (3) in, and Represent the weight factors of task scheduling delay and UAV task scheduling energy consumption, and satisfy In this embodiment, =0.4 and =0.6; and They represent the maximum task scheduling delay and the maximum UAV task scheduling energy consumption respectively.
[0020] S3: The task scheduling cost minimization problem is modeled as a discrete partially observable Markov decision process, which includes a local observation state, a global state, an action space, and a reward function. The specific steps are as follows: Each drone Considered as an agent, a discrete partially observable Markov decision process consists of the following elements: (1) Local observation state: in the time slot , the agent collects environmental information within the communication range and observes the local state is defined as: (4) in, For drones The computational load; is the service deployment set for the UAV ; is the task queue status of the UAV ; is the channel status of the UAV ; is the remaining computing resources of the UAV ; (2) Global state: At time slot , the global state includes all local observation states of the UAVs , the global unprocessed task queue and the global computing task request information . The global state is defined as: (5) (3) Action space: At time slot , the action of the UAV is represented by the action . The action is defined as: (6) Constraint conditions: and ; where is the target node, and the target node is a UAV or a low-earth orbit satellite; indicates that the target node has deployed the corresponding service, indicates that the occupancy rate of the buffer of the target node is lower than the load threshold ; Reward function: The goal of task scheduling is to find the optimal task scheduling strategy to minimize the scheduling cost. Therefore, at time slot , the reward obtained by the agent after interacting with the environment is defined as: (7) where is the resource overrun penalty coefficient, and its value ranges from 0 to 1. In this embodiment, = 0.1; is the computing resource allocated by the UAV to the computing task , is the available computing resource of the UAV .
[0021] S4: Simulate the environment data offline, and use the multi-agent proximal policy optimization algorithm to perform offline training on the task scheduling problem in the network to obtain an approximate optimal scheduling policy model; the offline simulation environment data is obtained through a digital twin simulation platform and contains state-action-reward tuples. , which is used as the data input for the training of the multi-agent proximal policy optimization algorithm. The specific steps are as follows: S4-1: Simulate the environment data offline: First, build a digital twin simulation platform. Based on the temporary network topology of the drone-Low Earth Orbit (LEO) satellite-ground users defined in steps S1 and S2, clarify the simulation parameters, including the number of drones , the number of LEO satellites , the number of ground users , the services deployed on each drone , the load threshold = 80%, the task generation rate tasks / second; during the simulation process, simulate the computational task flow dynamically generated by ground users, and use the Poisson distribution model to generate computational tasks . Each task carries service type requirements and records the task upload time , the migration time and the computing time to obtain the task scheduling delay data; record the task computing energy consumption of the drone , the migration energy consumption and the hovering energy consumption of the drone to obtain the task scheduling energy consumption data of the drone; meanwhile, generate local observation states at each time slot , including the computing load of the drone , the service deployment set of the drone , the task queue status of the drone , the channel status of the drone , the remaining computing resources of the drone ; form state-action-reward tuples , which is used as the data input for the training of the multi-agent proximal policy optimization algorithm; S4-2: Use the multi-agent proximal policy optimization algorithm to perform offline training on the task scheduling problem in the network to obtain an approximate optimal scheduling policy model: Using the state-action-reward tuple obtained in step S4-1 as the input, perform offline training on the task scheduling problem in the network. The loss function of the multi-agent proximal policy optimization algorithm is defined as: (8) where represents the action probability ratio, is the scheduling policy parameter, is the advantage function, is the clipping range parameter, ∈[0.1, 0.3], and in this embodiment, = 0.2; represents clipping the action probability ratio , and the clipping range is ; The definition of the advantage function is: (9) where is the value function at time slot , is the value function at time slot , is the discount factor, with a value between 0.9 and 0.99. In this embodiment, = 0.94; In this embodiment, through rounds of offline simulation iterative training, minimize the loss function and the scheduling cost , and maximize the reward ; where represents the number of offline simulation iterations. In this embodiment, = 200.
[0022] S5: Pre-deploy the model parameters obtained from offline training in the UAV. When a computing task request is received, directly call the model for online inference to obtain the optimal scheduling policy. The specific steps are as follows: S5-1: Pre-deploy the model parameters obtained from offline training in the UAV: Compress the approximate optimal scheduling policy model obtained from offline training in S4 to reduce the space occupied by the model; and pre-deploy the compressed model parameters to the edge server of the UAV so that the UAV can quickly call them when receiving a computing task request from a ground user; S5-2: When a computing task request is received, directly call the model for online inference to obtain the optimal scheduling policy: When the UAV receives a ground user Computing task When performing the computing task , the local observation state is first obtained , including the computing load of the UAV , the service deployment set of the UAV , the task queue state of the UAV , the channel state of the UAV , and the remaining computing resources of the UAV ; then the local observation state is input into the deployed policy model for forward inference, and the optimal scheduling policy
[0023] is output Figure 5 . The simulation comparison results of the task scheduling cost using the method provided by the present invention and the online-trained MAPPO method are as
[0024] shown. This embodiment completes the simulation in Matlab R2021a, uses a computer memory of 32GB, and uses an Intel Core i5-12400F CPU processor based on x64. The specific parameters of all simulations in this embodiment are listed in Table 1 , It can be seen from the simulation results Figure 5 that as the computing power of the UAV increases, the task scheduling cost of the method provided by the present invention is significantly reduced. This is because when the computing power of the UAV increases, the system can process more tasks within the same scheduling period, thereby reducing the load overflow caused by task accumulation and reducing the task processing cost at the same time; the online-trained MAPPO method performs relatively poorly. When the computing power of the UAV is set to 0.6 (billion cycles per second), the method provided by the present invention reduces the task scheduling cost by about 40.3% compared with the online-trained MAPPO method
[0025] In summary, the present invention effectively avoids the high computing load brought by online training, reduces the task scheduling cost, and improves the overall task scheduling efficiency of the system
[0026] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions required to be protected by the present invention
Claims
1. A method for scheduling edge tasks based on multi-agent proximal strategy optimization of offline training, characterized in that: The method comprises the following steps: S1: Deploy drones in areas without ground infrastructure, form a temporary network with low-orbit satellites, and add service deployment restrictions and load control mechanisms. In the temporary network, the edge server carried by each drone deploys only a part of the services, and the edge server carried by the low-orbit satellite deploys all services; S2: Define task scheduling delays separately and task scheduling energy consumption , and delay according to task scheduling and task scheduling energy consumption Defining the cost of scheduling tasks ; The task scheduling delay Including task upload time , task migration time and task calculation time ; Task scheduling energy consumption Including calculation of energy consumption , Migration energy consumption and hovering energy consumption ; S3: The task scheduling cost minimization problem is modeled as a discrete partially observable Markov decision process, wherein the discrete partially observable Markov decision process includes a local observation state, a global state, an action space, and a reward function; S4: Through offline simulation of environmental data, a multi-agent proximal strategy optimization algorithm is used to perform offline training on the task scheduling problem in the network to obtain an approximate optimal scheduling strategy model; the offline simulation environmental data is obtained through a digital twin simulation platform and includes a state-action-reward tuple , as the data input for training the multi-agent proximal policy optimization algorithm; S5: The model parameters obtained through offline training are pre-deployed in the UAV. When a computing task request is received, the model is directly called for online reasoning to obtain the optimal scheduling strategy.
2. The edge task scheduling method based on multi-agent proximal strategy optimization offline training as claimed in claim 1, characterized in that: The step S1 is specifically as follows: S1-1: Deploy drones in areas without ground infrastructure and establish a temporary network with low-orbit satellites: Deployment in areas without ground infrastructure drones, and Low Earth Orbit Satellites and Ground users form a temporary network; drones , low-orbit satellite and users The sets of these three types of nodes are represented as , and Among them, drones and low-orbit satellites are equipped with edge servers to provide computing services to ground users; S1-2: Add service deployment restrictions and load control mechanisms: make Indicates the number of services in the network, Represents the set of services in the network. Each edge server carried by the drone deploys only a part of the services, i.e. ,in, Indicates drone The edge servers on the LEO satellite deploy all services. , the computing tasks generated by ground users can only be scheduled to the edge servers with corresponding services for processing. For example, if the ground users generate image recognition computing tasks, the computing tasks need to be scheduled to the edge servers deployed with image recognition services for processing; Add a load control mechanism to the edge server on the drone. Received to ground user Computational tasks After that, it is first stored in the buffer; if the drone Have a service to handle the computing task, and the drone Check that the occupancy rate of its own buffer is lower than the load threshold , the computation task will be directly Otherwise, the load control mechanism is triggered to transfer the computing task to Forwarded to the load buffer below the load threshold And have a drone that handles the corresponding service for the computing task If the drone No qualified drones were found within the communication range , then the calculation task Forward to low-orbit satellite On processing.
3. The edge task scheduling method based on multi-agent proximal strategy optimization offline training as claimed in claim 1, characterized in that: The step S2 is specifically as follows: S2-1: Define task scheduling delays separately and task scheduling energy consumption : The UAV task processing modes are divided into three types: local processing, forwarding to other UAVs for processing, and forwarding to low-orbit satellites for processing; therefore, task scheduling delays Including task upload time , task migration time and task calculation time : (1) Drones Task scheduling energy consumption Including calculation of energy consumption , Migration energy consumption and hovering energy consumption : (2) S2-2: Delay according to task scheduling and task scheduling energy consumption Defining Scheduling Costs : The goal of the task scheduling problem is to find the optimal task scheduling strategy , minimize the scheduling cost; the cost of scheduling tasks Defined as: (3) in, and Represent the weight factors of task scheduling delay and UAV task scheduling energy consumption, and satisfy ; and They represent the maximum task scheduling delay and the maximum UAV task scheduling energy consumption respectively.
4. The edge task scheduling method based on multi-agent proximal strategy optimization offline training as claimed in claim 1, characterized in that: In step S3, the task scheduling cost minimization problem is modeled as a discrete partially observable Markov decision process, which includes a local observation state, a global state, an action space, and a reward function, as follows: Each drone Considered as an agent, a discrete partially observable Markov decision process consists of the following elements: (1) Local observation state: in the time slot , the agent collects environmental information within the communication range and observes the local state is defined as: (4) in, For drones The computational load; For drones A collection of service deployments; For drones The task queue status; For drones The channel status; For drones Remaining computing resources; (2) Global state: in the time slot , global state Contains the local observation status of all drones , Global unprocessed task queue and global computing task request information , global state is defined as: (5) (3) Action space: in time slots , use actions Indicates drone Action space, action is defined as: (6) Constraints: and ; in, is the target node, which is a drone or a low-orbit satellite; Indicates the target node Deployed the corresponding services. Indicates the target node Buffer occupancy Below load threshold ; Reward function: The goal of task scheduling is to find the optimal task scheduling strategy , minimize the scheduling cost, so in the time slot , the reward obtained by the agent after interacting with the environment is defined as: (7) in, is the resource overlimit penalty coefficient, which ranges from 0 to 1; For drones Assign to computing tasks of computing resources, For drones Available computing resources.
5. The edge task scheduling method based on multi-agent proximal strategy optimization offline training as claimed in claim 1, characterized in that: The step S4 is specifically as follows: S4-1: Simulate environmental data through offline simulation: First, a digital twin simulation platform is constructed. Based on the temporary network topology of UAV-LEO-ground users defined in steps S1 and S2, simulation parameters are clarified, including UAV Low Earth Orbit Satellite terrestrial users Services deployed per drone , Load Threshold , task generation rate During the simulation, the computing task flow generated dynamically by ground users is simulated, and the Poisson distribution model is used to generate computing tasks. , each task carries the service type requirement and records the task upload time , Migration time and calculation time , in order to obtain the task scheduling delay Data; Recording drone The task calculation energy consumption , Migration energy consumption and hovering energy consumption , in order to obtain the drone Task scheduling energy consumption data; at the same time, in each time slot Generate local observation state , including drones The computational load , drones A collection of service deployments , drones Task queue status , drones Channel status , drones Remaining computing resources ; Construct state-action-reward tuple , as the data input for training the multi-agent proximal policy optimization algorithm; S4-2: Use the multi-agent proximal strategy optimization algorithm to perform offline training on the task scheduling problem in the network and obtain an approximate optimal scheduling strategy model: Take the state-action-reward tuple obtained in step S4-1 As input, the task scheduling problem in the network is trained offline, and the loss function of the multi-agent proximal strategy optimization algorithm is Defined as: (8) in, represents the action probability ratio, is the scheduling strategy parameter, is the advantage function, is the clipping range parameter, ∈[0.1,0.3]; Represents the probability ratio of the action Crop, and the cropping range is [ ]; advantage function is defined as: (9) in, For time slot The value function when For time slot The value function when is the discount factor, ranging from 0.9 to 0.
99.
6. The edge task scheduling method based on multi-agent proximal strategy optimization offline training as claimed in claim 1, characterized in that: The step S5 is specifically as follows: S5-1: Pre-deploy the model parameters obtained from offline training in the drone: The approximate optimal scheduling strategy model obtained by offline training in S4 is compressed, and the compressed model parameters are pre-deployed to the edge server of the UAV; S5-2: When a computing task request is received, the model is directly called for online reasoning to obtain the optimal scheduling strategy: When a drone Receiving from ground users Computational tasks When , first obtain the local observation state , including drones The computational load , drones A collection of service deployments , drones Task queue status , drones Channel status , drones Remaining computing resources ; Then the local observation state Input the deployed policy model for forward reasoning and output the optimal scheduling policy .
Citation Information
Patent Citations
Task scheduling method and system based on service deployment updating
CN115696588A
MEC task scheduling method based on near-end strategy optimization in Internet of Vehicles
CN117014832A
Task load balancing scheduling method based on cloud side-end collaboration
CN118394520A
MAPPO-Att-based multi-unmanned aerial vehicle edge network task unloading and resource allocation method
CN119653373A
Multi-machine collaborative task scheduling method for aerospace network edge computing scene
CN119759587A
Cited By
Sea area sensing task-oriented task scheduling method and device
CN121092294A
A task scheduling method and device for sea area perception task
CN121092294B