An intelligent scheduling method for low-altitude economic activities
By building a task environment model and dividing task sub-regions, combining market bidding method and DDPG algorithm, the efficiency of drone task allocation and charging strategies is solved, and efficient drone scheduling and continuous operation are achieved.
Patent Information
- Application Number
- CN202510377558.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-28
AI Technical Summary
The prior art is difficult to accurately allocate UAV tasks according to rapidly changing tasks and environmental needs, resulting in inefficient resource utilization and lack of efficient charging strategies, making it difficult to ensure the continuity of tasks.
By building a task environment model, the task sub-regions are divided, and the task is allocated to the computing drone using the market bidding method. At the same time, the DDPG algorithm is used to generate real-time scheduling strategies for charging drones, and dynamically adjust the flight path and charging actions.
It improves the efficiency and resource utilization of drone mission scheduling, ensures the continuous and efficient operation of computing drones, and adapts to changes in the dynamic environment.
Smart Images

Figure CN119886771B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of UAV scheduling, and particularly to an intelligent scheduling method for low-altitude economic activities. Background Art
[0002] Traditional edge computing distributes data processing capabilities and storage resources at the network edge, close to the data source, to reduce the burden on remote data centers. This approach reduces transmission latency, improves data processing speed and response efficiency, and is very suitable for application scenarios that require real-time processing and low latency requirements, such as the Internet of Things and smart cities. However, fixed edge nodes are usually limited in distribution and cannot flexibly handle dynamic and temporary activities.
[0003] To address the above problems, applying UAVs to edge computing can improve the flexibility and coverage of the system. As mobile computing nodes, UAVs can quickly respond when task requirements change, thereby providing computing resources and data collection functions. By carrying computing devices, UAVs can provide mobile edge computing services in a wide geographical space, solve the problem of areas that cannot be covered by fixed nodes, and make the application of edge computing more flexible and extensive.
[0004] However, UAV scheduling planning faces many challenges. Existing technologies are difficult to accurately allocate tasks according to rapidly changing task and environmental requirements, resulting in low resource utilization efficiency. Moreover, due to the limited battery life of UAVs, current technologies have not been able to provide sufficiently efficient charging strategies to ensure the continuity of tasks, which makes it difficult to achieve efficient UAV scheduling in dynamic and uncertain environments. Summary of the Invention
[0005] To solve the problems in the above background art, the present invention adopts the following technical solutions:
[0006] An intelligent scheduling method for low-altitude economic activities, comprising the steps of:
[0007] Construct a task environment model, and obtain a task requirement set, a computing UAV set, and a charging UAV set; the task environment model includes a task area and several device positions; the computing UAV set includes several computing UAVs; the charging UAV set includes several charging UAVs; the task requirement set includes the ground devices and resource demand quantities corresponding to several tasks;
[0008] Divide several task sub-areas according to the task area, device positions, and the ground devices and resource demand quantities corresponding to the tasks;
[0009] Allocate computing UAVs to each task sub-area; allocate the tasks in the task sub-area to the computing UAVs through the market bidding method;
[0010] During the process of the computing UAV executing tasks, a real-time scheduling strategy for the charging UAV is generated through the DDPG algorithm.
[0011] As a preferred solution of the present invention, dividing several task sub-regions according to the task area, device location, and the required amount of ground devices and resources corresponding to the task includes the steps of:
[0012] S21. Obtain the total resource demand of each ground device according to the task and the corresponding resource demand
[0013] S22. Determine the number of task sub-regions, and initialize the position and velocity of each particle in the particle swarm optimization algorithm; the position of the particle is used to represent the division scheme of the task sub-regions;
[0014] S23. Calculate the fitness values of all particles, and update the global best position and the personal best position of each particle;
[0015] S24. Update the velocity and position of each particle according to the personal best position and the global best position;
[0016] S25. Repeat steps S23 to S24 until the maximum number of iterations is reached or the fitness value is less than the set threshold, and use the global best position as the division scheme of the task sub-regions.
[0017] As a preferred solution of the present invention, the total resource demand of the ground device is expressed as:
[0018] ;
[0019] Wherein, represents the total resource demand of the h-th ground device; represents the resource demand of the q-th task of the h-th ground device; represents the total number of tasks of the h-th ground device.
[0020] As a preferred solution of the present invention, the position of the particle is expressed as:
[0021] ,
[0022] Wherein, is the number of task sub-regions; represents the clustering center of the k-th task sub-region, , and respectively represent the longitude and latitude of the clustering center of the k-th task sub-region; represents the clustering radius of the k-th task sub-region.
[0023] As a preferred solution of the present invention, the fitness value is expressed as:
[0024] ,
[0025] Among them, represents the distance fitness function; represents the resource balance fitness function; and respectively represent the distance fitness function and the weight parameter of the distance fitness function; represents the distance between the hth ground device and the kth clustering center; represents the total number of tasks in the kth task sub-region.
[0026] As a preferred solution of the present invention, the task in the task sub-region is assigned to the computing UAV through the market bidding method, specifically: repeat the following steps until all tasks in the task demand set are assigned:
[0027] Obtain the task with the largest unassigned resource demand in the task sub-region of the task demand set, and record it as the current task to be assigned;
[0028] Calculate the quotation index according to the difference between the expected total revenue after the current task to be assigned is added to the task sequence and the current total revenue of the computing UAV, and generate a quotation request according to the quotation index;
[0029] Each computing UAV broadcasts the quotation request to other computing UAVs in all task sub-regions;
[0030] After each computing UAV receives the quotation request, it generates an evaluation result according to its own quotation index and transmits it to the scheduling center;
[0031] The scheduling center assigns the current task to be assigned to the corresponding computing UAV according to all evaluation results, and updates the task demand set and the task sequence of the computing UAV.
[0032] As a preferred solution of the present invention, the quotation index is expressed as:
[0033] ;
[0034] Among them, i represents the computing UAV number, represents the current task to be assigned in the kth task sub-region; represents the task sequence of the ith computing UAV; represents the expected total revenue after the current task to be assigned is added to the task sequence ; represents the current total revenue of the ith computing UAV.
[0035] As a preferred solution of the present invention, the method for obtaining the expected total revenue after the current task to be assigned is added to the task sequence is:
[0036] Determine the optimal position of the current task to be assigned in the task sequence according to the device position corresponding to the current task to be assigned, and generate a reference task sequence;
[0037] Calculate the expected total revenue according to the scheduling duration between adjacent tasks in the reference task sequence, the task revenue of each task, and the calculated task execution time.
[0038] As a preferred solution of the present invention, the method for generating a real-time scheduling strategy for a charging drone by using the DDPG algorithm includes the steps of:
[0039] Construct a Markov decision process for the charging drone; the Markov decision process consists of a state space, an action space, a transition probability, a reward function, and a discount factor;
[0040] Set an optimization objective function for implementing the scheduling strategy of the charging drone;
[0041] Use the DDPG algorithm based on the Markov decision process and the optimization objective function to generate a real-time scheduling strategy for the charging drone.
[0042] As a preferred solution of the present invention, the optimization objective function is set to minimize the task completion duration of all computing drones;
[0043] The reward function is expressed as:
[0044] ;
[0045] Wherein, represents the reward obtained by the charging drone within the time slot; represents the charging amount provided by the charging drone within the time slot; is a preset constant; represents a charging fairness index, , represents the number of computing drones in the k-th task sub-region; When it means that the charging drone does not perform a charging action within the time slot, When it means that the charging drone performs a charging action within the time slot.
[0046] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0047] By building a task environment model and dividing sub-regions according to task requirements, the present invention can provide clear task partitions for computing drones and charging drones, reducing unnecessary frequent interactions and resource waste between devices and improving the overall task scheduling efficiency. By subdividing the task area into multiple sub-regions and precisely allocating according to resource requirements, it ensures that each task sub-region can obtain corresponding computing resources and charging services according to actual needs. During the task execution process, a scheduling strategy for charging drones is dynamically generated through the DDPG algorithm, enabling the charging drones to adjust their flight paths and charging actions according to real-time tasks, thereby ensuring the continuous and efficient operation of computing drones in a dynamic environment.
[0048] In the embodiment of the present invention, based on the particle swarm optimization algorithm, several task sub-regions are divided according to the task area, device locations, and the ground devices and resource requirements corresponding to the tasks; based on the distance fitness function, the concentration degree of the positions of ground devices within each task sub-region is evaluated; based on the resource balance fitness function, the equilibrium state of resource requirements within each task sub-region can be effectively evaluated and optimized; the fitness function of the particle swarm optimization algorithm is set as the weighted sum of the distance fitness function and the resource balance fitness function, and the task area division is continuously optimized by updating the velocity and position of the particles, so as to comprehensively consider geographical locations and resource requirements, make the distribution of devices and tasks within the clustering region more reasonable, reduce the additional consumption of drones caused by distance problems and the drone scheduling conflicts caused by uneven resource allocation, and improve the overall task execution efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The drawings herein are incorporated into the specification and form a part of this specification, showing the embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.
[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0051] Figure 1 It is a schematic flowchart of an intelligent scheduling method for low-altitude economic activities provided by an embodiment of the present invention;
[0052] Figure 2 It is a schematic flowchart of dividing several task sub-regions according to the task area, device locations, and the ground devices and resource requirements corresponding to the tasks provided by an embodiment of the present invention;
[0053] Figure 3 It is a schematic flowchart of allocating tasks within a task sub-region to computing drones through the market bidding method provided by an embodiment of the present invention.
[0054] Figure 4 This is a schematic flowchart of generating a real-time scheduling strategy for a charging UAV through the DDPG algorithm provided by an embodiment of the present invention. Specific embodiments
[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0056] It should be noted that all directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention are only used to explain the relative position relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly.
[0057] In addition, the descriptions involving "first", "second", etc. in the present invention are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.
[0058] Traditional edge computing distributes data processing capabilities and storage resources at the network edge, close to the data source, to reduce the burden on the remote data center. This method reduces transmission latency, improves data processing speed and response efficiency, and is very suitable for application scenarios that require real-time processing and low latency requirements, such as the Internet of Things and smart cities. However, fixed edge nodes are usually limited in distribution and cannot flexibly handle dynamic and temporary activities.
[0059] To address the above problems, applying UAVs to edge computing can improve the flexibility and coverage of the system. As mobile computing nodes, UAVs can quickly respond when task requirements change, thereby providing computing resources and data collection functions. By carrying computing devices, UAVs can provide mobile edge computing services in a wide geographical space, solve the problem of areas that cannot be covered by fixed nodes, and make the application of edge computing more flexible and extensive.
[0060] However, the scheduling and planning of drones face many challenges. Existing technologies are difficult to accurately allocate tasks according to the rapidly changing task and environmental requirements, resulting in low resource utilization efficiency. Moreover, due to the limited battery life of drones, current technologies have not been able to provide sufficiently efficient charging strategies to ensure the continuity of tasks, which makes it difficult to achieve efficient drone scheduling in a dynamic and uncertain environment.
[0061] As Figure 1 shown, an intelligent scheduling method for low-altitude economic activities includes the steps of:
[0062] Construct a task environment model, obtain a task requirement set, a computing drone set, and a charging drone set; the task environment model includes a task area and several device locations; the computing drone set includes several computing drones; the charging drone set includes several charging drones; the task requirement set includes the ground devices and resource demand quantities corresponding to several tasks;
[0063] Divide several task sub-areas according to the task area, device locations, and the ground devices and resource demand quantities corresponding to the tasks;
[0064] Allocate computing drones to each task sub-area; allocate the tasks within the task sub-area to the computing drones through the market bidding method;
[0065] During the process of the computing drones executing tasks, generate a real-time scheduling strategy for the charging drones through the DDPG algorithm.
[0066] By constructing a task environment model and dividing sub-areas according to task requirements, the present invention can provide clear task partitions for computing drones and charging drones, reduce unnecessary frequent interactions and resource waste between devices, and improve the overall task scheduling efficiency; by subdividing the task area into multiple sub-areas and accurately allocating according to the resource demand quantities, it is ensured that each task sub-area can obtain corresponding computing resources and charging services according to actual needs; during the task execution process, a scheduling strategy for the charging drones is dynamically generated through the DDPG algorithm, enabling the charging drones to adjust the flight path and charging actions according to real-time tasks, thereby ensuring the continuous and efficient operation of the computing drones in a dynamic environment.
[0067] The following further elaborates on an intelligent scheduling method for low-altitude economic activities of the present invention in combination with embodiments:
[0068] An intelligent scheduling method for low-altitude economic activities includes the steps of:
[0069] S1. Construct a task environment model to obtain a task requirement set, a computing drone set, and a charging drone set; the task environment model includes a task area and several device locations; the computing drone set includes several computing drones; the charging drone set includes several charging drones; the task requirement set includes the ground devices corresponding to several tasks and the required resource quantities;
[0070] The device location is represented as ; where h represents the ground device number;
[0071] The task requirement set is represented as , where represents the task serial number, represents the required resource quantity for the q-th task of the h-th ground device;
[0072] The computing drone set is represented as ;
[0073] In an embodiment of the present invention, the charging drone and the computing drone are equipped with a wireless charging component, such as a magnetic resonance or inductive charging system, which can transmit electricity to the computing drone through an electromagnetic field. After the charging drone approaches the computing drone, it maintains a relatively stable and appropriate distance and angle to ensure the efficiency and effect of the wireless charging process. Generally, this distance is between several centimeters and dozens of centimeters. During the implementation of the present invention, the charging drone and the computing drone during the charging process can be regarded as having the same position.
[0074] After ensuring to maintain a relatively stable and appropriate distance and angle, the charging drone activates the wireless charging system to generate an electromagnetic field; the receiving device (such as a receiving coil) on the computing drone converts the electromagnetic field into electrical energy to charge its battery.
[0075] During the charging process, the charging drone continuously monitors the charging status, including voltage, current, and distance. If the computing drone moves or the environment changes, the charging drone will adjust its position and charging parameters in real time to optimize the charging efficiency and safety.
[0076] S2. Divide several task sub-areas according to the task area, device location, and the ground devices and required resource quantities corresponding to the tasks;
[0077] Further, please refer to Figure 2 , the dividing of several task sub-areas according to the task area, device location, and the ground devices and required resource quantities corresponding to the tasks includes the steps of:
[0078] S21. Obtain the total required resource quantity of each ground device according to the task and its corresponding required resource quantity;
[0079] The total resource requirement of the ground equipment is expressed as:
[0080] ;
[0081] Among them, represents the total resource requirement of the h-th ground equipment; represents the resource requirement of the q-th task of the h-th ground equipment; represents the total number of tasks of the h-th ground equipment.
[0082] In this step, the total resources required by each ground equipment are calculated first, aiming to understand the requirements of each ground equipment in the entire mission area and provide basic data for the subsequent division of mission sub-areas.
[0083] S22. Determine the number of mission sub-areas, and initialize the position and velocity of each particle in the particle swarm optimization algorithm; the position of the particle is used to represent the division scheme of the mission sub-areas; the position of the particle is expressed as:
[0084] ,
[0085] Among them, is the number of mission sub-areas; represents the clustering center of the k-th mission sub-area, , and respectively represent the longitude and latitude of the clustering center of the k-th mission sub-area; represents the clustering radius of the k-th mission sub-area.
[0086] In this step, the number of mission sub-areas can be reasonably determined according to factors such as the size of the mission area, the number of charging UAVs, the number and distribution of ground equipment, and the mission frequency. In particular, in an embodiment of the present invention, the number of mission sub-areas is set to the number of charging UAVs, that is, each mission sub-area is assigned a charging UAV to charge the computing UAVs therein.
[0087] The position of each particle represents a division scheme of the mission sub-areas, including the coordinates (longitude, latitude) of the clustering center of each mission sub-area and the clustering radius. The velocity of the particle is a randomly generated small value.
[0088] S23. Calculate the fitness values of all particles, and update the global best position and the personal best position of each particle;
[0089] The fitness value is expressed as: ,
[0090] Among them, represents the distance fitness function; represents the resource balance fitness function; and respectively represent the distance fitness function and the weight parameter of the distance fitness function; represents the distance between the h-th ground device and the k-th clustering center; represents the total number of tasks in the k-th task sub-region;
[0091] The fitness value consists of two parts: the distance fitness function and the resource balance fitness function. In one embodiment, the optimization objective of the particle is to minimize the fitness value, and is set, . In the distance fitness function, by constraining , the distances between all ground devices within each clustering region and the clustering center are selected and summed up to evaluate the clustering result through geographical location. The calculation formula of the resource balance fitness function is based on the Jain's Fairness Index equation, and its value range is , . The closer it is to 1, the more balanced the resource requirements of each clustering region are, that is, the gap between the total sum of the resource requirements of the ground devices in each clustering region is smaller.
[0092] S24. Update the velocity and position of each particle according to the personal best position and the global best position;
[0093] In the particle swarm optimization algorithm, the velocity and position of each particle are updated based on the personal best position and the global best position.
[0094] First, three factors are considered in the velocity update of the particle. The current velocity of the particle is decayed by an inertia weight to maintain a certain inertia; then, the particle adjusts its velocity according to its own best position, moves towards this personal best position, and at the same time combines a random factor to increase the diversity of exploration; finally, the particle also approaches the best position discovered by the group to understand the collective optimal solution. This velocity adjustment ensures that the particle can comprehensively consider the advantages of both the individual and the group to improve the search efficiency.
[0095] The position update of the particle is achieved by adding the updated velocity to the previous position. This means that the particle will move to a new position under the influence of the new velocity, so as to explore possible optimal solutions in the solution space.
[0096] S25. Repeat steps S23 to S24 until the maximum number of iterations is reached or the fitness value is less than the set threshold, and use the global best position as the division scheme of the task sub-region.
[0097] The speed and position are updated by repeatedly executing steps S22 to S23 until the number of iterations reaches a preset value or the fitness value is less than a set threshold, and the particle swarm gradually converges to the optimal solution of the objective function during this process. Finally, the global best position is used as the final task sub-region division scheme.
[0098] In this embodiment, based on the particle swarm optimization algorithm, several task sub-regions are divided according to the task area, equipment positions, and the required quantity of ground equipment and resources corresponding to the tasks; the concentration degree of the positions of the ground equipment in each task sub-region is evaluated based on the distance fitness function; based on the resource balance fitness function, the equilibrium state of the resource requirements in each task sub-region can be effectively evaluated and optimized; the fitness function of the particle swarm optimization algorithm is set as the weighted sum of the distance fitness function and the resource balance fitness function, and the task area division is continuously optimized by updating the speed and position of the particles, so as to comprehensively consider the geographical location and resource requirements, make the distribution of equipment and tasks in the clustering area more reasonable, reduce the extra consumption caused by the distance problem of the UAVs and the UAV scheduling conflicts caused by uneven resource allocation, and improve the overall task execution efficiency.
[0099] S3. Allocate computing UAVs for each task sub-region; allocate the tasks in the task sub-region to the computing UAVs through the market bidding method;
[0100] After the division of the task sub-regions is completed, the equipment positions and the total required quantity of resources in each task sub-region are planned, so as to allocate computing resources, that is, computing UAVs, for the task sub-regions. In this solution, the computing UAVs act as edge computing nodes, and their advantage lies in sinking the computing power to the network edge, approaching the data source, and realizing real-time data processing and analysis, thereby reducing latency and improving the response time. As mobile edge nodes, the computing UAVs are equipped with computing and storage capabilities and can provide computing services within the task sub-regions.
[0101] During the implementation process, the computing UAVs will move to suitable positions according to the division of the task sub-regions and resource requirements to ensure that sufficient computing resources are provided in the areas where computing is required. According to the resource requirements corresponding to the tasks in the task sub-regions and the equipment positions of the ground equipment corresponding to each task, an income evaluation index for executing the tasks can be generated, and the tasks are allocated to the appropriate computing UAVs according to the income evaluation index; after the tasks are allocated, the computing UAVs will directly process the data on-site using their own computing resources and execute the computing tasks.
[0102] In one embodiment, please refer to Figure 3 , the allocation of the tasks in the task sub-region to the computing UAVs through the market bidding method is specifically as follows: The following steps are repeatedly executed until all the tasks in the task requirement set are allocated:
[0103] S31. Obtain the task with the largest unallocated resource demand in the task sub-region of the task demand set, and denote it as the currently to-be-allocated task;
[0104] Obtain the task with the largest unallocated resource demand in the task sub-region from the task demand set, and denote it as the currently to-be-allocated task. In this way, it is ensured that tasks with larger resource demands are processed preferentially, so as to improve resource utilization with a lower computational complexity.
[0105] S32. Calculate a quotation index based on the difference between the expected total revenue after adding the currently to-be-allocated task to the task sequence and the current total revenue of the computing drone, and generate a quotation request according to the quotation index;
[0106] The said quotation index is expressed as:
[0107]
[0108] where i represents the computing drone number, represents the currently to-be-allocated task in the k-th task sub-region; represents the task sequence of the i-th computing drone; represents the expected total revenue after adding the currently to-be-allocated task to the task sequence ; represents the current total revenue of the i-th computing drone.
[0109] S33. Each computing drone broadcasts the quotation request to other computing drones within all task sub-regions;
[0110] In this step, each computing drone broadcasts the generated quotation request to all other computing drones within the task sub-region, enabling other drones to synchronize information and participate in subsequent bidding evaluations, so as to increase transparency and competitiveness, ensure that each task can be undertaken by the most suitable drone, and thus optimize the overall task allocation result.
[0111] S34. After receiving the quotation request, each computing drone generates an evaluation result according to its own quotation index and transmits it to the scheduling center;
[0112] After receiving the quotation request, each computing drone generates an evaluation result according to its own quotation index, and then transmits the evaluation result to the scheduling center. By integrating the evaluation results of each drone, it is ensured that each task can be reasonably allocated to the most suitable drone without affecting its own operation efficiency.
[0113] S35. The scheduling center allocates the currently to-be-allocated task to the corresponding computing drone according to all evaluation results, and updates the task demand set and the task sequence of the computing drone.
[0114] After the scheduling center receives all the evaluation results, it selects the computing UAV with the best evaluation result to assign the currently unassigned task, and simultaneously updates the task requirement set and the task sequence of the computing UAV.
[0115] In this embodiment, the market auction method is adopted to ensure that tasks can be efficiently and reasonably assigned to computing UAVs, so as to improve resource utilization rate, enhance the fairness of task assignment, and improve the overall operation efficiency of the system, thereby realizing the optimal configuration and execution of tasks within the task sub-region.
[0116] In one embodiment, the method for obtaining the expected total revenue after the currently unassigned task is added to the task sequence is as follows:
[0117] S321. Determine the optimal position of the currently unassigned task in the task sequence according to the device position corresponding to the currently unassigned task, and generate a reference task sequence;
[0118] Among them, determining the optimal position of the currently unassigned task in the task sequence according to the device position corresponding to the currently unassigned task aims to find an optimal position in the task sequence to insert the currently unassigned task, so that the execution efficiency of the entire task sequence is optimal (for example, minimizing the total movement time). The specific method for obtaining the optimal position can be: calculate the distance between the device position corresponding to the currently unassigned task and the device positions corresponding to the tasks in all task sequences; insert the currently unassigned task into the task corresponding to the device position with the smallest distance.
[0119] S322. Calculate the expected total revenue according to the scheduling duration between adjacent tasks in the reference task sequence, the task revenue of each task, and the calculated task execution time.
[0120] After generating the reference task sequence, obtain the scheduling duration between each pair of adjacent tasks in the reference task sequence, the task revenue of each task, and the calculated task execution time.
[0121] Among them, the scheduling duration is the movement duration caused by different device positions corresponding to tasks, and can be calculated according to the distance between device positions and the movement performance of UAVs.
[0122] In an edge computing environment, task revenue is a metric that represents the benefits obtained by the system after task execution. This benefit can be in various forms. In one embodiment, the evaluation influencing factors of task revenue include response time, resource utilization rate, user satisfaction, and task execution quality. Specifically, for response time, edge computing deploys computing resources close to the data source, such as user devices, Internet of Things devices, etc. This can significantly reduce the data transmission delay, thereby improving the system's response time; for resource utilization rate, edge computing can make full use of the computing resources of edge devices instead of sending all tasks to a remote server for processing, thus improving the resource utilization rate of the entire system; for user satisfaction, due to the fast response time and high reliability, the user experience is significantly improved; for task execution quality, edge computing can efficiently execute tasks while reducing latency and energy consumption, and at the same time improve the quality and reliability of task completion. When generating task information, task revenue is calculated through the quantified evaluation influencing factors, namely response time, resource utilization rate, user satisfaction, and task execution quality, and is stored in association with the resource requirements of the task. It is directly used to measure and calculate the revenue of edge computing tasks in the implementation of this step.
[0123] S4. During the process of the drone executing the task, generate a real-time scheduling strategy for the charging drone through the DDPG algorithm.
[0124] The DDPG (Deep Deterministic Policy Gradient) algorithm is a deep reinforcement learning algorithm suitable for continuous action spaces. It combines the advantages of policy gradient and Q-learning and is suitable for real-time decision-making in dynamic environments.
[0125] Furthermore, please refer to Figure 4 , the real-time scheduling strategy for the charging drone generated through the DDPG algorithm includes the steps:
[0126] S41. Construct a Markov decision process for the charging drone; the Markov decision process consists of a state space, an action space, a transition probability, a reward function, and a discount factor;
[0127] The Markov decision process (MDP) is a mathematical framework used to model decision-making problems, especially in decision-making processes faced with a stochastic environment. MDP is widely used in fields such as reinforcement learning, control theory, and decision analysis.
[0128] In the present invention, the state space is composed of states corresponding to several time slots; the states include the position, power, and task progress of the computing drone, as well as the position, power reserve, charging state, and moving state of the charging drone;
[0129] The action space consists of actions corresponding to several time slots of the charging UAV; the actions include the moving direction, the charging action, and the charging target.
[0130] The transition probability is used to represent the probability that the current state of the charging UAV transfers to the next possible state after performing an action in the action set.
[0131] The reward function is used to represent the reward obtained when the current state of the charging UAV transfers to the next state after performing an action.
[0132] The discount factor is used to determine the influence function of future rewards on the current reward.
[0133] Specifically, the reward function is expressed as:
[0134] ;
[0135] where is used to represent the reward obtained by the charging UAV within a time slot; represents the charging amount provided by the charging UAV within a time slot; is a preset constant, used to encourage the charging UAV to perform more charging actions rather than perform meaningless movements or hoverings; represents the charging fairness index, , represents the number of computing UAVs in the k-th task sub-region. The charging fairness index is used to encourage the charging UAV to charge each computing UAV fairly. In one embodiment, it is based on the Jain's Fairness Index equation; When, it represents the charging UAV does not perform a charging action within a time slot, When, it represents the charging UAV performs a charging action within a time slot.
[0136] In one embodiment, the number of task sub-regions is set to the number of charging drones, that is, each task sub-region is assigned a charging drone to provide charging services for the computing drones therein. By having each charging drone focus on one sub-region, the complexity of the state and action spaces that the DDPG algorithm needs to process is reduced, improving the learning efficiency and convergence speed of the algorithm. The algorithm only needs to optimize the policy for the state of the current sub-region, avoiding the complex deduction of dealing with the global state. Each charging drone performs tasks within a fixed sub-region, and can collect experience data more stably. These data play an effective role in continuously strengthening the same policy, reducing the frequency of policy adjustment caused by environmental changes, thereby improving the stability and convergence of the algorithm. On the premise of independent optimization in each sub-region, the charging strategies of each drone tend to local optimal solutions. This can ensure that the computing drones in each region can obtain charging services in a timely manner, and at the same time, these local optimal solutions jointly promote the improvement of the global performance, minimizing the overall task completion time.
[0137] S42. Set the optimization objective function for the implementation scheduling strategy of the charging drones;
[0138] The optimization objective function is set to minimize the task completion duration of all computing drones, expressed as:
[0139] ;
[0140] where, represents the task completion duration of the i-th computing drone, represents the number of computing drones in the k-th task sub-region; represents the set of numbers of the computing drones in the k-th task sub-region.
[0141] S43. Use the DDPG algorithm based on the Markov decision process and the optimization objective function to generate the real-time scheduling strategy of the charging drones.
[0142] Further, step S43 includes:
[0143] S431. Initialize the DDPG algorithm parameters: Initialize the parameters of the policy network and the Q network, and generate the corresponding target network; Initialize the experience replay pool for storing the state transitions after the charging drones execute actions in the environment.
[0144] S432. Calculate the reward function: Calculate the immediate reward by real-time evaluating the effects generated after the charging drones' actions.
[0145] The immediate reward function includes four parts: the charging amount provided by the charging drones within a time slot, the charging fairness index, the task completion duration of the computing drones in the task region, and the penalty / reward coefficient for whether to execute the charging action.
[0146] S433. The policy network generates actions: Using the policy network to input the current state and output the actions of the charging UAV in the current time slot, including the moving direction and the charging target. In the exploration phase, appropriate noise is added to increase the exploration of the policy.
[0147] S434. Execute the actions and observe the results: According to the actions output by the policy network, adjust the behavior of the charging UAV and execute the actions in the environment; observe the new state after executing the actions and store the state transition (current state, action, reward, and next state) in the experience replay pool.
[0148] S436. Update the policy and value networks: Randomly sample (state transition data) from the experience replay pool to update the policy and value networks; calculate the target Q value and update the Q network by minimizing the error of the Q value.
[0149] Optimize the policy network through policy gradient ascent so that the actions output by the policy network can maximize the output of the Q network.
[0150] In this embodiment, the scheduling strategy of the charging UAV can be continuously optimized through the DDPG algorithm, enabling it to provide charging services for the computing UAVs more efficiently in a complex dynamic environment, thereby minimizing the overall task completion time. The generation process of this real-time scheduling strategy not only considers the current states and task requirements of each UAV but also effectively balances charging fairness and movement, ensuring the achievement of the global optimization goal of the charging process.
[0151] In several embodiments provided by the present application, it should be understood that the disclosed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the modules can be electrical, mechanical, or other forms.
[0152] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0153] In addition, the functional modules in the various embodiments of the present application may be integrated into one processing module, may exist physically as individual modules, or two or more modules may be integrated into one module. The above-mentioned integrated modules may be implemented in the form of hardware or in the form of software functional modules.
[0154] If the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs.
[0155] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features claimed herein.
Claims
1. An intelligent dispatching method for low-altitude economic activities, characterized by: Includes steps: Construct a task environment model, and obtain a task requirement set, a computing drone set, and a charging drone set; the task environment model includes a task area and several equipment locations; the computing drone set includes several computing drones; the charging drone set includes several charging drones; the task requirement set includes ground equipment and resource requirements corresponding to several tasks; Divide the mission into several sub-areas according to the mission area, equipment location, and ground equipment and resource requirements corresponding to the mission; Allocate computing drones in each task sub-area; allocate tasks in the task sub-area to computing drones through market bidding; During the calculation of the UAV's mission execution, the DDPG algorithm is used to generate the real-time scheduling strategy for the charging UAV; The method of dividing the task sub-areas according to the task area, the equipment location, and the ground equipment and resource requirements corresponding to the task includes the following steps: S21. Obtain the total resource requirement of each ground device according to the task and its corresponding resource requirement; S22, determining the number of task sub-regions, initializing the position and velocity of each particle in the particle swarm optimization algorithm; the position of the particle is used to represent the division scheme of the task sub-region; S23, calculate the fitness values of all particles, update the group's best position and each particle's personal best position; S24, updating the speed and position of each particle according to the individual best position and the group best position; S25. Repeat steps S23 to S24 until the maximum number of iterations is reached or the fitness value is less than a set threshold, and use the optimal position of the group as a division scheme for the task sub-areas.
2. The intelligent dispatching method for low-altitude economic activities according to claim 1 is characterized by: The total resource demand of the ground equipment is expressed as: , in, represents the total resource demand of the hth ground equipment; represents the resource requirement of the qth task of the hth ground equipment; Represents the total number of tasks of the hth ground equipment.
3. The intelligent dispatching method for low-altitude economic activities according to claim 2 is characterized by: The position of the particle is expressed as: , in, is the number of task sub-areas; represents the cluster center of the k-th task sub-region, , and represent the longitude and latitude of the cluster center of the k-th task sub-region respectively; represents the clustering radius of the k-th task sub-region.
4. The intelligent dispatching method for low-altitude economic activities according to claim 3 is characterized by: The fitness value is expressed as: , in, represents the distance fitness function; represents the resource balance fitness function; and Respectively represent the distance fitness function and the weight parameter of the distance fitness function; represents the distance between the hth ground device and the kth cluster center; Represents the total number of tasks in the kth task sub-region.
5. The intelligent dispatching method for low-altitude economic activities according to claim 1 is characterized by: The method of allocating tasks in the task sub-area to computing drones by market bidding is as follows: Repeat the following steps until all tasks in the task requirement set are allocated: Obtain the task with the largest unallocated resource demand in the task sub-region in the task demand set, and record it as the current task to be allocated; Calculate the quotation index based on the difference between the expected total revenue of the current task to be assigned after it is added to the task sequence and the current total revenue of the calculated drone, and generate a quotation request based on the quotation index; Each computing drone broadcasts the quotation request to other computing drones in all task sub-areas; After receiving the quotation request, each computing drone generates an evaluation result based on its own quotation index and transmits it to the dispatch center; The dispatch center assigns the current tasks to be assigned to the corresponding computing drones based on all evaluation results, and updates the task requirement set and the task sequence of the computing drones.
6. The intelligent dispatching method for low-altitude economic activities according to claim 5 is characterized by: The quotation index is expressed as: , Among them, i represents the number of the calculated drone, represents the current tasks to be assigned in the kth task sub-area; represents the task sequence of the i-th computing drone; Indicates that the current task to be assigned is added to the task sequence The expected total return after Represents the current total benefit of the i-th calculated drone.
7. The intelligent dispatching method for low-altitude economic activities according to claim 5 is characterized by: The method for obtaining the expected total benefit after the current task to be assigned is added to the task sequence is as follows: Determine the optimal position of the current task to be assigned in the task sequence according to the device position corresponding to the current task to be assigned, and generate a reference task sequence; The expected total benefit is calculated based on the scheduling duration between adjacent tasks in the reference task sequence, the task benefit of each task, and the calculation task execution time.
8. The intelligent dispatching method for low-altitude economic activities according to claim 1 is characterized by: The real-time scheduling strategy for charging drones generated by the DDPG algorithm includes the following steps: Construct a Markov decision process for charging the drone; the Markov decision process consists of a state space, an action space, a transition probability, a reward function, and a discount factor; Set the optimization objective function of the charging drone implementation scheduling strategy; A real-time scheduling strategy for charging UAVs is generated using the DDPG algorithm based on the Markov decision process and the optimization objective function.
9. The intelligent dispatching method for low-altitude economic activities according to claim 8 is characterized by: The optimization objective function is set to minimize the task completion time of all computing drones; The reward function is expressed as: , in, express Rewards for charging drones during time slots; express The amount of charge provided by the charging drone in the time slot; is a preset constant; represents the charging fairness indicator, , represents the number of drones calculated in the kth task sub-area; When The charging drone did not perform charging action during the time slot. When The charging drone performs charging actions within the time slot.
Citation Information
Patent Citations
Air-ground integrated unmanned aerial vehicle cluster scheduling method, device and system
CN112631326A
Scheduling method for aerial charging of task unmanned aerial vehicle by charging unmanned aerial vehicle
CN114548663A
Multi-access edge computing network task scheduling and resource allocation method and system
CN117459951A
Cited By
Low-altitude economic intelligent scheduling method based on multi-dimensional data fusion and dynamic topological optimization
CN121305932A