An unmanned aerial vehicle group intelligent scheduling method and system based on edge computing

CN122653239APending Publication Date: 2026-08-28山东和贝网络科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610832522.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-10
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

然而,随着无人机节点数量的增加,集中式方法面临状态空间指数膨胀的问题,计算复杂度急剧上升,且要求各无人机向中心节点上传完整的信道与任务数据,难以保障数据隐私

Benefits of technology

[0018]To address the problems described in the background art, this invention receives scheduling instructions and, based on these instructions, determines the UAV group scheduling environment. This environment includes a set of UAV nodes, an edge computing server, and a base station. The UAV node set comprises multiple UAV nodes, each equipped with a task caching unit. Spatial location information is obtained from the UAV node set, and the set is dynamically partitioned into multiple UAV clusters based on this information. Each UAV cluster includes multiple UAV nodes and a cluster leader node. Therefore, this invention, through dynamic cluster partitioning based on polar coordinate parameters, achieves complementary channel conditions. Furthermore, grouping UAV nodes with similar spatial locations into the same cluster ensures transmission efficiency when multiple UAV nodes share communication resources. It also decomposes the large-scale scheduling problem into multiple smaller-scale UAV cluster-level decision problems. For each UAV cluster, the following operations are performed: A cluster state representation is constructed based on the UAV cluster; a pre-built local scheduling model in the cluster's head node is used to perform decision reasoning on the cluster state representation to obtain the cluster scheduling action set. Thus, this invention encodes cluster state information into a unified-dimensional input vector, combines it with a policy network of a near-end policy optimization algorithm for decision reasoning, and utilizes a radix transformation rule to efficiently group the cluster... The hierarchical action index is decomposed into independent scheduling instructions for each UAV node, realizing fine-grained joint scheduling decisions within the cluster. Task scheduling is performed on the UAV node set based on the cluster scheduling action set, resulting in a scheduling feedback dataset. This dataset is then used to update the local scheduling model, yielding an updated local scheduling model. Model parameters are obtained based on this updated model. It is evident that this invention constrains the policy update magnitude through the pruning mechanism of the near-end policy optimization algorithm and fully utilizes empirical data through multiple small-batch sampling, achieving stable and efficient parameter optimization of the local scheduling model. The model parameters are then aggregated to obtain a model parameter set, which is then processed using an edge computing server. The data sets are aggregated to obtain global scheduling model parameters. These parameters are then broadcast to the head node of each of the multiple drone clusters. The local scheduling model is updated with these global parameters, and the process is repeated for each drone cluster until the scheduling feedback dataset meets the preset convergence conditions. This achieves intelligent scheduling of the drone groups. It is evident that this invention determines the convergence state by using the average of iterative fluctuation values ​​within a sliding window. This avoids misjudgments caused by single fluctuations and promptly captures the stable trend of the scheduling strategy, ensuring that the model terminates training after full convergence. This balances the quality of the scheduling strategy with training efficiency. Therefore, this invention can optimize task scheduling between multiple drones and edge servers, reducing packet loss and energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122653239A_ABST
    Figure CN122653239A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of unmanned aerial vehicle communication scheduling, and relates to an unmanned aerial vehicle group intelligent scheduling method based on edge computing, which comprises the following steps: receiving a scheduling instruction, confirming an unmanned aerial vehicle group scheduling environment based on the scheduling instruction, obtaining a spatial position information set by using an unmanned aerial vehicle node set, performing dynamic cluster division on the unmanned aerial vehicle node set based on the spatial position information set to obtain a plurality of unmanned aerial vehicle clusters, performing task scheduling on the unmanned aerial vehicle node set based on a cluster scheduling action set to obtain a scheduling feedback data set, obtaining model parameters based on updating a local scheduling model, and aggregating the model parameter set by using an edge computing server to obtain global scheduling model parameters. The application can optimize task scheduling between a plurality of unmanned aerial vehicles and an edge server, reduce packet loss, and reduce energy consumption.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) communication and scheduling technology, and in particular to an intelligent scheduling method for UAV groups based on edge computing. Background Technology

[0002] With the large-scale deployment of drone technology in scenarios such as emergency rescue, environmental monitoring, and logistics delivery, the task processing requirements of drone clusters have evolved from single-machine independent computing to a hybrid processing mode of multi-machine collaboration and edge offloading. Edge computing, by deploying computing servers on the base station side, provides a feasible path for task offloading for drone nodes with limited computing resources, but it also introduces multi-dimensional scheduling challenges such as channel resource competition, inter-cluster interference management, and energy consumption constraints.

[0003] Existing UAV mission scheduling methods mostly employ centralized optimization strategies, where a central node collects the state information of all UAVs and calculates the globally optimal scheduling scheme. However, as the number of UAV nodes increases, centralized methods face the problem of exponential expansion of the state space, a sharp increase in computational complexity, and the requirement for each UAV to upload complete channel and mission data to the central node, making it difficult to guarantee data privacy. Although some studies have introduced reinforcement learning to achieve adaptive decision-making, they still mainly rely on single-agent modeling and fail to effectively decompose the joint scheduling problem of large-scale UAV groups.

[0004] Furthermore, existing methods typically neglect the complementary relationship of channel conditions among UAV nodes in scheduling decisions and fail to establish a clustered cooperation mechanism based on spatial location and channel characteristics, resulting in inefficient management of co-channel interference and low utilization of communication resources. At the model training level, the scheduling experience of each UAV node is isolated, lacking cross-node strategy sharing and collaborative optimization methods, leading to slow model convergence and insufficient generalization ability. Therefore, how to achieve dynamic clustering of UAV groups, refined scheduling decisions within the cluster, and cross-cluster collaborative model optimization while ensuring data privacy has become an urgent technical problem to be solved. Summary of the Invention

[0005] This invention provides an intelligent scheduling method for drone groups based on edge computing and a computer-readable storage medium. Its main purpose is to optimize task scheduling between multiple drones and edge servers, reduce packet loss, and lower energy consumption.

[0006] To achieve the above objectives, the present invention provides an intelligent scheduling method for unmanned aerial vehicle (UAV) groups based on edge computing, comprising: The system receives a scheduling instruction and determines the drone group scheduling environment based on the instruction. The drone group scheduling environment includes a drone node set, an edge computing server, and a base station. The drone node set includes multiple drone nodes, and each drone node has a task cache unit. A set of spatial location information is obtained by using a set of UAV nodes. Based on the set of spatial location information, the set of UAV nodes is dynamically divided into clusters to obtain multiple UAV clusters. Each UAV cluster includes multiple UAV nodes and a cluster head node. For each of the multiple drone clusters, the following operations are performed: a cluster state representation is constructed based on the drone cluster, and a decision reasoning is performed on the cluster state representation using a pre-built local scheduling model in the cluster head node to obtain a cluster scheduling action set; Task scheduling is performed on the UAV node set based on the cluster scheduling action set to obtain a scheduling feedback dataset. The local scheduling model is updated using the scheduling feedback dataset to obtain an updated local scheduling model. Model parameters are obtained based on the updated local scheduling model. The model parameters are summarized to obtain a model parameter set. The model parameter set is then aggregated using an edge computing server to obtain global scheduling model parameters. These global scheduling model parameters are broadcast to the first node of each of the multiple UAV clusters. The local scheduling model is updated with the global scheduling model parameters, and then the process is repeated for each of the multiple UAV clusters. This process continues until the scheduling feedback dataset is confirmed to meet the preset convergence conditions, thus achieving intelligent scheduling of the UAV groups.

[0007] Optionally, the dynamic clustering of the UAV node set based on the spatial location information set to obtain multiple UAV clusters includes: Based on the spatial location information set, the polar coordinate parameters of each UAV node relative to the base station are obtained, resulting in multiple polar coordinate parameters, including: channel gain sub-interval identifier and quantized azimuth angle. Each drone node in the drone node set is initialized to obtain multiple initial clusters, where each initial cluster corresponds one-to-one with the polar coordinate parameters; For each of the multiple initial clusters, perform the following operation: The initial cluster is removed from multiple initial clusters to obtain a set of remaining initial clusters. Based on multiple polar coordinate parameters, an initial cluster that matches the initial cluster is retrieved from the set of remaining initial clusters to obtain the retrieval results, where the retrieval results indicate whether the cluster exists or does not exist. If the search result indicates that a candidate initial cluster set exists, then a candidate initial cluster set is obtained from the remaining initial cluster set, and the following operation is performed on each candidate initial cluster set: Calculate the absolute difference in quantized azimuth angles between the initial cluster and the candidate initial clusters to obtain the absolute angle difference. Summarize the absolute angle differences and arrange them in ascending order to obtain the sequence of absolute angle differences. Select the candidate initial cluster corresponding to the absolute angle difference with the position number 1 as the matching initial cluster. Merge the initial cluster with the matching initial cluster into the target cluster; otherwise, mark the initial cluster as an unmerged initial cluster. By summarizing the target clusters, a target cluster set is obtained; Obtain the drone nodes in the unmerged initial cluster to obtain the set of drone nodes to be assigned; For each drone node in the set of drone nodes to be assigned, the following operations are performed: Candidate drone clusters are obtained based on the drone nodes to be assigned and the target cluster set; By aggregating the candidate drone clusters, multiple candidate drone clusters are obtained; After confirming that the number of drone nodes in each of the multiple candidate drone clusters meets the preset cluster capacity range, multiple drone clusters are obtained, and a cluster leader node is determined in each drone cluster.

[0008] Optionally, the construction of cluster state representation based on drone swarm includes: Obtain the cache queue vector, quantization channel gain, and battery energy level of each drone node in the drone cluster to obtain the cache queue vector set, quantization channel gain set, and battery energy level set. Associate the cache queue vector set, quantization channel gain set, and battery energy level set to obtain the cluster internal information group. The drone clusters are removed from the multiple drone clusters to obtain a set of remaining drone clusters. The set of remaining drone clusters includes multiple remaining drone clusters. The average channel gain of all drone nodes in the set of remaining drone clusters is obtained to obtain a set of statistical channel parameters. The set of statistical channel parameters includes multiple statistical channel parameters, and the statistical channel parameters correspond one-to-one with the remaining drone clusters. By associating the cluster's internal information group with the statistical channel parameter set, an initial state representation is obtained; The number of drone nodes in the drone cluster is counted to obtain the node count. If the node count is less than the preset node count limit, a zero-value filling operation is performed on the initial state representation to obtain the cluster state representation.

[0009] Optionally, the step of using a pre-built local scheduling model in the cluster head node to perform decision reasoning on the cluster state representation to obtain a cluster scheduling action set includes: The policy network and value network are obtained based on the local scheduling model. The cluster state representation is input into the policy network to obtain the scheduling policy distribution; Based on the distribution of the scheduling strategy, action sampling is performed to obtain the cluster scheduling action index. The cluster scheduling action index is decomposed into multiple scheduling action nodes using a preset radix conversion rule to obtain the cluster scheduling action set. The scheduling action node includes the execution type and the amount of execution data, and the execution type is idle, local processing, or unloading processing.

[0010] Optionally, the step of performing task scheduling on the UAV node set based on the cluster scheduling action set to obtain a scheduling feedback dataset includes: For each scheduling action node in the cluster scheduling action set, the following operations are performed: After confirming that the execution type corresponding to the scheduling action node is local processing, the local energy consumption consumed by the UAV node in processing the amount of execution data is obtained, and the task cache unit of the UAV node is reduced based on the amount of execution data to obtain the first reduced cache queue. After confirming that the execution type corresponding to the scheduling action node is offloading processing, the uplink transmission rate when the UAV node transmits the execution data to the edge computing server is obtained. Based on the uplink transmission rate, the transmission delay and offloading energy consumption are obtained. After confirming that the transmission delay meets the preset time slot constraint, the execution data is offloaded to the edge computing server via the base station for processing. Based on the execution data, the task cache unit is reduced to obtain the second reduced cache queue. Based on the number of data packets dropped in the task cache unit after the first or second reduction of the cache queue, a packet loss statistics node is obtained, wherein the packet loss statistics node includes the number of packet loss due to latency violation and the number of packet loss due to cache overflow. Based on the packet loss statistics node and the local energy consumption or offload energy consumption, the scheduling feedback data is calculated and summarized to obtain the scheduling feedback dataset.

[0011] Optionally, the uplink transmission rate when the drone node transmits the execution data to the edge computing server includes: The offloaded transmit power of the UAV node is obtained. The UAV node is removed from the UAV cluster corresponding to the UAV node to obtain the remaining UAV node set. The intra-cluster interference power is obtained based on the remaining UAV node set. The inter-cluster interference power is estimated based on the statistical channel parameter set. The uplink transmission rate is calculated using the offloaded transmit power, quantized channel gain, intra-cluster interference power, and inter-cluster interference power estimates. The calculation formula is as follows: in, This indicates the uplink transmission rate. This indicates the preset uplink transmission bandwidth. This indicates the offloaded transmission power. This represents the quantized channel gain. This indicates the interference power within the cluster. This represents the estimated inter-cluster interference power. This represents the preset noise power spectral density. Indicates the uplink transmission identifier. Indicates an identifier within the cluster. Indicates an inter-cluster identifier. This represents the unload identifier.

[0012] Optionally, updating the local scheduling model using the scheduling feedback dataset to obtain an updated local scheduling model includes: Obtain the cache tuple set based on the scheduling feedback dataset and the pre-built experience cache area; Randomly sample from the cached tuple set using a preset small-batch sampling ratio to obtain a small-batch tuple set. Calculate the policy ratio and pruning objective function value based on the small-batch tuple set. Update the parameters of the policy network and value network based on the policy ratio and pruning objective function value, and count the number of iterations to obtain the initial update scheduling model and the number of iterations. If the number of iterations is less than a preset threshold, then return to the step of randomly sampling the cached tuple set using a preset small batch sampling ratio until the number of iterations equals the threshold, and use the initial update scheduling model as the update local scheduling model.

[0013] Optionally, the step of aggregating the model parameter set using an edge computing server to obtain global scheduling model parameters includes: Obtain the model parameters of the updated local scheduling model of each cluster leader node in multiple drone clusters to obtain the model parameter set; The model parameter set is aggregated by using an edge computing server to obtain the global scheduling model parameters. The calculation formula is shown below: in, This represents the parameters of the global scheduling model. This indicates that there are a total of [number] drones in the aforementioned drone cluster. A drone swarm, The first parameter in the model parameter set Each model parameter.

[0014] Optionally, the confirmation scheduling feedback dataset satisfies a preset convergence condition, including: Extract scheduling feedback data sequentially from the scheduling feedback dataset, and perform the following operations on the extracted scheduling feedback data: Using the scheduling feedback data, the analytical scheduling feedback data is identified in the scheduling feedback dataset. The analytical scheduling feedback data is adjacent to and lags behind the scheduling feedback data. The absolute difference between the analytical scheduling feedback data and the scheduling feedback data is calculated to obtain the iterative fluctuation value. The iterative fluctuation values ​​are summarized to obtain an iterative fluctuation value set. An evaluation fluctuation value set is extracted from the iterative fluctuation value set using a preset fluctuation evaluation window. The mean of the evaluation fluctuation values ​​in the evaluation fluctuation value set is calculated to obtain the evaluation fluctuation mean. The average value of the evaluated fluctuation is compared with a preset convergence threshold. If the average value of the evaluated fluctuation is less than or equal to the convergence threshold, it is confirmed that the scheduling feedback dataset meets the convergence condition.

[0015] To achieve the above objectives, the present invention also provides an intelligent scheduling system for unmanned aerial vehicle (UAV) groups based on edge computing, comprising: The environment confirmation module is used to receive scheduling instructions and confirm the UAV group scheduling environment based on the scheduling instructions. The UAV group scheduling environment includes a UAV node set, an edge computing server and a base station. The UAV node set includes multiple UAV nodes and each UAV node has a task cache unit. The cluster partitioning module is used to obtain a set of spatial location information using the set of UAV nodes, and to dynamically partition the set of UAV nodes into multiple UAV clusters based on the set of spatial location information. Each UAV cluster includes multiple UAV nodes and a cluster head node. The scheduling decision module is used to perform the following operations on each of the multiple drone clusters: construct a cluster state representation based on the drone cluster, and use the pre-built local scheduling model in the cluster head node to make decision reasoning on the cluster state representation to obtain a cluster scheduling action set; Task scheduling is performed on the UAV node set based on the cluster scheduling action set to obtain a scheduling feedback dataset. The local scheduling model is updated using the scheduling feedback dataset to obtain an updated local scheduling model. Model parameters are obtained based on the updated local scheduling model. The global aggregation module is used to summarize model parameters to obtain a model parameter set. The model parameter set is aggregated using an edge computing server to obtain global scheduling model parameters. The global scheduling model parameters are broadcast to the first node of each of the multiple UAV clusters. After updating the local scheduling model with the global scheduling model parameters, the module returns to the previous state. The following steps are performed on each of the multiple UAV clusters until the scheduling feedback dataset is confirmed to meet the preset convergence conditions, thereby realizing intelligent scheduling of the UAV group.

[0016] To address the above problems, the present invention also provides an electronic device, the electronic device comprising: Memory, storing at least one instruction; The processor executes the instructions stored in the memory to implement the aforementioned intelligent scheduling method for unmanned aerial vehicle (UAV) groups based on edge computing.

[0017] To address the aforementioned issues, the present invention also provides a computer-readable storage medium storing at least one instruction, which is executed by a processor in an electronic device to implement the aforementioned edge computing-based intelligent scheduling method for unmanned aerial vehicle (UAV) groups.

[0018] To address the problems described in the background art, this invention receives scheduling instructions and, based on these instructions, determines the UAV group scheduling environment. This environment includes a set of UAV nodes, an edge computing server, and a base station. The UAV node set comprises multiple UAV nodes, each equipped with a task caching unit. Spatial location information is obtained from the UAV node set, and the set is dynamically partitioned into multiple UAV clusters based on this information. Each UAV cluster includes multiple UAV nodes and a cluster leader node. Therefore, this invention, through dynamic cluster partitioning based on polar coordinate parameters, achieves complementary channel conditions. Furthermore, grouping UAV nodes with similar spatial locations into the same cluster ensures transmission efficiency when multiple UAV nodes share communication resources. It also decomposes the large-scale scheduling problem into multiple smaller-scale UAV cluster-level decision problems. For each UAV cluster, the following operations are performed: A cluster state representation is constructed based on the UAV cluster; a pre-built local scheduling model in the cluster's head node is used to perform decision reasoning on the cluster state representation to obtain the cluster scheduling action set. Thus, this invention encodes cluster state information into a unified-dimensional input vector, combines it with a policy network of a near-end policy optimization algorithm for decision reasoning, and utilizes a radix transformation rule to efficiently group the cluster... The hierarchical action index is decomposed into independent scheduling instructions for each UAV node, realizing fine-grained joint scheduling decisions within the cluster. Task scheduling is performed on the UAV node set based on the cluster scheduling action set, resulting in a scheduling feedback dataset. This dataset is then used to update the local scheduling model, yielding an updated local scheduling model. Model parameters are obtained based on this updated model. It is evident that this invention constrains the policy update magnitude through the pruning mechanism of the near-end policy optimization algorithm and fully utilizes empirical data through multiple small-batch sampling, achieving stable and efficient parameter optimization of the local scheduling model. The model parameters are then aggregated to obtain a model parameter set, which is then processed using an edge computing server. The data sets are aggregated to obtain global scheduling model parameters. These parameters are then broadcast to the head node of each of the multiple drone clusters. The local scheduling model is updated with these global parameters, and the process is repeated for each drone cluster until the scheduling feedback dataset meets the preset convergence conditions. This achieves intelligent scheduling of the drone groups. It is evident that this invention determines the convergence state by using the average of iterative fluctuation values ​​within a sliding window. This avoids misjudgments caused by single fluctuations and promptly captures the stable trend of the scheduling strategy, ensuring that the model terminates training after full convergence. This balances the quality of the scheduling strategy with training efficiency. Therefore, this invention can optimize task scheduling between multiple drones and edge servers, reducing packet loss and energy consumption. Attached Figure Description

[0019] Figure 1This is a flowchart illustrating an intelligent scheduling method for unmanned aerial vehicle (UAV) groups based on edge computing, provided in an embodiment of the present invention. Figure 2 A functional block diagram of an intelligent scheduling system for unmanned aerial vehicle (UAV) groups based on edge computing, provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device that implements the intelligent scheduling method for unmanned aerial vehicle (UAV) groups based on edge computing, according to an embodiment of the present invention.

[0020] Explanation of reference numerals in the attached figures: 10. Electronic device; 11. Processor; 12. Memory; 13. Bus.

[0021] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0022] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0023] This application provides an intelligent scheduling method for unmanned aerial vehicle (UAV) squadrons based on edge computing. The executing entity of this intelligent scheduling method includes, but is not limited to, at least one of the following electronic devices that can be configured to execute the method provided in this application: a server, a terminal, etc. In other words, the intelligent scheduling method for UAV squadrons based on edge computing can be executed by software or hardware installed on a terminal device or a server device, and the software can be a blockchain platform. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster.

[0024] Reference Figure 1 The diagram shown is a flowchart illustrating an intelligent scheduling method for unmanned aerial vehicle (UAV) groups based on edge computing, according to an embodiment of the present invention. In this embodiment, the intelligent scheduling method for UAV groups based on edge computing includes: S1. Receive scheduling instructions and confirm the UAV group scheduling environment based on the scheduling instructions. The UAV group scheduling environment includes a set of UAV nodes, an edge computing server and a base station. The set of UAV nodes includes multiple UAV nodes and each UAV node has a task cache unit.

[0025] It should be explained that the scheduling command is a control command issued by the scheduling administrator or the upper-level task management system to initiate the intelligent scheduling process of the UAV group and trigger subsequent operations such as cluster partitioning, task unloading, and model training. The UAV group scheduling environment is the software and hardware working environment required to realize the intelligent scheduling function of the UAV group. The UAV node set is the collection of all UAVs participating in this scheduling task. The UAV node set includes multiple UAV nodes, each representing a UAV device with computing, communication, and flight capabilities. The edge computing server is a server device with computing and storage resources deployed near the base station, used to receive and execute the computing tasks unloaded by the UAVs. The base station is a relay device for wireless communication between the UAVs and the edge computing server, used to forward data and provide network access services. The task cache unit is a limited-capacity cache area in the UAV node used to store computing tasks to be processed. The task cache unit uses a queue structure to manage task data packets, that is, newly arriving task data packets enter the tail of the queue, and processed task data packets are removed from the queue.

[0026] For example, in a disaster relief scenario, the command center issues a dispatch instruction to the drone group and then confirms the drone group dispatch environment, which includes a drone node set consisting of 6 drones with image acquisition and data processing capabilities, a base station with a coverage radius of 500 meters, and an edge computing server deployed near the base station. Each drone is equipped with a task cache unit with a capacity of 4 data packets to cache the acquired image recognition tasks.

[0027] S2. Obtain a spatial location information set using the UAV node set, and dynamically divide the UAV node set into multiple UAV clusters based on the spatial location information set. Each UAV cluster includes multiple UAV nodes and a cluster head node.

[0028] It should be understood that the spatial location information set is a set of location data obtained by the positioning module onboard each drone node in the drone node set, containing the spatial orientation information of each drone node relative to the base station. The dynamic cluster partitioning is the process of dividing the drone node set into multiple functional clusters based on the spatial location information of the drone nodes. Drone nodes within each drone cluster have similar communication conditions, which can improve communication efficiency and reduce the complexity of scheduling decisions. The drone cluster is a group of drone nodes obtained after dynamic cluster partitioning, and each drone cluster participates in the subsequent scheduling process as an independent decision-making unit. The cluster leader node is a drone node randomly selected in each drone cluster, responsible for collecting the status information of all drone nodes within its cluster and executing local scheduling decisions.

[0029] Understandably, since the channel conditions and spatial locations of each drone node in a drone node set are different, directly scheduling all drone nodes in a unified manner would lead to an excessively large state space and difficulties in interference management. Therefore, it is necessary to dynamically partition the drone node set based on spatial location information to reduce the complexity of the problem. The dynamic partitioning of the drone node set based on the spatial location information set results in multiple drone clusters, including: Based on the spatial location information set, the polar coordinate parameters of each UAV node relative to the base station are obtained, resulting in multiple polar coordinate parameters, including: channel gain sub-interval identifier and quantized azimuth angle. Each drone node in the drone node set is initialized to obtain multiple initial clusters, where each initial cluster corresponds one-to-one with the polar coordinate parameters; For each of the multiple initial clusters, perform the following operation: The initial cluster is removed from multiple initial clusters to obtain a set of remaining initial clusters. Based on multiple polar coordinate parameters, an initial cluster that matches the initial cluster is retrieved from the set of remaining initial clusters to obtain the retrieval results, where the retrieval results indicate whether the cluster exists or does not exist. If the search result indicates that a candidate initial cluster set exists, then a candidate initial cluster set is obtained from the remaining initial cluster set, and the following operation is performed on each candidate initial cluster set: Calculate the absolute difference in quantized azimuth angles between the initial cluster and the candidate initial clusters to obtain the absolute angle difference. Summarize the absolute angle differences and arrange them in ascending order to obtain the sequence of absolute angle differences. Select the candidate initial cluster corresponding to the absolute angle difference with the position number 1 as the matching initial cluster. Merge the initial cluster with the matching initial cluster into the target cluster; otherwise, mark the initial cluster as an unmerged initial cluster. By summarizing the target clusters, a target cluster set is obtained; Obtain the drone nodes in the unmerged initial cluster to obtain the set of drone nodes to be assigned; For each drone node in the set of drone nodes to be assigned, the following operations are performed: Candidate drone clusters are obtained based on the drone nodes to be assigned and the target cluster set; By aggregating the candidate drone clusters, multiple candidate drone clusters are obtained; After confirming that the number of drone nodes in each of the multiple candidate drone clusters meets the preset cluster capacity range, multiple drone clusters are obtained, and a cluster leader node is determined in each drone cluster.

[0030] In detail, the polar coordinate parameters are parameters describing the spatial position of the UAV node using a polar coordinate system with the base station as the origin. They include two components: channel gain sub-interval identifier and quantized azimuth angle. The channel gain sub-interval identifier is the sub-interval number into which the UAV node's current channel gain falls after dividing the continuous range of channel gain values ​​into multiple discrete sub-intervals. Different sub-interval numbers correspond to different signal transmission quality levels. Optionally, the channel gain range can be divided into three sub-intervals, corresponding to low-quality, medium-quality, and high-quality signal transmission conditions, respectively. The channel gain sub-interval identifier is determined as follows: First, the channel gain value between the UAV node and the base station is obtained, measured by the communication module on the UAV node. Then, the channel gain range is divided into multiple sub-intervals, each defined by its minimum and maximum values. Finally, the measured channel gain value is compared with the range of each sub-interval to determine the sub-interval number into which the channel gain value falls, which is the channel gain sub-interval identifier. The quantized azimuth angle is the discrete angle value corresponding to the azimuth angle of the UAV node relative to the base station after discretizing the angle space. The method for confirming the quantized azimuth angle is as follows: First, the azimuth angle of the UAV node relative to the base station is calculated based on the spatial location information of the UAV node. Then, a pre-constructed quantization function is used to map the continuous azimuth angles to the nearest discrete angle value to obtain the quantized azimuth angle. The quantization function is a mapping rule that performs a rounding operation on the azimuth angle according to a preset angle quantization step size. The angle quantization step size is equal to the angle range divided by the number of discrete angle values. The calculation method of the quantization function is as follows: the azimuth angle is divided by the angle quantization step size and then rounded to the nearest integer. The rounded result is then multiplied by the angle quantization step size to obtain the quantized azimuth angle.

[0031] Furthermore, the initialization involves treating each drone node in the drone node set as an independent cluster, ensuring that the number of initial clusters equals the number of drone nodes, facilitating subsequent initial cluster merging. The initial cluster is a cluster composed of individual drone nodes after initialization. Matching occurs when drone nodes in two initial clusters are located in different channel gain sub-intervals and have the same or adjacent quantization azimuth angles, i.e., satisfying the conditions that the channel gain sub-interval identifiers are different and the absolute difference in quantization azimuth angles does not exceed the angle quantization step size. The matched initial cluster is the candidate initial cluster with the smallest absolute angle difference from the initial cluster in the remaining initial cluster set. When multiple candidate initial clusters satisfy the matching conditions, the one with the smallest absolute angle difference is selected for merging. The target cluster is the drone cluster formed after merging the initial cluster and the matched initial cluster. The unmerged initial cluster is the initial cluster that failed to find a matching object.

[0032] It should be understood that the set of drone nodes to be assigned is a collection of drone nodes from all unmerged initial clusters. The process of obtaining candidate drone clusters based on the drone nodes to be assigned and the target cluster set is as follows: traverse each target cluster in the target cluster set, determine whether the absolute difference between the quantized azimuth angle of the drone node to be assigned and all drone nodes in the target cluster does not exceed the angle quantization step size, and whether the number of drone nodes in the target cluster is less than the upper limit of the preset cluster capacity range. If both conditions are met, the drone node to be assigned is added to that target cluster, resulting in a candidate drone cluster. If the drone node to be assigned cannot be added to any target cluster, then a target cluster that meets the conditions is searched from the target cluster set, and a drone node whose angle is adjacent to that of the drone node to be assigned is separated from it. The two together form a new candidate drone cluster. This operation requires that the number of drone nodes in the original target cluster after separation is not less than the lower limit of the cluster capacity range. The cluster capacity range is the allowable range of the number of drone nodes in each cluster, for example, a lower limit of 2 and an upper limit of 4. The method for determining the first node of the cluster is: randomly select a drone node within each drone cluster as the first node of the cluster.

[0033] For example, assume a drone node set contains 6 drone nodes, denoted as drone node 1 to drone node 6. The channel gain range is divided into 3 sub-intervals (sub-interval 1, sub-interval 2, and sub-interval 3, corresponding to low, medium, and high signal transmission quality, respectively), and the angle space is quantized into 4 discrete angles (0 degrees, 90 degrees, 180 degrees, and 270 degrees). The polar coordinate parameters of each drone node are obtained after measurement as follows: drone node 1's polar coordinate parameter is (sub-interval 3, 0 degrees), drone node 2's is (sub-interval 1, 0 degrees), drone node 3's is (sub-interval 2, 90 degrees), drone node 4's is (sub-interval 3, 90 degrees), drone node 5's is (sub-interval 1, 180 degrees), and drone node 6's is (sub-interval 2, 180 degrees). First, the 6 drone nodes are initialized to obtain 6 initial clusters, each containing 1 drone node. Taking the initial cluster containing UAV node 1 as an example, candidate initial clusters with different channel gain sub-interval identifiers and an absolute difference in quantization azimuth angles not exceeding 90 degrees are searched in the remaining initial cluster set. The initial cluster containing UAV node 2 (sub-interval 1, 0 degrees) is found as the matching initial cluster, with an absolute angle difference of 0 degrees. The two are merged into target cluster A, which includes UAV node 1 and UAV node 2. Similarly, UAV node 3 (sub-interval 2, 90 degrees) and UAV node 4 (sub-interval 3, 90 degrees) are merged into target cluster B, and UAV node 5 (sub-interval 1, 180 degrees) and UAV node 6 (sub-interval 2, 180 degrees) are merged into target cluster C. All UAV nodes have been merged, and there are no UAV nodes to be assigned. It is confirmed that the number of UAV nodes in target clusters A, B, and C are all 2, which meets the cluster capacity range (lower limit 2, upper limit 4), resulting in 3 UAV clusters. A UAV node is randomly selected from each UAV cluster as the cluster leader node. This invention, through dynamic cluster partitioning based on polar coordinate parameters, groups UAV nodes with complementary channel conditions and similar spatial locations into the same cluster. This not only ensures transmission efficiency when multiple UAV nodes share communication resources, but also decomposes the large-scale scheduling problem into multiple small-scale UAV cluster-level decision problems.

[0034] S3. For each of the multiple drone clusters, perform the following operations: construct a cluster state representation based on the drone cluster, and use the pre-built local scheduling model in the first node of the cluster to perform decision reasoning on the cluster state representation to obtain a cluster scheduling action set.

[0035] It should be explained that the cluster state representation is a vector representation obtained by mathematically encoding the current operating state of the UAV cluster, and is used as input to the local scheduling model. The local scheduling model is a deep reinforcement learning model deployed in the cluster's head node, used to generate scheduling decisions based on the cluster state representation. Optionally, a near-end policy optimization algorithm can be used to achieve this. The decision reasoning is the process of inputting the cluster state representation into the local scheduling model for computation to obtain the scheduling decision. The cluster scheduling action set is the set of scheduling decisions output by the local scheduling model, containing scheduling instructions for each UAV node in the UAV cluster.

[0036] Understandably, in order to comprehensively reflect the operational status of the drone swarm and provide sufficient information for scheduling decisions, it is necessary to comprehensively consider the buffer status, channel conditions, battery status of each drone node within the swarm, as well as interference information between swarms. Therefore, the construction of a swarm status representation based on the drone swarm includes: Obtain the cache queue vector, quantization channel gain, and battery energy level of each drone node in the drone cluster to obtain the cache queue vector set, quantization channel gain set, and battery energy level set. Associate the cache queue vector set, quantization channel gain set, and battery energy level set to obtain the cluster internal information group. The drone clusters are removed from the multiple drone clusters to obtain a set of remaining drone clusters. The set of remaining drone clusters includes multiple remaining drone clusters. The average channel gain of all drone nodes in the set of remaining drone clusters is obtained to obtain a set of statistical channel parameters. The set of statistical channel parameters includes multiple statistical channel parameters, and the statistical channel parameters correspond one-to-one with the remaining drone clusters. By associating the cluster's internal information group with the statistical channel parameter set, an initial state representation is obtained; The number of drone nodes in the drone cluster is counted to obtain the node count. If the node count is less than the preset node count limit, a zero-value filling operation is performed on the initial state representation to obtain the cluster state representation.

[0037] Specifically, the buffer queue vector is the state vector of each slot in the task buffer unit of the UAV node. Each component of the vector records the waiting time of the data packet in the corresponding slot, and empty slots are marked with -1. The quantized channel gain is the value obtained by mapping the continuous channel gain value between the UAV node and the base station to a discrete value space through a quantization function. The quantized channel gain is obtained by first measuring the channel gain value between the UAV node and the base station through the communication module, and then using a preset quantization function to map the channel gain value to discrete values ​​in the corresponding sub-interval to obtain the quantized channel gain. When the quantized channel gain is expressed in dB, it needs to be converted into a linear quantized channel gain before being substituted into the uplink transmission rate formula. The battery energy level is the number of remaining energy units in the current battery of the UAV node. The buffer queue vector set is the set of buffer queue vectors of all UAV nodes in the UAV cluster. The quantized channel gain set is the set of quantized channel gains of all UAV nodes in the UAV cluster. The battery energy level set is the set of battery energy levels of all UAV nodes in the UAV cluster. The cluster internal information group is a data structure obtained by orderly associating the cache queue vector set, quantization channel gain set, and battery energy level set, which fully describes the task load, communication quality, and energy reserve status of all UAV nodes in the cluster.

[0038] Further, the remaining drone swarm set is the set of drone swarms obtained by removing the drone swarm currently constructing its state representation from multiple drone swarms. The statistical channel parameter is the arithmetic mean of the channel gain of all drone nodes in a given remaining drone swarm, used to estimate inter-swarm interference without sharing precise channel information. The statistical channel parameter set is a collection of statistical channel parameters corresponding to all remaining drone swarms. The initial state representation is comprehensive state data obtained by associating the cluster's internal information group with the statistical channel parameter set. The node number limit is the maximum number of drone nodes that each drone swarm can accommodate. The zero-value filling operation is the process of filling the missing drone nodes in the cluster state representation with specific default values ​​when the actual number of drone nodes in the drone swarm is less than the node number limit. Specifically, the filling method is as follows: the buffer queue vector is filled with all -1 vectors (indicating that the buffer is empty), the quantization channel gain is filled with the minimum value of the channel gain range (indicating the worst channel conditions), and the battery energy level is filled with 0 (indicating no remaining energy). For drone swarm information missing from the statistical channel parameter set, it is filled with the maximum average value of the channel gain range (to avoid being included in the interference estimation calculation). The purpose of the zero-value padding operation is to give the state representation of drone clusters of different sizes a unified dimension, thereby ensuring that the local scheduling model can accept input from different drone clusters without changing the network structure.

[0039] It should be understood that, in order to accurately generate scheduling instructions for each drone node, the cluster state representation needs to be input into a local scheduling model containing a policy network and a value network for decision-making and reasoning. Therefore, the step of using the pre-built local scheduling model in the cluster head node to perform decision-making and reasoning on the cluster state representation to obtain a cluster scheduling action set includes: The policy network and value network are obtained based on the local scheduling model. The cluster state representation is input into the policy network to obtain the scheduling policy distribution; Based on the distribution of the scheduling strategy, action sampling is performed to obtain the cluster scheduling action index. The cluster scheduling action index is decomposed into multiple scheduling action nodes using a preset radix conversion rule to obtain the cluster scheduling action set. The scheduling action node includes the execution type and the amount of execution data, and the execution type is idle, local processing, or unloading processing.

[0040] It should be explained that the policy network is a neural network component in the local scheduling model responsible for outputting the probability distribution of scheduling policies. Its function is to evaluate the merits of various scheduling actions based on the cluster state representation and output the corresponding probability values. Optionally, a 4-layer fully connected actor network structure (128 neurons per layer) and the ReLU activation function can be used to achieve this purpose. The value network is a neural network component in the local scheduling model responsible for evaluating the value of the current cluster state. Its function is to estimate the expected cumulative scheduling feedback data that can be obtained by executing scheduling under the current cluster state representation. Optionally, a critic network can be used to achieve this purpose. In this embodiment of the invention, the policy network and the value network share the parameters of the first two layers of the network to improve training efficiency and stability. The scheduling policy distribution is the set of all possible scheduling actions and their corresponding probability values ​​output by the policy network, and the sum of the probability values ​​is 1. The action sampling is the process of randomly selecting a scheduling action according to the probability value in the scheduling policy distribution. The cluster scheduling action index is the decimal number of the selected scheduling action in the action space.

[0041] Understandably, the radix conversion rule is a rule that converts the cluster scheduling action index from decimal representation to a multi-digit representation based on the number of optional actions for each drone node, thereby obtaining the scheduling actions for each drone node. Specifically, let the number of optional actions for each drone node be... ,in This represents the maximum number of data packets that can be processed locally. To determine the maximum number of packets to be processed for offloading, the cluster scheduling action index can be based on the number of available actions. The process is decomposed, with each decomposed element corresponding to a scheduling action node for a drone node. This scheduling action node is the basic unit in the cluster's scheduling action set, containing the action type (execution type) and the number of data packets to be processed (execution data volume) for the corresponding drone node. When the execution type is idle, the drone node does not process any data packets, and the execution data volume is 0. When the execution type is local processing, the drone node processes the execution data volume of data packets on its local computing resources. When the execution type is offload processing, the drone node offloads the execution data volume of data packets to an edge computing server for remote processing via the base station.

[0042] For example, suppose a drone cluster contains two drone nodes (drone node 1 and drone node 2), and the maximum number of data packets that can be processed locally is... The maximum number of data packets that can be processed by offloading is 1. If the value is 3, then the number of optional actions for each drone node is 3. (These represent idle, local processing of 1 data packet, unloading 1 data packet, unloading 2 data packets, and unloading 3 data packets, respectively). The current buffer queue vector for UAV node 1 is (0, 1, -1, -1), the quantization channel gain is -2.5dB, and the battery energy level is 2. The buffer queue vector for UAV node 2 is (0, 0, -1, -1), the quantization channel gain is 3.2dB, and the battery energy level is 3. The statistical channel parameter set corresponding to this UAV cluster includes the statistical channel parameters of the other two clusters, which are 1.8dB and -0.6dB, respectively. After assembling the above data into a cluster state representation, it is input into the policy network. The policy network outputs the probability distribution of 25 scheduling actions (action space size is...). Assuming the cluster scheduling action index obtained from action sampling is 8, it is decomposed according to a radix of 5 to obtain (1, 3), that is, the scheduling action node of UAV node 1 is (local processing, 1 data packet), and the scheduling action node of UAV node 2 is (offload processing, 3 data packets). This embodiment of the invention encodes the cluster state information into a unified-dimensional input vector, combines it with the policy network of the near-end policy optimization algorithm for decision reasoning, and utilizes the radix transformation rule to efficiently decompose the cluster-level action index into independent scheduling instructions for each UAV node, thus achieving refined joint scheduling decision-making within the cluster.

[0043] S4. Perform task scheduling on the UAV node set based on the cluster scheduling action set to obtain a scheduling feedback dataset. Update the local scheduling model using the scheduling feedback dataset to obtain an updated local scheduling model. Obtain model parameters based on the updated local scheduling model.

[0044] It should be explained that the task scheduling is the process of performing specific data processing operations on corresponding drone nodes in the drone node set based on the execution type and execution data volume of each scheduling action node in the cluster scheduling action set. This includes both local computation and offloading to an edge computing server. The scheduling feedback dataset is a set of feedback information collected after the task scheduling is completed, used to evaluate the effectiveness of the scheduling decision and guide model updates. The updated local scheduling model is obtained by optimizing the policy network and value network parameters of the local scheduling model using the scheduling feedback data. The model parameters are all the parameters of the policy network and value network in the updated local scheduling model.

[0045] Understandably, in order to accurately evaluate the effectiveness of each scheduling decision, it is necessary to record multi-dimensional feedback information such as energy consumption and packet loss during the task scheduling process. Therefore, the task scheduling based on the cluster scheduling action set for the UAV node set, to obtain the scheduling feedback dataset, includes: For each scheduling action node in the cluster scheduling action set, the following operations are performed: After confirming that the execution type corresponding to the scheduling action node is local processing, the local energy consumption consumed by the UAV node in processing the amount of execution data is obtained, and the task cache unit of the UAV node is reduced based on the amount of execution data to obtain the first reduced cache queue. After confirming that the execution type corresponding to the scheduling action node is offloading processing, the uplink transmission rate when the UAV node transmits the execution data to the edge computing server is obtained. Based on the uplink transmission rate, the transmission delay and offloading energy consumption are obtained. After confirming that the transmission delay meets the preset time slot constraint, the execution data is offloaded to the edge computing server via the base station for processing. Based on the execution data, the task cache unit is reduced to obtain the second reduced cache queue. Based on the number of data packets dropped in the task cache unit after the first or second reduction of the cache queue, a packet loss statistics node is obtained, wherein the packet loss statistics node includes the number of packet loss due to latency violation and the number of packet loss due to cache overflow. Based on the packet loss statistics node and the local energy consumption or offload energy consumption, the scheduling feedback data is calculated and summarized to obtain the scheduling feedback dataset.

[0046] It should be understood that the local energy consumption refers to the energy consumed by the UAV node when processing and executing a certain number of data packets locally. The calculation method is: multiply the amount of data processed by the preset local processing power and then multiply by the time slot duration. ,in To handle the amount of data, For local processing power, The time slot duration is specified. Data reduction is the operation of removing processed data packets from the task cache unit's queue and updating the queue status. The first reduced cache queue represents the queue status of the task cache unit after local processing is complete. The second reduced cache queue represents the queue status of the task cache unit after unloading processing.

[0047] Furthermore, the uplink transmission rate is the rate at which the UAV node transmits data to the base station when unloading data packets. Therefore, obtaining the uplink transmission rate when the UAV node transmits the execution data amount to the edge computing server includes: The offloaded transmit power of the UAV node is obtained. The UAV node is removed from the UAV cluster corresponding to the UAV node to obtain the remaining UAV node set. The intra-cluster interference power is obtained based on the remaining UAV node set. The inter-cluster interference power is estimated based on the statistical channel parameter set. The uplink transmission rate is calculated using the offloaded transmit power, quantized channel gain, intra-cluster interference power, and inter-cluster interference power estimates. The calculation formula is as follows: in, This indicates the uplink transmission rate. This indicates the preset uplink transmission bandwidth. This indicates the offloaded transmission power. This represents the quantized channel gain. This indicates the interference power within the cluster. This represents the estimated inter-cluster interference power. This represents the preset noise power spectral density. Indicates the uplink transmission identifier. Indicates an identifier within the cluster. Indicates an inter-cluster identifier. This represents the unload identifier.

[0048] In detail, the offloading transmission power is the wireless transmission power used by the UAV node when offloading data packets, and the offloading transmission power is limited by a preset maximum offloading power. The offloading transmission power is determined by: using the uplink transmission rate calculation formula and combining it with the time slot constraint, deriving the minimum transmission power required to complete transmission within the time slot, that is, ensuring that the uplink transmission completes the sending, waiting, receiving, and decoding of all data packets within the time slot duration, and solving the equation to obtain the offloading transmission power. The intra-cluster interference power is the sum of the interference power generated by other UAV nodes in the same UAV cluster whose channel gain is lower than that of the current UAV node during uplink transmission. Since the channel gain and offloading decision information of each UAV node in the cluster are known at the cluster leader node, the intra-cluster interference power can be calculated accurately. The inter-cluster interference power estimate is the result of estimating the uplink interference power generated by other UAV clusters based on the statistical channel parameter set. Since precise channel information and offloading decision information are not shared among clusters, estimation can only be made based on the statistical channel parameters (i.e., the average channel gain) of each remaining drone cluster. In the estimation process, it is assumed that all drone nodes in other clusters transmit signals at maximum offloading power, and whether the current drone node is affected by interference from that cluster is determined by whether its channel gain is higher than the statistical channel parameters of other clusters.

[0049] Understandably, the transmission latency is the time required for a UAV node to transmit a data packet to the base station via the uplink. The time slot constraint is the maximum duration of a single scheduling time slot; the sending, waiting, receiving, and decoding processes of the offloading operation must be completed within this duration. When the transmission latency exceeds the time slot constraint, it means that the data packet transmission has timed out, and the offloading operation will result in a transmission error. The offloading energy consumption is the total energy consumed by the UAV node during the offloading process, including four parts: sending energy consumption, waiting energy consumption, receiving energy consumption, and decoding energy consumption.

[0050] It should be explained that the packet loss statistics node is a data structure that records the packet dropping status in the task cache unit. The latency violation packet loss count is the number of packets forcibly dropped because the waiting time of the packets in the task cache unit exceeds the preset maximum strict latency. The maximum strict latency is the longest time slot allowed for a packet to wait from the time it enters the task cache unit until it is processed. The buffer overflow packet loss count is the number of packets dropped because the number of newly arriving packets exceeds the number of remaining empty slots in the task cache unit and thus cannot enter the cache. The scheduling feedback data is a dimensionless feedback value calculated based on the latency violation packet loss count, buffer overflow packet loss count, and local energy consumption or offloading energy consumption in the packet loss statistics node. It is equal to the weighted sum of the negative values ​​of packet loss count and the negative values ​​of normalized energy consumption. This definition method ensures that the smaller the packet loss, the larger the value of the scheduling feedback data, thereby guiding the local scheduling model to optimize its strategy towards reducing packet loss.

[0051] Furthermore, to improve the policy quality of the local scheduling model, it is necessary to iteratively update the model parameters using the collected scheduling feedback data. Therefore, updating the local scheduling model using the scheduling feedback dataset to obtain an updated local scheduling model includes: Obtain the cache tuple set based on the scheduling feedback dataset and the pre-built experience cache area; Randomly sample from the cached tuple set using a preset small-batch sampling ratio to obtain a small-batch tuple set. Calculate the policy ratio and pruning objective function value based on the small-batch tuple set. Update the parameters of the policy network and value network based on the policy ratio and pruning objective function value, and count the number of iterations to obtain the initial update scheduling model and the number of iterations. If the number of iterations is less than a preset threshold, then return to the step of randomly sampling the cached tuple set using a preset small batch sampling ratio until the number of iterations equals the threshold, and use the initial update scheduling model as the update local scheduling model.

[0052] It should be understood that the experience cache is a storage area used to store state-action-reward cache tuples generated during the scheduling process. Its capacity can be preset; optionally, the capacity of the experience cache can be set to 1024 cache tuples. The cache tuple set is the collection of all cache tuples stored in the experience cache. Each cache tuple includes the current cluster state representation, the index of the executed cluster scheduling action, the policy probability value of the previous round, the cluster state representation of the next time slot of the value network, and scheduling feedback data. The mini-batch sampling ratio is the proportion of random sampling from the cache tuple set, used to determine the number of samples used for each parameter update. The policy ratio is the ratio of the probability value of the current policy on a certain state-action pair to the probability value of the old policy on the same state-action pair, used to measure the magnitude of policy change. The pruning objective function value is the core objective function of the near-end policy optimization algorithm. By pruning the result of multiplying the policy ratio by the dominance function, the policy update magnitude is limited to... Within a range to ensure training stability, among which The pruning factor is used. The advantage function is used to evaluate the superiority or inferiority of a certain action relative to the average level, and a weighted sum of the multi-step reward error is performed by combining the discount factor and the advantage estimation factor. The number threshold is the maximum number of sampling iterations of the experience buffer data in each round.

[0053] For example, suppose a drone swarm executes task scheduling within a scheduling time slot, and drone node 1 selects to process one data packet locally, with a local processing power of... 180 microwatts, time slot duration If the value is 1 millisecond, then the local energy consumption is In nanojoules, after the first data reduction, the buffer queue changes from (0, 1, -1, -1) to (1, -1, -1, -1), meaning that data packets with a waiting time of 0 are processed and removed, while data packets with a waiting time of 1 continue to age out. Drone node 2 selects to offload two data packets to the edge computing server, with an offload transmission power of 1.5 milliwatts and an uplink transmission bandwidth of... The frequency is 4MHz, the quantization channel gain is 3.2dB, which is converted to a linear quantization channel gain of approximately 2.09. Substituting these values ​​into the uplink transmission rate formula, the uplink transmission rate is calculated. The intra-cluster interference power is 0 (no other drone nodes are simultaneously unloading). The inter-cluster interference power estimate is calculated based on the statistical channel parameters of other clusters. Substituting these values ​​into the formula yields the uplink transmission rate. The transmission delay meets the time slot constraint, and the unloading operation is complete. Packet loss is analyzed for two drone nodes. Drone node 1 has no delay violation packet loss or buffer overflow packet loss, and drone node 2 also has no packet loss. The scheduling feedback data is 0. During the training phase, the cached tuple is stored in the empirical buffer. After 1024 time slots, the buffer is full. 64 cached tuples are randomly sampled to form a mini-batch tuple set. The policy ratio and pruning objective function value are calculated, and the policy network and value network parameters are updated. This sampling and update is repeated 10 times to obtain the updated local scheduling model. This embodiment of the invention constrains the policy update magnitude through the pruning mechanism of the near-end policy optimization algorithm and fully utilizes empirical data through multiple mini-batch sampling, achieving stable and efficient parameter optimization of the local scheduling model.

[0054] S5. Summarize the model parameters to obtain the model parameter set. Use the edge computing server to aggregate the model parameter set to obtain the global scheduling model parameters. Broadcast the global scheduling model parameters to the first node of each of the multiple UAV clusters. Update the local scheduling model with the global scheduling model parameters and return to the previous step of performing the following operations on each of the multiple UAV clusters until the scheduling feedback dataset is confirmed to meet the preset convergence conditions, thereby realizing intelligent scheduling of the UAV group.

[0055] It should be explained that the model parameter set is a collection of model parameters extracted from the updated local scheduling model of each cluster leader node in multiple drone clusters. The aggregation is the process of merging and calculating the model parameters from multiple cluster leader nodes on an edge computing server to generate unified global model parameters. The global scheduling model parameters are the unified model parameters obtained after aggregation, representing the common scheduling strategy of all drone clusters. The broadcast is a communication operation that sends the global scheduling model parameters to all cluster leader nodes via a base station.

[0056] Understandably, the core idea of ​​the federated learning framework is that each cluster's head node only transmits model parameters rather than the original scheduling data. This protects data privacy while achieving cross-cluster policy sharing through global aggregation. Therefore, the aggregation of the model parameter set using an edge computing server to obtain global scheduling model parameters includes: Obtain the model parameters of the updated local scheduling model of each cluster leader node in multiple drone clusters to obtain the model parameter set; The model parameter set is aggregated by using an edge computing server to obtain the global scheduling model parameters. The calculation formula is shown below: in, This represents the parameters of the global scheduling model. This indicates that there are a total of [number] drones in the aforementioned drone cluster. A drone swarm, The first parameter in the model parameter set Each model parameter.

[0057] It should be understood that the mean aggregation is an operation that calculates the arithmetic mean of all model parameters in the model parameter set element by element according to their corresponding positions. This aggregation method assigns the same weight to each UAV cluster, enabling the global scheduling model to take into account the differences between different UAV clusters in terms of data arrival patterns, number of UAV nodes, and channel conditions. After the aggregated global scheduling model parameters are broadcast to the head node of each cluster, each cluster head node replaces the current parameters of its local scheduling model with the global scheduling model parameters, and then re-enters the scheduling loop to continue performing operations such as cluster state representation construction, decision reasoning, and task scheduling based on the updated model.

[0058] It should be explained that the scheduling loop is repeatedly executed in steps S3 to S5. In each loop, the first node of each cluster makes a decision based on the current model, collects feedback, updates the model, and uploads the weights to the edge computing server to complete global aggregation. Iteration stops when the scheduling feedback dataset meets the convergence condition. The global scheduling model at this point is the final intelligent scheduling strategy. The convergence condition is a preset criterion used to determine whether the scheduling model has been trained to a stable state.

[0059] Understandably, in order to determine whether the scheduling model has converged to a stable scheduling policy, it is necessary to quantitatively evaluate the changing trend of the scheduling feedback data. Therefore, confirming that the scheduling feedback dataset meets the preset convergence conditions includes: Extract scheduling feedback data sequentially from the scheduling feedback dataset, and perform the following operations on the extracted scheduling feedback data: Using the scheduling feedback data, the analytical scheduling feedback data is identified in the scheduling feedback dataset. The analytical scheduling feedback data is adjacent to and lags behind the scheduling feedback data. The absolute difference between the analytical scheduling feedback data and the scheduling feedback data is calculated to obtain the iterative fluctuation value. The iterative fluctuation values ​​are summarized to obtain an iterative fluctuation value set. An evaluation fluctuation value set is extracted from the iterative fluctuation value set using a preset fluctuation evaluation window. The mean of the evaluation fluctuation values ​​in the evaluation fluctuation value set is calculated to obtain the evaluation fluctuation mean. The average value of the evaluated fluctuation is compared with a preset convergence threshold. If the average value of the evaluated fluctuation is less than or equal to the convergence threshold, it is confirmed that the scheduling feedback dataset meets the convergence condition.

[0060] In detail, the analyzed scheduling feedback data is the next scheduling feedback data that is temporally adjacent to and follows the currently extracted scheduling feedback data. The absolute difference between the two reflects the degree of change in the effect of two adjacent scheduling decisions. The iterative fluctuation value is a measure of the absolute difference between two adjacent scheduling feedback data. The smaller the iterative fluctuation value, the more stable the scheduling strategy is. The iterative fluctuation value set is a collection of all calculated iterative fluctuation values, recording the fluctuation history of the scheduling feedback data throughout the training process. The fluctuation evaluation window is a sliding window used to extract recent iterative fluctuation values, focusing on the fluctuation trend during the most recent training period rather than the global fluctuation situation. In this embodiment, the fluctuation evaluation window can be set to the most recent 100 iterative fluctuation values. The evaluation fluctuation value set is a subset of recent iterative fluctuation values ​​extracted from the iterative fluctuation value set using the fluctuation evaluation window. The evaluation fluctuation mean is the arithmetic mean of all evaluation fluctuation values ​​in the evaluation fluctuation value set, representing the average stability of the recent scheduling strategy. The convergence threshold is a pre-set judgment criterion. When the evaluation fluctuation mean is less than or equal to the convergence threshold, the scheduling model is considered to have converged to a stable strategy, and training can be stopped. The convergence threshold is set according to the accuracy requirements of the scheduling task and the tolerance of the actual scenario.

[0061] For example, suppose that after multiple scheduling cycles, the five most recent scheduling feedback data in the scheduling feedback dataset are -2.3, -2.1, -2.05, -2.08, and -2.06. The iterative fluctuation values ​​between adjacent scheduling feedback data are calculated as follows: |-2.1-(-2.3)|=0.2, |-2.05-(-2.1)|=0.05, |-2.08-(-2.05)|=0.03, and |-2.06-(-2.08)|=0.02. Let the fluctuation evaluation window be the four most recent iterative fluctuation values; then the evaluation fluctuation value set is {0.2, 0.05, 0.03, 0.02}, and the average evaluation fluctuation is (0.2+0.05+0.03+0.02) / 4=0.075. If the convergence threshold is set to 0.05, and the average evaluation fluctuation of 0.075 is greater than the convergence threshold of 0.05, the scheduling model has not yet converged, and the scheduling loop continues. As training continues, assuming that the subsequent average evaluation fluctuation drops to 0.008, which is less than the convergence threshold of 0.05, the scheduling feedback dataset is confirmed to meet the convergence condition, and the scheduling training ends. At this point, the global scheduling model parameters are the optimal intelligent scheduling strategy. The first node of each cluster performs real-time scheduling of UAV nodes based on this strategy, realizing intelligent scheduling of UAV groups. This embodiment of the invention judges the convergence status by the average of the iterative fluctuation values ​​within a sliding window, which avoids misjudgments caused by single fluctuations and can capture the stable trend of the scheduling strategy in a timely manner, ensuring that the model terminates training after full convergence, thus balancing the quality of the scheduling strategy and training efficiency.

[0062] To address the problems described in the background art, this invention receives scheduling instructions and, based on these instructions, determines the UAV group scheduling environment. This environment includes a set of UAV nodes, an edge computing server, and a base station. The UAV node set comprises multiple UAV nodes, each equipped with a task caching unit. Spatial location information is obtained from the UAV node set, and the set is dynamically partitioned into multiple UAV clusters based on this information. Each UAV cluster includes multiple UAV nodes and a cluster leader node. Therefore, this invention, through dynamic cluster partitioning based on polar coordinate parameters, achieves complementary channel conditions. Furthermore, grouping UAV nodes with similar spatial locations into the same cluster ensures transmission efficiency when multiple UAV nodes share communication resources. It also decomposes the large-scale scheduling problem into multiple smaller-scale UAV cluster-level decision problems. For each UAV cluster, the following operations are performed: A cluster state representation is constructed based on the UAV cluster; a pre-built local scheduling model in the cluster's head node is used to perform decision reasoning on the cluster state representation to obtain the cluster scheduling action set. Thus, this invention encodes cluster state information into a unified-dimensional input vector, combines it with a policy network of a near-end policy optimization algorithm for decision reasoning, and utilizes a radix transformation rule to efficiently group the cluster... The hierarchical action index is decomposed into independent scheduling instructions for each UAV node, realizing fine-grained joint scheduling decisions within the cluster. Task scheduling is performed on the UAV node set based on the cluster scheduling action set, resulting in a scheduling feedback dataset. This dataset is then used to update the local scheduling model, yielding an updated local scheduling model. Model parameters are obtained based on this updated model. It is evident that this invention constrains the policy update magnitude through the pruning mechanism of the near-end policy optimization algorithm and fully utilizes empirical data through multiple small-batch sampling, achieving stable and efficient parameter optimization of the local scheduling model. The model parameters are then aggregated to obtain a model parameter set, which is then processed using an edge computing server. The data sets are aggregated to obtain global scheduling model parameters. These parameters are then broadcast to the head node of each of the multiple drone clusters. The local scheduling model is updated with these global parameters, and the process is repeated for each drone cluster until the scheduling feedback dataset meets the preset convergence conditions. This achieves intelligent scheduling of the drone groups. It is evident that this invention determines the convergence state by using the average of iterative fluctuation values ​​within a sliding window. This avoids misjudgments caused by single fluctuations and promptly captures the stable trend of the scheduling strategy, ensuring that the model terminates training after full convergence. This balances the quality of the scheduling strategy with training efficiency. Therefore, this invention can optimize task scheduling between multiple drones and edge servers, reducing packet loss and energy consumption.

[0063] like Figure 2The diagram shown is a functional block diagram of an intelligent scheduling system for unmanned aerial vehicle (UAV) groups based on edge computing, provided in an embodiment of the present invention.

[0064] The edge computing-based intelligent scheduling system 100 for unmanned aerial vehicles (UAVs) described in this invention can be installed in an electronic device. Depending on the functions implemented, the edge computing-based intelligent scheduling system 100 may include an environment verification module 101, a cluster partitioning module 102, a scheduling decision module 103, and a global aggregation module 104. The modules described in this invention can also be referred to as units, which are a series of computer program segments that can be executed by the processor of an electronic device and perform a fixed function, stored in the memory of the electronic device.

[0065] The environment confirmation module 101 is used to receive scheduling instructions and confirm the UAV group scheduling environment based on the scheduling instructions. The UAV group scheduling environment includes a UAV node set, an edge computing server and a base station. The UAV node set includes multiple UAV nodes and each UAV node has a task cache unit. The cluster partitioning module 102 is used to obtain a set of spatial location information using the set of UAV nodes, and to dynamically partition the set of UAV nodes into multiple UAV clusters based on the set of spatial location information. Each UAV cluster includes multiple UAV nodes and a cluster head node. The scheduling decision module 103 is used to perform the following operations on each of the multiple drone clusters: constructing a cluster state representation based on the drone cluster, and using the pre-constructed local scheduling model in the cluster head node to perform decision reasoning on the cluster state representation to obtain a cluster scheduling action set; Task scheduling is performed on the UAV node set based on the cluster scheduling action set to obtain a scheduling feedback dataset. The local scheduling model is updated using the scheduling feedback dataset to obtain an updated local scheduling model. Model parameters are obtained based on the updated local scheduling model. The global aggregation module 104 is used to summarize model parameters to obtain a model parameter set, aggregate the model parameter set using an edge computing server to obtain global scheduling model parameters, broadcast the global scheduling model parameters to the first node of each of the multiple UAV clusters, update the local scheduling model with the global scheduling model parameters, and then return to the step of performing the following operations on each of the multiple UAV clusters until it is confirmed that the scheduling feedback dataset meets the preset convergence conditions, thereby realizing intelligent scheduling of UAV groups.

[0066] In detail, the modules in the edge computing-based intelligent scheduling system 100 for unmanned aerial vehicle (UAV) groups described in this embodiment of the invention employ the same methods as described above. Figure 1The method uses the same technical means as the edge computing-based intelligent scheduling method for unmanned aerial vehicle (UAV) groups described in the article and can produce the same technical effect, so it will not be elaborated here.

[0067] like Figure 3 The diagram shown is a schematic representation of an electronic device that implements an intelligent scheduling method for unmanned aerial vehicle (UAV) groups based on edge computing, according to an embodiment of the present invention.

[0068] The electronic device 1 may include a processor 10, a memory 11 and a bus 12, and may also include a computer program stored in the memory 11 and capable of running on the processor 10, such as a method program for intelligent scheduling of unmanned aerial vehicle groups based on edge computing.

[0069] The memory 11 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 11 can be an internal storage unit of the electronic device 1, such as the portable hard drive of the electronic device 1. In other embodiments, the memory 11 can also be an external storage device of the electronic device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device 1. Furthermore, the memory 11 includes both internal storage units and external storage devices of the electronic device 1. The memory 11 can be used not only to store application software and various types of data installed on the electronic device 1, such as the code of a UAV group intelligent scheduling method program based on edge computing, but also to temporarily store data that has been output or will be output.

[0070] In some embodiments, the processor 10 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control unit of the electronic device, connecting various components of the entire electronic device through various interfaces and lines. It executes programs or modules stored in the memory 11 (e.g., an edge computing-based intelligent scheduling method for unmanned aerial vehicle groups) and calls data stored in the memory 11 to perform various functions of the electronic device 1 and process data.

[0071] The bus 12 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus 12 can be divided into an address bus, a data bus, a control bus, etc. The bus 12 is configured to realize the connection and communication between the memory 11 and at least one processor 10, etc.

[0072] Figure 3 Only electronic devices with components are shown; it will be understood by those skilled in the art that... Figure 3 The structure shown does not constitute a limitation on the electronic device 1, and may include fewer or more components than shown, or combine certain components, or have different component arrangements.

[0073] For example, although not shown, the electronic device 1 may also include a power supply (such as a battery) to power the various components. Preferably, the power supply can be logically connected to the at least one processor 10 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The electronic device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.

[0074] Furthermore, the electronic device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, a Bluetooth interface, etc.), which is typically used to establish communication connections between the electronic device 1 and other electronic devices.

[0075] Optionally, the electronic device 1 may further include a user interface, which may be a display, an input unit (such as a keyboard), and optionally, a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the electronic device 1 and to display a visual user interface.

[0076] The memory 11 in the electronic device 1 stores a program for intelligent scheduling of unmanned aerial vehicle (UAV) groups based on edge computing. This program is a combination of multiple instructions, which, when run in the processor 10, can achieve the following: The system receives a scheduling instruction and determines the drone group scheduling environment based on the instruction. The drone group scheduling environment includes a drone node set, an edge computing server, and a base station. The drone node set includes multiple drone nodes, and each drone node has a task cache unit. A set of spatial location information is obtained by using a set of UAV nodes. Based on the set of spatial location information, the set of UAV nodes is dynamically divided into clusters to obtain multiple UAV clusters. Each UAV cluster includes multiple UAV nodes and a cluster head node. For each of the multiple drone clusters, the following operations are performed: a cluster state representation is constructed based on the drone cluster, and a decision reasoning is performed on the cluster state representation using a pre-built local scheduling model in the cluster head node to obtain a cluster scheduling action set; Task scheduling is performed on the UAV node set based on the cluster scheduling action set to obtain a scheduling feedback dataset. The local scheduling model is updated using the scheduling feedback dataset to obtain an updated local scheduling model. Model parameters are obtained based on the updated local scheduling model. The model parameters are summarized to obtain a model parameter set. The model parameter set is then aggregated using an edge computing server to obtain global scheduling model parameters. These global scheduling model parameters are broadcast to the first node of each of the multiple UAV clusters. The local scheduling model is updated with the global scheduling model parameters, and then the process is repeated for each of the multiple UAV clusters. This process continues until the scheduling feedback dataset is confirmed to meet the preset convergence conditions, thus achieving intelligent scheduling of the UAV groups.

[0077] Specifically, the processor 10's implementation method for the above instructions can be found in [reference needed]. Figures 1 to 3 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.

[0078] Furthermore, if the modules / units integrated in the electronic device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).

[0079] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor of an electronic device, can perform the following: The system receives a scheduling instruction and determines the drone group scheduling environment based on the instruction. The drone group scheduling environment includes a drone node set, an edge computing server, and a base station. The drone node set includes multiple drone nodes, and each drone node has a task cache unit. A set of spatial location information is obtained by using a set of UAV nodes. Based on the set of spatial location information, the set of UAV nodes is dynamically divided into clusters to obtain multiple UAV clusters. Each UAV cluster includes multiple UAV nodes and a cluster head node. For each of the multiple drone clusters, the following operations are performed: a cluster state representation is constructed based on the drone cluster, and a decision reasoning is performed on the cluster state representation using a pre-built local scheduling model in the cluster head node to obtain a cluster scheduling action set; Task scheduling is performed on the UAV node set based on the cluster scheduling action set to obtain a scheduling feedback dataset. The local scheduling model is updated using the scheduling feedback dataset to obtain an updated local scheduling model. Model parameters are obtained based on the updated local scheduling model. The model parameters are summarized to obtain a model parameter set. The model parameter set is then aggregated using an edge computing server to obtain global scheduling model parameters. These global scheduling model parameters are broadcast to the first node of each of the multiple UAV clusters. The local scheduling model is updated with the global scheduling model parameters, and then the process is repeated for each of the multiple UAV clusters. This process continues until the scheduling feedback dataset is confirmed to meet the preset convergence conditions, thus achieving intelligent scheduling of the UAV groups.

[0080] In the several embodiments provided by this invention, it should be understood that the disclosed devices, systems, and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative, and actual implementations may have other classification methods.

[0081] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0082] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.

[0083] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0084] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for intelligent scheduling of unmanned aerial vehicle (UAV) groups based on edge computing, characterized in that, The method includes: The system receives a scheduling instruction and determines the drone group scheduling environment based on the instruction. The drone group scheduling environment includes a drone node set, an edge computing server, and a base station. The drone node set includes multiple drone nodes, and each drone node has a task cache unit. A set of spatial location information is obtained by using a set of UAV nodes. Based on the set of spatial location information, the set of UAV nodes is dynamically divided into clusters to obtain multiple UAV clusters. Each UAV cluster includes multiple UAV nodes and a cluster head node. For each of the multiple drone clusters, the following operations are performed: a cluster state representation is constructed based on the drone cluster, and a decision reasoning is performed on the cluster state representation using a pre-built local scheduling model in the cluster head node to obtain a cluster scheduling action set; Task scheduling is performed on the UAV node set based on the cluster scheduling action set to obtain a scheduling feedback dataset. The local scheduling model is updated using the scheduling feedback dataset to obtain an updated local scheduling model. Model parameters are obtained based on the updated local scheduling model. The model parameters are summarized to obtain a model parameter set. The model parameter set is then aggregated using an edge computing server to obtain global scheduling model parameters. These global scheduling model parameters are broadcast to the first node of each of the multiple UAV clusters. The local scheduling model is updated with the global scheduling model parameters, and then the process is repeated for each of the multiple UAV clusters. This process continues until the scheduling feedback dataset is confirmed to meet the preset convergence conditions, thus achieving intelligent scheduling of the UAV groups.

2. The intelligent scheduling method for unmanned aerial vehicle (UAV) groups based on edge computing as described in claim 1, characterized in that, The dynamic clustering of the UAV node set based on the spatial location information set yields multiple UAV clusters, including: Based on the spatial location information set, the polar coordinate parameters of each UAV node relative to the base station are obtained, resulting in multiple polar coordinate parameters, including: channel gain sub-interval identifier and quantized azimuth angle. Each drone node in the drone node set is initialized to obtain multiple initial clusters, where each initial cluster corresponds one-to-one with the polar coordinate parameters; For each of the multiple initial clusters, perform the following operation: The initial cluster is removed from multiple initial clusters to obtain a set of remaining initial clusters. Based on multiple polar coordinate parameters, an initial cluster that matches the initial cluster is retrieved from the set of remaining initial clusters to obtain the retrieval results, where the retrieval results indicate whether the cluster exists or does not exist. If the search result indicates that a candidate initial cluster set exists, then a candidate initial cluster set is obtained from the remaining initial cluster set, and the following operation is performed on each candidate initial cluster set: Calculate the absolute difference in quantized azimuth angles between the initial cluster and the candidate initial clusters to obtain the absolute angle difference. Summarize the absolute angle differences and arrange them in ascending order to obtain the sequence of absolute angle differences. Select the candidate initial cluster corresponding to the absolute angle difference with the position number 1 as the matching initial cluster. Merge the initial cluster with the matching initial cluster into the target cluster; otherwise, mark the initial cluster as an unmerged initial cluster. By summarizing the target clusters, a target cluster set is obtained; Obtain the drone nodes in the unmerged initial cluster to obtain the set of drone nodes to be assigned; For each drone node in the set of drone nodes to be assigned, the following operations are performed: Candidate drone clusters are obtained based on the drone nodes to be assigned and the target cluster set; By aggregating the candidate drone clusters, multiple candidate drone clusters are obtained; After confirming that the number of drone nodes in each of the multiple candidate drone clusters meets the preset cluster capacity range, multiple drone clusters are obtained, and a cluster leader node is determined in each drone cluster.

3. The intelligent scheduling method for unmanned aerial vehicle (UAV) groups based on edge computing as described in claim 2, characterized in that, The construction of cluster state representation based on drone swarm includes: Obtain the cache queue vector, quantization channel gain, and battery energy level of each drone node in the drone cluster to obtain the cache queue vector set, quantization channel gain set, and battery energy level set. Associate the cache queue vector set, quantization channel gain set, and battery energy level set to obtain the cluster internal information group. The drone clusters are removed from the multiple drone clusters to obtain a set of remaining drone clusters. The set of remaining drone clusters includes multiple remaining drone clusters. The average channel gain of all drone nodes in the set of remaining drone clusters is obtained to obtain a set of statistical channel parameters. The set of statistical channel parameters includes multiple statistical channel parameters, and the statistical channel parameters correspond one-to-one with the remaining drone clusters. By associating the cluster's internal information group with the statistical channel parameter set, an initial state representation is obtained; The number of drone nodes in the drone cluster is counted to obtain the node count. If the node count is less than the preset node count limit, a zero-value filling operation is performed on the initial state representation to obtain the cluster state representation.

4. The intelligent scheduling method for unmanned aerial vehicle (UAV) groups based on edge computing as described in claim 3, characterized in that, The decision-making reasoning based on the cluster state representation using the pre-built local scheduling model in the cluster leader node yields a set of cluster scheduling actions, including: The policy network and value network are obtained based on the local scheduling model. The cluster state representation is input into the policy network to obtain the scheduling policy distribution; Based on the distribution of the scheduling strategy, action sampling is performed to obtain the cluster scheduling action index. The cluster scheduling action index is decomposed into multiple scheduling action nodes using a preset radix conversion rule to obtain the cluster scheduling action set. The scheduling action node includes the execution type and the amount of execution data, and the execution type is idle, local processing, or unloading processing.

5. The intelligent scheduling method for unmanned aerial vehicle (UAV) groups based on edge computing as described in claim 4, characterized in that, The process of performing task scheduling on the UAV node set based on the cluster scheduling action set yields a scheduling feedback dataset, including: For each scheduling action node in the cluster scheduling action set, the following operations are performed: After confirming that the execution type corresponding to the scheduling action node is local processing, the local energy consumption consumed by the UAV node in processing the amount of execution data is obtained, and the task cache unit of the UAV node is reduced based on the amount of execution data to obtain the first reduced cache queue. After confirming that the execution type corresponding to the scheduling action node is offloading processing, the uplink transmission rate when the UAV node transmits the execution data to the edge computing server is obtained. Based on the uplink transmission rate, the transmission delay and offloading energy consumption are obtained. After confirming that the transmission delay meets the preset time slot constraint, the execution data is offloaded to the edge computing server via the base station for processing. Based on the execution data, the task cache unit is reduced to obtain the second reduced cache queue. Based on the number of data packets dropped in the task cache unit after the first or second reduction of the cache queue, a packet loss statistics node is obtained, wherein the packet loss statistics node includes the number of packet loss due to latency violation and the number of packet loss due to cache overflow. Based on the packet loss statistics node and the local energy consumption or offload energy consumption, the scheduling feedback data is calculated and summarized to obtain the scheduling feedback dataset.

6. The intelligent scheduling method for unmanned aerial vehicle (UAV) groups based on edge computing as described in claim 5, characterized in that, The uplink transmission rate when the drone node transmits the execution data to the edge computing server includes: The offloaded transmit power of the UAV node is obtained. The UAV node is removed from the UAV cluster corresponding to the UAV node to obtain the remaining UAV node set. The intra-cluster interference power is obtained based on the remaining UAV node set. The inter-cluster interference power is estimated based on the statistical channel parameter set. The uplink transmission rate is calculated using the offloaded transmit power, quantized channel gain, intra-cluster interference power, and inter-cluster interference power estimates. The calculation formula is as follows: in, This indicates the uplink transmission rate. This indicates the preset uplink transmission bandwidth. This indicates the offloaded transmission power. This represents the quantized channel gain. This indicates the interference power within the cluster. This represents the estimated inter-cluster interference power. This represents the preset noise power spectral density. Indicates the uplink transmission identifier. Indicates an identifier within the cluster. Indicates an inter-cluster identifier. This represents the unload identifier.

7. The intelligent scheduling method for unmanned aerial vehicle (UAV) groups based on edge computing as described in claim 6, characterized in that, The step of updating the local scheduling model using the scheduling feedback dataset to obtain the updated local scheduling model includes: Obtain the cache tuple set based on the scheduling feedback dataset and the pre-built experience cache area; Randomly sample from the cached tuple set using a preset small-batch sampling ratio to obtain a small-batch tuple set. Calculate the policy ratio and pruning objective function value based on the small-batch tuple set. Update the parameters of the policy network and value network based on the policy ratio and pruning objective function value, and count the number of iterations to obtain the initial update scheduling model and the number of iterations. If the number of iterations is less than a preset threshold, then return to the step of randomly sampling the cached tuple set using a preset small batch sampling ratio until the number of iterations equals the threshold, and use the initial update scheduling model as the update local scheduling model.

8. The intelligent scheduling method for unmanned aerial vehicle (UAV) groups based on edge computing as described in claim 7, characterized in that, The aggregation of the model parameter set using an edge computing server to obtain global scheduling model parameters includes: Obtain the model parameters of the updated local scheduling model of each cluster leader node in multiple drone clusters to obtain the model parameter set; The model parameter set is aggregated by using an edge computing server to obtain the global scheduling model parameters. The calculation formula is shown below: in, This represents the parameters of the global scheduling model. This indicates that there are a total of [number] drones in the aforementioned drone cluster. A drone swarm, The first parameter in the model parameter set Each model parameter.

9. The intelligent scheduling method for unmanned aerial vehicle (UAV) groups based on edge computing as described in claim 8, characterized in that, The confirmation scheduling feedback dataset satisfies preset convergence conditions, including: Extract scheduling feedback data sequentially from the scheduling feedback dataset, and perform the following operations on the extracted scheduling feedback data: Using the scheduling feedback data, the analytical scheduling feedback data is identified in the scheduling feedback dataset. The analytical scheduling feedback data is adjacent to and lags behind the scheduling feedback data. The absolute difference between the analytical scheduling feedback data and the scheduling feedback data is calculated to obtain the iterative fluctuation value. The iterative fluctuation values ​​are summarized to obtain an iterative fluctuation value set. An evaluation fluctuation value set is extracted from the iterative fluctuation value set using a preset fluctuation evaluation window. The mean of the evaluation fluctuation values ​​in the evaluation fluctuation value set is calculated to obtain the evaluation fluctuation mean. The average value of the evaluated fluctuation is compared with a preset convergence threshold. If the average value of the evaluated fluctuation is less than or equal to the convergence threshold, it is confirmed that the scheduling feedback dataset meets the convergence condition.

10. An intelligent scheduling system for unmanned aerial vehicle (UAV) groups based on edge computing, characterized in that, The system includes: The environment confirmation module is used to receive scheduling instructions and confirm the UAV group scheduling environment based on the scheduling instructions. The UAV group scheduling environment includes a UAV node set, an edge computing server and a base station. The UAV node set includes multiple UAV nodes and each UAV node has a task cache unit. The cluster partitioning module is used to obtain a set of spatial location information using the set of UAV nodes, and to dynamically partition the set of UAV nodes into multiple UAV clusters based on the set of spatial location information. Each UAV cluster includes multiple UAV nodes and a cluster head node. The scheduling decision module is used to perform the following operations on each of the multiple drone clusters: construct a cluster state representation based on the drone cluster, and use the pre-built local scheduling model in the cluster head node to make decision reasoning on the cluster state representation to obtain a cluster scheduling action set; Task scheduling is performed on the UAV node set based on the cluster scheduling action set to obtain a scheduling feedback dataset. The local scheduling model is updated using the scheduling feedback dataset to obtain an updated local scheduling model. Model parameters are obtained based on the updated local scheduling model. The global aggregation module is used to summarize model parameters to obtain a model parameter set. The model parameter set is aggregated using an edge computing server to obtain global scheduling model parameters. The global scheduling model parameters are broadcast to the first node of each of the multiple UAV clusters. After updating the local scheduling model with the global scheduling model parameters, the module returns to the previous state. The following steps are performed on each of the multiple UAV clusters until the scheduling feedback dataset is confirmed to meet the preset convergence conditions, thereby realizing intelligent scheduling of the UAV group.