A data driven resource allocation method

By constructing an air-space-ground network and optimizing resource allocation and energy consumption, the problems of data transmission timeliness and low spectrum energy efficiency in the connection between low-Earth orbit satellites and IoRT equipment have been solved, achieving efficient data transmission and spectrum utilization.

CN120150792BActive Publication Date: 2025-11-18CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510285402.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-11-18
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

In IoRT applications, when low-Earth orbit satellites are directly connected to IoRT devices, there are problems such as limited concurrent connections, limited resources, and lack of hardware facilities, resulting in poor data transmission timeliness and low spectrum efficiency.

Method used

A space-air-ground network is constructed, and the monitoring area is divided into multiple clusters through three-layer clustering. Each cluster is further subdivided into groups. A three-hop transmission process is designed from sensing point to convergence point, from convergence point to UAV, and from UAV to low-Earth orbit satellite. Non-orthogonal multiple access technology and interference cancellation methods are used, combined with deep reinforcement learning and PPO algorithm to optimize resource allocation and energy consumption.

Benefits of technology

It improves the timeliness and spectrum efficiency of data transmission, maximizes spectrum utilization, optimizes the energy harvesting and data transmission path of UAVs, and reduces data backlog and energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120150792B_ABST
    Figure CN120150792B_ABST
Patent Text Reader

Abstract

The embodiment of the present disclosure provides a data-driven resource allocation method, which belongs to the technical field of communication and specifically comprises the following steps: constructing a space-air-ground network; performing three-layer clustering on a monitoring area; merging and splitting clusters and groups according to a preset standard; designing a three-hop transmission process; designing a ground-sky transmission model according to a beam allocation and a convergence point rotation mechanism; designing a sky-ground transmission model according to a non-orthogonal frequency division multiple access technology; constructing a global optimization problem according to the ground-sky transmission model and the sky-ground transmission model, and splitting the global optimization problem into a first optimization problem corresponding to the ground-sky transmission model and a second optimization problem corresponding to the sky-ground transmission model; setting a solution scheme of the first optimization problem; setting a solution scheme of the second optimization problem; designing an ISS-PPO algorithm to solve the first optimization problem and the second optimization problem; and designing a global optimization problem solution scheme. Through the scheme of the present disclosure, the data transmission timeliness and spectrum energy efficiency are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The embodiment of the present disclosure relates to the technical field of communication, in particular to a data-driven resource allocation method. BACKGROUND

[0002] Currently, with the increasing demand for global coverage of the sixth generation (6th Generation, 6G) mobile communication system, the integration of space-air-ground (even underwater) networks is increasingly concerned, which provides support for Internet of Things applications. Remote Internet of Things (Internet of remote things, IoRT) applications include various remote monitoring systems for wildlife changes, vegetation degradation, natural disasters, and climate change, which can alleviate the difficulty of sending personnel and deploying a large number of expensive equipment in remote areas. However, a large number of IoRT real-time monitoring devices need to send the sensed data to a remote data center for further analysis. Therefore, in IoRT applications, it is of great significance to ensure the timeliness of data collection through reliable and efficient end-to-end transmission.

[0003] Compared with medium and high orbit satellites, low orbit satellites (Low Earth Orbit, LEO) have the advantages of small propagation loss and low delay, and are an indispensable infrastructure for connecting Internet of Things devices to remote data centers. However, relying solely on low-orbit satellites to directly connect IoRT devices still faces some challenges: 1) The number of concurrent connections of a single low-orbit satellite is limited, while the number of IoRT networking devices is huge; 2) IoRT devices have limited resources and cannot guarantee the performance of long-distance transmission; 3) Most IoRT devices lack the hardware and software facilities to directly connect to low-orbit satellites. Using unmanned aerial vehicles (unmanned aerial vehicles, UAV) as a bridge (or relay) between low-orbit satellites and IoRT devices can effectively solve the above challenges. However, new challenges have emerged: 1) How to plan the flight path of each unmanned aerial vehicle to collect as much data as possible during the flight process with limited battery reserves. 2) How to allocate resources (i.e., power and bandwidth) to match the communication between the sensing point, the aggregation point, and the unmanned aerial vehicle, to ensure the timeliness of end-to-end data transmission in an economical and efficient manner. 3) How to effectively use spectrum resources to meet specific transmission needs.

[0004] Therefore, there is an urgent need for a data-driven resource allocation method. SUMMARY

[0005] In view of this, the embodiment of the present disclosure provides a data-driven resource allocation method to solve the problems of poor data transmission timeliness and low spectrum energy efficiency in space-air-ground networks.

[0006] The embodiment of the present disclosure provides a data-driven resource allocation method, comprising:

[0007] Step 1: Combine the sensing points, drones, and low-orbit satellites in the monitoring area to form an air-space-ground network;

[0008] Step 2: Perform three-level clustering on the monitoring area to divide the monitoring area into multiple clusters, and each cluster is further subdivided into multiple groups, where each group contains multiple subgroups. Use drones to provide time-sharing services to each group.

[0009] Step 3: Merge and split clusters and groups according to preset standards;

[0010] Step 4: Design the three-hop transmission process from the sensing point to the convergence point, from the convergence point to the UAV, and from the UAV to the low-orbit satellite;

[0011] Step 5: Design the ground-to-sky transmission model based on beam allocation and convergence point rotation mechanism;

[0012] Step 6: Design a space-to-air transmission model, use non-orthogonal multiple access technology, share low-Earth orbit satellite sub-channels, and use interference cancellation methods to decode multiple UAV signals in the satellite segment;

[0013] Step 7: Construct a global optimization problem based on the ground-to-space transmission model and the space-to-air transmission model, and break it down into a first optimization problem corresponding to the ground-to-space transmission model and a second optimization problem corresponding to the space-to-air transmission model. The global optimization problem includes resource allocation, energy consumption control and data backlog management.

[0014] Step 8: Define the solution scheme for the first optimization problem;

[0015] Step 9: Define the solution scheme for the second optimization problem;

[0016] Step 10: Design the ISS-PPO algorithm to solve the first and second optimization problems;

[0017] Step 11: Design a solution to the global optimization problem. By combining the solutions to the first and second optimization problems using a dual-iteration algorithm, iteratively adjust the number of sub-channels in the system and the capacity of the UAV's energy harvesting board.

[0018] According to a specific implementation of an embodiment of this disclosure, step 2 specifically includes:

[0019] Step 2.1, first-level clustering: the sensing points are clustered into multiple clusters based on distance, and each cluster is handled by a drone for data transmission;

[0020] Step 2.2, the second layer of clustering, divides the sensing points in each cluster into multiple groups, and the number of sensing points in each group does not exceed the coverage area of ​​the drone;

[0021] Step 2.3, the third layer of clustering, uses a directional partitioning algorithm to divide each group into subgroups. Each subgroup is assigned a convergence point responsible for data collection and transmission. When the UAV hovers over each subgroup, the convergence points of each subgroup simultaneously send data to the UAV.

[0022] According to a specific implementation of an embodiment of this disclosure, step 3 specifically includes:

[0023] Step 3.1: For clusters or groups containing fewer than a preset number of sensing points, merge them with neighboring clusters or groups, selecting the nearest adjacent cluster to complete the merger.

[0024] Step 3.2: For clusters larger than the threshold, split them according to a preset standard. The splitting standard is based on the spatial distribution of the cluster. If the horizontal range is larger than the vertical range, split them according to the x-coordinate; otherwise, split them according to the y-coordinate.

[0025] According to a specific implementation of an embodiment of this disclosure, step 4 specifically includes:

[0026] Step 4.1, First hop: From the sensing point to the convergence point, the sensing point transmits data to the drone through the convergence point. A convergence point rotation mechanism is introduced between the convergence point and the sensing point, with multiple sensing points serving as candidate nodes and taking turns to act as the convergence point.

[0027] Step 4.2, Second hop: From the convergence point to the drone, the drone continuously collects data from ground sensing points throughout the mission cycle;

[0028] Step 4.3, the third hop: from the UAV to the low-Earth orbit satellite. The UAV selects the observed low-Earth orbit satellite for data transmission. The low-Earth orbit satellite uses a store-and-forward mode to temporarily store the data received from the UAV before finally transmitting it to the ground data center. The system adopts non-orthogonal multiple access technology to enable the UAV to share sub-channels for data transmission.

[0029] According to a specific implementation of an embodiment of this disclosure, step 5 specifically includes:

[0030] Communication between sensing points and drones is conducted in clusters. Each drone receives data from multiple convergence points simultaneously through multiple beam radio frequency chains. Through beam allocation and convergence point rotation mechanisms, drones utilize spectrum resources while hovering. The system uses time-division multiple access mode for data transmission.

[0031] According to a specific implementation of an embodiment of this disclosure, step 6 specifically includes:

[0032] Low-Earth orbit satellites use continuous interference cancellation to decode signals from multiple UAVs and sort them according to channel gain. The signals of UAVs with high channel gain are decoded first. During the continuous interference cancellation process, a partial continuous interference cancellation mechanism is used to simulate the impact of the residual power of the interference signal on the system, so as to eliminate the interference of undecoded signals.

[0033] According to a specific implementation of an embodiment of this disclosure, step 8 specifically includes:

[0034] Based on a multi-agent architecture using deep reinforcement learning, each cluster is assigned an agent to monitor the status of the sensing points and drones within the cluster. The agent's action decisions dynamically adjust the transmission power and scheduling order of the sensing points and drones, so that the data collection and transmission of the system in each task cycle can be optimized.

[0035] According to a specific implementation of an embodiment of this disclosure, step 9 specifically includes:

[0036] Deploying a single-agent deep reinforcement learning model on a low-Earth orbit satellite allows for real-time acquisition of the UAV's status, which in turn enables adjustments to the UAV's transmit power and sub-channel allocation.

[0037] According to a specific implementation of an embodiment of this disclosure, step 10 specifically includes:

[0038] By combining the PPO algorithm with independent action components, self-attention mechanisms, and strict hard constraints, the ISS-PPO algorithm is obtained, and the first and second optimization problems are solved accordingly.

[0039] According to a specific implementation of an embodiment of this disclosure, step 11 specifically includes:

[0040] Step 11.1, Sub-channel optimization: Combining the solutions to the first and second optimization problems, the number of sub-channels is dynamically adjusted according to the data backlog of the UAV within the mission cycle;

[0041] Step 11.2, Energy Harvesting Optimization: Combining the solutions to the first and second optimization problems, the most suitable capacity of the UAV energy harvesting board is found using the bisection method.

[0042] The data-driven resource allocation scheme in this embodiment includes: Step 1, combining sensing points, drones, and low-Earth orbit satellites in the monitoring area to form a space-air-ground network; Step 2, performing three-layer clustering on the monitoring area to divide it into multiple clusters, each cluster being further subdivided into multiple groups, where each group contains multiple subgroups, and drones provide time-sharing services to each group; Step 3, merging and splitting clusters and groups according to preset standards; Step 4, designing a three-hop transmission process from sensing point to aggregation point, aggregation point to drone, and drone to low-Earth orbit satellite; Step 5, designing a ground-to-space transmission model based on beam allocation and aggregation point rotation mechanism; Step 6, designing a space-to-air transmission model, using non-orthogonal multiple access technology, sharing low-Earth orbit satellite sub-channels, and employing interference cancellation in the satellite segment. The method decodes multiple UAV signals; Step 7: Construct a global optimization problem based on the ground-to-space transmission model and the space-to-air transmission model, and decompose it into a first optimization problem corresponding to the ground-to-space transmission model and a second optimization problem corresponding to the space-to-air transmission model. The global optimization problem includes resource allocation, energy consumption control, and data backlog management; Step 8: Set a solution scheme for the first optimization problem; Step 9: Set a solution scheme for the second optimization problem; Step 10: Design the ISS-PPO algorithm to solve the first and second optimization problems; Step 11: Design a solution to the global optimization problem, and iteratively adjust the number of sub-channels of the system and the energy harvesting board capacity of the UAV by combining the solutions of the first and second optimization problems through a dual-iteration algorithm.

[0043] The beneficial effects of this disclosure are as follows: The scheme of this disclosure considers transmitting data collected from ground sensing points to low-Earth orbit satellites in a timely manner via UAVs, while maximizing spectral energy efficiency. The main global problem includes three issues: data transmission from the aggregation node to the UAV, data transmission from the UAV to the satellite, and the overall system's UAV energy harvesting board capacity and UAV-to-satellite channel resource allocation. We first solve the first two problems, then use their solutions to solve the third problem, ultimately obtaining the global solution. For the first and second problems, an improved algorithm based on Proximal Policy Optimization (PPO)—ISS-PPO—is proposed. This algorithm combines Independent Action Components (IaaS), a self-attention mechanism, and strict hard constraints. For the first problem, it can maximize the data transmission volume from the aggregation node to the UAV, optimize the power allocation of the aggregation node and the path selection of the UAV, and minimize data backlog at the aggregation node and the energy consumption of the aggregation node transmission and the UAV movement. For the second problem, we can optimize the power control and sub-channel allocation of the UAV, minimize the data backlog and energy consumption of the UAV transmission, and maximize the total amount of data received by the low-Earth orbit satellite. Finally, a dual iteration algorithm (DIA) is proposed to solve the third problem. This algorithm iteratively calls the ISS-PPO algorithm to determine the appropriate UAV power harvesting board capacity and the number of UAV-satellite sub-channels. Attached Figure Description

[0044] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0045] Figure 1 A flowchart illustrating a data-driven resource allocation method provided in an embodiment of this disclosure;

[0046] Figure 2 This is a schematic diagram of a data collection scenario in a three-hop uplink integrated network provided by an embodiment of the present disclosure;

[0047] Figure 3 This is a schematic diagram of a data collection scenario in a three-hop uplink integrated network provided in an embodiment of the present disclosure. Detailed Implementation

[0048] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0049] The following specific examples illustrate the implementation of this disclosure. Those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0050] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.

[0051] It should also be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of this disclosure. The drawings only show the components related to this disclosure and are not drawn according to the number, shape and size of the components in actual implementation. In actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0052] Furthermore, specific details are provided in the following description to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0053] This disclosure provides a data-driven resource allocation method, which can be applied to the air-to-ground network communication process in mobile communication scenarios.

[0054] See Figure 1 This is a flowchart illustrating a data-driven resource allocation method provided in an embodiment of this disclosure. Figure 1As shown, the method mainly includes the following steps:

[0055] Step 1: Combine the sensing points, drones, and low-orbit satellites in the monitoring area to form an air-space-ground network;

[0056] In practice, in remote areas such as deserts and Gobi, numerous sensing points are deployed at arbitrary locations on the ground to monitor the ecological environment. Due to the limited number of personnel stationed and the limited ground communication infrastructure, it is difficult to establish or maintain a large number of concurrent end-to-end connections. This results in a large amount of ubiquitous data generated on the ground not being transmitted to data centers for processing in a timely manner, thus losing its timeliness. However, the deployability and high mobility of drones make it possible to establish communication paths for end-to-end data transmission. Since low-Earth orbit satellites periodically cover and provide services to ground areas, establishing a complete ground-air-space communication link is feasible. This involves how to allocate resources to improve the overall performance of the system. Figure 2 As shown, we consider using the Sink-UAV-LEO backbone network to collect ubiquitous data generated in IoT applications in a timely manner, and then transmitting the data to a remote data center through the selected Sink-UAV-LEO path.

[0057] Step 2: Perform three-level clustering on the monitoring area to divide the monitoring area into multiple clusters, and each cluster is further subdivided into multiple groups, where each group contains multiple subgroups. Use drones to provide time-sharing services to each group.

[0058] In practice, due to the extremely wide monitoring range and geographical isolation, a large number of sensing points are distributed across a certain number of non-overlapping areas (denoted as M). Within each area, the sensing points are clustered based on their distance, forming a cluster. Each cluster is then assigned a drone. Since the number of sensing points in a cluster is still large and widely distributed, while the drone's coverage area is limited, it's impossible for the drone to cover all sensing points simultaneously. Therefore, a cluster is subdivided into multiple groups, each no larger than the drone's coverage area. Members within a group send data to their aggregation point. Each drone provides time-sharing services to different groups within its cluster, receiving data collected by the aggregation point. To more effectively utilize the concurrency of each drone, each group needs to be divided into smaller subsets, each with its own aggregation point. When a drone hovers over a group, the aggregation points of the subsets simultaneously send data to the drone. This clustering process is called Three-Layer Clustering (TLC). Any classical clustering algorithm can be used for the first two layers of clustering, while our proposed Orientation-Based Area Division (OBAD) method is used for the third layer. OBAD will be introduced later. Before the start of each task cycle, each level of clustering in TLC can be performed based on the changes in the location of the sensing points, but it cannot be performed during the execution of each task cycle to avoid interference of clustering transformations with task transmission.

[0059] Step 3: Merge and split clusters and groups according to preset standards;

[0060] In practice, during the first two levels of clustering, some clusters or groups may contain only a few sensing points, and providing services to these areas individually would lead to an imbalance between cost and benefit. To address this issue, we merge smaller clusters / groups into adjacent clusters / groups. Taking a cluster as an example, the selection criteria for adjacent clusters for merging are as follows:

[0061]

[0062] In formula (1), C is the set of clusters, e represents the e-th cluster, and d e It is the distance between the center point of the small cluster and the center point of the e-th adjacent cluster. Compared with similar methods in [5], this method reduces the time complexity from O(n^2) to O(n^2). 2The complexity is reduced to O(n). Furthermore, clustering may result in large clusters (or groups) containing a large number of sensing points. Handling large-scale data transmission is unreasonable for limited UAV resources. Therefore, we divide large clusters into several smaller clusters based on their range and center point location. Specifically, if the horizontal range of a cluster is greater than its vertical range, it is divided according to the x-coordinate: nodes with x-coordinates less than the center point's x-coordinate form one cluster, and the remaining nodes form another cluster. Otherwise, it is divided according to the y-coordinate.

[0063] Step 4: Design the three-hop transmission process from the sensing point to the convergence point, from the convergence point to the UAV, and from the UAV to the low-orbit satellite;

[0064] In practice, the three-hop transmission process is designed as follows:

[0065] 1) First-hop transmission from the sensing point to the convergence point

[0066] In the first hop, the sensing points in each subset transmit data to the aggregation point within that subset. We assume the sensing points have energy harvesting devices capable of collecting renewable energy sources such as solar power and storing it in batteries. Therefore, battery limitations and dynamic energy consumption will affect system performance. The aggregation point, like the sensing points, has limited hardware resources, typically consisting of only one RF chain and one omnidirectional antenna. Therefore, the aggregation point can only transmit data when connected to the drone; at other times, it can only collect data.

[0067] To prevent severe energy imbalance between sensing points and sink points, and to ensure that data collection and transmission at the sink point can occur simultaneously, we select multiple sensing points for each subset to rotate the aggregation task. Specifically, we select several sensing points near the central area as candidate nodes in the subset. We then use a circular queue to record these candidate nodes and use pointers to indicate which sensing point is currently acting as the sink point for data collection. When a drone hovers over the swarm to provide service, the pointer is passed to the next candidate node, which begins data collection upon receiving the pointer. Simultaneously, the previous candidate node, as indicated by the pointer, transmits its collected data to the drone. This process is called the "Sink Rotation Mechanism" (SRM).

[0068] 2) Second-hop transmission from the aggregation point to the drone

[0069] The drones possess the same energy harvesting capabilities as the sensing points, providing power for their data transmission to low-Earth orbit satellites. We assume each drone is equipped with an energy harvesting plate. Additionally, the drones are equipped with replaceable batteries, which can be swapped out when depleted, allowing them to resume their mission. To efficiently utilize limited battery resources while ensuring seamless data transmission, the energy from the replaceable batteries is used only for non-communication operations (such as drone flight and hovering), while the harvested energy is used for communication operations.

[0070] We configured the drone's data collection as a periodic task, or mission cycle. At the start of a mission cycle, the drone receives instructions for that cycle from the satellite ground station. It takes off from the battery swapping point, then flies to a designated group, hovers, and collects data from the convergence point. This path selection and data collection process is repeated continuously until the drone has only enough power left to return to the battery swapping point. We limit the duration of this process to T.

[0071] 3) Third-hop transmission from drones to low-Earth orbit satellites

[0072] In the third hop of transmission, we utilize a super-dense constellation of low-Earth orbit (LEO) satellites, ensuring that each UAV can always select a suitable LEO satellite from its observations for data transmission. These LEO satellites can temporarily store the received data before ultimately transmitting it to a ground-based data center. The UAV simultaneously collects and transmits data within its mission cycle, forwarding any untransmitted data to the satellites at battery replacement points before entering the next mission cycle. We divide the time within each mission cycle for the UAV to transmit data to the satellites into several fixed-length time slots (τ). The UAV employs a store-and-forward mode, transmitting data collected in one time slot in the next.

[0073] In air-space communication scenarios, Non-Orthogonal Multiple Access (NOMA) is used to improve spectral efficiency and system capacity. To facilitate frequency band allocation and sharing, the frequency band is divided into multiple sub-channels, which are then allocated to different areas according to a strategy. UAVs within the same area share their allocated sub-channels using NOMA technology. The number of UAVs in Area m is represented by J. m At the same time, we use This represents the collection of all drones. For the set of all clusters, For the number of clusters, For cluster C m,i The group in the middle.

[0074] Step 5: Design the ground-to-sky transmission model based on beam allocation and convergence point rotation mechanism;

[0075] In practical implementation, for the sake of simplicity, Cluster C will be used. m,i Taking this as an example, we analyze the communication between the sensing point, the aggregation point, and the drone, and assume that drone U provides services to this cluster. m,i It is the i-th drone in Area m.

[0076] Unlike sensing points, UAVs typically possess sufficient hardware resources to support K+1 non-overlapping beam radio frequency chains. This enables them to simultaneously receive data from up to K convergence points within the same group-air frequency band during a mission cycle, while simultaneously communicating with low-Earth orbit satellites in the air-space satellite band (such as the Ka band). To fully utilize the concurrency capabilities of UAVs, we propose the OBAD method. Figure 3 As shown, we divide the group into K equal subsets centered at the group's center point, represented as follows: In each subset, an SRM method is used to select a convergence point, so that when the drone hovers over the group center point, the convergence point in each subset can be connected to the drone's beam. m,i,n The convergence point in the middle is represented as

[0077]

[0078] Subset The nodes in the array are represented as

[0079]

[0080] Each aggregation point selects several active sensing points using a fair algorithm and receives data in Time Division Multiple Access (TDMA) mode. The aggregation point then aggregates all received data into a single data packet for transmission. (Binary variables) express Whether the perception point r in the system is being served: 1 if it is being served, 0 otherwise.

[0081] Within a task cycle The data collected at the central convergence point is denoted as Its calculation formula is:

[0082]

[0083] In formulas (2) and (3), W represents the power spectral density (AWGN) of additive white Gaussian noise.RS It is the bandwidth between the sensing point and the convergence point. It is from the sensing point r to the convergence point The transmission power, It is from the sensing point r to the convergence point Channel gain. For each task cycle, cluster C... m,i The total amount of data collected by all aggregation points in the middle is

[0084]

[0085] Based on the omnidirectional transmission characteristics of each convergence point, its transmission gain and transmission interference gain are approximated as 1. Considering the directional reception characteristics of each UAV, UAV U m,i To the convergence point The gain of the k-th receiving beam is

[0086]

[0087] In formula (5), It is a drone U m,i The width of the k-th beam, It is a drone U m,i The centerline of the k-th receiving beam is parallel to the transmitting end. to the receiving end U m,i The angle between the lines. Furthermore, 0 ≤ ∈ < 1 is the sidelobe gain, where ∈ < < 1 for narrow beams. Correspondingly, the convergence point... To drone U m,i The transmission rate is

[0088]

[0089] In formula (6), W SU It is the bandwidth between the aggregation point and the drone. It is a convergence point To drone U m,i transmission power, yes The interference channel gain of the link's convergence point on the drone. It is a drone U m,i The k-th beam is The resulting interference gain.

[0090] Assign numbers to all groups within the same cluster, using a sequence of positive integers (L... m,i,1 ,…,L m,i,x ,…,L m,i,X To represent the unmanned aerial vehicle (U) m,i The group service sequence. Within a task cycle, the UAV U m,i Total data collected and mobile energy consumption E m,i Energy consumption can be expressed as

[0091]

[0092] In formulas (7) and (8), It is a convergence point Transmission time, It is a drone hover time, p h and p f These are the hovering power and flight power of the drone.

[0093] yes and The distance between the center points, G m,i,0 This is the location of the battery replacement point, v f It refers to the flight speed of the drone. and t m,i,n The calculation formula is:

[0094]

[0095] In formulas (9) and (10), and It is a convergence point Remaining battery energy and harvested energy within mission cycle t. yes The transmission power, From To drone U m,i The equation shows that significant differences in data collection volume between different aggregation points within the same subset lead to low data collection efficiency for UAVs. When some aggregation points have more data to transmit while others have already transmitted data, hovering energy is wasted and some idle receiving beams go unused. To address this issue, based on the normal distribution of aggregation point data volume, a Data Symmetric Balance (DSB) technique is proposed. Specifically, within the same group, the aggregation point with the largest data volume sends data to the aggregation point with the smallest data volume; the aggregation point with the second largest data volume sends data to the aggregation point with the second smallest data volume, and so on. DSB is executed periodically, with the time interval depending on the group size: the larger the group, the shorter the time interval. Ultimately, the data volume becomes more evenly distributed among the aggregation nodes.

[0096] Step 6: Design a sky-ground transmission model. Use non-orthogonal multiple access technology to share the sub-channels of low-earth orbit (LEO) satellites, and employ interference cancellation methods at the satellite segment to decode multiple unmanned aerial vehicle (UAV) signals.

[0097] For example, assume that multiple LEO satellites in one orbit provide uninterrupted services to UAVs in turn, and each UAV can only access the nearest LEO satellite at the same time. Compared with the beam coverage area of an LEO satellite with a coverage radius of 50 km, the ground sensing area studied in this invention (such as a coverage radius of 10 km) is much smaller. In addition, the arc length between adjacent LEO satellites sharing the same orbital plane is much larger than the radius of each ground sensing area. Therefore, within a given time slot, all UAVs share the same LEO satellite.

[0098] The LEO satellite receiver uses successive interference cancellation (SIC) to decode the signals of multiple users on the channel. The users decoded first will face interference from the signals of the undecoded users in the user set, while the subsequent users will benefit from the reduced interference and thus obtain a higher transmission rate. To mitigate the adverse effects of SIC, the decoding order should match the decreasing channel gain from the user to the LEO satellite. Therefore, the users with higher channel gains are decoded first. Assume that the channel gains of UAVs are sorted from large to small as Based on this, during the process of decoding the UAVU m,j signal, for U m,i where i > j, the signals of the UAVs are cancelled, while the signals of the UAVs with i < j are regarded as noise.

[0099] In actual situations, due to limitations such as inaccurate channel estimation, poor signal quality, and hardware, decoding errors of interference signals may occur. Therefore, there may be residual interference after SIC, that is, so-called incomplete SIC. The residual interference generated by incomplete SIC can be simulated as a linear function of the interference signal power, and the coefficient of this function can be determined through long-term measurements. Therefore, on the sub-channel m assigned to J m UAVs, the signal-to-noise ratio of the j-th UAV can be expressed as

[0100]

[0101] In formula (11), W B is the bandwidth of each sub-channel between the UAV and the LEO satellite, and the coefficient represents incomplete SIC, which ranges from [0, 1]. The larger the value, the more severe the interference. When is equal to 1, the interference is not eliminated at all, which is the most unfavorable situation. The instantaneous achievable rate of the UAV is

[0102] R m,i (t)=W B Dρ m log2(1+γ m,i (12)

[0103] Step 7: Construct a global optimization problem based on the ground-to-space transmission model and the space-to-air transmission model, and break it down into a first optimization problem corresponding to the ground-to-space transmission model and a second optimization problem corresponding to the space-to-air transmission model. The global optimization problem includes resource allocation, energy consumption control and data backlog management.

[0104] In practical implementation, given the multi-level transmission characteristics of a three-hop network, to optimize the overall performance of the system, we first need to clarify the data transmission and backlog situation of each hop, as well as the related energy consumption.

[0105] For clusterC m,i The data that may accumulate at all convergence points within a task cycle is represented by the following formula.

[0106]

[0107] ClusterC m,i The energy consumption of ground-air data transmission can be expressed as:

[0108]

[0109] For drones U m,i The amount of data that may accumulate and the energy consumption of data transmission within a task cycle can be expressed as follows:

[0110]

[0111]

[0112] In formula (16), T′ represents the unmanned aerial vehicle U. m,i The actual time for data collection during each mission cycle. (UAV) m,i The amount of data forwarded to low-Earth orbit satellites within a mission cycle is represented as follows:

[0113]

[0114] To assess the fairness of data transmission, prevent data congestion at any bottleneck point, and achieve fair and reasonable resource allocation, we adopted Jain's Fairness index, which is calculated using the following formula:

[0115]

[0116] In formulas (18) and (19), P m,i P and P' are U and U respectively, the drones m,i Jain's Fairness value for low-Earth orbit satellites.

[0117] The primary objective of this invention is to ensure the timeliness of data collection while minimizing energy consumption and maximizing spectrum utilization. To achieve this, we must minimize data backlog and resource waste at bottlenecks (such as aggregation points and UAVs) along the uplink transmission path, while maximizing the total amount of data uploaded to low-Earth orbit satellites. Therefore, we propose the following global optimization problem.

[0118]

[0119] In equation (20), β is the energy consumption coefficient of the system environment. It is the maximum battery capacity of the drone. and These are conservative empirical values. Constraint C1 states that the UAV's path selection for each mission cycle is limited by its maximum energy storage. We assume that the UAV's battery is full after each battery replacement at a battery replacement point. Constraints C2 and C3 ensure service fairness between the UAV and low-Earth orbit satellites.

[0120] This is a mixed-integer nonlinear programming problem, making it difficult to solve. Furthermore, the dynamic adjustment of resources and demands over time makes integrating these problems into a single optimization problem very challenging. Therefore, we further consider two local optimization problems based on the system's multi-hop characteristics. The formula for the first local optimization problem is:

[0121]

[0122] In equation (21), β' is the penalty weight for the energy consumption at the convergence point, and P max It is the maximum achievable transmission power of the sensing point and the convergence point. and These are the maximum battery capacity and maximum energy collected for each task cycle at the sensing point and the convergence point, respectively. and yes The remaining battery energy and harvesting energy of the sensing point r in the system. Constraint C4 restricts all UAVs to complete their mission cycle within a specified time to ensure system synchronization and avoid wasting time. Constraint C5 specifies the initial and final positions of the UAVs, which is of practical significance in periodic data collection scenarios

[12] . Constraints C6 and C7 are the minimum and maximum transmit power constraints for the sensing point and the convergence point, respectively. Constraints C8 and C9 indicate that the energy use of the sensing point and the convergence point cannot exceed the current battery capacity. Constraints C10 and C11 ensure that the battery capacity and harvesting energy do not exceed the maximum limit.

[0123] question The goal is to optimize data transmission in each mission cycle within a ground-air scenario. Specifically, the aim is to maximize the amount of data transmitted from the sensing point to the UAV while minimizing data backlog and transmission energy consumption at the aggregation point, as well as the UAV's motion energy consumption. However, this invention also focuses on communication in air-space scenarios. It involves UAV transmit power control and sub-channel allocation in different regions under NOMA technology. To this end, this invention proposes a second local optimization problem, namely...

[0124]

[0125] In equation (22), It is the energy consumption coefficient of the drone. This is the maximum achievable transmission power of the drone. and These represent the upper limit of the drone's energy harvesting plate capacity and the energy harvesting capacity of the drone per time slot, respectively. and U drones m,i The remaining energy and harvested energy. Constraint C12 ensures that the total allocated bandwidth of all areas does not exceed the available bandwidth between the UAV and the low-Earth orbit satellite. Constraints C13 and C14 are the minimum and maximum transmit power constraints for the UAV, respectively. Constraint C15 states that the UAV's harvested energy usage cannot exceed the current harvested energy storage. Constraints C16 and C17 ensure that the harvested energy and harvested energy storage are not unlimited.

[0126] The focus is on data transmission from UAVs to low-Earth orbit satellites, aiming to maximize data transmission volume while minimizing data backlog and transmission energy consumption for the UAVs. In optimization... and On this basis, The focus is solely on the bandwidth of low-Earth orbit satellites and the capacity of UAV energy harvesting panels. Therefore, we initially employ a DRL model to address this. and Then we combine their solutions, for Solve the problem.

[0127] Step 8: Define the solution scheme for the first optimization problem;

[0128] In practice, we use a multi-agent DRL model with a distributed training and distributed execution framework to solve the problem. At each satellite ground station in each area, each cluster is assigned an agent to make decisions regarding data transmission. Since each cluster is disjoint, and a drone serves only one cluster, we will use a single cluster as an example. Each agent participates in the dynamic environment by performing actions, making observations, and receiving rewards. The agent is capable of monitoring data transmission and developing strategies for the actions of rendezvous points and drones. In the air-to-ground scenario, each drone periodically combines its characteristics with rendezvous point information, working in conjunction with the agent to monitor fluctuations within the cluster. Within each mission cycle, the agent assigns actions to drones and rendezvous points in the corresponding cluster. After receiving instructions from the agent, the drone flies to the designated cluster to provide the corresponding service. All rendezvous points in the cluster adjust their transmission power according to the instructions. After executing each action, the agent receives an immediate reward indicating its contribution to the objective, and the cluster system enters the next state accordingly. Throughout the mission cycle, the agent continuously monitors environmental changes and updates the system state accordingly.

[0129] Given Given the mixed-integer nonlinear programming properties, it is necessary to reconstruct it as a discrete Markov decision process, represented by tuples as follows: To this end, we define the state, action, and reward in a tuple for each task cycle. Since the decisions of each group are independent and they share the same data transmission pattern, we use clusterC... m,i Let's take an example for analysis.

[0130] 1) State Space

[0131]

[0132] In formula (23), the agent obtains the convergence point. Remaining energy Energy harvesting groupG m,i,n coordinates of the center point Transmission task volume UAV m,i Remaining battery power U m,i coordinates and convergence point-UAV channel coefficient This calculation assumes the drone is hovering above the group's center point. Based on this information, the agent can select actions for the system.

[0133] 2) Space for Action

[0134]

[0135] The action includes the dispatch sequence number L m,i,n (Each group has a number) and number L m,i,n Convergence point power of the corresponding group Convergence point transmission power according to Adjustments will be made.

[0136] 3) Instant rewards

[0137]

[0138] The instant reward in formula (25) ensures timely data collection and minimizes power consumption in each task cycle.

[0139] Step 9: Define the solution scheme for the second optimization problem;

[0140] In practice, to avoid drone contention for sub-channels and to improve model training and decision-making efficiency, we use a single-agent DRL model to solve the problem. We deploy the agent on low-Earth orbit (LEO) satellites, with all satellites sharing their model parameters. During the model training phase, once a satellite completes its service, it transmits the agent's model parameters to the next satellite. During the execution phase, since the agent's model parameters are no longer updated, the satellites stop transmitting parameters. Specifically, all drones transmit their status to the agent on the current satellite, which processes the data and sends its decisions to each drone. The drones then adjust their transmission power accordingly and send data to the LEO satellites via their assigned sub-channels.

[0141] 1) State Space

[0142]

[0143] In formula (26), m,i represents a specific region m and the i-th UAV in that region. In each time slot, integers... U-shaped drone m,i area, U-shaped drone m,i Channel coefficient with the nearest low-Earth orbit satellite, It's U m,i The amount of transmission tasks. and U drones m,i The remaining harvested energy and the harvested energy. Based on the above environmental conditions, the agent can make decisions.

[0144] 2) Space for Action

[0145]

[0146] Actions include the sub-channel allocation ratio ρ of the area. m and U drones m,i Transmit power P m,i Agents receive immediate rewards after taking action.

[0147] 3) Instant rewards

[0148]

[0149] In formula (28), It is a drone U m,i The amount of data forwarded to low-Earth orbit satellites within time slot t. It is a drone U m,i The goal of this timely reward is to ensure the timeliness of data transmission from drones to satellites by maximizing the amount of data collected and minimizing the data backlog per time slot. Furthermore, it also aims to minimize the energy consumption of drone transmissions.

[0150] Step 10: Design the ISS-PPO algorithm to solve the first and second optimization problems;

[0151] In practice, the proposed three-hop network structure is intertwined, making the optimization problem extremely complex, and ordinary DRL algorithms cannot solve it directly. and Therefore, we improved the PPO algorithm by combining it with independent action components, self-attention mechanisms, and strict hard constraints, resulting in the improved PPO algorithm (ISS-PPO). Using ISS-PPO, we can efficiently solve... or Due to its low coupling and scalability, this algorithm is also applicable to other similar problems.

[0152] 1) Independent motion components

[0153] Since the action space contains multiple decisions, and each decision includes numerous actions, the combination of all actions forms a high-dimensional discrete action space. However, ordinary neural networks struggle to handle such action spaces, so we introduce independent action components. Independent action components divide the complex action space A into independent, composable sub-action spaces, which can be represented as A = A1 × … A x ×…×A X And A x This is the x-th independent sub-action space. The agent's policy π can be represented as the independent sub-policy π. x Combinations:

[0154]

[0155] In formula (29), s is the current state, a x ∈A x The agent combines individual actions to make Maximizing the total expected return, expressed as:

[0156]

[0157] In formula (30), γ is the discount factor. In state s t Choose Action The rewards achieved. This approach has three main advantages: it reduces the high-dimensional action space to a lower dimension, simplifying policy learning; it provides a framework that is easily scalable and adaptable to various tasks and environments; and it modularizes policies, facilitating individual debugging and optimization. Its main challenges lie in handling dependencies between actions and optimizing combined policies.

[0158] 2) Self-attention mechanism

[0159] Due to the large dimensionality of the state space, the varying degrees of dependency between state sequences, and the differing impacts of various states on action decisions across different scenarios, network models struggle to capture global and long-sequence dependencies when processing environmental feedback states. This leads to reduced feature representation and generalization capabilities, and also impacts computational efficiency. These limitations affect the decision-making accuracy and adaptability of network models. Self-attention mechanisms, however, are an effective tool for addressing these issues.

[0160] Self-attention mechanisms process sequential data by assigning different attention weights based on intra-sequence relationships. It enhances DRL models by capturing global dependencies, dynamically weighted states, and feature representations. The key steps in calculating attention scores can be represented as follows:

[0161]

[0162] In formula (31), A is the attention weight matrix, Q = SW Q It is a query matrix, K=SW K S is the key matrix, and S is the input matrix. and It is a learnable weight matrix, d model and d k It is the dimension of the input features and the query vector. This is a scaling factor to prevent the dot product value from becoming too large. After obtaining the attention weight matrix A, it is weighted and summed with the feature representation of the input state to generate richer and more expressive features.

[0163] 3) Strict hard constraints

[0164] Our system has many constraints, such as a drone's beam only being able to connect to one convergence point and the sum of bandwidth allocation ratios for each region not exceeding 1. Therefore, it's necessary to constrain the actions of the neural network output to avoid violating these constraints. The two most common methods are soft constraints and hard constraints. Soft constraints are typically used to handle invalid operations. When the network's output contains invalid operations, it provides the network with low or even negative rewards as a warning to reduce the output of invalid operations in this state. While soft constraints are simple and easy to implement, they do not necessarily prevent invalid operations from occurring.

[0165] Hard masking is a method that directly prohibits impossible or unavailable operations. It masks the logits of invalid operations before the network output, ensuring that the output layer avoids these operations. Hard constraints can completely avoid invalid operations, but require prior knowledge of which operations are invalid under specific conditions. They cannot prevent invalid operations that can only be identified after the output, making them unsuitable for our scheme. Therefore, we propose a novel constraint method—strict hard constraints—based on the serialization characteristics of the high-dimensional discrete action space. First, we randomly rearrange the logits sequence to ensure fairness. Then, we set the logits of actions that violate the constraints to extremely low values ​​and their gradients to zero. This prevents the softmax function from selecting these actions. Finally, we use modified logits and the softmax function to determine the actions. The description is as follows:

[0166] a x =softmax(l x,1 ,…,l x,j ,…,l x,J (32)

[0167]

[0168] In formulas (32), (33), and (34), a xRepresents the x-th dimension decision, l x,j Describes the j-th action on the x-th dimension. It is a large negative number (e.g.) ), f x This represents the number of constraints associated with the x-th dimension decision. It is a sequence representing whether the constraints in the x-th dimension have been violated, where 0 indicates that they have not been violated, and otherwise, they have been violated. By using strict hard constraints, we can impose constraints on actions in any dimension at any stage of neural network execution to shield invalid actions. The specific ISS-PPO algorithm is described below, where...

[0169] 4) ISS-PPO algorithm

[0170] Operating in any satellite ground station or low-Earth orbit satellite

[0171] Input: Initial policy network parameters θ, initial value network parameters φ, initial weight matrix W Q and W K Training iterations M (5000 rounds), scaling factor (0.088)

[0172] Output: Value function V φ Policy function π θ

[0173] Step 1: Decompose the action space into independent subspaces

[0174] Step 2: Determine whether the training rounds indicator variable t (initially 0) has reached the network model training time steps M. If yes, end the current training process; otherwise, go to step 2.

[0175] Step 3: Clear the training buffer

[0176] Step 4: Determine if the current data collection steps have reached the upper limit (128,000 steps). If yes, proceed to step 9; otherwise, proceed to step 4.

[0177] Step 5: Observe the environmental condition S t ;

[0178] Step 6: Calculate the query matrix Q = SW Q Bond matrix K = SW K ;

[0179] Step 7: Combine Q, K, and Substituting into formula (31) yields the weight matrix A;

[0180] Step 8: Calculate S t A obtained

[0181] Step 9: Determine if the current action count indicator variable x (initially 0) is equal to the action dimension. If yes, end the current loop; otherwise, go to step 10.

[0182] Step 10: Filter out invalid actions using formulas (32), (33), and (34);

[0183] Step 11: According to Select action a x ;

[0184] Step 12: Increment the action count indicator variable x by 1, then proceed to step 9.

[0185] Step 13: Combine a1, ..., a x ,…,a X get

[0186] Step 14: Perform the action And receive a reward

[0187] Step 15: The environment transitions to the next state S. t+1 ;

[0188] Step 16: Increment the data collection step count by 1, then proceed to step 4;

[0189] Step 17: Calculate the advantage estimate using formula (38) and experience Add to training buffer middle;

[0190] Step 18: Determine whether the current training step count indicator variable b (initial value is 0) has reached the upper limit (1000 steps). If yes, go to step 26; otherwise, go to step 19.

[0191] Step 19: Recalculate the advantage estimate using formula (38)

[0192] Step 20: Set the training buffer Divide into K (256) smaller batches;

[0193] Step 21: Determine whether the current batch indicator variable k (initially 0) has reached K. If yes, go to step 25; otherwise, go to step 22.

[0194] Step 22: Calculate the total PPO loss using formula (30);

[0195] Step 23: Update the actor network, critic network, and weight matrix W based on the total ISS-PPO loss. Q and W K ;

[0196] Step 24: Increment the batch indicator variable k by 1, then proceed to step 21;

[0197] Step 25: Increment the training step count indicator variable b by 1, then proceed to step 18;

[0198] Step 26: Increment the training round number indicator variable t by 1, then proceed to step 2;

[0199] Step 27: Output network parameters V φ , π θ .

[0200] Step 11: Design a solution to the global optimization problem. By combining the solutions to the first and second optimization problems using a dual-iteration algorithm, iteratively adjust the number of sub-channels in the system and the capacity of the UAV's energy harvesting board.

[0201] In specific implementation, in order to solve To address this problem, we propose a dual-iteration algorithm. The main purpose of this algorithm is to... and The solution is to find the optimal number of sub-channels and the capacity of the UAV energy harvesting board. The main purpose of lines 1-8 is to build a three-hop network and train the ISS-PPO model set to solve the problem. Lines 9-17 aim to optimize the number of sub-channels. This is solved by iteratively training another ISS-PPO model. And obtain environmental feedback. Based on this, optimize the number of sub-channels while meeting transmission efficiency requirements. The purpose of lines 18-28 is to optimize the capacity of the UAV energy harvesting board by using the idea of ​​binary search to quickly find the most suitable capacity of the UAV energy harvesting board with the current number of sub-channels.

[0202] 1) Double Iteration Algorithm

[0203] Operating on low-Earth orbit satellites

[0204] Input: Data backlog threshold Drone battery capacity range threshold φ

[0205] Output: Number of sub-channels D, UAV energy harvesting board capacity

[0206] Step 1: Collect the location information of ground sensing nodes;

[0207] Step 2: Use TLC, CDP, and OBAD to construct a three-hop network;

[0208] Step 3: Assign one drone to each cluster and establish a drone collection U;

[0209] Step 4: Locate a LEO satellite that currently covers you and mark its successor as the first LEO satellite in the same orbital plane;

[0210] Step 5: Constructing a LEO satellite ensemble based on satellite motion patterns (22) and assign them an ISS-PPO model M2;

[0211] Step 6: Assign an ISS-PPO model to each cluster and obtain the model set.

[0212] Step 7: Train the model set To solve the problem

[0213] Step 8: Assign an experience value to D (e.g., 20), and give... Assign a large value (e.g., 30J).

[0214] Step 9: Use the model set generated To train model M2, and use it to solve the problem.

[0215] Step 10: Call model M2 to obtain and

[0216] Step 11: Determine If both 1Mb and BD<0 are true, proceed to step 12; otherwise, proceed to step 15.

[0217] Step 12: Reduce the number of sub-channels using a preset step size;

[0218] Step 13: Determine Check whether BD>0 is true at the same time. If yes, go to step 14; otherwise, go to step 15.

[0219] Step 14: Increase the number of sub-channels by a preset step size;

[0220] Step 15: Determine If BD = 0, proceed to step 16; otherwise, proceed to step 9.

[0221] Step 16: Set variables

[0222] Step 17: Determine whether high-low≥φ and BD≠0 are both true. If yes, go to step 11; otherwise, go to step 18.

[0223] Step 18: Calculation

[0224] Step 19: Using the model set generated To train model M2, and use it to solve the problem.

[0225] Step 20: Call model M2 to obtain and

[0226] Step 21: Determine If both 1Mb and BD≠0 are true, proceed to step 22; otherwise, proceed to step 23.

[0227] Step 22: Settings

[0228] Step 23: Determine If both 1Mb and BD≠0 are true, proceed to step 24; otherwise, proceed to step 17.

[0229] Step 24: Settings

[0230] Step 25: Output parameter D,

[0231] This embodiment provides a data-driven resource allocation method by proposing a Sink-UAV-LEO backbone network that ensures spectral efficiency. This network consists of the proposed three-layer clustering (TLC), merging and partitioning process (CDP), direction-based region partitioning (OBAD), data symmetric balancing (DSB), and sink rotation mechanism (SRM). In remote areas with wide coverage and multiple regions, this backbone network can achieve reliable and spectrally efficient real-time data transmission. The joint power control and UAV path selection problems of sink points in the Sink-UAV layer, and the joint power control and sub-channel allocation problems of UAVs in the UAV-LEO layer are described. Then, an improved PPO algorithm (ISS-PPO) is designed by combining the PPO algorithm with independent action components, self-attention mechanisms, and strict hard constraints to solve the two local problems mentioned above. The global problem of the backbone network is described, and a dual-iteration algorithm (DIA) is proposed to solve this problem. This algorithm aims to match the bandwidth of low-Earth orbit satellites and the capacity of UAV energy harvesting boards with the data collection task load. In DIA, in addition to completing network construction, it is also necessary to repeatedly call two solutions based on ISS-PPO to obtain an approximate solution to the global problem, ensuring data transmission timeliness and spectral efficiency.

[0232] The method will be further illustrated below with a specific embodiment. The ground area studied is set at 10km × 10km, containing 6 non-overlapping areas. Each area has a radius between 1km and 2km and randomly distributed 100–300 sensing points. It is assumed that the sensing points and convergence points can collect up to 80% of their maximum energy storage capacity (e.g., 50J) per mission cycle, and the UAV can collect up to 80% of its maximum energy harvesting plate capacity (e.g., 10J) per time slot. 22 low-Earth orbit satellites take turns covering this area, providing seamless access. Other parameters are listed in Table 1.

[0233] Table 1

[0234]

[0235]

[0236] This invention implements the above scheme using Python and PyTorch. By categorizing the data collection task load into seven levels (A, B, C, D, E, F, G) from low to high, the ability of five integrated schemes to handle variable loads was compared. These five integrated schemes are as follows:

[0237] (1) SE-PSPC-NOPC: SE-PSPC (Sequential Path Selection with Power Control) uses the ISS-PPO algorithm to decide the next group for UAV service and the transmit power of the aggregation point in that group. NOPC (Non-Orthogonal Multiple Access with Power Control) uses the ISS-PPO algorithm to decide the proportion of sub-bands used for accessing the LEO satellite network in each area and the transmit power of the UAV. UAVs in the same area use NOMA to share sub-channels.

[0238] (2) SE-PSPC-FDPC: FDPC (Frequency Division Multiple Access with Power Control) uses the ISS-PPO algorithm to make decisions on the sub-channel proportion and power of each UAV, and the UAV uses FDMA (Frequency Division Multiple Access) on its own sub-channel.

[0239] (3) SE-PSPC-TDPC: TDPC (Time Division Multiple Access with Power Control) divides each time slot into equal hour slots based on the number of drones. Each drone occupies one hour slot for data transmission and uses all sub-channels during this period. The drone's power is determined by the ISS-PPO algorithm.

[0240] (4) SE-PSPC-FTPC: FTPC (Frequency and Time Division Multiple Access with Power Control) uses the ISS-PPO algorithm to determine the sub-channel proportion of each area and the UAV's transmit power. UAVs in different areas use FDMA, while UAVs in the same area use TDMA.

[0241] (5) SE-PSPC-TDFP: TDFP (Time Division Multiple Access with Fixed Power) The only difference between this scheme and SE-PSPC-TDFP is that the UAV uses a fixed power. It is mainly used to verify the effectiveness of power control.

[0242] We categorized the data collection mission load into seven levels (from low to high, A to G) to compare the resource requirements of the five integration schemes under different loads. For simplicity, we ensured that the total amount of data transmitted from the aggregation point to the UAV continuously decreased from level G to level A within each mission cycle. Level G load is based on the aggregation point's transmission capacity to the UAV under ideal ground communication conditions and maximum transmit power. We evaluated the number of sub-channels and the UAV energy harvesting board capacity at different load levels when the average absolute value of the UAV data backlog percentage was below 50%, 30%, and 10% of the total data received per time slot.

[0243] First, as the load increases, the demand for sub-channels and UAV energy harvesting board capacity increases accordingly for each scheme, which is consistent with reality. Since UAVs share sub-channels, SE-PSPC-NOPC requires the fewest sub-channels, but because data transmission on each shared sub-channel consumes energy, its demand for UAV energy harvesting board capacity is the highest. Compared to other schemes, SE-PSPC-NOPC and SE-PSPC-FTPC require more UAV energy harvesting board capacity to achieve better transmission with fewer channels, but their demand is relatively lower when there are more sub-channels. Therefore, as the load level or the percentage of UAV data backlog decreases, the difference in UAV energy harvesting board capacity required by SE-PSPC-NOPC and SE-PSPC-FTPC compared to the other three schemes will decrease; in some cases, SE-PSPC-FTPC may even require less capacity than SE-PSPC-TDPC and SE-PSPC-TDFP. Overall, the other three schemes have less flexibility in adjusting resource requirements under different scenarios.

[0244] Furthermore, it can be observed that, under the same load level but different drone data backlog ratios, the number of sub-channels required for each scheme increases exponentially as the backlog ratio decreases. This is because, when the load remains constant, data transmission capacity is not proportional to the number of sub-channels. Due to the limitations of the sub-channel allocation scheme, sufficient bandwidth resources are needed to ensure that the data backlog for each drone is sufficiently small to reduce the overall data backlog rate of the system. However, as the backlog ratio decreases, the bandwidth requirement of the SE-PSPC-NOPC scheme increases much less than other schemes, highlighting its superiority in sub-channel allocation. From the above analysis, it can be inferred that energy capacity determines the lower limit of system transmission performance, while bandwidth resources determine the upper limit of system transmission performance.

[0245] It should be understood that the various parts of this disclosure can be implemented in hardware, software, firmware, or a combination thereof.

[0246] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.

Claims

1. A data-driven resource allocation method, characterized in that, include: Step 1: Combine the sensing points, drones, and low-orbit satellites in the monitoring area to form an air-space-ground network; Step 2: Perform three-level clustering on the monitoring area to divide the monitoring area into multiple clusters, and each cluster is further subdivided into multiple groups, where each group contains multiple subgroups. Use drones to provide time-sharing services to each group. Step 3: Merge and split clusters and groups according to preset standards; Step 4: Design the three-hop transmission process from the sensing point to the convergence point, from the convergence point to the UAV, and from the UAV to the low-orbit satellite; Step 5: Design the ground-to-sky transmission model based on beam allocation and convergence point rotation mechanism; Step 6: Design a space-to-air transmission model, use non-orthogonal multiple access technology, share low-Earth orbit satellite sub-channels, and use interference cancellation methods to decode multiple UAV signals in the satellite segment; Step 7: Construct a global optimization problem based on the ground-to-space transmission model and the space-to-air transmission model, and break it down into a first optimization problem corresponding to the ground-to-space transmission model and a second optimization problem corresponding to the space-to-air transmission model. The global optimization problem includes resource allocation, energy consumption control and data backlog management. Step 8: Define the solution scheme for the first optimization problem; Step 9: Define the solution scheme for the second optimization problem; Step 10: Design the ISS-PPO algorithm to solve the first and second optimization problems; Step 10 specifically includes: By combining the PPO algorithm with independent action components, self-attention mechanisms, and strict hard constraints, the ISS-PPO algorithm is obtained, and the first and second optimization problems are solved accordingly. Step 11: Design a solution to the global optimization problem. By combining the solutions to the first and second optimization problems using a dual-iteration algorithm, iteratively adjust the number of sub-channels in the system and the capacity of the UAV's energy harvesting board.

2. The method according to claim 1, characterized in that, Step 2 specifically includes: Step 2.1, first-level clustering: the sensing points are clustered into multiple clusters based on distance, and each cluster is handled by a drone for data transmission; Step 2.2, the second layer of clustering, divides the sensing points in each cluster into multiple groups, and the number of sensing points in each group does not exceed the coverage area of ​​the drone; Step 2.3, the third layer of clustering, uses a directional partitioning algorithm to divide each group into subgroups. Each subgroup is assigned a convergence point responsible for data collection and transmission. When the UAV hovers over each subgroup, the convergence points of each subgroup simultaneously send data to the UAV.

3. The method according to claim 2, characterized in that, Step 3 specifically includes: Step 3.1: For clusters or groups containing fewer than a preset number of sensing points, merge them with neighboring clusters or groups, selecting the nearest adjacent cluster to complete the merger. Step 3.2: For clusters larger than the threshold, split them according to a preset standard. The splitting standard is based on the spatial distribution of the cluster. If the horizontal range is larger than the vertical range, split them according to the x-coordinate; otherwise, split them according to the y-coordinate.

4. The method according to claim 3, characterized in that, Step 4 specifically includes: Step 4.1, First hop: From the sensing point to the convergence point, the sensing point transmits data to the drone through the convergence point. A convergence point rotation mechanism is introduced between the convergence point and the sensing point, with multiple sensing points serving as candidate nodes and taking turns to act as the convergence point. Step 4.2, Second hop: From the convergence point to the drone, the drone continuously collects data from ground sensing points throughout the mission cycle; Step 4.3, the third hop: from the UAV to the low-Earth orbit satellite. The UAV selects the observed low-Earth orbit satellite for data transmission. The low-Earth orbit satellite uses a store-and-forward mode to temporarily store the data received from the UAV before finally transmitting it to the ground data center. The system adopts non-orthogonal multiple access technology to enable the UAV to share sub-channels for data transmission.

5. The method according to claim 4, characterized in that, Step 5 specifically includes: Communication between sensing points and drones is conducted in clusters. Each drone receives data from multiple convergence points simultaneously through multiple beam radio frequency chains. Through beam allocation and convergence point rotation mechanisms, drones utilize spectrum resources while hovering. The system uses time-division multiple access mode for data transmission.

6. The method according to claim 5, characterized in that, Step 6 specifically includes: Low-Earth orbit satellites use continuous interference cancellation to decode signals from multiple UAVs and sort them according to channel gain. The signals of UAVs with high channel gain are decoded first. During the continuous interference cancellation process, a partial continuous interference cancellation mechanism is used to simulate the impact of the residual power of the interference signal on the system, so as to eliminate the interference of undecoded signals.

7. The method according to claim 6, characterized in that, Step 8 specifically includes: Based on a multi-agent architecture using deep reinforcement learning, each cluster is assigned an agent to monitor the status of the sensing points and drones within the cluster. The agent's action decisions dynamically adjust the transmission power and scheduling order of the sensing points and drones, so that the data collection and transmission of the system in each task cycle can be optimized.

8. The method according to claim 7, characterized in that, Step 9 specifically includes: Deploying a single-agent deep reinforcement learning model on a low-Earth orbit satellite allows for real-time acquisition of the UAV's status, which in turn enables adjustments to the UAV's transmit power and sub-channel allocation.

9. The method according to claim 8, characterized in that, Step 11 specifically includes: Step 11.1, Sub-channel optimization: Combining the solutions to the first and second optimization problems, the number of sub-channels is dynamically adjusted according to the data backlog of the UAV within the mission cycle; Step 11.2, Energy Harvesting Optimization: Combining the solutions to the first and second optimization problems, the most suitable capacity of the UAV energy harvesting board is found using the bisection method.

Citation Information

Patent Citations

  • Satellite unmanned aerial vehicle fusion network resource allocation method and device

    CN111669758A

  • Data transmission method of space-air-ground integrated network and related equipment

    CN118381549A