Data-driven resource allocation method

By building a space and earth network in IoRT application, using data-driven resource allocation method and deep reinforcement learning algorithm, the problems of limited resources and low transmission efficiency when low-orbit satellites are directly connected to IoRT devices are solved, and efficient and reliable end-to-end data transmission is achieved.

CN120150792AActive Publication Date: 2025-06-13CENT SOUTH UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510285402.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-13
Estimated Expiration
2045-03-11

AI Technical Summary

Technical Problem

In IoRT applications, it is of great significance to ensure the timeliness of data collection through reliable and efficient end-to-end transmission, but low-orbit satellites face challenges such as limited concurrent connections, limited resources and lack of software and hardware facilities when directly connecting IoRT devices.

Method used

The data-driven resource allocation method is adopted, and the aerospace and earth network is formed by combining the perception points, drones and low-orbit satellites in the monitoring area to form a space-space network, and three-layer clustering and resource allocation optimization are carried out, the triple-hop transmission process and the earth-sky, sky-space transmission model are designed, and non-orthogonal multiple access technology and interference cancellation methods are used, and resource management is managed in combination with deep reinforcement learning algorithms.

Benefits of technology

It effectively improves the timeliness of data transmission and spectrum energy efficiency, solves the challenges of data backlog, energy consumption and resource allocation, and realizes reliable and efficient end-to-end data transmission in IoRT applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120150792A_ABST
    Figure CN120150792A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a data-driven resource allocation method, which belongs to the technical field of communication, and specifically comprises the following steps: forming an air-space-ground network; performing three-layer clustering on the monitoring area; merging and splitting the cluster and the group according to a preset standard; designing a three-hop transmission process; designing a ground-sky transmission model according to beam allocation and a rendezvous point rotation mechanism; designing a space-air transmission model according to a non-orthogonal frequency division multiple access technology; constructing a global optimization problem according to the ground-sky transmission model and the sky-air transmission model, and splitting the global optimization problem into a first optimization problem corresponding to the ground-sky transmission model and a second optimization problem corresponding to the sky-air transmission model; setting a solution scheme of the first optimization problem; setting a solution scheme of a second optimization problem; designing an ISS-PPO algorithm to solve the first optimization problem and the second optimization problem; and designing a global optimization problem solution. Through the scheme disclosed by the invention, the data transmission timeliness and the frequency spectrum energy efficiency are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of communication technologies, and in particular, to a data-driven resource allocation method. Background Art

[0002] Currently, with the increasing requirements of the 6th Generation (6G) mobile communication system for global coverage, the integration of the sky-air-ground (and even underwater) network has received increasing attention, which provides support for Internet of Things (IoT) applications. Remote Internet of Things (IoRT) applications include various remote monitoring systems for wildlife changes, vegetation degradation, natural disasters, and climate change, etc., which can alleviate the difficulties of dispatching personnel and deploying a large number of expensive devices in remote areas. However, a large number of IoRT real-time monitoring devices need to send the sensed data to a remote data center for further analysis. Therefore, in IoRT applications, it is of great significance to ensure the timeliness of data collection through reliable and efficient end-to-end transmission.

[0003] Compared with medium and high-earth orbit satellites, low earth orbit (LEO) satellites have the advantages of small propagation loss and low latency, and are an indispensable infrastructure for IoT devices to connect to remote data centers. However, relying solely on LEO satellites to directly connect to IoRT devices still faces some challenges: 1) The concurrent connection number of a single LEO satellite is limited, while the number of IoRT networked devices is huge; 2) The resources of IoRT devices are limited and cannot guarantee the performance of long-distance transmission; 3) Most IoRT devices lack the software and hardware facilities for directly connecting to LEO satellites. Using unmanned aerial vehicles (UAVs) as a bridge (or repeater) between LEO satellites and IoRT devices can effectively solve the above challenges. However, new challenges have emerged: 1) How to plan the flight path of each UAV so that it can collect as much data as possible during the flight with limited battery reserves. 2) How to match the resource (i.e., power and bandwidth) allocation with the communication between the sensing points, aggregation points, and UAVs to ensure the timeliness of end-to-end data transmission in an economical and efficient manner. 3) How to effectively utilize the spectrum resources to meet specific transmission requirements.

[0004] It can be seen that there is an urgent need for a data-driven resource allocation method. Summary of the Invention

[0005] In view of this, embodiments of the present disclosure provide a data-driven resource allocation method to solve the problems of poor timeliness of data transmission and low spectral energy efficiency in the sky-air-ground network.

[0006] Embodiments of the present disclosure provide a data-driven resource allocation method, including:

[0007] Step 1: Combine the sensing points, drones, and low-earth orbit satellites in the monitoring area to form a space-air-ground network;

[0008] Step 2: Perform three-layer clustering on the monitoring area, divide the monitoring area into multiple clusters, and each cluster is further divided into multiple groups, where each group contains multiple subgroups. The drones serve each group in a time-sharing manner;

[0009] Step 3: Merge and split the clusters and groups according to preset criteria;

[0010] Step 4: Design a three-hop transmission process for sensing point - aggregation point, aggregation point - drone, and drone - low-earth orbit satellite;

[0011] Step 5: Design a ground - to - space transmission model according to the beam allocation and aggregation point rotation mechanism;

[0012] Step 6: Design a space - to - sky transmission model, use non - orthogonal multiple access technology to share the sub - channels of low - earth orbit satellites, and use interference cancellation methods to decode multiple drone signals in the satellite segment;

[0013] Step 7: Construct a global optimization problem based on the ground - to - space transmission model and the space - to - sky transmission model, and split it into a first optimization problem corresponding to the ground - to - space transmission model and a second optimization problem corresponding to the space - to - sky transmission model. Among them, the global optimization problem includes resource allocation, energy consumption control, and data backlog management;

[0014] Step 8: Set the solution scheme for the first optimization problem;

[0015] Step 9: Set the solution scheme for the second optimization problem;

[0016] Step 10: Design the ISS - PPO algorithm to solve the first optimization problem and the second optimization problem;

[0017] Step 11: Design a solution for the global optimization problem, and through the dual - iteration algorithm, combine the solutions of the first optimization problem and the second optimization problem to iteratively adjust the number of sub - channels of the system and the capacity of the energy harvesting board of the drones.

[0018] According to a specific implementation manner of the embodiments of the present disclosure, step 2 specifically includes:

[0019] Step 2.1: First - layer clustering, cluster the sensing points into multiple clusters according to the distance, and one drone is responsible for data transmission within each cluster;

[0020] Step 2.2: Second - layer clustering, divide the sensing points within each cluster into multiple groups, and the number of sensing points in each group does not exceed the coverage range of the drone;

[0021] Step 2.3, Third-layer clustering: Use the azimuth division algorithm to divide each group into subgroups, and set a convergence point for each subgroup to be responsible for data collection and transmission. When the UAV hovers over each subgroup, the convergence points of each subgroup send data to the UAV simultaneously.

[0022] According to a specific implementation manner of the embodiment of the present disclosure, the step 3 specifically includes:

[0023] Step 3.1, For clusters or groups with the number of sensing points less than a preset value, merge them with adjacent clusters or groups, and select the nearest adjacent cluster to complete the merge;

[0024] Step 3.2, For clusters with a scale larger than the threshold, split them according to a preset standard. The splitting standard is based on the spatial distribution of the cluster. If the horizontal range is larger than the vertical range, split it according to the x coordinate; otherwise, split it according to the y coordinate.

[0025] According to a specific implementation manner of the embodiment of the present disclosure, the step 4 specifically includes:

[0026] Step 4.1, First hop: From the sensing point to the convergence point, the sensing point transmits data to the UAV through the convergence point. A convergence point rotation mechanism is introduced between the convergence point and the sensing point, and multiple sensing points take turns acting as the convergence point as candidate nodes;

[0027] Step 4.2, Second hop: From the convergence point to the UAV, the UAV continuously collects data from the ground sensing points during the mission cycle;

[0028] Step 4.3, Third hop: From the UAV to the low-earth orbit satellite, the UAV selects the observed low-earth orbit satellite for data transmission. The low-earth orbit satellite uses the store-and-forward mode to temporarily store the data received from the UAV and finally transmit it to the ground data center. The system adopts non-orthogonal multiple access technology to enable the UAV to share sub-channels for data transmission.

[0029] According to a specific implementation manner of the embodiment of the present disclosure, the step 5 specifically includes:

[0030] The communication between the sensing point and the UAV is in units of clusters. Each UAV receives data from multiple convergence points simultaneously through multiple beam radio frequency chains. Through the beam allocation and convergence point rotation mechanism, the UAV utilizes the spectrum resources during hovering. The system uses the time division multiple access mode for data transmission.

[0031] According to a specific implementation manner of the embodiment of the present disclosure, the step 6 specifically includes:

[0032] The low-earth orbit satellite uses the successive interference cancellation method to decode the signals of multiple drones, sorts them according to the level of channel gain, and first decodes the signals of drones with high channel gain. During the successive interference cancellation process, a partial successive interference cancellation mechanism is used to simulate the impact of the residual power of the interference signal on the system, so as to eliminate the interference of the undecoded signals.

[0033] According to a specific implementation manner of the embodiment of the present disclosure, step 8 specifically includes:

[0034] Based on the multi-agent architecture of deep reinforcement learning, an agent is assigned to each cluster to monitor the sensing points and the states of the drones within the cluster, and the transmission power and scheduling order of the sensing points and the drones are dynamically adjusted through the action decisions of the agents, so that the data collection and transmission in each task cycle of the system reach the optimal.

[0035] According to a specific implementation manner of the embodiment of the present disclosure, step 9 specifically includes:

[0036] Deploy the single-agent deep reinforcement learning model on the low-earth orbit satellite, and obtain the drone states in real time through the single-agent deep reinforcement learning model, and accordingly adjust the drone transmission power and sub-channel allocation.

[0037] According to a specific implementation manner of the embodiment of the present disclosure, step 10 specifically includes:

[0038] Combine the PPO algorithm with the independent action component, the self-attention mechanism and the strict hard constraint to obtain the ISS-PPO algorithm and solve the first optimization problem and the second optimization problem accordingly.

[0039] According to a specific implementation manner of the embodiment of the present disclosure, step 11 specifically includes:

[0040] Step 11.1, sub-channel optimization, combining the solutions of the first optimization problem and the second optimization problem, dynamically adjust the number of sub-channels according to the data backlog of the drones during the task cycle;

[0041] Step 11.2, energy harvesting optimization, combining the solutions of the first optimization problem and the second optimization problem, find the most suitable capacity of the drone energy harvesting board through the bisection method.

[0042] The data-driven resource allocation scheme in the disclosed embodiment includes: step 1, combining the sensing points, drones and low-orbit satellites in the monitoring area to form an air-ground-space network; step 2, performing three-layer clustering on the monitoring area, dividing the monitoring area into multiple clusters, each cluster is further subdivided into multiple groups, wherein the group contains multiple subgroups, and each group is served by drones in a time-sharing manner; step 3, merging and splitting clusters and groups according to preset standards; step 4, designing a three-hop transmission process of sensing point-aggregation point, aggregation point-drone, and drone-low-orbit satellite; step 5, designing a ground-to-space transmission model according to beam allocation and aggregation point rotation mechanism; step 6, designing a sky-to-air transmission model, using non-orthogonal multiple access technology to share low-orbit satellite sub-channels, and using interference elimination in the satellite segment. In addition, the method is used to decode multiple drone signals; step 7, constructing a global optimization problem according to the ground-to-sky transmission model and the sky-to-air transmission model, and splitting it into a first optimization problem corresponding to the ground-to-sky transmission model and a second optimization problem corresponding to the sky-to-air transmission model, wherein the global optimization problem includes resource allocation, energy consumption control and data backlog management; step 8, setting a solution to the first optimization problem; step 9, setting a solution to the second optimization problem; step 10, designing an ISS-PPO algorithm to solve the first optimization problem and the second optimization problem; step 11, designing a solution to the global optimization problem, combining the solutions of the first optimization problem and the second optimization problem through a dual-iteration algorithm, and iteratively adjusting the number of sub-channels of the system and the capacity of the energy collection board of the drone.

[0043] The beneficial effects of the embodiments of the present disclosure are as follows: Through the solution of the present disclosure, it is considered to transmit the data collected by the aggregation point from the ground sensing points to the low-earth orbit satellite in a timely manner by means of an unmanned aerial vehicle (UAV), while maximizing the spectral energy efficiency. The global main problem includes three problems, namely, the data transmission problem from the aggregation node to the UAV, the data transmission problem from the UAV to the satellite, and the capacity of the UAV energy harvesting board and the UAV-satellite channel resource allocation problem of the overall system. We first solve the first two problems, and then use their solutions to solve the third problem, and finally the global solution can be obtained. For the first and second problems, an improved algorithm based on Proximal Policy Optimization (PPO) - ISS-PPO is proposed. This algorithm combines Independent Action Components, self-attention mechanism, and Strict-Hard Constraint. For the first problem, the data transmission volume from the aggregation node to the UAV can be maximized, the power allocation of the aggregation node and the path selection of the UAV can be optimized, and the data backlog at the aggregation node, the transmission energy of the aggregation point, and the movement energy consumption of the UAV can be minimized. For the second problem, the power control and sub-channel allocation of the UAV can be optimized, the data backlog of the UAV and the transmission energy consumption of the UAV can be minimized, and the total data volume received by the low-earth orbit satellite can be maximized. Finally, a dual iteration algorithm (DIA) for solving the third problem is proposed. This algorithm iteratively calls the ISS-PPO algorithm to determine the appropriate capacity of the UAV energy harvesting board and the number of UAV-satellite sub-channels. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 FIG. is a schematic flowchart of a data-driven resource allocation method provided by an embodiment of the present disclosure;

[0046] Figure 2 FIG. is a schematic diagram of a data collection scenario in a three-hop uplink integrated network provided by an embodiment of the present disclosure;

[0047] Figure 3 FIG. is a schematic diagram of a data collection scenario in a three-hop uplink integrated network provided by an embodiment of the present disclosure. Detailed Implementation Modes

[0048] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0049] The following uses specific specific examples to illustrate the implementation modes of the present disclosure. Those skilled in the art can easily understand other advantages and effects of the present disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. The present disclosure can also be implemented or applied through other different specific implementation modes. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present disclosure. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts belong to the scope of protection of the present disclosure.

[0050] It should be noted that the following describes various aspects of the embodiments within the scope of the appended claims. It should be obvious that the aspects described herein can be embodied in a wide variety of forms, and any specific structure and / or function described herein is illustrative only. Based on the present disclosure, those skilled in the art should understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects described herein can be used to implement the device and / or practice the method. In addition, this device and / or this method can be implemented using other structures and / or functions in addition to one or more of the aspects described herein.

[0051] It should also be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present disclosure in a schematic manner. Only the components related to the present disclosure are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The type, quantity, and ratio of each component in its actual implementation can be an arbitrary change, and the component layout type may also be more complex.

[0052] In addition, in the following description, specific details are provided to facilitate a thorough understanding of the examples. However, those skilled in the art will understand that the described aspects can be practiced without these specific details.

[0053] The embodiments of the present disclosure provide a data-driven resource allocation method, and the method can be applied to the communication process of the space-air-ground network in the mobile communication scenario.

[0054] See Figure 1 , which is a flowchart of a data-driven resource allocation method provided by the embodiments of the present disclosure. As Figure 1As shown, the method mainly includes the following steps:

[0055] Step 1: Combine the sensing points, unmanned aerial vehicles (UAVs), and low Earth orbit (LEO) satellites in the monitoring area to form a space-air-ground network.

[0056] During specific implementation, in remote areas such as deserts and gobi, a large number of sensing points are deployed at arbitrary ground positions to monitor the ecological environment. Due to the small number of stationed personnel and limited ground communication infrastructure, it is difficult to establish or maintain a large number of concurrent end-to-end connections. This results in a large amount of ubiquitous data generated on the ground being unable to be transmitted to the data center for processing in a timely manner, thus losing timeliness. However, the deployability and high mobility of UAVs provide the possibility of establishing a communication path for end-to-end data transmission. Since LEO satellites periodically cover the ground area and provide services, it is feasible to establish a complete ground-air-space communication link. This involves how to allocate resources to improve the overall performance of the system. As Figure 2 shown, consider using the Sink-UAV-LEO backbone network to collect the ubiquitous data generated in Internet of Things applications in a timely manner, and then transmit the data to the remote data center through the selected Sink-UAV-LEO path.

[0057] Step 2: Perform three-layer clustering on the monitoring area, divide the monitoring area into multiple clusters, each cluster is further divided into multiple groups, where each group contains multiple subgroups, and the UAV serves each group in a time-sharing manner.

[0058] In specific implementation, due to the extremely wide monitoring range and geographical isolation, a large number of sensing points are distributed in a certain number of non-overlapping areas (area, denoted as M). In each area, the sensing points are clustered according to the distance to form a cluster, and then each cluster is assigned a drone. Since the number of sensing points in a cluster is still large and widely distributed, and the coverage of the drone is limited, the drone cannot cover all sensing points at the same time. Therefore, a cluster is subdivided into multiple groups, the size of each group does not exceed the coverage of the drone, and the members in the group send data to their convergence point. Each drone provides time-sharing services to different groups in the cluster to which it belongs, and it receives the data collected by the convergence point. In order to more effectively utilize the concurrent capabilities of each drone, it is necessary to divide each group into smaller subgroups, each with a convergence point. When the drone hovers over a group, the convergence point of the subsets will send data to the drone at the same time. The above clustering process is called three-layer clustering (Three-Layer Clustering, TLC). Any classical clustering algorithm can be used for the first two levels of clustering, while the Orientation-Based Area Division (OBAD) method we proposed is used for the third level of clustering. OBAD will be introduced later. Before each task cycle, each level of clustering in TLC can be implemented according to the position change of the sensing point, but it cannot be implemented during each task cycle to avoid the interference of clustering transformation on task transmission.

[0059] Step 3, merging and splitting clusters and groups according to preset criteria;

[0060] In the specific implementation, in the first two levels of clustering, some clusters or groups may contain only a few sensing points. Providing services to these areas alone will lead to an imbalance between costs and benefits. To solve the above problem, we merge small clusters / groups into adjacent clusters / groups. Taking cluster as an example, the selection criteria for adjacent clusters to be merged are

[0061]

[0062] In formula (1), C is the cluster set, e represents the e-th cluster, and d e is the distance between the center of the small cluster and the center of the e-th adjacent cluster. Compared with the similar method in [5], this method reduces the time complexity from O(n 2Reduce it to O(n). In addition, some large clusters (or groups) containing a large number of sensing points may be formed after clustering. For the limited UAV resources, it is unreasonable to process large-scale data transmission. Therefore, we divide a large cluster into several small clusters according to the range and center point position of the large cluster. Specifically, if the horizontal range of a cluster is greater than its vertical range, it is divided by the x coordinate: the nodes with x coordinate less than the x coordinate of the center point form one cluster, and the other nodes form another cluster. Otherwise, it is divided by the y coordinate.

[0063] Step 4, design the three-hop transmission process of sensing point - sink, sink - UAV, and UAV - low earth orbit satellite;

[0064] When specifically implemented, the design process of the three-hop transmission process is as follows:

[0065] 1) The first-hop transmission from the sensing point to the sink

[0066] In the first-hop transmission, the sensing points in each subset transmit data to the sink in that subset. We assume that the sensing points have energy harvesting devices that can collect renewable energy such as solar energy and store it in the battery. Therefore, battery limitations and dynamic energy consumption will affect the system performance. The sink, like the sensing points, has limited hardware resources, usually only one radio frequency chain and one omnidirectional antenna. Therefore, when the sink is connected to the UAV, it can only transmit data, and can only collect data at other times.

[0067] To prevent a serious imbalance in energy consumption between the sensing points and the sink, and to ensure that data collection and transmission at the sink can be carried out simultaneously, for each subset, we select multiple sensing points to implement the rotation of the aggregation task. Specifically, we select several sensing points near the central area as candidate nodes in the subset. Then, we use a circular queue to record these candidate nodes, and use a pointer to indicate the current sensing point to collect data as the sink. When the UAV hovers above the group to provide services, the pointer is passed to the next candidate node, which starts data collection after receiving the pointer. At the same time, the previous candidate node pointed to by the pointer transmits the collected data to the UAV. This process is called the "Sink Rotation Mechanism" (SRM).

[0068] 2) The second-hop transmission from the sink to the UAV

[0069] The UAV has the same energy harvesting function as the sensing points, providing energy for its data transmission to the low-earth orbit satellite. We assume that each UAV is equipped with an energy harvesting board. At the same time, the UAV is equipped with a replaceable battery. When the battery runs out, the battery can be replaced and the task can be executed again. To make reasonable use of the limited battery resources and ensure seamless data transmission of the UAV, the energy of the replaceable battery is only used for non-communication operations (such as the flight and hovering of the UAV), while the harvested energy is used for communication operations.

[0070] We set the data collection of the UAV as a periodic task, that is, the task cycle. When a task cycle starts, the UAV receives the instructions for this cycle from the satellite ground station. It takes off from the battery replacement point, then flies to the designated group for hovering and collects data from the convergence point. This path selection and data collection process are repeated continuously until the UAV only has enough power to return to the battery replacement point. We limit the duration of this process to T.

[0071] 3) The third-hop transmission from the UAV to the low-earth orbit satellite

[0072] In the third-hop transmission, we use a super-dense low-earth orbit satellite constellation. Therefore, each UAV can always select a suitable one from the multiple low-earth orbit satellites it observes for data transmission. These low-earth orbit satellites can temporarily store the received data and finally transmit it to the ground data center. The UAV collects and sends data simultaneously during the task cycle, and forwards the unsent data to the satellite at the battery replacement point, and then enters the next task cycle. We divide the time for the UAV to send data to the satellite in each task cycle into several time slots of fixed length (i.e., τ). The UAV adopts the store-and-forward mode, and the data collected in one time slot is transmitted in the next time slot.

[0073] In the air-space communication scenario, Non-Orthogonal Multiple Access (NOMA) is used to improve the spectral efficiency and system capacity. To facilitate the allocation and sharing of frequency bands, we divide the frequency band into multiple sub-channels, and then allocate the sub-channels to each area according to the strategy. The UAVs within the same area use NOMA technology to share the allocated sub-channels. The number of UAVs in area m is denoted as J m . At the same time, we use to represent the set of all UAVs, to represent the set of all clusters, to represent the number of clusters, to represent the group in cluster C m,i .

[0074] Step 5, design a ground-to-space transmission model according to the beam allocation and convergence point rotation mechanism;

[0075] In specific implementation, for simplicity, taking Cluster C m,i as an example, analyze the communication among the sensing points, the convergence points and the UAVs, and assume that the UAV U m,i providing services for this cluster is the i-th UAV in Area m.

[0076] Different from the sensing points, UAVs usually have sufficient hardware resources to support K + 1 non-overlapping beam radio frequency chains. This enables it to receive data from up to K convergence points within the same group-air frequency band simultaneously within one mission cycle, while communicating with low-orbit satellites in the air-space satellite frequency band (such as the Ka band). To make full use of the concurrency ability of UAVs, we propose the OBAD method. As Figure 3 shown, with the group center point as the center, the group is evenly divided into K subsets, denoted as and a convergence point is selected in each subset using the SRM method, so that when the UAV hovers above the group center point, the convergence points in each subset can be connected to the UAV's beam. The convergence points in group G m,i,n are denoted as

[0077]

[0078] Subset The nodes in are denoted as

[0079]

[0080] Each convergence point can select several active sensing points through a fairness algorithm and receive data in the Time Division Multiple Access (TDMA) mode. The convergence point aggregates all the received data into a data packet for transmission. The binary variable indicates whether the sensing point r in is served: 1 if served, otherwise 0.

[0081] Within one mission cycle, the data collected by the convergence points in is denoted as Its calculation formula is

[0082]

[0083] In formulas (2) and (3), represents the power spectral density of additive white Gaussian noise (AWGN), WRS is the bandwidth between the sensing point and the aggregation point, is the transmission power from the sensing point r to the aggregation point and is the channel gain from the sensing point r to the aggregation point . In each task cycle, the total amount of data collected by all aggregation points in cluster C m,i is

[0084]

[0085] Based on the omnidirectional transmission characteristics of each aggregation point, its transmission gain and transmission interference gain are approximated as 1. Considering the directional reception characteristics of each UAV, the gain of the k-th receiving beam of UAV U m,i to the aggregation point is

[0086]

[0087] In formula (5), is the width of the k-th beam of UAV U m,i , is the angle between the center line of the k-th receiving beam of UAV U m,i and the connection line from the sending end to the receiving end U m,i . In addition, 0 ≤ ∈ < 1 is the sidelobe gain, where ∈ << 1 for a narrow beam. Accordingly, the transmission rate from the aggregation point to UAV U m,i is

[0088]

[0089] In formula (6), W SU is the bandwidth between the aggregation point and the UAV, is the transmission power from the aggregation point to UAV U m,i , is the interference channel gain of the aggregation point to the UAV on the link, is the interference gain generated by m,i on the k-th beam of UAV U .

[0090] Assign numbers to all groups within the same cluster and use the positive integer sequence (L m,i,1 , …, L m,i,x , …, L m,i,X ) to represent the group service order of UAV U m,i . In a task cycle, the total data collected by UAV U m,i is and the mobile energy consumption E m,i The energy consumption can be expressed as

[0091]

[0092] In formulas (7) and (8), is the sink node The transmission time of is the hovering time of the UAV at p h and p f are the hovering power and flight power of the UAV respectively,

[0093] is and The distance between the center points, G m,i,0 is the position of the battery replacement point, v f is the flight speed of the UAV. and t m,i,n The calculation formulas of are

[0094]

[0095] In formulas (9) and (10), and are the remaining battery energy and the collected energy of the sink node within the mission cycle t, is The transmission power of is from to the UAV U m,i The transmission rate of. The equation shows that significant differences in the amount of data collected between different sink nodes within the same subset will lead to low data collection efficiency of the UAV. When some sink nodes have more data to transmit while some have already transmitted their data, it will result in waste of hovering energy and some idle receiving beams not being utilized. To solve this problem, according to the characteristics of the normal distribution of the data volume of the sink nodes, the Data Symmetric Balance (DSB) technology is proposed. Specifically, within the same group, the sink node with the largest data volume sends data to the sink node with the smallest data volume; the sink node with the second largest data volume sends data to the sink node with the second smallest data volume, and so on. DSB is executed every once in a while, and the length of the time interval depends on the size of the group: the larger the group, the shorter the time interval. Eventually, the data volume becomes more evenly distributed among the collection nodes.

[0096] Step 6, design a sky-ground transmission model, use non-orthogonal multiple access technology to share the sub-channels of low-earth orbit satellites, and adopt interference cancellation methods at the satellite segment to decode multiple UAV signals;

[0097] For example, assume that multiple low-earth orbit satellites in one orbit take turns to provide uninterrupted services for UAVs, and each UAV can only access the nearest low-earth orbit satellite at the same time. Compared with the beam coverage area of a low-earth orbit satellite with a coverage radius of 50 kilometers, the ground sensing area (such as a coverage radius of 10 kilometers) studied in the present invention is much smaller. In addition, the arc length between adjacent low-earth orbit satellites sharing the same orbital plane is much larger than the radius of each ground sensing area. Therefore, within a given time slot, all UAVs share the same low-earth orbit satellite.

[0098] The low-earth orbit satellite receiver uses Successive Interference Cancellation (SIC) to decode the signals of multiple users on the channel. The user decoded first will face interference from the signals of the undecoded users in the user set, while the subsequent users will benefit from the reduced interference and thus obtain a higher transmission rate. To mitigate the adverse effects of SIC, the decoding order should match the decreasing channel gain from the user to the low-earth orbit satellite. Therefore, the user with a higher channel gain is decoded first. Assume that the channel gains of UAVs are sorted from large to small as Accordingly, during the process of decoding the UAVU m,j signal, for U m,i , the signals of UAVs where i > j are cancelled, while the signals of UAVs where i < j are regarded as noise.

[0099] In actual situations, due to limitations such as inaccurate channel estimation, poor signal quality, and hardware, decoding errors of interference signals may occur. Therefore, there may be residual interference after SIC, that is, so-called incomplete SIC. The residual interference generated by incomplete SIC can be simulated as a linear function of the interference signal power, and the coefficient of this function can be determined through long-term measurement. Therefore, on the sub-channel m assigned to J m UAVs, the signal-to-noise ratio of the j-th UAV can be expressed as

[0100]

[0101] In formula (11), W B is the bandwidth of each sub-channel between the UAV and the low-earth orbit satellite, and the coefficient represents incomplete SIC, which ranges from [0, 1], and the larger the value, the more serious the interference. When is equal to 1, the interference is not eliminated at all, which is the most unfavorable situation. The instantaneous achievable rate of the UAV is

[0102] R m,i R(t) = W B Dρ m log 2 (1 + γ m,i ) (12)

[0103] Step 7: Construct a global optimization problem based on the ground-sky transmission model and the sky-air transmission model, and split it into a first optimization problem corresponding to the ground-sky transmission model and a second optimization problem corresponding to the sky-air transmission model, where the global optimization problem includes resource allocation, energy consumption control, and data backlog management;

[0104] In specific implementation, considering the multi-hop transmission characteristics of the three-hop network, to optimize the overall performance of the system, we first need to clarify the data transmission and backlog situations of each hop, as well as the related energy consumption.

[0105] For all aggregation points in cluster C m,i The data that may be backlogged within one task cycle can be expressed by the following formula

[0106]

[0107] ClusterC m,i The energy consumption of groud-air data transmission in ClusterC can be expressed as

[0108]

[0109] For the unmanned aerial vehicle U m,i The amount of data that may be backlogged and the energy consumption of data transmission within one task cycle can be expressed as

[0110]

[0111]

[0112] In formula (16), T′ represents the actual time for the unmanned aerial vehicle U m,i to collect data in each task cycle. The amount of data forwarded by the unmanned aerial vehicle U m,i to the low-earth orbit satellite within one task cycle is expressed as

[0113]

[0114] To evaluate the fairness of data transmission, avoid the occurrence of data congestion at any bottleneck point, and achieve fair and reasonable resource allocation, we adopt the Jain’s Fairness index, and its calculation formula is

[0115]

[0116] In formulas (18) and (19), P m,i and P' are the Jain’s Fairness values of the UAV U m,i and the LEO satellite, respectively.

[0117] The main objective of this invention is to ensure the timeliness of data collection, minimize energy consumption, and improve spectrum utilization. To achieve this goal, we must minimize data backlog and resource waste at bottleneck points (such as aggregation points and UAVs) on the uplink transmission path, while maximizing the total amount of data uploaded to the LEO satellite. Therefore, we propose the following global optimization problem.

[0118]

[0119] In equation (20), β is the energy consumption coefficient of the system environment, is the maximum battery capacity of the UAV, and are conservative empirical values. The constraint condition C1 indicates that the path selection of the UAV in each mission cycle is restricted by its maximum energy storage. We assume that the battery is full after the UAV replaces the battery at the battery replacement point each time. The constraint conditions C2 and C3 ensure the service fairness of the UAV and the LEO satellite.

[0120] Since this is a mixed-integer non-linear programming problem, it is difficult to solve. In addition, since resources and demands will be dynamically adjusted over time, it will bring great difficulties to integrate the above problems into a single optimization problem. Therefore, we further consider two local optimization problems according to the multi-hop characteristics of the system. The formula for the first local optimization problem is

[0121]

[0122] In equation (21), β' is the penalty weight for the energy consumption of the aggregation point, P max is the maximum achievable transmission power between the sensing point and the aggregation point, and are the upper limits of the battery capacity and the collected energy of the sensing point and the aggregation point in each mission cycle, respectively, and are The remaining battery energy and harvested energy of the sensing point r in it. The constraint C4 restricts all UAVs to complete their mission cycles within a specified time to ensure system synchronization and avoid time waste. The constraint C5 specifies the initial and final positions of the UAVs, which is of practical significance in the periodic data collection scenario

[12] . The constraints C6 and C7 are the minimum and maximum transmit power constraints of the sensing point and the sink point respectively. The constraints C8 and C9 indicate that the energy consumption of the sensing point and the sink point cannot exceed the current battery stock. The constraints C10 and C11 ensure that the battery stock and the harvested energy do not exceed the maximum limit.

[0123] Problem The objective is to optimize the data transmission in each mission cycle in the ground-air scenario. Specifically, its purpose is to maximize the data transmission volume from the sensing point to the UAV, while minimizing the data backlog and transmission energy consumption at the sink point, as well as the motion energy consumption of the UAV. However, the present invention also focuses on the communication in the air-space scenario. It involves the transmit power control of the UAV and the sub-channel allocation in each area under the NOMA technology. For this purpose, the present invention proposes a second local optimization problem, that is

[0124]

[0125] In Equation (22), is the energy consumption coefficient of the UAV, is the maximum achievable transmit power of the UAV, and are the upper limit of the capacity of the UAV's energy harvesting board and the amount of energy harvested by the UAV in each time slot respectively, and are the remaining energy and harvested energy of the UAV U m,i respectively. The constraint C12 ensures that the total allocated bandwidth of all areas does not exceed the available bandwidth between the UAV and the low-earth orbit satellite. The constraints C13 and C14 are the minimum and maximum transmit power constraints of the UAV respectively. The constraint C15 indicates that the amount of harvested energy used by the UAV cannot exceed the current harvested energy storage. The constraints C16 and C17 ensure that the harvested energy and the harvested energy storage are not infinite.

[0126] Focuses on the data transmission from the UAV to the low-earth orbit satellite, aiming to maximize the data transmission volume while minimizing the data backlog and transmission energy consumption of the UAV. In optimizing and On the basis of, Only need to focus on the size of the low-earth orbit satellite bandwidth and the capacity of the UAV's energy harvesting board. Therefore, we first adopt the DRL model to solve and Then we combine their solutions to solve for

[0127] Step 8, set the solution scheme for the first optimization problem;

[0128] In specific implementation, we use a multi-agent DRL model with a distributed training and distributed execution framework to solve the problem At the satellite ground station in each area, each cluster is assigned an agent to make decisions on data transmission. Since each cluster is disjoint and a single drone serves only one cluster, we take one cluster as an example. Each agent participates in the dynamic environment by performing actions, making observations, and receiving rewards. The agent has the ability to monitor data transmission and formulate strategies for the actions of the sink and drones. In the ground-air scenario, each drone regularly combines its features with the sink information and works with the agent to monitor the fluctuation information within the cluster. In each task cycle, the agent assigns actions to the drones and sinks in the corresponding cluster. After receiving the agent's instructions, the drones fly to the designated group to provide corresponding services. All sinks in the group will adjust their transmission power according to the instructions. After each action is executed, the agent receives an immediate reward indicating its contribution to the goal, and the cluster system will enter the next state accordingly. Throughout the task cycle, the agent continuously monitors the environmental changes and updates the system state accordingly.

[0129] Given the mixed-integer non-linear programming nature of it is necessary to reconstruct it into a discrete Markov decision process, represented by the tuple m,i For this purpose, we demarcate the state, action, and reward in the tuple for each task cycle. Since the decisions of each group are independent and have the same data transmission pattern, we take clusterC

[0130] 1) State space

[0131]

[0132] In formula (23), the agent obtains the remaining energy of the sink the harvested energy of the coordinates of the center point of groupG groupG m,i,n the transmission task volume the remaining power of drone U drone U m,i the remaining battery power of U m,i the coordinates of and the convergence point-UAV channel coefficient It is calculated assuming that the UAV hovers above the center point of the group. Based on the above information, the agent can select actions for the system.

[0133] 2) Action space

[0134]

[0135] The actions include scheduling the sequence number L m,i,n (each group has a number) and the power of the convergence point of the corresponding group with number L m,i,n The transmit power of the convergence point According to Adjust.

[0136] 3) Immediate reward

[0137]

[0138] The immediate reward in formula (25) can ensure timely data collection and minimize the power consumption of each task cycle.

[0139] Step 9, set the solution scheme for the second optimization problem;

[0140] In specific implementation, to avoid UAVs competing for sub-channels and improve the model training and decision-making efficiency, we use a single-agent DRL model to solve the problem We deploy the agent on a low-earth orbit satellite, and all satellites share its model parameters. During the model training phase, after a satellite finishes its service, it will pass the model parameters of the agent to the next satellite. During the execution phase, since the model parameters of the agent are no longer updated, the satellites stop transmitting the parameters. Specifically, all UAVs transmit their states to the agent on the current satellite, the agent processes them and sends the decisions to each UAV. The UAVs adjust their transmit powers accordingly and send data to the low-earth orbit satellite through the assigned sub-channels.

[0141] 1) State space

[0142]

[0143] In formula (26), m, i represent a specific area m and the i-th UAV in that area. In each time slot, the integer represents the area to which UAV U m,i belongs,[[]] represents the channel coefficient between UAV U m,i and the nearest low-earth orbit satellite,[[]]​ is the transmission task volume of U m,i . and are the remaining collection energy and the collected energy of the UAV U m,i respectively. According to the above environmental state, the agent can make a decision.

[0144] 2) Action space

[0145]

[0146] The actions include the sub-channel allocation ratio ρ of the aream m and the transmission power P m,i of the UAV U m,i . After the agent takes an action, an immediate reward feedback can be obtained.

[0147] 3) Immediate reward

[0148]

[0149] In formula (28), is the amount of data forwarded by the UAV U m,i to the LEO satellite within time slot t, is the data backlog of the UAV U m,i . The purpose of this immediate reward is to ensure the timeliness of data transmission from the UAV to the satellite by maximizing the amount of collected data and minimizing the data backlog in each time slot. In addition, it also minimizes the transmission energy consumption of the UAV.

[0150] Step 10, design the ISS-PPO algorithm to solve the first optimization problem and the second optimization problem;

[0151] In specific implementation, due to the intertwined three-hop network structure proposed, the optimization problem is very complex and cannot be directly solved by ordinary DRL algorithms and Therefore, we improved the PPO algorithm by combining it with an independent action component, a self-attention mechanism, and strict hard constraints to obtain an improved version of the PPO algorithm (i.e., ISS-PPO). By using ISS-PPO, we can efficiently solve or Due to the low coupling and scalability of this algorithm, it is also applicable to other similar problems.

[0152] 1) Independent action component

[0153] Due to the existence of multiple decisions in the above-mentioned action space, and each decision contains numerous actions, the combination of all actions will form a high-dimensional discrete action space. However, it is difficult for ordinary neural networks to handle such an action space, so we introduce independent action components. The independent action components divide the complex action space A into independent and combinable sub-action spaces, which can be expressed as A = A 1 ×…A x ×…×A X , and A x is the x-th independent sub-action space. The policy π of the agent can be expressed as the combination of independent sub-policies π x :

[0154]

[0155] In formula (29), s is the current state, a x ∈A x . The agent combines individual actions to maximize the total expected return of , which is expressed as

[0156]

[0157] In formula (30), γ is the discount factor, is the reward obtained by selecting action t in state s . This method has three major advantages: it reduces the high-dimensional action space to a lower dimension, simplifies policy learning; provides a framework that is easy to expand and adapt to various tasks and environments; and makes the policy modular, facilitating separate debugging and optimization. Its main challenge lies in dealing with the dependencies between actions and optimizing the combined policy.

[0158] 2) Self-attention mechanism

[0159] Due to the large dimension of the above-mentioned state space, different degrees of dependence between state sequences, and different degrees of influence of various states on action decisions in different scenarios, it is difficult for the network model to capture global dependencies and long-sequence dependencies, the feature representation and generalization ability are reduced, and the computational efficiency will also be affected when processing the states fed back by the environment. These limitations affect the decision-making accuracy and adaptability of the network model. The self-attention mechanism is an effective tool to solve the above problems.

[0160] The self-attention mechanism assigns different attention weights according to the relationships within the sequence, thereby processing sequence data. It enhances the DRL model by capturing global dependencies, dynamically weighting states, and feature representation. The key steps for calculating the attention scores can be expressed as

[0161]

[0162] In formula (31), A is the attention weight matrix, Q = SW Q is the query matrix, K = SW K is the key matrix, S is the input matrix, and are learnable weight matrices, d model and d k are the dimensions of the input features and the query vector, is the scaling factor to prevent the dot product value from being too large. After obtaining the attention weight matrix A, it is weighted and summed with the feature representation of the input state to generate richer and more expressive features.

[0163] 3) Strict hard constraints

[0164] There are many constraints in our system. For example, only one beam of a drone can access one aggregation point, and the sum of the bandwidth allocation ratios of each area is not greater than 1. Therefore, it is necessary to constrain the actions output by the neural network to avoid violating these constraints. The two most common methods are soft constraints and hard constraints. Soft constraints are usually used to handle invalid operations. When the output of the network contains invalid operations, it provides low or even negative rewards to the network as a warning to reduce the output of invalid operations in this state. Although soft constraints are simple and easy to implement, they do not necessarily prevent the occurrence of invalid operations.

[0165] Hard masking is a method that directly prohibits impossible or unavailable operations. It masks the logits of invalid operations before the network output to ensure that the output layer avoids these operations. Hard constraints can completely avoid invalid operations, but it is necessary to know in advance which operations are invalid in specific situations. They cannot prevent operations that are determined to be invalid only after the output, and are not suitable for our scheme. Therefore, based on the serialization characteristics of the high-dimensional discrete action space, we propose a new constraint method - strict hard constraints. First, we randomly rearrange the logits sequence to ensure fairness. Then, we set the logits of actions that violate the constraints to extremely low values and set their gradients to zero. This can prevent the softmax function from selecting these actions. Finally, we use the modified logits and the softmax function to determine the actions. The description is as follows

[0166] a x = softmax(l x,1 , …, l x,j , …, l x,J )(32)

[0167]

[0168] In formulas (32), (33), and (34), ax Denote the decision of the \(x\)-th dimension as \(l\). x,j Denote the \(j\)-th action on the \(x\)-th dimension as is a large negative number (e.g., ), \(f\) x Denote the number of constraints related to the decision of the \(x\)-th dimension as is a sequence indicating whether the constraints of the \(x\)-th dimension are violated, where the 0 element indicates no violation, otherwise it indicates violation. By using strict hard constraints, we can impose constraints on actions of any dimension at any stage of the neural network execution to mask invalid actions. The specific ISS-PPO algorithm is described as follows, where

[0169] 4) ISS-PPO algorithm

[0170] Run in any satellite ground station or low-earth orbit satellite

[0171] Input: Initial policy network parameters \(\theta\), initial value network parameters \(\varphi\), initial weight matrix \(W\) Q and \(W\) K , number of training iterations \(M\) (5000 rounds), scaling factor (0.088)

[0172] Output: Value function \(V\) φ , policy function \(\pi\) θ

[0173] Step 1: Decompose the action space into independent subspaces

[0174] Step 2: Determine whether the training round indicator variable \(t\) (initial value is 0) reaches the training time step \(M\) of the network model. If so, end this training process; otherwise, go to Step 2.

[0175] Step 3: Clear the training buffer

[0176] Step 4: Determine whether the current data collection step has reached the upper limit (128000 steps). If so, go to Step 9; otherwise, go to Step 4.

[0177] Step 5: Observe the environmental state \(S\) t ;

[0178] Step 6: Calculate the query matrix \(Q = SW\) Q and the key matrix \(K = SW\) K ;

[0179] Step 7: Substitute \(Q\), \(K\) and into formula (31) to obtain the weight matrix \(A\).

[0180] Step 8: Calculate \(S\)t A obtains

[0181] Step 9: Determine whether the current action number indication variable x (initial value is 0) is equal to the action dimension. If so, end this loop; otherwise, go to Step 10;

[0182] Step 10: Mask invalid actions through formulas (32), (33), and (34);

[0183] Step 11: According to Select action a x ;

[0184] Step 12: Increment the action number indication variable x by 1 and go to Step 9

[0185] Step 13: Combine a 1 , …, a x , …, a X to obtain

[0186] Step 14: Execute the action and obtain the reward

[0187] Step 15: The environment transitions to the next state S t+1 ;

[0188] Step 16: Increment the data collection step count by 1 and go to Step 4;

[0189] Step 17: Calculate the advantage estimate through formula (38) and add the experience to the training buffer ;

[0190] Step 18: Determine whether the current training step count indication variable b (initial value is 0) has reached the upper limit (1000 steps). If so, go to Step 26; otherwise, go to Step 19;

[0191] Step 19: Recalculate the advantage estimate through formula (38)

[0192] Step 20: Divide the training buffer into K (256) small batches;

[0193] Step 21: Determine whether the current batch indication variable k (initial value is 0) has reached K. If so, go to Step 25; otherwise, go to Step 22;

[0194] Step 22: Calculate the total PPO loss through formula (30);

[0195] Step 23: Update the actor network, critic network, and weight matrix W according to the total ISS-PPO loss Q and W K ;

[0196] Step 24: Increment the batch indicator variable k by 1, and go to Step 21;

[0197] Step 25: Increment the training step indicator variable b by 1, and go to Step 18;

[0198] Step 26: Increment the training epoch indicator variable t by 1, and go to Step 2;

[0199] Step 27: Output the network parameters V φ , π θ .

[0200] Step 11, design a solution to the global optimization problem, and iteratively adjust the number of sub-channels of the system and the capacity of the UAV's energy harvesting board by combining the solutions of the first optimization problem and the second optimization problem through a dual-iteration algorithm.

[0201] Specifically, to solve the problem, we propose a dual-iteration algorithm. The main purpose of this algorithm is to find the optimal number of sub-channels and the capacity of the UAV's energy harvesting board according to the and solutions. The main purpose of lines 1-8 is to establish a three-hop network and train the ISS-PPO model set to solve The purpose of lines 9-17 is to optimize the number of sub-channels. Another ISS-PPO model is iteratively trained to solve and obtain environmental feedback. On this basis, the number of sub-channels is optimized while meeting the transmission efficiency requirements. The purpose of lines 18-28 is to optimize the capacity of the UAV's energy harvesting board, and using the idea of the dichotomy method, quickly find the most suitable capacity of the UAV's energy harvesting board under the current number of sub-channels.

[0202] 1) Dual-iteration algorithm

[0203] Operating on a low-earth orbit satellite

[0204] Input: Data backlog threshold UAV battery capacity range threshold φ

[0205] Output: Number of sub-channels D, UAV energy harvesting board capacity

[0206] Step 1: Collect the position information of ground sensing nodes;

[0207] Step 2: Call TLC, CDP, OBAD to construct a three-hop network;

[0208] Step 3: Assign a drone to each cluster and establish a set of drones U;

[0209] Step 4: Find a LEO satellite that currently covers itself and mark its successor as the first LEO satellite on the same orbital plane;

[0210] Step 5: Construct a set of LEO satellites (22) and assign an ISS-PPO model M to them 2 ;

[0211] Step 6: Assign an ISS-PPO model to each cluster and obtain a set of models

[0212] Step 18: Train the set of models to solve the problem

[0213] Step 8: Assign an empirical value to D (e.g., 20), and assign a relatively large value to (e.g., 30J)

[0214] Step 9: Use the set of models generated to train model M 2 , so as to solve the problem

[0215] Step 10: Invoke model M 2 to obtain and

[0216] Step 11: Judge whether (e.g., 1Mb) and BD < 0 hold simultaneously. If so, go to Step 12; otherwise, go to Step 15;

[0217] Step 12: Decrease the number of sub-channels with a preset step size;

[0218] Step 13: Judge whether and BD > 0 hold simultaneously. If so, go to Step 14; otherwise, go to Step 15;

[0219] Step 14: Increase the number of sub-channels with a preset step size;

[0220] Step 15: Judge whether or BD = 0. If it holds, go to Step 16; otherwise, go to Step 9;

[0221] Step 16: Set the variable

[0222] Step 17: Determine whether high - low ≥ φ and BD ≠ 0 hold simultaneously. If so, go to Step 11; otherwise, go to Step 18;

[0223] Step 18: Calculate

[0224] Step 19: Use the model set generated to train model M 2 to solve the problem

[0225] Step 20: Invoke model M 2 to obtain and

[0226] Step 21: Determine whether (such as 1Mb) and BD ≠ 0 hold simultaneously. If so, go to Step 22; otherwise, go to Step 23;

[0227] Step 22: Set

[0228] Step 23: Determine whether (such as 1Mb) and BD ≠ 0 hold simultaneously. If so, go to Step 24; otherwise, go to Step 17;

[0229] Step 24: Set

[0230] Step 25: Output parameters D,

[0231] The data-driven resource allocation method provided in this embodiment proposes a Sink-UAV-LEO backbone network that ensures spectral energy efficiency. This network consists of methods such as the proposed three-layer clustering (TLC), combination and division process (CDP), orientation-based area division (OBAD), data symmetry balance (DSB), sink rotation mechanism (SRM), etc. In remote areas with a wide range and multiple regions, reliable and spectrally efficient real-time data transmission can be achieved through this backbone network. The joint power control of the sink and UAV path selection problems in the Sink-UAV layer, as well as the joint power control of the UAV and sub-channel allocation problems in the UAV-LEO layer, are described. Then, we combine the PPO algorithm with independent action components, self-attention mechanism, and strict hard constraints to design an improved PPO algorithm (ISS-PPO) that can solve the above two local problems. The global problem of the backbone network is described, and a dual-iteration algorithm (DIA) is proposed to solve this problem. This algorithm aims to match the low-earth orbit satellite bandwidth size and the UAV energy harvesting board capacity with the data collection task load. In DIA, in addition to completing network construction, two solutions based on ISS-PPO need to be repeatedly called to obtain an approximate solution to the global problem and ensure data transmission timeliness and spectral energy efficiency.

[0232] The method will be further described below with a specific embodiment. The ground area we studied is set to 10km×10km, which contains 6 non-overlapping areas, and the radius of each area is between 1km and 2km, with 100 - 300 sensing points randomly distributed. It is assumed that the sensing points and the sink can collect at most 80% of their maximum energy storage capacity (such as 50J) for each task cycle, and the UAV can collect at most 80% of its maximum energy harvesting board capacity (such as 10J) for each time slot. 22 low-earth orbit satellites take turns to cover this area to provide seamless access. Other parameters are listed in Table 1.

[0233] Table 1

[0234]

[0235]

[0236] The present invention implements the above solution using python + pytorch tools. By dividing the load of the data collection task into 7 levels from low to high, namely A, B, C, D, E, F, and G, the capabilities of 5 integration schemes to handle variable loads are compared. The 5 integration schemes are as follows:

[0237] (1) SE-PSPC-NOPC: SE-PSPC (Sequential Path Selection with Power Control) means making decisions on the selection of the next group served by the UAV and the transmission power of the convergence point in that group using the ISS-PPO algorithm. NOPC (Non-Orthogonal Multiple Access with Power Control) means ISS-PPO makes decisions on the proportion of sub-bands used by each area to access the LEO satellite network and the transmission power of the UAV. UAVs in the same area share sub-channels using NOMA.

[0238] (2) SE-PSPC-FDPC: FDPC (Frequency Division Multiple Access with Power Control) means making decisions on the sub-channel occupancy ratio and power of each UAV using the ISS-PPO algorithm, and UAVs use FDMA (Frequency Division Multiple Access) on their own sub-channels.

[0239] (3) SE-PSPC-TDPC: TDPC (Time Division Multiple Access with Power Control) divides each time slot into corresponding equal small time slots according to the number of UAVs. Each UAV occupies one small time slot for data transmission and uses all sub-channels during this period. The power of the UAV is decided by the ISS-PPO algorithm.

[0240] (4) SE-PSPC-FTPC: FTPC (Frequency and Time Division Multiple Access with Power Control) means making decisions on the sub-channel occupancy ratio of each area and the transmission power of the UAV using the ISS-PPO algorithm. UAVs in different areas use FDMA, while UAVs in the same area use TDMA.

[0241] (5) SE-PSPC-TDFP: TDFP (Time Division Multiple Access with Fixed Power) The only difference between this scheme and SE-PSPC-TDPC is that the UAV uses a fixed power. It is mainly used to verify the effectiveness of power control.

[0242] We divide the load of the data collection task into seven levels (from low to high, from A to G) to compare the amount of resources required by five integration schemes under different loads. For simplicity, we ensure that within each task cycle, the total amount of data transmitted from the sink point to the UAV continuously decreases from level G to level A. The load at level G is based on the transmission capacity of the sink point to the UAV under the most ideal ground communication conditions and the maximum transmission power. When the average absolute value of the UAV data backlog percentage is lower than 50%, 30%, and 10% of the total amount of data received in each time slot, we will evaluate the number of sub-channels and the capacity of the UAV energy harvesting board at different load levels.

[0243] First of all, as the load increases, the demand for sub-channels and the capacity of the UAV energy harvesting board for each scheme also increases accordingly, which is in line with the actual situation. Since the UAVs share sub-channels, SE-PSPC-NOPC requires the fewest sub-channels. However, since the data transmission on each shared sub-channel consumes energy, its requirement for the capacity of the UAV energy harvesting board is the highest. Compared with other schemes, the SE-PSPC-NOPC and SE-PSPC-FTPC schemes require more capacity of the UAV energy harvesting board to achieve better transmission with fewer channels. However, when the number of sub-channels is large, the demand for the capacity of the UAV energy harvesting board for both of them is relatively small. Therefore, as the load level or the UAV data backlog percentage decreases, the difference in the capacity of the UAV energy harvesting board required by SE-PSPC-NOPC and SE-PSPC-FTPC compared with the other three schemes will decrease, and even sometimes SE-PSPC-FTPC will be less than SE-PSPC-TDPC and SE-PSPC-TDFP. Generally speaking, the other three schemes have poor flexibility in adjusting resource requirements under different scenarios.

[0244] In addition, it can also be observed that under the same load level but different UAV data backlog ratios, the number of sub-channels required by each scheme will increase exponentially as the backlog ratio decreases. This is because when the load remains unchanged, the data transmission capacity is not proportional to the number of sub-channels. Due to the limitations of the sub-channel allocation scheme, sufficient bandwidth resources are required to ensure that the data backlog of each UAV is small enough to reduce the overall data backlog rate of the system. However, as the backlog ratio decreases, the growth rate of the bandwidth requirement of the SE-PSPC-NOPC scheme is much smaller than that of other schemes, highlighting its superiority in sub-channel allocation. From the above analysis, it can be inferred that the energy capacity determines the lower limit of the system transmission performance, while the bandwidth resources determine the upper limit of the system transmission performance.

[0245] It should be understood that various parts of the present disclosure can be implemented by hardware, software, firmware, or a combination thereof.

[0246] As described above, it is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present disclosure should be covered within the protection scope of the present disclosure. Therefore, the protection scope of the present disclosure shall be subject to the protection scope of the claims.

Claims

1. A data-driven resource allocation method, characterized in that: include: Step 1: Combine the sensing points, drones and low-orbit satellites in the monitoring area to form an air-ground network; Step 2: Perform three-layer clustering on the monitoring area, divide the monitoring area into multiple clusters, and each cluster is further divided into multiple groups, where a group contains multiple subgroups, and drones are used to provide time-sharing services to each group; Step 3, merging and splitting clusters and groups according to preset criteria; Step 4: Design a three-hop transmission process from sensing point to aggregation point, aggregation point to drone, and drone to low-orbit satellite. Step 5: Design the ground-to-space transmission model based on the beam allocation and convergence point rotation mechanism; Step 6: Design a sky-to-space transmission model, use non-orthogonal multiple access technology, share low-orbit satellite subchannels, and use interference cancellation methods to decode multiple drone signals in the satellite segment; Step 7, constructing a global optimization problem according to the ground-sky transmission model and the sky-space transmission model, and splitting it into a first optimization problem corresponding to the ground-sky transmission model and a second optimization problem corresponding to the sky-space transmission model, wherein the global optimization problem includes resource allocation, energy consumption control, and data backlog management; Step 8, setting a solution to the first optimization problem; Step 9, setting a solution to the second optimization problem; Step 10, designing an ISS-PPO algorithm to solve the first optimization problem and the second optimization problem; Step 11, design a solution to the global optimization problem, combine the solutions of the first optimization problem and the second optimization problem through a dual-iteration algorithm, and iteratively adjust the number of sub-channels of the system and the capacity of the energy harvesting board of the drone.

2. The method according to claim 1, characterized in that The step 2 specifically includes: Step 2.1, the first layer of clustering, clustering the sensing points into multiple clusters according to the distance, and one drone in each cluster is responsible for data transmission; Step 2.2, the second layer of clustering, divides the sensing points in each cluster into multiple groups, and the number of sensing points in each group does not exceed the coverage of the drone; Step 2.3, the third-level clustering, uses the azimuth partitioning algorithm to divide each group into subgroups. Each subgroup sets a convergence point responsible for data collection and transmission. When the drone hovers over each subgroup, each subgroup convergence point sends data to the drone at the same time.

3. The method according to claim 2, characterized in that The step 3 specifically includes: Step 3.1: For clusters or groups containing perception points less than a preset value, merge them with adjacent clusters or groups, and select the nearest adjacent cluster to complete the merger; Step 3.2: For clusters whose size is larger than the threshold, they are split according to the preset standard. The splitting standard is based on the spatial distribution of the cluster. If the horizontal range is larger than the vertical range, it is split according to the x-coordinate, otherwise it is split according to the y-coordinate.

4. The method according to claim 3, characterized in that The step 4 specifically includes: Step 4.1, first hop: from the sensing point to the sink point, the sensing point transmits data to the drone through the sink point, and a sink point rotation mechanism is introduced between the sink point and the sensing point. Multiple sensing points serve as candidate nodes and take turns to act as sink points; Step 4.2, second hop: from the sink to the drone, the drone continuously collects data from the ground sensing point during the mission cycle; Step 4.3, the third hop: from the UAV to the low-orbit satellite. The UAV selects the observed low-orbit satellite for data transmission. The low-orbit satellite uses the store-and-forward mode to temporarily store the data received from the UAV and finally transmits it to the ground data center. The system uses non-orthogonal multiple access technology to enable the UAV to share sub-channels for data transmission.

5. The method according to claim 4, characterized in that The step 5 specifically includes: The communication between the sensing points and the drones is based on clusters. Each drone receives data from multiple aggregation points simultaneously through multiple beam RF chains. Through the beam allocation and aggregation point rotation mechanism, the drone utilizes spectrum resources during hovering, and the system uses time division multiple access mode for data transmission.

6. The method according to claim 5, characterized in that The step 6 specifically includes: Low-orbit satellites use continuous interference cancellation methods to decode signals from multiple drones, and sort them according to channel gain, decoding drone signals with high channel gain first. During the continuous interference cancellation process, partial continuous interference cancellation mechanisms are used to simulate the impact of the residual power of interference signals on the system, so as to eliminate the interference of undecoded signals.

7. The method according to claim 6, characterized in that The step 8 specifically includes: Based on a multi-agent architecture of deep reinforcement learning, each cluster is assigned an agent to monitor the status of the sensing points and drones within the cluster. The transmission power and scheduling order of the sensing points and drones are dynamically adjusted through the action decisions of the agent to optimize the data collection and transmission of the system in each mission cycle.

8. The method according to claim 7, characterized in that The step 9 specifically includes: The single-agent deep reinforcement learning model is deployed on a low-orbit satellite. The UAV status is obtained in real time through the single-agent deep reinforcement learning model, and the UAV transmission power and sub-channel allocation are adjusted accordingly.

9. The method according to claim 8, characterized in that The step 10 specifically includes: The PPO algorithm is combined with independent action components, self-attention mechanism and strict hard constraints to obtain the ISS-PPO algorithm, which is used to solve the first optimization problem and the second optimization problem.

10. The method according to claim 9, characterized in that The step 11 specifically includes: Step 11.1, sub-channel optimization, combining the solutions of the first optimization problem and the second optimization problem, dynamically adjusting the number of sub-channels according to the data backlog of the UAV during the mission cycle; Step 11.2, energy harvesting optimization, combines the solutions of the first optimization problem and the second optimization problem, and finds the most suitable capacity of the drone energy harvesting board through the dichotomy method.

Citation Information

Patent Citations

  • Satellite unmanned aerial vehicle fusion network resource allocation method and device

    CN111669758A

  • Data transmission method of space-air-ground integrated network and related equipment

    CN118381549A

  • Unmanned aerial vehicle network routing method under spectrum denial based on low earth orbit satellite cooperation

    CN119316903A

  • Method of online sales of meal kits to prevent spoilage and enhance flavor of raw meat using raw parsnip

    KR1020250171927A

  • System and method for allocating resources within a communication network

    US20170063445A1