Low-altitude wireless access network resource scheduling method and system
By combining compression ratio optimization, contextual multi-armed gambling machine algorithm and load prediction model in a hierarchical collaboration, and integrating a multi-agent model, the compression ratio is dynamically adjusted to solve the problem that traditional network slicing algorithms cannot adapt to changes in air traffic and channels, thereby improving the network adaptability and service reliability of low-altitude radio access networks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGDONG UNIV OF TECH
- Filing Date
- 2026-02-09
- Publication Date
- 2026-05-01
AI Technical Summary
Traditional network slicing algorithms lack environmental awareness and cannot adapt to rapid changes in air traffic and channel conditions, resulting in low service reliability of low-altitude radio access networks.
The compression ratio optimization-context multi-arm gambling machine algorithm is used to determine the compression ratio set of the current time slot of the UAV. Combined with the load prediction model based on hybrid sampling and attention enhancement adaptive fusion, the predicted node load sequence is output. The set of resource scheduling actions is determined by the multi-agent model to realize dynamic adjustment of the compression ratio to reduce redundant data transmission and network load.
By dynamically adjusting the compression ratio and accurately predicting network load, the network adaptability and service reliability in low-altitude scenarios are improved, redundant data transmission is reduced, network load is lowered, and service latency and load balancing are enhanced.
Smart Images

Figure CN121968345A_ABST
Abstract
Description
A method and system for resource scheduling in low-altitude wireless access networks Technical Field
[0001] This invention relates to the field of communication technology, and in particular to a method and system for scheduling low-altitude wireless access network resources. Background Technology
[0002] The low-altitude economy encompasses a wide range of applications, including drone logistics, aerial inspection, and air taxis. These aerial applications rely on continuous sensing, real-time video streaming, and onboard decision-making capabilities, placing high throughput and low latency demands on wireless communication. With the emergence of embedding artificial intelligence (AI) into Radio Access Network (RAN) architectures, RANs can autonomously manage spectrum, optimize beamforming, and predict the resource needs of aerial users. For example, in drone parcel delivery scenarios, the RAN can dynamically adjust communication parameters based on flight trajectories and service priorities to ensure stable and energy-efficient connections.
[0003] To address the challenges of low-altitude service in AI-enabled RANs, network slicing algorithms have been introduced to dynamically allocate radio, computing, and storage resources based on UAV movement patterns and AI task priorities, becoming a key enabling technology for meeting these needs. However, traditional network slicing algorithms often lack environmental awareness and cannot adapt to rapid changes in air traffic and channel conditions, resulting in low service reliability for low-altitude radio access networks. Summary of the Invention
[0004] This invention provides a method and system for scheduling resources in low-altitude radio access networks, which solves the technical problem that traditional network slicing algorithms often lack environmental awareness and cannot adapt to rapid changes in air traffic and channel conditions, resulting in low service reliability of low-altitude radio access networks.
[0005] The first aspect of this invention provides a method for scheduling low-altitude radio access network resources, comprising:
[0006] The compression ratio set for the current time slot of any UAV is determined by the compression ratio optimization-context multi-arm gambling machine algorithm;
[0007] Based on the historical node load sequence of the low-altitude wireless access network, a load prediction model based on hybrid sampling and attention enhancement adaptive fusion is used to output the predicted node load sequence.
[0008] Based on the compression ratio set and the predicted node load sequence, with the goal of maximizing service latency reward and load balancing reward, a set of resource scheduling actions for the current time slot is determined through a multi-agent model.
[0009] Optionally, determining the set of compression ratios for the current time slot of the UAV through the compression ratio optimization-context multi-arm gambling machine algorithm includes:
[0010] Determine the contextual features of the compression ratio optimization-context multi-armed gambling machine algorithm, including the maximum allowable latency for each data frame, the minimum accuracy requirement for each data frame, and the weighting parameters used to balance compression efficiency and inference accuracy;
[0011] Determine the actions of the compression ratio optimization-context multi-armed gambler algorithm, including the compression ratio set;
[0012] Context-action pairs are constructed using the context features and the action, and the observation reward for the context-action pair is determined.
[0013] The joint posterior distribution of rewards is updated based on the observed rewards, and the mean and variance of the joint posterior distribution of rewards are determined.
[0014] Based on the mean and variance, the set of compression ratios for the current time slot of any UAV is determined by solving for the maximum value of the upper confidence bound acquisition function.
[0015] Optionally, the load prediction model includes a load input projection layer, a load encoder, a load decoder, and a load output projection layer; the historical node load sequence based on the low-altitude radio access network outputs a predicted node load sequence using a load prediction model based on hybrid sampling and attention enhancement adaptive fusion, including:
[0016] The historical node load sequence of the low-altitude wireless access network is projected into load input features through the load input projection layer.
[0017] After encoding and decoding based on the load input features using a load encoder and a load decoder, a load output projection layer is used to process and output the predicted node load sequence.
[0018] Both the load encoder and the load decoder include multiple spatiotemporal attention modules. Each spatiotemporal attention module includes a hybrid sampling submodule, a temporal attention submodule, a spatial attention submodule, and a gated fusion submodule. The processing procedure of the spatiotemporal attention module includes:
[0019] Based on the hybrid sampling submodule, the spatiotemporal input features are sampled at equal intervals along the time dimension from different offsets to determine equally spaced features, and the spatiotemporal input features are sampled in equal segments along the time axis to determine equally segmented features;
[0020] The equal-interval features and equal-segment features are respectively processed by the time attention submodule, and the corresponding outputs are equal-interval time-enhanced features and equal-segment time-enhanced features.
[0021] The equally spaced time enhancement features and equally segmented time enhancement features are input into the gated fusion submodule for adaptive fusion, and the time fusion features are output.
[0022] The spatial attention submodule is used to process the temporal fusion features and output spatial features.
[0023] The temporal and spatial features are fused by the gated fusion submodule to determine the spatiotemporal output features.
[0024] Optionally, the step of determining the set of resource scheduling actions for the current time slot based on each of the compression ratio sets and the predicted node load sequence, with the goal of maximizing service latency reward and load balancing reward, through a multi-agent model, includes:
[0025] Determine the joint observations for the multi-agent model. The joint observations include global observations and local observations. The global observations include the set of node connection relationships of the low-altitude radio access network, the set of maximum available computing resources of the nodes of the low-altitude radio access network, and the node load sequence. The local observations include the service requests of each UAV. The service requests include the number of frames to be transmitted, the uncompressed size of each data frame, the maximum allowable latency of each data frame, the minimum accuracy requirement of each data frame, and the set of compression ratios.
[0026] Determine the action embedding of the multi-agent model, including a set of resource scheduling actions, which includes the association decision variables between centralized nodes and drones, the association decision variables between distributed nodes and drones, the association decision variables between radio frequency nodes and drones, and the allocation decision variables between radio frequency nodes' resource blocks and drones.
[0027] A multi-agent model trained with the goal of maximizing the total reward of service latency reward and load balancing reward is adopted. The model input is the joint observation and historical action embedding determined according to the compression ratio set and the predicted node load sequence. The model outputs the set of resource scheduling actions associated with the current time slot of each UAV in a sequential decision.
[0028] Optionally, the calculation process for the total reward includes:
[0029] ;
[0030] In the formula, Indicates time slot Total reward Indicates time slot Service delay reward Indicates time slot Load balancing rewards Indicates the first A drone in a time slot The average service latency, Represents the logarithm with base e as the natural constant. Indicates the network node index. Represents a centralized node. Represents distributed nodes, Indicates radio frequency node, Represents network nodes The computational overhead incurred Represents network nodes The maximum acceptable computational cost, Indicates the first A drone in a time slot Total service latency Indicates the first A drone in a time slot Number of frames to be transmitted Indicates the data frame index. Indicates the first A drone in a time slot The Middle Inference latency per data frame, Indicates the first A drone in a time slot The maximum allowable latency for each data frame in the process. Indicates the first One drone, Indicates time slot A collection of drones.
[0031] Optionally, the total service latency includes the total transmission latency, the total processing latency, and the total forwarding latency.
[0032] A second aspect of the present invention provides a low-altitude wireless access network resource scheduling system, comprising:
[0033] The compression ratio determination module is used to determine the set of compression ratios for any UAV in the current time slot through the compression ratio optimization-context multi-arm gambling machine algorithm;
[0034] The node load prediction module is used to output the predicted node load sequence based on the historical node load sequence of the low-altitude wireless access network using a load prediction model based on hybrid sampling and attention enhancement adaptive fusion.
[0035] The network resource scheduling module is used to determine the set of resource scheduling actions for the current time slot based on the compression ratio set and the predicted node load sequence, with the goal of maximizing service latency reward and load balancing reward, through a multi-agent model.
[0036] A computer device provided in a third aspect of the present invention includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the steps of the low-altitude wireless access network resource scheduling method as described in any of the preceding claims.
[0037] The fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed, implements the low-altitude wireless access network resource scheduling method as described in any of the preceding claims.
[0038] The fifth aspect of the present invention provides a computer program product, comprising a computer program / instructions, wherein when the computer program / instructions are executed by a processor, they implement the low-altitude wireless access network resource scheduling method as described in any of the preceding claims.
[0039] As can be seen from the above technical solutions, the present invention has the following advantages:
[0040] The above-mentioned solution of the present invention provides a low-altitude wireless access network resource scheduling method, comprising: determining the compression ratio set of any UAV in the current time slot through a compression ratio optimization-context multi-armed gambler algorithm; based on the historical node load sequence of the low-altitude wireless access network, outputting a predicted node load sequence using a load prediction model based on hybrid sampling and attention enhancement adaptive fusion; determining the resource scheduling action set of the current time slot according to each compression ratio set and the predicted node load sequence, with the goal of maximizing service latency reward and load balancing reward, through a multi-agent model; based on the above solution, through the hierarchical collaboration of the compression ratio optimization-context multi-armed gambler algorithm and the load prediction model, the dynamic adjustment of the compression ratio effectively reduces redundant data transmission and lowers the network load, the hybrid sampling attention mechanism achieves accurate network load prediction, and the resource allocation based on the multi-agent model achieves the minimum average service latency and load balancing through a resource scheduling algorithm, thereby improving the network adaptability and service reliability in low-altitude scenarios. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0042] Figure 1 is a flowchart of the steps of a low-altitude wireless access network resource scheduling method provided in Embodiment 1 of the present invention;
[0043] Figure 2 is a schematic diagram of the HDT-RAN framework provided in Embodiment 1 of the present invention;
[0044] Figure 3 is a schematic diagram of the HDT-RAN algorithm provided in Embodiment 1 of the present invention;
[0045] Figure 4 is a schematic diagram of the simulation experiment results of reward and iteration number provided in Embodiment 1 of the present invention;
[0046] Figure 5 is a schematic diagram of the simulation experiment results of compression accuracy and weight provided in Embodiment 1 of the present invention;
[0047] Figure 6 is a schematic diagram of the simulation experiment results of latency and number of tasks provided in Embodiment 1 of the present invention;
[0048] Figure 7 is a structural block diagram of a low-altitude wireless access network resource scheduling system provided in Embodiment 2 of the present invention. Detailed Implementation
[0049] This invention provides a method and system for scheduling low-altitude radio access network resources, which addresses the technical problem that traditional network slicing algorithms often lack environmental awareness and cannot adapt to rapid changes in air traffic and channel conditions, resulting in low service reliability of low-altitude radio access networks.
[0050] To make the objectives, features, and advantages of this invention more apparent and understandable, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described below are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0051] Please refer to Figure 1. Embodiment 1 of the present invention provides a low-altitude radio access network resource scheduling method, which relates to a radio access network (RAN), and integrates the radio access network with a digital twin (Digital Twin). Combining the physical RAN (Internet Protocol) and digital twin (DT) approaches, a hierarchical digital twin-assisted RAN resource scheduling (HDT-RAN) framework, as shown in Figures 2 and 3, is proposed to support low-latency, low-altitude economic applications. By constructing a virtual replica of the physical RAN system, DT enables real-time synchronization, behavior prediction, and cross-domain resource visualization. Integrating the DT module into the slicing process enables hierarchical decision-making, where the physical layer provides real-time sensor data, while the twin layer performs global optimization and predictive control. This hierarchical digital twin-assisted slicing framework can continuously monitor data flow, adjust compression parameters, and balance the load of edge nodes and airborne nodes, thereby significantly improving network adaptability and service reliability. Based on the above HDT-RAN framework, two DT modules are designed: UAV digital twin (UAV-DT) and RAN digital twin (RAN-DT). UAV-DT can simulate UAV status through compression ratio optimization-context multi-arm algorithm to dynamically adjust data compression ratio, while RAN-DT can simulate RAN status through hybrid sampling attention mechanism to accurately predict network node load. This low-latency application uses discrete equal-length time slots, with a time slot index of [insert index here]. , Indicates time slot , This represents the total number of time slots. In this embodiment, it uses... This represents the set of nodes in a low-altitude wireless access network. Represents the set of centralized nodes (CUs). Represents the set of distributed nodes (DUs). Represents a set of radio frequency nodes (RUs); methods include:
[0052] Step 101: Determine the set of compression ratios for any UAV in the current time slot using the compression ratio optimization-context multi-arm gambling machine algorithm.
[0053] In one specific embodiment of this example, step 101 includes the following sub-steps:
[0054] S11. Determine the contextual features of the compression ratio optimization-context multi-armed gambling machine algorithm, including the maximum allowable latency for each data frame, the minimum accuracy requirement for each data frame, and the weight parameters used to balance compression efficiency and inference accuracy.
[0055] S12. Determine the actions of the compression ratio optimization-context multi-armed gambling machine algorithm, including the compression ratio set;
[0056] S13. Construct context-action pairs using context features and actions, and determine the observation reward for the context-action pairs;
[0057] S14. Update the joint posterior distribution of rewards based on observed rewards, and determine the mean and variance of the joint posterior distribution of rewards;
[0058] S15. Based on the mean and variance, determine the set of compression ratios for the current time slot of any UAV by solving for the maximum value of the upper confidence bound acquisition function.
[0059] It should be noted that this implementation models the online derivation process of the compression ratio of each data frame of the UAV as a Contextual Multi-Armed Bandit (MAB) problem, thus forming the compression ratio optimization-contextual multi-armed bandit algorithm:
[0060] The contextual features of the compression ratio optimization-context multi-armed gambling machine algorithm include: , Indicates the first A drone in a time slot Contextual features, Indicates the first A drone in a time slot The maximum allowable latency for each data frame in the process. Represents the normalization function. Indicates the first A drone in a time slot The minimum accuracy requirement for each data frame. Indicates the first A drone in a time slot Weighting parameters used to balance compression efficiency and inference accuracy;
[0061] Defining the action space of the compression ratio optimization-context multi-armed gambling machine algorithm Actions include , Indicates the first A drone in a time slot A set of compression ratios;
[0062] action The reward is from predict, The reward prediction function is represented (this embodiment does not limit the specific reward prediction function), and the observed noisy reward is defined as follows: , Indicates time slot Observation rewards Indicates time slot Context-action pairs Indicates time slot Observation noise, This indicates that the expression follows a pattern with a mean of 0 and a noise variance of . normal distribution This represents the deviation between the actual reward and the predicted reward;
[0063] As a sequential strategy for global optimization, the compression ratio optimization-context multi-armed gambling machine algorithm implemented in this paper learns the reward function online from historical observations. This algorithm models the objective function using a Gaussian process probability model and quantifies its uncertainty through posterior updates; time slots Previous historical observation reward set is , Indicates time slot Observation rewards This indicates transpose, and correspondingly, time slot. The previous set of historical context-action pairs was , Indicates time slot The context-action pairs, and the joint posterior distribution of the historical observation reward are: , Represents the mean function, Indicates time slot The covariance matrix, Represents the kernel function. Representing historical context-action pairs , Representing historical context-action pairs Used to indicate the historical context-action pair between any two historical time slots. Denotes the covariance term. Represents variance. Represents the identity matrix; when the time slot is determined Context-action pair, set time slot Observation reward The joint posterior distribution of rewards is updated using a Gaussian process, which is consistent with... The joint posterior distribution of the rewards is:
[0064] ;
[0065] In the formula, This represents the set of historical context-action pairs in the joint posterior distribution of rewards. The mean, Indicating the time slots in the joint posterior distribution of rewards The mean of the context-action pairs, Indicates time slot The covariance vector of the context-action pair and all historical context-action pairs. Indicates time slot The covariance between the context-action pair and itself;
[0066] The reward joint posterior distribution slot can be derived using the aforementioned reward joint posterior distribution. mean Joint posterior distribution slots of rewards variance For details, please refer to existing technologies, which will not be elaborated here; by solving for the maximum value of the acquisition function based on the upper confidence boundary, the optimal compression ratio set for the current time slot can be determined, that is: , Indicates time slot Used for balancing time slots mean and time slot variance The coefficients representing utilization and exploration rates, respectively, will be used for the current time slot. It is applied to network resource scheduling subsequent timeslots, and after executing this compression ratio set, the agent environment feedback of the compression ratio optimization-context multi-armed gambling machine algorithm provides a real reward. And form new data pairs Incorporating historical data can be used to update the joint posterior distribution of rewards for the next time slot.
[0067] Step 102: Based on the historical node load sequence of the low-altitude wireless access network, output the predicted node load sequence using a load prediction model based on hybrid sampling and attention enhancement adaptive fusion.
[0068] In one specific embodiment of this example, step 102 includes the following sub-steps:
[0069] S21. Project the historical node load sequence of the low-altitude wireless access network into load input features through the load input projection layer.
[0070] S22. After encoding and decoding based on the load input features using a load encoder and load decoder, the load output projection layer is used to process and output the predicted node load sequence.
[0071] In a more specific embodiment of this example, the processing procedure of the spatiotemporal attention module includes:
[0072] S31. Based on the hybrid sampling submodule, the spatiotemporal input features are sampled at equal intervals along the time dimension from different offsets to determine equally spaced features, and the spatiotemporal input features are sampled in equal segments along the time axis to determine equally segmented features.
[0073] S32. Attention processing is performed on the equally spaced features and equally segmented features through the time attention submodule, and the corresponding outputs are equally spaced time-enhanced features and equally segmented time-enhanced features.
[0074] S33. Input the equally spaced time enhancement features and equally segmented time enhancement features into the gated fusion submodule for adaptive fusion, and output the time fusion features;
[0075] S34. Use the spatial attention submodule to process the temporal fusion features and output spatial features;
[0076] S35. The temporal and spatial features are fused through the gating fusion submodule to determine the spatiotemporal output features.
[0077] It should be noted that this embodiment uses a designed load prediction model to predict the network load of the low-altitude radio access network, in order to predict future network load. The node load sequence of each time slot is denoted as . , This represents the predicted node load sequence. Indicates the predicted time slot index. Indicates the total number of predicted time slots. Indicates the predicted time slot Predicted node load, This represents the total number of nodes, including centralized nodes and distributed nodes. Indicates the dimension of the load data;
[0078] Referring to Figure 3, the load prediction model includes a load input projection layer, a load encoder, a load decoder, and a load output projection layer:
[0079] Obtaining historical node load sequences Then, it is projected through the load input projection layer to obtain , To represent the feature dimension, in practical implementation, a non-linear network can be considered as the load input projection layer;
[0080] Next, further feature processing is performed using an encoder-decoder combination, where both the load encoder and load decoder include... Each of the three spatiotemporal attention (STA) modules comprises a hybrid sampling submodule, a temporal attention submodule, a spatial attention submodule, and a gated fusion submodule. The decoder's processing procedure is essentially the same as the encoder's, employing... Indicates the index of the spatiotemporal attention module, starting with the first... Taking the processing of a spatiotemporal attention module as an example:
[0081] First, a hybrid sampling strategy combining equal-interval sampling and equal-segmentation sampling is designed in the hybrid sampling submodule. In the equal-interval sampling strategy, the first... Spatiotemporal input features of each spatiotemporal attention module Sampled along the time dimension Subsequences are combined to form equally spaced features. , This represents the number of samples, with each subsequence starting from a different offset. In a segmented sampling strategy, this is... Divide evenly along the time axis into There are segments, each with a segment length of . The characteristics of equal segmentation are: , Indicates a segmented index. Represents the equally segmented feature vector;
[0082] Secondly, two temporal attention submodules are used to perform attention processing on equally spaced features and equally segmented features, respectively. For example, taking the temporal multi-head attention submodule as an example, the equally spaced features are... After linear projection is represented as Query, Key, and Value, a multi-head temporal attention mechanism is applied, and the outputs of all attention heads are concatenated to obtain equally spaced temporally enhanced features. The same process is applied to equidistant segmentation features. To obtain equally segmented time augmentation features ;
[0083] Furthermore, a gated fusion submodule adaptively fuses equally spaced time enhancement features and equally segmented time enhancement features to determine the time fusion features, which are then input into the spatial attention module for attention processing to output spatial features. For example, spatial multi-head attention can be used in specific implementations.
[0084] Finally, a gated fusion submodule is used to fuse the temporal and spatial features to obtain the output features of the spatiotemporal attention module, i.e., the spatiotemporal output features.
[0085] Once the decoding output characteristics of the decoder are determined, the predicted node load sequence is generated by projection through the load output projection layer.
[0086] Based on the above load prediction model, the macro trend, local fluctuations, and spatial dependencies between network nodes of the historical load sequence are captured to achieve accurate network load prediction.
[0087] Step 103: Based on the set of compression ratios and the predicted node load sequence, with the goal of maximizing service latency reward and load balancing reward, determine the set of resource scheduling actions for the current time slot through a multi-agent model.
[0088] In one specific embodiment of this example, step 103 includes the following sub-steps:
[0089] S41. Determine the joint observations of the multi-agent model, wherein the joint observations include global observations and local observations. The global observations include the set of node connection relationships of the low-altitude radio access network, the set of maximum available computing resources of nodes in the low-altitude radio access network, and the node load sequence. The local observations include the service requests of each UAV. The service requests include the number of frames to be transmitted, the uncompressed size of each data frame, the maximum allowable latency of each data frame, the minimum accuracy requirement of each data frame, and the set of compression ratios.
[0090] S42. Determine the action embedding of the multi-agent model, including a set of resource scheduling actions, wherein the set of resource scheduling actions includes the association decision variables between centralized nodes and UAVs, the association decision variables between distributed nodes and UAVs, the association decision variables between radio frequency nodes and UAVs, and the allocation decision variables between radio frequency nodes' resource blocks and UAVs.
[0091] S43. A multi-agent model trained with the goal of maximizing the total reward of service latency reward and load balancing reward is adopted. The model input is the joint observation and historical action embedding determined according to the set of compression ratios and the predicted node load sequence. The sequential decision output is the set of resource scheduling actions associated with the current time slot of each UAV.
[0092] It should be noted that this embodiment proposes a multi-agent model based on the combination of Contextual MAB architecture and multi-agent, enabling multiple agents to independently manage the service requests of a single UAV (unmanned aerial vehicle) and generate corresponding action embeddings. Indicates time slot A collection of drones, Indicates the first One drone, Indicates the number of drones:
[0093] The state space of a multi-agent model is defined as including: Joint observations by multiple agents, where joint observations Including global observation With local observation , , Indicates time slot Global observation, Indicates the first A drone in a time slot Local observations Indicates in time slot Global observation, This represents the set of node connections in a low-altitude radio access network. The set of node connections includes the set of nodes. and link set , Represents a node With nodes The link, Represents node index , Represents node index , This represents the maximum set of available computing resources for a node in a low-altitude wireless access network. ,node Maximum available computing resources It is expressed in units of cycles per second (cycles / s). Indicates in time slot The node load sequence, , Represents a node The load (understandably, RU nodes do not need to undertake computing tasks like CU and DU nodes, so usually only the load of CU and DU nodes can be considered). Indicates the first A drone in a time slot Local observations Indicates the first A drone in a time slot The set of real-time service requests, consisting of service requests, is denoted as . , Indicates the first A drone in a time slot Number of frames to be transmitted Indicates the first The uncompressed size of each data frame from each drone (in bits). Indicates the first A drone in a time slot The maximum allowable latency for each data frame in the process. Indicates the first A drone in a time slot The minimum accuracy requirement for each data frame. Indicates the first A drone in a time slot The set of compression ratios, , Indicates the first A drone in a time slot No. Compression ratio of each data frame;
[0094] The action embedding of the action space defining a multi-agent model includes a set of resource scheduling actions. , Indicates time slot Action embedding, Indicates the first One drone time slot The action embedding, where the resource scheduling action set Including centralized nodes and the first One drone time slot Related decision variables Distribution nodes and the first One drone time slot Related decision variables RF nodes and the first One drone time slot Related decision variables and the resource blocks (RBs) of the radio frequency nodes and the first One drone time slot Allocation decision variables , Represents the action space, Indicates the number of nodes in the set. Indicates the number of distributed nodes. Indicates the number of radio frequency nodes. This represents the number of resource blocks; understandably, it defines the time slots of centralized nodes and drones. The set of related decision variables is , Represents the centralized node index. Indicates the first The central node and the first One drone time slot The associated decision variables, if the drone With centralized nodes Related Distributed nodes and UAV time slots The set of related decision variables and radio frequency nodes and drone time slots The set of related decision variables Similarly, Indicates the index of the distributed nodes. Indicates the first The distribution node and the first One drone time slot Related decision variables, Indicates the radio frequency node index. Indicates the first The radio frequency node and the first One drone time slot Let the associated decision variables be... This represents the set of resource blocks for a radio frequency node. This represents the resource block index, where the resource block of the radio frequency node corresponds to the drone's time slot. The set of allocation decision variables is , Indicates the first The first radio frequency node The resource block and the first One drone time slot The allocation decision variables, if Then drone Assigned from the The first radio frequency node One resource block, otherwise ;
[0095] In this embodiment, the total reward function designed in the algorithm includes service latency reward and load balancing reward (to promote uniform load distribution among nodes). After training the multi-agent model with the goal of maximizing the total reward, the joint observations for resource scheduling calculation are determined based on the compression ratio set determined in step 101 and the predicted node load sequence determined in step 102. In the multi-agent model, when generating the current... When embedding the actions of an agent, the model relies on the joint observations of that agent and the action embeddings already output by the preceding agent, i.e., the historical action embeddings, as inputs to the model. The model outputs the action embeddings associated with the current time slot of each UAV through sequential decision-making.
[0096] In a more specific implementation of this embodiment, the multi-agent model is a multi-agent Transformer.
[0097] It should be noted that, in practical implementation, a multi-agent Transformer can be used, leveraging a Transformer encoder-decoder architecture to capture latent state representations and generate agent actions. Specifically, the agent encoder uses joint observations as encoding input and embeds the encoded joint observations into... As the encoded output, it is processed by the agent decoder as a joint observation embedding and historical action embedding. As the decoding input, the action embedding is processed through masked self-attention, cross-attention, multilayer perceptron (MLP), residual connections, and layer normalization to obtain the decoding output. It is understood that the specific implementation of the multi-agent Transformer can be found in existing technologies; only a brief description of the general process is provided here.
[0098] In a more specific implementation of this embodiment, the calculation process for the total reward includes:
[0099] ;
[0100] In the formula, Indicates time slot Total reward Indicates time slot Service delay reward Indicates time slot Load balancing rewards Indicates the first A drone in a time slot The average service latency, Represents the logarithm with base e as the natural constant. Indicates the network node index. Represents a centralized node. Represents distributed nodes, Indicates radio frequency node, Represents network nodes The computational overhead incurred Represents network nodes The maximum acceptable computational cost, Indicates the first A drone in a time slot Total service latency Indicates the first A drone in a time slot Number of frames to be transmitted Indicates the data frame index. Indicates the first A drone in a time slot The Middle Inference latency per data frame, Indicates the first A drone in a time slot The maximum allowable latency for each data frame in the process. Indicates the first One drone, Indicates time slot The collection of drones; it is understood that the inference latency mainly depends on the hardware computing power of the edge server (relatively fixed) and the size of the input data (determined by the compression ratio). Therefore, once the compression ratio is determined, the inference latency can be obtained by querying historical data and is not affected by newly created slices. In order to meet the latency requirements, this embodiment constrains the total service latency of the RAN and the inference latency of the model to not exceed the sum of the maximum allowable latency.
[0101] More specifically, the uplink transmission delay of data frames in the RAN mainly includes the transmission delay from UAV to RU, the processing delay of CU and DU, and the forwarding delay of RU and DU. Therefore, the total service delay includes the total transmission delay, the total processing delay, and the total forwarding delay. In the formula, Indicates the first A drone in a time slot Total transmission delay, Indicates the first A drone in a time slot Total processing latency, Indicates the first A drone in a time slot Total forwarding latency;
[0102] The total transmission delay is mainly determined by the throughput and the amount of data. The calculation process includes:
[0103] ;
[0104] In the formula, Indicates the first A drone in a time slot Total transmission delay, Indicates the data frame index. Indicates the first A drone in a time slot Number of frames to be transmitted Indicates the first A drone in a time slot No. Compression ratio of each data frame Indicates the radio frequency node index. Represents a set of radio frequency nodes. Indicates the first The radio frequency node and the first One drone time slot Related decision variables, Indicates the first The radio frequency node and the first A drone in a time slot Uplink transmission rate between;
[0105] The total processing latency includes:
[0106] ;
[0107] In the formula, Indicates the first A drone in a time slot Total processing latency, Indicates the first A drone in a time slot Number of data packets, Represents the centralized node index. Represents the set of central nodes. Indicates the first The central node and the first One drone time slot Related decision variables, Indicates the index of the distributed nodes. Represents the set of distributed nodes. Indicates the first The distribution node and the first One drone time slot Related decision variables, Indicates a data packet in the 1st... Average processing latency of each centralized node Indicates a data packet in the 1st... Average processing latency of distributed nodes;
[0108] The process of dividing each data frame into several data packets and calculating the number of data packets includes:
[0109] ;
[0110] In the formula, Indicates the first A drone in a time slot Number of data packets, Indicates the data frame index. Indicates the first A drone in a time slot Number of frames to be transmitted Indicates the first The uncompressed size of each data frame from each drone. Indicates the first A drone in a time slot No. Compression ratio of each data frame Indicates the data packet size (in bits).
[0111] Based on the M / D / 1 model, the average processing latency of the centralized nodes includes:
[0112] ;
[0113] In the formula, Indicates a data packet in the 1st... Average processing latency of each centralized node Indicates the first The data packet processing rate of a centralized node, Indicates the first The data packet arrival rate of a centralized node. Indicates the first One drone, Indicates time slot A collection of drones, Indicates the first The central node and the first One drone time slot Related decision variables, Indicates the first The data packet arrival rate of each drone; understandably. Depends on the node Maximum available computing resources , , Represents a node The load, This represents the resource coefficient, reflecting the elasticity between computing power and system load.
[0114] Similarly, the first The data packet arrival rate of each distributed node is expressed as: The calculation process for the average processing latency of distributed nodes can be found in the section on the average processing latency of centralized nodes.
[0115] The total forwarding latency, taking into account the number of data packets, includes:
[0116] ;
[0117] In the formula, Indicates the first A drone in a time slot Total forwarding latency, Indicates the first A drone in a time slot Number of data packets, Indicates the radio frequency node index. Represents a set of radio frequency nodes. Represents the centralized node index. Represents the set of central nodes. Indicates the index of the distributed nodes. Represents the set of distributed nodes. Indicates the first The radio frequency node and the first One drone time slot Related decision variables, Indicates the first The distribution node and the first One drone time slot Related decision variables, Indicates the first The central node and the first One drone time slot Related decision variables, Indicates time slot No. The distribution node at the th th Average forwarding latency at each radio frequency node Indicates time slot No. The centralized node at the ... Average forwarding latency at each distributed node;
[0118] According to the M / D / 1 model, the average forwarding latency of a centralized node at a distributed node includes:
[0119] ;
[0120] In the formula, Indicates time slot No. The centralized node at the ... Average forwarding latency at each distributed node Indicates time slot No. The packet forwarding rate of each distributed node, Indicates the first The data packet arrival rate of each distributed node; the average forwarding delay of the distributed node at the radio frequency node can be found in the calculation process of the average forwarding delay of the centralized node at the distributed node.
[0121] To better illustrate the effectiveness of this implementation, a simulation experiment was conducted:
[0122] A low-altitude economic scenario was constructed, comprising 20 UAVs, 2 RUs, a fronthaul network, 1 BBU, and an edge server. A YOLOv5 architecture deployed on the edge server was used to evaluate inference accuracy at different compression ratios and measure the corresponding system latency. During data collection, UAVs captured real-time data frames at different compression ratios and transmitted them to the YOLOv5 network via the RUs and BBU. RAN-DT, deployed on the edge server, was used for traffic prediction, determining the appropriate compression ratio through context and weight configuration and sending it to the UAVs via the network. Key parameter settings are summarized in Table 1.
[0123] Table 1 Parameter Settings
[0124]
[0125] To compare performance, several benchmark algorithms were set up for comparison: LinUCB, a linear contextual MAB algorithm that selects actions based on a confidence upper bound of the estimated reward; Multi-agent Deep Deterministic Policy Gradient (MA-DDPG) algorithm, which is used to evaluate average service latency through distributed collaboration; and Branch DDQN (B-DDQN), which extends the Dual Deep Q Network (DDQN) with a branch architecture and incorporates Bayesian exploration to improve learning efficiency.
[0126] As shown in Figure 4, simulation experiment 1 regarding reward and iteration count demonstrates the convergence performance of the algorithm. The results show that the algorithm proposed in this embodiment has a faster convergence speed and higher reward stability. In the early stage of training (the first 100 iterations), the rewards of the three algorithms all show a gradual upward trend, indicating that the algorithms are effectively exploring the environment. However, the algorithm proposed in this embodiment shows a steeper growth rate, which means that its learning efficiency is higher and its policy optimization speed is faster. When the number of training iterations exceeds 250, the algorithm proposed in this embodiment can maintain a stable high reward with minimal fluctuation, demonstrating strong convergence stability and small policy fluctuation. In contrast, the MA-DDPG algorithm shows moderate fluctuation after convergence, while the B-DDQN algorithm shows obvious instability and frequent reward decreases, indicating that the algorithm is sensitive to non-stationary environments and has limited robustness.
[0127] As shown in Figure 5, simulation experiment 2 demonstrates the relationship between compression accuracy and weight parameters. The relationship between the two weights; under all weight configurations, the algorithm proposed in this embodiment can always achieve higher compression accuracy, which demonstrates its excellent adaptability in balancing compression efficiency and reconstruction fidelity; when When the weights are increased from 6 to 10, the accuracy of all schemes improves, indicating that assigning higher weights to the reconstruction loss enhances the model's ability to preserve feature integrity during compression. This result shows that the algorithm proposed in this embodiment can effectively balance compression ratio and accuracy by dynamically adjusting to the optimal weight range. Overall, the proposed algorithm outperforms the benchmark algorithm across all weight ranges, demonstrating its effectiveness in achieving adaptive and high-fidelity compression of real-time data transmission in a RAN environment.
[0128] As shown in Figure 6, simulation experiment 3 compared the system service latency under different numbers of tasks. The results show that, under all task numbers, the algorithm proposed in this embodiment always maintains the lowest latency, demonstrating excellent scheduling efficiency and scalability. As can be seen from Figure 6, when the number of tasks exceeds 20, the latency of all algorithms shows an upward trend as network congestion and computing demands increase. However, the growth rate of the algorithm proposed in this embodiment is significantly lower. This indicates that under high load scenarios, the algorithm has stronger robustness and higher resource utilization efficiency. In contrast, the B-DDQN algorithm exhibits the highest latency (exceeding 1.15 seconds with 30 tasks), reflecting its limited adaptability to dynamic task changes. This result proves the advantages of the algorithm proposed in this embodiment in managing real-time task scheduling and resource coordination in RAN systems, especially in highly dynamic and dense network environments.
[0129] In this embodiment of the invention, through the hierarchical collaboration of compression ratio optimization-context multi-armed gambling machine algorithm and load prediction model, the compression ratio is dynamically adjusted to effectively reduce redundant data transmission and lower network load. The hybrid sampling attention mechanism is used to achieve accurate network load prediction. The resource allocation based on multi-agent model is achieved to minimize average service latency and balance load, thereby improving network adaptability and service reliability in low-altitude scenarios.
[0130] Please refer to Figure 7. A low-altitude wireless access network resource scheduling system provided in Embodiment 2 of the present invention includes:
[0131] Compression ratio determination module 701 is used to determine the set of compression ratios for any UAV in the current time slot through the compression ratio optimization-context multi-arm gambling machine algorithm;
[0132] The node load prediction module 702 is used to output the predicted node load sequence based on the historical node load sequence of the low-altitude wireless access network using a load prediction model based on hybrid sampling and attention enhancement adaptive fusion.
[0133] The network resource scheduling module 703 is used to determine the set of resource scheduling actions for the current time slot based on the set of compression ratios and the predicted node load sequence, with the goal of maximizing service latency rewards and load balancing rewards, through a multi-agent model.
[0134] Embodiment 3 of the present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program; when the computer program is executed by the processor, the processor performs the steps of the low-altitude wireless access network resource scheduling method as described in Embodiment 1 of the present invention.
[0135] Embodiment 4 of the present invention also provides a computer-readable storage medium storing a computer program / instruction thereon, which, when executed by a processor, implements the steps of the low-altitude wireless access network resource scheduling method described in Embodiment 1 of the present invention.
[0136] Embodiment 5 of the present invention also provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the low-altitude wireless access network resource scheduling method as described in Embodiment 1 of the present invention.
[0137] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the system and modules described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0138] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the shown or discussed mutual couplings, direct couplings, or communication connections may be through some interfaces; indirect couplings or communication connections between systems or modules may be electrical, mechanical, or other forms.
[0139] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0140] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0141] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0142] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for scheduling low-altitude wireless access network resources, characterized in that, include: The compression ratio set for the current time slot of any UAV is determined by the compression ratio optimization-context multi-arm gambling machine algorithm; Based on the historical node load sequence of the low-altitude wireless access network, a load prediction model based on hybrid sampling and attention enhancement adaptive fusion is used to output the predicted node load sequence. According to the compression ratio set and the predicted node load sequence, with the goal of maximizing service latency reward and load balancing reward, a set of resource scheduling actions for the current time slot is determined through a multi-agent model.
2. The low-altitude wireless access network resource scheduling method according to claim 1, characterized in that, The step of determining the compression ratio set of the current time slot of the UAV using the compression ratio optimization-context multi-armed gambling machine algorithm includes: determining the context features of the compression ratio optimization-context multi-armed gambling machine algorithm, including the maximum allowable latency of each data frame, the minimum accuracy requirement of each data frame, and weight parameters used to balance compression efficiency and inference accuracy; determining the actions of the compression ratio optimization-context multi-armed gambling machine algorithm, including the compression ratio set; constructing context-action pairs using the context features and the actions, and determining the joint posterior distribution of the rewards for the context-action pairs; determining the mean and variance of the context-action pairs based on the joint posterior distribution of the rewards; and determining the compression ratio set of any UAV's current time slot by solving for the maximum value of the upper confidence bound acquisition function based on the mean and the variance.
3. The low-altitude wireless access network resource scheduling method according to claim 1, characterized in that, The load prediction model includes a load input projection layer, a load encoder, a load decoder, and a load output projection layer. The load prediction model, based on historical node load sequences from a low-altitude radio access network (LAN), outputs predicted node load sequences using a load prediction model that employs a hybrid sampling and attention-enhanced adaptive fusion approach. This includes: projecting the historical node load sequences from the LNA into load input features through the load input projection layer; encoding and decoding the load input features using the load encoder and load decoder; and then processing and outputting the predicted node load sequences using the load output projection layer. Each load encoder and load decoder includes multiple spatiotemporal attention modules, which in turn include a hybrid sampling submodule, a temporal attention submodule, and a spatial attention submodule. The spatiotemporal attention module's processing includes: based on the hybrid sampling submodule, sampling the spatiotemporal input features at equal intervals along the time dimension from different offsets to determine equally spaced features, and sampling the spatiotemporal input features in equal segments along the time axis to determine equally segmented features; performing attention processing on the equally spaced features and equally segmented features through the time attention submodule, respectively, to output equally spaced time-enhanced features and equally segmented time-enhanced features; inputting the equally spaced time-enhanced features and equally segmented time-enhanced features into the gated fusion submodule for adaptive fusion, outputting time-fusion features; using the spatial attention submodule to process the time-fusion features and output spatial features; and fusing the time-fusion features and spatial features through the gated fusion submodule to determine the spatiotemporal output features.
4. The low-altitude wireless access network resource scheduling method according to claim 1, characterized in that, The step of determining the resource scheduling action set for the current time slot based on the compression ratio sets and the predicted node load sequence, with the goal of maximizing service latency rewards and load balancing rewards, through a multi-agent model, includes: determining the joint observations of the multi-agent model, wherein the joint observations include global observations and local observations. Global observations include the set of node connection relationships in the low-altitude radio access network, the set of maximum available computing resources for each node in the low-altitude radio access network, and the node load sequence. Local observations include the service requests of each UAV, where each service request includes the number of frames to be transmitted, the uncompressed size of each data frame, the maximum allowable latency for each data frame, and the minimum required latency for each data frame. Accuracy requirements and compression ratio sets are defined. The action embeddings of the multi-agent model are determined, including a resource scheduling action set. This resource scheduling action set includes the association decision variables between centralized nodes and UAVs, distributed nodes and UAVs, radio frequency nodes and UAVs, and the allocation decision variables between radio frequency nodes' resource blocks and UAVs. A multi-agent model trained with the goal of maximizing the total reward of service latency reward and load balancing reward is used. The model input is the joint observation and historical action embeddings determined based on each compression ratio set and the predicted node load sequence. The model outputs the resource scheduling action set associated with each UAV's current time slot through sequential decision-making.
5. The low-altitude wireless access network resource scheduling method according to claim 4, characterized in that, The calculation process for the total reward includes: In the formula, Indicates time slot Total reward Indicates time slot Service delay reward Indicates time slot Load balancing rewards Indicates the first A drone in a time slot The average service latency, Represents the logarithm with base e as the natural constant. Indicates the network node index. Represents a centralized node. Represents distributed nodes, Indicates radio frequency node, Represents network nodes The computational overhead incurred Represents network nodes The maximum acceptable computational cost, Indicates the first A drone in a time slot Total service latency Indicates the first A drone in a time slot Number of frames to be transmitted Indicates the data frame index. Indicates the first A drone in a time slot The Middle Inference latency per data frame, Indicates the first A drone in a time slot The maximum allowable latency for each data frame in the process. Indicates the first One drone, Indicates time slot A collection of drones.
6. The low-altitude wireless access network resource scheduling method according to claim 5, characterized in that, Total service latency includes total transmission latency, total processing latency, and total forwarding latency.
7. A low-altitude wireless access network resource scheduling system, characterized in that, include: The compression ratio determination module is used to determine the set of compression ratios for any UAV in the current time slot through the compression ratio optimization-context multi-arm gambling machine algorithm; The node load prediction module is used to output a predicted node load sequence based on the historical node load sequence of the low-altitude wireless access network using a load prediction model based on hybrid sampling and attention enhancement adaptive fusion. The network resource scheduling module is used to determine the set of resource scheduling actions for the current time slot based on the compression ratio set and the predicted node load sequence, with the goal of maximizing service latency reward and load balancing reward, through a multi-agent model.
8. A computer device, characterized in that, The system includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor causes the processor to perform the steps of the low-altitude wireless access network resource scheduling method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the low-altitude wireless access network resource scheduling method as described in any one of claims 1-6.
10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the low-altitude wireless access network resource scheduling method as described in any one of claims 1-6.