NOMA-enabled air-ground content distribution network trajectory and resource optimization method
By decoupling discrete variables and continuous variables through the matching-deep reinforcement learning method and combining it with the DDPG framework to train the UAV's flight trajectory and power allocation, the complexity of UAV trajectory planning and resource allocation in the NOMA-assisted air-ground network is solved, and the content acquisition latency is minimized and the service quality is improved.
Patent Information
- Application Number
- CN202310334491.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-31
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2043-03-31
AI Technical Summary
Existing technologies are unable to efficiently solve the high computational complexity of UAV trajectory planning in NOMA-assisted air-ground networks, the multi-dimensional resource coupling of resource allocation problems, and the optimization challenges in dynamic network environments. In addition, the overhead of obtaining global state information in actual systems is high, resulting in low optimization decision-making efficiency.
The matching-deep reinforcement learning method is adopted to decouple discrete variables and continuous variables through the many-to-one matching theory. Combined with the DDPG framework, the UAV's flight trajectory and power allocation decision are trained to achieve joint optimization of UAV trajectory and resources, thereby reducing content acquisition latency.
It effectively reduces the user's content acquisition latency, improves service quality, optimizes the resource allocation and trajectory planning of the NOMA-assisted air-ground network, and improves system performance.
Smart Images

Figure CN118741532B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of air-ground integrated networks, and in particular to a trajectory and resource optimization method for a NOMA-enabled air-ground content distribution network. Background Art
[0002] In recent years, mobile edge caching has garnered significant attention as a solution to alleviate the challenges of backhaul congestion and increased core network burden caused by the massive influx of IoT devices. By extracting and deploying hotspot content at edge nodes and distributing it directly to users, content retrieval latency can be effectively reduced. Furthermore, air-ground content distribution networks, leveraging the high line-of-sight probability and flexible deployment of UAVs, can achieve enhanced coverage of terrestrial hotspots, providing reliable content distribution services to users even when base stations are overloaded.
[0003] NOMA technology allows multiple users to simultaneously share orthogonal channel resources, significantly improving system spectrum efficiency and throughput. Furthermore, successive interference cancellation is used to mitigate interference caused by channel reuse, and proper power control and channel allocation are crucial for user fairness and system performance in NOMA-assisted air-ground networks.
[0004] However, achieving efficient content distribution in NOMA-assisted air-ground networks still faces severe challenges. For mobile UAVs, the trajectory planning problem is a continuous non-convex optimization problem. Traditional optimization methods have high computational complexity and are difficult to adapt to highly dynamic network environments. The resource allocation problem in NOMA-assisted networks is usually an NP-hard mixed integer nonlinear programming problem with multi-dimensional resource coupling. Existing methods cannot guarantee the optimization results with polynomial complexity. What is even more difficult is that the overhead of obtaining global state information in actual systems is extremely high, so the trajectory and resource optimization decisions of UAVs can only be selected based on local observations. In order to effectively address the above challenges, the present invention proposes a trajectory and resource optimization method for NOMA-enabled air-ground content distribution networks, which can minimize the latency of user content acquisition. Summary of the Invention
[0005] The present invention discloses a trajectory and resource optimization method for a NOMA-enabled air-ground content distribution network, mainly for the joint optimization of flight trajectory and network resources when a NOMA-enabled UAV-assisted air-ground network distributes content to users. The method comprises the following steps: step 1, deriving a communication rate expression when the UAV distributes content to users based on a NOMA-enabled air-ground network model, and deriving delay formulas for edge, collaborative and cloud content transmission in combination with cache deployment conditions, to describe a joint UAV trajectory planning and resource allocation problem for minimizing network content distribution delay; step 2, further splitting the optimization problem into a channel allocation subproblem and a joint UAV trajectory and transmission power optimization subproblem based on the many-to-one matching theory, thereby decoupling discrete variables and continuous variables; step 3, adopting a matching model to describe the multiplexing relationship between subchannels and UAVs, setting the matching direction and convergence conditions according to the swap mechanism, and obtaining the optimal subchannel allocation result; step 4, adopting a DDPG framework to train the flight trajectory and power allocation decision of the UAV, so that the UAV follows the user's moving trajectory to distribute content, while alleviating the co-channel interference of the downlink NOMA. This paper decouples the hybrid action space through a matching-deep reinforcement learning method, thereby efficiently solving the trajectory planning and resource allocation problems of the air-ground content distribution network, reducing the user's content acquisition latency and effectively improving the service quality. The specific process is as follows:
[0006] The NOMA-enabled air-ground content distribution network proposed in this invention includes a high altitude platform station (HAPS), UAVs and users, total Content cache time slots Indicates that the length of each time slot is Assume that the UAV divides users into The clusters are represented as and Slave to UAV .Using spherical coordinates to represent UAV In the time slot Flight trajectory ,in is the flight speed, then UAV The change in coordinates can be expressed as
[0007]
[0008] Total in the system Each UAV uses an orthogonal channel to distribute content to subordinate users through NOMA. Terminals on the same channel will be interfered with by each other. Continuous interference cancellation technology can effectively alleviate the mutual interference of users in the same cluster. The channel reuse strategy of UAV is expressed as , then UAV Occupied channel Service users The communication rate is
[0009]
[0010] in Refers to the channel bandwidth; Refers to the channel gain of the UAV to the user, which can be calculated by substituting the real-time position information into the channel model of the ground-to-air link; Cluster Number of users; is the noise power.
[0011] UAV collaborative caching is used to improve content hit rate. Specifically, when the user's subordinate UAV caches the requested content, the UAV directly sends the content to the user, which is edge content acquisition. , otherwise it is 0; when the slave UAV has not cached the requested content, but other UAVs in the network cache the requested content, the UAV that caches the content sends the content to HAPS, which then forwards it to the slave UAV and transmits it to the user. This is called collaborative content acquisition. If all UAVs in the network do not cache the requested content, the content can only be sent from the cloud to HAPS, slave UAVs, and users in sequence. . Note 、 and It is determined by the user request and the UAV cache decision, and the user can only obtain the requested content in one way, that is, Based on the above content acquisition process, the edge, collaborative and cloud content distribution delays are calculated as
[0012]
[0013] in is the content data volume, and are the UAV-HAPS link and HAPS-cloud link rates, respectively, which are assumed to be constant.
[0014] The problem of minimizing the content delivery delay in NOMA-enabled air-ground networks can be described as follows:
[0015] : , : , , : , , : , , : , , , : , , : , , : , , ,
[0016] The optimization variables are channel allocation, UAV trajectory and power allocation; Obtaining the definition of indicator factors for content; is the channel multiplexing constraint; Refers to UAV trajectory constraints; Refers to power allocation constraints; Refers to the relationship between delay size.
[0017] The solution to this problem can be divided into the following steps:
[0018] 1) According to the matching theory, the problem is further decomposed into two sub-problems: channel allocation and joint UAV trajectory and power optimization. Specifically, the many-to-one matching function is defined For: From the collection To the collection The mapping, UAV Multiplexed sub-channel ,at the same time It must also be established, and Respectively represent and Therefore, the problem can be transformed into two sub-problems: channel allocation optimization and UAV trajectory planning and power control under fixed channel allocation.
[0019] 2) For the channel allocation subproblem, the matching utility function between the UAV and the subchannel is calculated as follows:
[0020]
[0021] Set the swap operation condition to 1) , both 2) ,make , and through the swap operation Swapping two UAVs and Matched subchannel and , until the above conditions are no longer met and converge to a stable matching result, that is, the optimal sub-channel allocation solution.
[0022] 3) For the joint UAV trajectory and power optimization sub-problem, the DDPG architecture is adopted to treat the UAV as agents, each agent Have an actor network Used to select decisions based on the current state space , critic network Responsible for evaluating state action values, where and are the neural network parameters. In addition, the state space in the Markov decision process component The specific definition of the open content delivery network is
[0023]
[0024] in and Refers to the coordinates of the user and UAV respectively; action space is defined as follows:
[0025]
[0026] Reward Function The design is as follows:
[0027]
[0028] It is noted that the reward functions of each UAV are correlated with each other through the interference term, which promotes the UAVs to maintain distance during content distribution and alleviate co-channel interference.
[0029] 4) The DDPG neural network consists of an input layer, a hidden layer, and an output layer, all of which adopt a fully connected structure. Specifically, the hidden layers of the network use the ReLU activation function, and the output layer of the UAV actor network uses the sigmoid activation function to obtain , using tanh to obtain , and output through softmax , the critic network of each agent takes global observations and actions as input and outputs state-action values. In addition, MADDPG uses experience replay technology to replay the experience tuples obtained from each interaction between the agent and the environment. Stored in the replay buffer, each training from the buffer random experience tuples are used to update the neural network weights, thereby breaking the correlation of training data; each agent has another target actor network With target critic network , its structure is the same as the original network, but the weight update is much slower than the original network, which can prevent training oscillations.
[0030] 5) During the training phase of DDPG, the agent accumulates historical experience through continuous interaction with the environment and backpropagates to update the neural network weights. At the same time, random noise needs to be added to the strategy to stimulate the agent's exploration, that is, ,in is the noise that follows a Gaussian distribution. After that, the neural network is trained using stochastic gradient descent, and the weights of the actor and critic networks are updated by minimizing the following loss functions:
[0031]
[0032] This results in a fully trained intelligent agent neural network, where each UAV can make flight trajectory and power allocation decisions based on real-time observations.
[0033] The technical method of the present invention has the following advantages:
[0034] First, the joint optimization of NOMA enables UAV trajectory planning and resource allocation decisions in air-ground content distribution networks, achieving the goal of minimizing content delivery latency. Second, a matching deep reinforcement learning framework is employed to decouple the mixed action space, effectively solving channel allocation results based on many-to-one matching theory, while also enhancing the convergence of DDPG training, thereby achieving joint optimization of trajectory and network resource allocation. Finally, the proposed method optimizes content delivery from drones to users, reducing user content access latency and effectively improving service quality.
[0035] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the embodiments of the present invention or the technical methods in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0037] Figure 1 This is the curve of the cumulative return of the agent relative to the round during the DDPG training process.
[0038] Figure 2 This is the curve of the average content delivery delay versus time slot.
[0039] Figure 3 is the flight trajectory curve of the UAV.
[0040] Figure 4 is the curve of average throughput versus the number of users. DETAILED DESCRIPTION
[0041] The present invention proposes a trajectory and resource optimization method for a NOMA-enabled air-ground content distribution network. The embodiments are described in detail below with reference to the accompanying drawings.
[0042] The implementation scenario of this invention is a simulated area of 1 km × 1 km, containing 3-6 UAVs and a large number of randomly distributed ground terminals. The UAVs fly at altitudes between 100 and 300 m and at a speed of 10 m / s. The maximum user speed is 0.5 m / s. The number of orthogonal channels is 6, and the upper limit of multiplexing for each channel is 3.
[0043] The simulation time dimension of the present invention is 200, the length of each time slot is 0.2 s, and the method of the present invention is executed in each time slot and the optimization result is obtained. The simulation parameters are selected as follows: the communication channel between the terminal and HAPS and UAV adopts a probabilistic path loss model, where the carrier frequency is 0.1 GHz, the environmental parameters are 4.88 and 0.43, the additional path loss distribution of line-of-sight and non-line-of-sight is 0.1 dB and 21 dB; the content data volume is [8, 20] MB; the maximum power of UAV is 29 dBm, the channel bandwidth is 20 MHz, and the noise power is -144 dBm. The hyperparameters of DDPG are selected as follows: the discount factor is 0.9, the learning rate is 10 -6 , the target network update rate is 0.005.
[0044] The specific implementation steps are as follows:
[0045] 1) Based on the NOMA-enabled air-ground network model, we derive the communication rate expression when UAV distributes content to users. Combined with the cache deployment, we derive the delay formulas for edge, collaborative, and cloud content transmission. We then describe the joint UAV trajectory planning and resource allocation problem to minimize the network content distribution delay.
[0046] 2) Define a many-to-one matching function and transform the problem into two sub-problems: channel allocation optimization and UAV trajectory planning and power control under fixed channel allocation.
[0047] 3) For the channel allocation subproblem, initialize a and Random matching , and then traverse any two UAVs in the system and ,according to Find the orthogonal channels they multiplex and , calculate the communication rate and obtain the matching utility function. Then calculate the swap operation The utility function below determines whether the conditions for the swap operation are met. If so, use Replace the original match , otherwise keep matching The process remains unchanged and the process repeats until there is no UAV in the system that meets the swap condition, and the sub-channel allocation solution is obtained.
[0048] 4) For the joint trajectory planning and power allocation sub-problem, the DDPG architecture is used for training. The agent interacts with the air-ground content distribution network environment. That is, the UAV selects the flight trajectory based on the actor network. After that, each agent receives a reward and the system state transitions to the next time slot. At the same time, each agent sends the experience tuple generated in this time slot to the next time slot. Store in experience buffer.
[0049] 5) Using experience replay technology, experience tuples are randomly extracted from the experience buffer during each training session. The target network outputs the target action and evaluation value, and the error function between the actor and critic networks is calculated. Subsequently, the neural network weights are updated using stochastic gradient descent. Training continues until the set number of training rounds is reached, resulting in a fully trained UAV trajectory planning and power allocation strategy.
[0050] 6) The above optimization method is executed in each time slot until the optimization time ends.
[0051] Figure 1The figure shows the evolution of the agent's cumulative reward over each round during DDPG training. The curve shows a fluctuating upward trend at the beginning of training, as the UAV explores how to provide content distribution services while avoiding co-channel interference. DDPG training converges when the agent learns the optimal strategy. From then on, the UAV can directly make flight direction and power allocation decisions based on the fully trained neural network.
[0052] Figure 2 The variation of average content delivery delay relative to time slots is shown, demonstrating that the proposed method outperforms all baseline algorithms. This is because the proposed method jointly optimizes channel allocation, UAV trajectory, and power allocation, effectively mitigating intra-cluster and inter-cluster interference caused by downlink NOMA, thereby improving user service quality. Numerical results show that the average content delivery delay of the proposed method is 34.80% lower than that of fixed UAV trajectory and random resource allocation methods, respectively.
[0053] Figure 3 The UAV flight trajectory curves are shown, showing that each UAV flies directly from its initial position to the corresponding user cluster, then follows the user's movement trajectory to provide content distribution services. Given the UAV flight altitude constraint, the UAV chooses to increase its altitude, thereby increasing the probability of line of sight and reducing path loss. These trajectory results demonstrate that the proposed method can effectively adjust UAV flight trajectories to provide users with better content distribution services.
[0054] Figure 4 The average throughput is plotted against the number of users, showing that throughput decreases with increasing user numbers. This is because user clusters become more dispersed, increasing the distance content must be transmitted. Furthermore, the increased amount of content required for distribution leads to more severe co-channel interference. Numerical results show that the throughput of the proposed method outperforms the random resource allocation and fixed UAV trajectory methods by 17.17% and 44.45% respectively.
[0055] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the present disclosure is not limited to the technical methods formed by a specific combination of the above-mentioned technical features, but also encompasses other technical methods formed by any combination of the above-mentioned technical features or their equivalents without departing from the inventive concept. For example, a technical method formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in the present disclosure.
Claims
1. A trajectory and resource optimization method for a NOMA-enabled air-ground content delivery network. This method aims to jointly optimize flight trajectories and network resources when a NOMA-enabled UAV assists in delivering content to users in an air-ground network. The method comprises the following steps: Step 1: Establish a NOMA-enabled air-ground network model consisting of a high-altitude platform station, K UAVs, and M users. Derive the communication rate expression when the UAV distributes content to users. Combined with the cache deployment, derive the delay formulas for edge, collaborative, and cloud content transmission. The joint UAV trajectory planning and resource allocation problem for minimizing the network content distribution delay is described as follows: in refers to the kth UAV, is the mth user, who is in the user cluster served by UAVk middle, refers to the i-th orthogonal channel, the total number of UAVs, users and channels are K, M and I respectively, t is the time slot index, and the total time dimension is T; A(t) = {a k,i (t)} refers to the channel allocation strategy, a k,i (t) = 1 means that UAVk occupies channel i in time slot t, otherwise a k,i (t) = 0; is the UAV flight trajectory, and are vertical and horizontal angles respectively; P(t)={P k,m (t)} is the UAV power allocation strategy, P k,m (t) represents the power allocated by UAVk to user m; and are indicator factors for edge, collaborative, and cloud content acquisition methods, respectively. When the user acquires content in the corresponding method, the value is 1, otherwise it is 0; and is the delay of the corresponding content acquisition method; C1 is the definition constraint of the content acquisition method indicator factor, and the user can only obtain the requested content in one way; C2, C3 and C4 refer to a k,i (t) Definition constraints: a UAV can only reuse one sub-channel and the same sub-channel can be used by at most UAVs are reused; C5 and C6 are trajectory constraints, where h min and h max They refer to the minimum and maximum altitudes of the UAV flight respectively; C7 is the power allocation constraint, where is the maximum power of the UAV; C8 refers to the latency ranking of the three content acquisition methods; Step 2: Solving the joint UAV trajectory planning and resource allocation problem for minimizing network content delivery latency is as follows: First, based on the many-to-one matching theory, the optimization problem is further split into a channel allocation subproblem with discrete optimization variables and a joint UAV trajectory and transmit power optimization subproblem with continuous optimization variables, thereby decoupling the hybrid action space; Step 3. Secondly, for the channel allocation sub-problem, a matching model is used to describe the multiplexing relationship between sub-channels and UAVs, and the utility function of the matching parties is calculated, where: The utility function of a subchannel is the sum of the communication rates of all UAVs occupying it, and the utility function of a UAV is the sum of the rates of its service users. The optimal subchannel allocation result is obtained by setting the matching direction and convergence conditions based on the swap mechanism. Step 4. Finally, for the UAV trajectory and power optimization sub-problems, the DDPG framework is used to define the state space s of each UAV separately. k (t), action space a k (t) and reward function r k (t), and then generate the actor network π k (s k (t);φ k ) makes decisions, critic network Q k (s k (t),a k (t);θ k ) is responsible for evaluation, where φ k and θ k For the neural network weights, the agent uses the experience replay technique to randomly extract experience tuples during each training, and then the weights of the actor and critic networks are updated by minimizing the following loss functions: where a′ k (t+1) and Q′ k (s k (t+1),a′ k (t+1);θ′ k ) refer to the output of the target actor and the critic network respectively, and γ is the discount factor, so as to obtain the optimization result of the joint UAV trajectory and network resources.
Citation Information
Patent Citations
Content deployment optimization method for unmanned aerial vehicle cooperative caching 6G network
CN118740929A
Unmanned aerial vehicle track and task unloading optimization method based on multi-agent learning
CN118741602A