An Energy Efficiency Maximization Method for UAV Cooperative NOMA Communication Networks

Through deep Q learning and Lagrangian dual method, the UAV collaborative NOMA communication network is optimized, which solves the problem of resource allocation in hardware damage and dynamic environments, maximizes the energy efficiency of UAV collaborative NOMA communication network, and provides a more practical resource allocation strategy.

CN116170824BActive Publication Date: 2025-07-22HANGZHOU FEISUAN TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310065613.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-13
Publication Date
2025-07-22
Estimated Expiration
2043-01-13

AI Technical Summary

Technical Problem

The existing UAV collaborative NOMA communication network fails to effectively consider the impact of hardware damage on communication performance, and traditional convex optimization methods are difficult to obtain the optimal solution, are highly complex, and are difficult to apply to actual networks.

Method used

The deep Q learning algorithm and the Lagrangian dual method are adopted, combined with reinforcement learning and convex optimization, and the energy efficiency maximization model of the UAV collaborative NOMA communication network is optimized, taking into account hardware damage and dynamic environment, and the resource allocation strategy is updated through iterative optimization and deep neural networks.

Benefits of technology

Taking into account hardware damage and dynamic environments, optimizing the energy efficiency of the UAV collaborative NOMA communication network provides a more practical resource allocation solution, avoiding the waste of computing resources and local optimal solutions of traditional methods, and improving system energy efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116170824B_ABST
    Figure CN116170824B_ABST
Patent Text Reader

Abstract

The present invention claims protection for a method for maximizing the energy efficiency of a non-orthogonal multiple access (NOMA) network in drone cooperative communication, belonging to the field of NOMA network resource allocation. Under the constraints of user quality of service, base station transmit power, drone transmit power, flight trajectory, and decoding order, the energy efficiency of the NOMA network system for UAV cooperative communication is maximized. Its innovation lies in that, different from traditional drone cooperative NOMA communication resource management, the present invention uses reinforcement learning to perform resource allocation and also considers the impact of actual hardware imperfections on system performance. The present invention designs a resource allocation scheme using methods such as deep Q-learning, Lagrangian duality, and quadratic transformation. The method provided by the present invention provides a solution for system energy efficiency and drone flight trajectory, which is more practical and easier to transplant to handle complex scenarios compared to the scheme that uses traditional convex optimization methods and only considers the drone deployment location.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of wireless resource management, and particularly to a method for maximizing the energy efficiency of an unmanned aerial vehicle (UAV) cooperative non-orthogonal multiple access (NOMA) communication network under the condition of hardware impairment. Background Art

[0002] The development of UAV communication technology has further improved the service range of NOMA networks. Therefore, the combination of the two is considered a promising solution for large-scale device access in future wireless networks. In addition to improving the communication coverage of the system, UAVs also naturally have the characteristics of high controllability and low overhead. Therefore, when dealing with emergency situations such as establishing emergency communication after a disaster, they can restore communication more quickly than traditional means.

[0003] By studying the wireless resource management method of UAV cooperative NOMA communication, it is found that the current scenarios of UAV combined with NOMA for communication are mainly divided into the following two types:

[0004] First, the actual communication scenario is not considered. In the actual communication scenario, the radio frequency generators at the transceiver ends are often under non-ideal conditions and are affected by actual hardware impairments such as in-phase quadrature imbalance, nonlinear amplification noise, and radio frequency circuit noise. These effects have a non-negligible negative impact on the communication performance of the system. For example, Ningyuan Wang et al. published an article titled "NOMA-Based Energy-Efficiency Optimization for UAV Enabled Space-Air-Ground Integrated Relay Networks" in "IEEE N. Wang, F. Li, D. Chen, L. Liu and Z. Bao,", "in IEEE Transactions on Vehicular Technology, 2022, 71(4): 4129-4141." which studied the energy efficiency problem of a UAV-assisted relay NOMA system. However, unfortunately, the performance impact of actual non-ideal factors on the system was not considered and it cannot be applied to the actual network. Secondly, the solution method is too complex, resulting in difficulty in transplantation to other scenarios.

[0005] Second, most resource allocation methods adopt traditional convex optimization methods or intelligent algorithms. It is true that the resource management scheme derived by modeling and analyzing a single scenario is feasible. However, the actual environment will be different from the model. Secondly, traditional convex optimization methods often cannot obtain the optimal solution due to the complexity of the target problem. In addition, using intelligent algorithms is also likely to fall into local optimal solutions. The method of reinforcement learning can avoid cumbersome mathematical operation processes and can be used as an online mechanism. When the actual environment changes, there is no need to recalculate the resource allocation scheme. Through reinforcement learning, the optimization goal can be changed to a long-term optimization goal, and the agent can obtain the best resource allocation management strategy according to learning.

[0006] Therefore, on the basis of considering the above two problems, the present invention designs an energy efficiency maximization method for a UAV cooperative NOMA network under the condition of hardware impairment, which has important practical significance and application value. The processing of the continuous power allocation coefficient for this scheme is a technical difficulty. Summary of the Invention

[0007] The present invention aims to solve the above problems of the prior art. It proposes an energy efficiency maximization method for a UAV cooperative NOMA communication network. The technical solution of the present invention is as follows:

[0008] An energy efficiency maximization method for a UAV cooperative NOMA communication network, which includes the following steps:

[0009] Step 1), establish an energy efficiency maximization model for a UAV cooperative NOMA communication network under the condition of hardware impairment;

[0010] Step 2), iteratively optimize and solve the energy efficiency maximization model, initialize the starting point, ending point positions of the UAVs and their transmission powers, decoding order variables, base station position and transmission power, user minimum rate threshold, base station power allocation coefficient, flight cycle size;

[0011] Step 3), set the parameters of the environment in the deep Q-learning algorithm, define the state space as {horizontal coordinate of the UAV, system energy efficiency}, define the action space as {base station transmission power, UAV transmission power, UAV flight direction, UAV flight speed}, and define the reward as {system energy efficiency}, where the UAV acts as an agent;

[0012] Step 4), solve the decoding order and power allocation coefficient according to the initial position of the UAV and its power, base station transmission power, and judge whether the power allocation coefficient converges. If it has converged, update the power allocation coefficient, otherwise repeat this step until the result converges;

[0013] Step 5): Input the horizontal position coordinates, decoding order, flight direction, speed, base station transmission power, and the power allocation coefficient after convergence in step 4) of the UAV at the current moment into the deep neural network, and update the Q value in the deep Q-learning algorithm;

[0014] Step 6): Through learning the samples, the deep neural network outputs the reward obtained from the environment at the current moment, changes the agent state to the next moment, and updates the base station transmission power. At the same time, it stores the learned samples in the experience replay pool and saves the output parameters;

[0015] Step 7): Determine whether the UAV flight cycle has ended. If it has not ended, go to step 4). If it has ended, first determine whether the average reward of the UAV flight cycle has converged. If it has converged, draw the trajectory of the UAV position at each moment recorded in step 6), and at the same time output the power of the UAV, flight speed, and base station transmission power in each time slot. If it has not converged, go to step 2) until convergence.

[0016] Furthermore, the establishment of the UAV cooperative NOMA communication network energy efficiency maximization model in step 1) is as follows:

[0017]

[0018] s.t.

[0019] C1a:

[0020] C1b:

[0021] C1c:

[0022] C1d:

[0023] C1e: 0 ≤ α m ≤ 1

[0024] C1f:

[0025] C1g:

[0026] C1h:

[0027] C1i:

[0028] C1j:

[0029]

[0030] C1k:

[0031] C1l:q uav [1]=q0

[0032] C1m:q uav [N]=q F

[0033] The optimization variable of this optimization problem is P bs Base station transmit power, P uav UAV transmit power, q uav The position of the UAV, V represents the flight speed of the UAV, α represents the power allocation coefficient vector; the considered system includes M users, and the flight period of the UAV is N, and the maximum flight speed of the UAV is V max , let t represent each individual moment, η EE represents the system energy efficiency at the current moment t and represent the maximum transmit powers of the base station and the UAV respectively, α m , a k represent the power allocation coefficients of users m and k respectively, A uav represents the flight altitude of the UAV represents the channel power gain from the base station to the UAV at time t represents the channel power gain from the UAV to user m at time t, φ k,m is a binary decoding variable, indicating decoding the signal of user m at the end of user k. Similarly, φ m,k and φ m,m represent decoding the signal of user k at the end of user m and decoding the signal of user m itself respectively. Its value is 0 or 1. When its value is 1, it means that user m needs to regard the signal of user k as interference when decoding its own signal. If it is 0, it means that user m can decode the signal of user k first before decoding its own signal and does not need to consider the interference brought by it when decoding itself. κ bs,uav , κ uav,m represent the overall hardware damage level between the base station and the UAV and the overall hardware damage level between the UAV and user m respectively, σ 2 represents the additive white Gaussian noise power, P com represents the power consumption of the UAV for communication, P0, P i represent the profile power and induction power of the UAV's rotor in the hovering state respectively, U tip represents the speed of the UAV's rotor, v0 represents the average rotational axis induction vector of the UAV in hovering, d0 represents the fuselage drag ratio, ρ represents the air density, s represents the rotational axis firmness, A represents the rotor rotation area, q0 represents the initial position of the UAV, q F represents the final position of the UAV

[0034] Among them, the constraint C1a is that the distance the UAV flies in each time slot does not exceed the constraint of flying in a straight line at the maximum speed. δ is the length of each time slot, in seconds. The constraint C1b is the base station transmission power constraint, the constraint C1c is the UAV transmission power constraint, the constraint C1d is the constraint on the sum of the power allocation coefficients, indicating that the sum is 1. The constraint C1e is that each individual constraint is between 0 and 1. C1f is the rate constraint limit for each user. C1g is the constraint that user m can decode its own information. C1h is the constraint that for any two users k and m, one of them must be able to decode the other. C1i is the value constraint of the decoding variable. C1j is the constraint that when the distance from user k to the UAV is farther than the distance from user m to the UAV, it can be decoded by user m. C1k represents the constraint that the distance the UAV flies in each time slot is equal to the product of the time slot length and the speed in that time slot. C1l represents the constraint on the starting position of the UAV. C1m represents the constraint on the ending position of the UAV.

[0035] Furthermore, in step 2), initialize the starting and ending positions of the UAV as and The UAV transmission power P uav , and the UAV adopts the amplify-and-forward protocol, and determine the decoding order variable as φ k,m ∈{0,1}, When φ k,m =1, it means that user k has a stronger channel power gain. When decoding user m, it can be regarded as noise. In other cases, φ k,m =0. Determine the base station position q bs =[x bs ,y bs T and its transmission power P bs , set the user minimum rate threshold as R min =1bps / Hz. The base station power allocation coefficient is defined as α={α1,α2,...,α M}, which satisfies and each term is non - negative. There are a total of M users in the system. Define the flight period of the UAV T = Nδ, where T is the entire flight period of the UAV and δ is the length of each time slot. One period is divided into N time slots.

[0036] Furthermore, in step 3), set the range of the UAV flight area in the deep Q - learning algorithm, define the state space s t as Define the action space a t as {P bs ,P uav ,V directi o n ,V}, where V​direction Define the reward function reward as the flight direction Use the drone as an agent in the environment.

[0037] Furthermore, in step 4), using the given drone position, its power, flight speed, and base station transmission power, introduce an approximation variable γ to the objective function in P1 m , then the power allocation sub-problem can be rewritten as follows:

[0038]

[0039] s.t.

[0040] C1a - C1m

[0041] C2a:

[0042] It can be found that the approximation variable γ m approximates the signal-to-interference-plus-noise ratio part in the objective function, so a new constraint C2a is introduced, which means the upper bound of the introduced approximation variable is the original signal-to-interference-plus-noise ratio at the current time t; for problem P2, use the Lagrangian dual method to handle the constraint by putting it into the objective function, and problem P2 can be rewritten as follows:

[0043] (P3)

[0044] s.t.

[0045] C1a - C1m

[0046] μ m represents the Lagrange multiplier of the m-th term. Take the derivative of the objective function of problem P3 with respect to μ m and set it equal to 0, we can get and substitute it into each term of the objective function of P3, then problem P3 can be rewritten as follows:

[0047]

[0048] s.t.

[0049] C1a - C1m

[0050] where there is

[0051]

[0052] Similarly, for problem P4, take the derivative with respect to γ m and set it equal to 0, we can get and substitute it into each term of the objective function of P4, then problem P4 can be rewritten as follows:

[0053]

[0054] s.t.

[0055] C1a - C1m

[0056] where For any m - th item, there is:

[0057]

[0058] Since the first term in the product term of the objective function f(α, γ, t, V) in P4 does not contain the power allocation coefficient variable α or the approximation variable γ here, it does not affect the solution of the optimal solution. In the objective function of problem P5, it can be removed and rewritten as follows:

[0059]

[0060] There is a fractional - sum part in the objective function of problem P5. By performing a quadratic transformation on this fractional - sum part, the power allocation coefficient optimization problem is finally transformed into a standard convex optimization problem, as shown below:

[0061]

[0062] s.t.

[0063] C1a - C1m

[0064] where

[0065]

[0066]

[0067] At this time, the power allocation optimization problem P6 is already a standard convex optimization problem. The interior - point method is used for iterative solution until the previous solution value is the same as the current solution value, which means convergence. Then the power allocation coefficient vector α is updated. Otherwise, continue to iteratively solve problem P6 until convergence.

[0068] Furthermore, in step 5), when the decoding order and the power allocation coefficient are fixed, the state s(t) of the UAV (agent) at this moment is {horizontal position coordinate, system energy efficiency}, and the action a(t) at this time is {base - station transmission power, UAV transmission power, UAV position, UAV flight speed, UAV flight direction}. According to the deep Q - learning algorithm, the Q - value is calculated by the following formula:

[0069] Q(s(t), a(t)) = E[r t + discount max Q(s(t + 1), a(t + 1))|s(t), a(t)]

[0070] where r t represents the reward obtained at the current moment, discount represents the discount factor, when it takes 0, it means paying more attention to the current reward, E[·] represents the operation of taking the expectation, and the Q value is updated.

[0071] Further, in step 6), the current state and action of the UAV input in step 6) are used as samples for learning. By making a greedy policy judgment on the current Q value, the action at the next moment is selected with a greedy probability. After that, the position of the UAV is updated and recorded, and the learned samples will be stored in the experience replay pool.

[0072] Further, in step 7), it is judged whether the flight cycle of the UAV has ended. If it has not ended, it will turn to step 4). If it has ended, it will first judge whether the average reward of the UAV flight cycle has converged. If it has converged, the UAV trajectory data saved in step 6) will be plotted, and at the same time, the UAV power, base station power, flight speed and direction of each time slot will be output. If it has not converged, it will turn to step 2) until the average reward in the last cycle converges.

[0073] The advantages and beneficial effects of the present invention are as follows:

[0074] Under the constraints of the UAV flight trajectory, transmission power, decoding order, base station transmission power, power distribution coefficient, and user service quality, the present invention maximizes the energy efficiency of the UAV cooperative NOMA communication network.

[0075] 1. Compared with traditional communication scenarios, the present invention considers the negative impact of hardware non-ideal factors in communication on communication performance. For practical considerations, the present invention also considers that the signal source node, UAV node, and user node often have different levels of noise intensity and hardware damage levels. It is worth noting that the present invention also considers the height of the base station tower, which is more in line with the actual scenario. In addition, different from solving the optimal position deployment of the UAV, the present invention considers solving the UAV trajectory during the flight cycle. When solving the energy efficiency, the transmission powers of the base station and the UAV are optimized considering power consumption. Compared with solving the system sum rate with the maximum transmission power, it has more practical value.

[0076] 2. Compared with the previous solutions for wireless resource management, the present invention applies the method of reinforcement learning. In a dynamic environment considering the change of UAV position, whenever the environment changes, the traditional convex optimization method needs to recalculate the resource allocation scheme. When the frequency of environmental change increases, this undoubtedly causes a problem of waste of a large amount of computing resources. In addition, the change of UAV position will cause the change of the direct channel power gain between the UAV and the user, so the decoding order will change accordingly. In the present invention, no assumption about the user distribution is made, but the change of the decoding order is considered. Compared with the assumptions made in previous inventions or researches, the present invention is more reasonable in this regard.

[0077] 3. For the continuous power allocation coefficient optimization problem P2, the present invention does not quantify it, but solves this continuous variable by the method of convex optimization, so as to avoid quantization error and inaccurate methods, thereby reducing the dimension of the action space in the deep Q-learning technology and achieving the effect of promoting the solution of each other. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] Figure 1 is the system model of the preferred embodiment provided by the present invention;

[0079] Figure 2 is the average reward convergence graph;

[0080] Figure 3 is the single-step reward convergence;

[0081] Figure 4 is the flight trajectory graph

[0082] Figure 5 is the flow schematic diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0083] Next, the technical solutions in the embodiments of the present invention will be clearly and detailedly described with reference to the accompanying drawings in the embodiments of the present invention. The described embodiments are only a part of the embodiments of the present invention.

[0084] The technical solution of the present invention to solve the above technical problems is:

[0085] The present invention Figure 5 discloses a method for maximizing the energy efficiency of a UAV cooperative NOMA communication network under hardware impairment conditions, which includes the following steps:

[0086] The first step: Establish a model for maximizing the energy efficiency of a UAV cooperative NOMA communication network under the condition of hardware impairment;

[0087] Step 2: Iteratively optimize and solve the energy efficiency maximization model, initialize the starting and ending positions of the UAV, its transmission power, decoding order variables, base station position and transmission power, user minimum rate threshold, base station power allocation coefficient, and flight period size;

[0088] Step 3: Set the parameters of the environment in the deep Q - learning algorithm. Define the state space as {horizontal coordinate of the UAV, system energy efficiency}, the action space as {base station transmission power, UAV transmission power, UAV flight direction, UAV flight speed}, and the reward as {system energy efficiency}, where the UAV acts as the agent;

[0089] Step 4: Solve the decoding order and power allocation coefficient based on the initial position and power of the UAV and the base station transmission power, and determine whether the power allocation coefficient converges. If it has converged, update the power allocation coefficient; otherwise, repeat this step until the result converges;

[0090] Step 5: Input the horizontal position coordinate, decoding order, flight direction, and speed of the UAV at the current moment, the base station transmission power, and the power allocation coefficient after convergence in Step 4 into the Deep Neural Networks, and update the Q - value in the deep Q - learning algorithm;

[0091] Step 6: Through learning from samples, the deep neural network will output the reward obtained from the environment at the current moment, change the agent state to the next moment and update the base station transmission power. At the same time, store the learned samples in the experience replay pool and save the output parameters;

[0092] Step 7: Determine whether the UAV flight period has ended. If it has not ended, go to Step 4; if it has ended, first determine whether the average reward of the UAV flight period has converged. If it has converged, plot the trajectory of the UAV position at each moment recorded in Step 6, and at the same time output the power and flight speed of the UAV in each time slot and the base station transmission power. If it has not converged, go to Step 2 until convergence.

[0093] Furthermore, in the energy efficiency maximization method for a UAV - cooperative NOMA communication network described in Step 1, it is characterized in that the energy efficiency maximization model of the UAV cooperative NOMA communication network established in Step 1 is:

[0094]

[0095] C1a:

[0096] C1b:

[0097] C1c:

[0098] C1d:

[0099] C1e: 0 ≤ α m ≤ 1

[0100] C1f:

[0101] C1g:

[0102] C1h:

[0103] C1i:

[0104] C1j:

[0105]

[0106] C1k:

[0107] C1l: q uav [1] = q0

[0108] C1m: q uav [N] = q F

[0109] The optimization variables of this optimization problem are P bs the base station transmission power, P uav the UAV transmission power, q uav the position of the UAV, V represents the flight speed of the UAV, and α represents the power allocation coefficient vector. The considered system includes M users, and the flight period of the UAV is N, and the maximum flight speed of the UAV is V max , let t represent each individual moment, η EE represents the system energy efficiency at the current moment t, and respectively represent the maximum transmission powers of the base station and the UAV, α m , a k respectively represent the power allocation coefficients of user m and user k, A uav represents the flight altitude of the UAV, represents the channel power gain from the base station to the UAV at time t, represents the channel power gain from the UAV to user m at time t, φ k,m is a binary decoding variable, indicating decoding the signal of user m at the end of user k. Similarly, φ m,k and φ m,mThey respectively represent decoding the signal of user k at the m - th user side and the m - th user decoding its own signal. Their values are 0 or 1. When the value is 1, it means that the m - th user needs to regard the signal of user k as interference when decoding its own signal. If it is 0, it means that the m - th user can decode the signal of user k first before decoding its own signal and does not need to consider the interference it brings during its own decoding, κ bs,uav and κ uav,m They respectively represent the overall hardware impairment level from the base station to the UAV and the overall hardware impairment level from the UAV to the m - th user, σ 2 represents the additive white Gaussian noise power, P com represents the power consumption of the UAV for communication, P0, P i They respectively represent the profile power and the induced power of the UAV's rotor in the hovering state, U tip represents the speed of the UAV's rotor, v0 represents the average rotational axis induction vector of the UAV in hovering, d0 represents the fuselage drag ratio, ρ represents the air density, s represents the shaft firmness, A represents the rotor rotation area, q0 represents the initial position of the UAV, q F represents the final position of the UAV;

[0110] Among them, the constraint C1a is that the distance the UAV flies in each time slot does not exceed the constraint of flying in a straight line at the maximum speed. δ is the length of each time slot, with the unit of seconds. The constraint C1b is the base station transmission power constraint. The constraint C1c is the UAV transmission power constraint. The constraint C1d is the constraint on the sum of the power allocation coefficients, indicating that the sum is 1. The constraint C1e is that each single constraint is between 0 and 1. C1f is the rate constraint limit for each user. C1g is the constraint that the m - th user can decode its own information. C1h is that for any two users k, m, one of them must be able to decode the other. The constraint C1i is the value constraint of the decoding variable. The constraint C1j is that when the distance from user k to the UAV is farther than the distance from user m to the UAV, the constraint that it can be decoded by user m. The constraint C1k represents that the distance the UAV flies in each time slot is equal to the product of the time slot length and the speed in that time slot. The constraint C1l represents the constraint on the starting position of the UAV. The constraint C1m represents the constraint on the ending position of the UAV.

[0111] Furthermore, in the method for maximizing the energy efficiency of a UAV - assisted NOMA communication network described in the second step, it is characterized in that, in the second step, the initial and ending positions of the UAV are initialized as and The UAV transmission power P uav , and the UAV adopts the amplify - and - forward protocol, and the decoding order variable is determined as φ k,m ∈{0,1}, When φ k,mWhen \(k = 1\), it means that user \(k\) has a stronger channel power gain and can be regarded as noise when decoding user \(m\). In other cases, \(\varphi\) k,m \(= 0\) to determine the base station location \(q\) bs \(= [x\) bs , y\) bs \) T and its transmission power \(P\) bs Set the minimum rate threshold of the user to \(R\) min \(= 1\) bps / Hz, and the base station power allocation coefficient is defined as \(\alpha=\{\alpha_1,\alpha_2,...,\alpha\) M \}\), which satisfies and each term is non - negative. Define the flight period \(T = N\delta\) of the UAV, where \(T\) is the entire flight period of the UAV and \(\delta\) is the length of each time slot. One period is divided into \(N\) time slots.

[0112] Furthermore, in a method for maximizing the energy efficiency of a UAV - assisted NOMA communication network described in the third step, it is characterized in that, in the third step, set the flight area range of the UAV in the deep Q - learning algorithm and define the state space \(s\) t as Define the action space \(a\) t as \(\{P\) bs , P\) uav , V\) directi _o\) n , V\}, where \(V\) direction is the flight direction, and define the reward function reward as Use the UAV as an agent in the environment.

[0113] Furthermore, in a method for maximizing the energy efficiency of a UAV - assisted NOMA communication network described in the fourth step, it is characterized in that, in the fourth step, using the given UAV position, its power, flight speed, and base station transmission power, introduce an approximation variable \(\gamma\) to the objective function in \(P1\) m , then the power allocation sub - problem can be rewritten as follows:

[0114]

[0115] \(C1_a - C1_m\)

[0116]

[0117] It can be found that the approximation variable \(\gamma\) m approximates the signal - to - interference - plus - noise ratio part in the objective function, so a new constraint \(C2_a\) is introduced, which means that the upper bound of the introduced approximation variable is the original signal - to - interference - plus - noise ratio at the current time \(t\). Using the Lagrangian dual method for problem \(P2\) and putting the constraint into the objective function for processing, problem \(P2\) can be rewritten as follows:

[0118]

[0119] such that

[0120] C1a - C1m

[0121] μ m denotes the Lagrange multiplier of the m-th term. Taking the derivative of the objective function of problem P3 with respect to μ m and setting it equal to 0, we obtain After substituting it into each term of the objective function of P3, problem P3 can be rewritten as follows:

[0122]

[0123] such that

[0124] C1a - C1m

[0125] wherein there is

[0126]

[0127] Similarly, taking the derivative of problem P4 with respect to γ m and setting it equal to 0, we obtain After substituting it into each term of the objective function of P4, problem P4 can be rewritten as follows:

[0128]

[0129] such that

[0130] C1a - C1m

[0131] wherein For any m-th term, there is:

[0132]

[0133] Since the first term in the product term of the objective function f(α, γ, t, V) in P4 does not contain the power allocation coefficient variable α or the approximation variable γ here, it does not affect the calculation of the optimal solution. In the objective function of problem P5, it can be removed and rewritten as follows:

[0134]

[0135] In the objective function of problem P5, there is a fractional sum part. By performing a quadratic transformation on this fractional sum part, the power allocation coefficient optimization problem is finally transformed into a standard convex optimization problem as follows:

[0136]

[0137] such that

[0138] C1a - C1m

[0139] wherein

[0140]

[0141]

[0142] At this time, the power allocation optimization problem P6 is already a standard convex optimization problem. The interior point method is used for iterative solution until convergence occurs when the previous solution value is the same as the current solution value. The power allocation coefficient vector α is updated. Otherwise, the problem P6 is continuously iteratively solved until convergence.

[0143] Furthermore, in a method for maximizing the energy efficiency of a UAV cooperative NOMA communication network described in the fifth step, it is characterized in that, in the fifth step, when the decoding order and the power allocation coefficient are fixed, the current state s(t) of the UAV (agent) is {horizontal position coordinates, system energy efficiency}, and the current action a(t) is {base station transmission power, UAV transmission power, UAV position, UAV flight speed, UAV flight direction}. According to the deep Q - learning algorithm, the Q - value is calculated by the following formula:

[0144] Q(s(t), a(t)) = E[r t + discount max Q(s(t + 1), a(t + 1))|s(t), a(t)]

[0145] where r t represents the reward obtained at the current moment, discount represents the discount factor. When it takes 0, it means more attention is paid to the current reward, and E[·] represents the expectation operation. Update the Q - value.

[0146] Furthermore, in a method for maximizing the energy efficiency of a UAV cooperative NOMA communication network described in the sixth step, it is characterized in that, in the sixth step, the current state and action of the UAV input in the sixth step are used as samples for learning. By making a greedy policy judgment on the current Q - value, the next - moment action is selected with a greedy probability. Then the position of the UAV is updated and recorded. At the same time, the learned samples will be stored in the experience replay pool.

[0147] Further, in the method for maximizing the energy efficiency of a drone cooperative NOMA communication network described in the seventh step, it is characterized in that, in the seventh step, it is judged whether the flight cycle of the drone has ended. If it has not ended, it is transferred to the fourth step. If it has ended, it is first judged whether the average reward of the drone flight cycle has converged. If it has converged, the drone trajectory data saved in the sixth step is plotted, and at the same time, the drone power, base station power, flight speed, and direction of each time slot are output. If it has not converged, it is transferred to the second step until the average reward in the last cycle converges.

[0148] Considering the constraints such as the power consumption of drones and base stations, decoding order, the worst quality of service of users, and drone flight trajectories, the present invention maximizes the energy efficiency of the drone cooperative NOMA communication network. Its innovation lies in that, firstly, compared with the research in previous scenarios, it considers the actual factors that the transceiver hardware of communication devices is not ideal and suffers from inconsistent hardware damage degrees, and the noises suffered by each communication node are not equal. It should be noted that the present invention also considers the height of the base station tower and studies the energy efficiency problem of the communication network, taking into account the transmission powers of the base station and the drone, and not fixing them to the maximum transmission power, which has more practical application value than the sum rate. Secondly, the present invention adopts the method of reinforcement learning. Compared with traditional convex optimization means, there is no need to redesign the resource allocation scheme for slight changes in the actual scenario, and it provides an offline learning strategy, which is easy to transplant and has better practicability and feasibility.

[0149] This embodiment is a method for maximizing the energy efficiency of a drone cooperative NOMA communication network. In a system where UAVs cooperate with NOMA for communication, there are 6 ground users in total. Since there are obstacles or severe shadow effects between the base station and the users, there is no direct link between the base station and the users. The height of the base station tower is 25m, the maximum flight height of the drone is 140m, the coordinates of the base station are (0,0), the horizontal flight range of the drone and the positions of the ground users are in a 300×300 square meter space of a horizontal Cartesian coordinate [0,300] and vertical [0,300], and the maximum transmission power of the base station is 10W, and the maximum transmission power of the drone is 1W, the channel power gain per unit distance β0 = -60dB, the noise powers at the drone side and the user side are 1×10 -12 W and 1×10 -11 W respectively, and the average hardware damage degree of the link from the base station to the drone is κ bs,uav = 0.01, and the average hardware damage degree of the link from the drone to the user side is κ uav,m = 0.02, and the minimum quality of service threshold R of the user min= 0.1 bps / Hz. The entire flight period T is divided into 60 time slots, and the length of each time slot δ is 1 second.

[0150] In this embodiment, Figure 1 is the system model in the UAV cooperative NOMA communication network provided by the present invention. In the figure, the UAV acts as an aerial relay to forward the signal of the base station to the ground users, and all communication nodes are equipped with single antennas; Figure 2 is the convergence graph of the average reward obtained by the UAV as an agent in each round of iteration; Figure 3 is the convergence graph of the reward obtained by the UAV as an agent in a single step in each iteration. Figure 4 is the flight trajectory graph of the UAV.

[0151] From Figure 2 it can be seen that in the method of this embodiment, the UAV as an agent does not perform very well in the initial stage. However, as the learning experience increases, the single-step reward that the proposed algorithm can obtain becomes higher and higher. After 1000 steps, as the model is trained, the UAV can almost obtain the maximum reward in each step, verifying the feasibility of the algorithm proposed by the present invention.

[0152] From Figure 3 it can be seen that the method of this embodiment not only converges in obtaining rewards in a single step, but also the cumulative average reward within each flight period of the specified UAV can reach convergence. This means that within the specified period, the UAV can obtain higher rewards for the flight trajectory in each time slot, and the setting of the reward is positively correlated with the system energy efficiency, verifying the effectiveness of the proposed algorithm.

[0153] From Figure 4 it can be seen that in the method of this embodiment, the UAV can not only finally reach the specified end range, but also serve each user in the middle process to ensure the minimum quality of service of the user, verifying the feasibility of the proposed algorithm.

[0154] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0155] It should also be noted that the terms "comprising", "including" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.

[0156] The above embodiments should be understood as being only for illustrative purposes of the present invention and not for limiting the protection scope of the present invention. After reading the content described in the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.

Claims

1. An energy efficiency maximization method for an unmanned aerial vehicle cooperative NOMA communication network, characterized in that, It includes the following steps: Step 1): Establish an energy efficiency maximization model for a UAV cooperative NOMA communication network under the condition of hardware damage; Step 2): Iteratively optimize and solve the energy efficiency maximization model, initialize the starting and ending positions of the UAV, its transmission power, decoding order variables, base station position and transmission power, user minimum rate threshold, base station power allocation coefficient, and flight cycle size; Step 3): Set the parameters of the environment in the deep Q-learning algorithm, define the state space as {horizontal coordinate of the UAV, system energy efficiency}, define the action space as {base station transmission power, UAV transmission power, UAV flight direction, UAV flight speed}, and define the reward as {system energy efficiency}, where the UAV acts as an agent; Step 4): Solve the decoding order and power allocation coefficient according to the initial position and power of the UAV and the base station transmission power, and determine whether the power allocation coefficient converges. If it has converged, update the power allocation coefficient; otherwise, repeat this step until the result converges; Step 5): Input the horizontal position coordinate, decoding order, flight direction, and speed of the UAV at the current moment, the base station transmission power, and the power allocation coefficient after convergence in Step 4) into the deep neural network, and update the Q value in the deep Q-learning algorithm; Step 6): Through learning the samples, the deep neural network outputs the reward obtained from the environment at the current moment, changes the agent state to the next moment and updates the base station transmission power, and at the same time stores the learned samples in the experience replay pool and saves the output parameters; Step 7): Determine whether the UAV flight cycle has ended. If it has not ended, go to Step 4). If it has ended, first determine whether the average reward of the UAV flight cycle has converged. If it has converged, plot the trajectory of the UAV position at each moment recorded in Step 6), and at the same time output the power and flight speed of the UAV in each time slot and the base station transmission power. If it has not converged, go to Step 2) until convergence; The energy efficiency maximization model for the UAV cooperative NOMA communication network established in Step 1) is: s.t. C1e: 0 ≤ α m ≤ 1 C1l:q uav [1] = q0 C1m:q uav [N] = q F The optimization variable of this optimization problem is P bs Base station transmit power, P uav UAV transmit power, q uav The position of the UAV, V represents the flight speed of the UAV, α represents the power allocation coefficient vector; the considered system includes M users, and the flight period of the UAV is N, and the maximum flight speed of the UAV is V max , let t represent each individual moment, η EE represents the system energy efficiency at the current moment t, and respectively represent the maximum transmit powers of the base station and the UAV, α m , a k respectively represent the power allocation coefficients of user m and k, A uav represents the flight altitude of the UAV, represents the channel power gain from the base station to the UAV at time t, represents the channel power gain from the UAV to user m at time t, φ k,m is a binary decoding variable, indicating decoding the signal of user m at the end of user k. Similarly, φ m,k and φ m,m respectively represent decoding the signal of user k at the end of user m and user m decoding its own signal. Its value is 0 or 1. When its value is 1, it means that user m needs to regard the signal of user k as interference when decoding its own signal. If it is 0, it means that user m can decode the signal of user k first before decoding its own signal and does not need to consider the interference brought by it when decoding itself. κ bs,uav , κ uav,m respectively represent the overall hardware damage level between the base station and the UAV and the overall hardware damage level between the UAV and user m. σ 2 represents the additive white Gaussian noise power, P com represents the power consumption of the UAV for communication, P0, P i respectively represent the profile power and induction power of the rotor of the UAV in the hovering state, U tip represents the speed of the UAV rotor, v0 represents the average rotational axis induction vector of the UAV in hovering, d0 represents the fuselage drag ratio, ρ represents the air density, s represents the shaft firmness, A represents the rotor rotation area, q0 represents the initial position of the UAV, q F represents the final position of the UAV; Among them, constraint C1a is the constraint that the distance the UAV flies in each time slot does not exceed the constraint of flying in a straight line at the maximum speed. δ is the length of each time slot, with the unit of second. Constraint C1b is the base station transmission power constraint. Constraint C1c is the UAV transmission power constraint. Constraint C1d is the constraint on the sum of the power allocation coefficients, indicating that the sum is 1. Constraint C1e is that each individual constraint is between 0 and 1. C1f is the rate constraint limit for each user. C1g is the constraint that user m can decode its own information. C1h is that for any two users k, m, one of them must be able to decode the other. Constraint C1i is the value constraint of the decoding variable. Constraint C1j is the constraint that when the distance from user k to the UAV is farther than the distance from user m to the UAV, it can be decoded by user m. Constraint C1k represents the constraint that the distance the UAV flies in each time slot is equal to the product of the time slot length and the time slot speed. Constraint C1l represents the constraint on the starting position of the UAV. Constraint C1m represents the constraint on the ending position of the UAV; In step 4), an approximate variable γ is introduced into the objective function in P1 by using the given UAV position, its power, flight speed, and base station transmission power. m , then the power allocation sub-problem can be rewritten as follows: s.t. C1a - C1m The approximate variable γ can be found m The dry - to - interference ratio part in the objective function is approximated, and thus a new constraint C2a is introduced, which means that the upper bound of the introduced approximate variable is the original signal - to - interference - plus - noise ratio at the current time t; for problem P2, the Lagrangian dual method is used to incorporate the constraint into the objective function for processing, and problem P2 can be rewritten as follows: s.t. C1a - C1m μ m denotes the Lagrange multiplier of the \(m\)th term. Taking the derivative of the objective function of problem P3 with respect to \(\mu\) m and setting it equal to 0, we can obtain After substituting it into each term of the objective function of P3, problem P3 can be rewritten as follows: s.t. C1a - C1m wherein there is Similarly, make problem P4 for γ m Take the derivative with respect to it and set it equal to 0, we can get After substituting each term of the objective function of P4, problem P4 can be rewritten as follows: s.t. C1a - C1m Among them For any m-th item, there is Since the first term in the product term of the objective function f(α, γ, t, V) in P4 does not contain the power allocation coefficient variable α or the approximation variable γ here, it does not affect the solution of the optimal solution. In the objective function of problem P5, it can be removed and rewritten as follows: In the objective function of problem P5, there is a fractional sum part. By using quadratic transformation for this fractional sum part, the power allocation coefficient optimization problem is finally transformed into a standard convex optimization problem as follows: s.t. C1a - C1m where At this time, the power allocation optimization problem P6 is already a standard convex optimization problem. The interior point method is used for iterative solution until the previous solution value is the same as the current solution value, which means convergence. The power allocation coefficient vector α is updated. Otherwise, the problem P6 is continuously iteratively solved until convergence.

2. The energy efficiency maximization method for an unmanned aerial vehicle cooperative NOMA communication network according to claim 1, wherein In step 2), initialize the initial and end positions of the UAV as and the transmission power P of the UAV uav , and the UAV adopts the amplify-and-forward protocol, and determine the decoding order variable as φ k,m ∈{0,1}, When φ k,m =1, it means that user k has a stronger channel power gain. When decoding user m, it can be regarded as noise. In other cases, φ k,m =0, determine the base station position q bs =[x bs ,y bs T and its transmission power P bs , set the minimum rate threshold of the user as R min =1bps / Hz, the base station power allocation coefficient is defined as α={α1,α2,...,α M} and it satisfies and each item is non-negative. There are M users in the system in total; define the flight period T of the UAV as T = Nδ, where T is the entire flight period of the UAV, δ is the length of each time slot, and one period is divided into N time slots.​ 3. The energy efficiency maximization method for a UAV cooperative NOMA communication network according to claim 2, wherein In step 3), set the range of the UAV flight area in the deep Q - learning algorithm and define the state space s t as define the action space a t as {P bs , P uav , V direction , V}, where V direction is the flight direction, and define the reward function reward as Use the UAV as an agent in the environment.

4. The energy efficiency maximization method of a UAV cooperative NOMA communication network according to claim 3, characterized in that, In step 5), when the decoding order and the power allocation coefficient are fixed, the current state s(t) of the UAV agent is {horizontal position coordinate, system energy efficiency}, and the action a(t) at this time is {base station transmission power, UAV transmission power, UAV position, UAV flight speed, UAV flight direction}. According to the deep Q - learning algorithm, the Q - value is calculated by the following formula: Q(s(t), a(t)) = E[r t + discount max Q(s(t + 1), a(t + 1)) | s(t), a(t)] where r t represents the reward obtained at the current moment, discount represents the discount factor. When it takes 0, it means paying more attention to the current reward. E[·] represents the expectation operation to update the Q value.

5. The energy efficiency maximization method for a UAV cooperative NOMA communication network according to claim 4, characterized in that, In step 6), the current state and action of the UAV input in step 6) are used as samples for learning. By making a greedy policy judgment on the current Q - value, the next - moment action is selected with a greedy probability. Then the position of the UAV is updated and recorded. At the same time, the learned samples will be stored in the experience replay pool.

6. The energy efficiency maximization method of a drone cooperative NOMA communication network according to claim 4, wherein In step 7), it is judged whether the flight period of the UAV has ended. If it has not ended, it turns to step 4). If it has ended, first judge whether the average reward of the UAV flight period has converged. If it has converged, the UAV trajectory data saved in step 6) is plotted, and at the same time, the UAV power, base station power, flight speed, and direction of each time slot are output. If it has not converged, it turns to step 2) until the average reward in the last period converges.

Citation Information

Patent Citations

  • Resource allocation method based on max-min fairness for unmanned aerial vehicle assisted backscatter communication system

    CN114124705A

  • Method of discrete digital signal recovery in noisy overloaded wireless communication systems in the presence of hardware impairments

    WO2021198406A1