A Resource Optimization Method for UAV Spectrum Sharing Networks
By constructing the input state information variables of agents and cognitive drones in the drone spectrum sharing network, and optimizing resource allocation using AC algorithm and DDPG algorithm, the problems of low resource allocation efficiency and high data dependence in the existing technology are solved, and more efficient data utilization and computing efficiency are achieved.
Patent Information
- Application Number
- CN202410040985.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-11
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2044-01-11
AI Technical Summary
The existing UAV communication resource allocation methods have problems such as low data efficiency, susceptibility to environmental changes, limited exploration capabilities and high dependence on data storage, and it is difficult to effectively deploy in actual systems.
A resource optimization method for the spectrum sharing network of drone is proposed. By constructing the input state information variables of sub-band agents and cognitive drones, the transmission rate of the primary and secondary users is determined, and the reward function is constructed. The AC algorithm and DDPG algorithm are used to select and update the action network and optimize resource allocation.
It significantly improves data utilization efficiency and computing efficiency, reduces dependence on data storage, can cope with non-convex problems, realizes automatic parameter updates, has stronger global search capabilities, and is suitable for drone spectrum sharing scenarios with anti-interference technology.
Smart Images

Figure CN117880821B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wireless communication, and particularly relates to a resource optimization method for an unmanned aerial vehicle (UAV) spectrum sharing network. Background Art
[0002] The low-altitude airspace is a national strategic resource and has become a new engine for national development. UAV communication is the core to support the full utilization of the low-altitude airspace and the rapid development of the low-altitude economy. Due to its advantages such as flexible deployment, high cost-effectiveness, and rapid deployment, it is widely used in various low-altitude practical applications, such as relay communication, emergency rescue, and network traffic control. However, most UAV communications occur in unlicensed spectrums, making UAVs vulnerable to security threats from interference sources; moreover, with the rapid increase in frequency-using devices in the ground frequency bands, the spectrum resources in the unlicensed bands are becoming increasingly scarce, making the problem of scarce UAV communication spectrum resources more and more serious; spectrum sharing enables secondary users to opportunistically access the primary user spectrum and is expected to alleviate the spectrum scarcity problem; however, the heterogeneity and high dynamic variability of the wireless communication environment pose great challenges to efficient and rapid spectrum allocation and UAV trajectory optimization in spectrum sharing.
[0003] The existing UAV communication resource allocation methods can be mainly divided into two categories, the resource allocation method based on online learning and the resource allocation method based on offline learning.
[0004] However, both the online learning and offline learning methods have limitations and disadvantages. The online learning method has low data efficiency, is easily affected by the strong correlation of data samples, has limited exploration ability, and is prone to falling into local optimum when the environmental state changes continuously. Although the experience replay method can break the data correlation and improve the data efficiency, the offline learning method depends on a large amount of training data and is difficult to be deployed in an actual system when the data collection is time-consuming and costly. Therefore, in order to address the above problems, it is urgent to develop a new resource allocation framework. Summary of the Invention
[0005] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a resource optimization method for an unmanned aerial vehicle spectrum sharing network, which improves the data utilization efficiency and computing efficiency, and at the same time reduces the dependence on data storage.
[0006] The present invention provides the following technical solutions:
[0007] In a first aspect, a resource optimization method for an unmanned aerial vehicle spectrum sharing network is provided, including:
[0008] Constructing the input state information variables of the sub-band agent and the input state information variables of the cognitive UAV;
[0009] Determining the transmission rates of the primary user and the secondary user, and constructing a reward function;
[0010] Initialize the number of rounds and steps of the training round. During training, based on the input state information of the current sub-band agent, use the action network of the AC algorithm to select the action output of the sub-band agent and determine the immediate reward of the sub-band agent; based on the input state information of the current cognitive UAV, use the action network of the DDPG algorithm to select the action output of the cognitive UAV and determine the immediate reward of the cognitive UAV.
[0011] When the number of steps in the training round reaches the total number of steps in this training round, update the network parameters of the AC algorithm and the network parameters of the DDPG algorithm.
[0012] According to the update results of the network parameters of the AC algorithm and the DDPG algorithm, continuously train until the set number of rounds in the training round is satisfied, output the optimized model, and use the output optimized model for resource optimization.
[0013] Furthermore, in the construction of the input state information variable of the sub-band agent and the input state information variable of the cognitive UAV:
[0014] The input state information variable of the sub-band agent is:
[0015]
[0016] The input state information variable of the cognitive UAV is:
[0017]
[0018] Among them, is the input state information of the k-th sub-band agent at the n-th time step, represents the frequency band allocation strategy of the k-th sub-band agent at the (n - 1)-th time step; I n-1,k is the interference received by the k-th secondary user in the current sub-band at the (n - 1)-th moment, H n,k is the channel power gain; is the position of the sub-band agent at the (n - 1)-th time step; is the position of the cognitive UAV at the (n - 1)-th time step;
[0019] The channel power gain H n,k is:
[0020]
[0021] Among them, is the channel power gain from the cognitive UAV to the k-th secondary user, is the channel power gain from the interfering UAV to the k-th secondary user, is the channel power gain from the primary base station to k secondary users.
[0022] Furthermore, in the step of determining the transmission rates of the primary users and the secondary users and constructing the reward function,
[0023] the transmission rate of the primary user is:
[0024]
[0025] the transmission rate of the secondary user is:
[0026]
[0027] where B is the bandwidth of each sub - band, is the signal - to - noise ratio of the k - th secondary user at time slot n and sub - band m, is the signal - to - noise ratio of the j - th primary user at time slot n and sub - band m;
[0028] the signal - to - noise ratio of the j - th primary user at time slot n and sub - band m is:
[0029]
[0030] the signal - to - noise ratio of the k - th secondary user at time slot n and sub - band m is:
[0031]
[0032] where, is the transmit power of the primary base station on the m - th sub - band, is the channel power gain between the primary base station and the j - th primary user; is the noise power of the j - th primary user, ρ k,n [m] indicates whether the m - th sub - band is occupied, P c is the transmit power of the cognitive UAV, is the channel power gain from the cognitive UAV to the j - th primary user at time slot n; indicates whether the m - th sub - band is occupied by the interfering UAV, P J is the transmit power of the interfering UAV, is the channel power gain from the interfering UAV to the j - th primary user at time slot n;
[0033] is the channel power gain from the cognitive UAV to the k - th secondary user, is the noise power of the k - th secondary user; is the channel power gain from the primary base station to the k-th secondary user, is the channel power gain from the n-slot interference UAV to the k-th secondary user.
[0034] Further, in the step of determining the transmission rates of the primary users and the secondary users and constructing the reward function,
[0035] the reward function is:
[0036]
[0037] where, is the transmission rate of the k-th primary user, w 1 is the weight of the sum of the primary user transmission rates, δ j is the penalty term when the minimum transmission rate of the primary user is not satisfied, w 2 is the weight of the sum of the penalty terms.
[0038] The penalty term δ when the minimum transmission rate of the primary user is not satisfied j is:
[0039]
[0040] where, is the transmission rate of the j-th secondary user at the n-th time slot, R min is the minimum transmission rate requirement of the j-th primary user.
[0041] Further, in the step of updating the AC algorithm network parameters and the DDPG algorithm network parameters when the number of steps in the training episode reaches the total number of steps in the current training episode,
[0042] the update of the AC algorithm network parameters includes:
[0043] The evaluation network of the AC algorithm generates the current value of the AC algorithm;
[0044] According to the input state information of the current sub-band agent and the action output of the action network in the AC algorithm, determine the policy gradient of the action network of the AC algorithm;
[0045] According to the policy gradient of the action network of the AC algorithm and the immediate reward of the sub-band agent, determine the target value of the AC algorithm;
[0046] According to the current value and the target value of the AC algorithm, determine the TD error of the AC algorithm;
[0047] According to the policy gradient of the action network of the AC algorithm, use the gradient ascent strategy to update the action network parameters, and according to the TD error of the AC algorithm, use the gradient descent strategy to update the evaluation network parameters.
[0048] Furthermore, the policy gradient of the action network of the AC algorithm is as follows:
[0049]
[0050] where is the advantage function, is the action network of the AC algorithm; is the output of the action network in the AC algorithm, is the input state of the sub-band agent, and ξ is the action network parameter of the action network in the AC algorithm;
[0051] The target value y of the AC algorithm k is as follows:
[0052]
[0053] where r n+1 is the immediate reward of the sub-band agent, γ 1 is the discount factor of the AC algorithm, is the current value of the AC algorithm, and ω is the evaluation network parameter of the evaluation network in the AC algorithm;
[0054] The TD error of the AC algorithm is:
[0055]
[0056] Updating the action network parameter using the gradient ascent strategy is:
[0057]
[0058] where α a is the learning rate parameter of the action network;
[0059] Updating the evaluation network parameter using the gradient descent strategy is:
[0060]
[0061] where α c is the learning rate parameter of the evaluation network.
[0062] Furthermore, when the number of steps in the training episode reaches the total number of steps in this training episode, in updating the network parameters of the AC algorithm and the DDPG algorithm,
[0063] Updating the network parameters of the DDPG algorithm includes:
[0064] The target evaluation network of the DDPG algorithm outputs the current value;
[0065] Determine the deterministic policy gradient of the DDPG algorithm according to the status information of the cognitive UAV and the action output result of the action network in the DDPG algorithm;
[0066] Determine the target value of the DDPG algorithm according to the immediate reward of the cognitive UAV and the current value of the target evaluation network;
[0067] Determine the TD error of the DDPG algorithm according to the current value of the target evaluation network and the target value of the DDPG algorithm;
[0068] Update the network parameters of the DDPG evaluation network by using the gradient descent algorithm according to the TD error of the DDPG algorithm;
[0069] Update the network parameters of the DDPG target action network and the network parameters of the target evaluation network by using the soft update strategy according to the network parameters of the DDPG evaluation network.
[0070] Furthermore, the deterministic policy gradient of the DDPG algorithm is:
[0071]
[0072] where N tr is the sample batch size, is the input state of the cognitive UAV, is the input value of the action network in the DDPG algorithm; is the output of the evaluation network in the DDPG algorithm when the input value is , are the action network parameters in the DDPG algorithm, and λ is the network parameter of the action network in the DDPG algorithm;
[0073] The target value y of the DDPG algorithm i is:
[0074]
[0075] where r i+1 is the immediate reward of the cognitive UAV; γ 2 is the discount factor of the DDPG algorithm, is the output of the target evaluation network in the DDPG algorithm, and λ - is the network parameter of the target evaluation network, is the network parameter of the target action network;
[0076] The TD error of the DDPG algorithm is:
[0077]
[0078] where When the input value is the output of the evaluation network in the DDPG algorithm;
[0079] is the output value of the action network in the DDPG algorithm;
[0080]
[0081] where η is the noise variance obeying the Gaussian distribution;
[0082] The network parameter λ for updating the DDPG evaluation network by using the gradient descent algorithm is:
[0083]
[0084] where β c is the learning rate parameter of the evaluation network in the DDPG algorithm;
[0085] The network parameters for updating the DDPG target action network and the target evaluation network by using the soft update strategy are:
[0086]
[0087]
[0088] where τ is the frequency factor for controlling the soft update strategy.
[0089] In a second aspect, a computer device is provided, including a processor and a memory; wherein, when the processor executes the computer program stored in the memory, the steps of the resource optimization method for the UAV spectrum sharing network described in the first aspect are implemented.
[0090] In a third aspect, a computer-readable storage medium is provided for storing a computer program; when the computer program is executed by a processor, the steps of the resource optimization method for the UAV spectrum sharing network described in the first aspect are implemented.
[0091] Compared with the prior art, the beneficial effects of the present invention are:
[0092] (1) The method of the present invention makes full use of the unique features between heterogeneous optimization variables, proposes a heterogeneous resource allocation learning framework, significantly improves the data utilization efficiency of a single online learning method, and at the same time, compared with the offline learning method, reduces the dependence on data storage and improves the network computing efficiency.
[0093] (2) The method proposed by the present invention does not depend on specific environment modeling, can handle non-convex problems, can realize automatic parameter update at the same time, reduces the time cost, and the algorithm is not sensitive to the initial solution and has stronger global search ability.
[0094] (3) The present invention is applied to the UAV spectrum sharing scenario using anti-interference technology, verifying the effectiveness of the proposed method. Moreover, this framework has strong generalization ability and is not limited to specific external environments and application scenarios, providing reliable support for various wireless communication requirements. Description of the Drawings
[0095] Figure 1 is the flowchart of the resource optimization method for the UAV spectrum sharing network of the present invention;
[0096] Figure 2 is the algorithm framework diagram of the resource optimization method for the UAV spectrum sharing network of the present invention;
[0097] Figure 3 is the comparison diagram of the secondary network transmission rates between the method of the present invention and other existing technologies under different numbers of users;
[0098] Figure 4 is the comparison diagram of the secondary network transmission rates between the method of the present invention and other existing technologies under different UAV transmission powers;
[0099] Figure 5 is the comparison of the convergence performance of the method of the present invention under different numbers of hidden layers;
[0100] Figure 6 is the comparison diagram of the convergence performance of the method of the present invention under different experience pool capacities. Detailed Embodiments
[0101] The present invention will be further described below with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solution of the present invention and should not be used to limit the protection scope of the present invention.
[0102] Embodiment 1
[0103] Please refer to Figure 1 and 2 shown. A resource optimization method for a UAV spectrum sharing network includes the following steps:
[0104] S1: Calculate the air-to-ground channel power gain and the ground-to-ground channel power gain.
[0105] Specifically, the air-to-ground channel power gain includes: the channel power gains from the cognitive UAV and the interfering UAV to the secondary user, and the channel power gains from the cognitive UAV and the interfering UAV to the primary user; the ground-to-ground channel power gain includes: the channel power gain between the primary base station and the secondary user, and the channel power gain between the primary base station and the primary user.
[0106] Specifically, S1-1: Set the channel power gain β at unit distance ref .
[0107] S1-2: Calculate the distance d c,k between the cognitive unmanned aerial vehicle (C-UAV) and the k-th secondary user (SU) at the n-th time slot; and the distance d p,k between the cognitive unmanned aerial vehicle (C-UAV) and the j-th primary user (PU) at the n-th time slot.
[0108]
[0109]
[0110] where H u is the flight altitude of the cognitive unmanned aerial vehicle, q c [n] is the position of the cognitive unmanned aerial vehicle, w b [n] is the position of the ground base station, w s,k [n] is the position of the secondary user, w p,j [n] is the position of the primary user.
[0111] S1-3: Calculate the channel power gain between the cognitive unmanned aerial vehicle and the secondary user, the channel power gain between the cognitive unmanned aerial vehicle and the primary user, the channel power gain between the interfering unmanned aerial vehicle and the secondary user, and the channel power gain between the interfering unmanned aerial vehicle and the primary user according to the line-of-sight transmission link; calculate the channel power gain between the primary base station (PBS) and the primary user and the secondary user according to the small-scale Rayleigh fading channel.
[0112] Specifically, an example is given:
[0113] The channel power gain between the cognitive unmanned aerial vehicle and the secondary user is:
[0114]
[0115] The calculation methods of the channel power gain between the cognitive unmanned aerial vehicle and the primary user, the channel power gain between the interfering unmanned aerial vehicle and the secondary user, and the channel power gain between the interfering unmanned aerial vehicle and the primary user are the same as those of the channel power gain between the cognitive unmanned aerial vehicle and the secondary user.
[0116] Specifically, the channel power gain between the primary base station and the primary user is:
[0117]
[0118] where ζ j is a random variable subject to an exponential distribution, is a constant exponent;
[0119] The calculation method of the channel power gain between the primary base station and the secondary user is the same as that between the primary base station and the primary user.
[0120] S2: Calculate the transmission rates of the PU and the SU.
[0121] S2-1: Calculate the signal-to-noise ratio of the k-th secondary user in the m-th sub-band and the signal-to-noise ratio of the j-th primary user in the m-th sub-band.
[0122] Specifically, the signal-to-noise ratio of the k-th secondary user in the n-th time slot and the m-th sub-band is:
[0123]
[0124] The signal-to-noise ratio of the j-th primary user in the n-th time slot and the m-th sub-band is:
[0125]
[0126] Where, is the transmit power of the primary base station in the m-th sub-band, is the channel power gain between the primary base station and the j-th primary user; is the noise power of the j-th primary user, ρ k,n [m] indicates whether the m-th sub-band is occupied, P c is the transmit power of the cognitive UAV, is the channel power gain from the cognitive UAV to the j-th primary user in the n-th time slot; is whether the m-th sub-band is occupied by the interfering UAV, P J is the transmit power of the interfering UAV, is the channel power gain from the interfering UAV to the j-th primary user in the n-th time slot; is the channel power gain from the cognitive UAV to the k-th secondary user, is the noise power of the k-th secondary user; is the channel power gain between the primary base station and the k-th secondary user, is the channel power gain between the interfering UAV and the k-th secondary user in the n-th time slot.
[0127] Specifically, when ρ k,n [m] = 1, the m-th sub-band is occupied by the communication link between the k-th SUs. When ρ k,n [m] = 0, the m-th sub-band is not occupied by the communication link between the k-th SUs, is the interfering UAV spectrum occupancy variable. When , the m-th sub-band is occupied by the interfering UAV. When At this time, the m-th sub-band is not occupied by interfering drones.
[0128] S2-2: Calculate the transmission rate of the primary user and the transmission rate of the secondary user.
[0129] The transmission rate of the primary user is:
[0130]
[0131] The transmission rate of the secondary user is:
[0132]
[0133] where B is the bandwidth of each sub-band, is the signal-to-noise ratio of the k-th secondary user in the m-th sub-band at the n-th time slot, is the signal-to-noise ratio of the j-th primary user in the m-th sub-band at the n-th time slot;
[0134] S3: Construct the input state information variable of the sub-band agent and the input state information variable of the cognitive drone agent.
[0135] S3-1: The input state information variable of the sub-band agent is:
[0136]
[0137] S3-2: The input state information variable of the cognitive drone is:
[0138]
[0139] where, in steps S3-1 and S3-2, is the input state information of the k-th sub-band agent at the n-th time step, represents the frequency band allocation strategy of the k-th sub-band agent at the (n-1)-th time step; I n-1,k is the interference received by the k-th secondary user in the current sub-band at time n-1, H n,k is the channel power gain; is the position of the sub-band agent at the (n-1)-th time step; is the position of the cognitive drone at the (n-1)-th time step;
[0140] The channel power gain H n,k is:
[0141]
[0142] where, To recognize the channel power gain from the cognitive UAV to the k-th secondary user, To recognize the channel power gain from the interfering UAV to the k-th secondary user, To recognize the channel power gain from the primary base station to the k secondary users.
[0143] S4: Construct the reward function.
[0144] The reward function is:
[0145]
[0146] Where, is the transmission rate of the k-th primary user, w 1 is the weight of the sum of the primary user transmission rates, δ j is the penalty term when the minimum transmission rate of the primary user is not satisfied, w 2 is the weight of the sum of the penalty terms.
[0147] w 1 and w 2 are both non-negative constants, both are constants greater than 0 and less than 1, and are determined according to experience.
[0148] The penalty term δ when the minimum transmission rate of the primary user is not satisfied j is:
[0149]
[0150] Where, is the transmission rate of the j-th secondary user at the n-th time slot, R min is the minimum transmission rate requirement of the j-th primary user.
[0151] S5: Initialize the training round ep to 0.
[0152] S6: Initialize the time step t in the ep round to 0.
[0153] S7: Complete a single environment interaction.
[0154] S7-1: The AC action network Selects and outputs according to the current sub-band agent state Where ξ are the network parameters of the action network of the AC algorithm.
[0155] S7-2: Obtain the new sub-band agent state and the immediate reward r feedback from the environment n+1 n+1 .
[0156] S7-3: DDPG outputs an action according to the action network Outputs an action Obtain the new cognitive UAV state And store it in the experience pool, where is the DDPG action network parameter, and η follows a Gaussian distribution σ 2 is the noise variance.
[0157] S8: Determine whether t < T is satisfied, where T is the total number of steps in the ep episode. If so, then t = t + 1, and return to step S7. If not, then enter step S9.
[0158] As Figure 2 shown, S9: Calculate the policy gradient and TD error of the AC action network.
[0159] S9-1: The evaluation network generates the current value ω is the evaluation network parameter.
[0160] S9-2: Calculate the action network gradient
[0161]
[0162] where is the advantage function is the action network of the AC algorithm is the output of the action network in the AC algorithm is the input state of the sub-band agent, and ξ is the action network parameter of the action network in the AC algorithm is the gradient calculation symbol.
[0163] S9-3: Calculate the target value through the AC evaluation network.
[0164] The target value y of the AC algorithm k is
[0165]
[0166] where r n+1 is the immediate reward of the sub-band agent, γ 1 is the discount factor of the AC algorithm is the current value of the AC algorithm, and ω is the evaluation network parameter of the evaluation network in the AC algorithm.
[0167] S9-4: Calculate the TD error of the AC network at this time according to the target value and the current value.
[0168] The TD error of the AC algorithm is
[0169]
[0170] S10: Update the action network and the evaluation network according to the gradient ascent and gradient descent strategies.
[0171] Update the parameters of the action network using the gradient ascent strategy as:
[0172]
[0173] where α a is the learning rate parameter of the action network;
[0174] Update the parameters of the evaluation network using the gradient descent strategy as:
[0175]
[0176] where α c is the learning rate parameter of the evaluation network.
[0177] S11: Calculate the deterministic policy gradient and TD error in DDPG.
[0178] S11-1: Calculate the deterministic policy gradient in DDPG.
[0179] Specifically, the deterministic policy gradient of the DDPG algorithm is:
[0180]
[0181] where N tr is the sample batch size, is the input state of the cognitive UAV, is the input value of the action network in the DDPG algorithm; is the output of the evaluation network in the DDPG algorithm when the input value is ; are the parameters of the action network in the DDPG algorithm, and λ is the network parameter of the action network in the DDPG algorithm;
[0182] S11-2: Calculate the output of the DDPG target evaluation network and obtain the corresponding target value by combining the reward.
[0183] The target value y of the DDPG algorithm i is:
[0184]
[0185] where r i+1 is the immediate reward of the cognitive UAV; γ 2 is the discount factor of the DDPG algorithm, is the output of the target evaluation network in the DDPG algorithm, and λ -are the network parameters of the target evaluation network, are the network parameters of the target action network.
[0186] S11-3: Calculate the TD error of the offline DDPG network.
[0187] The TD error of the DDPG algorithm is:
[0188]
[0189] where, is the output of the evaluation network in the DDPG algorithm when the input value is .
[0190] S12: Minimize the DDPG network loss to update the network parameters of the UAV agent network.
[0191] S12-1: Update the network parameters of the DDPG evaluation network using the gradient descent algorithm.
[0192] The network parameter λ for updating the DDPG evaluation network using the gradient descent algorithm is:
[0193]
[0194] where, β c is the learning rate parameter of the evaluation network in the DDPG algorithm;
[0195] S12-2: Update the network parameters of the DDPG target action network and the network parameters of the DDPG target evaluation network using the soft update strategy.
[0196] The network parameters for updating the DDPG target action network and the network parameters of the DDPG target evaluation network using the soft update strategy are:
[0197]
[0198]
[0199] where, τ is the frequency factor for controlling the soft update strategy.
[0200] S13: Determine whether the number of rounds ep < EP is satisfied, where EP is the total number of episodes. If so, then ep = ep + 1, and return to step S7. If not, then the optimization ends, and the optimized model is obtained, and the system transmission rate at this time is calculated.
[0201] The ground users of the present invention have mobility, and the interference drones will generate various dynamic interferences. With the goal of maximizing rewards, the present invention has low complexity. Moreover, due to better experience replay pool information and fewer neural networks, it can obtain faster decisions. In addition, this application can simultaneously model the discrete actions of sub-band agents and the continuous actions of cognitive drones to ensure decision-making accuracy.
[0202] Embodiment 2
[0203] The effects of the present invention will be further described below in combination with simulation experiments.
[0204] 1. Simulation conditions and parameter settings:
[0205] Assume that there are M = 4 sub-bands in the system, J = 4 primary users, K = 3 secondary users, and the channel power gain β when the reference distance is 1 meter ref = 1×10 -3 . Both the cognitive drone and the interference drone fly at a fixed height of 100 meters, and the background noise power is fixed at -169 dBm.
[0206] 2. Simulation content:
[0207] Appendix Figure 3 is a comparison chart of the sum transmission rate obtained by using the method of the present invention and the prior art method under different numbers of secondary users. The abscissa represents the number of secondary users, and the ordinate represents the sum transmission rate of the secondary network. It can be seen that the scheme proposed by the present invention can obtain the maximum sum transmission rate. When the number of users is one, the performance is similar to that of the single-agent method. However, when the number of users increases, the advantage of the method of the present invention becomes more obvious, which verifies the effectiveness of the method we proposed.
[0208] Appendix Figure 4 is a comparison chart of the sum transmission rate obtained by using the method of the present invention and the prior art method under different transmission powers of the drone. The abscissa represents the transmission power of the cognitive drone, and the ordinate represents the sum transmission rate of the secondary network. It can be seen that the method of the present invention can always obtain the maximum sum transmission rate of the secondary network, and as the transmission power increases, the sum of the transmission rates also increases. When the transmission power exceeds 230 mW, the transmission rate tends to be stable. This is because when the transmission power increases, the interference to the primary network also increases. In order to avoid the interference exceeding the threshold, the present invention can maximize the secondary network transmission rate as much as possible while ensuring the communication quality of the primary network.
[0209] Appendix Figure 5This is a comparison of the convergence performance of the method of the present invention under different numbers of hidden layers. The abscissa is the number of training epochs, and the ordinate is the average reward. It can be seen that as the number of hidden layers increases, the network convergence speed accelerates. However, more hidden layers will lead to unstable training. When using 6 hidden layers, serious training oscillations occurred after the 150th training epoch. When the number of hidden layers is insufficient, the convergence is slow, and convergence was not completed within 500 training epochs.
[0210] Figure 6 This is a comparison of the convergence performance of the method of the present invention under different experience pool capacities. The abscissa is the number of training epochs, and the ordinate is the average reward. It can be seen that when the capacity is 12000, a significant decline occurred at the 300th training epoch. This is because an overly large experience pool will retain more useless information, so it is easily interfered by useless information as training progresses. Compared with an experience pool capacity of 1000, when the experience pool capacity is 8000, the convergence speed is increased by approximately 75 training epochs.
[0211] Based on the above simulation results and analysis, the resource intelligent optimization method for the UAV spectrum sharing network based on heterogeneous MA2C-DDPG proposed by the present invention can enable the spectrum sharing system to obtain the maximum sum transmission rate, accelerate the network convergence speed, and the framework has strong generalization ability and can be extended to various wireless communication scenarios, which makes the invention better applied in practice.
[0212] Embodiment III
[0213] The present invention provides a computer device, including a processor and a memory; wherein, when the processor executes the computer program stored in the memory, the steps of the above-mentioned resource optimization method for the UAV spectrum sharing network are implemented.
[0214] For a more specific process of the above method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.
[0215] Embodiment IV
[0216] The present invention provides a computer-readable storage medium for storing a computer program; when the computer program is executed by a processor, the steps of the above-mentioned resource optimization method for the UAV spectrum sharing network are implemented.
[0217] For a more specific process of the above method, reference can be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.
[0218] In the present specification, the various embodiments are described in a progressive manner. Each embodiment focuses on the differences from other embodiments, and the same or similar parts among the embodiments can be referred to each other. For the systems, devices, and storage media disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0219] Those skilled in the art can clearly understand that the technologies in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the parts that contribute to the prior art can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments of the present invention.
[0220] The above are only the preferred embodiments of the present invention, and the protection scope of the present invention is not limited to the above embodiments. All technical solutions falling within the idea of the present invention belong to the protection scope of the present invention. It should be noted that for those of ordinary skill in the art, several improvements and refinements made without departing from the principle of the present invention should be regarded as within the protection scope of the present invention.
Claims
1. A resource optimization method for unmanned aerial vehicle spectrum sharing network, characterized in that: include: Construct the input state information variables of the sub-band intelligent agent and the input state information variables of the cognitive drone; Determine the transmission rates of primary and secondary users and construct a reward function; Initialize the number of rounds and steps of the training round. During training, based on the input state information of the current sub-band agent, use the action network of the AC algorithm to select the action output of the sub-band agent and determine the immediate reward of the sub-band agent. Based on the input state information of the current cognitive drone, the action network of the DDPG algorithm is used to select the action output of the cognitive drone and determine the immediate reward of the cognitive drone; When the number of steps in a training round reaches the total number of steps in this training round, update the network parameters of the AC algorithm and the DDPG algorithm; According to the update results of the AC algorithm network parameters and the DDPG algorithm network parameters, training is continued until the set number of training rounds is met, and the optimization model is output, and the output optimization model is used to optimize resources; In the input state information variables of constructing the sub-band intelligent agent and the input state information variables of the cognitive drone: The input state information variable of the sub-band agent for: The input state information variable of the cognitive drone is: in, is the input state information of the k-th sub-band agent at the n-th time step, I represents the frequency band allocation strategy of the k-th sub-band agent at the n-1th time step; n-1,k is the interference received by the kth secondary user in the current sub-band at time n-1, H n,k is the channel power gain; is the position of the sub-band agent at the n-1th time step; is the position of the cognitive drone at the n-1th time step; The channel power gain H n,k for: in, is the channel power gain from the cognitive drone to the kth secondary user, is the channel power gain from the jammer UAV to the kth secondary user, is the channel power gain from the primary base station to k secondary users; The transmission rates of the primary user and the secondary user are determined, and the reward function is constructed. The transmission rate of the primary user for: The transmission rate of the secondary user for: Where B is the bandwidth of each sub-band, is the signal-to-noise ratio of the kth secondary user in the nth time slot and the mth sub-band, is the signal-to-noise ratio of the jth primary user in the nth time slot and the mth sub-band; The signal-to-noise ratio of the jth primary user in the nth time slot and the mth sub-band for: The signal-to-noise ratio of the kth secondary user in the nth time slot and the mth sub-band for: in, is the transmit power of the main base station in the mth sub-band, is the channel power gain between the primary base station and the jth primary user; is the noise power of the jth primary user, ρ k,n [m] indicates whether the mth sub-band is occupied, P c is the transmit power of the cognitive drone, is the channel power gain from the cognitive drone to the jth primary user in the nth time slot; is whether the mth sub-band is occupied by the interfering drone, P J To interfere with the transmission power of the drone, is the channel power gain from the interfering UAV to the jth primary user in the nth time slot; is the channel power gain from the cognitive drone to the kth secondary user, is the noise power of the kth secondary user; is the channel power gain from the primary base station to the kth secondary user, is the channel power gain between the n-time slot interfering UAV and the k-th secondary user; The transmission rates of the primary user and the secondary user are determined, and the reward function is constructed. The reward function is: in, is the transmission rate of the kth primary user, w1 is the weight of the sum of the primary user transmission rates, δ j is the penalty term when the minimum transmission rate of the primary user is not met, w2 is the weight of the penalty term and, The penalty term δ when the minimum transmission rate of the primary user is not met j for: in, is the transmission rate of the jth secondary user in the nth time slot, R min is the minimum transmission rate requirement of the jth primary user; When the number of steps in the training round reaches the total number of steps in this training round, the AC algorithm network parameters and the DDPG algorithm network parameters are updated. The updating of AC algorithm network parameters includes: The AC algorithm evaluation network generates the current value of the AC algorithm; Determine the policy gradient of the action network of the AC algorithm based on the input state information of the current sub-band agent and the action output of the action network in the AC algorithm; Determine the target value of the AC algorithm based on the policy gradient of the AC algorithm's action network and the immediate reward of the sub-band agent; Determine the TD error of the AC algorithm based on the current value and target value of the AC algorithm; According to the policy gradient of the action network of the AC algorithm, the action network parameters are updated using the gradient ascent strategy, and according to the TD error of the AC algorithm, the evaluation network parameters are updated using the gradient descent strategy; The policy gradient of the action network of the AC algorithm for: in, is the advantage function, It is the action network of AC algorithm; is the output of the action network in the AC algorithm, is the input state of the sub-band agent, ξ is the action network parameter of the action network in the AC algorithm; The target value y of the AC algorithm k for: Among them, r n+1 is the instant reward of the sub-band agent, γ1 is the discount factor of the AC algorithm, is the current value of the AC algorithm, ω is the evaluation network parameter of the evaluation network in the AC algorithm; The TD error of the AC algorithm is: The gradient ascent strategy is used to update the action network parameters: Among them, α a is the learning rate parameter of the action network; The gradient descent strategy is used to update the evaluation network parameters: Among them, α c To evaluate the learning rate parameters of the network; When the number of steps in the training round reaches the total number of steps in this training round, the AC algorithm network parameters and the DDPG algorithm network parameters are updated. Update the DDPG algorithm network parameters including: The target evaluation network of the DDPG algorithm outputs the current value; Determine the deterministic policy gradient of the DDPG algorithm based on the state information of the cognitive drone and the action output results of the action network in the DDPG algorithm; Determine the target value of the DDPG algorithm based on the immediate reward of the cognitive drone and the current value of the target evaluation network; Determine the TD error of the DDPG algorithm based on the current value of the target evaluation network and the target value of the DDPG algorithm; According to the TD error of the DDPG algorithm, the network parameters of the DDPG evaluation network are updated using the gradient descent algorithm; According to the network parameters of the DDPG evaluation network, the network parameters of the DDPG target action network and the network parameters of the target evaluation network are updated using a soft update strategy; Deterministic Policy Gradient of the DDPG Algorithm for: Among them, N tr is the sample batch size, is the input state of the cognitive drone, It is the input value of the action network in the DDPG algorithm; The input value is When , the output of the evaluation network in the DDPG algorithm is, is the action network parameter in the DDPG algorithm, and λ is the network parameter of the action network in the DDPG algorithm; The target value y of the DDPG algorithm i for: Among them, r i+1 is the instant reward of the cognitive drone; γ2 is the discount factor of the DDPG algorithm, is the output of the target evaluation network in the DDPG algorithm, λ - Evaluate the network parameters of the network for the target, are the network parameters of the target action network; The TD error of the DDPG algorithm is: in, The input value is When , the output of the evaluation network in the DDPG algorithm; is the output value of the action network in the DDPG algorithm; Among them, η is the noise variance that follows Gaussian distribution; The network parameter λ of the DDPG evaluation network is updated using the gradient descent algorithm: Among them, β c The learning rate parameter for evaluating the network in the DDPG algorithm; The network parameters of the DDPG target action network and the target evaluation network are updated using the soft update strategy as follows: Among them, τ is the frequency factor that controls the soft update strategy.
2. A computer device, characterized in that: It includes a processor and a memory; wherein, when the processor executes the computer program stored in the memory, the steps of the resource optimization method for the drone spectrum sharing network described in claim 1 are implemented.
3. A computer-readable storage medium, characterized in that: Used to store computer programs; when the computer programs are executed by the processor, the steps of the resource optimization method for the drone spectrum sharing network described in claim 1 are implemented.