Method for maximizing safety and rate resource allocation for uav-assisted rsma communication system
By optimizing UAV trajectory and transmission power through deep reinforcement learning, the limitations of resource allocation in UAV-assisted RSMA communication systems are solved, achieving high user safety rates and improved system performance.
Patent Information
- Application Number
- CN202411038627.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2044-07-31
AI Technical Summary
Existing UAV-assisted RSMA communication systems have limitations in user safety rate optimization and resource allocation, including wasted resources for users with low information security requirements, insufficient power allocation, and the challenge of multi-agent trajectory optimization in high-dimensional variable spaces.
A deep reinforcement learning approach is adopted, which optimizes the UAV trajectory and transmission power through a multi-agent deep Q-learning algorithm. Combined with the RSMA communication protocol, the security and rate resource allocation of the UAV-assisted RSMA communication system is optimized. The Mixer network in the deep Q-learning algorithm is used for centralized learning to realize a distributed strategy.
In a dynamic environment, the system achieves efficient resource allocation for UAV-assisted RSMA communication, improves user security and system performance, reduces computational resource waste, and exhibits stronger robustness and faster convergence.
Smart Images

Figure CN119052748B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical field of rate segmented multiple access downlink systems, specifically relating to a precoding method for maximizing security and rate in rate segmented multiple access downlink systems and a method for optimizing UAV trajectories. Background Technology
[0002] Given the highly diverse application demands of future wireless networks, drone-assisted communication (UAV-assisted communication) is considered one of the most promising technologies due to its operational flexibility and wide coverage. By acting as temporary base stations, UAVs can perform a wide range of activities, including disaster relief support, real-time surveillance, and data collection. However, due to their wide coverage and large-scale connections, open networks are vulnerable to eavesdropping and various malicious attacks. Since UAVs are often used to transmit private data, designing secure UAV-assisted systems is crucial.
[0003] With the advent of the next generation of mobile internet, the massive access of IoT devices will inevitably cause traffic congestion. RSMA (Rate-Splitting Multiple Access) is a multi-antenna, multi-user communication multiple access scheme based on rate-splitting and linear precoding that multiplexes users in the power domain. RSMA divides user messages into common and private parts, encoding the common part into one or more common streams, while encoding the private part into separate streams. These streams are precoded using channel state information available at the transmitter. They are superimposed and transmitted through multiple-input multiple-output (MIMO) or multiple-input single-output (MISO) channels. At the receiver, all receivers decode the common stream, treat other users' messages as noise and perform continuous interference cancellation (SIC), and then decode their respective private streams. The main advantage of RSMA is its flexible management of interference by allowing the decoding of interfering parts.
[0004] Although some research has been conducted both domestically and internationally on the secure communication problems of RSMA-based UAV communication systems, and these studies have to some extent met certain optimization goals for RSMA-based UAV communication systems, existing research still has certain limitations in its analytical methods and optimization approaches for the secure communication problems of RSMA-based UAV communication systems. These limitations include:
[0005] (1) Current research on optimizing user security speed usually does not take into account that there are users in the user group who only have requirements for speed but low requirements for information security. Optimizing these users as confidential users may lead to a huge waste of resources.
[0006] (2) Due to the equipment precision and other reasons, the eavesdropping may be better than the user affected by the environmental noise, and the method of encoding the artificial noise with the user signal at the transmitting end will cause insufficient power allocated to the user, thereby reducing the information rate of the user.
[0007] (3) The trajectory optimization of the unmanned aerial vehicle is usually a non-convex problem, which can be approximately converted into a convex problem, and solved by means of a convex optimization tool, however, the traditional optimization method is difficult to solve the multi-agent trajectory optimization problem in a high-dimensional variable space. SUMMARY
[0008] The present application aims to solve the problems of the prior art. A method for maximizing security and rate resource allocation of an unmanned aerial vehicle assisted RSMA communication system is proposed. The technical scheme of the present application is as follows:
[0009] A method for maximizing security and rate resource allocation of an unmanned aerial vehicle assisted RSMA communication system, comprising the following steps:
[0010] 101. Establishing a security and rate maximization model of an unmanned aerial vehicle assisted RSMA communication network under the condition of existing eavesdroppers and secondary users;
[0011] 102. Initializing the starting point of the unmanned aerial vehicle, the transmission power allocation coefficient, the base station position and the transmission power, the minimum rate threshold of the secondary user, and the flight cycle size;
[0012] 103. Setting the parameters of the environment in the deep Q learning algorithm, defining the state space S as {the horizontal coordinates of the unmanned aerial vehicle, the user security rate}, defining the action space a as {the power allocation coefficient, the flight direction of the unmanned aerial vehicle, the flight speed of the unmanned aerial vehicle}, defining the reward r as {the user security rate}, wherein the unmanned aerial vehicle acts as an agent;
[0013] 104. Initializing the experience replay buffer D, the learning rate alpha, the target network update frequency tau, the Mixer network parameter theta, and the local Q network parameter theta i ;
[0014] 105. Randomly sampling a small batch of experience from the experience replay buffer D, calculating the local Q value function and the global Q value function for each agent separately, updating the Mixer network parameter, and selecting the agent action through the global Q value;
[0015] 106. Outputting the reward obtained from the environment at the current time, changing the agent state to the next moment, storing the learned samples in the experience replay pool, and saving the output parameters.
[0016] Further, the step 101 specifically comprises the following steps:
[0017] First, a model containing two unmanned aerial vehicles, an eavesdropper, a confidential user and a rate user is established. A transmitting base station with N antennas needs to communicate with M single-antenna users through a relay unmanned aerial vehicle with K antennas. An eavesdropper Eve eavesdrops the communication information of the main user. An aerially deployed jammer unmanned aerial vehicle transmits artificial noise to interfere with the eavesdropping channel from the base station to Eve. The secondary user without the risk of eavesdropping also communicates through the relay unmanned aerial vehicle. The common rate allocation is optimized to maximize the secure communication rate of the eavesdropped user under the condition that the rate of such user is higher than the threshold. Information transmission needs two stages. In the first time slot, the base station transmits a broadcast signal: where d is a data symbol and satisfies E[|d| 2 ]=1, w m is the precoding vector of the base station transmitting the mth user message, w c is the precoding vector of the base station transmitting the common message, and M is the total number of users. The signal received at the relay from the base station is: where w i ~CN(0,σ 2 ) is additive white Gaussian noise, σ 2 is the variance of the noise, W=[w c ,w1,…,w m ] represents the precoding matrix of the base station, P J is the power of the jammer unmanned aerial vehicle, G BR represents the channel gain between the base station and the relay, and h JR represents the channel gain between the jammer unmanned aerial vehicle and the relay.
[0018] In the second time slot, the relay transmits a signal to the user u using the DF protocol. The signal received by the user u is: where W represents the precoding matrix of the unmanned aerial vehicle, is the precoding vector of the relay unmanned aerial vehicle transmitting the mth user message, is the precoding vector of the relay unmanned aerial vehicle transmitting the common message, and G Ru represents the channel gain between the relay and the user u.
[0019] Further, the step 101 is based on the security and rate maximization objective function in the unmanned aerial vehicle assisted RSMA communication system:
[0020]
[0021] C3:X min ≤X J (t)≤X max ,X min ≤X R (t)≤X max
[0022] C4:Y min ≤Y J (t)≤Y max ,Y min ≤Y R (t)≤Y max
[0023] C5:H min ≤H J (t)≤H max ,H min ≤H R (t)≤H max
[0024] q J, q R represent the interfering UAV and the relay UAV trajectory respectively, T, t represent the UAV flight time, the time slot the system is in respectively, tr is the trace of the precoding matrix, P B , P R represent the base station power, the relay UAV power respectively, R S , R NS (t) represent the system safety and rate, the minimum rate of the secondary user, the rate of the secondary user, [X min ,X max ],[Y min ,Y max ],[H min ,H max ] are the UAV flight space limits, i.e. the minimum, maximum length, width, height of the flight space cuboid of the UAV R, the UAV J; l min is the minimum distance of two UAVs, less than this distance, the two UAVs may collide and cause loss; R S (t), R m (t), R m,E (t), R c (t), R c,E (t) represent the system safety and rate, the private message rate of the user m, the private message rate of the user m overheard by the eavesdropper, the public message rate of the user m, the public message rate of the user m overheard by the eavesdropper; for the relay UAV and the user, first eliminate artificial noise, then decode the public signal by regarding all private signals as noise, then delete the public message through SIC, and then decode the private signal of the receiver by regarding the private signals of other users as noise; in order to ensure that all users can correctly decode the public message, the public rate cannot exceed R' c (t) = min{R 1,c (t), R2,c (t),…,R m,c (t)}, and in the relay network, the user rate should satisfy R c (t)=min{R R,c (t),R u,c (t)},R 1,p (t)=min{R R,u,p (t),R' u,p (t)};
[0025] This indicates the common rate of the drone relay reception. This indicates the private rate at which the drone receives data from user u. This represents the common rate received by user u. V represents the rate at which user u receives private messages, V = [v c ,v1,…,v m ] represents the receive precoding matrix of the relay drone. This represents the private message precoding vector of user m;
[0026] For an eavesdropper, Eve can obtain confidential user information from base station transmissions or drone launch messages. u,E (t)≤max{R 1,u,E (t),R 2,u,E (t)},R c,E (t)≤max{R 1,c,E (t),R 2,c,E (t)} represent the public rate and private rate of the eavesdropper, respectively; in the first stage, the eavesdropper rate can be expressed as
[0027] G BE (t), h JE (t) represents the channel gain from the base station, the jamming drone, and the eavesdropper, respectively;
[0028] In the second phase, the eavesdropper rate can be expressed as
[0029] The precoding matrix of the UAV is represented by W = [w c ,w1,…,w m [] represents the precoding matrix of the base station. By optimizing the precoding matrix through power allocation of public and private signals and antenna incident angle, when the information rate received by the UAV from the base station is greater than the information rate received by the user, the power coefficient of the base station in the DF relay is set to be equal to the power allocation coefficient of the UAV. P B ,P R This indicates the transmission power of the drone and the base station. denotes the minimum rate threshold of the secondary user, G BR ,G BE ,G RU ,G RE ,h JR ,h JU ,h JE denote the channel gain from the base station to the drone relay, the eavesdropper, the drone relay to the user u, the eavesdropper, the interference drone to the drone relay, the user u, the eavesdropper, respectively;[X J (t),Y J (t),H J (t) is the position of the interference drone in the time slot, [X R (t),Y R (t),H R (t) is the position of the relay drone in the time slot;[X min ,X max ],[Y min ,Y max ],[H min ,H max ] is the flight space limit of the drone; C1 is the precoding matrix power constraint, C2 is the minimum transmission rate requirement of the secondary user, C3-C5 represent the horizontal limit of the drone flight area, and C6 is the limit of the drone collision avoidance.
[0030] Further, the step 102 initializes: the initial position of the relay drone is [0m, 500m], the position of the interference drone is [0m, 200m], the position of the base station is [0m, 0m], the common information power allocation coefficient is a c =0.4, the primary user power allocation coefficient is a1=0.4, a2=0.1, the secondary user power allocation coefficient is a3=0.1, and the base station power is 25dBm.
[0031] Further, the step 104 and the step 105 update the Mixer parameter, and the QMIX algorithm can realize centralized learning, but a distributed strategy is obtained, that is: First, according to the current parameter θ i , the local action value Q i (S i ,A i ; θ i ) of the agent is calculated. i (S i ,A i ; θ i ) + α [R + γmax a Q i (S i ',A i ; θ i)-Q i (S i A i ;θ i [), where the learning rate α = 2 × 10 -4 With a discount rate γ = 0.99, by performing action A in the current environment S, a new environment S' and reward R can be obtained; QMIX's training method is end-to-end training that minimizes the loss function. Where θ - The TD error in the DQN network; actions are selected using a greedy strategy, based on Q... tot With monotonicity constraints, update the local action value function parameter θ. i Then, in the hybrid network, each Q-value is mixed with the global state S to generate Q. tot The final Mixer is obtained. The parameters are repeated in the above steps until the maximum number of iterations is reached, where b represents the number of samples sampled from the experience pool.
[0032] Furthermore, it also includes the step of calculating the reward function of the system. In DRL, by designing a reward function, difficult-to-optimize objectives are optimized through the accumulation of rewards. When an agent performs an action, it transitions to another state and receives a reward value. The reward function consists of:
[0033] (1) Set a negative reward to penalize collisions between drones or sub-user rates falling below the threshold rate before training ends;
[0034]
[0035] (2) The setting of the reward function is not only related to the communication quality between the drone and the base station and the user, but also to the channel between the drone and the eavesdropper; in this setting, the drone will receive a positive reward for the user's safe rate according to the strategy selection.
[0036] Reward = R S
[0037] It further outputs the reward obtained from the environment at the current moment, changes the agent's state to the next moment, stores the learned samples in the experience replay pool, and saves the output parameters.
[0038] The advantages and beneficial effects of this invention are as follows:
[0039] This invention maximizes the sum and rate of a UAV-assisted RSMA communication system under constraints such as UAV flight trajectory, transmit power, base station transmit power, power allocation coefficient, and secondary user rate (step one). It considers solving for the UAV trajectory within the flight cycle and optimizes the transmit power allocation between the base station and the UAV, which is more practically valuable than solving for system safety and rate using maximum transmit power. Compared to previous solutions for wireless resource management, this invention applies deep reinforcement learning (step three). Considering the dynamic environment of UAV position changes, traditional convex optimization methods require recalculating the resource allocation scheme whenever the environment changes. When the frequency of environmental changes increases, this undoubtedly leads to a significant waste of computational resources. Furthermore, changes in UAV position will cause changes in the direct channel power gain between the UAV and user channels. The advantage of this invention is that the RSMA protocol can achieve high-quality communication for a large number of users, and the use of deep reinforcement learning for optimization provides stronger robustness. Moreover, this invention uses multi-agent deep reinforcement learning, which has faster convergence time and stronger convergence performance compared to other single-agent algorithms. Attached Figure Description
[0040] Figure 1 This invention provides a preferred embodiment of the system model.
[0041] Figure 2 Figures show the convergence curves of the average system user rate of the proposed rate division protocol and comparison method (NOMA and SDMA) under different secondary user rate constraints, where the secondary user rate constraints in Figures (a), (b), and (c) are 0, 0.2 Mbit / s, and 0.4 Mbit / s, respectively.
[0042] Figure 3 The different drone relay power consumption P c =8dbm,P c =10dBm,P c =12dbm,P c =14dBm, the system reward of the proposed invention and the comparison method (interference-free drone and drone in straight flight);
[0043] Figure 4 This is a graph showing the system reward convergence curves of the algorithm (QMIX) and the comparison methods (DoubleDQN, DuelingDQN) of this invention.
[0044] Figure 5 It is a drone trajectory map;
[0045] Figure 6 This is a diagram showing the power distribution changes during the relay flight of a UAV.
[0046] Figure 7 A flowchart of a method for maximizing safety and rate resource allocation in a UAV-assisted RSMA communication system. DETAILED DESCRIPTION
[0047] The technical solutions in the embodiments of the present application will be described clearly and in detail below with reference to the drawings in the embodiments of the present application. The described embodiments are only some of the embodiments of the present application.
[0048] The technical solutions of the present application to solve the above technical problems are:
[0049] The present application discloses a method for maximizing safety and rate resource allocation in a UAV-assisted RSMA communication system under the condition of a secondary user rate threshold. It includes the following steps:
[0050] The first step is to establish a UAV-assisted RSMA communication network safety and rate maximization model under the condition of the presence of eavesdroppers and secondary users;
[0051] The second step is to initialize the starting point of the UAV, the transmission power allocation coefficient, the base station location and the transmission power, the minimum rate threshold of the secondary user, and the flight cycle size;
[0052] The third step is to set the parameters of the environment in the deep Q learning algorithm, define the state space S as {the horizontal coordinates of the UAV, the user safety rate}, define the action space a as {the power allocation coefficient, the flight direction of the UAV, the flight speed of the UAV}, and define the reward r as {the user safety rate}, wherein the UAV acts as an intelligent agent;
[0053] The fourth step is to initialize the experience replay buffer D, the learning rate a, the target network update frequency τ, the Mixer network parameter θ, and the local Q network parameter θ i .
[0054] The fifth step is to randomly sample a small batch of experience from the experience replay buffer D, calculate the local Q value function for each intelligent agent separately, then calculate the global Q value function, update the Mixer network parameter, and select the intelligent agent action through the global Q value.
[0055] The sixth step is to output the reward obtained from the environment at the current time, change the state of the intelligent agent to the next moment, store the learned samples in the experience replay pool, and save the output parameters.
[0056] The safety and rate maximization objective function in the UAV-assisted rate split multiple access (RSMA) communication system is:
[0057]
[0058] C3:Xmin ≤X J (t)≤X max ,X min ≤X R (t)≤X max
[0059] C4:Y min ≤Y J (t)≤Y max ,Y min ≤Y R (t)≤Y max
[0060] C5:H min ≤H J (t)≤H max H min ≤H R (t)≤H max
[0061]
[0062] For safety and rate expression, for relay drones and users, artificial noise is first eliminated. Then, the public signal is decoded by treating all private signals as noise. After deleting the public message using SIC, the receiver's private signal is decoded by treating other users' private signals as noise. To ensure that all users can correctly decode the public message, the public rate cannot exceed R'. c (t)=min{R 1,c (t),R 2,c (t),…,R m,c In a relay network, information propagation must adhere to information causality; that is, the amount of information received by a user cannot exceed the amount of information transmitted. Therefore, the user rate should satisfy R. c (t)=min{R R,c (t),R u,c (t)},R 1,p (t)=min{R R,u,p (t),R' u,p (t)};
[0063] This indicates the common rate of the drone relay reception. This indicates the private rate at which the drone receives data from user u. This represents the common rate received by user u. R represents the rate at which user u receives private messages. From an eavesdropper's perspective, eavesdropper Eve can obtain confidential user information from information transmitted by base stations or drones. m,E (t)≤max{R1,u,E (t),R 2,u,E (t)},R c,E (t)≤max{R 1,c,E (t),R 2,c,E (t)} respectively represent the eavesdropper's common rate and private rate. In the first phase, the eavesdropper rate can be expressed as
[0064] In the second phase, the eavesdropper rate can be expressed as
[0065]
[0066] W represents the precoding matrix of the UAV, W = [w c ,w1,…,w m ] represents the precoding matrix of the base station, V = [v c ,v1,…,v m ] represents the receiving precoding matrix of the relay UAV, by optimizing the precoding matrix for the power allocation of the common signal and the private signal and the antenna incident angle, in order to simplify the algorithm complexity, in the case that the information rate of the UAV receiving the base station is greater than the user accepted information rate, the power coefficient of the base station in the DF relay can be equal to the power allocation coefficient of the UAV. B ,P R represents the maximum transmission power of the UAV and the base station, represents the minimum rate threshold of the secondary user, G BR ,G BE ,G RU ,G RE ,h JR ,h JU ,h JE respectively represent the channel gain from the base station to the UAV relay, the eavesdropper, the channel gain from the UAV relay to the user u, the eavesdropper, the channel gain from the interfering UAV to the UAV relay, the user u, the eavesdropper.[X J (t),Y J (t),H J (t)] is the position of the interfering UAV in the time slot, [X R (t),Y R (t),H R (t)] is the position of the relay UAV in the time slot.[X min ,X max ],[Y min ,Y max ],[H min ,H max] is the flight space restriction of the UAV. C1 is the precoding matrix power constraint, C2 is the minimum transmission rate requirement of the secondary user, C3-C5 represent the horizontal restriction of the flight area of the UAV, and C6 is the restriction of the UAV to avoid collision.
[0067] In the second step, the initial position of the UAV and the transmission power of the UAV are initialized, the precoding matrix is set, the minimum rate threshold of the secondary user is set as Rmin=1bps / Hz, the flight period of the UAV is defined as T=Nt, T is the entire flight period of the UAV, t is the length of each time slot, one period is divided into N time slots, and the unit is second. In each time slot, the distance between the UAV and any ground node in the research area is substantially unchanged, and the channel gain between the UAV and the user and the eavesdropper and the gain of the base station antenna to the UAV remain substantially unchanged.
[0068] The embodiment is a resource allocation method for maximizing the safety and rate of a UAV-assisted RSMA communication system under the rate constraint of a secondary user. Downlink communication is considered, and the path loss of propagation depends on the distance between the UAV and the ground user and the type of propagation environment. The UAV can hover above the target area, and one UAV acts as an aerial relay base station to provide communication services for k (k≥1, k∈K) ground users in a rate division multiple access mode, while one interfering UAV hinders potential eavesdroppers from eavesdropping on the legitimate users. In the first stage of relay communication, the base station transmits data to the UAV, and the initial coordinates of the base station, the UAV relay (Relay) and the UAV (Jammer) interference are (0, 0), (0, 200) and (200, 0) m respectively. The primary user is located at a random position in the range of [4000±1000, 4000±1000], the eavesdropper is located at a random position in the range of [4000±750, 2000±750], the secondary user is located at a random position in the range of [1500±1000, 4000±1000], the number of antennas of the base station is 4, the number of antennas of the UAV relay is 4, and the environmental noise is σ 2 =10 -17 W.
[0069] In this example, Figure 1 The system model provided by the present application for maximizing the safety and rate of a UAV-assisted RSMA communication system is provided. Figure 2 Under different rate constraints of the secondary user, the system user average rate convergence curves of the proposed rate division protocol and the comparative methods (NOMA and SDMA) are shown in the following figures. In the figures (a), (b) and (c), the rate constraints of the secondary user are 0, 0.2 Mbit / s and 0.4 Mbit / s respectively. Figure 3 Under different power consumptions P c =8dbm, P c= 10dbm, P c = 12dbm, P c = 14dbm, the system reward of the proposed method and the comparative method (no interference and no interference UAV and UAV straight flight) under the condition of Figure 4 is the system reward convergence curve of the algorithm (QMIX) and the comparative method (DoubleDQN, DuelingDQN) of the application, Figure 5 is the UAV trajectory diagram, Figure 6 is the power allocation change diagram in the process of UAV relay flight;
[0070] From Figure 2 It can be seen from the embodiment method that the RSMA protocol can obtain greater advantages than the SDMA protocol under three secondary user constraint conditions, and compared with the NOMA protocol, there is little difference when there is no secondary user constraint, at this time the RSMA protocol may degenerate into NOMA, and in the other two cases it can be significantly better than the NOMA protocol, verifying the superiority of the communication protocol used in the application.
[0071] From Figure 3 It can be seen from the embodiment method that the optimization of the relay UAV and the interference UAV trajectory can effectively improve the service quality of the user, and compared with the UAV straight flight and the no interference UAV scene, the primary user can have a larger safety rate under the secondary user service quality constraint, verifying the feasibility of the algorithm proposed in the application.
[0072] From Figure 4 It can be seen from the embodiment method that the QMIX algorithm can converge faster and have a larger convergence value than DuelingDQN and DoubleDQN in the same environment, and can better optimize the UAV trajectory and power allocation compared with the other two algorithms, verifying the feasibility of the algorithm proposed in the application.
[0073] From Figure 5 It can be seen from the embodiment method that the relay UAV will approach the user away from the eavesdropper and finally stay between the primary user and the secondary user, which may be to ensure the service quality of the secondary user and improve the public information rate, and the interference UAV will directly move to the eavesdropper directly above without interfering with the user, which maximizes the interference to the eavesdropper.
[0074] From Figure 6 It can be seen from the embodiment method that the relay UAV will first improve the power allocation coefficients of the public information and the secondary user during the flight process, which is to ensure the secondary user service quality constraint in the case of poor channel environment, and when the UAV approaches the secondary user, the channel improves, then the power allocation coefficient of the secondary user is reduced, and the power allocation coefficient of the primary user is improved, improving the service quality of the primary user.
[0075] The systems, apparatuses, modules or units illustrated by the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions.
[0076] It should be further noted that the terms "comprising" or "including" or any other variation thereof are intended to cover a non-exclusive inclusion, such that processes, methods, articles, or apparatuses that comprise a list of elements do not include only those elements but can also include other elements not expressly listed or inherent to such processes, methods, articles, or apparatuses. Without further limitation, an element defined by the phrase "comprising a" does not exclude the existence of additional identical elements in the process, method, article, or apparatus that includes the element.
[0077] The above embodiments are understood as merely illustrative of the present application and not restrictive of the scope of the present application. After reading the above description of the present application, those skilled in the art can make various modifications or changes to the present application, and these equivalent changes and modifications also fall within the scope of the present application defined by the claims.
Claims
1. A method for maximizing security and rate resource allocation in an unmanned aerial vehicle (UAV) assisted RSMA communication system, characterized in that, Includes the following steps:
101. Establish a network security and rate maximization model for UAV-assisted RSMA communication under the condition of the presence of eavesdroppers and secondary users; 102. Initialize the UAV's starting point, transmission power allocation coefficient, base station location and transmission power, minimum rate threshold for secondary users, and flight cycle size; 103. Set the parameters of the environment in the deep Q-learning algorithm. Define the state space S as {the horizontal coordinate of the drone and the safe speed of the user}, define the action space a as {the power allocation coefficient, the flight direction of the drone, and the flight speed of the drone}, and define the reward r as {the safe speed of the user}, where the drone acts as the intelligent agent.
104. Initialize the experience replay buffer D, learning rate α, target network update frequency τ, Mixer network parameters θ, and local Q-network parameters θ. i ; 105. Randomly sample a small batch of experiences from the experience replay buffer D, calculate the local Q-value function for each agent separately, then calculate the global Q-value function, update the Mixer network parameters, and select the agent action based on the global Q-value.
106. Output the reward obtained from the environment at the current moment, change the agent's state to the next moment, store the learned samples in the experience replay pool, and save the output parameters.
2. The method for maximizing security and rate resource allocation in an unmanned aerial vehicle-assisted RSMA communication system according to claim 1, characterized in that, Step 101 specifically includes the following steps: First, a model is established that includes two drones, an eavesdropper, a secure user, and a rate user. A base station with N antennas needs to communicate with M single-antenna users via a relay drone with K antennas. The eavesdropper Eve eavesdrops on the communication information of the primary user. An aerial jamming drone Jammer emits artificial noise to interfere with the eavesdropping channel from the base station to Eve. Secondary users without the risk of being eavesdropped also communicate through relay drones. While ensuring that the rate of such users is higher than a threshold, the secure communication rate of the eavesdropped users is maximized by optimizing the allocation of common rates and the trajectories of Jammer and Relay. Information transmission requires two stages. In the first time slot, the base station sends a broadcast signal: In the formula, d is a data symbol, and it satisfies E[|d| 2 ] = 1, w m w is the precoding vector for the m-th user message sent by the base station. c The precoded vector for sending public messages to the base station, where M is the total number of users, and the signal received from the base station at the relay is: Where w i ~CN(0,σ 2 ) is additive white Gaussian noise, σ 2 It is the variance of the noise, W = [w c ,w1,…,w m ] represents the precoding matrix of the base station, P J To interfere with the drone's power, G BR The channel gain, h, is expressed as the channel gain between the base station and the relay. JR This is expressed as the channel gain between the interfering drone and the relay; In the second time slot, the relay uses the DF protocol to send a signal to user u. The signal received by user u is: This represents the precoding matrix of the UAV. This is the pre-encoded vector for relaying the m-th user message from the drone. G is the precoded vector for relaying public messages from drones. Ru This represents the channel gain between the relay and user u.
3. A method for maximizing security and rate resource allocation in an unmanned aerial vehicle-assisted RSMA communication system according to claim 2, characterized in that, The objective function for maximizing security and rate in the UAV-assisted RSMA communication system, as described in step 101, is: C3:X min ≤X J (t)≤X max ,X min ≤X R (t)≤X max C4:Y min ≤Y J (t)≤Y max ,Y min ≤Y R (t)≤Y max C5:H min ≤H J (t)≤H max ,H min ≤H R (t)≤H max For the safety and rate expressions, q J, q R The trajectories of the jamming UAV and the relay UAV are represented respectively, where T and t represent the UAV flight time and the time slot in which the system is located, respectively, tr is the trace of the precoding matrix, and P is the trajectory of the relay UAV. B P R These represent the base station power, relay drone power, and R... S , R NS (t) represent system security and rate, minimum rate for secondary users, rate for secondary users, and rate for secondary users, respectively. [X min ,X max ],[Y min ,Y max ],[H min H max [ ] represents the flight space constraints for unmanned aerial vehicles (UAVs), specifically the minimum and maximum length, width, and height of the flight space cube for UAVs R and J; l min This is the minimum distance between two drones; if the distance is less than this, the two drones may collide and cause damage. S (t), R m (t), R m,E (t), R c (t), R c,E (t) represent the system security and rate, the rate of private messages of user m, the rate of private messages of user m eavesdropped by the eavesdropper, the rate of public messages of user m, and the rate of public messages of user m eavesdropped by the eavesdropper, respectively. For relay drones and users, artificial noise is first eliminated. Then, the public signal is decoded by treating all private signals as noise. After deleting the public message using SIC, the receiver's private signal is decoded by treating other users' private signals as noise. To ensure that all users can correctly decode the public message, the public rate cannot exceed R'. c (t)=min{R 1,c (t),R 2,c (t),…,R m,c (t)}, and in the relay network, the user rate should satisfy R c (t)=min{R R,c (t),R u,c (t)},R 1,p (t)=min{R R,u,p (t),R' u,p (t)}; This indicates the common rate of the drone relay reception. This indicates the private rate at which the drone receives data from user u. This represents the common rate received by user u. V represents the rate at which user u receives private messages, V = [v c ,v1,…,v m ] represents the receive precoding matrix of the relay drone. This represents the private message precoding vector of user m; For an eavesdropper, Eve can obtain confidential user information from base station transmissions or drone launch messages. u,E (t)≤max{R 1,u,E (t),R 2,u,E (t)},R c,E (t)≤max{R 1,c,E (t),R 2,c,E (t)} represent the public rate and private rate of the eavesdropper, respectively; in the first stage, the eavesdropper rate can be expressed as G BE (t), h JE (t) represents the channel gain from the base station, the jamming drone, and the eavesdropper, respectively; In the second phase, the eavesdropper rate can be expressed as The precoding matrix of the UAV is represented by W = [w c ,w1,…,w m [] represents the precoding matrix of the base station. By optimizing the precoding matrix through power allocation of public and private signals and antenna incident angle, when the information rate received by the UAV from the base station is greater than the information rate received by the user, the power coefficient of the base station in the DF relay is set to be equal to the power allocation coefficient of the UAV. P B ,P R This indicates the transmission power of the drone and the base station. G represents the minimum rate threshold for secondary users. BR G BE G RU G RE ,h JR ,h JU ,h JE These represent the channel gain from base station to drone relay and eavesdropper, drone relay to user u and eavesdropper, and interference drone to drone relay, user u, and eavesdropper, respectively; [X] J (t),Y J (t),H J [(t)] represents the position of the interfering UAV in the time slot, [X] R (t),Y R (t),H R [t] represents the position of the relay UAV in the time slot; [X] min ,X max ],[Y min ,Y max ],[H min H max C1 represents the flight space constraints for the drone; C2 represents the power constraints of the precoding matrix; C3-C5 represent the minimum transmission rate requirements for secondary users; C3-C5 represent the levels of drone flight area restrictions; and C6 represents the drone collision avoidance constraints.
4. The method for maximizing security and rate resource allocation in an unmanned aerial vehicle-assisted RSMA communication system according to claim 1, characterized in that, Step 102 initialization: The initial position of the relay drone is [0m, 500m], the position of the interfering drone is [0m, 200m], the position of the base station is [0m, 0m], and the public information power allocation coefficient is a. c =0.4, the primary user power allocation coefficient is a1=0.4, a2=0.1, the secondary user power allocation coefficient is a3=0.1, and the base station power is 25dBm.
5. A method for maximizing security and rate resource allocation in an unmanned aerial vehicle-assisted RSMA communication system according to claim 4, characterized in that, Steps 104 and 105 update the Mixer parameters, enabling the QMIX algorithm to achieve centralized learning but obtain a distributed strategy, namely: First, based on the current parameter θ i Calculate the local action value Q of the agent. i (S i A i ;θ i )←Q i (S i A i ;θ i )+α[R+γmax a Q i (S i ',A i ;θ i )-Q i (S i A i ;θ i [), where the learning rate α = 2 × 10 -4 With a discount rate γ = 0.99, by performing action A in the current environment S, a new environment S' and reward R can be obtained; QMIX's training method is end-to-end training that minimizes the loss function. Where θ - The TD error in the DQN network; actions are selected using a greedy strategy, based on Q... tot With monotonicity constraints, update the local action value function parameter θ. i Then, in the hybrid network, each Q-value is mixed with the global state S to generate Q. tot The final Mixer is obtained. The parameters are repeated in the above steps until the maximum number of iterations is reached, where b represents the number of samples sampled from the experience pool.
6. A method for maximizing security and rate resource allocation in an unmanned aerial vehicle-assisted RSMA communication system according to claim 4, characterized in that, It also includes the step of calculating the reward function of the system. In DRL, by designing a reward function, difficult-to-optimize objectives are optimized through the accumulation of rewards. When an agent performs an action, it transitions to another state and receives a reward value. The reward function consists of: (1) Set a negative reward to penalize collisions between drones or sub-user rates falling below the threshold rate before training ends; (2) The setting of the reward function is not only related to the communication quality between the drone and the base station and the user, but also to the channel between the drone and the eavesdropper; in this setting, the drone will receive a positive reward for the user's safe rate according to the strategy selection. Reward=R S It further outputs the reward obtained from the environment at the current moment, changes the agent's state to the next moment, stores the learned samples in the experience replay pool, and saves the output parameters.
Citation Information
Patent Citations
Energy efficiency maximization method of unmanned aerial vehicle cooperative NOMA communication network
CN116170824A
Method for maximizing minimum safety rate in downlink RSMA unmanned aerial vehicle communication system
CN118055403A