Multi-relay unmanned aerial vehicle assisted RSMA communication system maximum security and rate resource allocation method
By using a multi-relay UAV-assisted RSMA communication system, and combining K-Mediods and QMIX algorithms to optimize UAV trajectories and resource allocation, the problems of secure communication and high-dimensional optimization in multi-UAV systems are solved, achieving efficient user communication and enhanced security.
Patent Information
- Application Number
- CN202510787542.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-08-15
AI Technical Summary
Existing research has failed to effectively optimize secure communication among multiple drone relays in drone communication systems, and traditional methods are difficult to solve the multi-agent trajectory optimization problem in high-dimensional variable spaces, leading to reduced user information rates and increased eavesdropping risks.
A multi-relay UAV-assisted RSMA communication system is adopted, which combines K-Mediods and QMIX multi-agent deep reinforcement learning algorithms to optimize UAV trajectory planning and resource scheduling. Safety and speed are maximized through precoding and interference management.
Under the constraints of multiple UAV flight trajectories and transmission power, high-efficiency user communication quality and security are achieved, the robustness and computational efficiency of the system are improved, and it can adapt to dynamic environmental changes.
Smart Images

Figure CN120499819A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of RSMA downlink systems, and in particular to a method for precoding and optimizing unmanned aerial vehicle (UAV) trajectories for maximizing security and rate in a rate division multiple access (RSMA) downlink system. Background Art
[0002] Faced with the highly diverse application demands of future wireless networks, drone-assisted communications (DAVs) are considered a promising technology due to their flexible operation and wide coverage. By acting as temporary base stations, drones can perform a wide range of activities, including disaster relief support, real-time surveillance, and data collection. However, due to the wide coverage area and large-scale connectivity provided by drones, open networks are vulnerable to eavesdropping and various malicious attacks. Since drones are often used to transmit private data, designing secure DAV-assisted systems is crucial.
[0003] With the advent of the next-generation mobile internet era, the access of massive IoT devices is bound to cause traffic congestion. RSMA (Rate-Splitting Multiple Access) is a multi-antenna, multi-user communication multiple access scheme based on rate splitting and linear precoding that multiplexes users in the power domain. RSMA divides user messages into public and private parts, encoding the public part into one or more public streams while encoding the private part into separate streams. These streams are precoded using the channel state information available at the transmitter. They are superimposed and transmitted via a Multiple-Input Multiple-Output (MIMO) or Multiple-Input Single-Output (MISO) channel. At the receiving end, all receivers decode the public stream, treat other users' messages as noise, perform Successive Interference Cancellation (SIC), and then decode their respective private streams. The main advantage of RSMA is the flexible interference management by allowing decoding of the interfering part.
[0004] Although some research has been conducted on the secure communication issues of RSMA-based UAV communication systems at home and abroad, and these studies have met the optimization goals of certain aspects of RSMA-based UAV communication systems to a certain extent, the existing research on the analysis methods and optimization methods for the secure communication issues of RSMA-based UAV communication systems still has certain limitations, including:
[0005] (1) When optimizing user safety rates, current research usually only considers the optimization problem of a single drone relay, but rarely considers multiple drones relaying and multiple drones clustering users to provide services separately.
[0006] (2) Due to reasons such as equipment accuracy, the eavesdropper may be more affected by environmental noise than the user. The method of encoding artificial noise with the user signal at the transmitting end will result in insufficient power allocated to the user and reduce the user's information rate.
[0007] (3) UAV trajectory optimization is usually a non-convex problem, which can be approximately converted into a convex problem and solved with the help of convex optimization tools. However, traditional optimization methods are difficult to solve multi-agent trajectory optimization problems in high-dimensional variable spaces. Summary of the Invention
[0008] The present invention aims to solve the above problems in the prior art. It proposes a method for maximizing security and rate resource allocation in a multi-relay UAV-assisted RSMA communication system. The technical solution of the present invention is as follows:
[0009] A method for maximizing security and rate resource allocation in a multi-relay UAV-assisted RSMA communication system comprises the following steps:
[0010] Step 1) Establish a multi-relay UAV-assisted RSMA communication system to maximize security and rate resource allocation model;
[0011] Step 2) Initialize the drone position and transmit power, base station position and transmit power, user service threshold, and total number of iterations;
[0012] Step 3) Associate the drone with the user based on K-Mediods;
[0013] Step 4) QMIX-based multi-agent deep reinforcement learning algorithm solves the joint problem of UAV trajectory planning and resource scheduling;
[0014] Step 5) determines whether the preset number of iterations has been reached. If so, the method ends. If not, the parameters are updated using the back propagation algorithm until the number of iterations is reached, at which point the algorithm ends.
[0015] According to the multi-relay drone-assisted RSMA communication system maximizing security and rate resource allocation method, it is characterized in that in the step 1), a model including U relay drones, a single jamming drone Jammer, a single eavesdropper, multiple confidentiality users and rate users is first established. A transmitting base station with N antennas needs to communicate with M single-antenna users with the help of a drone relay with K antennas. The M users are divided into U clusters, each cluster is provided with communication services by one drone, and each user can only be provided with communication services by one drone at most. At the same time, in time slot t, the set of users served by drone u can be expressed as F u (t)={1,2,...f uEve, the eavesdropper, eavesdrops on the communication information of a confidential user. A jamming drone, Jammer, is deployed in the air to emit artificial noise to interfere with the eavesdropping channel from the base station to Eve. Rate users without eavesdropping risk also communicate through the relay drone. By optimizing the public rate allocation while ensuring that the rate of such users is above the threshold, the Jammer and Relay trajectories maximize the secure communication rate of the eavesdropped user. Information transmission requires two stages: In the first time slot, the base station sends a broadcast signal: Where d is the data symbol and satisfies E[|d| 2 ]=1,w Ru The base station sends the precoding vector of the u-th drone, w c is the precoding vector for the base station to send public messages, and M is the total number of users. Therefore, the signal received from the base station at the relay is where w i ~CN(0,σ 2 ) is additive white Gaussian noise, σ 2 is the variance of the noise, W=[w c ,w1,w2,...w U ] represents the precoding matrix of the base station, P J To interfere with the UAV power, G BRu Expressed as the channel gain between the base station and the relay, h JR It is represented as the channel gain between the interference drone and the relay. In the second time slot, after the drone receives the signal, it first decodes the signal to obtain the relevant information of the associated user and then forwards it. In the second time slot, the relay uses the DF protocol to send the signal to the user. The signal received by user m is expressed as:
[0016]
[0017] Precoding vector for sending the mth user message to the relay drone The precoding vector for the relay drone to send public messages, G Ru,m It is expressed as the channel gain between the relay and user u. Ri,m represents the channel gain of other relay drones to user m, Represents the precoding vector transmitted by other relay drones.
[0018] According to step 1), the multi-relay UAV-assisted RSMA communication system maximizes security and rate resource allocation. The multi-relay UAV-assisted RSMA communication system maximizes security and rate model P1 is:
[0019]
[0020] C1c:X min≤X J (t)≤X max ,X min ≤X R (t)≤X max
[0021] C1d:Y min ≤Y J (t)≤Y max ,Y min ≤Y R (t)≤Y max
[0022] C1e:H min ≤H J (t)≤H max ,H min ≤H R (t)≤H max
[0023]
[0024] C1g:F min ≤f u (t)≤F max
[0025] in, is the security and rate from the base station to the user at time t. Indicates the long-term fairness of the system, represents the average security throughput of confidential user m in time t. T represents the total communication time. H , They represent the base station beamforming covariance matrix and the UAV beamforming covariance matrix respectively. B Indicates the base station transmission power, P R R represents the transmission power at the UAV. NS (t) represents the minimum transmission rate required by the rate user within time t. p,N Indicates the private rate of the rate user. N R c Indicates the weighted public rate of rate users. Indicates the minimum rate requirement of the rate user. X J (t) represents the horizontal x-coordinate position of the jammer UAV at time t, X R (t) represents the horizontal x-coordinate position of the relay drone Relay at time t. J (t) represents the horizontal y coordinate position of the jammer UAV at time t. R (t) represents the horizontal coordinate y position of the relay drone Relay at time t. J(t) represents the vertical coordinate z position of the jammer UAV at time t. R (t) represents the vertical coordinate z position of the relay drone Relay at time t. [X min ,X max ],[Y min ,Y max ],[H min ,H max ] indicates the flight space limit of the drone. R (t) represents the three-dimensional position of the relay drone Relay at time t. J (t) represents the three-dimensional position of the jammer UAV at time t. min Indicates the minimum distance to avoid collision. u (t) represents the number of users assisted by drones. [F min ,F max ] represents the range limit of the UAV-assisted user. W represents the precoding matrix at the base station and the UAV.
[0026] According to the principles of information theory, the signal rate from the base station to the user can be equivalent to the smaller value of the rate from the base station to the relay and the rate from the relay to the user. That is, R c (t) = min{a u R R,c (t),R′ c (t)},R' c (t) = min{R′ 1,c (t),...R′ f,c (t)},
[0027] R R,c (t) = min{R R,1,c (t),R R,2,c (t)…R R,u,c (t)},R m,p (t) = min{a u,m R R,u,p (t),R′ u,p (t)}, where a u is the weight of public information, where a u,m is the weight of user m’s private information in drone u’s private information.
[0028] in, represents the public flow rate between the UAV and user f.
[0029] Indicates the public flow rate between the base station and the drone.
[0030] Indicates the private flow rate between the base station and the drone.
[0031] represents the private flow rate between the UAV and user m.
[0032] For an eavesdropper, it can obtain confidential user information from the information transmitted by the transmitting base station or the message transmitted by the drone. The eavesdropping rate is less than or equal to the maximum value of the two methods. The eavesdropper's public flow and private flow can achieve the following rates:
[0033] R m,E (t)≤max{R 1,m,E a u,m (t),R 2,m,E (t)}, R c,E (t)≤max{R 1,c,E a u (t),R 2,c,E (t)}
[0034]
[0035] Indicates the eavesdropping public rate obtained from the base station.
[0036]
[0037] represents the private rate of user m obtained from the base station.
[0038]
[0039] Represents the eavesdropping public rate obtained from the drone.
[0040]
[0041] represents the private rate of user m obtained from the drone.
[0042] The optimization variables of problem P1 are the precoding matrix W at the base station and the UAV, the UAV trajectory qR, and the interfering UAV trajectory qJ. The objective function of P1 is It represents the long-term security and rate of the system confidential users, constraint C1a is the precoding matrix power constraint, C1b is the minimum transmission rate requirement of the secondary user; C1c-C1d represent the level of drone flight area restriction; C1f is the restriction for drone to avoid collision; C1g is the restriction on the number of users served by the drone.
[0043] The proposed method for maximizing security and rate resource allocation in a multi-relay UAV-assisted RSMA communication system is characterized by initializing the UAV position and transmit power, base station position and transmit power, user service threshold, and total number of iterations in step 2). The problem is then converted into an MDP problem, and the objective function is maximized using a QMIX-based DRL algorithm, converting the problem into an MDP formulation. The state space S comprises the UAV's three-dimensional spatial position, user position, base station position, eavesdropper position, and user quality of service. The observation space O comprises other UAVs, users and their positions, some eavesdropper positions, and user quality of service. The action space A comprises the UAV flight distance vector and the power allocation coefficient increase or decrease. Rewards and penalties: In the system, agents adopt a cooperative game mechanism, with all agents sharing rewards. Furthermore, to ensure system constraints, if a UAV's action causes the inter-UAV distance to fall below the safe distance or fails to provide the minimum quality of service, the agent is penalized.
[0044] The proposed method for maximizing security and rate resource allocation in a multi-relay UAV-assisted RSMA communication system is characterized by the following: in step 3), drones are associated with users based on the K-Mediods method: First, K user groups are determined based on the number of users carried by the drone. The user location information is input into the drone's clustering algorithm. K locations are then randomly selected from the users as initial cluster centers. The distance between each user and the K cluster centers is then calculated, and each user is assigned to the cluster center closest to it, forming a cluster. Based on the number of users in each group, it is determined whether the maximum number of users that can be carried by the drone, MAX, is exceeded. If U (the number of users in the current group), exceeds MAX, the grouping needs to be adjusted. Some users farthest from the current cluster center are removed from the group. Then, users are reallocated: for each user, it is removed from the group of the current cluster center and reallocated to the cluster center closest to it, generating a new cluster. Steps 4 and 5 are repeated until the clustering is stable, i.e., no user's location changes. According to the proposed method for maximizing security and rate resource allocation for a multi-relay UAV-assisted RSMA communication system, it is characterized in that in step 4), a QMIX-based multi-agent deep reinforcement learning algorithm is used to solve the UAV trajectory planning and resource scheduling sub-problems.
[0045] The specific steps are as follows: First, the environment is initialized, resetting the positions of the user and eavesdropper. An agent is selected and the agent space is linked. When the flight training time is less than Nt, step 3 is executed. The base station sends a message to the relay drone, and the jammer drone emits artificial noise. The relay drone then encodes the message and sends it to the associated user. The jammer drone emits artificial noise, and the user and eavesdropper decode the signal received from the relay drone. When the flight time expires, the agent selects an action based on a greedy strategy to obtain a new reward. The experience (Φ(s), a, r, Φ(s')) is cached, and data is randomly sampled from the cache. The TD error is calculated using a mixing network and updated using the backpropagation algorithm.
[0046] The proposed method for maximizing security and rate resource allocation in a multi-relay UAV-assisted RSMA communication system is characterized in that, in step 5), a determination is made as to whether a preset number of iterations has been reached. If so, the method terminates. If not, parameters are updated using a backpropagation algorithm until the number of iterations is reached, at which point the algorithm terminates.
[0047] The advantages and beneficial effects of the present invention are as follows:
[0048] This invention maximizes the sum rate of a drone-assisted RSMA communication system under constraints such as multi-UAV flight trajectories, transmit power, base station transmit power, and secondary user rate. By taking into account the resolution of UAV trajectories within the flight cycle, the transmit power allocation between the base station and UAVs is optimized, offering greater practical value compared to solving for system security and rate based on maximum transmit power. Compared to previous approaches to wireless resource management, this invention utilizes deep reinforcement learning. In a dynamic environment where UAV positions vary, traditional convex optimization methods require recalculating resource allocation schemes every time the environment changes. This wastes significant computational resources as the frequency of environmental changes increases. Furthermore, changes in UAV positions will cause changes in the direct channel power gain between the UAV and user channels. The present invention is beneficial in that the RSMA protocol can achieve high-quality communication for a large number of users while utilizing deep reinforcement learning for optimization, resulting in enhanced robustness. Furthermore, the invention utilizes multi-agent deep reinforcement learning, resulting in faster convergence time and better convergence performance than other single-agent algorithms. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 The present invention provides a preferred embodiment of a multi-relay UAV-assisted RSMA communication system model;
[0050] Figure 2This is a comparison chart of the average user security rate of the RSMA of the present invention and the comparison methods (NOMA and SDMA) at different numbers of iterations;
[0051] Figure 3 is the average safety rate of primary users, long-term fairness index, and average reward graph of secondary user interruption probability in the last 50 rounds of the present invention and the comparison method (non-interference drone model);
[0052] Figure 4 It is the Qos diagram of the present invention and the comparative method (non-interference drone model and drone straight flight);
[0053] Figure 5 It is the UAV trajectory diagram of the present invention;
[0054] Figure 6 It is a schematic flow diagram of the present invention. DETAILED DESCRIPTION
[0055] The following will describe the technical solutions in the embodiments of the present invention in detail with reference to the accompanying drawings. The described embodiments are only a part of the embodiments of the present invention.
[0056] The technical solution of the present invention to solve the above technical problems is:
[0057] The present invention Figure 6 A method for maximizing security and rate resource allocation in a multi-relay UAV-assisted RSMA communication system is disclosed, which comprises the following steps:
[0058] Step 1: Establish a resource allocation model to maximize security and rate for a multi-relay UAV-assisted RSMA communication system.
[0059] Step 2: Initialize the drone position and transmit power, base station position and transmit power, power allocation coefficient, user service threshold, and total number of iterations;
[0060] Step 3: Associate drones with users based on K-Mediods;
[0061] Step 4: QMIX-based multi-agent deep reinforcement learning algorithm solves the joint problem of drone trajectory planning and resource scheduling;
[0062] Step 5: Determine whether the preset number of iterations has been reached. If so, the method ends. If not, update the parameters using the backpropagation algorithm until the number of iterations is reached, and the algorithm ends.
[0063] Furthermore, in the first step, the multi-relay UAV-assisted RSMA communication system maximizes security and rate resource allocation model (P1) as follows:
[0064]
[0065] C1c:X min ≤X J (t)≤X max ,X min ≤X R (t)≤X max
[0066] C1d:Y min ≤Y J (t)≤Y max ,Y min ≤Y R (t)≤Y max
[0067] C1e:H min ≤H J (t)≤H max ,H min ≤H R (t)≤H max
[0068]
[0069] C1g:F min ≤f u (t)≤F max
[0070] in, is the security and rate from the base station to the user at time t. Indicates the long-term fairness of the system, represents the average security throughput of confidential user m in time t. T represents the total communication time. H , They represent the base station beamforming covariance matrix and the UAV beamforming covariance matrix respectively. B Indicates the base station transmission power, P R R represents the transmission power at the UAV. NS (t) represents the minimum transmission rate required by the rate user within time t. p,N Indicates the private rate of the rate user. N R c Indicates the weighted public rate of rate users. Indicates the minimum rate requirement of the rate user. X J (t) represents the horizontal x-coordinate position of the jammer UAV at time t, X R (t) represents the horizontal x-coordinate position of the relay drone Relay at time t. J(t) represents the horizontal y coordinate position of the jammer UAV at time t. R (t) represents the horizontal coordinate y position of the relay drone Relay at time t. J (t) represents the vertical coordinate z position of the jammer UAV at time t. R (t) represents the vertical coordinate z position of the relay drone Relay at time t. [X min ,X max ],[Y min ,Y max ],[H min ,H max ] indicates the flight space limit of the drone. R (t) represents the three-dimensional position of the relay drone Relay at time t. J (t) represents the three-dimensional position of the jammer UAV at time t. min Indicates the minimum distance to avoid collision. u (t) represents the number of users assisted by drones. [F min ,F max ] represents the range limit of the UAV-assisted user. W represents the precoding matrix at the base station and the UAV.
[0071] According to the principles of information theory, the signal rate from the base station to the user can be equivalent to the smaller value of the rate from the base station to the relay and the rate from the relay to the user. That is, R c (t) = min{a u R R,c (t),R′ c (t)},R' c (t) = min{R′ 1,c (t),...R′ f,c (t)},
[0072] R R,c (t) = min{R R,1,c (t),R R,2,c (t)…R R,u,c (t)},R m,p (t) = min{a u,m R R,u,p (t),R′ u,p (t)}, where a u is the weight of public information, where a u,m is the weight of user m’s private information in drone u’s private information.
[0073] in, represents the public flow rate between the UAV and user f.
[0074] Indicates the public flow rate between the base station and the drone.
[0075] Indicates the private flow rate between the base station and the drone.
[0076] represents the private flow rate between the UAV and user m.
[0077] For an eavesdropper, it can obtain confidential user information from the information transmitted by the transmitting base station or the message transmitted by the drone. The eavesdropping rate is less than or equal to the maximum value of the two methods. The eavesdropper's public flow and private flow can achieve the following rates:
[0078] R m,E (t)≤max{R l,m,E a u,m (t),R 2,m,E (t)}, R c,E (t)≤max{R l,c,E a u (t),R 2,c,E (t)}
[0079]
[0080] Indicates the eavesdropping public rate obtained from the base station.
[0081]
[0082] represents the private rate of user m obtained from the base station.
[0083]
[0084] Represents the eavesdropping public rate obtained from the drone.
[0085]
[0086] represents the private rate of user m obtained from the drone.
[0087] The optimization variables of problem P1 are the precoding matrix W at the base station and the UAV, the UAV trajectory qR, and the interfering UAV trajectory qJ. The objective function of P1 is It represents the long-term security and rate of the system confidential users, constraint C1a is the precoding matrix power constraint, C1b is the minimum transmission rate requirement of the secondary user; C1c-C1d represent the level of drone flight area restriction; C1f is the restriction for drone to avoid collision; C1g is the restriction on the number of users served by the drone.
[0088] Furthermore, in the step 2), the drone position and transmission power, base station position and transmission power, power allocation coefficient, user service threshold, and total number of iterations are initialized. In the model, the initial conditions are set to BS coordinates (4000, 4000), and the user and eavesdropper are at a random position within the range. In this way, an eavesdropper hidden in a suspicious area and randomly distributed legitimate users in the environment are simulated. Similarly, the initial position of the drone relay and drone interference is a random position within the range of (4000±200, 4000±200), which means that the drone is controlled to take off through the base station. The environmental boundary is set to 8000×8000, the maximum height of the drone is 600m, the base station power is 30dbm, the interference drone power is 15dbm, and the total number of iterations is 400. The environmental noise is 10 -17 w.
Claims
1. A method for maximizing security and rate resource allocation in a multi-relay UAV-assisted RSMA communication system, characterized in that: The following steps are involved: Step 1) Establish a multi-relay UAV-assisted RSMA communication system to maximize security and rate resource allocation model; Step 2) Initialize the drone position and transmit power, base station position and transmit power, user service threshold, and total number of iterations; Step 3) Associate the drone with the user based on K-Mediods; Step 4) QMIX-based multi-agent deep reinforcement learning algorithm solves the joint problem of UAV trajectory planning and resource scheduling; Step 5) determines whether the preset number of iterations has been reached. If so, the method ends. If not, the parameters are updated using the back propagation algorithm until the number of iterations is reached, at which point the algorithm ends.
2. The method for maximizing security and rate resource allocation in a multi-relay UAV-assisted RSMA communication system according to claim 1, characterized in that: In step 1), a model is first established that includes U relay drones, a single jammer drone, a single eavesdropper, multiple confidential users, and rate users. A transmitting base station with N antennas needs to communicate with M single-antenna users via a drone relay with K antennas. The M users are divided into U clusters, each cluster is provided with communication services by one drone, and each user can only be provided with communication services by one drone at most. At the same time, within time slot t, the set of users served by drone u can be expressed as F u (t)={1,2,...f u Eve, the eavesdropper, eavesdrops on the communication information of a confidential user. A jamming drone, Jammer, is deployed in the air to emit artificial noise to interfere with the eavesdropping channel from the base station to Eve. Rate users without the risk of eavesdropping also communicate through the relay drone. By optimizing the public rate allocation while ensuring that the rate of such users is above the threshold, the Jammer and Relay trajectories maximize the secure communication rate of the eavesdropped user. Information transmission requires two stages: In the first time slot, the base station sends a broadcast signal: Where d is the data symbol and satisfies E[|d| 2 ]=1,w Ru The base station sends the precoding vector of the u-th drone, w c is the precoding vector for the base station to send a public message, and M is the total number of users; therefore, the signal received from the base station at the relay is where w i ~CN(0,σ 2 ) is additive white Gaussian noise, σ 2 is the variance of the noise, W=[w c ,w1,w2,...w U ] represents the precoding matrix of the base station, P J To interfere with the UAV power, G BRu Expressed as the channel gain between the base station and the relay, h JR It is represented by the channel gain between the interference drone and the relay. In the second time slot, after the drone receives the signal, it first decodes the signal to obtain the relevant information of the associated user, and then forwards it. In the second time slot, the relay uses the DF protocol to send the signal to the user. The signal received by user m is expressed as: Precoding vector for sending the mth user message to the relay drone The precoding vector for the relay drone to send public messages, G Ru,m It is expressed as the channel gain between the relay and user u; G Ri,m represents the channel gain of other relay drones to user m, Represents the precoding vector transmitted by other relay drones; The maximum security and rate model P1 of the multi-relay UAV-assisted RSMA communication system is: C1c:X min ≤X J (t)≤X max ,X min ≤X R (t)≤X max C1d:Y min ≤Y J (t)≤Y max ,Y min ≤Y R (t)≤Y max C1e:H min ≤H J (t)≤H max ,H min ≤H R (t)≤H max C1g:F min ≤f u (t)≤F max in is the security and rate from the base station to the user at time t; Indicates the long-term fairness of the system, represents the average security throughput of confidential user m in time t; T represents the total communication time; ww H , They represent the base station beamforming covariance matrix and the UAV beamforming covariance matrix respectively; the base station sends P B Indicates the base station transmission power, P R Represents the transmitting power at the UAV; R NS (t) represents the minimum transmission rate required by the rate user within time t; R p,N Indicates the private rate of the rate user; a N R c represents the weighted public rate of rate users; Indicates the minimum rate requirement of the rate user; X J (t) represents the horizontal x-coordinate position of the jammer UAV at time t, X R (t) represents the horizontal x-coordinate position of the relay drone Relay at time t; Y J (t) represents the horizontal y coordinate position of the jammer UAV at time t; Y R (t) represents the horizontal coordinate y position of the relay drone Relay at time t; H J (t) represents the vertical coordinate z position of the jammer UAV at time t; H R (t) represents the vertical coordinate z position of the relay drone Relay at time t; [X min ,X max ],[Y min ,Y max ],[H min ,H max ] indicates the flight space limit of the drone; q R (t) represents the three-dimensional position of the relay drone Relay at time t; q J (t) represents the three-dimensional position of the jamming UAV Jammer at time t; l min Indicates the minimum distance to prevent collision; f u (t) represents the number of users assisted by drones; [F min ,F max ] represents the range limit of the UAV-assisted user; W represents the precoding matrix at the base station and the UAV; According to the principles of information theory, the signal rate from the base station to the user can be equivalent to the smaller value of the rate from the base station to the relay and the rate from the relay to the user; that is, R c (t)=min{a u R R,c (t),R′ c (t)},R′ c (t)=min{R′ 1,c (t),...R′ f,c (t)}, R R,c (t) = min{R R,1,c (t), R R,2,c (t)…R R,u,c (t)}, R m,p (t) = min{a u,m R R,u,p (t), R' u,p (t)}, where a u is the weight of public information, where a u,m is the weight of user m’s private information in drone u’s private information; in, represents the public flow rate between the UAV and user f; Indicates the public flow rate between the base station and the drone; Indicates the private flow rate between the base station and the drone; represents the private flow rate between the UAV and user m; For an eavesdropper, it can obtain confidential user information from the information transmitted by the transmitting base station or the message transmitted by the drone. The eavesdropping rate is less than or equal to the maximum value of the two methods. The eavesdropper's public flow and private flow can achieve the following rates: R m,E (t)≤max{R l,m,E a u,m (t),R 2,m,E (t)},R c,E (t)≤max{R l,c,E a u (t),R 2,c,E (t)} represents the eavesdropping public rate obtained from the base station; represents the private rate of user m obtained from the base station; represents the eavesdropping public rate obtained from the drone; represents the private rate of user m obtained from the drone; The optimization variables of problem P1 are the precoding matrix W at the base station and the UAV, the UAV trajectory qR, and the interfering UAV trajectory qJ. The objective function of P1 is It represents the long-term security and rate of the system confidential users, constraint C1a is the precoding matrix power constraint, C1b is the minimum transmission rate requirement of the secondary user; C1c-C1d represent the level of drone flight area restriction; C1f is the restriction for drone to avoid collision; C1g is the restriction on the number of users served by the drone.
3. The method for maximizing security and rate resource allocation in a multi-relay UAV-assisted RSMA communication system according to claim 2, characterized in that: In the step 2), the drone position and transmission power, the base station position and transmission power, the user service threshold, and the total number of iterations are initialized; the problem is converted into an MDP problem, and the DRL algorithm based on QMIX is used to maximize the objective function, and the problem is converted into an MDP formula; the state space S includes the three-dimensional spatial position of the drone, the user position, the base station position, the eavesdropper position, and the user service quality; the observation space O includes other drones, users and positions, the positions of some eavesdroppers, and the user service quality; the action space A includes the drone flight distance vector and the increase or decrease of the power allocation coefficient; rewards and penalties: in the system, the intelligent agents adopt a cooperative game mechanism, and all intelligent agents can share rewards; at the same time, in order to ensure system constraints, if the drone's action causes the distance between drones to be less than the safe distance or the minimum quality service cannot be provided, the intelligent agent will be punished.
4. The method for maximizing security and rate resource allocation in a multi-relay UAV-assisted RSMA communication system according to claim 3, characterized in that: In step 3), the drones are associated with users based on the K-Mediods method: first, K user groups are determined based on the number of items carried by the drones; the user location information is input into the drone clustering algorithm; and then K locations are randomly selected from the users as initial cluster centers; Then calculate the distance between each user and the K cluster centers, and assign each user to the cluster center closest to it to form a cluster; based on the number of users in each group, determine whether it exceeds the maximum number MAX that the drone can carry; If U (the number of users in the current group) exceeds MAX, the grouping needs to be adjusted. Some users farthest from the current cluster center are removed from the group. Then, users are redistributed. For each user, they are removed from the group at the current cluster center and redistributed to the cluster center closest to them to generate a new cluster. Continue steps 4 and 5 until the cluster is stable, that is, no user's location changes; The method for maximizing security and rate resource allocation in a multi-relay UAV-assisted RSMA communication system is characterized in that, in step 4), a QMIX-based multi-agent deep reinforcement learning algorithm is used to solve the sub-problems of UAV trajectory planning and resource scheduling; The specific steps are as follows: first, initialize the environment and reset the positions of the user and eavesdropper; select the agent and link the agent space at the same time. When the flight training time is less than Nt, execute step 3), and the base station sends information to the relay drone, interfering with the drone to emit artificial noise. After that, the relay drone encodes and sends information to the associated user, interfering with the drone to emit artificial noise, and the user and eavesdropper decode and receive the signal sent by the relay drone; when the flight time ends, the agent selects the agent action according to the greedy strategy to obtain a new reward value; and stores the experience (Φ(s), a, r, Φ(s')) in the cache, randomly samples some data from the cache, passes through the Mixing network and calculates the TD error, and updates it through the backpropagation algorithm.
5. The method for maximizing security and rate resource allocation in a multi-relay UAV-assisted RSMA communication system according to claim 4, characterized in that: In step 5), it is determined whether the preset number of iterations has been reached. If so, the method ends. If not, the parameters are updated by the back propagation algorithm until the number of iterations is reached and the algorithm ends.