Physical layer security enhancement method for cooperative communication network
By building a multi-UAV downlink secure communication system with collaborative multi-point transmission technology, combined with Rice fading channel modeling and deep reinforcement learning algorithm, the problem of difficulty in taking into account security in resource allocation and trajectory planning in the UAV communication network is solved, and the system's confidentiality performance and dynamic adaptability are improved.
Patent Information
- Application Number
- CN202510459393.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-04
AI Technical Summary
In existing drone communication networks, resource allocation and trajectory planning in the dynamic environment of multiple drones are difficult to take into account security, traditional encryption technology is difficult to apply, and traditional collaborative multi-point technology cannot adapt to the motion characteristics of drones, resulting in a decrease in communication link reliability and an increase in the risk of eavesdropping.
A multi-UAV downlink secure communication system based on collaborative multi-point transmission technology is built, and the confidentiality rate is calculated through Rice fading channel modeling is split into beamforming optimization and drone trajectory optimization sub-problems. SCA and SDR algorithms are used to solve beamforming, and the drone trajectory planning is combined with the improved multi-agent deep reinforcement learning algorithm RES-QMIX to maximize the system's average confidentiality rate.
The confidentiality and dynamic adaptability of multi-UAV communication systems have been improved. Through collaborative multi-point transmission technology and beamforming technology, the system degree of freedom and spectrum efficiency have been enhanced, the eavesdropping signals have been suppressed, and the communication security and efficiency of legitimate users have been improved.
Smart Images

Figure CN120263232A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of UAV communication security, and particularly to a method for enhancing physical layer security for a cooperative communication network. Background Art
[0002] With the rapid development of 5G and 6G technologies, UAVs have been widely used in the fields of smart cities, emergency rescue, security monitoring, etc. due to their flexible deployment, low cost, and strong adaptability. By building line-of-sight links, UAVs can bring more economical and flexible communication solutions for wireless communication, supporting ubiquitous communication and computing services.
[0003] In existing UAV communication networks, security issues are particularly prominent: First, due to the limited power and coverage of a single UAV, it is easy to cause fluctuations in the user channel state and a decrease in the reliability of the communication link, providing potential attack opportunities for eavesdroppers; Second, limited by the computing resources and high mobility of UAVs, traditional encryption technologies are difficult to effectively apply to UAV communication networks; Third, traditional coordinated multi-point technologies rely on fixed base stations and cannot adapt to the motion characteristics of UAVs; Fourth, in the scenario of multi-UAV cooperation, the real-time requirements of dynamic trajectory planning and resource allocation make it difficult for traditional static optimization methods to meet the solution requirements. Therefore, there is an urgent need for a method for enhancing physical layer security for a cooperative communication network, which can ensure its communication security by jointly optimizing the trajectories and beamforming of UAVs. Summary of the Invention
[0004] The present invention provides a method for enhancing physical layer security for a cooperative communication network, mainly aiming to solve the problem that it is difficult to balance security in resource allocation and trajectory planning in the dynamic environment of multiple UAVs in the prior art.
[0005] The present invention provides a method for enhancing physical layer security for a cooperative communication network, including:
[0006] Step 1: Obtain the relevant configuration information of the UAV base station and ground users, and construct a multi-UAV downlink secure communication system based on the coordinated multi-point transmission technology according to the relevant configuration information;
[0007] Step 2: Use the Rice fading channel to model the links from multiple UAVs to ground users and eavesdroppers, and calculate the secrecy rate of the system according to the achievable data rate of the users and the eavesdropping rate of the eavesdroppers, as the research index of physical layer security;
[0008] Step 3: Based on the secrecy rate of the system, construct a mathematical optimization model for enhancing physical layer security for a cooperative communication network, and split the mathematical optimization model into a beamforming optimization sub-problem and a UAV trajectory optimization sub-problem;
[0009] Step 4: Based on the SCA and SDR algorithms, iteratively solve the beamforming optimization sub-problem, dynamically obtain the optimal beamforming method, and simultaneously obtain an approximate optimal solution of the beamforming matrix;
[0010] Step 5: Transform the UAV trajectory optimization sub-problem into a Markov decision process, calculate the reward function according to the approximate optimal solution of the beamforming matrix obtained in Step 4, and use the improved multi-agent deep reinforcement learning algorithm RES-QMIX to solve this sub-problem to achieve real-time UAV trajectory planning in complex environments;
[0011] Step 6: Substitute the UAV positions optimized in Step 5 into Step 4, repeat Step 4, and through the alternating optimization algorithm, combine the above two sub-problem solving methods to maximize the average secrecy rate of the system, and perform UAV trajectory planning and beamforming according to the optimal solution to achieve physical layer security enhancement in the communication network.
[0012] A physical layer security enhancement method for a cooperative communication network according to the present invention has at least the following beneficial effects:
[0013] (1) The method of the present invention is based on a multi-UAV downlink secure communication system assisted by cooperative multi-point transmission technology, and improves the system degrees of freedom and spectral efficiency through multiple input single output and beamforming technologies.
[0014] (2) The method of the present invention utilizes physical layer security technology during the communication process, processes the signal through the channel characteristics, forms a strong signal gain in the direction of the target user, and simultaneously suppresses the eavesdropping signal to reduce the eavesdropping risk from the source.
[0015] (3) The present invention adopts a multi-UAV cooperative transmission scheme, combines the CoMP-JP joint processing technology, supports the user equipment to simultaneously receive signals from multiple UAVs, and improves the data transmission rate and signal reliability.
[0016] (4) The present invention transforms the UAV trajectory optimization problem into a Markov decision process, uses the improved multi-agent deep reinforcement learning algorithm RES-QMIX for UAV trajectory optimization to achieve real-time trajectory planning in complex environments. And based on the successive convex approximation and semi-definite relaxation algorithms, iteratively solve the beamforming optimization sub-problem to obtain an approximate optimal solution. Through the alternating optimization algorithm, jointly optimize the UAV trajectory and beamforming to maximize the average secrecy rate of the system.
[0017] (5) The present invention significantly improves the secrecy performance and dynamic adaptability of the multi-UAV communication system, and provides a more secure and efficient communication service for legitimate users. Description of the Drawings
[0018] Figure 1It is a flowchart of a physical layer security enhancement method for a cooperative communication network according to the present invention;
[0019] Figure 2 It is a schematic diagram of a multi-UAV downlink secure communication system based on cooperative multi-point transmission technology;
[0020] Figure 3 It is a flowchart of a beamforming optimization algorithm based on SCA-SDR provided by an embodiment of the present invention;
[0021] Figure 4 It is a flowchart of a Gaussian randomization method for recovering a rank-one solution provided by an embodiment of the present invention;
[0022] Figure 5 It is a flowchart of UAV trajectory optimization based on RES-QMIX and SCA-SDR provided by an embodiment of the present invention;
[0023] Figure 6 It is a schematic diagram of an alternating optimization framework for joint UAV trajectory and beamforming provided by an embodiment of the present invention;
[0024] Figure 7 It shows the convergence comparison results of different trajectory algorithms provided by an embodiment of the present invention;
[0025] Figure 8 It shows the convergence comparison results of trajectories under different flight time constraints provided by an embodiment of the present invention;
[0026] Figure 9 It shows the comparison results of two different UAV trajectory algorithms provided by an embodiment of the present invention;
[0027] Figure 10 It shows the convergence performance comparison results of the SCA-SDR beamforming optimization algorithm provided by an embodiment of the present invention;
[0028] Figure 11 It shows the comparison results of the relationship between the system average secrecy rate and the flight altitude and the number of antennas provided by an embodiment of the present invention;
[0029] Figure 12 It shows the comparison results of the relationship between the system average secrecy rate and the maximum transmit power and the number of users provided by an embodiment of the present invention;
[0030] Figure 13 It shows the comparison results of the relationship between the system average secrecy rate and the minimum secrecy rate and the number of UAVs provided by an embodiment of the present invention. Detailed implementation manners
[0031] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0032] An embodiment of the present invention provides a method for enhancing physical layer security for a cooperative communication network, as Figure 1 shown, the method includes:
[0033] Step 1: Obtain relevant configuration information of the UAV base station and ground users, and construct a multi-UAV downlink secure communication system based on the coordinated multi-point transmission technology.
[0034] In an embodiment of the present invention, a multi-UAV downlink secure communication system based on the coordinated multi-point transmission technology is constructed, and its architecture diagram is as Figure 2 shown. This system constructs an air-to-ground multiple-input single-output communication scenario, which includes U UAVs equipped with multi-antenna arrays as aerial base stations, and broadcasts confidential signals to N legitimate users in the presence of E suspicious eavesdroppers. Specifically, the UAV set is defined as The legitimate user set is defined as The eavesdropper set is defined as ε = {1, 2,..., E}. Each UAV is equipped with Q antennas and uses beamforming technology for communication, while ground users and eavesdroppers are single-antenna devices. It is assumed that the UAVs can fully obtain the location information of legitimate users and eavesdroppers through an appropriate information exchange mechanism, so as to support subsequent beamforming and cooperative transmission. In the system, since all UAVs share the same spectrum resource, cell-edge users may face relatively high inter-channel interference, resulting in a decline in communication quality. To solve this problem, the present invention introduces the coordinated multi-point transmission technology and adopts a downlink joint processing mechanism. Through the coordinated multi-point joint transmission (CoMP-Joint Transmission, CoMP-JT) technology, multiple UAV base stations cooperate to serve ground legitimate users, convert interference signals into useful signals, and significantly improve communication quality and reliability. In addition, the coordinated multi-point (Coordinated Multi-Point, CoMP) transmission technology further enhances physical layer security through cooperative beamforming and UAV power allocation, effectively suppressing the signal reception quality of eavesdroppers. In the system, through cooperative beamforming optimization, a strong signal gain is formed in the direction of the target user, the signal reception quality of eavesdroppers is suppressed, and physical layer security is improved; multiple-input single-output and beamforming technologies are adopted to improve the system degrees of freedom and spectral efficiency.
[0035] Step 2: Use the Rice fading channel to model the links from multiple UAVs to the ground users and eavesdroppers. Calculate the secrecy rate of the system based on the achievable data rate of the users and the eavesdropping rate of the eavesdroppers, which is used as an indicator for physical layer security research.
[0036] Since UAVs are usually limited by battery capacity, this invention focuses on a specific communication cycle with a duration of K seconds. For ease of processing, the total flight time is discretized into T equal-length time slots, i.e., the time slot length △T = K / T, and the set of time slots is denoted as Use a three-dimensional Cartesian coordinate system, and all dimensions are in meters. Assume the UAV flies at a fixed height H, which is the lowest flight height to avoid ground obstacles. At time slot t, the position of UAV u is denoted as l u (t), the position of user n is denoted as l n (t), and the position of eavesdropper e is denoted as l e (t). During the duration of K seconds, the trajectory of UAV u can be approximately represented as a sequence of length T {l u (t)}.
[0037] In the downlink transmission scenario, set the UAV as the legitimate transmitter and the ground node as the legitimate receiver. At time slot t, the Euclidean distance from UAV u to user n is denoted as d u,n (t), and the Euclidean distance from UAV u to eavesdropper e is denoted as d u,e (t). At time slot t, the Euclidean distances from UAV u to user n and from UAV u to eavesdropper e can be respectively expressed as:
[0038]
[0039] In the UAV-to-ground communication scenario, the line-of-sight transmission link usually dominates, and direct link communication can be guaranteed. At this time, the amplitude of one multipath component of the signal is much higher than that of other components, and the Rice distribution is applicable. Therefore, the communication links from the UAV to the ground user and from the UAV to the eavesdropper can be modeled as Rice fading channel models. According to the 3GPP urban macrocell technical specification, the path loss model is:
[0040] PL(d u,n (t)) = 28 + 22log 10 (d u,n (t)) + 20log 10 (f c )
[0041] PL(d u,e (t)) = 28 + 22log 10 (d u,e (t)) + 20log 10 (fc )
[0042] Among them, f c represents the carrier frequency (unit: GHz).
[0043] Step 2.1: Use the Rician fading channel to model the links from multiple UAVs to the ground user and the eavesdropper, and synthesize the channel gain through the weighted combination of the line-of-sight component and the non-line-of-sight component:
[0044] The channel gain of the direct link from UAV u to ground user n is:
[0045]
[0046] The channel gain of the direct link from UAV u to eavesdropper e is:
[0047]
[0048] Among them, ζ represents the Rician factor, and both represent the line-of-sight components of the link, and both represent the non-line-of-sight components of the link.
[0049] Through the coordinated multi-point transmission technology, each user can be served by all cooperative UAV base stations in the CoMP-JT mode. Therefore, the received signal of the legitimate user n can be expressed as:
[0050]
[0051] Among them, represents the beamforming vector from UAV u to user n at time slot t, and s n (t) represents the target signal of user n, satisfying The first term on the right side of the equation is the target signal of the user, the second term is the interference signal received by the user, and the third term is the additive Gaussian white noise power, which follows the distribution
[0052] Similarly, the received signal of each eavesdropper e at time slot t is:
[0053]
[0054] Among them, is the additive Gaussian white noise power.
[0055] Step 2.2: The signal-to-interference-plus-noise ratio of user n at time slot t is expressed as:
[0056]
[0057] The signal-to-interference-plus-noise ratio of the eavesdropper e in time slot t is expressed as:
[0058]
[0059] Wherein, represents the beamforming vector from the UAV u to the user n in time slot t; and are the additive white Gaussian noise power; N is the total number of legitimate users, and U is the total number of UAVs.
[0060] According to the Shannon formula, the achievable data rate of user n and the eavesdropping rate of eavesdropper e are respectively:
[0061] R n (t)=log2(1 + γ n (t))
[0062]
[0063] Wherein, R n (t) is the data transmission rate of the legitimate user on the ground, the data transmission rate of the eavesdropper;
[0064] Then the secrecy rate of user n is:
[0065]
[0066] Wherein, R E (t) is the maximum eavesdropping rate, and the operator
[0067] Then the total secrecy rate of the system is:
[0068]
[0069] Then the average secrecy rate of the system is:
[0070]
[0071] Wherein, T is the total number of time slots.
[0072] Step 3: Based on the secrecy rate of the system, construct a physical layer security enhancement mathematical optimization model for the cooperative communication network, and split the mathematical optimization model into a beamforming optimization sub-problem and a UAV trajectory optimization sub-problem.
[0073] In the embodiment of the present invention, the energy of the UAV is mainly used for signal transmission and flight. Denote the maximum transmission power of the UAV u as To ensure fairness among users, the secrecy rate of each user needs to be greater than the secrecy rate threshold R min。The flight speed of the UAV is fixed at V ≤ V max , and the total flight time is discretized into T equal-length time slots. To maximize the system performance gain, each UAV adaptively adjusts its forward direction within each time slot. The UAVs fly in a low-altitude plane at a fixed height H, and its horizontal flight area is defined by the range of the horizontal axis and the range of the vertical axis. According to the maximum speed limit of the UAV, the moving distance of the UAV between adjacent time slots shall not exceed the maximum flight distance D = V max τ within a single time slot. Assume that the starting and ending positions of the UAVs are fixed, denoted as l u (1) and l u (T) respectively. In the multi-UAV cooperative transmission scenario, collisions between UAVs need to be avoided. For this purpose, the UAVs are equipped with on-board sensors, the sensing range radius of which is d s , and the minimum safety distance between UAVs is d min . To ensure that the UAVs can detect surrounding obstacles in a timely manner, it is necessary to satisfy d min < d s .
[0074] Step 3.1: Taking the trajectory of the UAV and the active beamforming matrix of the UAV as optimization variables, by jointly optimizing the trajectory and the active beamforming matrix, the optimization goal of maximizing the system average secrecy rate is achieved, and the expression is as follows:
[0075]
[0076] where R min is the minimum secrecy rate set, is the maximum value of the UAV transmission power; x u (t), y u (t) represent the position coordinates of UAV u; the two-dimensional area range of the flight of UAV u is represented by ; l u (t), l i (t) and l j (t) are the positions of UAV u, i, and j at time slot t respectively.
[0077] Among them, the constraints are, from top to bottom in turn: the constraint to ensure system fairness, ensuring that the secrecy rate of each user is not lower than the preset minimum value; the maximum communication power constraint of each UAV, restricting its transmission power not to exceed the preset maximum value; the constraint representing the flight range limit of the UAV in the horizontal direction of the map, restricting it to fly within the preset geographical boundary; the maximum flight distance constraint of the UAV within a time slot, and the constraint representing the need to maintain a safe distance between any two UAVs at any time to avoid collisions.
[0078] The above optimization problem is a typical non-convex optimization problem, and its complexity is mainly reflected in the following aspects. First, due to the operator [·] + , the objective function is non-smooth and difficult to handle directly. Second, the difference of logarithmic functions of the user communication rate and the eavesdropper's eavesdropping rate involved in the objective function and the constraint conditions results in the non-concavity of the objective function and the non-convexity of the constraint conditions. In addition, there is a high degree of coupling between the UAV trajectory and the beamforming variables , which further increases the complexity of the problem and cannot be directly solved by traditional optimization methods.
[0079] In addition, the above problem is a long-term dynamic optimization problem. The UAV position and related optimization variables change dynamically over time and need to be optimized in real time. Traditional mathematical iterative methods are difficult to establish an effective correlation between adjacent time slots, and the computational complexity is high, making it difficult to obtain the global or local optimal solution. Although traditional deep reinforcement learning has advantages in dynamic optimization, directly training all coupled variables will lead to an explosion in the dimensions of the state and action spaces, significantly reducing the algorithm convergence speed. Therefore, how to achieve efficient optimization in a high-dimensional space has become the key challenge to solve this problem.
[0080] Step 3.2: Since the variables in the original problem are highly coupled, first decouple the highly coupled problem into two sub-problems, and split the mathematical optimization model into sub-problems, namely the beamforming optimization sub-problem and the UAV trajectory optimization sub-problem.
[0081] Step 4: Based on the SCA and SDR algorithms, iteratively solve the beamforming optimization sub-problem to dynamically obtain the optimal beamforming method and simultaneously obtain the approximate optimal solution of the beamforming matrix.
[0082] Specifically, the UAV active beamforming sub-problem can be expressed as:
[0083]
[0084] For the UAV active beamforming optimization sub-problem, given the UAV trajectory , design a convex optimization algorithm to transform the complex quadratic constraint into a tractable convex constraint form and solve the problem.
[0085] The P2 problem contains the positive part operator [·] + , resulting in its non-smoothness. Remove the positive part operator and perform an equivalent problem transformation to transform the UAV active beamforming sub-problem P2 into:
[0086]
[0087] The P2 problem and the transformed P2.1 problem are equivalent. The proof is as follows:
[0088] Let and denote the optimal solution sets of problems P2 and P2.1, respectively. First, from the non-negativity of the positive part operator, for any x, we have [x] + ≥ x holds. Therefore, it is inevitable that Then let denote the optimal solution of problem P2, and define the function Furthermore, construct a feasible solution of problem P2.1: when f(W n ) ≥ 0, take Otherwise, set Let denote the objective function value of problem (P4) when using . This construction method ensures that while as a feasible solution of problem P2.1 satisfies Furthermore, we can obtain Thus, it is proved that Proof completed.
[0089] Sub-problem P2 is essentially a non-convex optimization problem. Its non-convex characteristics come from multiple aspects: First, the positive part operator [·] + is included in the objective function, resulting in its non-smoothness, and the system average security rate is essentially composed of the difference of logarithmic functions, forming a non-standard concave function structure. Second, the minimum user security rate constraint involves non-linear logarithmic operations, making traditional convex optimization methods not directly applicable and further optimization in combination with equivalent problem transformation is required.
[0090] Step 4.2: After the equivalent transformation, problem P2.1 is still an NP-hard problem that is difficult to solve directly. First, reconstruct the channel matrix of user n as H n = [h 1,n , h 2,n , …, h U,n , the channel matrix of eavesdropper e as H e = [h 1,e , h 2,e , …, h U,e , the beamforming matrix of user n as W n = [ω 1,n , ω 2,n , …, ω U,n , and rewrite the signal-to-interference-plus-noise ratios of user n and eavesdropper e as:
[0091]
[0092] Set auxiliary variables: where Also, because:
[0093]
[0094] For any matrix, there is an equivalent relationship between its second norm and the trace operation. This transformation preserves the mathematical essence of the optimization problem unchanged and only changes the expression form. The trace operation can transform the complex matrix norm constraint into the standard form of semidefinite programming (SDP). The trace operation has better compatibility with convex optimization tools such as CVX and can directly apply the semidefinite relaxation (SDR) technique to handle non-convex problems. Therefore, problem P2.1 is equivalently reconstructed as:
[0095]
[0096]
[0097] Step 4.3: Since problem P2.2 is still a non-convex problem, the optimization algorithms of SCA and SDR are used to solve this non-convex problem:
[0098] First, apply the DC decomposition principle to decompose the user data transmission rate R n into the difference between two concave functions:
[0099]
[0100] Use the SCA algorithm to convert R n into a concave function and implement it iteratively. In the r-th iteration, r≥1, use Taylor expansion to perform a local convex approximation on and write the beamforming matrix as Then we have:
[0101]
[0102] where, is the first derivative matrix;
[0103] Secondly, similarly, decompose the eavesdropping rate into the difference between two concave functions:
[0104]
[0105] Furthermore, use the SCA algorithm to convert into a convex function and implement it iteratively. In the r-th iteration, r≥1, use Taylor expansion to perform a local convex approximation on Then we have:
[0106]
[0107] where, is the first derivative matrix;
[0108] Finally, introduce the auxiliary variables τ n and ψ n , and relax the rank-one constraint This optimization problem is transformed into the standard semidefinite programming (SDP) form in the r-th iteration of SCA:
[0109]
[0110] The objective function of Problem P2.3 is a strictly concave function, and all constraint conditions form a convex set. According to the convex optimization theory, this problem belongs to the second-order cone programming problem. With the help of mathematical tools, through the SDR algorithm, a professional convex optimization solver is used to explore the optimal solution of this sub-problem; if the optimal solution after relaxation does not satisfy the rank-one constraint and a feasible beamforming vector cannot be directly obtained, the Gaussian randomization method needs to be further used to recover the rank-one solution of the problem.
[0111] As Figure 3 shown, it is the flowchart of the beamforming optimization algorithm based on SCA-SDR in the embodiment of the present invention. Since the solution after relaxation may not satisfy the rank-one constraint and a feasible beamforming vector cannot be directly obtained. Therefore, it is necessary to further use the Gaussian randomization method to recover the rank-one solution of the problem. Through the Gaussian randomization method, a feasible solution that satisfies the rank-one constraint can be recovered based on the relaxed solution, so as to obtain the optimal beamforming vector. In summary, this optimization sub-problem has been transformed into a convex optimization problem and can be solved by the CVX solver. The selection of the solver is not specifically limited in the embodiment of the present invention. The embodiment of the present invention provides a method for recovering the rank-one solution of the Gaussian randomization beamforming matrix, and the schematic diagram of the method flow is shown in Figure 4 .
[0112] Step 5: Transform the UAV trajectory optimization sub-problem into a Markov decision process, calculate the reward function according to the approximate optimal solution of the beamforming matrix obtained in Step 4, and use the improved multi-agent deep reinforcement learning algorithm RES-QMIX to solve this sub-problem to achieve real-time UAV trajectory planning in a complex environment.
[0113] The UAV dynamic trajectory planning sub-problem can be expressed as:
[0114]
[0115] For the sub-problem of UAV dynamic trajectory optimization, an improved multi-agent deep reinforcement learning algorithm is used to solve it. The trajectory planning problem is re-modeled as a distributed partially observable Markov decision process (Dec-POMDP), the beamforming optimization result is integrated into the design of its reward function, and through information sharing and collaborative decision-making among multiple UAVs, dynamic trajectory planning in complex environments is achieved. First, the problem is modeled using Dec-POMDP.
[0116] The Dec-POMDP model tuple is defined as where is the set of UAV agents; is the global state space; is the joint action space of all agents, is the discrete action space of agent u, i.e., the flight direction; r, γ are the reward function, local observation space, and discount factor respectively. The total flight time is discretized into T equal-length time slots, and the time slot set is denoted as If the actual flight time of the UAV exceeds T, the trajectory exploration is considered to fail. Each UAV u acts as an agent and obtains experience by continuously interacting with the environment s. The dynamic trajectory planning specifically includes the following steps:
[0117] As Figure 5 shown, the UAV trajectory optimization process based on RES-QMIX and SCA-SDR in the implementation of the present invention is as follows:
[0118] Step 5.1: Establish the state, observation, action, and reward in the distributed partially observable Markov decision process, which are specifically defined as:
[0119] State space
[0120] where, includes the current position l U (t), starting position l Start and end position l End ; includes the position l N of the ground user and the position l e of the eavesdropper; represents the remaining time difference of the UAV, ΔT u (t) represents the remaining time of UAV u from the current position to the end point and the minimum time difference; represents the average security rate from time slot 1 to t - 1; a(t - 1) represents the actions of all agents in the previous time slot.
[0121]
[0122] a(t - 1) = {a1(t - 1), a2(t - 1),..., a U (t - 1)
[0123] Local observation space
[0124] Z u (t) defines the remaining time difference ΔT of the UAV u u (t), the current position information and the action a of the previous time slot u (t - 1).
[0125] Action space The action a of the agent u u (t) is defined as the flight direction f(t) at time slot t + 1; the UAV flies at a fixed altitude H, so it is only discretized into F horizontal directions, that is, a u (t) = {f1(t), …, f F (t)}.
[0126] Reward function r: The reward r(t) combines the system average secrecy rate and the UAV obstacle avoidance constraint, and is specifically defined as:
[0127]
[0128] Among them, the constant C1 is the penalty for the UAV not arriving at the end point on time; C2 is the penalty for the UAV collision; C3 is the reward for all UAVs arriving at the end point on time; C4 is the basic reward in other cases.
[0129] Step 5.2: Initialize the parameters θ of the current value network and the target value network - and θ, the positions l of the user and the eavesdropper n and l e , the starting and ending positions l of the UAV Start and l End , the maximum number of training rounds E, the maximum flight time steps T in one training, the sampling size b, and the experience replay buffer size B.
[0130] Step 5.3: Start a new round of training and initialize the position of the UAV to the preset starting point.
[0131] Step 5.4: If the time step t does not reach the maximum flight time step T and the loss function of RES - QMIX does not reach the lowest threshold limit, the termination condition is not met, and the global state is obtained Each UAV u is used as an agent to obtain the local observation Z from the environment u(t) and the action a at the previous time step u (t - 1).
[0132] Step 5.5: The UAV u calculates the optimal action value according to the local observation and the global state By maximizing the action value function Q, select the current action a u (t), and form a joint action by combining the actions of all UAVs.
[0133] The calculation formula of the target Q value of the RES-QMIX algorithm is:
[0134]
[0135] where r is the reward function, γ ∈ [0, 1] is the discount factor, is the Softmax operator based on the action space , and its calculation formula is:
[0136]
[0137] where β is the inverse temperature parameter, which controls the smoothness of Softmax; when β → ∞, Softmax approaches the maximum operator; when β → 0, Softmax approaches the mean operator; Q tot (s, a) is the joint action value function:
[0138] Q tot (s, a) = f s (Q1(s, a1),..., Q U (s, a U ))
[0139] where Q tot (s, a) represents combining the values Q of multiple agents through a non-linear monotonic function f s , s represents the global environmental state, and a u represents the action selected by the agent u. u
[0140] Step 5.6: First, move the UAV according to the joint action. Secondly, according to the position of the UAV, call the SCA-SDR beamforming optimization algorithm to obtain the optimal beamforming matrix. Furthermore, according to the optimal beamforming matrix, calculate the system average secrecy rate and incorporate it into the reward function. Finally, return the joint reward r(s(t), a(t));
[0141] Step 5.7: The environment transfers from state s(t) to a new state s(t + 1) according to the state transition function, and stores the quadruple (s(t), a(t), r(t), s(t + 1)) in the experience replay pool.
[0142] Step 5.8: If the algorithm update condition is satisfied, sample a mini-batch of data of size b from the experience replay buffer, and calculate the loss function of the RES-QMIX algorithm:
[0143]
[0144] where, is the temporal difference error, R t (s, a) is the discounted cumulative return, θ is the parameter of the individual Q-value network, θ - is the parameter of the target Q-value network; λ is the regularization coefficient; r t+k is the sum of rewards accumulated starting from time slot t.
[0145] Step 5.9: After completing the training of one time step, update the time step t = t + 1, and return to Step 5.4 to enter the training of the next time step.
[0146] Step 5.10: After meeting the termination condition, if the maximum number of training rounds has not been reached, return to Step 5.3 to start a new round of training.
[0147] The network structure of the algorithm includes multiple agent networks, one hypernetwork, and one fusion network. The action value function of each agent is implemented by an agent network, whose input is the local observation Z u (t) and the action a u (t - 1) at the previous time slot, and the output is the optimal action value The structure of the agent network adopts the gated recurrent unit (GRU) and multi-layer perceptron (MLP) structure. The hybrid network generates non-negative weights through the hypernetwork and takes the global environmental state s as the input to ensure the flexible integration of global state information into the joint action value estimation.
[0148] The algorithm uses the experience replay mechanism to improve stability. During the algorithm learning process, all UAV agents are in the environment s(t), and select actions a according to the local observation u (t). The actions of all agents form the joint action a(t), the environment transfers from state s(t) to a new state s(t + 1) according to the state transition function, returns the joint reward r(s(t), a(t)), and stores the quadruple (s(t), a(t), r(t), s(t + 1)) in the experience replay pool. By continuously interacting with the environment, each agent learns the optimal policy The policies of all agents together form the joint policy π* , maximizing the long-term return of the system.
[0149] Finally, the RES-QMIX algorithm calculates the RES Loss by mixing the current value of the network and the output of the target value network, and updates the network parameters in reverse until convergence.
[0150] Step 6: Substitute the UAV positions obtained by optimizing in Step 5 into Step 4, repeat Step 4, and by means of the alternating optimization algorithm, combining the above two sub-problem solving methods, maximize the average secrecy rate of the system, and perform UAV trajectory planning and beamforming according to the optimal solution to achieve physical layer security enhancement in the communication network.
[0151] As Figure 6 shown, it is a schematic diagram of the alternating optimization framework for joint UAV trajectory and beamforming in the implementation of the present invention. This framework adopts a semi-distributed architecture. In the learning stage, the central unit trains the UAV agents and executes the SCA-SDR algorithm to solve the optimal beamforming vector. In the implementation stage, a centralized training and distributed execution (CTDE) scheme is adopted. The UAVs make real-time decisions according to local observation information and training strategies, and the central unit optimizes the beamforming and transmits the results to ensure efficient and stable multi-UAV cooperative communication. By integrating the advantages of traditional optimization and deep reinforcement learning, the present invention achieves a balance between global optimization and dynamic adaptability while reducing computational complexity and communication overhead.
[0152] The implementation of the alternating optimization framework includes:
[0153] First, decouple the highly coupled original problem into two sub-problems: the UAV trajectory optimization sub-problem, which is responsible for online dynamically optimizing the UAV trajectory given the beamforming; and the beamforming optimization sub-problem, which is responsible for optimizing the beamforming matrix given the UAV positions.
[0154] Secondly, for the UAV active beamforming optimization sub-problem, given the UAV trajectory, design a convex optimization algorithm based on SCA-SDR to transform the complex quadratic constraints into a tractable convex constraint form, and transform the original problem into a standard SDP problem for solution. For the UAV dynamic trajectory optimization sub-problem, use the improved multi-agent deep reinforcement learning algorithm RES-QMIX algorithm for solution. Re-model the trajectory planning sub-problem as a distributed partially observable Markov decision process (Dec-POMDP), integrate the beamforming optimization results into its reward function design, and through information sharing and collaborative decision-making among multiple UAVs, achieve dynamic trajectory planning in complex environments.
[0155] Finally, execute the joint optimization framework to achieve global optimization of the system average security rate.
[0156] Figures 7 - 13Shows the schematic diagrams of various parameter performances in this application.
[0157] Figure 7 Shows the comparison results of the convergence of different trajectory algorithms provided by the embodiments of the present invention. Specifically, the experimental results show that all six trajectory optimization algorithms can converge. However, RES-QMIX and RES-VDN are superior to other algorithms in terms of convergence speed and performance. This benefits from the introduction of the Reward Enhancement Strategy (RES), which significantly improves the convergence robustness by effectively utilizing environmental reward information and reducing the Q-value overestimation bias. Although RES-VDN also adopts the RES mechanism, its network structure based on linear summation limits the ability to model complex individual relationships, resulting in weaker performance than RES-QMIX. Although QMIX captures the non-linear relationships between individuals through a mixing network, due to the lack of the RES mechanism, it has disadvantages in the utilization rate of reward signals, with a slower convergence speed and larger fluctuations. QTRAN can reach a high target value during the training process, but due to its high computational complexity, its convergence speed is slow and its stability is not outstanding, and it does not achieve a significant improvement over the QMIX algorithm. QPLEX decouples action evaluation through a dual network structure, enhancing the expressive ability of the algorithm and improving the convergence speed to a certain extent, but its stability is poor. And VDN has the worst performance due to its simple structure and inability to model individual interactions.
[0158] Figure 8 Shows the comparison results of the trajectory convergence under different flight time limits provided by the embodiments of the present invention. Specifically, in the initial stage of training, there are zero values, which is because the agent lacks sufficient training in the initial stage and is difficult to make correct decisions and reach the end point within a limited time. As T increases, the algorithm convergence speed accelerates, but the optimal value and convergence are not affected by T. This is because the looser flight time constraint provides the agent with a more sufficient exploration space, but too long flight time will reduce the efficiency of communication task execution. Finally, T = 15s, which is a compromise between the convergence speed and task timeliness, is selected as the optimal parameter to ensure that the UAV completes trajectory optimization within a limited time.
[0159] Figure 9The comparison results of two different UAV trajectory algorithms provided by the embodiments of the present invention are shown. Specifically, both the RES-QMIX algorithm and the heuristic trajectory optimization algorithm can achieve dynamic trajectory adjustment. However, the heuristic algorithm explores the trajectory more slowly and is more inclined to conservatively avoid eavesdroppers, thereby reducing the eavesdropping rate of eavesdroppers. However, it cannot provide a good communication rate for users, limiting the system security performance. In contrast, the trajectory generated by the RES-QMIX algorithm can establish a more secure communication link in a shorter time, and its average system security rate is also significantly better than that of the heuristic algorithm. This shows that the RES-QMIX trajectory optimization algorithm realizes efficient dynamic trajectory strategy design through online decision optimization.
[0160] Figure 10 The comparison results of the convergence performance of the SCA-SDR beamforming optimization algorithm provided by the embodiments of the present invention are shown. Specifically, the experimental results show that the scheme combining the RES-QMIX algorithm and the CoMP technology is superior to the other two schemes in terms of convergence speed and final performance, significantly improving the system security rate and the user communication rate. This advantage mainly stems from two key factors: First, the RES-QMIX trajectory has improved spatial efficiency compared to the heuristic trajectory, creating more favorable geometric conditions for multi-UAV cooperative transmission. Second, through the CoMP technology, multiple UAVs can achieve symbol-level signal cooperation and beamforming, significantly enhancing the signal coverage and transmission capacity and improving the user communication rate; at the same time, the multi-path transmission and interference management mechanism of CoMP suppresses the eavesdropping risk and also significantly optimizes the system security rate. In contrast, the single-UAV benchmark group without CoMP cooperation obtains the worst performance due to the lack of spatial diversity gain and joint processing ability. Moreover, due to resource constraints, a single UAV needs a longer time to adjust parameters and has a slower speed of searching for the optimal strategy, so the convergence speed is also slower.
[0161] Figure 11The comparison results of the system average secrecy rate provided by the embodiments of the present invention with respect to the flight altitude and the number of antennas are shown. Specifically, the experimental results show that the increase in the flight altitude of the UAV will exacerbate the path loss of the communication link, resulting in a decrease in signal quality and a reduction in the system average security rate. At the same time, with the increase in the number of antennas, spatial diversity and beamforming gain can be used to suppress the eavesdropping channel capacity while increasing the user transmission rate, thereby improving the system average security rate. In addition, under the same conditions, the scheme combining the RES-QMIX algorithm and the CoMP technology can achieve the best physical layer security performance. On the one hand, the CoMP technology effectively expands the signal coverage range and reduces the co-channel interference through the cooperative beamforming mechanism, improving the system average security rate. Therefore, the performance of the CoMP-assisted optimization framework is significantly better than that of a single UAV without CoMP, fully verifying the performance gain of the CoMP technology. On the other hand, the RES-QMIX algorithm has a better trajectory optimization effect compared with the heuristic trajectory algorithm through dynamic trajectory optimization and resource allocation strategies, which can further improve the system average security rate.
[0162] Figure 12 The comparison results of the system average secrecy rate provided by the embodiments of the present invention with respect to the maximum transmit power and the number of users are shown. Specifically, the experimental results show that the increase in the transmit power of the UAV can enhance the beamforming gain, improve the signal coverage and anti-interference ability, thereby increasing the system average security rate. At the same time, with the increase in the number of users, resource allocation becomes scattered and interference intensifies, resulting in a decrease in the system average security rate. In addition, the increase in the number of users introduces additional interference, further reducing the communication quality and security. In addition, under the same conditions, the scheme combining the RES-QMIX algorithm and the CoMP technology can achieve the best physical layer security performance. On the one hand, the CoMP technology effectively expands the signal coverage range and reduces the co-channel interference through the cooperative beamforming mechanism, improving the system average security rate. On the other hand, the RES-QMIX algorithm has a better trajectory optimization effect compared with the heuristic algorithm, which can further improve the system average security rate.
[0163] Figure 13The comparison results of the relationship between the system average secrecy rate, the minimum secrecy rate, and the number of UAVs provided by the embodiments of the present invention are shown. Specifically, the system average security rate can remain relatively stable under different user security rate threshold conditions, which indicates that the framework has strong adaptability in resource allocation and interference management and has good robustness. At the same time, as the number of UAVs increases, the system average security rate increases, but when the number of UAVs increases to a certain value, the growth rate begins to slow down. This is because the increase in the number of UAVs provides more signal coverage and transmission paths, but due to the limitations of system resources and channel conditions, its marginal benefit will decrease with the further increase in the number. In addition, under the same conditions, the scheme combining the RES-QMIX algorithm with the CoMP technology can obtain the best physical layer security performance. On the one hand, the CoMP technology effectively expands the signal coverage range and reduces the co-channel interference through the cooperative beamforming mechanism, improving the system average security rate. On the other hand, the RES-QMIX algorithm has a better trajectory optimization effect compared with the heuristic algorithm, which can further improve the system average security rate.
[0164] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for enhancing physical layer security in a cooperative communication network, characterized in that Including: Step 1: Obtain the relevant configuration information of the UAV base station and ground users, and construct a multi-UAV downlink secure communication system based on the coordinated multi-point transmission technology according to the relevant configuration information; Step 2: Use the Rice fading channel to model the links from multiple UAVs to ground users and eavesdroppers, and calculate the secrecy rate of the system based on the achievable data rate of the users and the eavesdropping rate of the eavesdroppers as the physical layer security research index; Step 3: Based on the secrecy rate of the system, construct a physical layer security enhancement mathematical optimization model for the cooperative communication network, and split the mathematical optimization model into a beamforming optimization sub-problem and a UAV trajectory optimization sub-problem; Step 4: Based on the SCA and SDR algorithms, iteratively solve the beamforming optimization sub-problem to dynamically obtain the optimal beamforming method, and at the same time obtain an approximate optimal solution of the beamforming matrix; Step 5: Transform the UAV trajectory optimization sub-problem into a Markov decision process, calculate the reward function according to the approximate optimal solution of the beamforming matrix obtained in Step 4, and use the improved multi-agent deep reinforcement learning algorithm RES-QMIX to solve this sub-problem to achieve real-time UAV trajectory planning in complex environments; Step 6: Substitute the UAV positions optimized in Step 5 into Step 4, repeat Step 4, and through the alternating optimization algorithm, combine the above two sub-problem solving methods to maximize the average secrecy rate of the system, and perform UAV trajectory planning and beamforming according to the optimal solution to achieve physical layer security enhancement in the communication network.
2. The physical layer security enhancement method for a cooperative communication network according to claim 1, wherein The relevant configuration information of the UAV base station and ground users includes: Set multiple UAVs equipped with multi-antenna arrays as aerial base stations, responsible for broadcasting confidential signals to legitimate ground users; set multiple legitimate ground users and suspected eavesdroppers equipped with single-antenna devices. The legitimate ground users are used to receive confidential signals, and the suspected eavesdroppers conduct potential eavesdropping.
3. The physical layer security enhancement method for a cooperative communication network according to claim 1, characterized in that In the multi-UAV downlink secure communication system based on the coordinated multi-point transmission technology: Multiple UAV base stations cooperate to serve edge users in a coordinated multi-point joint transmission manner, converting inter-channel interference into useful signals; Through cooperative beamforming optimization, a strong signal gain is formed in the direction of the target user, suppressing the signal reception quality of eavesdroppers and enhancing physical layer security; Adopt multiple input single output and beamforming technologies to improve the system degrees of freedom and spectral efficiency.
4. The physical layer security enhancement method for a cooperative communication network according to claim 1, characterized in that, The specific content of Step 2 is: Step 2.1: Use the Rice fading channel to model the links from multiple UAVs to ground users and eavesdroppers, and synthesize the channel gain by weighting the line-of-sight component and non-line-of-sight component: The channel gain of the direct link from UAV u to ground user n is: The channel gain of the direct link from UAV u to eavesdropper e is: Among them, represents the beamforming vector from the UAV u to the user n at the t-th time slot; and is the additive white Gaussian noise power; N is the total number of legitimate users, and U is the total number of UAVs. Step 2.2: The signal-to-interference-plus-noise ratio of user n in time slot t is expressed as: The signal-to-interference-plus-noise ratio of eavesdropper e in time slot t is expressed as: Among them, represents the beamforming vector from the UAV u to the user n at time slot t; and is the additive white Gaussian noise power; N is the total number of legitimate users, and U is the total number of UAVs; Step 2.3: Calculate the data transmission rate of legitimate ground users and the data transmission rate of eavesdroppers through the Shannon formula: R n (t) = log2(1 + γ n (t)) Among them, R n (t) is the data transmission rate of legitimate users on the ground, and the data transmission rate of the eavesdropper; Then the secrecy rate of user n is: Among them, R E (t) is the maximum eavesdropping rate, and the operator Then the total secrecy rate of the system is: Average secrecy rate of the system is as follows: where T is the total number of time slots.
5. The physical layer security enhancement method for a cooperative communication network according to claim 1, wherein The specific content of Step 3 is: Step 3.1: Using the trajectory of the UAV and the active beamforming matrix of the UAV as optimization variables, by jointly optimizing the trajectory and the active beamforming matrix, the optimization objective of maximizing the system average secrecy rate is achieved, and the expression is as follows: Among them, R min is the minimum value of the set secrecy rate, is the maximum value of the UAV transmission power; x u (t), y u (t) represent the position coordinates of UAV u; Through represents the two-dimensional area range in which UAV u flies; l u (t), l i (t) and l j (t) are the positions of UAVs u, i, and j at time slot t, respectively. Step 3.2: Split the mathematical optimization model into sub-problems, which are split into a beamforming optimization sub-problem and a UAV trajectory optimization sub-problem.
6. The physical layer security enhancement method for a cooperative communication network according to claim 1, characterized in that The specific content of step 4 is as follows: Step 4.1: Represent the UAV active beamforming sub-problem as: The P2 problem contains the positive part operator [·] + , which results in its non-smoothness. Remove the positive part operator, perform equivalent problem transformation, and transform the UAV active beamforming sub-problem P2 into: Step 4.2: The problem P2.1 after equivalent transformation is still an NP-hard problem that is difficult to solve directly. First, reconstruct the channel matrix of user n as H n =[h 1,n , h 2,n , …, h U,n , the channel matrix of eavesdropper e is H e =[h 1,e , h 2,e , …, h U,e , the beamforming matrix of user n is W n =[ω 1,n , ω 2,n , …, ω U,n , and rewrite the signal-to-interference-plus-noise ratios of user n and eavesdropper e as follows: Set auxiliary variables: Wherein Also, because: For any matrix, there is an equivalent relationship between its second norm and trace operation. This transformation maintains the mathematical essence of the optimization problem unchanged and only changes the expression form. Therefore, problem P2.1 is equivalently reconstructed as: Step 4.3: Since problem P2.2 is still a non-convex problem, an optimization algorithm combining SCA and SDR is used to solve this non-convex problem: First, apply the DC decomposition principle to decompose the user data transmission rate R n into the difference between two concave functions: Convert R using the SCA algorithm n into a concave function and implement it iteratively. In the r-th, r ≥ 1 iteration, use Taylor expansion for to perform a local convex approximation, and write the beamforming matrix as Then we have: Among them, is the first derivative matrix; Secondly, similarly, decompose the wiretapping rate into the difference of two concave functions: Furthermore, the SCA algorithm is used to transform it into a convex function and implement it in an iterative manner. In the r-th iteration with r ≥ 1, the Taylor expansion is used to perform a local convex approximation, and then we have: Among them, is the first derivative matrix; Finally, introduce auxiliary variables τ n and ψ n , and relax the rank-one constraint This optimization problem is transformed into a standard semidefinite programming form in the r-th iteration of SCA: The objective function of problem P2.3 is a strictly concave function, and all constraint conditions form a convex set. This problem belongs to a second-order cone programming problem. Through the SDR algorithm, a professional convex optimization solver is used to explore the optimal solution of this sub-problem; if the optimal solution after relaxation does not satisfy the rank-one constraint and a feasible beamforming vector cannot be directly obtained, the Gaussian randomization method needs to be further used to recover the rank-one solution of the problem.
7. The physical layer security enhancement method for a cooperative communication network according to claim 1, characterized in that The specific content of step 5 is as follows: Step 5.1: Establish the states, observations, actions, and rewards in the distributed partially observable Markov decision process, and define them in detail as: State space Among them, including the current position l of the drone U (t), the starting position l Start and the end position l End ; including the position l of the ground user N and the position l of the eavesdropper e ; representing the remaining time difference of the drone, ΔT u (t) represents the remaining time from the current position of the drone u to the end point and the minimum time difference; representing the average security rate from time slot 1 to t - 1; a(t - 1) represents the actions of all agents in the previous time slot; a(t - 1) = {a1(t - 1), a2(t - 1),..., a U (t - 1)} Local Observation Space Z u (t) defines the remaining time difference ΔT of the UAV u u (t), the current position information and the action a of the previous time slot u (t - 1); Action space The action a of the agent u u (t) is defined as the flight direction f(t) at time slot t + 1; the UAV flies at a fixed altitude H, so it is only discretized into F horizontal directions, that is, a u (t) = {f1(t), …, f F (t)}; Reward function r: The reward r(t) incorporates the system's average secrecy rate and the UAV obstacle avoidance constraint, and is specifically defined as: Among them, the constant C1 is the penalty for the UAV not being able to reach the end point on time; C2 is the penalty for UAV collision; C3 is the reward for all UAVs reaching the end point on time; C4 is the basic reward in other cases; Step 5.2: Initialize the parameters θ of the current value network and the target value network - and θ, the positions l of the user and the eavesdropper n and l e the starting and ending positions l of the UAV Start and l End the maximum number of training rounds E, the maximum number of flight time steps T in one training, the sampling size b, and the size B of the experience replay buffer; Step 5.3: Start a new round of training and initialize the position of the UAV to the preset starting point; Step 5.4: If the time step t has not reached the maximum flight time step T and the loss function of RES-QMIX has not reached the lowest threshold limit, the termination condition is not satisfied, and the global state is obtained. Each drone u, as an agent, obtains the local observation Z from the environment. u (t) and the action a u (t - 1); Step 5.5: The drone u calculates the optimal action value based on local observations and the global state Select the current action a by maximizing the action value function Q u (t), and form a joint action by combining the actions of all drones; The calculation formula for the target Q value of the RES-QMIX algorithm is: where \(r\) is the reward function, \(\gamma\in[0,1]\) is the discount factor, is the Softmax operator based on the action space and its calculation formula is: where β is the inverse temperature parameter that controls the smoothness of Softmax; when β → ∞, Softmax approaches the maximum operator; when β → 0, Softmax approaches the mean operator; Q tot (s,a) is the joint action-value function: Q tot (s,a) = f s (Q1(s,a1),...,Q U (s,a U )) Among them, Q tot (s,a) represents passing through a non-linear monotonic function f s to combine the values Q of multiple agents u , s represents the global environmental state, a u represents the action selected by agent u; Step 5.6: First, move the UAV according to the joint action. Second, according to the position of the UAV, call the UAV beamforming optimization algorithm to obtain the optimal beamforming matrix. Third, calculate the system average secrecy rate according to the optimal beamforming matrix and incorporate it into the reward function. Finally, return the joint reward r(s(t), a(t)); Step 5.7: The environment transfers from state s(t) to a new state s(t + 1) according to the state transition function, and stores the quadruple (s(t), a(t), r(t), s(t + 1)) in the experience replay pool; Step 5.8: If the algorithm update condition is met, sample a mini-batch of sample data of size b from the experience replay buffer, and calculate the loss function of the RES-QMIX algorithm: Among them, is the temporal difference error, R t (s,a) is the discounted cumulative return, θ is the parameter of the individual Q-value network, θ - is the parameter of the target Q-value network; λ is the regularization coefficient; r t+k is the sum of rewards accumulated starting from time slot t; Step 5.9: After completing the training of one time step, update the time step t = t + 1, and return to step 5.4 to enter the training of the next time step; Step 5.10: After meeting the termination condition, if the maximum number of training rounds has not been reached, return to step 5.3 to start a new round of training.