Unmanned aerial vehicle physical layer transmission method for overall security
By calculating threat areas in the UAV communication system and combining deep reinforcement learning algorithms with LSTM and PER mechanisms, the interference strategy of interfering with the UAV is optimized, and the problems of accumulated security performance and low learning efficiency in UAV communication are solved, and the continuous security and efficient and secure transmission of the UAV communication process are achieved.
Patent Information
- Application Number
- CN202510540055.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-03-04
- Filing Date
- 2025-04-27
- Publication Date
- 2025-07-01
AI Technical Summary
The existing UAV physical layer security transmission method is difficult to optimize the cumulative security performance of the entire transmission process when the potential eavesdropper location is unknown, and the deep reinforcement learning method has problems such as slow convergence speed and low learning efficiency when dealing with decision-making problems in the UAV communication environment.
A drone physical layer transmission method for overall security is adopted. By acquiring the location information of the communication drone and the user, calculating the threat area and performing discrete processing, combining the deep reinforcement learning algorithm DDQN of the long and short-term memory network LSTM and the priority experience playback PER mechanism, the interference strategy of interfering with the drone is optimized to ensure the continuous security of the drone communication process.
It realizes the rapid establishment of secure transmission strategies in a dynamic drone communication environment, improves learning efficiency and convergence performance, and ensures the continuous security and efficient security performance of the drone communication process.
Smart Images

Figure CN120238235A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of UAV communication, and particularly relates to a UAV physical layer transmission method for overall security. Background Technique
[0002] UAV communication systems are vulnerable to security threats such as eavesdropping and interference due to their open wireless transmission environment. To address this issue, physical layer security technology has become an important security measure by leveraging the physical characteristics of the wireless channel to ensure communication security. Most previous studies on UAV physical layer security were based on the situation where the location of the potential eavesdropper Eves was known, and focused on the instantaneous security performance of a single time slot, such as maximizing the instantaneous secrecy capacity or instantaneous security rate. However, in the actual UAV communication process, the lack of information about the location of the potential eavesdropper Eves makes it impossible to calculate the eavesdropping channel, resulting in a lack of a system security measurement method for this situation. Therefore, it is more practical to optimize the cumulative security performance of the entire transmission process when the location of the potential eavesdropper Eves is unknown.
[0003] In addition, due to the dynamic and complex nature of the UAV communication environment, traditional optimization methods are difficult to effectively handle high-dimensional decision spaces. Although deep reinforcement learning methods have shown advantages in solving complex decision problems, existing deep Q-network algorithms have problems such as low learning efficiency and poor convergence performance when dealing with time-related decision sequences. Therefore, there is an urgent need for a UAV physical layer security transmission method that can optimize the overall security performance and has high learning ability. Summary of the Invention
[0004] Aiming at the above deficiencies in the prior art, a UAV physical layer transmission method for overall security provided by the present invention solves the following technical problems in the existing UAV physical layer security transmission methods: First, the existing methods only focus on the instantaneous security performance of a single time slot, ignoring the cumulative security performance of the entire transmission process, and it is difficult to ensure the continuous security of the UAV communication process; Second, the current research focuses on simple and easy-to-handle scenarios, ignoring the consideration of the location of the potential eavesdropper Eves and the movement of the UAV. Third, existing deep reinforcement learning methods have problems of slow convergence speed and low learning efficiency when dealing with continuous decision problems in the UAV communication environment.
[0005] To achieve the above objectives, the technical solution adopted by the present invention is: A UAV physical layer transmission method for overall security, comprising the following steps:
[0006] S1. Obtain the position and power information of each communication UAV at each moment, and the position of each user at each moment;
[0007] S2. Calculate the Euclidean distance between the communication UAV and the ground node when the instantaneous signal-to-noise ratio of the potential eavesdropper Eves is equal to the threshold of the signal-to-noise ratio of the potential eavesdropper Eves based on the position and power information of the communication UAV at the current moment, and obtain a threat area with the communication UAV as the center and R as the radius on the ground, where the Euclidean distance is denoted as R;
[0008] S3. Discretize the threat area in a grid manner and construct a vector of length n, where each element in the vector corresponds to a position, indicating whether the position is a safe area;
[0009] S4. Establish an optimization problem for the elements in the vector and the position security at each moment, sum up the optimization problems for all time periods, and use the summation result as the final optimization problem, where the objective function is to maximize the sum of the proportions of the safe areas in all time periods, and this position is the threatened area ω e The position of the point after discretization, representing the position of the potential eavesdropper Eves;
[0010] S5. According to the final optimization problem, use the long short-term memory network LSTM and the prioritized experience replay PER mechanism, combined with the deep deterministic policy gradient algorithm DDQN, to obtain the interference strategy of the interference UAV and complete the physical layer transmission of the UAV for overall security.
[0011] The beneficial effects of the present invention are as follows: The present invention uses an additional interference UAV as an auxiliary to interfere with potential eavesdroppers Eves in the environment by sending artificial noise to ensure the secure transmission of information between legitimate UAVs and users; takes the security performance of multiple time slots as the overall optimization goal, and combines the improved double deep Q network algorithm to achieve continuous security guarantee for the UAV communication process. This method improves the learning efficiency and convergence performance of the algorithm by introducing the long short-term memory network LSTM and the prioritized experience replay PER mechanism, enabling it to better adapt to the dynamically changing UAV communication environment.
[0012] Furthermore, the expression of the Euclidean distance is as follows: P NLoS =1 - P LoS ; Where represents the average path loss between UAV k and ground node u, P LoS represents the line-of-sight distance transmission probability between the UAV and the ground unit, d k,u represents the Euclidean distance between UAV k and ground node u, β L and β N respectively represent the path loss exponents of LoS and NLoS, NLoS represents the non-line-of-sight link, LoS represents the line-of-sight link, η LoS and ηNLoS Both represent the attenuation coefficient, P NLoS represents the transmission probability of NLoS, and both α and β represent constants, h k represents the vertical distance from the communication UAV k to the ground, r k,u represents the projected distance on the ground plane between the ground user u and the communication UAV k.
[0013] Furthermore, the expression of the instantaneous signal-to-noise ratio of the potential eavesdropper Eves is as follows: where, γ e represents the instantaneous signal-to-noise ratio of the potential eavesdropper Eves, P s and P j represent the transmission power of the communication UAV and the transmission power of the friendly jamming UAV, respectively, and represent the path losses from the communication UAV and the jamming UAV to the potential eavesdropper Eves, respectively, represents the noise power of the eavesdropper Eves.
[0014] Furthermore, the expression of the vector is as follows: ; where, m t represents the vector at time t, represents the i-th element in this vector, represents the n-th element, and T represents the vector transpose.
[0015] The beneficial effect of the above further scheme is that the vector m is used to represent the security of each discretized position in the threat area ω e . Specifically, each element m t of the vector m corresponds to a position in the threat area. m i >0 indicates that there is a potential eavesdropping threat at this position, and m i <0 indicates that this position is secure. Through this quantitative representation, the distribution of the threat area can be clearly described, providing a clear input for the subsequent optimization problem, not only simplifying the complexity of the problem, but also enabling the optimization problem to be efficiently solved by numerical methods (such as convex optimization, reinforcement learning, etc.).
[0016] Furthermore, the expression of the final optimization problem is as follows:
[0017]
[0018]
[0019]
[0020] E total ≤E max ; (xs,t -x j,t ) 2 +(y s,t -y j,t ) 2 +(h s,t -h j,t ) 2 ≥c 2 ;x j,min ≤x j,t ≤x j,max ;
[0021] y j,min ≤y j,t ≤y j,max ;h j,min ≤h j,t ≤h j,max ;P j,min ≤P j,t ≤P j,max ;
[0022] Among them, a represents the acceleration of the friendly jamming UAV, P represents the power of the friendly jamming UAV, represents the i-th element in the vector at time t, ω e,t represents the threat area at time t, T represents the total time, m t represents the vector at time t, n t represents the length of the vector at time t, γ e represents the instantaneous signal-to-noise ratio of the potential eavesdropper Eves, x j,t and y j,t respectively represent the x-coordinate and y-coordinate of the friendly jamming UAV at time t, P j,t and h j,t respectively represent the power and altitude of the friendly jamming UAV at time t, x e,i and y e,i represent the x-coordinate and y-coordinate of the potential eavesdropper Eves, represents the threshold of the signal-to-noise ratio of the potential eavesdropper Eves, represents the user signal-to-noise ratio threshold, γ b represents the user signal-to-noise ratio, represents the horizontal acceleration vector of the friendly jamming UAV at time t, represents the vertical acceleration vector of the friendly jamming UAV at time t, a max and E max respectively represent the upper limits of acceleration and total energy, represents the horizontal instantaneous velocity of the friendly jamming UAV at time t, represents the vertical instantaneous velocity of the friendly jamming UAV at time t, v maxDenote the maximum speed of the friendly jamming UAV, E total Denote the three-dimensional UAV propulsion energy model, x s,t , y s,t and h s,t respectively represent the x-coordinate, y-coordinate and altitude of the communication UAV at time t. c represents the minimum distance to prevent collision between communication UAVs. x j,min , x j,max , y j,min , y j,max , h j,min and h j,max represent the upper and lower limits of the flight range of the friendly jamming UAV. δ t Denote the time slot size. P0 and P1 respectively represent the profile power and induced power of the blade in the hovering state. P2 represents the continuous power during the entire descent and ascent process. Denote the average speed of the friendly jamming UAV at time t, U tip Denote the tip speed of the rotor blade. d0 represents the drag coefficient of the fuselage. ρ represents the air density. s represents the firmness of the rotor. A represents the area of the rotor disk. v0 represents the average induced speed of the rotor during hovering. Denote the vertical average speed of the friendly jamming UAV at time t.
[0023] The beneficial effects of the above further solution are as follows: The per-slot optimization only focuses on the optimal solution at a single time point, while the overall time optimization starts from a global perspective and considers the safety during the entire mission cycle. This global optimization can avoid local optimal solutions and ensure that the UAV always adopts the optimal safety strategy throughout the mission process, thereby improving the overall safety performance of the system.
[0024] Furthermore, the S5 includes the following steps:
[0025] Initialize the parameters and buffer of the long short-term memory network LSTM;
[0026] Obtain the initial environmental observation state, including the position information of the UAV and the user, etc.;
[0027] Use the long short-term memory network LSTM with the environmental observation state as the input to output the Q values of all possible system actions, representing the expected value of the long-term reward obtained by selecting a specific action in the current state;
[0028] Use the greedy strategy method e-greedy to select an action i.e., the acceleration and jamming power of the jamming UAV;
[0029] Execute the action Calculate the proportion of the safe area under the current action as the reward r t , and respectively receive the next state st+1 End identifier end t ;
[0030] Store it in the buffer according to the priority;
[0031] Extract inexperienced samples from the buffer according to the sample optimization level, where the expression of the sampling weight is as follows:
[0032]
[0033]
[0034] k where w represents the sampling weight, N represents the total number of samples in the experience pool, P represents the selection probability of the sample, β' represents the hyperparameter for controlling the correction amplitude, k represents the index of the sample, and p
[0035] Calculate the objective value according to the final optimization problem, combined with the extracted inexperienced samples, and update the long short-term memory network parameter θ using the loss function L(θ);
[0036] Through continuous iterative optimization, obtain the interference strategy for the interfering UAV, and complete the physical layer transmission of the UAV for overall security.
[0036] Furthermore, the expression of the priority is as follows: p = |δ| + ∈; where p represents the priority, δ represents the time difference, and ∈ represents a constant for adjusting the importance of the optimization level;
[0037] The update expression of the time difference δ is as follows:
[0038] The beneficial effect of the above further solution is that the Prioritized Experience Replay (PER) mechanism assigns priorities to each experience sample according to the Temporal Difference error (TD error), so that samples that have a greater impact on policy optimization are sampled and trained more frequently. This mechanism can accelerate model convergence, avoid inefficient learning caused by uniform sampling, and at the same time correct the bias brought by non-uniform sampling through importance sampling weights, ensuring the stability and accuracy of training. Brief Description of the Drawings
[0039] Figure 1 It is a flowchart of the method of the present invention.
[0040] Figure 2 It is a model diagram of a UAV communication system. Detailed Embodiments
[0041] The specific implementation manners of the present invention will be described below to facilitate those skilled in the art to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation manners. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.
[0042] Embodiment
[0043] As Figure 1 shown, the present invention provides a UAV physical layer transmission method for overall security, and its implementation method is as follows:
[0044] S1. Obtain the position and power information of the communication UAV at each moment, and the position of the user at each moment. Among them, the position of the user at each moment is used to calculate the user signal-to-noise ratio to ensure that the user signal-to-noise ratio is greater than the user signal-to-noise ratio threshold, which is reflected in the second constraint condition of the optimization problem;
[0045] S2. According to the position and power information of the communication UAV at the current moment, when calculating that the instantaneous signal-to-noise ratio of the potential eavesdropper Eves is equal to the threshold of the signal-to-noise ratio of the potential eavesdropper Eves, obtain the Euclidean distance between the communication UAV and the ground node before, and obtain a threat area with the communication UAV on the ground as the center and R as the radius, where the Euclidean distance is denoted as R;
[0046] S3. Discretize the threat area in a grid manner and construct a vector of length n, where each element in the vector corresponds to a position, indicating whether the position is a safe area;
[0047] S4. Establish an optimization problem for the element in the vector and the position security at each moment, sum up the optimization problems of all time periods, and use the summation result as the final optimization problem. Among them, the objective function is to maximize the sum of the proportions of the safe areas in all time periods, and this position is the threatened area ω e represents the position of the point after discretization, and represents the position of the potential eavesdropper Eves;
[0048] S5. According to the final optimization problem, use the long short-term memory network LSTM and the prioritized experience replay PER mechanism, combined with the deep reinforcement learning algorithm DDQN, to obtain the interference strategy of the interference UAV, and complete the UAV physical layer transmission for overall security. Its implementation method is as follows:
[0049] Initialize the parameters and buffer of the long short-term memory network LSTM;
[0050] Obtain the initial environmental observation state, including the position information of the UAV and the user, etc.;
[0051] Using the long short-term memory network (LSTM), taking the environmental observation state as the input, and outputting the Q-values of all possible system actions, which represent the expected value of the long-term reward that can be obtained by selecting a specific action in the current state;
[0052] Selecting actions using the greedy policy method e-greedy That is, the acceleration and interference power of the interfering UAV;
[0053] Executing the action Calculating the proportion of the safe area under the current action as the reward r t , and respectively receiving the next state s t+1 , and the end identifier end t ;
[0054] Storing in the buffer according to the priority;
[0055] Extracting inexperienced samples from the buffer according to the sample optimization level;
[0056] According to the final optimization problem, combining the extracted inexperienced samples, calculating the target value, and updating the long short-term memory network parameter θ using the loss function L(θ);
[0057] Through continuous iterative optimization, obtaining the interference strategy of the interfering UAV and completing the physical layer transmission of the UAV for overall security.
[0058] In this embodiment, as Figure 2 shown, the system is a UAV communication system composed of a communication UAV (S-UAV), a user (Bob), a friendly interfering UAV (J-UAV), and a potential eavesdropper Eves (Eve). In this system, the communication UAV is responsible for data collection and transmitting confidential information to the ground user, the ground user is responsible for executing tasks according to a predetermined trajectory, and the friendly interfering UAV interferes with the potential eavesdropper Eves by sending artificial noise.
[0059] Communication model:
[0060] The calculation formula for the transmission probability of the line-of-sight (LOS) link between the UAV and the ground unit is as follows:
[0061] Correspondingly, the NLoS transmission probability is: P NLoS = 1 - P LoS .
[0062] The average path loss is a key indicator for evaluating the link quality between the UAV and the ground device, and its mathematical expression is: Among them, represents the average path loss between UAV k and ground node u, PLoS Denotes the transmission probability of the line-of-sight distance between the UAV and the ground unit, d k,u Denotes the Euclidean distance between the UAV k and the ground node u, β L And β N Respectively denote the path loss exponents of LoS and NLoS. NLoS represents a non-line-of-sight link, and LoS represents a line-of-sight link, η LoS And η NLoS Both denote the attenuation coefficient, P NLoS Denotes the transmission probability of NLoS. Both α and β denote constants, h k Denotes the vertical distance from the communication UAV k to the ground, r k,u Denotes the projected distance on the ground plane between the ground user u and the communication UAV k.
[0063] The instantaneous signal-to-noise ratios of the user and the potential eavesdropper Eves are: Where, γ b Denotes the signal-to-noise ratio of the user, γ e Denotes the instantaneous signal-to-noise ratio of the potential eavesdropper Eves, P s And P j Respectively denote the transmission power of the communication UAV and the transmission power of the friendly jamming UAV, And Respectively denote the path losses from the communication UAV and the jamming UAV to the potential eavesdropper Eves, Denotes the noise power of the eavesdropper Eves.
[0064] Security measurement criteria:
[0065] To accurately measure the security of the system, two thresholds And Are set to evaluate the receiving performance of the user and the potential eavesdropper Eves. Specifically, if the signal-to-noise ratio of the receiver is higher than the threshold, it means that the receiver can decode the information; otherwise, it means that the receiver cannot obtain the content of the information. Therefore, in the absence of a friendly jamming UAV, the position set ω Satisfying e Is considered a threat area. Under the interference of the friendly jamming UAV, the position set ω e Becomes The positions called the secure region SR; the ratio of the number of grid points in SR at the current moment to the number of grid points in the position set ω e Is regarded as the secure region ratio at that moment.
[0066] Kinematics description of the friendly jamming UAV J-UAV:
[0067] Define the horizontal acceleration vector of the friendly jamming UAV J-UAV at time t as It can be expressed as: Wherein, and are the acceleration components in the x and y directions at time t respectively, and the acceleration component in the vertical direction is defined as Similarly, the instantaneous velocity of the friendly jamming UAV at time t is defined as the horizontal direction vector and in the vertical direction. Then the velocity at the next moment t + 1 can be calculated by the following formula: Wherein, represents the instantaneous velocity vector in the horizontal direction of the friendly jamming UAV at the next moment t + 1, and δ t represents the time slot size, represents the velocity in the vertical direction at the next moment t + 1.
[0068] The average velocity at time t is denoted as: Wherein, and represent the average velocities in the horizontal and vertical directions respectively.
[0069] According to the position of the friendly jamming UAV at time t and the above motion states, the position of the friendly jamming UAV at the subsequent moment t + 1 can be calculated by the following formula: Wherein, q j,t+1 and q j,t represent the positions of the friendly jamming UAV at moments t + 1 and t respectively, h j,t+1 represents the height of the jamming UAV at moment t + 1, h j,t+1 and h j,t represent the heights of the friendly jamming UAV at moments t + 1 and t respectively, represents the vertical instantaneous velocity of the friendly jamming UAV at time t, represents the acceleration vector in the vertical direction of the friendly jamming UAV at time t, and δ t represents the time slot size. q refers to x or y, which is used to simplify the coordinate position calculation formula of x and y, represents the velocity component of the friendly jamming UAV in the x or y direction at time t.
[0070] Combined with the above kinematic equations, it can be seen that by adjusting the acceleration of the friendly jamming UAV within each specified time period, the expected results of controlling the velocity and trajectory of the friendly jamming UAV can be achieved.
[0071] In order to represent the energy consumption of the entire communication UAV flight, the present invention adopts a three-dimensional UAV propulsion energy model:
[0072] Within a given spatial and temporal range, let the acceleration of the friendly jamming UAV be \(a\) and the power be \(p\), both of which are decision variables to be optimized. The optimization objective of the present invention aims to maximize the sum of the proportions of the safe regions accumulated over time by finely regulating the acceleration parameter and power output of the friendly jamming UAV. The specific optimization problem is formulated as follows: where \(x\) j,min , \(x\) j,max , \(y\) j,min , \(y\) j,max , \(h\) j,min , \(h\) j,max represent the upper and lower limits of the flight range of the friendly jamming UAV, and \(c\) represents the minimum distance to prevent collision between communication UAVs.
[0073] To improve the efficiency of the optimization program, the present invention introduces a vector: where \(m\) t represents the vector at time \(t\), represents the \(i\)-th element in this vector, represents the \(n\)-th element, and \(T\) represents the vector transpose.
[0074] This vector is used to show the safety performance of these discrete threat positions at time \(t\). Each element represents the safety of the \(i\)-th potential position within \(\omega\) e,t , and \(\omega\) e,t represents the threat area at time \(t\). The number of non-zero elements in the vector represents the number of positions that do not meet the safety conditions.
[0075] In summary, the present invention introduces a new objective function such that the number of non-zero elements in the vector \(m\) t represents \(SR\), thus transforming the original problem into the following problem:
[0076] \(E\) total \(\leq E\) max ;
[0077] \((x\) s,t - x\) j,t ) 2 + (y\) s,t - y\) j,t ) 2 + (h\) s,t - h\) j,t ) 2 \(\geq c\) 2 ; \(x\) j,min \(\leq x\) j,t \(\leq x\) j,max ; \(y\) j,min \(\leq y\) j,t \(\leq y\) j,max; h j,min ≤ h j,t ≤ h j,max ; P j,min ≤ P j,t ≤ P j,max ; where a represents the acceleration of the friendly jamming UAV, and P represents the power of the friendly jamming UAV. represents the i-th element in the vector at time t, ω e,t represents the threat area at time t, T represents the total time, m t represents the vector at time t, n t represents the length of the vector at time t, γ e represents the instantaneous signal-to-noise ratio of the potential eavesdropper Eves, x j,t and y j,t respectively represent the x-coordinate and y-coordinate of the friendly jamming UAV at time t, P j,t and h j,t respectively represent the power and altitude of the friendly jamming UAV at time t, x e,i and y e,i represent the x-coordinate and y-coordinate of the potential eavesdropper Eves. represents the threshold of the signal-to-noise ratio of the potential eavesdropper Eves. represents the user signal-to-noise ratio threshold, γ b represents the user signal-to-noise ratio. represents the horizontal acceleration vector of the friendly jamming UAV at time t. represents the vertical acceleration vector of the friendly jamming UAV at time t, a max and E max respectively represent the upper limits of acceleration and total energy. represents the horizontal instantaneous velocity of the friendly jamming UAV at time t. represents the vertical instantaneous velocity of the friendly jamming UAV at time t, v max represents the maximum speed of the friendly jamming UAV, E total represents the three-dimensional UAV propulsion energy model, x s,t 、y s,t and h s,t respectively represent the x-coordinate, y-coordinate and altitude of the communication UAV at time t, c represents the minimum distance to prevent collision between communication UAVs, x j,min 、x j,max 、y j,min 、y j,max 、h j,min and h j,max represent the upper and lower limits of the flight range of the friendly jamming UAV, δ t represents the time slot size, P0 and P1 respectively represent the profile power and induced power of the blade in the hovering state, and P2 represents the continuous power during the entire descent and ascent process. Denotes the average speed of the friendly interference drone at time t, U tip Denotes the tip speed of the rotor blade, d0 denotes the drag coefficient of the fuselage, ρ denotes the air density, s denotes the robustness of the rotor, A denotes the area of the rotor disk, v0 denotes the average induced speed of the rotor during hovering, Denotes the vertical average speed of the friendly interference drone at time t.
[0078] Since this problem has long-term goals and non-convex constraints, traditional optimization algorithms are difficult to handle. Therefore, the present invention utilizes a deep reinforcement learning algorithm, which can autonomously learn complex strategies and manage high-dimensional, non-linear features to solve complex problems in dynamic environments. Although the training process of the deep reinforcement learning DRL algorithm takes a certain amount of time, the trained model can use existing data to quickly generate optimized or feasible solutions, thus quickly adapting and providing reliable results in real-time or dynamic environments.
[0079] In view of the fact that the problem considered in the present invention is the safety of the overall time, the convolutional neural network (CNN) in the deep reinforcement learning algorithm DDQN cannot capture long-term dependencies when dealing with time series problems. Therefore, the present invention replaces the original convolutional neural network with a long short-term memory network (LSTM). The long short-term memory network is a special type of recurrent neural network (RNN) designed to address the long-term dependency problems faced by standard recurrent neural networks when dealing with continuous data. Standard recurrent neural networks are difficult to capture long-term dependencies and often lead to gradient vanishing or explosion problems when dealing with long sequences. The long short-term memory network solves this problem through a gating mechanism (including an input gate, a forget gate, and an output gate). This mechanism enables the long short-term memory network to effectively capture long-term dependencies and avoid the gradient vanishing or explosion problems that occur in traditional recurrent neural networks when dealing with long sequences. The core state update process of the long short-term memory network can be expressed as:
[0080] h t =o t *tanh(C t );where f t 、i t and o t
[0081] respectively denote the forget gate, the input gate, and the output gate that control the update of the storage unit. The current storage state and the candidate storage state calculated based on the current input state and the previous state are respectively denoted as C t and h t denotes the hidden state of the current time slot, and C t-1 denotes the storage state at time t-1.
[0082] Long short-term memory networks can effectively manage the information flow, ensuring that the model captures important long-term dependencies. When combined with the deep reinforcement learning algorithm DDQN, long short-term memory networks can improve the model's ability to capture temporal relationships and make accurate decisions based on long-term dependencies, especially when the model needs to process highly uncertain states during the decision-making process.
[0083] In the traditional experience replay mechanism of the deep reinforcement learning algorithm DDQN, all experience samples are selected for training with the same probability. However, this uniform sampling cannot distinguish the importance of experience samples, which may lead to the neglect of key experiences, reduce the training efficiency, and prolong the convergence time. The prioritized experience replay PER mechanism is an improved experience replay technique that assigns priorities to each experience sample according to the temporal difference (TD error) to measure its potential to improve the policy. The higher the priority of a sample, the greater the probability of being selected.
[0084] The priority of each experience is calculated as follows: p = |δ| + ∈; where p represents the priority, δ represents the temporal difference, and ∈ represents a constant used to adjust the importance of the optimization level.
[0085] Since the priority-based sampling method causes uneven sample selection, it can no longer be assumed that all samples contribute equally to policy optimization. This bias may reduce the representativeness of the samples, thereby affecting the accuracy of gradient updates and the learning effect. Therefore, it is necessary to correct the error in gradient updates. In this case, it is necessary to introduce importance sampling weights in the policy to adjust the bias caused by non-uniform sampling. The importance sampling weight is defined as: where N represents the total number of samples in the experience pool, represents the selection probability of the sample, and β' represents a hyperparameter used to control the correction amplitude. Therefore, when updating the gradient, the TD error is updated according to the following formula:
[0086] Then, the priority of the sample is re-evaluated according to the new TD error. This update mechanism enables the model to dynamically adjust the sample selection strategy. It can be seen from the above process that the prioritized experience replay strategy PER can effectively increase the sampling frequency of key experiences and make the training process of deep reinforcement learning more efficient.
[0087] In this embodiment, the action involved in the present invention refers to performing an action where, and are the acceleration components in the x and y directions at time t respectively, represents the acceleration component in the vertical direction, and the state refers to the environmental state s in this application t , The reward (the objective function for a single time slot) refers to (i.e., a new objective function cited in the present invention), where m t represents the vector at time t, and n t represents the length of the vector at time t, x s,t , y s,t and h s,t respectively represent the x-coordinate, y-coordinate and altitude of the communication UAV at time t, x j,t and y j,t respectively represent the x-coordinate and y-coordinate of the friendly jamming UAV at time t, P j,t and h j,t respectively represent the power and altitude of the friendly jamming UAV at time t, x b,t , y b,t represent the x-coordinate and y-coordinate of the ground user at time t, and represent the horizontal and vertical speeds of the jamming UAV.
[0088] In this embodiment, the optimization problem proposed by the present invention is to optimize the execution of actions t under the given environmental state s such that the objective function is maximized.
[0089] In this embodiment, the jamming strategy of the jamming UAV is obtained as follows:
[0090] 1) Receive the trajectory data of the predefined communication UAV and the ground user, the initial position of the jamming UAV, and the speed and acceleration limits;
[0091] 2) According to the position and power information of the communication UAV at each moment, obtain the threatened area ω e,t at that moment t; after discretizing the positions in the area, count the number of positions n t Construct a vector
[0092] 3) According to n t and m t at each moment, construct an overall security optimization problem;
[0093] 4) Construct a Markov decision process, s t ∈S represents the system state, the state of the system at time slot t is represented as s t , and the action taken by the jamming UAV at time slot t is represented as represented as Use the percentage of the safe area in each time slot as the reward value, and its expression is (i.e., a new objective function cited in the present invention);
[0094] (1) Initialize the network structure, including the current training network parameters θ, target network parameters θ', total number of iterations T, network parameter update step size N r , learning rate μ, sample batch size N b , discount factor γ; Replace the traditional convolutional neural network with a long short-term memory network (LSTM);
[0095] (2) Initialize the experience replay buffer D for storing the experience samples of the interaction between the agent and the environment;
[0096] (3) Solve the Q-values of all actions through the LSTM network;
[0097] (4) Interact with the environment and collect experiences, and select actions using the ε-greedy strategy , and the expression of this strategy is:
[0098]
[0099] where ε represents the probability of exploration. As the training progresses, the value of ε will gradually decrease, thus encouraging more utilization of actions with higher Q-values in the later stage of training;
[0100] (5) Execute the action , obtain the reward r feedback from the environment t and transition to the next state s t+1 , and store the experience sample in the experience replay buffer D, where i is the number of the sample in the experience pool, and p i is the corresponding priority of each sample, which is calculated by p = |δ| + ∈ and the TD error , where r i represents the system reward value of the i-th experience sample, represents the system action executed at the corresponding moment of this sample, s i+1 represents the environmental state transitioning to the next moment, end i represents whether the i-th experience sample is the last moment of a training cycle, Q' represents the goodness value of the combination of the next environmental state and a specific action, θ' represents the hyperparameters of the current training network in the deep reinforcement learning model, and θ represents the hyperparameters of the target network in the deep reinforcement learning model;
[0101] (6) Extract N b experience samples from the experience replay buffer for training, and give priority to training samples with high errors. For each experience sample, if end i is false, the target value is y i = r i , otherwise the target value is γ represents the discount factor, which is used to control the influence degree of future rewards on the current reward;
[0102] (7) Calculate the loss function according to the target value Update the network weights by minimizing the loss function;
[0103] (8) Every N r steps, copy the current network parameters to the target network to ensure the stability of Q-value estimation.
[0104] 5) When the above iteration is completed, output the optimal deep reinforcement learning neural network model for solving the optimization problem. The solution process of the optimization problem is as follows: at time slot t0, input the environmental state information s at time t0 t to the deep reinforcement learning neural network model. The deep reinforcement learning neural network model outputs the optimal action, executes the optimal action, and enters time slot t1; at time slot t1, input the environmental state information s at time slot t1 t to the deep reinforcement learning neural network model. The deep reinforcement learning neural network model outputs the optimal action, executes the optimal action, and enters time slot t2, and so on until all actions in all T time slots are completed, that is, the solution of the optimization problem is completed:
[0105] (1) The input of this solution process is the preset UAV trajectory and user trajectory, which are used as the known environmental state change information and are used to construct the environmental state at the current moment in the subsequent solution process and provide basic data for the deep reinforcement learning model;
[0106] (2) Initialize the action sequence A;
[0107] (3) Observe the trajectory of the preset UAV, the trajectory information of the user, the transmission power of the current interfering UAV, and the position information of the interfering UAV at the current t moment, format and encapsulate this information into a vector form to form the environmental state vector at the current moment
[0108] (4) Input the state vector s at the current moment t into the trained deep reinforcement learning model. In the input processing of the deep reinforcement learning model, the state vector at the current moment will be concatenated with the state information at the previous moment to form a state sequence and input into the LSTM network; this LSTM network further captures the influence of historical states on the current decision by learning the temporal dependence relationship in the state sequence and outputs a feature vector representation of the state at the current moment; next, this feature vector will be processed through multiple fully connected layers and undergo non-linear transformation through the activation function, and finally output the Q-values corresponding to all possible actions at the current moment; at the same time, a masking mechanism is used to screen the Q-values and assign infinitesimal values to the Q-values that do not meet the constraint conditions of the problem;
[0109] (5) Obtain the optimal action corresponding to the maximum Q value from the action space
[0110] (6) Execute the action
[0111] (7) Store the action in the action sequence A;
[0112] (8) Transfer the environmental state to s t+1 , and finally output the complete action sequence A to obtain the interference strategy for interfering with the UAV.
[0113] In this embodiment, the present invention uses an improved reinforcement learning method to solve the optimization problem, that is, train a reinforcement learning model. Inputting the environmental information into this model can quickly obtain the optimal action (i.e., the optimal optimization variable, including the power, acceleration, etc. of the UAV ------ the optimization goal of the present invention aims to maximize the sum of the safety region ratios accumulated in the time dimension by finely adjusting the acceleration parameter and power output of the friendly interfering UAV). The purpose of this optimal action is to maximize the reward of the action, that is, to maximize the objective function of a single time slot of the optimization problem. The objective function of the optimization problem
[0114] is the cumulative sum of all T time slots. The present invention uses the trained reinforcement learning neural network model to output the optimal action in each time slot t, and then enters the next time slot, and so on, to complete the solution of the optimization problem for all time slots.
[0115] In summary, compared with the prior art, the present invention has the following advantages:
[0116] By taking the security performance of multiple time slots as the overall optimization goal, the present invention can achieve higher long-term security. Specifically, the present invention not only considers the instantaneous secrecy capacity of a single time slot, but also ensures the continuous security of the entire transmission process through cumulative security optimization, effectively preventing the risk of information leakage caused by fluctuations in instantaneous security performance;
[0117] The present invention adopts an improved double deep Q network algorithm. By introducing a long short-term memory network, the system can effectively capture the time correlation characteristics in the UAV communication environment. This improvement enables the algorithm to make full use of historical information when making transmission parameter decisions, improving the accuracy of decision-making;
[0118] The present invention innovatively introduces a prioritized experience replay mechanism. By preferentially learning important samples, the convergence efficiency of the algorithm is significantly improved. This mechanism enables the system to learn effective secure transmission strategies faster and can better cope with sudden security threats;
[0119] The method proposed by the present invention has strong practicability and adaptability. By comprehensively considering the cumulative security performance and learning efficiency, this method can quickly establish a secure transmission strategy in various complex UAV communication scenarios and has a low computational complexity.
Claims
1. A physical layer transmission method for drones with overall security, characterized in that: The following steps are involved: S1. Obtain the location and power information of the communication drone at each moment, as well as the location of the user at each moment; S2. According to the current position and power information of the communication drone, the Euclidean distance between the communication drone and the ground node is calculated when the instantaneous signal-to-noise ratio of the potential eavesdropper Eves is equal to the threshold of the signal-to-noise ratio of the potential eavesdropper Eves, and the threat area with the communication drone on the ground as the center and R as the radius is obtained, where the Euclidean distance is recorded as R; S3, discretize the threat area in a grid manner and construct a vector of length n, where each element in the vector corresponds to a position, indicating whether the position is a safe area; S4. Establish an optimization problem for the safety of elements and positions in the vector at each moment, and sum the optimization problems of all time periods, and use the sum result as the final optimization problem, where the objective function is to maximize the sum of the safe area ratios of all time periods, and the position is the threatened area ω e The position of the discretized point in represents the position of the potential eavesdropper Eves; S5. According to the final optimization problem, the long short-term memory network LSTM and the priority experience replay PER mechanism are used in combination with the deep reinforcement learning algorithm DDQN to obtain the interference strategy for jamming drones and complete the drone physical layer transmission for overall security.
2. The overall security-oriented UAV physical layer transmission method according to claim 1 is characterized in that: The expression of the Euclidean distance is as follows: in, represents the average path loss between UAV k and ground node u, P LoS represents the line-of-sight transmission probability between the UAV and the ground unit, d k,u represents the Euclidean distance between UAV k and ground node u, β L and β N Respectively represent the path loss index of LoS and NLoS, NLoS represents non-line-of-sight link, LoS represents line-of-sight link, η LoS and η NLoS Both represent the attenuation coefficient, P NLoS represents the transmission probability of NLoS, α and β are constants, and h k represents the vertical distance from the communication UAV k to the ground, r k,u Represents the projection distance between the ground user u and the communication UAV k on the ground plane.
3. The overall safety-oriented UAV physical layer transmission method according to claim 1 is characterized in that: The expression of the instantaneous signal-to-noise ratio of the potential eavesdropper Eves is as follows: Among them, γ e represents the instantaneous signal-to-noise ratio of the potential eavesdropper Eves, P s and P j They represent the transmission power of the communication UAV and the transmission power of the friendly interference UAV respectively. and They represent the path losses from the communication UAV and the jamming UAV to the potential eavesdropper Eves, Represents the noise power of the eavesdropper Eves.
4. The overall safety-oriented UAV physical layer transmission method according to claim 1, characterized in that: The expression of the vector is as follows: Among them, m t represents the vector at time t, represents the i-th element in the vector, represents the nth element and T represents vector transpose.
5. The overall safety-oriented UAV physical layer transmission method according to claim 1, characterized in that: The final optimization problem is expressed as follows: AND total ≤E max (x s,t -x j,t ) 2 +(y s,t -y j,t ) 2 +(h s,t -h j,t ) 2 ≥c 2 x j,min ≤x j,t ≤x j,max and j,min ≤y j,t ≤y j,max h j,min ≤h j,t ≤h j,max P j,min ≤P j,t ≤P j,max Where a represents the acceleration of the friendly jamming drone, P represents the power of the friendly jamming drone, represents the i-th element in the vector at time t, ω e,t represents the threat area at time t, T represents the total time, m t represents the vector at time t, n t represents the length of the vector at time t, γ e represents the instantaneous signal-to-noise ratio of the potential eavesdropper Eves, x j,t and j,t They represent the x-coordinate and y-coordinate of the friendly jammer UAV at time t, P j,t and h j,t denote the power and height of the friendly jammer UAV at time t, respectively, and x e,i and e,i represents the x- and y-coordinates of the potential eavesdropper Eves, represents the threshold of Eves signal-to-noise ratio of potential eavesdroppers, represents the user signal-to-noise ratio threshold, γ b represents the user signal-to-noise ratio, represents the horizontal acceleration vector of the friendly jammer UAV at time t, represents the vertical acceleration vector of the friendly jammer UAV at time t, a max and E max denote the upper limits of acceleration and total energy, respectively. represents the horizontal instantaneous speed of the friendly jammer UAV at time t, represents the vertical instantaneous velocity of the friendly jammer UAV at time t, v max Indicates the maximum speed of the friendly jamming drone, E total represents the three-dimensional UAV propulsion energy model, x s,t ,y s,t and h s,t They represent the x-coordinate, y-coordinate and height of the communication UAV at time t, c represents the minimum distance between communication UAVs to prevent collision, x j,min 、x j,max ,y j,min ,y j,max 、h j,min and h j,max Indicates the upper and lower limits of the flight range of the friendly jamming drone, δ t represents the time slot size, P0 and P1 represent the profile power and induced power of the blade in the hovering state, respectively, and P2 represents the continuous power during the entire descent and ascent process. represents the average speed of the friendly jammer UAV at time t, U tip represents the tip speed of the rotor blade, d0 represents the drag coefficient of the fuselage, ρ represents the air density, s represents the strength of the rotor, A represents the area of the rotor disk, v0 represents the average induced speed of the rotor when hovering, represents the vertical average velocity of the friendly jammer UAV at time t.
6. The overall safety-oriented UAV physical layer transmission method according to claim 1 is characterized in that: The S5 comprises the following steps: Initialize the parameters and buffer of the long short-term memory network LSTM; Obtain the initial environment observation status, including the location information of the drone and the user; The long short-term memory network LSTM takes the environmental observation state as input and outputs the Q value of all possible system actions, which represents the expected value of the long-term reward that can be obtained by selecting a specific action in the current state; Select actions using the greedy strategy method e-greedy That is, the acceleration and interference power of the UAV are interfered; Execution Action Calculate the safe area ratio under the current action as the reward r t , and receive the next state s respectively t+1 , end identifier end t ; Will Store into the buffer according to priority; According to the sample optimization level, unexperienced samples are extracted from the buffer, where the expression of sampling weight is as follows: Among them, w represents the sampling weight, N represents the total number of samples in the experience pool, P represents the probability of sample selection, β' represents the hyperparameter used to control the correction amplitude, k represents the index of the sample, and p k Indicates the priority of the kth sample; According to the final optimization problem, combined with the extracted unexperienced samples, the target value is calculated, and the long short-term memory network parameter θ is updated using the loss function L(θ); Through continuous iterative optimization, the jamming strategy for jamming drones is obtained, and the drone physical layer transmission for overall security is completed.
7. The overall safety-oriented UAV physical layer transmission method according to claim 6, characterized in that: The priority expression is as follows: p=|δ|+∈ Among them, p represents the priority, δ represents the time difference, and ∈ represents the constant used to adjust the importance of the optimization level; The update expression of the time difference δ is as follows: