Non-Gaussian noise-assisted UAV covert communication method and system
By introducing the attention mechanism and multi-agent dual-delay deep deterministic policy gradient technology in drone communications, the drone trajectory is optimized, the problem of covert and secure transmission of drone communications under the assistance of non-Gaussian noise is solved, and more efficient and secure communication is achieved.
Patent Information
- Application Number
- CN202411237706.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-05
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-09-05
AI Technical Summary
Existing drone communication technology lacks effective trajectory optimization methods with the assistance of non-Gaussian noise, which increases the difficulty of designing transmission schemes. Traditional optimization algorithms have limitations and cannot meet the needs of covert and secure transmission in dynamic environments.
The attention mechanism and standard multi-agent double-delay deep deterministic policy gradient technology are used to optimize the trajectories of multiple UAVs. By establishing a non-Gaussian noise-assisted UAV covert security transmission model, combined with concealment and power constraints, the average secure covert throughput of multiple users is maximized.
It improves the stealth and security of drone communications, enhances the ability to resist eavesdroppers, optimizes drone trajectory design, and improves communication efficiency and security and covert throughput.
Smart Images

Figure CN119172774B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of drone communication technology, and in particular to a non-Gaussian noise-assisted drone covert communication method and system. Background Art
[0002] Drones are widely used in the IoT (Internet of Things) sector. Using drones for low-altitude flight to stabilize link connections can improve IoT data collection efficiency. Drones are also a viable solution for aeronautical communication services, providing support in scenarios where rapid deployment of communication infrastructure is required, thereby achieving reliable communication coverage and data transmission. Compared with traditional communication systems, drones offer greater flexibility and effectiveness. Communication between drones and legitimate users on the ground is carried out using LoS (Line of Sight) channels, which offers more reliable wireless transmission performance than ground-to-ground base station communications. The three-dimensional mobility of drones enables them to circumvent obstacles, providing ground users with faster transmission rates and greater capacity. Drones offer significant advantages in many application scenarios due to their excellent maneuverability, flexible deployment capabilities, and cost-effectiveness.
[0003] However, these advantages also bring new challenges. While drones' aerial mobility offers significant advantages over traditional ground base stations, hovering and moving require additional propulsion energy. Limited onboard energy resources restrict UAV mobility, necessitating trajectory optimization and power allocation to enhance communication performance and endurance. Furthermore, potential security risks must be considered in future applications of UAV communication systems. Due to the Loss of Segment (LoS) nature of UAV air-to-ground communications, ensuring communication security is a significant challenge.
[0004] Non-Gaussian noise-assisted technology uses non-Gaussian noise with impulsive and trailing characteristics, making it more difficult for eavesdroppers to separate the original signal from the mixed signal. By adding artificial jamming signals to the transmitted signal, legitimate users can correctly decode confidential information with minimal interference, while eavesdroppers find it difficult to obtain any valid information. Its main advantage lies in exploiting the asymmetry of wireless channels, meaning that legitimate communicating parties typically enjoy better channel conditions than illegal eavesdroppers. By deploying friendly jammers near the enemy, allowing the friendly jamming nodes to transmit interference and mix it with the useful information transmitted by the drone, this effectively hinders the activities of illegal eavesdroppers and monitors, ensuring the security and confidentiality of communications between the drone and legitimate users on the ground.
[0005] However, while introducing non-Gaussian noise technology improves security, it can also reduce communication efficiency. In drone network communication scenarios, ensuring that drones maintain high communication transmission quality while maximizing transmission throughput becomes a challenging task. In drone networks, this requires jointly optimizing the drone's flight trajectory and transmission power, as well as considering interference-tolerant transmission strategies. Therefore, to achieve covert and secure drone transmission, it is particularly important to study how to effectively manage the interference caused by non-Gaussian noise assistance and optimize the drone's trajectory to ensure communication security. Research on covert and secure drone transmission using non-Gaussian noise is urgently needed.
[0006] The use of drones as relay stations is also crucial in covert communications. Drones can effectively improve the reliability of long-distance communications while reducing the probability of detection by dynamically adjusting their position. Zhang et al. focused on drone-assisted covert communication networks, aiming to assist transmission and confuse wardens. By calculating the optimal detection threshold, an expression for the probability of interruption, and optimizing power, they achieved covert transmission. The results showed that cooperative interference and power trade-offs can effectively ensure the security and concealment of communications. Wang et al. proposed a covert communication scheme that utilizes drones and intelligent reflecting surfaces (IRSs). By determining the optimal detection threshold and optimizing the transmit power, IRS phase, and drone position, they achieved effective covert transmission, protecting it from observation by wardens. Chen et al. studied how drones can be used as relay nodes to assist in covert wireless networks, improving communication performance and avoiding detection by administrators.
[0007] In multi-UAV communication systems, a rational resource allocation strategy, including power regulation, spectrum allocation, and spatial layout, is crucial for improving covert communication performance. Deep reinforcement learning technology provides UAVs with the ability to dynamically adapt to environmental changes, enabling them to instantly optimize their communication strategies, thereby maintaining both covertness and efficiency under evolving conditions. Considering threat monitoring and detection scenarios, Li et al. designed UAV trajectories to maximize transmission rates to legitimate nodes. They employed a dual-delayed deep deterministic policy gradient method to effectively adaptively select speeds from a continuous action space. Numerical results demonstrated the significant advantages of this approach over baselines. Hu et al. studied covert communication in a single-UAV, single-user, single-monitor scenario. Aiming to maximize the average effective covert rate for the legitimate user, they employed a deep reinforcement learning algorithm to jointly optimize the UAV trajectory and transmit power. Experimental results demonstrated significant performance improvement over baseline schemes, while ensuring transmission covertness and algorithm convergence. Li et al. explored the application of cooperative UAV swarms in intelligent covert communication. By introducing a dynamic migration mechanism and a dual-deep Q-network-based strategy, they jointly optimized UAV scheduling, speed, and power, aiming to improve covertness and minimize the probability of false detection.
[0008] Through the above analysis, the inventors of this application discovered that the current drone communication technology has the following defects:
[0009] First, there is currently little research on covert and secure transmission of drones with the assistance of non-Gaussian noise, especially research on the dynamic environment of drones, covert safety requirements, and drone trajectory optimization, which increases the difficulty of designing transmission solutions.
[0010] Second, current research on optimization solutions for drone secure communication scenarios generally uses traditional optimization algorithms, but these algorithms have limitations. These include local optimality, poor adaptability to dynamic environments, inadequate decision-making timeliness, and a lack of transfer learning capabilities. These issues also pose difficulties for transmission solution design. Summary of the Invention
[0011] In view of this, the embodiments of the present application propose a non-Gaussian noise-assisted UAV covert communication method and system, develop a non-Gaussian noise-assisted UAV covert security transmission algorithm with better trajectory optimization performance and faster convergence speed, explore the use of attention mechanism and standard multi-agent dual-delay deep deterministic policy gradient technology, optimize the trajectories of multiple UAVs in non-Gaussian noise-assisted UAV covert security transmission, and effectively improve the average security and covert throughput of aerial UAV nodes in the non-Gaussian noise-assisted covert security communication scenario.
[0012] In the first aspect, an embodiment of the present application proposes a non-Gaussian noise-assisted UAV covert communication method, which includes the following steps: S1, establishing a UAV covert security communication scenario model assisted by a UAV jammer that emits non-Gaussian noise; S2, in the established model, with the goal of maximizing the average secure covert throughput of multiple users, combined with the concealment constraints and power limit constraints, establishing an optimization problem; S3, converting the optimization problem into a multi-UAV collaborative Markov decision process; S4, for the converted multi-UAV collaborative Markov decision process, using a combination of attention mechanism and standard multi-agent double-delay deep deterministic policy gradient to optimize the multi-UAV trajectory to obtain the optimal multi-UAV trajectory; S5, instructing the UAV to perform communication tasks based on the optimal multi-UAV trajectory.
[0013] Optionally, a drone covert secure communication scenario model assisted by a drone jammer that emits non-Gaussian noise is established, including: considering a drone covert secure communication scenario assisted by a drone jammer that emits non-Gaussian noise, the drone transmitter serves as an aerial mobile base station, transmitting signals to K assigned users in an environment where there are P eavesdroppers and a single monitor, and the drone jammer continuously transmits non-Gaussian noise while moving to protect the information transmission between the user and the drone transmitter from being eavesdropped and detected; the drone will optimize its own trajectory to ensure the service quality of the K assigned users, and the drone's trajectory is divided into N time slots of equal time length on the time axis. The drone maintains a constant altitude during the flight, and the drone transmitter and drone jammer are set at altitudes H, ... a With H j In the above example, a and j represent the UAV transmitter and UAV jammer, respectively. The user, monitor, and eavesdropper are regarded as ground nodes. It is assumed that each UAV knows the location of the monitor, eavesdropper, and user, but does not know the location of other UAVs. In the same time slot, at most one user receives the signal from the UAV transmitter. The eavesdropper steals information when the information is transmitted between the user and the UAV transmitter, and the monitor is responsible for monitoring the information transmission between the user and the UAV transmitter.
[0014] Alternatively, to consider adaptability, the transmission link between the UAV transmitter and the user is considered as a probabilistic LoS / NLoS channel link, and the path loss uses the probabilistic LoS / NLoS path loss model, which is expressed by the formula:
[0015]
[0016] Where PL(d,θ) represents the total path loss at distance d and elevation angle θ, where distance d is the straight-line distance between the UAV transmitter and the user, and elevation angle θ is the vertical angle between the UAV transmitter and the ground. represents the probability that the signal propagation path is a direct path at the elevation angle θ, τ and are environmental parameters adjusted based on the environment, represents the probability that the signal propagation path is a non-direct path at the elevation angle θ;
[0017] For user k, the received signal is expressed as:
[0018]
[0019] in, represents the signal received by user k on the i-th channel, P a [n] represents the transmission power of the UAV transmitter, PL(d ak ,θ a ) represents the path loss between the UAV transmitter and user k, d ak represents the straight-line distance between the UAV transmitter and user k, θ a Indicates the vertical angle between the drone transmitter and the ground, X a (i) represents the signal transmitted by the UAV transmitter, P j [n] represents the transmission power of the UAV jammer, PL(d jk ,θ j ) represents the path loss between the UAV jammer and user k, d jk represents the straight-line distance between the UAV jammer and user k, θ j represents the vertical angle between the UAV jammer and the ground, χ j (i) represents the non-Gaussian noise emitted by the UAV jammer, n k (i) represents the Gaussian white noise received by user k;
[0020] For the eavesdropper p, the received signal is expressed as:
[0021]
[0022] in, represents the signal received by the eavesdropper p on the i-th channel, PL(d ap ,θ a ) represents the path loss between the UAV transmitter and the eavesdropper p, d ap represents the straight-line distance between the drone transmitter and the eavesdropper p, PL(d jp ,θ j ) represents the path loss between the UAV jammer and the eavesdropper p, d jp represents the straight-line distance between the UAV jammer and the eavesdropper p, n p (i) represents the Gaussian white noise received by the eavesdropper p;
[0023] The monitor determines whether the drone transmitter is transmitting information based on the L signals received in each time slot. The signal received by the monitor is expressed as:
[0024]
[0025] Among them, H0 represents the null hypothesis that the UAV transmitter does not transmit information, H1 represents the alternative hypothesis that the UAV transmitter transmits information, and PL(d jw ,θ j ) represents the path loss between the UAV jammer and the monitor, d jw represents the straight-line distance between the UAV jammer and the monitor, PL(d aw ,θ a ) represents the path loss between the UAV transmitter and the monitor, d aw Indicates the straight-line distance between the UAV transmitter and the monitor, n w (i) represents the Gaussian white noise received by the monitor.
[0026] Optionally, in the established model, with the goal of maximizing the average secure covert throughput of multiple users, an optimization problem is established by combining the covertness constraint and the power limit constraint, including:
[0027] Assume that the monitor's decision is divided into two types: D0 supporting H0 and D1 supporting H1. Assume that the false alarm probability and missed detection probability are represented by P FA (t) and P MD (t) represents the total detection error probability ξ(t) is defined as ξ(t) = P FA (t)+P MD (t), the goal of the drone is to make ξ(t) in each time slot as close to 1 as possible, which can be expressed as:
[0028] ξ(t)=P FA (t)+P MD (t)≥1-ε
[0029] Among them, ε is a preset positive decimal close to 0 to ensure the confidentiality of communication;
[0030] Under the assumptions of H0 and H1, The likelihood functions are expressed as:
[0031]
[0032] Where CN(·) represents a complex function, represents the variance of the signal received by the monitor;
[0033] Assuming that the monitor knows the index information of the signal, the likelihood ratio of the best detection is expressed as:
[0034]
[0035] in, represents the likelihood ratio of the best detection;
[0036] The KL divergence is used to derive the lower limit of the minimum detection error probability. The lower limit of the minimum detection error probability is expressed as:
[0037]
[0038] Among them, ξ * [n] represents the lower limit of the minimum detection error probability, D 01 [n]≤2ε 2 Denotes the constraint of detection probability. When the constraint of detection probability is satisfied, the concealment constraint will also be satisfied. 01 [n] is expressed by the formula:
[0039]
[0040] Where I represents the identity matrix;
[0041] Based on the constraints of detection probability and stealth, the flight trajectory, transmission power and interference power of UAV transmitters and UAV jammers are optimized to maximize the average secure and covert throughput of multiple users. The UAV transmitter dynamically adjusts its position to provide high-quality services, and the UAV jammer adjusts its position to reduce the link quality of eavesdroppers and monitors, thus solving the optimization problem.
[0042] Optionally, the optimization problem is formulated as follows:
[0043]
[0044]
[0045] in, represents the average secure covert throughput of multiple users, C1 represents the concealment constraint, ε2 represents the power constraint of the UAV transmitter, represents the maximum transmission power of the UAV transmitter, C3 represents the power constraint of the UAV jammer, represents the maximum transmission power of the UAV jammer, C4 represents the flight speed constraint of the UAV, q[n] and q[n+1] represent the positions of the UAV at sampling points n and n+1 respectively, V maxrepresents the maximum flight speed of the UAV, T represents the time required for the UAV to fly from q[n] to q[n+1], C5 represents the flight range constraint of the UAV, and R D Indicates the scene space corresponding to the established model.
[0046] Optionally, the optimization problem is converted into a multi-UAV cooperative Markov decision process, including:
[0047] The UAV transmitter and UAV jammer are regarded as intelligent agents, forming a partially observable state space. The optimization problem is described by a five-tuple (S, A, Reward, PST, γ), where S represents the state space, A represents the action space, Reward represents the reward function, PST represents the state transition probability, and γ represents the discount factor.
[0048] The state space S is defined as the current horizontal position of the UAV transmitter and the current horizontal position of the drone jammer
[0049] The action space is defined as the actions of the UAV transmitter and the actions of drone jammers For the UAV transmitter, the action space includes the speed in the x-axis and y-axis directions and the magnitude of the transmission power. The action space of the UAV jammer includes the speed in the x-axis and y-axis directions and the magnitude of the jamming power.
[0050] In a collaborative drone environment, the decision-making process involves each agent receiving an immediate reward based on its current state, the action it performs, and the next state it transitions to, taking into account termination signals. The reward mechanism includes penalties for the drone's flight range, concealment, and distance. The reward and penalty terms are composed as follows:
[0051] Secure and concealed throughput rewards, Incentivize drones to increase safe and covert throughput;
[0052] Boundary penalty, p b =1, if q[n]>R D ,else0, is applied when the UAV exceeds the predetermined boundary to ensure that the UAV operates within the specified area;
[0053] Hidden punishment, p ∈ =1, if D 01 [n]>∈ 2 ,else0, is imposed when the minimum detection error probability of the monitor is lower than the threshold to ensure the concealment of communication;
[0054] Distance penalty,
[0055]
[0056] Among them, distance penalty and They are used to encourage drone transmitters and drone jammers to fly around the target, encourage drone transmitters to fly around users, and encourage drone jammers to fly around eavesdroppers and monitors respectively;
[0057] The overall reward function of the drone transmitter and drone jammer is expressed as:
[0058]
[0059] Among them, α and β are coefficients that weigh the impact of different rewards and penalties, and Ω1 and Ω2 are the weights of the corresponding penalty terms;
[0060] In a UAV collaborative environment, the decisions and actions of each UAV will affect the overall system state. The state transition function describes the probability of the system reaching the next state given the current state and the UAV action, which is expressed by the formula:
[0061]
[0062] In this collaborative environment, there is no clear endpoint. The goal of the UAVs is to maximize the total communication capacity during the flight, and the choice of the discount factor γ will reflect this long-term performance indicator.
[0063] Optionally, for the converted multi-UAV cooperative Markov decision process, the multi-UAV trajectory is optimized using a combination of an attention mechanism and a standard multi-agent double-delay deep deterministic policy gradient to obtain the optimal multi-UAV trajectory, including:
[0064] S41, initialize the network parameters of the Actor network, Critic network and attention network of the drone transmitter
[0065] S42, initialize the network parameters of the Actor network, Critic network and attention network of the drone jammer
[0066] S43, setting the network parameters of the target network to be the same as those of the main network, i.e.
[0067] S44: If the convergence condition is met, go to S48; otherwise, go to S44.1.
[0068] S44.1, UAV transmitter and UAV jammer observe their own status respectively and And select actions respectively and
[0069] S44.2, Execution Action in the environment;
[0070] S44.3, observe the next state The reward value r and the completion signal done are used to clarify s n+1 Whether it is a terminal state;
[0071] S44.4, Storage In playback buffer B;
[0072] S44.5, if the UAV flies out of the boundary, does not meet the concealment constraint, or the number of time slots reaches the upper limit, reset the environment state;
[0073] S44.6, when the preset number of updates is reached, set j = 1 and proceed to S45; otherwise, proceed to S47;
[0074] S45, if j≤k, then go to S45.1, otherwise go to S47;
[0075] S45.1, randomly sample a mini-batch of size SG from the replay buffer B;
[0076] S45.2, the target action of the UAV transmitter is calculated as:
[0077]
[0078] S45.3, the drone jammer calculates the target action as:
[0079]
[0080] S45.4, calculate the target value y of the UAV transmitter A and the target value y of the UAV jammer J for:
[0081]
[0082] S45.5, update the network parameters of the Critic network and the attention network of the drone transmitter and drone jammer by minimizing the loss as follows:
[0083]
[0084] S45.6, if j is the time point of the policy update, proceed to S46, otherwise proceed to S47;
[0085] S46, using the gradient ascent method to update the Actor network of the drone transmitter and drone jammer is as follows:
[0086]
[0087] Gradually update the target network as follows:
[0088]
[0089] S47, ending the policy update cycle;
[0090] S48, obtain the optimal multi-UAV trajectory.
[0091] Through the above approach, this application establishes a scenario model for covert and secure drone communication, assisted by a friendly drone jammer emitting non-Gaussian noise. With the goal of maximizing the average secure and covert throughput of multiple users, this application formulates an optimization problem under the constraints of stealth and power limits. This optimization problem is converted into a multi-drone collaborative Markov decision process. The multi-drone trajectories are optimized using a combination of an attention mechanism and a standard multi-agent double-delay deep deterministic policy gradient algorithm, filling a gap in this field. The use of multi-agent deep reinforcement learning methods can effectively address non-stationary problems, while the use of attention mechanisms can effectively capture key information and accelerate convergence, thereby improving the trajectory optimization performance of the algorithm. Compared to currently proposed technologies, this application explores the centralized decision-making problem in a multi-agent cooperative environment, emphasizing the importance of solving this centralized decision-making problem in non-Gaussian noise-assisted covert and secure drone transmission. Any non-Gaussian noise-assisted covert and secure drone transmission task can be accomplished using the technical solution proposed in this application, which has better generalization and wider versatility.
[0092] On the second aspect, an embodiment of the present application proposes a non-Gaussian noise-assisted UAV covert communication system, the system comprising: a scenario model establishment module for establishing a UAV covert secure communication scenario model assisted by a UAV jammer that emits non-Gaussian noise; an optimization problem establishment module for establishing an optimization problem in the established model with the goal of maximizing the average secure covert throughput of multiple users, combined with concealment constraints and power limit constraints; a Markov conversion module for converting the optimization problem into a multi-UAV collaborative Markov decision process; a multi-UAV trajectory optimization module for optimizing the multi-UAV trajectory based on the converted multi-UAV collaborative Markov decision process, using a combination of attention mechanism and standard multi-agent double-delay deep deterministic policy gradient to obtain the optimal multi-UAV trajectory; a communication task execution module for instructing the UAV to perform the communication task based on the optimal multi-UAV trajectory.
[0093] In a third aspect, an embodiment of the present application proposes an electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a non-Gaussian noise-assisted drone covert communication method as described in the first aspect above.
[0094] In a fourth aspect, an embodiment of the present application proposes a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement a non-Gaussian noise-assisted drone covert communication method as described in the first aspect above.
[0095] It can be understood that the beneficial effects of the second to fourth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0096] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the related technologies, the following is a brief introduction to the drawings required for use in the embodiments of the present application or the description of the related technologies. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0097] Figure 1 This is a flowchart of a non-Gaussian noise-assisted UAV covert communication method provided in one embodiment of the present application;
[0098] Figure 2 is a structural diagram of a non-Gaussian noise-assisted UAV covert communication system provided in another embodiment of the present application;
[0099] Figure 3 This is a schematic diagram of the simulation experimental results of different algorithm convergence curves of a non-Gaussian noise-assisted UAV covert security transmission system provided in another embodiment of the present application;
[0100] Figure 4 It is a structural diagram of an electronic device provided in another embodiment of the present application. DETAILED DESCRIPTION
[0101] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the embodiments of the present application will be described in detail below with reference to the accompanying drawings. In the various embodiments of the present application, many technical details are proposed to enable the reader to better understand the present application. However, even without these technical details and various changes and modifications based on the following embodiments, the technical solutions claimed in the present application can be implemented. The division of the following embodiments is only for the convenience of description and should not constitute any limitation on the specific implementation of the present application. The various embodiments can be combined with each other and referenced to each other under the premise of no contradiction.
[0102] An embodiment of the present application proposes a non-Gaussian noise-assisted UAV covert communication method, which is applied to electronic devices, wherein the electronic device can be a terminal or a server. In this embodiment and the following embodiments, the electronic device is described using a server as an example. The following describes the implementation details of the non-Gaussian noise-assisted UAV covert communication method proposed in this embodiment. The following content is only for the convenience of understanding the implementation details, and is not necessary for the implementation of this solution. The specific process of the non-Gaussian noise-assisted UAV covert communication method proposed in this embodiment is as follows: Figure 1 Shown, including:
[0103] S1, establishes a UAV covert secure communication scenario model assisted by a UAV jammer that emits non-Gaussian noise.
[0104] In the specific implementation, the server first establishes a drone covert secure communication scenario model assisted by a drone jammer that emits non-Gaussian noise. The drone jammer that emits non-Gaussian noise is considered to be friendly.
[0105] In one example, when establishing a drone covert secure communication scenario model assisted by a drone jammer that emits non-Gaussian noise, the server considers a drone covert secure communication scenario assisted by a drone jammer that emits non-Gaussian noise. The drone transmitter acts as an aerial mobile base station and transmits signals to K assigned users (also known as legal users, and users not assigned are illegal users) in an environment with P eavesdroppers and a single monitor. The drone jammer continuously transmits non-Gaussian noise while moving to protect the information transmission between users and the drone transmitter from being eavesdropped and detected. To ensure the service quality of the K assigned users, the drone will optimize its own trajectory. The drone's trajectory is divided into N time slots of equal length on the time axis. The drone maintains a constant altitude during flight, and the drone transmitter and drone jammer are set at altitudes H, ... a With H jIn this example, a and j represent the UAV transmitter and UAV jammer, respectively. The user, monitor, and eavesdropper are considered ground nodes. It is assumed that each UAV knows the location of the monitor, eavesdropper, and user, but does not know the location of other UAVs (i.e., the UAV transmitter and UAV jammer are mutually unaware of each other's locations). In the same time slot, at most one user receives a signal from a UAV transmitter. The eavesdropper steals information during transmission between the user and the UAV transmitter, and the monitor monitors the information transmission between the user and the UAV transmitter.
[0106] In one example, to consider adaptability, the transmission link between the drone transmitter and the user is considered as a probabilistic LoS / NLoS channel link, and the path loss uses the probabilistic LoS / NLoS path loss model, which is expressed by the formula:
[0107]
[0108] Where PL(d,θ) represents the total path loss at distance d and elevation angle θ, where distance d is the straight-line distance between the UAV transmitter and the user, and elevation angle θ is the vertical angle between the UAV transmitter and the ground. represents the probability that the signal propagation path is a direct path at the elevation angle θ, τ and ζ are environmental parameters adjusted based on the environment, It represents the probability that the signal propagation path is a non-direct path at the elevation angle θ.
[0109] Therefore, for user k, the received signal can be expressed as:
[0110]
[0111] in, represents the signal received by user k on the i-th channel, P a [n] represents the transmission power of the UAV transmitter, PL(d ak ,θ a ) represents the path loss between the UAV transmitter and user k, d ak represents the straight-line distance between the UAV transmitter and user k, θ a represents the vertical angle between the UAV transmitter and the ground, χ a (i) represents the signal transmitted by the UAV transmitter, P j [n] represents the transmission power of the UAV jammer, PL(d jk ,θ j ) represents the path loss between the UAV jammer and user k, d jk represents the straight-line distance between the UAV jammer and user k, θ j represents the vertical angle between the UAV jammer and the ground, χj (i) represents the non-Gaussian noise emitted by the UAV jammer, n k (i) represents the Gaussian white noise received by user k.
[0112] For the eavesdropper p, the received signal can be expressed as:
[0113]
[0114] in, represents the signal received by the eavesdropper p on the i-th channel, PL(d ap ,θ a ) represents the path loss between the UAV transmitter and the eavesdropper p, d ap represents the straight-line distance between the drone transmitter and the eavesdropper p, PL(d jp ,θ j ) represents the path loss between the UAV jammer and the eavesdropper p, d jp represents the straight-line distance between the UAV jammer and the eavesdropper p, n p (i) represents the Gaussian white noise received by the eavesdropper p.
[0115] In a covert communication scenario, the main task of the monitor is to detect whether the drone has established an information transmission connection with the ground user. To make a decision, the monitor determines whether the drone transmitter is transmitting information based on the L signals received in each time slot. In the nth time slot, the signal received by the monitor can be expressed as:
[0116]
[0117] Among them, H0 represents the null hypothesis that the UAV transmitter does not transmit information, H1 represents the alternative hypothesis that the UAV transmitter transmits information, and PL(d jw ,θ j ) represents the path loss between the UAV jammer and the monitor, d jw represents the straight-line distance between the UAV jammer and the monitor, PL(d aw ,θ a ) represents the path loss between the UAV transmitter and the monitor, d aw Indicates the straight-line distance between the UAV transmitter and the monitor, n w (i) represents the Gaussian white noise received by the monitor.
[0118] Non-Gaussian noise refers to noise whose probability distribution deviates from the Gaussian distribution. Unlike Gaussian noise, non-Gaussian noise is characterized by a heavy-tailed distribution and impulsiveness, which means that the noise may contain high-amplitude bursts of interference. This type of noise is common in many practical communication systems, especially in fields such as wireless communications and radar. Due to the two characteristics of heavy-tailed distribution and impulsiveness, the non-Gaussian noise model can make it more difficult for eavesdroppers to separate our signals from mixed signals. This embodiment uses the α-stable distribution model, which is more difficult to separate among non-Gaussian noise models.
[0119] The α-stable distribution model makes it more difficult for eavesdroppers to separate their own signals from mixed signals. The α-stable distribution has a heavy-tailed characteristic, meaning its probability density function may be high in the center, but generally does not approach 1 over most of the region. Because the α-stable distribution lacks a simple closed form, its characteristic function exhibits large uncertainty and a wide range of influence, making it difficult for its probability density function to approach 1 over a large range. Gaussian noise, on the other hand, has a uniform power spectral density and simple statistical properties, making it easier for eavesdroppers to obtain signals.
[0120] The probability density function of the α-stable distribution has no closed form, but can be expressed by the characteristic function as follows:
[0121]
[0122] Where α is the characteristic exponent, β is the skew parameter, γ is the scale parameter, and δ is the location parameter. When α = 2, the distribution degenerates into a Gaussian distribution. The probability density function of the α-stable distribution has a severe tailing property, and the smaller the impulse exponent α, the more severe the tailing property. Since the α-stable distribution noise has a heavy tail characteristic, its noise power P n It is affected by heavy tails and extreme values, which makes the signal-to-noise ratio of the signal after the matched filter expressed as follows:
[0123]
[0124] in, is the variance of α-stable distributed noise, which is typically high and unstable, especially when α is small. In a Gaussian noise background, a matched filter can effectively separate signals and achieve an optimal output signal-to-noise ratio. However, in an α-stable noise background, due to the heavy tails and extreme values of the noise, the performance of the matched filter degrades significantly, and the output signal-to-noise ratio is much lower than that of Gaussian noise. The statistical properties of Gaussian noise make signal separation relatively simple and efficient; however, the complex statistical properties and extreme values of α-stable distributed noise increase the difficulty of signal separation.
[0125] The heavy-tailed nature of non-Gaussian noise and its lack of a closed-form probability density function make it difficult to apply traditional signal processing methods. Extreme values and high-peak interference significantly impair signal separation. Using non-Gaussian noise in communication systems can significantly improve the security and stealth of communication links, making it more difficult for eavesdroppers to separate and steal information. The use of non-Gaussian noise, particularly α-stable distributed noise, can effectively enhance the anti-interference capability and information security of communication systems, preventing eavesdroppers from obtaining useful information.
[0126] S2, in the established model, with the goal of maximizing the average secure covert throughput of multiple users, combined with the concealment constraints and power limit constraints, an optimization problem is established.
[0127] In the specific implementation, after the server establishes a drone covert secure communication scenario model assisted by a drone jammer that emits non-Gaussian noise, it can establish an optimization problem in the established model with the goal of maximizing the average secure covert throughput of multiple users, combining the concealment constraints and power limit constraints.
[0128] In an example, suppose the monitor's decision is divided into two types: D0 supporting H0 and D1 supporting H1. Let the false alarm probability and missed detection probability be expressed as P FA (t) and P MD (t) represents the total detection error probability ξ(t) is defined as ξ(t) = P FA (t)+P MD (t), the goal of the drone is to make ξ(t) in each time slot as close to 1 as possible, which can be expressed as:
[0129] ξ(t)=P FA (t)+P MD (t)≥1-ε
[0130] Here, ε is a preset positive decimal close to 0 to ensure the confidentiality of communication.
[0131] Under the assumptions of H0 and H1, The likelihood functions are expressed as:
[0132]
[0133] Where CN(·) represents a complex function, Represents the variance of the signal received by the monitor.
[0134] Assuming that the monitor knows the index information of the signal, the likelihood ratio of the best detection is expressed as:
[0135]
[0136] in, represents the likelihood ratio of the best detection.
[0137] exist Based on this, the minimum detection error probability ξ can be derived * , however, due to the incomplete gamma function involved in ξ * This brings difficulties to further analysis. Therefore, we turn to using KL divergence to derive the lower limit of the minimum detection error probability. The lower limit of the minimum detection error probability is expressed as:
[0138]
[0139] Among them, ξ * [n] represents the lower limit of the minimum detection error probability, D 01 [n]≤2ε 2 It represents the constraint of detection probability. When the constraint of detection probability is satisfied, the concealment constraint will also be satisfied.
[0140] D 01 [n] is expressed by the formula:
[0141]
[0142] Here, I represents the identity matrix. Thus, an effective method for evaluating the effectiveness of covert communication is provided, which ensures the concealment of UAV communication and simplifies the analysis of the monitor's detection behavior.
[0143] Based on the constraints of detection probability and stealth, the server can maximize the average secure and covert throughput of multiple users by optimizing the flight trajectory, transmission power and interference power of the UAV transmitter and UAV jammer. The UAV transmitter dynamically adjusts its position to provide high-quality services, and the UAV jammer adjusts its position to reduce the link quality of the eavesdropper and monitor, thus solving the optimization problem.
[0144] In an example, the optimization problem can be expressed as follows:
[0145]
[0146] in, represents the average secure covert throughput of multiple users, C1 represents the concealment constraint, and C2 represents the power constraint of the UAV transmitter. represents the maximum transmission power of the UAV transmitter, C3 represents the power constraint of the UAV jammer, represents the maximum transmission power of the UAV jammer, C4 represents the flight speed constraint of the UAV, q[n] and q[n+1] represent the positions of the UAV at sampling points n and n+1 respectively, Vmax represents the maximum flight speed of the UAV, T represents the time required for the UAV to fly from q[n] to q[n+1], C5 represents the flight range constraint of the UAV, and R D It represents the scene space corresponding to the established model. It is worth noting that C2 and C3 are determined by the hardware and energy limitations of the UAV transmitter and UAV jammer.
[0147] S3, converts the optimization problem into a multi-UAV collaborative Markov decision process.
[0148] In a specific implementation, after establishing the optimization problem, the server can convert the optimization problem into a multi-UAV collaborative Markov decision process.
[0149] In one example, when converting the optimization problem into a multi-UAV collaborative Markov decision process, the server first regards the UAV transmitter and UAV jammer as intelligent agents, forming a partially observable state space, and describes the optimization problem through a five-tuple (S, A, Reward, PST, γ), where S represents the state space, A represents the action space, Reward represents the reward function, PST represents the state transition probability, and γ represents the discount factor.
[0150] The state space S is defined as the current horizontal position of the UAV transmitter and the current horizontal position of the drone jammer
[0151] The action space is defined as the actions of the UAV transmitter and the actions of drone jammers For the UAV transmitter, the action space includes the speed in the x-axis and y-axis directions and the magnitude of the transmission power. The action space of the UAV jammer includes the speed in the x-axis and y-axis directions and the magnitude of the jamming power.
[0152] In a collaborative drone environment, the decision-making process involves each agent receiving an immediate reward based on its current state, the action it performs, and the next state it transitions to, taking into account termination signals. The reward mechanism includes penalties for the drone's flight range, concealment, and distance. The reward and penalty terms are composed as follows:
[0153] Secure and concealed throughput rewards, Incentivize drones to improve security and covert throughput.
[0154] Boundary penalty, p b =1, if q[n]>R D ,else0, is applied when the UAV exceeds the predetermined boundary to ensure that the UAV operates within the specified area.
[0155] Hidden punishment, p ∈ =1, if D 01 [n]>∈ 2 ,else0, is imposed when the minimum detection error probability of the monitor is lower than the threshold to ensure the confidentiality of communication.
[0156] Distance penalty,
[0157]
[0158] Among them, distance penalty and They are used to encourage drone transmitters and drone jammers to fly around the target, encourage drone transmitters to fly around users, and encourage drone jammers to fly around eavesdroppers and monitors.
[0159] The overall reward function of the drone transmitter and drone jammer is expressed as:
[0160]
[0161] Among them, α and β are coefficients that weigh the impact of different rewards and penalties, and Ω1 and Ω2 are the weights of the corresponding penalty terms.
[0162] In a UAV collaborative environment, the decisions and actions of each UAV will affect the overall system state. The state transition function describes the probability of the system reaching the next state given the current state and the UAV action, which is expressed by the formula:
[0163] In this collaborative environment, there is no clear endpoint and a smaller discount factor γ will not be set. The goal of the drone is not to reach a certain location as quickly as possible, but to maximize the total communication capacity during the flight. Therefore, the choice of the discount factor γ will reflect this long-term performance indicator rather than the short-term fastest arrival time.
[0164] S4, for the converted multi-UAV collaborative Markov decision process, uses the combination of attention mechanism and standard multi-agent double-delay deep deterministic policy gradient to optimize the multi-UAV trajectory and obtain the optimal multi-UAV trajectory.
[0165] In the specific implementation, the server optimizes the multi-UAV trajectories based on the converted multi-UAV cooperative Markov decision process using a combination of attention mechanism and standard multi-agent double-delay deep deterministic policy gradient to obtain the optimal multi-UAV trajectory. This can be achieved through the following steps:
[0166] S41, initialize the network parameters of the Actor network, Critic network and attention network of the drone transmitter
[0167] S42, initialize the network parameters of the Actor network, Critic network and attention network of the drone jammer
[0168] S43, setting the network parameters of the target network to be the same as those of the main network, i.e.
[0169] S44: If the convergence condition is met, go to S48; otherwise, go to S44.1.
[0170] S44.1, UAV transmitter and UAV jammer observe their own status respectively and And select actions respectively and
[0171] S44.2, Execution Action In the environment.
[0172] S44.3, observe the next state The reward value r and the completion signal done are used to clarify s n+1 Whether it is a terminated state.
[0173] S44.4, Storage In playback buffer B.
[0174] S44.5, if the UAV flies out of the boundary, does not meet the concealment constraint, or the number of time slots reaches the upper limit, reset the environment state.
[0175] S44.6, when the preset number of updates is reached, set j=1 and go to S45, otherwise go to S47.
[0176] S45, if j≤k, then go to S45.1, otherwise go to S47.
[0177] S45.1, randomly sample a mini-batch of size SG from the replay buffer B.
[0178] S45.2, the target action of the UAV transmitter is calculated as:
[0179]
[0180] S45.3, the drone jammer calculates the target action as:
[0181]
[0182] S45.4, calculate the target value y of the UAV transmitter A and the target value y of the UAV jammer J for:
[0183]
[0184] S45.5, update the network parameters of the Critic network and the attention network of the drone transmitter and drone jammer by minimizing the loss as follows:
[0185]
[0186] S45.6, if j is the time point of the policy update, proceed to S46, otherwise proceed to S47.
[0187] S46, using the gradient ascent method to update the Actor network of the drone transmitter and drone jammer is as follows:
[0188]
[0189] Gradually update the target network as follows:
[0190]
[0191] S47, ending the policy update cycle;
[0192] S48, obtain the optimal multi-UAV trajectory.
[0193] S5 instructs the UAV to perform the communication task based on the optimal multi-UAV trajectory.
[0194] In the specific implementation, after the server obtains the optimal multi-UAV trajectory, it can instruct the UAV transmitter and UAV jammer to perform communication tasks based on the optimal multi-UAV trajectory.
[0195] This example establishes a scenario model for covert secure communication between drones (UAVs) assisted by a friendly UAV jammer emitting non-Gaussian noise. With the goal of maximizing the average secure and covert throughput of multiple users, an optimization problem is formulated under the constraints of stealth and power limits. This optimization problem is converted into a multi-UAV collaborative Markov decision process. Multi-UAV trajectories are optimized using a standard multi-agent double-delay deep deterministic policy gradient algorithm combined with an attention mechanism, thus filling the gap in non-Gaussian noise-assisted UAV covert communication. The use of a multi-agent deep reinforcement learning method effectively addresses non-stationary problems, while the use of an attention mechanism effectively captures key information and accelerates convergence, thereby improving the trajectory optimization performance of the algorithm. Compared to currently proposed technologies, this example investigates the centralized decision-making problem in a multi-agent cooperative environment, emphasizing the importance of solving this problem in non-Gaussian noise-assisted UAV covert secure transmission. The technical solution of this example can be used to accomplish any non-Gaussian noise-assisted UAV covert secure transmission task, demonstrating improved generalization and versatility.
[0196] The step division of the above various methods is only for the purpose of clear description. During implementation, they can be combined into one step or some steps can be split and decomposed into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this application; adding insignificant modifications or introducing insignificant designs to the algorithm or process without changing the core design of the algorithm and process are all within the scope of protection of this application.
[0197] Another embodiment of the present application proposes a non-Gaussian noise-assisted UAV covert communication system. The following is a detailed description of the implementation details of the non-Gaussian noise-assisted UAV covert communication system proposed in this embodiment. The following content is only for the convenience of understanding the implementation details and is not necessary for the implementation of this embodiment. Figure 2 This is a structural diagram of a non-Gaussian noise-assisted UAV covert communication system proposed in this embodiment. The system includes: a scene model establishment module 601, an optimization problem establishment module 602, a Markov conversion module 603, a multi-UAV trajectory optimization module 604 and a communication task execution module 605.
[0198] The scenario model building module 601 is used to build a UAV covert secure communication scenario model assisted by a UAV jammer that emits non-Gaussian noise.
[0199] The optimization problem establishment module 602 is used to establish an optimization problem in the established model with the goal of maximizing the average secure covert throughput of multiple users, combined with the concealment constraint condition and the power limit constraint condition.
[0200] The Markov transformation module 603 is used to transform the optimization problem into a multi-UAV collaborative Markov decision process.
[0201] The multi-UAV trajectory optimization module 604 is used to optimize the multi-UAV trajectory based on the converted multi-UAV collaborative Markov decision process using a combination of attention mechanism and standard multi-agent double-delay deep deterministic policy gradient to obtain the optimal multi-UAV trajectory.
[0202] The communication task execution module 605 is used to instruct the UAV to execute the communication task based on the optimal multi-UAV trajectory.
[0203] It is worth mentioning that all modules involved in this embodiment are logical modules. In actual applications, a logical unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, to highlight the innovation of this application, this embodiment does not include units that are not closely related to solving the technical problem proposed by this application. However, this does not mean that other units do not exist in this embodiment.
[0204] It is not difficult to find that this embodiment is a system embodiment corresponding to the above-mentioned method embodiment, and this embodiment can be implemented in conjunction with the above-mentioned method embodiment. The relevant technical details and technical effects mentioned in the above-mentioned embodiments are still valid in this embodiment, and to reduce repetition, they are not repeated here. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above-mentioned embodiments.
[0205] In another embodiment, in order to evaluate the performance of the non-Gaussian noise-assisted UAV covert communication method and system proposed in this application, we conducted relevant simulation experiments. In the simulation experiment, a non-Gaussian noise-assisted UAV covert security transmission system is considered. The specific parameters of the simulation experiment are as follows: Consider a UAV communication system consisting of three users, and assume that there are two eavesdroppers trying to intercept the communication content. The concealment constraint ∈ is set to 0.05. The maximum speed of the UAV transmitter and the UAV jammer are both set to To simulate a fast-moving aerial environment. The optimization algorithm used in training is the Adam algorithm, with the replay buffer pool set to 10,000, the discount factor set to 0.99, the learning rate set to 0.001, and the noise exploration space set to 5%. In the reward function, Ω1 is set to 1, Ω2 is set to 100, the reward coefficient α is set to 2, and β is set to 20. The range boundaries of the drone's flight are set from -100 to 600 on the x-axis and y-axis, indicating the flight area limits of the drone in these two axes. The noise is set to 5% of the action space;
[0206] Figure 3 The convergence curves of different algorithms are shown. Figure 3 As can be seen, the standard multi-agent double-delayed deep deterministic policy gradient algorithm (AM-MATD3) combined with the attention mechanism significantly outperforms other algorithms in both learning speed and final reward. It exhibits particularly rapid learning in the initial stages and stable reward growth with minimal fluctuations as training progresses, reflecting the stability of its learning process. This rapid and stable convergence demonstrates that AM-MATD3 is able to effectively learn effective policies in the environment and optimize these policies for higher long-term rewards. In multi-agent environments, collaboration between agents is key to improving overall performance. By introducing the attention mechanism, the AM-MATD3 algorithm allows each agent to more effectively learn how to interact and collaborate with other agents amidst the dynamics of the group. This stands in stark contrast to the single-agent TD3 algorithm, where agents are trained independently and are unable to capture the behavioral strategies of other agents. Consequently, in multi-agent settings, the performance of single-agent TD3 is limited, failing to achieve the same results as collaborative learning. The attention mechanism strengthens the interaction between agents, enabling the AM-MATD3 algorithm to achieve superior performance in complex multi-agent environments. By introducing an attention mechanism, each agent can focus more on environmental factors and the behaviors of other agents that are crucial to the success of its strategy. This focused learning approach is extremely valuable in multi-agent environments, fostering collaboration and accelerating the learning process. Therefore, the trajectory optimization performance of the proposed algorithm surpasses that of currently proposed algorithms, and simulation experiments have verified its superiority.
[0207] Another embodiment of the present application provides an electronic device, the structure of which can be as follows: Figure 4 As shown, it includes: at least one processor 701; and a memory 702 communicatively connected to the at least one processor 701; wherein the memory 702 stores instructions that can be executed by the at least one processor 701, and the instructions are executed by the at least one processor 701 to enable the at least one processor 701 to execute a non-Gaussian noise-assisted UAV covert communication method described in the above-mentioned method embodiments.
[0208] The memory and processor can be connected using a bus. The bus can include any number of interconnected buses and bridges, connecting various circuits of one or more processors and memories. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits. These are all well known in the art and will not be described in detail herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be a single component or multiple components, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on a transmission medium. Data processed by the processor is transmitted via an antenna on a wireless medium. Furthermore, the antenna receives data and transmits it to the processor.
[0209] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory can be used to store data used by the processor when performing operations.
[0210] Another embodiment of the present application relates to a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement the non-Gaussian noise-assisted drone covert communication method described in the above method embodiments.
[0211] That is, those skilled in the art will understand that all or part of the steps in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a program, which is stored in a storage medium and includes a number of instructions for causing a device (which may be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: a USB flash drive, a mobile hard drive, a ROM (Read-Only Memory), a RAM (Random Access Memory), a magnetic disk, or an optical disk, etc., various media that can store program code.
[0212] Those skilled in the art will appreciate that the above embodiments are specific embodiments for implementing the present application, and that in actual applications, various changes may be made thereto in form and detail without departing from the spirit and scope of the present application.
Claims
1. A non-Gaussian noise-assisted UAV covert communication method, characterized in that: include: S1, establishes a UAV covert secure communication scenario model assisted by a UAV jammer emitting non-Gaussian noise; S2, in the established model, with the goal of maximizing the average secure covert throughput of multiple users, combined with the concealment constraints and power limit constraints, establishes an optimization problem; S3, converts the optimization problem into a multi-UAV collaborative Markov decision process; S4, for the converted multi-UAV cooperative Markov decision process, uses the combined attention mechanism and the standard multi-agent double-delay deep deterministic policy gradient to optimize the multi-UAV trajectory and obtain the optimal multi-UAV trajectory; S5, instructs the UAV to perform the communication task based on the optimal multi-UAV trajectory; For the converted multi-UAV cooperative Markov decision process, the multi-UAV trajectory is optimized using a combination of attention mechanism and standard multi-agent double-delay deep deterministic policy gradient to obtain the optimal multi-UAV trajectory, including: S41, initialize the network parameters of the Actor network, Critic network and attention network of the drone transmitter 、 、 、 、 ; S42, initialize the network parameters of the Actor network, Critic network and attention network of the drone jammer 、 、 、 、 ; S43, setting the network parameters of the target network to be the same as those of the main network, i.e. , , , , , , , , , ; S44: If the convergence condition is met, go to S48; otherwise, go to S44.
1. S44.1, UAV transmitter and UAV jammer observe their own status respectively and , and select actions respectively and , ; ; S44.2, Execution Action in the environment; S44.3, observe the next state , reward value and completion signal , to clarify Whether it is a terminal state; S44.4, Storage In the playback buffer middle; S44.5, if the UAV flies out of the boundary, does not meet the concealment constraint, or the number of time slots reaches the upper limit, reset the environment state; S44.6, when the preset number of updates is reached, , go to S45, otherwise, go to S47; S45, if , then go to,S45.1, otherwise go to S47; S45.1, from the playback buffer The random sampling size is Small batch samples of S45.2, the target action of the UAV transmitter is calculated as: ; S45.3, the drone jammer calculates the target action as: ; S45.4, Calculate target values for drone transmitters and the target value of the drone jammer for: ; ; S45.5, update the network parameters of the Critic network and the attention network of the drone transmitter and drone jammer by minimizing the loss as follows: ; ; S45.6, if If it is the time point for policy update, then go to S46; otherwise, go to S47; S46, using the gradient ascent method to update the Actor network of the drone transmitter and drone jammer is as follows: ; ; Gradually update the target network as follows: ; ; ; ; ; ; S47, ending the policy update cycle; S48, obtain the optimal multi-UAV trajectory.
2. The non-Gaussian noise-assisted UAV covert communication method according to claim 1 is characterized in that: A UAV covert secure communication scenario model assisted by a UAV jammer emitting non-Gaussian noise is established, including: Consider a UAV covert secure communication scenario assisted by a UAV jammer that emits non-Gaussian noise. The UAV transmitter acts as an aerial mobile base station in the presence of The distribution of the environment with two eavesdroppers and a single monitor Each user transmits a signal, and the drone jammer keeps sending non-Gaussian noise while moving to protect the information transmission between the user and the drone transmitter from being eavesdropped and detected; Drones to ensure distribution The service quality of each user will optimize its own trajectory. The trajectory of the drone is divided into The UAV keeps the same altitude during the flight, and the UAV transmitter and UAV jammer are set at the same altitude. and Department, and denote the UAV transmitter and UAV jammer respectively, while the user, monitor, and eavesdropper are regarded as ground nodes. It is assumed that each UAV knows the location of the monitor, eavesdropper, and user, but does not know the location of other UAVs. In the same time slot, at most one user receives the signal from the drone transmitter. The eavesdropper steals information when information is transmitted between the user and the drone transmitter, and the monitor is responsible for monitoring the information transmission between the user and the drone transmitter.
3. The non-Gaussian noise-assisted UAV covert communication method according to claim 2, characterized in that: To consider adaptability, the transmission link between the drone transmitter and the user is regarded as a probabilistic LoS / NLoS channel link, and the path loss uses the probabilistic LoS / NLoS path loss model, which is expressed by the formula: ; ; ; ; in, Indicates distance and elevation angle Total path loss under distance That is, the straight-line distance between the drone transmitter and the user, the elevation angle That is, the vertical angle between the drone transmitter and the ground, Indicates elevation angle The probability that the signal propagation path is a direct path is: and are environmental parameters adjusted based on the environment, Indicates elevation angle The probability that the signal propagation path is a non-direct path; For users , the received signal is expressed as: ; in, Represents a user In the The signal received on the channel Indicates the transmission power of the UAV transmitter, Indicates the UAV transmitter to the user The path loss between Indicates the distance between the drone transmitter and the user The straight-line distance between Indicates the vertical angle between the drone transmitter and the ground, Indicates the signal transmitted by the drone transmitter, represents the transmit power of the UAV jammer, Indicates drone jammer to user The path loss between Indicates that the drone jammer and the user The straight-line distance between represents the vertical angle between the UAV jammer and the ground, represents the non-Gaussian noise emitted by the drone jammer, Represents a user Received Gaussian white noise; For eavesdroppers , the received signal is expressed as: ; in, Indicates eavesdropper In the The signal received on the channel Indicates the drone transmitter to the eavesdropper The path loss between Indicates drone transmitter and eavesdropper The straight-line distance between From drone jammer to eavesdropper The path loss between Represents drone jammers and eavesdroppers The straight-line distance between Indicates eavesdropper Received Gaussian white noise; The monitor is based on The signal received by the monitor is used to determine whether the drone transmitter is transmitting information. The signal received by the monitor is expressed as: ; in, represents the null hypothesis that the drone transmitter does not transmit information, represents the alternative hypothesis of the UAV transmitter transmitting information, represents the path loss between the UAV jammer and the monitor, Indicates the straight-line distance between the UAV jammer and the monitor, represents the path loss between the UAV transmitter and the monitor, Indicates the straight-line distance between the drone transmitter and the monitor, Represents the Gaussian white noise received by the monitor.
4. The non-Gaussian noise-assisted UAV covert communication method according to claim 3, characterized in that: In the established model, with the goal of maximizing the average secure and concealed throughput of multiple users, combined with the concealment constraints and power limit constraints, an optimization problem is established, including: The decision-making of the monitor is divided into support of and support of There are two types, assuming that the false alarm probability and missed detection probability are respectively and The total detection error probability is Defined as , the goal of the drone is to make As close to 1 as possible, expressed by the formula: ; in, It is a preset positive decimal close to 0 to ensure the confidentiality of communication; exist and Under the assumption that The likelihood functions are expressed as: ; ; ; ; in, represents a complex function, represents the variance of the signal received by the monitor; Assuming that the monitor knows the index information of the signal, the likelihood ratio of the best detection is expressed as: ; in, represents the likelihood ratio of the best detection; The KL divergence is used to derive the lower limit of the minimum detection error probability. The lower limit of the minimum detection error probability is expressed as: ; in, represents the lower bound of the minimum detection error probability, It represents the constraint of detection probability. When the constraint of detection probability is satisfied, the concealment constraint will also be satisfied. It is expressed by the formula: ; ; in, represents the identity matrix; Based on the constraints of detection probability and stealth, the flight trajectory, transmission power and interference power of UAV transmitters and UAV jammers are optimized to maximize the average secure and covert throughput of multiple users. The UAV transmitter dynamically adjusts its position to provide high-quality services, and the UAV jammer adjusts its position to reduce the link quality of eavesdroppers and monitors, thus solving the optimization problem.
5. The non-Gaussian noise-assisted UAV covert communication method according to claim 4, characterized in that: The optimization problem established is expressed by the following formula: ; ; ; ; ; ; in, represents the average security and concealment throughput of multiple users, represents the implicit constraint, represents the power constraint of the UAV transmitter, Indicates the maximum transmission power of the drone transmitter, represents the power constraint of the UAV jammer, Indicates the maximum transmit power of the UAV jammer, represents the flight speed constraint of the UAV, and Represents the UAV at the sampling point and The location, Indicates the maximum flight speed of the drone. Indicates that the drone Fly to The time required, Represents the flight range constraint of the UAV, Indicates the scene space corresponding to the established model.
6. The non-Gaussian noise-assisted UAV covert communication method according to claim 5, characterized in that: The optimization problem is converted into a multi-UAV cooperative Markov decision process, including: The UAV transmitter and UAV jammer are regarded as intelligent agents, forming a partially observable state space, and the optimization problem is expressed as a five-tuple. To describe, represents the state space, represents the action space, represents the reward function, represents the state transition probability, represents the discount factor; State Space Defined as the current horizontal position of the drone transmitter and the current horizontal position of the drone jammer ; The action space is defined as the actions of the UAV transmitter and the actions of drone jammers , for the UAV transmitter, the action space contains Axis and The speed in the axis direction and the size of the transmission power, the action space of the drone jammer is contained in Axis and The speed in the axial direction and the magnitude of the interference power; In a collaborative drone environment, the decision-making process involves each agent receiving an immediate reward based on its current state, the action it performs, and the next state it transitions to, taking into account termination signals. The reward mechanism includes penalties for the drone's flight range, concealment, and distance. The reward and penalty terms are composed as follows: Secure and concealed throughput rewards, , motivating drones to improve safe and covert throughput; Boundary punishment, ,applied when the drone exceeds the predefined boundary to ensure that the drone operates within the specified area; Hidden punishment, ,applied when the minimum detection error probability of the monitor is lower than the threshold,to ensure the concealment of communication; Distance penalty, ; ; Among them, distance penalty and They are used to encourage drone transmitters and drone jammers to fly around the target, encourage drone transmitters to fly around users, and encourage drone jammers to fly around eavesdroppers and monitors respectively; The overall reward function of the drone transmitter and drone jammer is expressed as: ; ; in, and To weigh the coefficients of the impact of different rewards and penalties, and is the weight of the corresponding penalty item; In a UAV collaborative environment, the decisions and actions of each UAV will affect the overall system state. The state transition function describes the probability of the system reaching the next state given the current state and the UAV action, which is expressed by the formula: ; In this collaborative environment, there is no clear endpoint. The goal of the drone is to maximize the total communication capacity during the flight. The discount factor The choice will reflect this long-term performance metric.
7. A non-Gaussian noise-assisted UAV covert communication system, characterized in that: include: A scenario model building module is used to build a UAV covert secure communication scenario model assisted by a UAV jammer that emits non-Gaussian noise; An optimization problem establishment module is used to establish an optimization problem in the established model with the goal of maximizing the average secure and concealed throughput of multiple users, combined with concealment constraints and power limit constraints; Markov transformation module, used to transform the optimization problem into a multi-UAV collaborative Markov decision process; The multi-UAV trajectory optimization module is used to optimize the multi-UAV trajectory based on the converted multi-UAV cooperative Markov decision process, using the attention mechanism and the standard multi-agent double-delay deep deterministic policy gradient to obtain the optimal multi-UAV trajectory; A communication task execution module, used to instruct the UAV to perform the communication task based on the optimal multi-UAV trajectory; For the converted multi-UAV cooperative Markov decision process, the multi-UAV trajectory is optimized using a combination of attention mechanism and standard multi-agent double-delay deep deterministic policy gradient to obtain the optimal multi-UAV trajectory, including: S41, initialize the network parameters of the Actor network, Critic network and attention network of the drone transmitter 、 、 、 、 ; S42, initialize the network parameters of the Actor network, Critic network and attention network of the drone jammer 、 、 、 、 ; S43, setting the network parameters of the target network to be the same as those of the main network, i.e. , , , , , , , , , ; S44: If the convergence condition is met, go to S48; otherwise, go to S44.
1. S44.1, UAV transmitter and UAV jammer observe their own status respectively and , and select actions respectively and , ; ; S44.2, Execution Action in the environment; S44.3, observe the next state , reward value and completion signal , to clarify Whether it is a terminal state; S44.4, Storage In the playback buffer middle; S44.5, if the UAV flies out of the boundary, does not meet the concealment constraint, or the number of time slots reaches the upper limit, reset the environment state; S44.6, when the preset number of updates is reached, , go to S45, otherwise, go to S47; S45, if , then go to,S45.1, otherwise go to S47; S45.1, from the playback buffer The random sampling size is Small batch samples of S45.2, the target action of the UAV transmitter is calculated as: ; S45.3, the drone jammer calculates the target action as: ; S45.4, Calculate target values for drone transmitters and the target value of the drone jammer for: ; ; S45.5, update the network parameters of the Critic network and the attention network of the drone transmitter and drone jammer by minimizing the loss as follows: ; ; S45.6, if If it is the time point for policy update, then go to S46; otherwise, go to S47; S46, using the gradient ascent method to update the Actor network of the drone transmitter and drone jammer is as follows: ; ; Gradually update the target network as follows: ; ; ; ; ; ; S47, ending the policy update cycle; S48, obtain the optimal multi-UAV trajectory.
8. An electronic device, characterized in that: include: at least one processor; and, a memory communicatively coupled to the at least one processor; In which, the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the non-Gaussian noise-assisted drone covert communication method as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the non-Gaussian noise-assisted UAV covert communication method as described in any one of claims 1 to 6 can be implemented.
Citation Information
Patent Citations
Perception-assisted intelligent unmanned aerial vehicle covert communication and path planning method
CN118413811A