A large-capacity low-energy consumption unmanned aerial vehicle covert communication method
By constructing a system model and a motion energy consumption model, and combining the Multi-Objective Deep Deterministic Policy Gradient Algorithm (MODDPG) to plan the trajectory and transmission power of the UAV, the problem of establishing a high-capacity secure link between the UAV and the ground base station in a dynamic environment is solved, thus realizing the security and energy efficiency of UAV communication.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIHANG UNIV
- Filing Date
- 2024-04-30
- Publication Date
- 2026-05-08
AI Technical Summary
Existing path planning algorithms and covert communication methods are not suitable for establishing high-capacity secure links between UAVs and ground base stations in dynamic environments, and cannot effectively prevent eavesdropping. In particular, how to plan trajectories and signal transmission strategies to protect communication security during UAV flight missions is a critical issue.
The Multi-Objective Depth Deterministic Policy Gradient Algorithm (MODDPG) is used to plan the trajectory and transmission power of the UAV. By constructing a system model, a communication model, and a motion energy consumption model, the transmission power and flight trajectory of the UAV are optimized to maximize the probability of eavesdropper detection error. The MODDPG algorithm is used to solve the optimization problem, and the UAV path and signal transmission strategy are designed in combination with covert communication technology.
It enables secure communication between UAVs and ground base stations in dynamic environments, maximizes communication throughput and minimizes UAV energy consumption, and provides triple protection of security, communication performance and energy saving performance, ensuring that the communication link cannot be detected by eavesdroppers.
Smart Images

Figure CN118678325B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of UAV path planning and covert communication technology, specifically relating to a high-capacity, low-energy-consumption UAV covert communication method. Background Technology
[0002] To advance the development of air-space-ground networks, next-generation space networks assisted by unmanned aerial vehicles (UAVs) have been widely applied in various scenarios, such as acquiring ecosystem monitoring data, emergency rescue, natural resource exploration, factory monitoring, providing navigation assistance, and communication coverage. However, the fragile wireless communication environment (time-varying and open channels, strong environmental noise, and severe signal attenuation) poses a significant challenge to establishing stable UAV-assisted communication systems. Unstable wireless communication links provide opportunities for potential eavesdroppers. In such cases, malicious eavesdroppers can easily detect signal leakage from UAVs and obtain sensitive information, including control commands, mission objectives, and transmitted content.
[0003] Most current research focuses on anti-eavesdropping mechanisms in static scenarios, without considering the eavesdropping threats faced by drones during flight missions. How to plan drone flight trajectories and signal transmission strategies to protect drones from eavesdropping throughout their flight is a key issue.
[0004] Existing path planning algorithms include Algorithms such as Rapidly Exploring RandomTree (RRT), Dijkstra's algorithm, genetic algorithms, and ant colony optimization are commonly used path planning algorithms. These algorithms are typically only suitable for obstacle avoidance or finding the shortest path and are not applicable to tasks with more stringent or non-convex constraints, such as optimizing drone energy consumption or communication throughput.
[0005] Existing covert communication methods include 1) covert power control; 2) covert waveform design; 3) covert signal modulation; and 4) covert frequency / time hopping. Covert power control adaptively changes the transmission power to blend with noise on the eavesdropper's channel, preventing eavesdropping. Covert waveform design includes techniques such as Direct Sequence Spread Spectrum (DSSS), which reduces the power spectral density by expanding the bandwidth, thus preventing eavesdropping. Covert signal modulation hides the communication link by expanding the bandwidth, for example, Orthogonal Frequency Division Multiplexing (OFDM). The design principle of time / frequency hopping technology is to dynamically change the transmission frequency or time during communication, using a shared hopping mode between transceivers to prevent eavesdropping.
[0006] While the aforementioned technologies offer different methods for achieving covert communication, they are not suitable for dynamic environments, especially for establishing high-capacity secure links between mobile drones and ground base stations. Summary of the Invention
[0007] In view of this, the purpose of this invention is to provide a high-capacity, low-power covert communication method for unmanned aerial vehicles (UAVs), which aims to plan the trajectory of UAVs and complete tasks such as detection, reconnaissance and measurement of target areas.
[0008] To achieve the above objectives, the present invention provides the following technical solution:
[0009] A high-capacity, low-power covert communication method for unmanned aerial vehicles (UAVs) includes the following steps:
[0010] S1: Construct a system model for covert communication of unmanned aerial vehicles;
[0011] S2: Construct a covert communication model for unmanned aerial vehicles;
[0012] S3: Construct a motion energy consumption model for the drone;
[0013] S4: Based on the system model, communication model and motion energy consumption model, the mission objective of the UAV is to find the optimal transmission power and flight trajectory to maximize the detection error probability of the eavesdropper, and to construct the optimization problem and constraints accordingly.
[0014] S5: Solve the optimization problem using the multi-objective depth deterministic policy gradient algorithm MODDPG.
[0015] Furthermore, in the system model for the covert communication of the UAV, let the starting point of the UAV be... The location of the target point is The location of the drone is The location of the eavesdropper is Record the total time taken for the drone to travel from the starting point to the target point as: Total time Average score The duration is Extremely short time slots, in each time slot The internal-viewing drone moves at a constant linear velocity; in the first... In each time slot, the distance between the drone and the ground base station is recorded as follows: The distance between the drone and the eavesdropper is .
[0016] Furthermore, the UAV covert communication model specifically includes:
[0017] The channel model between the UAV and the base station is set as a LoS channel, and the UAV adopts a block fading channel for signal transmission. The channel gain of the UAV signal remains unchanged in the same block.
[0018] During its movement, the drone will choose whether to communicate with the ground base station at different time slots. This represents the drone transmitting signals to the ground base station, using This indicates that the drone is not communicating with the ground base station;
[0019] When a drone communicates with a ground base station, the drone maps the messages it sends to codewords. ,in The number of channels used; in the... The first time slot ground base station received the first The signal of each channel is:
[0020]
[0021] in , and They represent the first In the first time slot The transmit power gain of each channel, the transmitted signal, and the channel gain from the UAV to the ground base station. The noise signal at the receiving point. and The magnitudes of the numbers follow a Gaussian distribution. and ;
[0022] The eavesdropper uses an energy meter to monitor signal energy within the area and determines whether the drone is communicating with a ground base station based on the signal-to-noise ratio of the received signal; in the... The eavesdropper received the first time slot eavesdropping message. The signal of each channel is:
[0023]
[0024] in, The noise signal at the eavesdropper's location follows a Gaussian distribution. ;
[0025] The eavesdropper determines whether the drone is communicating with a ground base station by detecting whether the signal-to-noise ratio (SNR) of the received signal exceeds a set power threshold. If the SNR is greater than the threshold, the drone is considered to be sending a signal to the ground base station; if the SNR is less than the threshold, the drone is considered not sending a signal. The eavesdropper uses the maximum likelihood ratio test (LRT) method to minimize its detection error, expressed as:
[0026]
[0027] in It is in the Within a time slot, the eavesdropper receives the sum of signals from all channels. It is the detection threshold set by the eavesdropper;
[0028] Using relative entropy Construct probabilistic constraints to ensure that drone communications are not detected by eavesdroppers, where and They are respectively and Assuming the maximum likelihood function of the signal received by the eavesdropper is expressed as follows:
[0029] ;
[0030] .
[0031] Relative Entropy It indicates and Assuming the distance between the probability distributions of signals received by an eavesdropper, by reducing the relative entropy, the eavesdropper can be unable to distinguish whether the sender is sending a signal, thus achieving covert communication.
[0032] Furthermore, assuming the eavesdropper has drone transmission power... and noise power Given the prior information, the optimal threshold set by the eavesdropper... for:
[0033]
[0034] in, The signal-to-noise ratio of the signal at the eavesdropper's location is: .
[0035] Furthermore, the probability constraint for constructing drone communication without being detected by eavesdroppers using relative entropy is specifically as follows:
[0036]
[0037] in for:
[0038]
[0039] in The probability that drone communications will be detected by an eavesdropper.
[0040] Furthermore, the motion energy consumption model of the UAV is as follows:
[0041] In each time slot, it is assumed that the UAV is in a quasi-static equilibrium state and that its speed remains constant in each time slot;
[0042] Record the drone in the The velocity within each time slot is The propulsion energy consumption of a UAV is a linear sum of horizontal propulsion energy consumption, vertical propulsion energy consumption, and profile energy consumption related to fluid resistance, where:
[0043] No. Horizontal propulsion energy consumption within each time slot Represented as:
[0044]
[0045] in For the weight of the drone, and This indicates the mass and gravitational acceleration of the drone; Let the cross-sectional area of the drone be the area along its direction of motion. air mass density, The length of each time slot;
[0046] No. Vertical propulsion energy consumption within each time slot Represented as:
[0047]
[0048] The profile energy consumption related to fluid resistance is:
[0049]
[0050] in To understand the relationship between the fluid resistance of a drone and its physical structure. The speed of the drone relative to the air. It's wind speed. Indicates the drag coefficient;
[0051] Drones in Total energy consumption within each time slot for:
[0052]
[0053] in, , and .
[0054] Furthermore, the optimization problem and constraints are as follows:
[0055]
[0056] and For the two optimization objectives in UAV path planning, let be maximizing UAV communication throughput and minimizing UAV motion energy consumption, respectively. , For decoding error probability; constraints The decoding error probability constraint at the receiver; constraint The transmit power of each channel of the drone was limited; constraints were imposed. The constraints imposed limitations on the stealth requirements between drones and ground base stations; and The maximum displacement and velocity variation of the UAV within each time slot were respectively limited, where This indicates the maximum achievable acceleration.
[0057] Furthermore, in step S5, the path planning and transmission power control of the UAV are modeled as a finite Markov Decision Process (MDP) problem. The UAV relies on its interaction with the environment to adjust its actions and learn the optimal policy. Its state space, action space, and reward function are as follows:
[0058] State space: ;
[0059] Action space: ;
[0060] Reward function: ,in, and The maximization of effective throughput and the minimization of UAV energy consumption are respectively expressed as:
[0061]
[0062] and Two auxiliary reward functions:
[0063]
[0064]
[0065] Indicates the communication security performance of drones. This reflects the penalty for longer path lengths;
[0066] Different weights are assigned to the four reward functions, denoted as . The complete reward is recorded as follows: .
[0067] Furthermore, the MODDPG algorithm comprises two network structures: an actor network and a critic network, each consisting of an online network and a target network; the online actor network specifies the main strategy. Mapping observed states to actions, online commentator network estimates ,in and These are parameters for two online networks;
[0068] Two target networks employing an actor-critic architecture are used. The target value is calculated by freezing the parameters of the target network before updating. and During the initialization phase, samples are copied from the online actor-critic network; when updating network parameters, a small batch of samples is randomly drawn from the experience replay pool.
[0069] The elements of the reward vector are transformed into a scalar weighted sum using a linear weighting method. Taking into account the preferences among multiple objectives and constraints, the weights are denoted as... Add a time-varying decaying noise to the actor strategy. Based on the transformation from the experience replay pool, the policy objective function is:
[0070]
[0071] The steps to optimize the online critic network are as follows: First, calculate the target value given by the online critic network and... The difference between the values is then calculated, and gradient descent is used to minimize the loss function, which is defined as the mean squared error (MSE) of the difference.
[0072]
[0073] The optimization objective of the online critic network is to minimize the MSE;
[0074] Using the online reviewer network The value is then used to calculate the strategy of the online actor network. gradient:
[0075]
[0076] The optimization objective of the online actor network is to maximize the gradient.
[0077] Furthermore, the MODDPG algorithm steps are as follows:
[0078] S51: Input weight parameter vector ;
[0079] S52: Randomly initialize online actor network parameters and online critic network parameters Initialize the target actor network parameters. and target commentator network parameters : , Initialize the experience replay pool Mini-batch size Discount factor Explore noise Learning factors of target actor network and critic network and ;
[0080] S53: Obtain initial observation status ;
[0081] S54: At each step, based on the current state and noise... Select and perform actions ;
[0082] S55: Perform the action And observe rewards Next moment state ;
[0083] S56: Will Store in the experience replay pool ;
[0084] S57: From Randomly selected Small batches of data;
[0085] S58: For each data point, calculate the objective function value. :
[0086]
[0087] S59: By minimizing the loss function Update the parameters of the online commentator network; by maximizing the policy gradient. Update online actor network parameters;
[0088] S510: Update target network parameters:
[0089]
[0090] S511: Exploring Noise Attenuation: ;
[0091] S512: Increment the step size by 1, return to step S54, and continue until the maximum step size is reached;
[0092] S513: Return to step S53 and retrain until the maximum number of training iterations is reached.
[0093] The beneficial effects of this invention are as follows: This invention designs a UAV path planning strategy based on multi-objective optimization that takes into account communication security. Under the premise that the communication link between the sender (UAV) and the receiver (ground base station) is not detected by eavesdroppers, this strategy achieves Pareto optimality of link throughput and UAV motion energy consumption, providing triple protection of security performance, communication performance and energy-saving performance for UAV to carry out missions.
[0094] Other advantages, objectives, and features of the invention will be set forth in the following description and will be apparent to those skilled in the art in some respects, or may be learned by practice of the invention. The objectives and other advantages of the invention can be realized and obtained through the following description. Attached Figure Description
[0095] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the following figures are provided for illustration:
[0096] Figure 1 This is the UAV covert communication trajectory planning model described in this invention. Detailed Implementation
[0097] This invention discloses a high-capacity, low-energy-consumption covert communication method for unmanned aerial vehicles (UAVs). First, a system model is constructed. In this covert communication model based on UAVs, the UAV has a defined starting point and target point. While collecting information, it maintains communication with a ground base station. Under the premise that the communication link cannot be detected by eavesdroppers, the path of the UAV from the starting point to the target point is planned to achieve Pareto optimality in communication throughput and UAV energy consumption. Figure 1 As shown, the positions of the starting point, the target point, the drone, and the eavesdropper are respectively determined by... , , and Let be the total time taken for the drone to travel from the starting point to the target point. Total time Average score The duration is Extremely short time slots, in each time slot The internal-viewing drone is moving at a constant velocity in a straight line. In the... During each time slot, the distances between the drone and the ground base station, and between the drone and the eavesdropper, are recorded as follows: and .
[0098] Then, a communication model is constructed. Although the air-to-ground communication link is a superposition of probabilities from line-of-sight (LoS) and non-line-of-sight (NLoS) channels, the probability of the LoS channel approaches 1 when the UAV's altitude is greater than 100 meters. Therefore, in this invention, the channel model between the UAV and the base station is set as a LoS channel, and the UAV is assumed to use a block fading channel for signal transmission. Since the duration of a block is shorter than the channel change rate, the channel gain of the UAV signal remains constant within the same block. During its movement, the UAV will choose whether to communicate with the ground base station in different time slots. This represents the drone transmitting signals to the ground base station, using This indicates that the drone is not communicating with the ground base station. When the drone communicates with the ground base station, it maps the messages it sends to codewords. ( (The number of channels used), while the eavesdropper uses an energy meter to monitor the signal energy in the area and determines whether the drone is communicating with the ground base station based on the signal-to-noise ratio of the received signal.
[0099] Therefore, in the first The first time slot ground base station received the first The signal of each channel is:
[0100] (1)
[0101] in , and They represent the first In the first time slot The transmit power gain of each channel, the transmitted signal, and the channel gain from the UAV to the ground base station. The noise signal at the receiving point. and The magnitudes of the numbers follow a Gaussian distribution. and .
[0102] Similarly, in the first The eavesdropper received the first time slot eavesdropping message. The signal of each channel is:
[0103] (2)
[0104] in, The noise signal at the eavesdropper's location follows a Gaussian distribution. .
[0105] The method eavesdroppers use to determine whether a drone is communicating with a ground base station is by detecting whether the signal-to-noise ratio (SNR) of the received signal exceeds a set power threshold. If the SNR is greater than the threshold, the drone is considered to be sending a signal to the ground base station; if the SNR is less than the threshold, the drone is considered not to be sending a signal. Based on the Neyman-Pearson criterion, eavesdroppers use the maximum likelihood ratio test (LRT) to minimize their detection error, which can be expressed by the following formula:
[0106] (3)
[0107] in, Furthermore, the above formula can be simplified to:
[0108] (4)
[0109] in It is in the Within a time slot, the eavesdropper receives the sum of signals from all channels. It is the detection threshold set by the eavesdropper, and the optimal threshold can be obtained by the following Theorem 1.
[0110] Theorem 1: Assume the eavesdropper has the transmission power of a drone. and noise power Given the prior information, the optimal threshold that the eavesdropper should set is... for:
[0111] (5)
[0112] in, The signal-to-noise ratio of the signal at the eavesdropper's location is: .
[0113] The proof of Theorem 1 is as follows:
[0114] because ,…, Since they are independent and identically distributed variables, we can obtain:
[0115] (6)
[0116] and
[0117] (7)
[0118] Substituting equations (6) and (7) into equation (6), we get:
[0119] (8)
[0120] Taking the logarithm of both sides of the above equation, we get:
[0121] (9)
[0122] Simplifying the above equation, we get:
[0123] (10)
[0124] Assumption and The prior probabilities are the same, that is Combining the above equation and equation (4), we can obtain:
[0125] (11)
[0126] The proof is complete.
[0127] After determining the optimal threshold, the false alarm rate of the eavesdropper is... and false negative rate They can be represented as:
[0128] (12)
[0129] and
[0130] (13)
[0131] Therefore, the eavesdropper's false detection rate is the sum of the two, expressed as:
[0132] (14)
[0133] remember Let be the probability that the drone's communication is detected by an eavesdropper. ,but and Since it is difficult to calculate, we use relative entropy to describe it here, as detailed in Theorem 2.
[0134] Theorem 2: Compared to It is a more stringent constraint. It can be represented as:
[0135] (15)
[0136] The proof of Theorem 2 is as follows:
[0137] First, prove... Compared to It is a more stringent constraint. Definition and The total variational distance between them is in the th Within the time slot , can be represented as:
[0138] (16)
[0139] Because the total variational distance satisfies Then there is According to Pinsker's inequality, the total variational distance and relative entropy satisfy the following equation:
[0140] (17)
[0141] therefore, Compared to It is a more stringent constraint.
[0142] Next, the derivation process of equation (15) is given. First, the derivation process of equation (15) is given. The expression is:
[0143] (18)
[0144] because ,…, Since the variables are independent and identically distributed, the above formula can be transformed into:
[0145] (19)
[0146] in, , and The proof is complete.
[0147] The motion energy consumption model is as follows: Due to the limited payload of the UAV, optimizing its energy consumption is crucial. Since the communication energy consumption of a UAV is typically two orders of magnitude smaller than its propulsion energy consumption, this invention primarily considers the propulsion energy consumption of the UAV. In each time slot, it is assumed that the UAV is in a quasi-static equilibrium state, meaning that the UAV flies smoothly with a small acceleration, and its speed remains constant in each time slot.
[0148] Record the drone in the The velocity within each time slot is Generally speaking, the propulsion energy consumption of a UAV can be viewed as a linear sum of horizontal propulsion energy consumption, vertical propulsion energy consumption, and profile energy consumption related to fluid resistance. Horizontal propulsion energy consumption within each time slot It can be represented as:
[0149] (20)
[0150] in For the weight of the drone, and This indicates the mass and gravitational acceleration of the drone. Let be the cross-sectional area of the drone in the direction of its motion. air mass density, The length of each time slot.
[0151] No. Vertical propulsion energy consumption within each time slot It can be represented as:
[0152] (twenty one)
[0153] Next, we calculate the energy consumption related to fluid resistance. First, we define the speed of the drone relative to the air as... ,in It's wind speed. Based on fluid dynamics methods, the relationship between the fluid resistance of a drone and its physical structure can be simulated as follows:
[0154] (twenty two)
[0155] in, Let represent the drag coefficient. According to equations (21) and (22), the energy consumption to overcome fluid resistance can be obtained as:
[0156] (twenty three)
[0157] According to equations (20)-(23), we can obtain the drone in the first... Total energy consumption within each time slot for:
[0158] (twenty four)
[0159] in, , and .
[0160] After establishing the above models, from the perspective of the drone, its mission objective is to find the optimal transmission power and flight trajectory to maximize the probability of detection error by the eavesdropper, which is equivalent to minimizing... The optimization problem and constraints are as follows:
[0161] (25)
[0162] and For the two optimization objectives in UAV path planning, let be maximizing UAV communication throughput and minimizing UAV motion energy consumption, respectively. , For decoding error probability; constraints The decoding error probability constraint at the receiver; constraint The transmit power of each channel of the drone was limited; constraints were imposed. The constraints imposed limitations on the stealth requirements between drones and ground base stations; and The maximum displacement and velocity variation of the UAV within each time slot were respectively limited, where This indicates the maximum achievable acceleration.
[0163] As an intelligent agent, a drone can stably learn trajectory and transmission power control strategies under constraints, and periodically determine its flight speed, direction, and transmission power for the next time slot. Furthermore, the drone's actions are determined only by its last state. However, traditional reinforcement learning algorithms may not be directly applicable to the model under consideration. To address the aforementioned multi-objective optimization problem, this invention proposes a Multi-objective Deep Deterministic Policy Gradient (MODDPG) algorithm. The drone's path planning and transmission power control can be modeled as a finite Markov Decision Process (MDP) problem, where the drone relies on its interaction with the environment to adjust its actions and learn the optimal policy. The state space, action space, and reward in the model will be described below.
[0164] ① State Space: Based on the master information exchange between the UAV and the ground base station, the UAV can periodically observe its position, wind speed, and noise power. Therefore, the state space is defined as:
[0165]
[0166] ② Action Space: After observing its state, the drone will react, and its action space is defined as:
[0167]
[0168] ③ Reward Functions: Based on the optimization problem (25), the reward designed in this invention includes four reward functions, expressed as follows: .in, and The maximization of effective throughput and the minimization of UAV energy consumption can be expressed as follows:
[0169] (26)
[0170] and The physical meaning is that paths with higher throughput will be rewarded more, while paths with higher energy consumption will be penalized. Furthermore, this invention also designs two auxiliary reward functions. and :
[0171] (27)
[0172] (28)
[0173] in, Indicates the communication security performance of drones. This reflects the penalty for longer path lengths.
[0174] Furthermore, based on the importance of these four reward functions, this invention designs different weights, denoted as... The complete reward can then be recorded as: .
[0175] MODDPG is an off-policy reinforcement learning algorithm that consists of two main network structures: an actor network and a critic network. Each network comprises an online network and a target network. The online actor network specifies the main policy. Mapping observed states to actions, while online commentator network estimates ,in and These are parameters for two online networks.
[0176] Furthermore, this invention designs two target networks using an actor-critic architecture. Calculating the target value by freezing the parameters of the target networks before updating improves algorithm convergence. Specifically, the parameters of the target networks... and During the initialization phase, samples are copied from an online actor-critic network. When updating network parameters, a mini-batch of samples is randomly drawn from the experience replay pool. This invention uses a linear weighting method to convert the elements of the reward vector into a scalar weighted sum. Considering the preferences among multiple objectives and constraints, the weights are denoted as... Meanwhile, this invention incorporates a time-varying decaying noise into the actor strategy. This addresses the problem of limited action space exploration when relying solely on learning algorithms. Based on transformations from the experience replay pool, the policy objective function is:
[0177] (29)
[0178] To optimize the online commentator network, this invention first calculates the target value given by the online commentator network and... The difference between the values is then calculated, and gradient descent is used to minimize the loss function, which is defined as the mean squared error (MSE) of the difference.
[0179] (30)
[0180] To optimize the strategy of online actor network This invention utilizes feedback from an online network of reviewers. Then calculate its gradient:
[0181] (31)
[0182] The optimization objectives for the online critic network and the online actor network are to minimize the MSE (Equation 30) and maximize the gradient (Equation 31), respectively. Unlike traditional deep reinforcement learning algorithms that use hard updates of target network parameters, the target network in MODDPG used in this invention updates parameters very slowly at each step. This parameter update method can greatly improve the stability of the algorithm. The specific algorithm flow is shown in Algorithm 1.
[0183]
[0184] In summary, this invention presents a multi-objective optimization-based UAV path planning strategy that considers communication security. This strategy, while ensuring that the communication link between the sender (UAV) and receiver (ground base station) is not detected by eavesdroppers, achieves Pareto optimality in link throughput and UAV energy consumption by designing the UAV's trajectory and signal transmission power. This provides triple protection for UAV missions, ensuring security, communication performance, and energy efficiency.
[0185] Finally, it should be noted that the above preferred embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail through the above preferred embodiments, those skilled in the art should understand that various changes can be made to it in form and detail without departing from the scope defined by the claims of the present invention.
Claims
1. A high-capacity, low-power covert communication method for unmanned aerial vehicles (UAVs), characterized in that: Includes the following steps: S1: Construct a system model for covert communication of unmanned aerial vehicles; S2: Construct a covert communication model for unmanned aerial vehicles; S3: Construct a motion energy consumption model for the drone; S4: Based on the system model, communication model and motion energy consumption model, the mission objective of the UAV is to find the optimal transmission power and flight trajectory to maximize the detection error probability of the eavesdropper, and to construct the optimization problem and constraints accordingly. S5: Solve the optimization problem using the Multi-Objective Deterministic Strategy Gradient Algorithm (MODDPG); In the system model of the UAV covert communication, let the starting point of the UAV be... The location of the target point is The location of the drone is The location of the eavesdropper is Record the total time taken for the drone to travel from the starting point to the target point as: Total time Average score The duration is Extremely short time slots, in each time slot The internal-viewing drone moves at a constant linear velocity; in the first... In each time slot, the distance between the drone and the ground base station is recorded as follows: The distance between the drone and the eavesdropper is ; The UAV covert communication model specifically includes: The channel model between the UAV and the base station is set as a LoS channel, and the UAV adopts a block fading channel for signal transmission. The channel gain of the UAV signal remains unchanged in the same block. During its movement, the drone will choose whether to communicate with the ground base station at different time slots. This represents the drone transmitting signals to the ground base station, using This indicates that the drone is not communicating with the ground base station; When a drone communicates with a ground base station, the drone maps the messages it sends to codewords. ,in The number of channels used; in the... The first time slot ground base station received the first The signal of each channel is: in , and They represent the first In the first time slot The transmit power gain of each channel, the transmitted signal, and the channel gain from the UAV to the ground base station. The noise signal at the receiving point. and The magnitudes of the numbers follow a Gaussian distribution. and ; The eavesdropper uses an energy meter to monitor signal energy within the area and determines whether the drone is communicating with a ground base station based on the signal-to-noise ratio of the received signal; in the... The eavesdropper received the first time slot eavesdropping message. The signal of each channel is: in, The noise signal at the eavesdropper's location follows a Gaussian distribution. ; The eavesdropper determines whether the drone is communicating with a ground base station by detecting whether the signal-to-noise ratio (SNR) of the received signal exceeds a set power threshold. If the SNR is greater than the threshold, the drone is considered to be sending a signal to the ground base station; if the SNR is less than the threshold, the drone is considered not sending a signal. The eavesdropper uses the maximum likelihood ratio detection method to minimize its detection error, expressed as: in It is in the Within a time slot, the eavesdropper receives the sum of signals from all channels. It is the detection threshold set by the eavesdropper; Using relative entropy Construct probabilistic constraints to ensure that drone communications are not detected by eavesdroppers, where and They are respectively and Assuming the maximum likelihood function of the signal received by the eavesdropper is expressed as follows: ; . Relative Entropy It indicates and Assuming the distance between the probability distributions of signals received by the eavesdropper, by reducing the relative entropy, the eavesdropper cannot distinguish whether the sender is sending a signal, thus achieving covert communication; Assuming the eavesdropper has drone transmission power and noise power Given the prior information, the optimal threshold set by the eavesdropper... for: in, The signal-to-noise ratio of the signal at the eavesdropper's location is: ; The probability constraint for constructing drone communication without being detected by eavesdroppers using relative entropy is as follows: in for: in The probability that drone communications will be detected by an eavesdropper; The motion energy consumption model of the UAV is as follows: In each time slot, it is assumed that the UAV is in a quasi-static equilibrium state and that its speed remains constant in each time slot; Record the drone in the The velocity within each time slot is The propulsion energy consumption of a UAV is a linear sum of horizontal propulsion energy consumption, vertical propulsion energy consumption, and profile energy consumption related to fluid resistance, where: No. Horizontal propulsion energy consumption within each time slot Represented as: in For the weight of the drone, and This indicates the mass and gravitational acceleration of the drone; Let be the cross-sectional area of the drone in the direction of its motion. air mass density, The length of each time slot; No. Vertical propulsion energy consumption within each time slot Represented as: The profile energy consumption related to fluid resistance is: in To understand the relationship between the fluid resistance of a drone and its physical structure. The speed of the drone relative to the air. It's wind speed. Indicates the drag coefficient; Drones in Total energy consumption within each time slot for: in, , and ; The optimization problem and constraints are as follows: and For the two optimization objectives in UAV path planning, let be maximizing UAV communication throughput and minimizing UAV motion energy consumption, respectively. , For decoding error probability; constraints The decoding error probability constraint at the receiver; constraint The transmit power of each channel of the drone was limited; constraints were imposed. The constraints imposed limitations on the stealth requirements between drones and ground base stations; and The maximum displacement and velocity variation of the UAV within each time slot were respectively limited, where This indicates the maximum achievable acceleration; In step S5, the path planning and transmission power control of the UAV are modeled as a finite Markov decision process (MDP) problem. The UAV relies on its interaction with the environment to adjust its actions and learn the optimal policy. Its state space, action space, and reward function are as follows: State space: ; Action space: ; Reward function: ,in, and The maximization of effective throughput and the minimization of UAV energy consumption are respectively expressed as: and Two auxiliary reward functions: Indicates the communication security performance of drones. This reflects the penalty for longer path lengths; Different weights are assigned to the four reward functions, denoted as . The complete reward is recorded as follows: ; The MODDPG algorithm comprises two network structures: an actor network and a critic network, each consisting of an online network and a target network. The online actor network specifies a master strategy. Mapping observed states to actions, online commentator network estimates ,in and These are parameters for two online networks; Two target networks employing an actor-critic architecture are used. The target value is calculated by freezing the parameters of the target network before updating. and During the initialization phase, samples are copied from the online actor-critic network; when updating network parameters, a small batch of samples is randomly drawn from the experience replay pool. The elements of the reward vector are transformed into a scalar weighted sum using a linear weighting method. Taking into account the preferences among multiple objectives and constraints, the weights are denoted as... Add a time-varying decaying noise to the actor strategy. Based on the transformation from the experience replay pool, the policy objective function is: The steps to optimize the online critic network are as follows: First, calculate the target value given by the online critic network and... The difference between the values is then calculated, and the loss function is minimized using gradient descent. The loss function is defined as the mean squared error (MSE) of the difference. The optimization objective of the online critic network is to minimize the MSE; Using the online reviewer network The value is then used to calculate the strategy of the online actor network. gradient: The optimization objective of the online actor network is to maximize the gradient; The steps of the MODDPG algorithm are as follows: S51: Input weight parameter vector ; S52: Randomly initialize online actor network parameters and online critic network parameters Initialize the target actor network parameters. and target commentator network parameters : , Initialize the experience replay pool Mini-batch size Discount factor Explore noise Learning factors of target actor network and critic network and ; S53: Obtain initial observation status ; S54: At each step, based on the current state and noise... Select and perform actions ; S55: Perform the action And observe rewards Next moment state ; S56: Will Store in the experience replay pool ; S57: From Randomly selected from Small batches of data; S58: For each data point, calculate the objective function value. : S59: By minimizing the loss function Update the parameters of the online commentator network; by maximizing the policy gradient. Update online actor network parameters; S510: Update target network parameters: S511: Exploring Noise Attenuation: ; S512: Increment the step size by 1, return to step S54, and continue until the maximum step size is reached; S513: Return to step S53 and retrain until the maximum number of training iterations is reached.
Citation Information
Patent Citations
Covert communication method based on unmanned aerial vehicle and intelligent reflecting surface
CN115442824A
Covert communication method in unmanned aerial vehicle assisted mobile edge computing system
CN117915409A