An unmanned aerial vehicle multi-hop relay stealth communication method based on trajectory and power joint design
By jointly optimizing the trajectory and transmission power of UAVs in a multi-hop relay network, the problem of insufficient transmission distance and coverage of single-hop networks in large-scale communication scenarios is solved, achieving efficient multi-hop covert communication and improving the system's throughput and robustness.
Patent Information
- Application Number
- CN202410991253.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-07-23
AI Technical Summary
Single-hop UAV communication networks suffer from limited transmission distance, insufficient coverage, and low network capacity in large-scale communication scenarios. Furthermore, they lack the complex routing and relay mechanisms found in multi-hop networks, making it difficult to meet the transmission needs of future wireless communication systems.
A multi-hop relay covert communication method based on trajectory and power joint design is adopted. By setting up a UAV multi-hop relay system, the transmission rate and covert constraints of each relay link are determined. The C-MAPPO algorithm is used to optimize the UAV's flight trajectory and transmission power, and a discrete-time Markov decision process is constructed to solve the optimization problem.
This improved the throughput of the UAV relay covert communication system, enhanced the system's robustness and flexibility, reduced the risk of collisions, ensured the successful completion of missions, and improved the utilization efficiency of communication resources.
Smart Images

Figure CN118714591B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of wireless communication technology, and in particular to a UAV multi-hop relay covert communication method based on trajectory and power joint design. Background Technology
[0002] Compared to traditional terrestrial communication networks, unmanned aerial vehicles (UAVs) exhibit significant advantages, particularly in terms of mobility, ease of deployment, and line-of-sight (LOS) links. These characteristics endow UAVs with unique capabilities, enabling them to directly act as communication base stations, providing users with immediate communication services. In addition to serving as direct base stations, UAVs can leverage their mobility to act as aerial relay stations, extending the coverage of communication networks. By deploying UAVs at key nodes, dispersed user groups can be effectively connected, improving the connectivity and efficiency of the entire network.
[0003] Leveraging the open nature of Loss links and wireless channels, UAV communication networks offer unparalleled flexibility and efficiency. However, these characteristics inevitably introduce a series of security challenges, especially in secure communication. The openness of wireless channels makes communication content susceptible to external interference and illegal interception, while the high mobility of UAVs, although providing advantages for rapid deployment and immediate response, can also make them targets for attackers, thus threatening communication security. Against this backdrop, the concept of covert communication has emerged, becoming an important research direction in the field of UAV communication network security. Covert communication aims to ensure the secrecy of information transmission, guaranteeing that communication content remains undetected or uninterpreted even in potentially hostile environments. In the design of covert communication transmission schemes, interference or artificial noise signals can be actively created as communication cover, while passively utilizing uncertainties such as noise and channel conditions to increase system uncertainty can help improve the performance of covert communication systems. The use of friendly jammers or artificial noise can improve the covert performance of communication systems, but it can also negatively impact the communication of friendly parties, thus requiring precise control and coordination. Using relay equipment can further improve communication efficiency by optimizing the drone's flight path and communication resource allocation. Furthermore, because multiple communication links exist within the network, these links effectively form a concealed communication network that provides mutual cover. By achieving covert communication, drones can perform critical missions undetected, providing a safer and more reliable communication environment for both parties.
[0004] In single-hop networks, drones, acting as aerial base stations, can leverage their air superiority and high maneuverability to achieve covert communication. In air-to-ground networks, optimizing the time slot allocation, power allocation, and trajectory of drones serving users can maximize the average covertness rate. Alternatively, deep reinforcement learning methods can be used to jointly optimize the drone's transmission power and trajectory to achieve covert communication. However, while the performance analysis and optimization of single-hop covert communication networks is relatively simple, they have limitations in transmission distance and coverage. As wireless network services evolve towards high-speed and large-scale access, single-hop networks are no longer adequate for the transmission demands of future wireless communication systems. For example, in ultra-dense networking technologies, the limited coverage of low-power base stations in single-hop networks leads to frequent user switching, reducing network capacity and user experience. In practical applications, single-hop covert communication networks may face several technical challenges, such as how to efficiently process massive amounts of network data without disrupting the semantics of communication protocols, how to provide sufficient communication capacity while maintaining covertness, and the lack of complex routing and relay mechanisms found in multi-hop networks, which can enhance the covertness and security of the communication network.
[0005] Compared to single-hop networks, relay networks enhance the communication link between source and destination nodes by introducing relay nodes, thereby ensuring transmission reliability. Due to the dynamic changes brought about by UAV mobility, the wireless connection of each UAV within a limited communication range is unstable. Furthermore, considering the energy limitations and stealth constraints of UAVs, it is necessary to jointly plan the UAV trajectory and transmission power. This approach can effectively improve the throughput of UAV relay covert communication systems. In addition, UAVs equipped with an Intelligent Reflecting Surface (IRS) can also act as relays to achieve covert communication between users. Alternating optimization of Alice's transmission power, IRS phase shift, and the UAV's horizontal position maximizes the covert transmission rate. Under the detection of an aerial observer, UAVs can also effectively achieve covert communication as relays. However, considering large-scale communication scenarios under covert constraints, single relay networks still have limitations. Therefore, using multiple UAVs to cooperate to achieve multi-hop relay communication has become a current research direction. Summary of the Invention
[0006] This application provides a UAV multi-hop relay covert communication method based on trajectory and power joint design, which can be used to solve the technical problems of single relay networks in large-scale communication scenarios under covert constraints.
[0007] This application provides a UAV multi-hop relay covert communication method based on joint trajectory and power design, the method including:
[0008] Step 1: Set up the UAV multi-hop relay system and define the transmission rate of each hop relay link;
[0009] Step 2: Set the concealment constraints for each hop of the communication link from the perspective of monitoring the drone;
[0010] Step 3: Set the concealment constraints for each hop of the communication link from the perspective of the communication drone;
[0011] Step 4: Define the initial optimization problem for covert communication of multi-hop relay UAVs based on the joint determination of trajectory and power;
[0012] Step 5: Give the optimized hidden constraints and rewrite the initial optimization problem;
[0013] Step 6: Construct the communication and optimization problem as a discrete-time Markov decision process and solve the optimization problem using the C-MAPPO method.
[0014] Furthermore, a multi-hop relay system for UAVs is established, and the transmission rate of each relay link is defined, including:
[0015] A multi-hop drone relay system is configured: the source drone S communicates with the target drone D with the assistance of multiple relay drones; S and D are assumed to be in fixed locations due to performing special tasks; S and D cannot communicate directly due to the distance between them, and a multi-hop drone relay network is used to avoid detection by monitoring drones. The multi-hop drone relay network includes M drones, and the index set of the multi-hop drone relay network is represented as follows. The first drone is designated S, and the Mth drone is designated D.
[0016] Assuming the drone's frequency bands are orthogonal and the system operates in full-duplex time-slot mode, i.e. And each time slot δ t If the duration is short enough, then the drone's coordinates are represented as q. m [t]=(x[t],y[t],z[t]), where S, D, the relay drone, and the surveillance drone W are each equipped with a single omnidirectional antenna. The communication drone estimates its relative distance to W by using radar or a camera. The initial position of the surveillance drone W is defined as q. w =(x w ,y w ,z w ), with a speed of v w (t)=(v wx ,v wy ,v wz If W = q, then the trajectory of W is represented as q. w [t] = q w +δ t·v w (t), where
[0017] Assume there are M drones in the network, let... This represents the set of indices for each hop in the linear relay path connecting the source UAV and the destination UAV; where ι1 is the first hop on the path, ι2 is the second hop, and ι3 is the third hop. M-1 This is the last hop before reaching the destination node; let d be the maximum communication distance between drones in each hop. c The minimum safe distance is d s ;
[0018] In UAV communication, the channel gain between any pair of communicating UAVs m and target UAVs m′ follows a free-space path loss model, where...
[0019]
[0020] Where ρ0 represents the path loss at the reference distance, which is defined as d0 = 1 meter;
[0021] p m (t) represents the transmission power of UAV m at time t; then the power received by UAV m from UAV m′ at UAV m is expressed as:
[0022]
[0023] Then the SNR received at point m+1 from drone m is:
[0024]
[0025] in, in Let B be the power spectral density of the noise, and B be the bandwidth.
[0026] The transmission rate between adjacent drones is expressed as:
[0027] R m,m+1 (t)=Blog2[1+α m,m+1 (t)],
[0028] That is, the ιth m The hop transmission rate is expressed as
[0029] Furthermore, from the perspective of monitoring drones, the covert constraints of each hop communication link are set, including:
[0030] Under the DF relay protocol, let the ιth... m The signal y received by the adjacent drone m+1 during the jump is from m.m,m+1 And the received signal of W is y m,w The corresponding channel gains are ||h m,m+1 (t)|| and ||h m,w (t)||, then:
[0031]
[0032] in For the ι m The symbol for jump transmission, n m+1 and Let m+1 and W represent the noise received by UAVs respectively. To detect the presence of signal transmission, the signal received by UAV W is monitored in the following two ways:
[0033]
[0034] L is the codeword length in the channel-related block, H1 represents the case where there is signal transmission between friendly nodes, and H0 represents the case where there is no signal transmission between friendly nodes. (Assume...) It is a unit amplitude phase shift keying signal, with phase denoted as ψ. The received signal distribution at W is as follows:
[0035]
[0036] P MD =P{H0|H1 is true} Missed detection, P FA =P{H1|H0 is true} False alarm;
[0037] Assuming the drone makes decisions based on all signals it receives during multi-hop transmission, to ensure stealth, the probability of detection error at each hop must satisfy P. MD +P FA ≥1-∈, ∈ reflects the strictness of the constraint on the bit error rate of the detection error;
[0038] The hidden constraints for each relay link are set as follows:
[0039] P MD +P FA ≥1-∈.
[0040] Furthermore, from the perspective of the communication drone, the concealment constraints of each hop of the communication link are set, including:
[0041] The lower bound of the probability of detection error is:
[0042]
[0043] If H1 or H0 is true, Q1(t) or Q0(t) represents the joint probability distribution of all W received signals at time t, and D(Q1(t)||Q0(t)) is the relative entropy between Q1 and Q0, i.e.:
[0044]
[0045] We obtain D(Q1(t)||Q0(t))≤2∈ 2 This is a constraint that is more in line with expectations;
[0046] Based on the definition of relative entropy and the independence of signals received on different hops, we can derive...
[0047] The hidden constraints of each relay link are rewritten as follows:
[0048]
[0049] Furthermore, an initial optimization problem for covert communication of multi-hop relay UAVs based on the joint determination of trajectory and power is defined, including:
[0050] Define the throughput of covert transmission This reflects the transmission efficiency, among which That is, the minimum transmission rate across all hops within each time slot. The optimization objective is to satisfy the stealth constraint and maximize the UAV's transmit power P. MAX Drone trajectory Maximize Φ under constraints C ,Right now:
[0051]
[0052]
[0053]
[0054] C3:||q m [t+1]-q m [t]||=Vδ t ,
[0055] C4:d s ≤||q m [t]-q m+1 [t]||≤d c
[0056] C1 represents the stealth requirement, C2 represents the maximum transmit power constraint of the UAV, and constraints C3 and C4 limit the speed and distance of the UAV, respectively.
[0057] Furthermore, the optimized hidden constraints are given, and the initial optimization problem is rewritten, including:
[0058] Derivation The expression for the t-th m Jump, and we arrive at:
[0059]
[0060] For two complex Gaussian distributions and Where μ p and μ q This corresponds to the mean of the complex Gaussian distribution. and To determine the variance corresponding to the complex Gaussian distribution, the following expression is derived:
[0061]
[0062] Substituting into the formula, we get:
[0063]
[0064] When x≥0, we have Therefore, we can deduce that:
[0065]
[0066] Write constraint C1 as:
[0067]
[0068] The optimization problem can be rewritten as:
[0069]
[0070]
[0071] C2:p m (t)≤P MAX ,
[0072] C3:||q m [t+1]-q m [t]||=Vδ t ,
[0073] C4:d s ≤||q m [t]-q m+1 [t]||≤d c .
[0074] Furthermore, the communication and optimization problem is constructed as a discrete-time Markov decision process, and the optimization problem is solved using the C-MAPPO method, including:
[0075] In a given scenario, UAVs must make collaborative decisions under given constraints. The covert transmission problem of UAV-assisted relay is transformed into a discrete-time Markov decision process, and then the C-MAPPO algorithm is used as the general solution to the corresponding problem.
[0076] The Markov Decision Process (MDP) model is as follows:
[0077] The optimization problem P2 is transformed into a quaternion MDP problem. Where O is the observation space. It is the action space. It is the state transition probability. The following are the definitions of the reward function, observation space, action space, reward function, and state transition probability:
[0078] The observation space of time slot t is composed of Given, where d m,m′ (t)=||q m [t]-q m' [t]|| represents the relative distance between the drones; considering that friendly drones may not be able to directly obtain the position of W, but can estimate the relative distance to W through onboard monitoring equipment, the distance d between the drone and W is included in the observation. m,w (t)=||q m [t]-q w [t]||;
[0079] The action space of a drone is defined as The discrete flight directions of the UAV are: at least two components are 0, and the non-zero components are ±1, i.e., stop, forward, backward, left, right, up, and down.
[0080] The reward function consists of the following two parts:
[0081] Part 1: Based on the optimization objective, calculate the minimum transmission rate R for all hops within each time slot. tr (t) is set as the reward of m, which encourages the drone to choose the appropriate action to achieve a higher transmission rate;
[0082] Part Two: Due to concealment limitation C1, set concealment parameter κ. ca ={0,1}, meaning κ = {0,1}, that is, when all drones satisfy the concealment constraint. ca =1, otherwise κ ca =0; additionally set the constant κ cp ={0,P c} is the penalty for concealment constraints; when drone m does not satisfy the concealment constraints, κ... cp =P c Otherwise κ cp =0;
[0083] The reward for drone m is:
[0084]
[0085] Perform maximum-min normalization on the drone's reward m:
[0086]
[0087] Where η max and η min These represent the maximum and minimum values that the reward can achieve;
[0088] Define state Pr(s) m (t+1)|s m (t),a m (t) represents the state s m Transition to state s m The probability of (t+1), where s m (t)=o m (t);
[0089] The communication network resources for UAV decision-making are divided into two parts: communication capability and spatial location, which are represented by the UAV's transmission power and flight trajectory. The transmission power is determined by the UAV based on the current wireless environment under the requirements of concealment constraints, while the flight trajectory is given by the MAPPO algorithm to provide the flight direction of each UAV. Subsequently, the environment provides the transmission rate of all current hops based on the concealment constraints, evaluates the current reward, and transmits the observation and state to MAPPO to make the next flight direction decision based on the reward function evaluation.
[0090] The UAV system based on the algorithm proposed in this application can effectively make power and trajectory decisions, exhibiting significant performance advantages. The optimization method provided in this application not only accelerates the learning speed of the agent and simplifies the complexity of the action space, but also improves the UAV's adaptive learning ability and real-time response capability. The UAV can optimize power allocation and flight trajectory in real time according to environmental changes and stealth requirements, thereby improving flight efficiency and communication resource utilization efficiency. Simultaneously, this method enhances the system's robustness, making the UAV more stable and reliable in the face of uncertainty and abnormal situations. Furthermore, the ability for multi-UAV collaborative operation and scalability support for large-scale UAV swarms further improves the stealth transmission throughput. The optimized power and trajectory decisions also help improve flight safety, reduce collision risks, and ensure the successful completion of missions. Attached Figure Description
[0091] Figure 1 A flowchart illustrating the implementation of the UAV multi-hop relay covert communication method based on trajectory and power joint design provided in this application embodiment.
[0092] Figure 2 A schematic diagram of a system for implementing the UAV multi-hop relay covert communication method based on trajectory and power joint design provided in this application embodiment.
[0093] Figure 3 This is a schematic diagram of the trajectories of various UAVs under different concealment constraints provided in the embodiments of this application.
[0094] Figure 4 This is a schematic diagram showing the power variations of various UAVs under different concealment constraints provided in the embodiments of this application.
[0095] Figure 5 The overall covert transmission throughput of the system under different algorithms with different covert constraints provided in the embodiments of this application.
[0096] Figure 6 The reward for each evaluation environment feedback in the different algorithms provided in the embodiments of this application. Detailed Implementation
[0097] To address the issue that current technologies struggle with covert communication due to excessively long communication distances between source and destination UAVs, making it difficult for relay networks composed of single relay UAVs to meet the requirements of covert communication, this application proposes a multi-hop relay covert communication network constructed by deploying multiple relay UAVs. In this network, data is transmitted hop-by-hop, with each relay node receiving and retransmitting the data to the next relay node or destination node. The wireless network's coverage area is significantly larger than that of a single relay node, effectively solving the problem of covert communication over large scales. Furthermore, multi-hop relay communication networks offer advantages in scalability, flexibility, and robustness, and are particularly effective in complex network environments and advanced detection technologies. For the problems of UAV multi-hop relay networks, multi-agent learning methods provide a new perspective on packet routing in UAV networks. In multi-hop relay covert communication systems, under aerial UAV surveillance scenarios, ground-based wireless networks utilize relay transmission to achieve covert communication, effectively improving system throughput. Furthermore, by utilizing efficient algorithms to find the optimal path and maximizing throughput and minimizing end-to-end latency under concealment constraints, covert communication can also be achieved in UAV ad hoc networks with multi-hop communication. These provide important theoretical and technical support for the construction of UAV multi-hop relay covert communication networks.
[0098] The embodiments of this application will now be described in conjunction with the accompanying drawings.
[0099] Step 1: Set up the UAV multi-hop relay system and define the transmission rate of each hop relay link;
[0100] Specifically, step 1 includes:
[0101] A covert multi-hop drone relay system is established; the source drone S communicates with the target drone D with the assistance of multiple relay drones, as shown in the attached diagram. Figure 2 As shown. Assume S and D are in fixed locations due to performing special tasks. Since S and D are too far apart to communicate directly, a multi-hop drone relay network is used to avoid detection by surveillance drones and ensure transmission security. The multi-hop drone relay network consists of M drones, and its index set is represented as... Let S be the first drone and D be the Mth drone. Assume the drones' frequency bands are orthogonal and the system operates in full-duplex, time-slot mode. And each time slot δ t If the duration is short enough, then the drone's coordinates are represented as q. m [t]=(x[t],y[t],z[t]), where S, D, the relay drone, and the surveillance drone W are each equipped with a single omnidirectional antenna. The communication drone estimates its relative distance to W by using radar or a camera. The initial position of the surveillance drone W is defined as q. w =(x w ,y w ,z w ), with a speed of v w (t)=(v wx ,v wy ,v wz If W = q, then the trajectory of W is represented as q. w [t] = q w +δ t ·v w (t), where
[0102] Assume there are M drones in the network, let... This represents the set of indices for each hop in the linear relay path connecting the source and destination drones. These indices represent specific drones in the drone network; they are nodes on the data transmission path. ι1 is the first hop on the path, ι2 is the second hop, and so on, up to ι M-1 This is the last hop before reaching the destination node. Let d be the maximum communication distance between drones in each hop. c The minimum safe distance is d s .
[0103] In UAV communication, the channel gain between any pair of communicating UAVs m and target UAVs m′ follows a free-space path loss model, where...
[0104]
[0105] Where ρ0 represents the path loss at the reference distance, which is defined as d0 = 1 meter.
[0106] p m Let (t) be the transmission power of UAV m at time t. Then the power received by UAV m from UAV m′ at point m is expressed as:
[0107]
[0108] Then the SNR received at point m+1 from drone m is:
[0109]
[0110] in, in Let B be the power spectral density of the noise, and B be the bandwidth.
[0111] The transmission rate between adjacent drones is expressed as:
[0112] R m,m+1 (t)=Blog2[1+α m,m+1 (t)],
[0113] That is, the ιth m The hop transmission rate is expressed as
[0114] Step 2: Set covert constraints for each hop of the communication link from the perspective of monitoring the drone;
[0115] Specifically, step 2 includes:
[0116] Under the DF (decode-and-forward, DF) relay protocol, let the ιth... m The signal y received by the adjacent drone m+1 during the jump is from m. m,m+1 And the received signal of W is y m,w The corresponding channel gains are ||h m,m+1 (t)|| and ||h m,w (t)||, then:
[0117]
[0118] in For the ι m The symbol for jump transmission, n m+1 and These represent the noise received by drones m+1 and W, respectively. To detect the presence of signal transmission, monitoring the signal received by drone W includes the following two cases:
[0119]
[0120] L is the codeword length in the channel-related block, H1 represents the case where there is signal transmission between friendly nodes, and H0 represents the case where there is no signal transmission between friendly nodes. For ease of analysis, we assume xι m It is a unit amplitude phase shift keying (PSK) signal, with phase denoted as ψ. The received signal distribution at W is as follows:
[0121]
[0122] P MD =P{H0|H1 is true} Missed detection, P FA =P{H1|H0 is true} False alarm.
[0123] Assuming the drone makes decisions based on all signals it receives during multi-hop transmission, to ensure stealth, the probability of detection error at each hop must satisfy P. MD +P FA ≥1-∈,∈ reflects the strictness of the constraint on the bit error rate of the detection error, and is generally small.
[0124] The hidden constraints for each relay link are set as follows:
[0125] P MD +P FA ≥1-∈.
[0126] Step 3: Set concealment constraints for each hop of the communication link from the perspective of the communication drone;
[0127] Specifically, step 3 includes:
[0128] The lower bound of the probability of detection error is:
[0129]
[0130] If H1 or H0 is true, Q1(t) or Q0(t) represents the joint probability distribution of all W received signals at time t, and D(Q1(t)||Q0(t)) is the relative entropy between Q1 and Q0, i.e.
[0131]
[0132] We obtain D(Q1(t)||Q0(t))≤2∈ 2 This is a constraint that is more in line with expectations.
[0133] Based on the definition of relative entropy and the independence of signals received on different hops, we can derive...
[0134] The hidden constraints of each relay link in step 2 are rewritten as follows:
[0135]
[0136] Step 4: Define the initial optimization problem for covert communication of multi-hop relay UAVs based on the joint determination of trajectory and power;
[0137] Specifically, step 4 includes:
[0138] Define the throughput of covert transmission This reflects the transmission efficiency, among which This refers to the minimum transmission rate across all hops within each time slot. The optimization objective is to satisfy step 3, under the concealment constraint, the maximum value P of the UAV's transmission power. MAX Drone trajectory Maximize Φ under constraints C ,Right now
[0139]
[0140]
[0141]
[0142] C3:||q m [t+1]-q m [t]||=Vδ t ,
[0143] C4:d s ≤||q m [t]-q m+1 [t]||≤d c .
[0144] C1 represents the stealth requirement, C2 represents the maximum transmit power constraint of the UAV, and constraints C3 and C4 limit the speed and distance of the UAV, respectively.
[0145] Step 5: Given the optimized hidden constraints, rewrite the initial optimization problem;
[0146] Specifically, step 5 includes:
[0147] Derivation The expression for the ιth m Jump, get out
[0148]
[0149] For two complex Gaussian distributions and Where μ p and μ qThis corresponds to the mean of the complex Gaussian distribution. and To determine the variance corresponding to the complex Gaussian distribution, the following expression is derived:
[0150]
[0151] Substituting into the formula, we get:
[0152]
[0153] When x≥0, we have Therefore, the derivation yields:
[0154]
[0155] Write constraint C1 as:
[0156]
[0157] The optimization problem in step 4 can be rewritten as:
[0158]
[0159]
[0160] C2:p m (t)≤P MAX ,
[0161] C3:||q m [t+1]-q m [t]||=Vδ t ,
[0162] C4:d s ≤||q m [t]-q m+1 [t]||≤d c .
[0163] Step 6: Construct the communication and optimization problem as a discrete-time Markov decision process and solve the optimization problem using the C-MAPPO method.
[0164] Specifically, step 6 includes:
[0165] In a given scenario, unmanned aerial vehicles (UAVs) must make collaborative decisions under specific constraints to determine the optimal flight trajectory and transmission power, thereby achieving the highest possible data transmission rate. Solving this complex optimization problem using traditional methods is challenging. Therefore, this paper proposes the C-MAPPO algorithm, based on the MAPPO (multi-agent proximal policy optimization) reinforcement learning method, to jointly optimize the UAV's flight trajectory and transmission power to address the proposed optimization problem.
[0166] The covert transmission problem of UAV-assisted relay is transformed into a discrete-time Markov decision process (MDP), and then the C-MAPPO algorithm is used as a general solution to the corresponding problem.
[0167] The Markov Decision Process (MDP) model is as follows:
[0168] The optimization problem P2 is transformed into a quaternion MDP problem. Where O is the observation space. It is the action space. It is the state transition probability. The following are the definitions of the reward function, observation space, action space, reward function, and state transition probability:
[0169] The observation space of time slot t is composed of Given, where d m,m′ (t)=||q m [t]-q m' [t]|| represents the relative distance between the drones; considering that friendly drones may not be able to directly obtain the position of W, but can estimate the relative distance to W through onboard monitoring equipment, the distance d between the drone and W is included in the observation. m,w (t)=||q m [t]-q w [t]||.
[0170] The action space of a drone is defined as The discrete flight directions of the UAV are: at least two components are 0, and the non-zero components are ±1, i.e., stop, forward, backward, left, right, up, and down.
[0171] The reward function consists of the following two parts:
[0172] Part 1: Based on the optimization objective, calculate the minimum transmission rate R for all hops within each time slot. tr (t) is set as the reward of m, which encourages the drone to choose the appropriate action to achieve a higher transmission rate;
[0173] Part Two: Due to concealment limitation C1, set concealment parameter κ. ca ={0,1}, meaning κ = {0,1}, that is, when all drones satisfy the concealment constraint. ca =1, otherwise κ ca =0; additionally set the constant κ cp ={0,P c} is the penalty for concealment constraints; when drone m does not satisfy the concealment constraints, κ... cp =P c Otherwise κ cp =0.
[0174] The reward for drone m is:
[0175]
[0176] To improve the stability and efficiency of the learning process and help the agent better adapt to the environment, the drone's reward m is processed using Min-Max Normalization:
[0177]
[0178] Where η max and η min These represent the maximum and minimum values that the reward can achieve, respectively.
[0179] Define state Pr(s) m (t+1)|s m (t),a m (t) represents the state s m Transition to state s m The probability of (t+1), where s m (t)=o m (t).
[0180] Combining the specific requirements of covert communication with the advantages of MAPPO in the field of multi-agent reinforcement learning, this application uses the C-MAPPO joint optimization algorithm to solve this problem. The method provided in this application divides the communication network resources for UAV decision-making into two parts: communication capability and spatial location, specifically represented by the UAV's transmit power and flight trajectory. The transmit power is derived by the UAV based on the current wireless environment under covert constraints, while the flight trajectory is determined by the MAPPO algorithm, which provides the flight direction for each UAV. Subsequently, the environment provides the transmission rate for all current hops based on the covert constraints, evaluates the current reward, and transmits the observations and state to MAPPO to make the next flight direction decision based on the reward function.
[0181] The C-MAPPO algorithm for dynamic multi-hop UAV covert communication based on multi-agent deep reinforcement learning provided in this application is as follows:
[0182] Step 61, Initialization phase: First, orthogonally initialize the parameters θ of the policy function π and φ and the parameters of the value function V to ensure that each parameter is independent at the start of training;
[0183] Step 62, Training Loop: Within no more than the maximum number of steps. max Under these conditions, the main training loop begins;
[0184] Step 63, Data Buffer Preparation: At the beginning of each iteration, clear the data buffer D to prepare for the collection of new training data;
[0185] Step 64, Batch Data Collection: For each batch size (batch_size), perform the following operations:
[0186] a. Initialize an empty list τ to store the trajectory data of a single agent;
[0187] b. Initialize the states of the recurrent neural network (RNN) for policy and value, respectively, as h 0,π and h 0,V ;
[0188] Step 65, Agent Interaction with Environment: At each time step t, from 1 to T, for all agents m:
[0189] a. Observe the local state o(t) of the agent in the environment;
[0190] b. Make a decision a(t) based on the policy function π;
[0191] c. Execute action a(t), and calculate the optimal transmission power p(t) under the premise of satisfying the concealment constraint;
[0192] d. Collect the signal strength χ(t) of all agents and the environmental state o(t+1) at the next time step;
[0193] Step 66, Data storage: Add the collected data τ to the data buffer D, including local state, policy and hidden state of the value network, actions, signal strength and environmental state at the next time step;
[0194] Step 67, Minimum Batch Data Processing: For each minimum batch k, from 1 to K, randomly sample a minimum batch b from the data buffer D, including data from all agents;
[0195] Step 68, Parameter Update: Using the randomly selected minimum batch b, update the parameters θ and φ using the Adam optimization algorithm to minimize their respective loss functions L(θ) and L(φ), including:
[0196] a. For each data block c, use the first hidden state in the min-batch b to update the hidden states of policy π and value V;
[0197] b. Use the entire batch of data to update parameters to optimize the performance of the strategy and value function.
[0198] The present application will be further described below with reference to the accompanying drawings.
[0199] Figure 1 This is a flowchart illustrating the implementation of the UAV multi-hop relay covert communication method based on trajectory and power joint design, as per the present invention. The method establishes the position and initial state of each UAV within the multi-hop relay covert communication system. Then, it defines the calculation method for the transmission rate of each relay link and sets covert constraints from the perspectives of both the detection and communication UAVs. Based on this, an optimization problem for UAV multi-hop relay covert communication based on trajectory and power joint design is established. In one step of each round, the UAV's flight actions are determined by MAPPO based on the current observations of each UAV and superimposed as the global state o(t) = [o1(t), o2(t), ..., o...]. m Make a decision a(t) = [a1(t), a2(t), ..., a m [(t)], then the drone performs action a(t). After the drone performs the action, the optimal transmission power is calculated based on the current wireless environment. The UAV determines the transmission power p(t) = [p1(t), p2(t), ..., p based on this decision. m [(t)], after which the environment evaluates the reward for the current step and returns the reward and observation at this point. Repeat the above steps until the end of the round.
[0200] Figure 2 This is a schematic diagram of a system implementing the UAV multi-hop relay covert communication method based on trajectory and power joint design proposed in this invention. Assume that S and D are in fixed positions due to performing special tasks. Since S and D are too far apart to communicate directly, to avoid detection of their communication behavior by the monitored UAV W and to ensure transmission security, the source UAV S communicates with the target UAV D with the assistance of multiple relay UAVs. The UAVs' frequency bands are orthogonal to each other, and the system operates in full-duplex time-slot mode, with the transmitting and receiving signals occupying the same bandwidth. D, the relay UAVs, and W are each equipped with a single omnidirectional antenna. The communicating UAV can estimate its relative distance to W by carrying radar or a camera.
[0201] Use respectively Figure 3 and Figure 4This invention explains the trajectory and power changes of each UAV in a multi-hop relay covert communication method based on trajectory and power joint design, where T = 100s. First, when ∈ = 0.5, the detection intensity of W is weak, and the relay UAV hovers near its initial position, actively seeking a suitable flight strategy to achieve the system's maximum covert transmission rate and maintaining a relatively high transmission power at certain times. Furthermore, UAV S quickly adjusts its power to maximum intensity after W moves away. Second, as the detection intensity gradually increases (from ∈ = 0.1 to ∈ = 0.01), the relay UAV gradually moves back to its initial position as W approaches. Once it leaves W's detection range, the relay UAV continues to search for a suitable position to increase the transmission rate. Figure 3 As shown, it will also adjust its transmission power to achieve the peak value required to meet concealment constraints, such as... Figure 4 As shown.
[0202] Figure 5 The figure demonstrates the overall covert transmission throughput of the system implemented using the UAV multi-hop relay covert communication method based on trajectory and power joint design proposed in this invention. As can be seen from the figure, the C-MAPPO joint optimization algorithm proposed in this invention has a higher covert transmission throughput than other comparative and baseline algorithms, exhibiting significant advantages. Furthermore, the system's covert transmission throughput also increases with the increase of ∈.
[0203] Figure 6 The figure illustrates the implementation of the proposed UAV multi-hop relay covert communication method based on joint trajectory and power design under different covert constraint strengths, using the rewards from each evaluation environment feedback in different algorithms. Since the OP-Optimize optimization algorithm does not employ reinforcement learning, it is not shown in the figure. As can be seen from the figure, the proposed C-MAPPO joint algorithm exhibits significant advantages in reward convergence speed and peak reward level compared to other single optimization algorithms.
[0204] This application constructs a joint trajectory and transmit power optimization problem in a dynamic multi-hop UAV covert communication system. Considering covertness requirements, maximum power, and trajectory constraints, a joint optimization method combining optimization and multi-agent reinforcement learning is proposed. This method combines the advantages of optimization algorithms and reinforcement learning, respectively used for policy analysis of UAV transmit power and flight trajectory, and makes joint decisions. The proposed scheme can effectively realize a cooperative covert communication strategy in a multi-UAV relay communication system and maximize system throughput.
[0205] The embodiments described above do not constitute a limitation on the scope of protection of this application.
Claims
1. A method for covert multi-hop relay communication of unmanned aerial vehicles based on joint trajectory and power design, characterized in that, The method includes: Step 1: Set up the UAV multi-hop relay system and define the transmission rate of each hop relay link; Step 2: Set the concealment constraints for each hop of the communication link from the perspective of monitoring the drone; Step 3: Set the concealment constraints for each hop of the communication link from the perspective of the communication drone; Step 4: Define the initial optimization problem for covert communication of multi-hop relay UAVs based on the joint determination of trajectory and power; Step 5: Give the optimized hidden constraints and rewrite the initial optimization problem; Step 6: Construct the communication and optimization problem as a discrete-time Markov decision process and solve the optimization problem using the C-MAPPO method; The initial optimization problem for covert communication of multi-hop relay UAVs based on the joint determination of trajectory and power is defined, including: A multi-hop drone relay network consists of M drones, and the index set of the multi-hop drone relay network is represented as follows: set up This represents the set of indices for each hop in the linear relay path connecting the source UAV and the destination UAV; where ι1 is the first hop on the path, ι2 is the second hop, and ι... M-1 It is the last hop before reaching the destination node; the maximum communication distance between drones in each hop is d. c The minimum safe distance is d s Assuming the UAV's frequency bands are orthogonal and the system operates in full-duplex time-slot mode, that is... And each time slot δ t If the duration is short enough, then the drone's coordinates are represented as q. m [t]=(x[t],y[t],z[t]), where Define the throughput of covert transmission. This reflects the transmission efficiency, among which That is, the minimum transmission rate among all hops in each time slot t, with the optimization objective being to satisfy the stealth constraint and maximize the UAV's transmit power P. MAX Drone trajectory Maximize Φ under constraints C ,Right now: C3:||q m [t+1]-q m [t]||=Vδ t , C4:d s ≤||q m [t]-q m+1 [t]||≤d c Where C1 is the concealment requirement, ∈ reflects the strictness of the constraint on the detection error rate; C2 is the maximum transmit power constraint of the UAV, p m (t) represents the transmit power of UAV m within time slot t; constraints C3 and C4 limit the speed and distance of the UAV, respectively.
2. The method according to claim 1, characterized in that, Configure a multi-hop relay system for UAVs and define the transmission rate of each relay link, including: A multi-hop drone relay system is configured: the source drone S communicates with the target drone D with the assistance of multiple relay drones; S and D are assumed to be in fixed locations due to performing special tasks; S and D cannot communicate directly due to the distance between them, and a multi-hop drone relay network is used to avoid detection by monitoring drones. The multi-hop drone relay network includes M drones, and the index set of the multi-hop drone relay network is represented as follows. The first drone is designated S, and the Mth drone is designated D. Assuming the drone's frequency bands are orthogonal and the system operates in full-duplex time-slot mode, i.e. And each time slot δ t If the duration is short enough, then the drone's coordinates are represented as q. m [t]=(x[t],y[t],z[t]), where S, D, the relay drone, and the surveillance drone W are each equipped with a single omnidirectional antenna. The communication drone estimates its relative distance to W by using radar or a camera. The initial position of the surveillance drone W is defined as q. w =(x w ,y w ,z w ), with a speed of v w (t)=(v wx ,v wy ,v wz If W = q, then the trajectory of W is represented as q. w [t] = q w +δ t ·v w (t), where Assume there are M drones in the network, let... This represents the set of indices for each hop in the linear relay path connecting the source UAV and the destination UAV; where ι1 is the first hop on the path, ι2 is the second hop, and ι... M-1 This is the last hop before reaching the destination node; let d be the maximum communication distance between drones in each hop. c The minimum safe distance is d s ; In UAV communication, the channel gain between any pair of communicating UAVs m and m′ follows a free-space path loss model, where... Where ρ0 represents the path loss at the reference distance, and the reference distance, i.e. the unit distance, is defined as d0 = 1 meter; p m (t) represents the transmission power of UAV m at time t; then the power received by UAV m from UAV m′ at UAV m is expressed as: Then the SNR received at point m+1 from drone m is: in, in Let B be the power spectral density of the noise, and B be the bandwidth. The transmission rate between adjacent drones is expressed as: R m,m+1 (t)=Blog2[1+α m,m+1 (t)] That is, the ιth m The hop transmission rate is expressed as 3. The method according to claim 1, characterized in that, From the perspective of monitoring drones, setting covert constraints for each hop of the communication link includes: Under the DF relay protocol, let the ιth... m The signal y received by the adjacent drone m+1 from drone m is... m,m+1 And the received signal of the monitoring drone W is y m,w The corresponding channel gains are ||h m,m+1 (t)|| and ||h m,w (t)||, then: in For the ι m The symbol for jump transmission, n m+1 and Let m and W represent the noise received by UAV m+1 and the monitoring UAV W, respectively. To detect the presence of signal transmission, the signal received by the monitoring UAV W includes the following two cases: L is the codeword length in the channel-related block, H1 represents the case where there is signal transmission between friendly nodes, and H0 represents the case where there is no signal transmission between friendly nodes. (Assume...) It is a unit amplitude phase shift keying signal, with phase denoted as ψ. The received signal distribution at the monitoring UAV W is as follows: P MD =P{H0|H1 is true} Missed detection, P FA =P{H1|H0 is true} False alarm; Assuming the drone makes decisions based on all signals it receives during multi-hop transmission, to ensure stealth, the probability of detection error at each hop must satisfy P. MD +P FA ≥1-∈, ∈ reflects the strictness of the constraint on the bit error rate of the detection error; The hidden constraints for each relay link are set as follows: P MD +P FA ≥1-∈。 4. The method according to claim 1, characterized in that, From the perspective of the communication drone, the covert constraints of each hop of the communication link are set, including: The lower bound of the probability of detection error is: If H1 or H0 is true, Q1(t) or Q0(t) represents the joint probability distribution of the signals received by all monitoring drones W at time t, and D(Q1(t)||Q0(t)) is the relative entropy between Q1 and Q0, i.e.: We obtain D(Q1(t)||Q0(t))≤2∈ 2 This is a constraint that is more in line with expectations; Based on the definition of relative entropy and the independence of signals received on different hops, we can derive... The hidden constraints of each relay link are rewritten as follows:
5. The method according to claim 1, characterized in that, Given the optimized hidden constraints, rewrite the initial optimization problem, including: Derivation The expression for the ιth m Jump, and we arrive at: For two complex Gaussian distributions and Where μ p and μ q This corresponds to the mean of a complex Gaussian distribution. and To determine the variance corresponding to the complex Gaussian distribution, the following expression is derived: Substituting into the formula, we get: When x≥0, we have Therefore, we can deduce that: Write constraint C1 as: The optimization problem can be rewritten as: C2:p m (t)≤P MAX C3:||q m [t+1]-q m [t]||=Vδ t C4:d s ≤||q m [t]-q m+1 [t]||≤d c 。 6. The method according to claim 1, characterized in that, The communication and optimization problems are conceived as discrete-time Markov decision processes, and the optimization problems are solved using the C-MAPPO method, including: In a given scenario, UAVs must make collaborative decisions under given constraints. The covert transmission problem of UAV-assisted relay is transformed into a discrete-time Markov decision process, and then the C-MAPPO algorithm is used as the general solution to the corresponding problem. The Markov Decision Process (MDP) model is as follows: The optimization problem P2 is transformed into a quaternion MDP problem. Where O is the observation space. It is the action space. It is the state transition probability. The following are the definitions of the reward function, observation space, action space, reward function, and state transition probability: The observation space of time slot t is composed of Given, where d m,m′ (t)=||q m [t]-q m' [t]|| represents the relative distance between the drones; considering that friendly drones may not be able to directly obtain the position of the surveillance drone W, but can estimate the relative distance to the surveillance drone W through onboard monitoring equipment, the distance d between the drone and the surveillance drone W is included in the observation. m,w (t)=||q m [t]-q w [t]||; The action space of a drone is defined as The discrete flight directions of the UAV are: stop, forward, backward, left, right, up, and down. The reward function consists of the following two parts: Part 1: Based on the optimization objective, calculate the minimum transmission rate R for all hops within each time slot. tr (t) is set as the reward for drone m, which encourages the drone to choose appropriate actions to achieve a higher transmission rate; Part Two: Due to concealment limitation C1, set concealment parameter κ. ca ={0,1}, meaning κ = {0,1}, that is, when all drones satisfy the concealment constraint. ca =1, otherwise κ ca =0; additionally set the constant κ cp ={0,P c } is the penalty for concealment constraints; when drone m does not satisfy the concealment constraints, κ... cp =P c Otherwise κ cp =0; The reward for drone m is: Perform maximum-min normalization on the drone's reward m: Where η max and η min These represent the maximum and minimum values that the reward can achieve; Define state Pr(s) m (t+1)|s m (t),a m (t) represents the state s m Transition to state s m The probability of (t+1), where s m (t)=o m (t); The communication network resources for UAV decision-making are divided into two parts: communication capability and spatial location, which are represented by the UAV's transmission power and flight trajectory. The transmission power is determined by the UAV based on the current wireless environment under the requirements of concealment constraints, while the flight trajectory is given by the MAPPO algorithm to provide the flight direction of each UAV. Subsequently, the environment provides the transmission rate of all current hops based on the concealment constraints, evaluates the current reward, and transmits the observation and state to MAPPO to make the next flight direction decision based on the reward function evaluation.