Track planning and resource allocation method for unmanned aerial vehicle assisted covert communication system
The MDP and FMADQN algorithms optimize the trajectory and resource allocation of drones, combined with the DQN algorithm, optimize the position and power control of interfering with drone, solve the problem of maximizing communication rates in the drone assisted hidden communication system, and maximize the system communication rate.
Patent Information
- Application Number
- CN202510500389.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-04
AI Technical Summary
In the drone-assisted concealed communication system, how to optimize the design of the drone's flight trajectory and system resource allocation strategy while meeting the concealment constraints to maximize the communication rate.
Markov decision-making process (MDP) and the federal multi-agent deep Q network (FMADQN) algorithm are used to jointly optimize the trajectory and resource allocation of drones, and combined with the deep Q network (DQN) algorithm to optimize the position and power control of interfering drones, and solve the problems of drone trajectory planning and resource allocation through alternating iteration methods.
Under the concealment constraint, the system communication rate is maximized and the performance of the drone-assisted concealed communication system is improved.
Smart Images

Figure CN120264470A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of wireless communication, and relates to a method for trajectory planning and resource allocation of an unmanned aerial vehicle (UAV)-assisted covert communication system. Background Art
[0002] With the rapid development of wireless communication technology, covert communication, as a key technology to ensure the security of information transmission, has important application values in fields such as military reconnaissance, emergency disaster relief, and privacy protection. Covert communication technology makes it difficult for potential eavesdroppers to effectively detect communication behaviors by means of controlling the transmitted signal power, optimizing the transmission strategy, etc. In recent years, due to its characteristics such as flexible deployment and wide-area coverage, UAV has become an aerial platform that can achieve efficient communication. However, in a UAV-assisted covert communication system, how to optimize the flight trajectory of UAVs and the system resource allocation strategy under the constraint of covertness has become an urgent problem to be solved.
[0003] Currently, there are already literatures studying the problems of UAV-assisted covert communication systems. For example, some literatures jointly optimize the transmission power of UAVs and the interference power based on the particle swarm optimization algorithm; some literatures introduce friendly interference nodes and adopt the successive convex approximation algorithm to jointly optimize the flight trajectory and transmission power of UAVs. However, existing studies rarely consider multi-UAV cooperation and improving the performance of covert communication systems through dynamic interference management. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a method for trajectory planning and resource allocation of a UAV-assisted covert communication system, which aims at a covert communication system including multiple transmitting UAVs, one interfering UAV, multiple users, multiple eavesdroppers, and one high-altitude platform, and realizes the UAV trajectory planning and resource allocation strategy under the constraint of covertness with the maximization of the system communication rate as the optimization goal.
[0005] To achieve the above purpose, the present invention provides the following technical solutions:
[0006] A method for trajectory planning and resource allocation of a UAV-assisted covert communication system, the method specifically includes the following steps:
[0007] S1: Model the scenario of the UAV-assisted covert communication system;
[0008] S2: Model the channel model;
[0009] S3: Model the system communication rate;
[0010] S4: Model the probability of incorrect detection of eavesdroppers;
[0011] S5: Model the constraint conditions of the UAV flight trajectory and resource allocation;
[0012] S6: Model the sub - problems of the sending UAV trajectory planning and resource allocation as a Markov decision process (MDP);
[0013] S7: Determine the flight trajectory and resource allocation strategy of the sending UAV based on the federated multi - agent deep Q - network (FMADQN) algorithm;
[0014] S8: Model the sub - problems of the jamming UAV position deployment and power control as an MDP;
[0015] S9: Determine the jamming UAV position deployment and power control strategy based on the deep Q - network (DQN) algorithm;
[0016] S10: Use the alternating iteration method to solve the sub - problems of the sending UAV flight trajectory and resource allocation, and the jamming UAV position deployment and power control in turn until the algorithm converges.
[0017] Furthermore, in step S1, model the scenario of the UAV - assisted covert communication system, specifically including: M sending UAVs, 1 friendly jamming UAV (UAV - J), K users, N eavesdroppers, and a high - altitude platform equipped with a parameter server; the sending UAVs start from the initial positions, fly over the area along a certain trajectory within the system time, send messages to the legitimate users within the covered area during the flight, and then return to the starting point; the ground eavesdroppers detect the communication between the sending UAVs and the ground users, and the UAV - J generates artificial noise to hinder the eavesdroppers' detection of the communication signals;
[0018] Let U m represent the m - th sending UAV, where 1 ≤ m ≤ M; let q k =(x k , y k , 0) represent the position of the k - th user, where 1 ≤ k ≤ K; represent the position coordinates of the n - th eavesdropper, where 1 ≤ n ≤ N. Let q J =(x J , y J , H) represent the position of the UAV - J, where H represents the altitude of the UAV; divide the system time into multiple time slots of equal size, let τ represent the length of each time slot, the total number of time slots is T, and let represent the flight trajectory of U m , where represents the position of U m at time slot t, where 0 ≤ t ≤ T; assume that all UAVs fully share the system spectrum resources, and the orthogonal frequency - division multiple access mechanism is adopted inside each UAV to communicate with users. Divide the spectrum into L orthogonal sub - channels of equal length, and let B represent the sub - channel bandwidth.
[0019] Further, in step S2, the channel model is modeled, specifically including: Modeling the air-to-ground link from the transmitting UAV and UAV-J to the user as a probabilistic line-of-sight channel model, let L m,k,l,t denote the average transmission loss (in dB) of the link corresponding to U m occupying sub-channel l to send a message to user k at time slot t, and is modeled as where and are the probabilities that the transmission link from U m to user k is a line-of-sight link and a non-line-of-sight link at time slot t, respectively, and satisfy is modeled as where a and b are coefficients related to the environment, and θ m,k,t denotes the elevation angle between user k and U m at time slot t, and is modeled as and denote the path losses of the LoS and NLoS links when U m occupies sub-channel l to send a message to user k at time slot t, and are modeled as where f l is the carrier frequency of sub-channel l, are the average additional losses of the LoS and NLoS links, respectively, 1 ≤ l ≤ L; denotes the distance between U m and user k at time slot t, and is modeled as Let h m,k,l,t denote the channel gain when U m occupies sub-channel l to communicate with user k at time slot t, and is modeled as Let denote the channel gain between U m and the nth eavesdropper on sub-channel l at time slot t, and let denote the channel gain between UAV-J and user k on sub-channel l at time slot t, denotes the channel gain between UAV-J and the nth eavesdropper on sub-channel l, and is modeled All are probabilistic line-of-sight channels.
[0020] Further, in step S3, the system communication rate is modeled, specifically including: Let γ m,k,l,t denote the signal-to-interference-plus-noise ratio (SINR) of the received signal when U m occupies sub-channel l to send a message to user k at time slot t, and is modeled as where β m,k,l,t denotes the sub-channel allocation variable of user k. If U m occupies sub-channel l to communicate with user k at time slot t, then β m,k,l,t = 1, otherwise βm,k,l,t = 0, p m,k,l,t represents the transmission power when the UAV transmits a message to user k on sub-channel l in time slot t, m represents the transmission power of UAV-J in time slot t, is a uniformly distributed random variable, i.e., where is the maximum AN peak transmission power adjustable by UAV-J, σ 2 represents the link noise power. Let R m,t represent the communication rate of UAV in time slot t, m modeled as
[0021] Furthermore, in step S4, the probability of eavesdropper's error detection is modeled, specifically including: Let s m represent the signal transmitted by the transmitting UAV U m to the users within its coverage area. Let s J represent the signal transmitted by UAV-J to the ground eavesdropper. E[|s m | 2 = 1, E[|s J | 2 = 1. Let y n,l,t represent the received signal of the nth eavesdropper on sub-channel l in time slot t, modeled as where z n,l,t represents the additive white Gaussian noise with a mean of 0 and a variance of σ 2 ; represents the null hypothesis, i.e., the UAV does not send a message to the user, represents the alternative hypothesis, i.e., the UAV sends a message to the user. Assuming that the eavesdropper performs signal detection based on the received signal energy, let T n,t represent the received power of the nth eavesdropper in time slot t, modeled as and respectively represent the corresponding decisions made by the nth eavesdropper for the hypotheses and ; represents the detection threshold of the nth eavesdropper. Assuming L→∞, then T n,t can be rewritten as Let represent the probability of error detection of the nth eavesdropper in time slot t, modeled as where and respectively represent the false alarm probability and the missed alarm probability of the nth eavesdropper in time slot t, modeled as modeled as
[0022] Furthermore, in step S5, the flight trajectory of the UAV and the resource allocation constraint conditions are modeled, specifically including:
[0023] (1) Modeling the UAV flight trajectory constraint conditions, including:
[0024] 1) Maximum flight distance constraint of the UAV in each time slot: v max is the maximum flight speed of the UAV;
[0025] 2) UAV safety distance constraint:
[0026] where d m,m′,t represents the distance between U m and U m′ at time slot t, is the distance between the transmitting UAV U m and UAV-J at time slot t, and d min is the minimum safety distance between UAVs;
[0027] 3) UAV starting position constraint:
[0028] (2) Modeling the resource allocation constraint conditions, including:
[0029] 1) UAV transmission power constraint: where p max is the maximum transmission power of the UAV;
[0030] 2) The user's received SINR needs to meet the minimum SINR constraint: γ m,k,l,t ≥β m,k,l,t γ th ; where γ th represents the minimum SINR threshold of the user;
[0031] 3) UAV-user association constraint: Each user can occupy at most one subchannel to communicate with one UAV in one time slot, expressed as: Each UAV can occupy at most one subchannel to communicate with one user in one time slot, expressed as:
[0032] 4) Subchannel allocation constraint: One time slot can allocate at most one subchannel for one UAV-user pair, expressed as:
[0033] 5) UAV concealment constraint: where ε represents the security threshold.
[0034] Furthermore, in step S6, under the constraints of concealment, transmission power selection, sub-channel allocation, etc., aiming to maximize the system communication rate, the UAV trajectory planning and resource allocation problem is decomposed into a transmission UAV trajectory and resource allocation optimization sub-problem, a UAV-J position deployment and power allocation sub-problem; and they are respectively modeled as MDPs.
[0035] Model the transmission UAV trajectory planning and resource allocation problem as an MDP, and its state, action, and reward functions are specifically expressed as follows:
[0036] (1) State: Let s t = {s 1,t , … s m,t , … s M,t} represent the system environment state of the transmission UAV at time slot t, where s m,t represents the state of U m at time slot t, and s m,t includes the trajectory position and channel gain of the transmission UAV, and is modeled as: where h m,k,t , represent the channel gain vectors, and are modeled as: h m,k,t = [h m,k,1,t , h m,k,2,t ,..., h m,k,L,t .
[0037] (2) Action: Let a t = {a 1,t , …, a m,t , … a M,t} represent the action of the transmission UAV at time slot t, where a m,t represents the action of U m at time slot t. To determine the flight trajectory, power allocation, and sub-channel allocation strategy of U m , a m,t is modeled as: where f0 represents the flight action space of the transmission UAV in each time slot, and is modeled as where d = v max τ; the UAV can only land directly above the landing position, which is expressed as: The position update formula of U m is β m,k,t = [β m,k,1,t , β m,k,2,t ,..., β m,k,L,t represents the sub-channel allocation vector of U m at time slot t, and p m,k,t = [p m,k,1,t , pm,k,2,t ,..., p m,k,L,t represents the power allocation vector of U in time slot t m ; the continuous power allocation variable is converted into discrete power levels by using a discretization mechanism. Specifically, the transmission power of the UAV is evenly divided into Q + 1 levels, and let p q represent the power of the q-th level, modeled as Let δ m,k,l,t,q represent the power level selection variable of U in time slot t m , 0 ≤ q ≤ Q. If the transmission power of U m occupying subchannel l to send a message to user k is p q , then δ m,k,l,t,q = 1, otherwise δ m,k,l,t,q = 0; p m,k,l,t can be modeled as The action a m,t can be rewritten as where δ m,k,l,t = [δ m,k,l,t,0 , δ m,k,l,t,1 ,..., δ m,k,l,t,Q represents the discrete power level selection vector of U in time slot t m ; if the selected action of U m does not meet the system concealment constraint, that is , then set the action to that the transmission power of the UAV is 0, the position remains hovering, and it does not communicate with the user, that is: β m,k,t = 0, δ m,k,l,t = 0;
[0038] (3) Reward function: Let r(s m,t , a m,t ) represent the immediate reward obtained when the state of U in time slot t m is s m,t and the action a m,t is selected to transfer to the next state s m,t+1 , modeled as where λ g , λ c , λ d are all positive numbers. If the UAV fails to return to its initial position within the system time, it will be punished by λ g ; if the distance between the transmitting UAVs is less than the minimum safe distance, it will be punished by λ c ; if the distance between the transmitting UAV and the interfering UAV is less than the minimum safe distance, it will be punished by λ d .
[0039] Further, in step S7, the flight trajectory and resource allocation strategy of the transmitting UAVs are determined based on the FMADQN algorithm; specifically: each transmitting UAV acts as a local client, and the high-altitude platform acts as a server to construct a federated learning framework; each transmitting UAV uses the DQN algorithm to determine the local strategy and uploads the local model parameters to the high-altitude platform; the high-altitude platform uses the average aggregation method to aggregate the model parameters uploaded by each client to update the global model parameters, and distributes the updated global model parameters to each client. The client performs iterative updates based on the received global model parameters until the algorithm converges; let denote the long-term reward of U m , modeled as where γ represents the reward discount factor, 0 ≤ γ ≤ 1, and U m obtains the policy by continuously interacting with the environment so as to make the optimal decision and maximize the long-term reward; specifically: let θ m , respectively denote the prediction network and target network model parameters of U m . During local model training, U m first initializes the prediction network, target network, and creates an experience replay pool At time slot t, U m adopts the ε-greedy strategy to select an action, that is where q represents a random number, 0 ≤ q ≤ 1, a r denotes the action randomly selected from the set of optional actions, and ε represents the exploration rate; U m interacts with the environment according to the selected action to obtain the immediate reward r m,t , transfers to the next state s m,t+1 , and stores the quadruple (s m,t , a m,t , r m,t , s m,t+1 ) in the experience replay pool Randomly samples a small batch of samples from the experience replay pool to train the model. Let Q(s m,t , a m,t , θ m ) represent the Q value of the prediction network, and y(s m,t , a m,t ) represent the Q value of the target network, modeled as Let denote the training loss function of U m . The mean squared error (MSE) is used as the loss function, then is modeled as The gradient descent algorithm is used to update the prediction network parameter θ m , modeled as where α represents the learning rate, 0 ≤ α ≤ 1, and after multiple rounds of iteration, the predicted network parameter θ m is assigned to the target network parameter Repeat the above steps until the algorithm converges to obtain the flight trajectory of the transmitting UAV and the resource allocation strategy; the transmitting UAV will send its updated parameter θ m to the high-altitude platform; let θ g represent the global model parameter, modeled as
[0040] Furthermore, in step S8, the sub-problem of the interference UAV position deployment and power control is modeled as an MDP. Specifically: the modeled MDP can be expressed as follows: (1) State: Let represent the state of the agent UAV-J at time slot t, which includes the position and channel gain of UAV-J, modeled as where represents the position of UAV-J at time slot t, represents the channel gain vector, modeled as (2) Action: Let represent the action of UAV-J at time slot t. To determine the deployment position and peak power control strategy of UAV-J, is modeled as represents the position deployment action space of the agent UAV-J, modeled as where d = v max τ, and the UAV-J position update formula is Using the discretization mechanism, the continuous peak power control variable of UAV-J is converted into a discrete power order. Let represent the power order selection variable of UAV-J at time slot t, 0 ≤ q ≤ Q. If the peak transmission power of UAV-J at time slot t is p q , then Otherwise The peak power control variable can be obtained as is Action can be rewritten as where represents the discrete power order selection vector of UAV-J at time slot t; if the action selected by UAV-J does not satisfy the concealment constraint, the action is changed to f t J = [0, 0, 0], (3) Reward function: Let represent the immediate reward obtained when the state of UAV-J at time slot t is and the action is selected, modeled as where λ w, λ s is a positive number. If the distance between any transmitting UAV and UAV-J is less than the minimum safety distance, a penalty of λ is imposed. w If the user receiving SINR is lower than the minimum SINR threshold, a penalty of λ is imposed. s .
[0041] Furthermore, in step S9, based on DQN, determine the interference UAV position deployment and power control strategy; specifically: Let represent the long-term reward of U m , modeled as UAV-J continuously interacts with the environment to obtain the policy and thus makes an optimal decision to maximize the long-term reward; specifically: Let θ J , represent the prediction network and target network model parameters of UAV-J respectively. When training the local model, UAV-J first initializes the prediction network, target network, and creates an experience replay pool At time slot t, UAV-J adopts an ε J greedy strategy to select an action, that is q J represents a random number, 0 ≤ q J ≤ 1, represents the action randomly selected from the set of optional actions, ε J represents the exploration rate. UAV-J interacts with the environment according to the selected action to obtain the immediate reward r t J , and transfers to the next state and stores the quadruple in the experience replay pool Randomly sample a small batch of samples from the experience replay pool to train the model. Let represent the Q value of the prediction network, represent the Q value of the target network, modeled as Let represent the training loss function of UAV-J. The mean squared error (MSE) is used as the loss function, then is modeled as Use the gradient descent algorithm to update the prediction network parameter θ J , modeled as where α J represents the learning rate, 0 ≤ α J ≤ 1. After multiple rounds of iteration, assign the prediction network parameter θ J to the target network parameter Repeat the above steps until the algorithm converges to obtain the interference UAV position deployment and power control strategy.
[0042] Further, in step S10, the alternating iteration method is used to sequentially solve the transmission UAV trajectory planning and resource allocation sub-problems, and the jamming UAV position deployment and power control sub-problems until the algorithm converges; specifically, it includes: given the peak transmission power and position deployment strategy of the initial UAV-J, each transmission UAV determines the flight trajectory, sub-channel, and power allocation strategy according to the FMADQN algorithm; based on the flight trajectory, sub-channel, and power allocation strategy of the transmission UAV, UAV-J determines the position deployment and peak transmission power according to the DQN algorithm; repeat the above process until the algorithm converges.
[0043] The beneficial effects of the present invention are as follows: In the UAV-assisted covert communication system of the present invention, under the covertness constraint, the flight trajectory and resource allocation strategy of the transmission UAV, and the position deployment and power control strategy of the jamming UAV are jointly optimized, achieving the maximization of the system communication rate.
[0044] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following specification. Brief Description of the Drawings
[0045] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0046] Figure 1 It is a schematic diagram of the trajectory planning and resource allocation scenario of the UAV-assisted covert communication system of the present invention;
[0047] Figure 2 It is a flowchart of the trajectory planning and resource allocation of the UAV-assisted covert communication system of the present invention. Detailed Embodiment
[0048] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0049] Among them, the attached drawings are only for illustrative purposes, showing only schematic diagrams rather than physical diagrams, and should not be construed as limiting the present invention; for better illustrating the embodiments of the present invention, some components in the attached drawings will be omitted, enlarged or reduced, which does not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the attached drawings may be omitted.
[0050] In the attached drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the attached drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the attached drawings are only for illustrative purposes and should not be construed as limiting the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0051] Please refer to Figures 1 to 2 , the present invention provides a method for trajectory planning and resource allocation of an unmanned aerial vehicle (UAV)-assisted covert communication system, Figure 1 which is a schematic diagram of the scenario of the UAV-assisted covert communication system constructed for the present invention. As Figure 1 shown, this communication system includes multiple transmitting UAVs, one interfering UAV, multiple users, multiple eavesdroppers, and a covert communication system with a high-altitude platform. Under the constraint of concealment and with the maximization of the system communication rate as the optimization goal, a UAV trajectory planning and resource allocation strategy is realized.
[0052] Figure 2 which is a schematic flowchart of the method for trajectory planning and resource allocation of the UAV-assisted covert communication system of the present invention. As Figure 2 shown, this method specifically includes the following steps:
[0053] Step 1: Modeling the scenario of the UAV-assisted covert communication system;
[0054] Modeling the scenario of the UAV-assisted covert communication system specifically includes: M transmitting UAVs, 1 friendly interfering UAV (UAV-J), K users, N eavesdroppers, and a high-altitude platform equipped with a parameter server; the transmitting UAVs start from the initial positions, fly over the area along a certain trajectory within the system time, send messages to the legitimate users within the covered area during the flight, and then return to the starting point; the ground eavesdroppers detect the communication between the transmitting UAVs and the ground users, and the UAV-J generates artificial noise to hinder the eavesdroppers from detecting the communication signals.
[0055] Let U mDenote the \(m\)-th transmitting UAV, where \(1\leq m\leq M\); let \(q\) k =(x k ,y k ,0) denote the location of the \(k\)-th user, where \(1\leq k\leq K\); Denote the location coordinates of the \(n\)-th eavesdropper, where \(1\leq n\leq N\). Let \(q\) J =(x J ,y J ,H) denote the location of UAV-J, where \(H\) represents the altitude of the UAV. Divide the system time into multiple time slots of equal size. Let \(\tau\) denote the length of each time slot, and the total number of time slots is \(T\). Let denote the flight trajectory of U m , where denote the location of U m at time slot \(t\), where \(0\leq t\leq T\). Assume that all UAVs fully share the system spectrum resources, and an orthogonal frequency division multiple access mechanism is adopted inside each UAV to communicate with users. Divide the spectrum into \(L\) orthogonal sub-channels of equal length. Let \(B\) denote the sub-channel bandwidth.
[0056] Step 2: Channel model modeling;
[0057] Model the channel as follows: Model the air-to-ground link from the transmitting UAV and UAV-J to the user as a probabilistic line-of-sight channel model. Let \(L\) m,k,l,t denote the average transmission loss (in dB) of the link corresponding to U m occupying sub-channel \(l\) to send a message to user \(k\) at time slot \(t\). Model it as where and are the probabilities that the transmission link from U m to user \(k\) is a line-of-sight link and a non-line-of-sight link at time slot \(t\), respectively, and satisfy Model it as where \(a\) and \(b\) are coefficients related to the environment, and \(\theta\) m,k,t denotes the elevation angle between user \(k\) and U m at time slot \(t\). Model it as and denote the path losses of the LoS and NLoS links when U m occupies sub-channel \(l\) to send a message to user \(k\) at time slot \(t\), respectively. Model them as where \(f\) l is the carrier frequency of sub-channel \(l\), are the average additional losses of the LoS and NLoS links, respectively, where \(1\leq l\leq L\); denotes the distance between U m and user \(k\) at time slot \(t\). Model it as Let \(h\) m,k,l,t denote U at time slot \(t\)m The channel gain when communicating with user k on subchannel l is modeled as Let denote the channel gain between U m and the nth eavesdropper on subchannel l at time slot t. Let denote the channel gain between UAV-J and user k on subchannel l at time slot t, denote the channel gain between UAV-J and the nth eavesdropper on subchannel l, and is modeled as Both are probabilistic line-of-sight channels.
[0058] Step 3: System communication rate modeling;
[0059] System communication rate modeling is specifically as follows: Let γ m,k,l,t denote the signal-to-interference-plus-noise ratio (SINR) of the received signal when U m occupies subchannel l to send a message to user k at time slot t, and is modeled as where β m,k,l,t denotes the subchannel allocation variable of user k. If U m occupies subchannel l to communicate with user k at time slot t, then β m,k,l,t = 1; otherwise, β m,k,l,t = 0. p m,k,l,t denotes the transmission power when U m occupies subchannel l to send a message to user k at time slot t, denotes the transmission power of UAV-J at time slot t, is a uniformly distributed random variable, that is where is the maximum AN peak transmission power adjustable by UAV-J, and σ 2 denotes the link noise power. Let R m,t denote the communication rate of U m at time slot t, and is modeled as
[0060] Step 4: Modeling the probability of error detection by eavesdroppers;
[0061] Modeling the probability of error detection by eavesdroppers is specifically as follows: Let s m denote the transmission signal of the transmitting UAV U m to the users within its coverage area. Let s J denote the transmission signal of UAV-J to the ground eavesdropper. E[|s m | 2 = 1, E[|s J | 2 = 1. Let y n,l,t denote the received signal of the nth eavesdropper on subchannel l at time slot t, and is modeled as where zn,l,t denotes additive white Gaussian noise with a mean of 0 and a variance of σ 2 ; denotes the null hypothesis, i.e., the UAV does not send a message to the user, denotes the alternative hypothesis, i.e., the UAV sends a message to the user. Assume that the eavesdropper performs signal detection based on the received signal energy. Let T n,t denote the received power of the n-th eavesdropper at the t-th time slot, modeled as and denote the corresponding decisions made by the n-th eavesdropper for the hypotheses and respectively, denote the detection threshold of the n-th eavesdropper. Assume L→∞, then T n,t can be rewritten as Let denote the probability of false detection of the n-th eavesdropper at the t-th time slot, modeled as where and denote the probability of false alarm and the probability of missed alarm of the n-th eavesdropper at the t-th time slot, respectively, modeled as modeled as
[0062] Step 5: Model the constraints on the UAV flight trajectory and resource allocation;
[0063] Model the constraints on the UAV flight trajectory and resource allocation, specifically: (1) Model the constraints on the UAV flight trajectory, including:
[0064] 1) The maximum flight distance constraint of the UAV for each time slot: v max is the maximum flight speed of the UAV;
[0065] 2) The safety distance constraint of the UAV:
[0066] where d m,m′,t denotes the distance between U m and U m′ at the t-th time slot, is the distance between the transmitting UAV U m and UAV-J at the t-th time slot, and d min is the minimum safety distance between UAVs;
[0067] 3) The starting position constraint of the UAV:
[0068] (2) Model the constraints on resource allocation, including:
[0069] 1) UAV transmission power constraint: where p max is the maximum transmission power of the UAV;
[0070] 2) The SINR received by the user needs to satisfy the minimum SINR constraint: γ m,k,l,t ≥β m,k,l,t γ th ; where γ th represents the minimum SINR threshold of the user;
[0071] 3) UAV-user association constraint: Each user can occupy at most one sub-channel to communicate with one UAV within one time slot, expressed as: Each UAV can occupy at most one sub-channel to communicate with one user within one time slot, expressed as:
[0072] 4) Sub-channel allocation constraint: At most one sub-channel is allocated to one UAV-user pair within one time slot, expressed as:
[0073] 5) UAV's secrecy constraint: where ε represents the security threshold.
[0074] Step 6: Model the UAV trajectory planning and resource allocation sub-problem as an MDP;
[0075] Model the UAV trajectory planning and resource allocation sub-problem as an MDP. Specifically: Under the constraints of secrecy, transmission power selection, sub-channel allocation, etc., with the goal of maximizing the system communication rate, split the UAV trajectory planning and resource allocation problem into the UAV trajectory and resource allocation optimization sub-problem, the UAV-J location deployment and power allocation sub-problem; and model them as MDPs respectively,
[0076] Model the UAV trajectory planning and resource allocation problem as an MDP. Its state, action, and reward functions are specifically expressed as follows:
[0077] (1) State: Let s t ={s 1,t ,…s m,t ,…s M,t} represent the system environment state of the transmitting UAV at time slot t, where s m,t represents the state of U m at time slot t, and s m,t includes the trajectory position and channel gain of the transmitting UAV, modeled as: where h m,k,t , represents the channel gain vector, modeled as: hm,k,t = [h m,k,1,t , h m,k,2,t ,..., h m,k,L,t ,
[0078] (2) Action: Let a t = {a 1,t , …, a m,t , … a M,t} represent the actions of the UAV sent in time slot t, where a m,t represents the action of U m at time slot t. To determine the flight trajectory, power allocation, and sub-channel allocation strategy of U m , a m,t is modeled as: Where f0 represents the flight action space of the UAV sent in each time slot and is modeled as where d = v max τ; The UAV can only land directly above the landing position, which is expressed as: U m 's position update formula is β m,k,t = [β m,k,1,t , β m,k,2,t ,..., β m,k,L,t represents the sub-channel allocation vector of U m at time slot t, and p m,k,t = [p m,k,1,t , p m,k,2,t ,..., p m,k,L,t represents the power allocation vector of U m at time slot t; The continuous power allocation variable is converted into discrete power levels using a discretization mechanism. Specifically, the transmission power of the UAV is evenly divided into Q + 1 levels, and let p q represent the power of the q-th level, which is modeled as Let δ m,k,l,t,q represent the power level selection variable of U m at time slot t, 0 ≤ q ≤ Q. If the transmission power of U m occupying sub-channel l to send a message to user k is p q , then δ m,k,l,t,q = 1, otherwise δ m,k,l,t,q = 0; p m,k,l,t can be modeled as Action a m,t can be rewritten as Where δ m,k,l,t = [δ m,k,l,t,0 , δ m,k,l,t,1 ,..., δ m,k,l,t,Q represents the δ of U mThe discrete power order selection vector; if U m The selected action does not satisfy the system concealment constraint, that is When, then set the transmission power of this action of the UAV to 0, keep the position hovering, and do not communicate with the user, that is: β m,k,t = 0, δ m,k,l,t = 0;
[0079] (3) Reward function: Let r(s m,t , a m,t ) represent the immediate reward obtained when the state of U m at time slot t is s m,t , select action a m,t and transfer to the next state s m,t+1 . It is modeled as where λ g , λ c , λ d are all positive numbers. If the UAV fails to return to its initial position within the system time, it will be punished by λ g ; if the distance between the transmitting UAVs is less than the minimum safety distance, it will be punished by λ c ; if the distance between the transmitting UAV and the interfering UAV is less than the minimum safety distance, it will be punished by λ d .
[0080] Step 7: Determine the flight trajectory and resource allocation strategy of the transmitting UAV based on the FMADQN algorithm;
[0081] Determine the flight trajectory and resource allocation strategy of the transmitting UAV based on the FMADQN algorithm. Specifically: Each transmitting UAV acts as a local client, and the high-altitude platform acts as a server to construct a federated learning framework; each transmitting UAV uses the DQN algorithm to determine the local strategy and uploads the local model parameters to the high-altitude platform; the high-altitude platform uses the average aggregation method to aggregate the model parameters uploaded by each client to update the global model parameters, and distributes the updated global model parameters to each client. The client performs iterative updates based on the received global model parameters until the algorithm converges; Let represent the long-term reward of U m . It is modeled as γ represents the reward discount factor, 0 ≤ γ ≤ 1. U m continuously interacts with the environment to obtain the strategy and thus makes the optimal decision to maximize the long-term reward; Specifically: Let θ m , respectively represent the prediction network and target network model parameters of U m . During local model training, U m first initializes the prediction network, target network, and creates an experience replay pool Time slot U m Select an action using the ε-greedy strategy, that is Let q denote a random number, where 0 ≤ q ≤ 1, and a r denotes the action randomly selected from the set of optional actions, and ε denotes the exploration rate; U m Interact with the environment according to the selected action to obtain an immediate reward r m,t , and transfer to the next state s m,t+1 , and store the quadruple (s m,t , a m,t , r m,t , s m,t+1 ) in the experience replay pool Randomly sample a small batch of samples from the experience replay pool to train the model. Let Q(s m,t , a m,t , θ m ) represent the predicted network Q value, and y(s m,t , a m,t ) represent the target network Q value, which is modeled as Let represent U m 's training loss function, and use the mean squared error (MSE) as the loss function. Then is modeled as Update the predicted network parameter θ using the gradient descent algorithm m , which is modeled as where α represents the learning rate, 0 ≤ α ≤ 1. After multiple rounds of iteration, assign the predicted network parameter θ m to the target network parameter Repeat the above steps until the algorithm converges to obtain the flight trajectory of the transmitting UAV and the resource allocation strategy; the transmitting UAV sends its updated parameter θ m to the high-altitude platform; let θ g represent the global model parameter, which is modeled as
[0082] Step 8: Model the interference UAV position deployment and power control sub-problem as an MDP;
[0083] Model the interference UAV position deployment and power control sub-problem as an MDP. Specifically, the modeled MDP can be represented as follows: (1) State: Let represent the state of the UAV-J agent at time slot t, which includes the position and channel gain of UAV-J, and is modeled as where represents the position of UAV-J at time slot t, represents the channel gain vector, and is modeled as (2) Action: Let represent the action of UAV-J in time slot t. To determine the deployment location and peak power control strategy of UAV-J, is modeled as represent the position deployment action space of agent UAV-J, which is modeled as where d = v max τ. The UAV-J position update formula is Using the discretization mechanism, the continuous peak power control variable of UAV-J is converted into discrete power levels. Let represent the power level selection variable of UAV-J in time slot t, 0 ≤ q ≤ Q. If the peak transmission power of UAV-J in time slot t is p q , then Otherwise The peak power control variable can be obtained as which is Action can be rewritten as where represents the discrete power level selection vector of UAV-J in time slot t; if the action selected by UAV-J does not satisfy the concealment constraint, the action is changed to f t J = [0, 0, 0], (3) Reward function: Let represent the immediate reward obtained when the state of UAV-J in time slot t is and the action is selected, which is modeled as where λ w 、λ s are positive numbers. If the distance between any transmitting UAV and UAV-J is less than the minimum safety distance, a penalty of λ w is imposed; if the user received SINR is lower than the minimum SINR threshold, a penalty of λ s is imposed.
[0084] Step 9: Determine the interference UAV position deployment and power control strategy based on the Deep Q-Network (DQN) algorithm;
[0085] Determine the interference UAV position deployment and power control strategy based on the DQN algorithm, specifically: Let represent the long-term reward of U m , which is modeled as UAV-J continuously interacts with the environment to obtain the policy and thus makes the optimal decision to maximize the long-term reward; specifically: Let θ J 、 respectively represent the prediction network and target network model parameters of UAV-J. During local model training, UAV-J first initializes the prediction network, target network, and creates an experience replay pool At time slot t, UAV-J adopts ε J greedy policy to select an action, that is q J represents a random number, 0 ≤ q J ≤ 1, represents the action randomly selected from the set of optional actions, and ε J represents the exploration rate. UAV-J interacts with the environment according to the selected action and obtains an immediate reward r t J and transfers to the next state and stores the quadruple in the experience replay pool Randomly sample a small batch of samples from the experience replay pool to train the model. Let represent the Q value of the prediction network, represent the Q value of the target network, and is modeled as Let represent the training loss function of UAV-J. The mean squared error (MSE) is used as the loss function, then is modeled as Use the gradient descent algorithm to update the prediction network parameter θ J and is modeled as where α J represents the learning rate, 0 ≤ α J ≤ 1. After multiple rounds of iteration, assign the prediction network parameter θ J to the target network parameter Repeat the above steps until the algorithm converges to obtain the interference UAV position deployment and power control strategy.
[0086] Step 10: Use the alternating iteration method to solve the problems of the sending UAV flight trajectory and resource allocation, and the interference UAV position deployment and power control sub-problems in turn until the algorithm converges;
[0087] Use the alternating iteration method to solve the problems of the sending UAV flight trajectory and resource allocation, and the interference UAV position deployment and power control sub-problems in turn until the algorithm converges. Specifically: Given the peak transmission power and position deployment strategy of the initial UAV-J, each sending UAV determines the flight trajectory, sub-channel, and power allocation strategy according to the FMADQN algorithm; Based on the flight trajectory, sub-channel, and power allocation strategy of the sending UAV, UAV-J determines the position deployment and peak transmission power according to the DQN algorithm; Repeat the above process until the algorithm converges.
[0088] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for trajectory planning and resource allocation of an unmanned aerial vehicle-assisted covert communication system, characterized in that, The method includes: S1: Model the scenario of an unmanned aerial vehicle (UAV)-assisted covert communication system; S2: Model the channel model; S3: Model the system communication rate; S4: Model the probability of the eavesdropper's error detection; S5: Model the UAV flight trajectory and resource allocation constraints; S6: Model the UAV trajectory planning and resource allocation sub-problem as a Markov decision process (MDP); S7: Determine the UAV flight trajectory and resource allocation strategy based on the federated multi-agent deep Q-network (FMADQN) algorithm; S8: Model the interfering UAV position deployment and power control sub-problem as an MDP; S9: Determine the interfering UAV position deployment and power control strategy based on the deep Q-network (DQN) algorithm; S10: Use the alternating iteration method to solve the UAV flight trajectory and resource allocation, and the interfering UAV position deployment and power control sub-problems in sequence until the algorithm converges.
2. The trajectory planning and resource allocation method for the UAV-assisted covert communication system according to claim 1, wherein, In step S1, modeling the scenario of the UAV-assisted covert communication system specifically includes: M transmitting UAVs, 1 friendly interfering UAV (UAV-J), K users, N eavesdroppers, and a high-altitude platform equipped with a parameter server; the transmitting UAVs start from the initial positions, fly over the area along a certain trajectory within the system time, send messages to the legitimate users within the covered area during the flight, and then return to the starting point; the ground eavesdroppers detect the communication between the transmitting UAVs and the ground users, and the UAV-J generates artificial noise to hinder the eavesdroppers' detection of the communication signals. Let U m represent the m-th transmitting UAV, where 1 ≤ m ≤ M; Let q k =(x k , y k , 0) represent the location of the k-th user, where 1 ≤ k ≤ K; represent the location coordinates of the n-th eavesdropper, where 1 ≤ n ≤ N. Let q J =(x J , y J , H) represent the location of UAV-J, where H represents the altitude of the UAV. Divide the system time into multiple time slots of equal size. Let τ represent the length of each time slot, and the total number of time slots is T. Let represent the flight trajectory of U m , where represents the location of U m at time slot t, where 0 ≤ t ≤ T. Assume that each UAV fully shares the system spectrum resources, and an orthogonal frequency division multiple access mechanism is adopted inside each UAV to communicate with users. Divide the spectrum into L orthogonal sub-channels of equal length. Let B represent the sub-channel bandwidth.
3. The method for trajectory planning and resource allocation of the UAV-assisted covert communication system according to claim 2, wherein In step S2, a channel model is built, specifically including: modeling the air-ground link from the transmitting UAV and UAV-J to the user as a probabilistic line-of-sight channel model, and letting L m,k,l,t represent the average transmission loss (in dB) of the link corresponding to U m occupying sub-channel l to send a message to user k at time slot t, which is modeled as where and are the probabilities that the transmission link from U m to user k is a line-of-sight link and a non-line-of-sight link at time slot t, respectively, and satisfy which is modeled as where a and b are coefficients related to the environment, and θ m,k,t represents the elevation angle between user k and U m at time slot t, which is modeled as and represent the path losses of the LoS and NLoS links when U m occupies sub-channel l to send a message to user k at time slot t, respectively, and are modeled as where f l is the carrier frequency of sub-channel l, are the average additional losses of the LoS and NLoS links, respectively, 1 ≤ l ≤ L; represents the distance between U m and user k at time slot t, which is modeled as Let h m,k,l,t represent the channel gain when U m occupies sub-channel l to communicate with user k at time slot t, which is modeled as Let represent the channel gain between U m and the nth eavesdropper on sub-channel l at time slot t. Let represent the channel gain between UAV-J and user k on sub-channel l at time slot t, represent the channel gain between UAV-J and the nth eavesdropper on sub-channel l, and is modeled as All are probabilistic line-of-sight channels.
4. The method for trajectory planning and resource allocation of the UAV-assisted covert communication system according to claim 3, wherein, In step S3, the communication rate of the modeling system specifically includes: Let γ m,k,l,t represent the signal-to-interference-plus-noise ratio (SINR) of the received signal when U m occupies sub-channel l to send a message to user k in time slot t, and is modeled as where β m,k,l,t represents the sub-channel allocation variable of user k. If U m occupies sub-channel l to communicate with user k in time slot t, then β m,k,l,t = 1; otherwise, β m,k,l,t = 0. p m,k,l,t represents the transmission power when U m occupies sub-channel l to send a message to user k in time slot t, represents the transmission power of UAV-J in time slot t, is a uniformly distributed random variable, that is where is the maximum AN peak transmission power adjustable by UAV-J, and σ 2 represents the link noise power. Let R m,t represent the communication rate of U m in time slot t, and is modeled as 5. The trajectory planning and resource allocation method for the UAV-assisted covert communication system according to claim 4, wherein In step S4, modeling the probability of eavesdropper error detection specifically includes: Let s m represent the signal sent by the transmitting UAV U m to the users within its coverage area. Let s J represent the signal sent by UAV-J to the ground eavesdropper. E[|s m | 2 = 1 and E[|s J | 2 = 1. Let y n,l,t represent the received signal of the nth eavesdropper on subchannel l at time slot t, which is modeled as where z n,l,t represents additive white Gaussian noise with a mean of 0 and a variance of σ 2 ; represents the null hypothesis, that is, the UAV does not send a message to the user, represents the alternative hypothesis, that is, the UAV sends a message to the user. Assuming that the eavesdropper performs signal detection based on the received signal energy, let T n,t represent the received power of the nth eavesdropper at time slot t, which is modeled as and represent the corresponding decisions made by the nth eavesdropper for the hypotheses and respectively. represents the detection threshold of the nth eavesdropper. Assuming L → ∞, then T n,t can be rewritten as Let represent the probability of error detection of the nth eavesdropper at time slot t, which is modeled as where and represent the false alarm probability and the missed alarm probability of the nth eavesdropper at time slot t respectively. is modeled as is modeled as 6. The method for trajectory planning and resource allocation of the UAV-assisted covert communication system according to claim 5, wherein, In step S5, modeling the UAV flight trajectory and resource allocation constraints specifically includes: (1) Model the UAV flight trajectory constraints, including: 1) Maximum flight distance constraint for each time slot of the UAV: v max is the maximum flight speed of the UAV; (2) The UAV safety distance constraint: Among them, d m,m′,t represents the distance between the U in the t time slot m and U m′ ; is the distance between the UAV U sent in the t time slot m and UAV-J, and d min is the minimum safe distance between UAVs; 3) UAV starting position constraint: (2) Model the resource allocation constraints, including: 1) UAV transmission power constraint: where p max is the maximum transmission power of the UAV; 2) The SINR received by the user needs to satisfy the minimum SINR constraint: γ m,k,l,t ≥β m,k,l,t γ th ; where γ th represents the minimum SINR threshold value of the user; 3) UAV - user association constraint: Each user can occupy at most one sub - channel to communicate with one UAV within a time slot, which is expressed as: Each UAV can occupy at most one sub - channel to communicate with one user within a time slot, which is expressed as: 4) Sub-channel allocation constraint: At most one sub-channel is allocated to a UAV-user pair in one time slot, which is expressed as: 5) Concealment constraint of the UAV: ζ n,t ≥ 1 - ε, where ε represents the safety threshold value.
7. The trajectory planning and resource allocation method for the UAV-assisted covert communication system according to claim 6, characterized in that In step S6, under the constraints of covertness, transmit power selection, sub-channel allocation, etc., with the goal of maximizing the system communication rate, split the UAV trajectory planning and resource allocation problem into the UAV trajectory and resource allocation optimization sub-problem, and the UAV-J position deployment and power allocation sub-problem; and model MDPs respectively. Model the UAV trajectory planning and resource allocation problem as an MDP, and its state, action, and reward functions are specifically expressed as follows: (1) State: Let s t = {s 1,t , … s m,t , … s M,t} represent the system environment state of the transmitting UAV in time slot t, where s m,t represents the state of U m at time slot t. s m,t includes the trajectory position of the transmitting UAV and the channel gain, modeled as: where h m,k,t , represents the channel gain vector, modeled as: h m,k,t = [h m,k,1,t , h m,k,2,t , …, h m,k,L,t , (2) Action: Let a t ={a 1,t ,…,a m,t ,…a M,t} represent the actions of the UAV sent in time slot t, where a m,t represents the action of U m in time slot t. To determine the flight trajectory, power allocation, and sub-channel allocation strategy of U m , a m,t is modeled as: where f0 represents the flight action space of the UAV sent in each time slot and is modeled as where d = v max τ; the UAV can only land directly above the landing position, which is expressed as: U m 's position update formula is represents the sub-channel allocation vector of U m in time slot t, and p m,k,t =[p m,k,1,t ,p m,k,2,t ,…,p m,k,L,t represents the power allocation vector of U m in time slot t; the continuous power allocation variable is converted into discrete power levels using a discretization mechanism. Specifically, the transmission power of the UAV is evenly divided into Q + 1 levels, and let p q represent the power of the q-th level, which is modeled as Let δ m,k,l,t,q represent the power level selection variable of U m in time slot t, 0 ≤ q ≤ Q. If the transmission power of U m occupying sub-channel l to send a message to user k is p q , then δ m,k,l,t,q = 1, otherwise δ m,k,l,t,q = 0; p m,k,l,t can be modeled as Action a m,t can be rewritten as where δ m,k,l,t =[δ m,k,l,t,0 ,δ m,k,l,t,1 ,...,δ m,k,l,t,Q represents the discrete power level selection vector of U m in time slot t; if the selected action of U m does not meet the system concealment constraint, that is , then set the action to the transmission power of the UAV being 0, the position remains hovering, and it does not communicate with the user, that is: β m,k,t = 0, δ m,k,l,t = 0; (3) Reward function: Let \(r(s m,t , a m,t )\) denote the immediate reward obtained when the state of U m at time slot \(t\) is \(s m,t \) and the action \(a m,t \) is selected to transfer to the next state \(s m,t+1 \). It is modeled as where \(\lambda g \), \(\lambda c \), and \(\lambda d \) are all positive numbers. If the UAV fails to return to its initial position within the system time, it will be penalized by \(\lambda g \); if the distance between the transmitting UAVs is less than the minimum safe distance, it will be penalized by \(\lambda c \); if the distance between the transmitting UAV and the interfering UAV is less than the minimum safe distance, it will be penalized by \(\lambda d \).
8. The method for trajectory planning and resource allocation of the unmanned aerial vehicle-assisted covert communication system according to claim 7, characterized in that In step S7, the flight trajectory and resource allocation strategy of the transmitting UAVs are determined based on the FMADQN algorithm; specifically: each transmitting UAV serves as a local client, and the high-altitude platform serves as a server to construct a federated learning framework; each transmitting UAV uses the DQN algorithm to determine the local strategy and uploads the local model parameters to the high-altitude platform; the high-altitude platform uses the average aggregation method to aggregate the model parameters uploaded by each client to update the global model parameters, and distributes the updated global model parameters to each client, and the client performs iterative updates based on the received global model parameters until the algorithm converges; let denote the long-term reward of U m , modeled as where γ represents the reward discount factor, 0 ≤ γ ≤ 1, and U m obtains the strategy π by continuously interacting with the environment: thus making the optimal decision to maximize the long-term reward; specifically: let θ m , respectively denote the prediction network and target network model parameters of U m . During local model training, U m first initializes the prediction network, target network, and creates an experience replay pool At time slot t, U m adopts an ε-greedy strategy to select an action, that is where q represents a random number, 0 ≤ q ≤ 1, a r denotes the action randomly selected from the set of optional actions, and ε represents the exploration rate; U m interacts with the environment according to the selected action to obtain the immediate reward r m,t , transfers to the next state s m,t+1 , and stores the quadruple (s m,t , a m,t , r m,t , s m,t+1 ) in the experience replay pool Randomly samples a small batch of samples from the experience replay pool to train the model. Let Q(s m,t , a m,t , θ m ) represent the Q value of the prediction network, and y(s m,t , a m,t ) represent the Q value of the target network, modeled as Let denote the training loss function of U m . The mean squared error (MSE) is used as the loss function, then is modeled as The gradient descent algorithm is used to update the prediction network parameter θ m , modeled as where α represents the learning rate, 0 ≤ α ≤ 1, and after multiple rounds of iteration, the predicted network parameters θ m are assigned to the target network parameters Repeat the above steps until the algorithm converges to obtain the flight trajectory of the transmitting UAV and the resource allocation strategy; the transmitting UAV sends its updated parameters θ m to the high-altitude platform; let θ g represent the global model parameters, modeled as 9. The trajectory planning and resource allocation method for the UAV-assisted covert communication system according to claim 8, wherein, In step S8, the sub-problem of jamming UAV position deployment and power control is modeled as an MDP. Specifically: The modeled MDP can be expressed as follows: (1) State: Let represent the state of the UAV-J at time slot t, which includes the position and channel gain of UAV-J, modeled as where represents the position of UAV-J at time slot t, and represents the channel gain vector, modeled as (2) Action: Let represent the action of UAV-J at time slot t. To determine the deployment position and peak power control strategy of UAV-J, is modeled as where d = v max τ. The UAV-J position update formula is Using the discretization mechanism, the continuous peak power control variable of UAV-J is converted into discrete power levels. Let represent the power level selection variable of UAV-J at time slot t, 0 ≤ q ≤ Q. If the peak transmit power of UAV-J at time slot t is p q , then Otherwise The peak power control variable is The action can be rewritten as where represents the discrete power level selection vector of UAV-J at time slot t; if the action selected by UAV-J does not satisfy the concealment constraint, the action is changed to (3) Reward function: Let represent the immediate reward obtained when the state of UAV-J at time slot t is and the action is selected, modeled as where λ w , λ s are positive numbers. If the distance between any transmitting UAV and UAV-J is less than the minimum safety distance, a penalty of λ w is incurred; if the user received SINR is lower than the minimum SINR threshold, a penalty of λ s is incurred.
10. The method for trajectory planning and resource allocation of the unmanned aerial vehicle-assisted covert communication system according to claim 9, wherein In step S9, the interference UAV position deployment and power control strategy is determined based on DQN; specifically: Let Table U represents U m The long-term reward, modeled as UAV-J continuously interacts with the environment to obtain the policy π: Thus, it makes an optimal decision to maximize the long-term reward; specifically: Let θ J , respectively represent the prediction network and target network model parameters of UAV-J. During local model training, UAV-J first initializes the prediction network, target network, and creates an experience replay pool At time slot t, UAV-J adopts an ε J greedy strategy to select an action, that is q J represents a random number, 0 ≤ q J ≤ 1, represents the action randomly selected from the set of optional actions, ε J represents the exploration rate. UAV-J interacts with the environment according to the selected action to obtain an immediate reward transfers to the next state and stores the quadruple in the experience replay pool Randomly sample a small batch of samples from the experience replay pool to train the model. Let represent the Q value of the prediction network, represent the Q value of the target network, modeled as Let represent the training loss function of UAV-J. The mean squared error (MSE) is used as the loss function, then is modeled as Adopt the gradient descent algorithm to update the prediction network parameter θ J , modeled as where α J represents the learning rate, 0 ≤ α J ≤ 1. After multiple rounds of iteration, assign the prediction network parameter θ J to the target network parameter Repeat the above steps until the algorithm converges to obtain the interference UAV position deployment and power control strategy.
11. The method for trajectory planning and resource allocation of the drone-assisted covert communication system according to claim 10, characterized in that, In step S10, use the alternating iteration method to solve the UAV trajectory planning and resource allocation sub-problem, and the interfering UAV position deployment and power control sub-problem in sequence until the algorithm converges; specifically including: given the peak transmit power and position deployment strategy of the initial UAV-J, each transmitting UAV determines the flight trajectory, sub-channel, and power allocation strategy according to the FMADQN algorithm; based on the flight trajectory, sub-channel, and power allocation strategy of the transmitting UAVs, the UAV-J determines the position deployment and peak transmit power according to the DQN algorithm; repeat the above process until the algorithm converges.
Citation Information
Cited By
Communication scheduling method and system
CN121442402A
Communication scheduling method and system
CN121442402B