Information age optimization method and system for UAV-assisted uplink covert transmission system

By building a drone-assisted uplink covert transmission system model, optimizing the drone position and the transmission power of IoT devices, the problem of information age optimization in the IoT is solved, more efficient covert and timely transmission is achieved, and the drone's flight time is extended.

CN119603709BActive Publication Date: 2025-09-30CHONGQING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411788580.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-09-30
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

The existing drone-assisted uplink covert transmission system fails to effectively optimize the information age in the Internet of Things, resulting in the security and timeliness of information transmission being limited. In particular, it is difficult to take into account the concealment performance of ground equipment under the nature of information broadcasting.

Method used

A UAV-assisted uplink covert transmission system model is constructed. By optimizing the UAV position and the transmission power of IoT devices, the Markov decision process and neural network are combined to optimize the parameters to minimize the average peak information age of IoT devices while meeting the constraints of concealment and communication quality.

Benefits of technology

It effectively reduces the information age of IoT devices, improves the concealment and timeliness of transmission, optimizes the energy efficiency of drones, and extends flight time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119603709B_ABST
    Figure CN119603709B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of Internet of Things (IoT), and specifically discloses an information age optimization method and system for a UAV-assisted uplink covert transmission system, which constructs a UAV-assisted uplink covert transmission system, including a covert IoT device (CIoTD), multiple public IoT devices (PIoTDs), an unmanned aerial vehicle (UAV) and a listener (Willie). The CIoTD and M PIoTDs perform uplink transmission to the UAV, and Willie attempts to detect the covert transmission of the CIoTD. The mobility of the UAV and the noise generated by the PIoTD further increase the concealment requirements of the CIoTD. The present invention also aims to minimize the average peak age of information (PAoI) of all PIoTDs (PIoTDs). While meeting the concealment requirements, an optimization problem is constructed, and the optimization problem is further solved based on a Markov decision process to obtain the position of the UAV and the transmission power of the CIoTD. Simulation results demonstrate the effectiveness and excellence of the present invention in minimizing the average PAoI of the PIoTD.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Internet of Things (IoT), and in particular to an information age optimization method and system for a drone-assisted uplink covert transmission system. Background Art

[0002] Current UAV-assisted covert uplink transmission systems utilize UAVs as relay nodes to enhance and improve the concealment of data transmission. By optimizing the flight trajectory and communication resource allocation of UAVs, this system can achieve optimal energy efficiency while ensuring the concealment of UAV uplink data transmission, significantly improving the flight time of energy-constrained small UAVs performing secure data transmission operations. For example, the UAV covert communication system trajectory optimization and communication resource allocation method and system disclosed in Publication No. CN115802494A. However, this scheme treats all ground devices as active devices, aiming to achieve better energy efficiency for all devices. Its covert communication model based on signal error detection rate must take all ground devices into account, resulting in limited concealment performance for ground devices. Furthermore, with the rapid development of the Internet of Things (IoT), the timeliness of information (characterized by information age) has become particularly important. However, the broadcast nature of IoT data transmission in UAV-assisted covert uplink transmission systems poses a significant challenge to ensuring the security of information transmission. Currently, there is no information age optimization solution for UAV-assisted covert uplink transmission systems. Summary of the Invention

[0003] The present invention provides an information age optimization method and system for a UAV-assisted uplink covert transmission system, and solves the technical problems of: how to provide a UAV-assisted uplink covert transmission system with better concealment performance, and how to effectively reduce the information age of the UAV-assisted uplink covert transmission system.

[0004] To solve the above technical problems, the present invention provides an information age optimization method for a UAV-assisted uplink covert transmission system, comprising the steps of:

[0005] S1. Construct a system model of UAV-assisted uplink covert transmission system;

[0006] The system model includes a covert IoT device (CIoTD), a listener (Willie), M public IoT devices (PIoTD), and an unmanned aerial vehicle (UAV). The CIoTD and the M PIoTDs perform uplink transmission to the UAV, and Willie attempts to detect the covert transmission of the CIoTD.

[0007] S2. Determine optimization parameters, optimization objectives, and constraints based on the system model to construct an optimization problem;

[0008] The optimization parameters are determined to be the UAV position and the CIoTD transmission power. The optimization goal is to minimize the average peak information age of M PIoTDs. The constraints include Willie's concealment constraint, the first transmission power constraint of the CIoTD, the flight altitude constraint of the UAV, the communication quality constraint, the data transmission quality constraint, and the final decoding constraint of the CIoTD.

[0009] S3. Solve the optimization problem to obtain solution values ​​of optimization parameters;

[0010] S4. Run the UAV-assisted uplink covert transmission system with the solution value of the optimization parameter.

[0011] Furthermore, during data transmission by CIoTD, each PIoTD generates N data packets with a time interval of Δt, and the average peak information age of M PIoTDs is the sum of the peak information ages of these MN data packets divided by M.

[0012] Furthermore, when Δτ mn When Δt is less than or equal to, the peak information age a of the nth data packet of the mth PIoTD mn Equal to Δτ mn , Δτ mn is the time to transmit the nth data packet of the mth PIoTD; when Δτ mn When a is greater than Δt, mn equal and Δτ mn The sum of is the waiting time from the generation to transmission of the nth data packet of the mth PIoTD.

[0013] Furthermore, the concealment constraint of Willie is expressed as Willie's total error rate ξ is greater than or equal to 1-δ, where δ is the concealment requirement; Willie's total error rate ξ is equal to the error response probability P FA and the probability of missed detection P MD The sum of the error response probability P FA Defined as the probability that CIoTD does not send information but Willie believes that CIoTD has sent it, the missed detection probability P MA It is defined as the probability that CIoTD has sent information but Willie believes that CIoTD has not.

[0014] Furthermore, assuming that Willie knows the transmission power of CIoTD and PIoTD, the channel gain from CIoTD to Willie, and Willie's noise distribution, Willie uses a binary hypothesis test H0, H1 based on the observed signal and the prior probabilities D0 and D1 of H0, H1 to calculate the probability of wrong response P FA=P(D1|H0) and the probability of missed detection P MA =P(D0|H1), where H0 represents the case where Willie determines based on the observed signal that CIoTD does not send any information to the UAV, while H1 represents the case where Willie determines based on the observed signal that CIoTD does send information. The corresponding relationships are:

[0015]

[0016] Among them, y w [l] represents the signal observed by Willie; n w [l] is the additive Gaussian white noise at Willie's location, with a mean of zero and a variance of x c [[l] and is a complex Gaussian signal, and represents a complex Gaussian distribution with a mean of 0 and a variance of 1; h cw 、 Represent the channel coefficients between CIoTD, the mth PIoTD and Willie, is the transmission power of the mth PIoTD, P c is the transmit power of CIoTD.

[0017] Furthermore, the first transmission power constraint of CIoTD is specifically the transmission power P of CIoTD. c Not less than the maximum transmit power P of CIoTD cmax ; The UAV flight altitude constraint is specifically that the UAV flight altitude is within the minimum altitude and the maximum height The communication quality constraint is γ j ≥γ th , γ j is the signal-to-interference-noise ratio of CIoTD and PIoTD, γ th is the signal-to-interference-noise ratio threshold required for non-interrupted communication; the data transmission quality constraint is specifically S c ≥S th , S c is the data transmission volume of CIoTD, S th is the data transmission volume threshold of CIoTD; the final decoding constraint of CIoTD is m=1,...,M,h cu 、 They represent the channel coefficients between CIoTD, the mth PIoTD and UAV respectively.

[0018] Furthermore, the signal-to-interference-noise ratio of CIoTD and PIoTD j=c,p m ,J=M+1 is the total number of IoT devices, P j |h ju | 2 is the part of the signal to be received, To affect the noise portion of the signal to be received, is the Gaussian white noise power of CIoTD and the mth PIoTD frequency band, j is the current device, i is all devices with worse channel quality than the current device; the data transmission volume of CIoTD Indicates the kth segment of data that needs to be transmitted in the nth time interval; h cu 、 equal j=c,p m ,ρ0 is the power gain at a reference distance of 1m, d ju represents the distance from CIoTD and the mth PIoTD to the UAV, represents the line-of-sight component of the air-to-ground channel, and α l is the Earth-to-space path loss exponent.

[0019] Furthermore, the step S3 specifically includes the steps of:

[0020] Transform the hidden constraint ξ≥1-δ into P c The second power constraint: represents the variance of the additive white Gaussian noise at Willie, x * represents the root of f(x)=-ln(1-x)-x, x must satisfy x≤x * , x is defined as

[0021] The UAV is considered as an agent that learns to interact with the environment, and the optimization problem is modeled as a Markov decision process. The state space of the Markov decision process is defined as S = {q u}, qu represents the rectangular three-dimensional coordinates of the UAV, the action space is defined as A = {d}, d represents the displacement of the UAV in each time step, and then the reward function r for moving the time step t t It is defined as the inverse of the average peak information age of M PIoTDs multiplied by a positive constant λ;

[0022] Based on the defined Markov decision process and the constraints of the optimization problem, an information age optimization network is constructed and trained. The network contains parameters θ oldThe old network, the actor network with parameters θ and the critic network with parameters φ, collects the state transition trajectory of T time steps, and the agent updates the parameters θ and φ. The old network only copies the parameters θ from the actor network for policy preservation and does not participate in training; outputs the current UAV position q u and the transmission power of CIoTD as the solution to the optimization problem.

[0023] Furthermore, during the training process, the agent's specific operation steps include:

[0024] F1, enter the location of CIoTD q c and transmit data S th , Willie's position q w , the channel uses L, the concealment requirement is δ, the location of PIoTD PIoTD transmit power The size of a single packet per PIoTD n , interval Δt;

[0025] F2, initialize the environment and state, and initialize the information age optimization network;

[0026] F3, train the information age optimization network;

[0027] In each round, the steps are performed:

[0028] F31, reset the initial state;

[0029] F32. In each time step, perform the following steps:

[0030] F321, UAV according to the current state s t Get the current position q u ;

[0031] F322, according to the current strategy π θ Select action a t ;

[0032] F323, in different transmission periods, make P c Satisfy the first power constraint and the second power constraint, and the CIoTD final decoding constraint, so that P c =min{P cv ,P cmax ,P ch}, P cv P calculated corresponding to the second power constraint c The maximum value that can be achieved, P ch P represents the final decoding constraint calculated by CIoTD c The maximum value that can be achieved; calculate the average PAoI and get the r of the current time stept ;

[0033] F324, execute a t , get the next state s t+1 ;

[0034] F325, experience t ,a t ,r t >Store to experience replay buffer;

[0035] F326, determine whether the time step T is reached, if not, return to step F321 to enter the next time step, if yes, enter step F327;

[0036] F327, according to the output V(s of the critic network t )calculate And copy the actor network's θ to the old network's θ old In the network, to maintain the old strategy;

[0037] F328. Update the parameters of the actor network and the critic network based on the experience data in the experience replay buffer;

[0038] F329, clear the experience replay buffer;

[0039] F330, return to the UAV's position q u , the transmission power of CIoTD and the average peak information age of M PIoTDs.

[0040] The present invention also provides an information age optimization system for a UAV-assisted uplink covert transmission system, the key of which is that it is provided with an intelligent agent, which is used to implement steps S1 to S4 in the information age optimization method of the UAV-assisted uplink covert transmission system.

[0041] ​The present invention provides an information age optimization method and system for a UAV-assisted uplink covert transmission system, which constructs a UAV-assisted uplink covert transmission system, including a covert Internet of Things device (CIoTD), multiple public Internet of Things devices (PIoTDs), a UAV (UAV) and a listener (Willie). The CIoTD and M PIoTDs perform uplink transmission to the UAV, and Willie attempts to detect the covert transmission of the CIoTD. The mobility of the UAV and the noise generated by the PIoTD further increase the concealment requirements of the CIoTD. The present invention also aims to minimize the average peak age of information (PAoI) of all PIoTDs (PIoTDs). While meeting the concealment requirements, an optimization problem is constructed, and the optimization problem is further solved based on the Markov decision process to obtain the position of the UAV and the transmission power of the CIoTD. The simulation results demonstrate the effectiveness and excellence of the present invention in minimizing the average PAoI of the PIoTD. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 This is a system model diagram of the UAV-assisted uplink covert transmission system provided by an embodiment of the present invention;

[0043] Figure 2 1 is a diagram showing the calculation principle of the peak information age of PIoTD provided by an embodiment of the present invention;

[0044] Figure 3 This is a diagram showing the calculation principle of the data transmission rate of CIoTD provided by an embodiment of the present invention;

[0045] Figure 4 is a comparison diagram of the training process of four algorithms provided by the embodiments of the present invention;

[0046] Figure 5 This is a graph showing the change in average PAoI values ​​of the four algorithms provided by the embodiments of the present invention over the number of training rounds;

[0047] Figure 6 FIG. 4 is a graph showing how the average PAoI of the six solutions provided in the embodiments of the present invention changes with the channel usage L. DETAILED DESCRIPTION

[0048] The following describes the embodiments of the present invention in detail with reference to the accompanying drawings. The embodiments are provided for illustrative purposes only and are not to be construed as limiting the present invention. The accompanying drawings are provided for reference and illustration only and do not constitute a limitation on the scope of protection of the present invention. Many changes may be made to the present invention without departing from the spirit and scope of the present invention.

[0049] The information age optimization method of the UAV-assisted uplink covert transmission system provided by the embodiment of the present invention is as follows: Figure 1 As shown, the steps include:

[0050] S1. Construct a system model of UAV-assisted uplink covert transmission system;

[0051] S2. Determine optimization parameters, optimization objectives, and constraints based on the system model to construct an optimization problem;

[0052] S3. Solve the optimization problem to obtain solution values ​​of optimization parameters;

[0053] S4. Run the UAV-assisted uplink covert transmission system with the solution value of the optimization parameter.

[0054] The embodiment of the present invention first constructs a UAV-assisted uplink covert transmission system, and its system model is as follows: Figure 1 As shown in the figure, there are a covert IoT device (CIoTD), a listener (Willie), M public IoT devices (PIoTDs), and an unmanned aerial vehicle (UAV). The CIoTD and M PIoTDs (PIoTDs) perform uplink transmission to the UAV, while Willie attempts to detect the covert transmission of the CIoTD.

[0055] Assume that the UAV has prior knowledge of the locations of all ground IoT devices (CIoTD and M PIoTDs) (i.e., it knows the specific locations of PIoTD and CIoTD). PIoTD and CIoTD use NOMA (non-orthogonal multiple access) for transmission, and the system bandwidth is B. The transmission from PIoTD can be regarded as noise, and CIoTD uses this noise as interference to send data packets to avoid being detected by Willie. In the data transmission of CIoTD, PIoTD generates data packets with a fixed time interval Δt until all CIoTD data packets are transmitted. UAV needs to search for a hovering position in the air to collect data from ground IoT devices. The rectangular three-dimensional coordinates of UAV, Willie, CIoTD and PIoTD are expressed as and m=1,...,M.

[0056] When the UAV altitude exceeds 120 meters, the line-of-sight (LoS) component of the channel between the UAV and the ground IoT device is significantly larger than the non-line-of-sight (NLoS) component. Therefore, the channel coefficient between the UAV and the ground IoT device is:

[0057]

[0058] Among them, h cu 、 denote the channel coefficients between CIoTD, the mth PIoTD and UAV, ρ0 is the power gain at a reference distance of 1m, and dju represents the distance from CIoTD and the mth PIoTD to the UAV, represents the line-of-sight component of the air-to-ground channel, and α l is the Earth-to-space path loss exponent.

[0059] Assume that the ground channel coefficient experiences quasi-static Rayleigh fading. Therefore, the channel coefficient between the ground IoT device and Willie is:

[0060]

[0061] Among them, h cw 、 They represent the channel coefficients between CIoTD, the mth PIoTD and Willie, d jw represents the distance from CIoTD, the mth PIoTD to Willie, represents the non-line-of-sight component of the channel between the ground IoT device and Willie, and α n is the ground path loss exponent.

[0062] Each PIoTD generates data packets with a time interval of Δt, t n is the time when the nth data packet is generated, ξ n The time when the transmission of the nth packet is completed is used to represent the timeliness of PIoTD. Figure 2 The calculation principle diagram of the peak information age of PIoTD. Specifically, there are two cases of the peak information age of PIoTD, as follows: Figure 2 As shown in (a) and (b), the peak information age in the two cases is calculated as follows:

[0063]

[0064] a mn represents the peak information age of the nth packet of the mth PIoTD, is the waiting time from the generation to transmission of the nth data packet of the mth PIoTD, Δτ mn is the time to transmit the nth data packet of the mth PIoTD.

[0065] With NOMA technology, the decoding order of the uplink starts from the better channel quality. Therefore, the reachability of IoT devices in NOMA can be expressed as:

[0066]

[0067] Where J is the total number of CIoTD and PIoTD (J=M+1), P c and are the transmission powers of CIoTD and the mth PIoTD respectively, is the Gaussian white noise power of the CIoTD and m-th PIoTD frequency bands, j is the current device, i is all devices with worse channel quality than the current device, and all i devices can be treated as noise interference and added to the noise formula. In order to enhance coverage, CIoTD is limited to the device with the worst channel quality, that is,

[0068] Figure 3 The calculation principle of CIoTD data transmission volume is shown. Since the transmission status of PIoTD is different at different times, which affects the transmission rate of each IoT device, it is necessary to calculate the data transmission of CIoTD and PIoTD in sections. The data transmission volume of CIoTD can be expressed as:

[0069] Before CIoTD transmits all data packets, a total of N intervals are required, and the PAoI of N intervals needs to be calculated. Therefore, the data transmission volume of CIoTD can be expressed as:

[0070]

[0071] Among them, s nk Indicates the kth segment of data that needs to be transmitted in the nth time interval.

[0072] When CIoTD is transmitting information, Willie attempts to detect these signals. Consider the worst case scenario, assuming that Willie knows the transmission power of CIoTD and PIoTD, the channel gain from CIoTD to Willie, and Willie's noise distribution. In this case, Willie will use a binary hypothesis test based on the observed signal:

[0073]

[0074] Wherein, l=1, 2, ..., L represents the index used by L channels. w [l] is the additive Gaussian white noise at Willie's location, with a mean of zero and a variance of x c [l] and x pm [l] is a complex Gaussian signal, and represents a complex Gaussian distribution with mean 0 and variance 1. w[l] represents the signal observed by Willie. H0 means that CIoTD did not send any information to the UAV, while H1 means that CIoTD did send information.

[0075] Assume that Willie's prior probabilities D0 and D1 of H0 and H1 are equal. If CIoTD wants to communicate covertly, this requires Willie to make two types of errors to ensure confidentiality. The first error is when CIoTD does not send a message, but Willie thinks she has sent it. The probability of this error is defined as the probability of wrong response P FA =P(D1|H0). The second error is when CIoTD sends a message, but Willie thinks she does not. The probability of this error is defined as the probability of missed detection P MA =P(D0|H1). Then Willie's total error rate can be expressed as ξ=P FA +P MD In order to minimize the total error rate, Willie will use the optimal test, the likelihood ratio test (LRT), which is:

[0076]

[0077] Among them, Ρ0 and Ρ1 represent the likelihood functions under H0 and H1 respectively. The sender's behavior is judged to be D1, that is, the sender is sending a message. When , the sender behavior is judged to be D0, that is, the sender is not sending a message. When , make any judgment. Formula (6) as a whole is making a decision to determine whether to send data. Function f(y w [l]|H0) and f(y w [l]|H1) is expressed as:

[0078]

[0079] It means that the mean is 0 and the variance is The complex Gaussian distribution of It means that the signal interference and Gaussian white noise generated by the information sent by M public devices are regarded as noise at the same time. The mean is 0 and the variance is The complex Gaussian distribution of It means that the interference generated by the information sent by the CIoTD hidden device, the signal interference generated by the information sent by M public devices, and Gaussian white noise are all regarded as noise.

[0080] During the operation of the UAV-assisted uplink covert transmission system, the optimization goal is to minimize the average PAoI of the PIoTD while meeting the data transmission requirements of the CIoTD. The CIoTD transmission power and the UAV's position affect the channel quality, which in turn affects the PIoTD transmission time. Therefore, the CIoTD transmission power and the UAV's position are used as optimization parameters. Based on the previous description, this optimization problem can be expressed as:

[0081]

[0082] Among them, δ is the concealment requirement, P cmax is the maximum transmission power of CIoTD, γ th is the signal-to-interference-and-noise ratio threshold required for non-interrupted communication, They represent the minimum and maximum altitudes of the UAV, respectively, and γj is the signal-to-interference-noise ratio of CIoTD and PIoTD, which are: j=c,p m ,P j h ju 2 is the part of the signal to be received, is the noise part that affects the received signal. th The size of the data packet to be transmitted can be customized. c1 is the stealth constraint, which ensures that Willie cannot detect the CIoTD. c2 sets the maximum transmit power limit for the CIoTD. c3 limits the UAV's flight altitude. The constraints in c4 ensure that communication is not interrupted. c5 ensures that all data required by the CIoTD is fully transmitted. c6 ensures that the CIoTD is ultimately decoded.

[0083] Solving problem P1 is a significant challenge due to the non-convexity and high dimensionality of the optimization problem. Constraint c1 is non-convex and needs to be handled in order to solve the optimization problem. KL divergence can be used as the minimum total error rate ξ * The lower bound of is expressed as:

[0084]

[0085] D(P0||P1) is the formula for calculating the divergence. By inputting P0 and P1, we can get the upper bound of the probability of being discovered by Willie.

[0086] Based on c1 and the above formula, the implicit constraint c1 is transformed into:

[0087] D(Ρ0||Ρ1)≤2δ 2 ,[[[[[[[[[[[[[[[[[[[[[[(11)

[0088] The KL divergence can be expressed as:

[0089] D(Ρ0||Ρ1)=L(-ln(1-x)-x),[[[[[[[[[[[[[[[(12)

[0090] in,

[0091]

[0092] Let f(x)=-ln(1-x)-x, f(x) is a monotonically increasing function in (0,1). Therefore, when f(x)=2δ 2 L, you can find the root x * , and x must satisfy x≤x * . The implicit constraint can then be transformed into:

[0093]

[0094] The UAV is considered as an agent that interacts with the environment to learn. Therefore, the optimization problem P1 can be modeled as a Markov decision process. The agent obtains the hovering position q of the UAV in each round by learning the Markov decision process. u And CIoTD transmission power. In combination with the situation of this embodiment, this embodiment defines the state space S, action space A and reward function R as follows:

[0095] (1) S: Based on P1, the three-dimensional hovering position of the UAV needs to be optimized. Therefore, the state space can be defined as S = {q u}, and has 3 dimensions.

[0096] (2)A: At each time step of the system, the UAV needs to choose the direction and speed of movement and then move. The displacement is defined as Therefore, the action space is defined as A = {d}.

[0097] (3) R: The agent's goal is to maximize the cumulative reward, and the optimization problem is to minimize the average PAoI of PIoTDs (all PIoTDs). Therefore, the reward function at a given time step t can be expressed as the inverse of the average PAoI:

[0098]

[0099] λ is a positive constant used to adjust the reward value to accelerate convergence.

[0100] Let π θ(a|s) represents the strategy adopted by the agent, where θ represents the parameters of the neural network. To find the optimal strategy as the optimization goal, we need to find the best parameters θ that describe the strategy to obtain the maximum cumulative reward expectation of the strategy. The information age optimization network contains three neural networks with parameters θ old The old network, the Actor network with parameters θ and the Critic network with parameters φ. After collecting the state transition trajectories for T time steps, the agent updates the parameters θ and φ. Parameter θ old Used to ensure that the new policy does not differ significantly from the old policy.

[0101] The role of the actor network is to approximate the agent’s policy π θ (a|s), in order to keep the old policy unchanged in each update cycle, the participant network is initialized to store the policy of the actor network. The participant network only copies the parameters θ from the actor network for policy preservation and does not participate in training. The actor network updates its parameters according to the following formula during training:

[0102]

[0103] in, Represents the policy under parameter θ and the old parameter θ old The ratio of the strategies below. Represents the advantage function estimated based on the Critic network, where R t is the reward at time step t, V(s t ) is the time step t state s t The status value of clip(r θ ,1-δ,1+δ) represents the formula for the reverse update parameter θ. The Critic network is used to estimate the state value V(s t ), the goal is to minimize the loss function based on the mean square error. The critic network parameter φ is updated according to the following formula:

[0104]

[0105] V φ (s t ) is the state value calculated under the old Critic network parameter φ, which is calculated as a general formula and will not be described here.

[0106] The specific operation steps of the agent include:

[0107] F1, enter the location of CIoTD q c and transmit data S th , Willie's position q w , the channel uses L, the concealment requirement is δ, the location of PIoTD PIoTD transmit power The size of a single packet per PIoTD n , interval Δt;

[0108] F2, initialize the environment and state, and initialize the information age optimization network;

[0109] F3, train the information age optimization network;

[0110] In each round, the steps are performed:

[0111] F31, reset the initial state;

[0112] F32. At each time step, perform the following steps:

[0113] F321, UAV according to the current state s t Get the current position q u ;

[0114] F322, according to the current strategy π θ Select action a t ;

[0115] F323, in different transmission periods, make P c Satisfy power constraints (14), c2, c6, so that P c =min{P cv ,P cmax ,P ch}, P cv P calculated by formula (14) c The maximum value that can be achieved is P ch Indicates P calculated by c6 c The maximum value that can be achieved is According to formula (15), the average PAoI is calculated and the r of the current time step is obtained. t ;

[0116] F324, execute a t , get the next state s t+1 ;

[0117] F325, experience t ,a t ,r t >Store to experience replay buffer;

[0118] F326, determine whether the time step T is reached, if not, return to step F321 to enter the next time step, if yes, enter step F327;

[0119] ​F327, according to the output V(s of the Critic network t )calculate And copy the Actor network's θ to the old Actor network's θ old In the network, to maintain the old strategy;

[0120] F328, based on the experience data in the experience replay buffer, update the parameters of the Actor network and the Critic network according to equations (16) and (17);

[0121] F329, clear the experience replay buffer;

[0122] F330, return to the UAV's position q u , the transmission power of CIoTD and the average PAoI of PIoTD.

[0123] After training, the current UAV position q is output u and the transmission power of CIoTD as the solution to the optimization problem.

[0124] Based on the above-mentioned information age optimization method for the drone-assisted uplink covert transmission system, an embodiment of the present invention also provides an information age optimization system for the drone-assisted uplink covert transmission system, which is provided with an intelligent agent, which is used to implement steps S1 to S4 in the above-mentioned information age optimization method for the drone-assisted uplink covert transmission system.

[0125] In order to obtain the implementation effect of the present invention, this example compares the optimization problem solving algorithm of this embodiment (abbreviated as PPO) with the solution algorithm based on deep Q network (DQN) and policy gradient (PG) in the training process. The comparison results are as follows: Figure 4 As shown. Figure 4 It can be observed that all algorithms converged, but it is obvious that PPO achieved the highest cumulative reward after 600 episodes of training.

[0126] Figure 5 The average PAoI of the four algorithms varies with the stealth requirement δ, where 2D-PPO refers to the two-dimensional hovering optimization plus PPO algorithm (the height is fixed at q u z =160 meters), 3D-PG refers to drone three-dimensional hovering optimization plus policy gradient, 3D-DQN refers to drone three-dimensional hovering optimization plus deep Q network, and 3D-PPO refers to drone three-dimensional hovering optimization plus PPO algorithm. Figure 5 It shows that as the value of δ increases, the average PAoI value of the four algorithms gradually decreases. This is because the increase of δ relaxes the requirements for secret communication and allows CIoTD to use a higher transmission power Pc This enables the CIoTD to transmit more data within each Δt, thereby reducing the number N. Furthermore, compared to two-dimensional optimization, three-dimensional optimization can reduce the average PAoI because 3D optimization allows the UAV to select more possible deployment options and find a superior deployment location. The average PAoI of the 3D-PPO algorithm used in this embodiment is consistently the lowest, demonstrating the effectiveness and superiority of the present invention in minimizing the average PAoI of the PIoTD when the concealment requirement δ varies.

[0127] Figure 6 The figure shows the average PAoI of six schemes as the channel usage L changes. These six schemes correspond to the PG algorithm, DQN algorithm and PPO algorithm when the PIoTD power is 0.1W and 0.2W. Figure 6 It can be clearly seen that the average PAoI of the six schemes increases with increasing channel usage L. This is because Willie can obtain more observations through extensive channel usage. Therefore, CIoTD needs to use lower power for packet transmission, which increases the number of packets N that PIoTD needs to transmit, resulting in an increase in the average PAoI. In addition, a comparison of PIoTD powers of 0.1[W] and 0.2[W shows that the average PAoI decreases significantly with increasing PIoTD power. Although increasing PIoTD power in NOMA reduces the transmission rate, the increased power also provides a more secure transmission environment, enabling CIoTD to transmit data faster and reduce N. Therefore, increasing power within a suitable range can effectively reduce the average PAoI. Most importantly, under the same power setting, the average PAoI of the present invention is at the lowest level as the channel usage L increases. This demonstrates the effectiveness and superiority of the present invention in minimizing the average PAoI of PIoTD when the channel usage L is increased.

[0128] In summary, the information age optimization method and system for the UAV-assisted uplink covert transmission system provided in the embodiments of the present invention construct a UAV-assisted uplink covert transmission system. The mobility of the UAV and the noise generated by the PIoTD further improve the concealment requirements of the CIoTD. The present invention also aims to minimize the average peak age of information (PAoI) of all PIoTDs (PIoTDs). While meeting the concealment requirements, it constructs an optimization problem and further solves the optimization problem based on the Markov decision process to obtain the position of the UAV and the transmission power of the CIoTD. The simulation results demonstrate the effectiveness and excellence of the present invention in minimizing the average PAoI of the PIoTD.

[0129] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.

Claims

1. The information age optimization method of the UAV-assisted uplink covert transmission system is characterized by: Including steps: S1. Construct a system model of UAV-assisted uplink covert transmission system; The system model includes a covert IoT device (CIoTD), a listener (Willie), M public IoT devices (PIoTD), and an unmanned aerial vehicle (UAV). The CIoTD and the M PIoTDs perform uplink transmission to the UAV, and Willie attempts to detect the covert transmission of the CIoTD. S2. Determine optimization parameters, optimization objectives, and constraints based on the system model to construct an optimization problem; The optimization parameters are determined to be the UAV position and the CIoTD transmission power. The optimization goal is to minimize the average peak information age of M PIoTDs. The constraints include Willie's concealment constraint, the first transmission power constraint of the CIoTD, the flight altitude constraint of the UAV, the communication quality constraint, the data transmission quality constraint, and the final decoding constraint of the CIoTD. S3. Solve the optimization problem to obtain solution values ​​of optimization parameters; S4. Run the UAV-assisted uplink covert transmission system using the solution value of the optimization parameter.

2. The information age optimization method for the UAV-assisted uplink covert transmission system according to claim 1 is characterized by: During data transmission by CIoTD, each PIoTD generates N data packets with a time interval of Δt. The average peak information age of M PIoTDs is the sum of the peak information ages of these MN data packets divided by M.

3. The information age optimization method for the UAV-assisted uplink covert transmission system according to claim 2 is characterized in that: When Δτ mn When Δt is less than or equal to, the peak information age a of the nth data packet of the mth PIoTD mn Equal to Δτ mn , Δτ mn is the time to transmit the nth data packet of the mth PIoTD; when Δτ mn When a is greater than Δt, mn equal and Δτ mn the sum of is the waiting time from the generation to transmission of the nth data packet of the mth PIoTD.

4. The information age optimization method for the UAV-assisted uplink covert transmission system according to claim 3 is characterized by: Willie's concealment constraint is expressed as Willie's total error rate ξ is greater than or equal to 1-δ, δ is the concealment requirement; Willie's total error rate ξ is equal to the error response probability P FA and the probability of missed detection P MD The sum of the error response probability P FA Defined as the probability that CIoTD does not send information but Willie believes that CIoTD has sent it, the missed detection probability P MA It is defined as the probability that CIoTD has sent information but Willie believes that CIoTD has not.

5. The information age optimization method for the UAV-assisted uplink covert transmission system according to claim 4 is characterized in that: Assuming that Willie knows the transmission power of CIoTD and PIoTD, the channel gain from CIoTD to Willie, and Willie's noise distribution, Willie uses a binary hypothesis test H0, H1 based on the observed signal and the prior probabilities D0 and D1 of H0, H1 to calculate the probability of wrong response P FA =P(D1|H0) and the probability of missed detection P MA =P(D0|H1), where H0 represents the case where Willie determines based on the observed signal that CIoTD does not send any information to the UAV, while H1 represents the case where Willie determines based on the observed signal that CIoTD does send information. The corresponding relationships are: Among them, y w [l] represents the signal observed by Willie; n w [l] is the additive Gaussian white noise at Willie's location, with a mean of zero and a variance of x c [l] and x pm [l] is a complex Gaussian signal, and represents a complex Gaussian distribution with a mean of 0 and a variance of 1; h cw 、 Represent the channel coefficients between CIoTD, the mth PIoTD and Willie, is the transmission power of the mth PIoTD, P c is the transmit power of CIoTD.

6. The information age optimization method for the UAV-assisted uplink covert transmission system according to claim 5 is characterized by: The first transmission power constraint of CIoTD is specifically the transmission power P of CIoTD. c Not less than the maximum transmit power P of CIoTD cmax ; The UAV flight altitude constraint is specifically that the UAV flight altitude is within the minimum altitude and the maximum height The communication quality constraint is γ j ≥γ th ,y j is the signal-to-interference-noise ratio of CIoTD and PIoTD, y th is the signal-to-interference-noise ratio threshold required for non-interrupted communication; the data transmission quality constraint is specifically S c ≥S th , S c is the data transmission volume of CIoTD, S th is the data transmission volume threshold of CIoTD; the final decoding constraint of CIoTD is h cu 、 They represent the channel coefficients between CIoTD, the mth PIoTD and UAV respectively.

7. The information age optimization method for the UAV-assisted uplink covert transmission system according to claim 6 is characterized in that: Signal-to-interference-noise ratio (SIR) of CIoTD and PIoTD j=c,p m ,J=M+1 is the total number of IoT devices, P j |h ju | 2 is the signal part to be received, To affect the noise portion of the signal to be received, is the Gaussian white noise power of CIoTD and the mth PIoTD frequency band, j is the current device, i is all devices with worse channel quality than the current device; the data transmission volume of CIoTD s nk Indicates the kth segment of data that needs to be transmitted in the nth time interval; h cu 、 equal j=c,p m ,ρ0 is the power gain at a reference distance of 1m, d ju represents the distance from CIoTD and the mth PIoTD to the UAV, represents the line-of-sight component of the air-to-ground channel, and α l is the Earth-to-space path loss exponent.

8. The information age optimization method of the UAV-assisted uplink covert transmission system according to claim 7 is characterized in that: The step S3 specifically includes the following steps: Transform the hidden constraint ξ≥1-δ into P c The second power constraint: represents the variance of the additive white Gaussian noise at Willie, x * represents the root of f(x)=-ln(1-x)-x, x must satisfy x≤x * , x is defined as The UAV is considered as an agent that learns to interact with the environment, and the optimization problem is modeled as a Markov decision process. The state space of the Markov decision process is defined as S = {q u }, qu represents the rectangular three-dimensional coordinates of the UAV, the action space is defined as A = {d}, d represents the displacement of the UAV in each time step, and then the reward function r for moving the time step t t It is defined as the inverse of the average peak information age of M PIoTDs multiplied by a positive constant λ; Based on the defined Markov decision process and the constraints of the optimization problem, an information age optimization network is constructed and trained. The network contains parameters θ old The old network, the actor network with parameters θ and the critic network with parameters φ, collects the state transition trajectory of T time steps, and the agent updates the parameters θ and φ. The old network only copies the parameters θ from the actor network for policy preservation and does not participate in training; outputs the current UAV position q u and the transmission power of CIoTD as the solution to the optimization problem.

9. The information age optimization method of the UAV-assisted uplink covert transmission system according to claim 8 is characterized in that: During training, the agent performs the following steps: F1, enter the location of CIoTD q c and transmit data S th , Willie's position q w , the channel uses L, the concealment requirement is δ, the location of PIoTD PIoTD transmit power The size of a single packet per PIoTD n , interval Δt; F2, initialize the environment and state, and initialize the information age optimization network; F3, train the information age optimization network; In each round, the steps are performed: F31, reset the initial state; F32. At each time step, perform the following steps: F321, UAV according to the current state s t Get the current position q u ; F322, according to the current strategy π θ Select action a t ; F323, in different transmission periods, make P c Satisfy the first power constraint and the second power constraint, and the CIoTD final decoding constraint, so that P c =min{P cv ,P cmax ,P ch }, P cv P calculated corresponding to the second power constraint c The maximum value that can be achieved, P ch P is the final decoding constraint calculated by CIoTD. c The maximum value that can be achieved; calculate the average PAoI and get the r of the current time step t ; F324, execute a t , get the next state s t+1 ; F325, experience t ,a t ,r t >Store to experience replay buffer;​ F326, determine whether the time step T is reached, if not, return to step F321 to enter the next time step, if yes, enter step F327; F327, according to the output V(s of the critic network t )calculate And copy the actor network's θ to the old network's θ old In the network, to maintain the old strategy; F328. Update the parameters of the actor network and the critic network based on the experience data in the experience replay buffer; F329, clear the experience replay buffer; F330, return to the UAV's position q u , the transmission power of CIoTD and the average peak information age of M PIoTDs.

10. The information age optimization system of the UAV-assisted uplink covert transmission system is characterized by: It is provided with an intelligent agent, which is used to implement steps S1 to S4 in the information age optimization method of the drone-assisted uplink covert transmission system according to any one of claims 1 to 9.