UAV-assisted ISAC system physical layer secure transmission method based on reinforcement learning

In the integrated perception and communication system of drone-assisted, TD3 algorithm integrated with reinforcement learning and generative diffusion model are used to optimize the transmitted beamforming, artificial noise and UAV trajectory, and the physical layer secure transmission problems under multi-perception targets and non-perfect CSI are solved, achieving efficient secure transmission and system robustness.

CN120017112AActive Publication Date: 2025-05-16CHONGQING UNIV OF POSTS & TELECOMM
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510159755.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-16
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the problem of physical layer security transmission under multi-perception targets and non-perfect CSI in the integrated perception and communication system assisted by drone, resulting in poor system security performance and high risk of information leakage.

Method used

Using reinforcement learning-based methods, transmission model, secure communication model and perception model of UAV-assisted ISAC system are constructed, and the transmission beamforming, artificial noise and UAV trajectory are jointly optimized to maximize user and safe rate. Specifically, the generative diffusion model (GDM) is integrated into the Actor network of the TD3 algorithm, and the GDMTD3 algorithm is designed to solve optimization problems.

Benefits of technology

It significantly improves the system's adaptability in a dynamic environment, improves decision-making efficiency and robustness, and realizes dynamic adjustment of UAV under multi-objective and non-perfect CSI conditions, reduces the risk of information leakage and improves system security performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017112A_ABST
    Figure CN120017112A_ABST
Patent Text Reader

Abstract

The invention relates to a UAV-assisted ISAC system physical layer secure transmission method based on reinforcement learning, and aims to cope with information leakage risks under the conditions of multiple perception targets and imperfect channel state information. Comprising the following steps: constructing a UAV-assisted ISAC system model, and further constructing a secure communication model and a sensing model based on the model; a target function for maximizing the user security rate is established; in order to solve the proposed optimization problem, a dual-delay depth deterministic strategy gradient algorithm of a generative diffusion model is proposed for solving. Compared with the traditional algorithm, the method has the advantages that the exploration capability is enhanced by utilizing the GDM, the exploration and utilization relationship is effectively balanced, the system performance is improved, and the robustness is enhanced. According to the method, the effectiveness of the method is verified in a simulation result by taking the user and the security rate as indexes, and excellent secrecy performance and adaptive capacity to the environment are shown.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of communication and information security, and specifically relates to a physical layer security transmission method for multi-sensing targets and imperfect CSI in an unmanned aerial vehicle-assisted integrated perception and communication system. Background Art

[0002] The statements in this section only provide background information related to the present disclosure, and these statements may constitute prior art. In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art.

[0003] With the rapid growth of demand for intelligent unmanned aerial vehicles (UAVs), their application scenarios with communication support and perception detection functions are becoming increasingly extensive. However, the current traditional UAV system based on the separation of communication and perception design can no longer meet the increasingly complex needs, and a more efficient solution is urgently needed. Integrated Sensing and Communication (ISAC) technology, as an emerging approach, can improve the spectrum utilization efficiency and hardware resource utilization of UAV systems by realizing communication and perception functions in the same frequency band, thereby effectively addressing the limitations of traditional designs.

[0004] In this context, UAVs, with their high mobility and rapid deployment capabilities, can flexibly act as temporary base stations or relay nodes, and work in conjunction with ISAC technology to not only provide efficient communication support for ground users, but also achieve extensive perception. However, due to the openness and broadcast characteristics of wireless channels, especially the significant characteristics of the Line-of-Sight (LoS) channel between UAVs and ground nodes, UAV communication links are more vulnerable to threats from ground eavesdroppers, which seriously reduces the security of information transmission. In addition, the transmission beam is pointed at the target during the perception process, which will further increase the risk of information leakage. Existing studies are mostly aimed at single perception target scenarios, assuming that the communication user and the perception target are stationary and have perfect channel state information (CSI). These idealized assumptions are difficult to correspond to the actual environment. In addition, traditional convex optimization methods cannot effectively solve the problem of secure transmission under dynamic environmental changes. Although deep reinforcement learning (DRL) is adaptable to dynamic environments, conventional algorithms have large variance and policy deviation problems in high-dimensional action spaces.

[0005] For example, the patent application number 202410215216.3 is named “A physical layer secure transmission method and system in a synesthesia integrated UAV network”, and its method is: construct a synesthesia integrated UAV scenario; model the synesthesia signal emission and perception process of the UAV; model the confidential communication process between the UAV and the ground user; model the UAV communication performance constraints, perception beam pattern constraints and transmission power constraints; establish a confidential communication beamforming optimization target, and use the semi-positive definite optimization algorithm of the Dinkelbach algorithm to solve it. However, this patent does not take into account the complex situation of multiple perception target scenarios and imperfect channel state information, and it is difficult to cope with the actual environment, resulting in poor system security performance in complex scenarios and increasing the risk of information leakage. In addition, due to the impact of imperfect CSI on synesthesia performance, the robustness of the system is also reduced.

[0006] How to develop a robust and intelligent physical layer security transmission method for the secure transmission problem in an unmanned aerial vehicle (UAV)-assisted integrated sensing (ISAC) system with multiple sensing targets and imperfect channel state information (CSI) is an urgent problem to be solved in this field. Summary of the invention

[0007] In view of the above problems, the object of the present invention is to solve part of the problems in the prior art, or at least alleviate these problems.

[0008] A physical layer secure transmission method for a UAV-assisted ISAC system based on reinforcement learning, comprising the following steps:

[0009] Constructing a UAV-assisted ISAC system transmission model, and further constructing a secure communication model and a perception model based on the UAV-assisted ISAC system transmission model;

[0010] Based on the transmission model, secure communication model and perception model of the UAV-assisted ISAC system, under the constraints of total transmit power, communication user signal-to-interference-noise ratio and perception target beam pattern gain performance, a system optimization model is established by jointly designing transmit beamforming, artificial noise and UAV trajectory to maximize user and security rates; the optimization problem is expressed as:

[0011]

[0012] Among them, the decision variable is the beamforming vector of the transmitted signal R v [n] is the artificial noise covariance matrix, q u [n]=(x u [n],y u [n]) is the horizontal position of the UAV itself (i.e., the trajectory of the UAV); is the safe rate in the worst case, Pmax is the maximum transmission power of UAV; P(θ m [n]) = a H (θ m [n])R x [n]a(θ m [n]) is the perception model, θ m [n] is the departure angle with error, R x [n] is the covariance matrix of the transmitted signal; a(θ m [n]) is the array steering vector, representing the steering vector of the mth sensing target; Γ m is the minimum beam pattern gain, Γ c The lowest communication signal-to-interference-to-noise ratio; is the signal-to-interference-noise ratio received at the legitimate user k; h k [n] is the channel vector from UAV to legitimate user k in time slot n, Δh k [n] represents the channel estimation error; V max is the maximum flight speed of the UAV in a time slot, The initial position set for the UAV; x min 、x max ,y min ,y max represents the boundary of the region, (x u [n],y u [n],H) represents the time-varying position of the UAV in time slot n;

[0013] The generative diffusion model (GDM) is integrated into the Actor network of the TD3 algorithm (Double Delay Deep Deterministic Policy Gradient), and the GDMTD3 algorithm (Double Delay Deep Deterministic Policy Gradient supporting the generative diffusion model) is designed to solve the optimization problem to generate the optimal action according to the current environment state.

[0014] Constructing a UAV-assisted ISAC system transmission model, and further constructing a secure communication model and a perception model based on the model, including the following steps:

[0015] The UAV-assisted ISAC system transmission model is constructed: one mobile unmanned aerial vehicle (UAV) acts as a dual-function base station to communicate with the legitimate user k; the sensing target m (eavesdropper) eavesdrops on the legitimate user's information; the UAV position changes within multiple time slots n;

[0016] Establish the transmission signal model: In the nth time slot, the downlink transmission signal of the UAV is In the formula represents the signal sent to the legitimate user k, vector is the precoding vector of the transmitted signal, represents the embedded artificial noise (ArtificialNoise, AN), is the AN covariance matrix, H represents the fixed flight altitude of the UAV;

[0017] Establish an imperfect CSI model for communication users: Model the CSI of legitimate users and perceived targets, including legitimate user channels and eavesdropping channels, as well as model channel errors and angle errors.

[0018] Furthermore, the UAV is equipped with L uniform linear array antennas for transmitting beamforming; the legitimate user and the sensing target are both single antennas.

[0019] A secure communication model is constructed, including the definition of the secure rate for legitimate users; the expression of the secure rate in the worst case is given as:

[0020]

[0021] Where [x] + =max{x,0}; is the eavesdropper’s eavesdropping rate, is the eavesdropping SINR of the mth target eavesdropping on the kth user; : The transmission rate that can be achieved by legitimate users;

[0022] Construct a perception model of the system, including perceiving the target by transmitting waveforms and using beam pattern gain to characterize the perception performance, which is expressed as:

[0023] P(θ m [n]) = a H (θ m [n])R x [n]a(θ m [n])

[0024] Among them, a(θ m [n]) represents the guidance vector of the mth perceived target.

[0025] The generative diffusion model (GDM) is integrated into the Actor network of the TD3 algorithm (Double Delay Deep Deterministic Policy Gradient) and the GDMTD3 algorithm (Double Delay Deep Deterministic Policy Gradient Supporting Generative Diffusion Model) is designed to solve the optimization problem, including the following steps:

[0026] The optimization problem is modeled as a Markov decision process and the state space , Action Space , reward function To define;

[0027] GDM is integrated into the Actor network of the TD3 algorithm to capture complex state characteristics and generate optimal actions based on the current environment state.

[0028] Furthermore, the state space :definition is the state of time slot n, including the position of the UAV, the position information of the legal user and the sensing target (eavesdropper) on the ground, and the channel state information; s[n] = {q u [n],u k [n],u m [n],h k [n],g m [n]}, m, where q u [n] is the position of the UAV, u k [n]、u m [n] are the location information of the legitimate user on the ground and the sensing target (eavesdropper), h k [n], g m [n] are the channel vectors from UAV to legitimate users and sensing targets (eavesdroppers), respectively;

[0029] The action space : is defined as the action in time slot n, including the transmit beamforming vector, artificial noise covariance matrix, and UAV flight angle and distance; a[n] = {w1[n], ...w K [n],R v [n],φ[n],dis[n]}, where w k [n] is the transmit signal beamforming vector, R v [n] is the covariance matrix of artificial noise, φ[n] and dis[n] represent the direction and distance of UAV displacement, respectively;

[0030] The reward function : In the nth time slot, the immediate reward can be defined as: where r p [n] is the penalty imposed on the UAV flying out of the boundary, and ω is the penalty factor.

[0031] Furthermore, the optimization problem is modeled as a Markov decision process, comprising the following steps:

[0032] UAV is an agent (intelligent agent); in each training set period, the agent first obtains the state s[n] from the environment;

[0033] Then select a strategy a[n] from the action space;

[0034] Again, the environment updates the current state to s[n+1] and obtains the corresponding reward r[n];

[0035] The experience tuple s[n], a[n], r[n], s[n+1] is then stored in the replay memory buffer middle.

[0036] Design of the GDMTD3 algorithm, including the use of parameterized models To approximate the conditional distribution The parameterized model is expressed as:

[0037]

[0038] in, is the mean, g is the conditional information; is the predetermined variance factor, expressed as α represents all steps r≤t r Cumulative product, where α t =1-β t ;

[0039] The mean value of the reverse process is calculated as:

[0040]

[0041] in, is the original data; t is the number of diffusion steps, is the distribution of the tth step, I is the standard Gaussian noise, β t is the variance factor;

[0042] The following steps are involved:

[0043] For the original data Make an estimate, expressed as:

[0044]

[0045] in, It is a deep neural network;

[0046] Denoising noise is generated according to the conditional information g, and then the mean is indirectly approximated, which is expressed as:

[0047]

[0048] Tracking from arrive The reverse transformation of as follows:

[0049]

[0050] in, represents the standard normal distribution; Indicates that in the known Conditional reverse denoising to The conditional distribution of

[0051] When generating distribution Successfully trained, continue to obtain samples from the above formula

[0052] The design of the GDMTD3 algorithm also includes a reparameterization process that is conducive to sampling, which is expressed as follows:

[0053]

[0054] Among them, s represents the current state of the environment in DRL, as the conditional variable in the parameterized function; ⊙ is the operator of the Hadamard product, represents standard Gaussian noise.

[0055] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the physical layer security transmission method of a UAV-assisted ISAC system based on reinforcement learning.

[0056] The present invention has the following beneficial effects:

[0057] 1. Joint optimization of beamforming and UAV trajectory: By jointly optimizing the UAV trajectory, transmit beamforming and artificial noise covariance matrix, the present invention maximizes the system and safety rates and significantly improves the system's adaptability in dynamic environments.

[0058] 2. Algorithm innovation and performance improvement: This paper combines GDM with TD3 algorithm, effectively enhancing the algorithm's exploration ability and optimizing the balance between exploration and utilization, thereby significantly improving decision-making efficiency. At the same time, the proposed GDMTD3 algorithm shows stronger robustness and higher learning efficiency in high-dimensional action space.

[0059] 3. Adaptability: The designed method can quickly respond to the dynamic changes of the user and perceived target positions, and realizes the dynamic adjustment of the UAV under multi-target and non-perfect CSI conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] Figure 1 A scene diagram of the UAV-assisted ISAC system of the present invention;

[0061] Figure 2 is a flow chart of the present invention;

[0062] Figure 3It is the flight trajectory diagram of the UAV in the dynamic scene of the present invention;

[0063] Figure 4 This is a comparison chart of the safety rate of the solution proposed in the present invention at different powers and other solutions;

[0064] Figure 5 This is a comparison chart of the safety rate of the solution proposed in the present invention and other solutions at different perception thresholds. DETAILED DESCRIPTION

[0065] The present invention is further described below in conjunction with the accompanying drawings. The embodiments of the present invention are only used to illustrate the present invention rather than to limit the present invention. Without departing from the technical idea of ​​the present invention, various substitutions and changes can be made according to common technical knowledge and customary means in the field, which should all be included in the scope of the present invention.

[0066] In order to solve the problems existing in the above prior art, the present invention proposes a physical layer security transmission method for multi-sensing targets and imperfect CSI in a UAV-assisted ISAC system (i.e., a physical layer security transmission method for a UAV-assisted ISAC system based on reinforcement learning). Under the constraints of total transmit power, communication user signal-to-interference-noise ratio, and perception target beam pattern gain performance, the user and security rate are maximized by jointly designing transmit beamforming, artificial noise, and UAV trajectory. To achieve the above purpose, the present invention adopts the following method: Figure 2 The technical solution shown achieves:

[0067] S1: Constructing a UAV-assisted ISAC system transmission model, and further constructing a secure communication model and a perception model based on the UAV-assisted ISAC system transmission model.

[0068] The following steps are involved:

[0069] S11: Constructing a UAV-assisted ISAC system transmission model

[0070] The system model structure consists of Figure 1 As shown, a UAV-ISAC system is considered, in which a mobile unmanned aerial vehicle (UAV) acts as a dual-function base station, broadcasting K independent confidential information to K legitimate ground users, and M sensing targets (potential eavesdroppers) eavesdropping on the information of all users; the UAV position changes in multiple time slots n.

[0071] The UAV is equipped with a uniform linear array of L antennas for transmitting beamforming. Both the user and the sensing target have single antennas. In order to analyze the system reasonably, the total service time is divided into N time slots, and the position of the UAV is assumed to be approximately unchanged in each time slot. In the nth time slot, the position of the legitimate user k is defined as u k [n]=(xk [n],y k [n],0). The position of the perceived target m is defined as u m [n]=(x m [n],y m [n],0), let (x u [n],y u [n],H) represents the time-varying position of the UAV in time slot n, and the horizontal position of the UAV itself is q u [n]=(x u [n],y u [n]), H represents its fixed flight altitude. The deployment position constraint of the UAV can be expressed as: min ≤x u [n]≤x max ,y min ≤y u [n]≤y max , where x min ,x max ,y min ,y max Indicates the boundary of the region. V max is the maximum flight speed of the UAV in a time slot, then there is a constraint ||q u [n]-q u [n-1]||≤V max ,

[0072] S12: Establishing transmission signal model

[0073] exist Figure 1 Under the system model shown, the signal model of the system is established. In time slot n, the downlink transmission signal of the UAV is given by the following formula:

[0074]

[0075] In the formula Represents the signal sent to user k. is the transmit beamforming vector in time slot n, represents the embedded AN. The covariance matrix of the transmitted signal is in, is the covariance matrix of the transmit beamforming vector, is the AN covariance matrix.

[0076] The average transmission power of the UAV in the nth time slot is The maximum transmission power of UAV is P max , so the total transmit power constraint is

[0077] S13: Establishing an imperfect CSI model for communication users

[0078] Construct imperfect CSI and statistical CSI models for communication users. Model the CSI of legitimate users and sensing targets, including legitimate user channels and eavesdropping channels, as well as model channel errors and angle errors. Since the altitude of UAV is relatively high, there is usually a strong LoS link between UAV and each ground user. Therefore, the present invention considers the LoS channel model, and on this basis, the channel vector from UAV to legitimate user k is recorded as:

[0079]

[0080] in, is the estimated channel, Δh k [n] represents the estimated error of the channel. The estimated channel is Among them, β c is the path loss at the legal channel reference distance of 1 m, η is the path loss factor from the UAV to the LOS link, is the distance from UAV to the legitimate user k; a(θ k [n]) is the array steering vector, expressed as:

[0081]

[0082] Where λ and d are the carrier wavelength and the distance between two adjacent antennas, respectively. represents the departure angle corresponding to user k; j represents the imaginary unit, L is the number of antennas, and T represents transpose.

[0083] Consider the common estimation error, i.e. bounded error, i.e. additive channel error model, modeled as ||Δh k [n]||≤μ k , where ||·|| represents the modulo operation, μ k The scalar represents the upper bound of the error, which is the bounded channel uncertainty of the communication user. Limited by factors such as hardware capabilities, UAV can obtain rough CSI, and the channel can be modeled as a bounded CSI error model, which is a worst-case assumption.

[0084] Then, the uncertainty model of the perceived target angle is established. In actual operation, it is unlikely to know the exact position of the target in advance. Therefore, a target position angle uncertainty region is defined. The perception channel is modeled as a pure LoS channel. Among them, β e is the path loss at the reference distance of 1m for the eavesdropping channel, is the distance from the UAV to the perceived target m. The angle error is: where Δθ m [n]≤εm , ε m represents the upper bound of the bounded error.

[0085] S14: Establish a secure communication model

[0086] In the nth time slot, the received signal of the legitimate user k is expressed as where z k [n] is additive Gaussian noise, and the signal-to-interference-plus-noise ratio (SINR) received by the legitimate user k is:

[0087]

[0088] in represents the communication additive Gaussian noise power, so the transmission rate that a legitimate user can achieve is w i [n] represents the transmit beamforming vector sent to the i-th user.

[0089] The signal received at the mth sensing target is Similarly, where z m [n] is additive Gaussian noise, and the eavesdropping SINR of the mth target eavesdropping on the kth legitimate user is given by:

[0090]

[0091] in represents the eavesdropping additive Gaussian noise power, tr(·) represents the trace of the matrix, so the eavesdropper’s eavesdropping rate is

[0092] The secure rate achievable at the legitimate user is defined as the difference between the achievable rates at the legitimate user and the eavesdropper. Therefore, the expression for the worst-case confidentiality rate is given as where [x] + =max{x,0}.

[0093] S15: Build a perceptual model of the system

[0094] The above transmission signals are used for both perception and communication functions, where each transmission signal is considered as a radar pulse. Then, the beam pattern gain is used to characterize the perception performance, which is expressed as:

[0095] P(θ m [n]) = a H (θ m [n])R x [n]a(θ m [n]) (6)

[0096] Among them, a(θ m [n]) represents the guidance vector of the mth perceived target.

[0097] S2: Establish system optimization model to maximize user and security rates

[0098] Based on the UAV-assisted ISAC system transmission model, secure communication model and perception model, the present invention proposes to establish a system optimization model by jointly designing transmit beamforming, artificial noise and UAV trajectory under the constraints of total transmit power, communication user signal-to-interference-noise ratio and perception target beam pattern gain performance to maximize user and security rates; the optimization problem is expressed as:

[0099]

[0100] Among them, the decision variable is the beamforming vector of the transmitted signal R v [n] is the artificial noise covariance matrix, q u [n]=(x u [n],y u [n]) is the horizontal position of the UAV itself (i.e., the trajectory of the UAV); is the safe rate in the worst case, P max is the maximum transmission power of UAV; P(θ m [n]) = a H (θ m [n])R x [n]a(θ m [n]) is the perception model, θ m [n] is the angle error, R x [n] is the covariance matrix of the transmitted signal; a(θ m [n]) is the array steering vector, representing the steering vector of the mth sensing target; Γ m is the minimum beam pattern gain, Γ c The lowest communication signal-to-interference-to-noise ratio; is the signal-to-interference-noise ratio received at the legitimate user k; h k [n] is the channel vector from UAV to legitimate user k in time slot n, Δh k [n] represents the channel estimation error; V max is the maximum flight speed of the UAV in a time slot, The initial position set for the UAV; x min 、x max ,y min ,y max represents the boundary of the region, (x u [n],y u[n], H) represents the time-varying position of the UAV in time slot n. Constraint C1 is the total transmit power constraint of the UAV at time n, and C2 is the constraint of the perception performance, which ensures that the minimum beam pattern gain that can meet multiple perception targets is Γ m C3 is the communication performance constraint, which ensures that each user can meet the minimum communication signal-to-interference-to-noise ratio Γ c C4-C5 are constraints on the UAV flight position, ensuring that the UAV flies from the set initial position Set off, and do not fly out of the designated area.

[0101] S3: Designing the GDMTD3 algorithm to solve optimization problems

[0102] The generative diffusion model (GDM) is integrated into the Actor network of the TD3 algorithm (Double Delay Deep Deterministic Policy Gradient), and the GDMTD3 algorithm (Double Delay Deep Deterministic Policy Gradient supporting the generative diffusion model) is designed to solve the optimization problem to generate the optimal action according to the current environment state.

[0103] In order to effectively solve the planned problem, the present invention uses the DRL method to solve the formulated problem. In order to be applicable to continuous state and action spaces, the TD3 learning method is selected among the many variants of DRL. However, conventional TD3 algorithms usually use stacked fully connected layers in the Actor network, which makes it difficult to capture deeper data features. Therefore, these algorithms may cause the learned policy distribution to deviate from the true data distribution. GDM can better capture the complex dynamics of the proposed problem scenario, thereby generating better solution strategies. Therefore, in response to the problems of low sampling efficiency, difficulty in modeling complex environments, and slow convergence speed in conventional TD3 algorithms, the present invention combines the TD3 algorithm with a generative diffusion model, and proposes a double-delayed deep deterministic policy gradient (Generative Diffusion Model-Enabled Twin Delayed Deep Deterministic Policy Gradient, GDMTD3) method that supports a generative diffusion model, which is designed to deal with the presence of mobile users, eavesdroppers, and imperfect CSI situations. Since GDM has a strong ability to model complex data distribution, it can reduce the number of samples required and more effectively represent the characteristics of complex and high-dimensional state and action spaces; and its unique structure involves a diffusion process, which can provide a more stable and effective learning process and effectively solve the above problems. Specifically, the designed GDMTD3 algorithm includes the following steps:

[0104] Firstly, the original optimization problem is modeled as a Markov decision process.

[0105] The original UAV-ISAC safety problem can be modeled as a Markov decision process. The UAV is an agent, where the state of the system at each time slot depends only on the current state of the system and the action taken by the UAV. The interaction between the agent and the environment is as follows. In each time period of the training set, the agent first obtains the state s[n] from the environment and then selects a strategy a[n] from the action space; then, the environment updates the current state to s[n+1] and obtains the corresponding reward r[n]. The experience tuple s[n], a[n], r[n], s[n+1] is then stored in the replay memory buffer. The corresponding elements are defined as follows:

[0106] State Space : The set of all possible observations of the system state by the UAV, i.e., the state space. Definition is the state of time slot n; s[n] = {q u [n],u k [n],u m [n],h k [n],g m [n]}, m, where q u [n] is the position of the UAV, u k [n],u m [n] are the location information of the legitimate user and the eavesdropper on the ground, h k [n],g m [n] are the channel vectors from UAV to user and eavesdropper respectively.

[0107] Action Space : The UAV selects the corresponding action based on the observed state information. Defined as the action in time slot n, a[n] = {w1[n], ...w K [n],R v [n],φ[n],dis[n]}, where w k [n] is the transmit signal beamforming vector, R v [n] is the covariance matrix of artificial noise, φ[n], dis[n] represent the direction and distance of UAV displacement, respectively.

[0108] Reward Function : The reward function is used to measure the performance of the agent and guide its decision-making process. In the optimization problem, the optimization goal is to maximize the system user and security rate under the constraints of perception performance, communication performance and transmission power. Therefore, in time slot n, the instantaneous reward is defined as where r p[n] The penalty imposed on the UAV flying out of the boundary, ω is the penalty factor, the degree of which is proportional to the gap between the agent and the constraint, and is defined as follows:

[0109]

[0110] The final reward is That is, the goal of the TD3 algorithm is to maximize the long-term cumulative discounted reward, where γ∈(0,1) is the discount factor.

[0111] Secondly, GDM is integrated into the Actor network of the TD3 algorithm to capture complex state characteristics and generate optimal actions based on the current environment state.

[0112] Aiming at the problems of low sampling efficiency, difficulty in modeling complex environments, and slow stability and convergence speed in the conventional TD3 algorithm, the present invention integrates GDM into the Actor network of the TD3 algorithm. The TD3 algorithm is an Actor-Critic based DRL algorithm, in which the Actor network μ(s|θ μ ) outputs a deterministic action, the Critic network Q(s,a|θ Q ) Evaluate the action-state value function. Among them, the Actor network is the "actor network" or "strategy network"; it is responsible for "decision-making", generating actions based on the current state and updating the strategy by maximizing the Q value of the Critic. The Critic network is the "evaluator network" or "value network"; it is responsible for evaluating the value of the actions generated by the Actor, providing feedback signals, and updating its own parameters by minimizing the temporal difference error.

[0113] The goal is to find the optimal strategy that maximizes the expected cumulative reward. The training process of TD3 involves updating the Actor and Critic networks based on a specific loss function. The update of the Critic network is achieved by minimizing the time difference error loss function, which is obtained from the experience replay buffer. A batch of samples are randomly selected from the , the number is B, and the loss function of the Critic network is approximately:

[0114]

[0115] in is the parameter related to the Critic network, δ b =r b +γmin i=1,2 Q' i (s b ,μ′(s b |θ' μ )+∈) is a recursive decomposition provided by the Bellman equation to update the action value function. Actor network μ(s|θ μ) should be updated less frequently than the Critic network to ensure stable learning. The goal of the Actor network is to maximize the expected Q value evaluated by the first Critic network. A batch of B samples are randomly selected from the Actor network, and the loss function of the Actor network is approximately:

[0116]

[0117] The target network is updated using a soft update mechanism that mixes the parameters of the main network and the target network using a weight factor. The update is defined as follows:

[0118]

[0119] as well as

[0120] θ' μ ←τθ μ +(1-τ)θ' μ , (12)

[0121] where τ is a small soft weight factor. It can be seen that the updated parameters of the target network are a weighted combination of its original parameters and the corresponding network parameters.

[0122] Secondly, the reverse denoising process of GDM is used. Specifically, the reverse denoising process reconstructs the original data by systematically removing noise; in the reverse process, the goal is to remove noise by iteration, from a standard Gaussian distribution. of Recovering original data from noisy samples However, the statistical distribution Requires computations involving data distribution, which is often intractable in practice. The strategy is to use a parameterized model To approximate the conditional distribution The model can be expressed as In the formula, is the mean, where g is the conditional information, is the predetermined variance factor, expressed as in, α represents all steps r≤t r Cumulative product, where α t =1-β t Using the Bayesian formula, the inverse process is reconstructed into a Gaussian probability density function. The average value of the reverse process is calculated as However, the parameterized model No access Therefore, it is necessary to estimate in is a deep neural network that generates denoised noise according to the condition g and then indirectly approximates the mean,

[0123]

[0124] Tracking from arrive The reverse transformation can establish the generating distribution as follows:

[0125]

[0126] in represents the standard normal distribution. Once the distribution is generated After successful training, you can continue to obtain samples from the above formula

[0127] The generative power of diffusion models enables the creation of complex action sets that are refined through the backward process of learning, allowing actions to be sampled directly from the generative distribution. A significant challenge in ensembling diffusion models is managing the stochastic component, which complicates the gradient descent methods commonly used for training. To overcome this issue, a sampling-friendly reparameterization process is employed, which is expressed as follows:

[0128]

[0129] Where s represents the current state of the environment in DRL, as the conditional variable in the parameterized function. ⊙ is the operator of the Hadamard product, Represents standard Gaussian noise. Based on the above design, a GDM-based motion sampling algorithm is designed and integrated into the existing TD3 algorithm.

[0130] The main steps of the action sampling process based on the generative diffusion model are detailed in Algorithm 1.

[0131] Table 1 Action sampling algorithm based on generative diffusion model

[0132]

[0133] GDM is integrated into the Actor network of the TD3 algorithm to capture complex state characteristics, generate optimal actions based on the current environment state, and enhance the decision-making ability of the Actor network. Algorithm 2 describes the specific implementation of this process.

[0134] Table 2 TD3 algorithm based on generative diffusion model

[0135]

[0136]

[0137] The computational complexity of GDMTD3 during the training phase is in represents the number of parameters in the two online critic networks, |θ d | represents the number of parameters in the online actor network supporting diffusion. M represents the number of training rounds, N is the number of steps per round, and d t is the number of denoising steps required to sample an action in the diffusion actor network. V represents the complexity of interaction with the environment, and d is the frequency of policy updates.

[0138] In order to demonstrate the effectiveness of the proposed scheme, a comprehensive evaluation of the proposed method was conducted, and the effectiveness and robustness of the proposed GDMTD3 in solving the UAV-ISAC security problem were verified. This example provides a description of the simulation setup, including environmental details, model design, and benchmarks for evaluating the performance of the proposed method. In this invention, a synaesthesia scenario consisting of 1 independent UAV (equipped with 8 antennas), 3 mobile legitimate users, and 2 mobile sensing targets (eavesdroppers) is considered, and the UAV flies at an altitude of 100m. In order to fully verify the advantages of the proposed scheme, the following comparison schemes are set up:

[0139] Comparison scheme 1: The UAV trajectory is fixed, and the beamforming of the proposed algorithm and the AN joint optimization strategy are used.

[0140] Comparison scheme 2: No AN is embedded, and the UAV trajectory and beamforming joint optimization strategy of the proposed algorithm are used.

[0141] Comparison scheme 3: UAV trajectory, beamforming and AN joint optimization strategy using conventional TD3 algorithm.

[0142] Through comparative simulation with the above schemes, this example gives specific simulation and conclusions.

[0143] Figure 3 The flight trajectory of the UAV in an environment with dynamic ground user changes is shown. The initial position of the UAV is [0,0] meters, and the flight altitude is 100 meters. The UAV adjusts the path based on real-time perception to optimize communication quality and avoid eavesdropping risks. As the positions of ground users and eavesdroppers change over time, the UAV flexibly adjusts the trajectory to cope with the dynamically changing environment. The UAV deployment location meets the perception requirements and improves the safety rate, reflecting its ability to adaptively adjust.

[0144] Figure 4The trend of user and security rate (Sum Security Rate, SSR) changing with power is shown. Under the condition that the sensing target position is determined and the sensing threshold is 3dB, the proposed scheme is compared with the baseline scheme. The results show that as the power increases, the SSR of all schemes increases, and the performance of the proposed scheme is significantly better than the baseline scheme, reflecting its advantages in optimizing channel conditions and avoiding eavesdropping. In addition, the proposed scheme requires less power to achieve the same SSR, proving its high efficiency. Under non-perfect CSI conditions, the proposed scheme can achieve an SSR close to perfect CSI, indicating that the algorithm has good robustness and can effectively reduce the impact of CSI errors.

[0145] Figure 5 The trend of SSR changing with the perception threshold is shown. The SSR performance of the proposed scheme is compared with that of the baseline scheme at different perception thresholds. The results show that as the perception threshold increases, the SSR of all schemes decreases. This is because a higher perception threshold requires more power to be allocated for the perception task, reducing communication power resources. The SSR of the proposed scheme is always better than the baseline scheme, reflecting its advantages in optimizing UAV paths, power allocation, and AN interference strategies. In addition, when there is uncertainty in the perceived target position, the SSR of the proposed algorithm is close to the perfect perception condition, indicating that it has good robustness and can adapt to situations where the perceived target position is slightly uncertain.

[0146] The present invention specifically includes:

[0147] Construct a transmission model of the UAV-assisted ISAC system: This model contains multiple sensing targets and takes into account the mobility of ground users and sensing targets. By jointly optimizing the transmit beamforming, artificial noise covariance matrix and UAV trajectory, the user and safety rates are maximized. First, consider the ideal scenario where the perfect communication user CSI and the precise location of the target are known, and obtain the ideal performance upper bound by designing a solution algorithm. Then consider the situation where the legitimate user CSI has errors and the sensing target has angle errors. In order to solve the formulated optimization problem, a Generative Diffusion Model-enabled Twin Delayed Deep Deterministic Policy Gradient (GDMTD3) method is proposed, which integrates GDM in the Actor network of the TD3 algorithm. Since GDM can model complex data distributions and generate high-quality samples, GDMTD3 can effectively solve the problems of low sampling efficiency, difficulty in modeling complex environments, and slow stability and convergence speed in the conventional TD3 algorithm. Finally, we verify the effectiveness of the proposed scheme through simulation.

[0148] The present invention combines the Generative Diffusion Model (GDM) and the Twin Delayed Deep Deterministic policy gradient (TD3) technology to solve the optimization problem under multi-sensory targets and imperfect CSI, maximize the system user and security rate, reduce the risk of information leakage, and improve the system security performance in complex scenarios. It also effectively handles the impact of imperfect CSI on synaesthesia performance and improves the robustness of the system.

[0149] The present invention is applicable to the field of UAV communication, especially to complex scenarios such as UAV operation and emergency communication in emergency situations. The technical solution ensures the dual efficiency of communication security and information perception by jointly optimizing the flight trajectory and wireless resources of the UAV. In the military and security fields, the technical solution can significantly improve the flexibility and safety of UAVs in military operations, border monitoring and other tasks, so that UAVs can dynamically adjust flight paths and resource management according to real-time battlefield situations, and achieve safe and efficient information transmission and situational awareness. More importantly, the present invention designs a highly adaptive physical layer security transmission strategy for the problems of imperfect CSI and dynamic environmental changes that are prevalent in actual communication environments, significantly enhancing the robustness and practicality of the system. Therefore, the present invention not only has significant advantages in the military and security fields, but is also widely applicable to multiple fields such as smart city management, disaster relief communications, and environmental monitoring, showing huge market demand and broad development prospects.

[0150] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the physical layer security transmission method of a UAV-assisted ISAC system based on reinforcement learning.

[0151] The conventional techniques and schemes not described in detail in the above embodiments are well known in the art, so they will not be described in detail here. The above embodiments and / or experimental examples describe the preferred embodiments of the present invention in detail, but the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, the technical scheme of the present invention can be subjected to a variety of simple modifications, and these simple modifications all belong to the protection scope of the present invention.

Claims

1. A physical layer security transmission method for a UAV-assisted ISAC system based on reinforcement learning, characterized in that: The following steps are involved: Constructing a UAV-assisted ISAC system transmission model, and further constructing a secure communication model and a perception model based on the UAV-assisted ISAC system transmission model; Based on the transmission model, secure communication model and perception model of the UAV-assisted ISAC system, under the constraints of total transmit power, communication user signal-to-interference-noise ratio and perception target beam pattern gain performance, a system optimization model is established by jointly designing transmit beamforming, artificial noise and UAV trajectory to maximize user and security rates; the optimization problem is expressed as: Among them, the decision variable is the beamforming vector of the transmitted signal R v [n] is the artificial noise covariance matrix, q u [n]=(x u [n],y u [n]) is the horizontal position of the UAV itself (i.e., the trajectory of the UAV); is the safe rate in the worst case, P max is the maximum transmission power of UAV; P(θ m [])=a H (θ m [])R x []a(θ m [n]) is the perception model, θ m [n] is the departure angle with error, R x [n] is the covariance matrix of the transmitted signal; a(θ m [n]) is the array steering vector, representing the steering vector of the mth sensing target; Γ m is the minimum beam pattern gain, Γ c The lowest communication signal-to-interference-to-noise ratio; is the signal-to-interference-noise ratio received at the legitimate user k; h k [n] is the channel vector from UAV to legitimate user k in time slot n, Δh k [n] represents the channel estimation error; V max is the maximum flight speed of the UAV in a time slot, The initial position set for the UAV; x min 、x max ,y min ,y max represents the boundary of the region, (x u [n],y u [n],H) represents the time-varying position of the UAV in time slot n; The generative diffusion model (GDM) is integrated into the Actor network of the TD3 algorithm (Double Delay Deep Deterministic Policy Gradient), and the GDMTD3 algorithm (Double Delay Deep Deterministic Policy Gradient supporting the generative diffusion model) is designed to solve the optimization problem to generate the optimal action according to the current environment state.

2. The physical layer secure transmission method of the UAV-assisted ISAC system based on reinforcement learning according to claim 1 is characterized in that: Constructing a UAV-assisted ISAC system transmission model, and further constructing a secure communication model and a perception model based on the UAV-assisted ISAC system transmission model, including the following steps: The UAV-assisted ISAC system transmission model is constructed: one mobile unmanned aerial vehicle (UAV) acts as a dual-function base station to communicate with the legitimate user k; the sensing target m (eavesdropper) eavesdrops on the legitimate user's information; the UAV position changes within multiple time slots n; Establish the transmission signal model: In the nth time slot, the downlink transmission signal of the UAV is In the formula represents the signal sent to the legitimate user k, vector is the precoding vector of the transmitted signal, represents the embedded artificial noise (ArtificialNoise, AN), is the AN covariance matrix, H represents the fixed flight altitude of the UAV; Establish an imperfect CSI model for communication users: Model the CSI of legitimate users and perceived targets, including legitimate user channels and eavesdropping channels, as well as model channel errors and angle errors.

3. The physical layer secure transmission method of the UAV-assisted ISAC system based on reinforcement learning according to claim 2 is characterized in that: The UAV is equipped with L uniform linear array antennas for transmitting beamforming; the legitimate user and the sensing target are both single antennas.

4. The physical layer secure transmission method of the UAV-assisted ISAC system based on reinforcement learning according to claim 1 is characterized in that: A secure communication model is constructed, including the definition of the secure rate for legitimate users; the expression of the secure rate in the worst case is given as: Where [x] + =max{x,0}; is the eavesdropper’s eavesdropping rate, is the eavesdropping SINR of the mth target eavesdropping on the kth user; The transfer rate achievable by legitimate users; Construct a perception model of the system, including perceiving the target by transmitting waveforms and using beam pattern gain to characterize the perception performance, which is expressed as: P(θ m [ ])=a H (i m [ ])R x [ ]a(θ m [n]) Among them, a(θ m [n]) represents the guidance vector of the mth perceived target.

5. According to the reinforcement learning-based UAV-assisted ISAC system physical layer secure transmission method of claim 1, a generative diffusion model (GDM) is integrated into the Actor network of the TD3 algorithm (double-delayed deep deterministic policy gradient), and a GDMTD3 algorithm (double-delayed deep deterministic policy gradient supporting a generative diffusion model) is designed to solve the optimization problem, comprising the following steps: The optimization problem is modeled as a Markov decision process and the state space Action Space Reward Function To define; GDM is integrated into the Actor network of the TD3 algorithm to capture complex state characteristics and generate optimal actions based on the current environment state.

6. The physical layer secure transmission method of the UAV-assisted ISAC system based on reinforcement learning according to claim 5 is characterized in that: The state space definition is the state of time slot n, including the position of the UAV, the position information of the legal user and the sensing target (eavesdropper) on the ground, and the channel state information; s[n] = {q u [n],u k [n],u m [n],h k [n],g m [n]}, where q u [n] is the position of the UAV, u k [n]、u m [n] are the location information of the legitimate user on the ground and the sensing target (eavesdropper), h k [n], g m [n] are the channel vectors from UAV to legitimate users and sensing targets (eavesdroppers), respectively; The action space is defined as the action in time slot n, including the transmit beamforming vector, artificial noise covariance matrix, and UAV flight angle and distance; a[n] = {w1[n], ...w K [n],R v [n],φ[n],dis[n]}, where w k [n] is the transmit signal beamforming vector, R v [n] is the covariance matrix of artificial noise, φ[n] and dis[n] represent the direction and distance of UAV displacement, respectively; The reward function In the nth time slot, the immediate reward is defined as: where r p [n] is the penalty imposed on the UAV flying out of the boundary, and ω is the penalty factor.

7. The physical layer secure transmission method of the UAV-assisted ISAC system based on reinforcement learning according to claim 5 or 6, characterized in that: The optimization problem is modeled as a Markov decision process, which includes the following steps: UAV is an agent. In each training set, the agent first obtains the state s[n] from the environment. Then select a strategy a[n] from the action space; Again, the environment updates the current state to s[n+1] and obtains the corresponding reward r[n]; The experience tuple s[n], a[n], r[n], s[n+1] is then stored in the replay memory buffer middle.

8. The physical layer secure transmission method of the UAV-assisted ISAC system based on reinforcement learning according to claim 1, characterized in that: Design of the GDMTD3 algorithm, including the use of parameterized models To approximate the conditional distribution The parameterized model is expressed as: in, is the mean, g is the conditional information; is the predetermined variance factor, expressed as α represents all steps r≤t r Cumulative product, where α t =1-β t ; The mean value of the reverse process is calculated as: in, is the original data; t is the number of diffusion steps, is the distribution of the tth step, I is the standard Gaussian noise, β t is the variance factor; The following steps are involved: For the original data Make an estimate, expressed as: in, It is a deep neural network; Denoising noise is generated according to the conditional information g, and then the mean is indirectly approximated, which is expressed as: Tracking from arrive The reverse transformation of as follows: in, represents the standard normal distribution; Indicates that in the known Reverse denoising under the condition The conditional distribution of When generating distribution Successfully trained, continue to obtain samples from the above formula 9. The physical layer secure transmission method of the UAV-assisted ISAC system based on reinforcement learning according to claim 8, characterized in that: The design of the GDMTD3 algorithm also includes a reparameterization process that is conducive to sampling, which is expressed as follows: Among them, s represents the current state of the environment in DRL, as the conditional variable in the parameterized function; ⊙ is the operator of the Hadamard product, represents standard Gaussian noise.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the physical layer security transmission method of a UAV-assisted ISAC system based on reinforcement learning are implemented as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Physical layer secure transmission method and system in communication and sensing integrated unmanned aerial vehicle network

    CN118042454A

  • Physical layer security-covert communication joint optimization method based on ISAC system

    CN117580029A

  • Reconfigurable intelligent surface auxiliary safety sensing and communication control method and system

    CN117880803A

  • Machine learning training method for unmanned aerial vehicle communication perception integrated network

    CN118778434A

  • Rate optimization method of unmanned aerial vehicle information acquisition system based on sensing integration

    CN119233285A