A ris-assisted uav-isac system security and energy efficiency optimization method

By constructing a multi-agent PPO decision network and a deep reinforcement learning model DRL, the reflection coefficient and beamforming matrix of the RIS-assisted UAV-ISAC system are optimized, solving the problems of poor CSI adaptability and insufficient energy consumption control in the UAV-ISAC system, and achieving efficient synergistic optimization of communication rate, energy efficiency and safety performance.

CN121441349BActive Publication Date: 2026-03-31SICHUAN UNIVERSITY OF SCIENCE AND ENGINEERING
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-30
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing technologies in UAV-ISAC systems suffer from poor CSI adaptability, weak multi-objective optimization capabilities, insufficient optimization efficiency and stability, and imperfect physical layer security mechanisms. In particular, they lack consideration for UAV energy consumption control, resulting in low energy efficiency during long-term operation.

Method used

The RIS-assisted UAV-ISAC system is adopted. By constructing a multi-agent PPO decision network and combining it with a deep reinforcement learning model DRL, the system optimizes the reflection coefficient and beamforming matrix of the reconfigurable smart surface RIS, dynamically adjusts the artificial noise power, and optimizes the UAV trajectory and channel state information, thereby achieving synergistic optimization of communication rate, energy efficiency and safety performance.

Benefits of technology

It improves the system's robustness in CSI imperfect scenarios, increases security rate and energy efficiency, reduces eavesdropper channel capacity, and ensures long-term efficient operation of UAV under constrained conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121441349B_ABST
    Figure CN121441349B_ABST
Patent Text Reader

Abstract

The application discloses a kind of RIS auxiliary UAV-ISAC system security energy efficiency optimization methods, comprising: using unmanned aerial vehicle UAV subsystem acquisition individual link channel information data;According to channel posterior distribution and channel transfer model, obtain channel posterior covariance;Multi-agent PPO decision network is constructed, and there is deep reinforcement learning model DRL in multi-agent PPO decision network, and multi-agent PPO decision network generates the beamforming matrix of artificial noise AN, and generates total transmission signal;Optimize the control parameter of multi-agent PPO decision network, output convergent multi-agent PPO decision network, to optimize the communication rate, energy efficiency and security performance of system.The present application introduces the mechanism of bayesian channel estimation, improves the robustness of system under the imperfect scene of CSI (channel state information), ensures the stable performance of channel uncertainty RMSE and SEE.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) communication technology, and more specifically to a RIS-assisted method for optimizing the safety and energy efficiency of a UAV-ISAC system. Background Technology

[0002] In the development of sixth-generation (6G) communication technology, Information-Sensing Integration (ISAC) has become one of the core technology directions due to its ability to improve spectrum utilization efficiency; Reconfigurable Smart Surfaces (RIS), with their programmable control of electromagnetic waves, can enhance the spectrum and energy efficiency of unmanned aerial vehicle (UAV) networks. The combination of the two provides a new path for improving the performance of ISAC systems.

[0003] The existing technology has the following drawbacks and shortcomings:

[0004] Poor adaptability of CSI (Channel State Information): Most studies are based on the assumption of perfect CSI, but in real-world scenarios, factors such as UAV (Unmanned Aerial Vehicle) maneuverability, user movement, and signal blockage can lead to imperfect CSI. Traditional fixed CSI estimation methods suffer significant performance degradation under channel uncertainty, resulting in insufficient system robustness.

[0005] Weak multi-objective optimization capability: Existing technologies mostly focus on optimizing a single objective (such as communication rate and sensing accuracy), failing to achieve multi-objective coordination of "communication security-sensing accuracy-energy efficiency optimization", making it difficult to balance the root mean square error of system positioning (RMSE) and security energy efficiency (SEE).

[0006] Insufficient optimization efficiency and stability: Traditional DRL algorithms (such as DDPG and TD3) suffer from problems such as overestimation bias, slow convergence speed, and poor adaptability to high-dimensional space; the single agent architecture cannot decouple multiple variables such as UAV trajectory, RIS (reconfigurable smart surface) reflection coefficient, beamforming, and AN (artificial noise) power ratio, and is prone to getting trapped in local optima.

[0007] The physical layer security mechanism is imperfect: Although some technologies introduce artificial noise (AN) to interfere with eavesdroppers, they are not coordinated with channel estimation technology. The power allocation of AN (artificial noise) lacks dynamic optimization, making it difficult to achieve targeted suppression of eavesdroppers, and the security protection effect is limited.

[0008] UAV (Unmanned Aerial Vehicle) energy consumption control is lacking: Existing research has not fully considered the energy consumption of UAV hovering and movement, and has not included energy consumption in the optimization objectives, resulting in low energy efficiency of the system during long-term operation, making it difficult to meet the needs of practical applications. Summary of the Invention

[0009] To address the aforementioned shortcomings of existing technologies, this invention provides a RIS-assisted UAV-ISAC system security and energy efficiency optimization method, which solves the robustness problem in scenarios with imperfect CSI (channel state information), the coordination problem between physical layer security and channel estimation, and the energy consumption optimization problem of UAVs (unmanned aerial vehicles).

[0010] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows:

[0011] A method for optimizing the safety and energy efficiency of a RIS-assisted UAV-ISAC system is provided, comprising:

[0012] Step S1: Collect time data using the UAV subsystem. t Link channel vector between UAV and legitimate users Link channel matrix of UAV-Reconfigurable Smart Surface RIS Reconfigurable Smart Surface RIS - Reflection Link Channel Vector for Legitimate Users ; Construction Time t UAV-eavesdropper link channel vector ;

[0013] Step S2: Based on time t -1 Channel posterior distribution and channel transition model, prediction time t Channel prior, and based on channel prior covariance The channel posterior distribution is calculated based on Kalman filtering to obtain the time... t The channel posterior covariance;

[0014] Step S3: Construct a multi-agent PPO decision network. The multi-agent PPO decision network includes a deep reinforcement learning model (DRL) and includes decision-making by communication beamforming agents, UAV agents, and perception agents.

[0015] Reflection coefficient matrix for optimizing reconfigurable smart surface RIS reflection units The beamforming matrix G[t] is calculated, the trajectory Q of the UAV and the update rate of the channel state information (CSI) are optimized, and the sensing beam is calculated; and the reward coefficient of the multi-agent PPO decision network is calculated.

[0016] Step S4: The multi-agent PPO decision network generates the beamforming matrix of the artificial noise AN, and combines it with the optimized beamforming matrix. Generate the total transmission signal;

[0017] Step S5: The multi-agent PPO decision network outputs the total security rate, energy efficiency, root mean square error of positioning, and eavesdropper capacity based on the total transmitted signal, and optimizes the control parameters of the multi-agent PPO decision network to output a converged multi-agent PPO decision network, thereby optimizing the communication rate, energy efficiency, and security performance of the system.

[0018] Further, step S2 includes:

[0019] Step S21: According to time t -1 Channel posterior distribution and channel transition model, prediction time t Channel priors;

[0020] ;

[0021] in, For a moment t The prior mean of the channel, For a moment t Channel prior covariance calculation F For the pose change of the unmanned aerial vehicle (UAV), This is the posterior estimate of the state at the previous time t–1. Let be the covariance matrix of the state estimate from the previous time step. Indicates the change in pose F Hermit transpose, For noise covariance;

[0022] Step S22: Based on the channel prior covariance Calculate the channel posterior distribution based on Kalman filtering;

[0023] ;

[0024] in, For Kalman gain, For a moment t The channel posterior mean, For a moment t Channel posterior covariance calculation R To measure the noise covariance, For the link channel vector Link channel vector Link channel matrix and reflection link channel vector The observation matrix formed Echo signals collected by the UAV subsystem For a moment t -1 is the channel posterior mean. Observation matrix Hermit transpose, It is the identity matrix. For a moment t The channel posterior covariance.

[0025] Furthermore, the communication beamforming agent decision is used to optimize the reflection coefficient matrix of the reconfigurable smart surface RIS reflector unit. and beamforming matrix G[t]; specifically:

[0026] State input vector for communication beamforming agent decision-making for:

[0027] ;

[0028] in, For the system at time t The true state vector, , For legitimate users or eavesdroppers at any time t Location, i The identifier is used to identify legitimate users or eavesdroppers as targets.

[0029] Reflection coefficient matrix of reconfigurable smart surface RIS reflective unit for:

[0030] ;

[0031] in, For the first M The amplitude reflection coefficient of each reflection unit, and the amplitude reflection coefficient of the reconstructed intelligent surface RIS reflection unit satisfy the following conditions: , m This refers to the number of the reflecting unit. For the first M Phase offset of each reflecting unit j Represents the imaginary unit. e The phase shift of the reflecting unit is the base of the natural logarithm. ;

[0032] The reflection coefficient matrix Decompose into real part With the imaginary part The beamforming matrix G[t] is decomposed into its imaginary part. Wajitsube The beamforming vector of the transmitted signal of a legitimate user in the beamforming matrix G[t] The constraints are satisfied:

[0033] ;

[0034] in, The power ratio of artificial noise. This represents the total transmission power.

[0035] The reward calculation model for communication beamforming agent decision-making is as follows:

[0036] ;

[0037] ;

[0038] ;

[0039] ;

[0040] in, The reward coefficient for decision-making by the communication beamforming agent. For the confidentiality rate of legitimate users, These are penalties for violating UAV (Unmanned Aerial Vehicle) mobility constraints, penalties for violating total power constraints, penalties for violating quality of service constraints, and energy consumption penalties. The coefficient for the penalty term. For legitimate users' communication rate, The communication rate between the eavesdropper and the legitimate user. p Number the eavesdroppers. k For legitimate users' IDs, For legitimate users, the signal-to-interference-plus-noise ratio (SIR) is... The noise-to-interference ratio is used to communicate with eavesdroppers.

[0041] Furthermore, the method for optimizing the beamforming matrix G[t] by the communication beamforming agent is as follows:

[0042] The communication beamforming agent decision-making process incorporates a deep reinforcement learning (DRL) model, which takes the state input vector as input. Features are extracted using the fully connected layer of the deep reinforcement learning model DRL, and the beamforming matrix G[t] is output. The dimensions of the beamforming matrix G[t] are... Each element in the beamforming matrix G[t] is represented as:

[0043] ;

[0044] in, For dimension Beamforming matrix, Beamforming matrix The real and imaginary parts;

[0045] Reassemble each element of the beamforming matrix G[t] into a complex beamforming matrix according to the transmit antenna index minus the valid user index. Complex beamforming matrix elements in Represented as;

[0046] ;

[0047] in, elements respectively The real and imaginary parts;

[0048] From complex beamforming matrix Extract the initial beamforming vector of the legitimate user and calculate the communication power between the transmitting antenna and the legitimate user. ;

[0049] ;

[0050] in, For the first k Initial beamforming vectors for each legitimate user;

[0051] like Then the beamforming matrix G[t] is scaled as follows:

[0052] ;

[0053] in, This represents the power ratio of artificial noise AN. This represents the total transmit power of the system.

[0054] Otherwise, retain the beamforming matrix G[t] and complete the optimization of the beamforming matrix G[t].

[0055] Furthermore, the communication beamforming agent makes decisions to optimize the reflection coefficient matrix. The method is as follows:

[0056] Communication beamforming agent decision output reflection coefficient matrix Reflection coefficient matrix Each element in the equation represents the complex reflection coefficient. ;

[0057] The output layer of the Deep Reinforcement Learning (DRL) model is designed as a 2M-dimensional real vector, and the output vector... ;

[0058] ;

[0059] in, The first M The real and imaginary parts of the beamforming vector corresponding to each reflection unit;

[0060] Using the real part of the beamforming vector of the reflecting unit and the virtual part Calculate the modulus ;

[0061] ;

[0062] like Then the real part of the beamforming vector is... and the virtual part Divide by the modulus respectively The real part of the correction is obtained. and the virtual part ;

[0063] ;

[0064] Using the real part of the correction and the virtual part Inverse phase offset Using the reverse phase offset Corrected reflection coefficient matrix The phase offset of each element in the matrix is ​​used to complete the reflection coefficient matrix. Optimization;

[0065] Otherwise, maintain the reflection coefficient matrix. constant.

[0066] Furthermore, the UAV agent's decision-making is used to optimize the UAV's trajectory Q;

[0067] State input vector for UAV agent decision-making for:

[0068] ;

[0069] in, For a moment t Location of UAVs For a moment t speed, For a moment t -1 is the root mean square error of UAV positioning. For a moment t -1 UAV energy efficiency.

[0070] The boundary constraints for the UAV's planar position are: ;in, B For planar position boundaries, These represent the horizontal and vertical positions of the UAV (Unmanned Aerial Vehicle).

[0071] The positional constraints of the UAV are: ;

[0072] The reward calculation model for UAV agent decision-making is as follows:

[0073] ;

[0074] in, The reward coefficient for UAV (Unmanned Aerial Vehicle) agent decision-making.

[0075] Furthermore, the perceptual agent's decision-making is used to optimize the update rate of Channel State Information (CSI) and the perceptual beam;

[0076] The state input vector for the perceptual agent's decision-making is:

[0077] ;

[0078] in, For legitimate users or eavesdroppers at any time t -1 target estimated location;

[0079] The multi-agent PPO decision network outputs Channel State Information (CSI) to update weights and sensing beam direction. This affects the amount of observational information and the quality of the Kalman update;

[0080] The reward calculation model for decision-making by a perceptual agent is as follows:

[0081] ;

[0082] ;

[0083] in, The reward coefficient for the decision-making of the perceptual agent. For the perception of resource penalty items, For the coefficient of the perceived resource penalty term, For the total security rate, For a moment t The root mean square error of UAV positioning. This represents the union of the set of legitimate users and the set of eavesdroppers. For legitimate users or eavesdroppers at any time t The position of +1 To perceive error, As a weight for the rate of secrecy, The weights for perceived error.

[0084] Further, step S4 includes:

[0085] Step S41: The deep reinforcement learning model DRL in the multi-agent PPO decision network is based on time... t channel posterior mean Location of the eavesdropper The power ratio of the output artificial noise AN ;

[0086] Step S42: The artificial noise module calculates the power of the artificial noise AN. Generate the null projection matrix of artificial noise And construct the beamforming matrix of artificial noise AN. ;

[0087] in, It is by K A matrix formed by concatenating the channel vectors of each legitimate user. For the first K The channel vector of a legitimate user.

[0088] Step S43: Based on the optimized beamforming matrix Beamforming matrix of artificial noise AN The UAV subsystem transmits communication signals. With artificial noise AN signal , Generate the total transmission signal for the information symbols. .

[0089] Further, step S5 includes:

[0090] Step S51: The multi-agent PPO decision network is based on the total transmitted signal. Output time t Total security rate Energy efficiency Root mean square error of positioning and eavesdropper capacity , and the security rate threshold Energy efficiency threshold Sensing accuracy threshold Compare;

[0091] like or or If the result is positive, proceed to step S52; otherwise, determine that the multi-agent PPO decision network has converged.

[0092] Step S52: Correct the control parameters of the agent PPO decision network and adjust the weights of the reward calculation model, then return to step S51. The multi-agent PPO decision network re-outputs the total security rate. Energy efficiency Root mean square error of positioning and eavesdropper capacity ;

[0093] Step S53: Iterate through step S52, calculating the total reward coefficient for each iteration of the agent PPO decision network. The total reward coefficient up to 100 consecutive iterations. Volatility If the PPO decision network of the agent is converged, the loop iteration is stopped.

[0094] , ;

[0095] Step S54: Output the control parameters when the intelligent agent PPO decision network converges, as the optimal control parameters. The optimal control parameters include the optimal phase of the reflection unit, the optimal trajectory of the UAV, the optimal beamforming matrix, and the optimal artificial noise AN power ratio, thereby optimizing the communication rate, energy efficiency, and safety performance of the system.

[0096] The beneficial effects of this invention are as follows: This invention introduces a Bayesian channel estimation mechanism to improve the robustness of the system in scenarios with imperfect CSI (channel state information) and ensure the stability of channel uncertainty RMSE and SEE.

[0097] A three-agent DRL architecture is constructed to decouple multiple optimization variables and achieve multi-objective collaboration of "communication security, perception accuracy, and energy efficiency optimization". Under the premise of ensuring RMSE≤2m (perception accuracy threshold), the system can improve SSR by more than 20% and SEE by more than 15%.

[0098] The PPO algorithm is adopted as the core of the intelligent agent to improve the convergence speed and stability of the algorithm, ensure rapid convergence, and avoid local optima in high-dimensional action space, so as to meet the millisecond-level real-time decision-making requirements in dynamic environments.

[0099] By achieving synergy between Bayesian CSI (Channel State Information) estimation and AN (Artificial Noise) mechanism, and dynamically optimizing the AN (Artificial Noise) power ratio and beam direction, targeted suppression of eavesdroppers is achieved, reducing the eavesdropper's channel capacity by more than 30% and improving the physical layer security protection effect.

[0100] Establish a UAV (unmanned aerial vehicle) energy consumption model, incorporate hovering and motion energy consumption into the optimization objective, maximize system SEE, and ensure the energy efficiency requirements of UAV (unmanned aerial vehicle) during long-term operation under the constraint of 30-50m altitude and 10m / s maximum speed. Attached Figure Description

[0101] Figure 1 A schematic diagram of the structure of a RIS-assisted UAV-ISAC system. Detailed Implementation

[0102] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0103] like Figure 1 As shown, a RIS-assisted UAV-ISAC system safety and energy efficiency optimization method is characterized by comprising:

[0104] Step S1: Collect time data using the UAV subsystem. t Link channel vector between UAV and legitimate users Link channel matrix of UAV-Reconfigurable Smart Surface RIS Reconfigurable Smart Surface RIS - Reflection Link Channel Vector for Legitimate Users ; Construction Time t UAV-eavesdropper link channel vector .

[0105] Step S1 specifically includes:

[0106] Time data collection using UAV subsystem t Link channel vector between UAV and legitimate users Link channel matrix of UAV-Reconfigurable Smart Surface RIS Reconfigurable Smart Surface RIS - Reflection Link Channel Vector for Legitimate Users ;

[0107] Link channel vector The latitude is , K For the number of legitimate users, Number of transmit antennas for UAVs; link channel matrix The latitude is , M The number of reflection units in the reconfigurable smart surface RIS; the reflection link channel vector. The latitude is ;

[0108] Build Time t UAV-eavesdropper link channel vector Link channel vector The latitude is , P The number of eavesdroppers;

[0109] Step S2: Based on time t -1 Channel posterior distribution and channel transition model, prediction time t Channel prior, and based on channel prior covariance The channel posterior distribution is calculated based on Kalman filtering to obtain the time... t The channel posterior covariance.

[0110] Step S2 specifically includes:

[0111] Step S21: According to time t -1 Channel posterior distribution and channel transition model, prediction time t Channel priors;

[0112] ;

[0113] in, For a moment t The prior mean of the channel, For a moment t Channel prior covariance calculation F For the pose change of the unmanned aerial vehicle (UAV), This is the posterior estimate of the state at the previous time t–1. Let be the covariance matrix of the state estimate from the previous time step. Indicates the change in pose F Hermit transpose, For noise covariance;

[0114] Step S22: Based on the channel prior covariance Calculate the channel posterior distribution based on Kalman filtering;

[0115] ;

[0116] in, For Kalman gain, For a moment t The channel posterior mean, For a moment t Channel posterior covariance calculation R To measure the noise covariance, For the link channel vector Link channel vector Link channel matrix and reflection link channel vector The observation matrix formed Echo signals collected by the UAV subsystem For a moment t -1 is the channel posterior mean. Observation matrix Hermit transpose, It is the identity matrix. For a moment t The channel posterior covariance.

[0117] Step S3: Construct a multi-agent PPO decision network, which includes a deep reinforcement learning model (DRL). The multi-agent PPO decision network comprises decision-making by communication beamforming agents, UAV agents, and perception agents; this is used to optimize the reflection coefficient matrix of the reconfigurable intelligent surface RIS reflective unit. And the beamforming matrix G[t]; and calculate the reward coefficient of the multi-agent PPO decision network.

[0118] State input vector for communication beamforming agent decision-making for:

[0119] ;

[0120] in, For the system at time t The true state vector, , For legitimate users or eavesdroppers at any time t Location, i The identifier is used to identify legitimate users or eavesdroppers as targets. The state estimate covariance after being updated by the Kalman filter is at time [time value missing]. t The channel posterior covariance;

[0121] Reflection coefficient matrix of reconfigurable smart surface RIS reflective unit for:

[0122] ;

[0123] in, For the first M The amplitude reflection coefficient of each reflection unit, and the amplitude reflection coefficient of the reconstructed intelligent surface RIS reflection unit satisfy the following conditions: , m This refers to the number of the reflecting unit. For the first M Phase offset of each reflecting unit j Represents the imaginary unit. e The phase offset of the reflecting unit is a natural constant. ;

[0124] The reflection coefficient matrix Decompose into real part With the imaginary part The beamforming matrix G[t] is decomposed into its imaginary part. Wajitsube The beamforming vector of the transmitted signal of a legitimate user in the beamforming matrix G[t] The constraints are satisfied:

[0125] ;

[0126] in, The power ratio of artificial noise. This represents the total transmission power.

[0127] The reward calculation model for communication beamforming agent decision-making is as follows:

[0128] ;

[0129] ;

[0130] ;

[0131] ;

[0132] in, The reward coefficient for decision-making by the communication beamforming agent. For the confidentiality rate of legitimate users, These are penalties for violating UAV (Unmanned Aerial Vehicle) mobility constraints, penalties for violating total power constraints, penalties for violating quality of service constraints, and energy consumption penalties. The coefficient for the penalty term. For legitimate users' communication rate, The communication rate between the eavesdropper and the legitimate user. p Number the eavesdroppers. k For legitimate users' IDs, For legitimate users, the signal-to-interference-plus-noise ratio (SIR) is... The noise-to-interference ratio is used to communicate with eavesdroppers.

[0133] The multi-agent PPO decision network introduces a complex variable decomposition mechanism on the basis of traditional algorithms, enabling it to output, reconstruct and optimize complex matrices. At the same time, power normalization and phase correction steps are added to the output layer to enable the application of deep reinforcement learning models in complex signal space.

[0134] The method for optimizing the beamforming matrix G[t] by a communication beamforming agent is as follows:

[0135] The communication beamforming agent decision-making process incorporates a deep reinforcement learning (DRL) model, which takes the state input vector as input. Features are extracted using the fully connected layer of the deep reinforcement learning model DRL, and the beamforming matrix G[t] is output. The dimensions of the beamforming matrix G[t] are... Each element in the beamforming matrix G[t] is represented as:

[0136] ;

[0137] in, For dimension Beamforming matrix, Beamforming matrix The real and imaginary parts;

[0138] Reassemble each element of the beamforming matrix G[t] into a complex beamforming matrix according to the transmit antenna index minus the valid user index. Complex beamforming matrix elements in Represented as;

[0139] ;

[0140] in, elements respectively The real and imaginary parts;

[0141] From complex beamforming matrix Extract the initial beamforming vector of the legitimate user and calculate the communication power between the transmitting antenna and the legitimate user. ;

[0142] ;

[0143] in, For the first k Initial beamforming vectors for each legitimate user;

[0144] like Then the beamforming matrix G[t] is scaled as follows:

[0145] ;

[0146] in, This represents the power ratio of artificial noise AN. This represents the total transmit power of the system.

[0147] Otherwise, retain the beamforming matrix G[t] and complete the optimization of the beamforming matrix G[t].

[0148] Communication beamforming agent decision optimization reflection coefficient matrix The method is as follows:

[0149] Communication beamforming agent decision output reflection coefficient matrix Reflection coefficient matrix Each element in the equation represents the complex reflection coefficient. ;

[0150] The output layer of the Deep Reinforcement Learning (DRL) model is designed as a 2M-dimensional real vector, and the output vector... ;

[0151] ;

[0152] in, The first M The real and imaginary parts of the beamforming vector corresponding to each reflection unit;

[0153] Using the real part of the beamforming vector of the reflecting unit and the virtual part Calculate the modulus ;

[0154] ;

[0155] like Then the real part of the beamforming vector is... and the virtual part Divide by the modulus respectively The real part of the correction is obtained. and the virtual part ;

[0156] ;

[0157] Using the real part of the correction and the virtual part Inverse phase offset Using the reverse phase offset Corrected reflection coefficient matrix The phase offset of each element in the matrix is ​​used to complete the reflection coefficient matrix. Optimization;

[0158] Otherwise, maintain the reflection coefficient matrix. constant.

[0159] The output beamforming matrix G[t] determines the directional focusing of the UAV (unmanned aerial vehicle) on the user signal, facilitating the calculation of the signal-to-dryness ratio and the reflection coefficient matrix. By controlling the reflection phase of RIS (Reconfigurable Smart Surface), signal enhancement and eavesdropping suppression can be achieved.

[0160] The UAV agent's decision-making is used to optimize the UAV's trajectory Q;

[0161] State input vector for UAV agent decision-making for:

[0162] ;

[0163] in, For a moment t Location of UAVs For a moment t speed, For a moment t -1 is the root mean square error of UAV positioning. For a moment t -1 UAV energy efficiency.

[0164] The boundary constraints for the UAV's planar position are: ;in, B For planar position boundaries, These represent the horizontal and vertical positions of the UAV (Unmanned Aerial Vehicle).

[0165] The positional constraints of the UAV are: ;

[0166] The reward calculation model for UAV agent decision-making is as follows:

[0167] ;

[0168] in, The reward coefficient for UAV (Unmanned Aerial Vehicle) agent decision-making.

[0169] The perceptual agent's decision-making is used to optimize the update rate of Channel State Information (CSI) and the perceptual beam.

[0170] The state input vector for the perceptual agent's decision-making is:

[0171] ;

[0172] in, For legitimate users or eavesdroppers at any time t -1 target estimated location;

[0173] The multi-agent PPO decision network outputs Channel State Information (CSI) to update weights and sensing beam direction. This affects the amount of observational information and the quality of the Kalman update;

[0174] The reward calculation model for decision-making by a perceptual agent is as follows:

[0175] ;

[0176] ;

[0177] in, The reward coefficient for the decision-making of the perceptual agent. For the perception of resource penalty items, For the coefficient of the perceived resource penalty term, For the total security rate, For a moment t The root mean square error of UAV positioning. This represents the union of the set of legitimate users and the set of eavesdroppers. For legitimate users or eavesdroppers at any time t The position of +1 To perceive error, As a weight for the rate of secrecy, The weights for perceived error.

[0178] Step S4: The multi-agent PPO decision network generates the beamforming matrix of the artificial noise AN, and combines it with the optimized beamforming matrix. Generate the total transmission signal. Step S4 specifically includes:

[0179] Step S41: The deep reinforcement learning model DRL in the multi-agent PPO decision network is based on time... t channel posterior mean Location of the eavesdropper The power ratio of the output artificial noise AN ;

[0180] Step S42: The artificial noise module calculates the power of the artificial noise AN. Generate the null projection matrix of artificial noise And construct the beamforming matrix of artificial noise AN. ;

[0181] in, It is by K A matrix formed by concatenating the channel vectors of each legitimate user. For the first K The channel vector of a legitimate user.

[0182] Step S43: Based on the optimized beamforming matrix Beamforming matrix of artificial noise AN The UAV subsystem transmits communication signals. With artificial noise AN signal , Generate the total transmission signal for the information symbols. .

[0183] Step S5: The multi-agent PPO decision network is based on the total transmitted signal. The system outputs the total security rate, energy efficiency, root mean square error of positioning, and eavesdropper capacity, and optimizes the control parameters of the multi-agent PPO decision network, outputting a converged multi-agent PPO decision network to optimize the system's communication rate, energy efficiency, and security performance.

[0184] Step S5 specifically includes:

[0185] Step S51: The multi-agent PPO decision network is based on the total transmitted signal. Output time t Total security rate Energy efficiency Root mean square error of positioning and eavesdropper capacity , and the security rate threshold Energy efficiency threshold Sensing accuracy threshold Compare;

[0186] like or or If the result is positive, proceed to step S8; otherwise, determine that the multi-agent PPO decision network has converged.

[0187] Step S52: Correct the control parameters of the agent PPO decision network and adjust the weights of the reward calculation model, then return to step S51. The multi-agent PPO decision network re-outputs the total security rate. Energy efficiency Root mean square error of positioning and eavesdropper capacity ;

[0188] Step S53: Iterate through step S52, calculating the total reward coefficient for each iteration of the agent PPO decision network. The total reward coefficient up to 100 consecutive iterations. Volatility If the PPO decision network of the agent is converged, the loop iteration is stopped.

[0189] , ;

[0190] The policy network and value network of each agent are updated based on the PPO algorithm: the advantage function is calculated by time-series differential error, and the clip mechanism is used to limit the policy update step size to avoid oscillation and improve convergence stability; after each round (5000 rounds), the Bayesian channel estimation is updated to optimize the filtering accuracy.

[0191] Step S54: Output the control parameters when the intelligent agent PPO decision network converges, as the optimal control parameters. The optimal control parameters include the optimal phase of the reflection unit, the optimal trajectory of the UAV, the optimal beamforming matrix, and the optimal artificial noise AN power ratio, thereby optimizing the communication rate, energy efficiency, and safety performance of the system.

[0192] This invention introduces a Bayesian channel estimation mechanism to improve the robustness of the system in scenarios with imperfect CSI (channel state information) and ensure the stability of channel uncertainty RMSE and SEE.

[0193] A three-agent DRL architecture is constructed to decouple multiple optimization variables and achieve multi-objective collaboration of "communication security, perception accuracy, and energy efficiency optimization". Under the premise of ensuring RMSE≤2m (perception accuracy threshold), the system can improve SSR by more than 20% and SEE by more than 15%.

[0194] The PPO algorithm is adopted as the core of the intelligent agent to improve the convergence speed and stability of the algorithm, ensure rapid convergence, and avoid local optima in high-dimensional action space, so as to meet the millisecond-level real-time decision-making requirements in dynamic environments.

[0195] By achieving synergy between Bayesian CSI (Channel State Information) estimation and AN (Artificial Noise) mechanism, and dynamically optimizing the AN (Artificial Noise) power ratio and beam direction, targeted suppression of eavesdroppers is achieved, reducing the eavesdropper's channel capacity by more than 30% and improving the physical layer security protection effect.

[0196] Establish a UAV (unmanned aerial vehicle) energy consumption model, incorporate hovering and motion energy consumption into the optimization objective, maximize system SEE, and ensure the energy efficiency requirements of UAV (unmanned aerial vehicle) during long-term operation under the constraint of 30-50m altitude and 10m / s maximum speed.

Claims

1. A method for RIS-aided UAV-ISAC system security and energy efficiency optimization, characterized in that, Comprise: Step S1: collecting time instants with the UAV subsystem t Link channel vector between UAV - legitimate user Link channel matrix between UAV - reconfigurable intelligent surface (RIS) Reflection link channel vector between RIS - legitimate user Constructing time instants t Link channel vector between UAV - eavesdropper ; Step S2: according to time t -1 channel posterior distribution and channel transfer model, predict time t channel prior, and according to the channel prior covariance Based on Kalman filtering to calculate the channel posterior distribution, get time t Channel posterior covariance of time Step S3: build a multi-agent PPO decision network, a deep reinforcement learning model DRL is built in the multi-agent PPO decision network, the multi-agent PPO decision network includes communication beamforming agent decision, unmanned aerial vehicle UAV agent decision and perception agent decision; Optimizing reflection coefficient matrix for reconfigurable intelligent surface, ris, reflecting elements and a beamforming matrix G[t], an update rate of channel state information, CSI, and a sensing beam for optimizing a trajectory Q of an unmanned aerial vehicle, UAV. And calculate the reward coefficient of the multi-agent PPO decision network; Step S4: The multi-agent PPO decision network generates a beamforming matrix of artificial noise AN, and combines the optimized beamforming matrix generating a total transmit signal; Step S5: the multi-agent PPO decision network outputs total secrecy rate, energy efficiency, positioning root mean square error and eavesdropper capacity based on total transmitted signal, and optimizes the control parameters of the multi-agent PPO decision network, outputs the converged multi-agent PPO decision network, and optimizes the communication rate, energy efficiency and security performance of the system; The communication beamforming agent decides for optimizing a reflection coefficient matrix of reconfigurable intelligent surface, RIS and a beamforming matrix G[t]; in particular: State input vector for communication beamforming agent decision making For: ; wherein, is the real state vector of the system at time t , is the position of the legitimate user or the eavesdropper at time t i is the number of the legitimate user or the eavesdropper as a target;​​ Reconfigurable intelligent surface, ris, reflection coefficient matrix is: ; wherein, is the amplitude reflection coefficient of the nth reflection unit, M , m is the number of the reflection unit, is the phase offset of the nth reflection unit, M denotes the imaginary unit, j is the base of the natural logarithm, e is the phase offset of the reflection unit ;​ The reflection coefficient matrix Decompose into real part With the imaginary part The beamforming matrix G[t] is decomposed into its imaginary part. Wajitsube The beamforming vector of the transmitted signal of a legitimate user in the beamforming matrix G[t] The constraints are satisfied: ; wherein, Pnoise is the power of the artificial noise, Ptotai is the total transmit power; The reward calculation model of the communication beamforming agent decision is: ; ; ; ; wherein, is a reward coefficient for the communication beamforming intelligent agent decision, is a secrecy rate of a legitimate user, are respectively a penalty term for violating a unmanned aerial vehicle (UAV) movement constraint, a penalty term for violating a total power constraint, a penalty term for violating a quality of service constraint, and an energy consumption penalty term, is a penalty term coefficient, is a communication rate of a legitimate user, is a communication rate of a legitimate user by an eavesdropper, p is an eavesdropper number, k is a legitimate user number, is a signal-to-interference-plus-noise ratio (SINR) of a legitimate user, is an eavesdropper SINR; The method for the communication beamforming agent decision to optimize the beamforming matrix G[t] is: The communication beamforming intelligent agent decision is built with a deep reinforcement learning model DRL, and a state input vector Features are extracted through a full connection layer of the deep reinforcement learning model DRL, and a beamforming matrix G[t] is outputted, the latitude of the beamforming matrix G[t] is Each element in the beamforming matrix G[t] is expressed as: ; wherein is a beamforming matrix of dimension is a beamforming matrix of dimension are the real and imaginary parts of the beamforming matrix are the real and imaginary parts of the beamforming matrix reorganizing each element in the beamforming matrix G[t] according to the transmit antenna index - the legitimate user index into a complex beamforming matrix , the elements in the complex beamforming matrix are denoted as ; ; wherein are the real and imaginary parts, respectively, of the element ​ extracting an initial beamforming vector of a legitimate user from a complex beamforming matrix computing a communication power between a transmitting antenna and the legitimate user ; ; wherein, is the initial beamforming vector for the k first legitimate user; If then the beamforming matrix G[t] is scaled to: ; wherein is a power ratio of artificial noise AN, is the total transmit power of the system; Otherwise, the beamforming matrix G[t] is retained, and the optimization of the beamforming matrix G[t] is completed. The communication beamforming agent decision optimization reflection coefficient matrix The method is: Communication beamforming agent decision output reflection coefficient matrix , reflection coefficient matrix each element is a complex reflection coefficient term ; The output layer of the deep reinforcement learning model DRL is designed as a 2M dimensional real-valued vector, the output vector ; ; wherein are the real and imaginary parts of the beamforming vector corresponding to the M first and second reflection unit, respectively. Utilizing real parts of beamforming vectors of reflection elements and imaginary parts calculating a modulus ; ; If , then the real part and the imaginary part of the beamforming vector are divided by the modulus , respectively, to obtain the corrected real part and the corrected imaginary part ; ; using the corrected real parts and imaginary parts back-calculate phase offsets , using the back-calculated phase offsets correct the reflection coefficient matrix the phase offsets of each element in the corrected reflection coefficient matrix optimization is completed Otherwise, keep the reflection coefficient matrix unchanged; The unmanned aerial vehicle UAV agent decision is used to optimize the trajectory Q of the unmanned aerial vehicle UAV; State input vector for unmanned aerial vehicle (UAV) agent decision making is: ; wherein, is the UAV position at time t is the velocity at time t is the UAV position root mean square error at time t is the UAV energy efficiency at time t -1.​​​ The boundary constraint of the planar position of the UAV is: ; wherein, B is the planar position boundary, respectively, is the planar lateral position and the planar vertical position of the UAV. The position constraint for the UAV is: ; The reward calculation model of the unmanned aerial vehicle UAV agent decision is: ; wherein, is a reward coefficient for the UAV agent decision; The perception agent decision is used to optimize the update rate of channel state information CSI and the perception beam; The state input vector of the perception agent decision is: ; wherein, the estimated position of the target at time t -1 for a legitimate user or an eavesdropper. Multi-agent PPO decision network outputs channel state information, CSI update weight and perception beam direction to affect the amount of observation information and Kalman update quality; The reward calculation model of the perception agent decision is: ; ; wherein, is a reward coefficient for the perception agent decision, is a perception resource penalty term, is a perception resource penalty term coefficient, is a total secrecy rate, is a time instant t is a UAV positioning root mean square error at time instant denotes the union of the set of legitimate users and the set of eavesdroppers, is the position of a legitimate user or eavesdropper at time instant t + 1, is a perception error, is a weight for the secrecy rate, is a weight for the perception error; The step S4 comprises: Step S41: The deep reinforcement learning model DRL in the multi-agent PPO decision network is based on time... t channel posterior mean Location of the eavesdropper The power ratio of the output artificial noise AN ; Step S42: The artificial noise module calculates the power of the artificial noise AN , generates the null space projection matrix of the artificial noise , and constructs the beamforming matrix of the artificial noise AN ; wherein is a matrix formed by concatenating the channel vectors of the K legal users, is the channel vector of the K th legal user; Step S43: generating a total transmit signal based on the optimized beamforming matrix and artificial noise AN beamforming matrix , the UAV subsystem transmits a communication signal with the artificial noise AN signal , for the information symbols ; The step S5 comprises: Step S51: The multi-agent PPO decision network is based on the total transmitted signal Output time t The total secrecy rate , energy efficiency , positioning root mean square error And the capacity of the eavesdropper , and the secrecy rate threshold , energy efficiency threshold , perception accuracy threshold Comparison; If or or Step S52 is executed, otherwise, it is determined that the multi-agent PPO decision network converges. Step S52: correct the control parameters of the agent PPO decision network and adjust the weights of the reward calculation model, and return to step S51, and the multi-agent PPO decision network re-outputs the total secrecy rate , energy efficiency , positioning root mean square error , and eavesdropper capacity ; Step S53: loop iteration step S52, each loop iteration process intelligent agent PPO decision network calculates total reward coefficient ; until the total reward coefficient of 100 consecutive iteration loops volatility , it is determined that the intelligent agent PPO decision network converges, and the loop iteration is stopped; , ; Step S54: output the control parameters of the converged agent PPO decision network as the optimal control parameters, the optimal control parameters include the optimal phase of the reflecting element, the optimal trajectory of the unmanned aerial vehicle, the optimal beamforming matrix and the optimal artificial noise AN power ratio, and the optimization of the communication rate, energy efficiency and security performance of the system is completed.

2. The RIS-assisted UAV-ISAC system safety and energy efficiency optimization method of claim 1, wherein, The step S2 comprises: Step S21: According to the time t -1 channel posterior distribution and channel transition model, predict the time t Channel prior; ; in, For a moment t The channel prior mean, For a moment t Channel prior covariance calculation F This represents the pose change of the unmanned aerial vehicle (UAV). This is the posterior estimate of the state at the previous time t–1. Let be the covariance matrix of the state estimate from the previous time step. Indicates the change in pose F Hermit transpose, For noise covariance; Step S22: Calculate the channel posterior covariance according to the channel prior covariance Calculate the channel posterior distribution based on Kalman filtering; ; wherein is the Kalman gain, is the channel posterior mean at time t is the channel posterior covariance operation at time t R is the measurement noise covariance, is the observation matrix composed of the link channel vector , the link channel vector , the link channel matrix and the reflected link channel vector is the echo signal collected by the UAV subsystem, is the channel posterior mean at time t -1, is the Hermitian transpose of the observation matrix is the identity matrix, is the channel posterior covariance at time t .​​​​

Citation Information

Patent Citations

  • Reconfigurable intelligent surface assisted communication and positioning integrated full duplex system

    CN115022146A

  • RIS-assisted unmanned aerial vehicle network-based communication and sensing integrated system and method

    CN119277319A