A secure communication method for unmanned aerial vehicles assisted by intelligent reflective surfaces

By optimizing the joint optimization model of intelligent reflective surface phase shift and drone trajectory, the problem of signal blocking and eavesdropping of drone communication systems in complex environments is solved, and safety and energy efficiency are improved.

CN116996867BActive Publication Date: 2025-08-22JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311049636.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-21
Publication Date
2025-08-22
Estimated Expiration
2043-08-21

AI Technical Summary

Technical Problem

UAV communication systems are susceptible to obstacles in dense user areas and lead to signal blockage, and wireless transmissions are susceptible to interference and eavesdropping. The existing encryption methods consume high computing resources, making it difficult to ensure communication security and efficiency in complex environments.

Method used

By optimizing the phase shift of the intelligent reflection surface and the drone movement trajectory, using reinforcement learning and alternating optimization algorithms to design a joint optimization model between the drone and the intelligent reflection surface, create a virtual line of sight link, enhance legitimate user signals, weaken eavesdropper signals, and reduce drone energy consumption.

Benefits of technology

It improves the confidentiality rate of legitimate users and the energy efficiency of drones, enhances the security and reliability of wireless networks, reduces the risk of eavesdropping, and reduces the energy consumption of drones propulsion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116996867B_ABST
    Figure CN116996867B_ABST
Patent Text Reader

Abstract

The present invention discloses a secure communication method for unmanned aerial vehicles (UAVs) assisted by an intelligent reflecting surface, comprising the following steps: first, obtaining the positions of the intelligent reflecting surface and each legitimate user, and obtaining the position of each eavesdropper; second, establishing an optimization objective function, and optimizing the phase shift of the intelligent reflecting surface and the movement trajectory of the UAV according to the optimization objective function, to obtain the optimal phase shift of the intelligent reflecting surface and the optimal position of the UAV at each moment; wherein the optimization objective function is: #imgabs0# wherein #imgabs1##imgabs2# represent the movement trajectory of the UAV, #imgabs3# represent the phase shift change of the intelligent reflecting surface, #imgabs4# represent the time series set, and q U [t] represents the position of the UAV at time t, Φ[t] represents the phase shift of the smart reflective surface at time t, R sec [t] represents the sum of the confidentiality rates of all legitimate users at time t; e[t] represents the flight energy consumption of the drone at time t; 3. At each moment, the intelligent reflective surface is adjusted to the optimal phase shift, the drone moves to the optimal position and sends communication data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of unmanned aerial vehicle (UAV) secure communication, and in particular relates to a UAV secure communication method assisted by an intelligent reflective surface. Background Art

[0002] Compared to traditional ground-based fixed base stations, communication systems using drones as aerial base stations offer advantages such as flexible deployment, controllable mobility, and low cost. They can provide more reliable and efficient wireless communication services to users on the ground. However, drone communications still face several challenges. First, in densely populated areas, obstacles and user capacity can block communication signals between drones and ground users, significantly degrading system performance. Second, due to the broadcast nature of wireless transmission, drone communication systems are vulnerable to interference and eavesdropping attacks. Upper-layer encryption can enable confidential communications in wireless networks, but the frequent encryption and decryption require high computing power, posing a significant challenge for drones with limited hardware resources. Artificial noise and traditional beamforming can achieve secure communications, but their effectiveness is significantly reduced in certain scenarios, such as when legitimate users and eavesdroppers are close to each other or facing the same direction. Summary of the Invention

[0003] The purpose of the present invention is to provide a method for safe communication of unmanned aerial vehicles (UAVs) assisted by intelligent reflective surfaces, which can improve the security of the communication system and reduce the propulsion energy consumption of the UAV by optimizing the phase shift of the intelligent reflective surface and the movement trajectory of the UAV in the communication system.

[0004] The technical solution provided by the present invention is:

[0005] A method for secure communication of unmanned aerial vehicles (UAVs) assisted by an intelligent reflective surface, comprising:

[0006] Step 1: Obtain the location of the intelligent reflective surface and each legitimate user, and obtain the location of each eavesdropper;

[0007] Step 2: Establish an optimization objective function, and optimize the phase shift of the smart reflective surface and the movement trajectory of the UAV according to the optimization objective function to obtain the optimal phase shift of the smart reflective surface and the optimal position of the UAV at each moment;

[0008] Wherein, the optimization objective function is:

[0009]

[0010] Where, Indicates the movement trajectory of the drone. represents the phase shift change of the smart reflective surface, represents a set of time series, q U[t] represents the position of the UAV at time t, Φ[t] represents the phase shift of the smart reflective surface at time t, R sec [t] represents the sum of the confidentiality rates of all legitimate users at time t; e[t] represents the flight energy consumption of the UAV at time t;

[0011] Step 3: At each moment, the intelligent reflective surface is adjusted to the optimal phase shift, and the drone moves to the optimal position and sends communication data.

[0012] Preferably, in step 2, the sum of the confidentiality rates of all legitimate users at time t is calculated using the following formula:

[0013]

[0014]

[0015] Where, represents the transmission rate of the nth legal user at time t, represents the transmission rate of the kth eavesdropper at time t, H UR [t] represents the channel between the UAV and the smart reflective surface at time t, represents the channel between the smart reflective surface and the nth legal user at time t, Represents the interference power at the legitimate user.

[0016] Preferably, in step 2, the flight energy consumption of the UAV at time t is calculated using the following formula:

[0017]

[0018] Where, t d Indicates the time slot length, P s and P m represent the blade profile and induced power in the hovering state, U r represents the tip speed of the rotor blade, v h represents the average rotor blade induced speed in the hovering state, d0 represents the fuselage drag ratio, ρ a represents the air density, z represents the hardness of the rotor blade, G represents the area of ​​the rotor disc; v h [t] represents the flight speed of the drone, a x [t] and a y [t] represents the distance the UAV moves in the horizontal and vertical directions at time t, respectively.

[0019] Preferably, in step 2, the movement trajectory of the UAV is optimized by a reinforcement learning algorithm, and the phase shift of the smart reflective surface is optimized by an alternating optimization algorithm.

[0020] Preferably, optimizing the moving trajectory of the UAV and the phase shift of the smart reflective surface comprises the following steps:

[0021] Step 1: Convert the optimization problem into a Markov decision process: state s t =[q U [t],e[t]], action a t =[a x [t],a y [t]], reward value

[0022] in, Indicates the penalty for the drone flying out of the monitoring area, ω sec ,ω e ∈{0,1} represents the weight of the target;

[0023] Step 2: Construct a reinforcement learning network based on the Markov decision process, which includes a critic network and a policy network.

[0024] Step 3: Use multivariate Gaussian to randomly generate N policy network parameter vectors as candidate solution sets based on the mean μ and covariance matrix ∑ of the policy network parameters;

[0025] Step 4: Randomly select multiple candidate solutions from the candidate solution set, apply the network parameters corresponding to the selected candidate solutions to the policy network, and output the corresponding UAV position; and use an alternating optimization algorithm to obtain the phase shift of the smart reflective surface corresponding to the UAV position output by each policy network;

[0026] And store the experience data set generated in the above process into the experience buffer;

[0027] The experience data set includes the current state, action, reward value and the next state;

[0028] Step 5: randomly extract multiple pieces of experience data from the experience buffer, and use a reinforcement learning algorithm to update the remaining candidate solutions in the candidate solution set based on the experience data;

[0029] Step 6: Apply the network parameters corresponding to the updated candidate solution to the policy network and output the corresponding UAV position; and use the alternating optimization algorithm to obtain the phase shift of the smart reflector corresponding to the UAV position output by each policy network;

[0030] The experience data set generated in the above process is stored in the experience buffer;

[0031] Step 7: Sort the candidate solutions in the updated candidate solution set according to the fitness function, select a specified number of candidate solutions with high fitness as elite candidate solutions, and update the mean μ and covariance matrix ∑ according to the elite candidate solutions;

[0032] Repeat steps 3 to 7 until the set number of iterations is reached and the optimal reinforcement learning network is obtained;

[0033] Step 8: After determining the position of the UAV at each moment through the optimal reinforcement learning network, the phase shift of the intelligent reflective surface at that moment is obtained through the alternating optimization algorithm.

[0034] Preferably, an alternating optimization algorithm is used to optimize the phase shift of the smart reflective surface, comprising the following steps:

[0035] Step A: Based on the discrete phase shift set Randomly initialize the phase shift of each reflective element on the smart reflective surface;

[0036] Where b is the number of quantization bits;

[0037] Step B: randomly selecting a reaction element on the intelligent reaction surface and fixing the phase shifts of other reflection elements; traversing each phase shift value in the discrete phase shift set for the selected reflection element, and selecting the phase shift value that maximizes the confidentiality rate of the system as the phase shift of the reflection element;

[0038] Use the above method to optimize the phase shifts of other reaction elements one by one until the phase shift optimization of all reaction elements is completed;

[0039] Repeat step B multiple times until the specified number of iterations is reached.

[0040] The beneficial effects of the present invention are:

[0041] The present invention provides a method for secure drone communication assisted by an intelligent reflective surface. A multi-objective joint optimization model is established to improve the security rate of legitimate users and reduce the propulsion energy consumption of the drone. The phase shift of the intelligent reflective surface at each moment is designed using an alternating optimization algorithm, and the position of the drone at each moment is designed using a reinforcement learning algorithm to solve the model. By controlling the amplitude and phase shift of each reflective unit on the intelligent reflective surface, a new virtual line-of-sight link is created between the drone with obstructed communication and the ground user, thereby improving the communication quality. Moreover, when there is an eavesdropper in the scene and its position is known, the intelligent reflective surface can reflect the drone's transmitted signal to an area with dense legitimate users, thereby enhancing the signal strength of the legitimate users and weakening the signal strength of the eavesdropper, causing the eavesdropper to be unable to intercept communication data due to insufficient signal strength. The propulsion energy consumption of the drone can be reduced by designing the movement trajectory of the drone within a time length T. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 Schematic diagram of smart reflective surfaces assisting drones in secure communications.

[0043] Figure 2 This is a flow chart of the method for secure communication of unmanned aerial vehicles assisted by intelligent reflective surfaces according to the present invention. DETAILED DESCRIPTION

[0044] The present invention will be described in further detail below in conjunction with the accompanying drawings so that those skilled in the art can implement the invention with reference to the description.

[0045] like Figure 1 As shown, during wireless network communication, there are many eavesdroppers who attempt to intercept communication data, resulting in the leakage of sensitive information, thus posing a threat to the security of wireless network communication.

[0046] The present invention provides a secure communication method for drones, assisted by intelligent reflective surfaces. The intelligent reflective surface creates a new virtual line-of-sight link between the drone and the user, mitigating severe path loss caused by obstructions and improving communication quality between the drone and legitimate users. Furthermore, based on the designed phase shift of each reflective element on the intelligent reflective surface, the drone's transmitted signal is redirected, thereby enhancing the signal strength received by legitimate users while simultaneously weakening the signal strength received by eavesdroppers. This insufficient signal strength allows eavesdroppers to intercept data, thereby improving the performance of secure wireless network communications. Furthermore, by designing the drone's trajectory over a time period T, the drone's propulsion energy consumption can be reduced. Therefore, this secure communication method, based on the directional enhancement and weakening of signal strength by intelligent reflective surfaces, can achieve the goal of ensuring the security and reliability of wireless network communications.

[0047] like Figure 2 As shown, the specific implementation process of the UAV secure communication method assisted by the intelligent reflective surface is as follows.

[0048] 1. Determine the location of the intelligent reflective surface and each legitimate user;

[0049] 2. Determine the service scope of drone communications;

[0050] 3. The drone uses its own equipment (such as radar) to detect the location of each eavesdropper;

[0051] 4. Design two objective functions that need to be optimized based on the secure communication requirements, and construct a multi-objective optimization problem based on multi-objective optimization theory;

[0052] First, the overall objective function is designed as follows:

[0053]

[0054]

[0055]

[0056] in,

[0057]

[0058]

[0059] Φ[t]=diag(φ[t]) (7)

[0060]

[0061]

[0062]

[0063]

[0064]

[0065]

[0066]

[0067]

[0068]

[0069]

[0070] Where, Indicates the movement trajectory of the drone. represents the phase shift change of the smart reflective surface, represents a set of time series, q U [t] represents the position of the UAV at time t, and Φ[t] represents the phase shift of the smart reflective surface at time t.

[0071] The first objective function f1 represents the sum of the confidentiality rates of all legitimate users within the time period T. sec [t] represents the sum of the confidentiality rates of all legitimate users at time t; represents the confidentiality rate of the nth legitimate user at time t, represents the transmission rate of the nth legal user at time t, represents the transmission rate of the kth eavesdropper at time t, P represents the transmission power of the drone, represents the interference power at the legitimate user, represents the interference power at the eavesdropper; M r and M c Respectively represent the number of reflective elements in each row and column on the smart reflective surface, Indicates the mth intelligent reflective surface at time t r Row m c The phase shift of the reflective element at the column, in order to meet the actual application requirements, the phase shift of the reflective element takes a discrete value, that is, b is the number of quantization bits; H UR [t] represents the channel between the UAV and the intelligent reflective surface at time t, α0 represents the standard channel power when the propagation distance d0 = 1m, d UR [t] represents the distance between the UAV and the smart reflective surface at time t; represents the line-of-sight link component between the UAV and the smart reflector at time t, λ represents the carrier wavelength, and d r and d c Respectively represent the distance between each row of elements and the distance between each column of reflective elements on the smart reflective surface, θ UR [t] and ξ UR [t] represents the vertical and horizontal angles of arrival, respectively; and They represent the channels between the smart reflective surface and the nth legitimate user and the kth eavesdropper at time t, and They represent the distances from the smart reflective surface to the nth legitimate user and the kth eavesdropper respectively, and They represent the non-line-of-sight link components between the smart reflector surface to the nth legitimate user and the kth eavesdropper at time t respectively; and denote the line-of-sight link components between the intelligent reflector and the nth legitimate user and the kth eavesdropper at time t, respectively. and They represent the vertical and horizontal arrival angles from the smart reflector to the nth legal user link, and denote the vertical and horizontal arrival angles of the link from the smart reflector to the kth eavesdropper, respectively;

[0072] The second objective function f2 represents the energy consumption of the drone during its movement. Among them, e[t] represents the flight energy consumption of the drone at time t, t d Indicates the time slot length, P s and P m are two constants, representing the blade profile and induced power in the hovering state, U r represents the tip speed of the rotor blade, v hrepresents the average rotor blade induced speed in the hovering state, d0 represents the fuselage drag ratio, ρ a represents the air density, z represents the hardness of the rotor blade, G represents the area of ​​the rotor disc; v h [t] represents the flight speed of the drone, a x [t] and a y [t] represents the distance the UAV moves in the horizontal and vertical directions at time t, respectively.

[0073] 5. The above objective function is affected by the phase shift of the intelligent reflective surface and the flight trajectory of the UAV. These two variables will change over time, so the scenario considered in the present invention is dynamic. Traditional optimization algorithms usually require prior knowledge about the environment as a basis, but in a highly dynamic environment, it is difficult to obtain prior knowledge. Unlike traditional optimization algorithms, reinforcement learning algorithms enable the intelligent agent to quickly adjust its own actions through trial and error learning, thereby achieving the effect of effectively adapting to environmental changes. In addition, since the optimization variables involve both continuous variables and high-dimensional discrete variables, if the action space of the intelligent agent in reinforcement learning contains all the optimization variables, the exploration performance of the intelligent agent will be reduced due to the huge action space, thereby affecting the performance of the reinforcement learning algorithm. Therefore, the present invention uses an alternating optimization algorithm to optimize the phase shift of the intelligent reflective surface, while the reinforcement learning algorithm is only used to optimize the trajectory of the UAV, thereby achieving dimensionality reduction of the action space of reinforcement learning. In addition, although the reinforcement learning algorithm has demonstrated excellent performance in many fields, it still faces problems such as limited exploration performance and poor robustness. Therefore, the present invention introduces an evolutionary algorithm into reinforcement learning to update the neural network parameters, fully combining the advantages of the natural exploration strategy of the evolutionary algorithm and the high sample efficiency of the reinforcement learning algorithm, thereby improving the solution accuracy and convergence speed of the reinforcement learning algorithm.

[0074] Based on the established objective function, the UAV's trajectory is first designed using a reinforcement learning algorithm. Then, the phase shift of the smart reflective surface is selected using an alternating optimization algorithm. The specific calculation process is as follows:

[0075] (1) First, the optimization problem is converted into a Markov decision process, that is, the state s t =[q U [t],e[t]], action a t =[a x [t],a y [t]], reward value in Indicates the penalty for the drone flying out of the monitoring area, ω sec ,ω e ∈{0,1} represents the weight of the target;

[0076] (2) The reinforcement learning algorithm includes a critic network and a policy network, where the critic network is used to estimate the Q value and the policy network outputs the action of the agent. In addition, the introduction of the experience replay mechanism in reinforcement learning can improve the training efficiency and stability of the agent. Therefore, it is necessary to initialize the agent neural network parameters and the experience replay buffer. The initialized agent neural network coefficients include the critic network and Target critic network and Temperature coefficient ξ and policy network π φ , and use the parameter φ as the initial value of the mean μ of the evolutionary algorithm;

[0077] (3) The evolutionary algorithm encodes the parameter vectors of N policy neural networks as candidate solutions (φ i ) i=1,…,N ;

[0078] (4) Select N / 2 candidate solutions, apply the network parameters corresponding to the candidate solutions to the policy neural network of the agent to output the position of the UAV, and then use the alternating optimization algorithm to obtain the phase shift of the intelligent reflective surface. The specific process is as follows:

[0079] ①First, according to the discrete phase shift set Randomly initialize the phase shift of each reflective element on the smart reflective surface;

[0080] ② Optimize the phase shift of each reflective element one by one. Specifically, when optimizing the phase shift of any reflective element, first fix the phase shifts of the other reflective elements. Then, for that reflective element, traverse each phase shift value in the discrete phase shift set and select the value that maximizes the system's confidentiality rate as the phase shift of that reflective element.

[0081] ③If the number of iterations is reached, stop; otherwise, repeat step ②.

[0082] After the above steps, the trajectory of the drone and the phase shift of the intelligent reflective surface are known, so the reward value r fed back to the agent by the environment can be obtained. t , then the fitness value is set to the cumulative reward value obtained by the agent in the interaction with the environment within the time T, and the experience data (s t ,a t ,r t ,s t+1 ) is stored in the experience buffer.

[0083] (5) Extract b pieces of historical data (usually between 1 and 512) from the experience buffer, denoted as B, and use the reinforcement learning algorithm to update the other half of the candidate solutions. The update formula is as follows:

[0084]

[0085]

[0086]

[0087] θ i′ =εθ i +(1-ε)θ i′ ,

[0088] Among them, J(θ i ) represents the Loss value of the critic network, represents the target value, r t (s t ,a t ) indicates that the agent is in state s t Next, perform action a t The reward obtained, γ is the discount factor, Represents the target critic network for state s t+1 and action a t+1 The output action value, Indicates that the current policy network is in state s t+1 The output action is: Represents the critic network for state s t and action a t The output action value; J(φ) represents the Loss value of the policy network, Indicates that in the current state s t Under this condition, the action output by the policy network; J(ξ) represents the Loss value of the temperature coefficient, represents the target entropy value, represents the action space set; ε represents the soft update parameter of the target critic network.

[0089] The other half (the remaining N / 2) candidate solutions after reinforcement learning are used as parameters to output the position of the drone through the policy neural network, and then the phase shift of the intelligent reflective surface is obtained by the alternating optimization algorithm. The cumulative reward value obtained by the agent from the environment within the time length T is used as the fitness function to evaluate the quality of the candidate solution. At the same time, the experience data (s t ,a t ,r t,s t+1 ) is stored in the experience buffer.

[0090] Through the above process, the candidate solution set is also updated, and the updated candidate solution set includes the initial N / 2 candidate solutions and the updated N / 2 candidate solutions.

[0091] (6) Sort each candidate solution in the updated candidate solution set according to the fitness function value, and select N e The elite candidate solution with the highest fitness value Update the mean μ and covariance matrix ∑. The update formula is as follows:

[0092]

[0093]

[0094] in, represents the weight coefficient, z i represents the elite candidate solution set, N e represents the number of elite candidate solutions, ∈ represents the additional variance, μ old Represents the mean of the population at the beginning of this round, μ new and ∑ new represents the mean and covariance of the population after the update.

[0095] (7) Then the mean μ new The corresponding network parameters are applied to the policy neural network of the agent, and combined with the alternating optimization algorithm, the output is the movement trajectory of the drone and the phase shift change of the intelligent reflective surface within the duration T during this round of iteration;

[0096] (8) If the iteration limit is reached, stop; otherwise, return to step (3) and repeat (3) to (7) until the termination condition is met.

[0097] 6. At each moment, based on the optimized phase shift change and the drone’s position, the intelligent reflective surface adjusts the phase shift, and the drone moves along the trajectory and sends communication data.

[0098] The present invention provides a method for secure drone communication assisted by a smart reflective surface. This method establishes a multi-objective joint optimization model to improve the confidentiality rate of legitimate users and reduce drone propulsion energy consumption. Within a discrete time series, the method first uses an alternating optimization algorithm to design the phase shift of the smart reflective surface. A reinforcement learning algorithm is then used to design the drone's trajectory. An evolutionary algorithm is then introduced into the reinforcement learning process to update neural network parameters, improving solution accuracy and convergence speed. The smart reflective surface creates a new virtual line-of-sight link between the drone and the user, alleviating the severe path loss caused by obstructions and improving communication quality between the drone and the legitimate user. Furthermore, the designed phase shift of each reflective element on the smart reflective surface redirects the drone's transmitted signal, thereby enhancing the signal strength received by the legitimate user and weakening the signal strength received by the eavesdropper. This insufficient signal strength allows the eavesdropper to intercept data, thereby improving the performance of secure wireless network communication. Furthermore, by designing the drone's trajectory over a time period T, the drone's propulsion energy consumption can be reduced. Therefore, this secure communication method based on the directional enhancement and weakening of signal strength by the smart reflective surface ensures the security and reliability of wireless network communications.

[0099] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the description and implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and illustrations shown and described herein.

Claims

1. A method for secure communication of unmanned aerial vehicles assisted by an intelligent reflective surface, characterized in that: The steps include: Step 1: Obtain the location of the intelligent reflective surface and each legitimate user, and obtain the location of each eavesdropper; Step 2: Establish an optimization objective function, and optimize the phase shift of the smart reflective surface and the movement trajectory of the UAV according to the optimization objective function to obtain the optimal phase shift of the smart reflective surface and the optimal position of the UAV at each moment; Wherein, the optimization objective function is: Where, Indicates the movement trajectory of the drone. represents the phase shift change of the smart reflective surface, represents a set of time series, q U [t] represents the position of the UAV at time t, Φ[t] represents the phase shift of the smart reflective surface at time t, R sec [t] represents the sum of the confidentiality rates of all legitimate users at time t; e[t] represents the flight energy consumption of the UAV at time t; Step 3: At each moment, the intelligent reflective surface is adjusted to the optimal phase shift, and the drone moves to the optimal position and sends communication data; Optimizing the movement trajectory of the UAV and the phase shift of the smart reflective surface includes the following steps: Step 1: Convert the optimization problem into a Markov decision process: state s t =[q U [t],e[t]], action a t =[a x [t],a y [t]], reward value in, Indicates the penalty for the drone flying out of the monitoring area, ω sec ,ω e ∈{0,1} represents the weight of the target; Step 2: Construct a reinforcement learning network based on the Markov decision process, which includes a critic network and a policy network. Step 3: Use multivariate Gaussian to randomly generate N policy network parameter vectors as candidate solution sets based on the mean μ and covariance matrix ∑ of the policy network parameters; Step 4: Randomly select multiple candidate solutions from the candidate solution set, apply the network parameters corresponding to the selected candidate solutions to the policy network, and output the corresponding UAV position; and use an alternating optimization algorithm to obtain the phase shift of the smart reflective surface corresponding to the UAV position output by each policy network; And store the experience data set generated in step 4 into the experience buffer; The experience data set includes the current state, action, reward value and the next state; Step 5: randomly extract multiple pieces of experience data from the experience buffer, and use a reinforcement learning algorithm to update the remaining candidate solutions in the candidate solution set based on the experience data; Step 6: Apply the network parameters corresponding to the updated candidate solution to the policy network and output the corresponding UAV position; and use the alternating optimization algorithm to obtain the phase shift of the smart reflector corresponding to the UAV position output by each policy network; The experience data set generated in step 6 is stored in the experience buffer; Step 7: sort the candidate solutions in the updated candidate solution set according to the fitness function, select a specified number of candidate solutions with high fitness as elite candidate solutions, and update the mean μ and covariance matrix ∑ according to the elite candidate solutions; Repeat steps 3 to 7 until the set number of iterations is reached and the optimal reinforcement learning network is obtained; Step 8: After determining the position of the UAV at each moment through the optimal reinforcement learning network, the phase shift of the intelligent reflective surface at that moment is obtained through an alternating optimization algorithm; The phase shift of the smart reflector is optimized using an alternating optimization algorithm, which includes the following steps: Step A: Based on the discrete phase shift set Randomly initialize the phase shift of each reflective element on the smart reflective surface; Where b is the number of quantization bits; Step B: randomly selecting a reaction element on the intelligent reaction surface and fixing the phase shifts of other reflection elements; traversing each phase shift value in the discrete phase shift set for the selected reflection element, and selecting the phase shift value that maximizes the confidentiality rate of the system as the phase shift of the reflection element; Optimize the phase shifts of other reaction elements one by one using the method in step B until all reaction elements have completed the phase shift optimization; Repeat step B multiple times until the specified number of iterations is reached.

2. The method for secure communication of unmanned aerial vehicles assisted by intelligent reflective surfaces according to claim 1, characterized in that: In step 2, the sum of the confidentiality rates of all legitimate users at time t is calculated using the following formula: Where, represents the transmission rate of the nth legal user at time t, represents the transmission rate of the kth eavesdropper at time t, H UR [t] represents the channel between the UAV and the smart reflective surface at time t, represents the channel between the smart reflective surface and the nth legal user at time t, Represents the interference power at the legitimate user.

3. The method for secure communication of unmanned aerial vehicles assisted by intelligent reflective surfaces according to claim 2, characterized in that: In step 2, the flight energy consumption of the UAV at time t is calculated using the following formula: Where, t d Indicates the time slot length, P s and P m represent the blade profile and induced power in the hovering state, U r represents the tip speed of the rotor blade, v h represents the average rotor blade induced speed in the hovering state, d0 represents the fuselage drag ratio, ρ a represents the air density, z represents the hardness of the rotor blade, G represents the area of ​​the rotor disc; v h [t] represents the flight speed of the drone, a x [t] and a y [t] represents the distance the UAV moves in the horizontal and vertical directions at time t, respectively.

Citation Information

Patent Citations

  • Multi-user cooperative transmission method driven by intelligent reflecting surface

    CN115037337A

  • Unmanned aerial vehicle-mounted RIS auxiliary vehicle network communication method and system

    CN115915069A