Reconfigurable intelligent surface-assisted secret communication optimization method based on deep reinforcement learning

By adopting deep reinforcement learning methods in RIS-assisted wireless communication systems, using two agents to optimize the parameters of the drone and RIS, the problem of insufficient confidentiality rate in traditional methods in high computational complexity and dynamic environments is solved, and more efficient confidential communication is achieved.

CN120433857AActive Publication Date: 2025-08-05BEIJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510632129.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-05
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

In the RIS-assisted wireless communication system, traditional optimization methods rely on accurate mathematical models and CSI, have high computational complexity, are difficult to deal with joint optimization problems in dynamic environments, and fail to effectively consider the existence of eavesdroppers and environmental changes, resulting in insufficient confidentiality rate.

Method used

Using a method based on deep reinforcement learning, two agents are used to optimize the active beamforming matrix of the drone, the phase shift matrix of RIS and the drone trajectory, and the deep reinforcement learning model is optimized through the TA-PPO algorithm to maximize the confidentiality rate.

Benefits of technology

In a dynamic environment, it improves the confidentiality rate of the drone communication system, enhances the flexibility and security of the system, reduces the computational complexity and power consumption, and is suitable for RIS-assisted drone communication systems and other confidential communication systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120433857A_ABST
    Figure CN120433857A_ABST
Patent Text Reader

Abstract

The invention discloses a reconfigurable intelligent surface-assisted secret communication optimization method based on deep reinforcement learning, and belongs to the technical field of wireless communication, and the method comprises the following steps: 1, an unmanned plane sends information to a user, and collects environment information; step 2, according to the collected environment information, respectively establishing a deep reinforcement learning model for optimizing the trajectory and beam forming of the unmanned aerial vehicle; step 3, optimizing the deep reinforcement learning model by using a TA-PPO algorithm; and step 4, obtaining an optimal joint beam forming scheme according to the optimized deep reinforcement learning model. According to the method, in the presence of an eavesdropper, an RIS-assisted NOMA unmanned aerial vehicle communication system is considered, and the active beam forming matrix of the unmanned aerial vehicle, the phase shift matrix of the RIS and the trajectory of the unmanned aerial vehicle are jointly optimized in a dynamic environment, so that the secrecy rate is maximized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of wireless communications, and in particular relates to a reconfigurable intelligent surface-assisted secure communication optimization method based on deep reinforcement learning. Background Art

[0002] Future 6G communication technology aims to provide communication services with higher speeds, lower latency, a larger number of connections, and higher reliability. In order to support large-scale connections and fully utilize spectrum resources, non-orthogonal multiple access (NOMA) technology remains a key technology in 6G communication systems. However, while the increase in the number of communication devices and the growth in communication needs have brought great convenience to daily life, they have also brought major security challenges to wireless communications. The inherent openness and wide propagation range of wireless signals make data easily intercepted by eavesdroppers, posing a serious threat to user privacy and data security. Therefore, it is crucial to protect wireless communications from unauthorized access and ensure communication security.

[0003] Reconfigurable Intelligent Surfaces (RIS) are considered a key technology with enormous potential in future wireless communications. They can improve signal propagation quality while enhancing the spectral efficiency, energy efficiency, and confidentiality of wireless communication systems. Specifically, RIS consists of a large number of low-cost passive reflective elements, forming a uniform planar array. Each element can adaptively adjust its reflection amplitude or phase, thereby precisely controlling the intensity and direction of the reflected electromagnetic wave and reconfiguring the wireless propagation environment. Beamforming is a key technology in RIS research. Its basic principle is to control the propagation direction of the incident signal by precisely adjusting the phase response of each element on the RIS surface. This technology can effectively concentrate or scatter wireless signals in a predetermined direction, enhancing signal strength at the target location, improving the quality of the communication link, and enhancing physical layer security.

[0004] Unmanned aerial vehicle (UAV) communication systems have become a crucial solution in modern networks due to their flexibility and efficiency. However, UAV communication systems are particularly susceptible to spatial obstructions. Buildings and other obstacles can cause signal attenuation or even complete interruption. Therefore, improving the adaptability of UAV communication systems in complex environments is a key issue. RIS, which can be integrated with UAV-assisted networks, not only enhances wireless communications, especially in situations where there is no line of sight between base stations and users, but also effectively reduces the strength of eavesdropping signals and improves link quality, thereby providing a more robust and reliable wireless network.

[0005] In RIS-assisted wireless communication systems, the dynamic nature of the communication environment poses significant challenges for network optimization design. Traditional optimization methods, such as alternating optimization algorithms and semidefinite relaxation algorithms, typically rely on precise mathematical models and have high computational complexity. In recent years, the rapid development of artificial intelligence (AI) has profoundly changed our way of working and living. Deep reinforcement learning (DRL), a key AI technology, combines the strengths of deep learning and reinforcement learning. It enables intelligent agents to optimize their decision-making strategies through trial and error through interaction with the environment, effectively solving dynamic decision-making problems. Commonly used DRL-based frameworks include Twin Delayed Deep Deterministic Policy Gradient (TD3) and Proximal Policy Optimization (PPO). DRL technology has extensive applications in communications and networking. Using DRL to optimize the design of RIS-assisted wireless communication systems can more efficiently adjust and optimize various variables, thereby improving system performance, which is of great practical significance.

[0006] However, the use of traditional iterative optimization algorithms to optimize RIS-assisted secure communication systems has the following disadvantages:

[0007] (1) Most traditional algorithms rely on accurate channel state information (CSI) and are optimized based on fixed models. However, due to the passive nature of RIS and the dynamic nature of wireless channels, it is very difficult to accurately estimate the cascaded channel to obtain perfect CSI.

[0008] (2) Traditional algorithms usually rely on precise mathematical models and have high computational complexity. In the case of limited CSI, these algorithms often have difficulty in effectively handling joint optimization problems and may cause large delays.

[0009] (3) This type of traditional algorithm lacks adaptive learning capabilities, making it difficult to reveal implicit information between users and to dynamically adjust optimization plans based on environmental changes.

[0010] Furthermore, current approaches using deep reinforcement learning algorithms for RIS beamforming optimization have the following drawbacks: 1) They fail to account for the presence of eavesdroppers and mostly focus solely on active and passive beamforming schemes, without optimizing other environmental variables. 2) They typically employ centralized deep reinforcement learning algorithms, employing only a single agent to perform the optimization task.

[0011] Therefore, it is necessary to provide a reconfigurable intelligent surface-assisted secure communication optimization method based on deep reinforcement learning to address the shortcomings of existing technologies. Summary of the Invention

[0012] In view of the shortcomings of the existing technology, the purpose of the present invention is to consider the RIS-assisted NOMA UAV communication system in the presence of eavesdroppers, and jointly optimize the UAV's active beamforming matrix, the RIS phase shift matrix and the UAV's trajectory in a dynamic environment, so as to maximize the confidentiality rate.

[0013] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is as follows:

[0014] A reconfigurable intelligent surface-assisted secure communication optimization method based on deep reinforcement learning utilizes two intelligent agents to optimize secure communication to maximize the secure rate. The method includes the following steps:

[0015] Step 1: The drone sends information to the user and collects environmental information;

[0016] Step 2: Based on the collected environmental information, a deep reinforcement learning model is established to optimize the drone trajectory and beamforming.

[0017] Step 3: Use the TA-PPO algorithm to optimize the deep reinforcement learning model;

[0018] Step 4: Based on the optimized deep reinforcement learning model, the optimal joint beamforming scheme is obtained.

[0019] As a further improvement to the above scheme, the realization of maximizing the confidentiality rate includes the following process:

[0020] There are K single-antenna legitimate users and one single-antenna eavesdropper; the UAV is equipped with an A-element uniform linear array, while the RIS uses a uniform planar array with M passive reflection units; the entire flight time T is evenly divided into N time slots, δ t Represents each individual time slot, then T=nδ t , n∈N; RIS is located at coordinate w R =(x R ,y R , z R ) T The coordinates of the user and the eavesdropper in time slot n are represented as w i =(x i [n],y i [n],z i [n]) T The coordinates of the drone at time slot n are q[n] = (x[n], y[n], H) T; The speed of the drone is expressed as:

[0021]

[0022] The definition of the UAV maneuverability constraints is as follows:

[0023] q[0]=(0,0,H),

[0024] |x[n]|,|y[n]|≤B,

[0025]

[0026] Among them, q[0] is the initial coordinate of the UAV, B is the moving boundary of the UAV, and D max represents the maximum maneuvering distance of the UAV at time slot n, and H represents the height of the UAV.

[0027] As a further improvement of the above scheme, the channel gains from the UAV to the RIS, from the UAV to the kth user, from the UAV to the eavesdropper, from the RIS to the kth user, and from the RIS to the eavesdropper are expressed as The channel is modeled as a Rayleigh fading channel, and the channel modeling is as follows:

[0028]

[0029] Where i∈{k,p}, ρ0 represents the path loss at a reference distance of 1m, β1, β2, and β3 represent the relevant path loss exponents, and K1, K2, and K3 are the Rayleigh factors; and represents the NLoS component, where and are all independent, identically distributed, and follow a circularly symmetric complex Gaussian distribution with zero mean and unit variance; d U,i d R,i and d U,R denote the distances from the UAV to the user / eavesdropper, from the RIS to the user / eavesdropper, and from the UAV to the RIS, respectively; and represents the LoS component given by the uniform linear antenna array response. The calculation process of the LoS component is as follows:

[0030]

[0031] is the ULA’s steering vector, denoted as in Expressed as AoD azimuth d is the antenna spacing, λ c is the carrier wavelength;

[0032] is the steering vector of UPA, expressed as:

[0033]

[0034] Among them, 0≤p, q≤n-1, and θ are the azimuth and altitude angles of the path between the RIS and the target object, respectively;

[0035] The cascade channel from the drone to the user or eavesdropper is expressed as The RIS passive beamforming matrix is defined as Here, ψ is a diagonal matrix whose main diagonal elements are given by Calculated, represents the phase shift of the nth reflective element;

[0036] The channel coefficients from the drone to all receivers are combined as follows:

[0037]

[0038] Therefore, the signal received by the i-th user from the drone is expressed as:

[0039] y i =H c,i Gs+n i ;

[0040] in, and Represent the beamforming matrix and transmission signal on the UAV respectively; n i represents background noise; g i represents the i-th column of G.

[0041] As a further improvement to the above solution, the decoding order of all users is set to Apply serial interference cancellation technology and set the decoding order from The first element to The last element of U K To U1; user The achievable data rate is expressed as:

[0042]

[0043] in, Represents user U i When decoding its own signal, user U i The decoding data rate; Indicates that when user U j Decoding user U i When the signal Decoding data rate for j < i; and Written as:

[0044]

[0045] Where x'=(U i-1 ,U i-2 ,...,U j , ..., U2, U1) is A subset of

[0046] The eavesdropper's eavesdropping rate on user k is:

[0047]

[0048] According to Weiner's eavesdropping code, the confidentiality rate is:

[0049] R sk =(R k -R ek ) + ,

[0050] Among them, (a) + =max(0,a);R K represents the information rate of user k.

[0051] As a further improvement to the above scheme, the confidentiality rate is maximized by adjusting the UAV trajectory Q, active beamforming G, and phase shift matrix ψ, which can be expressed as:

[0052]

[0053] As a further improvement of the above scheme, the two agents include a beamforming agent and a drone trajectory agent, which are used to independently optimize the active and passive beamforming of RIS and independently optimize the drone trajectory;

[0054] In step 1, the environmental information obtained by the beamforming agent at the nth moment is: the channel state information from the drone to the user and the eavesdropper at the n-1th moment;

[0055] The environmental information obtained by the UAV trajectory agent at the nth moment is: the current position of the UAV; where n∈N.

[0056] As a further improvement to the above scheme, in step 2, the establishment of the deep reinforcement learning model includes the following process:

[0057] Configure two agents and let them learn the optimal strategy π to maximize the cumulative discounted reward. The reward value of each round is as follows:

[0058]

[0059] Among them, r n is the reward value at step n, γ∈[0,1) is the discount factor used to weigh the importance of immediate rewards and future potential rewards; the agent’s actor network outputs the strategy, and the critic network is used to evaluate the strategy;

[0060] The state value function V needs to be evaluated during the evaluation process. π (s) and action-value function Q π (s, a) is evaluated; where the state value function V π (s) represents the expected value of the cumulative discounted reward obtained in the future when the agent is in state s under strategy π; the action value function Q π (s, a) represents the expected value of the cumulative discounted reward obtained in the future after the agent is in state s and performs action a under strategy π.

[0061] As a further improvement of the above scheme, in step 3, the TA-PPO algorithm uses two agents to decouple variables and models the communication scheme as a Markov decision process;

[0062] The beamforming agent is used to generate active and passive beamforming strategies. Its state is the CSI obtained through channel estimation, and its action is the UAV's active beamforming G and the phase shift matrix ψ of RIS. The UAV trajectory agent is used to generate the UAV's trajectory strategy. Its state is the UAV's position at the current moment, and its action is the UAV's position at the next moment.

[0063] The two agents share a reward, which is the sum of the confidentiality rates of all users at the current moment. Through multiple rounds of iteration, the strategy is continuously optimized to maximize the cumulative reward until the training process converges.

[0064] As a further improvement of the above scheme, in step 3, the optimization of the deep reinforcement learning model includes the following process: offline policy update is realized by using importance sampling; the importance sampling is to use the old policy π′ θ (a|s) generates samples to update the learned strategy π θ (a|s), where the gradient is:

[0065]

[0066] The TA-PPO algorithm directly updates the random policy neural network π θ , to approximate the probability distribution of the action given the state, that is, Pr θ (a|s)=π θ (a|s);Q π(s,a) is given by the advantage function Instead, the critic network uses to estimate how well an action performs relative to the average action for a given state;

[0067] Advantage function is defined as:

[0068]

[0069] in, is the baseline predicted by the critic network; The larger the value, the greater the probability that action a is selected by strategy π in state s;

[0070] Using the clipping method, the objective function is directly clipped to:

[0071]

[0072] The goal of the critic network is to make the actual value Minimize the gap between parameters Update is performed by gradient descent method, i.e. is the loss function of the critic network, which is the mean square error function and is defined as:

[0073]

[0074] As a further improvement to the above solution, in step 4, obtaining the optimal joint beamforming solution includes the following steps:

[0075] Step 4.1: Build a joint optimization framework based on dual-agent collaboration to simultaneously optimize the active and passive beamforming parameters of the RIS and the UAV trajectory through centralized training and distributed execution.

[0076] Step 4.2: During the centralized training phase, the beamforming agent and the drone trajectory agent share global state information, actions, and rewards. Based on the shared global state information, actions, and rewards, they estimate the advantage function and update the policy network to optimize the joint policy.

[0077] Step 4.3: In the distributed execution phase, the UAV trajectory agent independently generates a flight path based on the local observed UAV's current position; the beamforming agent independently generates a beamforming strategy based on the local observed channel state.

[0078] Compared with traditional communication methods, the present invention has the following advantages:

[0079] (1) Traditional physical layer security technologies typically rely on additional hardware, resulting in additional power consumption. Their effectiveness is often limited when the user and the eavesdropper are close together or in the same direction, i.e., when signals are strongly correlated. RIS, a revolutionary technology, transforms the controllability of traditional wireless channels by creating an intelligent, programmable environment. Unlike traditional technologies, RIS requires no additional power consumption and is easily deployed, demonstrating broad application prospects.

[0080] (2) RIS is deployed between users and drones to provide additional transmission paths, significantly enhancing the flexibility of the system.

[0081] (3) The present invention adopts the DRL algorithm to optimize the flight trajectory of the UAV, the beamforming of the UAV, and the RIS phase shift matrix, thereby improving the confidentiality rate of the entire system in the presence of channel uncertainty, eavesdropping risk, and outdated CSI.

[0082] (4) The present invention uses two agents to independently optimize the active and passive beamforming of the RIS and the UAV trajectory, using the CTDE execution paradigm. The beamforming agent is responsible for generating active and passive beamforming strategies. Its state is the CSI obtained through channel estimation, and its action is the UAV's active beamforming and the phase shift matrix of the RIS. The UAV trajectory agent is responsible for generating the UAV trajectory strategy. Its state is the UAV's position at the current moment, and its action is the UAV's position at the next moment. In this way, the two agents independently generate beamforming and trajectory strategies, thereby achieving more efficient joint optimization.

[0083] (5) The present invention is applicable to RIS-assisted UAV communication systems and other RIS-assisted secure communication systems. While the optimization target is the security rate, other physical layer security-related metrics, such as security energy efficiency, can also be optimized. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] Figure 1 A block diagram of the steps of the communication optimization method of the present invention;

[0085] Figure 2 Scenario diagram of the NOMA UAV communication system assisted by RIS;

[0086] Figure 3 Schematic diagram of the reinforcement learning interaction process;

[0087] Figure 4 This is the flow chart of the TA-PPO algorithm. DETAILED DESCRIPTION

[0088] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0089] Example:

[0090] like Figure 1 As shown, this embodiment provides a reconfigurable intelligent surface-assisted secure communication optimization method based on deep reinforcement learning, which uses two intelligent agents to optimize secure communication to maximize the secure rate; the method includes the following steps:

[0091] Step 1: The drone sends information to the user and collects environmental information;

[0092] Step 2: Based on the collected environmental information, a deep reinforcement learning model is established to optimize the drone trajectory and beamforming.

[0093] Step 3: Use the TA-PPO algorithm to optimize the deep reinforcement learning model;

[0094] Step 4: Based on the optimized deep reinforcement learning model, the optimal joint beamforming scheme is obtained.

[0095] This embodiment is applied to the NOMA drone communication system assisted by RIS, especially in scenarios where beamforming of RIS is required. This embodiment is applicable to complex communication transmission systems. In traditional methods, obtaining channel state information becomes more difficult, transmission overhead increases, and the system needs to spend more time iterating to obtain the optimal solution, resulting in non-negligible delay. The solution we proposed only transmits the scalar channel gain, and the action delay of the neural network output is negligible. Although this embodiment is explained using the NOMA drone communication system as an example, this communication optimization method is also applicable to other secure communication systems assisted by RIS.

[0096] The realization of maximizing the confidentiality rate includes the following specific processes:

[0097] like Figure 2 As shown in Figure 1, this embodiment considers a RIS-assisted NOMA drone communication system. In this system, there are K single-antenna legitimate users and one single-antenna eavesdropper. The drone is equipped with an A-element uniform linear array, while the RIS uses a uniform planar array with M passive reflective elements.

[0098] The entire flight time T is evenly divided into N time slots, δ t Represents each individual time slot, then T=nδt , n∈N. RIS is located at coordinate w R =(x R ,y R , z R ) T On the other hand, the coordinates of the user and the eavesdropper at time slot n are represented as w i =(x i [n],y i [n],z i [n]) T Finally, let q[n] = (x[n], y[n], H) T is the coordinate of the UAV at time slot n. The speed of the UAV can be expressed as:

[0099]

[0100] Let q[0] be the initial coordinate of the UAV, B be the moving boundary of the UAV, and D max represents the maximum maneuvering distance of the UAV at time slot n. The definition of the UAV maneuverability constraint is as follows:

[0101]

[0102] The channel gains from the drone to the RIS, from the drone to the kth user, from the drone to the eavesdropper, from the RIS to the kth user, and from the RIS to the eavesdropper are expressed as The channel is modeled as a Rayleigh fading channel, taking into account both large-scale fading and small-scale fading. The channel model is as follows:

[0103]

[0104]

[0105] where i∈{k,p}, ρ0 represents the path loss at a reference distance of 1 m, β1, β2, and β3 represent the relevant path loss exponents, and K1, K2, and K3 are the Rayleigh factors. and represents the NLoS component, where each element is independent and identically distributed and follows a circularly symmetric complex Gaussian distribution with zero mean and unit variance. U,i d R,i and d U,R denote the distances from the UAV to the user / eavesdropper, from the RIS to the user / eavesdropper, and from the UAV to the RIS, respectively. and Represent the LoS component given by the uniform linear antenna array response as follows:

[0106]

[0107] is the ULA's guidance vector, written as in Expressed as AoD azimuth d is the antenna spacing, λ c is the carrier wavelength. is the steering vector of UPA, expressed as:

[0108]

[0109] Among them, 0≤p, q≤n-1, and are the azimuth and altitude of the path between RIS and the target object, respectively.

[0110] The cascade channel from the drone to the user or eavesdropper is expressed as Likewise, the RIS passive beamforming matrix is defined as ψ is a diagonal matrix whose main diagonal elements are given by Calculated, represents the phase shift of the nth reflective element. In order to maximize the power of the reflected signal and simplify the problem, this embodiment sets β=1.

[0111] The channel coefficients from the drone to all receivers can be combined as follows:

[0112]

[0113] Therefore, the signal received by the i-th user from the drone is expressed as:

[0114] y i =H ci Gs+n i ,

[0115] in, and They represent the beamforming matrix and transmission signal on the UAV respectively. i Represents background noise. g i represents the i-th column of G.

[0116] It should be noted that for NOMA transmission scheme, the successive interference cancellation (SIC) technology must be applied to users. Assume that the decoding order of all users is Applying SIC technology, assuming the decoding order is from The first element to The last element of U Kto U1. Therefore, the user The achievable data rate can be expressed as:

[0117]

[0118] in, Represents user U i When decoding its own signal, user U i The decoding data rate. Indicates that when user U j Decoding user U i When the signal The decoding data rate for j < i. The min() of the data rate ensures that SIC can be applied smoothly. and can be written as:

[0119]

[0120] Where x′=(U i-1 ,U i-2 ,...,U j ,...,U2,U1) is a subset of x.

[0121] The eavesdropper's eavesdropping rate on user k is:

[0122]

[0123] According to Weiner's eavesdropping code, the confidentiality rate is:

[0124] R sk =(R k -R ek ) + ,

[0125] Where (a) + =max(0,a);R K represents the information rate of user k.

[0126] In practical systems, obtaining perfect instantaneous channel state information is challenging due to transmission and processing delays, as well as the mobility of drones. By the time drones send signals to RIS and users, the channel state information may be outdated. Transmitting with outdated CSI inevitably leads to performance degradation. Therefore, explicitly accounting for outdated CSI in system design is crucial.

[0127] This embodiment is dedicated to maximizing the confidentiality rate limited by the base station power. The confidentiality rate is maximized by adjusting the UAV's trajectory Q, active beamforming G, and phase shift matrix ψ. It can be expressed as:

[0128]

[0129] It should be noted that the two agents include a beamforming agent and a UAV trajectory agent, which are used to independently optimize the active and passive beamforming of RIS and independently optimize the UAV trajectory;

[0130] In step 1, the environmental information obtained by the beamforming agent at the nth moment is: the channel state information from the drone to the user and the eavesdropper at the n-1th moment;

[0131] The environmental information obtained by the UAV trajectory agent at the nth moment is: the current position of the UAV; where n∈N.

[0132] like Figure 3 As shown in Figure 2, reinforcement learning, an emerging paradigm for solving dynamic problems, seamlessly integrates with deep neural networks to provide an effective approach to addressing dynamic decision-making challenges. The goal of reinforcement learning is to enable an intelligent agent to learn an optimal policy through interaction with a given environment, thereby maximizing its accumulated rewards. Specifically, the agent needs to learn to choose the optimal action in each state to achieve the highest long-term reward. This process is often referred to as policy optimization.

[0133] In step 2, the establishment of the deep reinforcement learning model includes the following process: configuring two agents and letting them learn the optimal strategy π to maximize the cumulative discounted reward. The reward value of each round is as follows:

[0134]

[0135] where r n is the reward value at step n, and γ∈[0,1) is a discount factor that weighs the importance of immediate rewards against potential future rewards. The agent’s actor network outputs a policy, while the critic network is used to evaluate the policy.

[0136] The state value function V needs to be evaluated during the evaluation process. π (s) and action-value function Q π (s,a) is evaluated. ; Among them, the state value function V π (s) represents the expected value of the cumulative discounted reward obtained in the future when the agent is in state s under strategy π; the action value function Q π (s,a) represents the expected value of the cumulative discounted reward obtained in the future after the agent is in state s and performs action a under strategy π.

[0137] In step 3, the TA-PPO (i.e., Twin-agent PPO) algorithm uses two agents to decouple variables and models the communication scheme as a Markov decision process;

[0138] The beamforming agent is used to generate active and passive beamforming strategies. Its state is the CSI obtained through channel estimation, and its action is the UAV's active beamforming G and the phase shift matrix ψ of RIS. The UAV trajectory agent is used to generate the UAV's trajectory strategy. Its state is the UAV's position at the current moment, and its action is the UAV's position at the next moment.

[0139] The two agents share a reward, which is the sum of the confidentiality rates of all users at the current moment. Through multiple rounds of iteration, the strategy is continuously optimized to maximize the cumulative reward until the training process converges.

[0140] It should be noted that the PPO algorithm, as an important branch of reinforcement learning, is particularly suitable for processing continuous action spaces and does not require discretization. This embodiment uses importance sampling to implement offline policy updates, ensuring learning efficiency while improving the stability of policy updates.

[0141] Importance sampling refers to using the old strategy π′ θ (a|s) generates samples to update the learned strategy π θ (a|s), where the gradient is:

[0142]

[0143] The TA-PPO algorithm directly updates the random policy neural network π θ , to approximate the probability distribution of actions given a state, i.e., pr θ (a|s)=π θ (a|s).

[0144] In order to accelerate convergence, Q π (s, a) is usually represented by an advantage function Instead, the critic network uses To estimate how well an action performs relative to the average action in a given state. Advantage function is defined as:

[0145]

[0146] in, is the baseline predicted by the critic network. The larger the value, the greater the probability that action a is selected by strategy π in state s. On the other hand, since the large difference between the new and old strategies will lead to unstable training, the clipping method is used to directly clip the objective function to:

[0147]

[0148] The goal of the critic network is to make the actual value Minimize the gap between the parameters Update is performed by gradient descent method, i.e. is the loss function of the critic network, which is a mean squared error function defined as:

[0149]

[0150] Due to the strong coupling between the UAV trajectory Q and a large number of CSIs, it becomes very difficult to optimize all variables simultaneously, which may lead to poor convergence and poor performance. To address this issue, we adopt two agents to optimize the UAV trajectory and active and passive beamforming respectively.

[0151] like Figure 4 As shown, step 3 specifically includes the following processes:

[0152] Step 3.1, environment interaction: Set each time step t, the two agents are based on the current state s t Generate action a t,1 and a t,2 . Among them, the beamforming agent observes the current state s t,1 , i.e. channel state information, actor network Output action a t,1 , namely the beamforming matrix and phase shift matrix. The UAV trajectory agent observes the current state s t,2 , i.e. drone coordinates, actor network Output action a t,2 , i.e., flight direction and flight distance. The environment feedbacks the next state s t+1 and instant rewards t .

[0153] Step 3.2, Experience pool storage: store the quadruple (s t ,a t ,r t ,s t+1 ) is stored in an independent experience pool.

[0154] Step 3.3, offline strategy update:

[0155] Step 3.31, small batch sampling:

[0156] (1) Randomly extract small batches of samples from the experience pool.

[0157] (2) Importance Sampling: Using the Old Strategy The sample update new strategy πθ , through weight Correct the deviation.

[0158] Step 3.32, Advantage Function Calculation: Critic Network Outputs Baseline Value Calculating advantage estimates Advantage function Reflects the quality of an action relative to average performance.

[0159] Step 3.33, actor network update:

[0160] (1) Policy Gradient: Maximizing the Objective Function

[0161] (2) Truncation mechanism: Use the clip(·) function to limit the strategy update amplitude and constrain the probability ratio to be within [1-∈, 1+∈].

[0162] (3) Parameter update formula:

[0163] Step 3.34, Critic Network Update:

[0164] (1) Minimize the mean square error loss L(θ C ),make Approximate the actual cumulative reward.

[0165] (2) Parameter update formula:

[0166] In step 4, obtaining the optimal joint beamforming solution includes the following steps:

[0167] Step 4.1: Build a joint optimization framework based on dual-agent collaboration to simultaneously optimize the active and passive beamforming parameters of the RIS and the UAV trajectory through a centralized training with decentralized execution (CTDE) mechanism.

[0168] Step 4.2: During the centralized training phase, the beamforming agent and the drone trajectory agent share global state information, actions, and rewards. The global state information includes the state information obtained by each agent, the drone position, and the channel state information. Based on the shared global state information, actions, and rewards, the advantage function is estimated and the policy network is updated to optimize the joint policy.

[0169] Step 4.3: During the distributed execution phase, the UAV trajectory agent independently generates a flight path based on the local observation of the UAV's current position; the beamforming agent independently generates a beamforming strategy based on the local observation of the channel state. Through the CTDE approach, the algorithm effectively balances training efficiency and execution constraints.

[0170] As a powerful machine learning tool, DRL has proven its effectiveness in intelligent decision-making across numerous application scenarios. In this scheme, agents receive higher rewards if their chosen actions improve the overall rate, consistent with maximizing the objective function. Therefore, the agent continuously updates its strategy through interaction with the environment, gradually optimizing to maximize the objective function. The strategy the agent hopes to learn is a mapping from state space to action space, fitted by a neural network. By rationally designing actions, states, and reward functions, DRL can effectively solve the optimization problem of maximizing the confidentiality rate. The two agents independently generate active and passive beamforming strategies and drone trajectory strategies, enabling more efficient optimization.

[0171] Based on the disclosure and teachings of the above description, those skilled in the art may also make changes and modifications to the above embodiments. Therefore, the present invention is not limited to the specific embodiments disclosed and described above, and any modifications and variations of the invention should also fall within the scope of protection of the claims of the present invention. In addition, although certain specific terms are used in this description, these terms are for convenience of description only and do not constitute any limitation to the present invention.

Claims

1. A reconfigurable intelligent surface-assisted secure communication optimization method based on deep reinforcement learning, characterized by: Utilize two intelligent agents to optimize confidentiality communication to maximize confidentiality rate; including the following steps: Step 1: The drone sends information to the user and collects environmental information; Step 2: Based on the collected environmental information, a deep reinforcement learning model is established to optimize the drone trajectory and beamforming. Step 3: Use the TA-PPO algorithm to optimize the deep reinforcement learning model; Step 4: Based on the optimized deep reinforcement learning model, the optimal joint beamforming scheme is obtained.

2. The reconfigurable intelligent surface-assisted secure communication optimization method based on deep reinforcement learning according to claim 1 is characterized in that: The realization of maximizing the confidentiality rate includes the following processes: There are K single-antenna legitimate users and one single-antenna eavesdropper; the UAV is equipped with an A-element uniform linear array, while the RIS uses a uniform planar array with M passive reflection units; the entire flight time T is evenly divided into N time slots, δ t Represents each individual time slot, then T=nδ t , n∈N; RIS is located at coordinate w R =(x R ,y R , z R ) T The coordinates of the user and the eavesdropper in time slot n are represented as w i =(x i [n],y i [n],z i [n]) T The coordinates of the drone at time slot n are q[n] = (x[n], y[n], H) T ; The speed of the drone is expressed as: The definition of the UAV maneuverability constraints is as follows: q[0]=(0,0,H), |x[n]|,|y[n]|≤B, Among them, q[0] is the initial coordinate of the UAV, B is the moving boundary of the UAV, and D max represents the maximum maneuvering distance of the UAV at time slot n, and H represents the height of the UAV.

3. The reconfigurable intelligent surface-assisted secure communication optimization method based on deep reinforcement learning according to claim 2 is characterized in that: The channel gains from the drone to the RIS, from the drone to the kth user, from the drone to the eavesdropper, from the RIS to the kth user, and from the RIS to the eavesdropper are expressed as The channel is modeled as a Rayleigh fading channel, and the channel modeling is as follows: Where i∈{k,p}, ρ0 represents the path loss at a reference distance of 1m, β1, β2, and β3 represent the relevant path loss exponents, and K1, K2, and K3 are the Rayleigh factors; and represents the NLoS component, where and are all independent, identically distributed, and follow a circularly symmetric complex Gaussian distribution with zero mean and unit variance; d U,i d R,i and d U,R denote the distances from the UAV to the user / eavesdropper, from the RIS to the user / eavesdropper, and from the UAV to the RIS, respectively; and represents the LoS component given by the uniform linear antenna array response. The calculation process of the LoS component is as follows: is the ULA’s steering vector, denoted as in Expressed as AoD azimuth d is the antenna spacing, λ c is the carrier wavelength; is the steering vector of UPA, expressed as: Among them, 0≤p, q≤n-1, and are the azimuth and altitude angles of the path between the RIS and the target object, respectively; The cascade channel from the drone to the user or eavesdropper is expressed as The RIS passive beamforming matrix is defined as Here, ψ is a diagonal matrix whose main diagonal elements are given by Calculated, represents the phase shift of the nth reflective element; The channel coefficients from the drone to all receivers are combined as follows: Therefore, the signal received by the i-th user from the drone is expressed as: y i =H c,i Gs+n i ; in, and Represent the beamforming matrix and transmission signal on the UAV respectively; n i represents background noise; g i represents the i-th column of G.

4. The method for optimizing secure communication assisted by reconfigurable intelligent surface based on deep reinforcement learning according to claim 3 is characterized in that: Set the decoding order for all users to Apply serial interference cancellation technology and set the decoding order from The first element to The last element of U K To U1; user The achievable data rate is expressed as: Among them, represents user U i the decoding data rate of user U when decoding its own signal; i ; represents the decoding data rate of user j when user U i decodes the signal of user U, user j < i; and are written as: Where x′=(U i-1 , U i-2 ,...,U j , ..., U2, U1) is A subset of The eavesdropper's eavesdropping rate on user k is: According to Weiner's eavesdropping code, the confidentiality rate is: R sk =(R k -R ek ) + , Among them, (a) + =max(0,a);R K represents the information rate of user k.

5. The reconfigurable intelligent surface-assisted secure communication optimization method based on deep reinforcement learning according to claim 4 is characterized in that: The confidentiality rate is maximized by adjusting the UAV’s trajectory Q, active beamforming G, and phase shift matrix ψ, which can be expressed as: s.t. Tr{GG T }≤P t , 6. The reconfigurable intelligent surface-assisted secure communication optimization method based on deep reinforcement learning according to claim 5 is characterized in that: The two agents include a beamforming agent and a drone trajectory agent, which are used to independently optimize the active and passive beamforming of RIS and independently optimize the drone trajectory; In step 1, the environmental information obtained by the beamforming agent at the nth moment is: the channel state information from the drone to the user and the eavesdropper at the n-1th moment; The environmental information obtained by the UAV trajectory agent at the nth moment is: the current position of the UAV; where n∈N.

7. The method for optimizing reconfigurable intelligent surface-assisted secure communication based on deep reinforcement learning according to claim 6 is characterized in that: In step 2, the establishment of the deep reinforcement learning model includes the following processes: Configure two agents and let them learn the optimal strategy π to maximize the cumulative discounted reward. The reward value of each round is as follows: Among them, r n is the reward value at step n, γ∈[0,1) is the discount factor used to weigh the importance of immediate rewards and future potential rewards; the agent’s actor network outputs the strategy, and the critic network is used to evaluate the strategy; The state value function V needs to be evaluated during the evaluation process. π (s) and the action-value function Q π (s, a) is evaluated; where the state value function V π (s) represents the expected value of the cumulative discounted reward obtained in the future when the agent is in state s under strategy π; the action value function Q π (s, a) represents the expected value of the cumulative discounted reward obtained in the future after the agent is in state s and performs action a under strategy π.

8. The method for optimizing reconfigurable intelligent surface-assisted secure communication based on deep reinforcement learning according to claim 7 is characterized in that: In step 3, the TA-PPO algorithm uses two agents to decouple variables and models the communication scheme as a Markov decision process; The beamforming agent is used to generate active and passive beamforming strategies. Its state is the CSI obtained through channel estimation, and its action is the UAV's active beamforming G and the phase shift matrix ψ of RIS. The UAV trajectory agent is used to generate the UAV's trajectory strategy. Its state is the UAV's position at the current moment, and its action is the UAV's position at the next moment. The two agents share a reward, which is the sum of the confidentiality rates of all users at the current moment. Through multiple rounds of iteration, the strategy is continuously optimized to maximize the cumulative reward until the training process converges.

9. The method for optimizing reconfigurable intelligent surface-assisted secure communication based on deep reinforcement learning according to claim 8, characterized in that: In step 3, the optimization of the deep reinforcement learning model includes the following processes: Importance sampling is used to implement offline strategy update; the importance sampling is to use the old strategy π′ θ (a|s) generates samples to update the learned strategy π θ (a|s), where the gradient is: The TA-PPO algorithm directly updates the random policy neural network π θ , to approximate the probability distribution of the action given the state, that is, Pr θ (q|s)=π θ (a|s);Q π (s, a) is given by the advantage function Instead, the critic network uses to estimate how well an action performs relative to the average action for a given state; Advantage function is defined as: in, is the baseline predicted by the critic network; The larger the value, the greater the probability that action a is selected by strategy π in state s; Using the clipping method, the objective function is directly clipped to: The goal of the critic network is to make the actual value Minimize the gap between parameters Update is performed by gradient descent method, i.e. is the loss function of the critic network, which is the mean square error function and is defined as:

10. The reconfigurable intelligent surface-assisted secure communication optimization method based on deep reinforcement learning according to claim 9 is characterized in that: In step 4, obtaining the optimal joint beamforming solution includes the following steps: Step 4.1: Build a joint optimization framework based on dual-agent collaboration to simultaneously optimize the active and passive beamforming parameters of the RIS and the UAV trajectory through centralized training and distributed execution. Step 4.2: During the centralized training phase, the beamforming agent and the drone trajectory agent share global state information, actions, and rewards. Based on the shared global state information, actions, and rewards, they estimate the advantage function and update the policy network to optimize the joint policy. Step 4.3: In the distributed execution phase, the UAV trajectory agent independently generates a flight path based on the local observed UAV's current position; the beamforming agent independently generates a beamforming strategy based on the local observed channel state.

Citation Information

Patent Citations

  • Resource optimization method of STAR-RIS communication system based on deep reinforcement learning

    CN117615393A

  • RIS-assisted MISO system optimization method based on deep reinforcement learning

    CN118900143A

  • Unmanned aerial vehicle communication secrecy energy efficiency optimization method and system based on deep reinforcement learning

    CN119255227A

  • Reconfigurable intelligent surface enabled leakage suppression and signal power maximization system and method

    WO2023037243A1

  • Intelligent reflecting surface assistance-based task unloading and resource allocation method for unmanned aerial vehicle mobile edge computing network system

    WO2025020222A1

Cited By

  • RIS-NOMA resource allocation method based on deep reinforcement learning

    CN120730371A

  • Optimization method and device for multi-unmanned aerial vehicle assisted ocean communication, equipment and medium

    CN120896637A