Exploration enhanced DDPG algorithm for RIS-UAV network optimization

By constructing the joint communication optimization model of RIS-UAV network and exploring the enhanced DDPG algorithm, the problems of insufficient exploration capabilities and constraint optimization of existing algorithms in RIS-UAV network optimization are solved, and the information transmission rate is improved and the convergence stability is enhanced.

CN120474581APending Publication Date: 2025-08-12CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510509474.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing DDPG and TD3 algorithms lack the ability to explore in RIS-UAV network optimization, and it is difficult to converge to the global optimal solution in a highly dynamic environment. It is difficult to set the penalty coefficient when dealing with constraint optimization problems, and it is difficult to weigh the relationship between different constraints, resulting in unstable learning behavior.

Method used

Build a joint communication optimization model of RIS-UAV network, design and explore enhanced DDPG algorithm, and jointly optimize UAV trajectory, UAV beamforming and RIS phase shift, combine Markov decision-making process and exploration enhancement strategy network, optimize information transmission rate, and use learnable noise distribution parameters to improve exploration capabilities.

Benefits of technology

The information transmission rate of the RIS-UAV network has been improved, achieving an enhancement effect of 31.25%, and the convergence stability and exploration ability of the algorithm in a dynamic environment are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120474581A_ABST
    Figure CN120474581A_ABST
Patent Text Reader

Abstract

The invention provides an exploration enhanced DDPG algorithm for RIS-UAV network optimization, which comprises the following steps: step 1, an RIS-UAV network joint communication optimization model integrating UAV trajectory control, UAV beam forming and RIS phase shift optimization is provided to improve the network transmission performance; and step 2, designing an exploration enhancement strategy network to solve the problem of insufficient exploration capability of the DDPG algorithm caused by deterministic action output by the strategy network. And finally, organically fusing the exploration enhancement strategy network into the DDPG algorithm to form an exploration enhancement DDPG algorithm. According to the method disclosed by the invention, the model established in the step 1 is optimized and solved by utilizing the method in the step 2, so that the information transmission rate of the RIS-UAV network is optimized. According to the invention, not only is a new method developed for RIS-UAV network optimization research, but also the application range of the DDPG algorithm in the RIS-UAV network is widened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing technology, and in particular to an exploratory enhanced DDPG algorithm for RIS-UAV network optimization. Background Art

[0002] Reconfigurable Intelligent Surfaces (RIS)-assisted Unmanned Aerial Vehicle (UAV) communications demonstrate significant potential in network performance and have become a key technology in 6G communications. RIS-UAV networks involve multiple key parameters, such as the UAV trajectory, RIS reflection matrix, and UAV beamforming matrix. These optimization variables exhibit high dimensionality. Furthermore, to ensure stable and reliable communication services for users, RIS-UAV networks must meet a series of constraints, including UAV transmit power and spatial location. Furthermore, the channel parameters between the UAV and the user constantly change during flight, further increasing the complexity of network optimization. Consequently, RIS-UAV network optimization is characterized by high dimensionality, multiple constraints, and high dynamics, placing high demands on the network's real-time optimization capabilities.

[0003] Deep reinforcement learning (DRL) organically combines the powerful feature extraction capabilities of deep neural networks with the excellent decision-making capabilities of reinforcement learning, becoming an important means of solving complex dynamic optimization problems. In the DRL framework, the system no longer relies on pre-collected training data. Instead, it adaptively learns based on the reward signals provided by the environment through real-time interaction with the environment, gradually building a precise model of the environment and dynamically adjusting the optimization strategy. Based on the above analysis, DRL can effectively overcome the high-dimensional and highly dynamic optimization challenges in RIS-UAV network communications by leveraging the effective representation of high-dimensional state spaces by deep neural networks and the dynamic learning characteristics of reinforcement learning.

[0004] DDPG and TD3, as mainstream algorithms in the field of DRL, are widely used for RIS-UAV network optimization. However, the core policy networks of these algorithms all output deterministic actions. This means that under the same state, the agent lacks the possibility to explore other potential actions, which greatly limits the agent's ability to explore in unknown environments. Especially in highly dynamic environments, this limitation makes it difficult for the algorithm to converge to the global optimal solution. Furthermore, when dealing with constrained optimization problems, current DRL algorithms adopt a strategy of applying different static negative rewards to violations of various constraints. This strategy has two major problems: first, it is extremely difficult to appropriately set the penalty coefficient, and even small changes can significantly affect the agent's learning behavior; second, static weight assignment makes it difficult to balance the relationships between different constraints, resulting in over-learning of some constraints and neglect of others, which greatly limits the DRL algorithm's ability to search for a feasible optimal solution. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide an exploration-enhanced DDPG algorithm for RIS-UAV network optimization, aiming to address the limitations of DDPG and TD3 in exploration capabilities.

[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is: An exploration-enhanced DDPG algorithm for RIS-UAV network optimization includes the following steps: Step 1: Construct a RIS-UAV network joint communication optimization model, including system model, channel model, information transmission rate model, and optimization model: Step 1.1, build a system model; Step 1.2, establish a channel model; Step 1.3, establish information transmission rate model; Step 1.4, establish the optimization model; Step 2: Propose an exploration-enhanced DDPG algorithm to optimize and solve the model established in step 1.

[0007] The specific process of establishing the system model in Step 1.1 above is: The system model is equipped with M UAVs with uniform linear array antennas, deployed on the surface of buildings K Unit RIS, and N Single-antenna ground users access the system in single-antenna receiving mode, forming a dynamic service demand scenario; Beamforming Matrix of UAV Array Antenna B It can be expressed as: ;(1) ;(2) Where, Respectively m The amplitude and phase of the antenna unit; then according to the UAV beamforming matrix B , the UAV transmission power can be calculated by the following formula: ;(3) Where, is the Frobenius norm, which is used to measure the energy of the beam transmitted by the UAV array antenna; RIS independently adjusts the phase shift angle of each unit based on real-time channel state information, so that the incident signal emitted by the UAV array antenna produces controllable wavefront deformation during the reflection process, thereby focusing the beam energy to the target user area; RIS phase shift matrix It can be expressed as: ;(4) Where, express t Time Slot k The phase shift angle of each RIS reflection unit, , represents a diagonal matrix.

[0008] The specific process of establishing the channel model in Step 1.2 above is: The channel model includes: direct channel from UAV to user, and reflected channel from UAV to RIS and RIS to user; The direct connection channel from UAV to user is: ;(5) Where, for t Time slot UAV and n The distance between users; is the path loss at a reference distance of 1 m; is the path loss index; is a random scattering component that obeys a complex Gaussian distribution with a mean of 0 and a variance of 1; The channels from UAV to RIS and from RIS to user are: ;(6) ;(7) Where, is the path loss at a reference distance of 1 m; and Respectively expressed in t Time slot, distance between UAV and RIS, and distance between RIS and usern distance; and is the path loss index; is the Rice factor; and They represent the NLoS component of the channel between UAV and RIS, and the NLoS component of the channel from RIS to user. n The NLoS components of the channels between them are random scattering components that obey a complex Gaussian distribution with a mean of 0 and a variance of 1; and They represent the LoS component of the channel between UAV and RIS, and the LoS component of the channel from RIS to user. n The LoS component of the channel between In summary, the channel from UAU to user is: ;(8) Where, For the channel from UAV to RIS, For the channel from RIS to users, is the RIS phase shift matrix.

[0009] The specific process of establishing the information transmission rate model in Step 1.3 above is: exist t Time slot users n The received signal-to-interference-and-noise ratio can be calculated as: ;(9) Where, is the noise power, is the first in the UAV array antenna beamforming matrix i A quantity, N is the number of users; At the same time, according to Shannon’s formula, the total information transmission rate of the RIS-UAV network is: ;(10).

[0010] The specific process of establishing the optimization model in Step 1.4 above is: By jointly optimizing the UAV trajectory and UAV beamforming matrix B、 RIS phase shift matrix , to maximize the information transmission rate of the RIS-UAV network; for this purpose, the established mathematical model is as follows: ;(11) Where, Indicates the maximum value of the UAV transmission power, and UAV in tThe x-axis and y-axis positions of the time slot; C1 ensures that the actual transmission power of the UAV does not exceed its maximum value; C2 and C3 require the UAV to fly within the specified spatial range; C4 indicates that the phase shift angle of the RIS reflector unit needs to be limited to a fixed range.

[0011] The above Step 2 includes the following steps: Step 2.1, construction of Markov decision process; Step 2.2, design of state, action and reward functions; Step 2.3. Explore the construction of enhanced DDPG algorithm.

[0012] The construction of the Markov decision process in Step 2.1 above specifically includes: By jointly modeling the temporal coupling relationship between the UAV trajectory, the UAV array antenna beamforming matrix and the RIS phase shift matrix, it is concluded that the optimization problem in Equation (11) is equivalent to solving a reinforcement learning task with a continuous state-action space.

[0013] The design of the state, action, and reward function in Step 2.2 above specifically includes: Step 2.2.1, State Design: t Time slot, status It is composed of the UAV's position information and the channel information between the UAV and the user; therefore, t The state of a time slot can be defined as: ;(12) ;(13) Where, Represents UAV position information, Represents the cascade channel information between the UAV and the user; Step 2.2.2, Action Design: t Time slot, action It is composed of the UAV flight speed and heading angle, the UAV beamforming matrix, and the RIS phase shift matrix; therefore, t The actions of a time slot can be defined as: ;(14) Where, is the UAV flight speed, is the UAV heading angle, B is the UAV beamforming matrix, is the RIS phase shift matrix; Step 2.2.3, reward function design: t Time slot, the agent observes the state information and perform actions , and then calculate the immediate reward based on the reward function; the reward function is designed as follows: For constraint C1, the designed constraint violation is as follows: ;(15) Where, Indicates that the UAV is t The x-axis position of the time slot, and are the upper and lower bounds of the UAV on the x-axis respectively; For constraint C2, the designed constraint violation is as follows: ; (16) Where, Indicates that the UAV is t The y-axis position of the time slot, and are the upper and lower bounds of the UAV on the y-axis respectively; For constraint C3, the following transmit power constraint violation calculation formula is designed: ; (17) Where, Indicates the UAV transmission power, Indicates the maximum value of the UAV transmission power; Furthermore, the above three constraints C1, C2 and C3 are integrated into constraint violation functions, as follows: ; (18) Where, are different penalty coefficients; Finally, the constraint violation degree is integrated into the information transmission rate model to construct the agent’s reward function as follows: ; (19) Where, is the comprehensive penalty factor.

[0014] The exploration-enhanced DDGP algorithm in the above Step 2.3 specifically includes: agent setting, exploration-enhanced strategy network, value network, target strategy network and target value network.

[0015] Step 2.3.1, the agent is set as: ; (20) In the formula represents a parameterized exploration-enhanced policy network, where Expressed as A collection of represents the parameterized value network; meanwhile, the target policy network Directly copied from the exploration enhancement strategy network and target value network Directly replicated in the value network; Step 2.3.2. Explore the Enhanced Strategy Network The network operation process can be formally defined as: ;(twenty one) ;(twenty two) ;(twenty three) Where, and are the original weights and biases in the network, and are the learnable noise amplitude parameters, and are independent Gaussian vectors with zero mean, represents the Hadamard product operator; Furthermore, the noise amplitude parameter Incorporate into the set of learnable parameters , realizing the joint learning of noise intensity and policy optimization; in order to quantify the impact of noise on policy performance, the following loss function is defined: ;(twenty four) At the same time, the exploration enhancement strategy network uses the Monte Carlo method to approximate the gradient of the loss function, which is defined as follows: ; (25) Where, is the number of experiences randomly sampled from the experience replay pool, represents the derivative function; According to the gradient calculated above, the network parameters are updated as follows: ; (26) Where, is the learning rate, represents the derivative function; Step 2.3.3. Value Network The value network is based on random sampling from the experience replay pool. Based on this experience, the network is updated by minimizing the loss function, which can be approximated as: ; (27) ; (28) in, yt An approximate target generated by the target network based on randomly sampled experience Q Value. The value network parameters are updated by gradient descent, and its loss function For the value network parameters The gradient of is calculated as follows: ;(29) Furthermore, the value network parameters are updated by gradient descent as follows: ; (30) in, is the value network learning rate, represents the derivative function; Step 2.3.4, Target Strategy Network and Target Value Network Both the target strategy network and the target value network adopt soft update method, and their update processes are as follows: ; (31) ; (32) Where, Represents the learning rate of network update.

[0016] The present invention provides an exploration-enhanced DDPG algorithm for RIS-UAV network optimization. First, by integrating UAV trajectory control, UAV beamforming, and RIS phase shift optimization, a RIS-UAV network joint communication optimization model is constructed, thus laying the foundation for improving network performance. Secondly, an exploration-enhanced strategy network is designed to greatly enhance the exploration capability of the DDPG algorithm. Finally, the exploration-enhanced strategy network is organically integrated into the DDPG algorithm to form an exploration-enhanced DDPG algorithm, which ultimately significantly enhances the information transmission rate of the RIS-UAV network. Compared with the original DDPG algorithm, the method of the present invention increases the information transmission rate by 31.25%. The present invention not only opens up a new method for RIS-UAV network optimization research, but also expands the application scope of deep reinforcement learning algorithms in the field of wireless communications. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The present invention will be further described below with reference to the accompanying drawings and embodiments: Figure 1 This is a structural diagram of the RIS-UAV network system of the present invention; Figure 2 This is a network structure diagram of the exploration enhancement strategy of the present invention; Figure 3 This is a schematic diagram of the results of the embodiment of the present invention Figure 1 ; Figure 4This is a schematic diagram of the results of the embodiment of the present invention Figure 2 . DETAILED DESCRIPTION

[0018] The technical solution of the present invention is described in detail below with reference to the accompanying drawings and embodiments.

[0019] An exploration-enhanced DDPG algorithm for RIS-UAV network optimization includes the following steps: Step 1: Construct a RIS-UAV network joint communication optimization model, including system model, channel model, information transmission rate model, and optimization model: Step 1.1, build a system model; Step 1.2, establish a channel model; Step 1.3, establish information transmission rate model; Step 1.4, establish the optimization model; Step 2: Propose an exploration-enhanced DDPG algorithm to optimize and solve the model established in step 1.

[0020] The specific process of establishing the system model in Step 1.1 above is: The system model is equipped with M UAVs with uniform linear array antennas, deployed on the surface of buildings K Unit RIS, and N The system consists of single-antenna ground users. During flight, the Line of Sight (LoS) link between the UAV and the user may be interrupted by buildings. To address this, programmable RIS is deployed on the surface of buildings, leveraging the phase control capability of its reflective units to reconstruct the wireless propagation environment and effectively compensate for link interruptions. Ground users access the system in single-antenna reception mode, creating a dynamic service demand scenario. Beamforming Matrix of UAV Array Antenna B It can be expressed as: ;(1) ;(2) Where, Respectively m The amplitude and phase of the antenna unit; then according to the UAV beamforming matrix B , the UAV transmission power can be calculated by the following formula: ;(3) Where, is the Frobenius norm, which is used to measure the energy of the beam transmitted by the UAV array antenna; RIS independently adjusts the phase shift angle of each unit based on real-time channel state information, so that the incident signal emitted by the UAV array antenna produces controllable wavefront deformation during the reflection process, thereby focusing the beam energy to the target user area; RIS phase shift matrix It can be expressed as: ;(4) Where, express t Time Slot k The phase shift angle of each RIS reflection unit, , represents a diagonal matrix.

[0021] The specific process of establishing the channel model in Step 1.2 above is: The channel model includes: direct channel from UAV to user, and reflected channel from UAV to RIS and RIS to user; The direct connection channel from UAV to user is: ;(5) Where, for t Time slot UAV and n The distance between users; is the path loss at a reference distance of 1 m; is the path loss index; is a random scattering component that obeys a complex Gaussian distribution with a mean of 0 and a variance of 1; The channels from UAV to RIS and from RIS to user are: ;(6) ;(7) Where, is the path loss at a reference distance of 1 m; and Respectively expressed in t Time slot, distance between UAV and RIS, and distance between RIS and user n distance; and is the path loss index; is the Rice factor; and They represent the NLoS component of the channel between UAV and RIS, and the NLoS component of the channel from RIS to user. n The NLoS components of the channels between them are random scattering components that obey a complex Gaussian distribution with a mean of 0 and a variance of 1; and They represent the LoS component of the channel between UAV and RIS, and the LoS component of the channel from RIS to user. n The LoS component of the channel between In summary, the channel from UAU to user is: ;(8) Where, For the channel from UAV to RIS, For the channel from RIS to users, is the RIS phase shift matrix.

[0022] The specific process of establishing the information transmission rate model in Step 1.3 above is: exist t Time slot users n The received signal-to-interference-and-noise ratio can be calculated as: ;(9) Where, is the noise power, is the first in the UAV array antenna beamforming matrix i A quantity, N is the number of users; At the same time, according to Shannon’s formula, the total information transmission rate of the RIS-UAV network is: ;(10).

[0023] The specific process of establishing the optimization model in Step 1.4 above is: By jointly optimizing the UAV trajectory and UAV beamforming matrix B、 RIS phase shift matrix , to maximize the information transmission rate of the RIS-UAV network; for this purpose, the established mathematical model is as follows: ;(11) Where, Indicates the maximum value of the UAV transmission power, and UAV in t The x-axis and y-axis positions of the time slot; C1 ensures that the actual transmission power of the UAV does not exceed its maximum value; C2 and C3 require the UAV to fly within the specified spatial range; C4 indicates that the phase shift angle of the RIS reflector unit needs to be limited to a fixed range.

[0024] The above Step 2 includes the following steps: Step 2.1, construction of Markov decision process; Step 2.2, design of state, action and reward functions; Step 2.3. Explore the construction of enhanced DDPG algorithm.

[0025] The construction of the Markov decision process in Step 2.1 above specifically includes: The trajectory change of the UAV is only affected by its previous position and velocity, so the trajectory change of the UAV follows the characteristics of the MDP. In addition, since the UAV array antenna beamforming matrix and the RIS phase shift matrix are time-varying observations, their parameter space can be mapped to the action space of the MDP. At the same time, the optimization objective of the transmission rate defined in Equation (11) corresponds to the reward function of the MDP. Therefore, by jointly modeling the temporal coupling relationship between the UAV trajectory, the UAV array antenna beamforming matrix and the RIS phase shift matrix, it can be concluded that the optimization problem in Equation (11) is equivalent to solving a reinforcement learning task with a continuous state-action space.

[0026] The design of the state, action, and reward function in Step 2.2 above specifically includes: Step 2.2.1, State Design: t Time slot, status It is composed of the UAV's position information and the channel information between the UAV and the user; therefore, t The state of a time slot can be defined as: ;(12) ;(13) Where, Represents UAV position information, Represents the cascade channel information between the UAV and the user; Step 2.2.2, Action Design: t Time slot, action It is composed of the UAV flight speed and heading angle, the UAV beamforming matrix, and the RIS phase shift matrix; therefore, t The actions of a time slot can be defined as: ;(14) Where, is the UAV flight speed, is the UAV heading angle, B is the UAV beamforming matrix, is the RIS phase shift matrix; Step 2.2.3, reward function design: t Time slot, the agent observes the state information and perform actions , and then calculate the immediate reward based on the reward function; the reward function is designed as follows: For constraint C1, the designed constraint violation is as follows: ;(15) Where, Indicates that the UAV is t The x-axis position of the time slot, and are the upper and lower bounds of the UAV on the x-axis respectively; For constraint C2, the designed constraint violation is as follows: ; (16) Where, Indicates that the UAV is t The y-axis position of the time slot, and are the upper and lower bounds of the UAV on the y-axis respectively; For constraint C3, the following transmit power constraint violation calculation formula is designed: ; (17) Where, Indicates the UAV transmission power, Indicates the maximum value of the UAV transmission power; Furthermore, the above three constraints C1, C2 and C3 are integrated into constraint violation functions, as follows: ; (18) Where, are different penalty coefficients; Finally, the constraint violation degree is integrated into the information transmission rate model to construct the agent’s reward function as follows: ; (19) Where, is the comprehensive penalty factor.

[0027] The exploration-enhanced DDGP algorithm in the above Step 2.3 specifically includes: agent setting, exploration-enhanced strategy network, value network, target strategy network and target value network.

[0028] Step 2.3.1, the agent is set as: ; (20) In the formula represents a parameterized exploration-enhanced policy network, where Expressed as A collection of represents the parameterized value network; meanwhile, the target policy network Directly copied from the exploration enhancement strategy network and target value network Directly replicated in the value network; Step 2.3.2. Explore the Enhanced Strategy Network The core of the exploration enhancement strategy network is to embed the learnable noise distribution parameters into the network weight update process, and achieve adaptive optimization of exploration ability by dynamically adjusting the noise intensity. The network operation process can be formally defined as: ;(twenty one) ;(twenty two) ;(twenty three) Where, and are the original weights and biases in the network, and are the learnable noise amplitude parameters, and are independent Gaussian vectors with zero mean, represents the Hadamard product operator; Furthermore, the noise amplitude parameter Incorporate into the set of learnable parameters , realizing the joint learning of noise intensity and policy optimization; in order to quantify the impact of noise on policy performance, the following loss function is defined: ;(twenty four) At the same time, the exploration enhancement strategy network uses the Monte Carlo method to approximate the gradient of the loss function, which is defined as follows: ; (25) Where, is the number of experiences randomly sampled from the experience replay pool, represents the derivative function; According to the gradient calculated above, the network parameters are updated as follows: ; (26) Where, is the learning rate, represents the derivative function; Step 2.3.3. Value Network The value network is based on random sampling from the experience replay pool. Based on this experience, the network is updated by minimizing the loss function, which can be approximated as: ; (27) ; (28) in, y tAn approximate target generated by the target network based on randomly sampled experience Q Value. The value network parameters are updated by gradient descent, and its loss function For the value network parameters The gradient of is calculated as follows: ;(29) Furthermore, the value network parameters are updated by gradient descent as follows: ; (30) in, is the value network learning rate, represents the derivative function; Step 2.3.4, Target Strategy Network and Target Value Network Both the target strategy network and the target value network adopt soft update method, and their update processes are as follows: ; (31) ; (32) Where, Represents the learning rate of network update.

[0029] Example: The embodiment of the present invention provides an exploratory enhanced DDPG algorithm for RIS-UAV network optimization, comprising the following steps: Step 1: Construct a RIS-UAV network joint communication optimization model, which includes a system model, a channel model, an information transmission rate model, and an optimization model. Step 2: Propose an exploration-enhanced DDPG algorithm to optimize and solve the model established in step 1.

[0030] Step 1 includes the following steps: Step 1.1, establish the system model; Step 1.2, establish a channel model; Step 1.3, establish an information transmission rate model; Step 1.4, establish the optimization model.

[0031] Step 2 includes the following steps: Step 2.1, Construction of Markov Decision Process Step 2.2, design of state, action and reward functions; Step 2.3, explore the construction of the enhanced DDPG algorithm.

[0032] After 140 training rounds, the agent's reward value remained stable, indicating that the algorithm gradually converged. However, the reward value obtained by DDGP has been fluctuating, indicating that its convergence stability is insufficient. Figure 1 As shown in the figure, the horizontal axis represents the number of training rounds and the vertical axis represents the reward value. At the same time, compared with other algorithms - TD3 and DDPG, the method of the present invention achieves a higher information transmission rate and has better convergence stability. This shows that the method of the present invention effectively guides the agent to conduct more extensive exploration in the environment by introducing the exploration enhancement strategy network, thereby greatly improving the algorithm's exploration ability and convergence performance. The specific results are shown in the figure. Figure 2 As shown, the horizontal axis represents the training rounds and the vertical axis represents the information transmission rate.

Claims

1. A DDPG algorithm for RIS-UAV network optimization, characterized by: Including steps: Step 1: Construct a RIS-UAV network joint communication optimization model, including system model, channel model, information transmission rate model, and optimization model: Step 1.1, build a system model; Step 1.2, establish a channel model; Step 1.3, establish information transmission rate model; Step 1.4, establish the optimization model; Step 2: Propose an exploration-enhanced DDPG algorithm to optimize and solve the model established in step 1.

2. The exploratory enhanced DDPG algorithm for RIS-UAV network optimization according to claim 1 is characterized in that: The specific process of establishing the system model in Step 1.1 is as follows: The system model is equipped with M UAVs with uniform linear array antennas, deployed on the surface of buildings K Unit RIS, and N Single-antenna ground users access the system in single-antenna receiving mode, forming a dynamic service demand scenario; Beamforming Matrix of UAV Array Antenna B It can be expressed as: ;(1) ; (2) Where, Respectively m The amplitude and phase of the antenna unit; then according to the UAV beamforming matrix B , the UAV transmission power can be calculated by the following formula: ; (3) Where, is the Frobenius norm, which is used to measure the energy of the beam transmitted by the UAV array antenna; Based on real-time channel state information, RIS independently adjusts the phase shift angle of each unit, causing the incident signal transmitted by the UAV array antenna to produce controllable wavefront deformation during the reflection process, thereby focusing the beam energy on the target user area; RIS phase shift matrix It can be expressed as: ; (4) Where, express t Time Slot k The phase shift angle of each RIS reflection unit, , represents a diagonal matrix.

3. The exploratory enhanced DDPG algorithm for RIS-UAV network optimization according to claim 2, characterized in that: The specific process of establishing the channel model in Step 1.2 is as follows: The channel model includes: direct channel from UAV to user, and reflected channel from UAV to RIS and RIS to user; The direct connection channel from UAV to user is: ;(5) Where, for t Time slot UAV and n The distance between users; is the path loss at a reference distance of 1 m; is the path loss index; is a random scattering component that obeys a complex Gaussian distribution with a mean of 0 and a variance of 1; The channels from UAV to RIS and from RIS to user are: ;(6) ;(7) Where, is the path loss at a reference distance of 1 m; and Respectively expressed in t Time slot, distance between UAV and RIS, and distance between RIS and user n distance; and is the path loss index; is the Rice factor; and They represent the NLoS component of the channel between UAV and RIS, and the NLoS component of the channel from RIS to user. n The NLoS components of the channels between them are random scattering components that obey a complex Gaussian distribution with a mean of 0 and a variance of 1; and They represent the LoS component of the channel between UAV and RIS, and the LoS component of the channel from RIS to user. n The LoS component of the channel between In summary, the channel from UAU to user is: ;(8) Where, For the channel from UAV to RIS, For the channel from RIS to users, is the RIS phase shift matrix.

4. The exploratory enhanced DDPG algorithm for RIS-UAV network optimization according to claim 3 is characterized in that: The specific process of establishing the information transmission rate model in Step 1.3 is as follows: exist t Time slot users n The received signal-to-interference-and-noise ratio can be calculated as: ; (9) Where, is the noise power, is the first in the UAV array antenna beamforming matrix i A quantity, N is the number of users; At the same time, according to Shannon’s formula, the total information transmission rate of the RIS-UAV network is: ;(10)。 5. The exploratory enhanced DDPG algorithm for RIS-UAV network optimization according to claim 4 is characterized in that: The specific process of establishing the optimization model in Step 1.4 is as follows: By jointly optimizing the UAV trajectory and UAV beamforming matrix B、 RIS phase shift matrix , to maximize the information transmission rate of the RIS-UAV network; for this purpose, the established mathematical model is as follows: ; (11) Where, Indicates the maximum value of the UAV transmission power, and UAV in t The x-axis and y-axis positions of the time slot; C1 ensures that the actual transmission power of the UAV does not exceed its maximum value; C2 and C3 require the UAV to fly within a specified spatial range; C4 indicates that the phase shift angle of the RIS reflection unit needs to be limited to a fixed range.

6. The exploratory enhanced DDPG algorithm for RIS-UAV network optimization according to claim 5, characterized in that: The Step 2 includes the following steps: Step 2.1, construction of Markov decision process; Step 2.2, design of state, action and reward functions; Step 2.

3. Explore the construction of enhanced DDPG algorithm.

7. The exploratory enhanced DDPG algorithm for RIS-UAV network optimization according to claim 6, characterized in that: The construction of the Markov decision process in Step 2.1 specifically includes: By jointly modeling the temporal coupling relationship between the UAV trajectory, the UAV array antenna beamforming matrix and the RIS phase shift matrix, it is concluded that the optimization problem in Equation (11) is equivalent to solving a reinforcement learning task with a continuous state-action space.

8. The exploratory enhanced DDPG algorithm for RIS-UAV network optimization according to claim 7, characterized in that: The design of the state, action, and reward function in Step 2.2 specifically includes: Step 2.2.1, State Design: t Time slot, status It is composed of the UAV's position information and the channel information between the UAV and the user; therefore, t The state of a time slot can be defined as: ;(12) ;(13) Where, Represents UAV position information, Represents the cascade channel information between the UAV and the user; Step 2.2.2, Action Design: t Time slot, action It is composed of the UAV flight speed and heading angle, the UAV beamforming matrix, and the RIS phase shift matrix; therefore, t The actions of a time slot can be defined as: ;(14) Where, is the UAV flight speed, is the UAV heading angle, B is the UAV beamforming matrix, is the RIS phase shift matrix; Step 2.2.3, reward function design: t Time slot, the agent observes the state information and perform actions , and then calculate the immediate reward based on the reward function; the reward function is designed as follows: For constraint C1, the designed constraint violation is as follows: ;(15) Where, Indicates that the UAV is t The x-axis position of the time slot, and are the upper and lower bounds of the UAV on the x-axis respectively; For constraint C2, the designed constraint violation is as follows: ;(16) Where, Indicates that the UAV is t The y-axis position of the time slot, and are the upper and lower bounds of the UAV on the y-axis respectively; For constraint C3, the following transmit power constraint violation calculation formula is designed: ;(17) Where, Indicates the UAV transmission power, Indicates the maximum value of the UAV transmission power; Furthermore, the above three constraints C1, C2 and C3 are integrated into constraint violation functions, as follows: ;(18) Where, are different penalty coefficients; Finally, the constraint violation degree is integrated into the information transmission rate model to construct the agent’s reward function as follows: ;(19) Where, is the comprehensive penalty factor.

9. The exploratory enhanced DDPG algorithm for RIS-UAV network optimization according to claim 8, characterized in that: The exploration-enhanced DDGP algorithm in Step 2.3 specifically includes: agent setting, exploration-enhanced strategy network, value network, target strategy network and target value network. Step 2.3.1, the agent is set as: ;(20) In the formula represents a parameterized exploration-enhanced policy network, where Expressed as A collection of represents the parameterized value network; meanwhile, the target policy network Directly copied from the exploration enhancement strategy network and target value network Directly replicated in the value network; Step 2.3.

2. Explore the Enhanced Strategy Network The network operation process can be formally defined as: ;(21) ;(22) ;(23) Where, and are the original weights and biases in the network, and are the learnable noise amplitude parameters, and are independent Gaussian vectors with zero mean, represents the Hadamard product operator; Furthermore, the noise amplitude parameter Incorporate into the set of learnable parameters , realizing the joint learning of noise intensity and policy optimization; in order to quantify the impact of noise on policy performance, the following loss function is defined: ;(24) At the same time, the exploration enhancement strategy network uses the Monte Carlo method to approximate the gradient of the loss function, which is specifically defined as follows: ;(25) Where, is the number of experiences randomly sampled from the experience replay pool, represents the derivative function; According to the gradient calculated above, the network parameters are updated as follows: ;(26) Where, is the learning rate, represents the derivative function; Step 2.3.

3. Value Network The value network is based on random sampling from the experience replay pool. Based on this experience, the network is updated by minimizing the loss function, which can be approximated as: ;(27) ;(28) in, y t An approximate target generated by the target network based on randomly sampled experience Q Value. The value network parameters are updated by gradient descent, and its loss function For the value network parameters The gradient of is calculated as follows: ;(29) Furthermore, the value network parameters are updated by gradient descent as follows: ;(30) in, is the value network learning rate, represents the derivative function; Step 2.3.4, Target Strategy Network and Target Value Network Both the target strategy network and the target value network adopt soft update method, and their update processes are as follows: ;(31) ;(32) Where, Represents the learning rate of network update.