RIS auxiliary cognitive unmanned aerial vehicle system resource joint optimization method based on reinforcement learning

By adopting RIS-assisted resource optimization method based on reinforcement learning in cognitive drone systems, combined with deep reinforcement learning and reconstructible intelligent surface technology, the challenges of communication performance optimization in scenarios such as dynamic spectrum changes and multi-UAV collaborative work are solved, and efficient spectrum allocation and optimization are achieved, significantly improving the system's throughput and reliability.

CN119997213APending Publication Date: 2025-05-13HAINAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510097750.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the case of dynamic spectrum changes, fluctuations in user demands and collaborative work of multiple UAVs, it is difficult for the prior art to effectively optimize spectrum utilization and communication performance.

Method used

The RIS assisted cognitive UAV system resource joint optimization method based on reinforcement learning is adopted, and efficient spectrum allocation and optimization is achieved in a dynamic spectrum environment by combining cognitive radio technology, deep reinforcement learning and reconstructible intelligent surface technology. The specific steps include establishing a cognitive wireless communication system model, perception model, and transmission model, and optimizing the system transmission rate through a deep reinforcement learning framework.

Benefits of technology

It significantly improves the effective throughput and communication reliability of the system, can effectively respond to challenges in dynamic and uncertain environments, and improves spectrum efficiency and overall performance of the communication system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119997213A_ABST
    Figure CN119997213A_ABST
Patent Text Reader

Abstract

Aiming at a communication scene between a cognitive unmanned aerial vehicle and a ground mobile secondary user, the invention provides an intelligent reflecting surface (RIS) assisted cognitive unmanned aerial vehicle (CUAV) system resource joint optimization method based on reinforcement learning. The RIS-CUAV system comprises a cognitive wireless communication system model, a perception model, a transmission model and a channel model for RIS-assisted cognitive UAV network transmission. Based on the system, a mathematical model for maximizing the transmission rate of the RIS-CUAV system is established by jointly optimizing the spectrum sensing duration, the 3D trajectory of the UAV and the RIS phase. And then solving the mathematical model, and constructing a deep reinforcement learning framework based on the double-depth q-network. According to the method, RIS phase shift, spectrum sensing duration and unmanned aerial vehicle 3D trajectory can be adjusted in real time to adapt to dynamic environmental conditions, and simulation results show that the method can significantly improve the transmission rate of the RIS-CUAV system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of non-cognitive unmanned aerial vehicle network technology, and in particular to a RIS-assisted cognitive unmanned aerial vehicle system resource joint optimization method based on reinforcement learning. Background Art

[0002] With the widespread application of drones in commercial, civil and military fields, the efficient use of spectrum resources has become a key issue that needs to be solved urgently. Especially with the rapid development of emerging technologies such as 5G and the Internet of Things, the demand for communication bandwidth has increased dramatically, and traditional spectrum allocation methods are difficult to meet this growing demand. To meet this challenge, a cognitive unmanned aerial vehicle (UAV) system is proposed, which combines cognitive radio technology with drone networks to improve bandwidth utilization and spectrum resource efficiency.

[0003] Cognitive UAVs are able to operate in a shared spectrum environment and predict spectrum availability through machine learning techniques, thereby optimizing operating strategies, minimizing communication delays and improving overall performance. In critical missions such as disaster relief, public safety and military operations, cognitive UAV networks ensure uninterrupted communications in critical situations and guarantee the reliability of information flow. With the continuous development of UAV technology, the deep integration of cognitive radio and UAV networks is expected to play an important role in the next generation of wireless communication systems. However, in scenarios such as dynamic spectrum changes, fluctuations in user demand and collaborative work of multiple UAVs, the optimization of spectrum utilization and communication performance still faces huge challenges. To address these challenges, Deep Reinforcement Learning (DRL) technology has been introduced to autonomously learn and optimize the flight trajectory and communication service quality of UAVs.

[0004] In recent years, the Reconfigurable Intelligence Surface (RIS) technology has significantly improved spectrum sensing accuracy and communication quality by intelligently adjusting the phase of wireless signals to optimize the propagation path. The phase optimization of RIS has shown excellent communication performance improvement in the downlink of UAV networks and significantly improved the throughput of RIS-UAV relay communication. However, most existing technologies have not been able to fully consider complex environmental factors such as dynamic changes in spectrum, sudden user demands, and UAV collaborative work. The joint optimization of RIS phase, UAV 3D trajectory, and spectrum sensing strategy is still an open research area that needs to be solved in cognitive UAV networks. Summary of the invention

[0005] In view of this, the purpose of the present invention is to provide a RIS-assisted cognitive UAV system resource joint optimization method based on reinforcement learning, aiming to solve the problems of low spectrum utilization, poor communication service quality, and complex collaborative work in the prior art; by combining cognitive radio technology, deep reinforcement learning (DRL) and reconfigurable intelligent surface (RIS) technology, the present invention can achieve efficient spectrum allocation and optimization in a dynamic spectrum environment, maximize the effective throughput of the network and improve the overall system performance;

[0006] To achieve the above-mentioned invention object, the present invention provides a RIS-assisted cognitive UAV system resource joint optimization method based on reinforcement learning, comprising:

[0007] S1. Establish a cognitive wireless communication system model for RIS-assisted cognitive UAV network transmission;

[0008] S2. Establish a perception model. The UAV uses the energy detection method to perceive the spectrum occupancy. It is assumed that the PU spectrum state remains unchanged in each time slot. Based on the spectrum sensing results, the UAV transmits data to the ground user through the time division multiple access (TDMA) protocol.

[0009] S3. Establish a transmission model, including channel model; channel gain and reflection link; total channel gain; data transmission rate and total system throughput;

[0010] S4. Establish a system optimization problem with the goal of maximizing the user data transmission rate, and maximize the system throughput by combining spectrum sensing duration, RIS phase shift design, and UAV trajectory optimization;

[0011] S5. Solve the system transmission rate optimization problem by building a deep reinforcement learning framework;

[0012] Preferably, in step S1, the system involves an outdoor downlink communication scenario including a base station (PT), a reflective intelligent surface (RIS), an unmanned aerial vehicle (UAV) and multiple mobile users, and optimizes the spectrum sensing time to maximize the transmission time, thereby improving the management and allocation efficiency of spectrum resources and significantly improving the effective throughput of the system by combining the phase configuration of the RIS and the trajectory design of the UAV, thereby meeting the user's demand for high-speed data transmission; the IRS is uniformly distributed in an array and the number of reflective units is M=M x ×M y , the phase shift matrix of IRS in the nth time slot can be expressed as: The user moves according to the random walk model, and the user's movement direction and speed are completely random at any time t;

[0013] Preferably, in step S2, the system adopts a single antenna configuration, and the UAV performs spectrum sensing and data transmission by dividing the flight period T into N time slots. In each time slot, when the UAV detects that the spectrum is idle, the UAV senses the spectrum for a sensing duration τ. n Complete spectrum access preparation within t n -τ n Data transmission. False alarm probability P f,t Calculated by Gaussian Q function:

[0014]

[0015] Preferably, in step S3, the air-to-ground channel model between the drone and the user is provided according to the 3GPP version 15 specification, and the path loss depends on the line-of-sight (LoS) and non-line-of-sight (NLoS) states, expressed as P LOS / NOLS (t), line-of-sight probability P LOS , P NLOS Depending on the distance between the user and the UAV and the height of the UAV h u (t) Calculate the channel gain g between UAV and user uk (t) Calculated by direct link; channel gain g of the reflection link urk (t) is calculated from the path loss between RIS and UAV and user; the total channel gain of the reflection link is:

[0016] g urk (t)=|g rk (t) T Θg ur (t)| 2 (2)

[0017] The total channel gain between the drone and the user is:

[0018]

[0019] The signal-to-noise ratio and data rate are calculated as:

[0020]

[0021]

[0022] The optimized system throughput R(t) is the weighted average data rate in each time slot, expressed as

[0023]

[0024] The overall throughput is:

[0025]

[0026] Preferably, in step S4, under the constraints of UAV flight space, perception duration, maximum power, RIS phase shift and service quality, the continuous perception time, flight trajectory and RIS phase shift of the UAV are optimized, with the goal of maximizing the effective total throughput of the UAV during the service period; the position of the UAV during the service period is defined as H, and the perception duration is defined as Γ=τ n ; The flight speed is set to be constant, and the transmission power is P u ; The power allocation strategy is P = {P k (t), 0≤t≤T, k∈K}; the optimization problem is formalized as:

[0027]

[0028]

[0029]

[0030]

[0031] 0≤τ n ≤T n (8.2)

[0032]

[0033]

[0034]

[0035] Position constraint (8.1): ensure that the UAV is within the effective height range above the service area to avoid collision with ground obstacles and ensure stable signal coverage; Perception duration constraint (8.2): limit the spectrum perception duration, optimize the switching between transmission and perception, and reduce user delay; Power constraint (8.3): limit the maximum transmission power of each UAV to prevent power consumption from exceeding the maximum power limit of the device; Phase setting constraint (8.4): ensure that the phase of each reflection unit is within the control range to optimize the signal propagation path; Quality of service constraint (8.5): ensure that all users obtain a certain quality of service in the same service area to avoid uneven resource allocation;

[0036] Preferably, in step S5, the deep reinforcement learning framework divides the system transmission rate maximization problem into two parts: UAV spectrum sensing duration optimization and UAV trajectory and RIS phase optimization; the specific steps include:

[0037] S501. Optimization of sensing duration in UAV-assisted communication: In UAV-assisted communication systems, optimizing spectrum sensing time is crucial to balancing spectrum sensing accuracy and transmission time; the sensing time problem at time t can be expressed as (1-P f,t )(t n -τ n ), where P f,t represents the probability of spectrum sensing failure, t n is the total available time, τ n is the perception time; to optimize the perception time, a binary search algorithm is used to efficiently determine the optimal perception duration τ n , to minimize the failure probability P f,t while maximizing transmission opportunities;

[0038] The binary search gradually narrows down the possible sensing time range by iteratively evaluating the trade-off between sensing accuracy and transmission time. After determining the optimal sensing time, the time is passed as input to the neural network, which is used to dynamically adjust the flight path of the UAV according to the real-time spectrum feedback. This method combines the binary search algorithm with the trajectory design based on the dual deep Q network (DDQN) to form a feedback loop. By dynamically adjusting the sensing time according to the real-time channel conditions and optimizing the UAV flight path, the system maximizes the transmission opportunities and ensures the spectrum sensing efficiency. This method improves the reliability and throughput of the UAV-assisted communication system and can effectively cope with the challenges in dynamic and uncertain environments.

[0039] S502. DDQN-based UAV perception and communication optimization algorithm: This algorithm uses a dual deep Q network (DDQN) to optimize the 3D flight trajectory of the UAV and adjust the phase of the reconfigurable smart surface (RIS) to improve the performance of the UAV-assisted communication system. Its main goal is to maximize the communication throughput and minimize the false detection probability while considering the spectrum sensing efficiency;

[0040] S503 DDQN principle: DDQN solves the problem of overestimation of Q value in DQN by separating action selection and target value generation. DDQN uses two networks: the behavior network is used for action selection, and the target network is used for Q value evaluation. The Q value update follows the Bellman equation:

[0041] Q(S,A)=Q(S,A)+α[R+βQ t (S',argmaxQ b Q(S',A'))-Q(S,A)] (9)

[0042] Among them, α is the learning rate and β is the discount factor;

[0043] S504 State abstraction: The input state includes the position of the UAV, the user's channel gain, RIS phase, perception duration, and false detection probability. These variables are abstracted and normalized as the input of the neural network. The state is represented as follows:

[0044]

[0045] S505 Action Space: To reduce complexity, the action space is discretized. The UAV's motion is predefined as discrete actions, such as horizontal left, right, forward, backward, vertical up, down, and hovering. In addition, multiple fixed transmission power levels are defined for users to choose from to ensure stable communication.

[0046] S506 Neural Network and Experience Replay: Neural networks are used to estimate Q values ​​and select the best actions based on these Q values; in order to break temporal correlation and stabilize learning, experience replay technology is used to store past interaction experiences in a replay buffer and sample from it during training to improve generalization ability;

[0047] S507 RIS phase optimization: This algorithm combines the gradient-based RIS phase optimization method. First, the gradient of the reward function with respect to each RIS phase is estimated using the finite difference method, and then the RIS phase is updated using the gradient descent method:

[0048]

[0049] S508 Action Strategy: The agent adopts a decreasing ε-greedy strategy to balance exploration and exploitation. As time goes by, the probability of the agent exploring random actions gradually decreases, while the probability of choosing learned actions increases, thereby improving decision-making results;

[0050] S509 Reward Function: The goal of the reward function is to maximize the effective total throughput while ensuring user fairness. The reward function is defined as:

[0051]

[0052] Where R(t) is the total data rate and λ is the fairness penalty factor;

[0053] By combining the DDQN framework with RIS phase optimization, the proposed algorithm can dynamically adjust the RIS phase and UAV trajectory configuration in real time, thereby improving the throughput and reliability of the communication system and optimizing the overall performance of the UAV-assisted communication system in a dynamic environment;

[0054] Compared with the prior art, the beneficial effects of the present invention are:

[0055] The present invention uses the deep double Q network (DDQN) algorithm to jointly optimize the spectrum sensing duration, unmanned aerial vehicle (UAV) trajectory and reflective intelligent surface (RIS) phase, significantly improving the effective throughput and communication reliability of the system: the energy detection method ensures the reliability of communication, and the gradient descent method is combined with deep reinforcement learning to optimize the RIS phase and UAV trajectory, successfully reducing the estimation bias; simulation results show that the proposed algorithm can effectively improve the system throughput and overall performance, verifying the potential of deep reinforcement learning to achieve efficient and reliable communication in intelligent UAV networks. The advantages of this technology are that it improves spectrum efficiency, reduces interference, and significantly improves the overall performance and reliability of the communication system by jointly optimizing multiple key parameters. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings used in the embodiments. Obviously, the drawings only represent the preferred embodiments of the present invention, and ordinary technicians in this field can obtain other drawings based on these drawings without additional creative work;

[0057] Figure 1 This is a model diagram of RIS-assisted cognitive drone perception communication provided by an embodiment of the present invention;

[0058] Figure 2 It is a throughput comparison of an IRS-assisted cognitive UAV system under different learning rates and flight strategies provided by an embodiment of the present invention;

[0059] Figure 3 It is a drone trajectory and user movement path diagram of an IRS-assisted cognitive drone system under DDQN design provided by an embodiment of the present invention; DETAILED DESCRIPTION

[0060] The principles and features of the present invention are described below in conjunction with the accompanying drawings. The examples listed are only used to explain the present invention and are not used to limit the scope of the present invention.

[0061] This embodiment provides a UAV airborne IRS-assisted wireless recognition network physical layer security transmission method, the specific steps are as follows:

[0062] S1. Figure 1 As shown in FIG. 1 , the system involves an outdoor downlink communication scenario including a base station (PT), a reflective intelligent surface (RIS), an unmanned aerial vehicle (UAV), and multiple mobile users. The transmission time is maximized by optimizing the spectrum sensing time, thereby improving the management and allocation efficiency of spectrum resources; in addition, by combining the phase configuration of RIS and the trajectory design of UAV, the effective throughput of the system is significantly improved to meet the user's demand for high-speed data transmission; the IRS is uniformly distributed in an array and the number of reflective units is M = Mx ×M y , the phase shift matrix of IRS in the nth time slot can be expressed as: The user moves according to the random walk model, and the user's movement direction and speed are completely random at any time t;

[0063] S2. The system adopts a single antenna configuration. The UAV performs spectrum sensing and data transmission by dividing the flight period T into N time slots. In each time slot, when the UAV detects that the spectrum is idle, the UAV senses the spectrum for a duration of τ. n Complete spectrum access preparation within t n -τ n Transmit data; false alarm probability P f,t Calculated by Gaussian Q function:

[0064]

[0065] S3. The air-ground channel model between the drone and the user is provided according to the 3GPP Release 15 specification. The path loss depends on the line-of-sight (LoS) and non-line-of-sight (NLoS) states, expressed as P LOS / NOLS (t), line-of-sight probability P LOS , P NLOS Depending on the distance between the user and the UAV and the height of the UAV h u (t) Calculate the channel gain g between UAV and user uk (t) Calculated by direct link; channel gain g of the reflection link urk (t) is calculated from the path loss between RIS and UAV and user; the total channel gain of the reflection link is:

[0066] g urk (t)=|g rk (t) T Θg ur (t)| 2 (2)

[0067] The total channel gain between the drone and the user is:

[0068]

[0069] The signal-to-noise ratio and data rate are calculated as:

[0070]

[0071]

[0072] The optimized system throughput R(t) is the weighted average data rate in each time slot, expressed as

[0073]

[0074] The overall throughput is:

[0075]

[0076] S4. Under the constraints of UAV flight space, perception duration, maximum power, RIS phase shift and service quality, optimize the UAV's continuous perception time, flight trajectory, and RIS phase shift. The goal is to maximize the effective total throughput of the UAV during the service period; define the UAV's position during the service period as H, and the perception duration as Γ=τ n ; The flight speed is set to be constant, and the transmission power is P u , the power allocation strategy is P = {P k (t), 0≤t≤T, k∈K}; the optimization problem is formalized as:

[0077]

[0078]

[0079]

[0080]

[0081] 0≤τ n ≤T n (8.2)

[0082]

[0083]

[0084]

[0085] Position constraint (8.1): ensure that the UAV is within the effective height range above the service area to avoid collision with ground obstacles and ensure stable signal coverage; Perception duration constraint (8.2): limit the spectrum perception duration, optimize the switching between transmission and perception, and reduce user delay; Power constraint (8.3): limit the maximum transmission power of each UAV to prevent power consumption from exceeding the maximum power limit of the device; Phase setting constraint (8.4): ensure that the phase of each reflection unit is within the control range to optimize the signal propagation path; Quality of service constraint (8.5): ensure that all users obtain a certain quality of service in the same service area to avoid uneven resource allocation;

[0086] S5. The deep reinforcement learning framework is to divide the system transmission rate maximization problem into two parts: UAV spectrum sensing duration optimization and UAV trajectory and RIS phase optimization; the specific steps include:

[0087] S501. Optimization of sensing duration in UAV-assisted communication: In UAV-assisted communication systems, optimizing spectrum sensing time is crucial to balancing spectrum sensing accuracy and transmission time; the sensing time problem at time t can be expressed as (1-P f,t )(t n -τ n ), where P f,t represents the probability of spectrum sensing failure, t n is the total available time, τ n is the perception time. To optimize the perception time, a binary search algorithm is used to efficiently determine the optimal perception duration τ n , to minimize the failure probability P f,t At the same time, the transmission opportunities are maximized; the binary search gradually narrows the possible sensing time range by iteratively evaluating the trade-off between sensing accuracy and transmission time; after determining the optimal sensing time, the time is passed as input to the neural network, which is used to dynamically adjust the flight path of the UAV according to real-time spectrum feedback; this method combines the binary search algorithm with the trajectory design based on the dual deep Q network (DDQN) to form a feedback loop; by dynamically adjusting the sensing time according to the real-time channel conditions and optimizing the UAV flight path, the system maximizes the transmission opportunities and ensures the spectrum sensing efficiency; this method improves the reliability and throughput of the UAV-assisted communication system and can effectively cope with the challenges in dynamic and uncertain environments;

[0088] S502. DDQN-based UAV perception and communication optimization algorithm: This algorithm uses a dual deep Q network (DDQN) to optimize the 3D flight trajectory of the UAV and adjust the phase of the reconfigurable intelligent surface (RIS) to improve the performance of the UAV-assisted communication system. Its main goal is to maximize the communication throughput and minimize the false detection probability while considering the spectrum sensing efficiency;

[0089] S503.DDQN principle: DDQN solves the problem of overestimation of Q value in DQN by separating action selection and target value generation; DDQN uses two networks: the behavior network is used for action selection, and the target network is used for Q value evaluation; Q value update follows the Bellman equation:

[0090] Q(S,A)=Q(S,A)+α[R+βQ t (S',argmaxQ b Q(S',A'))-Q(S,A)] (9)

[0091] Among them, α is the learning rate and β is the discount factor;

[0092] S504. State abstraction: The input state includes the position of the UAV, the channel gain of the user, the RIS phase, the sensing duration, and the false detection probability; these variables are abstracted and normalized as the input of the neural network; the state is represented as follows:

[0093]

[0094] S505. Action space: To reduce complexity, the action space is discretized; the motion of the UAV is predefined as discrete actions, such as horizontal left, right, forward, backward, vertical up, down, and hovering; in addition, multiple fixed transmission power levels are defined for users to choose from to ensure stable communication;

[0095] S506. Neural Network and Experience Replay: Neural networks are used to estimate Q values ​​and select the best actions based on these Q values; in order to break the temporal correlation and stabilize learning, experience replay technology is used to store past interaction experiences in a replay buffer and sample from it during training to improve generalization ability;

[0096] S507.RIS Phase Optimization: This algorithm combines a gradient-based RIS phase optimization method; first, the gradient of the reward function with respect to each RIS phase is estimated using the finite difference method, and then the RIS phase is updated using the gradient descent method:

[0097]

[0098] S508. Action strategy: The agent adopts a decreasing ε-greedy strategy to balance exploration and exploitation. As time goes by, the probability of the agent exploring random actions gradually decreases, while the probability of selecting learned actions increases, thereby improving decision-making results.

[0099] S509. Reward function: The goal of the reward function is to maximize the effective total throughput while ensuring user fairness; the reward function is defined as:

[0100]

[0101] Where R(t) is the total data rate and λ is the fairness penalty factor;

[0102] By combining the DDQN framework with RIS phase optimization, the proposed algorithm can dynamically adjust the RIS phase and UAV trajectory configuration in real time, thereby improving the throughput and reliability of the communication system and optimizing the overall performance of the UAV-assisted communication system in a dynamic environment;

[0103] Figure 2The throughput changes of the dual deep Q network (DDQN) algorithm in multiple training cycles under different learning rates are shown. As can be seen from the figure, with the increase of training rounds, the system throughput gradually increases, especially after the introduction of reconfigurable intelligent surface (RIS) technology, the performance has been significantly improved. When the learning rate is 0.001, the throughput curve of the RIS integrated solution shows the best convergence and maintains a stable growth trend throughout the training process. In contrast, without the support of RIS technology, the system throughput is significantly lower, especially at the same learning rate, the gap is more obvious. These results fully demonstrate the importance of RIS technology in improving system performance and efficiency, especially in cognitive drone networks, through reasonable parameter settings, the system throughput and overall performance have been significantly improved.

[0104] Figure 3 A trajectory planning scenario in three-dimensional space is presented, highlighting the interaction between an unmanned aerial vehicle (UAV) and a user, who moves in a random motion pattern. In this scenario, the position of the reconfigurable smart surface (RIS) remains fixed, while the UAV approaches the user in a step-by-step zigzag trajectory starting from a starting point. The motion trajectory is carefully designed, taking into account the transmission rate at each step, allowing the UAV to intelligently adjust its trajectory. Ultimately, the UAV will stay in the best position to ensure effective communication with the user. This visualization demonstrates the algorithm's ability to adapt to the user's dynamic movement while optimizing the UAV's path, highlighting its potential to improve communication efficiency in practical applications.

[0105] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A RIS-assisted cognitive UAV system resource joint optimization method based on reinforcement learning, characterized in that: The following steps are involved: S1. Establish a cognitive wireless communication system model for RIS-assisted cognitive UAV network transmission; S2. Establish a perception model. The UAV uses the energy detection method to perceive the spectrum occupancy, assuming that the PU spectrum state remains unchanged in each time slot. Based on the spectrum sensing results, the UAV transmits data to the ground user through the time division multiple access (TDMA) protocol. S3. Establish a transmission model, including channel model; channel gain and reflection link; total channel gain; data transmission rate and total system throughput; S4. Establish a system optimization problem with the goal of maximizing the user data transmission rate, and maximize the system throughput by combining spectrum sensing duration, RIS phase shift design, and UAV trajectory optimization; S5. Solve the system transmission rate optimization problem by building a deep reinforcement learning framework.

2. The RIS-assisted cognitive UAV system resource joint optimization method based on reinforcement learning according to claim 1 is characterized in that: In step S1, the system involves an outdoor downlink communication scenario including a base station (PT), a reflective intelligent surface (RIS), an unmanned aerial vehicle (UAV) and multiple mobile users; the transmission time is maximized by optimizing the spectrum sensing time, thereby improving the management and allocation efficiency of spectrum resources. In addition, by combining the phase configuration of RIS and the trajectory design of UAV, the effective throughput of the system is significantly improved to meet the user's demand for high-speed data transmission. The IRS is uniformly distributed in an array and the number of reflective units is M = M x ×M y , the phase shift matrix of IRS in the nth time slot can be expressed as: The user moves according to the random walk model, and the user's movement direction and speed are completely random at any time t.

3. The RIS-assisted cognitive UAV system resource joint optimization method based on reinforcement learning according to claim 1 is characterized in that: In step S2, the system adopts a single antenna configuration, and the UAV performs spectrum sensing and data transmission by dividing the flight period T into N time slots. In each time slot, when the UAV detects that the spectrum is idle, the UAV n Complete spectrum access preparation within t n -τ n Data transmission. False alarm probability P f,t Calculated by Gaussian Q function:

4. The RIS-assisted cognitive UAV system resource joint optimization method based on reinforcement learning according to claim 1 is characterized in that: In step S3, the air-ground channel model between the UAV and the user is provided according to the 3GPP Release 15 specification. The path loss depends on the line-of-sight (LoS) and non-line-of-sight (NLoS) states, expressed as P LOS / NOLS (t), line-of-sight probability P LOS , P NLOS Based on the distance between the user and the UAV and the height of the UAV h u (t) Calculate the channel gain g between UAV and user uk (t) Calculated by direct link; channel gain g of the reflection link urk (t) is calculated from the path loss between RIS and UAV and user; the total channel gain of the reflection link is: g urk (t)=|g rk (t) T Θg ur (t)| 2 (2) The total channel gain between the drone and the user is: The signal-to-noise ratio and data rate are calculated as: The optimized system throughput R(t) is the weighted average data rate in each time slot, expressed as The overall throughput is:

5. The RIS-assisted cognitive UAV system resource joint optimization method based on reinforcement learning according to claim 1 is characterized in that: Under the constraints of UAV flight space, perception duration, maximum power, RIS phase shift and service quality, the continuous perception time, flight trajectory and RIS phase shift of UAV are optimized. The goal is to maximize the effective total throughput of UAV during the service period. The position of UAV during the service period is defined as H, and the perception duration is defined as Γ = τ n ; The flight speed is set to be constant; The transmission power is P u ; The power allocation strategy is P = {P k (t), 0≤t≤T, k∈K}; the optimization problem is formalized as: Position constraint (8.1): ensure that the UAV is in the effective height range above the service area to avoid collision with ground obstacles and ensure stable signal coverage; perception duration constraint (8.2): limit the spectrum perception duration, optimize the switching between transmission and perception, and reduce user delay; power constraint (8.3): limit the maximum transmission power of each UAV to prevent power consumption from exceeding the maximum power limit of the device; phase setting constraint (8.4): ensure that the phase of each reflection unit is within the control range to optimize the signal propagation path; quality of service constraint (8.5): ensure that all users obtain a certain quality of service in the same service area to avoid uneven resource allocation.

6. The RIS-assisted cognitive UAV system resource joint optimization method based on reinforcement learning according to claim 1 is characterized in that: In step S5, the deep reinforcement learning framework divides the system transmission rate maximization problem into two parts: UAV spectrum sensing duration optimization and UAV trajectory and RIS phase optimization. The specific steps include: S501. Optimization of sensing duration in UAV-assisted communication: In UAV-assisted communication systems, optimizing spectrum sensing time is crucial to balancing spectrum sensing accuracy and transmission time; the sensing time problem at time t can be expressed as (1-P f,t )(t n -τ n ), where P f,t represents the probability of spectrum sensing failure, t n is the total available time, τ n It is the perception of time; In order to optimize the perception time, a binary search algorithm is used to efficiently determine the optimal perception duration τ n , to minimize the failure probability P f,t The binary search gradually narrows down the possible sensing time range by iteratively evaluating the trade-off between sensing accuracy and transmission time. After determining the optimal sensing time, the time is passed as input to the neural network, which is used to dynamically adjust the flight path of the UAV according to real-time spectrum feedback. This method combines the binary search algorithm with the trajectory design based on the dual deep Q network (DDQN) to form a feedback loop. By dynamically adjusting the sensing time according to the real-time channel conditions and optimizing the UAV flight path, the system maximizes the transmission opportunities and ensures the spectrum sensing efficiency. This method improves the reliability and throughput of UAV-assisted communication systems and can effectively cope with challenges in dynamic and uncertain environments; S502. DDQN-based UAV perception and communication optimization algorithm: This algorithm uses a dual deep Q network (DDQN) to optimize the 3D flight trajectory of the UAV and adjust the phase of the reconfigurable smart surface (RIS) to improve the performance of the UAV-assisted communication system; its main goal is to maximize the communication throughput and minimize the false detection probability while considering the spectrum sensing efficiency; S503.DDQN principle: DDQN solves the problem of overestimation of Q value in DQN by separating action selection and target value generation. DDQN uses two networks: a behavior network for action selection and a target network for Q-value evaluation; the Q-value update follows the Bellman equation: Q(S,A)=Q(S,A)+α[R+βQ t (S',argmaxQ b Q(S',A'))-Q(S,A)] (9) in, α is the learning rate, β is the discount factor; S504. State abstraction: The input state includes the position of the UAV, the channel gain of the user, the RIS phase, the sensing duration and the false detection probability; These variables are abstracted and normalized as input to the neural network; the states are represented as follows: S505. Action space: To reduce complexity, the action space is discretized. The UAV's motion is predefined as discrete actions, such as horizontal left, right, forward, backward, vertical up, down, and hovering; in addition, multiple fixed transmission power levels are defined for users to choose from to ensure stable communication; S506. Neural Network and Experience Replay: Neural networks are used to estimate Q values ​​and select the best actions based on these Q values; in order to break the temporal correlation and stabilize learning, experience replay technology is used to store past interaction experiences in a replay buffer and sample from it during training to improve generalization ability; S507.RIS Phase Optimization: This algorithm combines a gradient-based RIS phase optimization method; first, the gradient of the reward function with respect to each RIS phase is estimated using the finite difference method, and then the RIS phase is updated using the gradient descent method: S508. Action strategy: The agent adopts a decreasing ε-greedy strategy to balance exploration and exploitation. As time goes by, the probability of the agent exploring random actions gradually decreases, while the probability of choosing learned actions increases, thereby improving decision-making results; S509. Reward function: The goal of the reward function is to maximize the effective total throughput while ensuring user fairness. The reward function is defined as: Where R(t) is the total data rate and λ is the fairness penalty factor; By combining the DDQN framework with RIS phase optimization, the proposed algorithm is able to dynamically adjust the RIS phase and UAV trajectory configuration in real time, thereby improving the throughput and reliability of the communication system and optimizing the overall performance of the UAV-assisted communication system in a dynamic environment.

Citation Information

Cited By

  • RIS-assisted UAV air-ground network joint optimization method based on reinforcement learning

    CN121284604A