Link and resource optimization method based on reinforcement learning in hybrid RF-FSO network

By constructing a three-layer transmission architecture and using the PPONT agent algorithm based on deep reinforcement learning, the dynamic optimization problem of UAV trajectory and link selection in complex atmospheric environments was solved, achieving high reliability and low latency communication of the system.

CN121815278APending Publication Date: 2026-04-07CHONGQING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-21
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies cannot optimize UAV trajectories, link selection, and communication resources in real time and dynamically in complex atmospheric environments, resulting in high system transmission latency and insufficient reliability.

Method used

A three-layer transmission architecture consisting of an aerial platform, a UAV relay layer, and a ground base station is constructed. The PPONT agent algorithm of deep reinforcement learning is used for joint optimization. By jointly optimizing link selection, UAV trajectory, power allocation, and resource allocation, a Markov decision process is established to minimize the total transmission delay of the system.

Benefits of technology

It significantly reduces the total system transmission latency, improves communication reliability and resource utilization, adapts to complex atmospheric environmental changes, and achieves efficient link switching and resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121815278A_ABST
    Figure CN121815278A_ABST
Patent Text Reader

Abstract

The invention discloses a link selection and resource optimization method based on reinforcement learning in a hybrid RF-FSO network, and the method comprises the steps: constructing a three-layer cooperative transmission architecture composed of a high-altitude platform, a plurality of unmanned aerial vehicles and a ground base station, and designing a weather perception dynamic three-level link model comprising a direct FSO link, a relay FSO link and a standby RF link; then establishing an optimization problem with the purpose of minimizing the total transmission delay of all base stations, and modeling the optimization problem into a Markov decision process by adopting a PPONT agent algorithm; and finally, outputting a joint optimization decision action including UAV track, link selection, sub-channel allocation and power allocation, thereby realizing end-to-end system performance optimization. According to the method, adverse weather influences such as cloud layer shielding and rain and fog can be intelligently avoided, the total transmission time delay of the system in different atmospheric environments is remarkably reduced through an efficient joint optimization strategy, and the communication reliability and the spectrum resource utilization rate are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of integrated air-space-ground network communication technology, specifically relating to a dynamic link selection and resource allocation method that combines free-space optical communication and radio frequency. Background Technology

[0002] The statements in this section are merely background information relating to this disclosure, and these statements may constitute prior art. In the process of developing this invention, the inventors discovered at least the following problems in the prior art.

[0003] With the development of smart city and other businesses, the demand for high-bandwidth data transmission, such as video, surges in high-density areas (e.g., urban residential communities) during peak hours, posing significant coverage and storage capacity bottlenecks for traditional terrestrial base stations (BS). High-Altitude Platforms (HAPs) have emerged as a solution due to their wide-area coverage and high-capacity advantages; however, traditional radio frequency spectrum resources are becoming increasingly scarce and unable to meet the rate requirements.

[0004] Free-space optical communication (FSO) technology offers advantages such as high bandwidth and interference resistance, making it an effective complement to RF. However, the performance of FSO links is highly susceptible to the effects of complex atmospheric environments, particularly cloud cover, rain, fog, and atmospheric turbulence. These factors can lead to severe signal attenuation or even link interruption, resulting in significantly increased communication latency or service unavailability.

[0005] Existing solutions typically employ unmanned aerial vehicles (UAVs) as relays or static link switching strategies. For example, patent application number 202510705372.2, entitled "A Multi-Dimensional Quality of Service-Driven Robust Resource Allocation Method for FSO / RF Air-Ground Integrated Systems," first establishes a SAGIN cloud-edge collaborative computing network model for FSO / RF communication, and then constructs an optimization problem OP1 based on this model, aiming to maximize the total processing load of long-duration computing tasks for all users. Finally, the DDPG deep learning method is used to solve optimization problem OP1, obtaining the optimal solution, which is the robust resource allocation scheme for FSO / RF air-ground integrated systems. However, such solutions often struggle to cope with complex and changing weather conditions and workloads in real time and dynamically. They typically cannot achieve joint optimization of UAV trajectories, link selection, and communication resources (such as power and channels), resulting in high total system latency and insufficient reliability even in complex environments.

[0006] The present invention aims to solve the problems of high system transmission latency and insufficient reliability caused by the inability to perform real-time joint optimization of UAV trajectory, link selection and communication resources in complex atmospheric environments. Summary of the Invention

[0007] In view of the above problems, the purpose of this invention is to solve some of the problems in the prior art, or at least alleviate these problems.

[0008] A reinforcement learning-based link and resource optimization method in a hybrid RF-FSO network includes the following steps:

[0009] A three-layer transmission architecture consisting of an aerial platform, a UAV relay layer, and a ground base station is constructed, and the status information of the system at the current decision moment is obtained; wherein, the three-layer transmission architecture is designed according to complex weather conditions for three optional communication links for data backhaul from the aerial platform to the ground base station, including a direct FSO link, a relay FSO link, and a backup RF link;

[0010] Establish a joint system optimization model: aiming to minimize the total transmission delay of all base stations, through joint optimization of link selection. Given the UAV-base station association ζ, UAV trajectory Q, matrix representations ξ, θ, and power allocation P, establish an optimization problem P1;

[0011] The optimization problem P1 is modeled as a Markov decision process;

[0012] The PPONT agent algorithm is designed to solve optimization problems. The agent model interacts extensively with the simulation environment. By continuously updating the network parameters, an agent model that can implement the joint optimization strategy is finally trained.

[0013] The state information is input into a pre-trained deep reinforcement learning agent model; the deep reinforcement learning agent model outputs a joint optimization action based on the input state information.

[0014] The system performs the joint optimization action to minimize the total transmission latency of all ground base stations and complete the data transmission.

[0015] The optimization problem is:

[0016]

[0017] C6:ζ m,n ∈{0,1}

[0018]

[0019] Where D is the total system delay, and t represents the time slot; These represent HAP to ground base station I respectively. n Through direct FSO link, UAV U m Communication is achieved via relay FSO links and backup RF links; These represent the transmit power of the HAP through the three links, P max (t) represents the maximum transmit power of the HAP; x min x m (t), x max U represents the unmanned aerial vehicle (UAV) m The minimum displacement position, current position, and maximum displacement position on the horizontal axis; y min y m (t), y max U represents the unmanned aerial vehicle (UAV) m The minimum displacement position, the current position, and the maximum displacement position on the vertical axis; ζ m,n The correlation coefficient between BS and UAV; The HAP connects to the ground base station I via three links. n Transmission delay, For ground base station I n Maximum tolerable delay; ξ n,k (t), θ n,s (t) are the allocation coefficients for the k-th RF channel and the s-th FSO channel, respectively, where K and S represent the number of sub-channels for the RF and FSO channels, respectively; v m (t) represents the velocity of the UAV, v max q represents the maximum speed at which a UAV can fly; m (t) represents the current position of the UAV.

[0020] Furthermore, the three-layer transmission architecture includes a high-altitude platform (HAP) that is stationary at high altitude, and its coordinates are represented as q. H =[x H ,y H ,z H [ ], serving as the core hub for data storage and distribution; M UAV relay layers, represented as As an FSO dynamic relay node, it flies between the ground base station and the HAP, and its position in time slot t is represented as q. m (t)=[x m (t),y m [(t),z]; N ground base stations BS, represented as The position is fixed at q n =[x n ,y n ,z n ] Request data return from HAP;

[0021] The direct FSO link is a direct free-space optical communication link between the high-altitude platform and the ground base station; the relay FSO link is a free-space optical communication link relayed by a UAV; the backup RF link is a radio frequency link between the high-altitude platform and the ground base station.

[0022] The status information includes at least: the current coordinates of all UAVs, the data load of all ground base stations, weather conditions characterizing the current atmospheric environment, and channel gain estimates for each potential link between the high-altitude platform and each ground base station.

[0023] Furthermore, OFDMA multiple access technology is used for hybrid resource allocation for the three optional communication links; considering various factors, channel models and delay calculation methods are established for each of the three links, including:

[0024] When weather conditions are good, the HAP transmits to the BS via a direct FSO link; when cloud cover with high liquid water content blocks the signal, a UAV is deployed to a location with negligible cloud cover to use a relay FSO link to reflect and forward the optical signal from the HAP to the BS; when weather conditions are very bad, a backup RF link is used.

[0025] The optimization problem P1 is modeled as a Markov decision process, and its main elements are defined as follows:

[0026] state space state It contains the environmental information needed to make a decision at time t. Specifically, it includes: the normalized two-dimensional coordinates of all UAVs, the normalized distance between UAV pairs (if M>1), and the normalized data payload e of all ground base stations. n The index of current weather conditions and its unique hot coding, the normalized index reflecting the potential impact of cloud thickness and rainfall rate on link performance, and the normalized channel gain estimation of the three links between HAP and each BS.

[0027] Action space action It is the set of control variables output by the agent at each decision step, including UAV movement, link selection and sub-channel allocation, and power allocation;

[0028] Reward function (R): Define the reward expression R t =-D total (t)-P penalty (t); where D total (t) represents the total transmission delay of the system; P penalty (t) represents the compound penalty term.

[0029] The PPONT agent algorithm is based on the PPO algorithm and introduces Noisy Networks and Target Networks techniques. Specifically, it includes a parameter... Actor Network An online Critic network V with parameter φ φ (s t A target Critic network V with parameter φ′ φ′ (s t The training data is generated through the interaction between the agent and the environment and stored in the experience replay buffer. The update of the Actor network depends on the pruning substitution objective function of PPO. The Critic network is updated by minimizing the value function estimation error. The parameters of the Actor and Critic networks are updated by the Adam optimizer.

[0030] Furthermore, the noise network adds parameterized noise to the network's weights and biases; a linear layer with noise parameters can be represented as:

[0031] y = (μ w +σ w ⊙∈ w )x+(μ b +σ b ⊙∈ b )

[0032] Among them, (μ w ,μ b ) and (σ w ,σ b ) are learnable mean and standard deviation parameters, while (∈ w ,∈ b The noise variable is sampled from a fixed distribution; the variance parameter of the noise (σ) w ,σ b The network weights μ are learned via gradient descent.

[0033] Two independent sets of network parameters are maintained: an online Critic network and a target Critic network. The online Critic network interacts with the environment and updates policies and values, while the target Critic network provides a relatively stable objective for calculating the TD error and advantage function. The parameters φ′ of the target Critic network are slowly and softly updated using the parameters φ of the online Critic network.

[0034] φ′←τφ+(1-τ)φ′

[0035] Here, τ is a small hyperparameter.

[0036] Furthermore, the PPONT agent algorithm is designed to solve the optimization problem to obtain a pre-trained deep reinforcement learning agent model, including the following steps:

[0037] The environment collects the current global state s at time t.t And send it to the Actor network; the state s t Includes the current location of all UAVs, the data load of all BSs, current weather conditions, and estimated link quality; the environment is a HAP-UAV-BS system;

[0038] The Actor uses a noisy network instead of traditional random sampling to output a complex joint action 'a'. t And send it back to the environment for execution; the action a t It is a decision package containing multiple parts, which determines the next movement vector of each UAV, allocates the best link to each BS, determines how much transmit power the HAP allocates to each link, and dynamically allocates sub-channel bandwidth to all FSO links and RF links;

[0039] Execute action a in the environment t It calculates the total transmission delay of the time slot and generates a reward r. t At the same time, a new state s is generated. t+1 ;

[0040] The complete empirical tuple s t a t r t s t+1 Store it in the Trajectory Buffer;

[0041] Once the trajectory buffer has collected enough data, a mini-batch of samples is randomly selected from the Trajectory Buffer and sent to the Actor, Critic, and Generalized Advantage Estimation (GAE) modules for computation.

[0042] The GAE module gathers the following three data sources for calculation and outputs the calculated values ​​of the advantage function and reward function; the three data sources include the current value estimated by the online Critic network, the target value estimated by the target Critic network, and the reward sampled from the Trajectory Buffer;

[0043] Network Update: Actors receive the advantage and update their parameters by optimizing the loss function; online critics receive the reward and update their parameters by minimizing the value loss.

[0044] After the online Critic completes the parameter update, the weights of the online Critic are slowly copied to the target Critic through soft updates.

[0045] The loop system clears the Trajectory Buffer, returns to step 1, and continues to interact with the environment using the updated Actor. The loop repeats continuously, and the agent eventually learns how to minimize latency under various weather and load conditions.

[0046] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the reinforcement learning-based link and resource optimization method in the hybrid RF-FSO network.

[0047] The present invention has the following beneficial effects:

[0048] 1. High Reliability and Environmental Adaptability: This invention constructs a dynamic three-level link architecture (FSO direct connection, FSO relay, and RF backup) that coordinates HAP and UAV. This architecture can intelligently sense atmospheric conditions (such as cloud cover and rainstorm interference) and automatically switch to the current optimal transmission link, effectively avoiding communication interruptions caused by weather factors and significantly improving communication reliability in non-line-of-sight and complex atmospheric environments.

[0049] 2. Highly efficient joint optimization capability: Compared to existing technologies that handle UAV trajectory, link selection, and resource allocation separately, this invention utilizes deep reinforcement learning to achieve joint optimization of four highly coupled variables: UAV trajectory, link selection, sub-channel allocation, and power control. This enables the system to break through local optima and find a better global strategy, thereby significantly reducing the total system transmission latency.

[0050] 3. High resource utilization: Through dynamic OFDMA sub-channel management and fine-grained power allocation, this invention can accurately allocate FSO and RF spectrum resources to the most needed links based on real-time service load and channel quality, avoiding resource waste and improving overall spectrum efficiency.

[0051] 4. Advanced Algorithm Design and Robustness: The PPO-Noisy-Target algorithm combines the efficient exploration capabilities of noisy networks with the training stability of target networks. This enables it to converge efficiently and robustly, responding quickly to dynamic environmental changes. It is suitable for solving complex optimization problems with high-dimensional, mixed action spaces (continuous + discrete), and has strong application value. Attached Figure Description

[0052] Figure 1 These are hybrid FSO-RF scene diagrams under different atmospheric conditions according to the present invention;

[0053] Figure 2 This is an overall flowchart of the present invention;

[0054] Figure 3 This is a framework diagram of the algorithm for solving the PPONT problem in this invention;

[0055] Figure 4 These are simulation diagrams showing the convergence of different algorithms in this invention;

[0056] Figure 5 This is a graph showing the trend of total system delay under different schemes of the present invention in clear weather as the FSO channel changes;

[0057] Figure 6 This is a graph showing the trend of total system delay under different schemes in cloudy weather as the FSO channel changes according to the present invention;

[0058] Figure 7 This is a graph showing the trend of total system delay under different schemes of the present invention in rainy weather as the FSO channel changes. Detailed Implementation

[0059] The present invention will be further described below with reference to the accompanying drawings. The embodiments of the present invention are only used to illustrate the present invention and not to limit the present invention. Various substitutions and modifications made based on ordinary technical knowledge and common practices in the art without departing from the technical concept of the present invention should be included within the scope of the present invention.

[0060] To address the problems of existing technologies, this invention proposes a hybrid FSO-RF transmission optimization method based on deep reinforcement learning, combining HAP and multiple UAVs, for complex atmospheric environments. This invention constructs a hybrid transmission architecture, models the channels of three potential links (direct FSO, relay FSO, and backup RF), and utilizes a near-end policy optimization (PPO)-based agent to jointly optimize UAV trajectories, link selection, sub-channel allocation, and power allocation, thereby minimizing the total system transmission delay.

[0061] like Figure 2 As shown, the present invention is implemented using the following technical solution:

[0062] S1, construct a three-layer transport architecture and obtain status information.

[0063] A three-layer transmission architecture consisting of a high-altitude platform, a UAV relay layer, and a ground base station is constructed, and the system's current state information at the moment of decision-making is obtained. The three-layer transmission architecture sets up three optional communication links for data backhaul from the high-altitude platform to the ground base station:

[0064] Link 1: Direct free-space optical communication (FSO) link between the high-altitude platform and the ground base station;

[0065] Link 2: FSO link relayed by drones;

[0066] Link 3: Radio frequency (RF) link between the high-altitude platform and the ground base station, serving as a backup link.

[0067] The status information includes at least: the current coordinates of all UAVs, the data load of all ground base stations, weather conditions characterizing the current atmospheric environment, and channel gain estimates for each potential link between the high-altitude platform and each ground base station.

[0068] This invention constructs a HAP-UAV-BS collaborative hybrid three-layer transmission architecture, such as... Figure 1 As shown, it includes one high-altitude platform (HAP) that is stationary at a high altitude (e.g., 20 km), and its coordinates are represented as q. H =[x H ,y H ,z H [This serves as the core hub for data storage and distribution. M unmanned aerial vehicle (UAV) relay layers are represented as...] As an FSO dynamic relay node, it flies between the ground base station and the HAP. Its position in time slot t is represented as q. m (t)=[x m (t),y m (t),z]. N ground base stations (BSs), represented as The position is fixed at q n =[x n ,y n ,z n ] Request data back from HAP.

[0069] This invention designs a dynamic three-level link model for weather sensing. Data transmission from the HAP to the nth BS has three selectable links: Link 1 is a direct FSO link. Link 2 is a relay FSO link That is, through drones U m FSO relay is performed; Link 3 is a backup RF link. To ensure unique link selection, binary variables were defined. And satisfy constraints

[0070] Furthermore, this invention employs OFDMA multiple access technology for hybrid resource allocation, dividing the RF channel into K channels with a bandwidth of B. RF The FSO channel is divided into S sub-channels with a bandwidth of B. FSO Sub-channels.

[0071] This invention further establishes channel models and delay calculation methods for the above three links.

[0072] When weather conditions are favorable, the high-altitude platform can transmit directly to the base station via the FSO link. This link primarily considers cloud attenuation. Atmospheric turbulence and beam propagation loss Using composite channel coefficients Therefore, the coefficient of the channel at time t is... At this point, the signal-to-noise ratio of the signal received by the nth BS can be expressed as: in Indicates a direct FSO link at base station I n The noise power at time slot t. According to Shannon's formula, we can calculate the information rate transmitted directly from the HAP to the nth BS via link FSO as:

[0073]

[0074] Among them, B FSO It is the bandwidth of each sub-channel of the FSO link. This is a coefficient determined based on the average optical power and peak optical power. Therefore, the transmission delay is... Among them, e n Indicates base station I n Data load.

[0075] When clouds with high liquid water content block high-speed transmission directly via the FSO link, it may completely disrupt the connection rate. In such cases, where the BS (Base Station) has high connection rate requirements, we deploy a UAV (Used Aerial Vehicle) at a location with negligible cloud cover. This UAV can reflect and forward optical signals from the HAP (Hydrogen Alternate Access Point) to the BS. At this point, the nth BS establishes a connection with the mth UAV, relaying the signal from the HAP. The path from the HAP to the BS is divided into two hops: from the HAP to the UAV, and then from the UAV to the BS. The channel modeling and delay calculation are the same as for a direct FSO connection.

[0076] When weather conditions are poor, rendering all FSO links unavailable, we use RF links. The channel losses we consider include path loss, cloud attenuation, and fading. Therefore, the channel coefficient of the RF link from HAP to BS at time t is expressed as:

[0077]

[0078] in, Channel fading coefficient, The signal attenuation in an RF link from transmitter to receiver is denoted as... The parameters in the formula are, respectively, the transmit antenna gain, the receive antenna gain, atmospheric attenuation, cloud attenuation, and free space loss. The instantaneous signal-to-noise ratio (SNR) of the BS received via the RF link at time t is:

[0079]

[0080] in Indicates a direct RF link at base station I n The noise power at time slot t.

[0081] The information rate through the link from HAP to the nth BS is calculated as follows: Among them, B RF The bandwidth of each sub-channel of the RF link gives the transmission delay.

[0082] In summary, the total system delay is expressed as:

[0083]

[0084] S2, Establish a joint optimization model for the system.

[0085] Establish a joint optimization model for the system to minimize the total transmission delay while satisfying all constraints.

[0086] Our goal is to jointly optimize link selection. The optimization problem considers UAV-base station association (ζ), UAV trajectory (Q), matrix form (ξ,θ), and power allocation (P), aiming to minimize the total transmission delay across all base stations. The problem is formulated as follows:

[0087]

[0088]

[0089] C6:ζ m,n ∈{0,1}

[0090]

[0091] Where D is the total system delay, and t represents the time slot; These represent HAP to ground BSI respectively. n Through direct FSO link, UAV U m Communication is achieved via relay FSO links and backup RF links; These represent the three links through which HAP connects to BS I. n The transmission power, P max (t) represents the maximum transmit power of the HAP; x min x m (t), x max They represent UAVU respectively m The minimum displacement position, current position, and maximum displacement position on the horizontal axis; y min y m (t), ymax They represent UAV and U respectively. m The minimum displacement position, current position, and maximum displacement position on the ordinate. ζ m,n The correlation coefficient between BS and UAV. HAP connects to BS I via three links respectively. n Transmission delay, ξ represents the maximum tolerable delay for the BS. n,k (t),θ n,s (t) represents the allocation coefficients for the nth RF channel and the sth FSO channel, respectively, and K and S represent the number of sub-channels for the RF and FSO channels, respectively. m (t) represents the velocity of the UAV, v max q represents the maximum speed at which a UAV can fly; m (t) represents the current position of the UAV.

[0092] Among these constraints, C1 and C2 ensure that link selection is associated with the UAV base station as a binary decision. Constraint C3 stipulates that the total transmit power of the Hybrid Access Point (HAP) must not exceed a threshold. Constraints C4 and C5 limit the maximum coverage area of ​​the UAV service, while C6 guarantees that each base station must select one and only one link. Constraint C7 ensures that the transmission delay of each base station does not exceed a maximum threshold. Constraints C8 and C9 respectively guarantee that the number of sub-channels of the allocated radio frequency links and free space optical communication (FSO) links does not exceed their respective available totals. Constraint C10 specifies the maneuverability restrictions for the UAV.

[0093] Given the mixed-integer nonlinear programming (MINLP) nature and NP-hardness of the joint optimization problem (P1), directly finding the optimal solution is generally infeasible. In particular, this problem involves high-dimensional continuous (UAV trajectory, power allocation) and discrete (link selection, subchannel allocation) decision variables, and requires real-time decisions in dynamically changing environments. Deep reinforcement learning (DRL) provides a powerful framework for solving such complex optimization problems. In this study, we specifically employ the Proximal Policy Optimization (PPO) algorithm and its variants because of its excellent performance in handling mixed action spaces and ensuring training stability.

[0094] S3 models the optimization problem as a Markov decision process.

[0095] We model problem P1 as a Markov Decision Process (MDP), whose main elements are defined as follows:

[0096] state space state It contains the environmental information needed to make a decision at time t. Specifically, it includes: normalized two-dimensional coordinates of all UAVs, normalized distance between UAV pairs (if M>1), normalized data payload e(t) of all ground base stations, index of current weather conditions and their one-hot coding, normalized indices reflecting the potential impact of cloud thickness and rainfall rate on link performance, and normalized channel gain estimates for three potential links (direct FSO, relay FSO, direct RF) between the HAP and each BS.

[0097] Action space action It is the set of control variables output by the agent at each decision step. This is a hybrid action space, containing:

[0098] Drone movement: The two-dimensional velocity vector of each drone m (continuous action) is parameterized by a Gaussian distribution output by the Actor network, used to update the position at the next moment, and is influenced by the maximum velocity v. max limit.

[0099] Link selection and sub-channel allocation: Select a transmission link (direct FSO) for each base station n. Relay FSO Or direct RF) and the corresponding FSO and RF subchannel allocation mode index (discrete action), parameterized by the classification probability distribution of the Actor network.

[0100] Power Allocation: The normalized power allocation factor (continuous action) for each link is parameterized to [0,1] by the Gaussian distribution output by the Actor network through hyperbolic tangent function (tanh) and scaling transformation, and is used to determine the actual transmit power. And subject to the total power P max limit.

[0101] Reward function (R): Define the reward expression R t =-D total (t)-P penalty (t); where D total (t) represents the total transmission delay of the system; P penalty (t) is a composite penalty term used to constrain conditions such as total transmit power, maximum transmission delay threshold, and UAV maneuver restrictions. The aim is to guide the agent to learn a strategy that minimizes the total system delay D.

[0102] S4. Design the PPONT (PPO-Noisy-Target) agent algorithm to solve the optimization problem.

[0103] Based on the above design, and building upon the PPO algorithm, we introduce noisy networks to promote more efficient and structured exploration. Simultaneously, drawing on the successful experiences of DQN and DDPG, we employ target networks. By combining a noisy network for efficient exploration and a target network for stable training, our proposed PPO-Noisy-Target (PPONT) algorithm aims to provide a robust and high-performance DRL solution for solving the complex joint optimization problem P1. Its algorithm architecture diagram is shown below. Figure 3 As shown.

[0104] Specifically, this includes: the Actor-Critic framework and PPO update: the agent contains an Actor network. (parameter is) The output action policy is used, along with an online Critic network V with parameter φ. φ (s t A target Critic network V with parameter φ′ φ′ (d t The value evaluation network (Actor Network) is used to evaluate the value of a state. Training data is generated through the interaction between the agent and the environment and stored in the Trajectory Buffer. Updates to the Actor Network rely on the core mechanism of the PPO—the Clipped Surrogate Objective:

[0105]

[0106] in, Empirical expectation over time slot t It is the probability ratio of the new strategy to the old strategy. is the advantage function estimate, ∈ is the clipping coefficient, and clip(·) is the clipping function. Its function is to force this ratio to be limited to a specific range.

[0107] Advantage function It is calculated using the Generalized Advantage Estimation (GAE), which utilizes the value estimate V of the online Critic network (or target Critic network). φ (s) and timing difference (TD) error δ t :

[0108]

[0109] δ t =R t +γVφ (s t+1 )-V φ (s t )

[0110] Where γ is the discount factor, V φ (s t ) is the current state value estimate, λ GAE These are GAE parameters. The Critic network minimizes the value loss function L. V (φ) Estimation error is updated:

[0111]

[0112] Among them, V t target =R t +γV φ′ (s t+1 (When using the target network) or The total loss function typically also includes an entropy reward term to encourage exploration. The Actor and Critic network parameters are updated via the Adam optimizer.

[0113] Traditional exploration methods, such as ∈-greedy or entropy regularization, which rely on adding simple random perturbations during the action selection phase, primarily depend on adding random perturbations to the action space. However, in such complex hybrid action spaces and dynamic environments, these simple "jittering" perturbations may be insufficient to drive agents to conduct effective and sustained exploration. To facilitate more efficient and structured exploration, we introduce NoisyNet. The core idea of ​​NoisyNet is to add parameterized noise to the network's weights and biases, rather than the action space.

[0114] A linear layer with noise parameters can be represented as:

[0115] y = (μ w +σ w ⊙∈ w )x+(μ b +σ b ⊙∈ b )

[0116] Among them, (μ w ,μ b ) and (σ w ,σ b ) are learnable mean and standard deviation parameters, while (∈ w ,∈ b This refers to noise variables sampled from a fixed distribution (e.g., a Gaussian distribution). The variance parameter of the noise (σ) w ,σ bThe noise parameter is learned along with the network weights μ via gradient descent. A key advantage of this parameter-space noise is that a single perturbation of the weights can guide consistent and potentially very complex state-dependent policy changes over multiple time slots, contrasting with methods that add irrelevant noise with each action. By learning the noise parameters, the agent can automatically adjust the degree of exploration to adapt to the needs of different task phases.

[0117] In deep reinforcement learning (DRL) algorithms based on value functions or actor-criticism, the network parameters used to compute the target value (e.g., the TD target) are the same as the network parameters being updated, which may cause the target value to change continuously, leading to instability and oscillations in the training process. To address this issue, we draw on the successful experience of Deep Q-Network (DQN) and Deep Deterministic Policy Gradient (DDPG) and adopt the Target Networks technique.

[0118] This technology maintains two independent sets of network parameters: the online Critic network and the target Critic network. and CriticV φ ) and Target Critic Network (CriticV) φ′ The online Critic network interacts with the environment and updates the policy and value, while the objective Critic network provides a relatively stable objective for calculating the TD error and advantage function. The parameters φ′ of the objective Critic network are not updated directly via gradient descent, but rather through a slow, "soft update" of the online Critic network parameters φ.

[0119] φ′←τφ+(1-τ)φ′

[0120] Here, τ is a small hyperparameter. This slow update mechanism ensures the target value V. t target =R t +γV φ′ (s t+1 The stability of V φ′ (s t+1 The next-state value estimation based on the target network reduces value estimation oscillations caused by bootstrapping, thereby stabilizing the entire learning process.

[0121] Detailed workflow:

[0122] Step 1: State Observation System: The environment (HAP-UAV-BS system) collects the current global state at time t. State s t Includes: current location of all UAVs, data load of all BSs, current weather conditions (sunny, cloudy, heavy rain), and estimated link quality (FSO direct, FSO relay, RF link channel gain). Status s t It is sent to the Actor network for decision-making.

[0123] Step 2: Joint Action Decision System: The Actor uses Noisy Layers instead of traditional random sampling to achieve more efficient autonomous exploration. The Actor outputs a complex joint action 'a'. t This action is a decision package comprising multiple parts: determining the next movement vector for each UAV, allocating the optimal link (FSO direct, FSO relay, or RF direct) for each BS, determining how much transmit power the HAP allocates to each link, and dynamically allocating sub-channel bandwidth for all FSO and RF links. Action a t It is sent back to the environment for execution.

[0124] Step 3: Environmental Execution and Feedback System: Environmental execution actions: HAP allocates power and channel, UAV begins movement, and data transmission begins. The system calculates the total transmission delay of the time slot and generates a reward r. t The environment simultaneously generates a new state s t+1 .

[0125] Step 4: Data storage system: complete empirical tuples t ,a t ,r t ,s t+1 It is stored in the Trajectory Buffer.

[0126] Step 5: Batch Sampling and System Training: Once the Trajectory Buffer has collected enough data, training begins. A "Mini-batch sample" is randomly selected from the Trajectory Buffer, and this small batch of data is simultaneously sent to the Actor, Critic, and GAE modules for computation.

[0127] Step 6: Advantage Calculation System: The GAE (Generalized Advantage Estimation) module is the core computational unit of PPO. It gathers three sources of data for calculation: the current value estimated by the online Critic network (also using Noisy Layers); the target value estimated by the target Critic network; and the reward sampled from the Trajectory Buffer. Output: Advantage: Measures action a t Good or bad relative to the average level. Return: For the current state s t A more accurate estimate of value.

[0128] Step 7: Network Update: The Actor receives the "advantage" and updates its parameters by optimizing the loss function. This loss function increases the probability of actions that produce a high "advantage" (e.g., a link choice resulting in extremely low latency). The online Critic receives the "reward." It updates its parameters by minimizing the "value loss." This process forces the online Critic's estimate to get closer and closer to the more accurate "reward" calculated by the GAE module.

[0129] Step 8: Target Network Update System: After the online Critic completes the parameter update, the weights of the online Critic are slowly copied to the target Critic through soft updates. This keeps the update of the "target value" smooth and stable, prevents drastic fluctuations during training, and greatly improves the convergence of the algorithm.

[0130] Step 9: The loop system clears the Trajectory Buffer, returns to Step 1, and continues to interact with the environment using the updated Actor. The entire "observation-decision-learning" loop is repeated continuously, and the agent eventually learns how to minimize latency under various weather and load conditions.

[0131] S5, simulation evaluation.

[0132] The state information is input into a pre-trained deep reinforcement learning agent model; the deep reinforcement learning agent outputs a joint optimization action based on the input state information. This action simultaneously determines the next trajectory of the UAV, the communication link selected for each ground base station, the sub-channel resources allocated to each link, and the transmission power allocated to each link; the system executes the joint optimization action to minimize the total transmission delay of all ground base stations and complete the data transmission.

[0133] The aforementioned UAV trajectory, communication link selection, sub-channel resource allocation, and transmit power allocation are jointly determined and output in the same decision-making step.

[0134] Simulation results show that:

[0135] like Figure 4 As shown, the convergence curves of different algorithms intuitively demonstrate the learning efficiency and final performance of various reinforcement learning algorithms in solving the complex communication resource scheduling problem constructed in this example. The graphs show that all four algorithms successfully learned effective policies, and their average reward (representing negative system latency) steadily increased with the number of training epochs, proving the feasibility of the DRL method for this problem. In the early stages of training, the performance curves of the various algorithms intertwined, reflecting the common phase of the agent exploring through numerous trial and error in an uncertain environment.

[0136] As training progresses, performance differences between algorithms become increasingly apparent. The standard PPO algorithm, serving as a baseline, is eventually surpassed by other improved algorithms, primarily due to its relatively basic exploration mechanism and network structure. Introducing a target network (PPO_Target) and a noise network (PPO_Noisy) significantly improves both convergence speed and final performance. Most importantly, our proposed PPO_Noisy_Target scheme exhibits optimal overall performance in the later stages of training, successfully combining the efficient exploration capabilities of the Noisy Net with the stable update mechanism of the Target Net, ultimately converging to the highest reward plateau. This strongly demonstrates that by combining advanced exploration and stabilization strategies, our approach can more effectively find near-optimal joint strategies for link selection, UAV deployment, and resource allocation under dynamically changing atmospheric environments and load demands.

[0137] As attached Figure 5 As shown, under clear weather conditions, all links are available, providing an ideal platform for evaluating the resource utilization efficiency of different strategies. As also shown, thanks to the high bandwidth of the FSO links, all strategies employing FSO significantly outperform the pure RF strategy in terms of latency. Furthermore, our complete agent strategy exhibits the strongest performance, consistently maintaining the lowest latency among all compared schemes. This is primarily due to its not only selection of the optimal FSO link but also joint optimization across multiple dimensions, including UAV trajectory, power, and sub-channel allocation, achieving the most efficient end-to-end management of communication resources. This result demonstrates that even under favorable channel conditions, our scheme's fine-grained resource scheduling capabilities simultaneously deliver significant performance gains.

[0138] As attached Figure 6As shown in the simulation results under cloudy weather, the advantages of our complete agent strategy in complex and dynamic environments are clearly demonstrated. In this scenario, the environment randomly switches between two conditions: "FSO direct connection interruption" and "FSO relay interference". As can be seen from the figure, any single fixed strategy performs poorly: the strategy of adhering to FSO direct connection suffers from latency spikes to 4000ms due to its inability to cope with cloud obstruction of the direct connection path; and other strategies relying on fixed links also perform far worse than our solution. Our complete agent strategy, with its precise perception of the environmental state, can dynamically make the optimal choice between FSO direct connection, relay, and RF, ultimately achieving the lowest average system latency. This strongly demonstrates that the core advantage of our solution lies in its flexible, real-time channel quality-based intelligent link switching capability.

[0139] As attached Figure 7 As shown, simulation results clearly reveal the robustness of different communication strategies under heavy rain conditions. All strategies relying on FSO (including direct connection and relay) experienced link interruption, resulting in system latency reaching the penalty limit of 8000ms. This is entirely consistent with physical reality, as high rainfall rates would cause devastating attenuation of FSO links in the 1550nm band. In contrast, strategies using RF links and our complete agent strategy demonstrated strong resilience. Notably, our solution can accurately identify the unavailability of FSO links and decisively switch to backup RF links, thereby stabilizing latency at a low level of approximately 500ms. This fully verifies the high reliability and intelligent link selection capability of our solution under extreme weather conditions.

[0140] This invention aims to address the problems existing in the background technology described above, and provides a dynamic transmission optimization method in hybrid RF-FSO networks. This method utilizes the collaboration of High Altitude Platforms (HAPs) and multiple Unmanned Aerial Vehicles (UAVs) to construct a dynamically reconfigurable transmission architecture. It also introduces a deep reinforcement learning algorithm to achieve joint optimization of UAV trajectories, link selection, sub-channels, and power resources. This allows for intelligent interference avoidance in complex atmospheric environments (especially cloud cover and rain / fog), minimizing the total system transmission latency and ensuring highly reliable, low-latency communication services.

[0141] This invention constructs a three-layer collaborative transmission architecture consisting of a High Altitude Platform (HAP), multiple Unmanned Aerial Vehicles (UAVs), and a Ground Base Station (BS), and designs a dynamic three-level link model for weather perception, including direct FSO links, UAV relay FSO links, and backup RF links. The core of this method lies in employing a deep reinforcement learning agent based on Proximal Policy Optimization (PPO) and incorporating Noisy Networks and Target Networks to model the complex optimization problem as a Markov decision process. Through interactive learning with the environment, this agent can output a joint optimization decision action including UAV trajectory, link selection, sub-channel allocation, and power allocation based on real-time weather conditions, channel quality, and traffic load, thereby achieving end-to-end system performance optimization.

[0142] This invention can intelligently avoid the adverse effects of cloud cover, rain, fog and other adverse weather conditions. Through an efficient joint optimization strategy, it can significantly reduce the total transmission latency of the system under different atmospheric conditions and improve communication reliability and spectrum resource utilization.

[0143] The conventional techniques and solutions not described in detail in the above embodiments are all well known in the art, and therefore will not be elaborated upon here. The above embodiments and / or experimental examples describe the preferred embodiments of the present invention in detail. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various simple modifications can be made to the technical solutions of the present invention, and these simple modifications all fall within the protection scope of the present invention.

Claims

1. A reinforcement learning-based link and resource optimization method in a hybrid RF-FSO network, characterized in that, Includes the following steps: A three-layer transmission architecture consisting of an aerial platform, a UAV relay layer, and a ground base station is constructed, and the status information of the system at the current decision moment is obtained; wherein, the three-layer transmission architecture is designed according to complex weather conditions for three optional communication links for data backhaul from the aerial platform to the ground base station, including a direct FSO link, a relay FSO link, and a backup RF link; Establish a joint system optimization model: aiming to minimize the total transmission delay of all base stations, through joint optimization of link selection. Given the UAV-base station association ζ, UAV trajectory Q, matrix representations ξ, θ, and power allocation P, establish an optimization problem P1; The optimization problem P1 is modeled as a Markov decision process; The PPONT agent algorithm is designed to solve optimization problems. The agent model interacts extensively with the simulation environment. By continuously updating the network parameters, an agent model that can implement the joint optimization strategy is finally trained. The state information is input into a pre-trained deep reinforcement learning agent model; The deep reinforcement learning agent model outputs joint optimization actions based on the input state information; The system performs the joint optimization action to minimize the total transmission latency of all ground base stations and complete the data transmission.

2. The link and resource optimization method based on reinforcement learning in the hybrid RF-FSO network according to claim 1, characterized in that, The optimization problem is: C6:g m,n ∈{0,1} Where D is the total system delay, and t represents the time slot; These represent HAP to ground base station I respectively n Through direct FSO link, UAV U m Communication is achieved via relay FSO links and backup RF links; These represent the transmit power of the HAP through the three links, P max (t) represents the maximum transmit power of the HAP; x min x m (t), x max U represents the unmanned aerial vehicle (UAV) m The minimum displacement position, current position, and maximum displacement position on the horizontal axis; y min y m (t), y max U represents the unmanned aerial vehicle (UAV) m The minimum displacement position, the current position, and the maximum displacement position on the vertical axis; ζ m,n The correlation coefficient between BS and UAV; The HAP connects to the ground base station I via three links. n Transmission delay, For ground base station I n Maximum tolerable delay; ξ n,k (t), θ n,s (t) are the allocation coefficients for the k-th RF channel and the s-th FSO channel, respectively, where K and S represent the number of sub-channels for the RF and FSO channels, respectively; v m (t) represents the velocity of the UAV, v max q represents the maximum speed at which a UAV can fly; m (t) represents the current position of the UAV.

3. The link and resource optimization method based on reinforcement learning in a hybrid RF-FSO network according to claim 1 or 2, characterized in that, The three-layer transmission architecture includes one high-altitude platform (HAP), which is stationary at high altitude and its coordinates are represented as q. H =[x H ,y H ,z H [ ], serving as the core hub for data storage and distribution; M UAV relay layers, represented as As an FSO dynamic relay node, it flies between the ground base station and the HAP, and its position in time slot t is represented as q. m (t)=[x m (t),y m [(t),z]; N ground base stations BS, represented as The position is fixed at q n =[x n ,y n ,z n ] Request data return from HAP; The direct FSO link is a direct free-space optical communication link between the high-altitude platform and the ground base station; the relay FSO link is a free-space optical communication link relayed by a UAV; the backup RF link is a radio frequency link between the high-altitude platform and the ground base station. The status information includes at least: the current coordinates of all UAVs, the data load of all ground base stations, weather conditions characterizing the current atmospheric environment, and channel gain estimates for each potential link between the high-altitude platform and each ground base station.

4. The link and resource optimization method based on reinforcement learning in the hybrid RF-FSO network according to claim 3, characterized in that, For the three selectable communication links, OFDMA multiple access technology is used for hybrid resource allocation; considering various factors, channel models and delay calculation methods are established for each of the three links, including: When weather conditions are good, the HAP transmits to the BS via a direct FSO link; when cloud cover with high liquid water content blocks the signal, a UAV is deployed to a location with negligible cloud cover to use a relay FSO link to reflect and forward the optical signal from the HAP to the BS; when weather conditions are very bad, a backup RF link is used.

5. The link and resource optimization method based on reinforcement learning in a hybrid RF-FSO network according to claim 1, characterized in that, The optimization problem P1 is modeled as a Markov decision process, and its main elements are defined as follows: state space state It contains the environmental information needed to make a decision at time t. Specifically, it includes: normalized two-dimensional coordinates of all UAVs, normalized distance between UAV pairs (if M>1), normalized data payload e(t) of all ground base stations, index of current weather conditions and its one-hot encoding, normalized indices reflecting the potential impact of cloud thickness and rainfall rate on link performance, and normalized channel gain estimates for the three links between HAP and each BS; Action space action It is the set of control variables output by the agent at each decision step, including UAV movement, link selection and sub-channel allocation, and power allocation; Reward function (R): Define the reward expression R t =-D total (t)-P penalty (t); where D total (t) represents the total transmission delay of the system; P penalty (t) represents the compound penalty term.

6. The link and resource optimization method based on reinforcement learning in a hybrid RF-FSO network according to claim 1, characterized in that, The PPONT agent algorithm is based on the PPO algorithm and introduces Noisy Networks and Target Networks techniques. Specifically, it includes a parameter... Actor Network An online Critic network V with parameter φ φ (s t A target Critic network V with parameter φ′ φ′ (s t The training data is generated through the interaction between the agent and the environment and stored in the experience replay buffer. The update of the Actor network depends on the pruning substitution objective function of PPO. The Critic network is updated by minimizing the value function estimation error. The parameters of the Actor and Critic networks are updated by the Adam optimizer.

7. The link and resource optimization method based on reinforcement learning in the hybrid RF-FSO network according to claim 6, characterized in that, The noisy network adds parameterized noise to the network's weights and biases; a linear layer with noise parameters can be represented as: y=(μ w +s w ⊙∈ w )x+(μ b +s b ⊙∈ b ) Among them, (μ w ,μ b ) and (σ w ,σ b ) are learnable mean and standard deviation parameters, while (∈ w ,∈ b The noise variable is sampled from a fixed distribution; the variance parameter of the noise (σ) w ,σ b The network weights μ are learned via gradient descent. Two independent sets of network parameters are maintained: an online Critic network and a target Critic network. The online Critic network interacts with the environment and updates policies and values, while the target Critic network provides a relatively stable objective for calculating the TD error and advantage function. The parameters φ′ of the target Critic network are slowly and softly updated using the parameters φ of the online Critic network. φ′←τφ+(1-τ)φ′ Here, τ is a small hyperparameter.

8. The link and resource optimization method based on reinforcement learning in a hybrid RF-FSO network according to claim 6 or 7, characterized in that, The PPONT agent algorithm is designed to solve the optimization problem to obtain a pre-trained deep reinforcement learning agent model, including the following steps: The environment collects the current global state s at time t. t And send it to the Actor network; the state s t Includes the current location of all UAVs, the data load of all BSs, current weather conditions, and estimated link quality; the environment is a HAP-UAV-BS system; The Actor uses a noisy network instead of traditional random sampling to output a complex joint action 'a'. t And send it back to the environment for execution; the action a t It is a decision package containing multiple parts, which determines the next movement vector of each UAV, allocates the best link to each BS, determines how much transmit power the HAP allocates to each link, and dynamically allocates sub-channel bandwidth to all FSO links and RF links; Execute action a in the environment t It calculates the total transmission delay of the time slot and generates a reward r. t At the same time, a new state s is generated. t+1 ; The complete empirical tuple s t a t r t s t+1 Store it in the Trajectory Buffer; Once the trajectory buffer has collected enough data, a mini-batch of samples is randomly selected from the Trajectory Buffer and sent to the Actor, Critic, and Generalized Advantage Estimation (GAE) modules for computation. The GAE module gathers the following three data sources for calculation and outputs the calculated values ​​of the advantage function and reward function; the three data sources include the current value estimated by the online Critic network, the target value estimated by the target Critic network, and the reward sampled from the Trajectory Buffer; Network Update: Actors receive the advantage and update their parameters by optimizing the loss function; online critics receive the reward and update their parameters by minimizing the value loss. After the online Critic completes the parameter update, the weights of the online Critic are slowly copied to the target Critic through soft updates. The loop system clears the Trajectory Buffer, returns to step 1, and continues to interact with the environment using the updated Actor. The loop repeats continuously, and the agent eventually learns how to minimize latency under various weather and load conditions.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the reinforcement learning-based link and resource optimization method in the hybrid RF-FSO network as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-dimensional quality of service driven robust resource allocation method for space-air-ground integrated fso / rf

    CN120239084B