Unmanned aerial vehicle relay fso, rf communication method and system based on hybrid reinforcement learning

By employing a hybrid reinforcement learning approach, the problem of strong coupling between discrete link switching and continuous power control in UAV relay communication systems under complex low-altitude environments was solved. This approach enabled collaborative optimization of the UAV FSO/RF link, improving system transmission efficiency and security, and meeting the optimization requirements of user fairness and multiple constraints.

CN121940037BActive Publication Date: 2026-06-09NANJING UNIV OF INFORMATION SCI & TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV OF INFORMATION SCI & TECH
Filing Date
2026-03-30
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing UAV relay communication systems struggle to effectively handle the coupling issues of discrete link switching, continuous power control, and three-dimensional maneuvering decision-making in complex low-altitude environments. This leads to reduced system capacity, limited transmission efficiency, and physical layer security risks. Furthermore, existing methods cannot balance user fairness with confidentiality.

Method used

A hybrid reinforcement learning-based approach is adopted to formulate the UAV relay communication problem as a finite-time Markov decision process. The solution is obtained by a hybrid soft actor-commentator algorithm, which realizes the joint optimization of UAV FSO/RF link switching, power allocation and 3D trajectory. A hybrid state space and action space are designed, and a time-slotted access mechanism combining steady-state mode and switching fallback mode is constructed to build a joint optimization model for multi-user end-to-end effective confidentiality throughput maximum minimum fairness.

Benefits of technology

In complex low-altitude obstacle environments, the system achieved coordinated optimization of UAV FSO/RF link switching, power allocation, and 3D trajectory, improving the end-to-end effective security rate and the fairness and robustness of user communication, and enhancing the coverage elasticity and service continuity of the integrated air-space-ground network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121940037B_ABST
    Figure CN121940037B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for UAV relay FSO and RF communication based on hybrid reinforcement learning, obtaining a joint optimization problem of hybrid RF and FSO UAV relay; the joint optimization problem of hybrid RF and FSO UAV relay is formulated as a finite-time Markov decision process; the finite-time Markov decision process is solved to obtain a joint optimization strategy for UAV FSO and RF link switching, power allocation, and three-dimensional trajectory. This invention provides an efficient solution for secure communication in UAV relay and achieves deep coupling optimization of discrete link switching and continuous power and trajectory control, breaking through the technical bottleneck of traditional methods in handling hybrid decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method and system for FSO and RF communication of unmanned aerial vehicles (UAVs) based on hybrid reinforcement learning, belonging to the field of UAV communication technology. Background Technology

[0002] With the rapid development of 6G mobile communication and low-altitude intelligent network technologies, unmanned aerial vehicles (UAVs) are gradually becoming important aerial nodes in integrated air-space-ground networks due to their advantages such as flexible deployment, high mobility, and strong three-dimensional coverage capabilities. In mountainous areas, densely built-up areas, disaster relief areas, and scenarios with temporary communication needs, where ground base stations are unable to provide continuous and stable coverage, UAVs can quickly establish aerial relay links to provide ground users with emergency access, enhanced edge coverage, and data forwarding services, thereby improving network coverage resilience and service continuity.

[0003] Traditional UAV relay systems mostly rely on a single radio frequency (RF) link for user access and data backhaul. However, in scenarios with multiple users accessing concurrently, a single RF link is susceptible to factors such as limited spectrum resources, increased co-channel interference, and intensified channel contention, leading to reduced system capacity and limited transmission efficiency. Simultaneously, RF signals have open broadcast propagation characteristics, making them vulnerable to eavesdropping by unauthorized nodes, thus posing physical layer security risks. In complex low-altitude environments, how to improve system throughput while ensuring transmission fairness and confidentiality has become a crucial technical challenge in UAV relay communication.

[0004] To improve backhaul link capacity, free-space light (FSO) communication is widely considered an important supplementary technology for UAV backhaul links due to its advantages such as high bandwidth, high directionality, resistance to electromagnetic interference, and no spectrum licensing. Under good line-of-sight conditions, FSO links can provide significantly higher data transmission capabilities than RF links, and due to their narrow beam and concentrated energy, they naturally possess better anti-eavesdropping capabilities, contributing to improved physical layer security. However, FSO communication is highly dependent on link alignment accuracy and line-of-sight conditions, and in low-altitude urban environments, it is susceptible to building obstruction, atmospheric turbulence, and platform attitude disturbances, leading to link quality fluctuations or even outages. Furthermore, when a UAV needs to switch FSO service targets between different users, it often needs to undergo a realignment process, introducing additional latency and link switching overhead.

[0005] Based on this, a hybrid UAV relay communication system can be constructed by combining RF and FSO links. This leverages the robust access capability of RF links under non-line-of-sight or weak line-of-sight conditions, and the high throughput and secure backhaul capability of FSO links under good line-of-sight conditions, thus maximizing the complementary advantages of the two types of links. However, in actual low-altitude complex scenarios, this type of system still faces several challenges: First, there is a significant coupling relationship between user access, FSO service object switching, power allocation, and UAV three-dimensional maneuvering decisions, involving both discrete mode selection and continuous control variables, resulting in a typical hybrid action characteristic of the optimization problem. Second, the system must simultaneously satisfy backhaul capacity constraints, flight boundary constraints, return-to-home reachability constraints, and long-term fairness constraints for both users, while also considering the improvement of effective secure throughput at the end-to-end level, further increasing the difficulty of joint optimization.

[0006] In existing technologies, some studies employ convex optimization, alternating optimization, or heuristic algorithms to optimize the communication resources and trajectories of UAV relay networks. However, these methods typically rely on strong model solvability assumptions and struggle to effectively handle issues such as rapid changes in link states, discrete handover decisions, and strong coupling between continuous power and motion control in low-altitude obstacle environments. Other studies have attempted to apply reinforcement learning methods to wireless resource management or UAV trajectory control, but most focus on single action types, single link modes, or single-objective optimization scenarios. For FSO / RF hybrid relay systems oriented towards physical layer security, a complete technical solution is still lacking that can simultaneously handle discrete FSO service object handover, continuous power control, UAV three-dimensional maneuver control, and fairness-oriented secure communication rate optimization.

[0007] Therefore, there is an urgent need to propose a hybrid FSO / RF relay cooperative optimization method for UAVs in low-altitude complex obstacle environments. This method should be used to jointly realize link switching decisions, power allocation and flight trajectory control while taking into account the long-term fairness of both users. This would improve the system's end-to-end effective confidential throughput, enhance the reliability of the backhaul link, and meet the application requirements of efficient, secure and robust communication in low-altitude intelligent networking and 6G scenarios. Summary of the Invention

[0008] Objective: To overcome the shortcomings of existing technologies, this invention provides a hybrid FSO and RF relay cooperative communication method for UAVs based on hybrid reinforcement learning. This invention provides an efficient solution for secure relay communication of UAVs and achieves deep coupling optimization of discrete link switching and continuous power and trajectory control, breaking through the technical bottleneck of traditional methods in handling hybrid decision-making.

[0009] Technical solution: To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0010] Firstly, a method for UAV relay FSO and RF communication based on hybrid reinforcement learning, specifically including:

[0011] To obtain the joint optimization problem of hybrid RF and FSO UAV relay.

[0012] The joint optimization problem of hybrid RF and FSO UAV relay is formulated as a finite-time Markov decision process.

[0013] By solving the finite-time Markov decision process, a joint optimization strategy for UAV FSO, RF link switching, power allocation, and 3D trajectory is obtained.

[0014] Optionally, the hybrid RF / FSO UAV relay joint optimization problem specifically includes:

[0015]

[0016] Where P1 represents the joint optimization problem model of the UAV relay communication system, Indicates user The cumulative end-to-end effective confidentiality throughput throughout the entire mission cycle Let the UAV be in the three-dimensional position of time slot n. Let n be the three-dimensional position of the UAV in time slot n+1. This indicates the actual access status of the communication system in time slot n. Indicates user At the RF transmit power in time slot n, Indicates user The maximum RF transmit power, Indicates user The FSO transmitted optical power in time slot n, Indicates user The maximum transmit power of FSO, This indicates the power of the drone to interfere with the eavesdropper's actions. This indicates the FSO (Flight-Oriented Response) transmit optical power from the drone to the base station. This indicates the maximum electrical power limit of the drone. Let x be the x-axis coordinate of the drone. Let y be the y-axis coordinate of the drone. Let Z be the z-axis coordinate of the UAV. , These represent the upper bounds of the system service area along the x-axis and y-axis, respectively. , These represent the minimum and maximum permissible flight altitudes for the drone, respectively. Represents the maximum flight speed of the drone. This represents the three-dimensional position of the drone in time slot 1. This represents the three-dimensional position of the UAV in time slot N. Represents the starting point of drones, Represents the endpoint of the drone. On behalf of users Using FSO uplink, users Using RF uplink, On behalf of users Using FSO uplink, users Using RF uplink, This indicates that the system is in an FSO switchover period, and FSO uplink is temporarily unavailable. Both users are transmitting uplink data via the RF link. Represents photoelectric conversion efficiency. Represents the maximum power of the UAV. This represents the length of each time slot n.

[0017] Optionally, the expression for the finite-time Markov decision process is as follows:

[0018]

[0019] in, For a finite-time Markov decision process, For state space, For a mixed action space, Represents the state transition probability. For the instantaneous reward of time slot n, As a discount factor, This represents the state of time slot n+1.

[0020] in, The state of time slot n The expression is as follows:

[0021]

[0022] in, Used to characterize the current geometric position, link mode, and key distance information of the UAV. Used to characterize the fairness of cumulative end-to-end effective confidentiality throughput for dual users. Used to characterize the return-to-home reachability of the drone in the current state. Used to introduce the actual execution displacement information of the previous time slot.

[0023] Action of time slot n The expression is as follows:

[0024]

[0025] in, This represents the discrete action of the nth time slot, used to determine the steady-state FSO service object selection for the current time slot. This indicates the continuous action of the nth time slot, used to jointly control the user's transmit power, UAV artificial interference power, UAV return optical power, and the three-dimensional displacement of the UAV.

[0026] The expression is as follows:

[0027]

[0028] in, As a feasibility penalty item, For indicator functions, This represents the terminal convergence interval, which is only when... , Otherwise, it is 0. For the sake of fairness, Due to return-to-base constraints, The indicator function representing the Nth time slot. Define a reward for the minimum cumulative throughput at each terminal. For the terminal fairness penalty item, Positive rewards are given to encourage terminals to achieve stricter fairness targets. This is a penalty item for terminal return error.

[0029] Optionally, the finite-time Markov decision process is solved using a hybrid soft actor-commentator algorithm.

[0030] Optionally, the finite-time Markov decision process is solved using a hybrid soft actor-commentator algorithm, specifically including:

[0031] A policy network for discrete and continuous actions is constructed. The policy network for discrete actions outputs discrete action branches, which are used to determine the FSO service object selection or link mode selection for the current time slot. The policy network for continuous actions outputs continuous action branches, which are used to jointly control the user's transmit power, the UAV's artificial interference power, the UAV's backhaul optical power, and the UAV's three-dimensional displacement.

[0032] Within the maximum entropy reinforcement learning framework, entropy regularization terms are introduced for both discrete and continuous action branches, and the policy networks for discrete and continuous actions are trained using an adaptive adjustment method based on learnable temperature parameters.

[0033] A dual-critic network is used to evaluate the joint action value of the policy network output. Combined with an experience replay mechanism and a target network update mechanism, the agent learns a joint optimization strategy for UAV FSO / RF link switching, power allocation and 3D trajectory under conditions of occlusion, backhaul bottleneck and fair constraints through continuous interaction and iteration with the environment.

[0034] Optionally, the The expression is as follows:

[0035]

[0036] in, This represents the proportion of data that can be transmitted back to the source. Represented as user The instantaneous security rate in time slot n, where N represents the number of time slots.

[0037] Optionally, the The expression is as follows:

[0038]

[0039] In the formula, For numerically stable terms, This represents the uplink transmission rate of the two users in time slot n. This indicates the transmission rate from the drone to the base station.

[0040] Among them, for Specifically, they can be categorized into the following situations:

[0041] In normal working mode, when At that time, the user Using FSO uplink, users Using RF uplink, users The RF transmission rate to the drone is expressed as:

[0042]

[0043] in, The bandwidth used for the RF channel. For users At the RF transmit power in time slot n, For users To the RF channel gain of the drone, This represents the noise power of the RF link.

[0044] user The FSO transmission rate to the drone is expressed as:

[0045]

[0046] in, For FSO uplink bandwidth, Indicates user SNR of the FSO link to the drone.

[0047] when At that time, the user Using FSO uplink, users Using RF uplink, users The RF transmission rate to the drone is expressed as:

[0048]

[0049] in, For users At the RF transmit power in time slot n, For users RF channel gain for drones.

[0050] user The FSO transmission rate to the drone is:

[0051]

[0052] in, Indicates user SNR of the FSO link to the drone.

[0053] In switching modes, when Two users simultaneously transmit uplink to the UAV via the RF link. Correspondingly, the user with the higher received power is selected based on their received power, sorted from highest to lowest. and smaller users The transmission rates at the UAV end are as follows:

[0054]

[0055] in, and Represented as users and users Signal-to-interference-plus-noise ratio (SIR) of non-orthogonal multiple access communication.

[0056] Optionally, the The expression is as follows:

[0057]

[0058] in, , For eavesdroppers to target current RF uplink users The rate of eavesdropping.

[0059] in, The expression is as follows:

[0060]

[0061] when or hour, Indicates a user using RF uplink Its equivalent signal-to-dryness ratio at the eavesdropping end is expressed as:

[0062]

[0063] in, This is represented as the equivalent power noise at the receiver's end. Indicates the user in the current time slot RF transmit power, The artificial interference power of the eavesdropper. Indicates user To the RF channel gain of the eavesdropper, This represents the channel gain of the interference link from the UAV to the eavesdropper within time slot n.

[0064] when At this time, the FSO uplink is temporarily unavailable, and both users simultaneously transmit uplink data via the RF link. In this situation, an eavesdropper will simultaneously receive RF signals from both users. Correspondingly, They are respectively , , Indicates to the user eavesdropping rate, Indicates to the user The eavesdropping rate, sorted by received power from highest to lowest, for users with higher power. and smaller users The eavesdropping rates are as follows:

[0065]

[0066] in, This indicates that the eavesdropper is decoding the user. At that time, the user The signal is considered interference, and the corresponding signal-to-interference-plus-noise ratio (SINNR) is... This indicates that the user has been successfully eliminated. After receiving the signal, the eavesdropper further decoded the user's signal. The corresponding signal-to-interference-plus-noise ratio at this time;

[0067]

[0068] in, and These represent users within the current time slot. and RF transmit power, and Each represents the corresponding user , RF channel gain for drones.

[0069] Optionally, the The expression is:

[0070]

[0071] in, , and Let represent the normalized three-dimensional position features of the UAV in the nth time slot. This represents the normalized time slot progress characteristics. This represents the normalized discrete link mode characteristics. This represents the normalized mode switching countdown feature. This represents the normalized switching coolant count characteristic. , and These represent the normalized distance characteristics from the user to the drone, from the drone to the base station, and from the drone to the eavesdropper, respectively.

[0072] The The expression is:

[0073]

[0074] in, Represents the normalized user Cumulative end-to-end secure throughput up to time slot n Represents the normalized user Cumulative end-to-end secure throughput up to time slot n Indicates the proportion of fairness gap. The lagging user index variable is used to assist discrete pattern decision-making by prioritizing the identification of users with low current cumulative end-to-end effective confidentiality throughput.

[0075] The The expression is:

[0076]

[0077] in, This represents the normalized distance characteristic from the UAV to the target endpoint. This represents the reachability margin feature of the normalized UAV in the current remaining time domain to reach the target endpoint.

[0078] The The expression is:

[0079]

[0080] in, , and They are respectively represented as the normalized values ​​at the th... The actual displacement components of the UAV in the x-axis, y-axis and z-axis directions within each time slot.

[0081] In a second aspect, a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a UAV relay FSO / RF communication method based on hybrid reinforcement learning as described in any of the first aspects.

[0082] Thirdly, a computer device comprising:

[0083] Memory is used to store instructions.

[0084] A processor is configured to execute the instructions, causing the computer device to perform operations of a hybrid reinforcement learning-based unmanned aerial vehicle relay FSO / RF communication method as described in any of the first aspects.

[0085] Beneficial effects: The UAV relay FSO and RF communication method and system based on hybrid reinforcement learning provided by this invention have the following advantages compared with existing methods:

[0086] 1. This invention addresses the challenges of strong coupling between discrete link switching and continuous power and trajectory control in complex low-altitude obstacle environments, where physical layer security, user communication fairness, and multiple constraints are difficult to optimize uniformly. It proposes a time-slotted access mechanism combining steady-state mode and a fallback switching mode. A joint optimization model is constructed with the objective of maximizing and minimizing fairness in end-to-end effective confidentiality throughput for multiple users, and this is transformed into a finite-time-domain Markov decision process for solution. Furthermore, a hybrid state space, hybrid action space, and constraint fusion reward mechanism are designed to achieve adaptive collaborative optimization of UAV trajectory, link switching, and power allocation. Compared to existing hybrid RF and FSO communication methods for UAV relays, this invention can achieve collaborative optimization of UAV FSO, RF link switching, power allocation, and 3D trajectory in complex low-altitude obstacle environments, effectively improving the end-to-end effective confidentiality rate of the system in multi-user scenarios, while effectively ensuring the fairness and robustness of user communication.

[0087] 2. This invention addresses the challenges of discrete link switching and strong coupling between continuous power and trajectory control in hybrid FSO / RF relay cooperative communication for UAVs in complex low-altitude obstacle environments, as well as the technical pain points of unifying and optimizing physical layer security, user communication fairness, and multiple constraints. It proposes a cooperative optimization method based on hybrid reinforcement learning. Compared to traditional methods such as convex optimization and alternating optimization that rely on precise models, this invention does not require a pre-set environment model. It generates the optimal control strategy through interactive learning with the environment and dynamically processes discrete and continuous quantities, as well as rationally switching between different modes. This results in stronger adaptability and generalization capabilities in dynamic and uncertain low-altitude communication environments, improving the overall utilization efficiency of communication resources and significantly enhancing the coverage elasticity, service continuity, and physical layer security performance of the integrated air-space-ground network. It has significant engineering application value and promotional significance.

[0088] 3. Most existing reinforcement learning methods for UAV relay communication focus on single action types, single link modes, or single-objective optimization scenarios, making it difficult to simultaneously handle discrete FSO service object switching, continuous power control, UAV three-dimensional maneuver control, and multi-constraint fairness and secure communication optimization. In contrast, this invention solves the problem of collaborative optimization of FSO service object switching, power allocation, and three-dimensional trajectory in low-altitude obstacle environments, aiming at improving secure rate and user communication fairness. Attached Figure Description

[0089] Figure 1 This is a structural diagram of the UAV relay FSO and RF communication system of the present invention.

[0090] Figure 2 This is a schematic diagram of the three-dimensional flight path of the UAV of the present invention.

[0091] Figure 3 This is a schematic diagram comparing the reward values ​​of the various algorithms in this invention.

[0092] Figure 4 This is a schematic diagram comparing the minimum confidential transmission volume for each algorithm in this invention.

[0093] Figure 5 This is a schematic diagram comparing the total confidential transmission volume of each algorithm in this invention.

[0094] Figure 6 This is a schematic diagram comparing the fair difference ratio of the amount of data transmitted between two users in each algorithm of the present invention.

[0095] Figure 7 This is a schematic diagram comparing the confidential transmission volume of each algorithm of the present invention for two users. Detailed Implementation

[0096] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0097] The present invention will be further described below with reference to specific embodiments.

[0098] Example 1:

[0099] This embodiment introduces a UAV relay FSO and RF communication method based on hybrid reinforcement learning, specifically including the following steps:

[0100] Step 1: Construct a drone relay FSO and RF communication system, such as Figure 1 As shown, the communication system includes: two legitimate users. and users A UAV (unmanned aerial vehicle) acting as an airborne relay, a ground base station, and an eavesdropper E.

[0101] The UAV is responsible for receiving uplink data from users and forwarding the received data to the base station via an FSO (Free Space Light) backhaul link. Due to building obstructions in complex urban environments, the propagation status of the RF (Radio Frequency) and FSO legitimate links dynamically changes with the UAV's three-dimensional position. Specifically, the RF link typically exhibits additional attenuation under obstruction, while the FSO link can be considered interrupted. Meanwhile, although the FSO link offers high bandwidth and strong anti-eavesdropping capabilities, it is highly sensitive to obstruction, pointing errors, and handover realignment. Therefore, the communication system needs to make a joint trade-off between link transmission performance, handover overhead, and user fairness.

[0102] Let the total duration of the task be T. Discretize the entire task into N equal-length time slots, where the length of each time slot n is... Therefore, the three-dimensional position of the UAV in time slot n can be represented as:

[0103]

[0104] in, Here is the x-axis coordinate of the UAV. Here is the y-axis coordinate of the UAV. Here is the z-axis coordinate of the UAV.

[0105] Define user The position is The location of the base station is The location of the eavesdropper is .in, users respectively x-axis coordinates, y-axis coordinates, z-axis coordinates. These are the x-axis, y-axis, and z-axis coordinates of the base station, respectively. These are the x-axis, y-axis, and z-axis coordinates of the eavesdropper.

[0106] In this invention, all ground nodes are considered stationary. Therefore, the time-varying geometric relationship in the communication system is determined only by the positional changes of the UAV.

[0107] Furthermore, the Euclidean distance between the UAV and each node is defined as follows:

[0108]

[0109] in, Indicates the distance from the user to the drone (UAV). This indicates the distance from the unmanned aerial vehicle (UAV) to the base station (BS). Indicates the distance between the drone (UAV) and the eavesdropper. This indicates the distance between the user and the eavesdropper. It should be noted that... , and For time slot related quantities, and This is the static distance.

[0110] Step 2: To ensure fairness in long-term service for both users, we cannot solely pursue higher secure transmission gains for one user. Instead, we should strive to balance the cumulative end-to-end effective secure throughput of both users throughout the entire task cycle. Therefore, this paper adopts the max-min fairness optimization criterion, using the smaller of the cumulative end-to-end effective secure throughput of the two users as the objective function. Thus, the joint optimization problem can be specifically expressed as:

[0111]

[0112] In the formula, P1 represents the joint optimization problem model of the UAV relay communication system. Indicates user The cumulative end-to-end effective confidentiality throughput throughout the entire mission cycle This indicates the actual access status of the communication system in time slot n. Indicates user At the RF transmit power in time slot n, Indicates user The maximum RF transmit power, Indicates user The FSO transmitted optical power in time slot n, Indicates user The maximum transmit power of FSO, This indicates the power of unmanned aerial vehicles (UAVs) to interfere with eavesdroppers. This indicates the FSO (Frequency-Oriented Response) transmit optical power from the UAV (Unmanned Aerial Vehicle) to the base station. This indicates the maximum electrical power limit of a drone (UAV). , These represent the upper bounds of the system service area along the x-axis and y-axis, respectively. , These represent the minimum and maximum permissible flight altitudes for UAVs, respectively. Represents the maximum flight speed of the UAV. This represents the three-dimensional position of the UAV in time slot 1. This represents the three-dimensional position of the UAV in time slot N. The status is represented as 0 (user) Using FSO uplink, users (using RF uplink) The status is represented as 1 (user) Using FSO uplink, users (using RF uplink) The status is 2 (the system is in the FSO switching period, FSO uplink is temporarily unavailable, and both users are transmitting uplink through the RF link). Represents photoelectric conversion efficiency. This represents the maximum power of the UAV.

[0113] Furthermore, the aforementioned user Cumulative end-to-end effective confidentiality throughput throughout the entire mission cycle The expression is as follows:

[0114]

[0115] In the formula, This represents the proportion of data that can be transmitted back to the source. Represented as user The instantaneous security rate in time slot n, specifically, It can be represented as:

[0116]

[0117] In the formula, For numerically stable terms, This represents the uplink transmission rate of the two users in time slot n. This indicates the transmission rate from the drone (UAV) to the base station.

[0118] Furthermore, for Specifically, these can be categorized into the following situations:

[0119] In normal operating mode, the communication system always maintains an access method where one user uses RF uplink and the other user uses FSO uplink. Therefore, when At that time, the user Using FSO uplink, users Using RF uplink, the user The RF transmission rate to the UAV is expressed as:

[0120]

[0121] in, The bandwidth used for the RF channel. For users At the RF transmit power in time slot n, For users RF channel gain to unmanned aerial vehicles (UAVs) This refers to the noise power of the RF link. Meanwhile, the user... The FSO transmission rate is expressed as:

[0122]

[0123] in, This refers to the uplink bandwidth of the FSO. Indicates user SNR of the FSO link to UAV.

[0124] Similarly, when At that time, the user Using FSO uplink, users Using RF uplink, at this time, the user The RF transmission rate to the UAV is expressed as:

[0125]

[0126] in, For users At the RF transmit power in time slot n, For users RF channel gain to UAV.

[0127] User 2's transmission rate is:

[0128]

[0129] in, Indicates user SNR of the FSO link to UAV.

[0130] In switching mode, i.e. Two users simultaneously transmit uplink to the UAV via the RF link. Correspondingly, the user with the higher received power is selected based on their received power, sorted from highest to lowest. and smaller users The transmission rates at the UAV end are as follows:

[0131]

[0132] in, and Represented as users and users The signal-to-interference-plus-noise ratio (SINR) for non-orthogonal multiple access communication is specifically expressed as:

[0133]

[0134] in, and These represent users within the current time slot. and RF transmit power, and Each represents the corresponding user and RF channel gain to unmanned aerial vehicles (UAVs).

[0135] Furthermore, It can be represented as:

[0136]

[0137] in, For FSO backhaul link bandwidth, The SNR (signal-to-noise ratio) of the UAV-to-FSO backhaul link can be expressed as:

[0138]

[0139] in, For the backhaul optical power of the unmanned aerial vehicle (UAV) in time slot n, For the backhaul gain of the unmanned aerial vehicle (UAV) in time slot n, This represents the equivalent noise power of the FSO backhaul link. Within each time slot, the system needs to simultaneously satisfy power constraints on both the UAV side and the user side. The UAV's transmit power in time slot n includes artificial interference power directed at eavesdroppers. and FSO backhaul transmit optical power towards base station Therefore, the total electrical power constraint for the unmanned aerial vehicle (UAV) in time slot n is:

[0140]

[0141]

[0142] in, For photoelectric conversion efficiency, This is the upper limit of the maximum electrical power of the UAV.

[0143] Furthermore, It can be represented as:

[0144]

[0145] in, , For eavesdroppers to target current RF uplink users The rate of eavesdropping.

[0146] Specifically, It can be represented as:

[0147]

[0148] The specific categories are as follows:

[0149] when or At that time, it indicates a user using RF uplink. Its equivalent signal-to-dryness ratio at the eavesdropping end can be expressed as:

[0150]

[0151] in, This is represented as the equivalent power noise at the receiver's end. Indicates the user in the current time slot RF transmit power, The artificial interference power of the eavesdropper. Indicates user To the RF channel gain of the eavesdropper, This represents the channel gain of the interference link from the UAV to the eavesdropper within time slot n.

[0152] when At this time, the FSO uplink is temporarily unavailable, and both users simultaneously transmit uplink through the RF link. In this situation, an eavesdropper will simultaneously receive RF signals from both users. Correspondingly, the user with the higher received power will be selected based on their received power, from highest to lowest. and smaller users The eavesdropping rates are as follows:

[0153]

[0154] in, This indicates that the eavesdropper is decoding the user. At that time, the user The signal is considered interference, and the corresponding signal-to-interference-plus-noise ratio (SINNR) is... This indicates that the user has been successfully eliminated. After receiving the signal, the eavesdropper further decoded the user's signal. The corresponding signal-to-interference-plus-noise ratio at this time is specifically expressed as:

[0155]

[0156] in, and These represent users within the current time slot. and RF transmit power, and Each represents the corresponding user , RF channel gain to UAV.

[0157] Step 3: Formulate the established hybrid RF / FSO UAV relay joint optimization problem as a finite-time Markov decision process, expressed as:

[0158]

[0159] in, For state space, For a mixed action space, Represents the state transition probability. For the instantaneous reward of time slot n, This is the discount factor. The randomness in state transitions mainly stems from environmental uncertainties such as link obstruction state changes, atmospheric turbulence disturbances, and FSO pointing errors. The agent's goal is to learn the optimal policy. To maximize the expected cumulative return from the discount, i.e.:

[0160]

[0161] Considering that system decision-making involves both discrete choice and continuous control, the action space is modeled as a hybrid form, specifically as follows:

[0162]

[0163] in, Represents the discrete action subspace. This represents the subspace of continuous actions.

[0164] Specifically, the joint optimization problem is modeled as a finite-time Markov decision process, with its state space designed as follows:

[0165]

[0166] in, Used to characterize the current geometric position, link mode, and key distance information of unmanned aerial vehicles (UAVs). Used to characterize the fairness of cumulative end-to-end effective confidentiality throughput for dual users. Used to characterize the return-to-home reachability of a drone (UAV) in the current state. This is used to incorporate the actual displacement information from the previous time slot to enhance the smoothness and stability of trajectory control.

[0167] Furthermore, the basic state Represented as:

[0168]

[0169] in, , and These represent the normalized three-dimensional position features of the unmanned aerial vehicle (UAV) in the nth time slot. This represents the normalized time slot progress characteristics. This represents the normalized discrete link mode characteristics. This represents the normalized mode switching countdown feature. This represents the normalized switching coolant count characteristic. , and These represent the normalized distance characteristics from the user to the drone (UAV), from the drone (UAV) to the base station, and from the drone (UAV) to the eavesdropper, respectively.

[0170] Furthermore, the aforementioned fairness state Represented as:

[0171]

[0172] in, Represents the normalized user Cumulative end-to-end secure throughput up to time slot n Represents the normalized user Cumulative end-to-end secure throughput up to time slot n Indicates the proportion of fairness gap. The lagging user index variable is used to assist discrete pattern decision-making in prioritizing the identification of users with low current cumulative end-to-end effective confidentiality throughput;

[0173] Preferably, the settlement user indicator variable Defined as:

[0174]

[0175] in, Indicates the user in the nth time slot Currently in a lagging state, otherwise it indicates the user Currently in a backward state;

[0176] Furthermore, the return-to-home reachability state Represented as:

[0177]

[0178] in, This represents the normalized distance characteristic from the unmanned aerial vehicle (UAV) to the target endpoint. This represents the reachability margin of a normalized unmanned aerial vehicle (UAV) to reach its target endpoint within the current remaining time domain. It is used to guide strategies to balance effective confidentiality throughput and fairness with terminal return constraints.

[0179] Furthermore, the control smooth state Represented as:

[0180]

[0181] in, , and They are respectively represented as the normalized values ​​at the th... The actual displacement components of the UAV in the x-axis, y-axis and z-axis directions within a time slot, by incorporating the control smoothing state into the state space, enable the strategy to combine the actual maneuvering information of the previous time slot when making decisions in the current time slot, thereby suppressing trajectory abrupt changes and improving the smoothness and physical feasibility of the flight trajectory;

[0182] Specifically, the action space is represented as:

[0183]

[0184] in, This represents the discrete action of the nth time slot, used to determine the steady-state FSO service object selection for the current time slot. This represents the continuous action of the nth time slot, used to jointly control the user's transmit power, UAV artificial interference power, UAV return optical power, and the three-dimensional displacement of the UAV. The discrete and continuous actions together constitute the joint control input for the current time slot, enabling hybrid FSO / RF access mode selection, power allocation, and UAV trajectory control.

[0185] Furthermore, the single-step slot reward for slot n is:

[0186]

[0187] While satisfying UAV backhaul bottleneck constraints, flight boundary constraints, and return-to-home reachability constraints, this approach also considers long-term fairness for both users and improves the system's end-to-end effective confidentiality throughput. To this end, multi-objective optimization is integrated into a single-step reward by combining throughput gain, fairness constraint penalty, bottleneck penalty, feasibility penalty, and terminal convergence term. This is to guide HSAC in learning stable policies in a hybrid discrete-continuous action space, where... As a feasibility penalty item, For indicator functions, This represents the terminal convergence interval, which is only when... , Otherwise, it is 0. For the sake of fairness, Due to return-to-base constraints, Define a reward for the minimum cumulative throughput at each terminal. For the terminal fairness penalty item, Positive rewards are given to encourage terminals to achieve stricter fairness targets. This is a penalty item for terminal return error.

[0188] Step 4: The hybrid soft actor-critic algorithm is used to solve the finite-time Markov decision process established in Step 3 to obtain the joint optimization strategy for UAV FSO, RF link switching, power allocation and three-dimensional trajectory.

[0189] Furthermore, step 4 specifically includes:

[0190] A policy network is constructed that simultaneously outputs discrete and continuous actions. The policy network for discrete actions outputs discrete action branches, which are used to determine the FSO service object selection or link mode selection for the current time slot. The policy network for continuous actions outputs continuous action branches, which are used to jointly control the user's transmit power, the UAV's artificial interference power, the UAV's return optical power, and the UAV's three-dimensional displacement.

[0191] Within the maximum entropy reinforcement learning framework, entropy regularization terms are introduced for both discrete and continuous action branches. The policy networks for discrete and continuous actions are trained using an adaptive adjustment method based on learnable temperature parameters to automatically adjust the exploration intensity, thereby improving policy search capability and training stability in complex dynamic environments.

[0192] A dual-critic network is used to evaluate the joint action value of the policy network output. The training stability is improved by combining the experience replay mechanism and the target network update mechanism. This allows the agent to learn a joint optimization strategy for UAV FSO / RF link switching, power allocation and 3D trajectory under conditions of occlusion, backhaul bottleneck and fair constraints through continuous interaction and iteration with the environment.

[0193] Example 2:

[0194] This embodiment describes a computer-readable storage medium storing a computer program that, when executed by a processor, implements a UAV relay FSO and RF communication method based on hybrid reinforcement learning as described in any of Embodiment 1.

[0195] Example 3:

[0196] This embodiment describes a computer device, including:

[0197] Memory is used to store instructions.

[0198] A processor is configured to execute the instructions, causing the computer device to perform operations as described in any of Embodiment 1 of a UAV relay FSO / RF communication method based on hybrid reinforcement learning.

[0199] Example 4:

[0200] This embodiment introduces a simulation experiment of a UAV relay FSO and RF communication method based on hybrid reinforcement learning, wherein the simulation scenario is set as follows: The system is a three-dimensional area where multiple users are randomly distributed, with base stations located at the edges of the ground area. A UAV takes off from a preset starting position, provides communication services to multiple users within the area according to the mission plan, and returns to the starting point after the service duration is completed, thus fulfilling its mission.

[0201] In the simulation experiment, the location coordinates of User 1 and User 2 were set to [400,200,0] and [600,100,0] respectively, the location coordinates of the eavesdropper were set to [900,980,0], the base station coordinates were [610,350,0], and the starting and ending coordinates were [50,100,130].

[0202] The RF link bandwidth is 1MHz, the FSO uplink bandwidth is 50MHz, the FSO downlink bandwidth is 60MHz, and the RF channel noise power is... The maximum UAV interference power is 0.5W, the maximum user RF transmit power is 0.1W, and the maximum user FSO transmit power is 0.02W.

[0203] Figure 2 The proposed method's UAV trajectory planning results in a 3D urban scene are presented. It can be seen that after starting from the origin, the UAV first climbs to a higher airspace, then performs 3D maneuvers based on obstacle distribution, and finally returns to the vicinity of the destination. The entire trajectory does not cross building areas, indicating that the proposed method can satisfy constraints such as spatial boundaries, UAV flight altitude, and terminal return, obtaining a feasible flight path. Furthermore, this trajectory does not simply pursue the shortest flight distance, but exhibits clear task-driven characteristics. Under the hybrid SAC (Soft Actor-Critic) framework, UAV 3D position adjustment and FSO / RF switching decisions are performed jointly; therefore, trajectory changes not only serve flight accessibility but also improve communication performance. UAV maneuvering in a higher airspace helps mitigate the adverse effects of obstacle occlusion on legitimate links and establishes more stable transmission conditions for RF access and FSO backhaul, thereby improving the system's effective and secure transmission capabilities. Therefore, it is evident that… Figure 2 This demonstrates that the proposed method can achieve synergistic optimization of 3D trajectory, link switching, and secure communication performance in complex congested environments.

[0204] Figure 3 The changes in reward values ​​during the training phase of the proposed method and four comparative methods are presented. Method 1 is for the user... Using FSO uplink, users Method 2 uses RF uplink; Method 2 is for users Using RF uplink, users Method 3 involves two users simultaneously transmitting data to the UAV via an RF link; Method 4 involves using a fixed circular trajectory reference.

[0205] It can be seen that all methods experienced a rapid increase in performance during the initial training phase (approximately the first 100–200 episodes), followed by a gradual stabilization. This indicates that the constructed training environment and reward function design effectively guide policy learning and achieve convergence. Compared to the baseline methods, the HSAC (Hybrid SoftActor-Critic) curve achieves a higher reward level within a shorter training phase and remains optimal after convergence, with a plateau value stabilizing at around 200,000, significantly higher than the other comparative strategies. This demonstrates that the joint decision-making of "discrete switching + continuous power / track" has a significant advantage in long-term returns, exhibiting stronger long-term decision-making ability and training stability. From the comparison results, Method 4's performance is inferior to the method proposed in this invention, with a plateau value of approximately 170,000. This indicates that while the fixed trajectory scheme can guarantee a certain level of system feasibility, its performance ceiling is limited due to the lack of adaptive adjustment capability to environmental changes. Methods 1 and 2 show similar overall performance, converging to around 90,000. Method 3, however, has the lowest reward value, with its platform value stabilizing only around 70,000. This reflects that in hybrid link scenarios with backhaul bottlenecks and obstruction disturbances, fixed communication modes struggle to balance throughput, confidentiality, and fairness. It indicates that neither fixed service objects nor consistently using a single RF transmission mode can fully leverage the advantages of hybrid RF / FSO link collaborative optimization. In summary, the proposed method demonstrates superior performance in training convergence speed, post-convergence reward level, and overall stability, validating its effectiveness in complex dynamic scenarios.

[0206] Figure 4The changes in minimum secure transmission volume during the training phase of the proposed method and four comparative methods are presented. This indicator reflects the system's ability to protect weaker users; a higher value indicates better system fairness. As shown in the figure, each method exhibits some fluctuations in the early training phase, subsequently stabilizing. Compared to the comparative methods, the proposed method quickly reached a high performance level after initial exploration fluctuations, remaining stable in the range of approximately 108-113 Mbit / s, consistently maintaining the optimal performance and significantly exceeding all baseline methods, indicating its more effective improvement in transmission performance for weaker users. Among the baseline methods, method 4 performs worse than the proposed method, with a platform of approximately 79-83 Mbit / s, while methods 1, 2, and 3 are generally at a lower level, around 5-9 Mbit / s. This is mainly because, without switching strategies, one user typically relies on the RF link for long-term transmission. The RF link not only carries the risk of eavesdropping, limiting secure transmission capabilities, but its bandwidth is also lower than that of the FSO link, further restricting the user's transmission performance. In contrast, FSO links are more directional, less susceptible to eavesdropping, and more conducive to achieving higher levels of secure transmission. Therefore, fixed service targets or fixed transmission modes are more likely to become system bottlenecks. Overall, the method proposed in this invention can effectively improve the system's lower limit performance and user fairness protection capabilities by dynamically switching FSO / RF service modes and jointly optimizing control variables.

[0207] Figure 5The figure shows the changes in total secure transmission volume during the training phase of the proposed method and four comparative methods. As can be seen from the figure, each method exhibits some fluctuations in the early stages of training, then gradually stabilizes, indicating that different strategies can converge under the current environment. Compared to the comparative methods, the proposed method quickly reaches a higher level within a shorter training phase and maintains its optimal performance throughout the later stages of training, consistently remaining in the 225-250 Mbit / s range, demonstrating its significant advantage in improving the overall secure transmission capability of the system. This is mainly because the proposed method can dynamically switch between FSO / RF service modes based on link status and user needs, and jointly optimize control variables such as trajectory and power, thereby making fuller use of hybrid link resources. From the baseline methods, method 4 performs worse than the proposed method, with a platform speed of approximately 175-200 Mbit / s, indicating that while a fixed trajectory can maintain a certain level of transmission stability, its total secure transmission volume is still limited due to a lack of adaptive adjustment capability to environmental changes. Methods 1 and 2 showed similar overall performance, with the platform eventually stabilizing at approximately 150 Mbit / s. Method 3, however, exhibited the lowest performance, with the platform reaching approximately 100-120 Mbit / s. This is because fixed service targets or a single transmission mode lead some users to rely on RF links for extended periods. RF links are not only susceptible to eavesdropping, but this also hinders further improvements in the overall secure transmission capability of the system. In summary, the method proposed in this invention achieves the best performance in terms of total secure transmission volume by more effectively coordinating link resources through dynamic switching and joint optimization.

[0208] Figure 6The changes in the terminal fairness gap ratio during the training phase of the proposed method and four comparative methods are presented. This index measures the difference in the cumulative confidential transmission volume between the two users at the end of the round; the smaller the value, the better the terminal fairness of the system. As shown in the figure, each method exhibits some fluctuations in the early stages of training, which then gradually stabilize. Compared to the comparative methods, the proposed method rapidly drops to a lower level after a brief exploration, approximately in the range of 0.00-0.02, and remains optimal throughout the later stages of training, indicating that it can more effectively suppress the imbalance of benefits between the two users. Looking at the baseline methods, method 4 also maintains a low terminal fairness gap ratio, approximately in the range of 0.02-0.04, but is still slightly higher than the proposed method overall, indicating that while the fixed trajectory has a certain balancing ability, its adaptability to environmental changes and link states is limited. Methods 1 and 2 consistently show high fairness gap ratios, ranging from approximately 0.85 to 0.95. This is primarily because the two users correspond to FSO and RF links respectively. FSO links have greater bandwidth, and eavesdroppers can only operate on RF links, not FSO links. Therefore, the difference in transmission capacity and security performance between the two types of links leads to a long-term imbalance in the cumulative secure transmission volume for the two users. Method 3 also exhibits a high fairness gap ratio. This is mainly because, in this baseline, to maintain consistency with the mixed link scenario, the two users employ asymmetrical bandwidth configurations, resulting in a significant difference in achievable secure transmission capabilities even under full RF transmission conditions. In summary, the method proposed in this invention performs best in the terminal fairness gap ratio metric, demonstrating a more significant advantage in balancing system performance and user fairness.

[0209] Figure 7The changes in the secure transmission volume between two users during the training phase of the proposed method and four comparative methods are presented to compare the differences in performance allocation between users. As shown in the figure, the two user curves corresponding to the proposed method remain at a high level of approximately 110-120 Mbit / s in the later stages of training, and the gap between them is relatively small. This indicates that the proposed method can effectively balance the performance of the two users while improving the overall secure transmission capability of the system. This demonstrates that the proposed method can dynamically adjust the FSO / RF service mode according to link status and environmental changes, and jointly optimize relevant control variables, thereby avoiding long-term system resource bias towards one user. In contrast, methods 1 and 2 both exhibit obvious user bias characteristics, with one side of the user consistently at a higher level (approximately 140 Mbit / s), while the other side remains consistently lower (approximately 10 Mbit / s). The main reason is that the two users correspond to FSO and RF links respectively. The FSO link has a larger bandwidth, and the eavesdropper only operates on the RF link, unable to eavesdrop on the FSO link. Therefore, the difference in transmission capacity and security performance between the two types of links directly leads to a significant imbalance in the secure transmission volume between the two users. In Method 3, although the curves of the two users are relatively close, the overall level is low, indicating that its "balance" stems more from overall performance limitations than from effective collaborative optimization. Furthermore, to ensure consistency with the mixed link scenario, the two users in this baseline use asymmetrical bandwidth configurations, thus some differences still exist between them. Method 4 can maintain a relatively balanced result, but the overall level is still lower than the proposed method, indicating that while the fixed trajectory scheme has a certain degree of stability, it is difficult to simultaneously achieve high performance and high fairness in dynamic scenarios. In summary, the method proposed in this invention can achieve a better balance effect while both users obtain a high secure transmission volume, further verifying its effectiveness in the collaborative optimization of fairness and overall performance.

[0210] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0211] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0212] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0213] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0214] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A UAV relay FSO and RF communication method based on hybrid reinforcement learning, characterized in that: Specifically, it includes: Obtain the joint optimization problem of hybrid RF and FSO UAV relay; The joint optimization problem of hybrid RF and FSO UAV relay is formulated as a finite-time Markov decision process. The finite-time Markov decision process is solved to obtain a joint optimization strategy for UAV FSO, RF link switching, power allocation and three-dimensional trajectory. The aforementioned joint optimization problem for hybrid RF and FSO UAV relay specifically includes: ; Where P1 represents the joint optimization problem model of the UAV relay communication system, Indicates user The cumulative end-to-end effective confidentiality throughput throughout the entire mission cycle Let the UAV be in the three-dimensional position of time slot n. Let n be the three-dimensional position of the UAV in time slot n+1. This indicates the actual access status of the communication system in time slot n. Indicates user At the RF transmit power in time slot n, Indicates user The maximum RF transmit power, Indicates user The FSO transmitted optical power in time slot n, Indicates user The maximum transmit power of FSO, This indicates the power of the drone to interfere with the eavesdropper's actions. This indicates the FSO (Flight-Oriented Response) transmit optical power from the drone to the base station. This indicates the maximum electrical power limit of the drone; Let x be the x-axis coordinate of the drone. Let y be the y-axis coordinate of the drone. Let Z be the z-axis coordinate of the UAV. , These represent the upper bounds of the system service area along the x-axis and y-axis, respectively. , These represent the minimum and maximum permissible flight altitudes for the drone, respectively. Represents the maximum flight speed of the drone. This represents the three-dimensional position of the drone in time slot 1. This represents the three-dimensional position of the UAV in time slot N. Represents the starting point of drones, Representing the endpoint of the drone, in C6 On behalf of users Using FSO uplink, users Using RF uplink, in C6 On behalf of users Using FSO uplink, users Using RF uplink, in C6 This indicates that the system is in an FSO switchover period, and FSO uplink is temporarily unavailable. Both users are transmitting uplink data via the RF link. Represents photoelectric conversion efficiency. Represents the maximum power of the UAV. This represents the length of each time slot n.

2. The UAV relay FSO and RF communication method based on hybrid reinforcement learning according to claim 1, characterized in that: The expression for the finite-time Markov decision process is as follows: ; in, For a finite-time Markov decision process, For state space, For a mixed action space, Represents the state transition probability. For the instant reward of time slot n, As a discount factor, This represents the state of time slot n+1; in, The state of time slot n The expression is as follows: ; in, Used to characterize the current geometric position, link mode, and key distance information of the UAV. Used to characterize the fairness of cumulative end-to-end effective confidentiality throughput for dual users. Used to characterize the return-to-home reachability of the drone in the current state. Used to introduce the actual execution displacement information of the previous time slot; Action of time slot n The expression is as follows: ; in, This represents the discrete action of the nth time slot, used to determine the steady-state FSO service object selection for the current time slot. It represents the continuous action of the nth time slot, used to jointly control the user's transmit power, UAV artificial interference power, UAV return optical power and UAV's three-dimensional displacement; The expression is as follows: ; in, As a feasibility penalty item, For indicator functions, This represents the terminal convergence interval, which is only when... , Otherwise, it is 0. For the sake of fairness, Due to return-to-base constraints, The indicator function representing the Nth time slot. Define a reward for the minimum cumulative throughput at each terminal. For the terminal fairness penalty item, Positive rewards are given to encourage terminals to achieve stricter fairness targets. This is a penalty item for terminal return error.

3. The UAV relay FSO and RF communication method based on hybrid reinforcement learning according to claim 1, characterized in that: The finite-time Markov decision process is solved using a hybrid soft actor-critic algorithm.

4. The UAV relay FSO and RF communication method based on hybrid reinforcement learning according to claim 1, characterized in that: The The expression is as follows: ; in, This represents the proportion of data that can be transmitted back to the source. Represented as user The instantaneous security rate in time slot n, where N represents the number of time slots.

5. The UAV relay FSO and RF communication method based on hybrid reinforcement learning according to claim 4, characterized in that: The The expression is as follows: ; In the formula, For numerically stable terms, This represents the uplink transmission rate of the two users in time slot n. This indicates the transmission rate from the drone to the base station; Among them, for Specifically, they can be categorized into the following situations: In normal working mode, when At that time, the user Using FSO uplink, users Using RF uplink, users The RF transmission rate to the drone is expressed as: ; in, The bandwidth used for the RF channel. For users At the RF transmit power in time slot n, For users To the RF channel gain of the drone, This represents the noise power of the RF link. user The FSO transmission rate to the drone is expressed as: ; in, For FSO uplink bandwidth, Indicates user SNR of the FSO link to the drone; when At that time, the user Using FSO uplink, users Using RF uplink, users The RF transmission rate to the drone is expressed as: ; in, For users At the RF transmit power in time slot n, For users RF channel gain to the drone; user The FSO transmission rate to the drone is: ; in, Indicates user SNR of the FSO link to the drone; In switching modes, when Two users simultaneously transmit uplink to the UAV via the RF link; correspondingly, the user with the higher received power is selected based on the received power from highest to lowest. and smaller users The transmission rates at the UAV end are as follows: ; in, and Represented as users and users Signal-to-interference-plus-noise ratio (SIR) of non-orthogonal multiple access communication.

6. The UAV relay FSO and RF communication method based on hybrid reinforcement learning according to claim 5, characterized in that: The The expression is as follows: ; in, , For eavesdroppers to target current RF uplink users The rate of eavesdropping; in, The expression is as follows: ; when or hour, Indicates a user using RF uplink Its equivalent signal-to-dryness ratio at the eavesdropping end is expressed as: ; in, This is represented as the equivalent power noise at the receiver's end. Indicates the user in the current time slot RF transmit power, The artificial interference power of the eavesdropper. Indicates user To the RF channel gain of the eavesdropper, This represents the channel gain of the interference link from the UAV to the eavesdropper within time slot n; when At this time, the FSO uplink is temporarily unavailable, and two users simultaneously transmit uplink data via the RF link; in this situation, an eavesdropper will simultaneously receive RF signals from both users; correspondingly, They are respectively , , Indicates to the user eavesdropping rate, Indicates to the user The eavesdropping rate, sorted by received power from highest to lowest, for users with higher power. and smaller users The eavesdropping rates are as follows: ; in, This indicates that the eavesdropper is decoding the user. At that time, the user The signal is considered interference, and the corresponding signal-to-interference-plus-noise ratio (SINNR) is... This indicates that the user has been successfully eliminated. After receiving the signal, the eavesdropper further decoded the user's signal. The corresponding signal-to-interference-plus-noise ratio at this time; ; in, and These represent users within the current time slot. and RF transmit power, and Each represents the corresponding user , RF channel gain for drones.

7. The UAV relay FSO and RF communication method based on hybrid reinforcement learning according to claim 2, characterized in that: The The expression is: ; in, , and Let represent the normalized three-dimensional position features of the UAV in the nth time slot. This represents the normalized time slot progress characteristics. This represents the normalized discrete link mode characteristics. This represents the normalized mode switching countdown feature. This represents the normalized switching coolant count characteristic. , and These represent the normalized distance characteristics from the user to the drone, from the drone to the base station, and from the drone to the eavesdropper, respectively. The The expression is: ; in, Represents the normalized user Cumulative end-to-end secure throughput up to time slot n Represents the normalized user Cumulative end-to-end secure throughput up to time slot n Indicates the proportion of fairness gap. The lagging user index variable is used to assist discrete pattern decision-making in prioritizing the identification of users with low current cumulative end-to-end effective confidentiality throughput; The The expression is: ; in, This represents the normalized distance characteristic from the UAV to the target endpoint. This represents the reachability margin feature of the normalized UAV in the current remaining time domain to reach the target endpoint; The The expression is: ; in, , and They are respectively represented as the normalized values ​​at the th... The actual displacement components of the UAV in the x-axis, y-axis and z-axis directions within each time slot.

8. A computer-readable storage medium, characterized in that: It stores a computer program, which, when executed by a processor, implements a UAV relay FSO and RF communication method based on hybrid reinforcement learning as described in any one of claims 1 to 7.

9. A computer device, characterized in that: include: Memory, used to store instructions; A processor is configured to execute the instructions, causing the computer device to perform the operation of a UAV relay FSO / RF communication method based on hybrid reinforcement learning as described in any one of claims 1 to 7.