A method for planning UAV relay paths based on hybrid optoelectronic switching
By optimizing UAV relay trajectory planning through hybrid optoelectronic switching and deep reinforcement learning, the communication stability and throughput issues of traditional UAV relay systems in low-altitude spectrum denial and complex obstacle environments are solved, achieving efficient communication in dynamic environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
- Filing Date
- 2025-03-06
- Publication Date
- 2026-06-30
AI Technical Summary
Traditional UAV relay communication systems struggle to achieve both improved communication stability and system throughput in low-altitude spectrum denial and complex obstacle environments. In particular, in dynamic spectrum denial and complex obstacle environments, existing methods cannot adjust the flight path in real time to avoid obstacles or optimize channel conditions, resulting in a sharp drop in communication performance.
A UAV relay trajectory planning method based on hybrid optoelectronic switching is adopted. Combined with a deep reinforcement learning framework, the DDQN algorithm is used to optimize UAV trajectory planning, optoelectronic switching strategy and user power allocation. A free space light and radio frequency channel model is established, and a hybrid obstacle and radio frequency denial model is constructed to achieve the coordinated optimization of dynamic resource allocation and intelligent trajectory planning.
In spectrum denial environments, it ensures link stability, improves multi-user access efficiency, generates efficient flight paths to avoid obstacles and adapt to dynamic interference, significantly improves communication performance and reliability, and provides technical support for low-altitude intelligent networking.
Smart Images

Figure CN120110492B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of UAV communication technology, specifically relating to a UAV relay trajectory planning method based on optoelectronic hybrid switching for low-altitude spectrum denial environments. Background Technology
[0002] With the rapid development of low-altitude intelligent networks, their core role in sixth-generation communication technology is becoming increasingly prominent, providing ubiquitous air-space connectivity for emerging services such as intelligent sensing and unmanned swarm collaboration. As a key carrier of low-altitude intelligent networks, the large-scale deployment of UAVs places stringent demands on communication systems for ultra-high capacity and ultra-low latency. However, traditional radio frequency communication is limited by scarce spectrum resources and frequent electromagnetic interference, making it difficult to support reliable interconnection of densely populated low-altitude nodes. Free-space optical communication, with its large bandwidth, unlicensed spectrum, and strong resistance to electromagnetic interference, has become an ideal supplementary solution for alleviating low-altitude spectrum congestion, especially suitable for high-speed data transmission between UAV swarms.
[0003] However, the practical application of free-space optical technology still faces significant challenges: on the one hand, atmospheric turbulence and complex weather conditions can cause severe fluctuations in optical signals, especially in long-distance transmission scenarios, where link stability drops sharply; on the other hand, densely distributed buildings, dynamic aircraft, and other obstacles in low-altitude airspace frequently block line-of-sight paths, making it difficult for pure free-space optical communication systems to maintain continuous coverage. Existing technologies can shorten single-hop transmission distance and reduce atmospheric attenuation by deploying UAV relay nodes, but traditional relay schemes often use fixed communication modes (free-space optical or radio frequency only), exhibiting significant shortcomings in environments with dynamic spectrum denial (such as sudden radio frequency interference) and complex obstacles: 1) A single free-space optical mode is prone to communication interruptions in line-of-sight obstructed or turbulent scenarios, lacking redundant link guarantees; 2) A single radio frequency mode is limited by spectrum resource competition and denial interference, failing to meet the demands of high-capacity transmission; 3) Static trajectory planning is not deeply coupled with dynamic switching of communication modes and power allocation, resulting in low resource utilization and limited system throughput.
[0004] Furthermore, existing UAV relay systems mostly employ heuristic algorithms or independent optimization strategies, which struggle to cope with the highly dynamic nature of low-altitude environments. For example, when sudden radio frequency interference forces a switch in communication mode, traditional methods cannot adjust the flight path in real time to avoid obstacles or optimize channel conditions, resulting in a sharp drop in communication performance. Therefore, there is an urgent need for a collaborative optimization method that integrates optoelectronic hybrid switching, dynamic resource allocation, and intelligent flight path planning to achieve a dual improvement in communication stability and system throughput in spectrum denial and complex obstacle environments, providing reliable technical support for the large-scale application of low-altitude intelligent networks. Summary of the Invention
[0005] Purpose of the invention: This invention addresses the problems existing in the prior art by proposing a UAV relay trajectory planning method based on optoelectronic hybrid switching. This invention can effectively cope with the challenges of low-altitude spectrum denial environments and ensure the robustness and adaptability of the system.
[0006] Technical solution: The UAV relay trajectory planning method based on photoelectric hybrid switching described in this invention specifically includes the following steps:
[0007] (1) Establish a hybrid photoelectric relay communication system for UAVs in low-altitude spectrum denial and complex obstacle environments, the system including a UAV relay equipped with free-space optical and radio frequency dual-mode communication equipment, a ground base station and multiple mobile users;
[0008] (2) Establish a free space optical channel model, considering the two main channel influencing factors of the free space optical link, namely the scintillation effect caused by turbulence and the turbulence attenuation caused by adverse weather conditions;
[0009] (3) Establish a radio frequency channel model that considers both line-of-sight and non-line-of-sight channels in large-scale fading, and considers the interference between multiple users in the downlink relay from UAV to user;
[0010] (4) Establish an energy consumption model for the UAV, considering only the propulsion energy required for the UAV to overcome resistance and support its own weight;
[0011] (5) Establish a low-altitude airspace hybrid obstacle model and calculate the feasible location set of UAVs that can achieve line-of-sight connection according to the Ray-AABB algorithm; the hybrid obstacle model includes ground fixed obstacles and time-varying aerial obstacles to simulate the real low-altitude airspace environment.
[0012] (6) Construct a time-dynamically changing radio frequency rejection interference model to demonstrate the feasibility of disrupting the radio frequency communication link at the transmitting end through random time-varying interference signals.
[0013] (7) Establish a user random walk movement model. After each movement, the movement node updates its speed and direction at a constant time interval or constant movement distance. If the movement node reaches the boundary of the simulation area, it will "bounce back" along the reflection angle of the incident direction and continue to move along the new path.
[0014] (8) Construct a joint optimization problem for UAV trajectory planning, photoelectric switching strategy and user power allocation;
[0015] (9) Based on the deep reinforcement learning framework, the DDQN algorithm is used to jointly optimize the UAV trajectory planning, photoelectric switching strategy and user power allocation.
[0016] Furthermore, the free-space optical channel model described in step (2) is as follows:
[0017] Average gain g of optical power at the receiver o for:
[0018]
[0019] Wherein, the first term represents the geometric loss due to beam divergence, and the second term represents the weather-related atmospheric loss caused by scattering and absorption; d is the receiver aperture diameter, φ is the beam divergence angle, and κ is the weather-related attenuation coefficient determined according to the Beer-Lambert law, the expression of which is:
[0020]
[0021] Where Vis represents the visibility range, λ is the optical wavelength, and p is the Mie scattering coefficient determined by the Kim model;
[0022] The intensity function under different turbulent conditions is described using the Gamma-Gamma distribution:
[0023]
[0024] Where Γ(·) represents the Gamma function, K p (·) represents the modified Bessel function of the second kind; α and β are the effective values of small-scale and large-scale eddies of the scattering environment related to atmospheric conditions, respectively, given by the following formulas:
[0025]
[0026] in, k is the refractive index structural parameter. o Optical wavenumber;
[0027] For a free-space optical channel using intensity modulation direct detection, the achievable rate under constraints of random channel gain and average optical transmission power is:
[0028]
[0029] Where e is the base of the natural logarithm, B FSO It is the bandwidth of the free-space optical link, P FSO σ represents optical transmission power. 2 This represents the noise variance.
[0030] Furthermore, the radio frequency channel model described in step (3) is as follows:
[0031] The free space path loss between the base station and the drone relay, and between the drone relay and the k-th user, is expressed as:
[0032]
[0033] Among them, L ω [n] represents the distance between the base station and the drone relay or between the drone relay and the k-th user, f c It is the carrier frequency, v c It's the speed of light;
[0034] The expression for the line-of-sight and non-line-of-sight path loss between the drone relay to the k-th user is:
[0035]
[0036] Where, η LOS and η NLOS These are the additional path losses for line-of-sight and non-line-of-sight links, respectively.
[0037] The probability model for the occurrence of a line-of-sight channel between the drone and the base station or the k-th user is as follows:
[0038]
[0039] Where, η a and η b It is a coefficient related to the operating environment, where θ is the elevation angle between the UAV and the user or base station; the probability of non-line-of-sight channel occurrence is expressed as ρ. NLOS,ω [n] = 1 - ρ LOS,ω [n];
[0040] The path loss between the base station and the drone relay, and between the drone relay and the k-th user, is expressed as follows:
[0041] PL ω [n] = ρ LOS,ω [n]×PL LOS,ω [n]+ρ NLOS,ω [n]×PL NLOS,ω [n]
[0042] =PL FS,ω [n]+ρ LOS,ω [n]η LOS +(1-ρ LOS,ω [n])η NLOS ,
[0043]
[0044] Therefore, the path loss factor between the UAV and the k-th user at time slot n is:
[0045]
[0046] Therefore, the radio frequency transmission rate from the base station to the drone relay is:
[0047]
[0048] Among them, B RF P represents the bandwidth of the radio frequency link. RF [n] is the radio frequency communication power, and n0 is the noise power spectral density;
[0049] The downlink transmission rate between the UAV relay and user k is expressed as:
[0050]
[0051] The total capacity of the system for multiple users is:
[0052]
[0053] Furthermore, the specific UAV energy consumption model described in step (4) is as follows:
[0054] The propulsion power consumption of a rotary-wing UAV is defined as:
[0055]
[0056] Where v represents the speed of the UAV, v0 is the average rotor-induced speed during hovering, and U tip Let P0 represent the blade tip velocity and P0 represent the blade power. The expression is:
[0057]
[0058] Where ζ represents the drag coefficient, ρ represents the air density, and s rotor Indicates rotor robustness, A represents turntable area, Ω rotor It is the blade angular velocity, R rotor It is the blade radius;
[0059] P i The induced power at zero velocity is expressed as:
[0060]
[0061] Where, k e W represents the incremental correction factor for induced power. uav Indicates the weight of the drone
[0062] The energy consumption of the system within time T is:
[0063]
[0064] Furthermore, the low-altitude airspace mixed obstacle model described in step (5) is as follows:
[0065] Q = Q obstacle [n]+Q NLOS [n]+Q available [n]
[0066] In the formula, Q represents the total space within the scene. obstacle [n] represents the total space occupied by all obstacles in time slot n, including fixed obstacles and time-varying obstacles. Ground obstacles are modeled based on the density and average height of urban buildings, and the proportion of space occupied by moving aerial obstacles is set; Q NLOS [n] represents the position where time slot n is not occupied by obstacles but a line-of-sight connection cannot be achieved, Q available [n] represents the path points that a UAV can choose in time slot n.
[0067] Furthermore, the radio frequency rejection interference model described in step (6) is as follows:
[0068] The feasibility of disrupting the transmitting end's radio frequency communication link through random time-varying interference signals; assuming that the state of whether the transmitting end base station is subject to radio frequency interference in each time slot n can be expressed by the following formula:
[0069]
[0070] Among them, RF fb [n] = 1 indicates that there is radio frequency interference at the base station in time slot n, which prevents the radio frequency communication link from being established. Conversely, RF fb If [n] = 0, it means that the radio frequency link can be used normally within this time slot.
[0071] Furthermore, the user random walk movement model described in step (7) is specifically as follows:
[0072] The user's direction of movement is determined by random angles uniformly distributed between [0, 2π), while the speed of movement ranges from a predefined range of [0, U]. max Randomly assigned in ], U max This represents the user's maximum movement speed; after each movement, the moving node updates its speed and direction at constant time intervals or constant movement distances; if the moving node reaches the boundary of the simulation area, it will "bounce back" along the reflection angle of the incident direction and continue moving along the new path.
[0073] Furthermore, the implementation process of step (8) is as follows:
[0074]
[0075] Among them, constraint C1 specifies the starting and returning positions of the UAV and restricts its trajectory points to a specified three-dimensional mission area; constraint C2 limits the total energy consumption of the UAV to ensure that it completes the mission within a limited energy budget; constraint C3 specifies the choice of communication method, using a binary variable w to specify whether the uplink in each time slot uses free space light or radio frequency communication; constraint C4 constrains the maximum power that the UAV can provide for downlink relay communication, ensuring that the total power allocated to all service users is less than the maximum power that the UAV relay can provide; constraint C5 limits the total throughput of users to not exceed the maximum capacity supported by the uplink through information causality constraints; constraint C6 ensures that the optoelectronic hybrid relay can provide communication coverage to all users in each time slot; constraint C7 stipulates that when using free space light for uplink communication, the location of the UAV must be able to achieve line-of-sight connection with the transmitting end; constraint C8 stipulates that when using radio frequency for uplink communication, the transmitting base station in that time slot must be free from radio frequency rejection interference.
[0076] Furthermore, the user power allocation process described in step (9) is as follows:
[0077] Power control is performed based on the user's instantaneous channel state information to compensate for unfairness caused by differences in channel conditions among users. The allocation strategy is as follows:
[0078]
[0079] Where, α ftpa (0≤α ftpa ≤1) represents the power allocation factor, g ftpa k Represents the channel gain from base station to user k, n k ftpa This represents the Gaussian white noise in the channel corresponding to user k; when α ftpa When α = 0, it indicates equal power distribution among users. ftpa The smaller the value, the less power is allocated to users with good channel conditions;
[0080] The initial transmit power is set to the maximum available power of the relay node. Based on the current transmit power, FTPA is used to calculate the power allocation for each user. It is verified whether the total throughput of the system satisfies the information causality constraint C2 under the current power allocation. If the current power allocation scheme cannot meet the constraint, the transmit power is adjusted using a bisection method to gradually approach the optimal solution that meets the condition. At the same time, attention is paid to the communication coverage problem of ground users. During the process of adjusting the transmit power, it is necessary to ensure the communication connection with ground users. An iterative optimization strategy is introduced. By repeatedly adjusting and approximating the transmit power and power allocation, the throughput and communication quality of the system are optimized while ensuring the information causality constraint.
[0081] Furthermore, the specific implementation process of step (9) is as follows:
[0082] The joint optimization problem of UAV trajectory planning and electro-optical switching strategy is modeled as a Markov decision process. By reasonably defining the state space, action space and reward function, the problem is transformed into a solution task under the reinforcement learning framework. The DDQN algorithm is adopted to continuously interact with the environment and learn, and gradually optimize the UAV trajectory planning and electro-optical switching strategy.
[0083] The state includes the drone's current position p. uav The location of the base station p BS User's location p users The speed of the drone, v uav Remaining energy of drones uav Ground obstacle information ground Information on aerial obstacles air Determining the line-of-sight distance (m) of the drone's location l and collision marker m c Atmospheric visibility data f vis The state space S[n] is specifically described as follows:
[0084] S[n]={p uav [n],p BS ,p k users [n],v uav [n],e uav [n],o ground ,o air [n],m l [n],m c [n],f vis [n]}
[0085] The action space A[n] of the UAV is the Cartesian product of the movement and communication actions:
[0086] A[n]=A move [n]×A com [n]
[0087] A move [n]={(x,y,z)∣x,y,z∈{-1,0,1}}
[0088] A com [n]∈{0,1}
[0089] Among them, A com [n] = 0 indicates that radio frequency communication is selected, A com [n] = 1 indicates that free-space optical communication is selected;
[0090] A complete action of a drone can be represented as a binary tuple:
[0091] a[n]=(a move [n],a com [n])
[0092] The reward function is set as follows:
[0093] The primary reward for slot n is the sum of the throughput of k users, denoted as:
[0094]
[0095] in, It is a normalization coefficient used to ensure that throughput rewards and penalties are set to the same order of magnitude.
[0096] When there are numerous obstacles in the mission area, a penalty and reward mechanism is used to constrain the drone's misbehavior. The specific settings are as follows:
[0097] R col [n]=κ col [n]·r col
[0098] R out [n]=κ out [n]·r out
[0099] R nlos [n]=κ nlos [n]·r nlos
[0100] Among them, κ col [n] is a binary collision constraint index used to guide the drone to avoid obstacles, κ col [n] = 1 indicates that the drone engaged in collision behavior in that time slot, κ col [n] = 0 indicates that the drone is flying normally and safely; κ out [n] and κ nlos [n] are binary variables representing whether the drone has crossed the boundary and whether it is a non-line-of-sight connection, respectively, which are used to encourage the drone to fly within the specified mission area and to prioritize line-of-sight links for transmission.
[0101] The total reward obtained by the agent in time slot n is expressed as:
[0102]
[0103] Beneficial Effects: Compared with existing technologies, the beneficial effects of this invention are as follows: This invention achieves dynamic switching between free-space optical and radio frequency communication in spectrum-denied environments by jointly optimizing trajectory, communication mode, and power allocation through deep reinforcement learning, ensuring link stability; it improves multi-user access efficiency by combining NOMA technology; and it generates efficient trajectories that avoid obstacles and adapt to dynamic interference through the DDQN algorithm. This invention can effectively cope with the challenges of low-altitude spectrum-denied environments, ensuring the robustness and adaptability of the system, while significantly improving communication performance and reliability under limited resource constraints, providing important technical support for the practical application of UAV optoelectronic hybrid relay communication. Attached Figure Description
[0104] Figure 1 This is a flowchart of the present invention;
[0105] Figure 2 This is a flowchart of the DDQN algorithm update proposed in this invention;
[0106] Figure 3 A relay flight path map of a hybrid electro-optical UAV when users are distributed in the corners of the mission area;
[0107] Figure 4 A relay flight path map of an electro-optical hybrid UAV when users are distributed in the center of the mission area;
[0108] Figure 5 A relay flight path map of an electro-optical hybrid UAV under mixed obstacle distribution;
[0109] Figure 6 This is a flight path map of a hybrid electro-optical UAV relay system with only ground obstacles present.
[0110] Figure 7 This represents the total user throughput and handover status under radio frequency interference.
[0111] Figure 8 This represents the total system throughput under different interference environments.
[0112] Figure 9 The throughput for users in the optoelectronic hybrid relay communication system. Detailed Implementation
[0113] The present invention will now be described in further detail with reference to the accompanying drawings.
[0114] like Figure 1 As shown, this invention proposes a UAV relay trajectory planning method based on electro-optical hybrid switching for spectrum denial, establishing a UAV electro-optical hybrid relay communication system in low-altitude spectrum denial and complex obstacle environments, including a UAV relay equipped with free-space optical and radio frequency dual-mode communication equipment, a ground base station, and multiple mobile users; specifically including the following steps:
[0115] Step S1: Establish a free-space optical communication channel model for UAV relay: By using adaptive pointing, acquisition and tracking coordination, the misalignment fading problem can be compensated. Only the two main channel influencing factors of the free-space optical link are considered, namely the scintillation effect caused by turbulence and the turbulence attenuation caused by adverse weather conditions.
[0116] Average gain g of optical power at the receiver o It can be represented as:
[0117]
[0118] The first term represents the geometric loss due to beam divergence, and the second term represents the weather-related atmospheric loss caused by scattering and absorption. Specifically, d is the receiver aperture diameter, φ is the beam divergence angle, and κ is the weather-related attenuation coefficient determined according to the Beer-Lambert law, expressed as:
[0119]
[0120] Where Vis represents the visibility range, λ is the optical wavelength, and p is the Mie scattering coefficient determined by the Kim model.
[0121] The intensity function under different turbulent conditions is described using the Gamma-Gamma distribution, and its expression is:
[0122]
[0123] Where Γ(·) represents the Gamma function, K p (·) represents the modified Bessel function of the second kind. α and β are the effective values of small-scale and large-scale eddies in the scattering environment related to atmospheric conditions, given by the following formulas:
[0124]
[0125] in, k is the refractive index structural parameter. o It is the optical wavenumber.
[0126] For a free-space optical channel using intensity modulation direct detection, its achievable rate (lower bound of channel capacity) under constraints of random channel gain and average optical transmission power can be expressed as:
[0127]
[0128] Where e is the base of the natural logarithm, B FSO It is the bandwidth of the free-space optical link, P FSO σ represents optical transmission power.2 This represents the noise variance.
[0129] Step S2: Establish the radio frequency communication channel model for UAV relay: both line-of-sight and non-line-of-sight channels are considered in large-scale fading, and interference between multiple users is considered in the downlink from UAV relay to the user.
[0130] The free space path loss between the base station and the drone relay, and between the drone relay and the k-th user, can be expressed as:
[0131]
[0132] Among them, L ω [n] represents the distance between the base station and the drone relay or between the drone relay and the k-th user, f c It is the carrier frequency, v c It's the speed of light.
[0133] The expression for the line-of-sight and non-line-of-sight path loss between the drone relay to the k-th user is:
[0134]
[0135] Where, η LOS and η NLOS These represent the additional path loss for line-of-sight and non-line-of-sight links, respectively.
[0136] The probability of a line-of-sight channel occurring between a drone and a base station or the k-th user can be modeled as follows:
[0137]
[0138] Where, η a and η b This is a coefficient related to the operating environment, where θ is the elevation angle between the drone and the user or base station. The probability of a non-line-of-sight channel occurring can be expressed as ρ. NLOS,ω [n] = 1 - ρ LOS,× [n].
[0139] Therefore, the path loss between the base station and the drone relay, and between the drone relay and the k-th user, can be expressed as:
[0140] PL × [n] = ρ LOS,ω [n]×PL LOS,ω [n]+ρ NLOS,ω [n]×PL NLOS,ω [n]
[0141] =PL FS,ω [n]+ρ LOS,ω [n]η LOS +(1-ρLOS,ω [n])η NLOS
[0142]
[0143] Therefore, the path loss factor between the UAV and the k-th user at time slot n can be obtained as follows:
[0144]
[0145] Therefore, the radio frequency transmission rate from the base station to the drone relay can be expressed as:
[0146]
[0147] Among them, B RF P represents the bandwidth of the radio frequency link. RF [n] is the radio frequency communication power, and n0 is the noise power spectral density.
[0148] In the downlink from UAV relay to user, interference between multiple users must be considered when calculating RF communication throughput. NOMA allows multiple users to share the same resource block. Through differential allocation in the power domain, users with better channel conditions can eliminate interference caused by users with poorer conditions, thereby improving system throughput and efficiency. Therefore, the downlink transmission rate between the UAV relay and user k can be expressed as:
[0149]
[0150] Therefore, the total capacity of the system for multiple users is:
[0151]
[0152] Step S3: Establish the UAV energy consumption model: Taking the rotary-wing UAV as the research object, and without considering communication energy consumption, we only focus on the propulsion energy required for the UAV to overcome drag and support its own weight.
[0153] The propulsion power consumption of a rotary-wing UAV is defined as:
[0154]
[0155] Where v represents the speed of the UAV, v0 is the average rotor-induced speed during hovering, and U tip Let P0 represent the blade tip velocity and P0 represent the blade power. The expression is:
[0156]
[0157] in, The drag coefficient is represented by ρ, and the air density is represented by s. rotorIndicates rotor robustness, A represents turntable area, Ω rotor It is the blade angular velocity, R rotor That is the blade radius.
[0158] P i The induced power at zero velocity is expressed as:
[0159]
[0160] Where, k e W represents the incremental correction factor for induced power. uav This indicates the weight of the drone.
[0161] Combining the above formula and the drone's trajectory, the system's energy consumption within time T can be expressed as:
[0162]
[0163] Step S4: Establish a low-altitude airspace hybrid obstacle model: Calculate the feasible location set of UAVs that can achieve line-of-sight connection based on the Ray-AABB algorithm. The hybrid obstacle model includes fixed ground obstacles and time-varying aerial obstacles to simulate the real low-altitude airspace environment.
[0164] The initial environment space is specifically represented as follows:
[0165]
[0166] Where Q represents the total space within the scene, Q obstacle [n] represents the total space occupied by all obstacles in time slot n, including fixed obstacles and time-varying obstacles. Ground obstacles are modeled based on the density and average height of city buildings, and the proportion of space occupied by moving aerial obstacles is set; Q NLOS [n] represents the position where time slot n is not occupied by obstacles but a line-of-sight connection cannot be achieved, calculated by the Ray-AABB algorithm; Q available [n] represents the path points that an n-slot UAV can choose; the relationship between them is Q = Q obstacle [n]+Q NLOS [n]+Q available [n].
[0167] Step S5: Construct a time-dynamically changing radio frequency rejection interference model to demonstrate the feasibility of disrupting the transmitter's radio frequency communication link through random time-varying interference signals.
[0168] This invention addresses the feasibility of disrupting the transmitting end's radio frequency (RF) communication link through random, time-varying interference signals. Unlike the spatial constraints imposed by low-altitude obstacles, RF denial-of-service interference is based on dynamic time-varying changes. It is assumed that the state of whether the transmitting base station is subject to RF interference in each time slot n can be represented by the following formula:
[0169]
[0170] Among them, RF fb [n] = 1 indicates that there is radio frequency interference at the base station in time slot n, which prevents the radio frequency communication link from being established. Conversely, RF fb If [n] = 0, it means that the radio frequency link can be used normally within this time slot.
[0171] Step S6: Establish a random walk model for users: After each move, the moving node updates its speed and direction at a constant time interval or a constant moving distance. If the moving node reaches the boundary of the simulation area, it will "bounce back" along the reflection angle of the incident direction and continue to move along the new path.
[0172] To accurately describe the movement characteristics of ground users, a random walk model is used to model their movement behavior. The random walk model was first proposed mathematically by Einstein in 1926 and is widely used to describe entities in nature with uncertain movement characteristics. In this model, the user's direction of movement is determined by a random angle uniformly distributed between [0, 2π), and the speed of movement ranges from [0, U...]. max Randomly assigned in ], where U max This represents the user's maximum movement speed. The moving node updates its speed and direction at constant time intervals or by a constant distance after each movement. If the moving node reaches the boundary of the simulation area, it will "bounce back" along the reflection angle of the incident direction and continue moving along a new path.
[0173] Step S7: Construct a joint design problem for UAV trajectory planning, electro-optical switching strategy, and user power allocation.
[0174] The problem is stated as follows:
[0175]
[0176] Constraint C1 specifies the UAV's starting and returning positions, limiting its trajectory points to a designated three-dimensional mission area. Constraint C2 limits the UAV's total energy consumption, ensuring it completes the mission within a limited energy budget. Constraint C3 specifies the communication method, using a binary variable `w` to indicate whether the uplink uses free-space light or radio frequency (RF) communication in each time slot. Constraint C4 constrains the maximum power that the UAV can provide for downlink relay communication, ensuring the total power allocated to all service users is less than the maximum power the UAV relay can provide. Constraint C5 uses information causality constraints to limit the total throughput of users to no more than the maximum capacity supported by the uplink. Constraint C6 ensures that the hybrid optoelectronic relay provides communication coverage to all users in each time slot. Constraint C7 requires that when using free-space light for uplink communication, the UAV's location must maintain a line-of-sight connection with the transmitting end. Constraint C8 requires that when using RF for uplink communication, the transmitting base station in that time slot must not be affected by RF denial interference.
[0177] Step S8: Create a user power allocation method.
[0178] The FTPA-based inverse information-constrained power allocation algorithm aims to achieve a dynamic balance between system throughput and communication coverage performance while satisfying information causality constraints. FTPA uses user channel gain as a reference and flexibly allocates power through fractional-order relationships. By adjusting the fractional-order parameters, it can achieve a dynamic trade-off between spectral efficiency and user fairness, while significantly reducing computational complexity. Specifically, the fractional-order power allocation strategy controls power based on the instantaneous channel state information of users, compensating for unfairness caused by differences in channel conditions among users, while ensuring the computational efficiency of the algorithm. Its allocation strategy can be expressed as follows:
[0179]
[0180] Where, α ftpa (0≤α ftpa ≤1) represents the power allocation factor, g ftpa k Represents the channel gain from base station to user k, n k ftpa This represents the Gaussian white noise in the channel corresponding to user k. When α... ftpa When α = 0, it indicates equal power distribution among users. ftpa The smaller the value, the less power is allocated to users with good channel conditions. Compared to traditional methods, FTPA is more adaptable to scenarios with dynamically changing resources, and its interference suppression effect is particularly significant for users with weak channels, which helps to improve overall communication quality and optimize system performance.
[0181] The basic process of the FTPA-based reverse information-constrained power allocation algorithm is as follows: First, the transmit power is initialized to the maximum available power of the relay node. Based on the current transmit power, FTPA is used to calculate the power allocation for each user. Next, it is verified whether the total throughput of the system satisfies the information causality constraint C2 under the current power allocation. If the current power allocation scheme does not meet the constraint, the transmit power is adjusted using a bisection method to gradually approach the optimal solution that satisfies the condition. During this process, the algorithm must pay special attention to the communication coverage problem of ground users. While adjusting the transmit power, it is necessary to ensure the communication connection with ground users. In addition, to improve the computational efficiency and stability of the algorithm, an iterative optimization strategy is introduced. By repeatedly adjusting and approximating the transmit power and power allocation, the system throughput and communication quality are optimized while ensuring the information causality constraint.
[0182] Step S9: Invent a method for obtaining the location of UAVs with the highest system throughput.
[0183] The joint optimization problem of UAV trajectory planning and electro-optical switching strategy is modeled as a Markov decision process. By reasonably defining the state space, action space, and reward function, the problem is transformed into a solution task within a reinforcement learning framework. To efficiently solve this complex problem, this invention employs the DDQN algorithm, which continuously interacts with and learns from the environment to progressively optimize the UAV's trajectory planning and electro-optical switching strategy. The key elements involved in the reinforcement learning process within the DDQN algorithm are described below:
[0184] State space S[n]: System throughput is related to the location of the transmitting base station, the UAV relay location, and the user's location. The performance of free-space optical communication is also related to line-of-sight connectivity and atmospheric conditions. Furthermore, the UAV's speed and remaining energy consumption affect the duration of the flight mission and can serve to remind the UAV not to exceed its speed limit and to return to base in time. The influence of air-to-ground obstacles and their potential collision risks jointly affect the UAV's trajectory. Therefore, the state space includes the following information: the UAV's current position p. uav The location of the base station p BS User's location p users The speed of the drone, v uav Remaining energy of drones uav Ground obstacle information ground Information on aerial obstacles air Determining the line-of-sight distance (m) of the drone's location l and collision marker m c Atmospheric visibility data f visThe above states cover the necessary decision-making factors for UAV trajectory planning and uplink communication in low-altitude airspace, laying the foundation for efficient decision-making by intelligent agents. The specific description is as follows:
[0185] S[n]={p uav [n],p BS ,p k users [n],v uav [n],e uav [n],o ground ,o air [n],m l [n],m c [n],f vis [n]}.
[0186] Action space A[n]: Since the problem involves both the movement of the UAV and the switching between photoelectric and electronic systems, the action space of the UAV consists of two parts: the movement actions of the UAV and the selection of relay communication.
[0187] Drone movement: The drone is programmed to move at a fixed speed along the x, y, and z axes within the mission area, with its specific trajectory determined by the chosen direction. To facilitate drone decision-making and planning, we discretized the selectable directions, dividing all possible directions in three-dimensional space into 27 types, and defined them through motion coding.
[0188] Specifically, the drone's motion space consists of 27 discrete actions, corresponding to all directional combinations in three-dimensional space, including single-axis motion, dual-axis composite motion, and tri-axis composite motion. Each direction has an increment of [-1, 0, 1], representing movement in the negative direction, no movement, and movement in the positive direction, respectively. For example, the action [1, 0, -1] indicates that the drone moves in the positive x-axis direction, remains stationary in the y-axis direction, and moves in the negative z-axis direction. Additionally, a hovering action [0, 0, 0] is included, allowing the drone to remain stationary at its current position. Therefore, the drone's motion space can be defined as follows:
[0189] A move [n]={(x,y,z)∣x,y,z∈{-1,0,1}}
[0190] Relay Communication Selection: The uplink is equipped with both radio frequency and free-space optical communication devices. The UAV relay can flexibly choose the appropriate device based on environmental characteristics. Therefore, the relay's communication action selection can be represented by a binary discrete variable A. com [n] represents the following:
[0191] A com [n]∈{0,1}
[0192] Among them, A com [n] = 0 indicates that radio frequency communication is selected, A com [n] = 1 indicates that free-space optical communication is selected. The total action space can be represented as the Cartesian product of movement actions and communication actions:
[0193] A[n]=A move [n]×A com [n]
[0194] A complete action of a drone can be represented as a binary tuple:
[0195] a[n]=(a move [n],a com [n])
[0196] Reward Design: In deep reinforcement learning, the reward function is used to evaluate the quality of an agent's behavior in a certain state, playing a crucial role in guiding the agent's behavior. Reward design can transform the optimization objective into maximizing cumulative rewards. The goal of this invention is to maximize the total user throughput of the system while considering constraints. Therefore, the reward is set as follows:
[0197] First, the main reward for slot n is the sum of the throughput of k users, denoted as:
[0198]
[0199] in, This is a normalization factor used to ensure that throughput rewards and penalties are set to the same order of magnitude. The reward guides the drone relay to choose a better communication method and move to a location where it can achieve higher system throughput.
[0200] In mission areas with numerous obstacles, we employ a penalty and reward mechanism to constrain drone misbehavior, thereby ensuring safety and communication quality. The specific settings are as follows:
[0201] R col [n]=κ col [n]·r col
[0202] R out [n]=κ out [n]·r out
[0203] R nlos [n]=κ nlos [n]·r nlos
[0204] Among them, κ col[n] is a binary collision constraint index used to guide the drone to avoid obstacles, κ col [n] = 1 indicates that the drone engaged in collision behavior in that time slot, κ col [n] = 0 indicates that the drone is flying normally and safely. Similarly, κ out [n] and κ nlos [n] are binary variables representing whether the drone has crossed the boundary and whether it is a non-line-of-sight (NLS) connection, respectively, which are used to encourage the drone to fly within the designated mission area and to prioritize line-of-sight links for transmission.
[0205] In summary, the total reward obtained by the agent in time slot n can be expressed as:
[0206]
[0207] During the training of DDQN, a Q-network Q(s,a;θ) and a target Q-network Q′(s,a;θ′) need to be initialized first, and different parameters θ and θ′ are set for them. Simultaneously, an experience replay pool is initialized to store the agent's experience interacting with the environment. In each training round, the agent starts from the current state s. t Initially, select an action a based on the ò-greedy strategy. t That is, choosing a random action with probability ò, or choosing the optimal action given by the Q network. After executing the action, the environment returns a reward r. t and the next state s t+1 This information is then stored in an experience replay pool. Next, at regular intervals, the algorithm randomly selects a small batch of experiences from the replay pool for training. For each experience, the target value y is first calculated using the target network Q′. t The calculation formula is as follows:
[0208]
[0209] Where γ is the discount factor, argmax a Q(s t+1 ,a;θ) represents the Q-network in the next state s t+1 The optimal action is selected. Then, the current prediction value Q(s) of the Q-network is calculated. t ,a t ;θ), and the difference between the target value and the predicted value is measured by the mean squared error loss function.
[0210]
[0211] Among them, S bitchThe size of the mini-batch experience is used to guide the Q-network to gradually approach the target value by minimizing this loss function. Finally, the parameters are updated using gradient descent, and the Adam optimizer is used to update the parameters θ of the Q-network. Furthermore, every fixed number of steps, the parameters of the target Q-network are synchronously updated to the parameters of the Q-network (θ′←θ) to reduce fluctuations in the target value and ensure training stability. This training process is repeated until training is complete, ultimately yielding a Q-network that can be used for decision-making. In this way, DDQN effectively reduces the instability caused by overestimation in standard DQN, thus achieving more stable and efficient training.
[0212] Figure 3 and Figure 4 This refers to the flight path of a UAV electro-optical hybrid relay under different obstacle environments and user location distribution conditions. Figure 3 This reflects how, in situations where ground users are widely distributed, drones dynamically adjust their positions based on changes in user movement. Guided by a joint optimization algorithm based on DDQN for photoelectric switching, power allocation, and trajectory planning, the drone plans a relatively smooth and efficient trajectory to ensure optimal communication coverage while effectively reducing path loss and interference. Figure 4 In the comparison, the user locations are more concentrated, and the drone trajectories are also more clustered, demonstrating the direct impact of user distribution on trajectory planning. This comparison fully illustrates that the proposed algorithm can flexibly adjust according to the different characteristics of user locations, optimize trajectory design, improve communication coverage quality, and demonstrate the system's adaptability and robustness in complex environments.
[0213] Figure 5 and Figure 6 The flight path of a hybrid electro-optical relay for a UAV under different obstacle distribution conditions. Figure 5 In this scenario, both dynamic aerial obstacles and fixed ground obstacles are present. Due to the obstruction of the line-of-sight link by aerial obstacles, the UAV adopts a low-altitude hovering strategy to avoid the impact of dynamic obstacles on free-space optical communication as much as possible, thereby maintaining the stability and reliability of the communication link. Figure 6 The study demonstrates a scenario with only fixed obstacles. In this scenario, the UAV searches for the optimal communication position by circling and ascending, resulting in a simpler and smoother flight path. This difference reflects the significant impact of obstacle type and distribution on the UAV's electro-optical hybrid relay flight path planning, while highlighting the adaptability and effectiveness of the proposed DDQN-based joint optimization algorithm for electro-optical switching, power allocation, and flight path planning in complex obstacle environments. By dynamically adjusting the flight path, the UAV can minimize path loss and communication interference in complex environments, providing solid technical support for improving the system's communication performance and environmental adaptability.
[0214] Figure 7 The graph illustrates the changes in system communication performance over time, including the changes in system throughput, the switching between free-space optical and radio frequency (RF) communication modes in the uplink of the UAV-based electro-optical hybrid relay communication system, and the situation where the transmitting base station is subjected to RF denial interference. In the throughput graph, the red line represents the total throughput of the electro-optical hybrid system, and the green line represents the throughput of RF communication. It can be clearly observed that the electro-optical hybrid system consistently maintains a higher total throughput compared to a single RF communication mode. Combining the changes in RF denial status with the switching of communication modes, the system's ability to flexibly switch communication modes according to environmental conditions can be intuitively demonstrated. Due to the high capacity advantage of free-space optical communication, it becomes the primary communication method in most time slots, while RF communication, as an auxiliary solution, is only activated in a few time slots to ensure communication continuity when the line-of-sight link is blocked. The proposed algorithm can quickly switch to radio frequency communication to maintain basic services when free space light is interrupted due to obstruction, and free space light communication can fill the performance gap in a timely manner when radio frequency is unavailable due to rejection. It fully leverages the high capacity advantage of free space light and the complementary characteristics of radio frequency's environmental adaptability, significantly improving the overall performance and reliability of communication in complex environments, and providing an important guarantee for achieving efficient and reliable communication services.
[0215] Figure 8 Communication performance was compared in four scenarios: fixed obstacles with radio frequency rejection, fixed obstacles without radio frequency rejection, dynamic obstacles with radio frequency rejection, and dynamic obstacles without radio frequency rejection. Figure 8 As can be seen, hybrid optoelectronic communication exhibits the best system throughput performance in all scenarios, significantly higher than single radio frequency (RF) or free-space optical communication modes. This indicates that, guided by the proposed DDQN-based joint optimization algorithm for optoelectronic switching, power allocation, and trajectory planning, the UAV hybrid optoelectronic relay combines the advantages of both technologies, enabling it to better adapt to complex obstacle environments and RF denial conditions. Due to limitations in RF bandwidth and spectrum resources, RF communication exhibits the worst system throughput in all scenarios. From an environmental perspective, in scenarios with RF denial, RF communication suffers severe interference, significantly reducing its throughput, while free-space optical communication, unaffected by RF interference, demonstrates a strong advantage in this scenario. In dynamic obstacle scenarios, regardless of the presence of RF denial conditions, the throughput of all three communication modes decreases compared to fixed obstacle scenarios. This is because the occlusion effect of dynamic obstacles increases the probability of non-line-of-sight free-space optical communication, leading the UAV hybrid optoelectronic relay to rely more on RF for uplink communication, resulting in an overall performance decline. Nevertheless, hybrid optoelectronic communication still maintains the highest throughput, indicating its superior adaptability to dynamic environments compared to single communication modes.
[0216] Figure 9 This study examines the throughput of a hybrid optoelectronic relay communication system for unmanned aerial vehicles (UAVs) across multiple users over time slots. It observes significant fluctuations in throughput across different time slots. High throughput phases correspond to user service using free-space optical communication, while low throughput phases rely on radio frequency (RF) communication. These variations in user throughput reflect the system's switching mechanism in low-altitude cluttered environments and complex spectrum conditions. It utilizes free-space optical communication to provide high data rates while supplementing with RF to maintain continuous user connectivity, thus achieving reliable service around the clock. Users with better channel quality (e.g., users 1 and 2) achieve higher data rates, while users with poorer channel quality experience lower throughput but still maintain stable network access. This demonstrates the system's ability to ensure multi-user communication coverage.
[0217] This invention also provides a UAV relay trajectory planning system for free-space optical communication, comprising:
[0218] Free-Space Optical Communication Channel Model Construction Module: This module is used to construct a free-space optical communication channel model for UAV relay, comprehensively considering four key influencing factors: atmospheric turbulence effects, weather-related atmospheric attenuation, optical link pointing error, and signal angle of arrival fluctuation. By dynamically compensating for geometric loss and atmospheric loss, it quantifies optical power gain and channel capacity, providing basic model support for communication link stability assessment.
[0219] RF Communication Channel Model Construction Module: This module establishes an RF channel model, distinguishing between line-of-sight and non-line-of-sight transmission scenarios, and combining path loss probability models with non-orthogonal multiple access technology. By dynamically adjusting the power allocation strategy through real-time evaluation of channel quality and multi-user interference, it ensures the reliability and spectral efficiency of the RF link in complex low-altitude environments.
[0220] Low-altitude airspace hybrid obstacle model construction module: This module constructs a 3D hybrid obstacle model that includes fixed ground obstacles (such as buildings and no-fly zones) and time-varying aerial obstacles (such as dynamic aircraft). Based on the Ray-AABB algorithm, it detects the line-of-sight connectivity between the UAV and the base station in real time. By combining offline pre-calculation of obstacle spatial distribution with online dynamic detection, it generates a set of feasible locations for the UAV, providing obstacle avoidance constraints for trajectory planning.
[0221] Radio Frequency Denial Model Construction Module: This module simulates the dynamic characteristics of time-varying radio frequency denial interference in low-altitude environments, using binary state variables to indicate the presence of interference. When radio frequency denial is detected, a forced switch to free-space optical communication mode is initiated. Furthermore, by incorporating statistical patterns from historical interference data, the communication mode switching strategy is optimized to reduce the impact of sudden interference on system performance.
[0222] User Mobility Model Construction Module: This module uses a random walk model to characterize the movement behavior of ground users, defines the random distribution rules of user movement direction and speed, and introduces a boundary reflection mechanism. When a user moves to the edge of the task area, the movement direction is automatically adjusted to simulate the dynamic trajectory in a real-world scenario, providing real-time user position prediction support for UAV trajectory planning.
[0223] Joint solution module for UAV trajectory planning, electro-optical switching strategy and user power allocation: This module uses a deep reinforcement learning framework to collaboratively optimize UAV trajectory, communication mode switching and user power allocation.
[0224] Specifically, it includes:
[0225] Path planning: Based on the DDQN algorithm, the system searches for the target path point that maximizes system throughput and adjusts the path in real time in combination with energy consumption constraints.
[0226] Optical-to-electric switching strategy: Dynamically select the optimal communication mode based on channel status and interference detection results.
[0227] Power allocation: A fractional power allocation algorithm is adopted. Under the premise of satisfying the causal constraints of reverse information, weak channel users are compensated first, thereby improving multi-user fairness and system throughput.
[0228] It should be understood that the UAV relay trajectory planning system for free-space optical communication provided in this embodiment can realize all the technical solutions in the above method embodiments. The functions of each functional module can be specifically implemented according to the methods in the above method embodiments. The specific implementation process can be referred to the relevant descriptions in the above embodiments, which will not be repeated here.
[0229] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for planning UAV relay paths based on photoelectric hybrid switching, characterized in that, Includes the following steps: (1) Establish a hybrid photoelectric relay communication system for UAVs in low-altitude spectrum denial and complex obstacle environments, the system including a UAV relay equipped with free space optical and radio frequency dual-mode communication equipment, a ground base station and multiple mobile users; (2) Establish a free space optical channel model, considering the two main channel influencing factors of the free space optical link, namely the scintillation effect caused by turbulence and the turbulence attenuation caused by severe weather conditions; (3) Establish a radio frequency channel model, consider both line-of-sight and non-line-of-sight channels in large-scale fading, and consider the interference between multiple users in the downlink relay from UAV to user; (4) Establish an energy consumption model for the UAV, considering only the propulsion energy required for the UAV to overcome resistance and support its own weight; (5) Establish a low-altitude airspace hybrid obstacle model and calculate the feasible location set of UAVs that can achieve line-of-sight connection according to the Ray-AABB algorithm; the hybrid obstacle model includes fixed ground obstacles and time-varying aerial obstacles to simulate the real low-altitude airspace environment; (6) Construct a time-dynamically changing radio frequency rejection interference model to demonstrate the feasibility of disrupting the radio frequency communication link at the transmitting end through random time-varying interference signals; (7) Establish a user random walk movement model. After each movement, the movement node updates its speed and direction at a constant time interval or constant movement distance. If the movement node reaches the boundary of the simulation area, it will "bounce back" along the reflection angle of the incident direction and continue to move along the new path. (8) Construct a joint optimization problem for UAV trajectory planning, photoelectric switching strategy, and user power allocation; (9) Based on the deep reinforcement learning framework, the DDQN algorithm is used to jointly optimize the UAV trajectory planning, photoelectric switching strategy and user power allocation.
2. The UAV relay trajectory planning method based on photoelectric hybrid switching according to claim 1, characterized in that, The free-space optical channel model described in step (2) is as follows: Average gain of optical power at the receiver for: The first term represents the geometric loss due to the divergence of the emitted beam, and the second term represents the weather-related atmospheric loss caused by scattering and absorption. It is the diameter of the receiver aperture. It is the beam divergence angle. The weather-related attenuation coefficient, determined according to the Beer-Lambert law, is expressed as follows: in, Indicates the visibility range. It is the optical wavelength. These are the Mie scattering coefficients determined by the Kim model; The intensity function under different turbulent conditions is described using the Gamma-Gamma distribution: in, Represents the Gamma function. This is a modified Bessel function of the second kind; and The effective values of small-scale and large-scale eddies in the scattering environment related to atmospheric conditions are given by the following formulas: in, , , For refractive index structural parameters, Optical wavenumber; For a free-space optical channel using intensity modulation direct detection, the achievable rate under constraints of random channel gain and average optical transmission power is: in, Let be the base of the natural logarithm. It is the bandwidth of the free-space optical link. Indicates optical transmission power. This represents the noise variance.
3. The UAV relay trajectory planning method based on photoelectric hybrid switching according to claim 1, characterized in that, The specific radio frequency channel model described in step (3) is as follows: Base station to drone relay and drone relay to the first The free space path loss between users is expressed as: in, This indicates a base station to a drone relay or a drone relay to a [other location / system]. The distance between users It is the carrier frequency. It's the speed of light; Drone relay to the The expression for line-of-sight and non-line-of-sight path loss between users is: in, and These are the additional path losses for line-of-sight and non-line-of-sight links, respectively. Drones and base stations or the first The probability modeling of line-of-sight channels between users is as follows: in, and It is a coefficient related to the operating environment. It is the elevation angle between the drone and the user or base station; the probability of a non-line-of-sight channel is expressed as... ; Base station to drone relay and drone relay to the first The path loss between users is expressed as: This leads to the conclusion in the time slot Time drones and the first The path loss factor between users is: Therefore, the radio frequency transmission rate from the base station to the drone relay is: in, Indicates the bandwidth of the radio frequency link. It is radio frequency communication power. It is the noise power spectral density; Drone relay and users The downlink transmission rate between them is expressed as: The total capacity of the system for multiple users is: 。 4. The UAV relay trajectory planning method based on photoelectric hybrid switching according to claim 1, characterized in that, The specific UAV energy consumption model described in step (4) is as follows: The propulsion power consumption of a rotary-wing UAV is defined as: in, Indicates the speed of the drone. It is the average rotor induced speed during hovering. This indicates the tip velocity of the propeller blade. The blade power is expressed as: in, Indicates the profile drag coefficient. Indicates air density, Indicates rotor robustness, Indicates the area of the turntable. It is the blade angular velocity. It is the blade radius; The induced power at zero velocity is expressed as: in, The incremental correction factor for induced power. Indicates the weight of the drone The system in time The energy consumption within is: 。 5. The UAV relay trajectory planning method based on photoelectric hybrid switching according to claim 1, characterized in that, The low-altitude airspace mixed obstacle model described in step (5) is as follows: In the formula, For the entire space within the scene, express The total space occupied by all obstacles in a time slot includes both fixed and time-varying obstacles. Ground obstacles are modeled based on the density and average height of urban buildings, and the proportion of space occupied by moving obstacles in the air is set. express A time slot that is not occupied by obstacles but where line-of-sight connection cannot be achieved. express Time-slot drones can choose their own waypoints.
6. The UAV relay trajectory planning method based on photoelectric hybrid switching according to claim 1, characterized in that, The radio frequency rejection interference model described in step (6) is as follows: The feasibility of disrupting the transmitter's radio frequency communication link using random time-varying interference signals; assuming that in each time slot In this context, the state of whether the transmitting base station is subject to radio frequency interference can be expressed by the following formula: in, Indicates in time slot The base station is experiencing radio frequency interference, which prevents the establishment of a radio frequency communication link. Conversely, This indicates that the radio frequency link can be used normally within that time slot.
7. The UAV relay trajectory planning method based on photoelectric hybrid switching according to claim 1, characterized in that, The user random walk movement model described in step (7) is as follows: The user's direction of movement is determined by The motion speed is determined by a random angle evenly distributed between the points, and ranges from a predefined range. Randomly assigned in the middle, This represents the user's maximum movement speed; after each movement, the moving node updates its speed and direction at constant time intervals or constant movement distances; if the moving node reaches the boundary of the simulation area, it will "bounce back" along the reflection angle of the incident direction and continue moving along the new path.
8. The UAV relay trajectory planning method based on photoelectric hybrid switching according to claim 1, characterized in that, The implementation process of step (8) is as follows: Among them, constraint C1 specifies the starting and returning positions of the UAV and restricts its trajectory points to a specified three-dimensional task area; constraint C2 limits the total energy consumption of the UAV to ensure that it completes the task within a limited energy budget; constraint C3 specifies the choice of communication method through binary variables. The constraints specify whether the uplink in each time slot uses free-space optical or radio frequency (RF) communication; constraint C4 constrains the maximum power that the UAV relay can provide for downlink communication, ensuring that the total power allocated to all service users is less than the maximum power that the UAV relay can provide; constraint C5 limits the total throughput of users to not exceed the maximum capacity supported by the uplink through information causality constraints; constraint C6 ensures that the hybrid optoelectronic relay can provide communication coverage to all users in each time slot; constraint C7 stipulates that when using free-space optical communication for uplink, it must be ensured that the UAV's location can achieve a line-of-sight connection with the transmitting end; constraint C8 stipulates that when using RF communication for uplink, it must be ensured that the transmitting base station in that time slot is not affected by RF rejection interference. For drone relay and users Downlink transmission rate between.
9. The UAV relay trajectory planning method based on photoelectric hybrid switching according to claim 1, characterized in that, The user power allocation process described in step (9) is as follows: Power control is performed based on the user's instantaneous channel state information to compensate for unfairness caused by differences in channel conditions among users. The allocation strategy is as follows: in, Represents the power allocation factor. Indicates base station to user Channel gain, Indicates user Gaussian white noise in the corresponding channel; when When, it indicates equal power distribution among users. The smaller the value, the less power is allocated to users with good channel conditions; The initial transmit power is set to the maximum available power of the relay node. Based on the current transmit power, FTPA is used to calculate the power allocation for each user. It is verified whether the total throughput of the system satisfies the information causality constraint C2 under the current power allocation. If the current power allocation scheme cannot meet the constraint, the transmit power is adjusted using a bisection method to gradually approach the optimal solution that meets the condition. At the same time, attention is paid to the communication coverage problem of ground users. During the process of adjusting the transmit power, it is necessary to ensure the communication connection with ground users. An iterative optimization strategy is introduced. By repeatedly adjusting and approximating the transmit power and power allocation, the throughput and communication quality of the system are optimized while ensuring the information causality constraint.
10. The UAV relay trajectory planning method based on photoelectric hybrid switching according to claim 1, characterized in that, The specific implementation process of step (9) is as follows: The joint optimization problem of UAV trajectory planning and electro-optical switching strategy is modeled as a Markov decision process. By reasonably defining the state space, action space and reward function, the problem is transformed into a solution task under the reinforcement learning framework. The DDQN algorithm is adopted to continuously interact with the environment and learn, and gradually optimize the UAV trajectory planning and electro-optical switching strategy. The status includes the drone's current location. Location of base stations User location The speed of drones Remaining energy of drones Ground obstacle information Information on aerial obstacles Determining the line-of-sight distance of the drone's location and collision markers Atmospheric visibility data State space The specific details are as follows: The action space of drones The Cartesian product of the movement action and the communication action: in, This indicates that radio frequency communication is selected. This indicates the use of free-space optical communication; A complete action of a drone can be represented as a binary tuple: The reward function is set as follows: Time slot The main reward is The total throughput of all users is denoted as: in, It is a normalization coefficient used to ensure that throughput rewards and penalties are set to the same order of magnitude. When there are numerous obstacles in the mission area, a penalty and reward mechanism is used to constrain the drone's misbehavior. The specific settings are as follows: in, It is a binary collision constraint index used to guide drones to avoid obstacles. This indicates that the drone collided during that time slot. This indicates that the drone is flying normally and safely. and These are binary variables representing whether the drone has crossed the boundary and whether it is a non-line-of-sight (LOS) connection, respectively, which are used to encourage drones to fly within the designated mission area and to prioritize line-of-sight links for transmission. The total reward obtained by the agent in time slot n is expressed as: 。
Citation Information
Patent Citations
Multi-unmanned aerial vehicle flight search and rescue trajectory optimization method for disaster area rescue
CN115866574A
Unmanned aerial vehicle relay route planning method and system for free space optical communication
CN118432695A