A Deep Reinforcement Learning-Based Method and System for Hybrid Optical-Electronic Secure Transmission of Unmanned Aerial Vehicles (UAVs)
By using a UAV three-dimensional acceleration control and multi-dimensional resource real-time joint scheduling network based on deep reinforcement learning, the problems of spectrum resource scarcity and rapid changes in link line distance in emergency scenarios of UAV relay secure transmission system are solved, realizing real-time optimization and secure transmission performance improvement of UAV relay communication system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANJING UNIV OF INFORMATION SCI & TECH
- Filing Date
- 2026-03-12
- Publication Date
- 2026-05-26
AI Technical Summary
Existing UAV relay secure transmission systems are susceptible to spectrum scarcity and co-channel interference in emergency scenarios, making it difficult to meet the requirements for communication security and high speed. Furthermore, the FSO/RF hybrid transmission method is difficult to adapt to the rapid changes in link line-of-sight caused by UAV three-dimensional maneuvering, and existing methods are difficult to optimize in real time in low-altitude dynamic environments.
A UAV 3D acceleration control and multi-dimensional resource real-time joint scheduling network based on deep reinforcement learning is adopted. By jointly optimizing UAV 3D acceleration and resources such as bandwidth and power in real time, an optimization framework for maximizing end-to-end confidentiality rate is constructed. The complementary characteristics of FSO high capacity and RF stability are utilized, and feasible mapping is achieved by combining hyperbolic tangent compression and Euclidean projection to generate real-time policies.
Real-time joint optimization of UAV three-dimensional acceleration and resource allocation is achieved in complex low-altitude dynamic environments, significantly improving the security and confidential transmission performance of FSO/RF hybrid UAV relay communication systems, and meeting the communication security and high-speed requirements of emergency scenarios.
Smart Images

Figure CN121842726B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method and system for secure transmission of photoelectric hybrid transmission in unmanned aerial vehicles (UAVs) based on deep reinforcement learning, belonging to the field of UAV communication technology. Background Technology
[0002] In typical application scenarios such as emergency rescue, public safety communications, and high-density urban mobile communications, traditional communication networks relying on fixed ground base stations often face problems such as limited infrastructure deployment, link interruptions due to terrain and building obstructions, and rapid changes in user distribution and service load, thus affecting service continuity and command and dispatch efficiency. Unmanned aerial vehicles (UAVs), with their advantages of high mobility, low cost, reusability, and rapid deployment and configuration, can serve as flexibly deployable aerial relay nodes in these scenarios. When fixed ground base stations are unable to maintain effective service due to mountainous terrain, building obstructions, or disaster damage, UAVs can be quickly deployed to establish air-ground collaborative relay channels. They can intelligently adjust their position or flight trajectory in three-dimensional space according to service needs, quickly establishing line-of-sight (LoS) links, providing more reliable data access and transmission support for ground terminals, thereby maintaining the communication continuity and reliability of critical services. Based on these advantages, UAV relay communication architecture plays a crucial role in mission communication assurance, expanding the effective coverage of terrestrial cellular networks, and rapidly rebuilding communication networks in emergency scenarios, possessing significant research value and broad application prospects.
[0003] In the aforementioned emergency rescue and public safety communication scenarios, communication links often carry sensitive content such as command and control instructions, situational awareness information, and critical business data. If this transmission is intercepted by unauthorized nodes, it will directly impact mission execution and command and dispatch security. Therefore, communication systems should not only focus on coverage and transmission capacity but also prioritize physical layer security and secure transmission as core performance indicators to ensure the secure transmission of critical business data and the stable availability of command and dispatch links. However, existing secure transmission research based on UAV relay architectures largely relies on a single radio frequency (RF) link and transmits data within a predetermined time-frequency resource allocation. This type of solution is susceptible to spectrum scarcity and co-channel interference, leading to link congestion and limited system capacity. This restricts the achievable rate of legitimate links and further increases the risk of eavesdropping, making it difficult to meet the security and high-speed secure transmission requirements of the aforementioned scenarios. Compared to RF communication, free-space light (FSO) communication offers advantages such as ultra-high bandwidth, unlicensed spectrum, and resistance to electromagnetic interference. In LoS (LoS) link mode, FSO links can provide significantly higher throughput than RF links, and their high directivity and narrow beam characteristics can significantly improve the security and secure transmission performance of communication systems. However, in low-altitude environments, buildings and various infrastructures can easily block the light beam, and atmospheric turbulence and weather conditions can cause light intensity fluctuations and severe fading, significantly affecting the availability and stability of FSO links, and even causing momentary outages. Therefore, to combine the advantages of FSO and RF links while compensating for their respective shortcomings, FSO / RF hybrid communication is considered the preferred solution for achieving highly reliable and secure communication. However, existing research on FSO / RF hybrid transmission for secure communication mainly focuses on non-UAV relay quasi-static application scenarios using fixed nodes or static platforms as relay nodes. When introducing FSO / RF hybrid transmission into UAV relay communication systems, more prominent engineering implementation and theoretical modeling challenges will be encountered. The three-dimensional maneuvering and flight trajectory changes of UAVs cause the geometry and obstruction status of the air-to-ground link to continuously evolve, which in turn leads to rapid changes in the communication link status over time. The FSO link is highly sensitive to the LoS condition and pointing alignment. When the link enters the non-line-of-sight (NLoS) state, the FSO transmission will experience a momentary interruption, which will cause the system transmission capacity to drop sharply and significantly weaken the continuity and reliability of service carrying.
[0004] Therefore, the link reliability and secure transmission performance of FSO / RF hybrid UAV relay secure transmission systems are no longer solely determined by resource allocation, but are closely related to UAV maneuver control. How to perform real-time joint optimization of UAV maneuver control and system resource allocation in complex dynamic environments, while focusing on security performance indicators, to ensure communication security and secure transmission performance, is a critical and significantly challenging problem.
[0005] Currently, a few studies have investigated secure transmission methods for UAV relays in FSO / RF hybrid systems. These methods introduce a continuous convex approximation approach, which alternately updates the UAV trajectory control variables and resource allocation variables such as power in blocks. In each iteration, the non-convex average security rate target and related non-convex constraints are made first-order convex at the current iteration point, thereby constructing a solvable sequence of convex subproblems to gradually approximate the approximate solution of the original problem. However, such methods are prone to getting trapped in local optima and are difficult to perform real-time joint optimization based on the real-time environmental conditions in low-altitude time-varying scenarios.
[0006] In summary, although some progress has been made in secure transmission via UAV relay, existing methods still have the following shortcomings:
[0007] (1) Most existing studies on secure transmission of UAV relays only consider RF single links, which are easily constrained by the scarcity of spectrum resources and co-channel interference, making it difficult to meet the needs of emergency scenarios for communication security and high-speed secure transmission.
[0008] (2) Existing FSO / RF hybrid secure communication methods mostly use static platforms as relay nodes, which are difficult to adapt to the rapid changes in link line-of-sight status and obstruction relationship caused by UAV three-dimensional maneuvering in UAV relay communication systems, and are difficult to guarantee communication continuity and security.
[0009] (3) Existing research on secure transmission of UAV relays for FSO / RF hybrid systems relies on offline iterative methods based on continuous convex approximation, which makes it difficult to respond quickly to changes in environmental conditions under low-altitude dynamic environments, thus significantly limiting the performance of secure transmission.
[0010] Therefore, there is an urgent need to improve the existing FSO / RF hybrid UAV relay transmission method. Summary of the Invention
[0011] Objective: Existing UAV relay secure transmission methods are mostly limited to RF links, which are susceptible to eavesdropping and have limited room for performance improvement. This invention provides a hybrid optoelectronic secure transmission method and system for UAV relay based on deep reinforcement learning. It leverages the complementary characteristics of high capacity, high directionality, and electromagnetic interference resistance of FSO (Flying Signal Socket) and strong stability and low interruption resistance of RF links to construct a joint optimization framework aimed at maximizing the end-to-end secure transmission rate at the time slot level. Real-time joint optimization of bandwidth, power, and UAV three-dimensional acceleration is performed. Simultaneously, a deep reinforcement learning-based UAV three-dimensional acceleration control and multi-dimensional resource real-time joint scheduling network with bounded and feasible mapping capabilities is introduced. Under multiple constraints, an instantly executable strategy is output through a single forward inference, thereby improving the system's end-to-end secure transmission rate and meeting the real-time deployment requirements of engineering projects.
[0012] Technical solution: To solve the above technical problems, the technical solution adopted by the present invention is as follows:
[0013] Firstly, a method for secure electronic-electrical hybrid transmission in UAV relays based on deep reinforcement learning, specifically including:
[0014] Step 1: Obtain the real-time joint optimization problem of UAV three-dimensional acceleration, bandwidth resources, user power, and UAV power configuration under multiple constraints.
[0015] Step 2: Construct a Markov decision process based on the real-time joint optimization problem, and obtain the state vector of each time slot through the Markov decision process.
[0016] Step 3: Input the state vector of each time slot into the real-time joint scheduling network to obtain the unconstrained original joint action vector.
[0017] Step 4: Perform bounded operation and feasible mapping on the unconstrained original joint action vector to obtain the actual executable joint action vector.
[0018] Optionally, it also includes: Step 5: Calculate the end-to-end confidentiality rate of the real-time joint optimization problem based on the actual executable joint action vector, and calculate the instant reward based on the end-to-end confidentiality rate of the real-time joint optimization problem.
[0019] Step 6: Input the actual executable joint action vector into the real-time joint scheduling network to obtain the updated state vector. Use the state vector, the actual executable joint action vector, the immediate report, and the updated state vector as an empirical sequence.
[0020] Step 7: Calculate the gradients of the value loss function and the total cost function of the real-time joint scheduling network based on the empirical sequence. Then, update the parameter vector of the real-time joint scheduling network synchronously based on the gradients of the value loss function and the total cost function to obtain the updated parameter vector of the real-time joint scheduling network.
[0021] Step 8: Replace the parameter vector in the real-time joint scheduling network in step 3 with the updated real-time joint scheduling network parameter vector, and repeat steps 1 to 4 to obtain the actual executable joint action vector.
[0022] Optionally, the expression for the real-time joint optimization problem is as follows:
[0023]
[0024] in, For discrete time slot indexes within a flight cycle, The length of each time slot, For the system's time slot-level end-to-end security rate, The security rate of the FSO / RF parallel link for ground users to send confidential information to UAVs. This refers to the RF link backhaul rate from the UAV to the ground base station. The total RF bandwidth of the system. For the RF bandwidth of the uplink from the ground user to the UAV, For the RF bandwidth of the UAV to ground base station backhaul link, This is the peak power of ground users. Electrical power allocated to the RF link for ground users. Electrical power allocated to FSO links for ground users The minimum transmit electrical power allocated to the RF link for ground users. The minimum transmit electrical power allocated to FSO links for ground users. Remaining allocable power for ground users For the peak power allowed by UAV, Return power allocated to UAV, Interference power allocated to UAVs For UAV, the current number Acceleration during flight in each time slot For UAV, the current number The speed of flight in each time slot For the speed of the UAV in the next time slot, The maximum acceleration of the UAV. This is the maximum speed of the UAV. For UAV No. The position of each time slot For UAV No. The position of each time slot For the first The remaining travel energy of the UAV in each time slot for Norm.
[0025] Optionally, the expression for the state vector is as follows:
[0026]
[0027] in, The coordinates of the UAV For UAV, the current number The speed of flight in each time slot The distance from the ground user to the UAV. The distance from the UAV to the eavesdropper. The distance from the ground user to the eavesdropper. This refers to the distance from the UAV to the ground base station. For the first The remaining travel energy of the UAV in each time slot For the line-of-sight status indicator variable of the link from the ground user to the UAV, This is a line-of-sight status indicator variable for the link from the UAV to the eavesdropper. For the line-of-sight status indicator variable of the link from the ground user to the eavesdropper, This is a line-of-sight status indicator variable for the UAV-to-ground base station link. When a Loss channel exists, , , , The line-of-sight status indicator variable is 1 if it is not 0 otherwise.
[0028] Optionally, step 3 specifically includes:
[0029] Step 3.1: Place the first The state vector of each time slot Input the policy network of the real-time joint scheduling network, and output the state. Down mean vector In state Down mean vector In state Down variance vector And in state Down variance vector .
[0030] Step 3.2: Obtain the policy network parameter vector Down about conditional probability density Its specific form is as follows:
[0031]
[0032] in, Indicates a Gaussian distribution. Represented by the variance vector The resulting diagonal covariance matrix.
[0033] Step 3.3: Obtain the policy network parameter vector Down about conditional probability density Its specific form is as follows:
[0034]
[0035] in, Represented by the variance vector The resulting diagonal covariance matrix.
[0036] Step 3.4: According to the information regarding conditional probability density Sampling was performed According to conditional probability density Sampling was performed ,according to and Get the first The unconstrained original joint action vector of each time slot ,in, .
[0037] Optionally, step 4 specifically includes:
[0038] Step 4.1: Based on the unconstrained original UAV three-dimensional acceleration Calculate the actual executable three-dimensional acceleration of the UAV. Its specific form is as follows:
[0039]
[0040] in, Indicated to Euclidean projection operation of the norm sphere This represents the hyperbolic tangent function.
[0041] Step 4.2: Based on the proportional coefficient vector formed by the unconstrained original bandwidth allocation continuous proportional coefficient and the original power allocation continuous proportional coefficient. Calculate the proportional coefficient vector consisting of the actual executable bandwidth allocation continuous proportional coefficient and the actual executable power allocation continuous proportional coefficient. Specifically, it includes: , and Its specific form is as follows:
[0042]
[0043] in, This represents the continuous proportionality coefficient of the total RF bandwidth allocated to the ground user-to-UAV link in the system.
[0044]
[0045] in, The continuous scaling factor for the actual executable power allocation to the backhaul link for UAVs.
[0046]
[0047] in, The continuous scaling factor for the actual executable power allocation to RF links for ground users. It is the hard hyperbolic tangent function.
[0048] Step 4.3, based on the actual executable UAV three-dimensional acceleration The proportional coefficient vector consisting of the actual executable bandwidth allocation continuous proportional coefficient and the actual executable power allocation continuous proportional coefficient. Obtain the actual executable joint action vector ,in, .
[0049] Optionally, step 5 specifically includes:
[0050] Step 5.1: Calculate the security rate of the FSO / RF parallel link for transmitting confidential information from the ground user to the UAV. Its specific form is as follows:
[0051]
[0052] in, For ground users Effective security rate of the RF link to UAV For ground users The rate of the FSO link to the UAV.
[0053] Step 5.2: Calculate the distance from the UAV to the ground base station RF link backhaul rate Its specific form is as follows:
[0054]
[0055] Step 5.3: Based on the security rate and return rate Calculate end-to-end security rate Its specific form is as follows:
[0056]
[0057] Step 5.4: Based on the end-to-end security rate Calculate instant returns Its specific form is as follows:
[0058]
[0059] in, and These are the weighting coefficients. For UAV in time slots With speed Flight power during flight Violation of confidentiality restrictions will result in penalties.
[0060] Optionally, step 6 specifically includes:
[0061] Step 6.1: Based on the actual executable UAV three-dimensional acceleration To obtain the kinematic relationship and the unprojected velocity, by going to The Euclidean projection of the norm sphere maps it to the velocity feasible region of the next time slot of the UAV. .
[0062] Step 6.2: Based on the current time slot position of the UAV To obtain the position of the next time slot of the UAV. .
[0063] Step 6.3: Based on the position of the next time slot of the UAV The distance and line-of-sight status indicators for each link, as well as the remaining flight energy of the UAV, are updated to obtain the state vector for the next time slot, which is then used as the updated state vector. The state vector, the actual executable joint action vector, the immediate report, and the updated state vector are then used as an empirical sequence.
[0064] Optionally, step 7 specifically includes:
[0065] Step 7.1: Summarize the preset number of experience sequences obtained from multiple runs to form the training sample batch for the current policy update, and perform multiple rounds of small-batch random sampling on them. For the sampled small batches, utilize the value network of the real-time joint scheduling network. Output state value assessment Based on the state value assessment quantity and immediate rewards in the experience sequence Constructing Discounted Cumulative Returns .
[0066] Step 7.2: Based on the state value assessment quantity Cumulative returns with discounts Constructing a value loss function During training, the value loss function is minimized. Parameter vectors for realizing value networks Update.
[0067] Step 7.3: Construct time-series differential residuals using the immediate returns in the empirical sequence and the state value assessments of two adjacent time slots. Based on time-series difference residuals Constructing the advantage function estimation .
[0068] Step 7.4: Utilize The ratio of the conditional probability density of the new and old strategies at the same sampled action Constructing the shearing objective function .
[0069] Step 7.5: Based on the shearing objective function Construct the total cost function of the policy network using the policy entropy regularization term. During training, the total cost function is minimized. Update the parameter vector of the policy network .
[0070] Step 7.6: After traversing all mini-batch samples within the training sample batch, one round of iteration is completed. Repeat the above mini-batch sampling and parameter vector update traversal process multiple times until the updated real-time joint scheduling network parameter vector is obtained.
[0071] In a second aspect, a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for secure electronic-electrical hybrid transmission of unmanned aerial vehicle relays based on deep reinforcement learning as described in any of the first aspects.
[0072] Thirdly, a computer device comprising:
[0073] Memory is used to store instructions.
[0074] A processor is configured to execute the instructions, causing the computer device to perform operations as described in any of the first aspects of a deep reinforcement learning-based UAV relay optoelectronic hybrid secure transmission method.
[0075] Beneficial Effects: This invention provides a method and system for UAV relay optoelectronic hybrid secure transmission based on deep reinforcement learning. Targeting FSO / RF hybrid UAV relay communication systems, it aims to maximize the time-slot-level end-to-end secure transmission rate. It constructs a real-time joint scheduling network for UAV three-dimensional acceleration control and multi-dimensional resources based on deep reinforcement learning, incorporating UAV three-dimensional acceleration with multi-dimensional resource configurations such as bandwidth and power into a unified real-time joint optimization framework. To ensure the executability of the strategy under multiple constraints, this invention employs hyperbolic tangent compression to achieve boundedness for UAV three-dimensional acceleration, and further... The Euclidean projection of the norm sphere maps it to a feasible region that satisfies the constraints, while performing interval bounded operations on the continuous scaling coefficients of bandwidth and power allocation, thus obtaining a practically executable joint action vector that satisfies multiple constraints. Furthermore, the velocity state is applied in the state recursion... The Euclidean projection of the norm sphere is used to satisfy velocity constraints. Furthermore, a UAV 3D acceleration control and multi-dimensional resource real-time joint scheduling network based on deep reinforcement learning is used for parameter vector training and online inference. In each time slot, a UAV 3D acceleration control and multi-dimensional resource allocation strategy satisfying multiple constraints is generated based on state changes, achieving real-time adaptive joint optimization of the above decision variables in each time slot to maximize the end-to-end secure rate in each time slot. This method can achieve real-time joint optimization of UAV 3D acceleration and resource allocation in complex low-altitude dynamic scenarios, significantly improving the security, secure transmission capability, and overall robustness of the FSO / RF hybrid UAV relay communication system. Compared with existing technologies, the advantages of this invention are as follows:
[0076] 1. This invention introduces an FSO / RF hybrid transmission mechanism into a UAV relay communication system, fully leveraging the advantages of FSO links such as high capacity, high directionality, and resistance to electromagnetic interference, while combining them with the high stability and uninterrupted nature of RF links, achieving complementary advantages between the two types of links. This further expands the performance enhancement potential of secure transmission, thereby meeting the needs of emergency response and other mission scenarios for communication security and high-speed secure transmission.
[0077] 2. This invention addresses the dynamic evolution of link line-of-sight status changes and FSO alignment conditions caused by UAV three-dimensional maneuvers. It integrates UAV three-dimensional acceleration control and resource allocation into the same real-time joint optimization framework, achieving adaptive adaptation to the dynamic evolution of link status and completing the joint optimization of UAV three-dimensional acceleration control and resource allocation. This improves communication continuity and enhances secure transmission performance in complex low-altitude environments.
[0078] 3. This invention provides a unified model for UAV three-dimensional acceleration control and resource allocation, and introduces hyperbolic tangent compression and... The Euclidean projection of the norm sphere enables feasible mapping of UAV 3D acceleration and velocity. Utilizing a real-time joint decision-making mechanism based on deep reinforcement learning, a real-time executable joint policy can be obtained with only one forward inference per time slot during the online deployment phase, thereby maximizing the end-to-end security rate of each time slot and significantly improving the system's security rate and security. Attached Figure Description
[0079] Figure 1 This is a system model diagram of the FSO / RF hybrid UAV relay secure transmission architecture of the present invention.
[0080] Figure 2 The image shows the three-dimensional flight trajectory of the UAV after using the proposed Deep Reinforcement Learning-based UAV Three-Dimensional Acceleration Control and Multi-Dimensional Resource Real-Time Joint Scheduling Network (DRL-ACMRNet) in this invention.
[0081] Figure 3 The training reward curve for using DRL-ACMRNet in this invention is shown.
[0082] Figure 4 This is a graph showing the change in the average end-to-end security rate during each round of training after using DRL-ACMRNet in this invention.
[0083] Figure 5 This is a comparison chart showing the interruption probability of the proposed real-time UAV three-dimensional acceleration and resource allocation joint optimization method based on DRL-ACMRNet under different rate thresholds, along with four benchmark schemes: no optimization of UAV three-dimensional acceleration, no optimization of power allocation, no optimization of RF bandwidth allocation, and only RF link.
[0084] Figure 6 The figure compares the average end-to-end security rate of the real-time UAV three-dimensional acceleration and resource allocation joint optimization method based on DRL-ACMRNet of this invention with four benchmark schemes: RF link only, no optimization of UAV three-dimensional acceleration, no optimization of power allocation, and no optimization of RF bandwidth allocation, within a complete flight cycle under different user total power conditions. Detailed Implementation
[0085] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.
[0086] The present invention will be further described below with reference to specific embodiments.
[0087] Example 1:
[0088] This embodiment introduces a UAV relay electro-optical hybrid secure transmission method based on deep reinforcement learning, which specifically includes the following steps:
[0089] Step 1: The FSO / RF hybrid UAV relay secure transmission architecture constructed in this invention, as follows... Figure 1 As shown, it includes a mobile ground user ( eavesdropper UAV, ground base stations deployed in fixed locations ( The system includes several fixed three-dimensional ground obstacles. To fully leverage the complementary advantages of FSO and RF, in this transmission architecture, ground users transmit confidential information to UAVs via an FSO / RF parallel link in each time slot. Eavesdroppers can only eavesdrop on the RF signals sent by ground users. While receiving confidential information, the UAV sends co-channel artificial interference signals to the eavesdropper to suppress its decoding capability and transmits the received information back to the ground base station via the RF link. Figure 1 This study demonstrates the impact of obstacles such as buildings on FSO and RF channels in urban scenarios and provides a system modeling foundation for integrating UAV 3D acceleration and resource allocation into a real-time joint optimization framework.
[0090] Confidential data is transmitted to the UAV via a parallel FSO / RF link, and minimum power constraints are set on the uplink FSO and RF links to keep the links active, facilitating continuous monitoring and channel state estimation, and reducing realignment overhead. tapping The UAV sends RF signals and simultaneously receives information and transmits them during flight. It sends an RF co-channel interference signal to prevent the decoding of confidential data and transmits the confidential information back via the RF link. .
[0091] Based on this secure transmission architecture, this invention aims to maximize the time-slot-level end-to-end secure rate of the system. It constructs a real-time joint optimization problem under multiple constraints, addressing UAV three-dimensional acceleration and bandwidth resources, user power, and UAV power configuration, expressed as:
[0092]
[0093] in, For discrete time slot indexes within a flight cycle, The length of each time slot, For the system's time slot-level end-to-end security rate, For ground users The security rate of the FSO / RF parallel link for sending confidential information to the UAV. For UAV to ground base station RF link backhaul rate, The total RF bandwidth of the system. For ground users RF bandwidth to UAV uplink, For UAV to ground base station RF bandwidth of the backhaul link Ground users peak power, For ground users Electrical power allocated to the RF link, For ground users Electrical power allocated to the FSO link, For ground users The minimum transmit electrical power allocated to the RF link. For ground users The minimum transmit electrical power allocated to the FSO link. For ground users Remaining allocable power, , For the peak power allowed by UAV, Return power allocated to UAV, Interference power allocated to UAVs For UAV, the current number Acceleration during flight in each time slot For UAV, the current number The speed of flight in each time slot For the speed of the UAV in the next time slot, The maximum acceleration of the UAV. This is the maximum speed of the UAV. For UAV No. The position of each time slot For UAV No. The position of each time slot For the first The remaining travel energy of the UAV in each time slot for Norm.
[0094] Step 2: Addressing the joint optimization problem of real-time UAV 3D acceleration and resource allocation under the aforementioned multi-constraint conditions, to avoid the frequent occurrence of infeasible decisions caused by traditional direct action generation and the problem of traditional hard truncation constraints weakening the discriminative power of advantageous signals and reducing training stability in multi-dimensional joint actions, this invention proposes a deep reinforcement learning-based real-time joint scheduling network for UAV 3D acceleration control and multi-dimensional resources (DRL-ACMRNet) as a policy solution mechanism. The real-time joint scheduling network introduces hyperbolic tangent compression in the joint action generation stage... The Euclidean projection of the norm sphere and the interval bounded operation enable feasible mapping of joint actions and improve training stability. DRL-ACMRNet consists of a policy network and a value network. In any time slot, both the policy network and the value network take the state vector of the current time slot as input.
[0095] Therefore, the real-time joint optimization problem under multiple constraints is first modeled as a Markov decision process, and the state vector of each time slot is constructed under the Markov decision process as the input of DRL-ACMRNet to drive the subsequent joint policy generation and value evaluation process. In three-dimensional space, the... During that brief moment, the eavesdropper The coordinates are The coordinates of the UAV are represented as Ground users The coordinates are Ground base station The coordinates are represented as Since the three-dimensional position and velocity of the UAV jointly determine the maneuverability and kinematic recursion of the UAV in the next time slot, the distance between each communication terminal reflects the spatial geometric relationship and path loss level, the line-of-sight status indicator variable of each link characterizes the Loss / NLoS state and obstruction conditions of the channel, and the remaining flight energy of the UAV limits the actual executable actions, therefore, the first... The state vector for each time slot is defined as:
[0096]
[0097] in, For ground users Distance to UAV From UAV to eavesdropper distance, For ground users To the eavesdropper distance, For UAV to ground base station distance, For the first The remaining travel energy of the UAV in each time slot For ground users Line-of-sight status indicator variable for the UAV link. From UAV to eavesdropper The line-of-sight status indicator variable of the link. For ground users To the eavesdropper The line-of-sight status indicator variable of the link. For UAV to ground base station The link's line-of-sight status indicator variable, when a Loss of Spectrum (LoS) channel exists. , , , The line-of-sight status indicator variable is 1 if it is not 0 otherwise.
[0098] Step 3: This invention controls the UAV maneuvering through UAV three-dimensional acceleration, while the system's bandwidth and power allocation are controlled through continuous proportional coefficients.
[0099] Specifically, the first The unconstrained raw UAV three-dimensional acceleration for each time slot is defined as:
[0100]
[0101] in, , and The first Each time-slot UAV moves along in a three-dimensional rectangular coordinate system axis, shaft and The unconstrained original three-dimensional acceleration components in the axial direction, and the three unconstrained original three-dimensional acceleration components are conditionally independent of each other.
[0102] The first The proportional coefficient vector, consisting of the continuous proportional coefficients of the original bandwidth allocation and the continuous proportional coefficients of the original power allocation for each unconstrained time slot, is defined as follows:
[0103]
[0104] in, This indicates the total RF bandwidth allocated to ground users. The continuous proportionality coefficient of the original bandwidth allocation to the UAV link. Ground users The continuous scaling factor of the original power allocation to the RF link. The original power allocation ratio coefficients assigned to the UAV for the backhaul link are continuous, and the three are mutually independent.
[0105] Based on this, the policy network parameter vector is represented as , will the The state vector of each time slot Input the policy network of DRL-ACMRNet, and output the state. Down mean vector In state Down mean vector In state Down variance vector And in state Down variance vector .
[0106] The policy network parameter vector Down about The conditional probability density is expressed as Its specific form is as follows:
[0107]
[0108] in, Indicates a Gaussian distribution. Represented by the variance vector The resulting diagonal covariance matrix.
[0109] Similarly, the policy network parameter vector Down about The conditional probability density is expressed as Its specific form is as follows:
[0110]
[0111] in, Represented by the variance vector The resulting diagonal covariance matrix.
[0112] Subsequently, according to the The conditional probability density is sampled to obtain the above. and and according to and Get the first The unconstrained original joint action vector of each time slot .
[0113] Step 4: Since the unconstrained original joint action vector is directly obtained from conditional probability density sampling, its values are often difficult to satisfy the kinematic constraints and bandwidth and power constraints of the UAV. Therefore, this invention introduces hyperbolic tangent compression... The Euclidean projection of the norm sphere and the interval boundedness operation generate a practically executable joint action vector that satisfies multiple constraints. Firstly, for the three-dimensional acceleration control of the UAV, to ensure that the acceleration output by the strategy satisfies the maximum acceleration constraint of the UAV, this invention employs hyperbolic tangent compression and... The Euclidean projection of the norm sphere performs a feasible mapping of acceleration, transforming the original continuous motion into actual executable acceleration. Specifically:
[0114] First of all Hyperbolic tangent compression is applied to each of the three dimensions, and then the process is executed to... The Euclidean projection of the norm sphere yields a practically executable three-dimensional UAV acceleration. ,Right now:
[0115]
[0116] in, Indicated to The Euclidean projection operation of the norm sphere ensures the acceleration constraint of the aforementioned real-time joint optimization problem. Furthermore, compared to the hard truncation method, it has better gradient transitivity and numerical stability, which can reduce the risk of policy oscillation caused by action out of bounds during the training phase.
[0117] Next, to ensure that the actual joint actions performed satisfy the interval constraints of the aforementioned real-time joint optimization problem, this invention allocates the total RF bandwidth of the system to... Apply a hyperbolic tangent compression-based interval bounding operation to the original bandwidth allocation continuous scaling factor to the UAV link and the original power allocation continuous scaling factor allocated by the UAV to the backhaul link. The original power allocation scaling factor allocated to the RF link is subjected to an interval bounded operation based on the hard hyperbolic tangent function. Specifically:
[0118] The total RF bandwidth of the system is allocated to Applying a hyperbolic tangent compression-based interval bounding operation to the original bandwidth allocation continuous scaling factor to the UAV link yields the actually executable bandwidth allocation continuous scaling factor. ,Right now:
[0119]
[0120] in, This indicates that the total RF bandwidth of the system is allocated to The actual bandwidth allocation ratio to the UAV link is continuously proportional, and the total RF bandwidth of the system is divided accordingly. Bandwidth allocation to UAV is .
[0121] At this time, UAV arrives The bandwidth allocation is .
[0122] Applying a hyperbolic tangent compression-based interval bounding operation to the original power allocation continuous scaling coefficients allocated by the UAV to the backhaul link yields the actually executable power allocation continuous scaling coefficients. ,Right now:
[0123]
[0124] in, A continuous proportional coefficient is assigned to the actual power allocation for the UAV to the backhaul link, and the total transmit power of the UAV is allocated accordingly. The power allocated to the backhaul link by the UAV is... Then UAV to The power of transmitting artificial interference signals is .
[0125] for The original power allocation continuous scaling coefficients allocated to the RF link are subjected to an interval bounded operation based on the hard hyperbolic tangent function to obtain the actually executable power allocation continuous scaling coefficients. ,Right now:
[0126]
[0127] in, It is a hard hyperbolic tangent function. express The continuous scaling factor of the actual power allocation to the RF link, and thus the... The total transmit power is allocated. The power allocated to the RF link is , The power allocated to the FSO link is .
[0128] Actual executable bandwidth allocation continuous scaling factor and actual executable power allocation continuous scaling factor Structure become The proportionality coefficient vector is .
[0129] No. The actual executable joint action vector for each time slot is: .
[0130] Step 5: In the In each time slot, the reinforcement learning agent based on DRL-ACMRNet executes the actual executable joint action according to the current state, based on the... , and Given a defined bandwidth and power configuration, the optimization objective of the aforementioned real-time joint optimization problem is calculated, namely, the system time-slot-level end-to-end security rate. And based on this, evaluate the system's performance in a two-hop relay structure. End-to-end secure transmission capability for each time slot.
[0131] This invention adopts the first The system's time-slot-level end-to-end security rate is defined by the minimum value of the uplink and backlink security rates for each time slot. ,in, for The security rate of the FSO / RF parallel link for sending confidential information to the UAV. For UAV to The RF link return rate. Since the uplink uses parallel transmission of FSO and RF, and the eavesdropper can only intercept the RF signal, the achievable rate of the FSO link can be directly included in the secure transmission rate. However, the RF link is subject to eavesdropping risk, and its effective secure rate is constrained by the achievable rate of the eavesdropping link.
[0132] therefore,
[0133] in, For ground users Effective security rate of the RF link to UAV For ground users The rate of the FSO link to the UAV.
[0134] in, Defined as and eavesdropping rate The positive part of the difference is:
[0135]
[0136] in, For ground users The uplink RF link rate to the UAV, eavesdropper For ground users The rate at which the transmitted RF signal can be eavesdropped. This is the attenuation coefficient for artificial interference signals. for Channel gain of the RF link to the UAV for To the eavesdropper The channel gain of the eavesdropping link, For UAV to The channel gain of the interference link, for Additive white Gaussian noise power in the RF link to the UAV for To the eavesdropper The additive white Gaussian noise power of the eavesdropping link, The positive part operator is defined as follows: the system has a positive secure transmission gain in the time slot if and only if the legitimate rate exceeds the eavesdropping rate; otherwise, the secure rate in the time slot is recorded as zero, corresponding to the secure interruption situation.
[0137] Rate of FSO link to UAV It can be represented as:
[0138]
[0139] in, for FSO link bandwidth to UAV It is the photoelectric conversion coefficient. for Channel gain of the FSO link to UAV yes The FSO signal transmit power, , The electro-optical conversion efficiency constant is... yes Additive white Gaussian noise power of the FSO link to UAV.
[0140] In the backhaul phase, UAV to ground base station The RF link backhaul rate is:
[0141]
[0142] in, For UAV to Channel gain of the RF link Is it UAV to The additive white Gaussian noise power of the RF link.
[0143] Since the objective of this invention is to maximize end-to-end security rate in real time, therefore, the first The instantaneous reward for each time slot is defined as:
[0144]
[0145] in, and These are the weighting coefficients. For UAV in time slots With speed Flight power during flight As a penalty item for navigation energy consumption, To avoid violating confidentiality restrictions, penalties apply. ,in, It is not the effective RF security rate itself, but a metric used to characterize the rate advantage of the eavesdropping end over the legitimate end on the RF link. Although the RF security rate in the uplink is defined using the positive part operator. The value is truncated to non-negative, but when the equivalent instantaneous reachable rate of the eavesdropping RF link is higher than that of the legitimate RF link, it still indicates that the RF link in that time slot cannot support secure transmission under eavesdropping constraints and is in a state of secure interruption. Therefore, relying solely on... The truncation can lead to insufficient differentiation in the immediate reward values of states with different levels of eavesdropping threat, making it difficult for the policy to effectively reduce the risk of RF link eavesdropping during training. To avoid insufficient RF security performance when the link switches to NLOS state and causes FSO transmission interruption, this penalty term can explicitly suppress the tendency of the eavesdropping end to choose actions that lead to a dominant position during policy learning, thereby guiding the policy to reduce the probability of security interruption.
[0146] Using the first The actual executable joint action vector for each time slot The system environment is recursively evolved, and the state vector is updated accordingly. Specifically:
[0147] Acceleration Under the influence of kinematics, the unprojected velocity is first obtained from the kinematic relationship. Then through to The Euclidean projection of the norm sphere maps it to the velocity feasible region, i.e.
[0148]
[0149] in, Indicated to Euclidean projection operation of the norm sphere.
[0150] At this point, the UAV location is updated as follows: Based on this, the distance and line-of-sight status indicator variables of each link and the remaining navigation energy of the UAV are updated to form the state vector for the next time slot. .
[0151] Subsequently, state transition samples for each time slot were recorded. Within each flight cycle, the state transition samples are recorded in ascending order of time slot index to form an empirical sequence.
[0152] Furthermore, in each time slot, the parameter vector is... The value network is based on the current state vector Synchronous calculation of state value assessment This is then cached as an auxiliary quantity for subsequently constructing the cumulative return of the discount and the estimation of the advantage function.
[0153] Step 6: During the training phase, in each policy update iteration, the preset number of experience sequences obtained from multiple runs are aggregated to form the training sample batch for that policy update. During the update phase, multiple rounds of iterative small-batch random sampling are performed on this training sample batch. For the sampled small batches, firstly based on the value network... Output state value assessment We construct discounted cumulative reward and advantage function estimates for policy evaluation, providing a stable learning signal for subsequent policy gradient updates.
[0154] Discount cumulative return is defined as
[0155] in, This represents the total number of time slots. For the summation index, As a discount factor, Indicates the first The final state is obtained after executing actions in each time slot and completing state recursion. For the first Instant returns per time slot Discount factor Power of 1 Discount factor Power of 1 The value term is the termination boundary value. The parameter vector update of the value network can be constructed as a value regression problem, where the value network accumulates returns with discounts. The goal is to make the network output state value assessment quantity... Approaching Discount Cumulative Returns .
[0156] Therefore, based on the least squares criterion, the value loss function is defined using the squared error between the state value assessment quantity and the target value.
[0157] in, This is a measure of state value. This indicates the operation of obtaining the desired result.
[0158] The parameter vector is achieved by minimizing this value loss function. The update provides a more accurate state value estimate for the advantage function estimation, thereby reducing the variance of the policy gradient estimate and improving the training convergence stability.
[0159] To depict the first To assess the relative advantage of joint actions across time slots relative to the state value assessment, this invention employs a trace decay recursive form based on time-series differential residuals to construct an advantage function estimate, defined as:
[0160]
[0161] in, This represents the total number of time slots. For the summation index, The trace attenuation coefficient is... The temporal differential residual is jointly determined by the immediate reward and the state value assessments of adjacent time slots, weighted by a discount factor. This advantage function estimate provides a relative merit signal for actions in each time slot, thereby driving policy gradient updates. Simultaneously, the policy network employs a shearing objective function to achieve stable policy updates. The ratio of the conditional probability density of the old and new strategies at the same sampled action With advantage function estimation With this as the core, constraints are placed on the range of changes between the old and new strategies to achieve more stable strategy updates. This can be written as follows:
[0162]
[0163] in, For the current strategy in action The conditional probability density value at point In action with a fixed reference strategy The conditional probability density value at point The ratio is used to characterize the new policy relative to the fixed reference policy in the same sampled action. The degree of change in conditional probability density at a given point. Each update saves and fixes the current policy parameter vector as a stable reference policy, and... Limited to To improve training stability within the range, Let represent the pruning coefficient. Furthermore, the policy network introduces a policy entropy regularization term to maintain necessary exploration; therefore, the total cost function of the policy network can be expressed as:
[0164]
[0165] in, Describes a random variable representing joint actions. In the state Down The conditional probability density, The entropy regularity coefficient is... Represents policy entropy, which is minimized during training. Iterative update of parameter vector .
[0166] In practice, calculations are performed separately for each mini-batch of samples. and The gradient is calculated, and the policy network parameter vector is determined accordingly. With value network parameter vector Synchronous updates are performed. After traversing all mini-batch samples within the training sample batch, one iteration is completed. To fully utilize the training sample batch and improve sample utilization efficiency, multiple rounds of the above mini-batch sampling and gradient update traversal process are repeated. After the update is completed, the training sample batch is cleared, and new experience sequences are collected again with the environment under the updated parameter vector to enter the next round of update iteration. During the training phase, DRL-ACMRNet gradually converges to a better-performing joint optimization strategy for bandwidth, power, and UAV three-dimensional acceleration by repeatedly sampling and alternately updating the policy network parameter vector and the value network parameter vector.
[0167] After training, the final policy network parameter vector and value network parameter vector are obtained for online deployment. During the online deployment phase, no complex numerical optimization or policy updates are required. Each time slot only needs to input the current system state into the policy network under the trained policy network parameter vector and value network parameter vector to complete one forward inference. This yields the combined actions of bandwidth, power, and UAV three-dimensional acceleration that satisfy the constraints, achieving real-time maximization of end-to-end security rate on a time-slot basis.
[0168] Example 2:
[0169] This embodiment describes a computer-readable storage medium storing a computer program that, when executed by a processor, implements a UAV relay optoelectronic hybrid secure transmission method based on deep reinforcement learning as described in any of Embodiment 1.
[0170] Example 3:
[0171] This embodiment describes a computer device, including:
[0172] Memory is used to store instructions.
[0173] A processor is configured to execute the instructions, causing the computer device to perform operations as described in any of Embodiment 1 of a UAV relay optoelectronic hybrid secure transmission method based on deep reinforcement learning.
[0174] Example 4:
[0175] This embodiment introduces a simulation experiment of a UAV relay optoelectronic hybrid secure transmission method based on deep reinforcement learning, wherein the simulation scenario is set as a... In a three-dimensional region, a moving ground user moves at a constant speed. The starting three-dimensional coordinates of the ground user are (120, 180, 0), and the ending three-dimensional coordinates are (860, 720, 0). The three-dimensional coordinates of the eavesdropper are (800, 300, 60), and the three-dimensional coordinates of the ground base station are (1000, 1000, 0). The UAV takes off from the preset starting position (180, 180, 120). During flight, it receives FSO / RF signals sent by the ground user while simultaneously sending artificial interference signals to the eavesdropper. It also transmits confidential information back to the ground base station via the RF link, and finally returns to the starting point (180, 180, 120). The maximum transmit power of the ground user is 0.1W, and the minimum transmit power allocated to the RF link and FSO link for maintaining the pilot is 0.01W. The maximum acceleration of the UAV is 6. The maximum speed of the UAV is 15. The UAV has a maximum transmit power of 4W, a total bandwidth of 20MHz for the uplink and backlink RF links, and dynamically allocates uplink and backlink RF bandwidth in each time slot. The RF channel noise power is... W, the bandwidth of the FSO link is 200MHz, and the FSO channel noise power is W.
[0176] like Figure 2 As shown, the proposed Deep Reinforcement Learning-based UAV 3D Acceleration Control and Multi-Dimensional Resource Real-Time Joint Scheduling Network (DRL-ACMRNet) provides the 3D flight trajectory of a UAV within a mission cycle, characterizing the flight path and spatial position evolution of the UAV under obstacle occlusion and multiple constraints. This trajectory is generated by the policy network in each time slot based on the system state, producing an executable 3D acceleration vector. It is then iteratively updated time-slot by time-slot-level kinematic recursion, satisfying constraints such as maximum speed, maximum acceleration, energy budget, and flight altitude. The relative positional relationship between the UAV trajectory and the spatial distribution of obstacles demonstrates that the UAV can proactively adjust its heading and flight altitude in low-altitude obstacle environments to improve link line-of-sight conditions while simultaneously ensuring secure transmission performance and balancing flight energy consumption. This figure visually illustrates that the method of this invention can achieve feasible and adaptive UAV 3D acceleration optimization under multiple constraints, obtaining a UAV 3D trajectory that satisfies the constraints and optimizes the target performance.
[0177] like Figure 3 As shown, the training reward curve of the DRL-ACMRNet used in this invention changes with the number of training epochs when the user power is 0.1W. This reward is obtained by accumulating the real-time rewards of each time slot within each epoch. The reward fluctuates greatly in the early stage of training and rises rapidly overall. The reward increases from approximately Upgraded to approximately This indicates that the strategy is in the exploratory stage. As training progresses, the overall reward increases and gradually stabilizes, reaching a stable level near 5000 rounds. to This demonstrates that the strategy can learn better UAV three-dimensional acceleration control and resource allocation behavior under complex obstacle environments and multiple constraints, reflecting that the method proposed in this invention can effectively suppress training oscillations and policy collapse risks, and effectively improve training rewards.
[0178] like Figure 4 The figure shows the average end-to-end security rate variation curve of DRL-ACMRNet during training when the user power is 0.1W. The horizontal axis represents the training round number, and the vertical axis represents the average end-to-end security rate over the complete flight cycle within that round. Specifically, it is obtained by averaging the end-to-end security rate of each time slot over one round. As shown in the figure, the average end-to-end security rate is below 50Mbps in the early stages of training. As training progresses, the average end-to-end security rate shows an upward trend, stabilizing around 150Mbps at the end of training. This curve characterizes the learning process of the policy from initial exploration to gradual convergence. As training progresses, the policy gradually converges to a better-performing joint optimization scheme of real-time UAV 3D acceleration and resource allocation, thereby continuously improving the end-to-end security rate and ultimately reaching a stable performance level.
[0179] like Figure 5As shown, the outage probabilities of the proposed real-time UAV 3D acceleration and resource allocation joint optimization method based on DRL-ACMRNet, compared with four baseline schemes—unoptimized UAV 3D acceleration, unoptimized power allocation, unoptimized RF bandwidth allocation, and only RF link—are compared under different rate thresholds throughout the complete flight cycle. The results show that the outage probability of each scheme increases overall with increasing rate threshold, while the proposed method maintains the lowest outage probability across the entire threshold range with relatively gradual changes. When the rate threshold increases from 90 Mbps to 120 Mbps, the outage probability of the proposed method remains within a narrow range of 0.155-0.160, only slightly increasing to 0.170 at a rate threshold of 130 Mbps. In contrast, the outage probability of the unoptimized UAV 3D acceleration scheme is significantly higher and continues to rise to 0.665 with increasing rate threshold. Similarly, the outage probability of the unoptimized user and UAV power allocation scheme deteriorates with increasing rate threshold, rising sharply to 0.845 at high rate thresholds. The outage probability of the unoptimized RF bandwidth allocation scheme also increases significantly at a rate threshold of 130 Mbps. Only the RF link scheme achieves an outage probability of 0.840 at 100 Mbps, remaining constant at 1.000 at rate thresholds of 110 Mbps and above, which is almost unacceptable for medium-to-high security rate requirements. In summary, the method proposed in this invention maintains high reliability even under strict security rate requirements, demonstrating the significant gains of the real-time joint optimization method based on DRL-ACMRNet in improving end-to-end security rate and system security.
[0180] like Figure 6As shown, this paper compares the average end-to-end security rate of the proposed method based on DRL-ACMRNet for real-time UAV three-dimensional acceleration and resource allocation with four benchmark schemes: RF link only, no optimization of UAV three-dimensional acceleration, no optimization of power allocation, and no optimization of RF bandwidth allocation, over a complete flight cycle under different total user power conditions. System simulation and performance evaluation show that the proposed method consistently outperforms the benchmark schemes in average end-to-end security rate under all user power configurations, with a more significant advantage in the low-to-medium power range. Retaining three decimal places for the simulation results, the proposed method achieves average end-to-end security rates of 143.201 Mbps, 227.484 Mbps, 216.197 Mbps, and 229.361 Mbps over a complete flight cycle when user power is 0.1W, 0.2W, 0.3W, and 0.4W, respectively. When user power is 0.1W and 0.2W, the proposed method improves performance by 33.649Mbps and 100.261Mbps compared to the optimal baseline scheme, respectively, demonstrating that the proposed method can achieve higher security gains by performing real-time joint optimization of UAV three-dimensional acceleration, bandwidth allocation, and power allocation under power constraints. When user power is 0.3W and 0.4W, the proposed method still improves performance by 61.294Mbps and 79.840Mbps compared to the optimal baseline scheme, respectively, demonstrating that the proposed method can still suppress eavesdropping rates and reduce the risk of end-to-end security rate limitations caused by link performance differences in the high power range through real-time joint optimization. In summary, the real-time UAV three-dimensional acceleration and resource allocation joint optimization method based on DRL-ACMRNet significantly improves system security performance and resource utilization efficiency.
[0181] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0182] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0183] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0184] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0185] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A method for secure optical-electrical hybrid transmission in UAV relays based on deep reinforcement learning, characterized in that: Specifically, it includes: Step 1: Obtain the real-time joint optimization problem of UAV three-dimensional acceleration, bandwidth resources, user power, and UAV power configuration under multiple constraints; Step 2: Construct a Markov decision process based on the real-time joint optimization problem, and obtain the state vector of each time slot through the Markov decision process; Step 3: Input the state vector of each time slot into the real-time joint scheduling network to obtain the unconstrained original joint action vector; Step 4: Perform boundedness and feasibility mapping on the unconstrained original joint action vector to obtain the actual executable joint action vector; The expression for the real-time joint optimization problem is as follows: ; in, For discrete time slot indexes within a flight cycle, The length of each time slot, For the system's time slot-level end-to-end security rate, The security rate of the FSO / RF parallel link for ground users to send confidential information to UAVs. This refers to the RF link backhaul rate from the UAV to the ground base station. The total RF bandwidth of the system. For the RF bandwidth of the uplink from the ground user to the UAV, For the RF bandwidth of the UAV to ground base station backhaul link, This is the peak power of ground users. Electrical power allocated to the RF link for ground users. Electrical power allocated to FSO links for ground users The minimum transmit electrical power allocated to the RF link for ground users. The minimum transmit electrical power allocated to FSO links for ground users. Remaining allocable power for ground users For the peak power allowed by UAV, Return power allocated to UAV, Interference power allocated to UAVs For UAV, the current number Acceleration during flight in each time slot For UAV, the current number The speed of flight in each time slot For the speed of the UAV in the next time slot, The maximum acceleration of the UAV. This is the maximum speed of the UAV. For UAV No. The position of each time slot For UAV No. The position of each time slot For the first The remaining travel energy of the UAV in each time slot for Norm.
2. The UAV relay optoelectronic hybrid secure transmission method based on deep reinforcement learning according to claim 1, characterized in that: Also includes: Step 5: Calculate the end-to-end security rate of the real-time joint optimization problem based on the actual executable joint action vector, and calculate the instant reward based on the end-to-end security rate of the real-time joint optimization problem. The actual executable joint action vector is input into the real-time joint scheduling network to obtain the updated state vector; the state vector, the actual executable joint action vector, the immediate report, and the updated state vector are used as an empirical sequence. The gradients of the value loss function and the total cost function of the real-time joint scheduling network are calculated based on the empirical sequence. The parameter vector of the real-time joint scheduling network is synchronously updated based on the gradients of the value loss function and the total cost function to obtain the updated parameter vector of the real-time joint scheduling network. Replace the parameter vector in the real-time joint scheduling network in step 3 with the updated real-time joint scheduling network parameter vector, and repeat steps 1 to 4 to obtain the actual executable joint action vector.
3. The UAV relay optoelectronic hybrid secure transmission method based on deep reinforcement learning according to claim 1 or 2, characterized in that: The expression for the state vector is as follows: ; in, The coordinates of the UAV For UAV, the current number The speed of flight in each time slot The distance from the ground user to the UAV. The distance from the UAV to the eavesdropper. The distance from the ground user to the eavesdropper. This refers to the distance from the UAV to the ground base station. For the first The remaining travel energy of the UAV in each time slot For the line-of-sight status indicator variable of the link from the ground user to the UAV, This is a line-of-sight status indicator variable for the link from the UAV to the eavesdropper. For the line-of-sight status indicator variable of the link from the ground user to the eavesdropper, This is a line-of-sight status indicator variable for the UAV-to-ground base station link. When a Loss channel exists, , , , The line-of-sight status indicator variable is 1 if it is not 0 otherwise.
4. The UAV relay optoelectronic hybrid secure transmission method based on deep reinforcement learning according to claim 1 or 2, characterized in that: Step 3 specifically includes: Step 3.1: Place the first The state vector of each time slot Input the policy network of the real-time joint scheduling network, and output the state. Down mean vector In state Down mean vector In state Down variance vector And in state Down variance vector ; Step 3.2: Obtain the policy network parameter vector Down about conditional probability density Its specific form is as follows: ; in, Indicates a Gaussian distribution. Represented by the variance vector The resulting diagonal covariance matrix; Step 3.3: Obtain the policy network parameter vector Down about conditional probability density Its specific form is as follows: ; in, Represented by the variance vector The resulting diagonal covariance matrix; Step 3.4: According to the information regarding conditional probability density Sampling was performed According to conditional probability density Sampling was performed ,according to and Get the first The unconstrained original joint action vector of each time slot ,in, .
5. The UAV relay optoelectronic hybrid secure transmission method based on deep reinforcement learning according to claim 1 or 2, characterized in that: Step 4 specifically includes: Step 4.1: Based on the unconstrained original UAV three-dimensional acceleration Calculate the actual executable three-dimensional acceleration of the UAV. Its specific form is as follows: ; in, Indicated to Euclidean projection operation of the norm sphere Represents the hyperbolic tangent function; Step 4.2: Based on the proportional coefficient vector formed by the unconstrained original bandwidth allocation continuous proportional coefficient and the original power allocation continuous proportional coefficient. Calculate the proportional coefficient vector consisting of the actual executable bandwidth allocation continuous proportional coefficient and the actual executable power allocation continuous proportional coefficient. Specifically, it includes: , and Its specific form is as follows: ; in, This represents the continuous proportionality coefficient of the total RF bandwidth allocated to the ground user-to-UAV link in the system. ; in, A continuous scaling factor for the actual executable power allocation to the backhaul link for the UAV; ; in, The continuous scaling factor for the actual executable power allocation to RF links for ground users. It is the hard hyperbolic tangent function; Step 4.3, based on the actual executable UAV three-dimensional acceleration The proportional coefficient vector consisting of the actual executable bandwidth allocation continuous proportional coefficient and the actual executable power allocation continuous proportional coefficient. Obtain the actual executable joint action vector ,in, .
6. The UAV relay optoelectronic hybrid secure transmission method based on deep reinforcement learning according to claim 2, characterized in that: The end-to-end confidentiality rate of the real-time joint optimization problem is calculated based on the actual executable joint action vector, and the instant reward is calculated based on the end-to-end confidentiality rate of the real-time joint optimization problem. The actual executable joint action vector is input into the real-time joint scheduling network to obtain the updated state vector; The state vector, the actual executable joint action vector, the immediate reward, and the updated state vector are used as an empirical sequence, specifically including: Calculate the security rate of the FSO / RF parallel link for transmitting confidential information from a ground user to a UAV. Its specific form is as follows: ; in, For ground users Effective security rate of the RF link to UAV For ground users The rate of the FSO link to the UAV; Calculate UAV to ground base station RF link backhaul rate Its specific form is as follows: ; According to the secrecy rate and return rate Calculate end-to-end security rate Its specific form is as follows: ; Based on end-to-end security rate Calculate instant returns Its specific form is as follows: ; in, and These are the weighting coefficients. For UAV in time slots With speed Flight power during flight To avoid violating confidentiality restrictions, penalties apply. Based on the actual executable UAV three-dimensional acceleration To obtain the kinematic relationship and the unprojected velocity, by going to The Euclidean projection of the norm sphere maps it to the velocity feasible region of the next time slot of the UAV. ; Based on the current time slot position of the UAV To obtain the position of the next time slot of the UAV. ; Based on the position of the next UAV time slot Update the distance and line-of-sight status indicators of each link and the remaining navigation energy of the UAV to obtain the state vector of the next time slot as the updated state vector; use the state vector, the actual executable joint action vector, the instantaneous report, and the updated state vector as an empirical sequence.
7. The UAV relay optoelectronic hybrid secure transmission method based on deep reinforcement learning according to claim 2, characterized in that: The step involves calculating the gradients of the value loss function and the total cost function of the real-time joint scheduling network based on the empirical sequence, and then synchronously updating the parameter vector of the real-time joint scheduling network based on the gradients of the value loss function and the total cost function to obtain the updated parameter vector of the real-time joint scheduling network. Specifically, this includes: The preset number of experience sequences obtained from multiple runs are aggregated to form the training sample batch for the current policy update, and multiple rounds of small-batch random sampling are performed on them; for the sampled small batches, the value network of the real-time joint scheduling network is utilized. Output state value assessment Based on the state value assessment quantity and immediate rewards in the experience sequence Constructing Discounted Cumulative Returns ; Based on the state value assessment Cumulative returns with discounts Constructing a value loss function During training, the value loss function is minimized. Parameter vectors for realizing value networks Update; Construct time-series differential residuals using immediate returns from empirical sequences and state value assessments from adjacent time slots. Based on time-series difference residuals Constructing the advantage function estimation ; use The ratio of the conditional probability density of the new and old strategies at the same sampled action Constructing the shearing objective function ; Based on the shearing objective function Construct the total cost function of the policy network using the policy entropy regularization term. During training, the total cost function is minimized. Update the parameter vector of the policy network ; After traversing all mini-batch samples within the training sample batch, one round of iteration is completed. The above mini-batch sampling and parameter vector update traversal process is repeated multiple times until the updated real-time joint scheduling network parameter vector is obtained.
8. A computer-readable storage medium, characterized in that: It stores a computer program, which, when executed by a processor, implements a method for secure transmission of UAV relay optoelectronics based on deep reinforcement learning as described in any one of claims 1-7.
9. A computer device, characterized in that: include: Memory, used to store instructions; A processor is configured to execute the instructions, causing the computer device to perform the operation of a UAV relay optoelectronic hybrid secure transmission method based on deep reinforcement learning as described in any one of claims 1-7.
Citation Information
Patent Citations
CN116193476A
CN117750524A