Unmanned aerial vehicle optical communication link tracking and pointing method based on reinforcement learning

Through reinforcement learning methods, the optical communication link of the UAV is optimized, and the problems of high interruption rate and bit error rate under high-speed motion of the UAV are solved, and the high-precision alignment and stable transmission of optical signals are achieved, which is suitable for communication needs in a variety of complex environments.

CN120357981AActive Publication Date: 2025-07-22NORTHEASTERN UNIV CHINA

Patent Information

Application Number
CN202510631633.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-07-22
Estimated Expiration
2045-05-16

AI Technical Summary

Technical Problem

UAV optical communication systems face problems with high communication interruption rate and bit error rate in complex environments. The existing traditional control methods are difficult to cope with high-speed motion and external interference of drones, resulting in insufficient stability and reliability of optical communication links.

Method used

Using reinforcement learning methods based on deep deterministic strategy gradient algorithm (DDPG) and near-end strategy optimization algorithm (PPO), we use the reinforcement learning method based on the drone optical communication link channel model and agent interaction environment, define the state space, action space and reward functions, train the agent to generate the best adjustment strategy, and optimize the drone pod attitude to achieve beam alignment.

Benefits of technology

It significantly reduces the interrupt rate and bit error rate of the optical communication link of the drone, improves the accuracy of optical signal alignment and data transmission quality, enhances the adaptability in complex environments, and is suitable for scenarios such as integrated aerospace and earth communication and emergency rescue.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120357981A_ABST
    Figure CN120357981A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle optical communication link tracking and pointing method based on reinforcement learning, and relates to the technical field of unmanned aerial vehicle optical communication. The method specifically comprises the following steps: acquiring characteristic parameters of an unmanned aerial vehicle optical communication link and constructing an unmanned aerial vehicle optical communication link channel model; defining an intelligent agent for adjusting the attitude of the unmanned aerial vehicle pod, and establishing an interaction environment of the intelligent agent based on an unmanned aerial vehicle optical communication link channel model; based on an interaction environment of the intelligent agent, respectively defining a state space, an action space and a reward function of the intelligent agent; and selecting a depth deterministic strategy gradient (DDPG) algorithm or a near-end strategy optimization (PPO) algorithm to train an intelligent agent according to a given flight mission demand, and generating an optimal adjustment strategy of the attitude of the unmanned aerial vehicle pod. Based on the depth deterministic strategy gradient algorithm and the near-end strategy optimization algorithm, the interruption rate and the bit error rate of unmanned aerial vehicle optical communication in a high-dynamic environment are effectively reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of UAV optical communication, and particularly to a method for tracking and aiming an UAV optical communication link based on reinforcement learning. Background Art

[0002] In recent years, the rapid development of communication and information technology has promoted the continuous evolution of wireless communication networks. Especially in the context of the increasing demand for high-speed data transmission and information security, free space optical communication (FSO) has become an important research direction for future wireless communication technologies due to its characteristics such as high bandwidth, low interference, and high security. However, FSO technology still faces many challenges in complex environments. Especially in UAV-assisted optical communication systems, affected by factors such as the high-speed movement of UAVs, attitude changes, and external interference, the communication interruption rate and bit error rate are relatively high. Existing research mainly uses traditional control theory algorithms to optimize the acquisition, pointing, and tracking (APT) system of the optical beam. For example, methods such as PID control, Kalman filtering, and particle filtering have improved the control accuracy of the APT system to a certain extent. However, these methods still have significant defects in complex environments. For example, the PID-based control method relies on error signal feedback for adjustment and is difficult to cope with the sudden errors caused by the rapid maneuvering of UAVs. When facing sudden interference, atmospheric turbulence, or occlusion, the filtering accuracy of traditional filtering methods drops significantly, resulting in communication interruption. Although methods such as particle filtering improve the tracking accuracy, the computational cost is relatively large, and for resource-constrained UAV platforms, the real-time performance of particle filtering is difficult to meet the requirements.

[0003] Traditional model-based methods mainly rely on physical modeling, mathematical optimization, and classical control theory, such as technologies like Kalman filtering and particle filtering, to achieve the tracking and prediction of optical communication links. However, the core assumption of these methods is that the system dynamics and environmental interference characteristics can be accurately modeled, and optimal control is carried out based on this. However, in a highly dynamic environment where UAVs move at high speeds, optical signals are severely interfered, and channel attenuation is highly unstable, the expression ability of the model is often limited, and it is difficult to accurately describe complex non-linear dynamic characteristics, thus affecting the stability and reliability of optical communication links. In addition, constructing a high-precision mathematical model usually relies on a large amount of prior knowledge and experimental data, and needs to be continuously adjusted in different application scenarios, making it difficult to meet the requirements of real-time performance and self-adaptability. Summary of the Invention

[0004] Aiming at the deficiencies of the above-mentioned existing technologies, based on the Deep Deterministic Policy Gradient (DDPG) algorithm and the Proximal Policy Optimization (PPO) algorithm, the present invention proposes a follow-up and aiming method for an unmanned aerial vehicle (UAV) optical communication link based on reinforcement learning to reduce the interruption rate and bit error rate of UAV optical communication in a high-dynamic environment.

[0005] A follow-up and aiming method for an unmanned aerial vehicle (UAV) optical communication link based on reinforcement learning proposed by the present invention includes the following processes:

[0006] Obtain the characteristic parameters of the UAV optical communication link and construct a UAV optical communication link channel model;

[0007] Define an agent for adjusting the attitude of the UAV pod, and establish an interaction environment for the agent based on the UAV optical communication link channel model;

[0008] Based on the interaction environment of the agent, define the state space, action space, and reward function of the agent respectively;

[0009] Select the Deep Deterministic Policy Gradient (DDPG) algorithm or the Proximal Policy Optimization (PPO) algorithm according to the given flight mission requirements to train the agent and generate the optimal adjustment strategy for the attitude of the UAV pod;

[0010] Furthermore, the specific content of obtaining the characteristic parameters of the UAV optical communication link and constructing a UAV optical communication link channel model is as follows:

[0011] Obtain the characteristic parameters of the UAV optical communication link, including: UAV coordinates, user coordinates, visibility, laser wavelength, fading caused by atmospheric turbulence, and pointing alignment error;

[0012] Construct a UAV optical communication link channel model including several links, and represent the channel gain of any link in the UAV optical communication link channel model as:

[0013]

[0014] where h k (t) represents the channel gain of the k-th link at time slot t; η is the responsivity of the user end; represents the atmospheric attenuation of the k-th link at time slot t; represents the fading caused by atmospheric turbulence of the k-th link at time slot t; represents the pointing alignment error of the k-th link at time slot t;

[0015] The calculation method of the atmospheric attenuation of the k-th link at time slot t is as follows: At time slot t, the distance between the UAV and the user is calculated using the obtained UAV coordinates and user coordinates, the attenuation coefficient is calculated using the obtained visibility and laser wavelength, and then, based on the calculated distance between the UAV and the user and the attenuation coefficient, the atmospheric attenuation of the k-th link at time slot t is calculated using the Beer-Lambert law.

[0016] Furthermore, the specific content of establishing the interaction environment of the agent based on the UAV optical communication link channel model is as follows:

[0017] Simulate the beam tracking process between the UAV and the user in a complex environment, establish the interaction environment of the agent, which is used to generate the state information of the agent, and update the state information of the agent according to the action selected by the agent;

[0018] The interaction environment of the agent includes three parts: the UAV optical communication link channel model, the optical imaging simulation system, and the dynamic interference simulation model; among them, the UAV optical communication link channel model is used to feedback the reward value according to the state information of the agent; the optical imaging simulation system is used to map the obtained user coordinates to the camera coordinate system of the camera carried in the UAV pod and perform distortion correction, and use the corrected user coordinates as the user coordinates in the characteristic parameters of the UAV optical communication link; the dynamic interference simulation model is used to simulate different interference scenarios in a complex environment.

[0019] Furthermore, the specific content of defining the state space, action space, and reward function of the agent based on the interaction environment of the agent is as follows:

[0020] Define the state space of the agent to include: UAV attitude information, beam footprint pixel coordinates, and beacon point coordinates; among them, the UAV attitude information includes: pitch angle, roll angle, and yaw angle;

[0021] Define the action space of the agent to include: pod pitch angle and pod yaw angle;

[0022] Define the reward function as:

[0023]

[0024] where R(t) represents the reward value of the agent at time slot t; represents the square of the horizontal distance between the beam footprint pixel coordinates and the beacon point coordinates; represents the square of the vertical distance between the beam footprint pixel coordinates and the beacon point coordinates.

[0025] Further, the specific method for selecting the Deep Deterministic Policy Gradient (DDPG) algorithm or the Proximal Policy Optimization (PPO) algorithm to train the agent according to the given flight mission and generate the optimal adjustment strategy for the UAV pod attitude is as follows:

[0026] Based on the interaction environment of the agent, a 2-DOF rotating platform is used to simulate the alignment process of the UAV optical communication link;

[0027] An Actor-Critic framework is used to construct the network architecture for training the agent, including: an APT capturer and an APT evaluator, where the APT capturer includes an evaluation network and a target network; the APT evaluator includes an evaluation network and a target network;

[0028] According to the given flight mission, if the requirement for tracking accuracy in this flight mission is higher than the requirement for the response speed of the adjustment strategy of the UAV pod attitude, then the Deep Deterministic Policy Gradient (DDPG) algorithm is selected to train the agent; otherwise, the Proximal Policy Optimization (PPO) algorithm is selected to train the agent;

[0029] The trained agent is used to generate the optimal adjustment strategy for the UAV pod attitude.

[0030] Further, the specific method for selecting the Deep Deterministic Policy Gradient (DDPG) algorithm to train the agent is as follows:

[0031] The network parameters and weights of the evaluation network in the APT capturer and the evaluation network in the APT evaluator are randomly initialized respectively, and the weights of the evaluation network in the APT capturer are copied to the target network in the APT capturer, and the weights of the evaluation network in the APT evaluator are copied to the target network in the APT capturer;

[0032] The experience replay pool, the state space of the agent, and the action space of the agent are initialized, and iterative training is started;

[0033] At time slot t, the current state S(t) of the agent is obtained, the optimal action A(t) is selected using the evaluation network of the APT capturer, and the immediate reward R(t) of the agent is calculated using the reward function. The next state S(t+Δt) of the agent is updated according to the immediate reward R(t) of the agent, and then (S(t), A(t), R(t), S(t+Δt)) is stored in the experience replay pool as a set of interaction data; continue to use the evaluation network of the APT capturer to select the optimal action for the next state S(t+Δt) of the agent until there are several sets of interaction data stored in the experience replay pool;

[0034] Set the total number of interactive data M for each round of iterative training, initialize the number of interactive data m for the current iteration round, randomly select a set of interactive data from the experience replay pool, increment the number of interactive data for the current iteration round by 1, and calculate the state-action value function in a way that maximizes the reward as the expected cumulative reward of the agent;

[0035] Based on the expected cumulative reward of the agent, construct the loss function of the evaluation network of the APT evaluator, and update the network parameters of the evaluation network of the APT evaluator by minimizing the value of the loss function of the evaluation network of the APT evaluator;

[0036] According to the chain rule, calculate the gradient of the objective function J with respect to the network parameters of the evaluation network of the APT capturer and update the network parameters of the evaluation network of the APT capturer by gradient ascent;

[0037] According to the updated network parameters of the evaluation network of the APT evaluator and the updated network parameters of the evaluation network of the APT capturer, update the network parameters of the target network of the APT evaluator and the network parameters of the target network of the APT capturer using the soft update mechanism;

[0038] Judge whether m > M holds. If not, return to the step of randomly selecting a set of interactive data from the experience replay pool to continue training, and increment the number of interactive data for the current iteration round by 1; if so, judge whether the current iteration round is greater than the preset number of training rounds. If so, stop the iterative training; if not, re-initialize the experience replay pool, the state space of the agent, and the action space of the agent, and start the next round of iterative training.

[0039] Furthermore, the method for using the evaluation network of the APT capturer to select the optimal action A(t) is as follows: first use the evaluation network of the APT capturer to select the current action, and then use the Ornstein-Uhlenbeck process (OU process) to add noise perturbation to the current action to generate the optimal action A(t) to be executed, expressed as:

[0040] A(t) = μ′(S(t)) = μ(S(t)|θ μ ) + N

[0041] where μ represents the evaluation network of the APT capturer; θ μ represents the network parameters of the evaluation network of the APT capturer; μ(S(t)|θ μ ) represents the policy for selecting actions using the evaluation network of the APT capturer; N represents the noise; μ′(S(t)) represents the policy for selecting actions using the evaluation network of the APT capturer after adding noise perturbation.

[0042] Furthermore, the state-action value function is expressed as:

[0043] Q(S(t), A(t)) = E R(t),S(t+Δt)~ E[R(S(t), A(t)) + γQ(S(t + Δt), μ(S(t + Δt)))]

[0044] Where Q(S(t), A(t)) represents the Q - value of selecting the optimal action in the current state S(t), that is, it is used to evaluate the expected cumulative reward of executing the optimal action A(t) in the current state S(t); R(S(t), A(t)) represents the immediate reward obtained after executing the optimal action A(t) in the current state S(t); γ is the discount factor, and γ ∈ [0, 1]; Δt represents the time interval; S(t + Δt) represents the next state of the agent; Q(S(t + Δt), μ(S(t + Δt))) represents the Q - value of selecting an action according to the evaluation network μ of the APT capturer in the next state of the agent; E R(t),S(t+Δt)~ E represents the expectation of the immediate reward and the next state;

[0045] The loss function of the evaluation network of the APT evaluator is expressed as:

[0046] L(θ Q ) = E[(Q(S(t), A(t)|θ Q ) - y t ) 2

[0047] Where L(θ Q ) represents the value of the loss function of the evaluation network of the APT evaluator; E[·] represents calculating the expected value; θ Q represents the network parameters of the evaluation network of the APT evaluator; Q(S(t), A(t)|θ Q ) represents the Q - value predicted by the evaluation network of the APT evaluator in the current state S(t) and the optimal action A(t); y t represents the fitting value function of the target network of the APT evaluator, and there is:

[0048] y t = R(S(t), A(t)) + γQ′(S(t + Δt), μ′(S(t + Δt)|θ μ′ )|θ Q′ )

[0049] Where θ Q′ represents the network parameters of the target network of the APT evaluator; θ μ′ represents the network parameters of the target network of the APT capturer; μ′ represents the target network of the APT capturer; Q' represents the target network of the APT evaluator; μ′(S(t + Δt)|θ μ′ ) represents the action selection of the target network of the APT evaluator according to the next state; ​

[0050] The calculation method of the gradient of the objective function J with respect to the network parameters of the evaluation network of the APT capturer is as follows:

[0051]

[0052] where denotes the gradient of the objective function J with respect to the network parameters of the evaluation network of the APT capturer; μ(S(t)|θ μ ) represents the optimal action A(t) selected by the evaluation network of the APT capturer with network parameters θ μ at the current state S(t).

[0053] Furthermore, the specific method for training the intelligent agent using the Proximal Policy Optimization (PPO) algorithm is as follows:

[0054] Randomly initialize the network parameters of the APT capturer and the network parameters of the APT evaluator;

[0055] Initialize the state space and action space of the intelligent agent, and start iterative training;

[0056] At time slot t, obtain the current state S(t) of the intelligent agent, select the optimal action A(t) using the evaluation network of the APT capturer, calculate the immediate reward R(t) of the intelligent agent using the reward function, update the next state S(t + Δt) of the intelligent agent according to the immediate reward R(t) of the intelligent agent, and use (S(t), A(t)) as a state-action pair. Repeat the processes of action selection, immediate reward calculation, and state update, and save several groups of state-action pairs;

[0057] For time slot t, according to the state-action pair (S(t), A(t)) at time slot t, construct the advantage function using the Generalized Advantage Estimation (GAE), and construct the loss function of the APT capturer using the advantage function;

[0058] Calculate the target state value according to the state-action pair (S(t), A(t)) at time slot t, and construct the loss function of the APT evaluator;

[0059] Update the network parameters of the APT capturer by maximizing the loss function of the APT capturer, and update the network parameters of the APT evaluator by minimizing the loss function of the APT evaluator;

[0060] Extract a fixed batch of state-action pairs, update the network parameters of the APT capturer using the clipped policy gradient method according to the loss function of the APT capturer, and at the same time update the network parameters of the APT evaluator by minimizing the loss function of the APT evaluator to complete one iteration of training;

[0061] Determine whether the current iteration round is greater than the preset number of training rounds. If so, stop the iterative training; if not, re-initialize the state space and action space of the agent, and start the next round of iterative training.

[0062] Furthermore, the advantage function is expressed as:

[0063]

[0064] where is the advantage estimator; δ(t) is the TD-error at time slot t; γ1 represents the discount factor; λ p is the GAE factor; δ(t+Δt) is the TD-error from time slot t to time slot t+Δt; T represents the termination time step; δ(T-Δt) represents the TD-error from T to time slot T-Δt;

[0065] The loss function of the APT capturer is expressed as:

[0066]

[0067] where represents the total loss of the APT capturer; represents the set of network parameters of the APT capturer; E t represents the expectation of the empirical sample; η1 is a fixed coefficient; represents the policy entropy; represents the clipped policy loss, and there is:

[0068]

[0069] where clip(·) is the clipping function; ∈ p is a hyperparameter; is the probability ratio, and there is:

[0070]

[0071] where is the policy adopted when selecting actions using the updated APT capturer; is the policy adopted when selecting actions using the APT capturer before update; note h t (θ old ) = 1;

[0072] The loss function of the APT evaluator is:

[0073]

[0074] where represents the value regression loss of the APT evaluator; E tDenote the expectation of the empirical sample; Denote the state-value function with parameter φ p ; Denote the target state value.

[0075] The beneficial effects produced by adopting the above technical solution are as follows:

[0076] Compared with traditional control methods such as PID and Kalman filtering, the APT algorithm based on reinforcement learning proposed by the method of the present invention has the following advantages in the beam alignment task of the UAV FSO communication link:

[0077] (1) Improve the alignment accuracy of optical signals: The method of the present invention uses deep reinforcement learning to dynamically adjust the beam direction, minimize the beam deviation, improve the alignment accuracy, and reduce the communication interruption rate. Experimental results show that the APT algorithm can reduce the interruption rate of the FSO optical communication link from 40.86% to 1%. Improve signal stability.

[0078] (2) Improve the data transmission quality: The method of the present invention optimizes the APT scheme through reinforcement learning and controls the bit error rate within 10 -6 .

[0079] (3) Enhance the adaptability to complex environments: The method of the present invention can adaptively adjust the beam alignment strategy and improve the stability of UAV FSO communication in complex scenarios.

[0080] (4) Wide application scenarios: The method of the present invention can be widely applied to scenarios such as space-air-ground integrated communication, emergency rescue, and UAV relay networks, significantly improving the reliability and practicality of FSO optical communication. Description of the Drawings

[0081] Figure 1 is the flowchart of the APT optimization method for the UAV optical communication link based on reinforcement learning in this embodiment;

[0082] Figure 2 is the flowchart of selecting the Deep Deterministic Policy Gradient (DDPG) algorithm to train the agent in this embodiment;

[0083] Figure 3 is the flowchart of selecting the Proximal Policy Optimization (PPO) algorithm to train the agent in this embodiment. Detailed Embodiments

[0084] To facilitate the understanding of the present application, the following further describes in detail the specific embodiments of the present invention in conjunction with the drawings and embodiments. The following embodiments are used to illustrate the present invention but are not used to limit the scope of the present invention. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thorough and comprehensive.

[0085] Compared with the traditional model - based method, this embodiment uses a model - free reinforcement learning algorithm to provide a new solution for the optimization of the UAV optical communication link. Since reinforcement learning does not rely on precise mathematical modeling, but autonomously explores the optimal strategy through dynamic interaction with the environment, it can achieve efficient alignment and adaptive optimization of the optical communication link in a complex environment. This characteristic makes the reinforcement learning method more advantageous in the face of non - linear, uncertain, and highly dynamic scenarios, providing a highly robust control scheme for the UAV optical communication system.

[0086] In a highly dynamic environment such as the UAV optical communication link, there are problems of difficult data acquisition, rapid environmental changes, and high real - time requirements. Therefore, there is an urgent need for an efficient reinforcement learning algorithm that can achieve accurate beam alignment in a complex, unstructured, and highly dynamic environment to meet the real - time application requirements of the UAV optical communication system.

[0087] A follow - up and aiming method for the UAV optical communication link based on reinforcement learning in this embodiment, as Figure 1 shown, this method includes the following processes:

[0088] Obtain the characteristic parameters of the UAV optical communication link and construct a channel model for the UAV optical communication link.

[0089] The specific content of obtaining the characteristic parameters of the UAV optical communication link and constructing a channel model for the UAV optical communication link is as follows:

[0090] Obtain the characteristic parameters of the UAV optical communication link, including: UAV coordinates, user coordinates, visibility, laser wavelength, fading caused by atmospheric turbulence, and pointing alignment error.

[0091] In this embodiment, the pixel coordinates of the beam footprint and the fiducial point coordinates are collected in real - time by using a camera or sensor mounted on the UAV pod. For example, the pixel coordinates of the beam footprint and the fiducial point coordinates at a certain moment can be obtained from the photos taken by the camera in real - time.

[0092] Construct a channel model for the UAV optical communication link including several links, and express the channel gain of any link in the UAV optical communication link channel model as:

[0093]

[0094] where h k (t) represents the channel gain of the k - th link at time slot t; η is the responsivity of the user end; represents the atmospheric attenuation of the k - th link at time slot t; represents the fading of the k - th link caused by atmospheric turbulence at time slot t; Denote the directivity alignment error of the k-th link at time slot t;

[0095] The calculation method of the atmospheric attenuation of the k-th link at time slot t is as follows: For time slot t, calculate the distance d between the UAV and the user using the obtained UAV coordinates and user coordinates k ( t ), calculate the attenuation coefficient β using the obtained visibility and laser wavelength, and then according to the calculated distance d between the UAV and the user k ( t ) and the attenuation coefficient β, use Beer-Lambert's law to calculate the atmospheric attenuation of the k-th link at time slot t, expressed as:

[0096]

[0097] where e is the base of the natural logarithm; d k ( t ) represents the distance between the UAV and the user, which is calculated using the UAV coordinates and user coordinates; β represents the attenuation coefficient, with the unit of [m -1 , and there is:

[0098] β = (β dB log(10)) / 10 4 [m -1 (3)

[0099] where β dB is the absorption loss, with the unit of [dB / km], and there is:

[0100] β dB = (3.91 / V)(λ / 550[nm]) -p [dB / km] (4)

[0101] where 550[nm] represents the reference wavelength of 550 nanometers; p represents the size distribution coefficient; V represents the visibility; λ represents the laser wavelength.

[0102] In this embodiment, the fading caused by atmospheric turbulence and the directivity alignment error can both be obtained by existing methods.

[0103] Define an agent for adjusting the attitude of the UAV pod, and establish an interactive environment for the agent based on the UAV optical communication link channel model.

[0104] In this embodiment, an agent is defined to generate control instructions for adjusting the attitude of the UAV pod. Therefore, the interaction environment of the agent is constructed with the UAV optical communication link channel model as the core. The channel model reflects the physical characteristics of the optical signal during its propagation in the air, including the channel bandwidth, the received - end noise power, and the path loss coefficient. The agent changes the link conditions by adjusting the pod attitude in the interaction environment, and the interaction environment feedbacks rewards according to the UAV optical communication link channel model, thereby guiding the agent to optimize the action strategy.

[0105] The specific content of establishing the interaction environment of the agent based on the UAV optical communication link channel model is as follows:

[0106] Simulate the beam tracking process between the UAV and the user in a complex environment, establish the interaction environment of the agent, which is used to generate the state information of the agent, and update the state information of the agent according to the actions selected by the agent.

[0107] The interaction environment of the agent includes three parts: the UAV optical communication link channel model, the optical imaging simulation system, and the dynamic interference simulation model. Among them, the UAV optical communication link channel model is used to feedback rewards according to the state information of the agent; the optical imaging simulation system is used to map the obtained user coordinates to the camera coordinate system of the camera carried in the UAV pod and perform distortion correction, and use the corrected user coordinates as the user coordinates in the characteristic parameters of the UAV optical communication link.

[0108] In this embodiment, simulate the beam tracking process of the UAV in a complex environment. Establish an optical imaging simulation system composed of camera internal parameters, camera external parameters, UAV motion trajectory data, and user motion trajectory data, which is used to map user coordinates to the camera coordinate system, considering factors such as lens distortion and imaging noise to enhance the authenticity of visual perception.

[0109] The dynamic interference simulation model is used to simulate different interference scenarios in a complex environment.

[0110] In this embodiment, the simulation environment needs to include different interference scenarios, such as atmospheric turbulence, loss of optical signals caused by building occlusion, UAV jitter, and high - dynamic movement of the user, etc., to enhance the adaptability of the algorithm and improve the robustness of the optical communication link in different scenarios.

[0111] The construction method of the optical imaging simulation system is as follows: for a UAV with a camera carried in the pod, establish an optical imaging simulation system composed of camera internal parameters, camera external parameters, UAV motion trajectory data, and user motion trajectory data.

[0112] In this embodiment, the UAV motion trajectory data is time-series information describing the position, attitude, and motion state of the UAV in three-dimensional space, including: the coordinate position of the UAV in the world coordinate system, pitch angle, roll angle, yaw angle, the speeds of the UAV in three directions, and the time stamp corresponding to each data point. The user motion trajectory data is time-series information describing the position and motion state of the target user in three-dimensional space, including: the coordinate position of the user in the world coordinate system, speed, and the time stamp corresponding to each data point.

[0113] Based on the agent-based interaction environment, the state space, action space, and reward function of the agent are defined respectively.

[0114] In this embodiment, the state space and action space of reinforcement learning are defined, and a reward function is designed to optimize the stability of the UAV optical communication link.

[0115] The specific contents of defining the state space, action space, and reward function of the agent in the agent-based interaction environment are as follows:

[0116] The defined state space of the agent includes: UAV attitude information, beam footprint pixel coordinates, and beacon point coordinates; where the UAV attitude information includes: pitch angle, roll angle, and yaw angle.

[0117] The acquisition methods of the beam footprint pixel coordinates and beacon point coordinates are: real-time acquisition of the photos taken by the camera carried in the UAV pod, and extraction of the beam footprint pixel coordinates and beacon point coordinates in each photo.

[0118] In this embodiment, by defining the state space of the agent, it is ensured that the agent can comprehensively perceive the communication link state.

[0119] The defined action space of the agent includes: pod pitch angle and pod yaw angle.

[0120] The reward function is defined as:

[0121]

[0122] where \(R(t)\) represents the immediate reward of the agent at time slot \(t\); represents the square of the horizontal distance between the beam footprint pixel coordinates and the beacon point coordinates; represents the square of the vertical distance between the beam footprint pixel coordinates and the beacon point coordinates.

[0123] In this embodiment, the defined reward function reduces the bit error rate and the communication interruption rate through training.

[0124] Select the Deep Deterministic Policy Gradient (DDPG) algorithm or the Proximal Policy Optimization (PPO) algorithm according to the given flight mission requirements to train the agent, and generate the optimal adjustment strategy for the attitude of the drone pod.

[0125] The specific method for selecting the Deep Deterministic Policy Gradient (DDPG) algorithm or the Proximal Policy Optimization (PPO) algorithm according to the given flight mission to train the agent and generate the optimal adjustment strategy for the attitude of the drone pod is as follows:

[0126] Based on the interaction environment of the agent, a 2-DOF rotating platform is used to simulate the alignment process of the drone optical communication link.

[0127] Use the Actor-Critic framework to construct the network architecture for training the agent, including: an APT capturer and an APT evaluator, where the APT capturer includes an evaluation network and a target network; the APT evaluator includes an evaluation network and a target network.

[0128] According to the given flight mission, if the requirement for tracking accuracy is higher than the requirement for the response speed of the adjustment strategy of the drone pod attitude, then select the Deep Deterministic Policy Gradient (DDPG) algorithm to train the agent; otherwise, select the Proximal Policy Optimization (PPO) algorithm to train the agent.

[0129] In this embodiment, the Deep Deterministic Policy Gradient algorithm (DDPG) is applicable to high-precision and low-error tasks; the Proximal Policy Optimization algorithm (PPO) is applicable to tasks with fast-changing environments and the need for rapid strategy adjustment.

[0130] In this embodiment, as Figure 2As shown in the figure, APT reinforcement learning training is carried out based on the DDPG algorithm. Specifically, the deep deterministic policy gradient algorithm DDPG is used to optimize the control of the APT process of the UAV optical communication link, and the decision-making ability of reinforcement learning in the continuous action space is utilized to enable the optical communication link to achieve adaptive adjustment in a high-dynamic environment. The training process of DDPG consists of an agent, an environment model, a state input, an action output, and a reward feedback. Among them, the state space includes the pixel coordinates of the beam footprint, the UAV attitude, and the beacon point coordinates, etc., and the action space covers the pod pitch angle and the pod yaw angle. The agent makes a control instruction decision according to the current state through the policy network, that is, the APT capturer, to adjust the optical axis direction of the optical communication link. After receiving this action, the environment updates the state and calculates the immediate reward value based on the picture pixel coordinates and the communication bit error rate. When the reward is high, the policy tends to be stable, and when the bit error rate rises or the link is interrupted, a penalty is given to optimize the decision-making direction. The value network, that is, the APT evaluator, is used to evaluate the long-term return of the current action, guide the policy to converge, and enable the agent to gradually learn the optimal beam tracking strategy. During the training process, DDPG uses experience replay to reduce data correlation and adopts a target network for stable update to prevent unstable policy convergence. In addition, the network parameters of the APT capturer are gradually adjusted through soft update to improve the learning stability and enhance the adaptability of the algorithm in a high-dynamic environment. Through this closed-loop reinforcement learning framework, the agent can gradually optimize the beam alignment strategy of the optical communication link, achieve the minimization of the bit error rate, the reduction of the link interruption rate, and improve the robustness of the communication system.

[0131] The specific method for training the agent by selecting the deep deterministic policy gradient DDPG algorithm is as follows:

[0132] Randomly initialize the network parameters and weights of the evaluation network in the APT capturer and the evaluation network in the APT evaluator respectively, and copy the weights of the evaluation network in the APT capturer to the target network in the APT capturer, and copy the weights of the evaluation network in the APT evaluator to the target network in the APT capturer.

[0133] Initialize the experience replay pool, the state space of the agent, and the action space of the agent.

[0134] At time slot t, obtain the current state S(t) of the agent, select the optimal action A(t) using the evaluation network of the APT capturer, and calculate the immediate reward R(t) of the agent using the reward function. Update the next state S(t + Δt) of the agent according to the immediate reward R(t) of the agent, and then store (S(t), A(t), R(t), S(t + Δt)) as a set of interaction data in the experience replay pool; continue to select the optimal action of the next state S(t + Δt) of the agent using the evaluation network of the APT capturer until there are several sets of interaction data stored in the experience replay pool.

[0135] The method for the evaluation network of the APT capturer to select the optimal action A(t) is as follows: first, use the evaluation network of the APT capturer to select the current action, and then use the Ornstein-Uhlenbeck (OU) process to add noise perturbation to the current action to generate the optimal action A(t) to be executed, which is expressed as:

[0136] A(t) = μ′(S(t)) = μ(S(t)|θ μ ) + N (6)

[0137] where μ represents the evaluation network of the APT capturer; θ μ represents the network parameters of the evaluation network of the APT capturer; μ(S(t)|θ μ ) represents the strategy for selecting actions using the evaluation network of the APT capturer; N represents noise; μ′(S(t)) represents the strategy for selecting actions using the evaluation network of the APT capturer after adding noise perturbation.

[0138] In this embodiment, in order to enhance the strategy exploration ability, the Ornstein-Uhlenbeck (OU) process is introduced to perturb the actions.

[0139] In this embodiment, the parameters of the APT capturer and the APT evaluator are randomly initialized. The target network is initialized so that its parameters are synchronized with the current network. The experience pool is cleared to prepare for storing sampled data. The environmental state is collected, and actions are selected according to the strategy. The actions are executed, rewards are obtained, and the environmental state is updated. The experience data is stored and mini-batch sampled, and the experience replay mechanism is adopted to enable the agent to learn effective strategies from historical data and improve the training stability. At the same time, the target network is used for stable strategy update to reduce the fluctuations during training, and soft update is adopted to reduce model oscillation and improve the learning efficiency. The parameters of the target network are adjusted by soft update until the average reward value reaches the expectation and then the training stops.

[0140] Set the total number of interaction data M for each round of iterative training, and initialize the number of interaction data m for the current iteration round. Randomly extract a set of interaction data from the experience replay pool, and increment the number of interaction data for the current iteration round by 1. Calculate the state-action value function in a way that maximizes the return as the expected cumulative reward of the agent.

[0141] The state-action value function is expressed as:

[0142] Q(S(t), A(t)) = E R(t),S(t+Δt)~ E[R(S(t), A(t)) + γQ(S(t + Δt), μ(S(t + Δt)))] (7)

[0143] Among them, Q(S(t), A(t)) represents the Q-value of selecting the optimal action in the current state S(t), that is, it is used to evaluate the expected cumulative reward of executing the optimal action A(t) in the current state S(t); R(S(t), A(t)) represents the immediate reward obtained after executing the optimal action A(t) in the current state S(t); γ is the discount factor, and γ ∈ [0, 1]; Δt represents the time interval; S(t + Δt) represents the next state of the agent; Q(S(t + Δt), μ(S(t + Δt))) represents the Q-value of selecting an action according to the evaluation network μ of the APT capturer in the next state of the agent; E R(t),S(t+Δt)~ E represents the expectation of the immediate reward and the next state, that is, taking the average of the two.

[0144] According to the expected cumulative reward of the agent, construct the loss function of the evaluation network of the APT evaluator, and update the network parameters of the evaluation network of the APT evaluator by minimizing the value of the loss function of the evaluation network of the APT evaluator;

[0145] The loss function of the evaluation network of the APT evaluator is expressed as:

[0146] L(θ Q ) = E[(Q(S(t), A(t)|θ Q ) - y t ) 2 (8)

[0147] Among them, L(θ Q ) represents the value of the loss function of the evaluation network of the APT evaluator; E[·] represents calculating the expected value; θ Q represents the network parameters of the evaluation network of the APT evaluator; Q(S(t), A(t)|θ Q ) represents the Q-value predicted by the evaluation network of the APT evaluator in the current state S(t) and the optimal action A(t); y t represents the fitting value function of the target network of the APT evaluator. Although y t also depends on θ Q , it is usually ignored, and there is:

[0148] y t = R(S(t), A(t)) + γQ'(S(t + Δt), μ'(S(t + Δt)|θ μ′ )|θ Q′ ) (9)

[0149] Among them, θ Q′ represents the network parameters of the target network of the APT evaluator; θ μ′The network parameters of the target network of the APT capturer; μ′ represents the target network of the APT capturer; Q' represents the target network of the APT evaluator; μ′(S(t+Δt)|θ μ′ ) represents that the target network of the APT evaluator makes an action selection according to the next state.

[0150] According to the chain rule, calculate the gradient of the objective function J with respect to the network parameters of the evaluation network of the APT capturer and update the network parameters of the evaluation network of the APT capturer through gradient ascent;

[0151] The calculation method of the gradient of the objective function J with respect to the network parameters of the evaluation network of the APT capturer is as follows:

[0152]

[0153] where represents the gradient of the objective function J with respect to the network parameters of the evaluation network of the APT capturer; μ(S(t)|θ μ ) represents that in the current state S(t), the optimal action A(t) selected by the evaluation network of the APT capturer with network parameters θ μ .

[0154] In this embodiment, for time slot t, the APT capturer can be updated to the expected return from the starting distribution for the APT capturer parameters according to the chain rule, that is, use the policy gradient theorem to optimize the network parameters of the APT capturer to maximize the long-term expected return J.

[0155] According to the updated network parameters of the evaluation network of the APT evaluator and the updated network parameters of the evaluation network of the APT capturer, use a soft update mechanism to update the network parameters of the target network of the APT evaluator and the network parameters of the target network of the APT capturer;

[0156] In this embodiment, in order to reduce the frequent update of online network parameters, DDPG adopts a soft update strategy and stores experience data.

[0157] The method of using a soft update mechanism to update the network parameters of the target network of the APT evaluator and the network parameters of the target network of the APT capturer is as follows:

[0158]

[0159] where θ μ0 represents the updated network parameters of the target network of the APT capturer; θ Q0 represents the updated network parameters of the target network of the APT evaluator.

[0160] Judge whether m > M holds. If not, return to the step of randomly extracting a set of interaction data from the experience replay pool to continue training, and increment the amount of interaction data for the current iteration round by 1. If so, judge whether the current iteration round is greater than the preset number of training rounds. If so, stop iterative training. If not, re-initialize the experience replay pool, the state space of the agent, and the action space of the agent, and start the next round of iterative training.

[0161] Through the above training process, the DDPG agent gradually learns to output alignment control actions that maximize the reward in different states, reducing the bit error rate and outage rate. The simulation results show that the DDPG algorithm can effectively solve the alignment tracking problem between the UAV and the ground user. During the training process, the total reward value gradually increases and tends to be stable, the cumulative reward curve is relatively smooth, and the convergence performance is good. Compared with other algorithms, the strategy obtained by DDPG shows higher link stability in a high-speed moving environment.

[0162] This embodiment uses the Proximal Policy Optimization algorithm to achieve APT control. The PPO algorithm also uses the Actor-Critic framework, but its policy network directly outputs the action probability distribution, which is applicable to discrete or continuous actions. The continuous actions are modeled by the Gaussian distribution, and the value network estimates the state value. The structure of the PPO agent is similar to that of the DDPG agent, but the method of updating the policy is different. As Figure 3 shown, the training process of the APT algorithm based on PPO includes the following loop: First, initialize the parameters of the policy network and the value network, as well as the environmental state. Then, repeatedly perform the following operations until convergence: (1) Interact with the environment to collect the state, action, and reward sequences for a certain number of time steps, and save the trajectory data. (2) In each round of update, calculate the advantage estimate and the target state value for each time step according to the collected data. (3) Use the above data to perform gradient updates on the value network and the policy network respectively, and iterate multiple times: the value network minimizes the state value loss, and the policy network updates the parameters by the clipped policy gradient method shown in the formula. (4) After updating a certain batch of data, enter the next round of interaction with the environment. Through this proximal policy constraint update method, PPO can gradually improve the policy performance and maintain stable convergence during training. The simulation results show that the PPO algorithm can also effectively achieve the optical link alignment control between the UAV and the terminal, and its cumulative reward continuously improves with the training iteration. Compared with DDPG, the fluctuation of PPO may be slightly larger in the initial stage of training, but after appropriately adjusting hyperparameters such as the batch size and learning rate, good convergence results and improved link stability can also be achieved.

[0163] The specific method for selecting the Proximal Policy Optimization PPO algorithm to train the agent is as follows:

[0164] Randomly initialize the network parameters of the APT capturer and the network parameters of the APT evaluator.

[0165] Initialize the state space and action space of the agent, and start iterative training.

[0166] At time slot t, obtain the current state S(t) of the agent, select the optimal action A(t) using the evaluation network of the APT capturer, calculate the immediate reward R(t) of the agent using the reward function, update the next state S(t+Δt) of the agent according to the immediate reward R(t) of the agent, and take (S(t),A(t)) as a state-action pair. Repeat the processes of action selection, immediate reward calculation, and state update, and save several groups of state-action pairs.

[0167] In this embodiment, according to the action executed by selecting according to the current policy, i.e., the evaluation network of the APT capturer, execute the action and obtain the reward and the new state. Store the state transition data, i.e., the state-action pair, for training.

[0168] For time slot t, according to the state-action pair (S(t),A(t)) at time slot t, use Generalized Advantage Estimation (GAE) to construct the advantage function, which is expressed as:

[0169]

[0170] where is the advantage estimator; δ(t) is the TD-error at time slot t; γ1 represents the discount factor; λ p is the GAE factor; δ(t+Δt) is the TD-error from time slot t to time slot t+Δt; T represents the termination time step; δ(T-Δt) represents the TD-error from T to time slot T-Δt.

[0171] In this embodiment, use Generalized Advantage Estimation (GAE) to optimize the policy gradient and improve the stability of beam tracking. Calculate the state value function and dynamically update the policy network to enhance the adaptability of the agent to different optical signal environments, and ensure that the reinforcement learning can operate stably in complex optical communication channels.

[0172] Use the advantage function to construct the loss function of the APT capturer composed of the clipped policy loss and the entropy reward, which is expressed as:

[0173]

[0174] where represents the total loss of the APT capturer; represents the set of network parameters of the APT capturer; E t represents the expectation of the empirical sample, i.e., take the average over all time steps in the current batch of data; η1 is a fixed coefficient; Represents the addition of policy entropy, mainly used for the action exploration of the agent to prevent entering the local optimum; Represents the clipped policy loss, used for more effective policy update, and there is:

[0175]

[0176] where clip(·) is the clipping function; ∈ p is a hyperparameter, usually set to ∈ p = 0.2; is the probability ratio, and there is:

[0177]

[0178] where is the policy adopted when selecting actions using the updated APT capturer; is the policy adopted when selecting actions using the APT capturer before update; note h t (θ old ) = 1.

[0179] In this embodiment, the second term in expression (14) restricts the large difference between the new policy update and the old policy update.

[0180] Calculate the target state value according to the state-action pair (S(t), A(t)) at time slot t, and construct the loss function of the APT evaluator as:

[0181]

[0182] where represents the value regression loss of the APT evaluator; E t represents the expectation of the empirical sample, that is, taking the average over all time steps in the current batch of data; represents the state-value function with parameter φ p ; represents the target state value.

[0183] Extract a fixed batch of state-action pairs, and update the network parameters of the APT capturer using the clipped policy gradient method according to the loss function of the APT capturer. At the same time, update the network parameters of the APT evaluator by minimizing the loss function of the APT evaluator to complete one iteration of training.

[0184] Judge whether the current iteration round is greater than the preset number of training rounds or whether the current training process meets the convergence condition. If the current iteration round is greater than the preset number of training rounds or the current training process meets the convergence condition, stop the iterative training; otherwise, re-initialize the state space and action space of the agent and start the next round of iterative training.

[0185] The convergence condition is the convergence of the loss function of the APT capturer or the convergence of the loss function of the APT evaluator.

[0186] In this embodiment, when the loss function converges and the policy update cannot significantly improve the performance, the training process needs to be terminated. Through the above training process, the PPO agent gradually learns how to select the optimal beam alignment policy in different states to maximize the cumulative reward. The simulation results show that the PPO algorithm can effectively improve the alignment accuracy and link stability of the UAV optical communication link, and has good robustness and convergence in a high-speed moving environment.

[0187] Use the trained agent to generate the best adjustment strategy for the UAV pod attitude.

[0188] In this embodiment, for the trained agent, the state of the agent is obtained in real time and input into the trained agent, and the selected optimal action is output, and this optimal action is used as the best adjustment strategy for the UAV pod attitude.

[0189] In this embodiment, two reinforcement learning algorithms, DDPG and PPO, are used for comparative training to evaluate the influence of different optimization methods on the APT performance, and a more suitable algorithm is selected according to different requirements in the actual application scenario. Specifically: apply the trained agent to the actual UAV optical communication pod platform, adjust according to the real-time flight data, and optimize the parameters. Through continuous learning in the real environment, the model can gradually adapt to external interference and complex flight conditions, and improve the success rate of aligning with the user and establishing a communication link.

[0190] Verification is carried out in the simulation environment and the actual test environment. According to the test results, the PPO algorithm shows significant advantages in rapid adaptation. Its average convergence step is 8×10 4 steps, which is 47% less than the average convergence step of 1.5×10 5 steps of the DDPG algorithm; its sample utilization rate is 68%, which is close to twice the sample utilization rate of 35% of the DDPG algorithm. Especially in a dynamic environment, such as in the case of sudden turbulence, it only takes 12 seconds to complete the policy adjustment, but its control accuracy (0.25±0.12mrad) has an error increase of about 52% compared to the control accuracy (0.12±0.03mrad) of the DDPG algorithm. Therefore, when the flight mission requires high precision and low error, the DDPG algorithm is selected for agent training; when the flight mission environment changes rapidly and rapid policy adjustment is required, the PPO algorithm is selected for agent training;

[0191] In addition, the algorithm performance is verified in the simulation environment and the actual test environment, and the stability and reliability of the intelligent agent in a high dynamic environment are tested through the UAV optical communication experiment. This implementation is compared with traditional APT schemes such as PID control and Kalman filtering, and the advantages of reinforcement learning in high dynamic optical communication are analyzed to ensure the effectiveness and feasibility of the designed intelligent control algorithm in practical applications. The test results show that PID control can provide stable control effects in low dynamic environments, such as static base stations to UAV links, but it is difficult to adapt to complex disturbances in high dynamic environments; Kalman filtering can perform better state estimation in low noise environments, but in UAV optical communication systems with strong nonlinearity and drastic environmental changes, the tracking accuracy is limited. In contrast, DDPG and PPO can automatically adapt to environmental changes based on reinforcement learning, and show better stability and accuracy in high dynamic scenarios, providing a more robust intelligent APT optimization solution for UAV optical communication links.

[0192] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A follow - up and tracking method for UAV optical communication links based on reinforcement learning, characterized in that, The method includes the following processes: Obtain the characteristic parameters of the UAV optical communication link and construct a UAV optical communication link channel model; Define an agent for adjusting the attitude of the UAV pod, and establish an interaction environment for the agent based on the UAV optical communication link channel model; Based on the interaction environment of the agent, define the state space, action space, and reward function of the agent respectively; Select the Deep Deterministic Policy Gradient (DDPG) algorithm or the Proximal Policy Optimization (PPO) algorithm according to the given flight mission requirements to train the agent and generate the optimal adjustment strategy for the UAV pod attitude.

2. The method for tracking and aiming of an unmanned aerial vehicle optical communication link based on reinforcement learning according to claim 1, wherein, The specific content of obtaining the characteristic parameters of the UAV optical communication link and constructing the UAV optical communication link channel model is as follows: Obtain the characteristic parameters of the UAV optical communication link, including: UAV coordinates, user coordinates, visibility, laser wavelength, fading caused by atmospheric turbulence, and pointing alignment error; Construct a UAV optical communication link channel model containing several links, and represent the channel gain of any link in the UAV optical communication link channel model as: where h k (t) represents the channel gain of the k-th link at time slot t; η is the responsivity of the user terminal; represents the atmospheric attenuation of the k-th link at time slot t; represents the fading of the k-th link caused by atmospheric turbulence at time slot t; represents the pointing alignment error of the k-th link at time slot t; The calculation method of the atmospheric attenuation of the k-th link at time slot t is as follows: For time slot t, calculate the distance between the UAV and the user using the obtained UAV coordinates and user coordinates, calculate the attenuation coefficient using the obtained visibility and laser wavelength, and then calculate the atmospheric attenuation of the k-th link at time slot t using the Beer-Lambert law based on the calculated distance between the UAV and the user and the attenuation coefficient.

3. The method for tracking and aiming of an unmanned aerial vehicle optical communication link based on reinforcement learning according to claim 2, wherein, The specific content of establishing the interaction environment for the agent based on the UAV optical communication link channel model is as follows: Simulate the beam tracking process between the UAV and the user in a complex environment, establish the interaction environment for the agent, which is used to generate the state information of the agent, and update the state information of the agent according to the action selected by the agent; The interaction environment of the agent includes three parts: the UAV optical communication link channel model, the optical imaging simulation system, and the dynamic interference simulation model; among them, the UAV optical communication link channel model is used to feedback the reward value according to the state information of the agent; the optical imaging simulation system is used to map the obtained user coordinates to the camera coordinate system of the camera carried in the UAV pod and perform distortion correction, and use the corrected user coordinates as the user coordinates in the characteristic parameters of the UAV optical communication link; the dynamic interference simulation model is used to simulate different interference scenarios in a complex environment.

4. The method for tracking and aiming of an unmanned aerial vehicle optical communication link based on reinforcement learning according to claim 3, wherein The specific content of defining the state space, action space, and reward function of the agent respectively based on the interaction environment of the agent is as follows: Define the state space of the agent to include: UAV attitude information, beam footprint pixel coordinates, and beacon point coordinates; among them, the UAV attitude information includes: pitch angle, roll angle, and yaw angle; Define the action space of the agent to include: pod pitch angle and pod yaw angle; Define the reward function as: where \(R(t)\) represents the reward value of the agent at time slot \(t\); represents the square of the horizontal distance between the pixel coordinates of the beam footprint and the fiducial point coordinates; represents the square of the vertical distance between the pixel coordinates of the beam footprint and the fiducial point coordinates.

5. The method for tracking and aiming of an unmanned aerial vehicle optical communication link based on reinforcement learning according to claim 4, wherein, The specific method of selecting the Deep Deterministic Policy Gradient (DDPG) algorithm or the Proximal Policy Optimization (PPO) algorithm according to the given flight mission to train the agent and generate the optimal adjustment strategy for the UAV pod attitude is as follows: Based on the interaction environment of the agent, use a 2-DOF rotating platform to simulate the alignment process of the UAV optical communication link; Use the Actor-Critic framework to construct a network architecture for training an agent, including: an APT capturer and an APT evaluator, where the APT capturer includes an evaluation network and a target network; the APT evaluator includes an evaluation network and a target network; According to the given flight mission, if the requirement for tracking accuracy of the flight mission is higher than the requirement for the response speed of the adjustment strategy of the UAV pod attitude, then select the Deep Deterministic Policy Gradient (DDPG) algorithm to train the agent; otherwise, select the Proximal Policy Optimization (PPO) algorithm to train the agent; Use the trained agent to generate the optimal adjustment strategy for the UAV pod attitude.

6. The method for tracking and aiming of an unmanned aerial vehicle optical communication link based on reinforcement learning according to claim 5, wherein The specific method for selecting the Deep Deterministic Policy Gradient (DDPG) algorithm to train the agent is as follows: Randomly initialize the network parameters and weights of the evaluation network in the APT capturer and the evaluation network in the APT evaluator respectively, and copy the weights of the evaluation network in the APT capturer to the target network in the APT capturer, and copy the weights of the evaluation network in the APT evaluator to the target network in the APT capturer; Initialize the experience replay pool, the state space of the agent, and the action space of the agent, and start iterative training; At time slot t, obtain the current state S(t) of the agent, use the evaluation network of the APT capturer to select the optimal action A(t), and calculate the immediate reward R(t) of the agent using the reward function. Update the next state S(t+Δt) of the agent according to the immediate reward R(t) of the agent, and then store (S(t), A(t), R(t), S(t+Δt)) as a set of interaction data in the experience replay pool; continue to use the evaluation network of the APT capturer to select the optimal action for the next state S(t+Δt) of the agent until there are several sets of interaction data stored in the experience replay pool; Set the total number of interaction data M for each round of iterative training, and initialize the number of interaction data m for the current iteration round. Randomly extract a set of interaction data from the experience replay pool, and increment the number of interaction data for the current iteration round by 1. Calculate the state-action value function in a way that maximizes the return as the expected cumulative reward of the agent; According to the expected cumulative reward of the agent, construct the loss function of the evaluation network of the APT evaluator, and update the network parameters of the evaluation network of the APT evaluator by minimizing the value of the loss function of the evaluation network of the APT evaluator; According to the chain rule, calculate the gradient of the objective function J with respect to the network parameters of the evaluation network of the APT capturer and update the network parameters of the evaluation network of the APT capturer by gradient ascent; According to the updated network parameters of the evaluation network of the APT evaluator and the updated network parameters of the evaluation network of the APT capturer, use the soft update mechanism to update the network parameters of the target network of the APT evaluator and the network parameters of the target network of the APT capturer; Judge whether m > M holds. If not, return to the step of randomly extracting a set of interaction data from the experience replay pool to continue training, and increment the number of interaction data for the current iteration round by 1; if so, judge whether the current iteration round is greater than the preset number of training rounds. If so, stop iterative training; Otherwise, re-initialize the experience replay pool, the state space of the agent, and the action space of the agent, and start the next round of iterative training.

7. The method for tracking and aiming an optical communication link of an unmanned aerial vehicle based on reinforcement learning according to claim 6, wherein The method of using the evaluation network of the APT capturer to select the optimal action A(t) is as follows: First, use the evaluation network of the APT capturer to select the current action, and then use the Ornstein-Uhlenbeck process (OU process) to add noise perturbation to the current action to generate the optimal action A(t) to be executed, which is expressed as: A(t) = μ′(S(t)) = μ(S(t)|θ μ ) + N where μ represents the evaluation network of the APT capturer; θ μ represents the network parameters of the evaluation network of the APT capturer; μ(S(t)|θ μ ) represents the policy of selecting actions using the evaluation network of the APT capturer; N represents noise; μ′(S(t)) represents the policy of selecting actions using the evaluation network of the APT capturer after adding noise perturbations.

8. The method for tracking and aiming of an unmanned aerial vehicle optical communication link based on reinforcement learning according to claim 7, wherein The state-action value function is expressed as: Q(S(t), A(t)) = E R(t),S(t+Δt)~ E[R(S(t), A(t)) + γQ(S(t + Δt), μ(S(t + Δt)))] Among them, Q(S(t), A(t)) represents the Q-value of selecting the optimal action in the current state S(t), that is, it is used to evaluate the expected cumulative reward of executing the optimal action A(t) in the current state S(t); R(S(t), A(t)) represents the immediate reward obtained after executing the optimal action A(t) in the current state S(t); γ is the discount factor, and γ ∈ [0, 1]; Δt represents the interval time; S(t + Δt) represents the next state of the agent; Q(S(t + Δt), μ(S(t + Δt))) represents the Q-value of selecting an action according to the evaluation network μ of the APT capturer in the next state of the agent; E R(t),S(t+Δt)~ E represents the expectation of the immediate reward and the next state; The loss function of the evaluation network of the APT evaluator is expressed as: L(θ Q ) = E[(Q(S(t), A(t)|θ Q ) - y t ) 2 ​ where L(θ Q ) represents the loss function value of the evaluation network of the APT evaluator; E[·] represents the calculation of the expected value; θ Q represents the network parameters of the evaluation network of the APT evaluator; Q(S(t), A(t)|θ Q ) represents the Q value predicted by the evaluation network of the APT evaluator under the current state S(t) and the optimal action A(t); y t represents the fitted value function of the target network of the APT evaluator, and there is: y t = R(S(t), A(t)) + γQ′(S(t + Δt), μ′(S(t + Δt)|θ μ′ )|θ Q′ ) where θ Q′ represents the network parameters of the target network of the APT evaluator; θ μ′ represents the network parameters of the target network of the APT capturer; μ′ represents the target network of the APT capturer; Q' represents the target network of the APT evaluator; μ′(S(t+Δt)|θ μ′ ) represents the action selection of the target network of the APT evaluator according to the next state; The calculation method of the gradient of the objective function J with respect to the network parameters of the evaluation network of the APT capturer is: where represents the gradient of the objective function J with respect to the network parameters of the evaluation network of the APT catcher; μ(S(t)|θ μ ) represents the optimal action A(t) selected by the evaluation network of the APT catcher with network parameters θ μ at the current state S(t).

9. The method for tracking and aiming of an unmanned aerial vehicle optical communication link based on reinforcement learning according to claim 8, characterized in that, The specific method of using the Proximal Policy Optimization (PPO) algorithm to train the agent is: Randomly initialize the network parameters of the APT capturer and the network parameters of the APT evaluator; Initialize the state space of the agent and the action space of the agent, and start iterative training; At time slot t, obtain the current state S(t) of the agent, use the evaluation network of the APT capturer to select the optimal action A(t), and use the reward function to calculate the immediate reward R(t) of the agent. Update the next state S(t+Δt) of the agent according to the immediate reward R(t) of the agent, and take (S(t), A(t)) as a set of state-action pairs. Repeat the processes of action selection, immediate reward calculation, and state update, and save several sets of state-action pairs; For time slot t, according to the state-action pair (S(t), A(t)) at time slot t, use the Generalized Advantage Estimation (GAE) to construct the advantage function, and use the advantage function to construct the loss function of the APT capturer; Calculate the target state value according to the state-action pair (S(t), A(t)) at time slot t, and construct the loss function of the APT evaluator; Update the network parameters of the APT capturer by maximizing the loss function of the APT capturer, and update the network parameters of the APT evaluator by minimizing the loss function of the APT evaluator; Extract a fixed batch of state-action pairs, and update the network parameters of the APT capturer using the clipped policy gradient method according to the loss function of the APT capturer. At the same time, update the network parameters of the APT evaluator by minimizing the loss function of the APT evaluator to complete one round of iterative training; Judge whether the current iteration round is greater than the preset number of training rounds. If so, stop the iterative training; if not, re-initialize the state space of the agent and the action space of the agent, and start the next round of iterative training.

10. The method for tracking and aiming of an unmanned aerial vehicle optical communication link based on reinforcement learning according to claim 9, wherein, The advantage function is expressed as: where is the advantage estimator; δ(t) is the TD-error at time slot t; γ1 represents the discount factor; λ p is the GAE factor; δ(t + Δt) is the TD-error from time slot t to time slot t + Δt; T represents the termination time step; δ(T - Δt) represents the TD-error from T to time slot T - Δt; The loss function of the APT capturer is expressed as: Among them represents the total loss of the APT capturer; represents the set of network parameters of the APT capturer; E t represents the expectation of the empirical sample; η1 is a fixed coefficient; represents the policy entropy; represents the clipped policy loss, and there is: where clip(·) is the clipping function; ∈ p is a hyperparameter; is the probability ratio, and we have: Among them is the strategy adopted when performing a selection action using the updated APT capturer; is the strategy adopted when performing a selection action using the APT capturer before the update; Note that h t (θ old ) = 1; The loss function of the APT evaluator is: where represents the value regression loss of the APT evaluator; E t represents the expectation of the empirical sample; represents the state-value function with parameter φ p ; represents the target state value.

Citation Information

Patent Citations

  • Unmanned aerial vehicle tracking and pointing communication system and method based on dual-light fusion

    CN113691781A

  • Unmanned aerial vehicle automatic trajectory planning method based on wireless optical communication

    CN116339370A

  • Aircraft control method based on reinforcement learning, terminal equipment and medium

    CN117311374A

  • Aircraft control system element gradient reinforcement learning design method

    CN117369268A

  • Multi-agent-based air security data acquisition method

    CN118249883A

Cited By

  • Heat exchanger real-time control method and system, electronic equipment and storage medium

    CN121326018A

  • Kalman filtering network Beidou satellite positioning method and system based on environment interaction

    CN121500352A