An automatic trajectory planning method for UAV based on wireless optical communication

By combining wireless optical communication and machine learning technology, the flight trajectory of drone is optimized, and the local convergence and unfair communication problems of drone trajectory planning in traditional methods are solved, and efficient communication capacity and fairness are improved in complex environments.

CN116339370BActive Publication Date: 2025-08-29TSINGHUA SHENZHEN INTERNATIONAL GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310063356.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-18
Publication Date
2025-08-29
Estimated Expiration
2043-01-18

AI Technical Summary

Technical Problem

Traditional drone trajectory planning methods are difficult to converge or only converge to the local optimal solution in complex environments, and the drone-assisted communication system has unfair communication problems, which cannot effectively improve the communication capacity and fairness of ground users.

Method used

Combining optical base stations and drones, a trajectory planning method based on wireless optical communication is adopted, and a machine learning technology is used to optimize the drone's flight trajectory. By setting communication thresholds and deep reinforcement learning, the optimal flight route is planned to improve the system's communication capacity, and the Jain fairness index is used to quantitatively measure communication fairness.

Benefits of technology

It realizes rapid convergence to the global optimal solution in complex environments, improves the system's communication capacity, and ensures the fairness of communication, solving the unfair communication problems in traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116339370B_ABST
    Figure CN116339370B_ABST
Patent Text Reader

Abstract

The present invention discloses an automatic trajectory planning method for a drone based on wireless optical communication, comprising: S1, establishing a drone-assisted wireless optical communication channel model based on a visible light communication link channel model, and calculating channel gain and channel capacity; S2, limiting the movement space between the drone and ground users and the minimum capacity threshold of the communication link based on the actual application environment and the channel capacity calculated in step S1, determining an optimization target and creating a problem statement for drone-assisted wireless optical communication; S3, automatically planning the drone's flight trajectory using deep reinforcement learning based on the drone-assisted wireless optical communication channel model established in step S1 and the restrictions in step S2, so as to achieve the optimization target.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of machine learning and communications, and in particular to an automatic trajectory planning method for a drone based on wireless optical communication. Background Art

[0002] Drones, due to their low construction and operation costs, high maneuverability, and low safety risks, are widely used in emergency communications and expanded communication capacity. Future wireless communication systems are expected to meet the unprecedented demand for high-quality wireless services. This poses a challenge to traditional terrestrial communication networks, especially in traffic hotspots such as football stadiums or rock concerts. Furthermore, in emergencies such as earthquakes and tsunamis, drones can serve as aerial base stations to supplement or support existing terrestrial communication infrastructure, extending the traditional two-dimensional base station distribution to three dimensions. Currently, Google's Loon project (S. Katikala, "Google project loon," InSight, Rivier Academic J., vol. 10, no. 2, pp. 1–6, 2014) and Facebook's Internet.org project (M. Zuckerberg, "Connecting the world from the sky," Facebook, Cambridge, MA, USA, Tech. Rep., 2014) are working to use aerial vehicles such as hydrogen balloons to address network access issues for users on the ground. However, traditional drone-assisted communication systems based on RF channels cannot effectively improve communication capacity for ground users due to issues such as limited frequency resources and narrow bandwidth allocation. Furthermore, in future ultra-dense wireless networks, drones deployed using RF technology as aerial base stations will interfere with ground equipment, significantly impacting terrestrial network performance. These issues can be addressed by equipping drones with high-bandwidth, unlicensed visible light communication (VLC) systems.

[0003] On the other hand, how to use drones for rapid network deployment has attracted widespread attention from researchers. Traditional trajectory planning systems based on optimization methods can only obtain suboptimal solutions or even fail to solve non-convex scenarios or NP-hard problems. However, machine learning technology, through continuous exploration and trial and error, learns optimal decisions based on the rewards obtained from interacting with the environment. This allows drones to make action decisions without complete environmental information, and can then make rapid decisions in response to environmental changes, thereby increasing the system's overall communication capacity in a short period of time. This makes it ideal for use in complex environmental systems.

[0004] Since a drone needs to provide communication services to multiple users on the ground at the same time, certain planning results can easily lead to unfair communication, that is, the communication capacity between the drone and a certain user is much larger than the communication capacity between the drone and other users, and even normal communication cannot be guaranteed for some users. Summary of the Invention

[0005] In view of this, the present invention combines optical base stations and drones and proposes an automatic trajectory planning method for drones based on wireless optical communication. By setting a communication threshold in the system, the drone can optimize different objective functions using machine learning technology to output the optimal flight trajectory planning route while meeting the communication threshold between each user, thereby improving the overall communication capacity of the system.

[0006] A method for automatic trajectory planning for a drone based on wireless optical communication comprises the following steps: S1, establishing a drone-assisted wireless optical communication channel model based on a visible light communication link channel model, and calculating the channel gain and channel capacity; S2, restricting the movement space between the drone and ground users and the minimum capacity threshold of the communication link based on the actual application environment and the channel capacity calculated in step S1, determining an optimization goal, and creating a problem statement for drone-assisted wireless optical communication; S3, automatically planning the drone's flight trajectory using deep reinforcement learning based on the drone-assisted wireless optical communication channel model established in step S1 and the restrictions in step S2, to achieve the optimization goal.

[0007] Furthermore, the method further includes: S4, using the Jain fairness index to quantitatively evaluate communication fairness under the flight trajectory planning obtained by optimization in step S3.

[0008] Furthermore, the step of calculating the channel gain in step S1 includes: considering the direct effect of visible light communication, assuming that a UAV in the air performs downlink communication with U random users on the ground, the channel gain h of the visible light communication link is ij Expressed as:

[0009]

[0010] Among them, h ij represents the channel gain of communication between UAV i and ground user j, m is the Lambert coefficient, A is the communication coverage area of ​​the UAV, and d ij represents the straight-line distance between UAV i and ground user j, is the visible light irradiation angle, g(φ ij ) is the gain of the optical concentrator, φ ij is the incident angle of visible light, Ψ c is the receiver field of view half angle; m and g (φ ij) are:

[0011]

[0012]

[0013] Among them, Φ 1 / 2 is the transmitter half angle, n e is the refractive index of the transmission medium.

[0014] Furthermore, the step of calculating the channel capacity in step S1 includes: the channel capacity C of the visible light communication link between the UAV i and the ground user j is:

[0015]

[0016] Where, e is the base of the natural logarithm function, σ w is the standard deviation of additive white Gaussian noise, P i is the optical power of the UAV’s optical device, ξ is the illumination target, and h ij is the channel gain for communication between UAV i and ground user j.

[0017] Furthermore, in step S2, the mobile space between the UAV and the ground user and the minimum capacity threshold of the communication link are restricted according to the actual application environment and the channel capacity calculated in step S1. Specifically, the following steps are performed: Assume that the position of the UAV in space at time t is (x t ,y t ,z t ), the random position coordinates of the jth ground user are (x j ,y j ,0), the simulation creation space is V, the minimum flight height of the drone is H, then the position of each user and drone on the ground satisfies: (x j ,y j ,0)∈V,(x t ,y t ,z t )∈V∩z t ≥H; The minimum capacity threshold of the communication link is the minimum communication capacity that ensures distortion-free communication between the UAV and the ground user; let the capacity threshold of the simulation be δ, then the downlink communication capacity C between the UAV and each ground user when it finally stops is j To meet: C j ≥δ.

[0018] Furthermore, the optimization objectives determined in step S2 include: the sum of the communication capacities of all ground users, the minimum communication capacity of each ground user, and the communication capacity with the balance parameter ρ.

[0019] Furthermore, in step S2, determining the optimization objective and creating a problem statement for UAV-assisted wireless optical communication specifically includes maximizing the sum of the communication capacities of all ground users, which is expressed as:

[0020] max∑ j∈U C j

[0021] st(x j ,y j ,0)∈V

[0022] (x t ,y t ,z t )∈V∩z t ≥H

[0023] C j ≥δ

[0024] Maximize the minimum communication capacity of each ground user, expressed as:

[0025] max(min j∈U C j )

[0026] st(x j ,y j ,0)∈V

[0027] (x t ,y t ,z t )∈V∩z t ≥H

[0028] C j ≥δ

[0029] Maximizing the communication capacity with the balancing parameter ρ is expressed as:

[0030] max(ρ×∑ j∈U C j +(1-ρ)×min j∈U C j )

[0031] st(x j ,y j ,0)∈V

[0032] (x t ,y t ,z t )∈V∩z t ≥H

[0033] C j ≥δ

[0034] ρ∈[0,1]

[0035] Among them, ∑ j∈U C j 、min j∈U C j 、(ρ×Σ j∈U C j +(1-ρ)×min j∈U C j ) are respectively the sum of the communication capacity of all ground users, the minimum communication capacity of each ground user, and the communication capacity with the balance parameter ρ.

[0036] Furthermore, in step S3, the DQN neural network architecture is adopted and the DQN algorithm is used to automatically plan the flight trajectory of the drone. During the planning process, the drone selects the action corresponding to the maximum Q value in the DQN neural network output with probability P when selecting an action, and randomly selects the next action with probability (1-P). At the same time, the probability P increases with the increase of the number of learning times n.

[0037] Where λ is the probability growth rate, P max is the upper limit of probability;

[0038] The planning process includes: when the drone receives an action instruction, it moves to the next state, and the system feeds back a reward value to the drone. If the drone exceeds the specified space after moving, a negative reward is output, and the learning is judged to have failed, and a new round of learning is restarted; if the communication capacity of the drone becomes smaller after moving, a negative reward is output, but learning continues; if the communication capacity of the drone remains unchanged after moving, the output reward is 0; if the communication capacity between the drone and a certain user after moving is less than the minimum capacity threshold, a negative reward is output; if the system determines that the learning has converged, but the communication capacity between the convergence coordinate and a certain user is less than the minimum capacity threshold, the learning is judged to have failed, and a new round of learning is restarted; if the system determines that the learning has converged, and the communication capacity between the convergence coordinate and each user is greater than the minimum capacity threshold, the learning is successful, and the system feeds back a positive reward.

[0039] Furthermore, the Jain fairness index in step S4 is:

[0040]

[0041] C k represents the communication link capacity between the UAV and the kth ground user.

[0042] On the other hand, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can implement the steps of the aforementioned method for automatic trajectory planning of a drone.

[0043] The beneficial effects of the technical solution of the present invention are as follows: most existing UAV trajectory planning solutions are based on traditional optimization methods, which may fail to converge or only converge to a local optimal solution when faced with complex environments. The present invention adopts a machine learning-based method to accelerate the convergence rate through continuous exploration and trial and error, and converge to a global optimal solution; most existing UAV-assisted communication systems aim to maximize communication capacity, but do not take into account the unfair communication problems caused by UAV movement, while the present invention can dynamically adjust the fairness of UAV path planning by setting balance parameters, while fully considering the communication fairness problem while improving the communication capacity of ground users; and in a further technical solution, the Jain fairness index is used to quantify and analyze fairness, verifying the effectiveness of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 2 is a diagram of the DQN neural network architecture in an embodiment of the present invention.

[0045] Figure 2 This is a curve diagram of single-user communication capacity change when the communication capacity is at the minimum value according to an embodiment of the present invention.

[0046] Figure 3 This is a curve diagram of single-user communication capacity change under the condition of the sum of communication capacities according to an embodiment of the present invention.

[0047] Figure 4 This is a curve diagram of single-user communication capacity variation when the balance parameter ρ is 0.5 according to an embodiment of the present invention.

[0048] Figure 5 1 is a curve showing how the fairness of the system changes with the balance parameters in an embodiment of the present invention. DETAILED DESCRIPTION

[0049] The present invention will be further described below with reference to the accompanying drawings and specific implementation methods. Due to the problems of tight frequency resources and narrow bandwidth allocation in radio frequency signals, providing high-speed communication services to ground users is currently in urgent need of improvement in the industry, and the application of visible light communication can alleviate the above problems to a certain extent. In the event of an emergency, such as an earthquake or tsunami, which causes damage to the ground base station and makes it impossible to communicate, the drone communication system can be quickly deployed to restore communication services in the damaged area in the shortest possible time; in addition, in places with a high density of people, such as concerts, where ground base station resources are tight, drones can be dispatched as extended communication systems to provide communication services to ground users. In this context, visible light-based drone communication systems are increasingly being used to provide communication services to ground users, and can provide lighting services while providing communication services.

[0050] The present invention addresses the problems of current visible light-based drone communication systems and proposes a method for automatic drone trajectory planning based on wireless optical communication. By setting a communication threshold in the system, the drone can meet the communication threshold between users while simultaneously using machine learning technology to optimize different objective functions to output the optimal flight trajectory planning route, thereby improving the overall communication capacity of the system. The method includes the following steps S1 to S3:

[0051] S1. Establish a UAV-assisted wireless optical communication channel model based on the visible light communication link channel model, and calculate the channel gain and channel capacity;

[0052] S2. Based on the actual application environment and the channel capacity calculated in step S1, the movement space between the UAV and the ground user and the minimum capacity threshold of the communication link are restricted, the optimization goal is determined, and the problem statement of UAV-assisted wireless optical communication is created;

[0053] S3. Based on the UAV-assisted wireless optical communication channel model established in step S1 and the constraints in step S2, deep reinforcement learning is used to automatically plan the UAV flight trajectory to achieve the optimization goal.

[0054] In order to quantify and analyze the communication fairness of the communication system based on the above-mentioned drone automatic trajectory planning method in the embodiment of the present invention, optionally, step S4 may be further included: under the flight trajectory planning optimized in step S3, the Jain fairness index is used to quantitatively evaluate the communication fairness.

[0055] Considering the direct effect of visible light communication, in this embodiment of the present invention, it is assumed that a UAV in the air performs downlink communication with U randomly located users on the ground. The channel gain h of the visible light communication link is ij It can be expressed as:

[0056]

[0057] Among them, h ij represents the channel gain of communication between UAV i and ground user j, m is the Lambert coefficient, A is the communication coverage area of ​​the UAV, and d ij represents the straight-line distance between UAV i and ground user j, is the visible light irradiation angle, g(φ ij ) is the gain of the optical concentrator, φ ij is the incident angle of visible light, Ψ c is the half angle of the receiver’s field of view. m and g(φ ij ) are:

[0058]

[0059] Among them, Φ 1 / 2 is the transmitter half angle, n e is the refractive index of the transmission medium.

[0060] The channel capacity C of the visible light communication link between UAV i and ground user j is:

[0061]

[0062] Where, e is the base of the natural logarithm function, σ w is the standard deviation of the additive white Gaussian noise, P i is the optical power of the UAV’s optical device, ξ is the illumination target, and h ij is the channel gain for communication between UAV i and ground user j.

[0063] In order to further determine the optimization problem of the entire system, it is necessary to limit the mobile space and the minimum capacity threshold of the communication link between the UAV and the ground user. Assume that the position of the UAV in space at time t is (x t ,y t ,z t ), the random position coordinates of the jth ground user are (x j ,y j ,0), the simulation creation space is V, the minimum flight height of the drone is H, then the position of each user and drone on the ground should satisfy:

[0064] (x j ,y j ,0)∈V and (x t ,y t ,z t )∈V∩z t ≥H.

[0065] The minimum capacity threshold of the communication link is the minimum communication capacity that ensures distortion-free communication between the UAV and the ground user. Assuming the capacity threshold of the simulation is δ, the downlink communication capacity C between the UAV and each ground user when it finally stops is j To meet: C j ≥δ.

[0066] To measure the fairness of the communication between the UAV and each ground user when the UAV is stationary, the embodiment of the present invention uses the Jain fairness index to quantify the fairness of the system. The Jain fairness index is defined as follows:

[0067]

[0068] Ck represents the communication link capacity between the UAV and the kth ground user.

[0069] The communication environment is a key factor influencing drone trajectory planning. Different optimization objectives can lead to significant differences in the drone's final resting position and service quality. This embodiment of the present invention optimizes three communication environments: the sum of the communication capacity of all ground users, the minimum communication capacity of each ground user, and the communication capacity with the balance parameter ρ. The resulting system optimization objectives are as follows:

[0070] Maximize the sum of the communication capacity of all ground users, expressed as:

[0071]

[0072] Maximize the minimum communication capacity of each ground user, expressed as:

[0073]

[0074] Maximizing the communication capacity with the balancing parameter ρ is expressed as:

[0075]

[0076] Among them, ∑ j∈U C j 、min j∈U C j 、(ρ×∑ j∈U C j +(1-ρ)×min j∈U C j ) are respectively the sum of the communication capacity of all ground users, the minimum communication capacity of each ground user, and the communication capacity with the balance parameter ρ.

[0077] Under the above channel model and environmental constraints, the embodiment of the present invention is based on the DQN (Deep Q-Network) neural network architecture and uses the DQN algorithm to automatically plan the flight trajectory of the drone. Figure 1 As shown, the pseudo code of the DQN algorithm is as follows Table 1:

[0078] Table 1 Pseudo code of automatic trajectory planning algorithm for UAV based on wireless optical communication

[0079]

[0080] The DQN algorithm is formed by adding a neural network system to the traditional Q-learning based on Q-table. It uses neural network fitting instead of Q-table for action decision-making, thus solving the problem of infinite Q-table in complex environments. Figure 1 The architecture of the DQN neural network in the embodiment of the present invention is shown. Please refer to Figure 1The current state s of the drone in the machine learning system is a coordinate in space. Since the simulation uses a rotorcraft, it has seven action options, namely "up, down, left, right, front, back, and airborne" ( Figure 1 In order to ensure that the system has a certain degree of exploratory power, the drone selects the action corresponding to the maximum Q value output by the Q network in the DQN with probability P when selecting an action, and randomly selects the next action with probability (1-P). At the same time, the probability P increases with the increase of the number of learning times n.

[0081] Where λ is the probability growth rate, P max is the upper limit of probability;

[0082] After receiving an action command, the drone moves to the next state, and the system provides a reward. If the drone moves beyond the specified space, a negative reward is output, learning is considered a failure, and a new learning cycle begins. If the communication capacity decreases after the drone moves, a negative reward is output, but learning continues. If the communication capacity remains unchanged after the drone moves, the reward is zero. If the communication capacity between the drone and a user after the drone moves is less than the minimum capacity threshold, a negative reward is output. If the system determines that learning has converged, but the communication capacity between the converged coordinate and a user is less than the minimum capacity threshold, learning is considered a failure, and a new learning cycle begins. If the system determines that learning has converged and the communication capacity between the converged coordinate and every user is greater than the minimum capacity threshold, learning is successful, and the system provides a positive reward. After the drone receives a reward, the system packages the current state s, the selected action a, the reward r, and the next state value s_ after executing the action as "experience" and stores it in the system's experience pool. After accumulating some experience in the experience pool, it is fed into the Q network and target Q network to calculate the predicted Q value and target Q value, and the loss function between them is calculated. After obtaining the loss function, the system continuously reduces the error between the predicted Q value and the target Q value through the gradient descent method and updates the parameters in the Q network. At the same time, the Q network also synchronizes its parameters to the target Q network every N steps.

[0083] The present invention optimizes three different communication environments and studies their fairness. Figure 2 The figure shows how the communication capacity of a single user changes with the number of learning times, based on the minimum communication capacity. It can be seen that after about 30 learning times, the communication capacity between the drone and each ground user approaches the same, and the Jain fairness index is calculated to be 0.9999. Figure 3The figure shows how the communication capacity of a single user changes with the number of learning times based on the sum of communication capacities. It can be seen that after the drone learns about 20 times, its communication capacity with user 3 is significantly higher than that with other users, while ensuring that the communication capacity with all users is greater than the threshold (Capacity_threshold). The Jain fairness index calculation result is 0.9. Figure 4 The figure shows how the single-user communication capacity with the balance parameter changes with the number of learning times when the balance parameter ρ = 0.5. It can be seen that the trajectory planning strategy of the drone in this case is different from the previous two, and the Jain fairness index calculation result is 0.9537. Figure 5 The paper shows the variation of the fairness index with the equilibrium parameter ρ. As ρ increases from 0 to 1, the communication fairness of the entire system decreases. Therefore, we can conclude that the trajectory planning strategy for drones based on the minimum communication capacity environment focuses more on the fairness of system communication, but does not significantly improve the overall communication capacity. The trajectory planning strategy for drones based on the sum communication capacity environment focuses more on improving the communication capacity of a single user, resulting in lower fairness in system communication. The introduction of the equilibrium parameter ρ balances fairness and the degree of improvement in system communication capacity to a certain extent. In practical applications, the value of ρ can be adjusted to guide the flight strategy of drones during trajectory planning according to different communication environments.

[0084] Another embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When executed by a processor, the computer program can implement the steps of the automatic trajectory planning method for a drone of the aforementioned embodiment. Based on this understanding, the technical solution of the present invention can be embodied in the form of a software product. The software product can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard drive, etc.), including a number of instructions for causing a computer device (such as a personal computer, server, or network device, etc.) to execute the method of the present invention.

[0085] The above is a further detailed description of the present invention in conjunction with specific preferred embodiments, and the specific implementation of the present invention should not be considered to be limited to these descriptions. Those skilled in the art will recognize that several equivalent substitutions or obvious variations can be made without departing from the scope of the present invention, and that any equivalent performance or application should be considered to fall within the scope of protection of the present invention.

Claims

1. A method for automatic trajectory planning of a UAV based on wireless optical communication, characterized in that: The steps include: S1. Establish a UAV-assisted wireless optical communication channel model based on the visible light communication link channel model, and calculate the channel gain and channel capacity of the visible light communication link; S2. Based on the actual application environment and the channel capacity calculated in step S1, the movement space between the UAV and the ground user and the minimum capacity threshold of the communication link are restricted, the optimization goal is determined, and the problem statement of UAV-assisted wireless optical communication is created; The minimum capacity threshold of the communication link is the minimum communication capacity that ensures distortion-free communication between the UAV and the ground user; The optimization objectives include: the sum of the communication capacities of all ground users, the minimum communication capacity of each ground user, and the communication capacity with a balancing parameter ρ; the problem statement includes: maximizing the sum of the communication capacities of all ground users, maximizing the minimum communication capacity of each ground user, and maximizing the communication capacity with a balancing parameter ρ; S3. Based on the UAV-assisted wireless optical communication channel model established in step S1 and the constraints in step S2, deep reinforcement learning is used to automatically plan the flight trajectory of the UAV to achieve the optimization goal.

2. The automatic trajectory planning method for a UAV according to claim 1, wherein: Also includes: S4. Under the flight trajectory planning optimized in step S3, the Jain fairness index is used to quantitatively evaluate the communication fairness.

3. The automatic trajectory planning method for a UAV according to claim 1, wherein: The step of calculating the channel gain in step S1 includes: Considering the direct effect of visible light communication, assuming that a UAV in the air performs downlink communication with U random users on the ground, the channel gain h of the visible light communication link is ij Expressed as: Among them, h ij represents the channel gain of communication between UAV i and ground user j, m is the Lambert coefficient, A is the communication coverage area of ​​the UAV, and d ij represents the straight-line distance between UAV i and ground user j, is the visible light irradiation angle, g(Φ ij ) is the gain of the optical concentrator, Φ ij is the incident angle of visible light, Ψ c is the half angle of the receiver’s field of view; m and g(Φ ij ) are: Among them, Φ 1 / 2 is the transmitter half angle, n e is the refractive index of the transmission medium.

4. The automatic trajectory planning method for a UAV according to claim 1, wherein: The step of calculating the channel capacity in step S1 includes: The channel capacity C of the visible light communication link between UAV i and ground user j is: Where, e is the base of the natural logarithm function, σ w is the standard deviation of additive white Gaussian noise, P i is the optical power of the UAV’s optical device, ξ is the illumination target, and h ij is the channel gain for communication between UAV i and ground user j.

5. The automatic trajectory planning method for a UAV according to claim 1, wherein: In step S2, based on the actual application environment and the channel capacity calculated in step S1, the movement space between the UAV and the ground user and the minimum capacity threshold of the communication link are restricted, specifically including: Assume that the position of the UAV in space at time t is (x t ,y t ,z t ), the random position coordinates of the jth ground user are (x j ,y j ,0), the simulation creation space is V, the minimum flight height of the drone is H, then the position of each user and drone on the ground satisfies: (x j ,y j ,0)∈V,(x t ,y t ,z t )∈V∩z t ≥H; Assuming the capacity threshold of the simulation is δ, the downlink communication capacity C between the UAV and each ground user when it finally stops is j To meet: C j ≥δ.

6. The automatic trajectory planning method for a UAV according to claim 5, characterized in that: In step S2, the optimization objective is determined and the problem statement of UAV-assisted wireless optical communication is created, which specifically includes: Maximize the sum of the communication capacities of all ground users, expressed as: max∑ j∈U C j s.t.(x j ,y j ,0)∈V (x t ,y t ,z t )∈V∩z t ≥H C j ≥δ Maximize the minimum communication capacity of each ground user, expressed as: max(min j∈U C j ) s.t.(x j ,y j ,0)∈V (x t ,y t ,z t )∈V∩z t ≥H C j ≥δ Maximizing the communication capacity with the balancing parameter ρ is expressed as: max(ρ×∑ j∈U C j +(1-ρ)×min j∈U C j ) s.t.(x j ,y j ,0)∈V (x t ,y t ,z t )∈V∩z t ≥H C j ≥δ ρ∈[0,1] Among them, ∑ j∈U C j 、min j∈U C j 、(ρ×∑ j∈U C j +(1-ρ)×min j∈U C j ) are respectively the sum of the communication capacity of all ground users, the minimum communication capacity of each ground user, and the communication capacity with the balance parameter ρ.

7. The automatic trajectory planning method for a UAV according to claim 1, wherein: In step S3, the DQN neural network architecture is adopted and the DQN algorithm is used to automatically plan the flight trajectory of the drone; During the planning process, the drone selects the action corresponding to the maximum Q value in the DQN neural network output with probability P, and randomly selects the next action with probability (1-P). At the same time, the probability P increases with the increase of the number of learning times n. Where λ is the probability growth rate, P max is the upper limit of probability; The planning process includes: When the drone receives an action instruction, it moves to the next state, and the system feeds back a reward value to the drone. If the drone exceeds the specified space after moving, a negative reward is output, and the learning is judged to have failed, and a new round of learning is restarted; if the communication capacity of the drone becomes smaller after moving, a negative reward is output, but learning continues; if the communication capacity of the drone remains unchanged after moving, the output reward is 0; if the communication capacity between the drone and a certain user after moving is less than the minimum capacity threshold, a negative reward is output; if the system determines that the learning has converged, but the communication capacity between the convergence coordinate and a certain user is less than the minimum capacity threshold, the learning is judged to have failed, and a new round of learning is restarted; if the system determines that the learning has converged, and the communication capacity between the convergence coordinate and each user is greater than the minimum capacity threshold, the learning is successful, and the system feeds back a positive reward.

8. The automatic trajectory planning method for a UAV according to claim 2, wherein: The Jain fairness index in step S4 is: C k represents the communication link capacity between the UAV and the kth ground user.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the automatic trajectory planning method for a drone according to any one of claims 1 to 8 can be implemented.

Citation Information

Patent Citations

  • Three-dimensional unmanned aerial vehicle communication network throughput and time delay balancing method

    CN114173304A

  • Unmanned aerial vehicle assisted air-ground communication optimization algorithm based on deep reinforcement learning algorithm

    CN114826380A