Unmanned aerial vehicle trajectory planning method and device, electronic equipment and storage medium

CN116225058BActive Publication Date: 2026-09-25BEIJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310181830.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-21
Publication Date
2026-09-25
Estimated Expiration
2043-02-21

AI Technical Summary

Technical Problem

[0003]本发明提供一种无人机轨迹规划方法、装置、电子设备及存储介质,用以解决现有技术中因为没有考虑无人机与对应用户终端簇通信过程中的能量损耗,导致整个通信系统的信号传输能力差,同时由于根据无人机的最大覆盖半径对整个区域进行点对点覆盖,造成信号能量损耗较大的缺陷

Benefits of technology

[0040]本发明还提供一种非暂态计算机可读存储介质,其上存储有计算机程序,所述计算机程序被处理器执行时实现如上述任一种所述无人机轨迹规划方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116225058B_ABST
    Figure CN116225058B_ABST
Patent Text Reader

Abstract

The application provides a UAV trajectory planning method and device, electronic equipment and storage medium. The UAV trajectory planning method comprises the following steps: clustering user terminals in a to-be-served area to obtain a plurality of user terminal clusters, each user terminal cluster corresponding to a hotspot area; obtaining a coordinate point sequence corresponding to an optimal energy consumption ratio of communication between a UAV and a user terminal cluster in the to-be-served area according to the distribution of the hotspot area, and taking the coordinate point sequence as a flight point sequence of the UAV; and obtaining a flight trajectory of the UAV in the to-be-served area based on the flight point sequence of the UAV. The optimal energy consumption ratio of communication between the UAV and the corresponding user terminal cluster in the to-be-served area is calculated, the signal transmission capacity of the entire communication system is improved, each hotspot area is quickly covered by dividing the hotspot area, the energy consumption of the UAV caused by exploring the hotspot area in the early stage is reduced, the signal transmission efficiency is improved, and finally the optimal trajectory planning of the UAV in the to-be-served area is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of unmanned aerial vehicle (UAV) technology, and in particular to a UAV trajectory planning method, apparatus, electronic device, and storage medium. Background Technology

[0002] In recent years, frequent disasters such as earthquakes and floods, as well as terrorist attacks, have disrupted terrestrial mobile networks, leading to communication outages in the region. Simultaneously, damaged transportation routes often hinder the movement of people and vehicles, and ground infrastructure is difficult to repair quickly. Search and rescue teams frequently rely on public safety communication networks, which only support voice services and cannot provide more advanced services, further complicating search and rescue efforts. Existing point-to-point and satellite communication methods are difficult to widely adopt due to challenges in route optimization and limited available site resources. Drones, with their advantages of small size, easy deployment, high flexibility, and low deployment cost, are considered an effective means of emergency communication. However, due to resource scarcity and limited drone battery capacity in emergency scenarios, proper drone trajectory planning is crucial. Existing drone trajectory planning methods do not consider energy loss during communication between the drone and corresponding user terminal clusters, resulting in poor signal transmission capabilities for the entire communication system. Furthermore, point-to-point coverage of the entire area based on the drone's maximum coverage radius leads to significant signal energy loss, reducing signal transmission efficiency. Summary of the Invention

[0003] This invention provides a method, apparatus, electronic device, and storage medium for planning the trajectory of unmanned aerial vehicles (UAVs), which solves the problems in the prior art where the energy loss during the communication process between the UAV and the corresponding user terminal cluster is not considered, resulting in poor signal transmission capability of the entire communication system. At the same time, the point-to-point coverage of the entire area based on the maximum coverage radius of the UAV causes significant signal energy loss.

[0004] This invention provides a method for planning the trajectory of an unmanned aerial vehicle (UAV), comprising:

[0005] The user terminals within the service area are clustered to obtain multiple user terminal clusters, and each user terminal cluster corresponds to a hotspot area.

[0006] Based on the distribution of hotspot areas, obtain the coordinate point sequence corresponding to the optimal energy consumption ratio of communication between the UAV and the user terminal cluster within the service area, and use the coordinate point sequence as the flight waypoint sequence of the UAV.

[0007] Based on the flight waypoint sequence of the UAV, the flight trajectory of the UAV within the service area is obtained.

[0008] According to a UAV trajectory planning method provided by the present invention, the step of obtaining the coordinate point sequence corresponding to the optimal energy consumption ratio of communication between the UAV and the user terminal cluster within the service area based on the distribution of hotspot areas includes:

[0009] In the area to be served, a deep reinforcement learning model covering the hotspot area is established, and the boundary points of the area to be served are used as the state limits of the deep reinforcement learning model.

[0010] Construct a reward function, which includes a ratio function of the amount of data transmitted by the UAV to the total energy consumption of the UAV and the coverage of the hotspot area;

[0011] The deep reinforcement learning model is solved based on the reward function to obtain the sequence of coordinate points corresponding to the optimal energy consumption ratio for communication between the UAV and the corresponding user terminal cluster within the service area.

[0012] According to a method for planning the trajectory of a UAV provided by the present invention, the method for acquiring the amount of data transmitted by the UAV includes:

[0013] The average link loss between the UAV and the corresponding user terminal cluster is obtained based on the channel model.

[0014] The average time-frequency resource ratio occupied by the user terminal in the hotspot area is obtained based on the beam scheduling model.

[0015] The average link loss, the average time-frequency resource ratio occupied by the user terminal in the hotspot area, and the communication duration between the UAV and the user terminal are input into the transmission model to obtain the amount of data transmitted by the UAV.

[0016] According to a UAV trajectory planning method provided by the present invention, the method for obtaining the total energy loss of the UAV includes:

[0017] The energy consumption model is used to obtain the energy loss of the UAV during uniform flight, the energy loss during accelerated flight, the energy loss during decelerated flight, and the energy loss during communication between the UAV and the user terminal.

[0018] The total energy loss of the UAV is the sum of the energy loss during uniform flight, the energy loss during accelerated flight, the energy loss during decelerated flight, and the energy loss during communication between the UAV and the user terminal.

[0019] According to a method for planning the trajectory of a drone provided by the present invention, the step of solving the deep reinforcement learning model based on the reward function includes:

[0020] Within the service area, a preset sequence of waypoints is used as a status value, where the status value is the current position coordinates.

[0021] Actions are determined based on the ε-greedy strategy;

[0022] Perform an action in the current state to obtain the next state and a reward value, wherein the reward value is obtained according to a reward function;

[0023] The reward value guides the action, resulting in the optimal next state value;

[0024] The deep reinforcement learning model is pushed forward to obtain the next state value at each time point in the future.

[0025] Based on the optimal next state value at each moment in the future, the optimal coordinate sequence and the corresponding optimal energy efficiency ratio for communication between the UAV and the corresponding user terminal cluster within the service area are obtained, and used as the output value of the deep reinforcement learning model.

[0026] According to the UAV trajectory planning method provided by the present invention, the reward function further includes constraints, the constraints including:

[0027] The signal transmission rates of the drone and the user terminal meet the signal transmission rate required by the user terminal.

[0028] The average link loss of the user terminal does not exceed the maximum link loss.

[0029] The displacement of the drone at any given time does not exceed the maximum coverage diameter of the drone;

[0030] The starting points of the drone's flight trajectory coincide.

[0031] According to a method for planning the trajectory of an unmanned aerial vehicle (UAV) provided by the present invention, the step of clustering user terminals within the service area includes:

[0032] Randomly select a user terminal and calculate the density value of the neighboring users of the user terminal within the cluster radius;

[0033] If the user density value in the search neighborhood is greater than or equal to a preset threshold, then the user terminal and the user terminals within the cluster radius are divided into a user terminal cluster.

[0034] Repeat the above steps until the search neighborhood user density value corresponding to a user terminal that is not in the user terminal cluster is less than the preset threshold.

[0035] The present invention also provides a drone trajectory planning device, comprising:

[0036] The hotspot detection module is used to cluster user terminals within the service area, obtain multiple user terminal clusters, and each user terminal cluster corresponds to a hotspot area.

[0037] The solution module is used to obtain the coordinate point sequence corresponding to the optimal energy consumption ratio of communication between the UAV and the user terminal cluster within the service area based on the distribution of hotspot areas, and to use the coordinate point sequence as the flight waypoint sequence of the UAV.

[0038] The trajectory planning module is used to obtain the flight trajectory of the UAV within the service area based on the sequence of flight waypoints of the UAV.

[0039] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement any of the above-described drone trajectory planning methods.

[0040] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the UAV trajectory planning method as described above.

[0041] This invention provides a method, apparatus, electronic device, and storage medium for UAV trajectory planning. It involves clustering user terminals within a service area to obtain multiple user terminal clusters, each corresponding to a hotspot area. Based on the distribution of hotspot areas, a sequence of coordinate points corresponding to the optimal energy consumption ratio for communication between the UAV and the user terminal clusters within the service area is obtained. This sequence of coordinate points is used as the UAV's flight path sequence. Based on the UAV's flight path sequence, the flight trajectory of the UAV within the service area is obtained. By calculating the optimal energy consumption ratio for communication between the UAV and the corresponding user terminal clusters within the service area, the signal transmission capability of the entire communication system is improved. Simultaneously, by dividing the area into hotspot areas, rapid coverage of communication in each hotspot area is achieved, reducing energy loss caused by the UAV exploring hotspot areas in the early stages, improving signal transmission efficiency, and ultimately realizing optimal trajectory planning for the UAV within the service area. Attached Figure Description

[0042] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0043] Figure 1 This is one of the flowcharts illustrating the UAV trajectory planning method provided by the present invention;

[0044] Figure 2 This is the second flowchart illustrating the UAV trajectory planning method provided by the present invention;

[0045] Figure 3 This is the third flowchart illustrating the UAV trajectory planning method provided by the present invention;

[0046] Figure 4 This is the fourth flowchart illustrating the UAV trajectory planning method provided by the present invention;

[0047] Figure 5 This is the fifth flowchart illustrating the UAV trajectory planning method provided by the present invention;

[0048] Figure 6 This is the sixth flowchart illustrating the UAV trajectory planning method provided by the present invention;

[0049] Figure 7 This is the seventh flowchart illustrating the UAV trajectory planning method provided by the present invention;

[0050] Figure 8 This is the eighth flowchart of the UAV trajectory planning method provided by the present invention;

[0051] Figure 9 This is a schematic diagram of the structure of the UAV trajectory planning device provided by the present invention;

[0052] Figure 10 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0054] Figure 1 This is a flowchart illustrating the UAV trajectory planning method provided by the present invention, as shown below. Figure 1 As shown, the UAV trajectory planning method provided by this invention includes:

[0055] Step 101: Cluster the user terminals in the area to be served to obtain multiple user terminal clusters, with each user terminal cluster corresponding to a hotspot area;

[0056] Step 102: Based on the distribution of hotspot areas, obtain the coordinate point sequence corresponding to the optimal energy consumption ratio of communication between UAVs and user terminal clusters within the service area, and use the coordinate point sequence as the flight path point sequence of the UAVs.

[0057] Step 103: Based on the flight waypoint sequence of the UAV, obtain the flight trajectory of the UAV within the service area.

[0058] The existing trajectory planning methods used by UAVs do not take into account the energy loss during the communication process between the UAV and the corresponding user terminal cluster, resulting in poor signal transmission capability of the entire communication system. At the same time, since the entire area is covered based on the maximum coverage radius of the UAV, the signal energy loss is large, which reduces the signal transmission efficiency.

[0059] This invention provides a method for UAV trajectory planning. It involves clustering user terminals within a service area to obtain multiple user terminal clusters, each corresponding to a hotspot area. Based on the distribution of hotspot areas, a sequence of coordinate points corresponding to the optimal energy consumption ratio for communication between the UAV and the user terminal clusters within the service area is obtained. This sequence of coordinate points is used as the UAV's flight path sequence. Based on the UAV's flight path sequence, the flight trajectory of the UAV within the service area is obtained. By calculating the optimal energy consumption ratio for communication between the UAV and the corresponding user terminal clusters within the service area, the signal transmission capability of the entire communication system is improved. Simultaneously, by dividing the system into hotspot areas, rapid coverage of communication in each hotspot area is achieved, reducing energy loss caused by the UAV exploring hotspot areas in the early stages, improving signal transmission efficiency, and ultimately realizing optimal trajectory planning for the UAV within the service area.

[0060] Based on any of the above embodiments, such as Figure 2 As shown, user terminals within the service area are clustered, including:

[0061] Step 201: Randomly select a user terminal and calculate the density value of the neighboring users within the cluster radius of the user terminal;

[0062] Step 202: If the user density value of the search neighborhood is greater than or equal to the preset threshold, then the user terminal and the user terminals within the cluster radius are divided into a user terminal cluster.

[0063] Step 203: Repeat the above steps until the search neighborhood user density value of the corresponding user terminal that is not in the user terminal cluster is less than the preset threshold.

[0064] Because the embodiments of the present invention are aimed at coverage compensation scenarios for hotspot areas in emergency communication, the UAV needs to dynamically cover the hotspot areas during flight to achieve additional coverage compensation and provide a basis for subsequent trajectory planning. Therefore, it is necessary to determine the hotspot areas of the area to be served. By determining the hotspot areas, the energy loss caused by the UAV in exploring the hotspot areas in the early stage is reduced. At the same time, it is conducive to the UAV to quickly cover the hotspot areas, improving signal recovery efficiency and quality.

[0065] This invention uses the DBSCAN algorithm to cluster ground user terminals based on their distribution density, with each cluster representing a hotspot area. DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is a typical density-based clustering method. It defines a cluster as the largest set of density-connected points, enabling the division of areas with sufficient density into clusters and the discovery of clusters of arbitrary shapes in noisy spatial datasets.

[0066] Based on the characteristics of emergency communication scenarios, key coverage is provided for hotspot areas where users gather after disasters. In the areas awaiting service, a hotspot detection algorithm based on the DBSCAN algorithm is used. Combined with historical user distribution information, the distribution of ground users is analyzed and clustered to determine the number and range of hotspot areas in this area. This facilitates trajectory planning for hotspot areas by UAV-based aerial base stations.

[0067] Based on any of the above embodiments, such as Figure 3 As shown, based on the distribution of hotspot areas, the sequence of coordinate points corresponding to the optimal energy consumption ratio for communication between UAVs and user terminal clusters within the service area is obtained, including:

[0068] Step 301: In the area to be served, establish a deep reinforcement learning model that covers the hotspot area, and use the boundary points of the area to be served as the state limits of the deep reinforcement learning model.

[0069] Step 302: Construct a reward function, which includes the ratio of the amount of data transmitted by the drone to the total energy consumption of the drone, as well as the coverage of hotspot areas;

[0070] Step 303: Solve the deep reinforcement learning model based on the reward function to obtain the coordinate point sequence corresponding to the optimal energy consumption ratio of communication between UAVs and user terminal clusters within the service area.

[0071] In this embodiment of the invention, the Deep Q Network (DQN) algorithm is used to establish a deep reinforcement learning model covering hotspot regions. The DQN algorithm is an improvement on the Q-learning algorithm by applying neural networks, directly calculating the optimal action using the neural network instead of a Q-table. The DQN algorithm is a value-based algorithm that relies on the interaction between the agent and the environment to complete the learning process, obtaining the state s by observing the environment. t The agent obtains all Q(s,a) values ​​for the given state using a value neural network, and then determines the action a based on an ε-greedy policy. t The environment will respond with a reward r based on the reward function. tand new state s t+1 The DQN algorithm iterates continuously according to the above process and the historical experience stored in the experience pool until it obtains the optimal state.

[0072] In this embodiment of the invention, user terminals are divided into multiple clusters based on their location and communication needs, with the center of each cluster serving as the coverage compensation communication target for the UAV. To effectively meet the emergency service needs of the service area and facilitate trajectory calculation and focused coverage of hotspot areas, the UAV's flight process is emphasized as several waypoints. The UAV's flight trajectory is obtained by presetting different numbers of waypoints and solving for these waypoints.

[0073] The drone adjusts its waypoint positions based on the current environment and the rewards it receives to complete its movement. Because each action choice is discrete and independent, the entire movement process can be modeled as a Markov Decision Process (MDP), which is a formal description of the environment in reinforcement learning, or a model of the environment in which the agent exists. In reinforcement learning, almost all problems can be formally represented as a Markov Decision Process. When planning its flight trajectory, the Deep Reinforcement Learning (DRL) algorithm can continuously change the waypoint positions according to the environment, finding its optimal position more quickly.

[0074] Based on any of the above embodiments, such as Figure 4 As shown, the methods for obtaining the amount of data transmitted by the drone include:

[0075] Step 401: Obtain the average link loss for communication between the UAV and the corresponding user terminal cluster based on the channel model;

[0076] In this embodiment of the invention, the deployment of drones is to address the problem of dense user areas. After the basic aerial base station completes basic coverage of the entire area, the drones provide coverage compensation for hotspot areas; therefore, the scenario is a dense urban environment. In this scenario, since the communication between the drone and the user terminal sometimes encounters obstacles, resulting in non-line-of-sight links, both line-of-sight and non-line-of-sight links are considered in the transmission link established between the drone and the user terminal.

[0077] The average link loss for communication between the unmanned aerial vehicle (UAV) and user terminal i within time slot t. It can be obtained through the following methods:

[0078]

[0079] Where P(LoS) is the line-of-sight propagation probability and P(NLoS) is the non-line-of-sight propagation probability, the line-of-sight propagation probability and the non-line-of-sight propagation probability can be obtained by the following methods:

[0080] P(LoS)=a(θ i (t)-θ o ) b

[0081] P(NLoS)=1-a(θ i (t)-θ o ) b

[0082] Where a is the first environmental coefficient, b is the second environmental coefficient, and a and b take different values ​​depending on the environment, θ i (t) represents the angle between the transmission link established between the UAV and the user terminal and the ground, θ. i ∈[θ o [90°]. θ i (t) can be obtained through the following method:

[0083]

[0084] Where (x(t), y(t), z(t)) are the coordinates of the UAV, (x i y i h i () represents the coordinates of user terminal i.

[0085] For example, (a,b)=(0.33,0.23), (PL NLoS ,PL LoS )=(2,2.65), θ o It is 15°.

[0086] Step 402: Obtain the average time-frequency resource ratio occupied by user terminals in the hotspot area based on the beam scheduling model;

[0087] In this embodiment of the invention, before data transmission, the UAV and the user terminal need to perform main beam alignment to ensure that the user receives a high-quality signal. In this embodiment, a two-stage beam search scheme is used for beam alignment. First, a small number of wide beams are used for coarse-grained scanning in the communication sector of the UAV to determine the alignment direction between the wide beams and the user. Then, multiple narrow beams are used to scan the sector covered by the wide beams. Since the narrow millimeter-wave beams are highly directional, beam-level scanning is required to cover the entire consideration area, ultimately achieving beam alignment.

[0088] Beam alignment incurs time overhead, primarily occurring in the second stage, as the time spent in the first stage is negligible. The beam alignment time τ can be obtained using the following method:

[0089]

[0090] Where, δ T,s δ represents the width of the sector at the drone location. R,s δ represents the width of the sector at the user terminal. T,i δ represents the beamwidth at the UAV end. R,i T represents the beamwidth at the user terminal, where T p The time for beam alignment as the beam travels across the entire sector and sends a guide signal at each location.

[0091] Within a hotspot area, when the number of user terminals receiving drone signals is greater than or equal to the number of beams transmitted by the drone, a round-robin scheme is required for beam scheduling. The average time-frequency resource ratio occupied by the user terminals is the ratio of the number of beams transmitted by the drone to the number of user terminals receiving drone signals within the hotspot area; otherwise, the average time-frequency resource ratio occupied by the user terminals is set to 1. The average time-frequency resource ratio η occupied by the user terminal (User Equipment, UE) is... u Approximately:

[0092]

[0093] Where, N b N represents the number of beams emitted by the drone. u This represents the number of user terminals receiving drone signals within the hotspot area.

[0094] Step 403: Input the average link loss and the ratio of average time-frequency resources occupied by user terminals in the hotspot area into the transmission model to obtain the amount of data transmitted by the UAV.

[0095] In this embodiment of the invention, the amount of data transmitted by the UAV is: the total amount of data transmitted to all user terminals covered by the UAV during the UAV's flight time T, wherein the total amount of data transmitted is in bits. The amount of data transmitted by the UAV. It can be obtained through the following methods:

[0096]

[0097] in, This represents the total data transmission volume of the i-th user terminal, expressed in bits. It can be obtained through the following methods:

[0098]

[0099] Among them, during the communication between the UAV and the user terminal, the information throughput of the i-th user terminal within the coverage area is given by R, where the information throughput is in nats and is obtained using Shannon's formula. i (t) can be obtained through the following method:

[0100]

[0101] Where, η u b represents the average time-frequency resource ratio occupied by user terminals. U τ represents the bandwidth allocated to all user terminals within the drone's connectivity range, τ is the beam alignment time, and t is the bandwidth allocated to all user terminals within the drone's connectivity range. tr The time during which the drone maintains communication with the i-th user terminal. Let be the signal-to-noise ratio when the UAV communicates with the i-th user terminal.

[0102] To simplify T tr The calculations are performed, and subsequent simulations take the cluster center of the entire user terminal cluster as the communication target. The average transmission time is the time it takes for the UAV to maintain communication with the user terminal corresponding to each cluster center. v represents the constant speed of the drone's flight. The maximum coverage radius of the drone. Among them, h U θ is the flight altitude of the drone. max To ensure that the average link loss does not exceed the maximum link loss PL max The maximum communication angle between the UAV and the cluster center of the user terminal cluster under the given conditions.

[0103] During communication between the user terminal and the drone, the user terminal receives service signals from other base stations and drones. Aside from the signals from the base station and drone serving the user terminal, the signals from other base stations and drones are considered noise. It can be obtained through the following methods:

[0104]

[0105] The signal-to-noise ratio is expressed in dB. This represents the signal power received by the drone from the i-th user terminal. Represents base station (BS) m The transmitted signal power, D is the set of all nearby ground base stations that may cause interference, h A For the free-space propagation path loss of the BS (Base Station), neglecting channel gain, σ 2This represents the power of Additive White Gaussian Noise (AWGN).

[0106] The signal strength received by the i-th user terminal is the transmit power of the UAV main lobe minus the path loss during signal transmission. Therefore, the received power of the i-th user terminal can be obtained by the following method:

[0107]

[0108] Pt UAV G represents the transmission power of the drone. M This represents the main lobe gain of the drone. This represents the path loss between the user terminal and the drone at any specified location.

[0109] Besides air-to-ground (A2G) propagation path loss, the gain of a directional millimeter-wave antenna is also a significant factor affecting power. For ease of calculation, we assume that a three-dimensional beam of the UAV has a uniform gain GM within its beamwidth and a small, constant sidelobe gain GS outside the beamwidth. The main lobe gain can be obtained as follows, where δ i It is the cone half-angle of the millimeter-wave beam of the i-th user terminal.

[0110]

[0111] Therefore, the amount of data transmitted by the drone Based on the above calculation results, the following method can be used to obtain:

[0112]

[0113] The communication process between the UAV and the user terminal was refined through channel modeling, beam scheduling modeling, and transmission modeling. Specifically, the channel model comprehensively considers line-of-sight (LAS) and non-LAS link losses, resulting in more accurate calculations of the average link loss. The beam scheduling model refines the communication duration between the UAV and the user terminal by removing the beam alignment time, yielding the actual communication duration. The transmission model incorporates average link loss and signal loss before communication between the UAV and the user terminal as influencing factors, leading to more accurate calculations of the UAV's transmitted data volume.

[0114] Based on any of the above embodiments, such as Figure 5 As shown, the methods for obtaining the total energy loss of a drone include:

[0115] Step 501: Based on the energy consumption model, obtain the energy loss of the UAV during uniform flight, the energy loss during accelerated flight, the energy loss during decelerated flight, and the energy loss during communication between the UAV and the user terminal.

[0116] In this embodiment of the invention, the UAV maintains data service throughout its flight. Energy loss primarily comprises two parts: flight energy, mobility energy, and communication energy. When a rotary-wing UAV communicates with a user terminal at a certain location, its initial speed is 0. It needs to accelerate to its maximum flight speed and then decelerate before reaching that location. Therefore, mobility energy includes energy from different states. The energy loss of the UAV during flight is E. F The hovering energy loss is E H The communication energy is E C The energy loss during accelerated flight is E A The energy loss during deceleration flight is E D .

[0117] P H Hovering power can be obtained in the following way:

[0118]

[0119] Where Ro is the number of rotors of the UAV, G represents gravity (G=Mg, where m is the weight of the UAV and g is the gravitational acceleration), ρ represents the air density, and β represents the rotor disk radius.

[0120] T F Let E be the flight time of the drone, and the energy loss during uniform flight is E. F It can be obtained in the following ways:

[0121]

[0122] Where P H denoted as hovering power, and f as air resistance.

[0123] T C The communication time between the drone and the user terminal, and the communication energy loss E C It can be obtained in the following ways:

[0124]

[0125] Where Pt UAV This refers to the transmission power of the drone.

[0126] Accelerated flight energy loss E A and energy loss during deceleration flight E D It can be obtained in the following ways:

[0127]

[0128] Where v represents the maximum speed that the drone can reach, that is, the speed at which the drone flies at a constant speed, m is the weight of the drone, and a is the acceleration of the drone.

[0129] Step 502: The sum of the energy loss of the UAV during uniform flight, the energy loss during accelerated flight, the energy loss during decelerated flight, and the energy loss during communication between the UAV and the user terminal is taken as the total energy loss of the UAV.

[0130] The total energy loss E of the drone can be obtained in the following way:

[0131]

[0132] The total energy loss E of the drone is expressed in joules.

[0133] Because the acceleration and deceleration time of a drone is very short and can be ignored during long-distance flight, the drone maintains a constant speed for most of the time. Air resistance can be considered constant, and communication is maintained throughout the flight. Therefore, the total energy loss E of the drone over all flight times T can be considered as the energy loss from constant speed flight and communication loss. Thus, the total energy loss E between the drone and the human-machine interface can be simplified as follows:

[0134]

[0135] Based on the characteristics of air-to-ground channels, the channel model, beam scheduling model, transmission model, and energy consumption model of the UAV are comprehensively considered to improve the energy efficiency ratio of the system and ultimately obtain the optimal trajectory of the UAV in the hotspot area.

[0136] Based on any of the above embodiments, such as Figure 6 As shown, solving the deep reinforcement learning model based on the reward function includes:

[0137] Step 601: Preset a sequence of waypoints in the area to be served as the status value. The status value is the current position coordinates.

[0138] In this embodiment of the invention, the state space S is used to represent the 3D position coordinate sequence of n points along the path of the UAV, S=[[x1,y1,h1],[x2,y2,h2],…,[x n ,y n ,h n ]], representing the x, y, and z coordinates of the n points along the route.

[0139] Step 602: Determine the action based on the ε-greedy strategy;

[0140] In this embodiment of the invention, the motion space A represents the 3D position change A = a at the nth path point of the UAV.n a n Indicates the drone's waypoints [x] n ,y n ,h n The particle size varies in the x-axis or y-axis direction, with a particle size of 1m.

[0141] In this embodiment of the invention, if the random action is less than ε, an action corresponding to the maximum value of Q(s,a) is selected through the value neural network of the DQN algorithm; if the random action is greater than or equal to ε, an action is randomly selected from the action space.

[0142] Step 603: Execute an action under the current state value to obtain the next state value and reward value. The reward value is obtained according to the reward function.

[0143] Step 604: Guide actions with reward values ​​to obtain the optimal next state value;

[0144] Step 605: Proceed the deep reinforcement learning model forward to obtain the next state value at each time step in the future.

[0145] Step 606: Based on the optimal next state value at each moment in the future, obtain the optimal coordinate point sequence and the corresponding optimal energy efficiency ratio for communication between the UAV and the corresponding user terminal cluster within the service area, and use them as the output value of the deep reinforcement learning model.

[0146] In this embodiment of the invention, the reward value calculated using a reward function guides the drone's actions. A larger reward value indicates that the aerial base station's actions are closer to the optimization target. To achieve the optimization goal of maximizing drone energy efficiency while simultaneously providing full coverage of hotspot areas, the reward function is defined as follows:

[0147]

[0148] This indicates the drone's energy efficiency at the current location, where ∈ is the proportionality coefficient, C represents all hotspot areas, and V represents the access status of each hotspot area.

[0149]

[0150] In this embodiment of the invention, the reward function further includes constraints, which include:

[0151] The signal transmission rates of the drone and the user terminal meet the signal transmission rate required by the user terminal.

[0152] The average link loss of the user terminal does not exceed the maximum link loss.

[0153] The displacement of the drone at any given moment shall not exceed the maximum coverage diameter of the drone;

[0154] The starting points of the drones' flight paths coincide.

[0155] In this embodiment of the invention, the optimal energy efficiency ratio and its constraints are as follows:

[0156]

[0157] in,

[0158]

[0159]

[0160]

[0161] u(0)=u(f)

[0162] Constraint 1: The transmission rates of the base station and the drone must meet the user's required rate. i (t) represents the signal transmission rate between the i-th user terminal and the UAV.

[0163] Constraint 2: The average path loss of user terminals shall not exceed the maximum path loss, that is, the user terminals shall be within the coverage area and meet the QoS (Quality of Service) requirements, and the coverage area shall be maximized as much as possible. For the path loss between the i-th user terminal and the drone, PL U max This represents the maximum path loss.

[0164] Constraint 3: The movement of the drone at any given moment shall not exceed the maximum diameter. This ensures continuous and complete coverage for users in the disaster area. u(t) represents the location coordinates of the drone at each time point.

[0165] Constraint 4: The starting points of the paths coincide, and the drone returns to the origin after flying one cycle, starting a new flight cycle.

[0166] In this embodiment of the invention, in the service area, the flight trajectory of the UAV is represented by a preset number of waypoint sequences. The first state value and action value of the waypoint sequence are randomly initialized. The first state value is the current 3D coordinates of the waypoint sequence, and the action value is the change of the current 3D coordinates of the waypoint sequence.

[0167] The action value is determined according to the ε-greedy policy. After the action is executed in the waypoint sequence, the reward value and the second state value are obtained. The reward value is the reward value after the environment receives the action value, and the second state value is the next 3D coordinate of the waypoint sequence.

[0168] The sample consisting of a first state value, an action value, a reward value, and a second state value is stored in the first sample set. A new sample is generated by iterating the first state value with the second state value and then stored in the sample set.

[0169] Randomly sample from the sample set, use the first state value and action value as input to the prediction value neural network to calculate the prediction Q value, and continuously update the weight parameters of the prediction value neural network based on the reward value and the prediction Q value;

[0170] Every so often, the weight parameters of the target value neural network are updated to the weight parameters of the prediction value neural network;

[0171] The loss value is calculated using the loss function based on the target Q value and the predicted Q value. The loss value is then updated using gradient descent until the target value neural network converges, at which point the optimal trajectory is obtained.

[0172] In this embodiment of the invention, by setting conditions such as the number of waypoints and user distribution, the drone continuously explores and learns from the environment, moves waypoints to obtain new information, and explores the coordinates of the waypoints with the optimal energy efficiency ratio of the drone. This can achieve global search, avoid getting stuck in local optima, and enable rapid deployment by migrating between different environments.

[0173] Figure 7 The flowchart of the overall method for UAV trajectory planning provided in the embodiments of the present invention is as follows: Figure 7 As shown, the overall method for UAV trajectory planning provided in this embodiment of the invention includes:

[0174] Step 701: Based on the historical distribution information of ground user terminals, use clustering algorithms to identify hotspot areas;

[0175] Step 702: Establish an optimization model for communication between the airborne base station and the user cluster;

[0176] Step 703: Based on the optimization model, the path sequence corresponding to the optimal energy efficiency ratio of the area to be served is solved by the aerial base station trajectory planning algorithm based on full coverage of hotspot areas for emergency communication, and finally the high-energy-efficiency trajectory of the aerial base station with full coverage of hotspots is obtained.

[0177] In this embodiment of the invention, by considering signal loss during signal transmission and the energy efficiency of airborne base stations, the DBSCAN algorithm is used to determine hotspot areas, thereby determining the key coverage areas for airborne base station trajectory planning. Then, an airborne base station trajectory planning algorithm based on deep reinforcement learning is used to calculate the high energy efficiency trajectory of the airborne base station, ensuring emergency communication services in the areas to be served.

[0178] Figure 8The flowchart of the high-efficiency airborne base station trajectory planning algorithm provided in the embodiments of the present invention is as follows: Figure 8 As shown, the high-efficiency aerial base station trajectory planning algorithm provided in this embodiment of the invention includes:

[0179] Step 801: Based on the historical distribution information of ground user terminals, user hotspot areas are divided using the DBSCAN algorithm;

[0180] In this embodiment of the invention, the input of the user hotspot area segmentation algorithm includes: J is the user location dataset, ε is the clustering radius parameter, and M is the neighborhood user density threshold; the output includes: user clusters and cluster centers based on density clustering.

[0181] The specific steps are as follows:

[0182] (1) Mark all user points as unvisited. Each user point is represented by the location coordinates in the user location dataset J. Randomly select an unvisited user object p from the set of users marked as unvisited and update p to visited.

[0183] (2) If p has at least M user objects in its ε-neighborhood, create a new user cluster C and add p to C;

[0184] (3) Let N denote the set of user points in the ε-neighborhood of p. For each p' in N, if p' is unvisited, mark p' as visited; if p' has at least M user points in its ε-neighborhood, add these user points to N; if p' is not yet a member of any user cluster, add p' to C.

[0185] (4) Calculate the average of the x and y coordinates of all points to obtain the cluster center coordinates of the user cluster.

[0186] Step 802: In the service area, the flight trajectory of the UAV is represented by a preset number of waypoint sequences. The first state value and action value of the waypoint sequence are randomly initialized. The first state value is the current 3D coordinates of the waypoint sequence, and the action value is the change of the current 3D coordinates of the waypoint sequence.

[0187] Step 803: Determine the action value according to the ε-greedy policy. After the action is performed in the path point sequence, the reward value and the second state value are obtained. The reward value is the reward value after the environment receives the action value, and the second state value is the next 3D coordinate of the path point sequence. Store the sample composed of the first state value, action value, reward value and second state value into the first sample set, and use the second state value to iterate the first state value to generate a new sample, and store it into the sample set.

[0188] Step 804: Randomly sample from the sample set, use the first state value and action value as input to the prediction value neural network to calculate the prediction Q value, and continuously update the weight parameters of the prediction value neural network based on the reward value and the prediction Q value.

[0189] Step 805: Every certain period of time, update the weight parameters of the target value neural network to the weight parameters of the prediction value neural network; calculate the loss value using the loss function based on the target Q value and the prediction Q value, and update the target neural network using gradient descent based on the loss value until the target value neural network converges, thus obtaining the optimal trajectory.

[0190] This invention divides user terminals into multiple clusters based on their location and communication needs, with each cluster center serving as the coverage compensation communication target for the UAV. To effectively meet the emergency service needs of the service area and facilitate trajectory calculation and focused coverage of hotspot areas, the UAV's flight process is emphasized as several waypoints. The UAV's flight trajectory is obtained by presetting different numbers of waypoints and solving for these waypoints. By setting conditions such as the number of waypoints and user distribution, and utilizing human-machine interaction to continuously explore and learn from the environment, the UAV continuously moves waypoints to acquire new information, exploring the coordinates of waypoints with the optimal energy efficiency ratio for the UAV. This allows for global search, avoiding getting trapped in local optima, and enables migration between different environments, facilitating rapid deployment.

[0191] In this embodiment of the invention, the emergency communication airborne base station trajectory planning algorithm based on deep reinforcement learning is as follows:

[0192] (1) Randomly select an initial state;

[0193] (2) Based on the state, an ε-greedy strategy is used to select action a. If the random number is less than ε, the action with the largest Q value is selected through the Q-Network. If the random number is greater than ε, an action is randomly selected from the action space.

[0194] (3) After the action selection is completed, the agent performs action a in the environment, and then the environment returns to the next state (S_) and reward (R).

[0195] (4) Store the quadruple (S, A, R, S_) into the experience pool.

[0196] (5) Treat the next state (S_) as the current state (S) and repeat step (2).

[0197] (6) Update the network in DQN. That is, randomly sample from the experience pool, feed the state (S) and action (A) in the sample into the value network to calculate the Q value, and update the value network according to the actual reward (R) and the Q value.

[0198] (7) After the value network is updated a certain number of times, the parameters are copied to the target network to update the network parameters.

[0199] (8) Repeat steps (5)-(7) until the target Q-Network converges.

[0200] The inputs to the DQN-based airborne base station trajectory optimization algorithm include: user set U, number of waypoints N, number of algorithm iterations Z, and UAV flight altitude H. In this embodiment of the invention, the UAV flight altitude H is a constant value and is used to calculate the total amount of transmitted data. The outputs include: airborne base station operating trajectory and energy efficiency.

[0201] The specific steps are as follows:

[0202] (1) Initialize the user distribution state and the sequence of aerial base station waypoints s; initialize the Q network and the target value network. Randomly generate weights θ; set target values ​​for network weights θ - =θ;

[0203] (2) Under the following constraints: the signal transmission rates of the UAV and the user terminal meet the signal transmission rate required by the user terminal; the average link loss of the user terminal does not exceed the maximum link loss; the displacement of the UAV at each moment does not exceed the maximum coverage diameter of the UAV; the starting points of the UAV's flight trajectory coincide, and the movement of the airborne base station is determined according to the ε-greedy strategy. t If the random action is less than ε, then the action corresponding to the maximum value of Q(s,a) is selected through the value neural network of the DQN algorithm; if the random action is greater than or equal to ε, then an action is randomly selected from the action space.

[0204] (3) Execute action a t Receive reward R t (Value is the reward function) and the new state s t+1 ;

[0205] (4) The sample (s) t ,a t ,R t ,s t+1 ) Stored in memory pool D to facilitate the elimination of sample correlation;

[0206] (5) Randomly draw samples from the memory pool (s) j ,a j ,R j ,s j+1 );

[0207] (6) If the preset number of training steps is reached in the (j+1)th step, let y j =r j Otherwise, let Where r is the reward value for this step, and γ is the discount factor;

[0208] (7) For (y) j -Q(s j ,a j ;θ)) 2 The gradient descent method is used to update θ;

[0209] (8) Every C steps, the network parameters are synchronized to the target value network parameters. After training for a specified number of steps, the optimal path point sequence corresponding to the optimal energy efficiency ratio is obtained.

[0210] This invention provides an aerial base station trajectory planning mechanism for emergency communication in unserved areas of a communication network. First, it clarifies a high-efficiency trajectory planning model for the aerial base station. Then, it proposes a two-step mechanism for trajectory planning in hotspot areas during emergency communication: First, based on the DBSCAN clustering algorithm, users are clustered according to the user distribution density in the unserved area to obtain hotspot areas, which are then used as key areas for subsequent trajectory planning. The aerial base station communicates with the user cluster heads, and its coverage radius is calculated according to the model, ensuring that the aerial base station covers the hotspot areas within its coverage area during flight. Finally, the flight trajectory is solved by finding waypoints using deep reinforcement learning methods.

[0211] The UAV trajectory planning device provided by the present invention is described below. The UAV trajectory planning device described below and the UAV trajectory planning method described above can be referred to in correspondence.

[0212] Figure 9 This is a schematic diagram of the structure of the UAV trajectory planning device provided in an embodiment of the present invention, as shown below. Figure 9 As shown, the UAV trajectory planning device provided in this embodiment of the invention includes:

[0213] The hotspot detection module 901 is used to cluster user terminals within the service area and obtain multiple user terminal clusters, with each user terminal cluster corresponding to a hotspot area.

[0214] The solution module 902 is used to obtain the coordinate point sequence corresponding to the optimal energy consumption ratio of communication between UAV and user terminal cluster within the service area based on the distribution of hotspot areas, and use the coordinate point sequence as the flight path point sequence of the UAV.

[0215] The trajectory planning module 903 is used to obtain the flight trajectory of the UAV within the service area based on the sequence of UAV waypoints.

[0216] The UAV trajectory planning device provided by this invention clusters user terminals within a service area, obtaining multiple user terminal clusters, each corresponding to a hotspot area. Based on the distribution of hotspot areas, it obtains a sequence of coordinate points corresponding to the optimal energy consumption ratio for communication between the UAV and the user terminal clusters within the service area, and uses this sequence as the UAV's flight path sequence. Based on the UAV's flight path sequence, it obtains the UAV's flight trajectory within the service area. By calculating the optimal energy consumption ratio for communication between the UAV and the corresponding user terminal clusters within the service area, the signal transmission capability of the entire communication system is improved. Simultaneously, by dividing the system into hotspot areas, rapid coverage of communication in each hotspot area is achieved, reducing energy loss caused by the UAV exploring hotspot areas in the early stages, improving signal transmission efficiency, and ultimately realizing optimal trajectory planning for the UAV within the service area.

[0217] Figure 10 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 10 As shown, the electronic device may include a processor 1010, a communications interface 1020, a memory 1030, and a communication bus 1040. The processor 1010, communications interface 1020, and memory 1030 communicate with each other via the communication bus 1040. The processor 1010 can call logical instructions in the memory 1030 to execute a UAV trajectory planning method. This method includes: clustering user terminals within the service area to obtain multiple user terminal clusters, each corresponding to a hotspot area; based on the distribution of hotspot areas, obtaining a sequence of coordinate points corresponding to the optimal energy consumption ratio for communication between the UAV and the user terminal clusters within the service area, and using this sequence of coordinate points as the UAV's flight path sequence; and based on the UAV's flight path sequence, obtaining the UAV's flight trajectory within the service area.

[0218] Furthermore, the logical instructions in the aforementioned memory 1030 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0219] On the other hand, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements the UAV trajectory planning method provided by the above methods. The method includes: clustering user terminals within a service area to obtain multiple user terminal clusters, each user terminal cluster corresponding to a hotspot area; obtaining a sequence of coordinate points corresponding to the optimal energy consumption ratio for communication between the UAV and the user terminal clusters within the service area according to the distribution of hotspot areas, and using the sequence of coordinate points as the flight path point sequence of the UAV; and obtaining the flight trajectory of the UAV within the service area based on the flight path point sequence of the UAV.

[0220] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0221] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.

[0222] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for planning the trajectory of an unmanned aerial vehicle (UAV), characterized in that, include: The user terminals within the service area are clustered to obtain multiple user terminal clusters, and each user terminal cluster corresponds to a hotspot area. Based on the distribution of hotspot areas, obtain the coordinate point sequence corresponding to the optimal energy consumption ratio of communication between the UAV and the user terminal cluster within the service area, and use the coordinate point sequence as the flight waypoint sequence of the UAV. Based on the flight waypoint sequence of the UAV, the flight trajectory of the UAV within the service area is obtained; The step of obtaining the coordinate point sequence corresponding to the optimal energy consumption ratio of communication between UAVs and user terminal clusters within the service area based on the distribution of hotspot areas includes: In the service area, a deep reinforcement learning model covering the hotspot area is established, and the boundary points of the service area are used as the state limits of the deep reinforcement learning model. The state value of the deep reinforcement learning model is the current position coordinate of the UAV in the service area based on a preset sequence of waypoints, and the action value is the change in the current position coordinate of the waypoint sequence. Construct a reward function, which includes a ratio function of the amount of data transmitted by the UAV to the total energy consumption of the UAV and the coverage of the hotspot area; The deep reinforcement learning model is solved based on the reward function to obtain the coordinate point sequence corresponding to the optimal energy consumption ratio of communication between UAVs and user terminal clusters within the service area; The total energy loss of the UAV is obtained based on an energy consumption model, which includes the energy loss of the UAV during uniform flight, the energy loss during accelerated flight, the energy loss during decelerated flight, and the energy loss during communication between the UAV and the user terminal. The amount of data transmitted by the UAV is obtained based on a channel model and a beam scheduling model. The channel model considers line-of-sight links and non-line-of-sight links. The beam scheduling model determines the beam alignment time based on a two-level beam search scheme and determines the average time-frequency resource ratio occupied by the user terminal based on the relationship between the number of beams emitted by the UAV and the number of user terminals receiving UAV signals in the hotspot area.

2. The UAV trajectory planning method according to claim 1, characterized in that, The method for obtaining the amount of data transmitted by the UAV includes: The average link loss between the UAV and the corresponding user terminal cluster is obtained based on the channel model. The average time-frequency resource ratio occupied by the user terminals within the hotspot area is obtained based on the beam scheduling model. The average time-frequency resource ratio is the ratio of the actual time-frequency resources occupied by the user terminals within the hotspot area when communicating with the UAV to the available time-frequency resources allocated to the UAV. When the number of user terminals receiving the UAV signal is greater than or equal to the number of beams transmitted by the UAV, the average time-frequency resource ratio is the ratio of the number of beams transmitted by the UAV to the number of user terminals receiving the UAV signal within the hotspot area; otherwise, the average time-frequency resource ratio is 1. The average link loss, the average time-frequency resource ratio occupied by the user terminal in the hotspot area, and the communication duration between the UAV and the user terminal are input into the transmission model to obtain the amount of data transmitted by the UAV. The transmission model is established based on Shannon's formula.

3. The UAV trajectory planning method according to claim 1, characterized in that, The method for obtaining the total energy loss of the UAV includes: The energy loss of the UAV during uniform flight, acceleration flight, deceleration flight, and communication between the UAV and the user terminal are obtained based on the energy consumption model. The total energy loss of the UAV is the sum of the energy loss during uniform flight, the energy loss during accelerated flight, the energy loss during decelerated flight, and the energy loss during communication between the UAV and the user terminal.

4. The UAV trajectory planning method according to claim 1, characterized in that, Solving the deep reinforcement learning model based on the reward function includes: Within the service area, a preset sequence of waypoints is used as a status value, where the status value is the current position coordinates. Actions are determined based on the ε-greedy strategy; Perform an action in the current state to obtain the next state and a reward value, wherein the reward value is obtained according to a reward function; The reward value guides the action, resulting in the optimal next state value; The deep reinforcement learning model is pushed forward to obtain the next state value at each time point in the future. Based on the optimal next state value at each moment in the future, the optimal coordinate sequence and the corresponding optimal energy efficiency ratio for communication between the UAV and the corresponding user terminal cluster within the service area are obtained, and used as the output value of the deep reinforcement learning model.

5. The UAV trajectory planning method according to claim 1, characterized in that, The reward function also includes constraints, which include: The signal transmission rates of the drone and the user terminal meet the signal transmission rate required by the user terminal. The average link loss of the user terminal does not exceed the maximum link loss. The displacement of the drone at any given time does not exceed the maximum coverage diameter of the drone; The starting points of the drone's flight trajectory coincide.

6. The UAV trajectory planning method according to claim 1, characterized in that, The process of clustering user terminals within the service area includes: Randomly select a user terminal and calculate the density value of the neighboring users of the user terminal within the cluster radius; If the user density value in the search neighborhood is greater than or equal to a preset threshold, then the user terminal and the user terminals within the cluster radius are divided into a user terminal cluster. Repeat the above steps until the search neighborhood user density value corresponding to a user terminal that is not in the user terminal cluster is less than the preset threshold.

7. A drone trajectory planning device, characterized in that, include: The hotspot detection module is used to cluster user terminals within the service area, obtain multiple user terminal clusters, and each user terminal cluster corresponds to a hotspot area. The solution module is used to obtain the coordinate point sequence corresponding to the optimal energy consumption ratio of communication between the UAV and the user terminal cluster within the service area based on the distribution of hotspot areas, and to use the coordinate point sequence as the flight waypoint sequence of the UAV. The trajectory planning module is used to obtain the flight trajectory of the UAV within the service area based on the sequence of flight waypoints of the UAV. The step of obtaining the coordinate point sequence corresponding to the optimal energy consumption ratio of communication between UAVs and user terminal clusters within the service area based on the distribution of hotspot areas includes: In the service area, a deep reinforcement learning model covering the hotspot area is established, and the boundary points of the service area are used as the state limits of the deep reinforcement learning model. The state value of the deep reinforcement learning model is the current position coordinate of the UAV in the service area based on a preset sequence of waypoints, and the action value is the change in the current position coordinate of the waypoint sequence. Construct a reward function, which includes a ratio function of the amount of data transmitted by the UAV to the total energy consumption of the UAV and the coverage of the hotspot area; The deep reinforcement learning model is solved based on the reward function to obtain the coordinate point sequence corresponding to the optimal energy consumption ratio of communication between UAVs and user terminal clusters within the service area; The total energy loss of the UAV is obtained based on an energy consumption model, which includes the energy loss of the UAV during uniform flight, the energy loss during accelerated flight, the energy loss during decelerated flight, and the energy loss during communication between the UAV and the user terminal. The amount of data transmitted by the UAV is obtained based on a channel model and a beam scheduling model. The channel model considers line-of-sight links and non-line-of-sight links. The beam scheduling model determines the beam alignment time based on a two-level beam search scheme and determines the average time-frequency resource ratio occupied by the user terminal based on the relationship between the number of beams emitted by the UAV and the number of user terminals receiving UAV signals in the hotspot area.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the UAV trajectory planning method as described in any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the UAV trajectory planning method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Trajectory optimization and resource allocation method for single unmanned aerial vehicle backscatter communication network

    CN112532300A

  • Path planning method based on multi-unmanned aerial vehicle auxiliary data collection

    CN114879726A