Multi-uav trajectory planning method and system

CN117873172BActive Publication Date: 2026-09-29UNIV OF SCI & TECH OF CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410059245.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-15
Publication Date
2026-09-29
Estimated Expiration
2044-01-15

AI Technical Summary

Technical Problem

传统的轨迹设计方法严重依赖于对环境的预先了解,包括地理信息和潜在的UE分布,而在真实场景中可能无法获得这些信息

Benefits of technology

[0021]由上述本发明提供的技术方案可以看出,具有与环境交互的能力,能够适应不同的UE分布,并且借助于深度强化学习中智能体与环境交互以优化决策的能力,能够实现在未知环境下无人机的最优飞行轨迹规划,在不同场景下与现有解决方案相比具有显著的性能提升。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117873172B_ABST
    Figure CN117873172B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-unmanned aerial vehicle trajectory planning method and system, multi-unmanned aerial vehicle three-dimensional trajectory planning problem is described as multi-objective optimization problem, to improve network coverage, access fairness, data transmission volume and energy efficiency simultaneously.The application is inspired by the decision-making ability of DRL, and the multi-unmanned aerial vehicle trajectory planning (Deep Transfer Reinforcement Learning Based Multi-UAV Trajectory Design, TL-RLMTD) scheme based on deep transfer learning can reduce the energy consumption of unmanned aerial vehicles, improve the UE access rate, fairness and the data transmission volume of the entire system while having the ability to explore UE clusters and find the optimal height.In addition, the unmanned aerial vehicle of the application has the ability to interact with the environment and can adapt to different user equipment distribution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of autonomous flight trajectory planning for multiple unmanned aerial vehicles (UAVs), and more particularly to a method and system for planning the trajectory of multiple UAVs. Background Technology

[0002] Traditional cellular communication networks are facing unprecedented challenges from the ever-increasing number of user equipment (UEs) demanding high-quality wireless network services. To address the high deployment and operational costs of terrestrial network infrastructure, unmanned aerial vehicle (UAV)-based communication networks have been proposed as a low-cost and flexible supplementary solution. Deploying UAV mobile base stations (UAV-BSs) in areas lacking communication infrastructure can effectively improve the coverage and performance of wireless communication networks. In scenarios where terrestrial networks are overloaded due to crowds, UAVs can be rapidly deployed as aerial base stations, flexibly adjusting their spatial location based on the movement of terrestrial user equipment to improve Quality of Service (QoS). Furthermore, when terrestrial network infrastructure is destroyed by natural disasters, deploying UAV-BSs can quickly establish emergency communication networks.

[0003] In drone applications, trajectory design is a fundamental problem in leveraging the maneuverability of drones. Traditional trajectory design methods rely heavily on prior knowledge of the environment, including geographic information and potential user unit (UE) distribution, which may not be available in real-world scenarios.

[0004] Furthermore, existing studies assume that UEs are uniformly distributed in the target area, and some studies only focus on the coverage of ground units without considering the distribution of UEs. Therefore, existing solutions lack the ability to adapt to different UE distributions. In addition, most studies use a centralized control method, where the control center collects global information, makes decisions, and sends commands to all UAVs, which introduces additional communication load to the system. Summary of the Invention

[0005] The purpose of this invention is to provide a multi-UAV trajectory planning method and system that enables autonomous decision-making by UAVs, distributed control of UAVs, and the ability to interact with the environment, adapting to different user equipment distributions.

[0006] The objective of this invention is achieved through the following technical solution:

[0007] A multi-UAV trajectory planning method includes:

[0008] A three-layer network architecture consisting of multiple ground user devices, a group of UAVs, and an aerial platform is constructed, and an intelligent agent is deployed in each UAV. The intelligent agent is pre-trained in the source domain environment and then migrates to the target domain environment. The distribution of user devices in the source domain environment and the target domain environment is the same, and the influence of user device movement is ignored during pre-training.

[0009] In the target domain environment, each agent in the UAV constructs its own local state of the current time slot by combining the global observation state of the current time slot provided by the high-altitude platform, and generates the action of the current time slot according to the action strategy. The action refers to the flight action of the UAV.

[0010] Based on the utility function of the drone after performing the corresponding action, the drone for service is selected in combination with the location of each user device;

[0011] The reward for each drone in the current time slot is calculated by combining the number of user devices for all drone services in the current time slot, the fairness index, the number of user devices and energy consumption for each drone service, and the legality of the action; the fairness index is calculated based on the access rate of user devices.

[0012] For each UAV's pre-trained agent, the state transition information of the current time slot is constructed using the local state, action, reward of the current time slot and the local state of the next time slot. After calculating the weight of the state transition information, the state transition information and the corresponding weight are stored in the replay buffer. At the same time, the absolute TD error of the state transition information of the current time slot is calculated. If it exceeds the threshold, it is sampled according to the weight of the state transition information, and then the sampled state transition information is used to train the agent.

[0013] A multi-UAV trajectory planning system, comprising:

[0014] The network architecture building unit is used to build a three-layer network architecture consisting of multiple ground user devices, a group of drones, and a high-altitude platform, and to deploy an intelligent agent in each drone.

[0015] The pre-training unit is used to pre-train the agent in the source domain environment; the pre-training is transferred to the target domain environment, and the distribution of user devices is the same in the source domain environment and the target domain environment. The influence of user device movement is ignored during pre-training.

[0016] The drone trajectory planning and agent training unit is used for:

[0017] In the target domain environment, each agent in the UAV constructs its own local state of the current time slot by combining the global observation state of the current time slot provided by the high-altitude platform, and generates the action of the current time slot according to the action strategy. The action refers to the flight action of the UAV.

[0018] Based on the utility function of the drone after performing the corresponding action, select the drone to serve each user device;

[0019] The reward for each drone in the current time slot is calculated by combining the number of user devices for all drone services in the current time slot, the fairness index, the number of user devices and energy consumption for each drone service, and the legality of the action; the fairness index is calculated based on the access rate of user devices.

[0020] For each UAV's pre-trained agent, the state transition information of the current time slot is constructed using the local state, action, reward of the current time slot and the local state of the next time slot. After calculating the weight of the state transition information, the state transition information and the corresponding weight are stored in the replay buffer. At the same time, the absolute TD error of the state transition information of the current time slot is calculated. If it exceeds the threshold, it is sampled according to the weight of the state transition information, and then the sampled state transition information is used to train the agent.

[0021] As can be seen from the technical solution provided by the present invention, it has the ability to interact with the environment, can adapt to different UE distributions, and can optimize decision-making by leveraging the ability of intelligent agents to interact with the environment in deep reinforcement learning. It can achieve optimal flight trajectory planning for UAVs in unknown environments and has significant performance improvement compared with existing solutions in different scenarios. Attached Figure Description

[0022] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram of a multi-UAV trajectory planning method provided in an embodiment of the present invention;

[0024] Figure 2 This is a schematic diagram of the network architecture provided in an embodiment of the present invention;

[0025] Figure 3 A flowchart of the hybrid priority experience replay scheme provided in an embodiment of the present invention;

[0026] Figure 4 This is a TL-RLMTD deployment architecture diagram provided in an embodiment of the present invention;

[0027] Figure 5 This is a schematic diagram of the experimental results provided in the embodiments of the present invention;

[0028] Figure 6This is a schematic diagram of a multi-UAV trajectory planning system provided in an embodiment of the present invention. Detailed Implementation

[0029] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0030] First, the following explanations are provided for the terms that may be used in this article:

[0031] The terms “including,” “comprising,” “containing,” “having,” or other similar semantic descriptions should be interpreted as non-exclusive inclusion. For example, “including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction conditions, processing conditions, parameter, algorithm, signal, data, product or article of manufacture, etc.)” should be interpreted as including not only the expressly listed technical feature element, but also other technical feature elements that are not expressly listed and are well-known in the art.

[0032] The term "composed of" excludes any technical features not expressly listed. When used in a claim, it closes the claim to exclude all technical features other than those expressly listed, except for associated conventional impurities. If the term appears only in a clause of a claim, it limits the claim to the elements expressly listed in that clause; elements recited in other clauses are not excluded from the overall claim.

[0033] The following is a detailed description of a multi-UAV trajectory planning method and system provided by the present invention. Contents not described in detail in the embodiments of the present invention are prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of the present invention, they should be performed according to conventional conditions in the art or conditions recommended by the manufacturer.

[0034] Example 1

[0035] This invention provides a multi-UAV trajectory planning method, such as... Figure 1 As shown, it mainly includes:

[0036] 1. Construct a three-layer network architecture and deploy intelligent agents.

[0037] In this embodiment of the invention, a three-layer network architecture consisting of multiple ground user devices, a group of drones, and a high-altitude platform is constructed, and an intelligent agent is deployed separately in each drone.

[0038] 2. Pre-training.

[0039] To accelerate the agent's exploration process and perceive environmental changes caused by user device movement in real time, a transfer learning mechanism is introduced. This mechanism allows for instant updates to the agent, ensuring it can adapt to the constantly changing environment. Inspired by transfer learning, the agent is pre-trained in the source domain environment and then transferred to the target domain environment. The distribution of user devices is the same in both the source and target domain environments, and the impact of user device movement is ignored during pre-training (i.e., all user devices remain stationary).

[0040] 3. Drone trajectory planning and training.

[0041] The main processes at this stage are as follows:

[0042] (1) In the target domain environment, each agent in the UAV constructs its own local state of the current time slot by combining the global observation state of the current time slot provided by the high-altitude platform, and generates the action of the current time slot according to the action strategy. The action refers to the flight action of the UAV.

[0043] In this embodiment of the invention, for drone B n In the current time slot t, UAV B n The intelligent agent in the middle will include its own position c n Observational information of (t) n (t) is uploaded to the high-altitude platform; the high-altitude platform obtains the positions of all UAVs from the observation information of all UAVs, and combines it with the average access rate of user equipment to form a global observation state s(t); UAV B n The agent in the data combines the global observation state s(t) with the energy consumption e of the previous time slot. n (t-1), construct its own local state s for the current time slot t. n (t)={e n (t-1),s(t)}.

[0044] In this embodiment of the invention, the flight actions of the UAV include: horizontal flight actions, vertical flight actions, and hovering; the horizontal flight actions include the horizontal flight direction and the flight distance in the horizontal direction, and the vertical flight actions include the vertical flight direction (ascending or descending) and the displacement in the vertical direction.

[0045] In this embodiment of the invention, the process of generating flight actions for a drone is the drone trajectory planning process.

[0046] (2) Select the drone to serve based on the utility function of the drone after performing the corresponding action, combined with the location of each user device.

[0047] In this embodiment of the invention, for user equipment u m Covering the user equipment u m The set of drones is represented as: in, Let N represent a drone. m To cover the user equipment u m The number of drones;

[0048] Calculate the utility function of the UAV in the current time slot t. Represented as:

[0049]

[0050] Where f1 and f2 are weight parameters, For UAV B in the current time slot t n The channel gain of M′ n′ (t) represents drone B n The number of user devices covered;

[0051] For user equipment u m The drone with the highest utility function is selected as the service drone, expressed as:

[0052]

[0053] Among them, B′ m For user equipment u m The service drones;

[0054] After the drone performs the corresponding action, for the user equipment u m Update and cover the user device u m Given a set of drones, recalculate the utility function for each drone. If drone B has the highest utility function at this point... m ≠B′ m If B... m =B′ m Then maintain contact with drone B′ m The connection.

[0055] (3) Calculate the reward for each drone in the current time slot by combining the number of user devices for all drone services in the current time slot, the fairness index, the number of user devices and energy consumption for each drone service, and the legality of the action; whereby the fairness index is calculated by the access rate of user devices.

[0056] In this embodiment of the invention, for the current time slot t, UAV B n The reward calculation method is as follows:

[0057]

[0058] Where J(t) is the fairness index of the current time slot t. The number of user devices serving all drones in current time slot t, and the number of user devices u in current time slot t. m For drones that access the service, then c m (t) = 1, otherwise c m (t) = 0, M is the number of user equipment; e n (t) represents the current time slot t for UAV B. n Energy consumption, K n (t) represents the current time slot t for UAV B. n The number of user devices serving the service, pλ is the penalty term, λ is the set value, and p is the characteristic function used to identify whether the action is legal.

[0059] The formula for calculating the fairness index is:

[0060]

[0061] Where J(t) is the fairness index of the current time slot t. A m (t) represents user equipment u m The access rate from time slot 1 to the current time slot t, and the user equipment u in time slot t' m For drones that access the service, then c m (t') = 1, otherwise c m (t') = 0, where M is the number of user equipment.

[0062] (4) For each UAV pre-trained agent, the state transition information of the current time slot is constructed by using the local state, action, reward of the current time slot and the local state of the next time slot. After calculating the weight of the state transition information, the state transition information and the corresponding weight are stored in the playback buffer. At the same time, the absolute TD error of the state transition information of the current time slot is calculated. If it exceeds the threshold, it is sampled according to the weight of the state transition information, and then the sampled state transition information is used to train the agent.

[0063] In this embodiment of the invention, the weight of the state transition information is a hybrid priority weight, and the calculation scheme is as follows:

[0064] Introducing p1(t) to represent the empirical priority weight of the state transition information e(t), expressed as:

[0065] p1(t) = 1 / rank(t)

[0066] Where rank(t) represents the order of samples in the replay buffer when sorted according to the absolute TD error, and e(t) is the state transition information of the current time slot t;

[0067] We introduce p2(t) to represent the time-sensitive weight of the state transition information e(t), denoted as:

[0068]

[0069] Where e is the natural constant, t sample and t generate These represent the time slots for the generation of state transition information and the previous sampling, respectively. When generating state transition information, t is initialized. sample =t generate .

[0070] The hybrid priority weight p(t) for calculating the state transition information e(t) based on p1(t) and p2(t) is expressed as:

[0071] p(t) = γ1 × p1(t) + γ2 × p2(t)

[0072] Among them, constants γ1 and γ2 are used to control the sampling strategy’s bias towards empirical priority and time sensitivity.

[0073] Finally, the state transition information e(t) and its mixed priority weight p(t) are stored together in the replay buffer.

[0074] In this embodiment of the invention, the intelligent agent includes a deep Q-network and a target network. The deep Q-network generates actions for each time slot according to the action strategy. The state transition information is sampled according to its weights, and the sampled state transition information is input into the deep Q-network and the target network respectively. The TD error is calculated by combining the outputs of the two networks. The calculated TD error is used to update the mixed priority weights of the corresponding state transition information. At the same time, the calculated TD error is used to update the parameters of the deep Q-network. The parameters of the target network are periodically updated using the parameters of the deep Q-network.

[0075] In this embodiment of the invention, the pre-training process of the agent in the source domain environment is similar to the training process in the target domain environment. The main difference is that at the beginning of pre-training, both the deep Q network and the target network use initialized parameters, while during training, the deep Q network uses the parameters obtained from pre-training, and the target network uses the parameters updated by the deep Q network.

[0076] The above (1) to (4) are repeated until the termination condition is met (e.g., the set number of training sessions is reached); thereafter, the agent in each UAV generates flight actions in the manner described in (1) above, thereby realizing autonomous flight trajectory planning of the UAV.

[0077] To more clearly demonstrate the technical solution and its effects provided by the present invention, the method provided by the embodiments of the present invention will be described in detail below with reference to specific examples.

[0078] I. Introduction to basic principles.

[0079] In this embodiment of the invention, the multi-UAV 3D trajectory planning problem is formulated as a multi-objective optimization problem to simultaneously improve network coverage, access fairness, data transmission volume, and energy efficiency. Inspired by the decision-making capabilities of DRL (Deep Transfer Reinforcement Learning), this invention proposes a multi-UAV trajectory planning (TL-RLMTD) scheme based on deep transfer learning. TL-RLMTD, while possessing the ability to explore UE clusters and find optimal altitudes, can reduce UAV energy consumption and improve UE access rate, fairness, and the overall system data transmission volume.

[0080] This section mainly introduces the basic principles; the next section will introduce specific drone trajectory planning schemes.

[0081] 1. Construct a network model.

[0082] This invention constructs a three-layer network, the overall architecture of which is as follows: Figure 2 As shown. Consists of multiple ground users (user equipment). A group of drones It consists of a high-altitude platform (HAP) composed of hot air balloons; where M represents the number of user equipment and N represents the number of drones. The drones fly at different altitudes and act as aerial mobile base stations to provide network services to the user equipment on the ground. Due to the limited communication range of drones, this invention uses a high-altitude platform to collect access information from all user equipment and synchronize local observation status information among the drones.

[0083] 2. Time slot model.

[0084] This invention divides each time slot into a flight state and a hovering state. At the start of each time slot, the UAV is in the flight state, the duration of which depends on the flight maneuvers performed by the UAV. In the current time slot t, UAV B... n In its flight state, it will first perform its flight maneuvers, including along direction θ in the horizontal direction. n (t)∈[0,2π) fly a specified distance l h,n (t), flying a specified distance l in the vertical direction v,n(t), or hovering at the current position (in which case the flight state lasts for 0 seconds). Each drone has the same horizontal velocity (denoted as v). h and the same vertical velocity (denoted as v) p Then drone B n The flight time in the current time slot t can be expressed as:

[0085]

[0086] Each time slot has a fixed duration T. s And by drone B n Flight time With drone B n Hovering time The composition is as follows:

[0087]

[0088] 3. Energy consumption model.

[0089] Assume that the energy consumption of each UAV in the current time slot t consists of hovering energy consumption, flight energy consumption, and communication-related energy consumption. For UAV B... n The energy consumption for hovering is Flight energy consumption is Energy consumption related to communication is Its energy consumption e n (t) is represented as:

[0090]

[0091] For drone B n The hovering time is ε1, and the hovering power is ε1.

[0092] Flight energy consumption of drones Composed of horizontal and vertical components, it can be calculated using the following formula:

[0093]

[0094] Where ε2 is the horizontal flight propulsion power, l h,n (t) represents the horizontal flight distance, v h For horizontal flight speed, l v,n (t) represents the vertical displacement, and ε3 is the energy consumption parameter during vertical flight.

[0095] This invention assumes that, at the same altitude, the energy consumption of horizontal flight is slightly higher than that of hovering.

[0096] Assume the communication-related power is a fixed value ε. c When drones hover as base stations, they also need to communicate. Therefore, we have:

[0097]

[0098] 4. Data transmission model.

[0099] Considering that when a user device is simultaneously covered by multiple drones, it will not access all of them at the same time, the current timeslot t is defined to be connected to drone B. n The user equipment set is And define K n (t) is The number of user devices in the network. For A2G (Air-to-Ground) links, the number of drones B... n and user equipment The channel between them consists of line-of-sight (LoS) components and non-line-of-sight (NLoS) components. According to existing research, the path loss can be expressed as follows:

[0100]

[0101] in,

[0102]

[0103] f represents the distance between the drone and the user equipment. c Let c represent the transmission frequency and the speed of light, respectively. n (t),y n (t),z n (t) represents the drone B n At time slot t, (x k (t),y k (t),z k (t) represents user equipment u k At time slot t, the three terms in the position are the x, y, and z coordinates, respectively. P LoS (t) is the Loss probability of the A2G connection, which can be expressed by the following formula:

[0104]

[0105] Among them, (a,b,η) LoS ,η NLoS ) is a constant related to the urban environment.

[0106] Using the OFDM (Orthogonal Frequency Division Multiplexing) scheme, the bandwidth between the U drone and the U user equipment it serves is divided into K... n There are (t) subchannels, each occupied by one user equipment. Assuming the A2A and A2G links share the same link bandwidth B, therefore, within time slot t, user equipment u... k The amount of data transmitted can be expressed as:

[0107]

[0108] Among them, E k σ represents the transmission power of the UE. 2 This represents the white Gaussian noise power.

[0109] To reasonably simulate the data transmission needs of user equipment, using Indicates user equipment u k The transmission requirement in time slot t. Therefore, UAV B n The amount of data transmitted in time slot t is:

[0110]

[0111] The backhaul connection between the UAV and HAP mainly consists of a Loss of Speed ​​(LoS) component, which can be represented using a free-space path loss model as follows:

[0112]

[0113] in

[0114]

[0115] Indicates drone B n The distance between the HAP and the backhaul connection. Therefore, the backhaul capability can be expressed as...

[0116]

[0117] Among them, E U For drone B n The transmission power.

[0118] A key optimization objective of this invention is to improve the data performance of all ground-based user equipment within the system and increase the data transmission volume of each user equipment in each time slot. A data transmission model is used to establish the optimization objective. The TL-RLMTD algorithm achieves this objective by controlling the UAV's interaction with the environment to obtain the optimal flight maneuvers, thereby increasing the number of user equipment served by the UAV in each time slot and extending the time the UAV acts as an aerial base station providing communication services, ultimately achieving the optimization effect.

[0119] 5. Mathematical modeling.

[0120] A fairness metric based on user equipment access rate is introduced to measure system fairness. In the current time slot t, if user equipment u... m If it is covered by at least one drone, and it makes an access choice, then c m (t) = 1, otherwise, let c m (t) = 0. Furthermore, to analyze the access status of each user equipment over a longer period, a metric called access rate is introduced:

[0121]

[0122] Therefore, in time slot T, the system's average access rate is:

[0123]

[0124] The widely accepted Jain fairness index is used to measure system fairness. Therefore, the cumulative fairness index of the system at time slot T can be expressed as:

[0125]

[0126] Clearly, the fairness index J(T) of time slot T ∈ [0,1] and J(T) is positively correlated with the fairness of the system.

[0127] To optimize the long-term performance of drones, the trajectory planning problem is formalized into a multi-objective optimization problem:

[0128]

[0129]

[0130]

[0131]

[0132]

[0133] Constraint c1 ensures that each drone has sufficient energy to perform the mission, where E represents the maximum battery capacity of each drone; constraint c2 ensures that the drone's flight time does not exceed the time slot length; constraint c3 ensures that the drone's flight range does not exceed a preset range, i.e., a length of L. x The width is L y The rectangular region. Constraint c4 is used to ensure that the drones maintain a minimum distance to avoid collisions, where Indicates drone B n and B n′The distance between them, R s It is a safety radius constant.

[0134] This multi-objective optimization problem is difficult to solve using classical methods. Therefore, this invention proposes the TL-RLMTD scheme, which solves the problem by training a neural network.

[0135] II. Multi-UAV trajectory planning scheme.

[0136] 1. Dynamic access switching solution.

[0137] This invention first proposes a dynamic access handover scheme to handle the frequent access handover process between UAVs and user equipment, and based on this dynamic access handover scheme, proposes a multi-UAV trajectory planning scheme based on deep transfer learning.

[0138] Under time slot t, consider covering user equipment u m A group of drones Define utility function Used to measure u m and each connected drone The effect between them. This function is determined by the channel gain. and the drone B′ n Number of user devices covered (M) n′ (t) determines this, and therefore can be expressed as:

[0139]

[0140] Here, f1 and f2 are weight parameters, and f1 and f2 ∈ (0, 1).

[0141] Next, user equipment u m Select the drone B with the largest utility function value. m To access, i.e.

[0142]

[0143] The dynamic access handover algorithm between drones and the user equipment they cover is shown in Table 1. In time slot t, drone B... n After completing its flight maneuver, it hovers as a base station and notifies the UEs it covers. Access. For each user device. First, record its current service drone B′ m B n Join the accessible drone ensemble And according to Calculate the new optimal drone B. m If B m ≠B′ mIf so, an access switch will be performed on drone B'. m It will notify the user's device to disconnect and switch to B. m When an access switch occurs, B m and B′ m They will update their collection of connected drones. Conversely, if B m =B′ m User equipment u m It will remain with drone B' m The above dynamic access handover scheme comprehensively considers channel conditions and UAV load, using a utility function to select the optimal UAV, enabling user equipment to choose a UAV with good channel conditions and less coverage for its users. Therefore, compared with existing solutions, the dynamic access handover scheme delivers better performance.

[0144] Table 1: Dynamic Access Switching Scheme Flow

[0145]

[0146] 2. A multi-UAV trajectory planning (TL-RLMTD) scheme based on deep transfer learning.

[0147] This invention presents a non-uniform experience replay algorithm to select the latest and most valuable experience data, thereby accelerating convergence. It also proposes the TL-RLMTD scheme to achieve distributed decision-making in multi-UAV systems, while simultaneously implementing a dynamic training mechanism. The design of the scheme's state space, action space, and reward function will be introduced first, followed by the key technology employed: hybrid priority experience replay. Finally, the implementation process of the scheme will be described.

[0148] 1) State space, action space, and reward function.

[0149] 1.1) Observation Space and state space

[0150] In time slot t, for drone B n Its observation o n (t) is determined by the position c of the drone n (t), energy consumption in the previous time slot e n (t-1) and the number of connected user devices K n (t) constitutes, thus o n (t)={c n (t),e n (t-1),K n (t)}, observation space Defined as:

[0151]

[0152] The global observation state s(t) consists of the locations of all UAVs (unmanned aerial vehicles) and the system's average UE access rate. in The state of each UAV consists of the global state received from the HAP and its own local observations, defined as s. n (t)={e n Therefore, the state space is defined as: (t), s(t)}.

[0153]

[0154] 1.2) Action Space

[0155] In time slot t, drone B n action a n (t) by horizontal flight motion f n (t)=(θ n (t),l n (t)) and vertical flight maneuvers h n (t) is composed of θ n (t) and l n (t) represents the direction and distance of horizontal flight, respectively.

[0156] 1.3) Reward function.

[0157] To comprehensively measure the benefits of each drone after an operation, B n The reward in time slot t is defined as

[0158]

[0159] Where J(t) is the fairness index, e represents the total number of user devices served by all drones in time slot t. n (t) is B n Energy consumption. The third term in the above formula is the penalty for illegal actions, which can be divided into three categories: flying out of the target area, exceeding the predetermined altitude. And causing a drone collision, once an illegal action occurs, the drone will not perform that action, and a fixed value λ will be subtracted from the reward as a penalty, where p is a characteristic function defined as

[0160]

[0161] In the reward function, e n (t) and K nJ(t) corresponds to the benefits of the drone itself, namely, reducing energy consumption and maximizing resource utilization, while J(t) and Corresponding to the overall global benefit of the system, namely maximizing the coverage and fairness of user devices, the existence of the penalty term ensures that the actions made by the agent do not violate the constraints of the multi-objective optimization problem proposed earlier. The TL-RLMTD scheme is based on reinforcement learning. During the training process, the agent maximizes the cumulative reward. Therefore, the actions given by the model will maximize the optimization of the multi-objective optimization problem without violating the constraints, in order to solve the optimization problem proposed earlier. More specifically, data transmission volume is a system evaluation indicator. Directly using it as the optimal reward for deep reinforcement learning is a coarse approach. Therefore, in this embodiment of the invention, by optimizing the benefits of the UAV itself and the global benefits of the system, the service time of the UAV and the service rate of all user devices are improved, indirectly achieving the effect of maximizing data transmission volume.

[0162] 2) Hybrid Priority Experience Replay Algorithm (HPER).

[0163] In this embodiment of the invention, a hybrid priority experience replay scheme is proposed to avoid selecting outdated data through uniform sampling in deep reinforcement learning. Figure 3 Both Table 2 and Table 2 show the process of the hybrid priority experience playback scheme.

[0164] Specifically: for a state transition e(t) =<s(t),a(t),r(t),s(t+1)> The absolute TD error |δ(t)| reflects the difference between the state-action value function in the real environment and the DQN estimate, and its expression is:

[0165] δ(t)=r(t)+γQ(s(t),a * *;w - )-Q(s(t-1),a(t-1);w)

[0166] In the formula, γ is the reward discount factor, r(t) is the reward, and Q(s(t), a**;w - ) indicates that the parameter w in the agent is - The target network outputs the state s(t) and action a * The corresponding Q-values, Q(s(t-1), a(t-1); w), represent the Q-values ​​of the state s(t-1) and action a(t-1) output by the DQN network with parameter w in the agent. The goal of DQN is to estimate the state-action value function of the real environment; therefore, samples with larger TD errors are more valuable for updating the value function and should be selected with higher priority. For this purpose, p1(t) is introduced to represent the empirical priority weight for state transitions, defined as:

[0167] p1(t) = 1 / rank(t)

[0168] Here, rank(t) represents the order of state transitions when the replay buffer is sorted according to |δ(t)|.

[0169] Furthermore, to avoid outdated state transitions, p2(t) is introduced to represent the time-sensitive weight of the state transition, which is defined as:

[0170]

[0171] Among them, t sample and t generate These represent the time slots where the state transition occurs and where the data is sampled, respectively.

[0172] Finally, the hybrid priority weight is defined as:

[0173] p(t) = γ1 × p1(t) + γ2 × p2(t)

[0174] The constants γ1 and γ2 are used to control the sampling strategy's bias towards empirical priority or time sensitivity.

[0175] Let γ1, γ2 ∈ [0, 1] and γ1 + γ2 = 1. Next, classify based on p1(t) and p2(t) and store {e(t); p(t)} in a specified region in B. When updating the DQN network, perform non-uniform sampling based on the mixed weights p(t). Use the sampled state transitions for gradient descent to optimize the DQN network parameters while updating the mixed weights of the corresponding state transitions. As for the target network, it periodically replicates the parameters of the DQN network. In summary, the mixed priority experience replay scheme increases the sampling probability of updating high-value state transition experiences while ensuring the freshness of the sampled data.

[0176] Table 2: Hybrid Priority Experience Replay Scheme Flow

[0177]

[0178]

[0179] 3. A multi-UAV trajectory planning scheme based on deep transfer learning.

[0180] In the TL-RLMTD architecture, each UAV makes decisions based on its own observations and globally shared information provided by HAP, and the agent updates using Scheme 2. To accelerate the agent's exploration process and perceive environmental changes caused by user device movement in real time, a transfer learning mechanism is introduced within the TL-RLMTD framework. This mechanism allows for instant updates to the agent, ensuring that it can adapt to the constantly changing environment. Inspired by transfer learning, the agent is first pre-trained in the source domain environment, where the distribution of user devices is the same as in the target domain environment, but the impact of user device movement is eliminated.

[0181] This invention employs an ∈-greedy strategy as the action strategy π for each agent. w The parameter ∈ is used to balance the trade-off between exploration and exploitation. During the pre-training phase in the source domain environment, the agent acquires prior knowledge about the environment, such as optimal altitude and illegal actions. This prior knowledge helps the agent converge faster in the target domain. When transferring the pre-trained agent model to the target domain, an error threshold δ is introduced. max This allows for online training to be triggered when a significant TD error occurs, thereby motivating the agent to explore in a dynamic environment. Table 3 illustrates the specific process of this scheme.

[0182] Table 3: Flowchart of Multi-UAV Trajectory Planning Scheme Based on Deep Transfer Learning

[0183]

[0184]

[0185] Specifically: (1) Pre-training in the source domain environment. Randomly initialize the DQN network parameters w for each unmanned aerial vehicle. n and target network parameters And the replay buffer Set to an empty set. Furthermore, initialize ∈ to 1 to ensure the model has sufficient exploratory power. The source domain pre-training phase consists of S training epochs, each lasting T time slots. In each time slot, drone B... n First, observe the observed value o n The drone's agent receives the global information s(t) integrated by the HAP and then uploads it to the HAP. The drone's agent receives the global information s(t) integrated by the HAP and constructs its own local state s. n (t). Based on behavioral strategies The drone's intelligent system can generate an action a n (t). Subsequently, evaluate action a. n The drone's agent verifies the legality of (t) and invokes Scheme 1 to complete the interaction with the user device and execute the legal action. Upon completion, the drone's agent determines the reward r. n(t) and the state s of the next time slot n (t+1). Finally, the UAV stores the generated transition in the replay buffer and calls scheme 2 for model update. (2) Training in the target domain environment. The UAV inherits the model obtained from the source domain pre-training. Unlike the source domain, the target domain allows the UE to have mobility. The target domain training phase includes T time slots, and the ∈ of each UAV is initialized to 0. The UAV's agent is based on the state s n (t) Generate action a n (t). Similarly, the drone verifies the legality of the action and executes it, then gains a state transition experience. If the absolute TD error of the real-time transition exceeds δ... max If the state transition is 0.3, it is considered an environmental change. In this case, the agent increases ∈ (examples of ∈ ← 0.3 are provided in Table 3) to enhance its ability to explore the environment, and samples the newly generated state transitions according to Scheme 2 to update the DQN network and the target network. Its deployment architecture is as follows: Figure 4 As shown.

[0186] III. Performance Description.

[0187] To illustrate the performance of the deep transfer learning-based multi-UAV trajectory planning scheme provided by this invention, the scheme was deployed and its performance in handling user equipment migration was recorded, such as... Figure 5 As shown. Figure 5 Part (a) shows the user's initial location before migration and the starting location of the drone swarm; Figure 5 Part (b) shows the movement trajectory of the drone swarm exploring the user's location as the user begins to move; Figure 5 Part (c) shows the process of the drone swarm gradually stabilizing as the user device migration ends, indicating that the drone swarm has basically found the user device's new location. Figure 5 Part (d) explains that the drone has completely located the user's relocated position and stabilized.

[0188] Through the above description of the embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software, or by using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.), including several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0189] Example 2

[0190] This invention also provides a multi-UAV trajectory planning system, which is mainly used to implement the methods provided in the foregoing embodiments, such as... Figure 6 As shown, the system mainly includes:

[0191] The network architecture building unit is used to build a three-layer network architecture consisting of multiple ground user devices, a group of drones, and a high-altitude platform, and to deploy an intelligent agent in each drone.

[0192] The pre-training unit is used to pre-train the agent in the source domain environment; the pre-training is transferred to the target domain environment, and the distribution of user devices is the same in the source domain environment and the target domain environment. The influence of user device movement is ignored during pre-training.

[0193] The drone trajectory planning and agent training unit is used for:

[0194] In the target domain environment, each agent in the UAV constructs its own local state of the current time slot by combining the global observation state of the current time slot provided by the high-altitude platform, and generates the action of the current time slot according to the action strategy. The action refers to the flight action of the UAV.

[0195] Based on the utility function of the drone after performing the corresponding action, select the drone to serve each user device;

[0196] The reward for each drone in the current time slot is calculated by combining the number of user devices for all drone services in the current time slot, the fairness index, the number of user devices and energy consumption for each drone service, and the legality of the action; the fairness index is calculated based on the access rate of user devices.

[0197] For each UAV's pre-trained agent, the state transition information of the current time slot is constructed using the local state, action, reward of the current time slot and the local state of the next time slot. After calculating the weight of the state transition information, the state transition information and the corresponding weight are stored in the replay buffer. At the same time, the absolute TD error of the state transition information of the current time slot is calculated. If it exceeds the threshold, it is sampled according to the weight of the state transition information, and then the sampled state transition information is used to train the agent.

[0198] Since the specific processing and calculation details involved in this system have been described in detail in the previous method embodiments, they will not be repeated here.

[0199] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.

[0200] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A multi-UAV trajectory planning method, characterized in that, include: A three-layer network architecture consisting of multiple ground user devices, a group of UAVs, and an aerial platform is constructed, and an intelligent agent is deployed in each UAV. The intelligent agent is pre-trained in the source domain environment and then migrates to the target domain environment. The distribution of user devices in the source domain environment and the target domain environment is the same, and the influence of user device movement is ignored during pre-training. In the target domain environment, each agent in the UAV constructs its own local state of the current time slot by combining the global observation state of the current time slot provided by the high-altitude platform, and generates the action of the current time slot according to the action strategy. The action refers to the flight action of the UAV. Based on the utility function of the drone after performing the corresponding action, the drone for service is selected in combination with the location of each user device; The reward for each drone in the current time slot is calculated by combining the number of user devices for all drone services in the current time slot, the fairness index, the number of user devices and energy consumption for each drone service, and the legality of the action; the fairness index is calculated based on the access rate of user devices. For each UAV's pre-trained agent, the state transition information of the current time slot is constructed using the local state, action, reward of the current time slot and the local state of the next time slot. After calculating the weight of the state transition information, the state transition information and the corresponding weight are stored in the replay buffer. At the same time, the absolute TD error of the state transition information of the current time slot is calculated. If it exceeds the threshold, it is sampled according to the weight of the state transition information, and then the sampled state transition information is used to train the agent. Multi-UAV trajectory planning is formulated as a multi-objective optimization problem, which is expressed as follows: ; in, For the current time slot drones Energy consumption, For time slots Fairness index, For time slots average access rate For the set of all drones, All represent a single drone, and N is the number of drones; Indicates drone In the time slot Data transmission volume, For drones In the time slot Hovering time, For drones In the time slot Flight time, The duration of each time slot; For drones In the time slot The position, where the three terms are the x, y, and z coordinates respectively; constraint c1 ensures that the UAV has sufficient energy to perform the mission. This represents the maximum battery capacity of the drone; constraint c2 ensures that the drone's flight time does not exceed the time slot length; constraint c3 ensures that the drone's flight range does not exceed the preset range, i.e., the length is... , width is The rectangular region; constraint c4 is used to ensure that the drones maintain a minimum distance to avoid collisions. Indicates drone and The distance between them It is a safety radius constant.

2. The multi-UAV trajectory planning method according to claim 1, characterized in that, The intelligent agent in each UAV constructs its own local state for the current time slot by combining the global observation state of the current time slot provided by the high-altitude platform, including: For drones In the current time slot drones The intelligent agent in the middle will include its own location. Observational information The data is uploaded to a high-altitude platform; the high-altitude platform obtains the positions of all UAVs from the observation information of all UAVs, and combines this with the average access rate of user equipment to form a global observation status. drones The intelligent agent in the system combines global observation state and the previous time slot Energy consumption Construct its own current time slot Local state .

3. The multi-UAV trajectory planning method according to claim 1, characterized in that, The step of selecting a service drone based on the drone's utility function after performing the corresponding action, combined with the location of each user device, includes: For user equipment Covering the user equipment The set of drones is represented as: ,in, This refers to a drone. To cover the user equipment The number of drones; Calculate the current time slot Utility function of drones , is represented as: ; in, For weight parameters, For the current time slot Unloading drones Channel gain, For drones The number of user devices covered; For user equipment The drone with the highest utility function is selected as the service drone, expressed as: ; in, For user equipment The service drones; After the drone performs the corresponding action, for the user equipment Update and cover the user device Given a set of drones, recalculate the utility function for each drone. If the drone with the highest utility function is found... If so, then an access switch will be performed. Then maintain contact with drones The connection.

4. The multi-UAV trajectory planning method according to claim 1, characterized in that, The calculation of the reward for each drone in the current time slot includes: For the current time slot drones The reward calculation method is as follows: ; in, For the current time slot Fairness index, Current time slot Number of user devices for all drone services, current time slot User equipment Drones that access the service, ,otherwise , Number of user devices; For the current time slot drones Energy consumption, For the current time slot drones The number of user devices served. As a penalty item, For setting value, It is a feature function used to identify whether an action is legal; and The benefits of drones themselves are reduced energy consumption and maximized resource utilization. and Corresponding to the global benefit, which is to maximize the coverage of user devices and the fairness of coverage, the existence of the penalty term ensures that the action made by the agent will not violate the constraints of the multi-objective optimization problem.

5. The multi-UAV trajectory planning method according to claim 1, characterized in that, Current time slot drones Energy consumption Energy consumption during hovering Flight energy consumption Energy consumption related to communication Composition, represented as: ; ; ; ; in, For drones Hovering time, This refers to hovering power; For horizontal flight propulsion power, The horizontal flight distance For horizontal flight speed, The vertical flight distance. and Energy consumption parameters during ascent and descent, respectively; This refers to communication-related power.

6. The multi-UAV trajectory planning method according to claim 1, characterized in that, The formula for calculating the fairness index is: ; in, For the current time slot Fairness index, , For user equipment From time slot 1 to the current time slot Access rate, time slot User equipment Drones that access the service, ,otherwise , This refers to the number of user devices.

7. The multi-UAV trajectory planning method according to claim 1, characterized in that, After calculating the weights of the state transition information, storing the state transition information and its corresponding weights together in the playback buffer includes: Introduction To represent state transition information The experience priority weight is expressed as: ; in, This indicates the order of samples in the playback buffer when sorted according to the absolute TD error. For the current time slot State transition information; Introduction To represent state transition information The time-sensitive weight is represented as: ; Where e is the natural constant, and These represent the time slots for the generation of state transition information and the previous sampling, respectively; when generating state transition information, initialization is performed. ; based on and Calculate state transition information Hybrid priority weights , is represented as: ; Among them, constants and Used to control the sampling strategy's bias towards experience priority and time sensitivity; Will Store in the playback buffer.

8. The multi-UAV trajectory planning method according to claim 7, characterized in that, The step of sampling based on the weights of the state transition information and then using the sampled state transition information to train the agent includes: The agent includes a deep Q-network and a target network, and the deep Q-network generates actions for each time slot according to the action strategy. Sampling is performed based on the weights of the state transition information. The sampled state transition information is then input into the deep Q-network and the target network, respectively. The TD error is calculated by combining the outputs of the two networks. The calculated TD error is used to update the hybrid priority weights of the corresponding state transition information; at the same time, the calculated TD error is used to update the parameters of the deep Q network. The parameters of the target network are periodically updated using the parameters of the deep Q-network.

9. A multi-UAV trajectory planning system, characterized in that, To implement the method according to any one of claims 1 to 8, comprising: The network architecture building unit is used to build a three-layer network architecture consisting of multiple ground user devices, a group of drones, and a high-altitude platform, and to deploy an intelligent agent in each drone. The pre-training unit is used to pre-train the agent in the source domain environment; the pre-training is transferred to the target domain environment, and the distribution of user devices is the same in the source domain environment and the target domain environment. The influence of user device movement is ignored during pre-training. The drone trajectory planning and agent training unit is used for: In the target domain environment, each agent in the UAV constructs its own local state of the current time slot by combining the global observation state of the current time slot provided by the high-altitude platform, and generates the action of the current time slot according to the action strategy. The action refers to the flight action of the UAV. Based on the utility function of the drone after performing the corresponding action, select the drone to serve each user device; The reward for each drone in the current time slot is calculated by combining the number of user devices for all drone services in the current time slot, the fairness index, the number of user devices and energy consumption for each drone service, and the legality of the action; the fairness index is calculated based on the access rate of user devices. For each UAV's pre-trained agent, the state transition information of the current time slot is constructed using the local state, action, reward of the current time slot and the local state of the next time slot. After calculating the weight of the state transition information, the state transition information and the corresponding weight are stored in the replay buffer. At the same time, the absolute TD error of the state transition information of the current time slot is calculated. If it exceeds the threshold, it is sampled according to the weight of the state transition information, and then the sampled state transition information is used to train the agent.

Citation Information

Patent Citations

  • Unmanned aerial vehicle auxiliary calculation migration method based on depth deterministic strategy gradient

    CN115640131A