Layered decision based energy consumption optimization method for communication and sensing of payload-constrained unmanned aerial vehicles
By employing a hierarchical decision-making framework, combined with adaptive large neighborhood search and deep reinforcement learning algorithms, the delivery grouping and flight trajectory of drones are optimized, addressing the challenges of load capacity, user satisfaction, and communication reliability in drone logistics delivery systems, and achieving energy-efficient and safe delivery.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- JIMEI UNIV
- Filing Date
- 2026-05-09
- Publication Date
- 2026-07-31
AI Technical Summary
In drone logistics delivery systems, how can we design an efficient collaborative framework to ensure flight safety and service quality while minimizing the flight energy consumption of the drone delivery system, given the complex real-world scenarios that consider load limits, user satisfaction, and communication reliability?
A hierarchical decision-making approach is adopted, which optimizes delivery grouping and access order at the upper level and flight trajectory at the lower level. By combining adaptive large neighborhood search algorithm and deep reinforcement learning algorithm, a heterogeneous hierarchical logistics decision-making framework is constructed to collaboratively optimize the delivery scheme of drones.
It effectively reduces the flight energy consumption of drone delivery systems, improves delivery service quality and flight safety, and meets the constraints of payload, user satisfaction and communication reliability.
Smart Images

Figure CN122155064B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of urban low-altitude logistics and distribution technology, and in particular to a method for optimizing the communication and sensing energy consumption of payload-constrained unmanned aerial vehicles (UAVs) based on hierarchical decision-making. Background Technology
[0002] With the rapid development of unmanned aerial vehicle (UAV) technology, it has shown broad application prospects in many fields such as logistics transportation, auxiliary communication, agricultural and forestry plant protection, and information collection. Especially against the backdrop of the rapid development of the "low-altitude economy," low-altitude airspace is gradually evolving into a key spatial carrier driving the innovation of logistics models. Due to the advantages of UAVs, such as flexibility, efficiency, and lack of terrain restrictions, they have shown great potential in "last-mile" delivery and have become a key technological means for low-altitude logistics delivery. However, the actual deployment of UAV low-altitude logistics delivery systems still faces multiple challenges. First, due to the structure and power system of UAVs, there is a clear upper limit to the maximum payload of a single UAV. When serving multiple users, if the total weight of the goods exceeds the maximum payload, users must be divided into multiple delivery batches, and UAVs need to make multiple round trips to perform tasks, which significantly increases the complexity of the delivery route planning problem. Second, users' requirements for the timeliness of delivery services are increasing, and optimizing energy consumption while meeting time window constraints is also a challenging combinatorial optimization problem. Finally, drones need to maintain stable communication with ground base stations (GBS) during flight to ensure real-time scheduling and flight safety. However, factors such as building obstruction and signal interference in urban environments can lead to decreased or even interrupted communication quality, affecting drone flight safety. (Here, "drone" refers to cellular-connected drones that communicate and are controlled via cellular networks (such as 4G and 5G).) Currently, existing research has explored solutions to these challenges from multiple perspectives. In terms of UAV path and energy consumption optimization, the existing literature: Wu K, Lu S, Chen H, et al. An energy-efficient logistic dronerouting method considering dynamic drone speed and payload. Sustainability, 16 (12), 4995 [EB / OL]. (2024) proposes an adaptive large neighborhood search algorithm that optimizes flight speed as a decision variable, aiming to minimize total energy consumption and achieve coordinated energy-saving optimization of speed and path. However, it lacks consideration for communication connectivity during flight.
[0003] Existing literature, including Peng H, Cao J, Yang D, et al. Balancing Energy Efficiency and Communication Quality in UAV Cargo Delivery Systems[J]. IEEE Internet of Things Journal, 2025. and Cao J, Xiao L, Yang D, et al. Energy consumption and communication quality tradeoff for logistics UAVs: A hybrid deep reinforcement learning approach[C] / / 2023 IEEE Wireless Communications and Networking Conference (WCNC). IEEE, 2023: 1-6, both aim to solve the tradeoff optimization problem between energy consumption and communication quality in logistics UAVs during delivery tasks. They jointly optimize the delivery sequence and flight trajectory of UAVs by combining traditional optimization algorithms (such as the improved ant colony algorithm) with deep reinforcement learning. However, using communication quality as a tradeoff objective still poses potential risks in safety-critical scenarios.
[0004] Existing literature: Cicek CT, Koç Ç, Gultekin H, et al. Communication-awaredrone delivery problem[J]. IEEE Transactions on Intelligent Transportation Systems, 2024, 25(8): 9168-9180. By introducing switching and interruption constraints, and developing a mixed integer programming model and an efficient genetic algorithm, the shortest delivery path with the total flight distance of the UAV under communication constraints is found.
[0005] Existing literature: Huang G, Cao J, Yang L, et al. Joint Optimization of Energy Efficiency and Stable Communication Quality for Cargo UAV-Enabled Multi-Package Pickup and Delivery[J]. IEEE Transactions on Cognitive Communications and Networking, 2025. This paper proposes a hybrid genetic algorithm and deep reinforcement learning (HGADRL) framework. By optimizing the delivery sequence and flight trajectory, it ensures that the probability of communication interruption during UAV flight remains below a set threshold while minimizing system energy consumption, thereby improving UAV flight safety. However, it fails to fully consider the group delivery needs arising from payload limitations in practice.
[0006] Although some progress has been made in research on drone delivery, designing an efficient collaborative framework to minimize the flight energy consumption of drone logistics delivery systems while ensuring flight safety and service quality remains a challenge, especially in complex real-world logistics scenarios that simultaneously consider drone payload limitations, user satisfaction, and communication reliability. Summary of the Invention
[0007] In view of this, the purpose of this invention is to propose a hierarchical decision-making method for optimizing the communication and sensing energy consumption of payload-constrained UAVs. This method optimizes delivery grouping and access order at the upper level and optimizes flight trajectory based on communication constraints at the lower level. This significantly reduces the flight energy consumption of the UAV delivery system and improves the quality of delivery services while meeting the requirements of payload, user satisfaction and communication reliability.
[0008] To achieve the above-mentioned technical objectives, the technical solution adopted by this invention is as follows: This invention provides a method for optimizing the communication and sensing energy consumption of payload-constrained unmanned aerial vehicles (UAVs) based on hierarchical decision-making, comprising the following steps: Step 1: Build a drone group delivery system; Step 2: Construct a communication model, a user satisfaction model, a group delivery sequence model, and an energy consumption model in the UAV group delivery system. Based on the communication model, user satisfaction model, group delivery sequence model, and energy consumption model, establish an optimization objective to minimize flight energy consumption. Step 3: In the upper-level decision-making, the problem of optimizing delivery grouping and access order is modeled as a vehicle routing problem with time windows. Based on the optimization objective, an adaptive large neighborhood search algorithm is used to solve the vehicle routing problem with time windows to obtain a global delivery sequence that satisfies the load constraint and the user satisfaction constraint. Step 4: In the lower-level decision-making, based on the global delivery sequence, the trajectory planning problem of each sub-segment is modeled as a Markov decision process, and a deep reinforcement learning algorithm is used to solve and train it to obtain a flight strategy network model. This flight strategy network model is used to output a flight direction decision that satisfies the communication interruption probability constraint according to the environmental state. Step 5: Deploy the global delivery sequence and the trained flight strategy network model to the UAV's onboard computing unit; during actual flight, the onboard computing unit loads and runs the flight strategy network model, which generates flight direction decisions in real time based on real-time perceived environmental information and executes the delivery task. Step 6: Through collaborative optimization between the upper and lower layers, output the optimal delivery plan that satisfies the constraints of load capacity, user satisfaction, and communication reliability, thereby minimizing the flight energy consumption of the drone.
[0009] Furthermore, step 1 specifically includes: Step 11: Based on the urban low-altitude logistics delivery needs, construct a cellular-connected drone group delivery system; the drone group delivery system defines the following elements: drones depart from the warehouse, deliver to each group of users and return to the warehouse, drones are subject to maximum load constraints, users have expected service time windows, and drones communicate with ground base stations through cellular networks. Step 12: Obtain delivery task data, communication network data, and drone performance parameters, including: The delivery task data includes the number of users, the location coordinates of each user, the required weight of goods, the user's expected service time window, and the warehouse location. The communication network data includes the number of ground base stations, the location coordinates of each base station, the transmission power, and the communication interruption threshold. ; The performance parameters of the UAV include the UAV's maximum payload, flight speed, flight altitude, UAV's own weight, and energy consumption model constant parameters.
[0010] Furthermore, step 2 specifically includes: Step 21: Establish a communication model, using the signal-to-interference-plus-noise ratio (SINR) as an indicator of communication quality and as a hard constraint for trajectory planning. Step 22: Establish a user satisfaction model, quantify user satisfaction based on the user's expected service time window, and use the average user satisfaction of all users as an indicator to measure user satisfaction. Step 23: Establish a group delivery sequence model. By grouping users and planning access routes for each group, optimization objects are provided for upper-level decision-making. Step 24: Based on the UAV's own weight, flight speed, and energy consumption model constant parameters, establish an energy consumption model, calculate the energy consumed by the UAV during flight, and use it as the core objective function for optimization. Step 25: Construct an optimization objective to minimize flight energy consumption based on the communication model, user satisfaction model, group delivery sequence model, and energy consumption model.
[0011] Furthermore, in step 3, an adaptive large neighborhood search algorithm is used to solve the vehicle routing problem with time windows based on the optimization objective, resulting in a global delivery sequence that satisfies both load constraints and user satisfaction constraints; specifically, this includes: Step 31: Based on the number of users and the location of the warehouse, the solution to the problem is represented by a multi-batch sequence encoding. The warehouse node is represented by 0. A solution is composed of multiple batch sequences with the warehouse node as the separator. The sequence between two adjacent 0s constitutes a delivery batch. The order within a delivery batch is the access order within the group, and the order of delivery batches is the execution order between groups. Step 32: Generate an initial solution using a construction method that incorporates heuristic rules; Step 33: Design various destruction and repair operators to perform destruction and repair operations on the current solution during the iteration process; Step 34: During the adaptive large neighborhood search iteration process, periodically perform a 2-opt local search on the user access sequence within each delivery batch. Select any two edges in the user access sequence that do not have a common user node. Each edge consists of two consecutive user nodes. Attempt to swap the end user nodes of the two edges to generate two new edges. If the swap can reduce the flight energy consumption of the delivery batch, the swap is accepted. If the swap cannot reduce the flight energy consumption of the delivery batch, the swap is not accepted. This process is repeated until no further improvement is possible or the preset number of iterations is reached. Step 35: Introduce the simulated annealing criterion to determine whether to accept the new solution. Let the current solution be... Current solution The objective function is The new interpretation is New interpretation The objective function is And the current temperature is The probability of the new solution being accepted is... for:
[0012] Step 36: Introduce an adaptive operator selection mechanism to dynamically adjust the selection probability of the destruction operator and the repair operator based on the historical performance of the operators; Step 37: Repeat steps 33 to 36. After the iteration is complete, output the currently found optimal solution that satisfies the load constraint and user satisfaction constraint as the global delivery sequence for use in the next layer trajectory planning.
[0013] Furthermore, step 32 specifically includes: Step 321: Calculate a comprehensive score based on the user's distance from the warehouse, the weight of the goods required, and the urgency of the user's expected service time window, and sort them in ascending order; Step 322: Assign users to the current delivery batch in the sorted order. If the total weight after loading exceeds the maximum payload of the drone, end the current delivery batch and start a new delivery batch. Step 323: For the user access sequence within each delivery batch, construct the nearest neighbor path based on the user's location coordinates, starting from each user, and select the user access sequence with the lowest flight energy consumption as the access order within the group for that delivery batch.
[0014] Furthermore, the destruction operators include the following five types: (1) Random destruction: randomly select q users to remove from the current solution; (2) Worst-case energy consumption destruction: Calculate the unit energy consumption contribution of each user and remove the q users with the largest unit energy consumption contribution; (3) Time window conflict disruption: Calculate the current satisfaction level of each user and remove the q users with the lowest satisfaction levels; (4) Batch destruction: Randomly select a complete delivery batch and remove all users within that delivery batch; (5) Neighbor destruction: Randomly select a reference user, calculate the spatial distance between other users and the reference user, and remove the q nearest users; The repair operators include the following four types: (1) Greedy insertion: Traverse all users to be inserted and all feasible insertion positions, and select the insertion method that minimizes the increase in flight energy consumption; (2) Time window priority insertion: Sort all users to be inserted in ascending order of their latest available service time, and then select the insertion position that maximizes their time window satisfaction for each user in turn; (3) Regret value insertion: For each user to be inserted, the difference between the energy consumption increment caused by the optimal insertion position and the second-best insertion position is calculated as the regret value, and the user with the largest regret value is inserted first. (4) Random insertion: Randomly select a user to be inserted and randomly select a feasible position for insertion; When performing an insertion operation, load constraints and time window constraints are checked simultaneously.
[0015] Furthermore, the adaptive operator selection mechanism in step 36 specifically includes: Step 361: Maintain a weight for each operator. The initial weights are set to be equal; Step 362: Before each iteration, select the destruction operator and the repair operator based on the weights using the roulette wheel method; Step 363: After the iteration is completed, the selected operators are scored based on the quality of the new solutions. The scoring principles are as follows: (1) If the new solution is better than the historical best solution, the score is ; (2) If the new solution is worse than the historical best solution but better than the current solution, the score is: ; (3) If the new solution is inferior to the current solution but is accepted, the score is: ; (4) If the new solution is inferior to the current solution and is rejected, the score is: ; in, The specific values are preset according to the application scenario; Step 364: Define each N iterations as a weight update cycle. At the end of each cycle, update the weights using the following formula based on the cumulative score and number of calls of the operator within that cycle:
[0016] in, Update the weighting factor. This represents the first operator before the update. Secondary weighting, Indicates the updated operator's first... Secondary weighting, This represents the cumulative score of the operator within the current period. This indicates the cumulative number of times the operator has been selected within the current period. This indicates that the operator has not been selected, and the cumulative score for the current period is... If it is 0, then The value is 0; Step 365: Update the weights The operator used for step 362 in the next iteration is repeated from step 362 to step 364 until the preset number of iterations is reached.
[0017] Furthermore, step 4 specifically includes: Step 41: After determining the global delivery sequence in the upper-level stage, the sub-problem is focused on the communication hard constraint trajectory planning problem in the delivery process between adjacent users, with the goal of minimizing the flight energy consumption of each sub-segment; Step 42: Based on the number of ground base stations, the location coordinates of each base station, the transmission power, and the communication interruption threshold, construct a low-altitude radio map of the city. Model the trajectory planning problem as a Markov decision process, defining the state, action, and reward as follows: (1) State s: The horizontal coordinate of the UAV at a certain moment; (2) Action a: Discretize the horizontal flight direction of the UAV into multiple uniformly distributed fixed directions, and each action corresponds to a fixed flight angle; (3) Reward r: Used to comprehensively reflect the satisfaction of energy consumption and communication constraints. The corresponding communication interruption probability is determined based on the current flight position and flight altitude of the UAV. When the probability of communication interruption of the UAV exceeds the communication interruption threshold, the reward r is awarded. When the area is [affected / injured], a penalty is imposed; Step 43: The deep reinforcement learning DQN algorithm is used to solve the Markov decision process. An action masking mechanism is introduced in the action selection stage to filter out invalid actions that cause the UAV to go out of the boundary or cause communication interruption. The deep reinforcement learning algorithm employs an experience replay mechanism to store and randomly sample historical experiences. By iteratively updating the policy network parameters through an approximation of the optimal action value network, the UAV learns a flight policy network model that satisfies communication constraints under a given state. This flight policy network model is used to minimize cumulative flight energy consumption under the premise of satisfying the communication interruption probability constraint, and outputs flight direction decisions based on the real-time environmental state. Step 44: Solve for all sub-segments in the global delivery sequence, execute the trained flight strategy network model for each sub-segment, record the flight energy consumption of each sub-segment and sum them up to obtain the flight energy consumption of the UAV to complete the entire delivery task.
[0018] Furthermore, the training process of the flight strategy network model is as follows: Step 431, Initialization: Initialize the network parameters of the action-value network Q. Initialize the target value network Network parameters Initialize the experience replay pool D, and set the starting point of the current sub-segment to the initial state. ; Step 432, Action Selection: In the current state The following is adopted - Greedy strategy selects actions based on probability. Randomly select an action with a probability of 1- choose At the same time, an action masking mechanism is introduced to filter out invalid actions that cause the drone to go beyond the boundary or cause communication interruption. Step 433, Perform Actions and Update the Environment: The UAV performs the selected actions. Transition to the next state Calculate the probability of communication disruption based on urban low-altitude radio maps and receive an instant reward. The reward Taking into account both energy consumption penalties and communication constraints, a penalty is imposed when the probability of communication interruption exceeds the communication interruption threshold; Step 434, Storing Experience: Store the experience samples ( Stored in the experience replay pool D; Step 435, Experience Replay and Network Update: Randomly sample a small batch of samples from the experience replay pool D, and calculate the target value network. The network parameters of the action value network Q are updated using gradient descent. ; Step 436: Update the target network: Every fixed number of steps, update the network parameters of the action value network Q. Replicating to the target value network, i.e. ; Step 437, Termination Judgment: Determine whether the sub-segment endpoint has been reached or the preset maximum number of steps has been reached. If not, then... If yes, return to step 432 to continue iterating; otherwise, end the training of the current sub-segment and output the flight strategy network model.
[0019] Furthermore, step 5 specifically includes: Step 51: Based on the environmental information, deploy the global delivery sequence and the trained flight strategy network model to the UAV's onboard computing unit; Step 52: The UAV's onboard computing unit loads the flight strategy network model and runs the flight strategy network model based on the perceived environmental information; Step 53: The flight strategy network model outputs flight direction decisions in real time and synchronizes them to the UAV flight control system; Step 54: The UAV flight control system controls the UAV to execute flight missions according to flight direction decisions and complete the delivery. The optimal delivery plan output in step 6 includes: delivery batch division, user access order within each group, and flight trajectory of each sub-segment.
[0020] By adopting the above technical solution, the present invention has the following beneficial effects compared with the prior art: (1) This invention proposes a heterogeneous hierarchical logistics delivery decision framework. The payload-constrained UAV group delivery problem involved in this invention needs to simultaneously satisfy payload constraints, user satisfaction constraints, and communication reliability constraints, and jointly optimize delivery grouping, access order, and flight trajectory. The decision variables are highly coupled, resulting in high problem complexity and difficulty in direct solution. This invention adopts a heterogeneous hierarchical logistics decision framework, decoupling the original joint optimization problem into two stages: upper-level task planning and lower-level trajectory planning, which are processed separately. The overall approximate optimum is achieved through the collaborative optimization of the upper and lower layers. The hierarchical strategy in this invention effectively reduces the solution complexity of the original problem and can minimize the system's flight energy consumption under the condition of simultaneously satisfying payload constraints, user satisfaction constraints, and communication reliability constraints. This solves the technical problem that existing methods cannot balance safety, service quality, and energy efficiency.
[0021] (2) In the optimization of upper-level grouping and delivery sequence, this invention combines the characteristics of UAVs having limited single-batch payload and needing to deliver in batches with the user's demand for time-based satisfaction. It models the delivery grouping and access order problem as a Capacitated Vehicle Routing Problem with Time Windows (CVRPTW). It adopts an adaptive large neighborhood search algorithm and designs a variety of targeted destruction operators (such as worst-case energy consumption destruction, time window conflict destruction, batch destruction, etc.) and repair operators (such as greedy insertion, time window priority insertion, regret value insertion, etc.). Combined with 2-opt local search and simulated annealing acceptance criteria, it effectively improves the search capability of the algorithm in the combinatorial optimization space, makes up for the lack of local finesse of ALNS global search, and further shortens the flight distance and reduces delivery energy consumption. In particular, this invention introduces an adaptive operator selection mechanism, which dynamically adjusts the probability of selection of operators based on their historical performance. This enables the algorithm to automatically focus on efficient neighborhood structures and solve the global delivery sequence with the lowest flight energy consumption under the premise of satisfying payload constraints and user satisfaction constraints.
[0022] (3) In the lower-level trajectory planning, this invention introduces communication reliability as a hard constraint into flight trajectory optimization and uses a deep reinforcement learning algorithm for solution. Most existing UAV path planning methods that consider communication constraints treat communication quality as a trade-off objective for energy consumption optimization (i.e., a soft constraint), which makes it difficult to guarantee absolute safety in critical scenarios where communication interruption may cause flight safety risks. To address this issue, this invention uses the probability of communication interruption as a hard constraint, constructs a low-altitude radio map of the city, and models the trajectory planning problem as a Markov decision process, using a deep reinforcement learning algorithm for solution. In the scenario of this invention, the communication constraint is a sparse reward (penalty is only received when entering a communication blind zone). The experience replay mechanism of the deep reinforcement learning algorithm stores historical experience and randomly samples it, effectively breaking the temporal correlation between data and improving the learning efficiency in sparse reward scenarios. Furthermore, this invention introduces an action masking mechanism in the action selection stage to filter invalid actions that may cause the UAV to exceed the boundary or cause communication interruption in real time, ensuring that each flight step meets the safety constraints from the decision-making level. Through the above design, the present invention can autonomously avoid communication blind spots in complex urban environments and plan the local flight trajectory with the lowest energy consumption while ensuring that the probability of communication interruption throughout the process is lower than the communication interruption threshold, which significantly improves the flight safety and mission reliability of UAVs in complex urban environments. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is an execution flowchart of a method for optimizing the communication and sensing energy consumption of a payload-constrained UAV based on hierarchical decision-making, provided in an embodiment of the present invention. Detailed Implementation
[0025] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be particularly noted that the following embodiments are for illustrative purposes only and do not limit the scope of the invention. Similarly, the following embodiments are only some, not all, embodiments of the present invention, and all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0026] Please see Figure 1 The present invention provides a method for optimizing the communication and sensing energy consumption of a payload-constrained unmanned aerial vehicle (UAV) based on hierarchical decision-making, comprising the following steps: Step 1, Delivery Scenario Initialization and Data Preparation: Constructing a drone group delivery system; In this embodiment, step 1 specifically includes: Step 11: Based on the urban low-altitude logistics delivery needs, construct a cellular-connected drone group delivery system; the drone group delivery system defines the following elements: drones depart from the warehouse, deliver to each group of users and return to the warehouse, drones are subject to maximum load constraints, users have expected service time windows, and drones communicate with ground base stations through cellular networks; the drones refer to cellular-connected drones.
[0027] Step 12: Obtain delivery task data, communication network data, and drone performance parameters, including: The delivery task data includes the number of users, the location coordinates of each user, the required weight of goods, the user's expected service time window, and the warehouse location. The communication network data includes the number of ground base stations, the location coordinates of each base station, the transmission power, and the communication interruption threshold. ; The performance parameters of the UAV include the UAV's maximum payload, flight speed, flight altitude, UAV's own weight, and energy consumption model constant parameters.
[0028] Step 2, Construction of a multi-constraint coupling model: Construct a communication model, a user satisfaction model, a group delivery sequence model, and an energy consumption model in the UAV group delivery system. Based on the communication model, user satisfaction model, group delivery sequence model, and energy consumption model, establish an optimization objective to minimize flight energy consumption. In this embodiment, step 2 specifically includes: Step 21: Establish a communication model to quantify the risk of communication interruption during UAV flight. Use the signal-to-interference-plus-noise ratio (SINR) as an indicator of communication quality and as a hard constraint for trajectory planning to ensure flight communication safety. Step 22: Establish a user satisfaction model, quantify user satisfaction based on the user's expected service time window, evaluate the impact of delivery service timeliness on service quality, and use the average user satisfaction of all users as an indicator to measure user satisfaction. Step 23: Establish a group delivery sequence model. This model is used to address the practical constraints of the limited payload of drones in a single trip. By grouping users and planning access routes for each group, it provides optimization objects for upper-level decision-making and forms the basis for global task planning. Step 24: Based on the UAV's own weight, flight speed, and energy consumption model constant parameters, establish an energy consumption model, calculate the energy consumed by the UAV during flight, and use it as the core objective function for optimization in order to achieve energy-minimized scheduling and trajectory planning. Step 25: Based on the communication model, user satisfaction model, group delivery sequence model and energy consumption model, construct an optimization objective to minimize flight energy consumption, which is used to optimize the flight energy consumption of the UAV to complete the delivery task.
[0029] Step 3, Upper-level grouping and delivery sequence optimization: In the upper-level decision-making, the problem of optimizing delivery grouping and access order caused by the coupling of load constraints and user timeliness requirements is modeled as a vehicle routing problem with time windows (CVRPTW). Based on the optimization objective, the adaptive large neighborhood search (ALNS) algorithm is used to solve the vehicle routing problem with time windows, and a global delivery sequence that satisfies the load constraints and user satisfaction constraints is obtained. This global delivery sequence is used as the input to the lower-level decision-making, providing a definite start point, end point and load status for each sub-route. In this embodiment, step 3 uses an adaptive large neighborhood search algorithm to solve the vehicle routing problem with time windows based on the optimization objective, obtaining a global delivery sequence that satisfies both load constraints and user satisfaction constraints; specifically, it includes: Step 31, Solution Representation: Based on the number of users and warehouse locations, the solution to the problem is represented using multi-batch sequence encoding. Warehouse nodes are represented by 0. A solution is composed of multiple batch sequences separated by warehouse nodes, in the form of: In this process, the sequence between two adjacent 0s constitutes a delivery batch, the order within a delivery batch is the intra-group access order, and the order of delivery batches is the inter-group execution order. This encoding method, which integrates grouping decision and sorting decision, facilitates subsequent neighborhood operations.
[0030] Step 32, Initial Solution Generation: An initial solution is generated using a construction method that incorporates heuristic rules. In this embodiment, step 32 specifically includes: Step 321, Comprehensive Score Ranking: Calculate a comprehensive score based on the user's distance from the warehouse, the weight of the goods required, and the urgency of the user's expected service time window, and sort them in ascending order, giving priority to users with lower service scores; Step 322, Greedy Packing: Assign users to the current delivery batch in the sorted order. If the total weight after loading exceeds the maximum payload of the drone, end the current delivery batch and start a new delivery batch. Step 323, Multi-starting point nearest neighbor optimization: For the user access sequence within each delivery batch, construct the nearest neighbor path based on the user's location coordinates, starting from each user, and select the user access sequence with the lowest flight energy consumption as the access order within the group for that delivery batch.
[0031] The above method, by incorporating heuristic rules based on the structural features of the problem, can quickly generate feasible solutions that satisfy the load and time window constraints, providing high-quality initial solutions for subsequent ALNS iterations.
[0032] Step 33: Design various destruction and repair operators to perform destruction and repair operations on the current solution during the iteration process; In this embodiment, the destruction operator is used to remove some users from the current solution during the iteration process to explore a new solution space, including the following five types: (1) Random destruction: randomly select q users to remove from the current solution; (2) Worst-case energy consumption destruction: Calculate the unit energy consumption contribution of each user and remove the q users with the largest unit energy consumption contribution; (3) Time window conflict disruption: Calculate the current satisfaction level of each user and remove the q users with the lowest satisfaction levels; (4) Batch destruction: Randomly select a complete delivery batch and remove all users within that delivery batch; (5) Neighbor destruction: Randomly select a reference user, calculate the spatial distance between other users and the reference user, and remove the q nearest users; The repair operators are used to reinsert users who have had their corrupted operators removed into the solution, and include the following four types: (1) Greedy insertion: Traverse all users to be inserted and all feasible insertion positions, select the insertion method that minimizes the increase in flight energy consumption, and select the insertion method that minimizes the energy consumption increment for possible insertion positions; (2) Time window priority insertion: Sort all users to be inserted in ascending order of their latest available service time, and then select the insertion position that maximizes their time window satisfaction for each user in turn; (3) Regret value insertion: For each user to be inserted, the difference between the energy consumption increment caused by the optimal insertion position and the second-best insertion position is calculated as the regret value, and the user with the largest regret value is inserted first. (4) Random insertion: Randomly select a user to be inserted and randomly select a feasible position for insertion; During the insertion operation, load constraints and time window constraints are checked simultaneously to ensure the feasibility of the generated solution.
[0033] Step 34, 2-opt search: To improve the precision of the solution, during the adaptive large neighborhood search iteration process, a 2-opt local search is periodically performed on the user access sequence within each delivery batch. Any two edges without common user nodes are selected in the user access sequence, and each edge consists of two consecutive user nodes. The endpoint user nodes of the two edges are swapped to generate two new edges. If the swap can reduce the flight energy consumption of the delivery batch, the swap is accepted; if the swap cannot reduce the flight energy consumption of the delivery batch, the swap is not accepted. This process is repeated until no further improvement is possible or the preset number of iterations is reached. Step 35, Solution Acceptance Criterion: To balance the algorithm's global exploration and local exploitation capabilities, a simulated annealing criterion is introduced to determine whether to accept a new solution. Let the current solution be... Current solution The objective function (energy consumption) is The new interpretation is New interpretation The objective function is And the current temperature is The probability of the new solution being accepted is... for:
[0034] temperature According to the cooling rate during the iteration process Gradually decreasing in temperature, the algorithm is more likely to accept inferior solutions in the early stages of the algorithm at high temperatures, thus exploring a wider solution space; as the temperature decreases, the algorithm gradually converges and becomes more inclined to accept superior solutions.
[0035] Step 36, Adaptive Operator Selection Mechanism: To improve algorithm efficiency, an adaptive operator selection mechanism is introduced, which dynamically adjusts the selection probability of the destruction operator and the repair operator based on the historical performance of the operator.
[0036] In this embodiment, the adaptive operator selection mechanism in step 36 specifically includes: Step 361: Maintain a weight for each operator. The initial weights are set to be equal; Step 362: Before each iteration, select the destruction operator and the repair operator based on the weights using the roulette wheel method; Step 363: After the iteration is complete, the selected operators are scored based on the quality of the new solution. The higher the score, the better the performance of the solution obtained this time. The scoring principle is as follows: (1) If the new solution is better than the historical best solution, the score is ; (2) If the new solution is worse than the historical best solution but better than the current solution, the score is: ; (3) If the new solution is inferior to the current solution but is accepted, the score is: ; (4) If the new solution is inferior to the current solution and is rejected, the score is: ; in, The specific values are preset according to the application scenario; Step 364: Define each N iterations as a weight update cycle. At the end of each cycle, update the weights using the following formula based on the cumulative score and number of calls of the operator within that cycle:
[0037] in, Update the weighting factor. This represents the first operator before the update. Secondary weighting, Indicates the updated operator's first... Secondary weighting, This represents the cumulative score of the operator within the current period. This indicates the cumulative number of times the operator has been selected within the current period. This indicates that the operator has not been selected, and the cumulative score for the current period is... If it is 0, then The value is 0; Step 365: Update the weights The operator used for step 362 in the next iteration is repeated from step 362 to step 364 until the preset number of iterations is reached.
[0038] This adaptive algorithm selection mechanism grants higher selection probabilities to operators with superior performance, thereby guiding the ALNS algorithm to automatically focus on more effective search strategies. Using the ALNS algorithm described above, the UAV group delivery sequence optimization problem considering load constraints and time window constraints can be effectively solved, yielding a global delivery sequence and providing a high-quality global delivery sequence for lower-level trajectory planning.
[0039] Step 37: Repeat steps 33 to 36. After the iteration is complete, output the currently found optimal solution that satisfies the load constraint and user satisfaction constraint as the global delivery sequence for use in the next layer trajectory planning.
[0040] A specific "break-and-repair" operator is designed for ALNS, which iteratively improves the current solution and introduces an adaptive mechanism to dynamically adjust operator weights, enabling the algorithm to automatically select efficient neighborhood structures. Simultaneously, simulated annealing is used to accept inferior solutions, preventing the algorithm from getting trapped in local optima. Furthermore, the algorithm introduces 2-opt local search as an independent optimization step, periodically refining the solution during ALNS iterations to obtain a global delivery sequence considering load constraints and time window constraints. This ALNS algorithm has been selected as a novel method for solving such NP-hard problems due to its powerful global search capability, flexible neighborhood structure design, and ability to embed problem-specific heuristics.
[0041] Step 4, Lower-level trajectory optimization and Markov decision process transformation: In the lower-level decision-making, based on the global delivery sequence, under the given delivery sequence conditions, for the fine trajectory planning problem under the communication reliability constraints in the complex urban wireless environment, the trajectory planning problem of each sub-segment is modeled as a Markov decision process, and the deep reinforcement learning (DQN) algorithm is used for solving and training to obtain the flight strategy network model. This flight strategy network model is used to output flight direction decisions that satisfy the communication interruption probability constraints according to the environmental state. In this embodiment, step 4 specifically includes: Step 41: After determining the global delivery sequence in the upper-level stage, the service order of the UAV is fixed throughout the mission. The sub-problem is focused on the hard-constraint trajectory planning problem of communication in the delivery process between two adjacent users, with the goal of minimizing the flight energy consumption of each sub-segment. Step 42: Based on the number of ground base stations, the location coordinates of each base station, the transmission power, and the communication interruption threshold, construct a low-altitude radio map of the city. Model the trajectory planning problem as a Markov Decision Process (MDP), defining the state, action, and reward as follows: (1) State s: The horizontal coordinate of the UAV at a certain moment; (2) Action a: Discretize the horizontal flight direction of the UAV into multiple uniformly distributed fixed directions, and each action corresponds to a fixed flight angle; (3) Reward r: Used to comprehensively reflect the satisfaction of energy consumption and communication constraints. The corresponding communication interruption probability is determined based on the current flight position and flight altitude of the UAV. When the probability of communication interruption of the UAV exceeds the communication interruption threshold, the reward r is awarded. When the area is in violation, penalties are imposed to improve communication reliability during drone flight. Step 43: Since the communication interruption probability and the UAV position exhibit randomness and non-convexity, traditional convex optimization or analytical methods are difficult to solve directly. Therefore, this invention adopts the Deep Q-Network (DQN) algorithm based on interactive learning to solve the Markov decision process, ensuring the feasibility and safety of the action. An action masking mechanism is introduced in the action selection stage to filter out invalid actions that cause the UAV to go beyond the boundary or cause communication interruption. In this scenario, communication constraints constitute sparse rewards. The deep reinforcement learning algorithm employs an experience replay mechanism, storing historical experience and randomly sampling it to break the temporal correlation between data, thereby improving learning efficiency in sparse reward scenarios and effectively enhancing learning stability in continuous state space. By iteratively updating the policy network parameters through an approximation of the optimal action value network, the UAV learns a flight policy network model that satisfies communication constraints under a given state. This flight policy network model is used to minimize cumulative flight energy consumption while satisfying the communication interruption probability constraint, and outputs flight direction decisions based on real-time environmental conditions. The training process of the flight strategy network model is as follows: Step 431, Initialization: Initialize the network parameters of the action-value network Q. Initialize the target value network Network parameters Initialize the experience replay pool D, and set the starting point of the current sub-segment to the initial state. ; Step 432, Action Selection: In the current state The following is adopted - Greedy strategy selects actions based on probability. Randomly select an action with a probability of 1- choose At the same time, an action masking mechanism is introduced to filter out invalid actions that cause the drone to go beyond the boundary or cause communication interruption. Step 433, Perform Actions and Update the Environment: The UAV performs the selected actions. Transition to the next state Calculate the probability of communication disruption based on urban low-altitude radio maps and receive an instant reward. The reward Taking into account both energy consumption penalties and communication constraints, a penalty is imposed when the probability of communication interruption exceeds the communication interruption threshold; Step 434, Storing Experience: Store the experience samples ( Stored in the experience replay pool D; Step 435, Experience Replay and Network Update: Randomly sample a small batch of samples from the experience replay pool D, and calculate the target value network. The network parameters of the action value network Q are updated using gradient descent. ; Step 436: Update the target network: Every fixed number of steps, update the network parameters of the action value network Q. Replicating to the target value network, i.e. ; Step 437, Termination Judgment: Determine whether the sub-segment endpoint has been reached or the preset maximum number of steps has been reached. If not, then... If yes, return to step 432 to continue iterating; otherwise, end the training of the current sub-segment and output the flight strategy network model.
[0042] Through the above training process, the UAV learns a flight strategy network model that satisfies communication constraints under a given state. This model can output flight direction decisions based on real-time environmental conditions, with the goal of minimizing cumulative flight energy consumption, while satisfying the communication interruption probability constraint.
[0043] Step 44: Solve for all sub-segments in the global delivery sequence, execute the trained flight strategy network model for each sub-segment, record the flight energy consumption of each sub-segment and sum them up to obtain the flight energy consumption of the UAV to complete the entire delivery task.
[0044] Step 5, Model Deployment and Online Decision Implementation: Deploy the global delivery sequence and the trained flight strategy network model to the UAV's onboard computing unit; during actual flight, the onboard computing unit loads and runs the flight strategy network model, which generates flight direction decisions in real time based on real-time perceived environmental information and executes the delivery task; In this embodiment, step 5 specifically includes: Step 51: Based on the environmental information, deploy the global delivery sequence and the trained flight strategy network model to the UAV's onboard computing unit; Step 52: The UAV's onboard computing unit loads the flight strategy network model and runs the flight strategy network model based on the perceived environmental information; Step 53: The flight strategy network model outputs flight direction decisions in real time and synchronizes them to the UAV flight control system; Step 54: The UAV flight control system controls the UAV to execute flight missions according to flight direction decisions and complete the delivery. Step 6, Collaborative Optimization and Solution Output: Through collaborative optimization between the upper and lower layers, an optimal delivery solution is output that satisfies load constraints, user satisfaction constraints, and communication reliability constraints, thereby minimizing the drone's flight energy consumption. The optimal delivery solution includes: delivery batch division, user access order within each group, and flight trajectory for each sub-segment.
[0045] The above description is only a part of the embodiments of the present invention and does not limit the scope of protection of the present invention. Any equivalent device or equivalent process transformation made based on the content of the present invention specification and drawings, or direct or indirect application in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A method for energy consumption optimization of communication and sensing of a payload-restricted unmanned aerial vehicle based on hierarchical decision-making, characterized in that, Includes the following steps: Step 1: Build a drone group delivery system; Step 2: Construct a communication model, a user satisfaction model, a group delivery sequence model, and an energy consumption model in the UAV group delivery system. Based on the communication model, user satisfaction model, group delivery sequence model, and energy consumption model, establish an optimization objective to minimize flight energy consumption. Step 3: In the upper-level decision-making process, the delivery grouping and access order optimization problem is modeled as a vehicle routing problem with a time window. Based on the optimization objective, an adaptive large neighborhood search algorithm is used to solve the vehicle routing problem with the time window, obtaining a global delivery sequence that satisfies the load constraint and the user satisfaction constraint; specifically including: Step 31: Based on the number of users and the location of the warehouse, the solution to the problem is represented by a multi-batch sequence encoding. The warehouse node is represented by 0. A solution is composed of multiple batch sequences with the warehouse node as the separator. The sequence between two adjacent 0s constitutes a delivery batch. The order within a delivery batch is the access order within the group, and the order of delivery batches is the execution order between groups. Step 32: Generate an initial solution using a construction method that incorporates heuristic rules; Step 33: Design various destruction and repair operators to perform destruction and repair operations on the current solution during the iteration process; Step 34: During the adaptive large neighborhood search iteration process, periodically perform a 2-opt local search on the user access sequence within each delivery batch. Select any two edges in the user access sequence that do not have a common user node. Each edge consists of two consecutive user nodes. Attempt to swap the end user nodes of the two edges to generate two new edges. If the swap can reduce the flight energy consumption of the delivery batch, the swap is accepted. If the swap cannot reduce the flight energy consumption of the delivery batch, the swap is not accepted. This process is repeated until no further improvement is possible or the preset number of iterations is reached. Step 35: Introduce the simulated annealing criterion to determine whether to accept the new solution. Let the current solution be... Current solution The objective function is The new interpretation is New interpretation The objective function is And the current temperature is The probability of the new solution being accepted is... for: Step 36: Introduce an adaptive operator selection mechanism to dynamically adjust the selection probability of the destruction operator and the repair operator based on the historical performance of the operators; Step 37: Repeat steps 33 to 36. After the iteration is complete, output the optimal solution that satisfies the load constraint and user satisfaction constraint as the global delivery sequence for use in the next layer trajectory planning. Step 4: In the lower-level decision-making, based on the global delivery sequence, the trajectory planning problem of each sub-segment is modeled as a Markov decision process, and a deep reinforcement learning algorithm is used to solve and train it to obtain a flight strategy network model. This flight strategy network model is used to output a flight direction decision that satisfies the communication interruption probability constraint according to the environmental state. Step 5: Deploy the global delivery sequence and the trained flight strategy network model to the UAV's onboard computing unit; during actual flight, the onboard computing unit loads and runs the flight strategy network model, which generates flight direction decisions in real time based on real-time perceived environmental information and executes the delivery task. Step 6: Through collaborative optimization between the upper and lower layers, output the optimal delivery plan that satisfies the constraints of load capacity, user satisfaction, and communication reliability, thereby minimizing the flight energy consumption of the drone.
2. The method for optimizing communication and sensing energy consumption of payload-constrained UAVs based on hierarchical decision-making as described in claim 1, characterized in that, Step 1 specifically includes: Step 11: Based on the urban low-altitude logistics delivery needs, construct a cellular-connected drone group delivery system; the drone group delivery system defines the following elements: drones depart from the warehouse, deliver to each group of users and return to the warehouse, drones are subject to maximum load constraints, users have expected service time windows, and drones communicate with ground base stations through cellular networks. Step 12: Obtain delivery task data, communication network data, and drone performance parameters, including: The delivery task data includes the number of users, the location coordinates of each user, the required weight of goods, the user's expected service time window, and the warehouse location. The communication network data includes the number of ground base stations, the location coordinates of each base station, the transmission power, and the communication interruption threshold. ; The performance parameters of the UAV include the UAV's maximum payload, flight speed, flight altitude, UAV's own weight, and energy consumption model constant parameters.
3. The method for optimizing communication and sensing energy consumption of payload-constrained UAVs based on hierarchical decision-making as described in claim 1, characterized in that, Step 2 specifically includes: Step 21: Establish a communication model, using the signal-to-interference-plus-noise ratio (SINR) as an indicator of communication quality and as a hard constraint for trajectory planning. Step 22: Establish a user satisfaction model, quantify user satisfaction based on the user's expected service time window, and use the average user satisfaction of all users as an indicator to measure user satisfaction. Step 23: Establish a group delivery sequence model. By grouping users and planning access routes for each group, optimization objects are provided for upper-level decision-making. Step 24: Based on the UAV's own weight, flight speed, and energy consumption model constant parameters, establish an energy consumption model, calculate the energy consumed by the UAV during flight, and use it as the core objective function for optimization. Step 25: Construct an optimization objective to minimize flight energy consumption based on the communication model, user satisfaction model, group delivery sequence model, and energy consumption model.
4. The energy consumption optimization method for communication and sensing of payload-constrained UAVs based on hierarchical decision-making as described in claim 1, characterized in that, Step 32 specifically includes: Step 321: Calculate a comprehensive score based on the user's distance from the warehouse, the weight of the goods required, and the urgency of the user's expected service time window, and sort them in ascending order; Step 322: Assign users to the current delivery batch in the sorted order. If the total weight after loading exceeds the maximum payload of the drone, end the current delivery batch and start a new delivery batch. Step 323: For the user access sequence within each delivery batch, construct the nearest neighbor path based on the user's location coordinates, starting from each user, and select the user access sequence with the lowest flight energy consumption as the access order within the group for that delivery batch.
5. The method for optimizing communication and sensing energy consumption of payload-constrained UAVs based on hierarchical decision-making as described in claim 1, characterized in that, The destruction operators include the following five types: (1) Random destruction: random selection One user is removed from the current solution; (2) Worst-case energy consumption destruction: Calculate the unit energy consumption contribution of each user and remove the q users with the largest unit energy consumption contribution; (3) Time window conflict disruption: Calculate the current satisfaction level of each user and remove the q users with the lowest satisfaction levels; (4) Batch destruction: Randomly select a complete delivery batch and remove all users within that delivery batch; (5) Neighbor destruction: Randomly select a reference user, calculate the spatial distance between other users and the reference user, and remove the q nearest users; The repair operators include the following four types: (1) Greedy insertion: Traverse all users to be inserted and all feasible insertion positions, and select the insertion method that minimizes the increase in flight energy consumption; (2) Time window priority insertion: Sort all users to be inserted in ascending order of their latest available service time, and then select the insertion position that maximizes their time window satisfaction for each user in turn; (3) Regret value insertion: For each user to be inserted, the difference between the energy consumption increment caused by the optimal insertion position and the second-best insertion position is calculated as the regret value, and the user with the largest regret value is inserted first. (4) Random insertion: Randomly select a user to be inserted and randomly select a feasible position for insertion; When performing an insertion operation, load constraints and time window constraints are checked simultaneously.
6. The energy consumption optimization method for communication and sensing of payload-constrained UAVs based on hierarchical decision-making as described in claim 1, characterized in that, The adaptive operator selection mechanism in step 36 specifically includes: Step 361: Maintain a weight for each operator. The initial weights are set to be equal; Step 362: Before each iteration, select the destruction operator and the repair operator based on the weights using the roulette wheel method; Step 363: After the iteration is completed, the selected operators are scored based on the quality of the new solutions. The scoring principles are as follows: (1) If the new solution is better than the historical best solution, the score is ; (2) If the new solution is worse than the historical best solution but better than the current solution, the score is: ; (3) If the new solution is inferior to the current solution but is accepted, the score is: ; (4) If the new solution is inferior to the current solution and is rejected, the score is: ; in, The specific values are preset according to the application scenario; Step 364: Define each N iterations as a weight update cycle. At the end of each cycle, update the weights using the following formula based on the cumulative score and number of calls of the operator within that cycle: in, Update the weighting factor. This represents the first operator before the update. Secondary weighting, Indicates the updated operator's first... Secondary weighting, This represents the cumulative score of the operator within the current period. This indicates the cumulative number of times the operator has been selected within the current period. This indicates that the operator has not been selected, and the cumulative score for the current period is... If it is 0, then The value is 0; Step 365: Update the weights The operator used for step 362 in the next iteration is repeated from step 362 to step 364 until the preset number of iterations is reached.
7. The method for optimizing communication and sensing energy consumption of payload-constrained UAVs based on hierarchical decision-making as described in claim 1, characterized in that, Step 4 specifically includes: Step 41: After determining the global delivery sequence in the upper-level stage, the sub-problem is focused on the communication hard constraint trajectory planning problem in the delivery process between adjacent users, with the goal of minimizing the flight energy consumption of each sub-segment; Step 42: Based on the number of ground base stations, the location coordinates of each base station, the transmission power, and the communication interruption threshold, construct a low-altitude radio map of the city. Model the trajectory planning problem as a Markov decision process, defining the state, action, and reward as follows: (1) State s: The horizontal coordinate of the UAV at a certain moment; (2) Action a: Discretize the horizontal flight direction of the UAV into multiple uniformly distributed fixed directions, and each action corresponds to a fixed flight angle; (3) Reward r: Used to comprehensively reflect the satisfaction of energy consumption and communication constraints. The corresponding communication interruption probability is determined based on the current flight position and flight altitude of the UAV. When the probability of communication interruption of the UAV exceeds the communication interruption threshold, the reward r is awarded. When the area is [affected / injured], a penalty is imposed; Step 43: The deep reinforcement learning DQN algorithm is used to solve the Markov decision process. An action masking mechanism is introduced in the action selection stage to filter out invalid actions that cause the UAV to go out of the boundary or cause communication interruption. The deep reinforcement learning algorithm employs an experience replay mechanism to store and randomly sample historical experiences. By iteratively updating the policy network parameters through an approximation of the optimal action value network, the UAV learns a flight policy network model that satisfies communication constraints under a given state. This flight policy network model is used to minimize cumulative flight energy consumption under the premise of satisfying the communication interruption probability constraint, and outputs flight direction decisions based on the real-time environmental state. Step 44: For all sub-segments in the global delivery sequence, execute the trained flight strategy network model to fly, record the flight energy consumption of each sub-segment and sum them up to obtain the flight energy consumption of the UAV to complete the entire delivery task.
8. The method for optimizing communication and sensing energy consumption of payload-constrained UAVs based on hierarchical decision-making as described in claim 7, characterized in that, The training process of the flight strategy network model is as follows: Step 431, Initialization: Initialize the network parameters of the action-value network Q. Initialize the target value network Network parameters Initialize the experience replay pool D, and set the starting point of the current sub-segment to the initial state. ; Step 432, Action Selection: In the current state The following is adopted - Greedy strategy selects actions based on probability. Randomly select an action with a probability of 1- choose At the same time, an action masking mechanism is introduced to filter out invalid actions that cause the drone to go beyond the boundary or cause communication interruption. Step 433, Perform Actions and Update the Environment: The UAV performs the selected actions. Transition to the next state Calculate the probability of communication disruption based on urban low-altitude radio maps and receive an instant reward. The reward Taking into account both energy consumption penalties and communication constraints, a penalty is imposed when the probability of communication interruption exceeds the communication interruption threshold; Step 434, Storing Experience: Store the experience samples ( Stored in the experience replay pool D; Step 435, Experience Replay and Network Update: Randomly sample a small batch of samples from the experience replay pool D, and calculate the target value network. The network parameters of the action value network Q are updated using gradient descent. ; Step 436: Update the target network: Every fixed number of steps, update the network parameters of the action value network Q. Replicating to the target value network, i.e. ; Step 437, Termination Judgment: Determine whether the sub-segment endpoint has been reached or the preset maximum number of steps has been reached. If not, then... If yes, return to step 432 to continue iterating; otherwise, end the training of the current sub-segment and output the flight strategy network model.
9. The method for optimizing communication and sensing energy consumption of payload-constrained UAVs based on hierarchical decision-making as described in claim 1, characterized in that, Step 5 specifically includes: Step 51: Based on the environmental information, deploy the global delivery sequence and the trained flight strategy network model to the UAV's onboard computing unit; Step 52: The UAV's onboard computing unit loads the flight strategy network model and runs the flight strategy network model based on the perceived environmental information; Step 53: The flight strategy network model outputs flight direction decisions in real time and synchronizes them to the UAV flight control system; Step 54: The UAV flight control system controls the UAV to execute flight missions according to flight direction decisions and complete the delivery. The optimal delivery plan output in step 6 includes: delivery batch division, user access order within each group, and flight trajectory of each sub-segment.