Rescue system path planning method and system based on drone-assisted communication
Through nested rollback strategy adaptation methods combined with Monte Carlo search, dynamically plan the flight path of the drone, solving the problem that the UAV communication system on the emergency rescue site is difficult to flexibly adjust the path and resources, improving system efficiency and response speed, and optimizing the average throughput.
Patent Information
- Application Number
- CN202510244389.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-03
AI Technical Summary
At the emergency rescue site, it is difficult for existing drone communication systems to flexibly adjust flight paths and communication resources, resulting in low mission execution efficiency and unstable communication quality. The battery life of drones is limited. How to optimize flight paths and computing resource allocation under limited power and resources has become a problem.
Introducing Nested Rollback Strategy Adaptation (NRPA) method, combining Monte Carlo search and rollback strategy adaptation, dynamically planning the flight path of the drone, optimizing computing resource allocation, and taking into account constraints such as battery energy consumption and task fairness.
It improves the overall performance and response speed of the system, optimizes the average throughput of the system, solves the problem of UAV flight path planning under dynamic requirements, and avoids the high computing complexity and local optimal traps of traditional algorithms.
Smart Images

Figure CN119737956B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of path planning, and in particular to a rescue system path planning method and system based on unmanned aerial vehicle (UAV) assisted communication. Background Art
[0002] With the rapid development of UAV technology, the application of UAVs in many fields has gradually increased. Especially in emergency rescue and disaster emergency response, UAVs have become an important tool for performing emergency tasks with their high mobility, rapid response capability and low cost. In complex post-disaster rescue scenarios, UAVs not only undertake tasks such as search, reconnaissance, and material delivery, but also need to provide real-time wireless communication support to ensure that ground rescue personnel can quickly obtain critical mission information and data. However, existing UAV communication systems face many challenges. First, the needs of emergency rescue sites are highly dynamic and uncertain. As the situation on the scene changes, UAVs need to flexibly adjust flight paths and communication resources to ensure efficient mission execution and communication quality. Second, the battery life of UAVs is limited, and the energy consumption of batteries is affected by multiple factors such as flight path, mission load and environmental conditions. Therefore, with limited power and resources, how to optimize the flight path of UAVs, reasonably allocate computing resources, and ensure the timely completion of tasks and efficient operation of the system has become a problem that needs to be solved urgently. Summary of the invention
[0003] In view of this, the present invention introduces the nested rollback policy adaptation (NRPA) method, combined with Monte Carlo search and rollback policy adaptation, to efficiently plan the flight path of the UAV in a dynamically changing environment. Through this method, the UAV can flexibly adjust the flight path according to the current mission requirements and environmental conditions, optimize the allocation of computing resources, and consider constraints such as battery energy consumption and task fairness, thereby improving the overall performance and response speed of the system.
[0004] To achieve the above purpose, the present invention provides the following technical solution: a rescue system path planning method based on drone-assisted communication, specifically:
[0005] S1. Based on the initial state of the UAV, set the initial UAV flight strategy as the starting point of path planning and establish a preliminary path planning plan, wherein the initial state of the UAV includes at least: UAV position, power, and communication range;
[0006] S2. Adopt the nested rollback strategy to adapt the NRPA method for path planning. The NRPA method is based on the Monte Carlo search principle. At each level of the flight path, a Monte Carlo simulation method is used to perform random path simulation, generate paths and calculate scores. The path score under the current strategy is compared with the existing best path score. If the current path score is better, the best path and the best score are updated.
[0007] S3. Adopt an adaptive adjustment strategy to adjust the flight action weights according to the current best path and best score, optimize the flight strategy, and make the flight strategy tend to high-quality paths.
[0008] In particular, the S1 step also sets the recursive depth of subsequent path planning according to the computing resources and battery power of the drone, and controls the number of levels of path search.
[0009] Furthermore, the nested rollback strategy is adapted to the NRPA method, specifically:
[0010] S201. Determine the current level. If the level is zero, perform a backtracking operation to obtain the path planning result in the current state;
[0011] S202. If the level is greater than zero, a recursive process is entered, and the NRPA method is called in sequence to obtain subpaths at lower levels through backtracking operations and evaluate the subpath scores;
[0012] S203. In each recursive process, the path score under the current strategy is evaluated and compared with the current best score. If the current path score is better than the existing best path score, the best score and path are updated.
[0013] Furthermore, the adaptive adjustment strategy uses the Adapt algorithm optimization strategy to adjust the flight action weights, and the optimized strategy is used to guide the next round of Monte Carlo simulation, namely: , where policy[m] represents the probability of selecting action m, m' represents an action among all actions, α represents the learning rate, and β represents the decay rate.
[0014] Furthermore, the nested rollback strategy adapts to the NRPA method and also collects environmental feedback information to optimize the flight path, wherein the environmental feedback indicators at least include: calculating energy consumption , flight energy consumption And the drone hovering energy consumption ,in, , o i (t) represents the number of tasks unloaded by the i-th rescue team at time t, represents the effective switching capacitance, C represents the number of CPU cycles required for each input bit, and S brepresents the total number of bits for a single task, and fc represents the CPU frequency; , k represents the parameters related to the drone mass M and acceleration a, V represents the drone speed, μ 1 , μ 2 represents the parameters of the UAV system, g represents the acceleration of gravity; , P h Indicates the hovering power, Indicates the upload rate.
[0015] Furthermore, when the nested rollback strategy is used to adapt to the NRPA method for path planning, the average throughput is maximized and the task allocation is fair. The average throughput objective function is: , T represents T time periods; energy consumption constraints: ; Fairness constraints: , λ is the battery charge threshold, E is the battery capacity, c i (t) The service status of each rescue team.
[0016] Furthermore, during the path planning process, the battery power of the drone is monitored in real time. When the battery power is greater than the set threshold, the path planning operation continues; when the battery power is less than the threshold, the path planning is stopped and the drone is guided to the nearest charging point or safe area.
[0017] Furthermore, the present invention also provides a multi-objective path optimization system in a UAV-assisted communication system, comprising:
[0018] A drone module, used to actually perform rescue missions;
[0019] Wireless communication system, which provides computing task offloading services for ground users and ensures that task data can be transmitted to the ground operation platform in a timely manner;
[0020] The control module uses the rescue system path planning method based on UAV-assisted communication to optimize and adjust the flight path of the UAV, ensuring that the UAV can perform tasks in complex rescue environments and return efficiently;
[0021] The battery management module monitors the battery power of the drone in real time and adjusts the flight mission execution strategy according to the remaining power.
[0022] Compared with the prior art, the present invention has the following beneficial effects:
[0023] (1) The present invention proposes a nested rollback strategy to adapt to the NRPA method to dynamically plan the flight path of the UAV, which improves the efficiency and response speed of the system. At the same time, under the premise of considering constraints such as fairness and energy consumption, the average throughput of the system is optimized, solving the problem of UAV flight path planning under dynamic requirements.
[0024] (2) Compared with traditional brute force search and depth-first search algorithms, the present invention avoids the high computational complexity of brute force search and the local optimal trap of depth-first search by combining strategy optimization and random simulation, and quickly approaches the global optimal solution in problems with a huge solution space. At the same time, NRPA can balance the computational time and solution quality by adjusting the recursive level and the number of iterations. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 An optional flow chart of the path planning method of the present invention;
[0026] Figure 2 Pseudo code of adaptive strategy update algorithm of the present invention;
[0027] Figure 3 A simulation scenario for an embodiment of the present invention;
[0028] Figure 4 is a comparison of average throughputs of embodiments of the present invention;
[0029] Figure 5 FIG. 4 is a comparison of throughputs at different hovering points according to embodiments of the present invention. DETAILED DESCRIPTION
[0030] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0031] The present invention is based on the method of unmanned aerial vehicle rescue communication path planning based on the nested rollback strategy adaptation method, which mainly uses a specific strategy adaptation method to plan the path of the unmanned aerial vehicle, especially in the rescue communication scenario. The nested rollback strategy adaptation method may be a recursive or hierarchical strategy adjustment method for dealing with path planning problems in complex environments, which can improve the efficiency of rescue while ensuring fairness, and can complete the computing task offloading service more quickly and accurately. Specifically, Figure 1 shown.
[0032] Step 1: First, based on the initial state of the drone, such as position, power, communication range, etc., initialize the flight strategy of the drone, which is used to guide the flight decision of the drone. The initial strategy provides a preliminary path planning scheme based on the environment and task requirements. Next, set the recursive depth of the path planning. The recursive depth is used to control the number of levels of path search to ensure that the algorithm can obtain sufficiently accurate path planning within limited computing resources and time. The core purpose of this step is to provide preliminary decision-making basis for subsequent path planning and ensure that the algorithm can be flexibly adjusted according to the complexity of different tasks.
[0033] Step 2: Use the nested rollback strategy adaptation method for path planning. This step is the core of the entire process. When planning the path, the drone will continuously adjust the strategy according to the current environmental information and mission requirements to find the optimal path. Specifically, by recursively evaluating the path planning effects at different levels, the NRPA method obtains the result of the current flight path through a backtracking operation at each level. If the current level is zero, the backtracking operation is performed directly to evaluate the path planning scheme under the current state; if the current level is greater than zero, the NRPA method is recursively called to enter a lower level and perform a backtracking operation to evaluate the effect of the sub-path. In each recursive process, the path score under the current strategy will be compared with the existing best path score. If the current path score is better, the best path and the best score will be updated. The goal of this step is to continuously optimize the flight path through multiple recursions and backtracking, and to ensure that each recursion is dynamically adjusted according to the current environment and mission requirements. The dynamic adjustment indicators are: Calculate energy consumption , flight energy consumption And the drone hovering energy consumption .
[0034] Step 3: Collect feedback information and adjust the strategy based on the feedback information. This step emphasizes the importance of feedback. During the process of executing path planning, the drone will collect feedback information from the environment and then adjust the strategy based on this information to ensure the accuracy and effectiveness of path planning. The specific strategy optimization method is: after each simulation, such as Figure 2 As shown in the figure, NRPA uses the Adapt algorithm to optimize the strategy and adjust the action weights to make the strategy more inclined to high-quality paths. The optimized strategy is used to guide the next round of simulation and gradually improve the path planning effect. , where P(m) is the probability of selecting action m, policy[m] is the weight of action m. m′ represents an action among all possible actions. , where α is the learning rate (used to increase the weight of the selected action) and β is the decay rate.
[0035] Step 4: In this process, the drone needs to constantly evaluate the effectiveness of the current strategy, such as by comparing the communication throughput of different paths and selecting the path with the highest throughput. At the same time, the drone also needs to consider energy consumption to ensure that it does not run out of power while completing the task. Through multiple iterations, the drone can find an optimal path that avoids obstacles and maintains high communication throughput. Specifically, in the path planning process, a variety of constraints are considered comprehensively, such as energy consumption constraints, task fairness constraints, real-time response capabilities, etc. By dynamically optimizing the flight path of the drone, the average throughput of the system is maximized, while ensuring the fairness of task allocation and reducing battery consumption. Among them, the objective function: , maximize the average throughput, that is, the average value of the user unloading data in T time periods, energy consumption constraints: , fairness constraints: , λ is the battery power threshold, in this embodiment λ is 10%, E is the battery capacity, c i (t) is the service status of each rescue team, that is, each user visits no more than twice. This step ensures that the drone can perform tasks efficiently in a variety of complex environments and meet the multiple needs of the rescue site.
[0036] Figure 3 The figure shows the background of a wireless communication scenario, in which a line-of-sight (LoS) communication link between an unmanned aerial vehicle (UAV) and a ground rescue team is envisioned. In this scenario, the channel gain follows the free space path loss model. In this configuration, a single UAV is responsible for providing services to multiple ground rescue teams distributed in a specific area. There are N fixed hovering points above the disaster area, and the UAV can only hover at these predefined locations to perform computing tasks for the rescue teams in the area. In addition, the UAV will fly horizontally at a fixed altitude H. The total service time T is divided into n independent time periods, and the time intervals are denoted by t=0,1,2,...,n-1. In each time interval, the UAV will move from one hovering point to another, giving priority to providing services to the nearest rescue team. At the beginning of each time interval, all the rescue teams will send a mission request containing their location information to the UAV so that the UAV can update the user's location information. At the initial time, t=0, the positions of the rescue teams are randomly determined. But at the beginning of the next time interval, their positions will be randomly adjusted. The height H is fixed, and at the same time, there are N fixed hovering points, which are expressed in three-dimensional coordinates.
[0037] Figure 4As shown in Figure 2, for a comprehensive evaluation, we compared the NRPA algorithm with the MCTS (Monte Carlo Tree Search) and NMCS (Nested Monte Carlo Search) baselines at levels 2 and 3 (i.e., the nesting level of NRPA and NMCS). Our results show that NRPA at level 2 outperforms the MCTS and NMCS baselines at all levels in terms of average reward. This advantage highlights the superior performance of NRPA in effectively balancing exploration and exploitation. The NMCS baseline shows better performance compared to MCTS, especially at level 2. Both NMCS and NRPA show the best results at level 2, indicating that this level achieves the best balance between exploring the path space and exploiting the available information for decision making. In contrast, increasing the search depth to level 3 introduces too much exploration, which may lead to missing some promising paths, resulting in lower average rewards than level 2. This observation emphasizes the effectiveness of NRPA at level 2 as it successfully finds a balance between exploration and reward optimization.
[0038] Figure 5 The average performance under different training rounds is shown. We evaluated the MCTS, NMCS and NRPA algorithms, recorded data every 50 rounds, and used 1000 simulations as the horizontal axis. As the number of training rounds increased, the average throughput gradually increased, among which the 2-level NRPA achieved the highest throughput. This highlights the advantages of NRPA strategy adaptation, indicating that it has more excellent performance than NMCS in effectively adjusting strategies and optimizing throughput. For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-mentioned exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive, and the scope of the present invention is defined by the attached claims rather than the above description, and it is intended to include all changes that fall within the meaning and scope of the equivalent elements of the claims in the present invention. Any figure mark in the claims should not be regarded as limiting the claims involved.
[0039] The above description of the disclosed embodiments enables one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A rescue system path planning method based on drone-assisted communication, characterized in that: The path planning method is specifically as follows: S1. Based on the initial state of the UAV, set the initial UAV flight strategy as the starting point of path planning and establish a preliminary path planning plan, wherein the initial state of the UAV includes at least: UAV position, power, and communication range; S2. Adopt the nested rollback strategy to adapt the NRPA method for path planning. The NRPA method is based on the Monte Carlo search principle. At each level of the flight path, a Monte Carlo simulation method is used to perform random path simulation, generate paths and calculate scores. The path score under the current strategy is compared with the existing best path score. If the current path score is better, the best path and the best score are updated. S3. Adopt an adaptive adjustment strategy to adjust the flight action weights according to the current best path and best score, optimize the flight strategy, and make the flight strategy tend to high-quality paths.
2. The method for rescue system path planning based on drone-assisted communication according to claim 1, characterized in that: The S1 step also sets the recursive depth of subsequent path planning according to the computing resources and battery power of the drone, and controls the number of levels of path search.
3. The method for rescue system path planning based on drone-assisted communication according to claim 1, characterized in that: The nested rollback strategy is adapted to the NRPA method, specifically: S201. Determine the current level. If the level is zero, perform a backtracking operation to obtain the path planning result in the current state; S202. If the level is greater than zero, a recursive process is entered, and the NRPA method is called in sequence to obtain subpaths at lower levels through backtracking operations and evaluate the subpath scores; S203. In each recursive process, the path score under the current strategy is evaluated and compared with the current best score. If the current path score is better than the existing best path score, the best score and path are updated.
4. The method for rescue system path planning based on drone-assisted communication according to claim 1, characterized in that: The adaptive adjustment strategy adopts the Adapt algorithm optimization strategy to adjust the flight action weights. The optimized strategy is used to guide the next round of Monte Carlo simulation, namely: , where policy[m] represents the probability of selecting action m, m ’ represents an action among all actions, α represents the learning rate, and β represents the decay rate.
5. The method for rescue system path planning based on drone-assisted communication according to claim 1, characterized in that: The nested rollback strategy adapts to the NRPA method and also collects environmental feedback information to optimize the flight path, where the environmental feedback indicators at least include: calculating energy consumption , flight energy consumption And the drone hovering energy consumption ,in, , o i (t) represents the number of tasks unloaded by the i-th rescue team at time t, represents the effective switching capacitance, C represents the number of CPU cycles required for each input bit, and S b represents the total number of bits of a single task, f c Indicates CPU frequency; , k represents the parameters related to the mass M and acceleration a of the UAV, V represents the speed of the UAV, μ1 and μ2 represent the system parameters of the UAV, and g represents the acceleration of gravity; , P h Indicates the hovering power, Indicates the upload rate.
6. The method for rescue system path planning based on drone-assisted communication according to claim 5, characterized in that: When the nested rollback strategy is used to adapt to the NRPA method for path planning, the average throughput is maximized and the task allocation is fair. The average throughput objective function is: , T represents T time periods; energy consumption constraints: ; Fairness constraints: , λ is the battery charge threshold, E is the battery capacity, c i (t) The service status of each rescue team, i (t) represents the number of tasks unloaded by the i-th rescue team at time t, S b Indicates the total number of bits for a single task, represents the computing energy consumption, Indicates flight energy consumption, Indicates the hovering energy consumption of the drone.
7. The method for rescue system path planning based on drone-assisted communication according to claim 1, characterized in that: During the path planning process, the battery power of the drone is also monitored in real time. When the battery power is greater than the set threshold, the path planning operation continues; when the battery power is less than the threshold, the path planning is stopped and the drone is guided to the nearest charging point or safe area.
8. A system using the rescue system path planning method based on drone-assisted communication according to any one of claims 1 to 7, characterized in that: The system includes: A drone module, used to actually perform rescue missions; Wireless communication system, which provides computing task offloading services for ground users and ensures that task data can be transmitted to the ground operation platform in a timely manner; The control module uses the rescue system path planning method based on UAV-assisted communication to optimize and adjust the flight path of the UAV, ensuring that the UAV can perform tasks in complex rescue environments and return efficiently; The battery management module monitors the battery power of the drone in real time and adjusts the flight mission execution strategy according to the remaining power.