Complex environment path planning system suitable for fire-fighting unmanned aerial vehicle and operation method

By combining deep reinforcement learning and fast-scaling random tree algorithm, an intelligent path planning system for unmanned systems in complex environments was designed, which solved the problem of insufficient dynamic adaptability and optimization of path planning in complex environments in the existing technology, and achieved rapid and safe path planning for drones in fire rescue scenarios.

CN119987390APending Publication Date: 2025-05-13QINGDAO SHITIAN INNOVATION AVIATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411920529.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to meet the needs of complex environment perception, multi-objective optimization and dynamic adaptation in the path planning of unmanned systems in complex environments. Especially in fire rescue scenarios, traditional algorithms lack dynamic adjustment capabilities and path optimization efficiency.

Method used

Combining deep reinforcement learning (DQN) algorithm and fast-scaling random tree (RRT**) algorithm, an intelligent path planning system for unmanned systems in complex environments is designed. The DQN algorithm is used to adaptively generate the optimal path in a dynamic environment, and the RRT** algorithm is used to optimize path length and energy consumption to achieve multi-objective balance.

Benefits of technology

It realizes dynamic adaptability and optimization of unmanned system paths in complex environments, ensures that the drone can quickly and safely reach the fire source location, and significantly improves the efficiency and adaptability of path planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987390A_ABST
    Figure CN119987390A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of unmanned system path planning, in particular to a complex environment path planning system suitable for a fire-fighting unmanned aerial vehicle and an operation method.The complex environment path planning system combines a deep reinforcement learning (DQN) algorithm and a fast extended random tree (RRT * *) algorithm, path generation, optimization and dynamic adjustment can be achieved in a complex dynamic environment, and path planning efficiency is improved. According to the technical scheme, the dynamic environment change can be sensed in real time, the path is rapidly generated and adjusted, dynamic obstacles are avoided, and meanwhile the global continuity of the planned path is kept.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of unmanned system path planning, and in particular to an unmanned system intelligent path planning system and an operation method in a complex environment. Background Art

[0002] With the widespread application of unmanned systems (such as drones, unmanned vehicles, unmanned ships, etc.) in logistics and transportation, inspection and monitoring, disaster relief and other fields, their autonomous path planning capabilities are particularly important. Especially in fire rescue, drones need to face complex environments such as smoke, high temperature, dynamic fire, etc., while quickly avoiding obstacles and accurately reaching the fire source area. These needs place extremely high demands on the path planning and dynamic adjustment capabilities of drones. However, in the existing technology, most path planning algorithms cannot simultaneously meet the needs of complex environment perception, multi-objective optimization and dynamic adaptation. When operating in complex environments, unmanned systems need to face the following technical difficulties: Environmental complexity: Complex environments usually contain a large number of static obstacles (such as buildings and trees) and dynamic obstacles (such as pedestrians and vehicles), which significantly increases the computational difficulty of path planning.

[0003] Dynamic challenges: The emergence of dynamic obstacles and real-time changes in the environment require the system to have a high degree of path adaptability, and traditional static path planning methods are difficult to meet the requirements.

[0004] Insufficient existing technology: Traditional rule-based algorithms, such as A* and Dijkstra algorithms, can provide the shortest path in a static environment, but lack dynamic adjustment capabilities and cannot adapt to real-time changes in complex environments.

[0005] Random sampling algorithms: such as the rapidly expanding random tree (RRT*) and its improved version RRT**, can quickly generate paths in high-dimensional space, but the path smoothness and dynamic adaptability are insufficient.

[0006] Reinforcement learning algorithms: such as Q-learning, can dynamically adjust paths through learning, but their performance in high-dimensional state spaces is limited.

[0007] Therefore, the existing technology still has great deficiencies in terms of dynamic adaptability of path planning, path optimization efficiency and multi-objective trade-off ability. How to combine the advantages of multiple algorithms to design an intelligent path planning system that can quickly generate paths and dynamically adjust them is the focus and difficulty of current technical research. Summary of the invention

[0008] This paper proposes an intelligent path planning system for unmanned systems in complex environments, which is particularly suitable for path planning and optimization of firefighting drones in fire scenes. By combining the deep reinforcement learning (DQN) algorithm and the rapidly expanding random tree (RRT**) algorithm, the system can dynamically generate paths at the fire scene, avoid high temperature areas and obstacles, and ensure that the firefighting drone reaches the fire source quickly and safely.

[0009] The technical solution of the complex environment path planning system applicable to firefighting drones of the present invention is as follows: A complex environment path planning system suitable for firefighting drones, comprising: The path generation module uses the deep Q learning DQN algorithm to generate paths and adaptively generates the optimal path in a dynamic environment through reinforcement learning; the DQN algorithm realizes path generation through the following Q value update formula: ; in: Represents the current state, including the system's current location and surrounding characteristics; represents the current action, i.e., the path selection of the unmanned system; represents the immediate reward, taking into account the path length, safety and energy consumption; is the learning rate, which controls the Q value update speed; is a discount factor used to measure the impact of future rewards.

[0010] The path optimization module combines the RRT** algorithm to optimize the path generated by the path generation module to achieve a multi-objective balance between path length and energy consumption; the RRT** algorithm generates an initial path and optimizes the path cost through the following formula: ; in: and They represent the adjacent nodes on the path respectively; Represents a slave node arrive The path cost, including path length and energy consumption; is the total optimized cost of the path; The control module is used to dynamically control the movement parameters of the unmanned system, including speed and direction, according to the path generation and optimization results to ensure the safe movement of the system in a complex environment.

[0011] Preferably, the deep Q learning algorithm of the path generation module learns the path generation strategy through the Q value function, and performs adaptive optimization based on the real-time status of the complex environment and the mobility requirements of the system.

[0012] Preferably, the RRT** algorithm of the path optimization module generates an initial path by searching for the nearest node according to the distance formula: ; in: ( )and( ) represent the positions of two nodes in the path respectively; Represents the distance between two nodes and is used to generate a connection path.

[0013] Preferably, the control module dynamically adjusts the motion parameters of the unmanned system according to the output of the path generation and optimization module, so that it can perform path selection and obstacle avoidance in a dynamic environment.

[0014] The operation method of the complex environment path planning system applicable to firefighting drones of the present invention is as follows: comprising the following steps: 1) Determine the starting point ( ), target point( ); 2) Perceiving the environment: Generate environmental maps through sensors to identify obstacles and open areas; 3) Generate a path: Use the Q learning DQN algorithm to generate a path from the starting point ( ) to the target point ( ) to avoid known obstacles as much as possible; 4) Optimize the route: adjust the route through the RRT** algorithm to reduce unnecessary turns or shorten the distance; 5) Dynamic obstacle avoidance: The dynamic obstacle avoidance module detects the location of obstacles based on the sensor data of the unmanned system, and adaptively adjusts the path through the DQN algorithm to avoid dynamic obstacles in real time; 6) Real-time adjustment: Dynamically update the path. During operation, the path generation module continuously receives environmental data feedback and dynamically adjusts the path nodes to adapt to sudden environmental changes, thereby achieving adaptive path optimization of the unmanned system.

[0015] 7) Arrive at the target point: Control the unmanned system to accurately reach the target point and complete the mission.

[0016] Preferably, the path generation step adaptively generates the optimal path through a deep Q learning algorithm, adjusts the generation strategy based on dynamic environmental changes, and the DQN algorithm adaptively adjusts the path through a Q value update function, and the algorithm formula is as follows: ; Preferably, the path optimization step generates a global path and optimizes the total path cost through the RRT** algorithm to achieve a multi-objective balance, and the algorithm formula is as follows: .

[0017] Preferably, the control step dynamically adjusts the speed and direction of the unmanned system according to the optimized path to avoid obstacles and travel along the optimal path.

[0018] Technical Effects Technical Effects of Algorithm Combination The present invention achieves the following technical effects by combining the deep reinforcement learning (DQN) algorithm with the rapidly expanding random tree optimization algorithm (RRT**): 1. Dynamic adaptability of path generation: By learning the optimal relationship between state and action, the DQN algorithm can adaptively generate paths in a dynamic environment, effectively respond to changes in obstacles, and adjust paths in real time.

[0019] The introduction of DQN enhances the system's ability to perceive and respond to dynamic environments, and is particularly suitable for complex environments with multiple obstacles.

[0020] 2. Globality of path optimization: The RRT** algorithm uses an efficient sampling strategy to quickly generate a global path from the starting point to the target point, and optimizes the path nodes during the path generation process to ensure the smoothness and shortest path.

[0021] RRT** provides a global reference path for the system, significantly improving the global performance of path planning.

[0022] 3. Algorithms work together to improve efficiency: RRT** and DQN work together: RRT** is responsible for quickly generating feasible preliminary paths, narrowing the learning and optimization space for DQN and improving the overall efficiency of the system.

[0023] In a dynamic environment, DQN uses real-time environmental feedback to further optimize the path generated by RRT** to ensure the robustness and adaptability of path planning.

[0024] Overall technical effect of the system The present invention has the following overall technical effects in unmanned system path planning: 1. High adaptability to dynamic environment: The system can perceive dynamic environmental changes in real time, quickly generate and adjust paths, avoid dynamic obstacles, and maintain the global coherence of the planned path.

[0025] 2. High optimization of path planning: By combining the RRT** and DQN algorithms, the generated path has shorter path length, lower energy consumption and higher security, which is significantly better than traditional methods.

[0026] 3. Significant improvement in system efficiency: By quickly generating preliminary paths through RRT**, the division of labor and cooperation of DQN dynamic optimization reduces the computational overhead required for path replanning and realizes real-time path planning in complex environments.

[0027] 4. Wide applicability: It is suitable for various unmanned systems such as drones, unmanned vehicles, and unmanned ships, and can meet the needs of multiple scenarios such as logistics transportation, disaster relief inspections, etc.

[0028] 5. Innovation and advancement: The innovative combination of deep learning and random sampling optimization algorithms overcomes the shortcomings of existing technologies in dynamic adaptability and path optimization capabilities, and provides an advanced solution for unmanned system navigation in complex environments. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] Figure 1 It is a flow chart of the operating method of the present invention. DETAILED DESCRIPTION

[0030] Embodiment 1: The complex environment path planning system applicable to a firefighting drone of the present invention comprises: The path generation module uses the deep Q learning DQN algorithm to generate paths and adaptively generates the optimal path in a dynamic environment through reinforcement learning; the DQN algorithm realizes path generation through the following Q value update formula: ;in: Represents the current state, including the system's current location and surrounding characteristics; represents the current action, i.e., the path selection of the unmanned system; represents the immediate reward, taking into account the path length, safety and energy consumption; is the learning rate, which controls the Q value update speed; is a discount factor used to measure the impact of future rewards.

[0031] The path optimization module combines the RRT** algorithm to optimize the path generated by the path generation module to achieve a multi-objective balance between path length and energy consumption; the RRT** algorithm generates an initial path and optimizes the path cost through the following formula: ; in: and They represent the adjacent nodes on the path respectively; Represents a slave node arrive The path cost, including path length and energy consumption; is the total optimized cost of the path; The control module is used to dynamically control the movement parameters of the unmanned system, including speed and direction, according to the path generation and optimization results to ensure the safe movement of the system in a complex environment.

[0032] Embodiment 2: The difference between this embodiment and embodiment 1 is that this embodiment further includes: The deep Q learning algorithm of the path generation module learns the path generation strategy through the Q-value function, and performs adaptive optimization based on the real-time status of the complex environment and the mobility requirements of the system.

[0033] Preferably, the RRT** algorithm of the path optimization module generates an initial path by searching for the nearest node according to the distance formula: ; in: ( )and( ) represent the positions of two nodes in the path respectively; Represents the distance between two nodes and is used to generate a connection path.

[0034] The control module dynamically adjusts the motion parameters of the unmanned system according to the output of the path generation and optimization module, so that it can perform path selection and obstacle avoidance in a dynamic environment.

[0035] Embodiment 3: The operation method of the complex environment path planning system applicable to the fire-fighting drone of the present invention is as follows: comprising the following steps: 1) Determine the starting point ( ), target point( ); 2) Perceiving the environment: Generate environmental maps through sensors to identify obstacles and open areas; 3) Generate a path: Use the Q learning DQN algorithm to generate a path from the starting point ( ) to the target point ( ) to avoid known obstacles as much as possible; 4) Optimize the path: adjust the path through the RRT** algorithm to reduce unnecessary turns or shorten the distance; 5) Dynamic obstacle avoidance: The dynamic obstacle avoidance module detects the location of obstacles based on the sensor data of the unmanned system, and adaptively adjusts the path through the DQN algorithm to avoid dynamic obstacles in real time; 6) Real-time adjustment: Dynamically update the path. During operation, the path generation module continuously receives environmental data feedback and dynamically adjusts the path nodes to adapt to sudden environmental changes, thereby achieving adaptive path optimization of the unmanned system.

[0036] 7) Arrive at the target point: Control the unmanned system to accurately reach the target point and complete the mission.

[0037] Embodiment 4: The difference between this embodiment and embodiment 3 is that this embodiment further includes: The path generation step adaptively generates the optimal path through the deep Q learning algorithm, adjusts the generation strategy based on dynamic environmental changes, and the DQN algorithm adaptively adjusts the path through the Q value update function. The algorithm formula is as follows: ; The path optimization step generates a global path and optimizes the total path cost through the RRT** algorithm to achieve a multi-objective balance. The algorithm formula is as follows: .

[0038] The control step dynamically adjusts the speed and direction of the unmanned system according to the optimized path to avoid obstacles and travel along the optimal path.

[0039] Embodiment 5: Path planning and dynamic obstacle avoidance application of firefighting drones at fire scenes Scenario description In a forest fire rescue, the firefighting drone needs to fly from the starting point (0, 0, 0) to the fire source location (50,80, 20). The high temperature range of the fire source area is [(45, 75), (50, 85)], the obstacle distribution on site is complex, and there are dynamic obstacles (such as fallen trees, the position is (48, 78, 18), the speed is (0.5, 0.2)). The drone is equipped with thermal imaging sensors, lidar and smoke detection sensors for real-time perception of the environment.

[0040] Detailed process

[0041] Determine the starting and destination points: The starting point is (0, 0, 0), and the target point is (50, 80, 20).

[0042] Perceiving the Environment Static obstacle point set acquired by LiDAR: ={(10,15,5),(25,45,10),(40,70,15)}.

[0043] Dynamic obstacle points: ={(48,78,18)}.

[0044] High temperature areas detected by thermal imaging: =[(45,75),(50,85)].

[0045] Path Generation: Deep Reinforcement Learning (DQN) Algorithm Define the Q value update formula: ; in: state : Current location and environmental characteristics (fire source location, high temperature area, obstacle distribution).

[0046] action : Current path selection (next moving direction).

[0047] Instant Rewards : Comprehensively consider the path length (penalty coefficient is -1), avoid obstacles and high temperature areas (reward is +5), and get close to the fire source (reward is +10).

[0048] Preliminary path calculation: starting point =(0,0,0), target point =(50,80,20).

[0049] Action selection: The drone moves to (20, 30, 10) (close to the fire source, avoiding obstacles).

[0050] Finally, the preliminary path is generated: =[(0,0,0),(20,30,10),(40,70,15),(50,80,20)].

[0051] Path optimization: RRT** algorithm Distance formula: distance between two nodes ; Optimize Path: Adjust nodes based on preliminary path and obstacle distribution.

[0052] Optimization results: =[(0,0,0),(22,32,12),(45,75,18),(50,80,20)].

[0053] Path length calculation: =d(0,22)+d(22,45)+d(45,50).

[0054] Calculation results: path length before optimization: 101 meters; path length after optimization: 95.4 meters.

[0055] Dynamic obstacle avoidance and real-time adjustment Dynamic obstacle prediction: obstacle (48, 78, 18) moves at a speed of (0.5, 0.2), predicted position: =(48+0.5t,78+0.2t,18).

[0056] The drone uses sensors to detect obstacle locations in real time and adjusts its path: =[(0,0,0),(22,32,12),(44,74,17),(50,80,20)].

[0057] Data Results and Performance Path optimization results: Optimized path length: 95.4 meters (5.6% less than before optimization).

[0058] Estimated reduction in energy consumption: 8.3%.

[0059] Dynamic obstacle avoidance performance: Average obstacle avoidance reaction time: 0.15 seconds.

[0060] Dynamic obstacle avoidance success rate: 98.6%.

[0061] Path planning time: Initial path generation time: 1.2 seconds.

[0062] Path optimization time: 0.8 seconds.

[0063] Fire source location success rate: 96.2%.

[0064] Technical Effects This embodiment demonstrates the significant advantages of the present invention in firefighting scenarios through a detailed calculation process: 1. Efficient path optimization: Combining deep reinforcement learning and random sampling algorithms to significantly shorten path length and optimize energy consumption.

[0065] 2. Real-time dynamic obstacle avoidance: Effectively avoid moving obstacles through adaptive adjustment of the dynamic obstacle avoidance module.

[0066] 3. Strong adaptability to fire scenarios: Combined with fire source location and high temperature area avoidance, efficient and accurate firefighting task planning can be achieved.

Claims

1. A complex environment path planning system suitable for firefighting drones, characterized in that: It includes: The path generation module uses the deep Q learning DQN algorithm to generate paths and adaptively generates the optimal path in a dynamic environment through reinforcement learning; the DQN algorithm realizes path generation through the following Q value update formula: ; in: ) indicates that in the state Take action The value function (Q value) is used to evaluate the quality of the current path selection; Indicates the current state, including the current location of the system and the characteristics of the surrounding environment; represents the current action, i.e., the path selection of the unmanned system; It represents the immediate reward, taking into account the path length, safety and energy consumption. The reward function is designed according to the actual situation. Represents the learning rate, which is used to control the speed of Q value update, and its value range is 0< ≤1; Represents the discount factor, which is used to measure the impact of future rewards. The value range is 0≤ <1,; is in the next state The best action among all possible actions. Indicates the next state Taking the best action The corresponding Q value is used to guide the path selection of the current state; The path optimization module combines the RRT** algorithm to optimize the path generated by the path generation module to achieve a multi-objective balance between path length and energy consumption; the RRT** algorithm generates an initial path and optimizes the path cost through the following formula: ; in: and They represent the adjacent nodes on the path respectively; Represents a slave node arrive The path cost, including path length and energy consumption; is the total optimized cost of the path; The control module is used to dynamically control the movement parameters of the unmanned system, including speed and direction, according to the path generation and optimization results to ensure the safe movement of the system in a complex environment.

2. The complex environment path planning system suitable for firefighting drones according to claim 1 is characterized in that: The deep Q learning algorithm of the path generation module learns the path generation strategy through the Q-value function, and performs adaptive optimization based on the real-time status of the complex environment and the mobility requirements of the system.

3. The complex environment path planning system suitable for firefighting drones according to claim 1 is characterized in that: The RRT** algorithm of the path optimization module generates an initial path by finding the nearest node according to the distance formula: ; in: ( )and( ) represent the positions of two nodes in the path respectively; Represents the distance between two nodes and is used to generate a connection path.

4. The complex environment path planning system suitable for firefighting drones according to claim 1 is characterized in that: The control module dynamically adjusts the motion parameters of the unmanned system according to the output of the path generation and optimization module, so that it can perform path selection and obstacle avoidance in a dynamic environment.

5. An operating method of a complex environment path planning system applicable to a firefighting drone according to any one of claims 1 to 4, characterized in that: The following steps are involved: 1) Determine the starting point ( ), target point( ); 2) Perceiving the environment: Generate environmental maps through sensors to identify obstacles and open areas; 3) Generate a path: Use the Q learning DQN algorithm to generate a path from the starting point ( ) to the target point ( ) to avoid known obstacles as much as possible; 4) Optimize the path: adjust the path through the RRT** algorithm to reduce unnecessary turns or shorten the distance; 5) Dynamic obstacle avoidance; 6) Real-time adjustment: dynamically update the path; 7) Arrive at the target point: Control the unmanned system to accurately reach the target point and complete the mission.

6. The operating method of the complex environment path planning system applicable to firefighting drones according to claim 5 is characterized in that: The path generation step adaptively generates the optimal path through the deep Q learning algorithm, adjusts the generation strategy based on dynamic environmental changes, and the DQN algorithm adaptively adjusts the path through the Q value update function. The algorithm formula is as follows: 。 7. The operating method of the complex environment path planning system applicable to firefighting drones according to claim 5 is characterized in that: The path optimization step generates a global path and optimizes the total path cost through the RRT** algorithm to achieve a multi-objective balance. The algorithm formula is as follows: 。 8. The operating method of the complex environment path planning system applicable to firefighting drones according to claim 5 is characterized in that: The steps described herein dynamically adjust the speed and direction of the unmanned system according to the optimized path to avoid obstacles and travel along the optimal path.