Hydraulic tunnel surveying and mapping unmanned aerial vehicle path planning method based on learning strategy

Through the improved Q-learning algorithm and raster map combined with heuristic reward function and B-spline function processing, the discontinuity and safety problems of drones in hydraulic tunnels are solved, and efficient and safe surveying and mapping tasks are achieved.

CN120385349AActive Publication Date: 2025-07-29ZHEJIANG HUADONG SURVEYING MAPPING & GEOINFORMATION
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510866172.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-29
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

The existing drone path planning algorithms have problems such as large calculation volume, poor adaptability of multiple evaluation indicators, difficulty in convergence, and discontinuous planning paths lead to safety hazards and low information collection accuracy in hydraulic tunnels.

Method used

The improved Q-learning algorithm is used to combine raster maps and heuristic reward functions, and path planning is performed through adaptive ε-greedy strategy and half-gradient time difference learning algorithm with qualification traces, and smoothing is performed using quasi-uniform B-spline function.

Benefits of technology

It improves the flight safety and mapping efficiency of drones in hydraulic tunnels, ensures the continuity and controllability of paths, and is suitable for tunnel environments with long distances and complex structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120385349A_ABST
    Figure CN120385349A_ABST
Patent Text Reader

Abstract

The invention relates to a hydraulic tunnel surveying and mapping unmanned aerial vehicle path planning method based on a learning strategy, and is suitable for the technical field of surveying and mapping geographic information acquisition. The method comprises the steps that a grid map is generated based on the environment of a target hydraulic tunnel, and the side length of a grid is determined based on the minimum safety distance between an unmanned aerial vehicle and an obstacle; regarding each passable grid in the grid map as a state, and constructing an action set based on the movement direction of the unmanned aerial vehicle moving from the grid center point to the adjacent grid center point in the grid map; based on the starting point and the ending point of the to-be-planned path, combining with the state and action set in the grid map, and adopting an improved Q-learning algorithm to carry out path planning to obtain a first path scheme; smoothing the first path scheme by using a quasi-uniform B-spline function to obtain a second path scheme; and the improved Q-learning algorithm adopts a heuristic reward function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a path planning method for a surveying and mapping UAV of a hydraulic tunnel based on a learning strategy, and is applicable to the technical field of surveying and mapping geographic information acquisition. Background Art

[0002] Hydraulic tunnels are widely used in scenarios such as water diversion, flood discharge, and drainage in hydropower projects. They have typical characteristics such as long and narrow spaces, insufficient lighting, high humidity, poor ventilation, and complex structures. Especially in inclined shaft and vertical shaft sections, there are often large inclinations and lengths, which pose great challenges to manual inspection and precise surveying and mapping. In such environments, it is difficult for staff to enter for a long time or even impossible to reach at all. Traditional surveying and mapping methods not only have low efficiency and limited accuracy, but also have obvious safety hazards.

[0003] The path planning and obstacle avoidance technology of UAVs is an important branch in the field of automation and control, and has been widely applied in industrial production and daily life. After being equipped with task-driven surveying and mapping instruments, UAVs can replace humans to perform information collection and processing tasks in complex or even harsh environments. In recent years, with the wide application of UAV platforms in the field of industrial inspection, it has become possible to carry out surveying and mapping of hydraulic tunnels by using sensors such as lidar and cameras with their flexible and efficient flight capabilities. However, as a closed indoor space, a hydraulic tunnel cannot receive GNSS signals, resulting in the UAV being unable to rely on a conventional navigation system for positioning and path control. In addition, there are often dense obstacles such as pipeline supports, temporary facilities, and variable cross-section structures inside the tunnel, which pose a threat to the flight safety of the UAV. Therefore, in such complex environments, achieving the autonomous path planning and obstacle avoidance flight capabilities of the UAV has become the key to ensuring the safe, efficient, and accurate completion of the tunnel surveying and mapping tasks.

[0004] Currently, many mature algorithms are used for the path planning of unmanned aerial vehicles (UAVs), or are improved to adapt to complex working spaces. These methods include, but are not limited to: algorithms based on search and sampling (ABSS), algorithms based on meta-heuristic intelligence (ABMHI), algorithms based on Q-learning (ABQL), and so on. Summarizing the literature and various path planning algorithms, it is found that compared with ABSS and ABMHI, reinforcement learning (RL) algorithms represented by Q-learning have powerful advantages and are a possible development direction. Whether it is classical Q-learning (CQL) or various improved Q-learning algorithms, they are all designed to solve the defects such as large computational complexity, poor adaptability to multiple evaluation indicators, and difficult convergence in path planning. However, many current algorithms based on Q-learning still plan paths that are only the shortest in distance while ignoring the energy consumption during the movement of the UAV. In addition, although the optimal path is obtained, it is composed of discontinuous line segments and lacks smoothness, resulting in a decline in the turning performance of the UAV. The existence of these defects leads to potential safety hazards when the surveying and mapping UAV flies in the hydraulic tunnel, and at the same time affects the acquisition accuracy and efficiency of environmental information. Summary of the Invention

[0005] The technical problem to be solved by the present invention is: in view of the above problems, to provide a path planning method for a hydraulic tunnel surveying and mapping UAV based on a learning strategy.

[0006] The technical solution adopted by the present invention is: a path planning method for a hydraulic tunnel surveying and mapping UAV based on a learning strategy, including: Generating a grid map based on the environment of the target hydraulic tunnel, where the side length of the grid in the grid map is determined based on the minimum safety distance between the UAV and the obstacle; Regarding each passable grid in the grid map as a state, and constructing an action set based on the movement direction of the UAV from the center point of the grid to the center point of the adjacent grid in the grid map; Based on the starting and ending points of the path to be planned, combining the states and action sets in the grid map, using an improved Q-learning algorithm for path planning to obtain a first path plan; Using a quasi-uniform B-spline function to smooth the first path plan to obtain a second path plan; The improved Q-learning algorithm uses a heuristic reward function, which is constructed based on a basic reward, a path length reward, and a path smoothness reward.

[0007] The side length of the grid in the grid map is determined based on the minimum safety distance between the UAV and the obstacle, including: Assume that the minimum safety distance between the UAV and the obstacle is a / 2, then the side length of the grid in the grid map is set to a.

[0008] The heuristic reward function is constructed based on a basic reward, a path length reward, and a path smoothness reward, including: ; Wherein, is the heuristic reward function, is the basic reward, is the path length reward, is the path smoothness reward, α1 and α2 are preset reward factors, 0 < α1, α2 < 1, and t is the iteration round

[0009] The basic reward is determined based on the policy function and value function of Q learning, and is designed based on the current coordinates and attitude parameters of the UAV; the path length reward is determined based on path safety and introduces a weighted obstacle avoidance cost parameter; the path smoothness reward is determined based on the angle change during UAV obstacle avoidance and incorporates the change in the path curvature radius.

[0010] The improved Q-learning algorithm uses an adaptive ε-greedy strategy as the action selection strategy, and the balance factor ε in this adaptive ε-greedy strategy increases with the increase of the iteration round.

[0011] The balance factor ε increases with the increase of the iteration round, including: ; Wherein, ε0, ε k and ε f represent the initial, k-th round, and final balance factors respectively, μ1 and μ2 are preset scale factors, M represents the total number of training rounds, 0 < ε0, ε f , μ1, μ2 < 1.

[0012] The improved Q-learning algorithm uses a semi-gradient temporal difference learning algorithm with eligibility traces to estimate the Q-value function.

[0013] The use of a semi-gradient temporal difference learning algorithm with eligibility traces to estimate the Q-value function includes: The Q function is decomposed into a state-value function and an action-value function for estimation; When estimating the Q function using the semi-gradient temporal-difference learning algorithm with eligibility traces, first define the eligibility trace z(t); Secondly, the state-value function adds a weight update mechanism to estimate the sum of rewards related to subsequent states; finally, the update of the action-value function is based on the optimized Bellman equation as shown below to implement the iteration of Q-values: ; Among them, is the heuristic reward function, represents the mean estimate, α and β are weight coefficients, represents at the current state under the state-value function with adjustable weight , and this policy represents the adaptive ε-greedy policy; is the eligibility trace at the current time t.

[0014] A path planning device for a water conveyance tunnel mapping unmanned aerial vehicle based on a learning strategy, comprising: A grid generation module, configured to generate a grid map based on the environment of the target water conveyance tunnel, and the side length of the grid in the grid map is determined based on the minimum safe distance between the unmanned aerial vehicle and the obstacle; A state-action definition module, configured to regard each passable grid in the grid map as a state, and construct an action set based on the movement direction of the unmanned aerial vehicle from the center point of the grid to the center point of the adjacent grid in the grid map; A first path module, configured to perform path planning using an improved Q-learning algorithm based on the start and end points of the path to be planned, combined with the states and action sets in the grid map, to obtain a first path plan; A second path module, configured to smooth the first path plan using a quasi-uniform B-spline function to obtain a second path plan; The improved Q-learning algorithm adopts a heuristic reward function, and the heuristic reward function is constructed based on a basic reward, a path length reward, and a path smoothness reward.

[0015] A storage medium, on which a computer program executable by a processor is stored, and when the computer program is executed, the steps of the path planning method for a water conveyance tunnel mapping unmanned aerial vehicle based on a learning strategy are implemented.

[0016] A path planning device, having a memory and a processor, and a computer program executable by the processor is stored on the memory, and when the computer program is executed, the steps of the path planning method for a water conveyance tunnel mapping unmanned aerial vehicle based on a learning strategy are implemented.

[0017] The beneficial effects of the present invention are: (1) In view of the characteristics of narrow internal channels, tortuous structures, and dense obstacles in hydraulic tunnels, the present invention adopts an improved Q-learning algorithm for path planning, enabling the mapping UAV to flexibly avoid obstacles such as walls, support structures, and equipment, effectively enhancing flight safety. (2) Through the optimization of the reinforcement learning strategy and the design of the reward function, the planned path not only meets the obstacle avoidance requirements but also takes into account the path length and curvature smoothness, ensuring that the UAV can move quickly and accurately in the tunnel, improving the mapping operation efficiency. (3) The present invention uses a quasi-uniform B-spline smoothing algorithm to post-process the path, greatly improving the continuity and controllability of the path, reducing the risk of collision or increased energy consumption due to difficult turning at complex corners of the UAV, and is particularly suitable for long-distance and fully enclosed hydraulic tunnel inspection tasks. (4) The path planning method proposed by the present invention shows good convergence speed and algorithm stability in multiple groups of simulation and physical experiments, verifying its practical value and promotion potential for autonomous navigation control of mapping UAVs in the hydraulic tunnel scenario. Description of the Drawings

[0018] Figure 1 It is a flowchart of the embodiment.

[0019] Figure 2 It is an effect diagram in a two-dimensional grid map of the embodiment. The red solid line represents the path planned by IQLA; the dotted lines of the remaining colors respectively represent the paths planned by common intelligent evolutionary algorithms, namely the wolf pack algorithm (blue), the sparrow search algorithm (yellow), and the particle swarm optimization algorithm (purple).

[0020] Figure 3 It is an effect diagram in a three-dimensional grid map, where (a) is the 3D simulation result and (b) is the horizontal sectional view. Detailed Embodiment

[0021] The embodiments of the present invention are described in detail below. The examples of the embodiments are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The embodiments described below with reference to the drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation of the present invention. For the step numbers in the following embodiments, they are only set for the convenience of description and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0022] In the description of the present invention, "a plurality of" means two or more. If the first and second are described, it is only for the purpose of distinguishing technical features and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art of this technology.

[0023] Embodiment 1: This embodiment is a method for path planning of a surveying and mapping unmanned aerial vehicle (UAV) for a hydraulic tunnel based on a learning strategy, specifically including the following steps: S100. Generate a grid map based on the environment of the target hydraulic tunnel. The side length of the grid in the grid map is determined based on the minimum safe distance between the UAV and the obstacle.

[0024] In view of the environmental characteristics of hydraulic tunnels, such as narrow channels, frequent cross-section changes, and diverse types of obstacles (such as pipelines, brackets, accumulated water, and equipment residues), this embodiment uses the two-dimensional grid method to model the working area of the tunnel. The internal space of the tunnel is discretized into N x ×N y grid cells, and a mapping relationship between the grid index and the physical coordinates is established through a numbering system. This modeling method facilitates the identification and shielding of obstacle areas and provides a constraint boundary for subsequent path search.

[0025] To avoid collisions caused by approaching obstacles in space, in this embodiment, the side length of the grid in the grid map is determined based on the minimum safe distance between the UAV and the obstacle. The minimum safe distance is set to a / 2, and the side length of the grid in the grid map is set to a.

[0026] S200. Consider each passable grid in the grid map as a state, and construct an action set based on the movement direction of the UAV from the center point of the grid to the center point of the adjacent grid in the grid map.

[0027] Reinforcement learning learns knowledge in the interaction between the agent and the environment, selects actions based on the evaluation feedback of the environment, and iterates the state. In view of the characteristics of tortuous paths, multiple obstacles, and high accuracy requirements in the surveying and mapping of hydraulic tunnels, this embodiment defines each accessible grid as a state S(t) that can be explored by the surveying and mapping UAV. Combining the flight physical constraints and safety distance requirements of the surveying and mapping UAV, an 8-direction movement method is used to design the action space, enabling the UAV to have a flexible turning response ability to adapt to the actual needs of frequent turning and obstacle avoidance in the tunnel. In the configured working area designed in this embodiment, the surveying and mapping UAV can move in any of the 8 directions (front, back, left, right, right front, right rear, left front, left rear) to the next feasible grid.

[0028] S300. Based on the start and end points of the path to be planned, combined with the state and action sets in the grid map, the improved Q-learning algorithm is used for path planning to obtain the first path plan.

[0029] In this embodiment, in the process of path learning and updating, the improved Q-learning algorithm adopts the following strategies to improve the path quality and algorithm convergence: (1) Action selection strategy: A reasonable action selection strategy can achieve the balance between exploration and exploitation of learning algorithms. The widely used ε-greedy strategy is prone to cause confusion between exploration and exploitation, making it difficult for the algorithm to converge, resulting in the planned path being only a local optimum.

[0030] To overcome the defects of the ε-greedy strategy, this embodiment adopts the adaptive ε-greedy policy as the action selection strategy. In this adaptive ε-greedy strategy, the balance factor ε increases with the increase of the iteration rounds, so as to encourage extensive exploration of the entire area of the hydraulic tunnel in the initial stage and gradually converge to the optimal path decision in the later stage of training, taking into account both globality and convergence speed, and effectively avoiding the exploration obstacles brought by the complex internal structure of the tunnel.

[0031] In this embodiment, the balance factor ε increases with the increase of the iteration rounds, including: ; where ε 0, ε k and ε f represent the initial, k-th round and final balance factors respectively (ε f ≈ 0), μ1 and μ2 are preset proportionality factors, M represents the total number of training rounds; 0 < ε0, ε f , μ1, μ2 < 1; ε k affects the final action selection strategy.

[0032] In this embodiment, the adaptive ε-greedy strategy is used to realize the strategy evolution of "extensive exploration in the initial stage + approaching the optimum in the later stage", enabling the mapping UAV to quickly identify the feasible area in the complex tunnel topology and gradually converge to the shortest path plan.

[0033] (2) Heuristic reward function:

[0034] The reward function in the Q-learning algorithm is used to determine the value of each action, and is the key to reducing ineffective search and improving the convergence speed. In CQL, the reward function is defined as positive when reaching the target, negative when colliding with obstacles, and zero in other states. Using a constant form of the reward function leads to an increase in the number of blind steps in the UAV search and low search efficiency.

[0035] Based on the comprehensive requirements of the shortest path length, smooth turning, and low flight energy consumption for the task of hydraulic tunnels, in this embodiment, a heuristic reward function R is constructed based on the basic reward, path length reward, and path smoothness reward. heu (t). Compared with the constant reward design in CQL, the improved Q-learning algorithm can effectively suppress the search for invalid paths in a complex obstacle-dense environment, accelerate the convergence speed, improve the path quality, and better meet the dual goals of "both accuracy and safety" in actual surveying and mapping operations.

[0036] In this embodiment, the basic reward is determined based on the policy function and value function of Q-learning and designed based on the current coordinates and attitude parameters of the UAV; the path length reward is determined based on path safety and introduces a weighted obstacle avoidance cost parameter; the path smoothness reward is determined based on the angle change during UAV obstacle avoidance and incorporates the change in the path curvature radius.

[0037] In this embodiment, the heuristic reward function is constructed based on the basic reward, path length reward, and path smoothness reward, including: ; where is the heuristic reward function, is the basic reward, is the path length reward, is the path smoothness reward, α1 and α2 are preset reward factors, 0 < α1, α2 < 1, and t is the iteration round.

[0038] In this example, a reward function model combining path length and turning smoothness is constructed to guide the UAV to prefer decisions with shorter flight paths and gentler turning angles, thus meeting the requirements of low energy consumption, stable positioning, and continuous image acquisition in tunnel surveying.

[0039] (3) Q-value update mechanism: In CQL, when using temporal difference learning to iteratively approximate the Q function, there are defects such as unstable processes and large approximation errors.

[0040] Considering the problems such as the large state space dimension and the easy fluctuation of the search process caused by the complex structure of hydraulic tunnels, in this embodiment, the semi-gradient temporal-difference (SGTD) learning algorithm with eligibility traces is used to replace the traditional temporal difference method. This method dynamically evaluates the rewards of subsequent states, significantly improves the convergence accuracy and training stability of the Q function, and ensures that the UAV makes more reliable path choices in variable-section tunnels.

[0041] In this embodiment, the Q function is decomposed into a state-value function and an action-value function for estimation. When using SGTD to estimate the Q function, first, the eligibility trace z(t) is defined; second, the state-value function is added with a weight update mechanism for estimating the sum of rewards related to subsequent states; finally, the update of the action-value function is based on the optimized Bellman equation shown below to implement the iteration of Q-values: ; where is the heuristic reward function, represents the mean estimation, and α, β are weight coefficients, denotes the state-value function with adjustable weight under the current state , and this policy represents the adaptive ε-greedy policy; is the eligibility trace at the current time t.

[0042] This embodiment introduces semi-gradient temporal difference learning with eligibility traces to estimate the Q-value function, overcomes the problem of unstable training of traditional methods when the state space is large, and improves the robustness of the algorithm under long tunnels and multi-branch structures.

[0043] S400. Smooth the first path plan using the quasi-uniform B-spline function to obtain the second path plan.

[0044] To further improve the turning fluency and flight stability of the mapping UAV in the narrow space of the tunnel, this embodiment performs quasi-uniform B-spline (QUBS) optimization on the first path plan.

[0045] In this embodiment, the control points are decoupled by the quasi-uniform B-spline function. By introducing a control node vector knots with higher degrees of freedom, the continuity and smoothness of the path are significantly improved while maintaining the validity of the path, reducing the flight deviation or instability caused by excessive path turning angles, and is particularly suitable for long-distance and multi-curved tunnel space scenarios.

[0046] This embodiment uses the quasi-uniform B-spline function (QUBS) to smooth the original planned path, enhancing the continuity and executability of the path, and is particularly suitable for operation scenarios with multiple sharp turns or asymmetric structures in the tunnel.

[0047] Store the second path plan generated in this embodiment in the navigation module, and perform trajectory tracking control in combination with the on-board positioning system of the mapping UAV (such as IMU, laser SLAM, etc.) to achieve efficient mapping inside the tunnel in the autonomous flight state. Experiments show that the planned path can not only effectively avoid structural obstacles, but also has excellent geometric characteristics and flight stability, and is suitable for typical operation scenarios of hydraulic tunnels such as long distances, weak light, and GNSS signal shielding.

[0048] Embodiment 2: This embodiment is a path planning device for a mapping UAV for a hydraulic tunnel based on a learning strategy, including: A grid generation module for generating a grid map based on the environment of the target hydraulic tunnel, where the side length of the grid in the grid map is determined based on the minimum safe distance between the UAV and the obstacle; A state-action definition module for regarding each passable grid in the grid map as a state and constructing an action set based on the movement direction of the UAV from the center point of the grid to the center point of the adjacent grid in the grid map; A first path module for performing path planning using an improved Q-learning algorithm based on the start and end points of the path to be planned, in combination with the states and action sets in the grid map, to obtain a first path plan; A second path module for smoothing the first path plan using a quasi-uniform B-spline function to obtain a second path plan; The improved Q-learning algorithm uses a heuristic reward function, which is constructed based on a basic reward, a path length reward, and a path smoothness reward.

[0049] Embodiment 3: This embodiment is a storage medium on which a computer program executable by a processor is stored, and when the computer program is executed, the steps of the path planning method for a mapping UAV for a hydraulic tunnel based on a learning strategy described in Embodiment 1 are implemented.

[0050] Embodiment 4: This embodiment is a path planning device having a memory and a processor, and a computer program executable by the processor is stored on the memory, and when the computer program is executed, the steps of the path planning method for a mapping UAV for a hydraulic tunnel based on a learning strategy described in Embodiment 1 are implemented.

[0051] In addition, although the present invention has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the above-described functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present invention. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Thus, those skilled in the art can implement the present invention as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.

[0052] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the above methods in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0053] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a predefined sequence of executable instructions for implementing logical functions, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0054] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which the above programs can be printed, since the above programs can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation or, if necessary, other suitable processing, and then stored in a computer memory.

[0055] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc.

[0056] In the above description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example" or "certain embodiments / examples", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0057] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the claims and their equivalents.

[0058] The above has specifically described the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A path planning method for a surveying and mapping UAV of a hydraulic tunnel based on a learning strategy, characterized in that, Including: Generate a grid map based on the environment of the target hydraulic tunnel, where the side length of the grid in the grid map is determined based on the minimum safe distance between the drone and the obstacle; Regard each passable grid in the grid map as a state, and construct an action set based on the movement direction of the drone from the center point of the grid to the center point of the adjacent grid in the grid map; Based on the start and end points of the path to be planned, combine the states and action sets in the grid map, and use an improved Q-learning algorithm for path planning to obtain the first path plan; Use the quasi-uniform B-spline function to smooth the first path plan to obtain the second path plan; The improved Q-learning algorithm uses a heuristic reward function, which is constructed based on the basic reward, path length reward, and path smoothness reward.

2. The method for path planning of a surveying and mapping unmanned aerial vehicle for hydraulic tunnels based on a learning strategy according to claim 1, wherein The side length of the grid in the grid map is determined based on the minimum safe distance between the drone and the obstacle, including: Assume that the minimum safe distance between the drone and the obstacle is a / 2, then the side length of the grid in the grid map is a.

3. The method for path planning of a surveying and mapping unmanned aerial vehicle for a hydraulic tunnel based on a learning strategy according to claim 1, wherein The heuristic reward function is constructed based on the basic reward, path length reward, and path smoothness reward, including: ; Among them, is a heuristic reward function, is the basic reward, is the path length reward, is the path smoothness reward, α1 and α2 are preset reward factors, 0 < α1, α2 < 1, and t is the number of iteration rounds.

4. The method for path planning of a surveying and mapping UAV for hydraulic tunnels based on a learning strategy according to claim 1, wherein The basic reward is determined according to the policy function and value function of Q learning, and is designed based on the current coordinates and attitude parameters of the drone; the path length reward is determined based on path safety and introduces a weighted obstacle avoidance cost parameter; the path smoothness reward is determined according to the angle change when the drone avoids obstacles and integrates the change of the path curvature radius.

5. The method for path planning of a surveying and mapping UAV for hydraulic tunnels based on a learning strategy according to claim 1, wherein, The improved Q-learning algorithm uses an adaptive ε-greedy strategy as the action selection strategy, and the balance factor ε in this adaptive ε-greedy strategy increases with the increase of the iteration rounds.

6. The method for path planning of a UAV for surveying and mapping of a hydraulic tunnel based on a learning strategy according to claim 5, wherein, The balance factor ε increases with the increase of the iteration rounds, including: ; Among them, ε0, ε k and ε f respectively represent the initial, the k-th round and the final balance factors, μ1 and μ2 are preset proportionality factors, M represents the total number of training rounds; 0 < ε0, ε f , μ1, μ2 < 1.

7. The method for path planning of a surveying and mapping UAV for hydraulic tunnels based on a learning strategy according to claim 1, characterized in that The improved Q-learning algorithm uses a semi-gradient temporal difference learning algorithm with eligibility traces to estimate the Q value function.

8. The method for path planning of a surveying and mapping unmanned aerial vehicle for a hydraulic tunnel based on a learning strategy according to claim 7, wherein The use of a semi-gradient temporal difference learning algorithm with eligibility traces to estimate the Q value function includes: The Q function is decomposed into a state-value function and an action-value function for estimation; When using a semi-gradient temporal difference learning algorithm with eligibility traces to estimate the Q function, first define the eligibility trace z(t); Secondly, the state-value function adds a weight update mechanism to estimate the total reward related to subsequent states; finally, the update of the action-value function is based on the following optimized Bellman equation to implement the iteration of Q-values: ; wherein, is a heuristic reward function, represents the mean estimation, and α, β are weight coefficients, represents at the current state under which the state value function with adjustable weights is adopted, and this policy represents an adaptive ε-greedy policy; is the eligibility trace at the current time t.

9. An unmanned aerial vehicle path planning device for hydraulic tunnel surveying and mapping based on a learning strategy, characterized in that, Including: A grid generation module for generating a grid map based on the environment of the target hydraulic tunnel, where the side length of the grid in the grid map is determined based on the minimum safe distance between the drone and the obstacle; A state-action definition module for regarding each passable grid in the grid map as a state and constructing an action set based on the movement direction of the drone from the center point of the grid to the center point of the adjacent grid in the grid map; A first path module for performing path planning using an improved Q-learning algorithm based on the start and end points of the path to be planned, combining the states and action sets in the grid map, and obtaining the first path plan; A second path module, configured to smooth the first path scheme by using a quasi-uniform B-spline function to obtain a second path scheme; The improved Q-learning algorithm adopts a heuristic reward function, and the heuristic reward function is constructed based on a basic reward, a path length reward, and a path smoothness reward.

10. A storage medium having stored thereon a computer program executable by a processor, characterized in that, When the computer program is executed, the steps of the method for path planning of a hydraulic tunnel surveying and mapping unmanned aerial vehicle based on a learning strategy according to any one of claims 1 to 8 are implemented.

11. A path planning device, having a memory and a processor, wherein a computer program capable of being executed by the processor is stored on the memory, characterized in that When the computer program is executed, the steps of the method for path planning of a hydraulic tunnel surveying and mapping unmanned aerial vehicle based on a learning strategy according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Mobile robot local path planning method based on value distribution deep reinforcement learning

    CN117470244A

  • AGV path planning method based on improved Q learning

    CN118293940A

  • Intention-driven reinforcement learning-based path planning method

    US20240219923A1