UAV Path Planning Method, Device and Storage Medium for Steel Bar Binding Site

By constructing a Markov-based path planning environment and dual strategy learning network model, the path planning of the drone at the steel bar binding site is optimized, and the problem of not the shortest path length and the planning time is solved, and the path planning is achieved with fast and intelligent obstacle avoidance.

CN119151110BActive Publication Date: 2025-07-29ZHONGSHAN POWER SUPPLY BUREAU OF GUANGDONG POWER GRID
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411622214.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-14
Publication Date
2025-07-29
Estimated Expiration
2044-11-14

AI Technical Summary

Technical Problem

The existing UAV path planning method is not the shortest and optimal path length in the steel bar binding site, and the planning time is relatively long.

Method used

A drone path planning environment with Markov properties is built, combined with obstacle evaluation model, drone position model and reward function model, a dual strategy learning network model is used for path planning, and a drone action strategy is trained through historical data to optimize path selection.

Benefits of technology

It realizes that the drone plans the optimal flight path in a shorter time during the steel bar binding site, intelligently avoids obstacles, and is suitable for real-time mission needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119151110B_ABST
    Figure CN119151110B_ABST
Patent Text Reader

Abstract

The present application provides a method, device and storage medium for UAV path planning at a steel bar binding site. The method includes: constructing a UAV path planning environment at the steel bar binding site with Markov properties, obtaining a historical initial position and a historical termination position, inputting the historical initial position, the historical termination position and the UAV path planning environment model into an action selection strategy model to obtain a historical UAV action strategy, using the historical UAV action strategy to control the movement of the UAV and obtaining historical real-time operation data during the movement, determining the historical real-time operation data as sample data, obtaining a dual policy learning network model, and inputting the initial position and the termination position of the UAV into the dual policy learning network model for prediction to obtain the flight path of the UAV at the steel bar binding site. It solves the problems that the path length planned by the existing UAV path planning method is not the shortest and optimal, and the planning time often needs to wait for a long time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of UAV path planning at the steel bar binding site. Specifically, it relates to a UAV path planning method, device, storage medium, and electronic device for the steel bar binding site. Background Technique

[0002] UAVs can quickly and flexibly perform binding quality inspection tasks at the steel bar binding site. Through path planning, it can be ensured that the UAV completes the task according to the optimal path, reducing time and resource waste. Good path planning not only needs to comprehensively cover the inspection points but also helps the UAV avoid obstacles at the construction site, reducing the risk of collision with other construction equipment or personnel and ensuring the safety of the construction site. However, the environment at the steel bar binding site is usually complex, with a large number of obstacles such as columns, beams, and walls and dynamically changing working conditions, which increases the difficulty of path planning.

[0003] The path planning methods of UAVs are mainly divided into classical algorithms and intelligent algorithms. Classical algorithms include the A* algorithm, artificial potential field method, Rapidly-Exploring Random Tree (RRT), and cell decomposition method. However, traditional global path planning algorithms are characterized by fast planning speed but lack of optimality. Planning algorithms that rely on prior environmental information can solve general static scenario problems but have limitations in the case of moving obstacles.

[0004] Therefore, in response to the challenges of UAV path planning in complex scenarios, scholars have proposed bio-inspired intelligent algorithms that simulate the foraging behavior of organisms in dynamic environments. These algorithms include genetic algorithms, ant colony algorithms, and particle swarm algorithms. With the increasing complexity and uncertainty of the task environment, UAV path planning faces greater challenges. In recent years, research on UAV path planning based on artificial intelligence has increased significantly, further expanding the application of Deep Reinforcement Learning (DRL) in UAV path planning. The existing UAV path planning methods cannot avoid obstacles while planning the path, and the path lengths planned by the UAV path planning methods are not the shortest and optimal, and the planning time often requires a long wait. Summary of the Invention

[0005] The main purpose of this application is to provide a UAV path planning method, device, storage medium, and electronic device for the steel bar binding site, so as to at least solve the problem that the path lengths planned by the existing UAV path planning methods are not the shortest and optimal, and the planning time often requires a long wait.

[0006] To achieve the above object, according to one aspect of the present application, a method for path planning of an unmanned aerial vehicle (UAV) at a steel bar binding site is provided, including: constructing a UAV path planning environment for the steel bar binding site with Markov properties, where the UAV path planning environment includes: an obstacle evaluation model, a UAV position model, a UAV movement direction model, and a reward function model; obtaining the historical initial position and the historical termination position within a historical time period, and inputting the historical initial position, the historical termination position, and the UAV path planning environment model into an action selection strategy model for operation to obtain a historical UAV action strategy, and using the historical UAV action strategy to control the movement of the UAV and obtaining historical real-time operation data during the movement, and determining the historical real-time operation data as sample data, where the sample data includes: UAV current moment position information, UAV current moment direction information, UAV next moment predicted position information, UAV current moment reward data, and the priority of the sample data; obtaining a dual policy learning network model, where the dual policy learning network model is trained using the sample data; inputting the initial position and the termination position of the UAV into the dual policy learning network model for prediction to obtain the flight path of the UAV at the steel bar binding site.

[0007] Optionally, constructing a UAV path planning environment model for the steel bar binding site with Markov properties includes: determining the obstacle evaluation model according to the distance between the UAV and the obstacle, the set value of the maximum radius of the danger zone, and the set value of the maximum radius of the obstacle within the danger zone; constructing the UAV position model and the UAV movement direction model, where the UAV position model is used to determine the position of the UAV in a two-dimensional space, and the UAV movement direction model is used to divide the area of the steel bar binding site to determine the movement direction of the UAV; constructing a reward function model, and using the reward function model to reward the operation state of the UAV; constructing the UAV path planning environment for the steel bar binding site with the Markov properties according to the obstacle evaluation model, the UAV position model, the UAV movement direction model, and the reward function model.

[0008] Optionally, determining the obstacle evaluation model according to the distance between the UAV and the obstacle, the set value of the maximum radius of the danger zone, and the set value of the maximum radius of the obstacle within the danger zone includes: according to the first formula: , determining the probability that the UAV is crashed, where in the formula, is the probability that the UAV is crashed, D is the distance between the UAV and the obstacle, represents the set value of the maximum radius of the danger zone, is the set value of the maximum radius of the obstacle within the danger zone, where ; According to the second formula: Evaluate the overall risk of damage to the UAV. In the formula, is the overall risk of damage to the UAV, k represents the number of dangerous areas, is the probability that the UAV is destroyed when passing through the i-th dangerous area; determine the obstacle evaluation model according to the first formula and the second formula.

[0009] Optionally, construct a reward function model and use the reward function model to reward the operating state of the UAV, including: using the fifth formula: Construct the reward function model. In the formula, R is the reward function, is the penalty term for the UAV flying into the dangerous area, and the penalty value, is the reward term for the UAV flying to the target position, and the reward A value, is the penalty term for the UAV flying out of the steel bar binding site, and the penalty B value, is the penalty term for the UAV's flight time exceeding the flight time threshold, and the penalty C value, is the penalty term for the UAV in other set states, and the penalty D value; among them, according to the sixth formula determine value, where is the overall risk of damage to the UAV, is the risk threshold.

[0010] Optionally, input the historical initial position, the historical termination position, and the UAV path planning environment model into the action selection strategy model for operation to obtain the historical UAV action strategy, including: dividing the UAV motion direction model according to the historical initial position and the historical termination position by using a heuristic search algorithm to obtain multiple sub-operation directions; using a greedy strategy algorithm to control the UAV action according to the multiple sub-operation directions to generate multiple historical UAV action strategies.

[0011] Optionally, after controlling the movement of the drone using the historical drone action strategy and obtaining historical real-time operation data during the movement, and determining the historical real-time operation data as sample data, the method further includes: constructing a priority flight experience storage pool, and inputting the sample data into the priority flight experience storage pool; selecting a preset number of the experience sample data from the sample data in the priority flight experience storage pool in descending order of the priority, and inputting them into the initial dual-policy learning network model to generate an expected Q value of the evaluation network and a target Q value of the target network, where the initial dual-policy learning network model includes the evaluation network and the target network; according to the target Q value and the expected Q value, using the third formula: , determining the TD target error value, where represents the TD target error value, represents the parameters of the target network, is the target Q value, Q is the expected Q value, represents the parameters of the evaluation network, represents the discount factor, which is used to represent the importance the drone attaches to future rewards, A is the movement direction of the drone, and the subscript j represents the jth experience sample data, represents the reward obtained by the drone, represents the feature vector of the drone position S, which is used to approximate the state value, The function represents finding an action such that the target Q value of the evaluation network Q is maximized; according to the TD target error value and the expected Q value of the evaluation network, using the fourth formula: , determining the loss function of the initial dual-policy learning model, where is the loss function, represents the priority weight of the jth experience sample data, is the position of the drone; optimizing the parameters of the target network and the parameters of the evaluation network of the dual-policy learning network according to the loss function to obtain the dual-policy learning model.

[0012] Optionally, after optimizing the parameters of the target network and the parameters of the evaluation network of the dual-policy learning network according to the loss function, the method further includes: using the seventh formula: , recalculating the priority of the sample data in the priority flight experience storage pool, where in the formula, represents the priority of the sample data, represents the TD target error value, is the expected Q value of the sample data j, represents the exploration rate.

[0013] According to another aspect of the present application, there is provided a UAV path planning device for a steel bar binding site, including: a construction unit for constructing a UAV path planning environment for the steel bar binding site with Markov properties, wherein the UAV path planning environment includes: an obstacle evaluation model, a UAV position model, a UAV movement direction model, and a reward function model; a first acquisition unit for acquiring a historical initial position and a historical termination position within a historical time period, and inputting the historical initial position, the historical termination position, and the UAV path planning environment model into an action selection strategy model for operation to obtain a historical UAV action strategy, and controlling the UAV to move using the historical UAV action strategy and acquiring historical real-time operation data during the movement, and determining the historical real-time operation data as sample data, where the sample data includes: UAV current moment position information, UAV current moment direction information, UAV next moment predicted position information, UAV current moment reward data, and the priority of the sample data; a second acquisition unit for acquiring a dual policy learning network model, where the dual policy learning network model is trained using the sample data; a prediction unit for inputting the initial position and the termination position of the UAV into the dual policy learning network model for prediction to obtain the flight path of the UAV at the steel bar binding site.

[0014] According to yet another aspect of the present application, there is provided a computer-readable storage medium, where the computer-readable storage medium includes a stored program, and when the program runs, it controls the device where the computer-readable storage medium is located to execute any one of the UAV path planning methods for the steel bar binding site.

[0015] According to still another aspect of the present application, there is provided an electronic device, including: one or more processors, a memory, and one or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include those for executing any one of the UAV path planning methods for the steel bar binding site.

[0016] Applying the technical solution of the present application, a UAV path planning environment with Markov property for the steel bar binding site is constructed. The UAV path planning environment includes: an obstacle evaluation model, a UAV position model, a UAV movement direction model, and a reward function model; the historical initial position and historical end position within a historical time period are obtained, and the historical initial position, historical end position, and UAV path planning environment model are input into the action selection strategy model for operation to obtain the historical UAV action strategy. The UAV is controlled to move using the historical UAV action strategy, and historical real-time operation data is obtained during the movement. The historical real-time operation data is determined as sample data, and the sample data includes: UAV current moment position information, UAV current moment direction information, UAV next moment predicted position information, UAV current moment reward data, and the priority of the sample data; a dual policy learning network model is obtained, where the dual policy learning network model is trained using the sample data; the initial position and end position of the UAV are input into the dual policy learning network model for prediction to obtain the flight path of the UAV at the steel bar binding site. By establishing a UAV path planning environment with Markov property for the steel bar binding quality inspection task using a UAV, and using the dual policy learning network model to enable the UAV agent to better select an action to execute, the UAV can make an optimal strategy in a shorter time to plan the inspection flight path to the target position, which is suitable for real-time tasks such as steel bar binding quality inspection, and can also intelligently avoid obstacles that may exist at the steel bar binding site; it solves the problem that the path planned by the existing UAV path planning method is not the shortest and optimal, and the planning time often requires a long wait. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The specification drawings forming a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:

[0018] Figure 1 The hardware structure block diagram of a mobile terminal showing a method for UAV path planning at a steel bar binding site provided in an embodiment of the present application is shown;

[0019] Figure 2 The flowchart showing a method for UAV path planning at a steel bar binding site provided in an embodiment of the present application is shown;

[0020] Figure 3 The flowchart showing a specific method for UAV path planning at a steel bar binding site provided in an embodiment of the present application is shown;

[0021] Figure 4The evaluation network and the target network structure diagram of the dual strategy learning network model DSL-NetWork provided according to the embodiments of the present application are shown;

[0022] Figure 5 The structural block diagram of a drone path planning device at a steel bar binding site provided according to the embodiments of the present application is shown. Detailed implementation manners

[0023] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0024] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances for the embodiments of the present application described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0026] As introduced in the background art, the path length planned by the existing drone path planning method is not the shortest and optimal, and the planning time often requires a long wait. To solve the problem that the path length planned by the existing drone path planning method is not the shortest and optimal, and the planning time often requires a long wait, the embodiments of the present application provide a drone path planning method, device, storage medium and electronic device at a steel bar binding site.

[0027] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0028] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example,Figure 1 It is a hardware structure block diagram of a mobile terminal for the UAV path planning method at the steel bar binding site in an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, processing devices such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown.

[0029] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the UAV path planning method at the steel bar binding site in an embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, the above method is implemented. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the mobile terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0030] In this embodiment, a UAV path planning method at the steel bar binding site running on a mobile terminal, a computer terminal, or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from here.

[0031] Figure 2 is a flowchart of the UAV path planning method at the steel bar binding site according to an embodiment of the present application. As Figure 2 shown, the method includes the following steps:

[0032] Step S201, construct a UAV path planning environment for the steel bar binding site with Markov property, where the UAV path planning environment includes: an obstacle evaluation model, a UAV position model, a UAV movement direction model, and a reward function model;

[0033] The path planning environment includes a UAV, a target position, a danger area, and an environmental boundary area. The obstacle evaluation model, the UAV position model, the UAV movement direction model, and the reward function model are used to determine the UAV, the target position, the danger area, and the environmental boundary area in the path planning environment. Among them, the UAV can only move inside the environmental boundary. The danger area is the area such as beams, columns, and walls in the construction site that the UAV path planning needs to avoid. Usually, the UAV performs the steel bar binding identification and detection task at a relatively low constant height. Therefore, the path planning environment scenario designed in this solution is a two-dimensional space.

[0034] Among them, the Markov property refers to the characteristic that a stochastic process has no aftereffect, that is, the future state only depends on the current state and is independent of the past state. Specifically, for a stochastic process, if the state at the current moment is given, the conditional probability distribution of the future state only depends on the state at the current moment and is independent of the past state sequence, then this stochastic process has the Markov property. The Markov property is an important assumption in many stochastic processes, which simplifies the modeling and analysis of reinforcement learning.

[0035] Step S202, obtain the historical initial position and the historical termination position within the historical time period, and input the historical initial position, the historical termination position, and the UAV path planning environment model into the action selection strategy model for operation to obtain the historical UAV action strategy. Then, use the historical UAV action strategy to control the movement of the UAV and obtain the historical real-time operation data during the movement process. Determine the historical real-time operation data as sample data, and the sample data includes: the UAV's current moment position information, the UAV's current moment direction information, the UAV's next moment predicted position information, the UAV's current moment reward data, and the priority of the sample data;

[0036] Specifically, in each time step during the path planning process, the intelligent agent outputs an action to make a decision. This solution combines heuristic rules and the ε-greedy strategy to let the UAV make action decisions and obtain sample data to optimize the action selection during the training process and improve the learning efficiency.

[0037] Step S203, obtaining a dual strategy learning network model, wherein the dual strategy learning network model is obtained by training with sample data;

[0038] Specifically, a dual strategy learning network DSL-NetWork (Duality Strategy Learning Network) is constructed, which consists of an evaluation network and a target network. The network structure of the dual strategy learning network DSL-NetWork is as follows: Figure 3 As shown, the evaluation network and the target network have the same neural network structure: a fully connected layer L1 with 128 neurons and a ReLu activation function layer, and a fully connected layer L2 with 64 neurons and a ReLu activation function layer. The fully connected layers act as approximations. w is a learnable parameter of the evaluation network, and the target network replicates w every C rounds as its own running parameter w'.

[0039] In step S204, the initial position and the final position of the UAV are input into the dual strategy learning network model for prediction to obtain the flight path of the UAV at the steel bar binding site.

[0040] Specifically, when planning the drone path for a rebar binding quality inspection site, the dangerous area is delineated according to the on-site environment, and the starting and target positions are set. The drone intelligent agent will use the trained dual strategy learning network to select a reasonable inspection flight path that covers all routes and intelligently avoids dangerous areas.

[0041] Through this embodiment, a UAV path planning environment for the steel bar binding site with Markov property is constructed. The UAV path planning environment includes: an obstacle evaluation model, a UAV position model, a UAV movement direction model, and a reward function model. The historical initial position and historical termination position within a historical time period are obtained, and the historical initial position, historical termination position, and UAV path planning environment model are input into the action selection strategy model for operation to obtain the historical UAV action strategy. The historical UAV action strategy is used to control the movement of the UAV, and historical real-time operation data is obtained during the movement process. The historical real-time operation data is determined as sample data, and the sample data includes: UAV current moment position information, UAV current moment direction information, UAV next moment predicted position information, UAV current moment reward data, and the priority of the sample data. A dual policy learning network model is obtained, where the dual policy learning network model is trained using the sample data. The initial position and termination position of the UAV are input into the dual policy learning network model for prediction to obtain the flight path of the UAV at the steel bar binding site. By establishing a UAV path planning environment for the steel bar binding site with Markov property according to the task of using the UAV for steel bar binding quality inspection, and using the dual policy learning network model to enable the UAV agent to better select an action to execute, the UAV can make an optimal strategy in a shorter time to plan the detection flight path to reach the target position, which is suitable for real-time tasks such as steel bar binding quality inspection, and can also intelligently avoid obstacles that may exist at the steel bar binding site; it solves the problem that the path planned by the existing UAV path planning method is not the shortest and optimal, and the planning time often requires a long wait.

[0042] In the specific implementation process, a UAV path planning environment model for the steel bar binding site with Markov property is constructed, including: determining the obstacle evaluation model according to the distance between the UAV and the obstacle, the set value of the maximum radius of the dangerous area, and the set value of the maximum radius of the obstacle in the dangerous area; constructing a UAV position model and a UAV movement direction model. Among them, the UAV position model is used to determine the position of the UAV in the two-dimensional space, and the UAV movement direction model is used to divide the area of the steel bar binding site to determine the movement direction of the UAV; constructing a reward function model, and using the reward function model to reward the operation state of the UAV; constructing a UAV path planning environment for the steel bar binding site with Markov property according to the obstacle evaluation model, UAV position model, UAV movement direction model, and reward function model.

[0043] Since the path planning environment scenario is a two-dimensional space, coordinate points are used to describe the position of the UAV in the environment to obtain the UAV position model, and this position is used as the position variable S.

[0044] To improve the efficiency of path planning, the environmental area is divided into nine grids arranged in three rows and three columns. The size of the grids can be designed according to the size of the specific steel bar binding quality inspection area. Assuming that the drone can fly freely to any one of the surrounding eight grids, the discrete action space consists of eight actions; the action space A of the drone motion direction model is expressed as: ; The digital codes (0, 1, 2, 3, 4, 5, 6, 7) in the action space A of the drone motion direction model are used to represent the corresponding moving directions, that is, moving from the current grid position to the adjacent grid position, which can be changed by programming according to the convention. For example: 0 represents moving north, 1 represents moving south, 2 represents moving west, 3 represents moving east, 4 represents moving northwest, 5 represents moving northeast, 6 represents moving southwest, and 7 represents moving southeast.

[0045] Specifically, an obstacle evaluation model is determined according to the distance between the drone and the obstacle, the set value of the maximum radius of the danger zone, and the set value of the maximum radius of the obstacle in the danger zone, including: According to the first formula: , determine the probability that the drone is destroyed, where, is the probability that the drone is destroyed, D is the distance between the drone and the obstacle, represents the set value of the maximum radius of the danger zone, is the set value of the maximum radius of the obstacle in the danger zone, where, ; According to the second formula: evaluate the overall risk of damage to the drone, where, is the overall risk of damage to the drone, k represents the number of danger zones, is the probability that the drone is destroyed when passing through the i-th danger zone; determine the obstacle evaluation model according to the first formula and the second formula.

[0046] Among them, and The relationship can be set as .

[0047] This method calculates the probability that the drone is destroyed in real time and the probability that the drone is destroyed when passing through the th danger zone, which can reduce the probability of the drone being destroyed.

[0048] More specifically, a reward function model is constructed, and the reward function model is used to reward the running state of the drone, including: Using the fifth formula: Construct a reward function model, where, R is the reward function, is the penalty term for the drone flying into the danger zone, and the penalty value, It is the reward item for the drone to fly to the target position, with a reward value of A. It is the penalty item for the drone to fly out of the steel bar binding site, with a penalty value of B. It is the penalty item for the flight time of the drone exceeding the flight time threshold, with a penalty value of C. It is the penalty item for the drone in other set states, with a penalty value of D; among them, according to the sixth formula Determine The value, where Is the overall risk degree of the drone being damaged, Is the risk degree threshold.

[0049] Among them, A can be set to 200, B can be set to 50, C can be set to 50, other states can be the normal flight state or the stationary state of the drone, and D can be set to 0.5. A smaller penalty value is used to encourage the drone to reach the target area as soon as possible and prevent the drone from adopting the strategy of staying still.

[0050] This method Is related to the overall risk degree of the drone being damaged When Exceeds the threshold The probability that the drone hits obstacles such as beams and columns at the steel bar binding site is 100%, and the penalty item is set to -50. When Is lower than the threshold, there is a certain probability that the drone will collide. When the drone meets , , , Any one of the four required conditions, the current flight round ends; otherwise, the drone remains in other states.

[0051] Furthermore, the historical initial position, historical termination position, and the drone path planning environment model are input into the action selection strategy model for operation to obtain the historical drone action strategy, including: dividing the drone motion direction model using the heuristic search algorithm according to the historical initial position and historical termination position to obtain multiple sub-motion directions; using the greedy strategy algorithm to control the drone action according to the multiple sub-motion directions to generate multiple historical drone action strategies.

[0052] This method further divides a smaller action space A' (sub-motion direction) from the action space A of the drone motion direction model according to the historical initial position and historical termination position of the drone. The smaller action space consists of three actions selected in the target direction and two actions at the boundaries of the target direction. In this way, the drone selects an action within the smaller action space, improving the training efficiency, shortening the training cycle, and preventing the problem of falling into local optimality.

[0053] For example: represents the current position of the drone, and represents the position of its target. If the target position is southeast of the drone, satisfying and , then the included moving directions are east, south, southeast, southwest, and northeast; if the target position is northeast of the drone, satisfying and , then the included moving directions are northwest, north, northeast, east, and southeast; if the target position is northwest of the drone, satisfying and , then the included moving directions are northeast, north, northwest, west, and southwest; if the target position is southwest of the drone, satisfying and , then the included moving directions are northwest, west, southwest, south, and southeast.

[0054] Furthermore, after controlling the movement of the drone using the historical drone action strategy and obtaining historical real-time operation data during the movement, and determining the historical real-time operation data as sample data, the method further includes: constructing a priority flight experience storage pool, and inputting the sample data into the priority flight experience storage pool; selecting a preset number of experience sample data from the sample data in the priority flight experience storage pool in descending order of priority and inputting them into the initial dual policy learning network model to generate the expected Q value of the evaluation network and the target Q value of the target network, where the initial dual policy learning network model includes an evaluation network and a target network; according to the target Q value and the expected Q value, using the third formula: , determining the TD target error value, where represents the TD target error value, represents the parameters of the target network, is the target Q value, Q is the expected Q value, represents the parameters of the evaluation network, represents the discount factor, which is used to represent the importance the drone attaches to future rewards, A is the movement direction of the drone, and the subscript j represents the jth experience sample data, represents the reward obtained by the drone, represents the feature vector of the drone position S, which is used to approximate the state value, The function represents finding an action such that the target Q value of the evaluation network Q is maximized; according to the TD target error value and the expected Q value of the evaluation network, using the fourth formula: , determining the loss function of the initial dual policy learning model, where is the loss function, represents the priority weight of the j-th empirical sample data, is the position of the UAV; the parameters of the target network and the evaluation network of the dual policy learning network are optimized according to the loss function to obtain the dual policy learning model.

[0055] Among them, the priority flight experience storage pool assigns priorities to empirical samples through a priority function to improve the learning efficiency. Different from the traditional experience replay pool, the priority flight experience storage pool in this solution does not sample empirical samples uniformly, but preferentially samples those with higher priorities. This can make the learning algorithm pay more attention to those experiences that are more important for learning, thereby improving the learning efficiency.

[0056] Specifically, after optimizing the parameters of the target network and the evaluation network of the dual policy learning network according to the loss function, the method further includes: adopting the seventh formula: , recalculate the priority of the sample data in the priority flight experience storage pool, where in the formula, represents the priority of the sample data, represents the TD target error value, is the expected Q value of the sample data j, represents the exploration rate.

[0057] Among them, the priority function takes the TD error of the empirical sample as the input of the priority function for priority assignment. For an empirical sample with a large TD error, it means that the UAV agent has not well learned this empirical sample, and the priority function will assign a relatively large value to this empirical sample, so that this empirical sample is more likely to be sampled by the dual policy learning network.

[0058] ;

[0059] ;

[0060] ;

[0061] represents the probability that the t-th experience is sampled, represents the priority of the sample data, is a hyperparameter used to control the attenuation speed of the priority. represents the TD error, which is used to measure the estimated value of the expected reward calculated after the UAV agent executes an action and the TD target the difference between them.

[0062] When the empirical samples stored in the experience replay pool exceed After that, some old experience samples need to be deleted. This solution adopts a priority-based deletion strategy, that is, those experience samples with lower priorities are deleted first. To improve the training speed, after the experience replay pool is constructed, the priority of each experience sample is no longer dynamically updated during the training process.

[0063] In order to enable those skilled in the art to more clearly understand the technical solution of the present application, the implementation process of the UAV path planning method at the steel bar binding site of the present application will be described in detail below in conjunction with specific embodiments.

[0064] This embodiment relates to a specific UAV path planning method at the steel bar binding site, as Figure 4 shown, including the following steps:

[0065] Step S1: Build a UAV path planning environment at the steel bar binding site with Markov properties;

[0066] The path planning environment includes a UAV, target positions, dangerous areas, and environmental boundary areas. The UAV can only move inside the environmental boundary. The dangerous areas are areas such as beams, columns, and walls in the construction site that need to be avoided in UAV path planning. Usually, the UAV performs steel bar binding identification and detection tasks at a relatively low constant altitude. Therefore, the path planning environment scenario designed in this embodiment is a two-dimensional space.

[0067] Step S1.1: Construct an obstacle evaluation method; construct an obstacle evaluation method based on the distance between the UAV and the center of the dangerous area. The degree of danger of the center of the dangerous area to the UAV is defined by the following formula, indicating the probability that the UAV is crashed.

[0068] ; where D represents the distance between the UAV and the center of the obstacle. represents the maximum radius of the dangerous area. represents the maximum radius of the obstacle entity in the dangerous area. differs from by 1 meter.

[0069] Assume that the crash probabilities of the UAV in the dangerous areas formed by different column-beams and other obstacles are independent of each other. Use the following formula to calculate the overall risk of the UAV being damaged when flying through each dangerous area .

[0070] ;

[0071] where k represents the number of dangerous areas. is the probability that the UAV is destroyed when passing through the th dangerous area.

[0072] Step S1.2: Construct the state space;

[0073] The path planning environment scenario is a two-dimensional space. We use the coordinate point (x, y) to describe the position of the UAV in the environment and take this position as the state variable S of the Markov decision process (MDP) model. .

[0074] Step S1.3: Construct the action space;

[0075] To improve the efficiency of path planning, the environmental area is divided into nine grids arranged in three rows and three columns. The size of the grids can be designed according to the size of the specific steel bar binding quality inspection area. Assume that the UAV can fly freely to any one of the surrounding eight grids, then the discrete action space consists of eight actions.

[0076] Action space is represented by the following equation: ;

[0077] Action space The digital coding (0, 1, 2, 3, 4, 5, 6, 7) in is used to represent the corresponding moving direction, that is, moving from the current grid position to the adjacent grid position, and the convention representation can be changed through programming.

[0078] For example: 0 represents moving north, 1 represents moving south, 2 represents moving west, 3 represents moving east, 4 represents moving northwest, 5 represents moving northeast, 6 represents moving southwest, and 7 represents moving southeast.

[0079] Step S1.4: Construct the reward function to guide the UAV agent to learn the optimal policy;

[0080] Construct the reward function R, which consists of five reward terms and is represented by the following equation:

[0081] ;

[0082] Among them, the reward terms , , , and represent the reward terms and penalty terms of the UAV in different states. represents the penalty term for flying into the dangerous area, and the penalty value. represents the reward term for flying back to the specified target point, and the reward value is 200. represents the penalty term for flying out of the environmental boundary, and the penalty value is -50. represents the penalty term for the UAV's flight time exceeding the specified maximum flight duration, and the penalty value is -50. Represents the penalty term for the drone in other states, with a penalty value of -0.5. A smaller penalty value is used to encourage the drone to reach the target area as soon as possible and prevent the drone from adopting a strategy of staying still.

[0083] Is related to the total risk level When Exceeds the threshold The probability that the drone hits obstacles such as beams and columns at the steel bar binding site is 100%. The penalty term is set to -50. When Is lower than the threshold, there is a certain probability that the drone will collide, which is expressed as: ;

[0084] When the drone's state satisfies , , , Any one of the four required conditions, the current flight round terminates; otherwise, the drone remains in other states, and other states are the states where the drone flies normally or stays still.

[0085] Among them, in reinforcement learning, the agent (drone) refers to the entity that participates in learning and decision-making, and it is at the forefront of interacting with the environment. The agent learns and improves its behavior by observing the state of the environment, selecting actions, and receiving rewards to achieve a certain predetermined goal. Its main task is to select the optimal action according to the feedback (reward signal) provided by the environment to maximize the long-term cumulative reward.

[0086] Among them, the Markov property means that a stochastic process has the property of no aftereffect, that is, the future state only depends on the current state and is independent of the past state. Specifically, for a stochastic process, if the state at the current moment is given, the conditional probability distribution of the future state only depends on the state at the current moment and is independent of the past state sequence, then this stochastic process has the Markov property. The Markov property is an important assumption in many stochastic processes, which simplifies the modeling and analysis of reinforcement learning.

[0087] Step S2: Develop an action selection strategy;

[0088] In each time step during the path planning process, the agent outputs an action to make a decision. In this embodiment, a heuristic rule and an ε-greedy policy are combined to let the agent make action decisions to optimize the action selection during training and improve the learning efficiency.

[0089] Step S2.1: Use a heuristic search algorithm to divide the smaller action space;

[0090] Construct the action space from Step 1.2 according to the current position and target position of the UAV to further divide it into smaller action spaces . The smaller action space consists of three actions selected in the target direction and two actions at the boundaries of the target direction. In this way, the UAV selects an action within the smaller action space, which can improve the training efficiency, shorten the training cycle, and prevent getting stuck in the local optimum problem.

[0091] Let represent the current position of the UAV, and represent the position of its target. If the target position is in the southeast of the UAV, satisfying and , then contains the moving directions of east, south, southeast, southwest, and northeast; if the target position is in the northeast of the UAV, satisfying and , then contains the moving directions of northwest, north, northeast, east, and southeast; if the target position is in the northwest of the UAV, satisfying and , then contains the moving directions of northeast, north, northwest, west, and southwest; if the target position is in the southwest of the UAV, satisfying and , then

[0092] For example:

[0093] According to the action space defined in Step S1.3, if the target position is in the southeast of the UAV, satisfying and , then the UAV agent will try all possible operations from the action space A = (1, 3, 5, 6, 7).

[0094] Similarly, if the target position is in the northeast of the UAV, satisfying and , then the UAV agent will try all possible operations from the action space A = (0, 3, 4, 5, 7).

[0095] If the target position is in the northwest of the UAV, satisfying and , then the UAV agent will try all possible operations from the action space A = (0, 2, 4, 5, 6).

[0096] If the target position is in the southwest of the UAV, satisfying and Then the UAV agent will try all possible operations from the action space A = (0, 2, 4, 6, 7).

[0097] Step S2.2: Select an action to execute in the action space using the greedy strategy;

[0098] Let represent the exploration rate, indicating the probability that the UAV agent randomly selects an action to execute in the action space. It is defined by the following equation, where represents the initial exploration rate of the UAV agent. represents the final exploration rate. i represents the current training round, represents the pre-training iteration to generate a specified amount of data, represents the exploration rate decay steps, represents the maximum exploration rate decay steps.

[0099] ;

[0100] During the entire actual training process, first give a relatively high initial value , so that the UAV can be biased to randomly select different actions to execute in the smaller action space obtained in step S2.1. As the number of training iterations increases, gradually decrease until it stabilizes, so that the agent is biased to select the action with a higher Q-value in the smaller action space obtained in step S2.1.

[0101] Among them, in reinforcement learning, the Q-value is a metric used to measure the value of taking a certain action in a given state. Specifically, for state s and action a, the Q-value is expressed as Q(s, a), representing the expected cumulative reward that can be obtained by taking action a in state s. The Q-value function is a key concept in reinforcement learning, which can help the agent (UAV) select the optimal action that should be taken in each state, so that the agent can obtain the maximum cumulative reward in the environment.

[0102] Step S3: Construct a dual policy learning network;

[0103] Step S3.1: Construct the Duality Strategy Learning Network (DSL-NetWork), which consists of an evaluation network and a target network. The neural network structures of the evaluation network and the target network are the same. Build a fully connected layer L1 with 128 neurons, build a ReLu activation function layer, build a fully connected layer L2 with 64 neurons, and build a ReLu activation function layer. The role of the fully connected layer is the approximate value function. w is the learnable parameter of the evaluation network, and the target network copies w as its running parameter w' every C rounds. The network structure of DSL-NetWork is as Figure 3 shown.

[0104] Step S3.2: Both the evaluation network and the target network take the state S as input and output the value function value and the advantage function value . The value function depends entirely on the state S and is independent of the action A, denoted as , is the parameter of the value function V, The output of has only one value, indicating the expected value of the long-term cumulative reward that the agent can obtain in the input state S. The advantage function is related to both the state S and the action A, denoted as , represents the parameter of the advantage function . The output of the advantage function has eight values, indicating that when the agent selects different actions in the same input state S, it will affect the immediate reward obtained in the short term. The eight output values represent the advantages and disadvantages of eight actions in the action space, and the higher the value, the better the action.

[0105] Step S3.3: Calculate the expected Q value of the evaluation network Q and the target Q value of the target network , and the calculation formula is as follows:

[0106] ;

[0107] ;

[0108] Among them, |A| is equal to the size of the action space. Among them, the learnable parameters in the neural network will be updated through optimization algorithms (such as gradient descent) during the training process to enable the network to adapt to the training data and improve performance.

[0109] For example: The value function value is 100, and the advantage function value is the 8 values of [5, 2, 4, 3, 1, 5, 6, 7, 8]. Therefore is calculated as (5 + 2 + 4 + 3 + 4 + 1 + 6 + 7 + 8) / 8 = 4.5. Then according to Calculation formula: 100 + (5 - 4.5) = 100.5; 100 + (2 - 4.5) = 97.5; 100 + (4 - 4.5) = 99.5; 100 + (3 - 4.5) = 98.5; 100 + (1 - 4.5) = 96.5; 100 + (5 - 4.5) = 100.5; 100 + (6 - 4.5) = 101.5; 100 + (7 - 4.5) = 102.5; 100 + (8 - 4.5) = 103.5;

[0110] The expected Q value of the evaluation network Q is [100.5, 97.5, 99.5, 98.5, 96.5, 101.5, 102.5, 103.5].

[0111] Step S4: Construct a priority flight experience storage pool;

[0112] The priority flight experience storage pool assigns priorities to experience samples through a priority function to improve the learning efficiency. Different from the traditional experience replay pool, the priority flight experience storage pool in this embodiment does not sample experience samples evenly, but preferentially samples those with higher priorities. This can make the learning algorithm pay more attention to those experiences that are more important for learning, thereby improving the learning efficiency.

[0113] Step S4.1: Initialize the priority flight experience storage pool;

[0114] Create an empty priority flight experience storage pool that can store N pieces of experience to store experience samples. Each experience sample contains the following information: the state at this moment , action , reward , the state at the next moment , whether the task is completed d, and priority P. Among them, the storage capacity "N" of the priority flight experience pool in this embodiment is 1,000,000 pieces.

[0115] Step S4.2: Define a priority function P(t) to measure the importance of each experience sample;

[0116] The priority function takes the TD error of the experience sample as the input of the priority function for priority assignment. For an experience sample with a large TD error, it means that the drone agent has not well learned this experience sample. The priority function will assign a relatively large value to this experience sample, so that this experience sample is more likely to be sampled by the dual policy learning network.

[0117] ;

[0118] ;

[0119] ;

[0120] represents the probability that the $t$-th experience is sampled, represents the priority of the sample data, is a hyperparameter used to control the decay rate of the priority. represents the TD error, which is used to measure the estimated value of the expected reward calculated after the UAV agent executes an action and the TD target the difference between.

[0121] When the number of experience samples stored in the experience replay pool exceeds After a certain number, some old experience samples need to be deleted. In this embodiment, a priority-based deletion strategy is adopted, that is, those experience samples with lower priorities are preferentially deleted. To improve the training speed, after the experience replay pool is constructed, the priority of each experience sample is no longer dynamically updated during the training process.

[0122] Among them, in reinforcement learning, the TD (Temporal Difference) error is an error metric used to evaluate the prediction model or value function. The TD error can be calculated by taking the difference between the estimated value of the prediction model and the actually observed reward value or the estimated value of the next state.

[0123] Among them, the experience replay pool is an important technique in deep reinforcement learning, usually used to train algorithms such as the deep Q-network (DQN). Its main purpose is to store the experiences of the agent's interaction with the environment and randomly sample from it during training to smoothly train the neural network and improve the efficiency and stability of training. The working principle of the experience replay pool is as follows:

[0124] 1. Store experiences: When the agent interacts with the environment, it saves the experiences of each step (including state, action, reward, next state, etc.) into the experience replay pool.

[0125] 2. Random sampling: During the training process, instead of directly using the latest collected experiences to update the neural network parameters, a batch of experiences is randomly sampled from the experience replay pool.

[0126] 3. Stable training: By random sampling, the training samples can be made more stable, reducing the data correlation during the training process and preventing the situation where the correlation between training samples is too high, resulting in unstable training of the neural network.

[0127] 4. Reuse experiences: The experience replay pool allows the agent to reuse previous experiences multiple times, which can improve the utilization efficiency of data and is particularly important in the case of scarce samples.

[0128] Among them, fully connected layers play an important role in deep learning. They can learn the complex relationships between input data and map these relationships into the output space. During the forward propagation of the neural network, the fully connected layers are responsible for performing linear and non-linear transformations on the input data to generate the output of the network.

[0129] Step S5: Train the dual policy learning network applicable to UAV path planning at the steel bar binding site;

[0130] Step S5.1: Initialize the UAV path planning environment. Set the initial state of the UAV agent , that is, the starting coordinates of the UAV; set the coordinates of the target position; calibrate the positions of dangerous areas such as columns, beams, and walls at the steel bar binding site.

[0131] Step S5.2: Randomly initialize the dual policy learning network and its network parameters w and w', set the maximum number of epochs for network training , the number of epochs C for copying the target network, the optimizer used by the network and the learning rate ; set the discount factor of the environment ; set the greedy policy parameter in the action selection policy , calculate the number of epochs parameter required for the exploration rate and ; set the maximum capacity of the prioritized flight experience pool and the capacity of a single sampled experience sample ;

[0132] Step S5.3: The UAV agent selects an action to execute and enters the next state . The environment calculates the reward value obtained by this action according to the coordinate state of the UAV agent , assigns a value d to whether the task is completed according to whether the target position is reached, initializes the priority P of the experience sample to 1, and stores the six-tuple in the prioritized flight experience pool constructed in S5. Repeat S5.3 until the number of stored experience samples in the prioritized flight experience pool exceeds pieces, and execute S5.4.

[0133] Step S5.4: Sample pieces of experience samples from the prioritized flight experience pool according to the priority as the input of the dual policy learning network. After network calculation, obtain the expected Q value of the evaluation network Q and the target Q value of the target network .

[0134] Step S5.5: Use the target network The target Q-value calculates the TD target, denoted as , and the calculation formula is as follows:

[0135] ;

[0136] where represents the parameters of the target network , represents the discount factor, which is used to represent the degree of importance that the UAV agent attaches to future rewards, and its value range is [0, 1]. The subscript j represents this round. represents the reward obtained according to S2.3 in this round of state. represents the feature vector of state S, which is used to approximate the state value. The function represents finding an action A to maximize the target Q-value of the evaluation network Q.

[0137] Step S5.6: Use the expected Q-value of the evaluation network Q and the TD target , and calculate the loss function Loss by the following formula to optimize the dual policy learning network: ;

[0138] where, represents the priority weight of the j-th sample.

[0139] Step S5.7: Update the parameters w and of the dual policy learning network according to the calculated Loss value, so that the subsequent calculated Loss value continues to decrease. Every C times of updating w, execute

[0140] Step S5.8: Recalculate the priority of the experience samples in the experience pool according to the priority function of S4.2 .

[0141] Step S5.9: When the state of the UAV agent is the target state, the UAV path planning task for this round is completed. If the number of training rounds reaches episodes times, the training ends, and a UAV agent suitable for UAV path planning at the steel bar binding site is obtained. Otherwise, transfer to S5.1 for a new round of training process. When S is not the target state, transfer to step S5.3 for execution.

[0142] Among them, the following hyperparameter values can be set during training: the maximum number of rounds episodes is set to 20000, the number of rounds C for copying the target network is set to 8, the optimizer used is the Adam optimizer, the learning rates of the evaluation network and the target network are set to 5e-4 and 4e-4 respectively, the discount factor γ of the environment is set to 0.99, and the greedy policy parameter Set to 0.95, the number of rounds parameter required for the exploration rate and Set to 500 and 12000, the maximum capacity of the priority flight experience pool Set to 1,000,000, the sample capacity of a single sampling experience Set to 256.

[0143] Step S6: Apply the dual policy learning network to perform UAV path planning at the steel bar binding quality inspection site;

[0144] When performing UAV path planning at the steel bar binding quality inspection site that requires the use of a UAV, a dangerous area is delimited according to the on-site environment, and the starting position and target position are set. The UAV agent will select a reasonable, fully covered route, and intelligently avoid the detection flight path in the dangerous area through the trained dual policy learning network.

[0145] In the embodiment of the present application, a Markov decision model of the steel bar binding on-site environment is established according to the task of using a UAV for steel bar binding quality inspection, the state space, action space, and reward function of the UAV agent are constructed, and a heuristic search algorithm and a greedy strategy are used to enable the UAV agent to better select an action to execute. A dual policy learning network is designed to calculate the true and expected Q values for path search, and combined with the constructed priority flight experience storage pool for training, it can comprehensively learn the advantages and disadvantages of executing each action, so as to ensure that the UAV agent selects the optimal flight strategy.

[0146] This embodiment can enable the UAV to make an optimal strategy in a shorter time to plan the detection flight path to reach the target position, which is suitable for real-time tasks such as steel bar binding quality inspection, and at the same time intelligently avoid obstacles that may exist at the steel bar binding site. Compared with the traditional DQN method and DDAN method, it realizes a more stable cumulative reward, and has a faster convergence speed, can approximate the Q function more effectively and accurately, and can better adapt to the situations of different steel bar binding sites.

[0147] The embodiment of the present application also provides a UAV path planning device for the steel bar binding site. It should be noted that the UAV path planning device for the steel bar binding site in the embodiment of the present application can be used to execute the UAV path planning method for the steel bar binding site provided by the embodiment of the present application. The device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0148] The following introduces the UAV path planning device at the steel bar binding site provided by the embodiments of the present application.

[0149] Figure 5 It is a schematic diagram of the UAV path planning device at the steel bar binding site according to the embodiments of the present application. As Figure 5 shown, the device includes: a construction unit 51, a first acquisition unit 52, a second acquisition unit 53, and a prediction unit 54.

[0150] The construction unit 51 is used to construct a UAV path planning environment for the steel bar binding site with Markov properties. Among them, the UAV path planning environment includes: an obstacle evaluation model, a UAV position model, a UAV movement direction model, and a reward function model;

[0151] Among them, the Markov property means that a stochastic process has the property of no aftereffect, that is, the future state is only related to the current state and has nothing to do with the past state.

[0152] The first acquisition unit 52 is used to acquire the historical initial position and historical termination position within a historical time period, and input the historical initial position, historical termination position, and UAV path planning environment model into the action selection strategy model for operation to obtain the historical UAV action strategy, and use the historical UAV action strategy to control the UAV to move and obtain historical real-time operation data during the movement, and determine the historical real-time operation data as sample data. The sample data includes: UAV current moment position information, UAV current moment direction information, UAV next moment predicted position information, UAV current moment reward data, and the priority of the sample data;

[0153] Specifically, at each time step during the path planning process, the agent outputs an action to make a decision. This solution combines heuristic rules and ε-greedy strategies to let the UAV make action decisions and obtain sample data to optimize the action selection during training and improve the learning efficiency.

[0154] The second acquisition unit 53 is used to acquire the dual policy learning network model, where the dual policy learning network model is trained using sample data;

[0155] Among them, the dual policy learning network DSL-NetWork (Duality Strategy Learning Network) consists of an evaluation network and a target network.

[0156] The prediction unit 54 is used to input the initial position and termination position of the UAV into the dual policy learning network model for prediction to obtain the flight path of the UAV at the steel bar binding site.

[0157] Specifically, when planning the UAV path for the quality inspection of steel bar binding on-site that requires the use of a UAV, a dangerous area is delimited according to the on-site environment, and the starting position and target position are set. The UAV agent will select a reasonable, fully covered, and intelligent detection flight path that avoids the dangerous area through the trained dual-policy learning network.

[0158] In this embodiment, a construction unit is used to construct a UAV path planning environment for the steel bar binding site with Markov properties. The UAV path planning environment includes: an obstacle evaluation model, a UAV position model, a UAV movement direction model, and a reward function model; a first acquisition unit is used to acquire the historical initial position and historical termination position within a historical time period, and input the historical initial position, historical termination position, and the UAV path planning environment model into the action selection strategy model for operation to obtain the historical UAV action strategy. The historical UAV action strategy is used to control the movement of the UAV, and historical real-time operation data is acquired during the movement. The historical real-time operation data is determined as sample data, and the sample data includes: UAV current moment position information, UAV current moment direction information, UAV next moment predicted position information, UAV current moment reward data, and the priority of the sample data; a second acquisition unit is used to acquire the dual-policy learning network model, where the dual-policy learning network model is trained using the sample data; a prediction unit is used to input the initial position and termination position of the UAV into the dual-policy learning network model for prediction to obtain the flight path of the UAV at the steel bar binding site. By establishing a UAV path planning environment for the steel bar binding site with Markov properties according to the task of using a UAV for the quality inspection of steel bar binding, and using the dual-policy learning network model to enable the UAV agent to better select an action to execute, the UAV can make an optimal strategy in a shorter time to plan the detection flight path to reach the target position, which is suitable for real-time tasks such as the quality inspection of steel bar binding, and at the same time intelligently avoids possible obstacles at the steel bar binding site; it solves the problem that the path length planned by the existing UAV path planning method is not the shortest and optimal, and the planning time often requires a long wait.

[0159] As an alternative solution, the building unit includes a first determination module, a first building module, a second building module, and a third building module; the first determination module is used to determine an obstacle evaluation model according to the distance between the drone and the obstacle, the set value of the maximum radius of the danger zone, and the set value of the maximum radius of the obstacle in the danger zone; the first building module is used to build a drone position model and a drone movement direction model, where the drone position model is used to determine the position of the drone in the two-dimensional space, and the drone movement direction model is used to divide the area of the steel bar binding site to determine the movement direction of the drone; the second building module is used to build a reward function model and use the reward function model to reward the running state of the drone; the third building module is used to build a drone path planning environment for the steel bar binding site with Markov properties according to the obstacle evaluation model, the drone position model, the drone movement direction model, and the reward function model. Since the path planning environment scenario is a two-dimensional space, coordinate points are used to describe the position of the drone in the environment to obtain the drone position model, and this position is used as the position variable .

[0160] In an alternative solution, the first determination module includes a first determination sub-module, an evaluation sub-module, and a second determination sub-module; the first determination sub-module is used to determine the probability of the drone being crashed according to the first formula: , where, is the probability of the drone being crashed, D is the distance between the drone and the obstacle, represents the set value of the maximum radius of the danger zone, is the set value of the maximum radius of the obstacle in the danger zone, where, ; according to the second formula: evaluate the overall risk of the drone being damaged, where, is the overall risk of the drone being damaged, k represents the number of danger zones, is the probability that the drone is destroyed when passing through the i-th danger zone; determine the obstacle evaluation model according to the first formula and the second formula.

[0161] Specifically, by calculating in real time the probability of the drone being crashed and the probability that the drone is destroyed when passing through the th danger zone, the probability of the drone being crashed can be reduced.

[0162] In an alternative solution, the second building module includes a building sub-module, which is used to build a reward function model by using the fifth formula: , where, R is the reward function, is the penalty term for the drone flying into the danger zone, and the penalty value, is the reward term for the drone flying to the target position, and the reward A value, It is a penalty item for the UAV to fly out of the steel bar binding site, with a penalty value of B. It is a penalty item for the UAV's flight time exceeding the flight time threshold, with a penalty value of C. It is a penalty item for the UAV in other set states, with a penalty value of D; among them, according to the sixth formula Determine value, where is the overall risk of the UAV being damaged, is the risk threshold.

[0163] Among them, A can be set to 200, B can be set to 50, C can be set to 50, other states can be the normal flight state or the stationary state of the UAV, and D can be set to 0.5. A smaller penalty value is used to encourage the UAV to reach the target area as soon as possible and prevent the UAV from adopting the strategy of staying still.

[0164] An optional solution is that the first acquisition unit includes a division module and a control module; the division module is used to divide the UAV motion direction model according to the historical initial position and historical end position by using a heuristic search algorithm to obtain multiple sub-motion directions; the control module is used to control the UAV actions according to multiple sub-motion directions by using a greedy strategy algorithm to generate multiple historical UAV action strategies.

[0165] Further divide a smaller action space A' (sub-motion direction) from the action space A of the UAV motion direction model according to the historical initial position and historical end position of the UAV. The smaller action space consists of three actions selected in the target direction and two actions at the boundary of the target direction. In this way, the UAV selects an action within the smaller action space, which can improve the training efficiency, shorten the training cycle, and prevent falling into the local optimum problem.

[0166] An optional solution is that the device further includes an input unit, a selection unit, a first determination unit, a second determination unit, and an optimization unit; the input unit is used to obtain historical real-time operation data during the process of controlling the UAV to move by using the historical UAV action strategy. After determining the historical real-time operation data as sample data, a priority flight experience storage pool is constructed, and the sample data is input into the priority flight experience storage pool; the selection unit is used to select a preset number of experience sample data from the sample data in the priority flight experience storage pool in descending order of priority and input them into the initial dual policy learning network model to generate the expected Q value of the evaluation network and the target Q value of the target network. Among them, the initial dual policy learning network model includes an evaluation network and a target network; the first determination unit is used to use the third formula according to the target Q value and the expected Q value: , determine the TD target error value, where represents the TD target error value. represent the parameters of the target network is the target Q value, and Q is the expected Q value represent the parameters of the evaluation network represents the discount factor, which is used to indicate the importance that the drone attaches to future rewards. A is the movement direction of the drone, and the subscript j represents the j-th empirical sample data represents the reward obtained by the drone represents the feature vector of the drone's position S, which is used to approximate the state value The function represents finding an action A to maximize the target Q value of the evaluation network Q; according to the TD target error value and the expected Q value of the evaluation network, the fourth formula is adopted , to determine the loss function of the initial dual policy learning model, where is the loss function represents the priority weight of the j-th empirical sample data is the position of the drone; according to the loss function, optimize the parameters of the target network and the evaluation network of the dual policy learning network to obtain the dual policy learning model

[0167] Among them, the priority flight experience storage pool assigns priorities to empirical samples through a priority function to improve the learning efficiency. Different from the traditional experience replay pool, the priority flight experience storage pool in this solution does not sample empirical samples evenly, but preferentially samples those samples with higher priorities. This can make the learning algorithm pay more attention to those experiences that are more important for learning, thereby improving the learning efficiency

[0168] An optional solution is that the device further includes a calculation unit, which is used to, after optimizing the parameters of the target network and the evaluation network of the dual policy learning network according to the loss function, adopt the seventh formula , to recalculate the priorities of the sample data in the priority flight experience storage pool. In the formula, represents the priority of the sample data represents the TD target error value is the expected Q value of the sample data j represents the exploration rate

[0169] Among them, the priority function takes the TD error of the empirical sample as the input of the priority function for priority assignment. For empirical samples with large TD errors, it means that the drone agent has not well learned this empirical sample, and the priority function will assign a relatively large value to this empirical sample, so that this empirical sample is more likely to be sampled by the dual policy learning network

[0170] The UAV path planning device at the steel bar binding site includes a processor and a memory. The above-mentioned construction unit, first acquisition unit, second acquisition unit, prediction unit, etc. are all stored in the memory as program units, and the corresponding functions are realized by the processor executing the above program units stored in the memory. The above modules are all located in the same processor; or, the above modules are respectively located in different processors in any combination form.

[0171] The processor contains a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the problems that the path length planned by the existing UAV path planning method is not the shortest and optimal, and the planning time often needs to wait for a long time can be solved.

[0172] The memory may include non-permanent memory in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.

[0173] An embodiment of the present invention provides a computer-readable storage medium, which includes a stored program. When the program runs, it controls the device where the computer-readable storage medium is located to execute the UAV path planning method at the steel bar binding site.

[0174] An embodiment of the present invention provides a processor, which is used to run a program. When the program runs, it executes the UAV path planning method at the steel bar binding site.

[0175] An embodiment of the present invention provides a device, which includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it realizes the UAV path planning method at the steel bar binding site. The device in this article can be a server, a PC, a PAD, a mobile phone, etc.

[0176] This application also provides a computer program product, which is suitable for initializing the UAV path planning method at the steel bar binding site when executed on a data processing device.

[0177] Obviously, those skilled in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. They can be implemented by program code executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.

[0178] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0179] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of multiple flows and / or blocks.

[0180] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of multiple flows and / or blocks.

[0181] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide means for implementing the functions in the flowFigure 1 one or more processes and / or blocks Figure 1 steps of functions specified in one or more blocks

[0182] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0183] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0184] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0185] It should also be noted that the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity, or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity, or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, commodity, or device comprising the element.

[0186] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for path planning of an unmanned aerial vehicle at a steel bar binding site, characterized in that, Including: Construct a UAV path planning environment for the steel bar binding site with Markov property, where the UAV path planning environment includes: an obstacle evaluation model, a UAV position model, a UAV movement direction model, and a reward function model; Obtain the historical initial position and historical end position within a historical time period, and input the historical initial position, the historical end position, and the UAV path planning environment model into the action selection strategy model for operation to obtain the historical UAV action strategy. Then, use the historical UAV action strategy to control the movement of the UAV and obtain historical real-time operation data during the movement. Determine the historical real-time operation data as sample data, where the sample data includes: UAV current moment position information, UAV current moment direction information, UAV next moment predicted position information, UAV current moment reward data, and the priority of the sample data; Obtain the dual policy learning network model, where the dual policy learning network model is trained using the sample data; Input the initial position and end position of the UAV into the dual policy learning network model for prediction to obtain the flight path of the UAV at the steel bar binding site; Among them, inputting the historical initial position, the historical end position, and the UAV path planning environment model into the action selection strategy model for operation to obtain the historical UAV action strategy includes: According to the historical initial position and the historical end position, use the heuristic search algorithm to divide the UAV movement direction model to obtain multiple sub-movement directions; Use the greedy strategy algorithm to control the UAV action according to the multiple sub-movement directions to generate multiple historical UAV action strategies.

2. The method according to claim 1, wherein Construct a UAV path planning environment model for the steel bar binding site with Markov property, including: Determine the obstacle evaluation model according to the distance between the UAV and the obstacle, the set value of the maximum radius of the danger zone, and the set value of the maximum radius of the obstacle within the danger zone; Construct the UAV position model and the UAV movement direction model, where the UAV position model is used to determine the position of the UAV in the two-dimensional space, and the UAV movement direction model is used to divide the area of the steel bar binding site to determine the movement direction of the UAV; Construct a reward function model and use the reward function model to reward the operation state of the UAV; Construct the UAV path planning environment for the steel bar binding site with the Markov property according to the obstacle evaluation model, the UAV position model, the UAV movement direction model, and the reward function model.

3. The method according to claim 2, wherein Determine the obstacle evaluation model according to the distance between the UAV and the obstacle, the set value of the maximum radius of the danger zone, and the set value of the maximum radius of the obstacle within the danger zone, including: According to the first formula: , determine the probability that the drone is crashed. In the formula, is the probability that the drone is crashed, D is the distance between the drone and the obstacle, represents the set value of the maximum radius of the danger zone, is the set value of the maximum radius of the obstacle in the danger zone, where, ; According to the second formula: evaluate the overall risk of damage to the UAV, where is the overall risk of damage to the UAV, k represents the number of dangerous areas, is the probability that the UAV is destroyed when passing through the i-th dangerous area; Determine the obstacle evaluation model according to the first formula and the second formula.

4. The method according to claim 2, wherein Construct a reward function model and use the reward function model to reward the operation state of the UAV, including: Adopt the fifth formula: Construct the reward function model, where R is the reward function, is the penalty term for the UAV flying into the dangerous area, and the penalty value, is the reward term for the UAV flying to the target position, and the reward E value, is the penalty term for the UAV flying out of the steel bar binding site, and the penalty B value, is the penalty term for the UAV's flight time exceeding the flight time threshold, and the penalty C value, is the penalty term for the UAV in other set states, and the penalty D value; Among them, according to the sixth formula determine value, where is the overall risk degree of damage to the UAV, is the risk degree threshold.

5. The method according to claim 1, characterized in that, After using the historical UAV action strategy to control the movement of the UAV and obtaining historical real-time operation data during the movement, and determining the historical real-time operation data as sample data, the method further includes: Construct a priority flight experience storage pool and input the sample data into the priority flight experience storage pool; Select a preset number of experience sample data from the sample data in the priority flight experience storage pool in descending order of the priority and input them into the initial dual-policy learning network model to generate the expected Q value of the evaluation network and the target Q value of the target network, where the initial dual-policy learning network model includes the evaluation network and the target network; According to the target Q value and the expected Q value, the third formula is adopted: , to determine the TD target error value, where represents the TD target error value, represents the parameters of the target network, is the target Q value, Q is the expected Q value, represents the parameters of the evaluation network, represents the discount factor, which is used to represent the importance that the drone attaches to future rewards. A is the action of the drone in the moving direction, and the subscript j represents the jth experience sample data, represents the reward obtained by the drone, represents the feature vector of the drone position S, which is used to approximate the state value, The function means to find an action A to maximize the target Q value of the evaluation network Q; Adopt the fourth formula according to the TD target error value and the expected Q value of the evaluation network: , determine the loss function of the initial dual policy learning model, where is the loss function, represents the priority weight of the j-th empirical sample data, is the position of the UAV; optimize the parameters of the target network and the parameters of the evaluation network of the dual policy learning network according to the loss function to obtain the dual policy learning model.

6. The method according to claim 5, wherein After optimizing the parameters of the target network and the parameters of the evaluation network of the dual-policy learning network according to the loss function, the method further includes: Adopt the seventh formula: , recalculate the priority of the sample data in the priority flight experience storage pool, where represents the priority of the sample data, represents the TD target error value, is the expected Q value of the sample data j, represents the exploration rate.

7. An unmanned aerial vehicle path planning device at a steel bar binding site, characterized in that, Including: A construction unit for constructing a UAV path planning environment for the steel bar binding site with Markov properties, where the UAV path planning environment includes: an obstacle evaluation model, a UAV position model, a UAV movement direction model, and a reward function model; A first acquisition unit for acquiring the historical initial position and the historical termination position within a historical time period, and inputting the historical initial position, the historical termination position, and the UAV path planning environment model into the action selection strategy model to calculate the historical UAV action strategy, and using the historical UAV action strategy to control the movement of the UAV and obtaining historical real-time operation data during the movement, and determining the historical real-time operation data as sample data, where the sample data includes: UAV current moment position information, UAV current moment direction information, UAV next moment predicted position information, UAV current moment reward data, and the priority of the sample data; A second acquisition unit for acquiring the dual-policy learning network model, where the dual-policy learning network model is trained using the sample data; A prediction unit for inputting the initial position and the termination position of the UAV into the dual-policy learning network model for prediction to obtain the flight path of the UAV at the steel bar binding site; The first acquisition unit includes a division module and a control module. The division module is used to divide the UAV movement direction model according to the historical initial position and the historical termination position using a heuristic search algorithm to obtain a plurality of sub-movement directions; the control module is used to control the UAV action according to the plurality of sub-movement directions using a greedy strategy algorithm to generate a plurality of the historical UAV action strategies.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, where when the program runs, it controls the device where the computer-readable storage medium is located to execute the UAV path planning method for the steel bar binding site according to any one of claims 1 to 6.

9. An electronic device, characterized in that, Including: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include a method for unmanned aerial vehicle path planning at a steel bar binding site according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Unmanned aerial vehicle autonomous path planning method in unknown environment based on deep reinforcement learning

    CN116225046A