A UAV inspection scheduling method based on ant colony algorithm and related equipment

Through the drone inspection and scheduling method based on the ant colony algorithm, the drone inspection path is adjusted in real time, which solves the problems of waste of resources and poor environmental adaptability in traditional systems, and improves patrol efficiency and resource utilization.

CN118674227BActive Publication Date: 2025-05-06FOSHAN NANHAI DISTRICT BIG DATA INVESTMENT & CONSTR CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410893037.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-04
Publication Date
2025-05-06
Estimated Expiration
2044-07-04

AI Technical Summary

Technical Problem

In traditional drone inspection systems, resource waste is serious, resource utilization is low, and inspection routes cannot be adjusted in real time according to environmental changes, resulting in poor environmental adaptability.

Method used

The drone patrol and scheduling method based on the ant colony algorithm is adopted to obtain the status information of each grid area through grid division, and the pre-trained action planning model is used to calculate the reward function value of the action, plan the next action in real time, and adjust the patrol path adaptively.

Benefits of technology

It reduces the waste of aircraft nests and drone resources, improves urban inspection efficiency and resource utilization, and enhances environmental adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118674227B_ABST
    Figure CN118674227B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of unmanned aerial vehicle dispatching, and discloses a method for dispatching unmanned aerial vehicle inspection based on an ant colony algorithm and related equipment. On the basis of dividing the inspected city into grids, an action planning model pre-trained based on an ant colony algorithm is used to guide the unmanned aerial vehicle to perform inspection flights; each unmanned aerial vehicle does not need to be paired with each machine nest one by one, and each unmanned aerial vehicle does not need to patrol in a fixed area, which is beneficial to reducing the number of machine nests and unmanned aerial vehicles, improving the utilization rate of machine nests and unmanned aerial vehicles, and reducing the waste of machine nests and unmanned aerial vehicle resources; in addition, the unmanned aerial vehicle does not need to return to a fixed machine nest for charging during the patrol process, which saves time, improves the efficiency of urban inspection, and reduces the waste of unmanned aerial vehicle resources caused by the inability to effectively utilize the unmanned aerial vehicle due to returning for charging, thereby improving resource utilization; in addition, each unmanned aerial vehicle plans the next action to be executed in real time according to the current grid status information of the grid area. When the grid status information of each grid area changes, the inspection path of the unmanned aerial vehicle can be adaptively adjusted in time according to such changes, and has good environmental adaptability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of drone scheduling, and more specifically, to a drone inspection scheduling method and related equipment based on an ant colony algorithm. Background Art

[0002] The integrated air-space-ground social governance work requires the use of drones to patrol the entire city and promptly discover various urban illegal, irregular, and unauthorized construction activities.

[0003] In the traditional drone inspection system, drones and nests are paired one by one. Drones only patrol in fixed areas near the corresponding nests, and can only return to the corresponding nests for charging when they need to be charged. In order to complete the patrol and aerial photography tasks within the specified period, more resources need to be invested in building nests and purchasing drones, which causes serious waste. In addition, due to the limited battery life of drones, they need to be charged multiple times. Patrolling in a fixed area causes drones to fly back and forth, wasting a lot of time and battery life (generally, the battery life of returning to charge cannot be effectively used for inspection work, so the battery life is wasted). In addition, the drone resources that can be mobilized in case of emergencies are limited. will also be restricted, resulting in low resource utilization; in addition, the existing UAV scheduling methods are mostly static scheduling methods, that is, the inspection route is planned and assigned to the UAV before the UAV starts the inspection, and the UAV then inspects according to the assigned inspection route. It is impossible to adjust the inspection route in real time according to environmental changes (for example, during the UAV inspection process, due to weather changes or other reasons, some areas change from the original allowed flight status to temporarily not allowed flight status, but the UAV can only continue to fly according to the original inspection route into the temporarily not allowed flight area, which is easy to cause damage to the UAV or cause other accidents), and the environmental adaptability is poor. .

[0004] Therefore, it is necessary to find a UAV inspection scheduling method to reduce the waste of machine nests and UAV resources, improve urban inspection efficiency and resource utilization, and have good environmental adaptability. Summary of the invention

[0005] The purpose of this application is to provide a drone inspection scheduling method and related equipment based on an ant colony algorithm, which can reduce the waste of machine nests and drone resources, improve urban inspection efficiency and resource utilization, and have good environmental adaptability.

[0006] In a first aspect, the present application provides a method for scheduling drone inspection based on an ant colony algorithm, which is applied to a drone of a drone inspection system, wherein the drone inspection system includes multiple drones and multiple machine nests; the method for scheduling drone inspection based on an ant colony algorithm includes the following steps:

[0007] A1. Obtain the current grid status information of each grid area of ​​the inspected city after being divided into grids; the grid status information includes non-fly-free status information indicating that the drone is allowed to fly in the corresponding grid area and no-fly status information indicating that the drone is not allowed to fly in the corresponding grid area;

[0008] A2. Obtain the current reference state information, and obtain the current action space of the UAV according to the grid state information; the reference state information includes the inspection task completion information, the grid position of the UAV and the remaining endurance distance of the UAV; the action space includes the allowed actions, and the allowed actions refer to flying toward the center of the grid area adjacent to the grid area where the UAV is currently located and the current grid state information is non-no-fly state information;

[0009] A3. According to the current reference state information and the current action space of the drone, the reward function value of each action in the current action space of the drone is calculated using the action planning model pre-trained based on the ant colony algorithm;

[0010] A4. Select the action with the largest reward function value in the current action space of the drone as the target action to be executed next by the drone.

[0011] By adopting the above method to conduct drone inspection and scheduling, each drone does not need to be paired with each machine nest one by one, and each drone does not need to patrol in a fixed area, which is conducive to reducing the number of machine nests and drones, improving the utilization rate of machine nests and drones, and reducing the waste of machine nests and drone resources; in addition, the drone does not need to return to the fixed machine nest for charging during the patrol process, which saves time, improves the efficiency of urban inspection, and reduces the waste of drone resources caused by the inability to effectively utilize drones due to returning for charging, thereby improving resource utilization; in addition, each drone plans the next action to be executed in real time according to the current grid status information of the grid area. When the grid status information of each grid area changes, the drone's inspection path can be adaptively adjusted in time according to such changes, and has good environmental adaptability.

[0012] Preferably, the inspected city is divided into grids in the following manner:

[0013] After obtaining the minimum circumscribed rectangle of the border line of the inspected city, the minimum circumscribed rectangle is grid-divided to obtain a plurality of grid areas;

[0014] The no-fly status information includes first status information indicating that the corresponding grid area is a no-fly zone or an out-of-city area and second status information indicating that the corresponding grid area is temporarily not suitable for flying. The non-no-fly status information includes third status information indicating that the corresponding grid area has completed inspection and fourth status information indicating that the corresponding grid area has not completed inspection.

[0015] By dividing the inspected city into grids and planning the inspection route based on the grid division results, the inspection route planning process can be simplified and the planning efficiency can be improved. At the same time, inspection route planning can be carried out according to the grid status information of each grid area. On the one hand, it can prevent drones from entering areas where flight is not allowed, and on the other hand, it is conducive to improving the rationality of the inspection route planning results.

[0016] Preferably, the action planning model is trained in the following manner:

[0017] C1. Initialize the network parameters of the target network and the prediction network, initialize the update number Up to 0, and initialize the iteration number Dn to 0; the network structure of the target network and the prediction network is the same;

[0018] C2. Randomly extract multiple historical sample data from the experience database to form a historical sample data set; the historical sample data includes historical first reference state information, historical actions, historical reward function values, historical second reference state information and historical back-step action space, the historical first reference state information is the reference state information before executing the historical action, the historical second reference state information is the reference state information after executing the historical action, and the historical back-step action space includes the next action allowed to be executed after executing the historical action;

[0019] C3. Input the historical second reference state information in each of the historical sample data and the next action allowed to be executed in the historical next-step action space into the target network, obtain the reward function value output by the target network, record it as the initial target value, and use it to calculate the target value;

[0020] C4. Inputting the historical first reference state information and the historical action in each of the historical sample data into the prediction network, obtaining the reward function value output by the prediction network, recorded as the prediction value;

[0021] C5. Calculate the loss function according to the target value and the predicted value to update the network parameters of the prediction network, and set Up=Up+1, Dn=Dn+1;

[0022] C6. If Up reaches the preset update number threshold, the updated network parameters of the predicted network are copied to the target network to update the network parameters of the target network, and Up=0;

[0023] C7. If Dn reaches the preset maximum number of iterations, stop training and use the updated target network as the trained action planning model; otherwise, return to step C2.

[0024] The action planning model trained according to the above method, when used, only needs to input the current reference state information and the action of the current action space of the UAV to quickly obtain the reward function value of each action, so as to quickly determine which action is the optimal action and ensure that the UAV flies along the optimal inspection path.

[0025] Preferably, in step C3, the target value is calculated in the following manner:

[0026] If the grid position in the historical second reference state information is the grid position corresponding to the machine nest, then let , otherwise, let ,in, is the target value, is the discount factor, is the historical second reference state information, is the next action allowed to be executed in the historical next-step action space, To correspond to and The initial target value of is the historical reward function value.

[0027] Preferably, in step C5, the loss function is calculated according to the following formula:

[0028] ;

[0029] in, is the loss function, is the predicted value, is the averaging function.

[0030] Preferably, before step C1, the method further comprises the following steps:

[0031] C0. Collect historical sample data based on ant colony algorithm to build an experience database.

[0032] Preferably, step C0 comprises:

[0033] D1. Obtain the initial grid position of each UAV;

[0034] D2. Establishing an ant colony including multiple ant groups, each ant group includes multiple ants, and the ants in each ant group correspond to each drone one by one;

[0035] D3. Initialize the heuristic pheromone, pheromone factor, heuristic information factor, pheromone increment coefficient, pheromone volatility coefficient and the pheromone concentration of the ants corresponding to each drone, and initialize the number of iterations T to 0;

[0036] D4. Initializing the position of each ant according to the initial grid position of each drone;

[0037] D5. planning the inspection path of each ant in each ant group according to the heuristic pheromone, the pheromone factor, the heuristic information factor and the pheromone concentration, and forming an inspection plan corresponding to each ant group;

[0038] D6. Calculate the fitness function of the inspection scheme corresponding to each ant group;

[0039] D7. According to the fitness function of the inspection scheme corresponding to each ant group, select a number of elite ant groups from each ant group, and update the pheromone concentration of the elite ant group according to the pheromone increment coefficient and the pheromone volatility coefficient;

[0040] D8. Select the two inspection schemes with the largest and smallest fitness functions as sampling inspection schemes, and trace back the execution process of each inspection path in the sampling inspection scheme to obtain multiple historical sample data and record them in the experience database;

[0041] D9. Let T = T + 1. If T reaches the preset upper limit of the number of iterations, stop the iteration; otherwise, return to execute step D4.

[0042] Preferably, step D5 includes sequentially taking each of the ant groups as a target ant group, taking the ants in the target ant group as target ants, and executing:

[0043] D501. Obtaining the real-time action space of each target ant;

[0044] D502. Calculate the state transition probability of each of the target ants according to the heuristic pheromone, the pheromone factor, the heuristic information factor, the pheromone concentration and the real-time action space;

[0045] D503. According to the state transition probability of each target ant, select the target grid area of ​​each target ant, and move each target ant to the corresponding target grid area;

[0046] D504. If the target grid area to which the target ant moves is the grid area where the machine nest is located, then the heuristic pheromone on the local path traveled by the corresponding target ant from the grid area where the last machine nest was located to the grid area where the current machine nest is located is updated;

[0047] D505. If all inspection tasks are completed, the grid areas that each target ant has walked through are used to form the inspection paths of each target ant, forming the inspection plan of the target ant group and ending the path planning. Otherwise, return to step D501.

[0048] In a second aspect, the present application provides an electronic device comprising a processor and a memory, wherein the memory stores a computer program executable by the processor, and when the processor executes the computer program, it runs the steps in the drone inspection scheduling method based on the ant colony algorithm as described in the preceding item.

[0049] In a third aspect, the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps in the drone inspection scheduling method based on the ant colony algorithm as described above are executed.

[0050] Beneficial effects: The drone inspection scheduling method and related equipment based on the ant colony algorithm provided in the present application do not need to be paired with each nest one by one, and each drone does not need to patrol in a fixed area, which is beneficial to reducing the number of nests and drones, improving the utilization rate of nests and drones, and reducing the waste of nests and drone resources; in addition, the drone does not need to return to a fixed nest for charging during the patrol process, which saves time, improves the efficiency of urban inspections, and reduces the waste of drone resources caused by the inability to effectively utilize drones due to returning for charging, thereby improving resource utilization; in addition, each drone plans the next action to be executed in real time according to the current grid status information of the grid area. When the grid status information of each grid area changes, the drone's inspection path can be adaptively adjusted in time according to such changes, and has good environmental adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 A flowchart of a drone inspection scheduling method based on an ant colony algorithm provided in an embodiment of the present application.

[0052] Figure 2 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0053] Description of reference numerals: 301, processor; 302, memory; 303, communication bus. DETAILED DESCRIPTION

[0054] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the application claimed for protection, but merely represents the selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative work belong to the scope of protection of the present application.

[0055] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.

[0056] Please refer to Figure 1 , Figure 1 A method for scheduling drone inspection based on an ant colony algorithm in some embodiments of the present application is applied to a drone in a drone inspection system, wherein the drone inspection system includes multiple drones and multiple nests (these nests are distributed in different locations of the inspected city); the method for scheduling drone inspection based on an ant colony algorithm includes the following steps:

[0057] A1. Obtain the current grid status information of each grid area of ​​the inspected city after being divided into grids; the grid status information includes non-fly-free status information indicating that the drone is allowed to fly in the corresponding grid area and no-fly status information indicating that the drone is not allowed to fly in the corresponding grid area;

[0058] A2. Obtain the current reference state information, and obtain the current action space of the UAV based on the grid state information; the reference state information includes the inspection task completion status information, the grid position of the UAV, and the remaining endurance distance of the UAV; the action space includes the actions that are allowed to be executed, and the actions that are allowed to be executed are to fly toward the center of the grid area adjacent to the grid area where the UAV is currently located and the current grid state information is not the no-fly state information;

[0059] A3. According to the current reference state information and the current action space of the UAV, the action planning model based on the ant colony algorithm pre-training is used to calculate the reward function value of each action in the current action space of the UAV;

[0060] A4. Select the action with the largest reward function value in the current action space of the drone as the target action for the next step of the drone.

[0061] By adopting the above method to conduct drone inspection and scheduling, each drone does not need to be paired with each nest one by one, and each drone does not need to patrol in a fixed area, which is conducive to reducing the number of nests and drones, improving the utilization rate of nests and drones, and reducing the waste of nests and drone resources; in addition, the drone does not need to return to the fixed nest for charging during the patrol process, which saves time, improves the efficiency of urban inspection, and reduces the waste of drone resources caused by the inability to effectively utilize the drone due to returning for charging, thereby improving resource utilization; in addition, each drone plans the next action to be executed in real time according to the current grid status information of the grid area. When the grid status information of each grid area changes (for example, the local area changes from the original non-flying ban state to the flying ban state due to weather changes), the drone's inspection path can be adaptively adjusted in time according to this change, and has good environmental adaptability.

[0062] The inspected cities are divided into grids in the following ways:

[0063] After obtaining the minimum bounding rectangle of the border line of the inspected city, the minimum bounding rectangle is grid-divided to obtain a plurality of grid areas.

[0064] By dividing the inspected city into grids and planning the inspection route based on the grid division results, the inspection route planning process can be simplified and the planning efficiency can be improved.

[0065] Among them, all grid areas include grid areas located outside the city under inspection; grid division by the above method is conducive to ensuring that the sizes of various grid areas are the same, so that the center distances between adjacent grids are equal, which is conducive to reducing the amount of calculation in the path planning process.

[0066] When the minimum enclosing rectangle is meshed, the minimum enclosing rectangle can be meshed using horizontal grid lines parallel to the horizontal sides of the minimum enclosing rectangle and vertical grid lines parallel to the vertical sides of the minimum enclosing rectangle, and the spacing between the horizontal grid lines is equal to the spacing between the vertical grid lines (the side length of the minimum enclosing rectangle can be appropriately increased to ensure that all mesh areas obtained by the division are square areas). The spacing between the horizontal grid lines and the spacing between the vertical grid lines can be set according to actual needs. By dividing, the minimum enclosing rectangle is divided into multiple grid areas, and the grid matrix composed of all grid areas can be called a full grid matrix, in which each grid area contains at most one machine nest.

[0067] The center point of each grid area is the shooting point of the aircraft inspection (also called inspection point or mission point).

[0068] Among them, the grid position of the drone refers to the center point position of the grid area where the drone is located.

[0069] Before executing the inspection task, each drone is docked in the corresponding nest. The nest where the drone is docked before starting the inspection task is called the initial docked nest. The number of drones that can be docked in each nest is determined by the specific model of the nest. Therefore, the initial docked nests of each drone can be the same or different. When starting the inspection task, the grid position of the drone in the current reference state information refers to the center point of the grid area where the initial docked nest of the drone is located.

[0070] Among them, the performance (such as flight speed, maximum cruising range, etc.) of each drone in the drone inspection system can be the same or different.

[0071] In some embodiments, the no-fly status information includes first status information indicating that the corresponding grid area is a no-fly zone (e.g., an airport area, a military facility area, etc.) or an out-of-city area and second status information indicating that the corresponding grid area is temporarily unsuitable for flying (e.g., temporarily unsuitable for flying due to bad weather), and the non-no-fly status information includes third status information indicating that the corresponding grid area has completed inspection and fourth status information indicating that the corresponding grid area has not completed inspection. Inspection path planning is performed based on the grid status information of each grid area, which can prevent the drone from entering an area where flight is not allowed, and is conducive to improving the rationality of the inspection path planning results.

[0072] The grid status information may be represented by symbols, letters or numerical values, for example, the first status information is represented by -2, the second status information is represented by -1, the third status information is represented by 0, and the fourth status information is represented by 1, but not limited thereto; thus, the inspection task completion information may be represented by the following task status matrix:

[0073] ;

[0074] Among them, X is the task state matrix, It is the grid state information of the grid area in the i-th row and j-th column in the minimum bounding rectangle.

[0075] Utilizing the above grid status information can help guide drones to avoid no-fly zones and areas that are temporarily unsuitable for flying, making inspection plans more reasonable and safer.

[0076] It should be noted that since all drones plan the next action based on the same inspection task completion information, the drone path planning process is not independent of each other, but is interactive through the inspection task completion information, thereby ensuring the optimization effect of the inspection path of all drones. Among them, the inspection task completion information can be updated by communication between all drones to achieve consistency, or it can be uniformly issued by a central control device (the central control device is connected to all drones and all machine nests) to achieve consistency.

[0077] Among them, the remaining cruising distance of the UAV can be calculated based on the maximum cruising distance of the UAV and the executed actions of the UAV. Specifically, the flight distance corresponding to each action is known (for example, the flight distance of the action of moving horizontally and the action of moving vertically is the side length of a grid area, and the flight distance of the action of moving diagonally in the grid area is the diagonal length of a grid area). The flight distances corresponding to the executed actions are accumulated to obtain the total flight distance, and then the total flight distance is subtracted from the maximum cruising distance to obtain the remaining cruising distance.

[0078] Specifically, the actions allowed to be executed in the action space may include part or all of the actions of moving one grid horizontally, moving one grid vertically, and moving one grid diagonally in the grid area.

[0079] For example, assuming that the horizontal grid lines and the vertical grid lines extend in the east-west direction and the north-south direction respectively, when there are 8 adjacent grid areas, and all of them are grid areas that are allowed to enter (i.e., grid areas whose grid state information is the third state information or grid areas whose grid state information is the fourth state information), then the action space can be expressed as {Nth, Sth, Wst, Est, NE, NW, SE, SW}, where action Nth means that the drone moves from the center point of the current grid area to the north to the center point of an adjacent grid area (belonging to the action of moving in the longitudinal direction); action Sth means that the drone moves from the center point of the current grid area to the south to the center point of an adjacent grid area (belonging to the action of moving in the longitudinal direction); action Wst means that the drone moves from the center point of the current grid area to the west to the center point of an adjacent grid area (belonging to the action of moving in the lateral direction). Action Est means that the drone moves from the center point of the current grid area to the east to the center point of an adjacent grid area (a horizontal movement); Action NE means that the drone moves from the center point of the current grid area to the northeast to the center point of an adjacent grid area (a diagonal movement); Action NW means that the drone moves from the center point of the current grid area to the northwest to the center point of an adjacent grid area (a diagonal movement); Action SE means that the drone moves from the center point of the current grid area to the southeast to the center point of an adjacent grid area (a diagonal movement); Action SW means that the drone moves from the center point of the current grid area to the southwest to the center point of an adjacent grid area (a diagonal movement). If the adjacent grid area in the north is not allowed to enter, action Nth is eliminated from the action space.

[0080] Preferably, the action planning model is trained in the following manner (i.e., before step A1, the following steps are also included):

[0081] C1. Initialize the network parameters of the target network and the prediction network, initialize the number of updates Up to 0, and initialize the number of iterations Dn to 0; the network structures of the target network and the prediction network are the same;

[0082] C2. Randomly extract multiple historical sample data from the experience database to form a historical sample data set; the historical sample data includes the historical first reference state information, historical actions, historical reward function values, historical second reference state information and historical back-step action space. The historical first reference state information is the reference state information before executing the historical action, the historical second reference state information is the reference state information after executing the historical action, and the historical back-step action space includes the next action allowed to be executed after executing the historical action;

[0083] C3. Input the historical second reference state information in each historical sample data and the next allowed action in the historical next-step action space into the target network, obtain the reward function value output by the target network, record it as the initial target value, and use it to calculate the target value;

[0084] C4. Input the historical first reference state information and historical actions in each historical sample data into the prediction network to obtain the reward function value output by the prediction network, which is recorded as the prediction value;

[0085] C5. Calculate the loss function based on the target value and the predicted value to update the network parameters of the prediction network and set Up=Up+1, Dn=Dn+1;

[0086] C6. If Up reaches the preset update number threshold, the updated network parameters of the predicted network are copied to the target network to update the network parameters of the target network (at this time, the network parameters of the predicted network and the target network are the same), and Up=0;

[0087] C7. If Dn reaches the preset maximum number of iterations, stop training and use the updated target network as the trained action planning model; otherwise, return to step C2.

[0088] The action planning model trained according to the above method only needs to input the current reference state information and the action of the current action space of the UAV to quickly obtain the reward function value of each action, so as to quickly determine which action is the optimal action and ensure that the UAV flies along the optimal inspection path.

[0089] In step C1, when initializing the network parameters of the target network and the prediction network, the network parameters can be made to obey a Gaussian distribution with an expected value of 0 and a standard deviation of 0.2. The target network and the prediction network are deep learning networks, and the specific network structure type can be selected according to actual needs.

[0090] Among them, the experience database records a large amount of historical sample data, which can be recorded as , It is the first historical reference status information. For historical action, is the historical reward function value (indicated in Execute historical actions under the conditions Rewards received), is the historical second reference status information, is the historical next-step action space, Include at least one , After executing the historical action, the next action allowed to be executed (i.e. is the action allowed to be executed in the next step in the historical back-step action space).

[0091] Among them, in step C3, the actions allowed to be executed in the next step in the historical next step action space in the historical sample data and the historical second reference state information are used to form input data, and the input data is input into the target network to obtain the corresponding initial target value. Further, the target value is calculated by the following method:

[0092] If the grid position in the historical second reference state information is the grid position corresponding to the machine nest (that is, the center point position of the grid area where the machine nest is located), then let , otherwise, let ,in, is the target value, is the discount coefficient (can be set according to actual needs), is the historical second reference status information, is the next action allowed in the historical back-step action space, To correspond to and The initial target value of is the historical reward function value. Among them, Indicates that the drone is performing historical actions After that, the maximum reward that can be obtained in the next step is calculated by the above method, which comprehensively considers the reward brought by the current action step and the maximum reward that can be obtained in the next action step caused by the current action step, so that the cumulative reward of the drone's inspection path can be larger, and the drone can be more effectively guided to fly along a better path.

[0093] In step C5, the loss function is calculated according to the following formula:

[0094] ;

[0095] in, is the loss function, is the predicted value, is the averaging function.

[0096] In step C5, the network parameters of the prediction network may be updated using a gradient descent method, but is not limited thereto.

[0097] Furthermore, before step C1, the method further includes the following steps:

[0098] C0. Collect historical sample data based on ant colony algorithm to build an experience database.

[0099] Before the action planning model is trained, an ant colony algorithm based on pheromone concentration and heuristic information is used to schedule drone inspections, and historical sample data is collected to build an experience database.

[0100] Specifically, step C0 includes:

[0101] D1. Obtain the initial grid position of each UAV (i.e. the center point of the grid area where the initial docking nest is located);

[0102] D2. Establishing an ant colony including multiple ant groups, each ant group includes multiple ants, and the ants in each ant group correspond to each drone one by one;

[0103] D3. Initialize the pheromone inspiration, pheromone factor, pheromone information factor, pheromone increment coefficient, pheromone volatility coefficient and pheromone concentration of the ants corresponding to each drone, and initialize the number of iterations T to 0;

[0104] D4. Initialize the position of each ant according to the initial grid position of each drone;

[0105] D5. Plan the inspection path of each ant in each ant group according to the heuristic pheromone, pheromone factor, heuristic information factor and pheromone concentration, and form an inspection plan corresponding to each ant group (that is, the inspection plan of each ant group includes the inspection path of each ant in the ant group);

[0106] D6. Calculate the fitness function of the inspection plan corresponding to each ant group;

[0107] D7. According to the fitness function of the inspection scheme corresponding to each ant group, select a number of elite ant groups (the specific number can be set according to actual needs) from each ant group, and update the pheromone concentration of the elite ant group according to the pheromone increment coefficient and the pheromone volatility coefficient;

[0108] D8. Select the two inspection plans with the largest and smallest fitness functions as sampling inspection plans, and trace back the execution process of each inspection path in the sampling inspection plan to obtain multiple historical sample data and record them in the experience database;

[0109] D9. Let T = T + 1. If T reaches the preset upper limit of the number of iterations, stop the iteration; otherwise, return to execute step D4.

[0110] Wherein, step D5 includes taking each ant group as the target ant group, taking the ants in the target ant group as the target ants, and executing:

[0111] D501. Obtain the real-time action space of each target ant;

[0112] D502. Calculate the state transition probability of each target ant according to the heuristic pheromone, pheromone factor, heuristic information factor, pheromone concentration and real-time action space; for example, the state transition probability of each target ant can be calculated according to the following formula:

[0113] ;

[0114] in, is the state transition probability of the kth target ant moving from the i-th grid area (representing the grid area where the kth target ant is currently located) to the j-th grid area, is the residual pheromone concentration of the k-th target ant between the i-th grid area and the j-th grid area, is the residual pheromone concentration of the k-th target ant between the i-th grid area and the s-th grid area, is the heuristic information from the i-th grid area to the j-th grid area, is the inspiration information from the i-th grid area to the s-th grid area, Pheromone factor, is the heuristic information factor, is the set of grid areas pointed to by the action of the kth target ant in the real-time action space, is the remaining endurance distance of the kth target ant, is the minimum flight distance from the i-th grid area to the j-th grid area (which can be calculated based on the Chebyshev distance calculation formula);

[0115] D503. According to the state transition probability of each target ant, select the target grid area of ​​each target ant, and move each target ant to the corresponding target grid area; for example, according to the state transition probability of each target ant, use the blocking roulette method to select the target grid area of ​​each target ant;

[0116] D504. If the target grid area to which the target ant moves is the grid area where the nest is located, then the heuristic pheromone on the local path that the corresponding target ant travels from the grid area where the last nest is located to the grid area where the current nest is located is updated; for example, the heuristic pheromone on the local path is updated according to the following formula:

[0117] ;

[0118] in, is the inspiration pheromone from the rth grid area to the r+1th grid area on the local path, R is the number of grid areas on the local path, is the number of completed tasks corresponding to the local path (equal to the number of grid areas on the local path whose grid state information is the fourth state information), is the waiting time of the corresponding target ant on the current docked nest (including the time for data transmission and battery replacement);

[0119] D505. If all inspection tasks are completed (i.e., the grid status information of all grid areas does not contain the fourth status information), the inspection paths of each target ant are composed of the grid areas passed by each target ant, forming an inspection plan for the target ant group and ending the path planning. Otherwise, return to step D501.

[0120] Preferably, in step D6, the fitness function of the inspection scheme corresponding to each ant group can be calculated according to the following formula:

[0121] ;

[0122] in, is the fitness function of the jth inspection plan, The task time required for the i-th UAV to complete the corresponding inspection task in the j-th inspection plan (i.e., to complete the inspection task of the corresponding inspection path in the inspection plan) (mainly including flight time, data transmission time and battery replacement time, the flight time can be calculated according to the length of the inspection path, the data transmission time can be calculated according to the number of grid areas whose grid status information on the inspection path is the fourth status information, and the battery replacement time is the sum of the battery replacement time of each nest parked on the inspection path, and the battery replacement time of each nest is a fixed known value), is a proportional constant (which can be set according to actual needs). The larger the fitness function is, the shorter the total time required to complete all inspection tasks in the inspected city is, which is conducive to minimizing the total time required for the final inspection plan.

[0123] Among them, in step D7, the pheromone concentration of the elite ant group can be updated according to the following formula:

[0124] ;

[0125] in, is the updated pheromone concentration remaining between the ith grid area and the jth grid area of ​​the kth ant in the dth elite ant group, is the pheromone concentration before update of the kth ant in the dth elite ant group between the ith grid area and the jth grid area, is the pheromone volatility coefficient, is the pheromone increment coefficient, is the fitness function of the inspection plan of the dth elite ant group, represents an action in the inspection path of the kth ant belonging to the dth elite ant group moving from the i-th grid area to the j-th grid area ( is the action set of the inspection path of the kth ant in the dth elite ant group).

[0126] Among them, in step D8, each action in each inspection path in the sampling inspection plan, the reference state information before executing each action, the reference state information after executing each action, and the next action allowed to be executed after executing each action are traced back, and the reward function value corresponding to each action in each inspection path in the sampling inspection plan is calculated. Each action obtained by backtracking, as well as the reference state information before executing the corresponding action, the reference state information after executing the corresponding action, the corresponding reward function value, and the set of actions allowed to be executed in the next step after executing the corresponding action are respectively used as historical actions, historical first reference state information, historical second reference state information, historical reward function values, and historical back-step action space, to form historical sample data and record them in the experience database.

[0127] Among them, the reward function value corresponding to each action in each inspection path in the sampling inspection scheme can be calculated by the following formula:

[0128] ;in, is the reward function value corresponding to the action executed at the tth time step in the inspection path in the sample inspection plan, is the fitness function of the sampling inspection scheme, e is the base of the natural logarithm, is a preset exponential parameter; by combining the above-mentioned reward function value constructed with the exponential penalty function based on the time step (i.e., t), a different reward function can be assigned to each step of the ant, and the workload between the ants is encouraged to be balanced, so as to achieve the effect of completing more inspection tasks in the shortest time.

[0129] Among them, in step D8, the two second inspection plans with the largest and smallest fitness functions are selected as sampling inspection plans to generate historical sample data, which can ensure the diversity of historical sample data in the experience database and is beneficial to improving the robustness of the action planning model when used to train the action planning model.

[0130] Specifically, step A3 includes:

[0131] The actions in the current action space of the UAV and the current reference state information are used as input data in sequence, and the input data is input into the action planning model pre-trained based on the ant colony algorithm to obtain the reward function value of each action in the current action space of the UAV outputted by the action planning model.

[0132] Furthermore, after step A4, the method further comprises the following steps:

[0133] A5. After executing the target action, update the grid status information of each grid area.

[0134] Specifically, the grid state information of the grid area where the drone is located after executing the target action is updated to the third state information, and the update result is broadcast to other drones so that other drones can update the grid state information, or the update result is sent to the central control device, which is forwarded by the central control device to other drones so that other drones can update the grid state information.

[0135] After step A5, if all inspection tasks are completed (ie, the grid status information of all grid areas does not contain the fourth status information), the inspection is stopped. At this time, each drone returns to the nearest nest to wait for the next inspection.

[0136] As can be seen from the above, the drone inspection scheduling method based on the ant colony algorithm obtains the current grid state information of each grid area of ​​the inspected city after grid division; the grid state information includes non-no-fly state information indicating that the drone is allowed to fly in the corresponding grid area and no-fly state information indicating that the drone is not allowed to fly in the corresponding grid area; obtains the current reference state information, and obtains the current action space of the drone based on the grid state information; the reference state information includes the completion status information of the inspection task, the grid position of the drone and the remaining endurance distance of the drone; the action space includes the actions that are allowed to be executed, and the actions that are allowed to be executed refer to flying toward the center of the grid area adjacent to the grid area where the drone is currently located and whose current grid state information is non-no-fly state information; according to the current reference state information and the current action space of the drone, the reward function value of each action in the current action space of the drone is calculated using the action planning model pre-trained based on the ant colony algorithm; the action with the largest reward function value in the current action space of the drone is selected as the target action to be executed by the drone next; thereby reducing the waste of machine nests and drone resources, improving the efficiency and resource utilization of urban inspections, and having good environmental adaptability. In addition, it also has the following advantages:

[0137] The present invention uses an ant colony algorithm to simulate the process of interaction between an intelligent agent and the environment, and collects environmental information, action and reward data. In the process of collecting data, the ant colony algorithm guides the decision-making process through pheromones, thereby solving the problems of slow learning speed and unstable learning of the intelligent agent. The ant colony algorithm continuously optimizes the overall inspection plan to generate a better strategy. The environmental information, action and reward data generated by the decision-making process of the better strategy are back-traced to build an experience database. The high-quality data in the experience database is used to train the path planning model, so that the path planning model can make better decisions.

[0138] Please refer to Figure 2 , Figure 2 This is a structural schematic diagram of an electronic device provided in an embodiment of the present application. The present application provides an electronic device, including: a processor 301 and a memory 302. The processor 301 and the memory 302 are interconnected and communicate with each other through a communication bus 303 and / or other forms of connection mechanisms (not shown). The memory 302 stores a computer program executable by the processor 301. When the electronic device is running, the processor 301 executes the computer program to execute the drone inspection scheduling method based on the ant colony algorithm in any optional implementation of the above-mentioned embodiment to achieve the following functions: obtaining the current grid status information of each grid area of ​​the inspected city after grid division; the grid status information includes non-no-fly status information indicating that the drone is allowed to fly in the corresponding grid area and non-no-fly status information indicating that the drone is not allowed to fly in the corresponding grid area. The method comprises the following steps: obtaining the no-fly status information of the grid area where the drone is currently located, obtaining the current reference status information, and obtaining the current action space of the drone according to the grid status information; the reference status information includes the completion status information of the inspection task, the grid position of the drone, and the remaining cruising range of the drone; the action space includes actions that are allowed to be executed, and the actions that are allowed to be executed refer to flying toward the center of the grid area adjacent to the grid area where the drone is currently located and whose current grid status information is non-no-fly status information; according to the current reference status information and the current action space of the drone, the action planning model pre-trained based on the ant colony algorithm is used to calculate the reward function value of each action in the current action space of the drone; and selecting the action with the largest reward function value in the current action space of the drone as the target action to be executed by the drone next.

[0139] The embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the unmanned aerial vehicle inspection scheduling method based on the ant colony algorithm in any optional implementation of the above embodiment is executed to achieve the following functions: obtaining the current grid state information of each grid area of ​​the inspected city after grid division; the grid state information includes non-no-fly state information indicating that the unmanned aerial vehicle is allowed to fly in the corresponding grid area and no-fly state information indicating that the unmanned aerial vehicle is not allowed to fly in the corresponding grid area; obtaining the current reference state information, and obtaining the current action space of the unmanned aerial vehicle according to the grid state information; the reference state information includes the completion status information of the inspection task, the grid position of the unmanned aerial vehicle and the remaining endurance distance of the unmanned aerial vehicle; the action space includes the actions that are allowed to be executed, and the actions that are allowed to be executed refer to flying toward the center of the grid area adjacent to the grid area where the unmanned aerial vehicle is currently located and whose current grid state information is non-no-fly state information; according to the current reference state information and the current action space of the unmanned aerial vehicle, the reward function value of each action in the current action space of the unmanned aerial vehicle is calculated using the action planning model pre-trained based on the ant colony algorithm; and selecting the action with the largest reward function value in the current action space of the unmanned aerial vehicle as the target action to be executed by the unmanned aerial vehicle next.

[0140] Among them, the computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable red-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0141] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0142] In addition, the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, and may be located in one place or distributed on multiple network model units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0143] Furthermore, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.

[0144] In this document, relational terms such as first and second, etc. are used merely to distinguish one entity or operation from another entity or operation, but do not necessarily require or imply any such actual relationship or order between these entities or operations.

[0145] The above description is only an embodiment of the present application and is not intended to limit the protection scope of the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for dispatching drone inspection based on ant colony algorithm, applied to drones in a drone inspection system, wherein the drone inspection system includes a plurality of drones and a plurality of machine nests, and the drones are not paired one-to-one with the machine nests; characterized in that: The UAV inspection scheduling method based on ant colony algorithm comprises the following steps: A1. Obtain the current grid status information of each grid area of ​​the inspected city after being divided into grids; the grid status information includes non-no-fly status information indicating that drones are allowed to fly in the corresponding grid area and no-fly status information indicating that drones are not allowed to fly in the corresponding grid area; the no-fly status information includes first status information indicating that the corresponding grid area is a no-fly zone or an out-of-city area and second status information indicating that the corresponding grid area is temporarily unsuitable for flying, and the non-no-fly status information includes third status information indicating that the corresponding grid area has completed inspection and fourth status information indicating that the corresponding grid area has not completed inspection; A2. Obtain the current reference state information, and obtain the current action space of the UAV according to the grid state information; the reference state information includes the inspection task completion information, the grid position of the UAV and the remaining endurance distance of the UAV; the action space includes the allowed actions, and the allowed actions refer to flying toward the center of the grid area adjacent to the grid area where the UAV is currently located and the current grid state information is non-no-fly state information; A3. According to the current reference state information and the current action space of the drone, the reward function value of each action in the current action space of the drone is calculated using the action planning model pre-trained based on the ant colony algorithm; A4. Select the action with the largest reward function value in the current action space of the drone as the target action to be executed next by the drone; The action planning model is trained in the following way: C1. Initialize the network parameters of the target network and the prediction network, initialize the update number Up to 0, and initialize the iteration number Dn to 0; the network structure of the target network and the prediction network is the same; C2. Randomly extract multiple historical sample data from the experience database to form a historical sample data set; the historical sample data includes historical first reference state information, historical actions, historical reward function values, historical second reference state information and historical back-step action space, the historical first reference state information is the reference state information before executing the historical action, the historical second reference state information is the reference state information after executing the historical action, and the historical back-step action space includes the next action allowed to be executed after executing the historical action; C3. Input the historical second reference state information in each of the historical sample data and the next action allowed to be executed in the historical next-step action space into the target network, obtain the reward function value output by the target network, record it as the initial target value, and use it to calculate the target value; C4. Inputting the historical first reference state information and the historical action in each of the historical sample data into the prediction network, obtaining the reward function value output by the prediction network, recorded as the prediction value; C5. Calculate the loss function according to the target value and the predicted value to update the network parameters of the prediction network, and set Up=Up+1, Dn=Dn+1; C6. If Up reaches the preset update number threshold, the updated network parameters of the predicted network are copied to the target network to update the network parameters of the target network, and Up=0; C7. If Dn reaches the preset maximum number of iterations, stop training and use the updated target network as the trained action planning model, otherwise, return to step C2; Before step C1, the method further includes the following steps: C0. Collect historical sample data based on ant colony algorithm to build an experience database; Step C0 includes: D1. Obtain the initial grid position of each UAV; D2. Establishing an ant colony including multiple ant groups, each ant group includes multiple ants, and the ants in each ant group correspond to each drone one by one; D3. Initialize the heuristic pheromone, pheromone factor, heuristic information factor, pheromone increment coefficient, pheromone volatility coefficient and the pheromone concentration of the ants corresponding to each drone, and initialize the number of iterations T to 0; D4. Initializing the position of each ant according to the initial grid position of each drone; D5. planning the inspection paths of the ants in each ant group according to the heuristic pheromone, the pheromone factor, the heuristic information factor and the pheromone concentration, and forming an inspection plan corresponding to each ant group; D6. Calculate the fitness function of the inspection scheme corresponding to each ant group; D7. According to the fitness function of the inspection scheme corresponding to each ant group, select a number of elite ant groups from each ant group, and update the pheromone concentration of the elite ant group according to the pheromone increment coefficient and the pheromone volatility coefficient; D8. Select the two inspection schemes with the largest and smallest fitness functions as sampling inspection schemes, and trace back the execution process of each inspection path in the sampling inspection scheme to obtain multiple historical sample data and record them in the experience database; D9. Let T = T + 1. If T reaches the preset upper limit of the number of iterations, stop the iteration. Otherwise, return to step D4. In step D8, each action in each inspection path in the sampling inspection scheme, the reference state information before executing each action, the reference state information after executing each action, and the action allowed to be executed next after executing each action are backtracked, and the reward function value corresponding to each action in each inspection path in the sampling inspection scheme is calculated. Each action obtained by backtracking, the reference state information before executing the corresponding action, the reference state information after executing the corresponding action, the corresponding reward function value, and the set of actions allowed to be executed next after executing the corresponding action are respectively used as historical actions, historical first reference state information, historical second reference state information, historical reward function values, and historical back-step action space, forming historical sample data and recorded in the experience database; Among them, the reward function value corresponding to each action in each inspection path in the sampling inspection plan is calculated by the following formula: ;in, is the reward function value corresponding to the action executed at the tth time step in the inspection path in the sample inspection plan, is the fitness function of the sampling inspection scheme, e is the base of the natural logarithm, is the preset index parameter.

2. The method for unmanned aerial vehicle inspection and scheduling based on ant colony algorithm according to claim 1 is characterized in that: The inspected cities are divided into grids in the following manner: After obtaining the minimum circumscribed rectangle of the border line of the inspected city, the minimum circumscribed rectangle is grid-divided to obtain a plurality of grid areas.

3. The method for unmanned aerial vehicle inspection and scheduling based on ant colony algorithm according to claim 1 is characterized in that: In step C3, the target value is calculated in the following manner: If the grid position in the historical second reference state information is the grid position corresponding to the machine nest, then let , otherwise, let ,in, is the target value, is the discount factor, is the historical second reference state information, is the next action allowed to be executed in the historical next-step action space, For the corresponding and The initial target value of is the historical reward function value.

4. The method for dispatching unmanned aerial vehicle inspection based on ant colony algorithm according to claim 3 is characterized in that: In step C5, the loss function is calculated according to the following formula: ; in, is the loss function, is the predicted value, is the averaging function.

5. The method for unmanned aerial vehicle inspection and scheduling based on ant colony algorithm according to claim 1 is characterized in that: Step D5 includes sequentially taking each of the ant groups as a target ant group, taking the ants in the target ant group as target ants, and executing: D501. Obtaining the real-time action space of each target ant; D502. Calculate the state transition probability of each of the target ants according to the heuristic pheromone, the pheromone factor, the heuristic information factor, the pheromone concentration and the real-time action space; D503. According to the state transition probability of each target ant, select the target grid area of ​​each target ant, and move each target ant to the corresponding target grid area; D504. If the target grid area to which the target ant moves is the grid area where the machine nest is located, then the heuristic pheromone on the local path traveled by the corresponding target ant from the grid area where the last machine nest was located to the grid area where the current machine nest is located is updated; D505. If all inspection tasks are completed, the grid areas that each target ant has walked through are used to form the inspection paths of each target ant, forming the inspection plan of the target ant group and ending the path planning. Otherwise, return to step D501.

6. An electronic device, characterized in that: It includes a processor and a memory, the memory stores a computer program executable by the processor, and when the processor executes the computer program, it runs the steps in the drone inspection scheduling method based on the ant colony algorithm as described in any one of claims 1 to 5.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, the steps in the drone inspection scheduling method based on the ant colony algorithm as described in any one of claims 1 to 5 are executed.

Citation Information

Patent Citations

  • Multi-intelligent agent reinforcement learning path planning method based on ant colony algorithm

    CN112286203A

  • Multi-unmanned aerial vehicle edge calculation path optimization and dependent task scheduling optimization method and system

    CN116451934A

  • Multi-electric logistics vehicle scheduling method fusing charging constraint and capacity constraint

    CN117522088A