Unmanned aerial vehicle group dynamic optimization scheduling method for sudden inspection task

Through the dynamic optimization scheduling method of drone clusters combining deep reinforcement learning and multi-objective genetic algorithms, the complexity of multi-UAV task allocation in the dynamic environment of smart grids is solved, rapid response and efficient execution are achieved, and the overall efficiency and resource utilization of smart grid patrols are improved.

CN120069409APending Publication Date: 2025-05-30BEIJING UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510119739.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-25
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the dynamic environment of smart grids, the allocation of multi-UAV tasks faces complexity and uncertainty, especially in burst tasks and dynamic environments, it is difficult for the existing technology to achieve rapid response and efficient execution.

Method used

Deep reinforcement learning (DRL) combined with multi-objective genetic algorithm (NSGA-II) is used to design a dynamic optimization scheduling method for emergency patrol tasks. This method uses real-time perception of the drone and mission status, intelligently selects the drone that is most suitable for performing new tasks, and optimizes the task list to achieve fast response and efficient execution.

Benefits of technology

It realizes rapid and efficient inspection of the smart grid in a dynamic environment, reduces time and energy consumption costs, avoids resource waste, and improves the overall efficiency and resource utilization of drone inspection tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069409A_ABST
    Figure CN120069409A_ABST
Patent Text Reader

Abstract

The invention provides an unmanned aerial vehicle group dynamic optimization scheduling method for a sudden inspection task. According to the method, deep reinforcement learning and a multi-objective optimization algorithm are combined, and the multiple unmanned aerial vehicles distribute and execute tasks in parallel and synchronously according to constraint conditions, so that rapid and efficient task redistribution in a dynamic inspection environment is realized, and repeated task execution is avoided. Firstly, unmanned aerial vehicle and task states are sensed in real time by using a deep reinforcement learning algorithm, an unmanned aerial vehicle executing a new task is intelligently selected, and effective utilization of unmanned aerial vehicle resources is ensured. And then, a genetic algorithm NSGA-II is adopted to carry out optimization sorting on the task list of the selected unmanned aerial vehicle so as to minimize task conflicts and time energy consumption. According to the invention, the overall efficiency and resource utilization rate of the unmanned aerial vehicle inspection task in the smart grid scene are effectively improved, rapid and efficient inspection of the smart grid in a dynamic environment is realized, the time and energy consumption cost are minimized, and resource waste caused by long-time waiting or repeated execution of a new task is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of unmanned aerial vehicles. Background Art

[0002] As an important part of modern power systems, the efficient, stable, and safe operation of the smart grid is crucial for ensuring power supply and meeting user demands. As an important operation and maintenance method, inspection has profound significance for the maintenance and management of the smart grid. It is not only an important means to ensure the safe operation of the smart grid, improve reliability and stability, but also can provide data support for the optimization of the smart grid. Therefore, in the operation and maintenance management of the smart grid, great importance needs to be attached to the inspection work to ensure the timeliness and effectiveness of the inspection.

[0003] The high efficiency and wide coverage of unmanned aerial vehicles (UAVs) enable them to complete large-scale inspection tasks in a short time, significantly improving the operation efficiency and reducing costs. At the same time, UAV inspections have high safety and can perform inspection tasks in complex or dangerous environments, reducing the safety risks of inspection personnel. Therefore, UAV inspections are gradually becoming the preferred method for inspection operations in many fields, and their application scope and influence are continuously expanding.

[0004] With the continuous development of UAV inspection technology and the increasing application requirements, how to efficiently and reasonably allocate the inspection tasks of multiple UAVs has become a very important issue. The multi-UAV task allocation not only needs to consider the performance and capabilities of individual UAVs, but also needs to comprehensively consider the coordination and cooperation of the overall system to achieve the optimization of the overall inspection effect. UAV task allocation is to coordinate UAVs to complete designated partial tasks, which plays a crucial role in whether the tasks can be effectively completed. Multiple factors such as the performance of UAVs, task requirements, and environmental conditions need to be considered to ensure that the tasks can be completed efficiently and accurately. It can be regarded as allocating to each part of the system from a high-level perspective, so as to execute concurrent tasks through their collective behavior, which is a combinatorial optimization problem. The goal of UAV task allocation is to allocate tasks to UAVs to achieve the optimization objective function. Facing the multi-UAV task allocation problem, many current related studies are mostly aimed at static task allocation problems. For example, the auction algorithm is used to execute an auction process with multiple rounds of bidding to determine the allocation scheme, and some heuristic algorithms have also been applied to the UAV task planning problem.

[0005] However, during the actual inspection process of drones, due to the existence of many uncertain factors, such as weather changes, the suddenness and unpredictability of power grid failures, it may be necessary to optimize the dynamic task scheduling of drones during the inspection process. These sudden dynamic inspection tasks will bring higher difficulties to the task allocation problem. Dynamic task allocation mainly solves the problem of resource reallocation after the emergence of new tasks, enabling the multi-drone system to quickly respond to sudden information and targets. This is a dynamic decision-making problem that changes with time and environment, and its complexity is further enhanced. Therefore, the allocation of sudden tasks in a dynamic environment is one of the most challenging aspects of the task planning problem. More appropriate and scalable algorithms must be developed in terms of the selection and execution among drones to handle dynamic and sudden events. There are already some related studies on the dynamic task allocation problem of drones. For example, currently, some researchers have proposed a method for dynamic task allocation of drones based on deep reinforcement learning. First, the drone that discovers a new task sends a task request to all other drones. Based on the DRL request network, the capabilities of the drones and the task requirements are sensed, and a certain number of responsive drones are selected. Then, through the drone response network based on Q-learning, it is judged whether each of the selected responsive drones finally participates in the execution of the new task. Multiple drones jointly participate in and complete the current new task at the same time, thus completing the processing of the new task.

[0006] The dynamic allocation of drone tasks faces a complex and ever-changing flight environment, diverse task requirements, and limitations in the capabilities of the drones themselves, especially in dynamic scenarios that require rapid response and efficient execution. Traditional allocation methods based on fixed strategies require a series of allocation rules to be set in advance. These methods lack flexibility when facing uncertainties and dynamic changes and are increasingly difficult to meet the requirements of modern drone task execution. With the deepening of the research on drone swarm technology and the rapid development of artificial intelligence algorithms, current research tends to construct a more intelligent, flexible, and adaptable drone task dynamic allocation system that can learn the distribution laws of tasks, the performance characteristics of drones, and the interaction relationships between them from a large amount of data. However, the current methods have the following problems: 1. The existing technologies do not fully consider the drone task list after adding new tasks, do not represent the current task stage, only consider the current moment of adding new tasks, or only consider a single location factor for the insertion position of new tasks, or cannot change the order of the original task list. 2. The existing methods have poor adaptability to different numbers of drones and task scenarios, and since the number of tasks in different drone task lists is different, the state and action dimensions are not fixed. Therefore, the present invention introduces the deep reinforcement learning (DRL) algorithm combined with the multi-objective genetic algorithm as the direction for solving the dynamic task planning of drones, which is more suitable for application in the environment of sudden inspection tasks and can better solve problems.

[0007] Aiming at the problems faced by task replanning in the dynamic environment of smart grid, especially the problem that unreasonable task allocation is likely to lead to time delay and resource waste, the present invention proposes a dynamic optimization scheduling method for a swarm of unmanned aerial vehicles (UAVs) facing sudden inspection tasks. This method combines deep reinforcement learning and multi-objective optimization algorithms. Multiple UAVs allocate and execute tasks in parallel and synchronously according to the constraint conditions, aiming to achieve fast and efficient task reallocation in the dynamic inspection environment and avoid repeated execution of tasks. Specifically, first, the deep reinforcement learning algorithm is used to perceive the states of UAVs and tasks in real time, and intelligently select the most suitable UAV to execute the new task, ensuring the effective utilization of UAV resources. Subsequently, the Non-dominated Sorting Genetic Algorithm II (NSGA-II) with an elite strategy is adopted to optimize and sort the task lists of the selected UAVs to minimize task conflicts and time and energy consumption. This dynamic optimization decision-making process for a swarm of UAVs facing sudden tasks not only achieves high scalability in terms of the number and type of UAVs, but also realizes fast response and efficient execution of new tasks by comprehensively considering task priorities, UAVs, and task phases, thereby effectively improving the overall efficiency and resource utilization rate of UAV inspection tasks in the smart grid scenario. It realizes fast and efficient inspection of the smart grid in a dynamic environment, minimizes time and energy consumption costs, avoids waste of resources caused by long waiting or repeated execution of new tasks, and provides better support for smart grid inspection. Summary of the Invention

[0008] Assume that multiple UAVs perform tasks on multiple fixed targets in the smart grid. At the initial moment, the ground control center has developed a set of current optimal task allocation schemes according to the known environmental information, task information, and UAV information, including the tasks that each UAV needs to execute and the execution order, and the UAVs execute tasks according to this scheme. Now assume that during the flight, a sudden task appears, that is, the task state changes, and it is necessary to reallocate UAVs to execute the tasks, that is, it is impossible to execute tasks according to the initial scheme. Under this condition, the present invention designs a dynamic environment UAV swarm inspection scheduling optimization method for sudden inspection tasks that comprehensively considers the states of UAVs and tasks, specifically including the following functions:

[0009] 1. The present invention designs a dynamic optimization scheduling method for emergency inspection tasks. When the UAV swarm is performing the planned tasks and an emergency inspection task appears, and task re-planning is required, the present invention optimizes the UAV selection and task order, and completes the UAV optimization scheduling in the dynamic situation of the emergency task through task re-planning, so as to achieve a rapid response to new tasks and improve the overall system stability and efficiency of the inspection. This method includes a UAV selection decision module based on deep reinforcement learning and an execution order optimization module based on NSGA-II. By perceiving relevant states such as UAV position, remaining power, and task requirements, and comprehensively considering time energy consumption and priority factors, UAVs are selected to execute new tasks, and the remaining task lists of the selected UAVs are re-sorted and optimized. After continuous training, the final planning result is given, so as to achieve a rapid adjustment of the system after the appearance of new tasks, reduce the overall time energy consumption in the dynamic environment, improve the overall system performance, provide guarantee for the automation and intelligence of UAV power grid inspection, and the dynamic optimization scheduling process is as Figure 1 shown.

[0010] (1) The task starts, and the UAVs execute according to the existing task allocation and observe the environmental status;

[0011] (2) Determine whether a new task appears. If so, execute step (3); if not, jump to step (9);

[0012] (3) Obtain the new task status information and environmental information;

[0013] (4) Input the information into the DRL UAV selection decision module for training;

[0014] (5) Output the UAV number for executing the new task;

[0015] (6) Obtain the task list of the selected UAV from the existing task allocation results of the UAVs and encode this task list;

[0016] (7) Use NSGA-II to optimize the task execution order of the selected UAVs;

[0017] (8) Output the final task scheduling strategy for emergency inspection tasks;

[0018] (9) Complete all inspection tasks and execute the tasks according to the new planning scheme;

[0019] (10) The task is completed.

[0020] 2. The present invention designs a UAV selection decision module based on deep reinforcement learning. This module selects a UAV to execute the emerging new task based on the current environmental information. Assume that there are M UAVs in the intelligent power grid inspection area, and the UAV set U = {U1 , U 2 , …, U m , …, U M}, the unmanned aerial vehicle U m 's position coordinates Distribute N task points within the service area of the unmanned aerial vehicle. The task set P = {P 1 , P 2 , …, P n , …, P N}, the task P n 's position coordinates At a certain moment during the inspection process, the task status changes, and an emergency task P new appears, with position coordinates and the required execution time T new . The task priority of this task is Imp. Assume that there are no interdependencies between tasks, that is, each task can be executed independently of other tasks, and each task is completed by one unmanned aerial vehicle. The state s U observed by the DRL module at the current moment Q U is the remaining battery power of the unmanned aerial vehicle, and s T is the current state of the unmanned aerial vehicle. The state of each unmanned aerial vehicle can be divided into three types: idle, executing a task, and having a task to be done. The output of the module is the action a t = (a 1 , a 2 , …, a M ), that is, the optimal unmanned aerial vehicle selected for the current new task. The value of the selected unmanned aerial vehicle is 1, and the values of other unmanned aerial vehicles are 0. The dimension of this vector is M. In addition, a reward function that comprehensively considers time energy consumption and task priority is designed:

[0021]

[0022] In the formula, α, β, and γ are weight parameters, and E total is the sum of the overall energy consumption of all unmanned aerial vehicles during the inspection process, T total is the total task completion time, defined as the time when the last unmanned aerial vehicle returns to the starting point. T LAT is the waiting time for the new task, that is, the time from the appearance of the new task to the time when an unmanned aerial vehicle starts to execute this task. The larger the ratio of the priority to the waiting time, the faster the task with a high priority is executed, reducing the time delay of high-priority tasks, reflecting the differential processing of tasks with different priorities by the system, and enabling more reasonable task allocation. Using this reward function, the DRL model is guided to reduce time energy consumption and prioritize the completion of emergency tasks, avoiding the problem of a single optimization goal in traditional algorithms.

[0023] 3. The present invention designs a training process for a drone selection decision DRL model. It describes how the DRL model generates a state vector s by observing the environmental information and the real-time status of the drone and the task at each time step t. t , generates action a according to its own strategy t , when performing action a t Then get reward r from the environment t , and transfer to the next state s t+1 . Get the four-tuple training sample t ,a t ,r t ,s t+1 >Exists in the experience pool. After completing the assigned tasks, randomly select samples from the experience pool to train the DRL model. The training process is as follows Figure 2 shown.

[0024] (1) The DRL model obtains the current new mission and the status information of the UAV, including location, remaining battery power, required time, etc.

[0025] (2) Form the input vector of the DRL model at time t

[0026] (3) Generates t Action t =π(s t )=(a 1 ,a 2 ,…,a M );

[0027] (4) Execute the action and optimize the order of the drone task list;

[0028] (5) Get the current reward r t , and enter the next state s t+1 ;

[0029] (6) Forming training samples t ,a t ,r t ,s t+1 >Save to experience pool;

[0030] (7) Determine whether all tasks are completed. If so, proceed to step (8); otherwise, return to step (1);

[0031] (8) Randomly sample from the experience pool to train the DRL model.

[0032] ​​4. The present invention designs an execution order optimization module based on NSGA-II, which is used to optimize the execution order of the remaining task list of the selected UAV. The NSGA-II algorithm has the following advantages: fast non-dominated sorting based on the dominance relationship, crowding distance comparison operator based on distance, and elitist strategy of merging parent and offspring. The task list of the UAV is encoded as a chromosome, each gene represents a task, and the position of the gene represents the execution order of the task. For the newly added task, it is numbered 0 and inserted as a new gene into different positions in the chromosome to form multiple initial solutions. Therefore, the length of the chromosome is equal to the number of all unfinished tasks of the UAV plus the new task. Since the UAV executes tasks in the order of the task list, the single-point crossover is changed to order crossover, as shown in Figure 3 (a), to increase the population diversity and retain the partial orderliness of excellent chromosomes. At the same time, since the execution order of the UAV tasks has a greater impact on the optimization objective, a combination of reverse mutation and insertion mutation can effectively change the chromosome structure, as shown in Figure 3 (b), to increase the population diversity and is applicable to sorting problems. The fitness function is defined as:

[0033]

[0034] In the formula, η, θ, and μ are weight parameters. Fitness reflects the quality of the individual's solution to the problem, and individuals with high fitness have a greater chance of being selected in subsequent steps.

[0035] Select the optimal solution from the final population, which is the best solution for the UAV task execution order.

[0036] 5. The present invention designs a UAV task execution order optimization process based on NSGA-II, which describes generating several random solutions as the initial population first, evaluating the fitness to judge the quality of the solutions, performing selection, crossover, and mutation operations in sequence, updating the population, and finally obtaining the optimal solution. The flowchart of the NSGA-II algorithm is as shown in Figure 4 .

[0037] (1) Obtain the task list of the UAV selected by the DRL module;

[0038] (2) Encode the task list of the UAV as a chromosome and initialize the population;

[0039] (3) Judge whether it is the first-generation population. If not, go to step (4); if so, go to step (5);

[0040] (4) After non-dominated sorting, obtain the first-generation offspring population through the three operations of selection, crossover, and mutation of the genetic algorithm;

[0041] (5) The current generation of evolution is 2;

[0042] (6) Combine the parent and offspring populations;

[0043] (7) Determine whether a new parent population is generated. If not, proceed to steps (8) and (9). If so, proceed to step (10);

[0044] (8) Perform fast non - dominated sorting and calculate crowding distance;

[0045] (9) Select individuals to form a new parent population according to the non - dominated relationship and the crowding distance of individuals; Give priority to selecting individuals with a lower non - dominated level. If the non - dominated levels are the same, select individuals with a larger crowding distance.

[0046] (10) Generate a new generation of offspring population through selection, crossover, and mutation operations;

[0047] (11) Determine whether the termination condition is reached. If so, continue to execute step (12). If not, increment the generation number by 1 and jump to step (6);

[0048] (12) Output the execution order.

[0049] According to the above description, the present invention designs a dynamic optimization scheduling method for an unmanned aerial vehicle (UAV) swarm facing sudden inspection tasks to achieve fast response and optimization for sudden tasks, considering the overall time and energy consumption. When considering task priorities, the system will rank tasks according to their urgency and importance, which helps to first process tasks that have a greater impact on the operation of the smart grid, improving the overall operation safety and efficiency. The UAV selection decision - making module based on deep reinforcement learning obtains all state information, represents the UAV task - stage state in more detail, and selects UAVs to execute new tasks by this module. The execution - order optimization module based on NSGA - II obtains the task list of the selected UAVs, optimizes the crossover and mutation operations, and optimizes the task execution order according to time and energy consumption and priorities. Finally, it realizes the optimized scheduling of UAV swarm inspection for sudden tasks in a dynamic environment with the overall time and energy consumption and task - priority execution as the optimization objectives, enabling the system to cope with different environmental conditions and task changes, and improving the self - adaptability and robustness of the smart grid. In summary, the dynamic - environment inspection - task scheduling optimization method combining deep reinforcement learning and multi - objective genetic algorithm provides an innovative and reliable solution for actual dynamic inspections in the smart - grid environment. Description of the Drawings

[0050] Figure 1 Dynamic optimization scheduling process for sudden inspection tasks

[0051] Figure 2 DRL model training flowchart

[0052] Figure 3 (a) Crossover operator of genetic algorithm

[0053] Figure 3 (b) Mutation operator of genetic algorithm

[0054] Figure 4 NSGA-II training flowchart Specific implementation manner

[0055] According to the above description, the following is a specific implementation process, but the scope protected by this patent is not limited to this implementation process.

[0056] Multiple drones perform tasks on multiple static targets in the smart grid. Each drone needs to have the ability of autonomous flight, and each task can be executed independently of other tasks. At the initial moment, the ground central station has formulated a set of current optimal task allocation schemes according to the known environmental information, task information and drone information, including the tasks that each drone needs to execute and the execution order, and the drones execute tasks according to this scheme.

[0057] Now assume that during the flight, at a certain moment, suddenly the task status changes, and the drones must be reallocated to execute tasks, that is, they cannot execute tasks according to the initial scheme. At the same time, assume that the resources of the drones are redundant for each task during the initial allocation, which means that even after reallocating new tasks, the drone resources can still meet the requirements of all tasks. Now the drones need to request a re-planning of tasks.

[0058] Assume that there are M drones in the power grid inspection area, and the drone set U = {U 1 , U 2 , …, U m , …, U M}, and the position coordinates of drone U m There are N task points distributed within the service area of the drones, and the task set p = {P 1 , P 2 , …, P n , …, P N}, and the position coordinates of task P n At a certain moment during the inspection process, the task status changes, and an emergency task P new appears, with position coordinates and the required execution time T new . All tasks can be inspected by the drones and only need to be inspected by the drones once. There is no interdependence between tasks, that is, each task can be executed independently of other tasks.

[0059] Step 1: Determination of task priority​​

[0060] Smart grids emphasize real-time monitoring and control of all links in the power system, and voltage fluctuation is an important real-time parameter in grid operation. Voltage fluctuations can accelerate equipment aging, lead to a decline in equipment insulation performance, and even cause equipment failures. Therefore, by determining the inspection priority of smart grid equipment based on the amplitude of voltage fluctuations, potential equipment problems can be discovered and addressed in a timely manner, ensuring the stable operation of the grid. Achieving dynamic assessment and timely response to the status of grid equipment is in line with the development trend of smart grids. The voltage fluctuation VF is expressed as:

[0061]

[0062] where Voltage max is the maximum voltage, Voltage min is the minimum voltage, and Voltage avg is the average voltage.

[0063] Based on the power load fluctuation, the priority of the inspection task is determined. Combining the voltage fluctuation amplitude and voltage flicker, the task priority PI is determined. The inspection tasks are divided into three levels in total. A high priority indicates that the inspection needs to be carried out first to determine the grid stability in this area. When the voltage fluctuation exceeds 10% or the voltage flicker Pst exceeds 1.0, the priority is the highest; when the voltage fluctuation is between 5% and 10% or the voltage flicker Pst is between 0.5 and 1.0, the priority is relatively high; when the voltage fluctuation is below 5% or the voltage flicker Pst is below 0.5, the priority is relatively low.

[0064] Step 2: Select drones to participate in new tasks

[0065] To achieve the priority task allocation of drones under the conditions of minimizing energy consumption and time, DRL is used to train the model.

[0066] The DRL module first obtains the inspection time T new of the sudden new task, the priority Imp of the inspection task, the location of the new task the location of the drone the remaining battery power Q of the drone U , and the status where s U represents the current positions of all drones s T is the current state of the drone. The state of each drone can be divided into three types: idle, executing a task, and having a task to be done, corresponding to s T = 0, 1, 2 respectively. The dimension of this vector is 3M + 5.

[0067] Generate an action a t for s t, perform this action to obtain the reward r of the environmental feedback t , the reward function r t , including the energy consumption reward r E = -E total , the time reward r S = -T total and the task reward E total is the sum of the overall energy consumption of all drones during the inspection process, T total is the total task completion time, defined as the time when the last drone returns to the starting point, T LAT is the waiting time for a new task, that is, the time from the appearance of a new task to the start of a drone to execute this task. The reward function is expressed as:

[0068] r t = αr E + βr s + μr p

[0069] Among them, α, β, and γ are weight parameters. Since it is necessary to execute efficiently and prioritize key tasks with higher power load fluctuations, that is, higher-priority tasks, there are high requirements for time and priority. Therefore, α = 1, β = 1.2, and μ = 1.2 are taken.

[0070] and enter the next state s t+1 , obtaining the quadruple training sample <s t , a t , r t , s t+1 > exists in the experience pool, and then repeat this training process for the DRL module to obtain the final output.

[0071] The output of the module is the action a executed by the DRL t = (a 1 , a 2 , …, a M ), that is, the optimal drone selected for the current new task. The value of the selected drone is 1, and the values of other drones are 0. The dimension of this vector is M, that is, the selected drone.

[0072] The reward function of this module considers the impact of priority on inspection. Using the self-learning mechanism of the DRL algorithm, after training, the model will adjust the parameters in the direction of maximizing the reward function, select the best drone to execute the emergency task, and achieve the expected effect of the present invention.

[0073] Step 3: Optimize the subsequent task list of the drone

[0074] After selecting a drone to perform a new task, obtain the remaining task list of the selected drone, and then the order of the task list of the selected drone needs to be re-optimized to complete the task with less time and energy consumption.

[0075] (1) Encoding and Decoding

[0076] Encoding: Encode the task list of the drone into a chromosome, where each gene represents a task and the value range is from 0 to k (k is the number of unfinished tasks in the initial task list of the selected drone). For the newly added task, number it as 0 and insert it as a new gene into different positions in the chromosome to form multiple initial solutions. Therefore, the length of the chromosome is equal to the number of all unfinished tasks of the drone plus the new task.

[0077] Decoding: Update the original task list according to the chromosome to obtain the complete task execution order.

[0078] (2) Fitness Evaluation

[0079] After obtaining the complete task execution order by decoding the chromosome, calculate its fitness value. In each generation of the algorithm, first, the fitness of each individual in the population needs to be evaluated. The fitness function is used to evaluate the quality of each insertion position. The greater the fitness, the better the solution. Finally, the optimal solution is selected based on the value of the fitness function.

[0080] The fitness function is defined as:

[0081]

[0082] (3) Genetic Operations

[0083] Selection operation: According to the fitness value, use the roulette wheel method to select excellent individuals from the population as parents.

[0084] Crossover operation: Since the drone executes tasks in the order of the task list and the value range is limited, the order crossover is selected to maintain the legality of the chromosome after crossover, so as to increase the population diversity and retain part of the order of the excellent chromosomes.

[0085] Mutation operation: Since the order of task execution by the drone has a great influence on the optimization goal, a combination of reverse mutation and insertion mutation can effectively change the chromosome structure and increase the population diversity, which is suitable for sorting problems.

[0086] After continuous iteration reaches the maximum iteration number of 500, it stops. The optimal solution selected from the final population is the best solution for the UAV mission execution sequence. Thus, it ultimately realizes the optimal scheduling of UAV swarm patrol for emergency tasks in a dynamic environment with the overall time energy consumption and task priority execution as the optimization objectives, enabling the system to cope with different environmental conditions and task changes, and improving the adaptability and robustness of the smart grid. To sum up, the optimization method for patrol task scheduling in a dynamic environment combining deep reinforcement learning and multi-objective genetic algorithm provides an innovative and reliable solution for the actual dynamic patrol in the smart grid environment.

[0087] 1. What are the key points and protected points of the present invention?

[0088] 1. The present invention designs an automated method for dynamic optimization scheduling of UAV swarms for emergency patrol tasks. It describes that after the present invention first determines that a new task appears during the UAV smart grid patrol, it senses relevant states such as the UAV position, remaining battery power, and task requirements, selects a UAV to execute the new task based on deep reinforcement learning, and then re-optimizes and sorts the tasks of the UAVs based on NSGA-II, comprehensively considering the UAV selection problem and the task order problem, and completes the optimal dynamic task reallocation through training iteration to adapt to the actual problems of UAV smart grid patrol in a dynamic environment, with strong robustness.

[0089] 2. The present invention designs a UAV selection decision-making module based on deep reinforcement learning. After a new task appears, this module first obtains information such as the UAV and task positions, remaining battery power, priority, and required time, and classifies the current state of the UAV into three types: idle, executing a task, and having a task to be done, which can better match the UAV and task requirements, form a state vector, output the optimal UAV for the current new task according to the current state, and transmit the result to the task optimization and sorting module. The reward function of this module comprehensively considers the time energy consumption and priority factors. This reward function will guide the model to adjust the output strategy in the direction of maximizing the reward, realize the patrol task with the minimum time energy consumption, and avoid the problem of single optimization objective of traditional algorithms.

[0090] 3. The present invention designs a training process for the DRL selection decision-making module. This process describes the training process of the selection decision-making module. It includes obtaining the current UAV and task information, forming a state vector, and then using DRL to generate corresponding actions. The environment feedbacks rewards according to this action and transfers to the next state, completing a time step, and at the same time generating a training sample and storing it in the experience pool. When the task is completed, the module extracts samples from the experience pool for model training to realize the autonomous learning of the DRL selection decision-making module to improve its performance.

[0091] 4. The present invention designs an execution order optimization module based on NSGA-II. This module obtains the DRL selection decision result, encodes and converts the execution order of the remaining task list of the UAV into a chromosome, and through non-dominated sorting, selection, crossover, and mutation operations on the solutions in the population, improves the crossover and mutation operations, increases the population diversity, and the elite retention strategy ensures that the optimal solution will not be eliminated, and selects the optimal solution, that is, the best solution for the UAV task execution order. At the same time, this execution result is the basis for the training of the DRL model.

[0092] 5. The present invention designs a training process for the NSGA-II order optimization module. This process describes the training process of the order optimization module. It includes obtaining the selected UAV task list, randomly generating an initial population, obtaining the first-generation offspring population through selection, crossover, and mutation after non-dominated sorting, merging the parent population and the offspring population, performing fast non-dominated sorting and crowding degree calculation to form a new parent population, generating a new offspring population, and continuously iterating until the optimal sorting is output to complete the final task reallocation.

Claims

1. A method for dynamic optimization and scheduling of drone swarms for sudden inspection tasks, characterized by: (1) When the mission starts, the drone executes the existing mission assignment and observes the environmental status; (2) Determine whether a new task appears. If yes, execute step (3); if no, jump to step (9); (3) Obtain new task status information and environmental information; (4) Input the information into the DRL UAV selection decision module training; (5) Output the number of the drone that will perform the new mission; (6) obtaining a task list of the selected UAV from the existing task assignment results of the UAV and encoding the task list; (7) Use NSGA-II to optimize the task execution sequence of the selected UAVs; (8) Output the final task scheduling strategy for emergency inspection tasks; (9) Complete all inspection tasks and execute tasks according to the new planning scheme; (10) Mission completed; The drone selection decision module selects a drone to perform a new task based on current environmental information; Assume that there are M drones in the smart grid inspection area, and the drone set U = {U1, U2, …, U m ,…,U M }, UAV U m Location coordinates N mission points are distributed in the service area of ​​the drone, and the mission set P = {P1, P2, ..., P n ,…,P N }, Task P n Location coordinates At a certain moment in the inspection process, the task status changes and an emergency task P appears. new , location coordinates Required execution time T new , the task priority of this task is Imp; it is assumed that there is no interdependence between tasks, that is, each task can be executed independently of other tasks, and each task is completed by one drone; the current state observed by the DRL module s U Contains the current positions of all drones Q U is the remaining battery power of the drone, s T is the current state of the drone. Each drone can be divided into three states: idle, executing a mission, and having a mission to do. The output of the module is the action a executed by DRL. t =(a1,a2,…,a M ), that is, the optimal drone selected for the current new task, the selected drone has a value of 1, and the other drones have a value of 0, and the dimension of this vector is M. In addition, a reward function that comprehensively considers time, energy consumption, and task priority is designed: In the formula, α, β, and γ are weight parameters, E total is the total energy consumption of all drones during the inspection process, T total is the total mission completion time, defined as the time when the last UAV returns to the starting point; T LAT The waiting time for a new task is the time from when a new task appears to when a drone starts to execute the task.

2. The method according to claim 1, characterized in that The specific training of the drone selection decision module is as follows: Describes the DRL model at each time step t, observing the environment information and the real-time status of the drone and the task to generate a state vector s t , generates action a according to its own strategy t , when performing action a t Then get reward r from the environment t , and transfer to the next state s t+1 ; Get the four-tuple training sample t ,a t ,r t ,s t+1 >Exists in the experience pool; after completing the assigned task, randomly extract samples from the experience pool to train the DRL model;​ (1) The DRL model obtains the current new mission and the status information of the UAV, including the location, remaining power, and required time; (2) Form the input vector of the DRL model at time t (3) Generates t Action t =π(s t )=(a1,a2,…,a M ); (4) Execute the action and optimize the order of the drone task list; (5) Get the current reward r t , and enter the next state s t+1 ; (6) Forming training samples t ,a t ,r t ,s t+1 >Save to experience pool;​ (7) Determine whether all tasks are completed. If so, proceed to step (8); otherwise, return to step (1); (8) Randomly sample from the experience pool to train the DRL model.

3. The method according to claim 1, characterized in that The optimization of the task execution sequence of the selected UAV using NSGA-II is as follows: The task list of the drone is encoded as a chromosome, where each gene represents a task and the position of the gene indicates the execution order of the tasks. For the newly added tasks, they are numbered as 0 and inserted into different positions of the chromosome as a new gene to form multiple initial solutions. Therefore, the length of the chromosome is equal to the number of all unfinished tasks of the drone plus the new tasks. Since the drone executes tasks in the order of the task list, the single-point crossover is changed to a sequential crossover. The fitness function is defined as: In the formula, η, θ, and μ are weight parameters; the optimal solution is selected from the final population, which is the optimal solution for the execution order of UAV tasks.

4. The method according to claim 1, characterized in that: The NSGA-II algorithm has the following features: (1) Obtain the mission list of the drone selected by the DRL module; (2) The drone’s task list is encoded as a chromosome and the population is initialized; (3) Determine whether it is the first generation population. If not, proceed to step (4); if yes, proceed to step (5); (4) After non-dominated sorting, the first generation of offspring population is obtained through the three operations of selection, crossover, and mutation of the genetic algorithm; (5) The current evolutionary generation is 2; (6) Merging the parent and offspring populations; (7) Determine whether to generate a new parent population. If not, proceed to steps (8) and (9). If yes, proceed to step (10). (8) Perform fast non-dominated sorting and calculate the congestion degree; (9) Select individuals to form a new parent population based on the non-dominant relationship and the crowding degree of the individuals; (10) Produce a new generation of offspring population through selection, crossover, and mutation operations; (11) Determine whether the termination condition is met. If yes, continue to execute step (12). If not, increase the evolutionary generation by 1 and jump to step (6); (12) Output execution order.