A Multi-UAV Cooperative Power Grid Detection Method and Device
By establishing a multi-UAV collaborative task planning model based on environmental factors and combining the Australian wild dog algorithm for iterative optimization, the problems of low task allocation accuracy and slow convergence speed in multi-UAV collaborative power grid detection in the existing technology are solved, and more efficient task allocation and better grid detection effects are achieved.
Patent Information
- Application Number
- CN202510166874.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-02-14
AI Technical Summary
The prior art has problems in the detection of multiple drone collaborative power grids such as slow convergence speed, insufficient optimization capability, low task allocation accuracy, and failure to fully consider communication quality and detection conditions.
By establishing a multi-UAV collaborative task planning model based on environmental factors, combining the Australian wild dog algorithm to iteratively find the nonlinear multi-objective problem, determine the optimal task allocation for multi-UAV collaborative power grid detection, and use the dynamic adjustment mechanism of the objective function and the grid detection evaluation index to optimize the task allocation plan.
The efficiency of multi-UAV collaborative power grid detection task allocation planning and calculation is improved, the algorithm is avoided from falling into local optimization, and the accuracy and efficiency of task allocation are improved.
Smart Images

Figure CN119624069B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of power grid detection, and particularly relates to a multi-UAV collaborative power grid detection method and device. Background Art
[0002] In power system maintenance, the multi-UAV collaborative task power grid detection planning plays a crucial role. This technology needs to comprehensively consider the terrain where the UAVs are located, the communication signal strength, and environmental factors to formulate the best allocation and task planning scheme for multiple UAVs to perform complex multi-power grid inspection tasks. Especially in complex communication and large-scale power grid detection environments, the multi-UAV collaborative power grid detection planning technology has become a key technology to improve the power grid inspection efficiency and reduce the maintenance cost. Since this problem essentially belongs to a multi-constraint optimization combination problem, in the face of the instability of communication signals and a large number of power grid objects to be detected, how to quickly and accurately perform collaborative task planning has become a major challenge in the current power inspection field.
[0003] Although there are currently a small number of task allocation schemes for multiple UAVs, they still have the following defects: 1. Using conventional optimization algorithms, their convergence speed and optimization ability cannot be satisfied; 2. Directly applying conventional optimization algorithms without adjusting the algorithm parameters during the optimization process or with insufficient adjustment accuracy of the algorithm parameters, resulting in relatively low task allocation accuracy in the corresponding scenarios; 3. Only considering the UAV flight speed and task type to determine the model constraints, while not considering communication quality, detection conditions, etc., and being unable to be well applied to the power grid detection field. For example, the invention patent application CN117170405A discloses a UAV task allocation method based on multi-objective particle swarm. This method first constructs a heterogeneous UAV collaborative multi-task allocation model, and then adopts a constraint-based initialization strategy to ensure that the particles generated by the initial particle swarm are all feasible solutions that meet the constraint conditions; afterwards, a two-stage optimal particle selection strategy is adopted to retain the individual optimal particles and population optimal particles found; next, a task-based small-module particle update strategy is adopted, which can make the particles quickly learn from the optimal particles, and at the same time perform a task-based mutation operation with a certain probability to improve the comprehensive performance of the algorithm; finally, the updated particles are corrected to ensure that they are feasible solutions that meet the constraints. Summary of the Invention
[0004] Aiming at the defects existing in the above-mentioned prior art, the present invention provides a multi-UAV collaborative power grid detection method and device, which can endow the UAVs with more intelligence and improve the solution efficiency of the task allocation planning for multi-UAV collaborative power grid detection.
[0005] In a first aspect, the present invention provides a multi-UAV collaborative power grid detection method, including:
[0006] Based on the dynamic cost and dynamic benefit required for UAV detection considering environmental factors, a multi-UAV collaborative task planning model is established;
[0007] Based on the multi-UAV collaborative task planning model and the multi-UAV collaborative decision-making mechanism, iterative optimization of the non-linear multi-objective problem is carried out to determine the optimal task allocation for multi-UAV collaborative power grid detection;
[0008] Based on the optimal task allocation, coordinate multiple UAVs to conduct power grid detection.
[0009] Furthermore, based on the dynamic cost and dynamic benefit required for UAV detection considering environmental factors, a multi-UAV collaborative task planning model is established, including:
[0010] Based on the detection cost and flight range cost generated by the UAV executing the detection task, determine the dynamic cost corresponding to UAV detection;
[0011] According to the detection accuracy of the UAV and the detection value of the power grid target, determine the dynamic benefit of the UAV;
[0012] According to the dynamic cost and dynamic benefit of the UAV, give the objective function of the multi-UAV collaborative task planning;
[0013] Determine the constraint conditions and the dynamic adjustment mechanism of the objective function of the UAV considering environmental factors, and combine with the objective function of the multi-UAV collaborative task planning to establish a multi-UAV collaborative task planning model.
[0014] Furthermore, determine the dynamic cost corresponding to UAV detection, including:
[0015] Obtain the distance between the UAV and the power grid target and the minimum distance to meet the detection quality, and give the interference distance;
[0016] Based on the detection interference of the environmental factors on the UAV and the detection power of the UAV, and combined with the interference distance, determine the detection cost of the UAV;
[0017] Obtain the real-time wind speed and terrain undulation parameters at the power grid target and perform normalization processing to give the terrain complexity coefficient at the power grid target;
[0018] Obtain the maximum Euclidean distance of all UAVs relative to the power grid target, and combine with the terrain complexity coefficient at the power grid target and the distance between the UAV and the power grid target to give the flight range cost of the UAV;
[0019] Superimpose the detection cost and flight range cost of the UAV to give the dynamic cost corresponding to UAV detection.
[0020] Furthermore, the dynamic cost corresponding to UAV detection satisfies the following relationship:
[0021]
[0022] In the formula, represents the dynamic cost of UAV i detecting power grid target j at time t;
[0023] The detection cost generated by the drone performing the detection task satisfies the following relationship:
[0024]
[0025] In the formula, represents the detection cost of UAV i detecting power grid target j at time t, α is the detection weight factor, is the detection power of drone i, d ij ( t ) represents the distance from the detection grid target j at time t, d min To meet the minimum distance for detection quality, is the interference coefficient of environmental factors on UAV i’s detection of power grid target j at time t. Environmental factors include weather attenuation, terrain shielding, and obstacle density. is the preset interference threshold, x ij Indicates whether the power grid target j is detected by UAV i. When the power grid target j is detected by UAV i, x ij =1, otherwise x ij =0;
[0026] The range cost generated by the drone performing the inspection task satisfies the following relationship:
[0027] ;
[0028] In the formula, represents the range cost of UAV i detecting power grid target j, β is the cost weight factor, τ j ( t ) represents the terrain complexity coefficient of the location of the power grid target j, γ 1, γ 2 is the normalization coefficient, W j ( t ) is the real-time wind speed at the location of the power grid target j, T j is the terrain undulation parameter where the power grid target j is located, d max represents the maximum Euclidean distance of all UAVs relative to the grid target j.
[0029] Further, the dynamic benefit of the UAV includes:
[0030] Obtain the current heading angle of the UAV and the optimal approach angle considering environmental factors, and give the deviation coefficient of the UAV.
[0031] Obtain the detection accuracy of the UAV and the inspection value of the power grid target, and determine the dynamic benefit of the UAV detecting the power grid target in combination with the deviation coefficient of the UAV.
[0032] Further, the dynamic benefit of the UAV satisfies the following relationship:
[0033]
[0034] In the formula, R ij ( t ) represents the dynamic benefit of UAV i detecting power grid target j at time t, η 0 is the equity weight factor, λ i represents the detection accuracy of the sensor carried by UAV i, F j is the inspection value of power grid target j, θ i ( t ) is the current heading angle of UAV i, is the optimal approach angle under environmental constraints.
[0035] Further, the objective function of multi-UAV cooperative mission planning satisfies the following relationship:
[0036]
[0037] In the formula, F ( t ) is the objective function of multi-UAV cooperative mission planning, Ω( t ) represents the dynamic weight vector, which is updated in real time through the Kalman filter, , ω1, ω2, ω3 are all dynamic weights, ω1 + ω2 + ω3 = 1, T is the transpose operation, Δ E ( t ) is the change amount of environmental parameters, K, Q is the adjustment matrix in the Kalman algorithm.
[0038] Further, the constraint conditions of the UAV considering environmental factors satisfy the following relationship:
[0039]
[0040] In the formula, represents the current maximum flight speed of UAV i, is the default maximum flight speed of the drone, κ is the wind speed influence factor, W max is the maximum ambient wind speed, d safe ( t ) is the maximum safe flight distance, μ is the safety margin adjustment factor, ρ(t) is the real-time obstacle density, and T is the power grid target set of the detection task;
[0041] Furthermore, the objective function dynamic adjustment mechanism satisfies the following relationship:
[0042]
[0043] In the formula, δ(t) is the environmental parameter change rate, δ threshold is the environmental parameter change rate threshold, E(t) is the environmental parameter at time t, Δt is the time change amount, ω'1 and ω'2 are the dynamic weights for updating and adjusting ω1 and ω2 respectively, ΔΦ is the communication interference change amount, ΔV is the power grid value fluctuation range, V max is the maximum power grid value, sigmoid() is the activation function, and sigmoid(ΔΦ)=1 / (1+e^-ΔΦ).
[0044] Furthermore, based on the multi-UAV cooperative task planning model and the multi-UAV cooperative decision-making mechanism, perform iterative optimization of the non-linear multi-objective problem to determine the optimal task allocation for multi-UAV cooperative power grid detection, including:
[0045] S21. Establish the mapping between the positions of the dingo individuals in the dingo algorithm and the corresponding power grid target allocation schemes of the multi-UAVs, determine the evolutionary strategies and coefficient decision factors of different hunting methods in the dingo algorithm, and initialize the positions of the dingoes in the dingo algorithm;
[0046] S22. Based on the multi-UAV cooperative decision-making mechanism at the current dingo position, select the corresponding coefficient decision factor to update the evolutionary strategy, generate a new dingo position according to the updated evolutionary strategy, and calculate the corresponding power grid detection evaluation index;
[0047] S23. Repeat the iterative step S22 until the preset conditions are met, output the dingo position with the optimal power grid detection evaluation index, and obtain the optimal task allocation for multi-UAV cooperative power grid detection.
[0048] Furthermore, based on the multi-UAV cooperative decision-making mechanism at the current dingo position, select the corresponding coefficient decision factor to update the evolutionary strategy, including:
[0049] Based on the multi-UAV collaborative mission planning model and environmental factors, determine the state space of the reinforcement learning algorithm, and determine the action space of the reinforcement learning algorithm based on the coefficient decision factor;
[0050] Determine the reward function of the reinforcement learning based on the multi-UAV collaborative mission planning model;
[0051] Based on the current dingo position, combined with the state space, action space and reward function, determine the probability of selecting an action in the action space under the current state space, and based on the maximum probability, determine the coefficient decision factor corresponding to the selected action;
[0052] Update the evolutionary strategy according to the coefficient decision factor corresponding to the selected action.
[0053] Furthermore, the hunting methods of the dingo algorithm include collective hunting, solitary predation, and scavenging;
[0054] The evolutionary strategy of collective hunting satisfies the following relationship:
[0055]
[0056] The evolutionary strategy of solitary predation satisfies the following relationship:
[0057]
[0058] The evolutionary strategy of scavenging satisfies the following relationship:
[0059]
[0060] In the formula, are the new positions of the dingo individuals when participating in collective hunting, solitary predation, and scavenging of dingoes respectively. α(t), β(t), and γ(t) are the hunting intensity coefficient of collective hunting, the exploration rate parameter of solitary predation, and the scavenging trigger threshold of scavenging respectively. k ∈ {1, 2,..., P}, where P is the number of the dingo population, X k represents the initial position of the selected dingo, is the time-varying coupling coefficient, , a3 is the time-varying coupling weight, T max is the maximum duration of power grid detection, X best is the optimal individual among all dingoes, that is, having the current optimal allocation scheme. F1 represents the adaptive hunting scaling factor, X m , X n are two random dingo individuals participating in collective hunting respectively. λ(t) is the dynamic learning rate, , are two random individuals among all dingo individuals respectively, is a random individual among all dingo individuals, a1 and a2 are the weights of individual predation and the scavenging position update threshold respectively, and rand is a random number between (0, 1).
[0061] Furthermore, the coefficient decision factors include the hunting intensity coefficient of collective hunting, the exploration rate parameter of individual predation, and the scavenging trigger threshold of scavenging, satisfying the following relationship:
[0062]
[0063] In the formula, Δ f is the fitness variance of the dingo population, σ is the preset variance, β min and β max are the minimum exploration degree and the maximum exploration degree respectively, γ base is the basic scavenging trigger value, and ε is the greedy coefficient.
[0064] Furthermore, based on the current dingo position, combined with the state space, action space, and reward function, determine the probability of selecting an action in the action space under the current state space, and based on the maximum probability, determine the coefficient decision factor corresponding to the selected action, including;
[0065] Based on the current dingo position, evaluate the expected long-term rewards of selecting each action in the action space under all states in the current state space before update through the reinforcement learning algorithm, and determine the corresponding expected long-term rewards at the next moment;
[0066] Select the maximum value from the expected long-term rewards corresponding to the next moment, and update the expected long-term rewards of selecting each action in the action space under all states in the current state space in combination with the reward function;
[0067] Based on the updated expected long-term rewards of selecting each action in the action space under all states in the current state space, determine the probability of selecting each action in the action space under all states in the current state space;
[0068] Screen and determine the action corresponding to the maximum probability, and based on the relationship between the coefficient decision factor and the action space, determine the coefficient decision factor of the hunting method corresponding to the action selected when the probability is the maximum.
[0069] Furthermore, the probability of selecting each action in the action space under all states in the current state space satisfies the following relationship:
[0070]
[0071] In the formula, π(a|s) is the probability of selecting action space A t under state s in the current state space S tThe probability of the middle action a, and Q'(s,a) is the updated current state space S t Select the action space A in the state s t The expected long-term reward of the action a, and Q(s,a) is the current state space S before update t Select the action space A in the state s t The expected long-term reward of the action a, T is the total number of grid targets, l r Is the learning rate, r(t) is the reward function, η is the discount factor, s', a' are the state at the next moment and the action selected for the state at the next moment reached by performing the action a in the current state s respectively, Δ F ( t ) is the difference between the multi-UAV cooperative mission planning objective function at time t and time t-1, Ω opt Is the small perturbation value of the dynamic weight vector, a small value close to 0, collision_rate is the collision rate, a4 and a5 are the revenue weights for minimizing the objective weight and minimizing the dynamic weight vector respectively, F ( t ) is the multi-UAV cooperative mission planning objective function, W j ( t ) is the real-time wind speed at the location of grid target j, Is the interference coefficient of environmental factors on UAV i detecting grid target j at time t, ρ(t) is the real-time obstacle density, α(t), β(t), γ(t) are the hunting intensity coefficient of collective hunting, the exploration rate parameter of individual predation, and the scavenging trigger threshold of scavenging respectively, Δα, Δβ, Δγ are the discrete adjustment amounts of the hunting intensity coefficient, exploration rate parameter, and scavenging trigger threshold respectively.
[0072] In a second aspect, the present invention also provides a multi-UAV cooperative grid detection device, which adopts the above multi-UAV cooperative grid detection method. The device includes:
[0073] A model construction module, which is used to establish a multi-UAV cooperative mission planning model based on the dynamic cost and dynamic revenue required for UAV detection considering environmental factors;
[0074] A task allocation module, which is used to perform iterative optimization of non-linear multi-objective problems based on the multi-UAV cooperative mission planning model and the multi-UAV cooperative decision-making mechanism to determine the optimal task allocation for multi-UAV cooperative grid detection;
[0075] A grid detection module, which is used to coordinate multi-UAVs for grid detection based on the optimal task allocation.
[0076] Further, the model construction module includes:
[0077] Determine the dynamic cost corresponding to the UAV detection based on the detection cost and the range cost generated by the UAV performing the detection task;
[0078] Determine the dynamic benefit of the UAV according to the detection accuracy of the UAV and the detection value of the power grid target;
[0079] Give the objective function of the multi-UAV collaborative task planning according to the dynamic cost and the dynamic benefit of the UAV;
[0080] Determine the constraint conditions and the dynamic adjustment mechanism of the objective function of the UAV considering environmental factors, and establish a multi-UAV collaborative task planning model in combination with the objective function of the multi-UAV collaborative task planning.
[0081] Further, the task allocation module includes:
[0082] S21. Establish the mapping between the position of the dingo individual in the dingo algorithm and the allocation scheme of the multi-UAV corresponding power grid target, determine the evolution strategy and the coefficient decision factor of different hunting methods of the dingo algorithm, and initialize the position of the dingo in the dingo algorithm;
[0083] S22. Select the corresponding coefficient decision factor to update the evolution strategy based on the multi-UAV collaborative decision-making mechanism at the current dingo position, generate a new dingo position according to the updated evolution strategy, and calculate the corresponding power grid detection evaluation index;
[0084] S23. Repeat the iterative step S22 until the preset condition is satisfied, output the dingo position with the optimal power grid detection evaluation index, and obtain the optimal task allocation for the multi-UAV collaborative power grid detection.
[0085] Further, the task allocation module includes:
[0086] Determine the state space of the reinforcement learning algorithm based on the multi-UAV collaborative task planning model and environmental factors, and determine the action space of the reinforcement learning algorithm based on the coefficient decision factor;
[0087] Determine the reward function of the reinforcement learning based on the multi-UAV collaborative task planning model;
[0088] Based on the current dingo position, combine the state space, the action space and the reward function to determine the probability of selecting an action in the action space in the current state space, and determine the coefficient decision factor corresponding to the selected action based on the maximum value of the probability;
[0089] Update the evolution strategy according to the coefficient decision factor corresponding to the selected action.
[0090] A multi-UAV collaborative power grid detection method and device provided by the present invention have at least the following beneficial effects:
[0091] (1) The distribution scheme evolution strategy established by the present invention in combination with the diverse predation mechanisms of the dingo algorithm includes various evolution strategies such as collective hunting, individual predation, and scavenging, replacing the single evolution method of traditional swarm intelligence algorithms, endowing the individuals in the population with more autonomy, and avoiding the algorithm from falling into local optima.
[0092] (2) By modeling the characteristics of power grid detection, the present invention designs a dynamic adjustment mechanism for the objective function and evaluation indexes for power grid detection, ensuring that the individuals in the population optimize towards the correct distribution scheme and improving the optimization efficiency of the algorithm.
[0093] (3) The present invention can quickly and effectively solve the problem of multi-UAV collaborative power grid detection task allocation based on the dingo algorithm, providing a new solution idea for solving the problem of multi-UAV collaborative power grid detection task allocation in complex environments. Brief Description of the Drawings
[0094] Figure 1 is a flowchart of a multi-UAV collaborative power grid detection method provided by the present invention;
[0095] Figure 2 is a schematic diagram of the principle of the multi-UAV collaborative power grid detection method provided by a certain embodiment of the present invention;
[0096] Figure 3 is a flowchart of obtaining the optimal task allocation provided by a certain embodiment of the present invention;
[0097] Figure 4 is a flowchart of updating the evolution strategy provided by a certain embodiment of the present invention;
[0098] Figure 5 is a schematic diagram of a multi-UAV collaborative power grid detection device provided by the present invention. Detailed Embodiments
[0099] In order to better understand the above technical solutions, the following will describe the above technical solutions in detail in combination with the accompanying drawings of the specification and specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0100] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms of "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. "Plural" generally includes at least two.
[0101] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a commodity or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the commodity or device comprising said element.
[0102] Based on the construction of a multi-UAV collaborative mission planning model, this invention combines the dingo algorithm for optimization and solution to give the optimal mission allocation plan for multi-UAVs, and finally realizes the collaborative power grid detection planning for multi-UAVs considering environmental factors. Although there is an existing application of the dingo algorithm in the path planning of power inspection UAVs, it only belongs to a simple application, using the dingo algorithm for single-objective path optimization, without the detailed cost and dynamic adjustment mechanism of the objective function and the power grid detection evaluation index for the application of the dingo optimization algorithm in the power grid detection task allocation in this invention. Through the objective function considering environmental factors, as well as the corresponding evolutionary strategy and coefficient decision factor of the dingo algorithm, the collaborative power grid detection planning of multiple objectives is finally realized, which is better applicable to the application scenario of multi-UAV power grid detection.
[0103] As Figure 1 shown, this invention provides a multi-UAV collaborative power grid detection method, including:
[0104] Based on the dynamic cost and dynamic benefit required for UAV detection considering environmental factors, establish a multi-UAV collaborative mission planning model;
[0105] Based on the multi-UAV collaborative mission planning model and the multi-UAV collaborative decision-making mechanism, conduct iterative optimization of the non-linear multi-objective problem to determine the optimal mission allocation for multi-UAV collaborative power grid detection;
[0106] Based on the optimal mission allocation, coordinate multi-UAVs to conduct power grid detection.
[0107] In the actual application scenario, as Figure 2 shown, it is the principle structure block diagram of this invention, and the detailed steps of this invention are as follows:
[0108] Step 1: Establish a multi-UAV collaborative mission planning model;
[0109] Step 1.2: Determine the dynamic cost:
[0110] Obtain the distance between the UAV and the power grid target and the minimum distance to meet the detection quality, and give the interference distance;
[0111] Based on the detection interference of the UAV by environmental factors and the detection power of the UAV, and combined with the interference distance, determine the detection cost of the UAV;
[0112] Obtain the real-time wind speed and terrain undulation parameters at the power grid target and perform normalization processing to give the terrain complexity coefficient at the power grid target;
[0113] Obtain the maximum Euclidean distance of all drones relative to the power grid target, and combine the terrain complexity coefficient at the power grid target and the distance between the drone and the power grid target to give the flight range cost of the drone;
[0114] Superimpose the detection cost and flight range cost of the drone to give the dynamic cost corresponding to the drone detection.
[0115] That is, the cost required for the drone detection of the present invention mainly includes the detection cost and flight range cost generated by the drone performing the detection task:
[0116] 1) Detection cost generated by the drone performing the detection task: It is determined according to the distance between the drone detection target and environmental factors, specifically:
[0117]
[0118] In the formula, represents the detection cost of drone i detecting power grid target j at time t, α is the detection weight factor, is the detection power of drone i, d ij ( t ) represents the distance from the detection of power grid target j at time t, d min is the minimum distance to meet the detection quality, that is, the minimum distance required to ensure that the detection quality meets the predetermined requirements, is the interference coefficient of environmental factors on drone i detecting power grid target j at time t, and environmental factors include weather attenuation, terrain occlusion, and obstacle density, is the preset interference threshold, x ij represents whether power grid target j is detected by drone i. When power grid target j is detected by drone i, x ij = 1, otherwise x ij = 0;
[0119] 2) Flight range cost generated by the drone performing the detection task: The smaller the flight range cost of the drone, the greater the probability that the power grid target is assigned to the drone. Then the flight range cost is:
[0120]
[0121] In the formula, represents the flight range cost of drone i detecting power grid target j, βis the cost weight factor, τ j ( t ) represents the terrain complexity coefficient at the location of grid target j, γ 1, γ 2 is the normalization coefficient, W j ( t ) is the real-time wind speed at the location of grid target j, T j is the terrain undulation parameter at the location of grid target j, d max represents the maximum Euclidean distance of all UAVs relative to grid target j;
[0122] Step 1.3. Determine the dynamic benefit: The dynamic benefit of the UAV refers to the detection benefit of the UAV considering environmental factors when performing detection tasks through sensors, including:
[0123] Obtain the current heading angle of the UAV and the optimal approach angle considering environmental factors, and give the deviation coefficient of the UAV;
[0124] Obtain the detection accuracy of the UAV and the inspection value of the grid target, and combine the deviation coefficient of the UAV to determine the dynamic benefit of the UAV detecting the grid target, satisfying the following relationship:
[0125]
[0126] In the formula, R ij ( t ) represents the dynamic benefit of UAV i detecting grid target j at time t, η 0 is the benefit weight factor, λ i represents the detection accuracy of the sensor carried by UAV i, F j is the inspection value of grid target j, θ i ( t ) is the current heading angle of UAV i, is the optimal approach angle under environmental constraints;
[0127] Step 1.4. Determine the objective function of multi-UAV collaborative task planning: Use the minimization of the cost paid for detection considering environmental factors to determine the final allocation plan. The smaller the objective function, the better the allocation plan. Then the objective function is:
[0128]
[0129] In the formula, F ( t) is the objective function for multi-UAV cooperative mission planning, and Ω( t ) represents the dynamic weight vector, which is updated in real time through the Kalman filter. , where ω1, ω2, and ω3 are all dynamic weights, corresponding respectively to the dynamic weight of the dynamic cost, the dynamic weight of the dynamic benefit, and the dynamic weight of the dynamic weight. ω1 + ω2 + ω3 = 1. T is the transpose operation, and Δ E ( t ) is the change amount of the environmental parameter. K, Q is the adjustment matrix in the Kalman algorithm.
[0130] Step 1.5: Determine the constraints considering environmental factors. The maximum flight speed and the maximum safe flight distance of all UAVs are the same, and each power grid target must be detected. Then the constraints are:
[0131]
[0132] In the formula, represents the current maximum flight speed of UAV i. is the default maximum flight speed of the UAV. κ is the wind speed influence factor. W max is the maximum environmental wind speed. d safe ( t ) is the maximum safe flight distance, μ is the safety margin adjustment factor, ρ(t) is the real-time obstacle density, and T is the set of power grid targets for the detection task.
[0133] Then the multi-UAV cooperative power grid detection task planning model of the present invention is:
[0134]
[0135] Step 1.6: Determine the dynamic adjustment mechanism of the objective function considering environmental factors. The weights of the dynamic cost and the dynamic benefit are adjusted respectively through the communication situation during UAV detection and the environmental parameters. Specifically:
[0136] By comparing the environmental parameter change rate with the corresponding environmental parameter change rate threshold, when the environmental parameter change rate is large, that is, the environmental parameter change rate is greater than the environmental parameter change rate threshold, the dynamic weights of the dynamic cost and the dynamic benefit in the objective function are adjusted to satisfy the following relationship:
[0137]
[0138] In the formula, δ(t) is the environmental parameter change rate, and δ thresholdis the threshold of the environmental parameter change rate, E(t) is the environmental parameter at time t, Δt is the time change amount, ω'1 and ω'2 are the dynamic weights for updating and adjusting ω1 and ω2 respectively, ΔΦ is the communication interference change amount, ΔV is the grid value fluctuation amplitude, and V max is the maximum grid value.
[0139] Step 2: Establish an Australian wild dog diverse evolutionary strategy:
[0140] Different wild dog individuals select different hunting methods based on probability, and solve the non-linear multi-constraint multi-optimization objectives based on different hunting methods to determine the optimal task allocation for multi-UAV collaborative power grid detection;
[0141] According to different hunting strategies, corresponding evolutionary strategies are designed. The hunting methods of the Australian wild dog algorithm include collective hunting, solitary hunting, and scavenging;
[0142] The evolutionary strategy of collective hunting satisfies the following relationship:
[0143]
[0144] The evolutionary strategy of solitary hunting satisfies the following relationship:
[0145]
[0146] The evolutionary strategy of scavenging satisfies the following relationship:
[0147]
[0148] In the formula, are the new positions of the wild dog individuals when participating in collective hunting, solitary hunting, and scavenging of wild dogs respectively. α(t), β(t), and γ(t) are the hunting intensity coefficient of collective hunting, the exploration rate parameter of solitary hunting, and the scavenging trigger threshold of scavenging respectively. k ∈ {1, 2,..., P}, where P is the number of the Australian wild dog population, and X k represents the initial position of the selected wild dog, is the time-varying coupling coefficient, , a3 is the time-varying coupling weight, and T max is the maximum duration of power grid detection, and X best is the optimal individual among all wild dogs, that is, it has the current optimal allocation scheme. F1 represents the adaptive hunting scaling factor, and X m , X n are two random wild dog individuals participating in collective hunting respectively, λ(t) is the dynamic learning rate, , are two random individuals among all Australian wild dog individuals respectively, is a random individual among all dingo individuals, a1 and a2 are the weights of individual predation and the threshold for updating the scavenging position respectively, and rand is a random number between (0, 1).
[0149] Among them means that when rand is greater than a2, calculate X best -X k , otherwise calculate .
[0150] Step 3: Nonlinear multi-objective optimization solution based on the dingo algorithm, as Figure 3 shown, including:[[]]
[0151] Establish the mapping between the positions of dingo individuals in the dingo algorithm and the corresponding grid target allocation scheme of multiple UAVs, determine the evolutionary strategies and coefficient decision factors of different hunting methods in the dingo algorithm, and initialize the positions of dingoes in the dingo algorithm;
[0152] Based on the multi-UAV collaborative decision-making mechanism, select the corresponding coefficient decision factor to update the evolutionary strategy at the current dingo position, generate a new dingo position according to the updated evolutionary strategy, and calculate the corresponding grid detection evaluation index;
[0153] Repeat the iterative update and calculation process until the preset conditions are met, output the dingo position with the optimal grid detection evaluation index, and obtain the optimal task allocation for multi-UAV collaborative grid detection.
[0154] Specifically:
[0155] Step 3.1: Encoding of dingo individuals
[0156] The present invention adopts the mapping real number vector encoding method to encode the individuals in the dingo algorithm, establish the mapping between the dingo individuals and the multi-UAV collaborative task planning model, the number of UAVs is M, and the number of tasks executed, that is, the number of grid target detections, is T;
[0157] Step 3.2: Population initialization
[0158] The initial position of the dingo population is initialized as follows:
[0159]
[0160] Among them, represents the initial position of the dingo, and represent the upper and lower boundary constraints of, represents a random number between 0 and 1;
[0161] Step 3.3: Hunting stage
[0162] During the hunting stage, the wild dog population conducts a neighborhood search (pack hunting and solitary hunting) or a global random search (scavenging) at the current wild dog position through the predation evolution strategy established in step 2 to develop new predation positions. Among them, in order to achieve precise predation position update under dynamic changes in environmental factors, the evolution strategy also needs to be updated.
[0163] Specifically, as Figure 4 shown, updating the evolution strategy by selecting the corresponding coefficient decision factor based on the multi-UAV collaborative decision-making mechanism at the current wild dog position may include:
[0164] Determine the state space of the reinforcement learning algorithm based on the multi-UAV collaborative mission planning model and environmental factors, and determine the action space of the reinforcement learning algorithm based on the coefficient decision factor;
[0165] Determine the reward function of the reinforcement learning based on the multi-UAV collaborative mission planning model;
[0166] Based on the current wild dog position, combine the state space, action space, and reward function to determine the probability of selecting an action in the action space under the current state space, and based on the maximum probability, determine the coefficient decision factor corresponding to the selected action;
[0167] Update the evolution strategy according to the coefficient decision factor corresponding to the selected action.
[0168] Among them, based on the current wild dog position, combine the state space, action space, and reward function to determine the probability of selecting an action in the action space under the current state space, and based on the maximum probability, determine the coefficient decision factor corresponding to the selected action, including;
[0169] Based on the current wild dog position, evaluate the expected long-term rewards of selecting each action in the action space under all states in the current state space before update through the reinforcement learning algorithm, and determine the corresponding expected long-term rewards at the next moment;
[0170] Select the maximum value from the corresponding expected long-term rewards at the next moment, and update the expected long-term rewards of selecting each action in the action space under all states in the current state space in combination with the reward function;
[0171] Based on the updated expected long-term rewards of selecting each action in the action space under all states in the current state space, determine the probability of selecting each action in the action space under all states in the current state space;
[0172] Screen and determine the action selected when the probability is the maximum, and based on the relationship between the coefficient decision factor and the action space, determine the coefficient decision factor of the hunting method corresponding to the action selected when the probability is the maximum.
[0173] The probabilities of selecting each action in the action space for all states in the current state space satisfy the following relationship:
[0174]
[0175] In the formula, π(a|s) is the probability of selecting action a in action space A in state s in the current state space S t ; Q'(s,a) is the expected long-term reward of selecting action a in action space A in state s in the updated current state space S t ; the corresponding max indicates that the selected action is the estimated best action, Q(s,a) is the expected long-term reward of selecting action a in action space A in state s in the current state space S before update t ; T is the total number of grid targets, l t is the learning rate, r(t) is the reward function, η is the discount factor, s', a' are respectively the state at the next moment and the action selected at the next moment state reached by executing action a in the current state s, Δ t ( t ) is the difference between the multi-UAV cooperative mission planning objective function at time t and time t-1, Ω r is the small perturbation value of the dynamic weight vector, a small value close to 0. For example, [0.01, 0.01, 0.01] can avoid problems such as gradient disappearance or gradient explosion caused by Ω(t) shrinking to 0. The collision_rate is the collision rate, a4 and a5 are respectively the benefit weights for minimizing the objective weight and minimizing the dynamic weight vector, a4 = 0.6, a5 = 0.3 F ( t ) is the multi-UAV cooperative mission planning objective function opt ( F ) is the real-time wind speed at the location of grid target j t ; W j ( t ) is the interference coefficient of environmental factors on UAV i detecting grid target j at time t, ρ(t) is the real-time obstacle density, α(t), β(t), γ(t) are respectively the hunting intensity coefficient of collective hunting, the exploration rate parameter of individual predation, and the scavenging trigger threshold of scavenging, and Δα, Δβ, Δγ are respectively the discrete adjustment amounts of the hunting intensity coefficient, exploration rate parameter, and scavenging trigger threshold ;
[0176] Among them, for the current state space S before update t in state s, selecting action a in action space A tThe expected long-term reward Q(s, a) for action a can be determined using the DQN network algorithm, specifically including: 1. Initializing the DQN network: Initialize the DQN network, with its input being the state s and action a, and the output being Q(s, a); the parameters of the DQN network are randomly initialized; 2. Observing the current state and action: At each time step t, observe the current state s and select an action a; 3. Forward propagation: Use the current state s and action a as inputs, and perform forward propagation through the DQN network to obtain the output Q(s, a) value, which is the expected long-term reward before update.
[0177] The present invention also includes establishing a dual-channel coding mechanism:
[0178] Main coding channel: Represents the task assignment relationship between the UAV and the target;
[0179] Auxiliary coding channel: Represents the set of coefficient decision factors, Φ α 、Φ β 、Φ γ are the networks for the hunting intensity coefficient of collective hunting, the exploration rate parameter of individual predation, and the scavenging trigger threshold of scavenging. That is, the policy u actually controls the hunting intensity coefficient of collective hunting, the exploration rate parameter of individual predation, and the scavenging trigger threshold of scavenging through these three networks.
[0180] By establishing a dual-channel coding mechanism, a mapping from S d to A d can be achieved through X t and Y t , that is, At = f(X d , Y d , S t ), and this mapping function is an MLP neural network.
[0181] After the present invention updates the evolutionary strategy and generates a new wild dog position, the calculated power grid detection evaluation index is specifically as follows:
[0182]
[0183] In the formula, P u (t) is the power grid detection evaluation index, Q u (t) is the long-term cumulative reward value of the policy u, R u (t) is the current immediate environment return. The policy u selects a certain action in the action space at a certain state in the state space. η = 0.8 is the discount factor. Q v (t), R v (t) represent the cumulative reward value and the immediate environment return under the recorded historical policy v. The historical policy v is all the policies u adopted in the historical position update;
[0184] In addition, the power grid detection evaluation index can be established through the short-term target value and long-term target value of the predation quality, and the specific steps are as follows:
[0185] Determine the fitness value calculated from the allocation scheme corresponding to the individual in the previous iteration and the fitness value calculated from the allocation scheme corresponding to the individual in the current iteration; select the minimum fitness value from the fitness value calculated from the allocation scheme corresponding to the individual in the previous iteration and the fitness value calculated from the allocation scheme corresponding to the individual in the current iteration, and calculate the ratio of the selected minimum fitness value to the fitness value calculated from the allocation scheme corresponding to the individual in the previous iteration; based on the difference between the benchmark threshold of the allocation scheme and the result of the ratio calculation, determine the short-term target value under the allocation scheme corresponding to the individual adopted in the current iteration; wherein, the benchmark threshold of the allocation scheme is 1;
[0186] Determine the adjustment value based on the number of iterations, and combine the number of times that the fitness value obtained by the individual adopting the corresponding allocation scheme in the current iteration and previous iterations is less than the benchmark threshold to determine the balance term value; calculate the ratio of the number of times that the fitness value obtained by the individual adopting the corresponding allocation scheme in the current iteration and previous iterations is less than the benchmark threshold to the total number of times that the individual adopts the corresponding allocation scheme in the current iteration and previous iterations, and superimpose the balance term value to give the long-term target value under the allocation scheme corresponding to the individual adopted in the current iteration;
[0187] Determine the weight parameters of the short-term target value and the long-term target value, and combine the short-term target value and the long-term target value to give the power grid detection evaluation index of the current iteration;
[0188] Compare the power grid detection evaluation index of the previous iteration with the power grid detection evaluation index of the current iteration, and update the parameters of the evolutionary strategy in the current iteration when the power grid detection evaluation index of the previous iteration is better than that of the current iteration.
[0189] The power grid detection evaluation index in this way is specifically:
[0190]
[0191] In Equation (11), represents the comprehensive evaluation value under the allocation scheme corresponding to the th iteration using the u th strategy, , are both strategy selection probability weight parameters, , respectively represent the short-term target value and the long-term target value under the allocation scheme corresponding to the th iteration using the u th strategy, represents the fitness value calculated from the allocation scheme corresponding to the individual in the previous iteration, represents the fitness value calculated from the allocation scheme corresponding to the individual in the current iteration, represents the individual at the t th iteration and before, using the u th strategy, the number of times the obtained fitness value is less than 1 for the corresponding allocation scheme, represents at the t th iteration and before, using the u th strategy, the total number of times for the corresponding allocation scheme, is the balance coefficient, and log(t) represents the natural logarithm with t as the independent variable.
[0192] In the present invention, the preset condition for stopping the update iteration is the maximum number of iterations, or the change in the power grid detection evaluation index after several consecutive iteration updates is small or gets worse. After completing the update iteration, by comparing all the power grid detection evaluation indexes, the wild dog position with the largest power grid detection evaluation index is selected, and the optimal task allocation of the corresponding multi - UAVs is determined according to the wild dog, and the multi - UAVs are coordinated to realize power grid detection.
[0193] The present invention has at least the following beneficial effects:
[0194] (1) The present invention combines the diverse predation mechanisms of the Australian wild dog algorithm. The established evolutionary strategy of the allocation scheme includes various evolutionary strategies such as collective hunting, individual predation, and scavenging, replacing the single evolutionary method of traditional swarm intelligence algorithms, endowing the individuals in the population with more autonomy, and avoiding the algorithm falling into local optimum.
[0195] (2) By modeling the characteristics of power grid detection, the present invention designs a dynamic adjustment mechanism for the objective function and power grid detection evaluation indexes, ensuring that the individuals in the population optimize towards the correct allocation scheme and improving the optimization efficiency of the algorithm.
[0196] (3) Based on the Australian wild dog algorithm, the present invention can quickly and effectively solve the problem of multi - UAV cooperative power grid detection task allocation, providing a new solution idea for solving the problem of multi - UAV cooperative power grid detection task allocation in complex environments.
[0197] As Figure 5 shown, the present invention also provides a multi - UAV cooperative power grid detection device, adopting the above - mentioned multi - UAV cooperative power grid detection method. The device includes:
[0198] A model construction module, which is used to establish a multi - UAV cooperative task planning model based on the dynamic cost and dynamic benefit required for UAV detection considering environmental factors;
[0199] A task allocation module, which is used to perform iterative optimization of a non-linear multi-objective problem based on a multi-UAV collaborative task planning model and a multi-UAV collaborative decision-making mechanism, and determine the optimal task allocation for multi-UAV collaborative power grid detection;
[0200] A power grid detection module, which is used to coordinate multiple UAVs to perform power grid detection based on the optimal task allocation.
[0201] Furthermore, the model construction module includes:
[0202] Determine the dynamic cost corresponding to UAV detection based on the detection cost and flight range cost generated by the UAV performing the detection task;
[0203] Determine the dynamic benefit of the UAV according to the detection accuracy of the UAV and the detection value of the power grid target;
[0204] Give the objective function of the multi-UAV collaborative task planning according to the dynamic cost and dynamic benefit of the UAV;
[0205] Determine the constraint conditions and the dynamic adjustment mechanism of the objective function of the UAV considering environmental factors, and establish a multi-UAV collaborative task planning model in combination with the objective function of the multi-UAV collaborative task planning.
[0206] Furthermore, the task allocation module includes:
[0207] S21. Establish a mapping between the position of the dingo individual in the dingo algorithm and the power grid target allocation scheme corresponding to the multi-UAV, determine the evolution strategy and coefficient decision factor of different hunting methods in the dingo algorithm, and initialize the position of the dingo in the dingo algorithm;
[0208] S22. Select the corresponding coefficient decision factor at the current dingo position to update the evolution strategy based on the multi-UAV collaborative decision-making mechanism, generate a new dingo position according to the updated evolution strategy, and calculate the corresponding power grid detection evaluation index;
[0209] S23. Repeat the iterative step S22 until the preset condition is met, output the dingo position with the optimal power grid detection evaluation index, and obtain the optimal task allocation for multi-UAV collaborative power grid detection.
[0210] Furthermore, the task allocation module includes:
[0211] Determine the state space of the reinforcement learning algorithm based on the multi-UAV collaborative task planning model and environmental factors, and determine the action space of the reinforcement learning algorithm based on the coefficient decision factor;
[0212] Determine the reward function of the reinforcement learning based on the multi-UAV collaborative task planning model;
[0213] Based on the current position of the wild dog, combined with the state space, action space, and reward function, determine the probability of selecting an action in the action space under the current state space, and based on the maximum probability, determine the coefficient decision factor corresponding to the selected action;
[0214] Update the evolutionary strategy according to the coefficient decision factor corresponding to the selected action.
[0215] Furthermore, the task allocation module includes;
[0216] Determine the fitness value calculated from the allocation scheme corresponding to the individual in the previous iteration and the fitness value calculated from the allocation scheme corresponding to the individual in the current iteration; select the minimum fitness value from the fitness value calculated from the allocation scheme corresponding to the individual in the previous iteration and the fitness value calculated from the allocation scheme corresponding to the individual in the current iteration, and perform a ratio calculation between the selected minimum fitness value and the fitness value calculated from the allocation scheme corresponding to the individual in the previous iteration; based on the difference between the benchmark threshold of the allocation scheme and the result of the ratio calculation, determine the short-term target value under the allocation scheme corresponding to the individual in the current iteration; wherein, the benchmark threshold of the allocation scheme is 1;
[0217] Determine the adjustment value based on the number of iterations, and combined with the number of times the individual obtained a fitness value less than the benchmark threshold using the allocation scheme corresponding to the individual in the current iteration and previous iterations, determine the balance term value; perform a ratio calculation between the number of times the individual obtained a fitness value less than the benchmark threshold using the allocation scheme corresponding to the individual in the current iteration and previous iterations and the total number of times the individual used the allocation scheme corresponding to the individual in the current iteration and previous iterations, and superimpose the balance term value to give the long-term target value under the allocation scheme corresponding to the individual in the current iteration;
[0218] Determine the weight parameters of the short-term target value and the long-term target value, and combined with the short-term target value and the long-term target value, give the power grid detection evaluation index for the current iteration.
[0219] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention. Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention also intends to include these changes and variations.
Claims
1. A multi-UAV collaborative power grid detection method, characterized in that: include: S1. Determine the dynamic cost corresponding to the drone detection based on the detection cost and flight cost generated by the drone performing the detection task; According to the detection accuracy of the UAV and the detection value of the power grid target, the dynamic benefits of the UAV are determined; according to the dynamic cost and dynamic benefits of the UAV, the objective function of multi-UAV collaborative task planning is given; the constraints of the UAV taking into account environmental factors and the dynamic adjustment mechanism of the objective function are determined, and the multi-UAV collaborative task planning model is established in combination with the multi-UAV collaborative task planning objective function; S21. Establish a mapping between the position of individual wild dogs in the dingo algorithm and the corresponding power grid target allocation scheme of multiple drones, and determine the evolutionary strategies and coefficient decision factors of different hunting methods of the dingo algorithm, where the hunting methods include collective hunting, individual hunting and scavenging; S22. Determine the state space of the reinforcement learning algorithm based on the multi-drone collaborative task planning model and environmental factors, and determine the action space of the reinforcement learning algorithm based on the coefficient decision factor; determine the reward function of reinforcement learning based on the multi-drone collaborative task planning model; based on the current wild dog position, combine the state space, action space and reward function, determine the probability of selecting an action in the action space under the current state space, and determine the coefficient decision factor corresponding to the selected action based on the maximum probability; update the evolutionary strategy according to the coefficient decision factor corresponding to the selected action, generate a new wild dog position according to the updated evolutionary strategy, and calculate the corresponding power grid detection evaluation index; S23. Repeat the iterative step S22 until the preset conditions are met, output the wild dog position with the best power grid detection evaluation index, and obtain the optimal task allocation for multi-drone collaborative power grid detection; S3. Based on the optimal task allocation, coordinate multiple UAVs to perform power grid inspection.
2. The multi-UAV collaborative power grid detection method according to claim 1, characterized in that: Determine the dynamic cost corresponding to drone detection, including: Obtain the distance between the UAV and the power grid target and the minimum distance that meets the detection quality, and give the interference distance; Based on the interference of environmental factors on drone detection and the detection power of drones, combined with the interference distance, the detection cost of drones is determined; The real-time wind speed and terrain undulation parameters at the power grid target are obtained and normalized to give the terrain complexity coefficient at the power grid target; Obtain the maximum Euclidean distance of all UAVs relative to the power grid target, and combine the terrain complexity coefficient at the power grid target and the distance between the UAV and the power grid target to give the UAV's range cost; The detection cost and flight cost of the drone are superimposed to give the dynamic cost corresponding to the drone detection.
3. The multi-UAV collaborative power grid detection method according to claim 2, characterized in that: The dynamic cost corresponding to drone detection satisfies the following relationship: ; In the formula, represents the dynamic cost of UAV i detecting power grid target j at time t; The detection cost of the drone satisfies the following relationship: ; In the formula, represents the detection cost of UAV i detecting power grid target j at time t, α is the detection weight factor, is the detection power of drone i, d ij ( t ) represents the distance from the detection grid target j at time t, d min To meet the minimum distance for detection quality, is the interference coefficient of environmental factors on UAV i’s detection of power grid target j at time t. Environmental factors include weather attenuation, terrain shielding, and obstacle density. is the preset interference threshold, x ij Indicates whether the power grid target j is detected by UAV i. When the power grid target j is detected by UAV i, x ij =1, otherwise x ij =0; The range cost of a drone satisfies the following relationship: ; In the formula, represents the range cost of UAV i detecting power grid target j, β is the cost weight factor, represents the terrain complexity coefficient of the location of the power grid target j, γ 1, γ 2 is the normalization coefficient, W j ( t ) is the real-time wind speed at the location of the power grid target j, T j is the terrain undulation parameter where the power grid target j is located, d max represents the maximum Euclidean distance of all UAVs relative to the grid target j.
4. The multi-UAV collaborative power grid detection method according to claim 1, characterized in that: The dynamic benefits of drones satisfy the following relationship: ; In the formula, represents the dynamic benefit of UAV i detecting power grid target j at time t, η 0 is the equity weight factor, λ i Indicates the detection accuracy of the sensor carried by drone i, F j is the inspection value of power grid target j, θ i ( t ) is the current heading angle of UAV i, is the optimal approach angle under environmental constraints.
5. The multi-UAV collaborative power grid detection method according to any one of claims 1 to 4, characterized in that: The objective function of multi-UAV collaborative mission planning satisfies the following relationship: ; In the formula, F ( t ) is the objective function for multi-UAV collaborative mission planning, represents the detection cost of UAV i detecting power grid target j at time t, represents the range cost of UAV i detecting power grid target j, R ij ( t ) represents the dynamic benefit of UAV i detecting power grid target j at time t, Ω( t ) represents the dynamic weight vector, which is updated in real time through the Kalman filter. , ω1, ω2, ω3 are all dynamic weights, Δ E ( t ) is the change of environmental parameters, K,Q is the adjustment matrix in the Kalman algorithm, [] T is the transpose operation; The dynamic adjustment mechanism of the objective function satisfies the following relationship: ; In the formula, is the rate of change of environmental parameters, is the environmental parameter change rate threshold, E(t) is the environmental parameter, ω'1 and ω'2 are the dynamic weights for updating and adjusting ω1 and ω2 respectively, is the change of communication interference, ΔV is the fluctuation amplitude of power grid value, V max is the maximum grid value.
6. The multi-UAV collaborative power grid detection method according to claim 1, characterized in that: The UAV takes into account the constraints of environmental factors and satisfies the following relationship: ; In the formula, Indicates the current maximum flight speed of drone i. is the default maximum flight speed of the drone. κ is the wind speed influencing factor, W j ( t ) is the real-time wind speed at the location of the power grid target j, W max is the maximum ambient wind speed, d safe ( t ) is the maximum safe flight distance, d min To meet the minimum distance for detection quality, μ is the safety margin adjustment factor, ρ(t) is the real-time obstacle density, x ij Indicates whether the power grid target j is detected by UAV i. When the power grid target j is detected by UAV i, x ij =1, otherwise x ij =0, T is the power grid target set of the detection task.
7. The multi-UAV collaborative power grid detection method according to claim 5, characterized in that: The evolutionary strategy of collective hunting satisfies the following relationship: ; The evolutionary strategy of single predation satisfies the following relationship: ; The evolutionary strategy of scavenging satisfies the following relationship: ; In the formula, are the new positions of wild dogs when participating in collective hunting, individual predation and scavenging, respectively. α(t), β(t) and γ(t) are the hunting intensity coefficient of collective hunting, the exploration rate parameter of individual predation and the scavenging trigger threshold of scavenging, i.e., the coefficient decision factors of different hunting methods. k∈{1,2,...,P}, P is the number of Australian wild dogs, X k represents the initial position of the selected wild dog, is the time-varying coupling coefficient, , a3 is the time-varying coupling weight, T max is the maximum duration of power grid detection, X best is the best individual among all wild dogs, that is, it has the current optimal allocation plan, F1 represents the adaptive hunting scaling factor, X m , X n are two random wild dogs participating in collective hunting, λ(t) is the dynamic learning rate, are two random individuals among all dingo individuals, is a random individual among all dingo individuals, a1 and a2 are the weight of individual predators and the threshold for updating scavenging positions, respectively, and rand is a random number between (0,1).
8. The multi-UAV collaborative power grid detection method according to claim 7, characterized in that: Based on the current position of the wild dog, combined with the state space, action space and reward function, determine the probability of selecting an action in the action space under the current state space, and based on the maximum probability, determine the coefficient decision factor corresponding to the selected action, including; Based on the current position of the wild dog, the reinforcement learning algorithm is used to evaluate the expected long-term benefits of each action in the action space under all states in the current state space before the update, and the expected long-term benefits corresponding to the next moment are determined; Select the maximum value from the expected long-term benefits corresponding to the next moment, and update the expected long-term benefits of each action in the action space under all states in the current state space in combination with the reward function; Determine the probability of selecting each action in the action space under all states in the current state space based on the expected long-term benefits of selecting each action in the action space under all states in the updated current state space; Screen and determine the action selected when the probability is the highest, and based on the relationship between the coefficient decision factor and the action space, determine the coefficient decision factor of the hunting method corresponding to the action selected when the probability is the highest.
9. The multi-UAV collaborative power grid detection method according to claim 8, characterized in that: The probability of selecting each action in the action space under all states in the current state space satisfies the following relationship: ; In the formula, π(a|s) is the current state space S t Select action space A in state s t The probability of action a in , Q'(s,a) is the updated current state space S t Select action space A in state s t The expected long-term benefit of action a in the above example is Q(s,a), and Q(s,a) is the current state space S before the update. t Select action space A in state s t The expected long-term benefit of action a, T0 is the total number of grid targets, l r is the learning rate, r(t) is the reward function, η is the discount factor, s', a' are the state at the next moment reached by executing action a in the current state s and the action selected by the state at the next moment, Δ F ( t ) is the difference between the objective function of multi-UAV collaborative mission planning at time t and time t-1, Ω opt is the small perturbation value of the dynamic weight vector, collision_rate is the collision rate, a4 and a5 are the benefit weights of minimizing the target weight and minimizing the dynamic weight vector respectively. F ( t ) is the objective function for multi-UAV collaborative mission planning, W j ( t ) is the real-time wind speed at the location of the power grid target j, is the interference coefficient of environmental factors on UAV i’s detection of power grid target j at time t, ρ(t) is the real-time obstacle density, α(t), β(t), and γ(t) are the hunting intensity coefficient of collective hunting, the exploration rate parameter of individual predation, and the scavenging trigger threshold of scavenging, respectively. Δα, Δβ, and Δγ are the discrete adjustment values of the hunting intensity coefficient, the exploration rate parameter, and the scavenging trigger threshold, respectively.
10. A multi-UAV collaborative power grid detection device, characterized in that: The multi-UAV collaborative power grid detection method according to any one of claims 1 to 9 is adopted, and the device comprises: The model building module is used to determine the dynamic cost of drone detection based on the detection cost and range cost generated by the drone performing the detection task; determine the dynamic benefit of the drone based on the detection accuracy of the drone and the detection value of the power grid target; give the multi-drone collaborative task planning objective function based on the dynamic cost and dynamic benefit of the drone; determine the constraints of the drone taking into account environmental factors and the dynamic adjustment mechanism of the objective function, and establish a multi-drone collaborative task planning model in combination with the multi-drone collaborative task planning objective function; The task allocation module is used to perform iterative optimization of nonlinear multi-objective problems based on a multi-UAV collaborative task planning model and a multi-UAV collaborative decision-making mechanism, and determine the optimal task allocation for multi-UAV collaborative power grid detection, including: S21, establishing a mapping between the positions of individual wild dogs in the dingo algorithm and the corresponding power grid target allocation schemes of multiple UAVs, and determining the evolutionary strategies and coefficient decision factors of different hunting methods of the dingo algorithm, where the hunting methods include collective hunting, individual hunting and scavenging; S22, determining the state space of the reinforcement learning algorithm based on the multi-UAV collaborative task planning model and environmental factors, and determining the coefficient decision factors of the reinforcement learning algorithm based on the coefficient decision factors. Action space; determine the reward function of reinforcement learning based on the multi-UAV collaborative task planning model; based on the current wild dog position, combine the state space, action space and reward function to determine the probability of selecting an action in the action space under the current state space, and based on the maximum probability, determine the coefficient decision factor corresponding to the selected action; update the evolutionary strategy according to the coefficient decision factor corresponding to the selected action, and generate a new wild dog position according to the updated evolutionary strategy, and calculate the corresponding power grid detection evaluation index; S23, repeat the iterative step S22 until the preset conditions are met, output the wild dog position with the best power grid detection evaluation index, and obtain the optimal task allocation for multi-UAV collaborative power grid detection; The power grid detection module is used to coordinate multiple UAVs to perform power grid detection based on optimal task allocation.
Citation Information
Patent Citations
Unmanned aerial vehicle task allocation method based on multi-target particle swarm
CN117170405A
Combined strategy algorithm for channel inspection and fine inspection of unmanned aerial vehicle
CN113342034A
Power inspection unmanned aerial vehicle path planning method, model training method and system
CN117742367A
Unmanned ship formation dynamic formation adjusting method considering switching cost
CN118859930A