Unmanned aerial vehicle group collaboration method based on combination of fuzzy ant colony algorithm and near-end strategy optimization

By combining fuzzy ant colony algorithm, near-end strategy optimization and particle swarm optimization, the problems of inefficient and insufficient adaptability in multi-drone collaborative tasks are solved, and efficient and intelligent task planning and execution are achieved.

CN119937592AActive Publication Date: 2025-05-06ZHEJIANG UNIV OF TECH

Patent Information

Application Number
CN202510102361.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-06
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The prior art is inefficient when multiple drones perform tasks in collaboratively, difficult to meet real-time requirements, and difficult to have adaptability in dynamic environments, resulting in a sharp increase in task planning complexity.

Method used

The UAV swarm collaboration method based on the combination of fuzzy ant colony algorithm and near-end strategy optimization is adopted. The task area division is optimized through adaptive clustering method, and the fuzzy ant colony algorithm performs path planning, and combined with near-end strategy optimization and particle swarm optimization, dynamic adjustment and global optimization are achieved.

Benefits of technology

It improves the intelligent control and efficiency of the drone group during the task execution process, realizes more efficient task allocation and path planning, and enhances the adaptability to the dynamic environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119937592A_ABST
    Figure CN119937592A_ABST
Patent Text Reader

Abstract

An unmanned aerial vehicle group cooperation method based on the combination of a fuzzy ant colony algorithm and near-end strategy optimization comprises the following steps: 1) distributing task areas for unmanned aerial vehicles through adaptive clustering, and narrowing a task range; 2) formulating a fuzzy rule and a membership function, solving a first-level task plan among different subgroups in the target group by adopting a fuzzy ant colony algorithm, and ensuring that the cost of each unmanned aerial vehicle is minimized when executing subgroup tasks; and 3) near-end strategy optimization and a particle swarm optimization method are combined to be used for task planning optimization in the target subgroup, so that the planning precision and efficiency are improved. According to the invention, intelligent control and efficiency of the unmanned aerial vehicle in a task execution process can be improved through collaborative task planning fusing time domain and space domain features, and the task execution efficiency of an unmanned aerial vehicle group is further enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of intelligent decision-making optimization, and in particular provides a drone group collaboration method based on the combination of fuzzy ant colony algorithm and proximal strategy optimization. Background Art

[0002] The rapid development of unmanned aerial vehicle (UAV) technology has led to its widespread application in military, logistics, environmental monitoring and other fields. In complex mission environments, the capabilities of a single UAV are limited and it is difficult to complete the mission efficiently. Therefore, studying the collaborative mission planning of UAV swarms has become the key to improving mission efficiency and reliability. When multiple UAVs perform tasks collaboratively, it is necessary to reasonably allocate tasks and plan paths to ensure efficient completion of the tasks. Traditional methods are often inefficient when dealing with complex environments and multiple tasks, and it is difficult to meet real-time requirements. In the process of performing tasks, UAV swarms may encounter dynamically changing environments, such as sudden obstacles and weather changes. How to enable UAV swarms to have adaptive capabilities to cope with environmental changes is an important challenge in collaborative mission planning.

[0003] In recent years, with the rapid development of artificial intelligence, communication technology and drone manufacturing technology, multi-drone systems and swarm technology have gradually become a hot topic of research in various countries. For example, the "Quail" project led by the US Department of Defense (USAIR Force) successfully demonstrated the ability of more than 100 drones to conduct coordinated reconnaissance and target strikes in a battlefield environment (SUAS, 2016). China has also made significant progress in drone swarm technology. China Electronics Technology Group Corporation has experimentally proved the possibility of coordinated flight of hundreds of drones for military and civilian fields (Wei Yunfeng et al., New Reading, 2017). In addition, the Dutch Army (defensie.nl, 2018), one of the applications of the European "HORNET" project, is exploring intelligent decision-making algorithms for using drone groups to perform complex tasks. Multi-drone systems are constantly promoting the development of swarm intelligence technology by drawing on the wisdom of social creatures in nature.

[0004] Proximal Policy Optimization (PPO) is a deep reinforcement learning algorithm with the characteristics of fast convergence speed and high stability. In UAV mission planning, PPO can learn the optimal strategy by interacting with the environment to achieve autonomous decision-making and path planning of the UAV. For example, Jiao Weidong (Aeronautical Computing Technology, 2024) and others used the PPO algorithm for UAV attitude control and achieved good results.

[0005] Ant Colony Optimization (ACO) is a swarm intelligence optimization algorithm that simulates the foraging behavior of ants and is suitable for solving path planning problems. In UAV mission planning, the ant colony algorithm can be used to find the optimal flight path, avoid obstacles, and reduce the probability of being discovered. For example, Shao Changxu (Ordnance Automation, 2018) et al. proposed a UAV route planning method based on the ant colony algorithm and verified its effectiveness. To this end, Zhang Lin (Northwestern Polytechnical University, 2020) introduced fuzzy logic into the ant colony algorithm to form a fuzzy ant colony algorithm to improve the global search capability and convergence speed of the algorithm. In UAV mission planning, the fuzzy ant colony algorithm can more effectively handle uncertainty and complexity and improve the efficiency and reliability of mission planning. In a dynamic environment, task allocation and path planning need to consider multiple factors such as task priority, time constraints, and environmental changes at the same time, which makes the complexity of problem solving increase dramatically.

[0006] Most existing evaluation methods find it difficult to obtain the global optimal solution in the complex environment of large-scale drone groups. They may have problems such as slow convergence and easy falling into local optimality, and may cause performance degradation in some interference environments. Summary of the invention

[0007] In order to overcome the above-mentioned shortcomings of the prior art and improve the intelligent control and efficiency of drones during mission execution, the present invention proposes a drone group collaborative mission method based on the combination of fuzzy ant colony algorithm and proximal strategy optimization.

[0008] When the UAV performs ground tasks, the collaborative path planning before the task is optimized to achieve efficient execution of multi-UAV collaborative tasks. First, the clustering results are optimized through an adaptive clustering method to determine the optimal task area division scheme. The present invention uses a fuzzy ant colony algorithm to solve the first-level task planning between different subgroups in the target group to ensure that the cost of each UAV is minimized when performing subgroup tasks. Secondly, the proximal strategy optimization is combined with the particle swarm optimization method to optimize the task planning within the target subgroup, thereby improving the planning accuracy and efficiency. Finally, through the collaborative task planning that integrates the time domain and airspace characteristics, the overall collaborative combat capability and task execution efficiency of the UAV group are further improved.

[0009] In order to solve the technical problem, the present invention adopts the following technical solution:

[0010] The first aspect of the present invention relates to a drone group collaboration method based on a combination of a fuzzy ant colony algorithm and proximal strategy optimization, comprising:

[0011] S1, an adaptive optimization method combining hierarchical clustering and DBSCAN clustering is used to adjust the task area allocation according to the location of the drone group and the target;

[0012] S2. Based on the allocation results of adaptive clustering, the fuzzy ant colony algorithm is introduced, and the pheromone intensity is dynamically adjusted in combination with fuzzy rules to intelligently allocate task target points to individual drones and optimize their path planning;

[0013] S3. Based on the preliminary path planning of fuzzy ant colonies, the proximal strategy optimization algorithm is used to dynamically adjust the path selection of individual drones. At the same time, the particle swarm optimization algorithm is combined to perform global search and optimization of the path, thereby generating the optimal task path for the drone swarm.

[0014] Wherein, step S1 comprises:

[0015] The coordinate matrix of the mission target point and the coordinate matrix of the UAV's departure point are taken as input, and the Euclidean distance matrix D is constructed based on the mission target point set; hierarchical clustering and DBSCAN clustering are applied to the target points respectively to generate different clustering results; then, the comprehensive quality score is calculated according to the clustering results, including the silhouette coefficient, Calinski-Harabasz index and Davies-Bouldin index, and the clustering scheme with the highest score is selected as the basis for UAV mission area allocation to achieve the optimization of task allocation.

[0016] Further, step S1 includes the following specific steps:

[0017] S1.1 uses the Ward method to gradually merge independent clusters initialized by mission target points to minimize the square error increment between clusters until the number of drones is reached and hierarchical clustering is completed. The Ward method merging formula is:

[0018]

[0019]

[0020] Among them, (x i ,y i ) and (x j ,y j ) are clusters C i and C j The center of the initial cluster is each task target point;

[0021] S1.2 uses the DBSCAN method to cluster the task target points, where the elements of the Euclidean distance matrix D are calculated using the following formula:

[0022]

[0023] Where (x p ,y p ) and (x q ,y q) are the coordinates of the target points p and q; by setting the neighborhood radius and the minimum number of samples, the high-density area is divided into clusters, and the sparse area is marked as noise points; the noise points are assigned to the nearest base station according to the Euclidean distance to the base station to ensure that all target points are effectively covered;

[0024] S1.3 evaluates the quality of clustering schemes S1.1 and S1.2 by using silhouette coefficient, cluster density and cluster compactness. The comprehensive scoring formula is as follows:

[0025]

[0026]

[0027] Among them, a(i) represents the average distance between sample i and other points in the same cluster, b(i) represents the average distance between sample i and other points in the nearest cluster, and μ is the mean of the entire data set. is the jth cluster C j The mean of n is the number of data points, k is the number of clusters, and s i represents the average divergence within cluster i, d ij Represents the distance between cluster i and cluster j; finally, the clustering scheme with the highest comprehensive score is selected as the optimal basis for the allocation of UAV mission areas.

[0028] Wherein, step S2 comprises:

[0029] In step S1, the mission area is divided into multiple smaller mission areas and assigned to the drone group; each drone starts from the starting base and selects a path based on the pheromone concentration on the path and the distance to the target point; in path planning, by introducing fuzzy rules, the target distance, obstacle density, energy consumption and mission urgency factors are comprehensively considered to dynamically adjust the update intensity of pheromones and optimize path selection; fuzzy reasoning converts input variables into precise values ​​for adjusting pheromones, and a shorter path obtains more information pheromones, thereby increasing the probability of being selected; after the maximum number of iterations, the pheromones gradually converge and eventually find the global optimal path.

[0030] Further, step S2 includes the following specific steps:

[0031] S2.1 distributes the task areas assigned by S1 to adjacent drone clusters and regards each task cluster as an ant population, where the number of samples in the cluster corresponds to the number of drones;

[0032] S2.2 Each drone starts from its starting base and randomly selects the next target node based on the pheromone concentration on the path and the distance between each target point, and gradually builds the path; the probability calculation formula for path selection is:

[0033]

[0034] Among them, T ij represents the pheromone concentration from node i to node j, d ij Represents the distance between nodes;

[0035] S2.3 introduces fuzzy rules in path selection, comprehensively considers target distance, obstacle density, energy consumption and task urgency, dynamically adjusts pheromone concentration through fuzzy reasoning, and optimizes the path selection process;

[0036] S2.4 After each round of iteration, the pheromones on the path are updated, and the shorter path accumulates more information pheromones, increasing the probability of being selected; within the maximum number of iterations, the pheromones are continuously optimized, and the path gradually converges to the global optimum, completing the efficient allocation and planning of UAV tasks.

[0037] Furthermore, in step S2.3, the specific process of fuzzy reasoning is as follows:

[0038] S2.3.1 The target distance, obstacle density, energy consumption, and task urgency are converted into fuzzy set membership through membership functions, as follows:

[0039] The target distance is divided into three fuzzy sets: "Near", "Medium" and "Far", and their membership functions are defined as follows:

[0040]

[0041] The obstacle density ρ is divided into three fuzzy sets: “Low”, “Medium” and “High”, and its membership function is defined as follows:

[0042] μ Low (ρ)=-e -5ρ (11)

[0043]

[0044] μ High (ρ)=1-e -5ρ (13)

[0045] Energy consumption E is divided into three fuzzy sets: "Low", "Medium" and "High", and its membership function is defined as follows:

[0046]

[0047] The task urgency U is divided into three fuzzy sets: "Low", "Medium" and "High", and its membership function is defined as follows:

[0048] μ Low (U) = -e-5U (17)

[0049]

[0050] μ High (U) = 1-e -5U (19)

[0051] S2.3.2 According to the established fuzzy rules, fuzzy reasoning is performed in combination with the degree of membership. The fuzzy rule table is as follows:

[0052]

[0053]

[0054] Among them, "+", "o" and "-" represent increasing, moderate and decreasing the pheromone update amount respectively; in general, when the target distance is close, the obstacle density is low, the energy consumption is low and the task urgency is high, the pheromone is increased and the current path is preferred; when the target distance is far, the obstacle density is high, the energy consumption is high and the task urgency is low, the pheromone is reduced and the exploration of other paths is encouraged; when there is a contradiction between the variables (such as the task is urgent but the obstacle density is high), the pheromone update amount is dynamically adjusted according to the task priority and resource conditions to achieve intelligent optimization of path selection;

[0055] S2.3.3 Defuzzify the fuzzy reasoning results by the centroid method and convert them into precise values ​​for dynamic adjustment of the pheromone concentration on the path; the pheromone on the path is updated according to the following formula:

[0056]

[0057] T ij (t+1)=(1-ρ)·T ij (t)+ΔT ij (twenty one)

[0058] Among them, A k is the set of ants in the ant colony, L k is the path length of ant k, Q is the total task weight, and ρ is the volatility coefficient, which is used to control the decay rate of pheromone;

[0059] Wherein, step S3 comprises:

[0060] Based on step S2, the path planning is modeled as a Markov decision process. The policy network is optimized by the proximal policy optimization method. The path length, energy consumption and task priority are used as reward functions to generate the optimal path strategy that can balance the time and energy constraints. Then, the global search capability of the particle swarm algorithm is combined to further optimize the path and make up for the shortcomings of local optimization, ultimately achieving more efficient and intelligent UAV path planning.

[0061] Further, step S3 includes the following specific steps:

[0062] S3.1 models the path planning problem as a Markov decision process, where the state represents the current path information of the drone, including the current position, visited target points, and remaining mission target points; the action corresponds to the next flight node selected by the drone; a clear reward function is constructed through path length, energy consumption, and task priority to evaluate the quality of each step of path selection while balancing time and energy constraints; by sampling path trajectories, calculating cumulative rewards, and continuously adjusting strategies, the drone gradually generates a path planning strategy that can maximize the reward function, thereby achieving more efficient path optimization and ensuring mission completion;

[0063] S3.2 Based on the path generated in step S3.1, the global search capability of particle swarm optimization is used to further optimize the path; finally, the global optimal path that minimizes the fitness function is found to achieve global optimization of the UAV path planning.

[0064] The second aspect of the present invention relates to a drone swarm collaboration device based on a combination of fuzzy ant colony algorithm and proximal strategy optimization, comprising a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the drone swarm collaboration method based on a combination of fuzzy ant colony algorithm and proximal strategy optimization of the present invention.

[0065] The third aspect of the present invention relates to a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the drone group collaboration method of the present invention based on the combination of fuzzy ant colony algorithm and proximal strategy optimization.

[0066] The working principle of the present invention is:

[0067] The entire UAV swarm collaborative task planning process can be divided into four stages: "task area division → initial task allocation and preliminary path planning → local dynamic adjustment → global optimization". Each stage has corresponding algorithms and methods to achieve it. Combination of hierarchical clustering and DBSCAN: Adaptively divide the task area into multiple reasonable sub-areas by considering density and hierarchy at the same time, reducing the repeated flight or conflict problems of UAVs in disordered scheduling over a large range. Fuzzy ant colony algorithm: Through the fusion of bionics + fuzzy theory, first obtain a relatively optimal initial solution in the local (sub-area), and can take into account multiple constraints and environmental uncertainties. Proximal strategy optimization: Through the "trial and error + feedback" mechanism of reinforcement learning, the current UAV's flight path is perceived and corrected in real time, providing adaptive adjustments for local emergencies. Particle swarm optimization: Use swarm intelligence algorithms to iteratively update the overall layout and routes of UAVs in a larger search space, avoid fragmentation or under-optimization caused by local dynamic decisions, and ensure better overall efficiency of the UAV swarm.

[0068] The advantages of the present invention are:

[0069] The UAV group task planning method based on the combination of "hierarchical clustering + DBSCAN adaptive partitioning", "fuzzy ant colony algorithm preliminary path planning" and "PPO + PSO dynamic and global optimization" proposed in the patent has the following advantages:

[0070] 1. By combining hierarchical clustering with DBSCAN, it can adaptively identify dense areas, sparse distributions, and noise points, and effectively deal with complex scenarios such as uneven distribution of task points or sudden increase or decrease of task points; the partitioning is more reasonable, and the search space for subsequent task allocation and path planning is effectively narrowed, reducing repeated flights and conflicts.

[0071] 2. Fuzzy logic is introduced on the basis of traditional ant colony algorithm, and comprehensive evaluation is performed through "fuzzy ant colony algorithm"; the generated initial allocation and path plan are more robust than the pure ant colony algorithm, and perform better in uncertain or multi-constrained environments.

[0072] 3. Proximal Policy Optimization (PPO) can perceive and adapt to environmental changes in real time during the mission execution, and quickly fine-tune the local path of the UAV; Particle Swarm Optimization (PSO) performs periodic or triggered re-optimization of the overall path or task allocation from a global perspective, effectively avoiding local optimality and ensuring overall efficiency.

[0073] 4. Combining the "real-time online learning capability" of reinforcement learning (PPO) with the "global search" capability of swarm intelligence algorithms (PSO, ant colony) can not only cope with instantaneous disturbances and emergencies, but also ensure the optimization of long-term planning goals; the training process converges more smoothly and the deployment is safer and more reliable by limiting the difference between new and old strategies.

[0074] 5. Each submodule (clustering method, task allocation algorithm, local / global optimization algorithm) has a certain degree of interchangeability: the clustering method (K-means++, etc.) can be replaced according to the scenario, the fuzzy rules or parameters of the fuzzy ant colony algorithm can be adjusted, and other reinforcement learning algorithms and optimization algorithms can be combined; it is suitable for various drone group sizes and various types of mission scenarios, such as large-scale target patrols, disaster monitoring, logistics distribution, and military reconnaissance.

[0075] 6. Through adaptive partitioning, intelligent allocation and local-global dual optimization, mutual interference and redundant flights between drones are significantly reduced; at the same time, a better balance can be achieved in low energy consumption, low latency and high safety. BRIEF DESCRIPTION OF THE DRAWINGS

[0076] Figure 1 It is a flowchart of the UAV group collaboration method based on the combination of fuzzy ant colony algorithm and proximal strategy optimization.

[0077] Figure 2 It is a flowchart of the UAV swarm task allocation based on the adaptive clustering method.

[0078] Figure 3 It is a flowchart of UAV mission path allocation based on fuzzy ant colony.

[0079] Figure 4 It is a flow chart of UAV mission path optimization based on proximal strategy optimization. DETAILED DESCRIPTION

[0080] The technical solution of the present invention is further described below in conjunction with the accompanying drawings.

[0081] Example 1

[0082] Reference Figure 1-4 ,The UAV group collaboration method based on the combination of fuzzy ant colony algorithm and ,proximal strategy optimization includes the following steps:

[0083] S1, an adaptive optimization method combining hierarchical clustering and DBSCAN clustering is used to adjust the task area allocation according to the location of the drone group and the target;

[0084] In order to solve the problem that the traditional single clustering algorithm is difficult to adapt to various scenarios due to the complex distribution of mission target points in the UAV mission area division, this paper proposes a method that combines hierarchical clustering and DBSCAN clustering. This method optimizes the clustering results by calculating the comprehensive scoring index to determine the optimal mission area division scheme.

[0085] Specifically, the coordinate matrix of the mission target point and the coordinate matrix of the UAV's departure point are used as input, and hierarchical clustering and DBSCAN clustering are performed on the mission target point set to obtain different clustering results. Subsequently, for each clustering result, the comprehensive score of its clustering quality is calculated to select the optimal clustering result as the basis for the UAV mission area allocation. Finally, the optimal target point cluster is associated with the UAV, and the task allocation matrix is ​​constructed by analyzing the inter-cluster distance and the attributes of the UAV, thereby completing the division of the mission area.

[0086] This method significantly improves the efficiency and rationality of task allocation in complex scenarios by introducing a clustering result optimization strategy, and provides a high-quality basic area division scheme for subsequent UAV path planning and task execution. The specific process of step S1 is as follows:

[0087] S1.1. Hierarchical cluster analysis

[0088] In a certain mission scenario, the radar system detects several target points, forming a target point set Target = {Target1, Target2, …, Target N These target points may be stationary targets on the ground or dynamic areas to be monitored. In order to achieve effective task allocation and path planning of drones, the relative positions of the target points must be clarified first.

[0089] In this case, we first need to calculate the Euclidean distance between the target points to obtain the distance matrix D. The construction of the distance matrix D can intuitively describe the spatial distribution relationship between the target points and provide basic data support for subsequent clustering analysis and path planning. ij The calculation formula is:

[0090]

[0091] Among them, (x j ,y j ) represents the coordinates of the target point j, d ij Represents the distance between the i-th target and the j-th target.

[0092] Next, each task target point is considered as an independent cluster. The number of initial clusters is N, which is the number of target points. At the same time, the center coordinates (x c ,y c ), the center of each cluster is the coordinate of its corresponding target point at the beginning. Then the Ward method is used to gradually merge clusters until the preset target number of clusters k≤N is reached. The Ward method is a clustering method based on variance minimization. It selects the merged clusters by minimizing the sum of squared deviations within the merged clusters to ensure that each merged cluster has the best compactness. The specific merging process is as follows:

[0093]

[0094]

[0095]

[0096] in, and Cluster C i and C j The centroid (mean) of the clusters. Formula 2 describes the calculation method of the inter-cluster distance based on the Ward method, Formula 3 describes the calculation formula of the centroid, Formula 4 describes the criterion for selecting the two clusters with the smallest distance to merge, and Formula 4 describes the process of merging two clusters. Through this step-by-step merging method, the required number of target clusters can be obtained, and the target points in the cluster can be ensured to be as close as possible.

[0097] S1.2. DBSCAN cluster analysis

[0098] For any target point, its ε-domain is defined as follows:

[0099] N ε (p)={q|d(p,q)≤ε} (6)

[0100] Among them, d(p,q) represents the distance between point p and point q (usually using Euclidean distance), and its calculation formula is:

[0101]

[0102] Where (x p ,y p ) and (x q ,y p ) are the coordinates of the target point p and the target point q.

[0103] Set the minimum number of samples MinPts. If the number of points contained in the ε-domain of point p is |N ε If (p)|≥MinPts, then point p is called a core target point; if point p is located in the ε-domain of a core target point, but the number of points in its own neighborhood is less than MinPts, then point p is called a boundary target point; if point p is neither a core target point nor in the neighborhood of any core target point, then point p is called a noise target point.

[0104] Starting from any unvisited core target point p, mark it as visited and expand the cluster from it. First, take the ε-domain N of point p. εFor all points in (p), if a point in its neighborhood is a core point, continue to add points in its neighborhood to the current cluster until no new points can be added. For point p, if its neighborhood N ε If there are other core points in (p) (i.e., the number of points in the neighborhood is greater than or equal to MinPts), then the points in the neighborhood of these core points are recursively added to the current cluster until all points are visited and assigned to a cluster.

[0105] This process ensures that all core points and points in their neighborhood are correctly classified into the same cluster, thus achieving effective clustering of target points.

[0106] S1.3 Comprehensive analysis of clustering results

[0107] In order to select the best clustering result, it is necessary to score the results of hierarchical clustering and DBSCAN clustering. The comprehensive scoring indicators include the following three items: Silhouette Coefficient, Calinski-Harabasz (CH) index and Davies-Bouldin (DB) index.

[0108] The silhouette coefficient is used to measure the difference between the similarity of a sample to the samples in its cluster and the difference between a sample and the samples in other clusters. Its value is between [-1,1], and the larger the value, the better the clustering effect. The calculation formula of the silhouette coefficient is:

[0109]

[0110] Among them, a(i) represents the average distance between sample i and other points in the same cluster, and b(i) represents the average distance between sample i and other points in the nearest cluster.

[0111] The Calinski-Harabasz (CH) index measures the compactness and separation of clusters. The larger the value, the better the clustering effect. The calculation formula of the CH index is:

[0112]

[0113] in, represents the trace of the inter-class distance difference matrix, represents the trace of the intra-class dispersion matrix, μ is the mean of the entire data set, is the jth cluster C j , n is the number of data points, and k is the number of clusters.

[0114] The Davies-Bouldin (DB) index is used to measure the similarity between clusters. The smaller the value, the better the clustering effect. The calculation formula of the DB index is:

[0115]

[0116] Among them, s i represents the average divergence within cluster i, d ij represents the distance between cluster i and cluster j.

[0117] Finally, the comprehensive score of the clustering results can be calculated based on the above three indicators. The calculation formula is:

[0118] Comprehensive score = silhouette coefficient + CH index - DB index (11) By scoring the clustering results of hierarchical clustering and DBSCAN, the method with the highest comprehensive score is selected as the final clustering solution for the task area division. This process ensures the best balance of clustering results in terms of compactness, separation and overall effect, so as to provide a high-quality division solution for the reasonable allocation of UAV tasks.

[0119] S2. Initial assignment of UAV mission paths based on fuzzy ant colony

[0120] By learning from the behavior of ants in finding food paths in nature, the path planning problem of drones can be transformed into a process of finding the optimal path. The Fuzzy Ant Colony Optimization (FACO) algorithm combines the traditional ant colony algorithm with fuzzy logic, and introduces a fuzzy control mechanism to improve the robustness and adaptability of the algorithm, especially in complex and dynamic environments.

[0121] In step S1, the larger task area has been divided into several smaller task areas, and the adjacent drone groups are assigned their responsible spaces, forming a gridded task layout. In this case, each task cluster can be regarded as an ant population, where the number of samples in the cluster corresponds to the number of ants in the population, and each ant represents a drone. Each drone starts from its starting base position, randomly selects a neighboring node as the target node, and then continues to select the next node from this node. The probability of path selection depends on the pheromone concentration on the current path and heuristic information (such as the distance between nodes and the importance of the task). The probability calculation formula for path selection is:

[0122]

[0123] Among them, T ij represents the pheromone concentration from node i to node j, η ij =1 / d ij is the heuristic factor, indicating the distance between nodes; α and β are the weights of pheromone concentration and heuristic factor, respectively.

[0124] Through this mechanism, the ant colony algorithm simulates the process of a drone searching for the optimal path in a complex environment. Pheromone concentration plays a key role in path selection, while heuristic factors such as distance and mission importance help the drone make more reasonable decisions. This approach effectively balances global search with local exploration, providing a reliable solution for efficient path planning for drone missions.

[0125] In the process of path selection, fuzzy rules are introduced to adjust the update intensity of pheromones, so that ants can more intelligently avoid obstacles and choose better paths when searching for paths. The following fuzzy rules are formulated based on the actual situation for the four key influencing factors of target distance, obstacle density, energy consumption and task urgency:

[0126] (1) When the target distance is far, the obstacle density is high, and the task urgency is low, the pheromone update amount should be “reduced” to encourage ants to explore other paths.

[0127] (2) When the target distance is close, the obstacle density is low, and the task urgency is high, the pheromone update amount should be “increased” to give priority to the current best path.

[0128] (3) When energy consumption is high and the task urgency is high, paths with higher energy efficiency should be prioritized to reduce energy consumption.

[0129] (4) When both the obstacle density and the mission urgency are high, a trade-off needs to be made between risk and efficiency, and a complex but faster path may be chosen.

[0130] The formulation of these fuzzy rules enables the UAV to dynamically adjust the path selection strategy according to different environmental conditions, thereby achieving more intelligent and efficient path planning. Specifically, the target distance d is divided into three fuzzy sets: "Near", "Medium", and "Far", and the corresponding membership functions are defined as follows:

[0131]

[0132] Among them, formula (13) adopts the inverse s-type function, so that the membership degree quickly approaches 1 in a relatively close distance range; formula (14) adopts the Gaussian function, and the membership degree is the highest in a medium distance range; formula (15) adopts the S-type function, so that the membership degree quickly approaches 1 in a relatively large distance range.

[0133] The obstacle density ρ is divided into three fuzzy sets: "Low Density", "Medium Density", and "High Density", and their membership functions are as follows:

[0134] μ Low (ρ)=-e -5ρ (16)

[0135]

[0136] μ High (ρ)=1-e -5ρ (18)

[0137] Among them, formula (16) adopts the inverse exponential function, and the membership is higher in the low density range; formula (17) adopts the triangular membership function, ρ1, ρ2, and ρ3 are 0.3, 0.5, and 0.7 respectively, and the peak is located at the medium density; formula (18) adopts the exponential function, and the membership increases rapidly in the high density range.

[0138] The energy consumption E is divided into three fuzzy sets: "Low Consumption", "Medium Consumption", and "High Consumption", and the membership functions are as follows:

[0139]

[0140] Among them, formula (19) adopts the inverse S-type function; formula (20) adopts the Bell-shaped membership function to achieve a smooth transition to the middle area; formula (21) adopts the S-type function.

[0141] The task urgency U is divided into three fuzzy sets: "Low Urgency", "Medium Urgency", and "High Urgency". The specific membership functions are:

[0142] μ Low (U) = -e -5U (twenty two)

[0143]

[0144] μ High (U) = 1-e -5U (twenty four)

[0145] Among them, formula (22) adopts the inverse exponential function; formula (23) adopts the bell-shaped membership function, which reflects the high membership degree of the middle value; formula (24) adopts the exponential function.

[0146] Through the above membership function, the input variables actually detected by the radar (such as target distance, obstacle density, etc.) are converted into corresponding fuzzy set membership. Fuzzy reasoning is performed by combining these memberships with the formulated fuzzy rules, and finally the reasoning results are converted into accurate values ​​through the centroid defuzzification method, which is used to dynamically adjust the pheromone update strategy of the ants.

[0147] Under the preset maximum number of iterations k, the pheromone on the path will be updated after each iteration. Shorter paths will accumulate more pheromones, thereby increasing the probability of being selected by other ants and gradually finding the optimal path. The pheromone update formula is:

[0148]

[0149] Among them, A k represents the set of ants in the ant colony, ΔT ij represents the pheromone increment left by ant k on path i→j. The calculation method of incremental pheromone is:

[0150]

[0151] Among them, L k is the length of the path taken by ant k, and Q is the total weight of the task.

[0152] After each pheromone update, the pheromone will gradually evaporate, simulating the natural dissipation process of pheromones in reality. The pheromone volatilization formula is:

[0153] T ij (t+1)=(1-ρ)·T ij (t)+ΔT ij (27)

[0154] Among them, ρ is the volatility coefficient, which controls the decay rate of the pheromone.

[0155] When the set maximum number of iterations is reached, the ants stop searching for paths and select the path with the highest pheromone concentration as the preliminary optimal path.

[0156] S3. Re-allocation of UAV mission paths based on proximal strategy optimization

[0157] S3.1 Path optimization based on PPO

[0158] Based on the fuzzy ant colony algorithm, a preliminary optimization path was obtained, and the path planning problem was further modeled as a Markov decision process (MDP) to improve the intelligence of path optimization. Specifically, the state represents the current path of the drone, and the action corresponds to the selection of the next flight node. In order to more comprehensively measure the path quality, the reward function is defined as a weighted combination of path length, energy consumption, and task priority factors.

[0159] In this model, the PPO (Proximal Policy Optimization) algorithm is used to further optimize the path selection strategy, whose goal is to minimize the loss function. The objective function form of PPO is:

[0160]

[0161] Among them, r t (θ) represents the ratio of the current strategy to the old strategy; is the advantage function, which is used to measure the pros and cons of the current action; ∈ is the clipping range to prevent the strategy from being updated too large.

[0162] By setting the number of iterations, PPO continuously optimizes the policy network and gradually finds the path planning strategy that can maximize the reward function. After training, the strategy generated by PPO can predict the optimal path, thereby further optimizing the flight path of the drone on the basis of combining the fuzzy ant colony algorithm and improving the intelligence and overall efficiency of path selection.

[0163] S3.2 Path optimization based on PSO

[0164] Preliminary work has trained the policy network through PPO (Proximal Policy Optimization) to optimize path selection, so that the UAV can more effectively balance time, energy consumption and other constraints in path planning. In order to further refine the path optimization, the particle swarm optimization (PSO) algorithm can be combined to perform global search and adjustment of the path to find a better path solution.

[0165] At this stage, the idea of ​​particle swarm optimization is to represent each particle as a possible path solution, and the dimension of the particle corresponds to each flight node in the path. The position of the particle represents a possible solution to the path, usually represented by the coordinates of each node, while the speed of the particle represents the change in the path.

[0166] In each iteration, the particle updates its speed and position based on its own historical best position and the global best position among all particles. The update formula is as follows:

[0167] v i (t+1)=ω·v i (t)+c1·r1·(p best -x i (t))+c2·r2·(g best -x i (t)) (17)

[0168] This formula is used to calculate the velocity of a particle at the next moment, which includes three parts: inertial term, cognitive term and social term. i(t) ensures that the particle continues to move in the current direction; the cognitive term c1·r1·(p best -x i (t)) guides the particle to move toward its best historical position; the social term c2·r2·(g best -x i (t)) guides the particle to approach the global optimal position. i (t) is the velocity of particle i at time t, x i (t) is the position (i.e. path) of particle i at time t, p best is the best historical position of particle i, g best is the global best position (i.e. the best path among all particles), c1 and c2 are learning factors that control the degree to which particles learn about their own historical best positions and the global best positions, r1 and r2 are random numbers that ensure the randomness of the algorithm, and ω is the inertia weight that controls the degree of particle adaptation during the search. In this way, particles find a balance between the global and local best positions and gradually approach the optimal solution.

[0169] Then, according to the updated velocity, the particle's position will also be updated, and the update formula is:

[0170] x i (t+1)=x i (t)+v i (t+1) (29)

[0171] This formula is used to update the position of the particle, that is, to determine the new position of the particle at the next moment based on the current position and the latest calculated speed. This allows each particle to continuously move in the solution space and explore new path solutions.

[0172] In order to measure the quality of each path, a fitness function is set, which usually takes into account factors such as path length, flight time and energy consumption. The form of the fitness function is:

[0173] F(x i )=α·L(x i )+β·E(x i )+γ·T(x i ) (30)

[0174] Among them, L(x i ) is the path length, E(x i ) is the energy consumption of the path, T(x i ) is the task completion time, α, β, γ are weight coefficients used to balance various factors.

[0175] In the whole optimization process, the PPO optimization strategy network is first used to determine an effective path. On this basis, the particle swarm optimization (PSO) is further applied to iteratively search and finally find the path that can minimize the fitness function, which is the global optimal path, thus achieving the overall optimization of the UAV path planning.

[0176] Example 2

[0177] The present embodiment relates to a drone swarm collaboration device based on a combination of a fuzzy ant colony algorithm and proximal strategy optimization, including a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the drone swarm collaboration method based on a combination of a fuzzy ant colony algorithm and proximal strategy optimization of Example 1.

[0178] Example 3

[0179] The present embodiment relates to a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, the drone group collaboration method based on the combination of fuzzy ant colony algorithm and proximal strategy optimization of embodiment 1 is implemented.

[0180] The embodiments of this specification are merely examples of implementations of the invention and are for illustrative purposes only. The scope of protection of the present invention should not be considered limited to the specific forms described in this embodiment, but also includes equivalent technical means that can be thought of by ordinary technicians in this field based on the invention.

Claims

1. A drone group collaboration method based on the combination of fuzzy ant colony algorithm and proximal strategy optimization is characterized by: include: S1, an adaptive optimization method combining hierarchical clustering and DBSCAN clustering is used to adjust the task area allocation according to the location of the drone group and the target; S2. Based on the allocation results of adaptive clustering, the fuzzy ant colony algorithm is introduced, and the pheromone intensity is dynamically adjusted in combination with fuzzy rules to intelligently allocate task target points to individual drones and optimize their path planning; S3. Based on the preliminary path planning of fuzzy ant colonies, the proximal strategy optimization algorithm is used to dynamically adjust the path selection of individual drones. At the same time, the particle swarm optimization algorithm is combined to perform global search and optimization of the path, thereby generating the optimal task path for the drone swarm.

2. The drone group collaboration method based on the combination of fuzzy ant colony algorithm and proximal strategy optimization according to claim 1 is characterized in that: Step S1 includes: The coordinate matrix of the mission target point and the coordinate matrix of the UAV's departure point are taken as input, and the Euclidean distance matrix D is constructed based on the mission target point set; hierarchical clustering and DBSCAN clustering are applied to the target points respectively to generate different clustering results; then, the comprehensive quality score is calculated according to the clustering results, including the silhouette coefficient, Calinski-Harabasz index and Davies-Bouldin index, and the clustering scheme with the highest score is selected as the basis for UAV mission area allocation to achieve the optimization of task allocation.

3. The drone group collaboration method based on the combination of fuzzy ant colony algorithm and proximal strategy optimization according to claim 2 is characterized in that: Step S1 includes the following specific steps: S1.1 uses the Ward method to gradually merge independent clusters initialized by mission target points to minimize the square error increment between clusters until the number of drones is reached and hierarchical clustering is completed. The Ward method merging formula is: Among them, (x i ,y i ) and (x j ,y j ) are clusters C i and C j The center of the initial cluster is each task target point; S1.2 uses the DBSCAN method to cluster the task target points, where the elements of the Euclidean distance matrix D are calculated using the following formula: Where (x p ,y p ) and (x q ,y q ) are the coordinates of the target points p and q; by setting the neighborhood radius and the minimum number of samples, the high-density area is divided into clusters, and the sparse area is marked as noise points; the noise points are assigned to the nearest base station according to the Euclidean distance to the base station to ensure that all target points are effectively covered; S1.3 evaluates the quality of clustering schemes S1.1 and S1.2 by using silhouette coefficient, cluster density and cluster compactness. The comprehensive scoring formula is as follows: Among them, a(i) represents the average distance between sample i and other points in the same cluster, b(i) represents the average distance between sample i and other points in the nearest cluster, and μ is the mean of the entire data set. is the jth cluster C j The mean of n is the number of data points, k is the number of clusters, and s i represents the average divergence within cluster i, d ij Represents the distance between cluster i and cluster j; finally, the clustering scheme with the highest comprehensive score is selected as the optimal basis for the allocation of UAV mission areas.

4. The drone group collaboration method based on the combination of fuzzy ant colony algorithm and proximal strategy optimization according to claim 1 is characterized in that: Step S2 includes: In step S1, the mission area is divided into multiple smaller mission areas and assigned to the drone group; each drone starts from the starting base and selects a path based on the pheromone concentration on the path and the distance to the target point; in path planning, by introducing fuzzy rules, the target distance, obstacle density, energy consumption and mission urgency factors are comprehensively considered to dynamically adjust the update intensity of pheromones and optimize path selection; fuzzy reasoning converts input variables into precise values ​​for adjusting pheromones, and a shorter path obtains more information pheromones, thereby increasing the probability of being selected; after the maximum number of iterations, the pheromones gradually converge and eventually find the global optimal path.

5. The drone group collaboration method based on the combination of fuzzy ant colony algorithm and proximal strategy optimization according to claim 4 is characterized in that: Step S2 includes the following specific steps: S2.1 distributes the task areas assigned in step S1 to adjacent drone clusters, and regards each task cluster as an ant population, where the number of samples in the cluster corresponds to the number of drones; S2.2 Each drone starts from its starting base and randomly selects the next target node based on the pheromone concentration on the path and the distance between each target point, and gradually builds the path; the probability calculation formula for path selection is: Among them, T ij represents the pheromone concentration from node i to node j, d ij Represents the distance between nodes; S2.3 introduces fuzzy rules in path selection, comprehensively considers target distance, obstacle density, energy consumption and task urgency, dynamically adjusts pheromone concentration through fuzzy reasoning, and optimizes the path selection process; S2.4 After each round of iteration, the pheromones on the path are updated, and the shorter path accumulates more information pheromones, increasing the probability of being selected; within the maximum number of iterations, the pheromones are continuously optimized, and the path gradually converges to the global optimum, completing the efficient allocation and planning of UAV tasks.

6. The drone group collaboration method based on the combination of fuzzy ant colony algorithm and proximal strategy optimization according to claim 5 is characterized in that: In step S2.3, the specific process of fuzzy reasoning is as follows: S2.3.1 The target distance, obstacle density, energy consumption, and task urgency are converted into fuzzy set membership through membership functions, as follows: The target distance is divided into three fuzzy sets: "Near", "Medium" and "Far", and their membership functions are defined as follows: The obstacle density ρ is divided into three fuzzy sets: "Low", "Medium" and "High", and its membership function is defined as follows: m Low (p)=-e -5ρ (11) m High (ρ)=1-e -5ρ (13) Energy consumption E is divided into three fuzzy sets: "Low", "Medium" and "High", and its membership function is defined as follows: The task urgency U is divided into three fuzzy sets: "Low", "Medium" and "High", and its membership function is defined as follows: μ Low (U)=-e -5U (17) μ High (U)=1-e -5U (19) S2.3.2 According to the established fuzzy rules, fuzzy reasoning is performed in combination with the degree of membership. The fuzzy rule table is as follows: Among them, "+", "o" and "-" represent increasing, moderate and decreasing the pheromone update amount respectively; in general, when the target distance is close, the obstacle density is low, the energy consumption is low and the task urgency is high, the pheromone is increased and the current path is preferred; when the target distance is far, the obstacle density is high, the energy consumption is high and the task urgency is low, the pheromone is reduced and the exploration of other paths is encouraged; when there is a contradiction between the variables, the pheromone update amount is dynamically adjusted according to the task priority and resource conditions to achieve intelligent optimization of path selection; S2.3.3 Defuzzify the fuzzy reasoning results by the centroid method and convert them into precise values ​​for dynamic adjustment of the pheromone concentration on the path; the pheromone on the path is updated according to the following formula: T ij (t+1)=(1-ρ)·T ij (t)+ΔT ij (21) Among them, A k is the set of ants in the ant colony, L k is the path length of ant k, Q is the total task weight, and ρ is the volatility coefficient, which is used to control the decay rate of pheromone.

7. The drone group collaboration method based on the combination of fuzzy ant colony algorithm and proximal strategy optimization according to claim 1 is characterized in that: Step S3 includes: Based on step S2, the path planning is modeled as a Markov decision process. The policy network is optimized by the proximal policy optimization method. The path length, energy consumption and task priority are used as reward functions to generate the optimal path strategy that can balance the time and energy constraints. Then, the global search capability of the particle swarm algorithm is combined to further optimize the path and make up for the shortcomings of local optimization, ultimately achieving more efficient and intelligent UAV path planning.

8. The drone group collaboration method based on the combination of fuzzy ant colony algorithm and proximal strategy optimization according to claim 7 is characterized in that: Step S3 includes the following specific steps: S3.1 models the path planning problem as a Markov decision process, where the state represents the current path information of the UAV, including the current position, visited target points, and remaining mission target points; The action corresponds to the next flight node selected by the drone. A clear reward function is constructed through path length, energy consumption and task priority to evaluate the quality of each path selection while balancing time and energy constraints. By sampling path trajectories, calculating cumulative rewards and continuously adjusting strategies, the drone gradually generates a path planning strategy that can maximize the reward function, thereby achieving more efficient path optimization and ensuring task completion. S3.2 Based on the path generated in step S3.1, the global search capability of particle swarm optimization is used to further optimize the path; finally, the global optimal path that minimizes the fitness function is found to achieve global optimization of the UAV path planning.

9. The UAV group collaboration device based on the combination of fuzzy ant colony algorithm and proximal strategy optimization is characterized by: It includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the drone group collaboration method based on the combination of fuzzy ant colony algorithm and proximal strategy optimization as described in any one of claims 1-8.

10. A computer-readable storage medium, characterized in that: A program is stored thereon, and when the program is executed by a processor, the drone group collaboration method based on the combination of fuzzy ant colony algorithm and proximal strategy optimization described in any one of claims 1-8 is implemented.

Citation Information

Patent Citations

  • Realization of keyword optimization based on fuzzy c-mean algorithm of ant colony

    CN106897376A

  • A fuzzy clustering method based on improved ant colony algorithm for tongue diagnosis image segmentation

    CN109509196A

  • Multi-unmanned aerial vehicle layout and task unloading decision-making method based on multi-target joint optimization

    CN117519252A

  • Autonomous data collection and system control for material recovery facilities

    EP4385632A1

  • Dynamic contextual road occupancy map perception for vulnerable road user safety in intelligent transportation systems

    WO2021194590A1

Cited By

  • Unmanned aerial vehicle flight control method and device supporting electric power emergency communication

    CN120215397A

  • Near-end strategy enhanced ant colony optimization path coverage method

    CN121187309A

  • Multi-AUV (Autonomous Underwater Vehicle) task area division and allocation method based on limited resources

    CN121436589A

  • Hybrid optimization algorithm for collaborative scheduling and dynamic obstacle avoidance of robot cluster

    CN121560033A

  • Particle swarm test case generation method of unmanned aerial vehicle training parallel program

    CN121658390A