Distributed unmanned aerial vehicle cluster collaborative search method and system based on evolutionary game theory
Through evolutionary game theory and information fusion technology, the search strategy of drone clusters is dynamically optimized, which solves the problem of autonomous decision-making of distributed drone clusters in complex environments and achieves efficient collaborative search and robustness.
Patent Information
- Application Number
- CN202510778511.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-06-11
AI Technical Summary
In existing technologies, distributed drone swarms find it difficult to fully utilize the advantages of autonomous decision-making in collaborative search tasks, and fixed strategies cannot adapt to complex dynamic environments, and there is a problem of cumbersome parameter optimization.
A method based on evolutionary game theory is adopted to dynamically update the search strategy through strategy optimization and adaptive learning between drones. Monte Carlo prediction and information fusion technology are used to optimize the collaborative search process of drone clusters.
It realizes efficient collaborative search of drone clusters in complex dynamic environments, improves mission efficiency and robustness, has strong adaptability, and reduces the complexity of parameter optimization.
Smart Images

Figure CN120670679A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of drone clusters, and in particular to a collaborative search method and system for distributed drone clusters based on evolutionary game theory. Background Art
[0002] Research on drone swarms collaborating to complete search missions is becoming increasingly important. For complex tasks, it's crucial to improving efficiency and success rates. Compared to a single drone, multiple drones can significantly improve mission efficiency, scenario coverage, and robustness by interacting and understanding environmental information, enabling rational task allocation, collaborative decision-making, and adaptive adjustments. Swarms of multiple drones have high practical application value for collaborative search for multiple dynamic targets within an area, and collaborative search within swarms is fundamental research for many practical applications.
[0003] For drone swarms, a distributed architecture is a common approach, delegating decision-making power to each individual drone. Compared to centralized architectures, drones in distributed architectures can make fully autonomous decisions based on their own sensory information and communication information with other drones. Furthermore, distributed control architectures offer a wider range of application scenarios and greater development potential (especially in large-scale swarms or highly dynamic environments). However, ensuring efficient and effective information exchange and understanding between drones in a distributed swarm, as well as ensuring the global nature of the distributed swarm system, remains a hot topic.
[0004] Furthermore, existing technologies for collaborative search within distributed drone swarms often rely on fixed collaborative strategies, employing pre-set rules or algorithms. While these approaches are simple to implement, they often fail to fully leverage the advantages of autonomous decision-making for fully autonomous individuals. Furthermore, fixed strategies not only make it difficult for drones to adapt to diverse and complex dynamic environments, but also present cumbersome parameter optimization challenges.
[0005] Therefore, a new collaborative search method for drone clusters under a distributed architecture is urgently needed, which can fully utilize the advantages of fully autonomous decision-making of drones under a distributed architecture, thereby dynamically optimizing strategies, while also ensuring that the group decisions made by the distributed cluster system have a certain degree of globality. Summary of the Invention
[0006] In view of this, an embodiment of the present invention provides a collaborative search method and system for distributed drone clusters based on evolutionary game theory, which can dynamically optimize the search strategy when the drone cluster performs collaborative search tasks, thereby giving full play to the advantages of the drone's fully autonomous decision-making, and the strategy optimization can also enable the drone to adapt to complex dynamic environments more quickly.
[0007] One aspect of the present invention provides a collaborative search method for a distributed drone cluster based on evolutionary game theory. The method comprises: For the strategy optimization phase, during the planning step, the following steps are performed: Drone Networking: Based on the maximum communication distance between drones, the current drone communicates with other drones in the cluster to form a connected subnet. The current drone applies a strategy in the strategy space, and the connected subnet includes multiple drones applying different strategies. The current drone also stores the corresponding relationship between each strategy and expected reward in the strategy space, as well as the search information graph. UAVs fuse information and determine the optimal decision path: The current UAV receives the search information graph stored by other UAVs in its connected subnet and uses the received search information graph to update its own stored search information graph. Based on the updated search information graph and the current UAV's applied strategy, it selects an optimal decision path within the rolling planning time domain from the mission search area and calculates the expected execution benefit of the optimal decision path. The drone updates the correspondence between each strategy and expected benefit: based on the expected execution benefit, the expected benefit corresponding to the strategy applied by the current drone is updated. The current drone receives the strategies and updated expected benefits applied by other drones in its connected subnet, and thus updates the stored correspondence according to the set expected benefit update rules. The drone calculates the strategy evaluation value: Based on the expected execution benefit and the updated expected benefit corresponding to the current drone application strategy, the relative evaluation value of the strategy currently applied by the drone is calculated. Then, based on the relative evaluation value at the current planning step and the relative evaluation value at a specific planning step, the weighted relative evaluation value of the strategy currently applied by the drone is calculated. The specific planning step refers to the planning step of the current drone's strategy during the period from the historical planning step with a set sliding window length to the current planning step. The drone updates its current strategy: The current drone determines the maximum value of the weighted relative evaluation value within its connected subnet, calculates the strategy learning probability corresponding to the current drone based on the weighted relative evaluation value of the strategy currently applied by the drone and the maximum value of the weighted relative evaluation value, and updates the strategy currently applied by the drone to the strategy corresponding to the maximum value of the weighted relative evaluation value based on the strategy learning probability; During the strategy optimization phase, the following steps are performed in the execution domain: UAVs execute tasks and update search information: In the execution domain, the current UAV executes the search task according to the optimal decision path, captures the target and outputs the target position when the target search conditions are met, and updates its own stored search information graph based on the detection information during the execution of the search task; wherein, the execution domain and the planning domain are both the planning step time of selecting the optimal decision path and one or more consecutive planning step times thereafter, and the planning domain is greater than or equal to the execution domain; after the steps in the execution domain are completed, if the search stop condition is not met, it jumps to the next planning step time and repeats all the steps in the above planning step time and execution domain, and finally completes the collaborative search of the UAV cluster.
[0008] In some embodiments of the present invention, for each strategy in the corresponding relationship, the expected benefit update rule set includes: If the expected return of the strategy in the corresponding relationship is not 0, then the updated expected return of the strategy is the original expected return in the corresponding relationship; If the expected return of the strategy in the corresponding relationship is 0, and the strategy is not the strategy currently applied by any drone in the connected subnet to which the current drone belongs, the updated expected return corresponding to the strategy is the default value; If the expected return of the strategy in the corresponding relationship is 0, and the strategy is the strategy currently applied by other drones in the connected subnet to which the current drone belongs, then the updated expected return corresponding to the strategy is the average of the expected returns obtained by other drones in the connected subnet to which the drone belongs that adopt this strategy.
[0009] In some embodiments of the present invention, the weighted relative evaluation value of the strategy currently applied by the UAV is calculated based on the relative evaluation value at the current planning step and the relative evaluation value at a specific planning step, including: In the case of a specific planning step, the average value of the relative evaluation value of the specific planning step is calculated; the weighted relative evaluation value of the strategy currently applied by the UAV is obtained by weighted summing the relative evaluation value of the current planning step and the average value; In the absence of a specific planning step, the relative evaluation value of the strategy currently applied by the UAV is used as the weighted relative evaluation value of the strategy currently applied by the UAV.
[0010] In some embodiments of the present invention, the correspondence between each strategy and expected benefit in the strategy space stored by the drone is obtained through a Monte Carlo prediction operation; For each drone in the drone cluster, the Monte Carlo prediction operation includes the following steps: at the planning step, the drones form a network, fuse information, and determine the optimal decision path; in the execution time domain, the drones execute tasks and update search information; and the drones calculate expected returns; Among them, the drone calculates the expected return by performing the following operations: after reaching the prediction stop condition, calculate the average value of the actual returns obtained by the current drone applying the strategy to perform the search task at all planning steps of the Monte Carlo prediction operation, and use it as the expected return corresponding to the strategy currently applied by the drone, and then obtain the corresponding relationship between each strategy stored in the current drone and the expected return.
[0011] In some embodiments of the present invention, the search information graph stored by the drone is determined based on the target existence probability, environmental uncertainty, and pheromone information detected by the drone when performing a search mission; and a two-dimensional Gaussian distribution is used to initialize and model the target existence probability in the search information graph stored by the drone; The search stop condition is that the planning step reaches the set number of search planning steps or the number of targets captured by the drone cluster reaches the set number of targets; the prediction stop condition is that the planning step reaches the set number of prediction planning steps; and The policy learning probability is calculated using the Fermi function.
[0012] In some embodiments of the present invention, based on the updated search information graph and the current drone application strategy, an optimal decision path within the rolling planning planning horizon is selected from the mission search area, and the expected execution benefit of the optimal decision path is calculated, including: In the planning domain, a traversal algorithm is used to calculate all feasible paths for the current UAV within the mission search area. The path benefits of each feasible path are calculated based on the updated search information graph. Based on the calculated path benefits, a set number of paths are selected from all feasible paths as the UAV's pre-decision paths. The path benefits of the feasible paths are determined based on the search value benefits and coordination benefits of the feasible paths. Based on the current UAV application strategy and the path benefit of the pre-decision path, the expected execution benefit of the pre-decision path is calculated, and the optimal decision path is selected from the pre-decision paths based on the calculated expected execution benefit.
[0013] In some embodiments of the present invention, the expected execution benefit of the pre-decision path is determined based on the strategy currently applied by the drone, the search value benefit of the pre-decision path, and the coordination benefit of the pre-decision path; Among them, the search value benefit of the pre-decision path is the sum of the pheromone concentration differences in all task sub-areas passed by the pre-decision path; the coordination benefit of the pre-decision path is the sum of the target existence probability and environmental uncertainty in all task sub-areas passed by the pre-decision path; among them, the pheromone concentration difference is the difference between the attraction pheromone concentration and the repulsion pheromone concentration.
[0014] In some embodiments of the present invention, a current drone receives a search information graph stored by another drone in its connected subnet, and uses the received search information graph to update its own stored search information graph, including: For the target existence probability map and environmental uncertainty map in the search information map, the UAV receives the target existence probability map and environmental uncertainty map stored by other UAVs in its connected subnet; For each task sub-area within the task search area, when the sub-area is within the search range of the current UAV, the target existence probability map and environmental uncertainty map stored by the current UAV are used as the target existence probability map and environmental uncertainty map in the search information map updated by the current UAV respectively; For each task sub-area in the task search area, when the sub-area is outside the search range of the drone, the average value of the target existence probability of the sub-area in the target existence probability map stored by other drones in the connected subnet to which the current drone belongs is calculated, and the square root of the product of the environmental uncertainty of the sub-area in the environmental uncertainty map stored by other drones in the connected subnet to which the current drone belongs is calculated. The calculation results are respectively used as the target existence probability map and environmental uncertainty map in the updated search information map of the current drone.
[0015] Another aspect of the present invention provides a collaborative search system for a distributed drone cluster based on evolutionary game theory, comprising a processor, a memory, and a computer program / instruction stored in the memory, wherein the processor is used to execute the computer program / instruction, and when the computer program / instruction is executed, the system implements the steps of the method described in any of the above embodiments.
[0016] Another aspect of the present invention provides a computer-readable storage medium having a computer program / instruction stored thereon, which implements the steps of the method described in any of the above embodiments when executed by a processor.
[0017] This paper proposes a collaborative search method and system for distributed drone swarms. Based on evolutionary game theory, this system adaptively learns and optimizes strategies during the collaborative search process. Through this dynamic evolutionary process, drones within a swarm can adaptively select the optimal search strategy, achieving efficient swarm collaboration within a distributed architecture and significantly improving collaborative search efficiency.
[0018] Additional advantages, objects, and features of the present invention will be set forth in part in the following description and will become apparent to those skilled in the art upon examination of the following or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained by the structures particularly pointed out in the description and drawings.
[0019] Those skilled in the art will understand that the purposes and advantages that can be achieved by the present invention are not limited to the above specific descriptions, and the above and other purposes that can be achieved by the present invention will be more clearly understood based on the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The drawings described herein are used to provide a further understanding of the present invention, constitute a part of this application, and do not constitute a limitation of the present invention. In the drawings: Figure 1 Schematic diagram of a drone cluster collaboratively searching for moving targets in one embodiment of the present invention.
[0021] Figure 2 Schematic diagram of dividing the task search area using a rasterization method in one embodiment of the present invention.
[0022] Figure 3 Schematic diagram of a drone route in one embodiment of the present invention.
[0023] Figure 4 Schematic diagram of the search range of the drone-mounted sensor in one embodiment of the present invention.
[0024] Figure 5 Schematic diagram of the process of a collaborative search method for a distributed drone cluster based on evolutionary game theory in one embodiment of the present invention.
[0025] Figure 6 Schematic diagram of initial distribution of search information in one embodiment of the present invention.
[0026] Figure 7 This is a schematic diagram of initializing partitions of a task area in one embodiment of the present invention.
[0027] Figure 8 Schematic diagram of a flow chart of a collaborative search method for a distributed drone cluster based on evolutionary game theory in another embodiment of the present invention. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments and the accompanying drawings. Here, the exemplary embodiments of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0029] It should also be noted that, in order to avoid obscuring the present invention due to unnecessary details, the accompanying drawings only show structures and / or processing steps closely related to the solutions according to the present invention, while other details that are not closely related to the present invention are omitted.
[0030] It should be emphasized that the term "include / comprises" when used herein refers to the existence of features, elements, steps or components, but does not exclude the existence or addition of one or more other features, elements, steps or components.
[0031] It should also be noted that, unless otherwise specified, the term "connection" herein may refer not only to a direct connection but also to an indirect connection involving an intermediate.
[0032] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In the accompanying drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0033] In a distributed architecture, each drone in a swarm can make fully autonomous decisions. However, current technologies often rely on fixed collaborative strategies to perform collaborative search tasks, which often fail to fully utilize the advantages of autonomous decision-making. Furthermore, fixed collaborative strategies not only make it difficult for drones to adapt to diverse and complex dynamic environments, but also present cumbersome parameter optimization issues.
[0034] Evolutionary Stable Strategy (ESS) is a core concept in evolutionary game theory. Specifically, it means that in the process of the game, since the players are boundedly rational, it is impossible for them to find the optimal strategy and the optimal equilibrium point at the beginning. Therefore, the players need to constantly learn, gradually correct strategic errors, and constantly imitate and improve the most advantageous strategies of themselves and other players in the past. After a period of imitation and error correction, all players will tend to a stable strategy.
[0035] Based on this, the present application proposes a distributed collaborative search method for drone clusters based on evolutionary game theory, which dynamically learns and optimizes strategies during the collaborative search process, giving full play to the advantages of autonomous decision-making of drones. Specifically, as the search progresses, each drone in the cluster updates the expected benefit of the current corresponding strategy, and evaluates the current corresponding strategies of each drone in the connected subnet based on the updated expected benefit; each drone optimizes and updates its strategy based on the strategy evaluation value using the evolutionary game mechanism, so that strategies with excellent performance are accepted by more drones. In addition, in the method proposed in the present application, drones can make completely autonomous decisions, that is, drones can continuously update their search strategies based on their own sensor information and communication information with other drones during the game phase, but make their own decisions without considering the decisions of other drones in the cluster during the decision phase.
[0036] The current UAV cluster collaborative search process is as follows: Figure 1 As shown in the figure, within a specific mission search area (also referred to as the mission area), there are m moving targets to be detected. The number, location, and direction of movement of these targets are unknown a priori. A swarm of n drones, without any prior information, conducts a collaborative search mission for the moving targets within the specific mission search area. The goal of the collaborative search mission is to ensure that each drone captures as many targets as possible within the set mission time domain through efficient collaborative search, while minimizing the number of false positives and ensuring its own safety, taking into account the detection performance of onboard sensors and errors caused by environmental factors.
[0037] In order to better describe environmental information, simplify the solution space of search decisions, and improve the efficiency of solving the UAV cluster collaborative search problem, a rasterization method can be introduced. That is, the rasterization method can be used to divide the task search area during the unmanned collaborative search process. By dividing the task search area into multiple task sub-areas, each UAV searches one or more sub-areas, thereby making the information detected by the UAV more refined. For example, Figure 2 As shown, the task area can be divided into N x ×N y Grid cells of equal size (also called task sub-regions), and each grid cell can be used as an independent information carrier. For the convenience of subsequent representation, each sub-region can be numbered according to its position in the task search area. For example, the position of each sub-region can be represented as (in this application, (x, y) represents the task sub-region in the xth column and yth row), then each sub-region can be recorded as G xy . Figure 2 The task sub-area division method shown is only an example. This application requires that the task sub-areas do not overlap and have equal areas, and does not specifically limit the division method and shape of the task sub-areas.
[0038] Furthermore, compared to the mission area, drones conducting a clustered collaborative search within a larger mission area can be viewed as point masses moving in a two-dimensional plane, ignoring the effects of their altitude variations. Assuming that each drone in the cluster is located within a mission sub-area at each planning step, the drones can search for targets within the mission sub-area within their search range, acquiring information about the mission area through onboard sensors, and their flight paths are subject to boundary and velocity constraints. Assume that the two-dimensional motion model of drone i can be described as follows: Among them, (x i (t), y i (t)) represents the sub-area where the projection of UAV i in the mission area is located at the planning step t, d i (t) represents the heading angle of UAV i at planning step t. For example, Figure 3 As shown, when the drone type is a rotary-wing drone, it can be assumed that the drone has nine possible movement directions, corresponding to nine different heading angles di(t) = {45°·v|v = 0, ±1, ±2, ±3, ±4}.
[0039] In addition, the search range of the UAV (the search range can also be called the detection range) is affected by many parameters, including the altitude of the UAV, the physical parameters of the sensor itself, and the uncertainty of the detection environment. Figure 4 As shown, it is assumed that the detection range of the drone's onboard sensors does not change with the drone's altitude, and that the drones in the cluster of this application can be homogeneous drones. Therefore, the detection range of each drone within the mission area is fixed. Furthermore, the search range of drone i in this application can be defined as: the grid cell where drone i is located, as well as the grid cells that share a common edge with the grid cell where drone i is located, belong to the search range of drone i within the mission area.
[0040] The drone cluster will perform collaborative search at each planning step. The process of executing collaborative search at each planning step is the same, and the steps for each drone in the cluster to execute collaborative search are the same. Therefore, in the following, the collaborative search method proposed in this application takes the planning step t as an example, and uses a single drone in the cluster as the execution entity to describe the collaborative search method proposed in this application.
[0041] The collaborative search method proposed in this application mainly performs strategy optimization operations in the strategy optimization stage, such as Figure 5 As shown, the method may include steps S110 to S160 in the strategy optimization stage.
[0042] Step S110: UAV networking. Based on the maximum communication distance between UAVs, the current UAV establishes communication connections with other UAVs in the cluster to form a connected subnet. Because the positions of UAVs in this application are dynamic, step S110 must be executed at each planning step to ensure the effectiveness of UAV communication. This means that the connected subnet is continuously updated during the collaborative search task.
[0043] Specifically, if multiple drones within a cluster can communicate with each other directly or indirectly (for example, drone i and drone j can communicate directly or indirectly through other drones), then these drones can form a connected subnet. A drone cluster can form at least one connected subnet, and no two connected subnets contain the same drone.
[0044] As an example, existing solutions usually form connected subnets based on the maximum communication distance between drones. In the process of executing a collaborative search task, the drone cluster can be networked based on the communication distance threshold at each planning step, and a directed graph G = (U, D) is used to represent the communication connection relationship between drones, where U represents the set of all drones in the cluster and D represents the set of Euclidean distances between drones in the task search area (also referred to as the task area). If the communication connection relationship between drones in the cluster is stored in matrix form, the adjacency matrix A of the drone cluster can be defined as: Among them, a ij represents the communication connection status between drone i and drone j, n represents the number of drones in the drone set U, d ij represents the Euclidean distance between UAV i and UAV j, C r It can be seen from formula (2) that when the distance between two drones is less than or equal to the threshold C r When , it is considered that a communication link can be established between them, that is, a ij =1 (UAV i and UAV j are in communication connection state), when the distance between the two UAVs is greater than the threshold C r When , it is considered that no communication link can be established between them, that is, a ij = 0. In addition, it can be set by default that drone i cannot establish a communication link with itself, that is, a ii =0.
[0045] Furthermore, based on the above adjacency matrix A, the communication topology matrix C of the entire drone cluster (also known as the reachability matrix of the directed graph G = (U, D)) can be calculated. The calculation formula of the communication topology matrix c is:
[0046] For example, suppose the total number of drones in the cluster is n = 4, and the adjacency matrix at planning step t is: It can be calculated that: Element c in the communication topology matrix C ij Indicates whether there is at least one communication path from UAV i to UAV j. If c ij ≠0, it means that there is a communication path from UAV i to UAV j. Therefore, the communication topology matrix C calculated above can be further simplified to a binary matrix C′: The above C′ results show that in this example, the four drones in the drone cluster are in a fully connected state at the planning step t, that is, there is a direct or indirect communication path between any two drones in the cluster. At this time, the drone cluster contains only one connected subnet, and the drones in this subnet are the cluster itself.
[0047] From the above example, we can see that the specific process of cluster formation of connected subnets is as follows: based on the maximum communication distance between drones, the adjacency matrix A of multiple drones is determined, and the communication topology matrix C is constructed based on the adjacency matrix A; if each element c in the communication topology matrix C is ij Both are c ij ≠0, then these multiple drones can form a connected subnet, so that these drones in the same connected subnet can communicate. A drone cluster can form one or more connected subnets. For example, at the planning step t, the connected subnet set N(t) of the drone cluster can be expressed as: N(t)={(UVA1, UVA2, ..., UVA K1 ),(UVA1,UVA2,...,UVA K2 ),...,(UVA1,UVA2,...,UVA Kw )};...(6) Among them, UVA Kw It represents the Kw-th UAV in a connected subnetwork at the current planning step t.
[0048] In addition, if at planning step time t, there is a UAV that cannot communicate directly or indirectly with other UAVs in the UAV cluster, then the UAV still performs the collaborative search task at planning step time t, but does not execute steps S140 and S150.
[0049] In some embodiments of the present invention, while executing a collaborative search mission, a drone collects information within the mission search area, integrates its own stored historical information, information obtained through detection, and information from other drones obtained through communication on a connected subnet, and optimizes strategies and makes decisions based on this integrated information. In the collaborative search method proposed in this application, step S120 includes: Step S121, the drone integrates the information to obtain an updated search information graph: The current drone receives the search information graph stored by other drones in its connected subnet, and updates its currently stored search information graph with the received search information graph.
[0050] In some embodiments of the present invention, the search information stored by each drone in a cluster can be presented in the form of a search information graph. Because each drone's stored search information graph may differ, the search information graphs of other drones obtained through communication via the connected subnet may differ, and the search areas and collected environmental information of each drone may differ, the updated search information graphs of each drone in the cluster may differ at the same planning step. Furthermore, the types of search information stored by drones typically include various types, such as target presence probability, and can be customized based on the collaborative search task. To specifically illustrate the subsequent information fusion process, this application uses the example of three types of search information: target presence probability, environmental uncertainty, and pheromone concentration. Therefore, the search information graph may include a target probability map (TPM), an environmental uncertainty map (EUM), and a digital pheromone map (DPM). Specifically, the search information graph stored by a drone is determined based on the target presence probability, environmental uncertainty, and pheromone information detected by the drone during the search task.
[0051] The search information graph is a collection of historical information about the search area stored by drones. Using individual sub-regions as information carriers, it records information about each sub-region within the mission area based on drone detection results and communication information. This reflects the drone swarm's understanding of the current mission area during the collaborative search mission. The search information graph is continuously updated during the drone swarm's detection process. The updated graph further deepens the swarm's understanding of the mission area, guiding the drone swarm to further plan search paths and make decisions based on the updated search information graph. This is a positive feedback process.
[0052] More specifically, at planning step t, the UAV cluster is based on the maximum communication distance C rAfter forming a connected subnet, drones within the same subnet need to access the search information maps of other drones in addition to their own stored search information maps. In other words, to enhance drones' global perception of the mission search area, deepen their understanding of the environment, and guide the drone cluster to make better decisions, each drone within the cluster needs to fuse information.
[0053] As an example, the target existence probability information mainly includes the probability of the target existing in the mission area (the target existence probability designed in this application is [0,1], where the target existence probability of 0 indicates that the drone believes that there is definitely no target in the sub-area, and the target existence probability of 1 indicates that the drone believes that there is definitely a target in the sub-area). The environmental uncertainty information mainly includes environmental uncertainty (used to reflect the degree of information mastery of the drone cluster on the mission area). The pheromone information is used to characterize the concentration of attractive information and the concentration of repulsive pheromones in the mission area to guide the movement direction of the drone.
[0054] In some embodiments of the present invention, each drone in a cluster selectively adopts a fusion strategy when performing information fusion, that is, it only accepts and fuses environmental information outside the detection range of the drone, while retaining local information within the detection range. This fusion strategy is designed to fully utilize global information while retaining the detection information of the single drone itself, and to perform focused classification processing on the information. Therefore, in this application, the process of a drone receiving a search information graph stored by other drones in its connected subnet and updating its own stored search information graph with the received search information graph includes: For the target existence probability map and environmental uncertainty map in the search information map, the UAV receives the target existence probability map and environmental uncertainty map stored by other UAVs in its connected subnet; For each task sub-area in the task search area, when the sub-area is within the search range of the current UAV, the target existence probability map and environmental uncertainty map stored by the current UAV are used as the target existence probability map and environmental uncertainty map in the search information map updated by the current UAV, respectively. When the sub-area is outside the search range of the UAV, the average value of the target existence probability of the sub-area in the target existence probability map stored by other UAVs in the connected subnet to which the current UAV belongs is calculated, and the square root of the product of the environmental uncertainty of the sub-area in the environmental uncertainty map stored by other UAVs in the connected subnet to which the current UAV belongs is calculated. The average value of the target existence probability and the square root of the product of the environmental uncertainty are used as the target existence probability map and environmental uncertainty map in the search information map updated by the current UAV, respectively.
[0055] Specifically, the fusion mechanism of the target existence probability map and the environment uncertainty map can be expressed as follows: Among them, p i (x, y, t) represents the target existence probability map stored by UAV i at planning step t, and subregion G xy The probability of target existence, p i (x, y, t) co Indicates that for sub-region G xy , the target existence probability updated by UAV i at planning step t, G i (t) represents the set of sub-areas included in the search range of UAV i at planning step t, Represents the total number of drones in the connected subnet to which drone i belongs.
[0056] Among them, η i (x, y, t) represents the environmental uncertainty map stored by UAV i at planning step t, and sub-region G xy The environmental uncertainty, η i (x, y, t) co Indicates that for sub-region G xy , the environmental uncertainty updated by UAV i at planning step t.
[0057] According to formula (7) and formula (8), if the sub-region G xy Located within the search range of UAV i (i.e. G xy in G i (t)), the target existence probability map and the environmental uncertainty map stored by the UAV are retained separately, and no information fusion is required; on the contrary, if the sub-region G xy Outside the detection range of drone i, information is received from other drones in the connected subnet and integrated to obtain the target presence probability and environmental uncertainty in the subregion corresponding to drone i. This fusion mechanism ensures the integrity of information within the detection range. While maintaining the integrity of information in key areas, drones can use this fused information to expand their global perception capabilities, thereby achieving efficient decision-making.
[0058] Furthermore, since the pheromone concentrations of each sub-area within the mission area are determined and independent (for example, each sub-area within the mission area may store pheromones at different concentrations), this application does not fuse the pheromone information stored by each drone. Furthermore, although the pheromone information stored by each drone is not fused, the pheromone concentration gradually evaporates over time within each decision cycle (in this application, executing steps S110 to S160 once is considered a decision cycle), as shown in the following equations (11) to (15).
[0059] As can be seen from the above, after executing step S121, the search information map stored in the drone is updated to the fused target existence probability map, the fused environmental uncertainty map, and the pheromone information map of concentration changes over time.
[0060] As an example, a search infographic could be defined as: ① The target probability map is used to represent the probability of the target existing in each task sub-area. TPM can be defined as: P i (t) = {p i (x,y,t)|x≤N x , y≤N y};……………………………………(9) Among them, P i (t) represents the probability of the target existing in the mission area stored by UAV i at the planning step t, p i (x, y, t) represents the sub-region G stored by drone i at planning step t xy The probability of the target existing at {x≤N x , y≤N y} represents the sub-region G xy Located in the mission area. i (x, y, t) = 0 means that at planning step t, UAV i thinks that G xy There can be no target, p i (x, y, t) = 1 means that at planning step t, UAV i believes that G xy There must be a goal.
[0061] ② The environmental uncertainty map is used to reflect the degree of information mastery of the drone in each mission sub-area, usually measured by information entropy. EUM can be defined as: U i (t) = {η i (x,y,t)|x≤Nx , y≤Ny};……………………………………(10) Among them, U i (t) represents the information mastery degree of UAV i on the mission area at planning step t, η i (x, y, t) represents the position of UAV i in sub-region G at planning step t. xy The degree of information mastery, η i (x, y, t) = 1 means that at planning step t, UAV i has a certain range in sub-region G. xy The information inside is completely unknown, η i (x, y, t) = 0 means that at planning step t, UAV i has a certain range in sub-region G. xyFully master the information within.
[0062] ③ The digital pheromone map is used to represent the pheromone concentration within the mission area at each planning step. Each drone maintains a separate digital pheromone map. Because different drones may not agree on the same location, the pheromone map stored by each drone can be different.
[0063] Pheromones, also known as external hormones, are a key medium for information transmission between organisms in nature. They are secreted by individuals and spread through media such as air and water. They are detected by conspecifics through olfaction, triggering behavioral, emotional, psychological, or physiological changes. Pheromones are widely used in nature for collaboration among social organisms such as ants and bees. For example, ants release pheromones to mark paths, guiding their companions to food sources or avoiding dangerous areas. In collaborative search missions involving drone swarms, drawing on the communication and interaction mechanisms between organisms in nature, the concept of pheromones can be introduced to mark search areas and target locations, thereby simulating and optimizing swarm collaborative behavior. Specifically, attractive and repulsive pheromones can be designed to guide drone swarms to better complete collaborative search missions, thereby improving algorithm performance and robustness. When a target is detected in a subarea, given the uncertainty of onboard sensors, multiple reconfirmations are required to confirm the target's presence. Therefore, pheromones are also called revisit pheromones.
[0064] For drone i, DPM can be defined as: Among them, s(t) represents the pheromone information in the task area at the planning step time t, s a (x, y, t) represents the sub-region G at the planning step time t. xy The concentration of attractive pheromones in r (x, y, t) represents the sub-region G at the planning step time t xy The concentration of repulsive pheromones within.
[0065] (1) Regarding attraction pheromones: attraction pheromones guide the drone to fly to the key area that has been detected before through gravity, and revisit the sub-areas within the area to confirm the existence of the target. In order to simplify the model, the action process of attraction pheromones is described as: release - when certain conditions are met (such as the sub-area has not been visited for more than the revisit time threshold; the detection result is that there is a target in the sub-area, etc.), the sub-area releases pheromones; volatilization - the pheromones in the sub-area are reduced at a certain ratio in each decision cycle. If the drone cluster detects the sub-area G xy When there is a target in the memory, the sub-region G xyThe attraction pheromone switch will be turned on, releasing new attraction pheromones to guide the drone to revisit the possible target area as soon as possible.
[0066] At planning step t, the subregion G xy The concentration of attraction pheromones within can be defined as: Where, ρ represents the volatility of pheromone, s a (x, y, t-1) represents the sub-region G at planning step time t-1 xy The concentration of attractive pheromone in the a Indicates the incremental unit of attracting pheromone, It represents the attraction pheromone switch coefficient, which can be adjusted according to the probability of the detected target existing. It is defined as follows: Among them, q(x, y, t-1) represents the time when UAV i detects the sub-area G at the planning step t-1. xy Is there a target inside? Indicates that the attraction pheromone switch is turned on. Indicates that the attraction pheromone switch is off.
[0067] In this application, planning step moments can be discrete. Planning step moments refer to the time at which the drone performs a collaborative search mission. Therefore, the time interval between adjacent planning step moments is a decision cycle. For simplicity, planning step moment t-1 is used in this application to represent the previous planning step moment before planning step moment t.
[0068] (2) For repulsive pheromones, at planning step t, sub-region G xy The probability of the internal target existing is greater than the set target existence probability threshold p t (p t is an experience value), it is considered that the drone has captured the target in the sub-area. At this time, the probability of the target existing in the sub-area needs to be updated to 0, and the sub-area will not be visited again in a short period of time. In order to improve the globality of the search and reduce invalid searches, repulsive pheromones are introduced. The function process is described as follows: Release - when certain conditions are met (for example, after the drone executes the capture command for the corresponding sub-area, it can release repulsive pheromones), the sub-area releases pheromones; Volatility - the pheromones in the sub-area are reduced at a certain ratio in each decision cycle. When the sub-area G xy When the target existence probability in the sub-area G exceeds the set target existence probability threshold, it is considered that the target has been captured by the drone, that is, within a short period of time, the sub-area G xy No new targets will appear in the area, the repulsion pheromone switch is turned on, and the sub-area releases new repulsion pheromones, driving the drone to other potential areas for search.
[0069] At planning step t, the subregion G xy The repulsive pheromone in r (x,y,t), can be defined as follows: Among them, s r (x,y,t-1) represents the sub-region G at planning step t-1 xy The repulsive pheromone concentration within the r represents the incremental unit of repulsive pheromone, It represents the repulsive pheromone switching coefficient, which can be adjusted according to the probability of the detected target existing, and is defined as follows: in, Indicates that the repulsive pheromone switch is turned on. Indicates that the repulsive pheromone switch is off, p i (x, y, t-1) represents the sub-region G stored by drone i at planning step t-1. xy The probability of the target existing.
[0070] In some embodiments of the present invention, when the drone cluster is not searching the mission area, it is necessary to initialize the search information graph of each drone. In addition, since the environmental uncertainty and pheromone information can be updated according to the target existence probability (the environmental uncertainty is updated according to equation (31), and the pheromone information is updated according to equations (12) to (15)), the initial setting process of the search information is described below using the target existence probability as an example. The environmental uncertainty and pheromone information can be initialized according to the initial target existence probability.
[0071] Specifically, under local communication conditions, limited by the communication distance, drone clusters usually conduct effective collaborative searches through connected subnets in the initial stage. If a uniform distribution is used to initialize the probability of target existence in the mission area, the drone cluster may fall into a local optimal solution in the early stages of the search, resulting in a waste of search resources. Therefore, in the absence of prior information, the present application is designed based on a rational assumption and can use a two-dimensional Gaussian distribution to initialize the probability of target existence in the mission area (that is, the initial target existence probability map in the search information map stored by the drone in the present application is designed to present a two-dimensional Gaussian distribution). This initialization method can give the drone group a key search area where the target may be present in the initial stage by reasonably setting the center point and variance of the Gaussian distribution (for example, the center point and variance can be set by rapid networking and interaction between the drone and environmental information).
[0072] As an example, the initial modeling process for target existence probability is as follows: At the initial planning step (assuming t = 0), each sub-region G in the task area xy The probability density function of the target existence probability can be expressed as: Among them, [x tar ,y tar ] T Represents the peak position of the two-dimensional Gaussian distribution, which is used to reflect the most likely initial position of the target in the task area. For example, the peak position can be defined as the center position of the task area. 2 Represents the variance of the two-dimensional Gaussian distribution, whose size determines the dispersion of the target distribution, σ 2 When it is larger, it means that the target position uncertainty is higher, σ 2 When it is smaller, it means the uncertainty of the target position is lower, which can be obtained by the following formula: σ 2 =(H / k) 2 ;……………………………………………………………………(17) Where H is the length or width of the task area, and k is the scale factor.
[0073] Considering that this application only needs to characterize the relative probability of the existence of the target in the mission area, the relative value of the probability density can be directly used as the representation of the probability of the target existence. By retaining the relative size relationship of the probability density (for example, the maximum value can be normalized to 1), the possibility distribution of the target in different locations in the mission area can be effectively reflected. Specifically, the area with higher probability density indicates that the possibility of the target existence is greater, while the area with lower probability density indicates that the possibility of the target existence is less. Initialize the target existence probability (see Figure 6 (a) in the original text) and the initial environment is uncertain (see Figure 6 The distribution diagram of (b)) is as follows Figure 6 shown.
[0074] The initial pheromone concentration of each sub-area within the design task area of this application is 0.
[0075] Furthermore, in order to maximize the dispersion of the UAV cluster in the initial stage and thus improve the efficiency of information collection in the mission area, the mission search area can be divided into multiple areas (each divided area contains one or more sub-areas), and the target existence probability and environmental uncertainty of each sub-area are independently initialized. Figure 7 As shown, the task search area can be divided into four areas evenly.
[0076] In some embodiments of the present invention, each drone in the drone cluster applies a strategy in the strategy space at each planning step. During the strategy optimization phase, the strategy applied by each drone may change at different planning steps.
[0077] As an example, in order to realize the collaborative search function of drones, a policy space Θ is set in this application. There are multiple strategies in the policy space Θ, and the strategies adopted by the drones in the cluster at each planning step can be obtained from the policy space (that is, whether it is the strategy before optimization or the strategy after optimization, the strategy applied by the drone is a strategy in the policy space). The policy space can be expressed as: Θ={θ1,θ2,...,θ b ,...,θ B};…………………………………………………………(18) Among them, θ b represents the bth strategy, and B represents the number of strategies in the strategy space (for subsequent optimization strategies, B ≥ 2). The strategies in the strategy space can provide a way to calculate the benefits when the drone performs the search task. For example, in step S120, the expected execution benefits can be calculated using the strategy currently applied by the drone.
[0078] In some embodiments of the present invention, after obtaining updated search information in step S121, the drone is further required to determine the optimal decision path in step S122, including: based on the updated search information graph and the current drone application strategy, selecting an optimal decision path within a rolling planning horizon (hereinafter referred to as the planning horizon) from the mission search area, and calculating the expected execution benefit of the optimal decision path. Step S122 specifically includes: Single-unit pre-decision: In the planning domain, a traversal algorithm is used to select all feasible paths for the current UAV from the mission search area. The path benefits of each feasible path are calculated based on the updated search information graph. Based on the calculated path benefits, a set number of paths are selected from the feasible paths as the pre-decision paths for the current UAV. Swarm Optimal Decision: Based on the current drone application strategy and the path benefits of the pre-decision paths, the expected execution benefits of the pre-decision paths are calculated. The optimal decision path is selected from the pre-decision paths based on the calculated expected execution benefits. A feasible path is a path that drones can travel when performing a collaborative search mission, and the optimal decision path is the path among the pre-decision paths that maximizes the expected execution benefits of the drone swarm.
[0079] As an example, the planning time domain refers to the time domain consisting of the current planning step (i.e., the planning step at which the optimal decision path is currently selected) and one or more subsequent consecutive planning step times, i.e., a rolling planning time domain is formed with the current planning step as the starting point and a planning step selected from all planning step times after the current planning step as the end point. In addition, in this application, a pruning strategy can be used to select a pre-decision path from all feasible paths based on path benefits (for example, all feasible paths can be sorted in descending order according to path benefits, and a pre-decision path can be selected from all feasible paths in descending order of path benefits).
[0080] More specifically, the path benefit of a feasible path can be determined based on the search value benefit and coordination benefit of the feasible path. Since the optimization process of each drone in a distributed architecture aims to maximize the benefits obtained by performing the search task on a single machine and improve decision-making efficiency through fully autonomous decision-making, this application can model the calculation method of the path benefit of the trajectory, and the constructed model can be used to represent the search planning process of each drone. The path benefit calculation model can be expressed using the following formula: Among them, [t s ,t s +T s ] represents the planning time domain of the UAV cluster to perform the search mission, T s is the length of the planning time domain, R i (t) is the decision variable, which means that UAV i is in the planning step t (at this time t∈[t s ,t s +T s ]) feasible path, J Si (t,R i (t)) represents the time when UAV i makes decision R at planning step t. i (t) Single machine path benefit, J Vi (t, R i (t)) and J Ci (t, R i (t)) represent the time when UAV i makes decision R i Search value benefits and coordination benefits under (t).
[0081] As an example, search for the value of gain J Vi (t, R i(t)) is a mathematical representation of the most valuable search area at the moment. It intuitively reflects the search priority of the drone cluster in the mission area at planning step t, and can be determined based on the pheromone concentration in the pheromone map of drone i. Specifically, due to the existence of sensor detection probability and false alarm probability, the drone needs to revisit the target multiple times during the search process to confirm the existence of the target and capture it. When the drone detects the presence of a target in a sub-area, the sub-area will release attractive pheromones to guide the drone to revisit the area, indicating that the area has a higher search value; conversely, if the target has been hit, the potential benefits of continuing to search the area are significantly reduced. At this time, the sub-area will release repulsive pheromones to drive the drone to search other potential areas. Therefore, the search value benefit can be defined as: Where (x, y)∈G i (R i (t)) represents the path R i (t) contains sub-region G xy If part or all of the path Ri of UAV i ( The projection of t) is located in the sub-region G xy If the path R i (t) contains sub-region G xy In this application, if the drone can detect the sub-area G xy If the information in the sub-area G is xy , which can also be considered as the drone’s path containing the sub-area G xy .
[0082] Moreover, the coordination payoff J Ci (t,R i The design of (t)) aims to reasonably balance local search and global search. On the one hand, it effectively avoids falling into the local optimum problem caused by relying solely on the target existence probability, and on the other hand, it limits excessive search in low-probability areas. Therefore, the target existence probability and uncertainty in the TPM and EUM of drone i are integrated to design the coordination benefit, which can be defined as follows: According to formula (20) and formula (21), the coordination benefit of the feasible path (or pre-decision path) is the sum of the target existence probability and environmental uncertainty (the information in the target existence probability map and environmental uncertainty map updated in step S121) in all task sub-areas passed by the feasible path; the search value benefit of the feasible path (or pre-decision path) is the sum of the pheromone concentration differences (the difference between the attraction pheromone concentration and the repulsion pheromone concentration) in all task sub-areas passed by the feasible path.
[0083] Furthermore, the expected execution benefit of the pre-decision path can be determined based on the current UAV application strategy, the search value benefit of the pre-decision path, and the coordination benefit of the pre-decision path. Similar to the coordination benefit and search value benefit of the available trajectory, the coordination benefit of the pre-decision path is the sum of the target existence probability and environmental uncertainty in all task sub-areas passed by the pre-decision path; the search value benefit of the pre-decision path is the sum of the pheromone concentration differences in all task sub-areas passed by the pre-decision path; where the pheromone concentration difference is the difference between the attraction pheromone concentration and the repulsion pheromone concentration. In other words, the path benefit of the pre-decision path can also be calculated using Equations (20) and (21).
[0084] Since the strategy applied by the drone is used to provide a calculation strategy for the drone's benefits, the calculation formula for the expected execution benefit of the pre-decision path can be expressed as: Among them, θ b,i (t) represents the strategy θ applied by UAV i at planning step t b , you can also use represents θ b,i (t), At the planning step t, UAV i is in strategy θ b,i (t) Execute the track R i (t)(At this time R i (t) can be expressed as the expected execution benefit of the pre-decision path).
[0085] In some embodiments of the present invention, the present application designs an update of the current drone application strategy based on evolutionary game theory during the strategy optimization phase. However, this process requires the use of the corresponding relationship between each strategy and the expected benefit in the strategy space (hereinafter simply referred to as the strategy-expected benefit relationship or the corresponding relationship) determined during the prediction phase (the present application performs Monte Carlo prediction operations during the prediction phase, so the prediction phase can also be referred to as the Monte Carlo prediction phase). The corresponding relationship between the strategy and the expected benefit is used to indicate the expected benefit obtained by adopting each strategy in the strategy space before the current planning step.
[0086] More specifically, the strategy-expected benefit relationship stored by the drone in this application can be set automatically based on empirical values, or it can be determined by performing Monte Carlo prediction operations through prediction experiments. Because the accuracy of the strategy-expected benefit relationship affects the subsequent collaborative search steps of the drones, this application specifically defines the steps for determining the strategy-expected benefit relationship through Monte Carlo prediction operations in the prediction phase: For each drone in the drone cluster, at the planning step, the strategy-expected benefit relationship can be obtained by executing steps S210 to S250.
[0087] As an example, before executing steps S210 to S250 in the prediction phase, some parameters need to be initialized and set, including: the maximum communication distance C between drones; r , detection probability P D , false alarm probability P F , probability modifier Pheromone increment unit Δs a and Δs r , pheromone volatility ρ, set number (that is, the number of pre-decision paths retained in the feasible path), trust factor λ, decay factor β, set sliding window length Z, set target existence probability threshold p t , planning time domain [t s , t s +T s ]、Execution time domain[t s , t s +T e ], the initial search information graph and the set task duration [t0, T 预测 ].
[0088] When performing Monte Carlo prediction operations, the strategy applied by each drone is fixed (i.e., no strategy learning and optimization is performed during the prediction phase). That is, each drone can be randomly assigned a strategy in the strategy space, and since drones make completely autonomous decisions under a distributed architecture, the strategy-expected benefit relationship stored by each drone can be different.
[0089] In some embodiments of the present invention, the Monte Carlo prediction operation process for each drone in the drone cluster is as follows: At each planning step in the forecast phase, the following steps are performed: Step S210, UAV networking: Based on the maximum communication distance between UAVs, the current UAV communicates with other UAVs in the cluster to form a connected subnet.
[0090] In step S220, the UAV integrates information and determines the optimal decision path: the current UAV receives the search information graph stored by other UAVs in the connected subnet to which it belongs, and uses the received search information graph to update its own stored search information graph; based on the updated search information graph and the strategy applied by the current UAV, an optimal decision path within the planning time domain is selected from the task search area, and the expected execution benefit of the optimal decision path (the optimal decision path is selected based on the principle of maximizing benefit) is calculated.
[0091] In the execution domain, perform the following steps: Step S230, the UAV executes the task and updates the search information: the current UAV executes the search task according to the optimal decision path selected in step S220, captures the target and outputs the target position when the target search conditions are met; and updates its currently stored search information map according to the information detected during the execution of the search task (see the following formulas (29) and (30) for details); wherein, the condition for issuing the capture command can be q i (x, y, t) ≥ p t (UAV i believes that the probability q of the target existing in a certain grid based on the detected information at the planning time t i (x, y, t) is greater than the set target existence probability threshold p t ).
[0092] The planning time domain in step S220 refers to the current planning step in the prediction phase, and refers to the current planning step and one or more consecutive planning step moments thereafter; the execution time domain in step S230 also refers to the current planning step in the prediction phase, and refers to the current planning step and one or more consecutive planning step moments thereafter; and the execution time domain length is less than or equal to the planning time domain length.
[0093] In the prediction stage, after the steps in the execution time domain are completed, if the prediction stop condition is not met, jump to the next planning step and repeat the above steps S210 to S230; wherein, the prediction stop condition can be that the current prediction planning step reaches the set number of prediction planning steps.
[0094] Step S240: The drone calculates the expected return: After the prediction stop condition is reached, the average of the actual returns obtained by the current drone applying the strategy to perform the search task at all planning steps in the prediction phase is calculated and used as the expected return corresponding to the strategy currently applied by the drone, thereby obtaining the corresponding relationship between each strategy stored in the current drone and the expected return. That is, the average of the actual returns of the drone at all planning steps in the Monte Carlo prediction operation is taken to obtain the preset strategy θ b (i.e., the fixed strategy to which the drone is assigned).
[0095] In addition to calculating the expected return after the prediction stop condition is reached, other existing technologies can also be used to calculate the expected return in this application, for example, the incremental mean method can be used for calculation, but the present invention is not limited thereto.
[0096] Specifically, during the strategy evaluation process, due to the inherent differences in returns between different strategies, the size of the returns cannot directly reflect the quality of the strategy. This application design uses relative returns to indicate the quality of the strategy. In the prediction stage, the expected return of the strategy currently applied by drone i can be estimated using the following formula: Among them, [t0, T 预测 ] is the task domain for performing Monte Carlo prediction operations. Indicates that at the prediction planning step t, UAV i is in strategy θ b The actual income under is the strategy θ for drone i b Under the Monte Carlo operation, the expected return is finally determined, that is, in the strategy optimization stage, the initial strategy-expected return relationship θ b,i and The correspondence between them.
[0097] Furthermore, if in the Monte Carlo prediction stage, each drone is randomly assigned only one strategy, although each drone stores the corresponding relationship between all strategies and expected returns in the strategy space, each drone only uses formula (23) to calculate the expected return of the assigned strategy, and the expected returns of other strategies can be set to 0 by default. Moreover, the steps of designing the Monte Carlo operation in this application are to estimate the expected returns of multiple strategies in the strategy space. Although strategy learning is not performed in the prediction stage, if all drones in the drone cluster adopt the same strategy, it is impossible to learn and optimize the strategy in the subsequent strategy optimization stage. Therefore, although the strategies applied by each drone are randomly assigned in the prediction stage, it is required that there are at least two strategies in the drone cluster so that the cluster can carry out the strategy evolution game process.
[0098] As an example, the prediction phase precedes the strategy optimization phase. If the task time domain of the UAV cluster used to perform the collaborative search task in the strategy optimization phase includes more planning steps, the strategy-expected benefit relationship can be determined first according to the actual situation of the task area when the search task is actually performed, and then the strategy optimization can be further performed. That is, in the overall task time domain [0, T r ] in [0, T l ] performs Monte Carlo prediction operations for collaborative search, in [T l , T r ]Perform collaborative search using strategy optimization based on evolutionary game theory.
[0099] For distributed systems, the premise of a game is the ability to communicate. Therefore, the participants in the game are limited to drones that form the same connected subnet at each planning step. That is, in the method proposed in this application, the players are all drones in the same connected subnet at each game. Therefore, each connected subnet can also be considered a game subnet. During the game, the information shared by each drone is true and accurate. From the perspective of each game subnet, all the playing drones within the subnet are in a cooperative relationship. The collaborative search method proposed in this application is based on evolutionary game theory. Because the game and decision-making are carried out at the same planning step, each drone cannot receive the strategies of other drones in a timely manner when executing the decision using the currently applied strategy in the subsequent step S160. Therefore, the strategy optimization stage of this application is designed to consider the strategies of other drones and update its own currently applied strategy during the game process (mainly including steps S130-S150). During the decision-making process (including step S160), only the updated strategy of the previous planning step is executed.
[0100] Furthermore, when executing strategy optimization operations in steps S130-S150 during the current planning step, drones within the same game subnet can engage in game play based on the expected payoffs obtained during the previous planning step, thereby achieving strategy optimization. To enable subsequent strategy learning and optimization, each connected subnet should contain multiple drones with different current strategies (for example, a connected subnet can contain drones with different strategies as well as drones with the same strategy).
[0101] Step S130, the drone updates the correspondence between the strategy and the expected benefit: based on the expected execution benefit, the expected benefit corresponding to the strategy of the current drone application is updated, and the current drone receives the strategies and updated expected benefits of other drone applications in the connected subnet to which it belongs, thereby updating the correspondence according to the set expected benefit update rules.
[0102] In a distributed architecture, all drones in the drone cluster store the strategy-expected benefit relationship and can maintain the expected benefit corresponding to each strategy in the strategy space through their own game and decision-making process. That is, for each strategy in the strategy space, each drone maintains the expected benefit corresponding to all strategies in the strategy space. For drone i, the currently applied strategy θ b The expected return at planning step t can be expressed as or
[0103] More specifically, the strategy θ applied by drone i at planning step t b,i or Based on the expected execution benefit, the expected benefit corresponding to the strategy currently applied by drone i can be updated using the following formula: in, Indicates that at the end of planning step t-1, UAV i is in strategy θ b The expected return under (can be determined from the corresponding relationship before the update), Indicates that at planning step t, drone i is in strategy θ b That is, the expected return corresponding to the strategy currently applied by the UAV is updated based on the expected return, including: for the strategy currently applied by the UAV at the current planning step, taking the average of the expected return corresponding to the strategy and the expected return in the corresponding relationship to obtain the updated expected return.
[0104] As an example, at planning step t, if the strategy currently applied by drone i is θ b , then the strategy applied by UAV i can be calculated based on formula (24) as θ b The corresponding expected returns for other strategies (such as strategy θ a ), the expected execution benefits of drone i under these strategies can be defaulted to 0. That is, this application does not use formula (24) to update the expected benefits corresponding to other strategies that are not currently adopted, but uses the following formula (25) to update them.
[0105] In some embodiments of the present invention, during the strategy game, each drone will update the expected payoff corresponding to its own strategy at the current planning step, and will also update the expected payoffs corresponding to other strategies in the corresponding relationship. That is, each drone will update the stored strategy-expected payoff relationship at each planning step. For each strategy in the corresponding relationship, the expected payoff update rules (i.e., the rules for updating the strategy-expected payoff relationship) include: If the expected return corresponding to the strategy in the corresponding relationship is not 0, then the updated expected return corresponding to the strategy is the original expected return in the corresponding relationship; if the expected return corresponding to the strategy in the corresponding relationship is 0, and the strategy is not the strategy currently applied by any drone in the connected subnet to which the current drone belongs, then the updated expected return corresponding to the strategy defaults to 0; if the expected return corresponding to the strategy in the corresponding relationship is 0, and the strategy is the strategy currently applied by any other drone in the connected subnet to which the current drone belongs, then the updated expected return corresponding to the strategy is the average of the expected returns updated by other drones in the connected subnet to which the drone belongs that adopt this strategy.
[0106] Specifically, for the strategy θ applied to drone i b,i Other strategies (such as the strategy θ in the corresponding relationship stored by drone i) a,i), at planning step t, drone i can update its own stored strategy-expected benefit relationship based on the updated expected benefits of other drones in the connected subnet. a,i For example, at the planning step t, the strategy θ in the corresponding relationship stored by drone i a,i The update formula of the expected return is as follows: Among them, stg j (t) = θ a The strategy applied by UAV j at planning step t is θ a , Indicates that UAV j and UAV i are in the same connected subnet (and the number of UAVs in the subnet is ), I represents the connected subnet where UAV i is located, and the strategy applied is θ a The number of drones.
[0107] Step S140, the drone calculates the strategy evaluation value: based on the expected execution benefit and the updated expected benefit corresponding to the strategy currently applied by the drone, the relative evaluation value of the strategy currently applied by the drone is calculated, and then the weighted relative evaluation value of the strategy applied by the drone is calculated according to the relative evaluation value at the current planning step and the relative evaluation value at the specific planning step.
[0108] During the strategy evaluation process, since there are inherent differences in the size of benefits when calculating different strategies, it is not reasonable to directly compare the size of benefits. Therefore, this application uses relative expected benefits to express it. However, in order to remove the impact of inherent differences, this application further uses strategy evaluation values to evaluate the size of the benefits of the strategies applied by the drone.
[0109] Specifically, suppose that at planning step t, UAV i applies the strategy θ b,i Make decisions, relative evaluation of strategies The following formula can be used for calculation: Furthermore, when the policy θ b When the expected return of strategy θ is low, a single high return will lead to a higher evaluation value; when strategy θ b When the expected return is high, a low single return will penalize a potentially high-quality strategy. If the single return evaluation value is directly used as the criterion for judging the quality of the strategy, significant random noise will be introduced. In order to reduce the noise while retaining the dynamic characteristics of the strategy game, a secondary calculation of the evaluation value is performed based on the sliding window weighting method. At planning step t, the drones in the connected subnet use the following formula to calculate the currently applied strategy θ b,i The weighted relative evaluation value of: in, Indicates the strategy θ applied by UAV i at planning step t b The weighted relative evaluation value of , β(β∈(0,1)) represents the attenuation factor, Z represents the set sliding window length (greater than the number of steps at a specific planning step), and F represents the strategy applied by UAV i within the set window as θ b The total number of steps at a specific planning moment, stg(T) = θ b Indicates that the current application strategy of UAV i at planning step t is θ b .
[0110] In some embodiments of the present invention, according to formula (27), the weighted relative evaluation value of the current application strategy of the drone is calculated based on the relative evaluation value of the current planning step moment and the relative evaluation value of the specific planning step moment, including: when there is a specific planning step moment, calculating the average value of the relative evaluation value of the specific planning step moment; obtaining the weighted relative evaluation value of the strategy currently applied by the drone by weighted summing the relative evaluation value of the current planning step moment and the calculated average value; when there is no specific planning step moment, taking the relative evaluation value of the strategy currently applied by the drone as the weighted relative evaluation value of the strategy currently applied by the drone; wherein, the specific planning step moment refers to the planning step moment at which the drone adopts the application strategy within the selected time domain (taking the planning step moment with the sliding window length set from the current planning step moment as the starting point, and the selected historical time domain before the current planning step moment).
[0111] Step S150, the UAV updates the current strategy: the current UAV determines the maximum value of the weighted relative evaluation value in the connected subnet to which it belongs, calculates the strategy learning probability corresponding to the current UAV based on the maximum value of the weighted relative evaluation value and the weighted relative evaluation value of the strategy applied by the current UAV, and updates the strategy applied by the current UAV to the strategy corresponding to the maximum value of the weighted relative evaluation value based on the strategy learning probability.
[0112] Traditional game theory requires that all parties in the game are completely rational and play under conditions of complete information. Therefore, if traditional game theory is used, when all drones estimate the optimal strategy for the next planning step, they will tend to choose the strategy that maximizes their own benefits. However, evolutionary game theory assumes that all game participants are "bounded rational individuals", that is, when all drones estimate the optimal strategy for the next planning step, they will tend to choose the strategy that maximizes their own benefits. Therefore, game participants are not necessarily able to choose the strategy that maximizes their own benefits, but based on the evolutionary stable strategy, all game parties tend to a certain stable strategy. It can be further understood that at each iteration, the drones in the game subnet will compare their own strategy and the strategies with the highest benefits of all members in the subnet, and learn the strategy of the player with the highest benefit with a certain probability. For example, the process of strategy learning in this application is strategy replacement, that is, replacing the strategy currently applied by the drone with the strategy corresponding to the maximum weighted relative evaluation value.
[0113] As an example, the strategy learning probability can be calculated using the Fermi function. Assume that the relative strategy evaluation value of the game drone i in the current connected subnet is Among all members in the connected subnet to which drone i belongs, drone j has the largest relative strategy evaluation value, which is expressed as The probability that drone i learns the strategy corresponding to drone j at the current planning step can be calculated by the following formula: Among them, λ is the trust factor, which indicates the drone's acceptance of external strategies. When λ→0, it means that regardless of the benefit, the strategy will be updated completely randomly, that is, the current game drone will randomly choose whether to learn the strategy of the member with the largest benefit. Conversely, if λ→∞, it indicates a deterministic update rule, that is, when the maximum benefit in the subnet is higher than the benefit brought by its own current strategy, the drone will definitely adopt the same strategy at the next moment.
[0114] The above-mentioned process of calculating the probability of strategy learning by using the Fermi function is only an example, and other methods can also be used, and the present invention is not limited thereto.
[0115] After executing steps S110 to S150 at the current planning step, step S160 needs to be continued to be executed within the execution time domain including the current planning step.
[0116] Step S160, the UAV executes the mission and updates the search information: the current UAV executes the search mission according to the optimal decision path, captures the target and outputs the target position when the target search conditions are met, and updates its own stored search information map based on the information detected during the execution of the search mission.
[0117] As an example, the execution time domain and the planning time domain are relative. The execution time domain refers to the current planning step and one or more consecutive planning step times thereafter. In addition, the execution time domain length is less than or equal to the planning time domain length. In this application, the planning time domain and the execution time domain both include the current planning step.
[0118] The target search condition in step S160 of this application can be q i (x, y, t) ≥ p t , that is, UAV i is planning for sub-area G at time t xy The probability of the detected target existing q i (x, y, t) is greater than or equal to the set target existence probability threshold p t The target search conditions in the above steps S230 and S160 are examples, and this application does not specifically limit the target search conditions.
[0119] In addition, for a certain sub-area, if the sub-area is outside the detection range of all drones in the connected subnet, the detection information of the drone can be defaulted to 0.
[0120] During the collaborative search mission, the drone cluster can update its stored search information map based on the information detected by the onboard sensors at the planning step t. That is, at the planning step t, the drone updates the search information map updated in step S120 based on the information detected in step S160. Since the environmental uncertainty and pheromone information can be updated based on the target existence probability, the following assumes that the target existence probability q detected by drone i in step S160 is t. i (x, y, t) as an example to describe the search information update process in step S160 (the search information can also be updated directly based on information such as the target existence probability and environmental uncertainty detected by the drone).
[0121] Specifically, in order to more accurately reflect the dynamic update process of search information, taking the target existence probability as an example, consider the sensor detection probability P D and false alarm probability P F , dynamic update of information can be achieved based on Bayesian Criterion. Assume that for sub-region G xy , the probability of the target detected by UAV i at planning step t is q i (x, y, t), then at planning step time t, the prior updated target existence probability can be updated according to the posterior target existence probability to obtain the target existence probability map stored by UAV i at planning step time t. The specific formula is: Among them, q i(x, y, t) represents the sub-area G detected by drone i at planning step t xy Is there a target (q i (x, y, t) takes the value of 0 or 1), q i (x, y, t-1) = 1 means that UAV i detects sub-area G xy Have a goal, i (x, y, t-1) = 0 means that UAV i detects sub-area G xy No target exists.
[0122] Furthermore, at the planning step t, if the sub-region G xy The probability of the internal target existing p i (x, y, t) is greater than the target existence probability threshold Then it is considered that sub-region G xy There is a target in the memory and a capture command is issued. After the target is captured, in order to more accurately reflect the actual situation of the mission area, the sub-area G xy The probability of the target existing within is modified. The modification mechanism is as follows: in, Indicates the probability correction value, which is an experience value.
[0123] As an example, the detection probability P in equation (29) D and false alarm probability P F It can be customized or determined by actual engineering application. Specifically, during the UAV detection process, the UAV will update its stored search information map in real time based on the detection information obtained by the onboard sensor, so as to make further decisions. However, due to the influence of factors such as the sensor's own performance, environmental occlusion, and resource constraints, there may be cases where the target is missed or falsely reported. It is necessary to set the detection probability P D and false alarm probability P F They represent the probability of the sensor finding the target or making an incorrect judgment within the search range (or detection range), respectively, and are defined as follows: Detection probability P D =P(detection result is target|actual target), P D ∈[0, 1]; False alarm probability P F =P(detection result is target|actually no target), P F ∈[0, 1].
[0124] Furthermore, for sub-region G xy , the environmental uncertainty of UAV i at planning step t can be updated by the information entropy of the target existence probability updated by formula (31), as follows: ηi (x,y,t)=-p i (x,y,t)log2p i (x,y,t)-(1-p i (x,y,t))log2(1-p i (x,y,t));…(31) The pheromone information is updated based on the target existence probability updated by equation (30) and according to equations (12) to (15).
[0125] After the steps in the execution time domain are completed (i.e., after step S160 is executed), if the search stop condition is not met, jump to the next planning step and repeat the above steps S110 to S160; when the search stop condition is met, stop the collaborative search of the drone cluster.
[0126] As an example, the search stop condition can be that the current planning step reaches a set number of search planning steps, the drone cluster captures a set number of targets, or the proportion of drones in the drone cluster using a certain strategy reaches a set ratio. The target search conditions and search stop conditions in this application can be set as needed, and the present invention is not limited to this.
[0127] As an example, in the prediction stage, before executing steps S110 to S160 in the strategy optimization stage, some parameters need to be initialized and set, including: the maximum communication distance C between drones r , detection probability P D , false alarm probability P F , probability modifier Pheromone increment unit Δs a and Δs r , pheromone volatility ρ, set quantity, trust factor λ, decay factor β, set sliding window length Z, set target existence probability threshold (in the strategy optimization stage, the first set target existence probability threshold ), planning time domain [t s , t s +T s ]、Execution time domain[t s , t s +T e ], the initial search information graph, the strategy-expected return relationship determined in the prediction phase, and the task time domain set in the strategy optimization phase [T l , T r ].
[0128] As shown in Table 1 and Figure 8 As shown, the collaborative search method proposed in this application adopts an evolutionary game method to randomly assign multiple pre-designed collaborative strategies to cluster members, and realizes strategy optimization through the evolutionary game process.
[0129] Table 1 Pseudo code of the distributed collaborative search method for UAV swarms based on evolutionary game theory The UAV swarm collaborative search method proposed in this application based on evolutionary game theory in a distributed architecture has the following significant advantages: ① This application organically integrates evolutionary game theory and distributed collaborative search of drone clusters. During the search process, the survival of the fittest of search strategies is continuously achieved through strategy evaluation and game. At the same time, each drone makes completely autonomous decisions, significantly improving the effectiveness of collaborative search. The method proposed in this application can divide the entire collaborative search process into a Monte Carlo prediction stage and a strategy optimization stage based on evolutionary games. In these two stages, each drone in the cluster performs collaborative search tasks through completely autonomous decision-making, maintains its own search information graph and strategy-expected benefit relationship, and integrates a path benefit calculation model and pruning strategy to determine the optimal search trajectory for each drone. In the Monte Carlo prediction stage, each drone in the cluster is assigned an initial strategy and conducts collaborative search according to the initial strategy, thereby estimating the expected benefit of the initial strategy and quantifying the search effectiveness of the initial strategy. After initially determining the strategy-expected reward relationship, the strategy optimization phase begins. As the search progresses, the expected reward of each drone's current corresponding strategy is continuously updated according to certain rules. Based on the expected reward, the current corresponding strategies of each drone in the subnet are evaluated. Each drone optimizes and updates its strategy based on the strategy evaluation value using an evolutionary game mechanism. The principle of strategy optimization and update is that high-performing strategies will be adopted by more drones, while poor-performing strategies will be gradually eliminated. Through this dynamic evolutionary process, the drone cluster can adaptively select the optimal search strategy within the subnet, thereby achieving efficient cluster collaboration under a distributed architecture and significantly improving the efficiency and effectiveness of collaborative search.
[0130] This method provides a feasible solution to the problem of collaborative search for moving targets by drone clusters under a distributed architecture in complex scenarios, while avoiding the tedious parameter adjustment process and has important practical application value.
[0131] ② Based on the concept of evolutionary game theory, this application designs a new strategy evaluation mechanism and expected payoff update rules, thereby constructing a new evolutionary game mechanism for drone swarms in a distributed architecture. Specifically, connected subnetworks are set up based on communication distance, and drones within the connected subnetworks perform information fusion. Search strategies are evaluated and played within the connected subnetworks based on the optimal decision path and the drones' current corresponding strategies. This enables autonomous evolutionary optimization of drone search strategies, ultimately allowing each drone to converge to its own optimal strategy.
[0132] ③ This application uses Gaussian distribution to initialize the search information graph. In the absence of prior information, based on the rationality assumption, a two-dimensional Gaussian distribution is used to initialize the target existence probability and environmental uncertainty in the mission area. This initialization method can initially give the drone a key search area where the target may exist by reasonably setting the center point and variance of the Gaussian distribution. Moreover, in order to maximize the dispersion of the drone cluster in the initial stage and improve the efficiency of environmental information collection, this application can also divide the mission area into multiple areas and independently initialize the search information of each divided area.
[0133] Corresponding to the above method, the present invention also provides a collaborative search system for a distributed drone cluster, which includes a computer device, wherein the computer device includes a processor and a memory, wherein the memory stores a computer program / instructions, and the processor is used to execute the computer program / instructions stored in the memory. When the computer program / instructions are executed by the processor, the system implements the steps of the method described above.
[0134] An embodiment of the present invention further provides a computer-readable storage medium having a computer program / instruction stored thereon, which, when executed by a processor, implements the steps of the aforementioned edge computing server deployment method. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a floppy disk, a hard disk, a removable storage disk, a CD-ROM, or any other form of storage medium known in the art.
[0135] It should be understood by those skilled in the art that the various exemplary components, systems and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software or a combination of the two. Whether it is specifically performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present invention are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link via a data signal carried in a carrier.
[0136] It should be understood that the present invention is not limited to the specific configurations and processes described above and illustrated in the figures. For the sake of brevity, a detailed description of known methods is omitted. In the above embodiments, several specific steps are described and illustrated as examples. However, the method of the present invention is not limited to the specific steps described and illustrated. Those skilled in the art may make various changes, modifications, and additions, or change the order of the steps after understanding the spirit of the present invention.
[0137] In the present invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or replace features of other embodiments.
[0138] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations to the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A collaborative search method for distributed drone swarms based on evolutionary game theory, characterized in that: For each drone in the drone cluster, the method includes: For the strategy optimization phase, during the planning step, the following steps are performed: Drone Networking: Based on the maximum communication distance between drones, the current drone communicates with other drones in the cluster to form a connected subnet. The current drone applies a strategy in the strategy space, and the connected subnet includes multiple drones applying different strategies. The current drone also stores the corresponding relationship between each strategy and expected reward in the strategy space, as well as the search information graph. The drone integrates information and determines the optimal decision path: The current drone receives the search information graph stored by other drones in its connected subnet and uses the received search information graph to update its own stored search information graph. Based on the updated search information graph and the current drone's applied strategy, it selects an optimal decision path within the rolling planning time domain from the mission search area and calculates the expected execution benefit of the optimal decision path. The drone updates the correspondence between each strategy and the expected benefit: based on the expected execution benefit, the expected benefit corresponding to the strategy applied by the current drone is updated, and the current drone receives the strategies and updated expected benefits applied by other drones in the connected subnet to which it belongs, thereby updating the stored correspondence according to the set expected benefit update rule; The drone calculates a strategy evaluation value: based on the expected execution benefit and the updated expected benefit corresponding to the current drone application strategy, the drone calculates a relative evaluation value of the strategy currently applied, thereby calculating a weighted relative evaluation value of the strategy currently applied by the drone based on the relative evaluation value at the current planning step and the relative evaluation value at a specific planning step; wherein the specific planning step refers to the planning step at which the strategy currently applied by the drone is the strategy currently applied by the drone within the period from the historical planning step with a set sliding window length from the current planning step to the current planning step; The drone updates the current strategy: the current drone determines the maximum value of the weighted relative evaluation value in the connected subnet to which it belongs, calculates the strategy learning probability corresponding to the current drone based on the weighted relative evaluation value of the strategy applied by the current drone and the maximum value of the weighted relative evaluation value, and updates the strategy applied by the current drone to the strategy corresponding to the maximum value of the weighted relative evaluation value based on the strategy learning probability; During the strategy optimization phase, the following steps are performed in the execution domain: The UAV executes the mission and updates the search information: During the execution time domain, the current UAV executes the search mission according to the optimal decision path, captures the target and outputs the target position when the target search conditions are met, and updates its stored search information graph based on the detection information during the execution of the search mission; wherein the execution time domain and the planning time domain are both the planning step time when the optimal decision path is selected and one or more consecutive planning step times thereafter, and the planning time domain is greater than or equal to the execution time domain; After the steps in the execution time domain are completed, if the search stop condition is not met, it jumps to the next planning step and repeats all the steps in the above planning step and execution time domain to finally complete the collaborative search of the UAV cluster.
2. The method according to claim 1, characterized in that For each strategy in the corresponding relationship, the set expected return update rules include: If the expected return of the strategy in the corresponding relationship is not 0, then the updated expected return of the strategy is the original expected return in the corresponding relationship; If the expected return of the strategy in the corresponding relationship is 0, and the strategy is not the strategy currently applied by any drone in the connected subnet to which the current drone belongs, the updated expected return corresponding to the strategy is the default value; If the expected return of the strategy in the corresponding relationship is 0, and the strategy is the strategy currently applied by other drones in the connected subnet to which the current drone belongs, then the updated expected return corresponding to the strategy is the average of the expected returns obtained by other drones in the connected subnet to which the drone belongs that adopt this strategy.
3. The method according to claim 1, characterized in that The calculating, based on the relative evaluation value at the current planning step and the relative evaluation value at the specific planning step, a weighted relative evaluation value of the strategy applied by the current UAV, comprises: when a specific planning step exists, calculating an average value of the relative evaluation values at the specific planning step; and obtaining the weighted relative evaluation value of the strategy applied by the current UAV by performing a weighted summation on the relative evaluation value at the current planning step and the average value; In the absence of a specific planning step, the relative evaluation value of the strategy currently applied by the UAV is used as the weighted relative evaluation value of the strategy currently applied by the UAV.
4. The method according to claim 1, wherein The corresponding relationship between each strategy and expected benefit in the strategy space stored by the drone is obtained through Monte Carlo prediction operation; For each drone in the drone cluster, the Monte Carlo prediction operation includes the following steps: at the planning step, executing drone networking, drone information fusion, and determining the optimal decision path; in the execution time domain, executing drone mission execution and updating search information; and drone calculation of expected returns; Among them, the drone calculates the expected benefit by performing the following operations: after reaching the prediction stop condition, calculate the average value of the actual benefits obtained by the current drone applying the strategy to perform the search task at all planning steps of the Monte Carlo prediction operation, and use it as the expected benefit corresponding to the strategy applied by the current drone, and then obtain the corresponding relationship between the various strategies stored in the current drone and the expected benefits.
5. The method according to claim 4, characterized in that The search information graph stored in the drone is determined based on the target existence probability, environmental uncertainty, and pheromone information detected by the drone when performing a search mission; and a two-dimensional Gaussian distribution is used to initialize and model the target existence probability in the search information graph stored in the drone; The search stop condition is that the planning step reaches the set number of search planning steps or the number of targets captured by the drone cluster reaches the set number of targets; the prediction stop condition is that the planning step reaches the set number of prediction planning steps; and The strategy learning probability is calculated using the Fermi function.
6. The method according to claim 5, characterized in that The step of selecting an optimal decision path within the rolling planning time domain from the mission search area based on the updated search information graph and the current drone application strategy, and calculating the expected execution benefit of the optimal decision path, includes: In the planning time domain, all feasible paths of the current UAV are calculated from the mission search area using a traversal algorithm, the path benefits of each feasible path are calculated based on the updated search information graph, and a set number of paths are selected from all feasible paths as pre-decision paths of the UAV based on the calculated path benefits; wherein the path benefits of the feasible paths are determined based on the search value benefits and coordination benefits of the feasible paths; Based on the current drone application strategy and the path benefits of the pre-decision paths, the expected execution benefits of the pre-decision paths are calculated, and the optimal decision path is selected from the pre-decision paths based on the calculated expected execution benefits.
7. The method according to claim 5, characterized in that The expected execution benefit of the pre-decision path is determined based on the current drone application strategy, the search value benefit of the pre-decision path, and the coordination benefit of the pre-decision path; Among them, the search value benefit of the pre-decision path is the sum of the pheromone concentration differences in all task sub-areas passed by the pre-decision path; the coordination benefit of the pre-decision path is the sum of the target existence probability and environmental uncertainty in all task sub-areas passed by the pre-decision path; among them, the pheromone concentration difference is the difference between the attraction pheromone concentration and the repulsion pheromone concentration.
8. The method according to claim 5, characterized in that The current drone receives the search information graph stored by other drones in the connected subnet, and updates its own stored search information graph using the received search information graph, including: For the target existence probability map and environmental uncertainty map in the search information map, the UAV receives the target existence probability map and environmental uncertainty map stored by other UAVs in its connected subnet; For each task sub-area within the task search area, when the sub-area is within the search range of the current UAV, the target existence probability map and environmental uncertainty map stored by the current UAV are used as the target existence probability map and environmental uncertainty map in the search information map updated by the current UAV respectively; For each task sub-area in the task search area, when the sub-area is outside the search range of the drone, the average value of the target existence probability of the sub-area in the target existence probability map stored by other drones in the connected subnet to which the current drone belongs is calculated, and the square root of the product of the environmental uncertainty of the sub-area in the environmental uncertainty map stored by other drones in the connected subnet to which the current drone belongs is calculated. The calculation results are respectively used as the target existence probability map and environmental uncertainty map in the updated search information map of the current drone.
9. A collaborative search system for distributed drone swarms based on evolutionary game theory, comprising a processor, a memory, and a computer program / instruction stored in the memory, characterized in that: The processor is configured to execute the computer program / instructions. When the computer program / instructions are executed, the system implements the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Dynamic path planning method and system in unmanned aerial vehicle group collaborative target search
CN113848987A
Heterogeneous unmanned aerial vehicle cluster collaborative search optimization method and system based on multi-situation map fusion
CN116301043A
Game behavior dynamic evolution and strategy deduction optimization method and system in space field
CN119539090A
Highly adaptive unmanned aerial vehicle cluster collaborative target search method based on deep reinforcement learning
CN119882820A