An asynchronous coordination method in a distributed space-air-ground multi-agent system
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SCHOOL OF SOFTWARE ZHEJIANG UNIV (NINGBO) MANAGEMENT CENT (NINGBO SOFTWARE EDUCATION CENT)
- Filing Date
- 2025-07-23
- Publication Date
- 2026-06-02
Smart Images

Figure CN120952386B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of artificial intelligence, distributed machine learning and edge computing, and is applied to the task collaboration of heterogeneous multi-agent systems in the air-space-ground context of disaster emergency scenarios. In particular, it relates to an asynchronous collaboration method in a distributed air-space-ground multi-agent system. Background Technology
[0002] In responding to large-scale disasters such as earthquakes and floods, it is often necessary to mobilize multiple types of intelligent agents to work together to complete tasks such as search and rescue and material distribution. These intelligent agents include mobile and flexible drones (aerial intelligent agents), satellites or high-altitude platforms with wide coverage (space-based intelligent agents), and ground robots or rescue personnel capable of penetrating complex terrain (ground-based intelligent agents).
[0003] However, disaster sites are complex and vast, and traditional methods often rely on centralized control or synchronous coordination mechanisms for task allocation and scheduling. For example, a unified command center plans the paths and task assignments for all agents. This centralized approach has significant limitations in disaster emergency scenarios. For instance, communication networks may become unstable due to disaster damage, easily leading to single points of failure and communication delays; synchronous coordination requires all agents to wait for unified instructions, resulting in slow response times and difficulty in meeting the needs of emergency tasks.
[0004] Furthermore, the topology and resource status of agents at disaster sites frequently change (e.g., the addition of temporary reinforcement drones or equipment failures), and existing technologies struggle to adapt quickly to these dynamic changes. Therefore, there is an urgent need for a method that can efficiently coordinate heterogeneous agents in unreliable network environments, fully leveraging the strengths of various agents while dynamically adapting to the addition or removal of agents and changes in tasks. Summary of the Invention
[0005] This invention aims to propose an asynchronous collaborative method for distributed air-space-ground multi-agent systems to address the problems of centralized control dependence, global synchronization delay, and lack of dynamic adaptability in existing technologies during disaster emergency scenarios. By dividing the disaster site into multiple autonomous sub-regions and introducing asynchronous game theory for task allocation, this invention enables efficient multi-agent collaboration under communication-constrained conditions. This method not only ensures the timeliness and comprehensiveness of task coverage but also possesses adaptive capabilities to the dynamic addition or removal of agents, thereby significantly improving the overall efficiency of disaster emergency response.
[0006] To achieve the above objectives, the present invention provides an asynchronous cooperation method in a distributed air-space-ground multi-agent system, comprising the following steps:
[0007] Step 1: Divide the disaster site into several sub-regions. Use the weighted Voronoi algorithm or manually divide the area to generate a set of sub-regions. Assign a regional command node to each sub-region to manage it.
[0008] Step 2: Introduce a boundary consistency handling mechanism to avoid task duplication or omission, and ensure the fairness and rationality of task allocation.
[0009] Step 3: Each regional command node generates an initial task list within its respective region, models the task allocation of the agent set as a potential game, and calculates the utility function; the agents asynchronously execute the best response, select the optimal task set through asynchronous policy updates, and execute their respective assigned tasks after the asynchronous policy updates converge.
[0010] Step 4: Design a regional state detection mechanism to monitor the state of agents within the region in real time. When the game state changes, the agents within the region update their strategies asynchronously until a new stable state is reached. This process is repeated until all tasks within the region have been completed.
[0011] Compared with the prior art, the present invention has the following beneficial effects:
[0012] 1. Distributed asynchronous collaboration enhances reliability and real-time performance: By employing a mechanism of regional autonomy and asynchronous policy updates, the system completely eliminates reliance on a central control node and global synchronization. Even when communication networks at disaster sites are damaged or unstable, each region can still independently and efficiently advance its tasks, avoiding task stagnation or system crashes due to single points of failure. Asynchronous game theory algorithms enable agents to dynamically adjust decisions based on the latest local information, without waiting for global synchronization instructions, thus significantly improving system response speed and real-time performance. Furthermore, this distributed architecture reduces reliance on high-bandwidth communication, decreases the possibility of network congestion, and further enhances system reliability.
[0013] 2. Boundary Coordination Ensures Complete Task Coverage: A boundary consistency coordination mechanism is introduced for cross-regional tasks. Task assignment is determined by calculating a task execution efficiency score, avoiding duplicate or omitted tasks. Boundary tasks are allocated to the optimal sub-region or completed collaboratively by multiple sub-regions, ensuring the integrity of the task sets in each region. Furthermore, when the task load in a region exceeds its processing capacity, neighboring regions can appropriately share the tasks through the boundary negotiation mechanism, achieving cross-regional load balancing. This mechanism not only improves the fairness and efficiency of task allocation but also enhances the collaborative capabilities between multiple regions, providing strong support for responding to complex disaster scenarios.
[0014] 3. Enhanced Flexibility and Adaptability of Task Allocation: By using a weighted Voronoi algorithm for autonomous region partitioning and combining it with a real-time region status detection mechanism, the system can dynamically adjust region partitioning and task allocation strategies based on the task distribution and agent status at the disaster site. This flexibility allows the system to quickly adapt to complex and changing environments, such as changes in task priority, differences in agent capabilities, and geographical limitations. Simultaneously, the utility function design fully considers factors such as task value, duplicate coverage penalties, and execution costs, ensuring the rationality and efficiency of task allocation and further enhancing the system's adaptability.
[0015] 4. Enhanced System Anti-interference Capability: In disaster emergency scenarios, communication networks may be severely interfered with or even completely interrupted. This invention constructs a highly anti-interference distributed system through autonomous region division and asynchronous strategy updates. Even if some regions or agents lose contact due to communication interruption or equipment failure, the remaining regions and agents can continue to operate independently, ensuring overall task coverage and execution efficiency. This anti-interference capability significantly improves the system's stability and survivability in extreme environments.
[0016] 5. Facilitating Efficient Collaboration of Heterogeneous Intelligent Agents Across Platforms: This invention is applicable to various types of intelligent agents, including drones, ground robots, high-altitude platforms, and satellites, fully leveraging the strengths of each type. For example, drones are suitable for rapid maneuver search, ground robots are suitable for rescue operations in complex terrains, and satellites are suitable for large-scale monitoring. Through a unified asynchronous collaborative framework, different types of intelligent agents can collaborate efficiently in their respective areas of expertise, complementing each other's strengths and maximizing overall task execution efficiency. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the asynchronous collaborative architecture of air-space-ground multi-agent system according to an embodiment of the present invention. It shows a distributed training system composed of heterogeneous nodes such as regional command nodes, ground agents, and air agents, as well as a central coordination unit. The interaction relationship between the modules is shown in the figure.
[0018] Figure 2 This is a flowchart illustrating the asynchronous collaboration process of multiple intelligent agents in space, air, and ground according to an embodiment of the present invention.
[0019] Figure 3 This is a schematic diagram of an asynchronous collaborative architecture of multiple intelligent agents in a single region according to an embodiment of the present invention, which illustrates the interaction relationship between heterogeneous nodes and tasks in the region.
[0020] Figure 4 This is a schematic diagram of the module interaction relationship in an embodiment of the present invention. Detailed Implementation
[0021] To better understand the present invention, the invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be emphasized that the following embodiments are intended to illustrate the technical principles and preferred implementations of the present invention, and are not intended to limit the scope of protection of the present invention. Based on the ideas of the present invention, those skilled in the art can make various modifications and substitutions in specific applications, and all such modifications and substitutions, as long as they do not deviate from the principles of the present invention, should fall within the scope of protection of the present invention.
[0022] This application first provides an asynchronous cooperation method in a distributed air-space-ground multi-agent system, including the following steps:
[0023] Step 1: As Figure 1 As shown, to overcome the limitations of centralized control, this embodiment divides the disaster site into several autonomous sub-regions, each managed independently by a regional command node. During the region division process, considering both task importance and geographical location information, a weighted Voronoi algorithm is used to generate a set of sub-regions, ensuring a balanced task load across all sub-regions. By calculating task weights and geographical coordinates, a set of sub-regions satisfying the constraints is generated.
[0024] For special scenarios (such as complex terrain or human-defined priority requirements), manual intervention is permitted for task partitioning to adapt to complex geographical environments and task distribution characteristics. Each sub-region is configured with a regional command node, responsible for managing task information and agent states within its region. This distributed architecture effectively reduces global communication dependencies, avoids the risk of single points of failure, and lays the foundation for subsequent task allocation and dynamic adjustments. See [link to relevant documentation]. Figure 3 .
[0025] Step 2: Since adjacent sub-regions may have overlapping tasks or ambiguous boundaries, this application embodiment designs a boundary consistency processing mechanism to ensure the consistency and efficiency of boundary task allocation. Specifically, boundary tasks are defined as tasks within a certain range from the sub-region boundary that meet the following conditions:
[0026]
[0027] Where D th This is a preset boundary distance threshold. To determine task ownership, the task execution efficiency score Score(a,j) is calculated using the following formula:
[0028]
[0029] Where: d(a,j) is the distance from agent a to task j; v a C represents the velocity of agent a. a The cost of performing a task for agent a; ρ a denoted as the ratio of the remaining energy of agent a; ∈ is a small constant to avoid a denominator of zero.
[0030] Ultimately, tasks are assigned to the sub-region of the agent with the lowest score, or completed collaboratively by agents in multiple sub-regions. This mechanism avoids the problems of task duplication or omission, while improving the execution efficiency of boundary tasks.
[0031] Step 3: To achieve efficient task allocation within the region, this embodiment employs asynchronous latent game theory to model the task allocation process of the agents. Each agent calculates a utility function based on its own capabilities and task value, and updates and selects the optimal task set through an asynchronous policy. The design of the utility function fully considers factors such as task value, duplicate coverage penalty, and execution cost, and its mathematical expression is:
[0032]
[0033] Wherein: S a V represents the set of tasks for agent a; j The value of task j; I bj ∈{0,1} indicates whether agent b performs task j; α jb C is the penalty coefficient for repeated coverage; a (S a Let agent a execute a set of tasks S. a The cost; β is the cost weighting coefficient.
[0034] Agent A asynchronously executes the optimal response strategy, selecting the set of tasks that maximizes the utility function: (S -a (Represents the set of tasks for other agents). The region command node records the latent function φ, which is used to evaluate the quality of the overall task allocation:
[0035]
[0036] Where |T| is the total number of tasks, and |N| is the total number of tasks. i | represents the total number of agents in the region. The convergence condition is met for K = 3 consecutive rounds:
[0037] |φ t -φ t-1 |<ε=0.01
[0038] The strategy is deemed to have reached a stable state. This method enables efficient distributed task allocation under communication constraints, while also possessing strong dynamic adaptability.
[0039] Step 4: In disaster emergency scenarios, the states of agents and tasks may change at any time. To address this, this embodiment designs a regional state detection mechanism to monitor in real time the access and exit of agents within the region, the addition and completion of tasks, and other related information. When a state change is detected, the agents within the region are triggered to asynchronously update their strategies until a new stable state is reached. Through this iterative detection and adjustment mechanism, this embodiment can dynamically adapt to changes in agent topology and task requirements, ensuring the continuous and efficient operation of the system.
[0040] Furthermore, step 1 specifically includes:
[0041] Step 1.1: Based on the actual situation at the disaster site, collect the weight and coordinate information of the set of tasks to be executed. The weight is used to represent the importance or urgency of the task, and the coordinate information is used to determine the location of the task in geographic space.
[0042] Step 1.2: Use the weighted Voronoi algorithm to divide the task set into regions, generate a set of sub-regions, and ensure that the total weight of the tasks in each sub-region meets the constraints.
[0043] If the weighted Voronoi algorithm cannot meet the actual needs, sub-regions are divided by manual intervention to ensure that the division results conform to the characteristics of task distribution and geographical constraints.
[0044] Step 1.3: Configure a regional command node in each sub-region to manage the task information and agent status within its region. The regional command node works in collaboration with other regional command nodes through the communication network to form the basic architecture for global task allocation.
[0045] Furthermore, the boundary consistency processing mechanism is specifically as follows:
[0046] For tasks located near the boundaries of adjacent sub-regions, they are defined as boundary tasks, satisfying the condition that the distance is no greater than a preset boundary distance threshold.
[0047] Command nodes in adjacent sub-regions exchange task information and agent status data through a preset communication protocol, and calculate the execution efficiency score of the boundary task;
[0048] The boundary task is divided into a sub-region to which the agent with the lowest score belongs, and the agent in that region executes the task independently; or, when cooperation among multiple agents is more efficient, the task is completed jointly by agents in multiple regions.
[0049] Furthermore, step 3 specifically includes:
[0050] Step 3.1: Based on the disaster information provided by the rescue command center, each regional command node generates an initial task list within its respective region, and the regional command node initially assigns the tasks within its region to its affiliated intelligent agents.
[0051] Step 3.2: Within a sub-region, the task allocation of the agent set is modeled as a potential game, and the utility function of the agent is defined; the agent executes the best response asynchronously, and the region command node records the potential function.
[0052] Step 3.3: Agents within the same region perform multiple rounds of asynchronous policy updates. When it is detected that the task allocation of each agent remains unchanged in multiple consecutive rounds of updates, the cooperative process is determined to have converged and the policy update loop ends.
[0053] Step 3.4: After the asynchronous policy update of the agents converges, they execute their respective assigned tasks. During this process, the regional agents report the task execution status information and agent status information to the regional command node in real time.
[0054] Furthermore, step 3.1 specifically includes:
[0055] Step 3.11: Each regional command node generates an initial task list for its region based on the disaster information provided by the rescue command center.
[0056] Step 3.12: Based on the importance of the task and the capabilities of the agents, the regional command node initially assigns the task to its assigned agents within the region.
[0057] Step 3.13: During the task allocation process, the regional command node can dynamically adjust the task allocation plan based on real-time information to ensure the rationality and efficiency of task allocation.
[0058] Furthermore, step 3.2 specifically includes:
[0059] Step 3.21: Model the task allocation of the agent set as a potential game and define the utility function of the agent.
[0060] Step 3.22: The agent asynchronously executes the optimal response strategy and selects the set of tasks that maximizes the utility function; the regional command node records the potential function, which is used to evaluate the quality of the overall task allocation.
[0061] Furthermore, step 3.3 specifically includes:
[0062] Step 3.31: Within the same region, the agents perform multiple rounds of asynchronous policy updates. When it is detected that the task allocation of each agent remains unchanged in continuous updates, the cooperative process is determined to have converged.
[0063] Step 3.32: Set a threshold for the change of the latent function. When the value is less than the set threshold for multiple consecutive rounds, the policy is considered to have converged, and the policy update loop ends.
[0064] Furthermore, step 3.4 specifically includes:
[0065] Step 3.41: After the policy update converges, the agents execute their respective tasks according to the assigned task list. The agents report the task execution status information and their own status information to the regional command node in real time.
[0066] Step 3.42: The regional command node receives and processes the state information of the intelligent agent to ensure the transparency and controllability of the task execution process.
[0067] Furthermore, step 4 specifically involves:
[0068] Step 4.1: The regional command node performs real-time detection of changes in the game state of its region. When the following situations occur, it is determined that the game state has changed: a new agent joins the region, an agent in the region leaves, a new task appears in the region, or an agent completes a certain task.
[0069] Step 4.2: When the game state changes, the agents in the region re-update their asynchronous policies until a new near-Nash equilibrium convergence point is reached.
[0070] Step 4.3: After the asynchronous strategy converges, continue to perform task allocation and execution and region status detection, repeating the process until all tasks in the region have been completed.
[0071] like Figure 2 As shown, this application also provides a specific embodiment of a multi-agent collaborative method for air-space-ground operations in earthquake rescue, including the following steps:
[0072] S1: Division of Autonomous Regions:
[0073] This paper establishes a foundation for distributed task management. In disaster emergency scenarios, the disaster area is vast, the tasks are complex and frequently change dynamically, and traditional centralized control methods are insufficient to address the challenges of limited communication and single points of failure. Therefore, this embodiment first decomposes the disaster area into multiple sub-regions based on geographical location or task density through autonomous region partitioning, with each sub-region managed by an independent regional command node. The weights w of the task set T = {j} to be executed are then used to further manage the tasks. j The coordinate information is used to generate a set of sub-regions {R} using either a weighted Voronoi algorithm or manual partitioning. i}, and ensure that each sub-region satisfies
[0074] For example, in an earthquake rescue operation, the rescue command center collected a set of tasks T = {j} to be performed in the disaster area based on preliminary reconnaissance data. These tasks included locations of rubble, suspected survivors, and road blockages. Each task j was assigned a weight w. j The weight is used to indicate the importance or urgency of a location. For example, the task weight for a suspected survivor location is set to 10, while the weight for a road blockage is 5.
[0075] To achieve task load balancing, a weighted Voronoi algorithm is used to generate the sub-region set {R}. i The algorithm comprehensively considers the task weights w. j Including geographic coordinate information, ensure that the total task weight of each sub-region does not exceed the preset maximum task load threshold W. max =30. For example, the disaster area was ultimately divided into three sub-regions R1, R2, and R3. R1 contains two ruins and one suspected survivor location, with a total weight of 26; R2 contains two road blockages, with a total weight of 10; and R3 contains two ruins, with a total weight of 16. For some special terrains (such as mountainous areas or river-separated areas), manual intervention is allowed for fine-tuning to adapt to complex geographical environments and task distribution characteristics.
[0076] Each sub-region is configured with a regional command node, responsible for managing task information and agent states within its region. This distributed architecture not only reduces global communication dependencies but also avoids the risk of single points of failure, laying a solid foundation for subsequent task allocation and dynamic adjustments.
[0077] S2: Boundary Consistency Handling
[0078] To ensure efficient task allocation across regions, this application introduces a boundary consistency processing mechanism to avoid task duplication or omission, as there may be task overlap or ambiguous boundaries between adjacent sub-regions.
[0079] Specifically, for tasks located near the boundaries of adjacent sub-regions, the relevant regional command nodes communicate and coordinate through a pre-defined protocol to determine task ownership and assign tasks to either an agent in one region for independent execution or to agents in multiple regions for collaborative completion, ensuring consistency in task allocation results. A boundary task is defined as task j within a certain range from the sub-region boundary, satisfying... (D th The condition is a preset boundary distance threshold. For example, at the boundary between R1 and R2, there is a suspected survivor location for task j, which is no more than 500 meters away from the boundary of either sub-region.
[0080] To determine task ownership, command nodes in adjacent sub-regions exchange task information and agent state data through a pre-defined communication protocol, and calculate a task execution efficiency score (Score(a,j)).
[0081]
[0082] Where d(a,j) is the distance from agent a to task j, and v a For the speed of the agent, C a The cost of an agent performing a task, ρ a Let be the ratio of the remaining energy of the agents, and ∈ be a small constant to avoid a denominator of zero. For example, suppose a drone a1 in R1 and a ground robot a2 in R2 can both perform task j, calculate their scores respectively:
[0083] For a1, its distance to the task is 1000 meters, its speed is 10 meters / second, its execution cost is 5, and its remaining energy ratio is 0.7. Therefore, Score(a1,j)≈107.14;
[0084] For a2, its distance to the task is 800 meters, its speed is only 2 meters per second, its execution cost is 3, and its remaining energy ratio is 0.9. Therefore, Score(a2,j)≈403.33.
[0085] Clearly, a1's lower score indicates it is more suitable for the task. Ultimately, task j is assigned to R1's drone for independent execution. If collaboration among multiple agents is more efficient, the task can be assigned to multiple sub-regions for collaborative completion.
[0086] Through the above protocols and calculations, it is ensured that the allocation results of all boundary tasks are consistent and conflict-free, avoiding duplicate or omitted tasks, improving the execution efficiency of boundary tasks, and ensuring the fairness and rationality of task allocation.
[0087] S3: Agent game task allocation: Dynamic task optimization based on asynchronous latent game.
[0088] After completing the regional division and boundary task processing, each regional command node generates an initial task list for its region based on disaster information provided by the rescue command center (such as the location of ruins and the reported locations of the injured). This task list includes information such as the geographical location of the task points, the task value, and the task cost. Based on the importance of the tasks and the capabilities of the agents, the regional command nodes initially assign tasks to their respective agents within the region. These agents include aerial agents, space-based agents, and ground-based agents, each undertaking different tasks according to its characteristics. During the task allocation process, the regional command nodes can dynamically adjust the task allocation scheme based on real-time information (such as changes in task priority and agent status updates) to ensure the rationality and efficiency of task allocation.
[0089] However, the environment and task requirements at disaster sites are often dynamically changing, thus requiring a flexible and efficient mechanism to dynamically adjust task allocation. To this end, this application employs asynchronous latent game theory to model the task allocation process of intelligent agents. Each agent calculates a utility function U based on its own capabilities and task value. a The optimal task set is selected by updating it using an asynchronous strategy.
[0090] The utility function is designed with full consideration of factors such as task value, duplicate coverage penalty, and execution cost. Its mathematical expression is as follows:
[0091]
[0092] Among them, S a Let V be the set of tasks for agent a. j For the value of task j, I bj ∈{0,1} indicates whether agent b performs task j, α jb C is the penalty coefficient for repeated coverage. a (S a Let agent a execute a set of tasks S. a The cost is β, where β is the cost weighting coefficient. For example, suppose that when drone a1 of R1 is performing tasks at ruins A and B, it finds that the task at ruins A has a higher value and is not covered by other agents, so it chooses to perform the task first.
[0093] The agent asynchronously executes the optimal response strategy, selecting the set of tasks that maximizes the utility function: Where S -a This represents the set of tasks for other intelligent agents.
[0094] At the same time, the regional command node records the potential function. It is used to evaluate the quality of the overall task allocation. When the change of the latent function is less than 0.01 for three consecutive rounds, the policy is considered to have reached a stable state. At this time, the agents perform their respective duties according to the assigned tasks and report the task execution status and their own status information to the regional command node in real time.
[0095] S4: Area State Detection: Dynamically Adapting to Changes in Agent Topology and Task Requirements. In disaster emergency scenarios, the states of agents and tasks may change at any time. Therefore, this application embodiment designs an area state detection mechanism to monitor in real time the access and exit of agents within the area, the addition and completion of tasks, etc. For example, during the search and rescue process in R1, assuming a new drone joins the area, the area command node immediately detects this change and notifies all agents to re-update their asynchronous policies. Similarly, if an agent exits due to equipment failure, or a new task suddenly appears, the system will also trigger a mechanism to re-update the policy.
[0096] Through a cyclical detection and adjustment mechanism, the system can dynamically adapt to changes in agent topology and task requirements, ensuring the flexibility and stability of task allocation. Finally, the collaborative process ends when all tasks within the region are completed.
[0097] This ability to detect and dynamically adjust in real time enables the system to maintain the flexibility and stability of task allocation in complex and ever-changing environments, thereby improving overall rescue efficiency. Through mechanisms such as autonomous region division, boundary consistency processing, agent game-based task allocation, and region state detection, the system can not only reduce communication dependencies and avoid single points of failure, but also respond in real time to changes in agent and task states, ensuring the flexibility, stability, and efficiency of overall task allocation, ultimately improving the overall efficiency and performance of rescue operations.
[0098] like Figure 4 As shown in the embodiments, this application also proposes a distributed air-space-ground multi-agent collaborative system. The system adopts a hierarchical architecture with regional autonomy, including two levels: regional command nodes and agent nodes. Regional command nodes can be high-altitude communication platforms (such as stratospheric high-altitude balloons or long-endurance UAVs) or low-Earth orbit satellites, used to schedule agents within their jurisdiction; agent nodes include ground robots, UAVs, small satellites, and other individuals performing specific perception tasks. The entire disaster site is divided into several adjacent areas, each configured with a regional command node responsible for task allocation and agent coordination within its area. When the tasks or agents in a certain area exceed its processing capacity, the regional command node collaborates with neighboring regional nodes through neighborhood communication, forming a distributed command network without centralized control. The main component functions are as follows:
[0099] Autonomous Region Partitioning Module: This module overcomes the limitations of centralized control by dividing the disaster site into several autonomous sub-regions. Each sub-region is managed by an independent regional command node, responsible for coordinating and scheduling task information and agent states within the region. During the partitioning process, the importance of tasks and geographical location information are comprehensively considered, and a weighted Voronoi algorithm is used to generate a set of sub-regions that meet load balancing requirements, ensuring that the total task weight of each sub-region does not exceed a preset maximum task load threshold. For complex terrain or special requirement scenarios, manual partitioning is allowed to adapt to complex geographical environments and task distribution characteristics. This distributed architecture not only effectively reduces global communication dependencies but also avoids the risk of single points of failure, providing a solid foundation for subsequent task allocation and dynamic adjustment.
[0100] Boundary Consistency Processing Module: To address potential task overlap or boundary ambiguity issues between adjacent sub-regions, this module designs a boundary consistency processing mechanism to ensure the consistency and efficiency of boundary task allocation. Boundary tasks are defined as those within a certain range of the sub-region boundary, and task assignment is determined by calculating a task execution efficiency score. This score comprehensively considers factors such as the distance from the agent to the task, speed, execution cost, and remaining energy ratio. Ultimately, the task is assigned to the sub-region belonging to the agent with the lowest efficiency score, or it can be completed collaboratively by agents from multiple sub-regions.
[0101] The agent-based task allocation module models the task allocation process of agents based on asynchronous latent game theory, achieving efficient task allocation within a region. Each agent calculates a utility function based on its own capabilities and task value, and selects the optimal task set through asynchronous policy updates. The utility function design fully considers factors such as task value, duplicate coverage penalty, and execution cost to ensure the rationality and efficiency of task allocation. The agent executes the optimal response strategy asynchronously, selecting the task set that maximizes the utility function. Simultaneously, the region command node records the latent function to evaluate the overall quality of task allocation. When the change in the latent function is less than a preset threshold in multiple consecutive updates, the policy is considered to have reached a stable state.
[0102] Area Status Detection Module: This module designs an area status detection mechanism to monitor the status changes of agents and tasks within the area in real time, including agent access and exit, task addition and completion, etc. When a status change is detected, it triggers the agents within the area to re-update their asynchronous policies until a new stable state is reached. Through a cyclical detection and adjustment mechanism, this module can dynamically adapt to changes in agent topology and task requirements, ensuring the continuous and efficient operation of the system in disaster emergency scenarios. This real-time detection and dynamic adjustment capability enables the system to maintain the flexibility and stability of task allocation in complex and ever-changing environments, thereby improving overall rescue efficiency.
[0103] The aforementioned methods, steps, and module designs work together to form a distributed collaborative task allocation scheme suitable for heterogeneous multi-agent environments. This scheme can be effectively deployed in integrated air-space-ground disaster emergency scenarios. For example, in the collaborative execution of tasks involving satellites, drones, and ground-based agents, each submodule collaborates according to its functional division of labor, dynamically adapting to changes in complex environments and task requirements, thereby achieving efficient task allocation and execution. Through mechanisms such as autonomous region division, boundary consistency processing, agent game-based task allocation, and regional state detection, the system not only reduces communication dependencies and avoids single points of failure but also responds in real time to changes in agent and task states, ensuring the flexibility, stability, and efficiency of overall task allocation, ultimately improving the overall efficiency and performance of rescue operations.
[0104] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0105] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.
Claims
1. An asynchronous cooperation method in a distributed air-space-ground multi-agent system, characterized in that, Includes the following steps: Step 1: Divide the disaster site into several sub-regions, using the weighted Voronoi algorithm or manually to generate a set of sub-regions, and assign a regional command node to each sub-region to manage it; Step 2: Introduce a boundary consistency handling mechanism to avoid task duplication or omission, and ensure the fairness and rationality of task allocation; Step 3: Each regional command node generates an initial task list within its respective region, models the task allocation of the agent set as a potential game, and calculates the utility function; the agents asynchronously execute the best response, select the optimal task set through asynchronous policy updates, and execute their assigned tasks after the asynchronous policy updates converge. Step 4: Design a regional state detection mechanism to monitor the state of agents within the region in real time. When the game state changes, the agents within the region update their strategies asynchronously until a new stable state is reached. This process is repeated until all tasks within the region have been completed. The boundary consistency processing mechanism is specifically as follows: For tasks located near the boundaries of adjacent sub-regions, they are defined as boundary tasks, satisfying the condition that the distance is no greater than a preset boundary distance threshold. Command nodes in adjacent sub-regions exchange task information and agent state data through a preset communication protocol, and calculate the execution efficiency score of the boundary task. ; in, For intelligent agents To the mission The distance; For intelligent agents speed; For intelligent agents The cost of performing the task; For intelligent agents The ratio of remaining energy; To avoid constants with a denominator of zero; Divide the boundary tasks into the smallest The value belongs to a sub-region of the agent, and the agent in that region performs the task independently; or, when cooperation among multiple agents is more efficient, the agent in multiple regions completes the task together. The utility function considers factors such as task value, repetition penalty, and execution cost, and its mathematical expression is: ; in, For intelligent agents The task set, For the task value, Represents intelligent agents Should the task be executed? , The penalty coefficient for repeated coverage. For intelligent agents Execution task set The cost, This is the cost weighting coefficient.
2. The method according to claim 1, characterized in that, Step 1 specifically involves: Step 1.1: Based on the actual situation at the disaster site, collect the weight and coordinate information of the set of tasks to be executed. The weight is used to represent the importance or urgency of the task, and the coordinate information is used to determine the location of the task in geographic space. Step 1.2: Use the weighted Voronoi algorithm to divide the task set into regions, generate a set of sub-regions, and ensure that the total weight of the tasks in each sub-region meets the constraints. If the weighted Voronoi algorithm cannot meet the actual needs, sub-regions are divided by manual intervention to ensure that the division results conform to the characteristics of task distribution and geographical constraints. Step 1.3: Configure a regional command node in each sub-region to manage the task information and agent status within its region. The regional command node works in collaboration with other regional command nodes through the communication network to form the basic architecture for global task allocation.
3. The method according to claim 1, characterized in that, Step 3 specifically involves: Step 3.1: Based on the disaster information provided by the rescue command center, each regional command node generates an initial task list within its respective region, and the regional command node initially assigns the tasks within its region to its affiliated intelligent agents; Step 3.2: Within a sub-region, the task allocation of the agent set is modeled as a potential game, and the utility function of the agent is defined; the agent executes the best response asynchronously, and the region command node records the potential function; Step 3.3: Agents within the same region perform multiple rounds of asynchronous policy updates. When it is detected that the task allocation of each agent remains unchanged in multiple consecutive rounds of updates, the cooperative process is determined to have converged and the policy update loop ends. Step 3.4: After the asynchronous policy update of the agents converges, they execute their respective assigned tasks. During this process, the regional agents report the task execution status information and agent status information to the regional command node in real time.
4. The method according to claim 3, characterized in that, Step 3.1 specifically involves: Step 3.11: Each regional command node generates an initial task list for its region based on the disaster information provided by the rescue command center; Step 3.12: Based on the importance of the task and the capabilities of the agents, the regional command node initially assigns the task to its assigned agents within the region; Step 3.13: During the task allocation process, the regional command node can dynamically adjust the task allocation plan based on real-time information to ensure the rationality and efficiency of task allocation.
5. The method according to claim 3 or 4, characterized in that, Step 3.2 specifically involves: Step 3.21: Model the task allocation of the agent set as a potential game and define the utility function of the agent; Step 3.22: The agent asynchronously executes the optimal response strategy and selects the set of tasks that maximizes the utility function; the regional command node records the potential function, which is used to evaluate the quality of the overall task allocation.
6. The method according to claim 3, characterized in that, Step 3.3 specifically involves: Step 3.31: Within the same region, the agents perform multiple rounds of asynchronous policy updates. When it is detected that the task allocation of each agent remains unchanged in continuous updates, the cooperative process is determined to have converged. Step 3.32: Set a threshold for the change of the latent function. When the value is less than the set threshold for multiple consecutive rounds, the policy is considered to have converged, and the policy update loop ends.
7. The method according to claim 3 or 6, characterized in that, Step 3.4 specifically involves: Step 3.41: After the policy update converges, the agents execute their respective tasks according to the assigned task list. The agents report the task execution status information and their own status information to the regional command node in real time. Step 3.42: The regional command node receives and processes the state information of the intelligent agent to ensure the transparency and controllability of the task execution process.
8. The method according to claim 1, characterized in that, Step 4 specifically involves: Step 4.1: The regional command node performs real-time detection of changes in the game state of its region. When the following situations occur, it is determined that the game state has changed: a new agent joins the region, an agent in the region leaves the region, a new task appears in the region, or an agent completes a certain task. Step 4.2: When the game state changes, the agents in the region re-update their asynchronous policies until a new near-Nash equilibrium convergence point is reached; Step 4.3: After the asynchronous strategy converges, continue to perform task allocation and execution and region status detection, repeating the process until all tasks in the region have been completed.