Multi-agent task-role real-time matching method and system in dynamic environment
Patent Information
- Application Number
- CN202610890731.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-18
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2046-06-18
AI Technical Summary
[0007]基于背景技术存在的技术问题,本发明提出了动态环境下多智能体任务-角色实时匹配方法及系统,旨在解决现有技术在处理复杂时序逻辑任务时可控性差、以及动态环境下资源分配效率低的问题
[0018] The advantages of the multi-agent task-role real-time matching method and system in dynamic environments provided by this invention are as follows: The multi-agent task-role real-time matching method and system in dynamic environments provided by the structure of this invention...
Smart Images

Figure CN122414758B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of collaborative control and decision-making technology for multi-agent systems, and particularly to a method and system for real-time matching of multi-agent tasks and roles in dynamic environments. Background Technology
[0002] With the rapid development of artificial intelligence and robotics, multi-agent systems, with their parallelism, robustness, and flexibility, have demonstrated enormous application potential in complex scenarios such as drone swarm inspections, smart logistics and warehousing scheduling, wilderness disaster search and rescue, and military swarm operations. In these applications, how to achieve efficient matching between agents and tasks, and how to adjust role allocation in real time in dynamically changing environments, are key issues determining the overall system performance.
[0003] However, existing multi-agent task allocation and cooperative control technologies still have many limitations when facing complex environments with high dynamics and strong constraints: First, traditional task allocation methods primarily handle simple point-to-point movement or single-target attack tasks, making it difficult to handle tasks with complex temporal constraints (e.g., patrolling area A first, collecting data, then moving to area B, while avoiding dangerous area C throughout the process). Existing reinforcement learning methods typically employ scalar reward signals. When faced with such long-sequence, sparse-reward logical tasks, the agent often struggles to converge, and may even violate temporal constraints, leading to task failure.
[0004] Secondly, most existing task-role matching algorithms rely on centralized combinatorial optimization methods (such as the Hungarian algorithm and genetic algorithms). While these methods perform reasonably well in static environments, their computational complexity increases exponentially with the number of agents in large-scale clusters and dynamic environments, making it difficult to meet real-time requirements. Furthermore, when the environment undergoes sudden changes (such as the introduction of new tasks or threats), traditional algorithms often need to recalculate the global optimum from scratch, resulting in a lag in system response and a lack of ability to adjust strategies in real time during dynamic interactions.
[0005] Furthermore, in terms of communication topology, existing multi-agent collaborative architectures typically employ a fixed hierarchical structure (such as a pre-defined Leader-Follower pattern). This rigid topology is extremely vulnerable to communication link disruptions caused by strong electromagnetic interference or physical damage. Once a critical leader node fails or communication is blocked, the collaborative function of the entire subsystem may be instantly paralyzed, lacking self-healing capabilities and role reorganization mechanisms based on local perception information.
[0006] Therefore, there is an urgent need to propose a multi-agent task-role real-time matching system and method that can understand complex temporal logic reductions, has distributed matching capabilities with low computational complexity, and exhibits strong robustness and self-healing in dynamic environments, in order to solve the aforementioned technical bottlenecks. Summary of the Invention
[0007] Based on the technical problems existing in the background technology, this invention proposes a method and system for real-time matching of multi-agent tasks and roles in dynamic environments, aiming to solve the problems of poor controllability and low resource allocation efficiency in dynamic environments when the existing technology is dealing with complex sequential logic tasks.
[0008] The present invention proposes a real-time multi-agent task-role matching method in dynamic environments, comprising: S1: Transform the linear sequential logic task constraints into a nondeterministic Butch automaton, and construct a product Markov decision process with the environmental Markov decision process to provide the agent with an enhanced state space that includes physical state and logical automaton state. S2: Based on the current state of each agent in the enhanced state space and the dynamic task set acquired in real time, construct a dynamic heterogeneous bipartite graph of agents and tasks, use graph attention network to extract feature embedding and quantify matching benefits, and then solve the Nash equilibrium point through a potential game model to output the probability distribution of task selection. S3: Based on the probability distribution and the logic-guided dense reward derived from the product Markov decision process, a centralized training-distributed execution architecture and task course learning are adopted to update the policy network online until convergence, thereby obtaining the optimal task allocation. S4: Based on dynamic spectrum clustering and distributed negotiation mechanism, analyze the communication topology of multi-agent system in real time, realize role scheduling and task execution in dynamic environment, and trigger step S2 to rematch when the topology breaks or the task changes.
[0009] Further, step S1 includes: The temporal constraints and safety specifications of multi-agent collaborative tasks are described using linear sequential logic and transformed into nondeterministic Butch automata. The environmental Markov decision process is combined with the Butch automaton to construct a product Markov decision process; In the product Markov decision process, a state transition function is defined such that the agent's state vector includes both physical state and logical automaton state.
[0010] Further, step S2 includes: The multi-agent set and the dynamic task set are modeled as a dynamic heterogeneous bipartite graph, where nodes include agent nodes and task nodes, edges represent potential executable relationships, and physical state and logical automaton state serve as the initial features of agent nodes. The ability feature embedding of the agent and the requirement feature embedding of the task are extracted using a graph attention network, and the similarity of the feature vectors is calculated as the matching reward value. Construct a potential game model to transform the problem of maximizing global payoff into finding the Nash equilibrium point of the game and output the probability distribution of the agent's choice of tasks.
[0011] Furthermore, by using the matching payoff value output by the graph attention network as the input to the game model, the generated probability selection distribution follows a Boltzmann distribution.
[0012] Further, step S3 includes: A logic-guided potential energy reward function is constructed to provide positive incentives when the agent's action triggers the spawn automaton to transition to the accepting state, in order to solve the sparse reward problem. Design a course learning mechanism with increasing task complexity, prioritizing training on simple tasks that satisfy linear sequential logic formulas, and gradually transitioning to complex sequential logic tasks; A centralized training-distributed execution network architecture is adopted. A gated communication mechanism is used to filter neighbor information whose importance weight exceeds a preset weight threshold, and update the policy network parameters until the multi-agent system converges to the cooperative optimal solution.
[0013] Furthermore, the formula for the potential energy reward function is as follows: ; in, In physical state Next action Then, the physical state transitions to Simultaneously, the state of the logical automaton changes from... Transfer to At that time, the system provides additional reward signals. These are the weighting coefficients. As a discount factor, They are time points The state of the logical automaton, They are time points The physical state, To accept the set of states, Rewards for achieving the goal For indicator functions, when belong The value is 1 if it is true, and 0 otherwise.
[0014] Furthermore, a gating communication mechanism is used to filter neighbor information whose importance weight exceeds a preset weight threshold, specifically as follows: Introducing communication gating units into distributed execution networks, intelligent agents Real-time calculation of the importance weights of the acquired observed features ,when At that time, intelligent agent Broadcast the intent vector encoded by the local policy network to neighbors. The preset weight threshold.
[0015] Further, step S4 includes: Maintain the agent communication topology graph in real time, and calculate the graph Laplacian matrix and its Fiedler vector; Fast spectral clustering using Fiedler vectors is used to divide the agent cluster into multiple functional subclusters, and roles are dynamically elected within each subcluster based on the contract network protocol. When a topology break or task change is detected, a renegotiation mechanism is triggered to switch roles and reorganize resources based on the latest local perception information. Step S2 is then triggered to rematch based on the latest state in the enhanced state space and the latest acquired dynamic task set.
[0016] Furthermore, the dynamic role election within each sub-cluster based on the contract network protocol specifically involves: Based on intelligent agents Remaining energy and computing power Calculate the overall score : ; in, These are the weight parameters; The agent with the highest overall score is elected as the temporary cluster leader, while the other agents act as execution nodes, following the local instructions of the cluster leader, forming a hybrid control architecture of global distributed and local hierarchical.
[0017] A computer system includes a memory, a processor, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the method described above.
[0018] The advantages of the multi-agent task-role real-time matching method and system in dynamic environments provided by this invention are as follows: The multi-agent task-role real-time matching method and system in dynamic environments provided by the structure of this invention... Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the process of the present invention. Detailed Implementation
[0020] The technical solution of the present invention will now be described in detail through specific embodiments. Many specific details are set forth in the following description to provide a thorough understanding of the invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar modifications without departing from the spirit of the invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0021] like Figure 1 As shown, the multi-agent task-role real-time matching method proposed in this invention for dynamic environments includes: S1: Transform the linear sequential logic task constraints into a nondeterministic Butch automaton, and construct a product Markov decision process with the environmental Markov decision process to provide the agent with an enhanced state space that includes physical state and logical automaton state. S2: Based on the current state of each agent in the enhanced state space and the dynamic task set acquired in real time, construct a dynamic heterogeneous bipartite graph of agents and tasks, use graph attention network to extract feature embedding and quantify matching benefits, and then solve the Nash equilibrium point through a potential game model to output the probability distribution of task selection. S3: Based on the probability distribution and the logic-guided dense reward derived from the product Markov decision process, a centralized training-distributed execution architecture and task course learning are adopted to update the policy network online until convergence, thereby obtaining the optimal task allocation. S4: Based on dynamic spectrum clustering and distributed negotiation mechanism, analyze the communication topology of multi-agent system in real time, realize role scheduling and task execution in dynamic environment, and trigger step S2 to rematch when the topology breaks or the task changes.
[0022] In this embodiment, a task refers to the target behavior that the multi-agent system needs to complete, including target area inspection, data acquisition, target search, material transportation, anomaly handling, and composite tasks that satisfy temporal logic constraints. A role refers to the functional identity that an agent assumes during task execution, including temporary cluster leader, execution node, collaborative perception node, and communication relay node. The real-time task-role matching includes: determining the matching relationship between agents and tasks based on agent states and dynamic task sets, and determining the role division of agents during the corresponding task execution process based on communication topology and agent capability states. This embodiment aims to solve the problems of poor controllability and low resource allocation efficiency in dynamic environments when handling complex temporal logic tasks in existing technologies.
[0023] Enhancing System Controllability and Standardization under Complex Task Constraints: This embodiment transforms complex linear temporal logic (LTL) tasks described in natural language into mathematically rigorous topological constraints by introducing a nondeterministic Büchi automata and a product Markov decision process. Compared to traditional methods, this embodiment utilizes a logic-guided potential reward function to address the sparse reward problem, ensuring that the multi-agent system strictly adheres to temporal logic and safety standards when executing long sequence tasks, thus avoiding task failures due to logic violations.
[0024] Achieving real-time and accurate matching of tasks and roles in dynamic environments: This embodiment abandons the traditional centralized combinatorial optimization solver and innovatively combines Graph Attention Network (GAT) with game theory. By calculating feature embeddings and Nash equilibrium probabilities in a dynamic heterogeneous bipartite graph, it achieves fast decision-making with O(1) time complexity, significantly improving the system's response speed and resource allocation efficiency under task mutations or environmental disturbances.
[0025] Optimizing the multi-agent collaboration mechanism and enhancing the system's self-healing capability: This embodiment uses spectral graph theory (Fiedler vectors) for dynamic sub-cluster partitioning and combines it with a distributed contract network protocol to achieve real-time role scheduling. This hybrid architecture of global distributed and local hierarchical structure eliminates the risk of single-point failure, endows the system with strong self-healing capabilities in the event of communication interruptions or node failures, and significantly improves the robustness of large-scale clusters.
[0026] In one embodiment, in order for the multi-agent system to understand and strictly follow complex natural language task instructions (such as patrolling area A first, then moving to area B if an anomaly is found, and avoiding area C throughout), this embodiment introduces a formal method for step S1, specifically including steps S1.1 to S1.3.
[0027] S1.1: Use linear temporal logic (LTL) to describe the temporal constraints and safety specifications of multi-agent cooperative tasks, and transform the LTL formula into a nondeterministic Büchi automaton (NBA), where the Büchi automaton is a type of Büchi automaton.
[0028] In this embodiment, the linear sequential logic (LTL) formula is used. To describe the task specification. Linear temporal logic formulas are derived from a set of atomic propositions. Boolean logic operators ( No, and, (or) and timing operators ( The next moment, always, final, Until) constitutes.
[0029] For example, the mission eventually reaches the target area. And always avoids obstacles It can be represented as: .
[0030] The above logic formula is transformed using a linear sequential logic transformation algorithm (such as the LTL2BA algorithm). This is transformed into a nondeterministic Büchi automaton (NBA). The automaton is defined as a quintuple. ,in: It is a set of states of a logical automaton, representing different logical stages in the task execution process (e.g., logical states such as: key not obtained, key obtained, task completed). It is an input alphabet representing the set of atomic events in the environment; It is the state transition function of a logical automaton, which describes how the automaton state jumps when a specific atomic proposition is satisfied; It is the initial state. It is a set of accepting states.
[0031] The acceptance condition for a Büchi automaton is: for an infinite path, if it visits the set of accepting states an infinite number of times. If the state in the trajectory is such that the trajectory satisfies the linear time-series logic formula, then the trajectory is said to satisfy the linear time-series logic formula. .
[0032] Step S1.2: Define the Environment Markov Decision Process (MDP), and perform a Cartesian product operation on the environment state space and the Büchi automaton state space to construct a product Markov Decision Process containing the automaton states.
[0033] First, the physical environment in which the multi-agent group exists is modeled as an environmental Markov decision process (MDP), denoted as ,in This is the physical state space (such as the position, velocity, orientation, charge, and load status of each agent). For an action space (for example, for a single intelligent agent, it can include basic actions such as movement direction, acceleration and deceleration, and task switching). This is a label function used to determine which atomic propositions the current physical state satisfies (e.g., whether the agent is currently located in an obstacle area). Let be the basic reward function for the physical environment, used to characterize the agent's basic behavioral gains in the physical environment. γ is a discount factor, and P is the state transition probability function, describing the probability distribution of the environment transitioning to the next state after performing a joint action in the current state. Unlike the logic-guided potential reward function built on the state of the logic automaton, which is used to provide additional evaluation of the degree of satisfaction of the temporal logic of the task during the reinforcement learning training phase.
[0034] To embed logical constraints into the agent's decision-making process, this embodiment constructs a Product Markov Decision Process (Product MDP), denoted as... ,in, For environmental Markov decision processes, For product Markov decision processes, This is a nondeterministic Büchi automaton. The state space of the product Markov decision process. Defined as the Cartesian product of the physical state and the state of the logical automaton: ; in, It is a set of states of a logical automaton. For physical state space, For a single physical state, This represents a single logical automaton state.
[0035] The formula indicates that an extended state of an agent includes not only its physical state. It also includes its current logical automaton state. The progress.
[0036] Step S1.3: Define a state transition function in the product Markov decision process so that the agent's state vector contains both physical state and logical automaton state, providing a state space with temporal logic awareness for subsequent reinforcement learning.
[0037] In the product Markov decision process, state transition The transition probability is determined by both the physical environment and the logical transition of the automaton. The specific rules are as follows: Physical state transition: It is in state Execute action Then based on environmental probability The physical state obtained at the next moment.
[0038] Logic automaton state transitions: It must be a nondeterministic Büchi automaton in the current logical automaton state. Under these conditions, the physical state is received. Corresponding label function Then, the next state of the logical automaton that allows the jump. ,Right now ,in This is the state transition function for a logic automaton.
[0039] Through this state enhancement, the agent can perceive the impact of its current action on task progress. If the agent performs an action that violates safety constraints (such as entering a restricted area), the labeling function... This will cause the nondeterministic Büchi automaton to fail to make a valid transition (or transition to a deadlock state), thus being identified and penalized in subsequent reinforcement learning steps. This design provides a low-level data structure with complete temporal logic information for subsequent distributed learning algorithms.
[0040] In one embodiment, to address the combinatorial explosion problem of task allocation in dynamic environments, this embodiment abandons the traditional centralized solver and instead proposes step S2, which employs a graph neural network for feature extraction and combines game theory to achieve distributed decision-making. Step S2 includes sub-steps S2.1 to S2.3: Step S2.1: Model the multi-agent set and dynamic task set as a dynamic heterogeneous bipartite graph. The nodes include agent nodes and task nodes, the edges represent potential executable relationships, and the physical state and logical automaton state serve as the initial features of the agent nodes.
[0041] At every decision moment The system maintains a dynamic heterogeneous bipartite graph. .
[0042] Node set: Represents the current set of idle agents; Represents the set of tasks currently awaiting assignment, where, For the first An intelligent agent. For the first One task.
[0043] Edge set: one edge in Indicates the first A smart agent Capable of executing the Task The agent's basic capabilities (such as sufficient power, compatible equipment, and being within reach) are considered. If an agent cannot perform a task (e.g., a drone cannot perform an underwater mission), then there is no corresponding edge in the dynamic heterogeneous bipartite graph. This pre-screening mechanism significantly reduces the search space for subsequent computations.
[0044] Step S2.2: Use a graph attention network (GAT) to extract the ability feature embeddings of the agent and the requirement feature embeddings of the task, and calculate the dot product of the feature vectors as the matching reward value.
[0045] To accurately quantify the matching degree between the agent and the task, this embodiment uses a graph attention network to process the dynamic heterogeneous bipartite graph. For each node in the bipartite graph, its original feature vector is input (features of agent nodes include position, velocity, remaining battery power, and payload type; features of task nodes include target position, deadline, and priority). The graph attention network layer updates the node embedding by aggregating neighbor node information: ; in, It is the updated node embedding vector. It is the set of neighboring nodes. It is a learnable weight matrix. Embed the node vector before the update. It is a non-linear activation function. Attention coefficients. Represents a node For nodes The importance of is calculated by normalization using the Softmax function.
[0046] go through After aggregation by the layered graph attention network, an intelligent agent is obtained. High-dimensional embedding representation and tasks High-dimensional embedding representation At this moment, the intelligent agent Select task Matching benefit value Defined as the similarity between two feature embedding vectors (such as dot product or cosine similarity): ; in, Let be the cosine similarity.
[0047] The matching benefit value reflects the agent's performance considering the global topology. Execute the task The overall effect.
[0048] Step S2.3: Construct a potential game model, transforming the global payoff maximization problem into finding the Nash equilibrium point of the game, and outputting the probability distribution of the agent's choice of tasks.
[0049] This embodiment models the task allocation problem as a potential game. In this game model, each agent is considered a participant, whose policy space is a set of available tasks, and whose utility function is the matching payoff value calculated above. A key property of the potential game is the existence of a global state function. The change in utility caused by a unilateral change in policy by an agent and the overall situation function The changes are completely consistent. This means that when all agents reach Nash equilibrium through game theory, the system's overall situation function (i.e., the overall task allocation efficiency) is also at a local maximum or global maximum.
[0050] To achieve a fast solution, this embodiment does not employ a time-consuming iterative optimal response algorithm, but instead directly utilizes the matching reward value output by the graph attention network in step S2.2 to generate a probabilistic strategy. (Agent) Select task probability Follows Boltzmann distribution (Softmax): ; in, For temperature parameters, For task indexing, For intelligent agents The set of neighboring tasks, For intelligent agents Select task The matching reward value is used to adjust the balance between exploration and exploitation. This probability distribution represents the agent's optimal hybrid strategy at the current moment, achieving... Fast matching decision with minimal time complexity.
[0051] In one embodiment, after determining the probabilistic strategy for task allocation, the agent needs to learn specific action control strategies to complete the task. To address the problem of sparse rewards and difficulty in convergence for complex logical tasks, this embodiment introduces step S3, which specifically includes sub-steps S3.1 and S3.3.
[0052] Step S3.1: Construct a logic-guided potential energy reward function to provide positive incentives when the agent's action triggers the nondeterministic Büchi automaton to transition to the accepting state, thus solving the sparse reward problem.
[0053] Traditional reinforcement learning typically rewards agents only upon task completion (+1), making it difficult for agents to find the correct path in long-sequence logical tasks. This embodiment utilizes the nondeterministic Büchi automaton (NBA) constructed in step S1 to design a potential energy reward function based on the automaton's state potential.
[0054] Define the potential energy reward function State of a logical automaton Up to the latest acceptance status The reciprocal of the graph theory distance. When the agent performs an action. Caused the system state to change Transferred to At that time, logic-guided reward The calculation is as follows: ; in, In physical state Next action Then, the physical state transitions to Simultaneously, the state of the logical automaton changes from... Transfer to At that time, the system provides additional reward signals. These are the weighting coefficients. As a discount factor, They are time points The state of the logical automaton, They are time points The physical state, To accept the set of states, A reward for achieving the goal, namely, a large additional reward for reaching the acceptance state. For indicator functions, when belong The value is 1 if it is true, and 0 otherwise.
[0055] First item The potential energy difference deformation reward is used to guide the agent towards the accepting state. Furthermore, if the agent triggers a transition that violates the linear temporal logic safety rules (i.e., the nondeterministic Büchi automaton has no valid successor state), a large penalty is imposed. The final total reward function is: ,in Rewards are based on the physical environment (such as obstacle avoidance and energy consumption).
[0056] Step S3.2: Design a course learning mechanism with increasing task complexity, prioritizing training on simple tasks that satisfy the atomic linear sequential logic formula, and gradually transitioning to composite sequential logic tasks.
[0057] To accelerate the convergence speed of neural networks, this embodiment draws on the gradual learning approach of humans and designs a task curriculum.
[0058] Beginner courses (atomic tasks): only include simple reachability tasks (such as...) (i.e., ultimately reaching point A). This stage primarily trains the agent's basic navigation and obstacle avoidance capabilities.
[0059] Intermediate Course (Sequential Task): Introducing Sequence Constraints (e.g.) (That is, go to A first and then go to B). This stage trains the agent to understand the state transition logic of the state machine.
[0060] Advanced Courses (Complex Constraint Tasks): Introducing complex safety constraints and cyclical tasks (such as...) This means patrolling A an infinite number of times while always avoiding C. During training, when the agent's average task success rate at the current difficulty level exceeds a set threshold (e.g., 85%), it is automatically promoted to the next difficulty level. This mechanism prevents the agent from blindly exploring high-difficulty logic in the early stages.
[0061] Step S3.3: Adopt a centralized training-distributed execution (CTDE) architecture, use a gated communication mechanism to filter neighbor information whose importance weights exceed a preset weight threshold, update the policy network parameters, until the multi-agent system converges to the cooperative optimal solution.
[0062] Specifically, this embodiment employs an improved algorithm based on Multi-Agent Proximal Policy Optimization (MAPPO). This algorithm follows a centralized training-distributed execution (CTDE) architecture: during the training phase, a central Critic network is trained using global state information. This network evaluates the value of joint actions and guides the updates of the Actor networks for each agent; during the execution phase, each agent relies solely on local observations to input the Actor network's decisions.
[0063] To enhance collaboration and reduce communication bandwidth, communication gating units are introduced into the Actor network. (Agent) Real-time calculation of the importance weights of its observed features Only when (in When a threshold is set (e.g., a pre-defined threshold), the agent broadcasts its intent vector, encoded by the local policy network, to its neighbors. This ensures that bandwidth resources are used only to transmit critical cooperative information (such as an impending collision or the discovery of a target), rather than redundant background noise. Through continuous online iteration of this algorithm, the agent can learn a set of joint policies that satisfy Nash equilibrium, achieving optimal cooperation under complex linear-sequential logic timing constraints.
[0064] In one embodiment, this embodiment addresses network topology changes and sudden task changes caused by node movement in a dynamic environment by implementing adaptive role scheduling through sub-steps S4.1 to S4.3.
[0065] Step S4.1: Maintain the agent communication topology graph in real time, and calculate the graph Laplacian matrix and its Fiedler vector.
[0066] The system monitors the communication connection status between each agent in real time (e.g., based on communication distance or signal strength thresholds) and constructs a real-time communication topology graph. , Let be the set of all nodes in the communication topology graph, with each node corresponding to an agent. For the combination of all edges in the communication topology graph, if there is a direct communication link between two agents (e.g., the distance is less than the communication radius or the signal strength is higher than the threshold), then there exists an undirected edge, which reflects the information reachability between agents.
[0067] Based on this topological graph, calculate its Laplace matrix. ,in Degree matrix ( ), It is an adjacency matrix.
[0068] Furthermore, the Laplace matrix is calculated using a distributed eigenvalue decomposition algorithm (such as the distributed power iteration method). The eigenvector corresponding to the second smallest eigenvalue, i.e., the Fiedler vector. The numerical distribution of Fiedler vectors can reflect the algebraic connectivity of a graph and the clustering characteristics of its nodes, and is a key mathematical basis for network segmentation in graph theory.
[0069] Step S4.2: Use Fiedler vectors to perform fast spectral clustering, divide the agent cluster into several functional sub-clusters, and dynamically elect roles within the sub-clusters based on the contract network protocol.
[0070] Based on the numerical features in the Fiedler vectors (such as the sign of the values or the Eigengap interval), large-scale intelligent agent clusters can be quickly divided into several functional subclusters with strong connectivity. ,in For the first This partitioning method, based on spectral theory, minimizes communication cuts between subclusters, ensuring efficient and stable communication within subclusters.
[0071] Within each predefined sub-cluster, an improved contract network protocol is used for role election. The system comprehensively considers the intelligent agents. Remaining energy and computing power Calculate the overall score .in, These are weight parameters, preset based on prior knowledge or system requirements. and The range of values for are respectively , And satisfy ; This indicates the weight of the remaining energy in the role election. This indicates the weight of computing power in the role election.
[0072] The node with the highest score is elected as the temporary cluster leader, responsible for coordinating local tasks, converging states, and scheduling resources within that subcluster; the remaining agents act as execution nodes, following the cluster leader's local instructions. This forms a hybrid control architecture of global distributed and local hierarchical structure in the system.
[0073] Step S4.3: When a topology break or task change is detected, a renegotiation mechanism is triggered. Role switching and resource reorganization are performed based on the latest local perception information. Step S2 is triggered to rematch based on the latest state in the enhanced state space and the latest acquired dynamic task set, thereby ensuring the self-healing and connectivity of the system in a dynamic environment.
[0074] When an agent detects a break in the communication link with its neighbor (i.e., a drastic change in the topology graph, resulting in a sharp drop in algebraic connectivity) or receives a sudden high-priority task instruction, it immediately triggers the renegotiation mechanism (i.e., triggers step S2 for rematching).
[0075] Relevant nodes recalculate the Fiedler vector components based on the latest local topology information and update the state of their respective subclusters. If the original cluster leader node fails or goes offline, the remaining nodes in the cluster automatically trigger a new round of contract network protocol election, electing a new leader node to take over command within milliseconds, achieving a smooth role switch. This dynamic reorganization mechanism endows the system with strong self-healing capabilities, ensuring that even in extreme environments where some nodes fail or communication is disrupted, the multi-agent system can still maintain overall connectivity and the continuity of task execution.
[0076] This embodiment enhances the controllability and standardization of the system under complex task constraints, realizes real-time and accurate matching of tasks and roles in dynamic environments, optimizes the multi-agent collaboration mechanism and enhances the system's resource allocation capabilities, and provides strong theoretical support and technical guarantee for the actual deployment of large-scale heterogeneous multi-agent systems in complex dynamic environments.
[0077] The above are basic embodiments of the present invention. The technical solution of the present invention will be further described below through a preferred embodiment.
[0078] Example 1 A method for real-time role matching in multi-agent tasks under dynamic environments, including: This embodiment uses a post-disaster drone swarm collaborative inspection and emergency search and rescue operation within a disaster-stricken industrial park as a specific application scenario. The multi-agent system includes... A drone, denoted as Each drone is equipped with an image acquisition module, a positioning module, a communication module, and an onboard computing unit. In Example 1, the intelligent agent is specifically a drone; the tasks include target area inspection, disaster image acquisition, abnormal target search, danger area avoidance, and communication link maintenance; the roles include temporary cluster leader, task execution node, collaborative perception node, and communication relay node.
[0079] T1: Transforms linear sequential logic task constraints into nondeterministic Butch automata, and constructs a product Markov decision process with the environmental Markov decision process, providing the agent with an enhanced state space that includes physical state and logical automaton state.
[0080] In this embodiment, the linear temporal logic task constraints are used to describe the inspection and search and rescue tasks of the UAV swarm in the post-disaster area. For example, πA represents the UAV arriving at inspection area A, πD represents completing disaster image acquisition, πE represents discovering an abnormal target, πB represents arriving at the abnormal target verification area B, and πC represents entering the danger area C. The task specification, which describes "first completing area inspection and data acquisition, then proceeding to the verification area after discovering an abnormal target, while avoiding danger areas throughout the process," is described by a linear temporal logic formula and transformed into a nondeterministic Butch automaton. The physical states in the environmental Markov decision process include the UAV's position, speed, remaining battery power, communication status, payload type, and task execution status, so that the augmented state of each UAV simultaneously reflects its physical motion state and task logic progress.
[0081] T2: Based on the current state of each agent in the augmented state space and the dynamic task set acquired in real time, a dynamic heterogeneous bipartite graph of agents and tasks is constructed. Feature embedding and matching benefits are extracted using a graph attention network. Then, the Nash equilibrium point is solved through a latent game model, and the probability distribution of task selection is output.
[0082] In this embodiment, the dynamic task set includes target area inspection tasks, image acquisition tasks, abnormal target search tasks, verification and confirmation tasks, and communication link maintenance tasks. Agent nodes correspond to various drones, and task nodes correspond to the aforementioned tasks to be executed. The characteristics of drone nodes include current location, remaining battery power, communication quality, payload type, and computing power; the characteristics of task nodes include task location, task priority, deadline, required payload type, and communication requirements. When a drone meets the basic conditions for executing a task, an edge is established between the corresponding drone node and the task node. A graph attention network is used to calculate the matching reward between drones and tasks, and the task selection probability of each drone for different tasks is output to determine the real-time matching relationship between drones and tasks.
[0083] T3: Based on the probability distribution and the logic-guided dense reward derived from the product Markov decision process, a centralized training-distributed execution architecture and task-based learning are adopted to update the policy network online until convergence, thereby obtaining the optimal task allocation.
[0084] In this embodiment, for UAVs performing inspection tasks, logical guidance dense rewards are used to encourage them to reach the designated inspection area and complete image acquisition; for UAVs performing abnormal target search or verification tasks, logical guidance dense rewards are used to encourage them to reach the abnormal target area while avoiding dangerous areas; for UAVs undertaking communication link maintenance tasks, logical guidance dense rewards are used to encourage them to maintain effective communication connections with neighboring UAVs or temporary cluster leaders. Through a centralized training-distributed execution architecture, each UAV updates its policy using the global task state during the training phase, and performs corresponding tasks based on local observations and neighbor communication information during the execution phase.
[0085] T4: Based on dynamic spectrum clustering and distributed negotiation mechanism, analyze the communication topology of multi-agent system in real time, realize role scheduling and task execution in dynamic environment, and trigger step S2 to rematch when the topology breaks or the task changes.
[0086] In this embodiment, the system maintains a real-time communication topology map between drones and divides the drone cluster into multiple functional subclusters based on the Fiedler vector of the graph Laplacian matrix. Within each subcluster, a comprehensive score is calculated based on the drone's remaining battery power and computing power. The drone with the highest score is elected as the temporary cluster leader, responsible for local task coordination, status aggregation, and resource scheduling within the subcluster. The remaining drones serve as task execution nodes, collaborative sensing nodes, or communication relay nodes according to their own capabilities and task requirements. When a drone is detected to be offline, a communication link is broken, a dangerous area expands, or a new high-priority search and rescue task is added, the system triggers a renegotiation mechanism and re-executes step S2 based on the latest drone status, task set, and communication topology status to achieve real-time updates of tasks and roles.
[0087] Through the above steps, this embodiment can transform constraints such as area arrival, data acquisition, anomaly verification, and danger zone avoidance in UAV swarm collaborative inspection and emergency search and rescue missions into computable logical states, and realize task allocation and role scheduling based on the state characteristics, payload capacity, and communication topology of each UAV. When a UAV goes offline, a communication link breaks, or a new high-priority search and rescue mission is added, the system can trigger a renegotiation mechanism to redefine the UAV's pending tasks and its roles such as temporary cluster head, task execution node, collaborative sensing node, or communication relay node, thereby improving the collaborative execution capability and robustness of the UAV swarm in dynamic post-disaster environments.
[0088] This embodiment also provides a real-time matching system for multi-agent tasks and roles in a dynamic environment. The real-time matching system for multi-agent tasks and roles in a dynamic environment can be implemented by executing the process steps of the real-time matching method for multi-agent tasks and roles in a dynamic environment. That is, those skilled in the art can understand the real-time matching method for multi-agent tasks and roles in a dynamic environment as a preferred implementation of the real-time matching system for multi-agent tasks and roles in a dynamic environment.
[0089] Specifically, a real-time multi-agent task-role matching system in a dynamic environment includes: Module M1: A formal behavioral model building module based on task automata; used to build a product Markov decision process that transforms linear temporal logic task constraints into nondeterministic Butch automata and constructs a product Markov decision process with the environmental Markov decision process, providing the agent with an enhanced state space that includes physical state and logical automaton state.
[0090] Module M2: Game theory matching module based on graph attention mechanism; Based on the current state of each agent in the enhanced state space and the dynamic task set acquired in real time, a dynamic heterogeneous bipartite graph of agents and tasks is constructed. The graph attention network is used to extract features, embed them and quantify the matching benefits. Then, the Nash equilibrium point is solved through the latent game model to output the probability distribution of task selection.
[0091] Module M3: Distributed multi-agent reinforcement learning solution module; used to obtain the optimal task allocation by adopting a centralized training-distributed execution architecture and task course learning based on the probability distribution and the logic-guided dense reward derived from the product Markov decision process, and updating the policy network online until convergence.
[0092] Module M4: Role scheduling module based on dynamic spectrum clustering; based on dynamic spectrum clustering and distributed negotiation mechanism, it analyzes the communication topology of multi-agent system in real time, realizes role scheduling and task execution in dynamic environment, and triggers step S2 to re-match when the topology breaks or the task changes.
[0093] The module M1 includes the following sub-modules: Module M1.1: Uses Linear Temporal Logic (LTL) to describe the temporal constraints and safety specifications of multi-agent cooperative tasks, and transforms the LTL formulas into nondeterministic Büchi automata (NBA), where the Büchi automata is a type of Büchi automaton; Module M1.2: Defines an Environment Markov Decision Process (MDP), which performs a Cartesian product operation on the environment state space and the Büchi automaton state space to construct a product Markov decision process containing the automaton states; Module M1.3: Defines a state transition function in the product Markov decision process so that the agent's state vector contains both physical state and logical automaton state, providing a state space with temporal logic awareness for subsequent reinforcement learning.
[0094] Module M2 includes the following sub-modules: Module M2.1: Models the multi-agent set and dynamic task set as a dynamic heterogeneous bipartite graph. Nodes include agent nodes and task nodes, edges represent potential executable relationships, and physical states and logical automaton states serve as the initial features of agent nodes. Module M2.2: Utilizes a graph attention network (GAT) to extract the ability feature embeddings of the agent and the requirement feature embeddings of the task, and calculates the dot product of the feature vectors as the matching reward value; Module M2.3: Constructs a potential game model, transforming the global payoff maximization problem into finding the Nash equilibrium point of the game, and outputs the probability distribution of the agent's choice of tasks.
[0095] Preferably, module M3 includes the following sub-modules: Module M3.1: Constructs a logic-guided potential energy reward function to provide positive incentives when the agent's action triggers the nondeterministic Büchi automaton to transition to the accepting state, thus solving the sparse reward problem; Module M3.2: Design a course learning mechanism with increasing task complexity, prioritizing training on simple tasks that satisfy atomic linear sequential logic formulas, and gradually transitioning to complex sequential logic tasks; Module M3.3: It adopts a centralized training-distributed execution (CTDE) architecture, uses a gated communication mechanism to filter neighbor information whose importance weight exceeds a preset weight threshold, updates the policy network parameters, until the multi-agent system converges to the cooperative optimal solution.
[0096] Module M4 includes the following sub-modules: Module M4.1: Maintains the communication topology graph of intelligent agents in real time, and calculates the graph Laplacian matrix and its Fiedler vector. Module M4.2: Utilizes Fiedler vectors for fast spectral clustering to divide the agent cluster into several functional subclusters, and dynamically elects roles within the subclusters based on the contract network protocol; Module M4.3: When a topology break or task change is detected, a renegotiation mechanism is triggered. Based on the latest local perception information, role switching and resource reorganization are performed. Step S2 is triggered to rematch based on the latest state in the enhanced state space and the latest acquired dynamic task set, thereby ensuring the self-healing and connectivity of the system in a dynamic environment.
[0097] Based on the above description of the embodiments, those skilled in the art will understand that the multi-agent task-role real-time matching method and system in a dynamic environment described in this embodiment can be implemented in pure software or deployed and run on a general-purpose or dedicated computing hardware platform. Based on this essence, the technical solution of this embodiment can be specifically implemented in the form of a software product containing program instructions. This software product can be stored on various non-volatile storage media or directly deployed as a local or cloud service. The program instructions are used to cause computer devices with processing capabilities—including but not limited to personal computers, server clusters, mobile terminals, or other network devices—to execute the steps described in this embodiment.
[0098] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A method for real-time matching of multi-agent tasks and roles in a dynamic environment, characterized in that, include: S1: The linear temporal logic task constraints are transformed into nondeterministic Butch automata, and a product Markov decision process is constructed with the environmental Markov decision process to provide the agent with an enhanced state space containing physical state and logical automaton state. The linear temporal logic task constraints are used to describe the inspection and search and rescue tasks of the UAV swarm in the disaster-stricken park. S2: Based on the current state of each agent in the enhanced state space and the dynamic task set acquired in real time, a dynamic heterogeneous bipartite graph of agents and tasks is constructed. Feature embedding and matching benefits are extracted using a graph attention network. Then, the Nash equilibrium point is solved through a latent game model, and the probability distribution of task selection is output. The dynamic task set includes target area inspection tasks, image acquisition tasks, abnormal target search tasks, verification and confirmation tasks, and communication link maintenance tasks. Agent nodes correspond to each UAV, and task nodes correspond to the above-mentioned types of tasks to be executed. The characteristics of UAV nodes include current position, remaining power, communication quality, payload type, and computing power. The characteristics of task nodes include task position, task priority, deadline, required payload type, and communication requirements. S3: Based on the probability distribution and the logically guided dense reward derived from the product Markov decision process, a centralized training-distributed execution architecture and task course learning are adopted to update the policy network online until convergence, thereby obtaining the optimal task allocation. For UAVs performing inspection tasks, the logically guided dense reward is used to encourage them to reach the designated inspection area and complete image acquisition. For UAVs performing abnormal target search or verification tasks, the logically guided dense reward is used to encourage them to reach the abnormal target area while avoiding dangerous areas. For UAVs undertaking the task of maintaining communication links, the logically guided dense reward is used to encourage them to maintain an effective communication connection with neighboring UAVs or temporary cluster heads. S4: Based on dynamic spectral clustering and distributed negotiation mechanisms, the communication topology of the multi-agent system is analyzed in real time to realize role scheduling and task execution in dynamic environments. When the topology breaks or the task changes, step S2 is triggered for re-matching. Specifically, the system maintains the communication topology map between drones in real time and divides the drone cluster into multiple functional sub-clusters based on the Fiedler vector of the graph Laplacian matrix. Within each sub-cluster, a comprehensive score is calculated based on the drone's remaining battery power and computing power. The drone with the highest score is elected as the temporary cluster leader, responsible for local task coordination, state aggregation, and resource scheduling within the sub-cluster. The remaining drones serve as task execution nodes, collaborative perception nodes, or communication relay nodes according to their own capabilities and task requirements. When a drone is detected to be offline, the communication link is broken, the dangerous area expands, or a new high-priority search and rescue task is added, the system triggers the renegotiation mechanism and re-executes step S2 based on the latest drone status, task set, and communication topology status to achieve real-time updates of tasks and roles.
2. The method of claim 1, wherein, Step S1 includes: The temporal constraints and safety specifications of multi-agent collaborative tasks are described using linear sequential logic and transformed into nondeterministic Butch automata. The environmental Markov decision process is combined with the Butch automaton to construct a product Markov decision process; In the product Markov decision process, a state transition function is defined such that the agent's state vector includes both physical state and logical automaton state.
3. The method of claim 1, wherein, Step S2 includes: The multi-agent set and the dynamic task set are modeled as a dynamic heterogeneous bipartite graph, where nodes include agent nodes and task nodes, edges represent potential executable relationships, and physical state and logical automaton state serve as the initial features of agent nodes. The ability feature embedding of the agent and the requirement feature embedding of the task are extracted using a graph attention network, and the similarity of the feature vectors is calculated as the matching reward value. Construct a potential game model to transform the problem of maximizing global payoff into finding the Nash equilibrium point of the game and output the probability distribution of the agent's choice of tasks.
4. The method of claim 3, wherein, By using the matching payoff value output by the graph attention network as the input to the game model, the generated probability selection distribution follows a Boltzmann distribution.
5. The method of claim 1, wherein, Step S3 includes: A logic-guided potential energy reward function is constructed to provide positive incentives when the agent's action triggers the spawn automaton to transition to the accepting state, in order to solve the sparse reward problem. Design a course learning mechanism with increasing task complexity, prioritizing training on simple tasks that satisfy linear sequential logic formulas, and gradually transitioning to complex sequential logic tasks; A centralized training-distributed execution network architecture is adopted. A gated communication mechanism is used to filter neighbor information whose importance weight exceeds a preset weight threshold, and update the policy network parameters until the multi-agent system converges to the cooperative optimal solution.
6. The method of claim 5, wherein, The formula for the potential energy reward function is as follows: ; in, In physical state Next action Then, the physical state transitions to Simultaneously, the state of the logical automaton changes from... Transfer to At that time, the system provides additional reward signals. These are the weighting coefficients. As a discount factor, They are time points The state of the logical automaton, They are time points The physical state, To accept the set of states, Rewards for achieving the goal For indicator functions, when belong The value is 1 if it is true, and 0 otherwise.
7. The method according to claim 5, characterized in that, The gating communication mechanism is used to filter neighbor information whose importance weight exceeds a preset weight threshold, specifically as follows: Introducing communication gating units into distributed execution networks, intelligent agents Real-time calculation of the importance weights of the acquired observed features ,when At that time, intelligent agent Broadcast the intent vector encoded by the local policy network to neighbors. The preset weight threshold.
8. The method according to claim 1, characterized in that, Step S4 includes: Maintain the agent communication topology graph in real time, and calculate the graph Laplacian matrix and its Fiedler vector; Fast spectral clustering using Fiedler vectors is used to divide the agent cluster into multiple functional subclusters, and roles are dynamically elected within each subcluster based on the contract network protocol. When a topology break or task change is detected, a renegotiation mechanism is triggered to switch roles and reorganize resources based on the latest local perception information. Step S2 is then triggered to rematch based on the latest state in the enhanced state space and the latest acquired dynamic task set.
9. The method according to claim 8, characterized in that, The dynamic role election within each sub-cluster based on the contract network protocol is specifically as follows: Based on intelligent agents Remaining energy and computing power Calculate the overall score : ; in, These are the weight parameters; The agent with the highest overall score is elected as the temporary cluster leader, while the other agents act as execution nodes, following the local instructions of the cluster leader, forming a hybrid control architecture of global distributed and local hierarchical.
10. A computer system comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the method according to any one of claims 1-9.
Citation Information
Patent Citations
Distributed multi-agent task cooperation method based on linear sequential logic
CN111340348A
Intelligent recommendation decision-making method and system based on graph neural network and cognitive architecture
CN121094131A