A multi-agent cooperation and execution method based on a task graph structure
Patent Information
- Application Number
- CN202611328933.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-31
- Publication Date
- 2026-09-25
AI Technical Summary
但工作流图的边与节点定义需要人工预先设计,无法从历史执行轨迹中自动发现原子操作及其依赖关系;且工作流引擎关注的是确定性流程的执行,缺乏对运行时条件动态判断和失败智能恢复的支持
本发明基于原子任务有向无环图(AT-DAG)对复杂自然语言指令进行结构化建模,通过原子任务提取、图结构重组与多层级任务生成流水线,驱动多个功能专属智能体协同完成具有顺序依赖、并行约束、条件分支和层级嵌套等复杂逻辑结构的复合任务;可广泛应用于企业办公流程自动化、软件开发流程编排、供应链多环节协同调度、科研实验流程管理及政务服务多窗口协同办理等场景;
Smart Images

Figure CN122820148A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence and multi-agent collaboration technology, specifically a multi-agent collaboration and execution method based on a task graph structure. Background Technology
[0002] In recent years, agent technology, centered on large language models (LLM), has made groundbreaking progress. From the execution of instructions by a single agent to the collaborative cooperation of multi-agent systems (MAS), academia and industry are jointly promoting the evolution of agent technology from "being able to execute simple instructions" to "being able to autonomously complete complex tasks".
[0003] Complex tasks in real-world business scenarios generally exhibit the following characteristics: (1) Multidimensionality of structure: A complete business process is often not linearly connected. Taking supply chain management as an example, the seemingly simple task of "supplier evaluation and procurement contract generation" actually includes three parallel sub-tasks: supplier qualification verification, historical delivery data evaluation, and contract template matching. Moreover, the result of "supplier qualification verification" determines whether to follow the "standard contract process" or the "supplementary review process", which naturally contains conditional branching logic.
[0004] (2) Cross-system collaboration: The execution of enterprise-level tasks usually spans multiple independent systems. A complete "customer onboarding" process requires creating a profile in the Customer Relationship Management System (CRM), establishing a financial account in the Enterprise Resource Planning System (ERP), and sending a welcome notification in the email system. These systems each manage independent data spaces, and there are data dependencies between operations (CRM output serves as ERP input), as well as windows that can be executed in parallel (ERP profile creation and email sending can be performed simultaneously).
[0005] (3) Execution fault tolerance: In complex tasks involving dozens or even hundreds of steps, failure of a certain step is the norm rather than the exception. Current mainstream solutions usually adopt a full retry strategy when encountering failure, resulting in the loss of all completed work from the preceding steps. Enterprise-level scenarios require the system to have local fault tolerance capabilities of "where it fails, it recovers from".
[0006] (4) Explainability requirement: In fields with stringent compliance requirements such as finance, healthcare, and government, every decision made by an intelligent agent needs to be traceable and auditable. However, existing intelligent agent systems generally lack the ability to record process steps and output the rationale for decisions, which fails to meet regulatory and auditing requirements.
[0007] The current technological roadmap for multi-agent systems is mainly evolving along two directions: One approach is an LLM-based conversational multi-agent framework: This type of framework achieves collaboration between agents through predefined agent roles (such as "planner," "executor," and "reviewer") and conversational messaging mechanisms. However, its subtask division relies on the predefined boundaries of roles rather than the inherent structure of the tasks themselves. This means that "what a role can do" determines "how the task is decomposed," rather than "what the task needs" determining "how the agents cooperate." When a task contains parallel subtasks, the conversational framework naturally serializes them into dialogue rounds, making it impossible to utilize parallel execution windows.
[0008] Option 2: Task orchestration framework based on workflow engines. This type of framework defines task execution paths through directed graphs, naturally supporting branching and parallelism. However, the definition of edges and nodes in the workflow graph needs to be pre-designed manually, and it cannot automatically discover atomic operations and their dependencies from historical execution trajectories; moreover, workflow engines focus on the execution of deterministic processes and lack support for dynamic judgment of runtime conditions and intelligent recovery from failures.
[0009] There is a clear gap between the two directions mentioned above, namely, the lack of a unified method framework that can automatically extract atomic operations from historical execution trajectories and construct dependency graphs, dynamically generate multi-level complex tasks based on graph structures, and inject graph structure information into multi-agent systems during execution to achieve topology-aware scheduling. This framework would systematically address the four core defects in existing technologies: coarse task modeling granularity, weak support for conditional branches, non-reusable atomic operations, and agents' lack of awareness of subtask structure boundaries. Summary of the Invention
[0010] This invention addresses the needs and shortcomings of current technological development by providing a multi-agent collaboration and execution method based on a task graph structure. This method can automatically extract atomic operations from historical execution trajectories and construct dependency graphs, dynamically generate multi-level complex tasks based on graph structures, and inject graph structure information into the multi-agent system during execution to achieve topology-aware scheduling.
[0011] The present invention provides a multi-agent collaboration and execution method based on a task graph structure, and the technical solution adopted to solve the above-mentioned technical problems is as follows: A multi-agent cooperation and execution method based on a task graph structure is implemented in the following two stages: Phase 1: Offline graph construction and complex task instruction generation; The execution trajectory of the target task system is collected through automated execution engine or manual operation records. The hierarchical feature hashing scheme is used to identify the task execution status and construct a state transition graph containing nodes and directed edges. Extract an atomic task quadruple containing the operation type, target object description, pre-state condition, and post-state effect from each directed edge in the state transition graph. Construct a directed acyclic graph of atomic tasks, using atomic tasks as nodes and the dependencies between operations as directed edges. Perform multidimensional structure parsing and natural language instruction synthesis on the directed acyclic graph of the atomic task, including: performing a systematic structure scan on the directed acyclic graph of the atomic task to identify linear path segments, identify concurrent branching points and merging points, and label conditional predicate nodes to generate a structure labeling table; recursively combining the structural elements in the structure labeling table into composite task instructions expressed in natural language according to the topological order of the directed acyclic graph of the atomic task; performing semantic consistency verification and path reachability verification on the composite task instructions; and writing them into the target task instruction set after passing the verification. Phase Two: Online Multi-Agent Collaborative Execution; The system distributes complex task instructions from the target task instruction set to a multi-agent system. This system includes a task parsing agent, an execution agent, an evaluation agent, and a coordination agent. The task parsing agent maps complex task instructions to corresponding atomic task directed acyclic graphs (DAGs). The coordination agent schedules different execution agents based on the topological order and inter-node dependencies of the atomic task DAGs. The evaluation agent performs operation quality assessments and can trigger error correction. This multi-agent system relies on the coordination agent to achieve concurrent execution of parallel tasks and local retries of failed nodes, without altering the execution results of completed predecessor nodes.
[0012] Optionally, the task execution status includes structural features and dynamic features; among which, structural features are stable features reflecting the task execution stage, including task type identifier, current stage label and key resource identifier; dynamic features are instantaneous features that change with the execution instance, including timestamp and temporary data content; When using a hierarchical feature hashing scheme to identify the task execution status, complete values are extracted for structural features to participate in hash calculations, while only category labels are extracted for dynamic features, ignoring specific values. This generates a unique identifier for the task status that is robust to dynamic changes, thus solving the problem of state space explosion caused by dynamic data.
[0013] Alternatively, the constructed state transition graph can be represented as: G=(V,E), where node V represents an intermediate state during task execution, and directed edge E represents a state transition triggered by an operation.
[0014] Optionally, in the atomic task quadruple, the operation type includes invocation, query, computation, writing, verification, and transformation; the target object is described as the unique identifier and attribute of the resource or module targeted by the operation; the preconditions are the task state constraints required to execute the operation; and the postconditions are the description of the changes in the task state after the operation is executed.
[0015] Optionally, the dependencies between operations are categorized into four types: sequential dependency, parallel independence, conditional mutual exclusion, and hierarchical containment; among which: Sequential dependency means that the preconditions of operation B are satisfied by the postconditions of operation A. Construct a directed edge from operation A to operation B and set the weight of the directed edge to 1 to represent that operation A and operation B are executed in a strict sequential order. Parallel independence means that the preconditions of operation A and operation B are independent of each other, and the postconditions do not affect each other. No directed edge is built between operation A and operation B, which represents that operation A and operation B are concurrently executable. Conditional mutual exclusion means that both operation B1 and operation B2 are preceded by operation A, but the triggering conditions of operation B1 and operation B2 are mutually exclusive (e.g., B1 is executed when the data volume exceeds the threshold, otherwise B2 is executed). The forking edges in AT-DAG are marked with conditional predicates. The hierarchy refers to the logical subtasks that are formed by the aggregation of the operation sequence C1→C2→C3 composed of atomic tasks. The local graph corresponding to the operation sequence is used as a subgraph, and it is transformed into a single composite node on the directed acyclic graph of atomic tasks through the subgraph shrinking method, which represents the logical subtask as an execution step in the higher-level operation.
[0016] Optionally, a systematic structural scan is performed on the directed acyclic graph of the atomic task to identify linear path segments, concurrent bifurcation points and merging points, and conditional predicate nodes to generate a structural annotation table. This process specifically includes: (1) Identify linear path segments: Perform a topological traversal on the directed acyclic graph of the atomic task, detect all consecutive node sequences that satisfy the condition that "the current node has an out-degree of 1 and its direct successor has an in-degree of 1", and mark them as linear execution segments. Each linear execution segment represents a set of atomic operation chains that can be described as a single target subprocess; perform semantic aggregation on the operation type and target object description of each node in the linear execution segment, and generate a structural summary of the linear execution segment according to the "action-object-result" template; (2) Identify concurrent branching points and merging points: Scan the directed acyclic graph of the atomic task, mark all nodes that satisfy "out-degree ≥ 2 and no directed path connection between successors" as concurrent branching points, and mark all nodes that satisfy "in-degree ≥ 2 and each predecessor node comes from different branches of the same branching point" as concurrent merging points; multiple non-intersecting paths between concurrent branching points and concurrent merging points constitute parallel segment groups. The preconditions of each path in the parallel segment group are independent of each other, and the postconditions are respectively merged into the merged input of the concurrent merging point; (3) Labeling condition predicate nodes: Traverse all outgoing edges of the directed acyclic graph of the atomic task, detect nodes that satisfy "multiple outgoing edges of the same source node carry mutually exclusive triggering conditions", and mark them as condition predicate nodes; extract the predicate description set of each condition predicate node, where the source of the condition predicate is divided into three categories: a) state attribute threshold comparison, which is directly read from the attribute of the task context object; b) external query result matching, which is obtained by matching the current business data through interface calls; c) compound logical operation, which is obtained by combining multiple sub-conditions with AND / OR / NOT; label the corresponding activation condition predicate and target path segment for each branch outgoing edge; for compound nodes under the hierarchical containment structure, expand its internal subgraph structure, and recursively perform the processing of identifying linear path segments, identifying concurrent branching points and merging points, and labeling condition predicate nodes on the internal subgraph; After completing the above operations, output the structure annotation table corresponding to the directed acyclic graph of the atomic task. Each row in the structure annotation table corresponds to a structural element, recording the semantic summary, activation condition and subsequent dependencies of the structural element.
[0017] Further, optionally, the composite task instruction includes four logical structures: sequential constraint expression, concurrent constraint expression, conditional branch expression, and nested sub-procedure expression, wherein: Sequential constraint expression refers to the use of the sentence structure "After completing [the summary of linear execution segment 1], execute [the summary of linear execution segment 2]" to connect two linear execution segments that have data dependencies. Concurrency constraint expression refers to the combined expression of each parallel path within a parallel segment group using the sentence structure "execute [path 1 summary] and [path 2 summary] simultaneously, and execute [concurrent merging point successor summary] after both are completed"; Conditional branching refers to the expression of multiple branches of a conditional predicate node using the sentence structure "Execute [branch path A summary] when [predicate description] is true, otherwise execute [branch path B summary]". Nested sub-procedure expression refers to the use of a sentence structure like "while executing [macro task objective], repeatedly execute [inner subgraph summary] until [termination condition] is met" to embed the composition instruction of the inner subgraph into a composite node under a hierarchical containment structure.
[0018] Further, optionally, semantic consistency verification and path reachability verification are performed on the composite task instructions. This verification process specifically includes: Check the source of all conditional predicates in the composite task instruction to confirm that they have corresponding state attributes or external interface support in the structure annotation table, and ensure that each condition has a basis for judgment. Based on the dependencies recorded in the structure annotation table, a graph path coverage check is performed on the composite task instructions to confirm that each sub-process step contained in the composite task instructions has a reachable execution path in the atomic task directed acyclic graph. Composite task instructions that pass both types of verification are written into the target task instruction set; composite task instructions that fail at least one type of verification trigger a local reorganization, re-executing the combinational logic of the composite task instructions only on the failed structural segments until verification is passed, or marked as deprecated after reaching the maximum number of retries.
[0019] Alternatively, after distributing complex task instructions from the target task instruction set to the multi-agent system, the multi-agent system completes the entire process in a fixed time sequence: First, the task parsing agent performs reverse mapping on the input complex task instructions to generate the corresponding atomic task directed acyclic graph. At the same time, it appends the atomic task directed acyclic graph structure information of the task nodes to the corresponding task instructions in structured JSON format. Secondly, the coordinating agent constructs and maintains the execution state graph of the directed acyclic graph of atomic tasks, and uniformly marks all task nodes as pending execution, in execution, completed or failed states; Furthermore, the coordinating agent dynamically completes task scheduling and allocation for the executing agents based on the topological order and node dependencies of the directed acyclic graph of atomic tasks: a) Subtask nodes without predecessor dependencies are directly marked as schedulable and allocated to idle executing agents for concurrent execution; b) For subtask nodes with predecessor dependencies, the node is sent to the scheduling queue to wait for execution only after all its predecessor nodes have been updated to complete execution status; c) For conditional branch decision nodes, the evaluation agent first reads the current task execution context, calculates the truth value of the conditional predicate, and then selects subtask nodes matching the branch based on the truth value results to complete the distribution and scheduling. Finally, the executing agent receives the task quadruple information and performs the corresponding task operation. After a single execution, it automatically compares the actual post-state effect with the expected post-state effect. Simultaneously, the evaluation agent conducts an operation quality evaluation from two dimensions: task progress and decision rationality. If the evaluation result does not reach the preset threshold, the metacognitive error correction process is automatically triggered to complete the closed-loop verification of task execution.
[0020] Alternatively, during the process of completing the entire process in a fixed time sequence, the multi-agent system synchronously generates step-level process supervision data and records the following four-tuple data for each operation: Observation: A structured representation of the context snapshot of the current task execution state; Analysis: The text describing the reasoning process of the agent describes the basis for locating the target and the reasons for choosing the operation. Action: The type of operation to be performed, the operation parameters, and the target object; Objective: To determine which atomic task node in the directed acyclic graph corresponds to this operation, and to explain the role of this operation in advancing the overall task objective.
[0021] The multi-agent cooperation and execution method based on task graph structure of the present invention has the following advantages compared with the prior art: This invention uses atomic task directed acyclic graphs (AT-DAG) to structurally model complex natural language instructions. Through atomic task extraction, graph structure reorganization, and a multi-level task generation pipeline, it drives multiple functionally dedicated intelligent agents to collaboratively complete composite tasks with complex logical structures such as sequential dependencies, parallel constraints, conditional branches, and hierarchical nesting. It can be widely applied to scenarios such as enterprise office process automation, software development process orchestration, multi-link collaborative scheduling of supply chains, scientific research experimental process management, and multi-window collaborative processing of government services. This invention can automatically extract atomic operations from historical execution trajectories and construct dependency graphs, dynamically generate multi-level complex tasks based on graph structures, and inject graph structure information into multi-agent systems during execution to achieve topology-aware scheduling. It also systematically solves the four core defects of existing technologies: coarse task modeling granularity, weak support for conditional branches, non-reusable atomic operations, and agents' lack of awareness of subtask structure boundaries. Attached Figure Description
[0022] Appendix Figure 1 This is a flowchart of stage one of the methods described in this invention; Appendix Figure 2 This is a flowchart of stage two in the method described in this invention. Detailed Implementation
[0023] To make the technical solution, the technical problem solved, and the technical effect of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with specific embodiments.
[0024] Example 1: Refer to Appendix Figure 1 , 2 This embodiment proposes a multi-agent collaboration and execution method based on a task graph structure, which includes the following two stages.
[0025] Phase 1: Offline graph construction and complex task instruction generation.
[0026] (1.1) The execution trajectory of the target task system is collected through automated execution engine or manual operation record, and the task execution status is identified by hierarchical feature hashing scheme. A state transition diagram containing nodes and directed edges is constructed.
[0027] The task execution status includes structural features and dynamic features. Structural features are stable characteristics reflecting the task execution stage, including task type identifiers, current stage labels, and key resource identifiers. Dynamic features are transient characteristics that change with the execution instance, including timestamps and temporary data content. When using a hierarchical feature hashing scheme to identify the task execution status, complete values are extracted from structural features for hash calculation, while only category labels are extracted from dynamic features, ignoring specific values. This generates a unique task status identifier that is robust to dynamic changes, thus solving the state space explosion problem caused by dynamic data.
[0028] The constructed state transition graph is represented as: G=(V,E), where node V represents an intermediate state during task execution, and directed edge E represents a state transition triggered by an operation.
[0029] (1.2) Extract an atomic task quadruple containing the operation type (ActionType), target object description (TargetObject), precondition (PreCondition), and posteffect (PostEffect) for each directed edge (i.e., single-step operation) in the state transition graph; where: ActionType includes invocation, query, calculation, writing, verification, and transformation; TargetObject is the unique identifier and attribute of the resource or module targeted by the operation; PreCondition is the task state constraint required to execute the operation; PostEffect is the description of the change in task state after the operation is executed.
[0030] (1.3) Construct an Atomic-Task Directed Acyclic Graph (AT-DAG) with atomic tasks as nodes and the dependencies between operations as directed edges.
[0031] Dependencies between operations are categorized into four types: sequential dependency, parallel independence, conditional mutual exclusion, and hierarchical containment; among which: Sequential dependency means that the preconditions of operation B are satisfied by the postconditions of operation A. Construct a directed edge from operation A to operation B (A→B) and set the weight of the directed edge to 1 to represent that operation A and operation B are executed in a strict order. Parallel independence means that the preconditions of operation A and operation B are independent of each other and their post-conditions do not affect each other. No directed edge is built between operation A and operation B, which represents that operation A and operation B are concurrently executable. Conditional mutual exclusion means that both operation B1 and operation B2 are preceded by operation A, but the triggering conditions of operation B1 and operation B2 are mutually exclusive (e.g., B1 is executed when the data volume exceeds the threshold, otherwise B2 is executed). The fork edges in AT-DAG are marked with conditional predicates. Nested refers to the aggregation of an operation sequence C1→C2→C3 composed of atomic tasks into a logical subtask belonging to a higher-level operation. The local graph corresponding to this operation sequence is taken as a subgraph and transformed into a single composite node on the atomic task directed acyclic graph (AT-DAG) through subgraph shrinking. This represents the logical subtask as an execution step in the higher-level operation.
[0032] (1.4) Perform multidimensional structure parsing and natural language instruction synthesis on the atomic task directed acyclic graph (AT-DAG), including: performing a systematic structure scan on the atomic task directed acyclic graph (AT-DAG) to identify linear path segments, identify concurrent branching points and merging points, and label conditional predicate nodes to generate a structure annotation table (SAT); according to the topological order of the atomic task directed acyclic graph (AT-DAG), recursively combine the structural elements in the structure annotation table (SAT) into composite task instructions expressed in natural language; perform semantic consistency verification and path reachability verification on the composite task instructions; and write them into the target task instruction set after passing the verification.
[0033] The process of performing multidimensional structure parsing and natural language instruction synthesis on the atomic task directed acyclic graph (AT-DAG) is a closely linked process. The process takes the constructed AT-DAG as input and finally outputs a verified set of multi-structured complex natural language instructions.
[0034] A systematic structural scan is performed on the directed acyclic graph (AT-DAG) of the atomic task to identify linear path segments, concurrent bifurcation points and merging points, and conditional predicate nodes, generating a structural annotation table (SAT). This process specifically includes: (1) Identify linear path segments: Perform topological traversal on the atomic task directed acyclic graph (AT-DAG) and detect all consecutive node sequences that satisfy the condition that "the current node has an out-degree of 1 and its direct successor has an in-degree of 1". Mark them as linear execution segments (LS). Each linear execution segment (LS) represents a chain of atomic operations that can be described as a single target subprocess. Semantically aggregate the operation type and target object description of each node in the linear execution segment (LS) and generate a structural summary of the linear execution segment (LS) according to the "action-object-result" template. (2) Identify concurrent fork nodes and join nodes: Scan the atomic task directed acyclic graph (AT-DAG), mark all nodes that satisfy "out-degree ≥ 2 and no directed path connection between successors" as concurrent fork nodes (FN), and mark all nodes that satisfy "in-degree ≥ 2 and each predecessor node comes from a different branch of the same fork node" as concurrent join nodes (JN); multiple non-intersecting paths between concurrent fork nodes (FN) and concurrent join nodes (JN) constitute parallel segment groups (PSG). The preconditions of each path in the parallel segment group (PSG) are independent of each other, and the postconditions are respectively merged into the merged input of the concurrent join node (JN); (3) Labeling condition predicate nodes: Traverse all outgoing edges of the atomic task directed acyclic graph (AT-DAG), detect nodes that satisfy "multiple outgoing edges of the same source node carry mutually exclusive triggering conditions", and label them as condition predicate nodes (PN); extract the predicate description set of each condition predicate node (PN), where the source of condition predicates is divided into three categories: a) state attribute threshold comparison, which is directly read from the attribute of the task context object; b) external query result matching, which is obtained by interface call to obtain the current business data for matching; c) compound logical operation, which is obtained by combining multiple sub-conditions with AND / OR / NOT; label the corresponding activation condition predicate and target path segment for each branch outgoing edge; for compound nodes under the hierarchical containment structure, expand its internal subgraph structure, and recursively perform the processing of identifying linear path segments, identifying concurrent branching points and merging points, and labeling condition predicate nodes on the internal subgraph; After completing the above operations, the Structural Labeling Table (SAT) corresponding to the Directed Acyclic Graph (AT-DAG) of the atomic task is output. Each row in the SAT corresponds to a structural element (LS / PSG / PN), which records the semantic summary, activation condition and subsequent dependencies of the structural element.
[0035] Using the Structure Annotation Table (SAT) as input, and based on the topological order of the atomic task directed acyclic graph (AT-DAG), when recursively combining the structural elements into a composite task instruction expressed in natural language, the composite task instruction contains four logical structures: sequential constraint expression, concurrent constraint expression, conditional branch expression, and nested sub-procedure expression. Sequential constraint expression refers to the use of the sentence structure "After completing [the summary of LS1], execute [the summary of LS2]" to connect two linear execution segments (LS) that have data dependencies (the latter's preconditions depend on the former's postconditions). Concurrency constraint expression refers to the combined expression of each parallel path within a parallel segment group (PSG) using the sentence structure "execute [path 1 summary] and [path 2 summary] simultaneously, and execute [concurrent merge point (JN) successor summary] after both are completed"; Conditional branching refers to the expression of multiple branches of a conditional predicate node (PN) using the sentence structure "execute [branch path A summary] when [predicate description] is true, otherwise execute [branch path B summary]". Nested sub-procedure expression refers to the use of a sentence structure like "while executing [macro task objective], repeatedly execute [inner subgraph summary] until [termination condition] is met" to embed the composition instruction of the inner subgraph into a composite node under a hierarchical containment structure.
[0036] This step performs semantic consistency verification and path reachability verification on the composite task instructions. The verification process specifically includes: Check the source of all conditional predicates in the composite task instruction to confirm that they have corresponding state attributes or external interface support in the Structure Label Table (SAT) to ensure that each condition has a basis for judgment. Based on the dependencies recorded in the Structure Annotation Table (SAT), a graph path coverage check is performed on the composite task instructions to confirm that each sub-process step contained in the composite task instructions has a reachable execution path in the atomic task directed acyclic graph (AT-DAG). Composite task instructions that pass both types of verification are written into the target task instruction set; composite task instructions that fail at least one type of verification trigger a local reorganization, re-executing the combinational logic of the composite task instructions only on the failed structural segments until verification is passed, or marked as deprecated after reaching the maximum number of retries.
[0037] Phase Two: Online Multi-Agent Collaborative Execution.
[0038] The system distributes complex task instructions from the target task instruction set to a multi-agent system. This system includes a Task Parser Agent, an Executor Agent, an Evaluator Agent, and a Coordinator Agent. The Task Parser Agent maps complex task instructions to corresponding atomic task directed acyclic graphs (AT-DAGs). The Coordinator Agent schedules different Executor Agents based on the topological order and inter-node dependencies of the AT-DAG. The Evaluator Agent performs operation quality assessment and can trigger error correction. This multi-agent system relies on the Coordinator Agent to handle concurrent execution conflicts of parallel tasks and executes local retry logic for failed nodes. The scope of local retry is limited to the failed node and its successor nodes, without affecting the execution results of successfully completed predecessor nodes.
[0039] Specifically, multi-agent systems complete the entire process in a fixed time sequence: First, the task parsing agent performs reverse mapping on the input complex task instructions to generate the corresponding atomic task directed acyclic graph (AT-DAG). At the same time, it appends the atomic task directed acyclic graph structure information of the task nodes to the corresponding task instructions in structured JSON format. Secondly, the coordinating agent constructs and maintains the execution state graph of the atomic task directed acyclic graph (AT-DAG), and marks all task nodes uniformly as pending execution, in execution, completed or failed states; Furthermore, the coordinating agent dynamically completes task scheduling and allocation for the executing agents based on the topological order and node dependencies of the atomic task directed acyclic graph (AT-DAG): a) Subtask nodes without predecessor dependencies are directly marked as schedulable and allocated to idle executing agents for concurrent execution; b) For subtask nodes with predecessor dependencies, the node is sent to the scheduling queue to wait for execution only after all its predecessor nodes have been updated to complete execution status; c) For conditional branch decision nodes, the evaluation agent first reads the current task execution context, calculates the truth value of the conditional predicate, and then selects subtask nodes matching the branch based on the truth value results to complete the distribution and scheduling. Finally, the executing agent receives the task quadruple information and performs the corresponding task operation. After a single execution, it automatically compares the actual post-state effect with the expected post-state effect. Simultaneously, the evaluation agent conducts an operation quality evaluation from two dimensions: task progress and decision rationality. If the evaluation result does not reach the preset threshold, the metacognitive error correction process is automatically triggered to complete the closed-loop verification of task execution.
[0040] Specifically, the Task Parser Agent uses a large language model with general natural language understanding, accurate user task intent recognition, fine-grained semantic parsing, unstructured text structured information extraction, and retrieval augmented generation (RAG) capabilities. It also supports highly consistent and strongly constrained JSON structured output to ensure the standardization and machine-parsable transformation of natural language tasks, such as OpenAIGPT-4o. The Task Parser Agent supports establishing a semantic mapping index with a pre-defined atomic task directed acyclic graph (AT-DAG) database offline. Online, it performs nearest neighbor retrieval, semantically aligning instruction fragments with atomic task entries in the database, with a matching success rate exceeding 85%.
[0041] The ExecutorAgent utilizes a native large language model with strong instruction compliance, standard toolcalling, functioncalling, standardized API call parameter generation, constrained structured data output, and end-to-end task execution result parsing capabilities. This meets the core requirements of automated agent scheduling, tool orchestration, and closed-loop task execution, such as Anthropic Claude 3.5 Sonnet. The ExecutorAgent supports receiving a subtask quadruple containing the operation type, target object description, and preconditions. It then calls the corresponding functional module or API to execute the operation sequence corresponding to a specific subtask node. After each step, the agent records the post-execution state effect and compares it with the expected post-execution effect. If they are inconsistent, a metacognitive error correction process is triggered.
[0042] The Evaluator Agent is preferably a general-purpose large language model with full-chain capabilities including long context understanding, refined semantic matching, multi-dimensional state comparison, rigorous logical reasoning, standardized result evaluation, and structured text evaluation generation. It can stably support the entire process of autonomous evaluation, result verification, and evaluation output, such as DeepSeek-V3. The Evaluator Agent supports an adjacent state chain-based evaluation strategy, extracting the task execution states from the current time step (t) and the previous time step (t-1) to construct an adjacent state chain. It outputs an operation quality evaluation score from two dimensions: task progress (the proportion of currently completed sub-task nodes to the total number of nodes) and decision rationality (the semantic similarity between the current operation and the actions labeled in the directed acyclic graph (AT-DAG) of atomic tasks). When the evaluation score is below a threshold of 0.6, an error correction signal is sent to the Executor Agent, along with deviation analysis text.
[0043] The Coordinator Agent employs a high-order large language model capable of complex task global planning, hierarchical task decomposition, task dependency resolution, long-chain multi-step logical reasoning, multi-tool combination invocation, and multi-agent collaborative orchestration and scheduling to meet the core requirements of multi-agent system overall scheduling, such as Google Gemini-1.5-Pro. The Coordinator Agent supports maintaining the execution state graph of the atomic task directed acyclic graph (AT-DAG) (each node is marked as pending / in execution / completed / failed), dynamically scheduling the task allocation of the Executor Agents based on the state graph, handling concurrent execution conflicts of parallel tasks, and handling local retry logic for failed nodes; the scope of local retries is limited to the failed node and its successor nodes, without affecting the execution results of successfully completed predecessor nodes.
[0044] During the process of a multi-agent system completing the entire process in a fixed time sequence, step-level process supervision data is generated synchronously, and the following four-tuple data is recorded for each step: Observation: A structured representation of the context snapshot of the current task execution state; Analysis: The text describing the reasoning process of the agent, which explains the basis for locating the target and the reasons for choosing the operation. Action: The type of operation to be performed, the operation parameters, and the target object; Purpose: This operation corresponds to which atomic task node in the directed acyclic graph (AT-DAG) of the atomic task, and explains the role of this operation in advancing the overall task objective.
[0045] The above quadruple data is used to construct a step-level process supervision dataset, providing fine-grained supervision signals for the interpretable training of subsequent agents.
[0046] Based on the method described in this embodiment, complex natural language instructions are structurally modeled using atomic task directed acyclic graphs (AT-DAG). Through atomic task extraction, graph structure reorganization, and multi-level task generation pipelines, multiple functionally dedicated intelligent agents are driven to collaboratively complete composite tasks with complex logical structures such as sequential dependencies, parallel constraints, conditional branches, and hierarchical nesting. This method can be widely applied to scenarios such as enterprise office process automation, software development process orchestration, multi-link collaborative scheduling of supply chains, scientific research experimental process management, and multi-window collaborative processing of government services.
[0047] In addition, based on the method described in this embodiment, its local retry mechanism improves the completion rate of complex tasks: when a subtask node fails during the execution of a complex task, the coordinating agent only retryes the failed node and its successor subgraph, and the execution results of the completed predecessor node are retained. In complex multi-step task scenarios, this mechanism improves the overall task completion rate by more than 10 percentage points compared with the full restart strategy. Based on the method described in this embodiment, the concurrent execution of parallel subtasks improves overall efficiency: compared with the fully serial execution method, the atomic task directed acyclic graph (AT-DAG) aware concurrent scheduling in complex tasks containing more than two parallel independent subtasks shows that the reduction in total task completion time is positively correlated with the proportion of parallel tasks; in a typical office automation scenario containing 30% parallel subtasks (such as "querying order data in database A while calling an external logistics interface to obtain delivery status, and then generating a comprehensive report"), the overall task time is shortened by about 25% compared with serial execution.
[0048] The above specific examples illustrate the principles and implementation methods of the present invention in detail. These embodiments are merely for the purpose of helping to understand the core technical content of the present invention. Based on the above specific embodiments of the present invention, any improvements and modifications made to the present invention by those skilled in the art without departing from the principles of the present invention should fall within the patent protection scope of the present invention.
Claims
1. A multi-agent cooperation and execution method based on a task graph structure, characterized in that, Its implementation includes the following two stages: Phase 1: Offline graph construction and complex task instruction generation; The execution trajectory of the target task system is collected through automated execution engine or manual operation records. The hierarchical feature hashing scheme is used to identify the task execution status and construct a state transition graph containing nodes and directed edges. Extract an atomic task quadruple containing the operation type, target object description, pre-state condition, and post-state effect from each directed edge in the state transition graph. Construct a directed acyclic graph of atomic tasks, using atomic tasks as nodes and the dependencies between operations as directed edges. Perform multidimensional structure parsing and natural language instruction synthesis on the directed acyclic graph of the atomic task, including: performing a systematic structure scan on the directed acyclic graph of the atomic task to identify linear path segments, identify concurrent branching points and merging points, and label conditional predicate nodes to generate a structure labeling table; recursively combining the structural elements in the structure labeling table into composite task instructions expressed in natural language according to the topological order of the directed acyclic graph of the atomic task; performing semantic consistency verification and path reachability verification on the composite task instructions; and writing them into the target task instruction set after passing the verification. Phase Two: Online Multi-Agent Collaborative Execution; The system distributes complex task instructions from the target task instruction set to a multi-agent system. This system includes a task parsing agent, an execution agent, an evaluation agent, and a coordination agent. The task parsing agent maps complex task instructions to corresponding atomic task directed acyclic graphs (DAGs). The coordination agent schedules different execution agents based on the topological order and inter-node dependencies of the atomic task DAGs. The evaluation agent performs operation quality assessments and can trigger error correction. This multi-agent system relies on the coordination agent to achieve concurrent execution of parallel tasks and local retries of failed nodes, without altering the execution results of completed predecessor nodes.
2. The multi-agent collaboration and execution method based on a task graph structure according to claim 1, characterized in that, The task execution status includes structural features and dynamic features; among them, structural features are stable features reflecting the task execution stage, including task type identifier, current stage label, and key resource identifier; dynamic features are instantaneous features that change with the execution instance, including timestamps and temporary data content. When using a hierarchical feature hashing scheme to identify the task execution status, complete values are extracted for structural features to participate in hash calculations, while only category labels are extracted for dynamic features, ignoring specific values. This generates a unique identifier for the task status that is robust to dynamic changes, thus solving the problem of state space explosion caused by dynamic data.
3. The multi-agent collaboration and execution method based on a task graph structure according to claim 2, characterized in that, The constructed state transition graph is represented as: G=(V,E), where node V represents an intermediate state during task execution, and directed edge E represents a state transition triggered by an operation.
4. The multi-agent collaboration and execution method based on a task graph structure according to claim 1, characterized in that, In an atomic task quadruple, the operation types include invocation, query, computation, writing, verification, and transformation; the target object is described as the unique identifier and attribute of the resource or module targeted by the operation; the preconditions are the task state constraints required to execute the operation; and the postconditions are the description of the changes in the task state after the operation is executed.
5. The multi-agent cooperation and execution method based on a task graph structure according to claim 1, characterized in that, Dependencies between operations are categorized into four types: sequential dependency, parallel independence, conditional mutual exclusion, and hierarchical containment; among which: Sequential dependency means that the preconditions of operation B are satisfied by the postconditions of operation A. Construct a directed edge from operation A to operation B and set the weight of the directed edge to 1 to represent that operation A and operation B are executed in a strict sequential order. Parallel independence means that the preconditions of operation A and operation B are independent of each other, and the post-conditions do not affect each other. No directed edge is built between operation A and operation B, which represents that operation A and operation B are concurrently executable. Conditional mutual exclusion means that both operations B1 and B2 require operation A as their predecessor, but the triggering conditions of operations B1 and B2 are mutually exclusive. The forking edges in the AT-DAG are labeled with conditional predicates. The hierarchy refers to the aggregation of operation sequences composed of atomic tasks into logical subtasks belonging to higher-level operations. The local graph corresponding to the operation sequence is taken as a subgraph and transformed into a single composite node on the directed acyclic graph of atomic tasks through subgraph shrinking. This represents the logical subtask as an execution step in the higher-level operation.
6. The multi-agent cooperation and execution method based on a task graph structure according to claim 5, characterized in that, A systematic structural scan is performed on the directed acyclic graph of the atomic task to identify linear path segments, concurrent bifurcation and merging points, and conditional predicate nodes, generating a structural annotation table. This process specifically includes: (1) Identify linear path segments: Perform topological traversal on the directed acyclic graph of the atomic task, detect all consecutive node sequences that satisfy the condition "the current node has an out-degree of 1 and its direct successor has an in-degree of 1", and mark them as linear execution segments. Each linear execution segment represents a set of atomic operation chains that can be described as a single target subprocess; perform semantic aggregation on the operation type and target object description of each node in the linear execution segment, and generate a structural summary of the linear execution segment according to the "action-object-result" template; (2) Identify concurrent branching points and merging points: Scan the directed acyclic graph of the atomic task, mark all nodes that satisfy "out-degree ≥ 2 and no directed path connection between successors" as concurrent branching points, and mark all nodes that satisfy "in-degree ≥ 2 and each predecessor node comes from a different branch of the same branching point" as concurrent merging points; multiple non-intersecting paths between concurrent branching points and concurrent merging points constitute parallel segment groups. The preconditions of each path in the parallel segment group are independent of each other, and the postconditions are respectively merged into the merged input of the concurrent merging point; (3) Labeling condition predicate nodes: Traverse all outgoing edges of the directed acyclic graph of the atomic task, detect nodes that satisfy "multiple outgoing edges of the same source node carry mutually exclusive triggering conditions", and mark them as condition predicate nodes; extract the predicate description set of each condition predicate node, where the source of the condition predicate is divided into three categories: a) state attribute threshold comparison, which is directly read from the attribute of the task context object; b) external query result matching, which is obtained by interface call to obtain the current business data for matching; c) compound logical operation, which is obtained by combining multiple sub-conditions with AND / OR / NOT; label the corresponding activation condition predicate and target path segment for each branch outgoing edge; for compound nodes under the hierarchical containment structure, expand its internal subgraph structure, and recursively perform the processing of identifying linear path segments, identifying concurrent branching points and merging points, and labeling condition predicate nodes on the internal subgraph; After completing the above operations, output the structure annotation table corresponding to the directed acyclic graph of the atomic task. Each row in the structure annotation table corresponds to a structural element, recording the semantic summary, activation condition and subsequent dependencies of the structural element.
7. A multi-agent cooperation and execution method based on a task graph structure according to claim 6, characterized in that, Composite task instructions include four logical structures: sequential constraint expression, concurrent constraint expression, conditional branch expression, and nested sub-procedure expression. Sequential constraint expression refers to the use of the sentence structure "After completing [the summary of linear execution segment 1], execute [the summary of linear execution segment 2]" to connect two linear execution segments that have data dependencies. Concurrency constraint expression refers to the combined expression of each parallel path within a parallel segment group using the sentence structure "execute [path 1 summary] and [path 2 summary] simultaneously, and execute [concurrent merging point successor summary] after both are completed"; Conditional branching refers to the expression of multiple branches of a conditional predicate node using the sentence structure "Execute [branch path A summary] when [predicate description] is true, otherwise execute [branch path B summary]". Nested sub-procedure expression refers to the use of a sentence structure like "while executing [macro task objective], repeatedly execute [inner subgraph summary] until [termination condition] is met" to embed the composition instruction of the inner subgraph into a composite node under a hierarchical containment structure.
8. The multi-agent cooperation and execution method based on a task graph structure according to claim 6, characterized in that, Semantic consistency verification and path reachability verification are performed on composite task instructions. This verification process specifically includes: Check the source of all conditional predicates in the composite task instruction to confirm that they have corresponding state attributes or external interface support in the structure annotation table, and ensure that each condition has a basis for judgment. Based on the dependencies recorded in the structure annotation table, a graph path coverage check is performed on the composite task instructions to confirm that each sub-process step contained in the composite task instructions has a reachable execution path in the atomic task directed acyclic graph. Composite task instructions that pass both types of verification are written into the target task instruction set; composite task instructions that fail at least one type of verification trigger a local reorganization, re-executing the combinational logic of the composite task instructions only on the failed structural segments until verification is passed, or marked as deprecated after reaching the maximum number of retries.
9. A multi-agent cooperation and execution method based on a task graph structure according to claim 6, characterized in that, After distributing complex task instructions from the target task instruction set to the multi-agent system, the multi-agent system completes the entire process according to a fixed time sequence: First, the task parsing agent performs reverse mapping on the input complex task instructions to generate the corresponding atomic task directed acyclic graph. At the same time, it appends the atomic task directed acyclic graph structure information of the task nodes to the corresponding task instructions in structured JSON format. Secondly, the coordinating agent constructs and maintains the execution state graph of the directed acyclic graph of atomic tasks, and uniformly marks all task nodes as pending execution, in execution, completed or failed states; Furthermore, the coordinating agent dynamically completes task scheduling and allocation for the executing agents based on the topological order and node dependencies of the directed acyclic graph of atomic tasks: a) Subtask nodes without predecessor dependencies are directly marked as schedulable and allocated to idle executing agents for concurrent execution; b) For subtask nodes with predecessor dependencies, the node is sent to the scheduling queue to wait for execution only after all its predecessor nodes have been updated to complete execution status; c) For conditional branch decision nodes, the evaluation agent first reads the current task execution context, calculates the truth value of the conditional predicate, and then selects subtask nodes matching the branch based on the truth value results to complete the distribution and scheduling. Finally, the executing agent receives the task quadruple information and performs the corresponding task operation. After a single execution, it automatically compares the actual post-state effect with the expected post-state effect. Simultaneously, the evaluation agent conducts an operation quality evaluation from two dimensions: task progress and decision rationality. If the evaluation result does not reach the preset threshold, the metacognitive error correction process is automatically triggered to complete the closed-loop verification of task execution.
10. A multi-agent collaboration and execution method based on a task graph structure according to claim 9, characterized in that, During the process of a multi-agent system completing the entire process in a fixed time sequence, step-level process supervision data is generated synchronously, and the following four-tuple data is recorded for each step: Observation: A structured representation of the context snapshot of the current task execution state; Analysis: The text describing the reasoning process of the agent describes the basis for locating the target and the reasons for choosing the operation. Action: The type of operation to be performed, the operation parameters, and the target object; Objective: To determine which atomic task node in the directed acyclic graph corresponds to this operation, and to explain the role of this operation in advancing the overall task objective.