A multi-level intelligent agent orchestration system and a dynamic routing method thereof
Patent Information
- Application Number
- CN202610013065.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-07
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-01-07
AI Technical Summary
这类方案通常以自然语言规划结果作为执行依据,但在实际运行中往往缺乏严格的任务结构约束,规划结果的可执行性和一致性难以保证
[0015]采用以上技术方案,本发明产生了以下有益效果:首先,通过在编排层中引入基于HTN分层任务网络的递归分解机制,将原本以自然语言形式存在的复杂用户请求转化为结构明确、边界清晰的原子任务序列,使任务从一开始就具备可执行、可调度和可验证的工程属性,有效避免了现有技术中任务拆解随意、粒度不稳定和执行路径不可控的问题。
Smart Images

Figure CN122019081B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a multi-level intelligent agent orchestration system and its dynamic routing method. Background Technology
[0002] With the development of artificial intelligence technology, intelligent agent systems based on large language models have gradually evolved from a single dialogue mode to a collaborative execution mode for complex tasks. In early applications, intelligent agents typically operated as single instances, generating results directly through one or more model inferences. This approach could meet basic needs in simple scenarios such as question-answering retrieval and text generation. However, as application scenarios have evolved towards enterprise-level business processes, cross-system collaboration, and complex decision support, the limitations of single intelligent agents in terms of capability coverage, execution stability, and controllability have become increasingly apparent.
[0003] To address the limitations of individual agent capabilities, existing technologies have gradually introduced the concept of multi-agent collaboration, configuring multiple agents with different strengths to jointly complete tasks. Some systems employ rule-based routing, mapping user requests to pre-defined agents according to keywords or regular expressions; others use fixed processes or graph structures to sequentially execute multiple agents in a predetermined order. These solutions improve system functionality to some extent, but still rely on manual pre-design of routing rules or execution paths. When the task structure changes or business requirements adjust, rules or processes need to be reconfigured, resulting in insufficient system flexibility and high maintenance costs. Further exploration has led to the introduction of automatic planning capabilities, breaking down complex tasks through model reasoning and then having different agents execute the sub-tasks. These solutions typically use natural language planning results as the basis for execution, but in practice, they often lack strict task structure constraints, making it difficult to guarantee the executability and consistency of the planning results. On one hand, the granularity of task decomposition is unstable, easily leading to over- or under-decomposition; on the other hand, the lack of clear termination conditions and status indicators results in uncontrollable execution processes, making stable operation in engineering systems difficult. Summary of the Invention
[0004] In view of this, the present invention provides a multi-level intelligent agent orchestration system and its dynamic routing method. By recursively decomposing user requests into atomic task sequences based on the HTN hierarchical task network in the orchestration layer, and combining an intelligent agent registration center, a conformal confidence gating mechanism and a contextual bandit dynamic routing strategy, reliable screening, adaptive delegation and orderly execution of atomic tasks are achieved. Thus, without relying on historical training data, the multi-agent collaboration process has clear structural boundaries, interpretable decision logic and stable execution results, effectively improving the execution determinism, system controllability and engineering implementation capability in complex task scenarios.
[0005] The technical solution adopted in this invention is as follows: A multi-level intelligent agent orchestration system, comprising an application layer, an orchestration layer, an intelligent agent registration center, a confidence gating module, a dynamic routing module, and an intelligent agent layer; the application layer is used to receive user requests and generate session objects; the orchestration layer is used to create root task nodes based on the session objects and perform recursive decomposition of the HTN hierarchical task network to generate atomic task sequences; the intelligent agent registration center is used to store intelligent agent configuration records; the confidence gating module is used to read the initial candidate set from the intelligent agent registration center for each atomic task in the atomic task sequence and generate a gated candidate set based on Conformal confidence gating; the dynamic routing module is used to select a selected routing arm from the gated candidate set based on ContextualBandit and send atomic tasks to the intelligent agents corresponding to the selected routing arms through the task delegation interface of the orchestration layer to obtain execution results; the intelligent agent layer is used to execute atomic tasks and return execution results, and the orchestration layer is used to summarize the execution results of each atomic task and output them to the application layer.
[0006] Furthermore, the root task node created by the orchestration layer includes a task identifier field, a task description field, a task type field, and a subtask list field. The orchestration layer writes the original text of the user request into the task description field and initializes the subtask list field to an empty list. The orchestration layer performs a recursive decomposition operation on the root task node. The recursive decomposition operation includes: calling the large language model and sending a decomposition instruction containing the content of the task description field of the current task node, and receiving the subtask description list returned by the large language model; creating a subtask node for each subtask description in the subtask description list and adding it to the subtask list field of the current task node; continuing to perform the recursive decomposition operation for subtask nodes with a composite task type field, and ending the decomposition for subtask nodes with an atomic task type field; after the recursive decomposition operation is completed, traversing the leaf task nodes with an atomic task type field to generate an atomic task sequence, and passing the atomic task sequence to the confidence gating module.
[0007] Furthermore, each agent configuration record stored in the agent registry center includes an agent identifier field, a capability tag list field, and a tool binding list field. After receiving the atomic task sequence, the confidence gating module reads all agent configuration records from the agent registry center for the current atomic task and forms an initial candidate set. The confidence gating module calls the large language model and sends a capability extraction instruction containing the task description field of the current atomic task. It receives the capability tag list returned by the large language model and stores the capability tag list as the required capability set for the current atomic task.
[0008] Furthermore, the confidence gating module creates a confidence evaluation table within the current session object. Each row of the confidence evaluation table corresponds to an agent in the initial candidate set, and each row contains an agent identifier column, a capability coverage count column, a capability missing count column, and a confidence pass flag column. The confidence gating module initializes the capability coverage count column and the capability missing count column to preset initial values and initializes the confidence pass flag column to a pending state. The confidence gating module performs a capability matching operation to update the capability coverage count column and the capability missing count column. The capability matching operation includes traversing each capability tag in the required capability set and checking the capability tag in the corresponding agent. The existence of the capability label list field of the entity; the confidence gating module performs a confidence judgment operation, which includes: when the capability missing count column meets the preset missing condition, the confidence pass flag column is set to pass status; when the capability missing count column does not meet the preset missing condition, the confidence pass flag column is set to pass status or rejection status according to the preset ratio relationship between the capability coverage count column and the total number of elements in the required capability set; the confidence gating module filters rows with the confidence pass flag column in the pass status from the confidence evaluation table to form a gating candidate set, and passes each atomic task and its corresponding gating candidate set to the dynamic routing module.
[0009] Furthermore, the dynamic routing module defines each agent in the gated candidate set as a routing arm for the current atomic task, and creates an arm state record for each routing arm within the current session object. The arm state record includes an arm identifier field, a trial count field, and a success count field, with the lifecycle of the trial count field and the success count field limited to the execution cycle of the current user request. The dynamic routing module also creates a context feature record within the current session object. This context feature record includes four feature slots: the first feature slot stores the position index of the current atomic task in the atomic task sequence, the second feature slot stores the number of elements in the required capability set, the third feature slot stores the number of elements in the gated candidate set, and the fourth feature slot stores the current session... The number of atomic tasks completed within the object; the dynamic routing module performs an arm score calculation operation on each routing arm and selects the routing arm with the highest final score as the selected routing arm. The arm score calculation operation includes generating a capability matching number based on the number of elements in the intersection of the requirement capability set and capability tag list fields, generating a tool available number based on the number of elements in the tool binding list field, summing the capability matching number and the tool available number to obtain the base score, and determining the exploration reward value based on the value of the trial count field and combining it with the base score to form the final score; the dynamic routing module sends atomic tasks through the task delegation interface and receives the execution results. When the execution result meets the preset valid output conditions, the success count field is updated and the execution result is written to the output buffer of the atomic task. After all atomic tasks are completed, the contents of the output buffer are summarized and returned to the application layer.
[0010] A dynamic routing method for a multi-level intelligent agent orchestration system includes: receiving user requests and generating session objects; creating root task nodes based on the session objects and performing recursive decomposition of the HTN hierarchical task network to obtain an atomic task sequence; for each atomic task in the atomic task sequence, calling a large language model to generate a set of required capabilities, and using a confidence gating module to filter from the initial candidate set based on Conformal confidence gating to obtain a gated candidate set; using a dynamic routing module to map the gated candidate set to routing arms based on ContextualBandit and calculating and selecting the selected routing arm based on the arm score, delegating the atomic tasks to the intelligent agents corresponding to the selected routing arms to obtain execution results and writing them to the output buffer; summarizing the contents of the output buffer and outputting them.
[0011] Furthermore, the recursive decomposition operation includes: sending a decomposition instruction to the large language model for the current task node and receiving a list of subtask descriptions; creating subtask nodes one by one for each subtask description and writing them into the subtask list field; classifying the subtask nodes into composite types or atomic types based on the task type field, and continuing to perform the recursive decomposition operation for composite type subtask nodes; after the recursive decomposition operation is completed, traversing all atomic type leaf task nodes to generate an atomic task sequence, the order of which is determined by the traversal order.
[0012] Furthermore, the agent selection operation for the current atomic task includes: reading all agent configuration records from the agent registry and forming an initial candidate set; calling the large language model and sending a capability extraction command to obtain a capability tag list and storing it as a required capability set; creating a confidence evaluation table within the session object and creating a corresponding row for each agent in the initial candidate set, initializing the capability coverage count column, capability missing count column, and confidence pass flag column; traversing the required capability set for each agent and updating the capability coverage count column or capability missing count column based on the existence of capability tags in the capability tag list field.
[0013] Furthermore, the confidence determination operation includes: reading the capability coverage count column and capability missing count column of the current row in the confidence assessment table; setting the confidence pass flag column to pass status when the capability missing count column is at a preset missing threshold value; comparing the capability missing count column with a preset ratio threshold of the total number of elements in the required capability set when the capability missing count column is greater than the preset missing threshold value, and setting the confidence pass flag column to pass status or rejection status accordingly; filtering rows with the confidence pass flag column set to pass status from the confidence assessment table and forming a gated candidate set.
[0014] Furthermore, the routing operation includes: defining each agent in the gated candidate set as a routing arm and creating an arm state record for each routing arm within the session object, and initializing the trial count field and success count field; creating a context feature record containing the first to fourth feature slots; calculating the capability matching number and tool availability number for each routing arm and generating a base score, determining the exploration reward value based on the trial count field and synthesizing the final score with the base score; selecting the routing arm with the highest final score as the selected routing arm and updating the trial count field; sending atomic tasks to the agent corresponding to the selected routing arm and receiving the execution results; updating the success count field based on whether the execution results meet the preset valid output conditions and writing the execution results to the output buffer; repeating the routing operation sequentially for the atomic task sequence until all atomic tasks are completed, and summarizing the contents of the output buffer before returning.
[0015] By adopting the above technical solution, the present invention has produced the following beneficial effects: First, by introducing a recursive decomposition mechanism based on HTN hierarchical task network into the orchestration layer, the complex user requests that originally existed in the form of natural language are transformed into atomic task sequences with clear structure and boundaries, so that the tasks have the engineering attributes of being executable, schedulable and verifiable from the beginning, effectively avoiding the problems of arbitrary task decomposition, unstable granularity and uncontrollable execution path in the prior art.
[0016] Secondly, by introducing a Conformal confidence gating mechanism during the agent selection phase, the capability coverage of agents is verified and gating item by item without relying on historical statistical data. This ensures that agents entering the execution phase have a clear lower limit guarantee in terms of capability, thereby reducing the probability of execution failure and improving the certainty and stability of task completion at the overall system level. Thirdly, this invention introduces the ContextualBandit decision-making concept into the dynamic routing process and strictly limits it to the lifecycle of a single session object. Through contextual features and immediate execution feedback, the optimal execution path for the current task is selected. This enables the system to explore new execution paths without introducing cross-session policy drift and unexplainable decision-making problems, significantly improving the system's adaptability and controllability in complex scenarios.
[0017] Furthermore, this invention manages and integrates the execution results of each atomic task in an orderly manner through an output buffer and a unified aggregation mechanism. This avoids common problems in multi-agent collaborative scenarios, such as result fragmentation, semantic conflicts, and disordered output order, ensuring that the final output remains consistent in semantics, structure, and business logic. In summary, this invention achieves collaborative optimization of multi-agent systems in task planning, capability selection, dynamic routing, and result integration without relying on long-term training or introducing complex external control logic. It significantly improves the system's engineering feasibility, operational stability, and business adaptability, making it particularly suitable for enterprise-level applications with high requirements for reliability, interpretability, and scalability. Attached Figure Description
[0018] Figure 1 This is a schematic diagram illustrating the score composition of the dynamic routing module performing arm score calculation operations in the multi-level intelligent agent orchestration system of this invention.
[0019] Figure 2 This is a logical diagram of the confidence evaluation table and its internal capability matching state constructed by the confidence gating module in the multi-level intelligent agent orchestration system of this invention.
[0020] Figure 3This is a schematic diagram of the evolution curves of the scores of each routing arm in the execution process of the ContextualBandit dynamic routing module in the multi-level intelligent agent orchestration system of this invention. Detailed Implementation
[0021] Any feature disclosed in this specification, unless otherwise stated, may be replaced by other equivalent or similar features. That is, unless otherwise stated, each feature is merely one example of a series of equivalent or similar features.
[0022] A multi-level agent orchestration system includes an application layer, an orchestration layer, an agent registry, a confidence gating module, a dynamic routing module, and an agent layer. The application layer receives user requests and generates session objects. The orchestration layer creates root task nodes based on the session objects and performs recursive decomposition of the HTN hierarchical task network to generate atomic task sequences. The agent registry stores agent configuration records. The confidence gating module reads an initial candidate set from the agent registry for each atomic task in the atomic task sequence and generates a gated candidate set based on Conformal confidence gating. The dynamic routing module selects a chosen routing arm from the gated candidate set based on ContextualBandit and sends atomic tasks to the agents corresponding to the selected routing arms through the task delegation interface of the orchestration layer to obtain execution results. The agent layer executes atomic tasks and returns execution results. The orchestration layer summarizes the execution results of each atomic task and outputs them to the application layer.
[0023] When the application layer receives a user request, it first converts requests from different entry points into a unified message content, and then generates a session object on this message content. Taking WeChat Work and web applications as examples, the application layer extracts the user identifier, channel identifier, request arrival time, original text, and any file references that may be carried from the original request. It then normalizes the original text to a unified encoding and cleans up invisible control characters, ensuring that the decomposition results are not unstable due to line break contamination or encoding differences during subsequent parsing by a large language model. When creating a session object, at least the session identifier, user identifier, channel identifier, message content, and metadata are written. The metadata may include client type, language preference, and idempotency flags. To allow the same session object to be reused in subsequent processes, the application layer typically sets the session identifier to a fixed-length string, such as a 32-bit hexadecimal string, and retains the session object in memory until the current user request is processed. When cross-process collaboration is required, the session object can optionally be serialized into a JavaScript object representation string and written to a key-value store with a 24-hour expiration time, so that the same session object can be directly read when the same session is triggered again within a short period. The advantage of doing this is that the root task node, atomic task sequence, and confidence evaluation table are all attached to the same session object. Each step in the process only needs to pass a reference to the session object, which reduces parameter assembly and avoids the problem of "task decomposition results not matching candidate set screening results" when the same request is processed concurrently.
[0024] Upon receiving the session object, the orchestration layer immediately creates the root task node. The root task node uses structured fields for storage, including a task identifier field, a task description field, a task type field, and a subtask list field. The task identifier field is typically obtained by concatenating the session identifier with an incrementing sequence number, such as "session identifier-0001," facilitating a stable mapping in the logs and output buffer. The task description field contains the original text of the user request, ensuring that subsequent recursive decomposition always revolves around the user's original message. The task type field is set to a composite type when the root task node is created because it plays the role of "accepting the entire user request and allowing further decomposition." If it were set to an atomic type from the beginning, it would prevent the generation of sufficiently fine-grained atomic task sequences, thus affecting agent selection and dynamic routing. The subtask list field is initialized to an empty list to hold the subtask nodes generated during the recursive decomposition phase.
[0025] During the recursive decomposition of the HTN hierarchical task network, the orchestration layer expands the "current task node" layer by layer. Each round of expansion sends a decomposition instruction to the large language model. The decomposition instruction includes at least: the task description field content of the current task node, output format constraints, and the constraint that "the subtask description list must be directly executable and semantically non-overlapping." It is recommended that the output format be fixed as a numbered list of subtask descriptions or a JavaScript object representation array string. This is because the orchestration layer needs to stably parse the subtask description list and create subtask nodes one by one. If free text output is allowed, the parsing will fluctuate with the model's expression habits, easily resulting in multiple subtasks being written on the same line or interspersed with explanatory text, leading to unclear subtask node boundaries. The orchestration layer can adopt the following executable parsing strategy: when the returned content is an array string, directly split it by array elements to obtain the subtask description; when the returned content is a numbered list, scan line by line, identify the prefixes such as "1.", "2.", "3.", etc., and extract the text after the prefix as the subtask description. Each subtask description generates a subtask node. The subtask node also contains a task identifier field, a task description field, a task type field, and a subtask list field. The task description field contains the subtask description, the subtask list field is initialized to an empty list, and the task identifier field follows the generation rule of "session identifier - parent node number - child number", for example, "session identifier - 0001 - 0003".
[0026] The determination of the task type field directly determines whether recursive decomposition continues. To avoid the uncertainty caused by "judging the atomicity based on experience," the orchestration layer uses a repeatable decision process to assign a value to the task type field after creating each subtask node: it sends a decision instruction containing the task description field content of the subtask node to the large language model, requiring the model to return only one of "atomic type" or "composite type"; the orchestration layer writes the return value into the task type field. The benefit of this is that the same subtask description yields a consistent task type field in different runtime environments, making the recursive decomposition path more stable. To prevent the recursive decomposition from expanding continuously in extreme requests, the orchestration layer optionally introduces two constraints without changing the field semantics: setting a maximum recursion level, such as 6 levels; and setting a maximum number of subtasks for a single task node, such as 8. When the maximum recursion level is reached, the orchestration layer writes the task type field of the current task node as an atomic type and stops further decomposition, ensuring that the entire recursive decomposition operation converges and produces an atomic task sequence.
[0027] After the recursive decomposition operation is completed, the orchestration layer traverses all leaf task nodes with the "atomic" task type field to generate atomic task sequences. The order of the atomic task sequences is determined by the traversal order. It is recommended to use a depth-first approach, traversing from front to back according to the storage order of the subtask list fields. This is because the order of the subtask list fields is usually consistent with the order of the subtask description list returned by the large language model, and this order implicitly implies prerequisite dependencies in many scenarios, such as "collecting information" before "generating conclusions." Through a fixed traversal strategy, atomic task sequences can be repeatedly generated on the same user request, thus ensuring a stable processing order for subsequent Conformal confidence gating and dynamic routing modules. The orchestration layer writes the atomic task sequences to the session object and passes them to the confidence gating module.
[0028] When storing agent configuration records, the agent registry ensures that each record contains at least an agent identifier field, a capability tag list field, and a tool binding list field. In practice, relational database tables or key-value stores can be used to persist these fields: the agent identifier field serves as the primary key; the capability tag list field is stored as a string array, for example, each agent configuration record contains 5 to 30 capability tags; and the tool binding list field is stored as an array of tool names, for example, 0 to 40 tool names. To reduce runtime read costs, the agent registry can optionally maintain an inverted index from capability tags to agent identifiers. However, when a perfect fit to the initial candidate set definition is required, the confidence gating module still prioritizes reading all agent configuration records and assembling the initial candidate set.
[0029] When the confidence gating module processes an atomic task, it first reads all agent configuration records from the agent registry and forms an initial candidate set. Then, it generates a required capability set for the atomic task. The generation of the required capability set is also accomplished through a large language model: the confidence gating module sends a capability extraction command containing the task description field of the current atomic task and requests the return of "capability tag list fields matching tags in the thesaurus". To ensure the results can be directly used for subsequent capability matching operations, an engineering practice typically provides a capability tag thesaurus, listing the standard spelling and common aliases for each capability tag. When the capability tag list returned by the large language model contains aliases, the confidence gating module normalizes the aliases to the standard spelling according to the thesaurus, and then stores the normalized capability tag list as the required capability set. The advantage of this approach is that different agent configuration records may be entered by different maintenance personnel. Without normalization, string differences in capability tags could be misjudged as "non-existent," leading to an unnecessary increase in the capability missing count column and an excessively narrowed candidate set after gating.
[0030] Subsequently, the confidence gating module creates a confidence evaluation table within the current session object. Each row of the confidence evaluation table corresponds to an agent in the initial candidate set and includes an agent identifier column, a capability coverage count column, a capability missing count column, and a confidence pass flag column. The confidence gating module initializes the capability coverage count column and the capability missing count column to preset initial values, typically 0 in engineering practice; it initializes the confidence pass flag column to a pending state to avoid premature pass or rejection decisions before the capability matching operation is completed, thus preventing inconsistencies caused by concurrent readings in intermediate states. The capability matching operation uses "each capability tag in the required capability set" as the smallest loop granularity: for the agent corresponding to the current row, it checks whether the capability tag exists in the capability tag list field of that agent. If it exists, the capability coverage count column is incremented by 1; if the search fails, the capability missing count column is incremented by 1. Since the search operation is essentially a string set inclusion judgment, it can be directly implemented using a hash set. The capability tag list field is pre-converted into a set to obtain a stable linear time complexity, and the execution latency remains acceptable even when the number of multi-agents reaches 500 or 1000.
[0031] refer to Figure 2 This diagram, presented in matrix form, details how the system uses the Conformal confidence gating algorithm to filter out the post-gating candidate set from the initial candidate set before dynamic routing. The horizontal axis corresponds to the capability labels in the capability set generated by the current atomic task. Taking atomic task T2_1 in the example, the horizontal axis lists the four key capability dimensions: "Logistics Query," "Interface Call," "Exception Handling," and "Order Parsing." The vertical axis corresponds to all agent configuration records read from the agent registry center, i.e., the five agents in the initial candidate set: A_order, A_logistics, A_writer, A_policy, and A_general. Each cell in the matrix represents a Boolean judgment result indicating whether a specific agent possesses a specific capability. This two-dimensional grid structure is represented in system memory as a confidence evaluation table data structure mounted within the session object.
[0032] like Figure 2As shown, a "√" sign in a cell indicates that the capability label in the corresponding column is present in the capability label list field of the agent in that row, while an "×" sign indicates that it is not present. This binary matching relationship is generated by the confidence gating module performing capability matching operations. For example, observing the row containing A_logistics, it is marked with "√" in the "Logistics Query", "Interface Call", and "Exception Handling" columns, while it is marked with "×" in the "Order Parsing" column. This intuitively corresponds to the capability coverage count column value being 3 and the capability missing count column value being 1. Similarly, the row A_order is marked in "Order Parsing" and "Exception Handling", indicating that its coverage count is 2 and its missing count is 2. For completely unrelated agents such as A_writer, all columns in this row are "×", indicating that its coverage count is 0 and its missing count is 4, meaning it is completely unable to respond to the current task. Figure 2 The core value lies in demonstrating how the decision boundary of confidence gating is formed.
[0033] Specifically, although both A_logistics and A_order have missing items (i.e., rows containing "×"), this does not cause them to be directly eliminated. According to the confidence determination operation described in the claim, the system first checks if there are any perfectly matched rows (i.e., rows with 0 missing items), which do not exist in this example. Subsequently, the system settles for a decision based on a preset proportional relationship. Since the total number of elements N in the demand capability set is 4, the system sets the passing threshold H to be half of N rounded up, i.e., 2. Figure 2 In the visualization logic, any row is considered to have passed if the number of "√" marks is greater than or equal to 2. Therefore, the rows containing A_logistics (3 √) and A_order (2 √) are marked as passed, thus forming the post-gated candidate set to be passed to the next stage. Conversely, although A_policy has "anomaly handling" capabilities, its row has only 1 "√", which is below the threshold of 2, so it is blocked; A_writer and A_general are rejected because they have no matches at all. This matrix coverage-based screening mechanism effectively strikes a balance between "overly strict leading to no path" and "overly lenient leading to computational waste", ensuring that the agent entering the expensive dynamic routing and model execution stage has at least half of the core capabilities to complete the task in probability. This figure not only shows the static matching results, but also reveals how the system transforms fuzzy text matching into precise numerical gating logic through structured capability mapping when dealing with complex, multi-dimensional natural language tasks.
[0034] The confidence determination operation is performed after the capability matching operation is completed. The confidence gating module reads the capability coverage count column and capability missing count column of the current row of the confidence evaluation table, and sets the confidence pass flag column according to the preset missing condition and preset ratio relationship. In the default implementation, the preset missing condition is set to "capability missing count column is 0", which means that the capability tag list field of the agent covers all capability tags in the required capability set. At this time, the confidence pass flag column is set to the pass state. When the capability missing count column does not meet the preset missing condition, the confidence gating module further compares the capability coverage count column with half of the total number of elements in the required capability set. If the capability coverage count column is greater than or equal to half of the total number of elements in the required capability set, the confidence pass flag column is set to the pass state; otherwise, it is set to the rejection state. Adopting the order of "checking missing capabilities first, then coverage ratio" has two direct benefits: First, agents with a missing capability count of 0 belong to the best-matching set, and prioritizing their entry into the gating candidate set can reduce the exploration cost of the subsequent dynamic routing module. Second, in actual business, there are often situations where "atomic tasks require multiple capabilities, but not all of them must be fully possessed by the same agent." For example, an atomic task may involve both "data reading" and "text generation." If only agents with full coverage are retained, the gating candidate set may be compressed to an excessively small size, leaving the dynamic routing module with few alternative routing arms. Introducing the rule of "passing even if half-coverage is achieved" allows us to retain a set of agents with the main capabilities as candidates, ensuring that the gating candidate set is neither too large nor too small, thus making it easier to complete the delegation within a limited execution cycle.
[0035] To ensure that the Conformal confidence gating is configurable, interpretable, and reproducible, the preset missing condition and preset ratio are written into the configuration file during deployment and released with each version, remaining unchanged during operation. The confidence gating module simultaneously writes the capability coverage count column, capability missing count column, and final confidence via a marker column back to the session object for each row, facilitating review of why a particular agent was included or excluded during problem localization. In an optional implementation, the preset missing condition can be set to "capability missing count column not greater than 1," to tolerate insufficient coverage of the same vocabulary or overly fine-grained individual capability tags. Correspondingly, the preset ratio remains "capability coverage count column greater than or equal to half the total number of elements in the required capability set," thus ensuring that the size of the candidate set after gating does not significantly expand due to missing tolerance.
[0036] The confidence gating module finally selects rows from the confidence evaluation table whose confidence pass flag column is set to pass, and forms a gated candidate set of agents corresponding to these rows. The mapping between atomic tasks and gated candidate sets is then written into the session object. When the gated candidate set is passed forward, the agent identifier field remains unchanged to avoid the same agent being referenced by different names in subsequent processing. Up to this point, the specific implementation processes of the application layer generating the session object, the orchestration layer generating the atomic task sequence, the agent registry providing the initial candidate set, and the confidence gating module generating the gated candidate set are all closed-loop within the same session object. Subsequently, the dynamic routing phase can proceed simply by retrieving the corresponding gated candidate set according to the atomic task sequence.
[0037] After receiving the atomic task sequence and the gated candidate set corresponding to each atomic task, the dynamic routing module first establishes a read-only mapping of "atomic tasks to gated candidate sets" within the session object, and then processes them one by one in the order of the atomic task sequence. The reason for executing the atomic tasks sequentially, rather than throwing all atomic tasks to different agents in parallel, is that the atomic task sequence is generated by the recursive decomposition of the HTN hierarchical task network, which usually implicitly contains a dependency relationship of "first producing intermediate conclusions that can be reused by subsequent tasks, and then generating the final expression." Sequential execution allows subsequent atomic tasks to carry key output fragments of completed atomic tasks in their task description fields, thereby reducing repeated queries and lowering the probability of contradictory execution results. To avoid session object bloat, when the dynamic routing module writes the execution result of each completed atomic task to the output buffer of that atomic task, it also writes a concise summary fragment. This summary fragment, for example, extracts the first 200 characters, and the original complete execution result is kept in another storage slot in the same output buffer. This allows the orchestration layer to prioritize using the summary fragment during aggregation and review the complete content when needed.
[0038] When processing a current atomic task, the dynamic routing module defines each agent in the gated candidate set of that current atomic task as a routing arm. Each routing arm corresponds to an arm state record within the session object. The arm state record includes an arm identifier field, a trial count field, and a success count field. The arm identifier field can be directly written to the agent identifier field of the corresponding agent, ensuring that the routing arm has a stable reference throughout the entire execution cycle. The trial count field and the success count field are both initialized to 0 and exist only within the execution cycle of the current user request. They are cleared when the session object is released after the execution cycle ends, thus avoiding the introduction of historical data across requests. To ensure reproducible and easily troubleshootable routing selection, the dynamic routing module also creates a context feature record within the session object. This record includes four feature slots: Slot 1 (1), Slot 2 (2), Slot 3 (3), and Slot 4 (4). Slot 1 stores the position index of the current atomic task within the atomic task sequence (e.g., 1 for the first atomic task); Slot 2 stores the number of elements in the required capability set (e.g., 6 for 6 capability tags); Slot 3 stores the number of elements in the gated candidate set (e.g., 12 for 12 agents); and Slot 4 stores the number of completed atomic tasks within the current session object (e.g., 3 for 3 completed). The value of recording these features lies in two aspects: firstly, explaining why multiple routing arms may have the same score when the sequence is relatively small, the required capability set is small, and the gated candidate set is large; secondly, in optional implementations, it can be used for tie-breaking decisions, ensuring that the same request receives a consistently selected routing arm across different server instances.
[0039] refer to Figure 3 , Figure 3 The horizontal axis represents the number of task execution rounds, ranging from 0 to 20, and the vertical axis represents the routing arm score, ranging from 0 to 14. The graph contains five curves, each corresponding to one of the five different routing arms: routing arm A_logistics, routing arm A_order, routing arm A_writer, routing arm A_policy, and routing arm A_general.
[0040] exist Figure 3In the illustrated embodiment, the five routing arms have different base scores. The agent corresponding to routing arm A_logistics has a high number of capability matches and available tools, with a base score of 8, represented by the curve marked with a blue circle in the figure. Routing arm A_policy has a base score of 7, represented by the curve marked with a green diamond in the figure. Routing arm A_order has a base score of 6, represented by the curve marked with a purple square in the figure. Routing arm A_writer has a base score of 4, but gains dynamic enhancement as tasks are executed, represented by the curve marked with an orange triangle in the figure. Routing arm A_general has the lowest base score of 3, represented by the curve marked with a red inverted triangle in the figure.
[0041] According to the dynamic routing method of the present invention, when the task execution round is 0, the trial count field of all routing arms is 0, so each routing arm can obtain a fixed exploration constant K as an exploration reward value. In this embodiment, the fixed exploration constant K is configured as 3. Therefore, in round 0, the final score of routing arm A_logistics is the base score 8 plus the exploration reward value 3, which equals 11; the final score of routing arm A_policy is 7 plus 3, which equals 10; the final score of routing arm A_order is 6 plus 3, which equals 9; the final score of routing arm A_writer is 4 plus 3, which equals 7; and the final score of routing arm A_general is 3 plus 3, which equals 6.
[0042] In the first round of task execution, the dynamic routing module selects the routing arm A_logistics with the highest final score as the selected routing arm and updates its trial count field from 0 to 1. At this point, routing arm A_logistics has already been tested, and its exploration reward value changes from 3 to 0. Therefore, from the first round, the final score of routing arm A_logistics drops to its base score of 8. The other four routing arms, since they have not yet been tested, retain their exploration reward value of 3. In the first round, routing arm A_policy becomes the routing arm with the highest final score of 10, and is subsequently selected, with its trial count field updated to 1.
[0043] In the second round, the exploration reward value of routing arm A_policy also becomes 0, and its final score drops to the base score of 7. At this time, routing arm A_order has a score of 9, becoming the highest. After being selected, its exploration reward value also disappears in the third round. According to this exploration mechanism, in the first 5 rounds, each routing arm that has not yet been tried will have the opportunity to be selected in turn due to the increase in its exploration reward value. This mechanism ensures that even agents with lower base scores have the opportunity to participate in task execution as long as they have not yet been tried, thus preventing some agents that may be more suitable for a specific task from never being called upon due to slightly lower initial scores.
[0044] Of particular note is the evolution curve of the routing arm A_writer. This routing arm started with a base score of only 4, and its final score remained at 7 during rounds 0 to 3 thanks to the exploration reward value. After being selected and completing the task in round 4, the exploration reward value disappeared, but the routing arm exhibited dynamic enhancement characteristics. From Figure 3 It can be observed that after the routing arm A_writer was tested, its score curve began to show an upward trend, reflecting that the agent gained improved capabilities through learning or context accumulation as tasks were performed. By the 20th round, the score of the routing arm A_writer had risen to about 10, surpassing several routing arms with higher initial base scores.
[0045] Figure 3 The text box indicating an exploration reward value of +3 is located at approximately x1 on the horizontal axis and y1 on the vertical axis, clearly indicating the value of the fixed exploration constant. The selection of this parameter directly affects the balance between exploration and utilization in the system. If the exploration constant is too small, for example, set to 1, the difference in base scores will dominate the selection process, potentially causing some routing arms to never get a trial opportunity. If the exploration constant is too large, for example, set to 10, it will over-encourage exploration, causing routing arms that have been proven inefficient to be repeatedly selected, reducing the overall efficiency of the system. In this embodiment, K equals 3, which achieves a good balance between exploration and utilization in most scenarios.
[0046] from Figure 3 The overall trend shows that after the initial exploration phase, the scores of each routing arm gradually stabilize. Routing arms with high base scores, such as A_logistics and A_policy, maintain high scores in subsequent rounds, while routing arms with low base scores, such as A_general, maintain low scores after losing exploration rewards. This evolutionary pattern is consistent with the expected behavior of the ContextualBandit algorithm, which quickly identifies high-performance routing arms through limited exploration and then mainly utilizes these routing arms to complete subsequent tasks, thereby achieving a high cumulative success rate throughout the execution cycle.
[0047] The arm score calculation then proceeds. For each routing arm, the dynamic routing module first calculates the capability matching count: it reads the capability tag list field of the agent corresponding to the routing arm, converts it into a set structure, and performs an intersection check with the required capability set of the current atomic task. The number of elements in the intersection is recorded as the capability matching count. This implementation does not require any "weight allocation"; essentially, it counts whether capability tags are covered item by item. Under common scales, this can be completed in milliseconds. For example, when the capability tag list field contains 18 capability tags and the required capability set contains 6 capability tags, only 6 existence checks are needed to obtain the capability matching count. Next, the number of available tools is calculated: it reads the tool binding list field of the agent corresponding to the routing arm and directly takes the number of elements as the number of available tools. For example, if the tool binding list field contains 7 tool names, the number of available tools is 7. The base score is obtained by adding the capability matching number and the tool availability number. The reason for this design is that the capability matching number describes "the coverage of capabilities required to complete the current atomic task", and the tool availability number describes "the range of external capabilities that the agent can call". Both are discrete counts that can be obtained directly from the agent's configuration record, without relying on training, historical data, or manually setting weight coefficients. Assuming that the candidate set has been shrunk by Conformal confidence gating after gating, this count superposition can push obviously unsuitable agents to the back with extremely low implementation cost.
[0048] refer to Figure 1 , Figure 1 This is a schematic diagram illustrating the score composition of the dynamic routing module performing arm score calculation operations in a multi-level intelligent agent orchestration system according to an embodiment of this application. Figure 1 As shown in the figure, this diagram visually illustrates how the dynamic routing module utilizes a context-based multi-armed slot machine algorithm to quantify and score different agents in the gated candidate set for a specific atomic task (e.g., atomic task T2_1 in the embodiment, i.e., "obtaining the return value of the logistics interface"). The horizontal axis lists the agent identifiers retained in the gated candidate set after being filtered by the confidence gating module. In this embodiment, this specifically includes agents A_order and A_logistics. It also schematically illustrates the scoring status if other agents (such as A_policy or A_writer) are considered, for comparison and explanation. The vertical axis represents the final score value of each routing arm. The bar chart corresponding to each agent in the figure consists of three stacked parts, representing the number of capability matches, the number of available tools, and the exploration reward value. These three together constitute the scalar basis used for the final routing decision.
[0049] Specifically, the bottom layer of the bar chart represents the capability matching count component of the base score. This value is not a statically set weight, but is calculated in real-time by the dynamic routing module at runtime. The system first reads the required capability set for the current atomic task. Taking the example of T2_1, the required capability set includes four tags: "logistics query," "interface call," "exception handling," and "order parsing." Simultaneously, the system reads the capability tag list field for each agent. For agent A_logistics, its capability tag list includes "logistics query," "interface call," and "exception handling," and the number of elements intersecting with the required capability set is 3. Therefore, the height of the capability matching count portion at the bottom of its bar chart is 3. For agent A_order, its intersection only includes "order parsing" and "exception handling," so the height of its capability matching count portion is 2. This layer intuitively reflects the degree of fit between the agent's self-defined business domain and the current task requirements. The middle layer of the bar chart represents the tool availability count component of the base score. This value comes from the number of elements in the tool binding list field of the agent configuration record. As shown in the figure, agent A_logistics is bound to two tools, logistics_api and geo_time_service, so its middleware layer height is 2; while agent A_order is only bound to one tool, order_db_query, so its height is 1. The sum of these two parts forms the base score for each routing arm: A_logistics has a base score of 5, and A_order has a base score of 3.
[0050] The shaded area at the top of the bar chart represents the exploration reward value. This is the key mechanism by which this application utilizes the ContextualBandit algorithm to solve the "cold start" and "exploration-exploitation" dilemma. At the time of this embodiment, it is assumed that the trial count fields for both agents for the current atomic task are both 0, meaning they have not yet been attempted. The system assigns equal exploration reward values to both agents based on a preset fixed exploration constant K (set to 3 in this example). Therefore, agent A_logistics' final score is the sum of its base score of 5 and its exploration reward value of 3, which is 8 points; agent A_order's final score is the sum of its base score of 3 and its exploration reward value of 3, which is 6 points. Figure 1The diagram clearly shows that, due to its superior capability matching and tool richness, A_logistics achieves a significantly higher final score than A_order with the same exploration reward, thus being selected as the chosen route arm by the dynamic routing module. Furthermore, the diagram implicitly reveals a dynamic change logic over time: once A_logistics is selected and executed, its trial count field will no longer be 0. When calculating the score for the same atomic task again, its top-level exploration reward value will disappear (reset to zero). At this point, if other agents with high base scores exist that have not yet been explored, the system will tend to allocate opportunities to those unexplored paths, thereby automatically covering unknown possibilities. This component-stacking-based score visualization fully reveals how this application achieves efficient and robust agent routing selection by combining deterministic rules with a dynamic exploration mechanism without relying on a large amount of pre-trained data.
[0051] The exploration reward value is determined strictly based on the trial count field in the arm state record. The dynamic routing module reads the trial count field: when the value of the trial count field is 0, the exploration reward value is set to a fixed exploration constant; when the value of the trial count field is greater than 0, the exploration reward value is set to 0. The fixed exploration constant is pre-configured as a positive integer during deployment, such as 3 or 5. The effect of this is that each untested routing arm is given a deterministic "first-try boost," giving it a chance to be selected and produce an execution result, thus preventing agents with slightly lower base scores but actually more suitable for the current atomic task from never getting a chance to be invoked. Since the trial count field only accumulates within the execution cycle of the current user request, this exploration only occurs within the same request and does not carry over the behavior of one request to another, thus avoiding cross-session bias. The final score is obtained by adding the base score to the exploration reward value, and the dynamic routing module selects the routing arm with the highest final score from all routing arms as the selected routing arm. If multiple routing arms have the same final score, the optional implementation provides two unambiguous tie-score rules: one is to prioritize the routing arm with a larger number of capability matches, because the matching of the requirement capability set and the capability label list field more directly determines whether valid output content can be produced; the other is to select the smallest one according to the lexicographical order of the arm identifier field, so as to ensure that the selection result is consistent in the case of tie scores in multi-instance deployment.
[0052] After selecting a routing arm, the dynamic routing module first increments the test count field of the corresponding arm status record by 1, and then sends the current atomic task to the agent corresponding to the selected routing arm through the task delegation interface of the orchestration layer. The task payload sent at the time includes at least: the task identifier field of the atomic task, the task description field, the required capability set, the session identifier of the session object, and the write position identifier of the output buffer. The reason for sending the required capability set together is that the task description field is often in natural language. If the agent only looks at the natural language during execution, it may have inconsistent understanding of "which capability tags must be covered". Explicitly passing in the required capability set allows the agent to internally check whether its tool binding list field contains tools related to these capability tags, thereby reducing invalid calls. To control the execution time of a single atomic task, an execution time limit field can be optionally included in the task payload, such as 60 seconds or 120 seconds. If the agent fails to complete before the execution time limit expires, it returns an execution result with a status flag, which is then processed by the dynamic routing module as "not meeting the preset valid output conditions".
[0053] After receiving an atomic task, the agent layer looks up the corresponding agent configuration record based on its agent identifier field, reads the capability tag list field and tool binding list field, and assembles the task description field and required capability set of the atomic task into input prompt content. To ensure execution is feasible, the input prompt content is recommended to use a fixed paragraph structure: the first paragraph contains the original text of the atomic task's task description field; the second paragraph contains the capability tag list of the required capability set; the third paragraph contains the expected output format, such as "output given as bullet points, containing no less than 3 key points"; the fourth paragraph contains constraints, such as "only return content related to the atomic task". During execution, the agent can decide whether to call external tools based on the tool binding list field; when the tool binding list field is empty, the agent directly generates and returns the text execution result; when the tool binding list field is not empty, the agent first generates a tool call plan, then calls each tool one by one and incorporates the tool return values into the final execution result. Regardless of whether tools are called, the final execution result returned by the agent maintains a consistent structure, at least including the execution result text content and an execution status field, such as "completed". To facilitate subsequent determination of valid output content, it is recommended to remove extra whitespace from the execution result text and ensure that it is a serializable string.
[0054] After receiving the execution result, the dynamic routing module determines whether it contains valid output content based on preset valid output conditions. These preset valid output conditions can be set as a set of directly achievable checks, such as simultaneously satisfying: the execution result text content, after removing leading and trailing whitespace, is at least 20 characters long; the execution result text content contains at least one period or semicolon to indicate a complete expression; and the execution status field is "completed". The advantages of this setting are: length checks can filter out extremely short responses like "I don't know," punctuation checks can filter out fragmented responses that only return keywords, and status checks can filter out execution interruptions or timeouts. If the execution result meets the preset valid output conditions, the dynamic routing module increments the success count field by 1 and writes the execution result to the output buffer of the current atomic task. Simultaneously, it records a mapping of "which routing arm completed the current atomic task" within the session object, facilitating the inclusion of execution link descriptions during orchestration layer summarization. If the execution result does not meet the preset valid output conditions, the optional implementation allows the dynamic routing module to perform a remedy without changing the order of the atomic task sequence: select the routing arm with the second highest final score from the remaining routing arms as the new selected routing arm, and send the same atomic task again through the task delegation interface; the number of remedies is set to, for example, 1 or 2 times, to avoid excessive waiting when the candidate set after gating is large. The remedy still only uses the arm state record within the current request execution cycle and does not introduce historical data.
[0055] Once all atomic tasks in the atomic task sequence have been processed, the orchestration layer aggregates the execution results of each atomic task and outputs them to the application layer. During aggregation, the orchestration layer reads the output buffer content of each atomic task in the order of the atomic task sequence and concatenates them into the final output. To make the final output more readable, the orchestration layer can insert separators between adjacent atomic task results, writing the task identifier field of the corresponding atomic task into the separator. If the output buffer of an atomic task is empty, the orchestration layer can optionally insert a placeholder paragraph and write the task identifier field of that atomic task along with a status message indicating "no valid output content produced," thus ensuring a stable output structure and preventing loss of execution traces. Finally, the orchestration layer writes the aggregated output into a response payload that the application layer can directly return. The response payload includes at least the session identifier of the session object and the final output content. Based on this, the application layer sends the response to the corresponding channel and ends the execution cycle of this user request.
[0056] In one specific implementation, the application layer receives the user's original request text: "Please check the logistics progress of order 12345, specify the estimated delivery date," and generates an explanation and delay handling suggestion that can be directly sent to the customer. Based on this, the application layer generates a session object containing a session identifier with the value S202512230001. An output buffer is reserved within the session object, using an "atomic task write" method, with each atomic task corresponding to one write slot, initially empty. The orchestration layer creates a root task node within the session object. The task identifier field of the root task node has a value of T0, the task description field contains the aforementioned user request text, the task type field has a composite type value, and the subtask list field is initialized to an empty list. The orchestration layer performs a recursive decomposition operation on the root task node: sending a decomposition instruction containing the task description field content of the root task node to the large language model. The large language model returns a list of subtask descriptions, which in this example contains three subtask descriptions: parsing order information, querying logistics information, and generating a customer reply. The orchestration layer creates subtask nodes T1, T2, and T3 for these three subtask descriptions, and adds T1, T2, and T3 to the subtask list field of the root task node.
[0057] The orchestration layer continues to determine the task type field of each sub-task node. To ensure subsequent implementation, the orchestration layer sets the task type field according to the standard of "whether further decomposition is needed to extract independently delegable and separately verifiable outputs": T1 is set to atomic type, T2 to composite type, and T3 to composite type. Then, a recursive decomposition operation is performed on T2: a decomposition instruction containing the task description field of T2 is sent to the large language model, and the large language model returns two sub-task descriptions: obtaining the logistics interface return value and summarizing key logistics conclusions. The orchestration layer creates sub-task nodes T2_1 and T2_2, and writes them into the sub-task list field of T2, while setting T2_1 and T2_2 to atomic type. Then, a recursive decomposition operation is performed on T3: a decomposition instruction containing the task description field of T3 is sent to the large language model, and the large language model returns two sub-task descriptions: calculating the expected delivery date expression and generating sendable scripts and delay handling suggestions. The orchestration layer creates subtask nodes T3_1 and T3_2, and writes them to the subtask list field of T3. Simultaneously, it sets T3_1 and T3_2 to atomic type. At this point, there are no longer any nodes with a composite task type field that have not yet been decomposed, and the recursive decomposition operation is complete.
[0058] The orchestration layer traverses all leaf task nodes whose task type field is atomic, and generates an atomic task sequence in traversal order. In this example, the traversal order is T1, T2_1, T2_2, T3_1, T3_2, so the atomic task sequence length is 5, and this atomic task sequence is passed to the confidence gating module.
[0059] In this example, the agent registry has stored 5 agent configuration records. Each record contains an agent identifier field, a capability tag list field, and a tool binding list field, as follows: Agent identifier field: A_order; Capability tag list field: Order parsing, order query, exception handling; Tool binding list field: order_db_query. Agent identifier field: A_logistics; Capability tag list field: Logistics query, API call, exception handling, time estimation; Tool binding list field: logistics_api, geo_time_service. Agent identifier field: A_writer; Capability tag list field: Copywriting generation, tone control; Tool binding list field: empty. Agent identifier field: A_policy; Capability tag list field: After-sales strategy, compliance script, exception handling; Tool binding list field: policy_kb_search. The agent identifier field is A_general, the capability tag list field contains two capability tags: general dialogue and summary, and the tool binding list field is an empty list.
[0060] The confidence gating module performs agent filtering for each atomic task in the atomic task sequence. Taking atomic task T2_1 as an example, the task description field of T2_1 is "retrieving the return value of the logistics interface". The confidence gating module reads all 5 agent configuration records from the agent registry and forms an initial candidate set with 5 elements. Then, the confidence gating module calls the large language model and sends a capability extraction command containing the content of the task description field of T2_1. The large language model returns a list of capability tags; in this example, it returns 4 capability tags: logistics query, interface call, exception handling, and order parsing. The confidence gating module stores this list of capability tags as the required capability set for T2_1.
[0061] A confidence assessment table is created within the session object. This table is a structured collection of records, with each row corresponding to one agent in the initial candidate set. Each row contains an agent identifier column, a capability coverage count column, a capability missing count column, and a confidence pass flag column. The confidence gating module initializes the capability coverage count and capability missing count columns to 0, and initializes the confidence pass flag column to a pending state. Then, a capability matching operation is performed: each capability tag in the required capability set of T2_1 is iterated through, and it is checked whether the capability tag exists in the capability tag list field of the corresponding agent. If it exists, the capability coverage count column for that row is incremented by 1; otherwise, the capability missing count column is incremented by 1.
[0062] Let the set of demand capabilities be Its number of elements is ,in This represents the total number of elements in the demand capacity set. In this example... ,therefore For any intelligent agent Let its capacity coverage count column be... The count of missing abilities is as follows ,in This represents the number of capability tags in the required capability set that can be covered by this intelligent agent. This indicates the number of capability tags in the required capability set that are not covered by this agent.
[0063] For A_order: the capability label list fields are {order parsing, order query, exception handling}. After comparing each item, it covered two items: order parsing and exception handling. Two items are missing: logistics query and API call. For A_logistics: the capability label list fields are {logistics query, API call, exception handling, time estimation}. This covers three items: logistics query, API call, and exception handling. One order parsing item is missing, therefore... For A_writer: the capability tag list field is {copy generation, tone control}. It covers 0 items, therefore... ; 4 items are missing, therefore For A_policy: the capability tag list field is {after-sales strategy, compliance scripts, exception handling}. One exception handling item is covered, therefore... ; 3 items are missing, therefore For A_general: the ability tag list field is {General Dialogue, Summary}. It covers 0 items, therefore... ; 4 items are missing, therefore .
[0064] The confidence gating module then performs a confidence determination operation. This is to avoid the situation where "half the total number of elements" occurs. When the threshold is odd, execution ambiguity arises. In this example, the threshold for determining "half" is defined as... ,in Indicates the minimum coverage threshold used for the decision, symbol This indicates rounding up to the nearest integer. In this example... ,therefore The decision-making rules strictly adhere to the following executable logic: If Then the confidence pass flag column is set to pass status; if Then compare and ,when Set to pass status if the condition is met, otherwise set to reject status.
[0065] Based on this calculation: A_order satisfies and Pass; A_logistics satisfies and Passed; A_writer satisfies , Reject; A_policy Satisfies , Reject; A_general satisfies The confidence gating module filters rows from the confidence evaluation table whose confidence pass column is set to pass, forming the gating candidate set for T2_1. This gating candidate set contains two agents: A_order and A_logistics. T2_1 and its gating candidate set are then passed to the dynamic routing module. The remaining atomic tasks T1, T2_2, T3_1, and T3_2 also obtain their respective gating candidate sets following the same process. This example only demonstrates two more atomic tasks related to the crucial routing and computation.
[0066] The dynamic routing module begins processing T2_1. It defines each agent in the gated candidate set of T2_1 as a routing arm; in this example, two routing arms are formed: arm A_order and arm A_logistics. The dynamic routing module creates an arm state record for each routing arm within the session object. This record includes an arm identifier field, a trial count field, and a success count field. Initially, both arms have a trial count field of 0 and a success count field of 0. The dynamic routing module creates a context feature record within the session object: the first feature slot is filled with the position index 2 of T2_1 in the atomic task sequence; the second feature slot is filled with the number of elements in T2_1's required capability set of 4; the third feature slot is filled with the number of elements in the gated candidate set of 2; and the fourth feature slot is filled with the number of completed atomic tasks of 1, because T1 has already been processed and completed.
[0067] Proceed to the arm score calculation. To fully illustrate the calculation process, this example provides the formula and explains the meaning of each symbol. For any routing arm... The number of capability matches is defined as follows: ,in This represents the number of elements in the intersection of the capability label list field and the required capability set of the agent corresponding to the routing arm. Indicates the number of elements in the set. This represents the set of capabilities required for the current atomic task. This represents the set of capability label list fields for the agent corresponding to the routing arm, after being set up. The available tool count is defined as follows: ,in This indicates the number of elements in the tool binding list field for the agent corresponding to the routing arm. This represents the collection of bound list fields for this tool. The base score is defined as... ,in This represents the base score, obtained by adding the number of matching abilities to the number of available tools. The exploration reward value is defined as follows: ,in Indicates the exploration reward value. This represents a fixed exploration constant; the configuration in this example is... The final score is defined as follows: ,in This represents the final score used for route selection.
[0068] Substitute the values into: for arm A_order Because the intersection is {order parsing, exception handling}; Because the tool's binding list field only contains order_db_query; therefore Since the test count field is 0, ,therefore For arm A_logistics, The intersection is {logistics query, API call, exception handling}; The tools are logistics_api and geo_time_service; therefore The test count field is 0. ,therefore .
[0069] The dynamic routing module selects the routing arm with the highest final score as the selected routing arm, therefore, the selected routing arm is arm A_logistics. The dynamic routing module increments the trial count field of arm A_logistics from 0 to 1, and then sends the task description field content of T2_1 to A_logistics through the task delegation interface of the orchestration layer, while attaching the session identifier S202512230001 and the task identifier T2_1 to the payload. After execution, A_logistics returns the following execution result text: Logistics found: Carrier A, current status: in transit, next station: distribution center, expected delivery in 2 days. The dynamic routing module judges based on preset valid output conditions. For executable purposes, this example defines the preset valid output conditions as three conditions being met simultaneously: the text length after removing leading and trailing whitespace is no less than 20 characters; the text contains at least one period or semicolon; the execution result text content is not empty. The text length is defined as... ,in Indicates the number of characters, symbols This function represents the number of characters in a string. This represents the text content of the execution result. In this example, the text content of the execution result... Since it contains a period, it meets the valid output condition. Therefore, the dynamic routing module increments the success count field of arm A_logistics from 0 to 1 and writes the execution result into the output buffer of T2_1.
[0070] Next, atomic task T3_2 is processed, whose task description field is "generate sendable scripts and delay handling suggestions." After the confidence gating module executes the capability extraction instruction on T3_2, the required capability set returned in this example is {copywriting generation, tone control, after-sales strategy, compliance scripts}. and After performing capability matching on the five agents, the results were as follows: A_writer (covered 2, missing 2, passed); A_policy (covered 2, missing 2, passed); A_general (covered 0, rejected); A_order (covered 0, rejected); A_logistics (covered 0, rejected). Therefore, the gated candidate set of T3_2 only contains A_writer and A_policy.
[0071] The dynamic routing module uses A_writer and A_policy as two routing arms and creates arm status records, initializing both the trial count and success count fields to 0. In this example, the fixed exploration constant remains unchanged. When calculating the score, the number of ability matches for A_writer is... The number of available tools is ,so , , The capability match count for A_policy is The number of available tools is ,so , , Therefore, arm A_policy is selected first. The dynamic routing module updates the trial count field of arm A_policy to 1 and sends T3_2 through the task delegation interface. The execution result text returned by A_policy is: "Suggested explanation and apology." The text length is... Calculated Since the output does not meet the valid output condition of "length not less than 20 characters", the dynamic routing module does not update the success count field, nor does it write to the output buffer of T3_2. Instead, it initiates an optional remedial execution, with the number of remedial attempts configured as 1 in this example.
[0072] During the remedial action, the dynamic routing module recalculates the final score for each routing arm. At this point, the trial count field for arm A_policy is 1, therefore... ,That The test count field of arm A_writer is still 0, therefore ,That Therefore, the remedial measure is to select arm A_writer as the chosen routing arm. The dynamic routing module updates the trial count field of arm A_writer to 1 and sends T3_2 to A_writer. Simultaneously, it appends the key logistics conclusion fragment from the T2_1 output buffer as context to the task description field, ensuring A_writer doesn't lack necessary information. A_writer returns the following execution result text: "Hello, we have checked the logistics information for order 12345 for you. The package is currently in transit, and the next stop is the distribution center. It is expected to arrive in 2 days. If there is a delay due to weather or carrier scheduling, we suggest you keep your phone accessible and monitor logistics updates. If the package is still not signed for after the estimated delivery date, we can register it to urge the carrier and assist in initiating after-sales processing." The dynamic routing module calculates... The text contains periods and semicolons, which meets the valid output conditions. Therefore, the success count field of arm A_writer is incremented from 0 to 1, and the execution result is written to the output buffer of T3_2.
[0073] The remaining atomic tasks in the atomic task sequence are completed in the same manner. After all atomic tasks are completed, the orchestration layer reads the output buffer contents of T1, T2_1, T2_2, T3_1, and T3_2 in the order of the atomic task sequence and performs aggregation. The final aggregated output is written to the application layer's response payload. The application layer returns the response payload to the user, the request processing ends, and the output buffer formed by the session object during this execution cycle is retained in the session object for front-end display and auditing. Meanwhile, the arm state record is no longer used for subsequent requests because its lifecycle is limited to the execution cycle of the current user request.
[0074] This invention is not limited to the specific embodiments described above. The invention extends to any new feature or combination disclosed in this specification, as well as any new method or process step or combination disclosed herein.
Claims
1. A multi-level intelligent agent orchestration system, characterized in that, It includes an application layer, an orchestration layer, an agent registration center, a confidence gating module, a dynamic routing module, and an agent layer; the application layer is used to receive user requests and generate session objects; the orchestration layer is used to create root task nodes based on the session objects and perform recursive decomposition of the HTN hierarchical task network to generate atomic task sequences; The agent registration center is used to store agent configuration records. Each agent configuration record includes an agent identifier field, a capability tag list field, and a tool binding list field. The confidence gating module is used to perform a confidence gating process for the current atomic task in the atomic task sequence. The confidence gating process includes: reading all agent configuration records from the agent registry and forming an initial candidate set; calling a large language model to obtain the capability tag list for the current atomic task; normalizing the aliases in the capability tag list to a standardized spelling based on a capability tag lexicon and forming a required capability set; creating a confidence evaluation table within the current session object, where each row of the confidence evaluation table corresponds to an agent in the initial candidate set, and each row includes an agent identifier column, a capability coverage count column, a capability missing count column, and a confidence pass flag column; assuming the total number of elements in the required capability set is N, the process is as follows: For each agent in the initial candidate set, the capability coverage count and capability missing count are initialized to 0, and the required capability set is traversed. If a capability tag in the traversed required capability set exists in the capability tag list field of the agent, the capability coverage count is incremented by 1; otherwise, the capability missing count is incremented by 1. When the capability missing count is 0, or the capability missing count is greater than 0 and the capability coverage count is greater than or equal to the integer obtained by rounding up half of N, the confidence pass flag column is set to pass status; otherwise, the confidence pass flag column is set to rejection status. Rows with the confidence pass flag column set to pass status are filtered from the confidence evaluation table to form the gated candidate set; the dynamic routing module For performing a dynamic routing process for the current atomic task, the dynamic routing process includes: defining each agent in the gated candidate set as a routing arm, and creating an arm state record for each routing arm within the session object. The arm state record includes an arm identifier field, a trial count field, and a success count field. The trial count field and the success count field are initialized to 0, and their lifecycles are limited to the execution cycle of the current user request. For each routing arm, the number of elements in the intersection of the required capability set and the capability tag list field of the corresponding agent is used as the capability matching number, and the number of elements in the tool binding list field of the corresponding agent is used as the matching number. The quantity is used as the number of available tools, and the sum of the capability matching number and the number of available tools is used as the base score; when the trial count field is 0, the exploration reward value is set to a pre-configured positive integer K, and when the trial count field is greater than 0, the exploration reward value is set to 0, and the sum of the base score and the exploration reward value is used as the final score; the routing arm with the largest final score value is selected as the selected routing arm, the trial count field of the selected routing arm is incremented by 1, and the current atomic task is sent to the agent corresponding to the selected routing arm through the task delegation interface of the orchestration layer; the agent layer is used to execute the current atomic task and return the execution result containing the execution result text content and the execution status field;After receiving the execution result, the dynamic routing module uses the following three conditions as preset valid output conditions: the length of the execution result text content after removing leading and trailing whitespace is not less than 20 characters; the execution result text content contains at least one period or semicolon; the value of the execution status field is "completed"; when the execution result meets the preset valid output conditions, the success count field of the selected routing arm is incremented by 1 and the execution result is written to the output buffer of the current atomic task; when the execution result does not meet the preset valid output conditions, a remedy is performed without changing the order of the atomic task sequence: the routing arm with the second highest final score is selected as the new selected routing arm, and the same atomic task is sent again through the task delegation interface; the orchestration layer is used to read the output buffer content of each atomic task in the atomic task sequence according to the order of the atomic task sequence after all atomic tasks in the atomic task sequence have been processed, summarize the results, and output the summary results to the application layer.
2. The system according to claim 1, characterized in that, The root task node created by the orchestration layer contains a task identifier field, a task description field, a task type field, and a subtask list field. The orchestration layer writes the original text of the user request into the task description field and initializes the subtask list field to an empty list. The orchestration layer performs a recursive decomposition operation on the root task node. The recursive decomposition operation includes: calling the large language model and sending a decomposition instruction containing the task description field content of the current task node, and receiving the subtask description list returned by the large language model. For each subtask description in the subtask description list, create a subtask node and add it to the subtask list field of the current task node; for subtask nodes with a task type field of composite type, continue to perform recursive decomposition operation, and end the decomposition for subtask nodes with a task type field of atomic type; after the recursive decomposition operation is completed, traverse the leaf task nodes with a task type field of atomic type to generate an atomic task sequence, and pass the atomic task sequence to the confidence gating module.
3. The system according to claim 2, characterized in that, The confidence gating module calls the large language model and sends a capability extraction instruction containing the task description field of the current atomic task. It receives the capability tag list returned by the large language model, normalizes the aliases in the capability tag list into the standard spelling according to the capability tag synonym list, and stores the normalized capability tag list as the required capability set of the current atomic task.
4. The system according to claim 3, characterized in that, The confidence gating module initializes the capability coverage count column and the capability missing count column to preset initial values and initializes the confidence to a pending state through the tag column; the confidence gating module performs a capability matching operation to update the capability coverage count column and the capability missing count column. The capability matching operation includes traversing each capability tag in the required capability set and checking the existence of the capability tag in the capability tag list field of the corresponding agent. The confidence gating module performs a confidence determination operation, which includes: setting the confidence pass flag column to the pass state when the capability missing count column meets the preset missing condition; and setting the confidence pass flag column to the pass state or the rejection state according to the preset ratio between the capability coverage count column and the total number of elements in the required capability set when the capability missing count column does not meet the preset missing condition. The confidence gating module filters rows with the confidence pass flag column in the pass state from the confidence evaluation table to form a gating candidate set, and passes each atomic task and its corresponding gating candidate set to the dynamic routing module.
5. The system according to claim 4, characterized in that, The dynamic routing module creates a context feature record within the current session object. The context feature record contains a first feature slot, a second feature slot, a third feature slot, and a fourth feature slot. The first feature slot stores the position index of the current atomic task in the atomic task sequence. The second feature slot stores the number of elements in the required capability set. The third feature slot stores the number of elements in the gating candidate set. The fourth feature slot stores the number of atomic tasks that have been completed within the current session object.
6. A dynamic routing method for a multi-level intelligent agent orchestration system, characterized in that, The system according to any one of claims 1 to 5 performs the following steps: receiving a user request and generating a session object; creating a root task node based on the session object and performing recursive decomposition of the HTN hierarchical task network to obtain an atomic task sequence; for the current atomic task in the atomic task sequence, a confidence gating module reads all agent configuration records from the agent registration center and forms an initial candidate set, wherein the agent registration center stores agent configuration records, and each agent configuration record includes an agent identifier field, a capability tag list field, and a tool binding list field; the confidence gating module calls a large language model to obtain the capability tag list of the current atomic task, and sorts the capability tags according to the capability tag lexicon. The aliases in the list are normalized to a standardized notation and form a set of required capabilities. A confidence evaluation table is created within the current session object. Each row of the confidence evaluation table corresponds to an agent in the initial candidate set, and each row contains an agent identifier column, a capability coverage count column, a capability missing count column, and a confidence pass flag column. Assuming the total number of elements in the required capability set is N, for each agent in the initial candidate set, the capability coverage count and capability missing count are initialized to 0, and the required capability set is iterated. If a capability tag in the iterated required capability set exists in the capability tag list field of the agent, the capability coverage count is incremented by 1; otherwise, the capability missing count is incremented by 1. Confidence gating is performed by the confidence gating module. The confidence gating process includes: when the capability missing count is 0, or when the capability missing count is greater than 0 and the capability coverage count is greater than or equal to the integer obtained by rounding up half of N, setting the confidence pass flag column to a pass state; otherwise, setting the confidence pass flag column to a rejection state; filtering rows with the confidence pass flag column set to a pass state from the confidence evaluation table to form a gating candidate set; and executing a dynamic routing process by the dynamic routing module, which includes: defining each agent in the gating candidate set as a routing arm, and creating an arm state record for each routing arm within the session object, the arm state record including an arm identifier field, a trial count field, and a success count field. The trial count field and the success count field are initialized to 0, and their lifecycles are limited to the execution cycle of the current user request. For each routing arm, the number of elements in the intersection of the required capability set and the capability tag list field of the corresponding agent is used as the capability matching number, the number of elements in the tool binding list field of the corresponding agent is used as the tool available number, and the sum of the capability matching number and the tool available number is used as the base score. When the trial count field is 0, the exploration reward value is set to a pre-configured positive integer K; when the trial count field is greater than 0, the exploration reward value is set to 0, and the sum of the base score and the exploration reward value is used as the final score.The routing arm with the highest final score is selected as the selected routing arm. The trial count field of the selected routing arm is incremented by 1, and the current atomic task is sent to the agent corresponding to the selected routing arm. After the dynamic routing module receives the execution result returned by the agent, which includes the execution result text content and the execution status field, it uses the following three conditions as preset valid output conditions: the length of the execution result text content after removing leading and trailing whitespace is no less than 20 characters; the execution result text content contains at least one period or semicolon; the value of the execution status field is "completed". When the execution result meets the preset valid output conditions, the success count field of the selected routing arm is incremented by 1, and the execution result is written to the output buffer of the current atomic task. When the execution result does not meet the preset valid output conditions, a remedial action is performed without changing the order of the atomic task sequence: the routing arm with the second highest final score from the remaining routing arms is selected as the new selected routing arm, and the same atomic task is sent again through the task delegation interface. After all atomic tasks in the atomic task sequence have been processed, the output buffer contents of each atomic task are read in the order of the atomic task sequence, summarized, and then output.
7. The method according to claim 6, characterized in that, The recursive decomposition operation includes: sending a decomposition instruction to the large language model for the current task node and receiving a list of subtask descriptions; creating subtask nodes one by one for each subtask description and writing them into the subtask list field; classifying the subtask nodes into composite types or atomic types based on the task type field, and continuing to perform the recursive decomposition operation for composite type subtask nodes; after the recursive decomposition operation is completed, traversing all atomic type leaf task nodes to generate an atomic task sequence, the order of which is determined by the traversal order.
8. The method according to claim 6, characterized in that, The agent selection operation for the current atomic task includes: calling the large language model and sending a capability extraction command to obtain a capability tag list; normalizing the aliases in the capability tag list to a standard spelling based on the capability tag lexicon; and storing the normalized capability tag list as a required capability set; initializing the capability coverage count column, the capability missing count column, and the confidence pass flag column; and for each agent, traversing the required capability set and updating the capability coverage count column or the capability missing count column based on the existence of the capability tag in the capability tag list field.
9. The method according to claim 8, characterized in that, The confidence determination operation includes: reading the capability coverage count column and capability missing count column of the current row in the confidence assessment table; setting the confidence pass flag column to pass status when the capability missing count column is at a preset missing threshold value; comparing the capability missing count column with a preset ratio threshold of the total number of elements in the required capability set when the capability missing count column is greater than the preset missing threshold value, and setting the confidence pass flag column to pass status or rejection status accordingly; filtering rows with the confidence pass flag column in pass status from the confidence assessment table and forming a gated candidate set.
10. The method according to claim 6, characterized in that, The routing operation includes creating a context feature record containing feature slots 1 through 4.
Citation Information
Patent Citations
Agent-driven pulverized coal boiler combustion optimization closed-loop control method and system
CN120386252A
Task processing method and device based on multiple Agents and related medium
CN121029370A
Intelligent agent automatic arrangement method and system based on large language model
CN121212278A