An agent task decision method, device, equipment and medium
Patent Information
- Application Number
- CN202611134197.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-29
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]然而,纯形式化验证要求规格完整表达,仅适用于世界可精确建模的场景,对规格缺失的开放性任务无法处理,验证失败时仅能拒绝而无替代方案;主流Agent框架完全依赖LLM自身判断保障安全性,不具备可证明的硬性底线约束;两类方案彼此割裂,形式化验证与神经网络输出之间存在语义鸿沟,无法在同一框架下实现统一裁决与协同决策
[0024]本发明实施例的技术方案,通过多维评估和任务象限分类实现任务差异化处理,实现精准策略匹配;通过目标导向引擎的高阶规则集保证智能体对模糊开放任务的灵活适应性,通过底线保障引擎的分层校验提供可数学证明的硬性行为保障,再通过仲裁机制实现底线约束与价值目标的协同统一,解决了现有技术中形式化验证与神经网络决策相互割裂、底线与目标无法兼容的技术问题,实现了高风险场景下兼顾安全与效能的智能体决策。
Smart Images

Figure CN122838560A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a task decision-making method, apparatus, device, and medium for an intelligent agent. Background Technology
[0002] With the rapid maturation of large language model technology, AI agents have demonstrated powerful capabilities in autonomous planning and multi-step execution, and are rapidly expanding from routine tasks to high-risk professional fields such as healthcare, finance, and industrial control. However, AI agents face a fundamental challenge in these scenarios: how to maintain flexible adaptability to complex, dynamically changing high-level value objectives while ensuring that behavioral bottom lines are absolutely inviolable.
[0003] The closest existing technologies fall into three categories: purely formal verification methods, which focus on model detection, theorem proof, and runtime verification. These require specifications to be fully described in a formal language, followed by an exhaustive search by the verification engine to output a binary judgment. Mainstream AI Agent frameworks, represented by ReAct and AutoGPT, use LLM to generate action plans through thought chain reasoning and call tools for execution, relying on model self-reflection to correct errors. Alignment methods such as RLHF encode safety preferences into the reward model and inject alignment targets into the model parameters during the training phase.
[0004] However, pure formal verification requires complete specification expression and is only applicable to scenarios where the world can be accurately modeled. It cannot handle open tasks with missing specifications, and can only reject verification without alternatives when it fails. Mainstream agent frameworks rely entirely on the LLM's own judgment to ensure security and do not have provable hard bottom-line constraints. The two types of solutions are disconnected from each other, and there is a semantic gap between formal verification and neural network output, making it impossible to achieve unified adjudication and collaborative decision-making under the same framework. Summary of the Invention
[0005] This invention provides a task decision-making method, apparatus, device, and medium for intelligent agents, enabling intelligent agent decision-making that balances safety and efficiency in high-risk scenarios.
[0006] According to a first aspect of the present invention, a task decision-making method for an intelligent agent is provided, comprising:
[0007] Obtain the original task request, perform multi-dimensional evaluation on the original task request, and generate a task context record;
[0008] Based on the task evaluation results, the task quadrant category to which the original task request belongs and the initial weight of the dual engines corresponding to the task quadrant category are determined. The initial weight of the dual engines includes the first weight value of the bottom-line guarantee engine and the second weight value of the goal-oriented engine.
[0009] Based on the task quadrant categories, the goal-oriented engine generates a set of candidate execution actions in combination with a preset high-order rule set;
[0010] Based on the task context record, the candidate action set is subjected to hierarchical verification through the bottom-line protection engine to obtain the verification decision result.
[0011] The final execution action is determined based on the verification decision result, the initial weights of the dual engines, and the set of candidate execution actions.
[0012] According to a second aspect of the present invention, a task decision-making device for an intelligent agent is provided, comprising:
[0013] The task evaluation module is used to obtain the original task request, perform multi-dimensional evaluation on the original task request, and generate a task context record.
[0014] The quadrant determination module is used to determine the task quadrant category to which the original task request belongs and the initial weight of the dual engines corresponding to the task quadrant category based on the task evaluation results. The initial weight of the dual engines includes a first weight value of the bottom-line guarantee engine and a second weight value of the goal-oriented engine.
[0015] The action generation module is used to generate a set of candidate execution actions based on the task quadrant category, through the goal-oriented engine and a preset high-order rule set;
[0016] The hierarchical verification module is used to perform hierarchical verification on the candidate execution action set based on the task context record and through the bottom-line guarantee engine to obtain the verification decision result.
[0017] The action determination module is used to determine the final action to be executed based on the verification decision result, the initial weights of the dual engines, and the set of candidate actions to be executed.
[0018] According to a third aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0019] At least one processor; and
[0020] A memory communicatively connected to the at least one processor; wherein,
[0021] The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the task decision-making method of the intelligent agent according to any embodiment of the present invention.
[0022] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the task decision-making method of an intelligent agent according to any embodiment of the present invention.
[0023] According to a fifth aspect of the present invention, embodiments of the present invention also provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the task decision-making method of an intelligent agent according to any embodiment of the present invention.
[0024] The technical solution of this invention achieves differentiated task processing and precise strategy matching through multi-dimensional evaluation and task quadrant classification; it ensures the agent's flexible adaptability to fuzzy open tasks through the high-order rule set of the goal-oriented engine; it provides mathematically provable hard behavioral guarantees through the hierarchical verification of the bottom-line guarantee engine; and it achieves the synergistic unity of bottom-line constraints and value goals through an arbitration mechanism. This solves the technical problems of formal verification and neural network decision-making being mutually exclusive and bottom-line and goals being incompatible in the prior art, and realizes agent decision-making that balances safety and efficiency in high-risk scenarios.
[0025] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0026] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0027] Figure 1 This is a flowchart of a task decision-making method for an intelligent agent according to Embodiment 1 of the present invention;
[0028] Figure 2 This is a flowchart of a task decision-making method for an intelligent agent according to Embodiment 2 of the present invention;
[0029] Figure 3 This is a schematic diagram of the structure of a task decision-making device for an intelligent agent according to Embodiment 3 of the present invention;
[0030] Figure 4 This is a schematic diagram of the structure of an electronic device that implements an embodiment of the present invention. Detailed Implementation
[0031] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0032] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0033] Example 1
[0034] Figure 1 This is a flowchart illustrating a task decision-making method for an intelligent agent according to Embodiment 1 of the present invention. This embodiment is applicable to intelligent decision-making situations of intelligent agents. The method can be executed by a task decision-making device of the intelligent agent, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:
[0035] S110. Obtain the original task request, perform multi-dimensional evaluation on the original task request, and generate a task context record.
[0036] In this embodiment, the original task request refers to the task to be processed submitted by the user or upstream system to the intelligent agent system, including natural language descriptions or structured specification documents, optionally with accompanying specification documents and domain knowledge base references. The task context record can be understood as a structured data package output after the multidimensional evaluation is completed, containing fields such as task semantic information, scores for each dimension, comprehensive formalizable index, and risk level, serving as a unified input basis for subsequent dual-engine execution.
[0037] Specifically, the intelligent agent can receive raw task requests submitted by users or upstream systems. The task parser calls a large language model to perform structured information extraction on the task text, extracting a five-tuple of information including task type label, input / output constraint description, success criterion description, domain label, and risk sensitivity annotation, forming a task semantic record. A multidimensional evaluator quantifies and scores the task's formalizability from multiple dimensions. The multidimensional scores are weighted and summed according to preset weights to output a comprehensive formalizability index, which, along with the task semantic record and multidimensional scores, is written into the task context storage to generate a task context record.
[0038] S120. Based on the task evaluation results, determine the task quadrant category to which the original task request belongs and the initial weight of the dual engines corresponding to the task quadrant category. The initial weight of the dual engines includes the first weight value of the bottom-line guarantee engine and the second weight value of the goal-oriented engine.
[0039] In this embodiment, the task quadrant category refers to the classification result of assigning tasks to one of the four quadrants based on specification clarity score and world modelability score, used for differentiated configuration processing strategies. The dual-engine initial weights refer to the initial decision weight values allocated to the Formal Baseline Engine and the High-Order Objective Engine, respectively. Specifically, the Formal Baseline Engine (FBE) is a computational subsystem with formal four-layer hierarchical verification at its core, performing provable constraint verification on candidate actions. The High-Order Objective Engine (HOE) is a computational subsystem with computational encoding based on a high-order generation mechanism at its core, generating candidate action sequences based on task objectives.
[0040] Specifically, the agent can compare scores with preset thresholds, and based on the comparison results, perform four-quadrant classification to determine the task quadrant category to which the original task request belongs. The agent can then output initial weight pairs for the dual engines based on the quadrant classification results and the risk sensitivity labels in the task evaluation results, obtaining the first weight value for the bottom-line guarantee engine and the second weight value for the goal-oriented engine. Before subsequent dual-engine execution, the input task undergoes structured parsing and a quantitative assessment of its formalizability, outputting the task quadrant assignment and initial dual-engine weight configuration, providing a basis for differentiated resource allocation in subsequent stages.
[0041] S130. Based on the task quadrant category, a set of candidate execution actions is generated by combining a goal-oriented engine with a preset high-level rule set.
[0042] In this embodiment, the preset high-order rule set refers to the set of executable rules formed by computerizing and encoding the nine high-order order generation mechanisms that maintain effective order in human society under conditions of fuzzy specifications and incomplete modeling. These may include, for example, progressive specification emergence (online learning + specification refinement), social contracts and consensus (distributed protocols + implicit rule encoding), fault tolerance and self-repair (transaction rollback + degradation strategy + alternative path generation), hierarchical delegation and trust chains (graph structure trust propagation + layered verification delegation), fuzzy reasoning and satisfaction (Satisficing threshold setting + approximate reasoning), feedback loops and continuous adaptation (reinforcement learning framework + online updates), narrative and meaning construction (intent understanding + value goal representation), multi-agent game and dynamic equilibrium (game theory equilibrium calculation + emergent coordination), and embodiedness and situational intelligence (long-term memory + situational encoding + experience accumulation). The candidate action set refers to multiple candidate action sequences generated by the goal-oriented engine based on high-order rules and task objectives, with each candidate action carrying a satisfaction score for arbitration selection.
[0043] Specifically, the intelligent agent can encode the user's high-level value objective (fuzzy description in natural language) into a structured objective decomposition tree, including three levels: top-level value objective, mid-level sub-objectives, and bottom-level verifiable sub-constraints. Based on task quadrant categories, the agent can determine a subset of rules to be activated from a pre-set set of higher-order rules that match the current task. Based on the objective decomposition tree and the selected subset of rules to be activated, the agent can invoke a large language model to generate multiple candidate action sequences while satisfying the value objective. It can then calculate a satisfaction score for each candidate action and sort them according to the satisfaction score to obtain a set of candidate actions to be executed.
[0044] S140. Based on the task context record, the candidate execution action set is checked in layers through the bottom-line guarantee engine to obtain the check decision result.
[0045] In this embodiment, the verification decision result refers to the verification conclusion output by the bottom-line protection engine after performing layered verification on the candidate actions.
[0046] Specifically, the agent can determine the combination of verification layers corresponding to the current task based on the risk level recorded in the task context, start the goal-oriented engine to perform serial verification of the candidate execution actions in the candidate execution action set of the verification layer combination, and then determine the verification decision result for each candidate execution action.
[0047] S150. Based on the verification decision results, the initial weights of the dual engines, and the set of candidate execution actions, determine the final execution action.
[0048] In this embodiment, the final execution action refers to the action with the best comprehensive score selected within the feasible domain that satisfies all bottom-line constraints after being adjudicated by the dual-engine arbitration bus according to a strict priority protocol, which is the final output of the present invention.
[0049] Specifically, the agent can execute the absolute coverage rule of the bottom-line guarantee engine, comparing each action in the candidate action set output by the goal-oriented engine with the verification and decision results of the bottom-line guarantee engine. It then identifies the candidate actions that were not rejected in the verification and decision results. A comprehensive score is calculated based on the satisfaction score of the candidate actions, the confidence level in the verification and decision results, and the initial weights of the two engines. The candidate action with the highest comprehensive score is selected as the final action. If there are no unrejected actions in the candidate action set, the bottom-line guarantee engine generates a set of alternative actions based on the verification and decision results. A comprehensive score is then calculated on the alternative action set, and the alternative action with the highest comprehensive score is selected as the final action.
[0050] The technical solution of this invention achieves differentiated task processing and precise strategy matching through multi-dimensional evaluation and task quadrant classification; it ensures the agent's flexible adaptability to fuzzy open tasks through the high-order rule set of the goal-oriented engine; it provides mathematically provable hard behavioral guarantees through the hierarchical verification of the bottom-line guarantee engine; and it achieves the synergistic unity of bottom-line constraints and value goals through an arbitration mechanism. This solves the technical problems of formal verification and neural network decision-making being mutually exclusive and bottom-line and goals being incompatible in the prior art, and realizes agent decision-making that balances safety and efficiency in high-risk scenarios.
[0051] As a first optional embodiment of this embodiment, after determining the final execution action based on the verification decision result, the initial weights of the dual engines, and the candidate execution action set, the following is also included:
[0052] The execution process of the final action is monitored, historical execution records are generated, and the preset dual-engine weights and preset high-order rule sets are calibrated.
[0053] In this embodiment, historical execution records refer to structured historical data generated after the agent executes the task according to the final execution action, including fields such as task type, action sequence, FBE adjudication result, HOE satisfaction score, user feedback, and execution result status (success / abnormal termination + reason). Preset dual-engine weights refer to the pre-defined dual-engine weights under different quadrants.
[0054] Specifically, the agent can acquire the runtime state flow of the final action during execution, continuously monitor key state variables during action execution, provide real-time warnings for state changes that trigger L2-level constraints, and immediately interrupt and trigger FBE re-verification for state changes that trigger L3-level constraints. When a violation of the bottom-line constraint is detected during runtime, the handling is escalated according to the severity of the violation: low-level violations are handled by logging and dynamic correction; medium-level violations are handled by pausing execution and generating alternative paths; high-level violations are handled by immediate termination, full state rollback, and manual intervention, obtaining the execution result (success / abnormal termination and reason) and writing it into the historical execution record. The agent can aggregate historical records using a sliding time window (default 90 days) to identify action patterns that repeatedly trigger the same FBE rejection rule (frequency threshold: occurrence ≥ 10 times in the same type of task); for the identified high-frequency patterns, common constraints are extracted to generate rule candidate records, with the structure schema being {pattern_id, trigger scenario description, constraint set, confidence score, number of source samples}, resulting in a set of rule candidate records. The agent can receive rule candidate records (input) through the Rule Formalizer, invoke the Natural Language Transformation Formalization (NLT) pipeline to transform the constraint set into executable formal rules (LTL / STL formulas or SMT constraints) at the L2 or L3 layer, and generate rule records with {formal rule expression, applicable level, applicable quadrant, and confidence level}. Rule candidates with a confidence level below the threshold (default 0.75) are submitted to the manual review queue. The agent can also use the FBE rule integrator to receive rule records (input) that pass the threshold, perform rule deduplication (equivalence with the existing FBE rule set computational logic), and rule conflict detection (whether the new rule contradicts the existing rules). After these steps, the agent writes the new rule into the corresponding level's FBE rule base; simultaneously, it updates the quadrant weight configuration (if the new rule significantly changes the Q-quadrant distribution, it triggers weight recalibration), resulting in an updated FBE rule base and calibrated dual-engine weights.
[0055] The technical solution of this invention automatically extracts stable behavior patterns from historical execution records, formally transforms them, integrates them into the FBE rule base, and dynamically recalibrates the dual-engine weights based on changes in rule coverage, thus achieving self-evolution of the system during continuous operation. This module effectively solves the technical problem of existing technologies lacking specification lifecycle management and self-evolution capabilities, enabling the system to gradually reduce its reliance on LLM self-judgment and increase its reliance on formal rules when handling similar tasks, achieving a continuous evolution loop of "experience → rules → bottom line".
[0056] Example 2
[0057] Figure 2This is a flowchart of a task decision-making method for an intelligent agent provided in Embodiment 2 of the present invention. This embodiment is a further refinement of the above embodiment. Figure 2 As shown, the method includes:
[0058] S201. Obtain the original task request, extract information from the original task request, and determine the task semantic record.
[0059] In this embodiment, the task semantic record refers to the structured quintuple information output after information extraction, which includes task type label, input and output constraint description, success judgment criterion description, domain label and risk sensitivity annotation, and is a structured semantic representation of the original task request.
[0060] Specifically, the intelligent agent invokes the Large Language Model (LLM) to perform structured information extraction on the task text, extracting a five-tuple of {task type label, input / output constraint description, success judgment criterion description, domain label, and risk sensitivity label}; if the task request is a structured specification document, it is directly parsed to obtain the task semantic record.
[0061] S202. Determine the first score of the task semantic record in the specification clarity dimension.
[0062] In this embodiment, the specification clarity dimension refers to a quantitative dimension for evaluating the accuracy of task specifications, used to measure the degree to which constraints, input / output interfaces, and success criteria in the task description are clearly defined. The first score refers to the score obtained after quantitatively evaluating the task semantic record on the specification clarity dimension, with a value ranging from 0 to 1, reflecting the accuracy of the task specifications.
[0063] Specifically, the intelligent agent can quantitatively evaluate the formalizability of the task from the dimension of specification clarity to obtain the first score.
[0064] For example, an agent can evaluate the precision of the task semantic record specification according to a preset specification clarification metric: Hoare triple / type signature format = 1.0, JSON Schema / checklist format = 0.75, detailed natural language requirements = 0.5, fuzzy natural language instructions = 0.25, no specification at all = 0.0; Output: Specification Clarity Score .
[0065] S203. Determine the second score of the task semantic record in the world modelability dimension.
[0066] In this embodiment, the world modelability dimension refers to a quantitative dimension that assesses the degree to which the domain involved in the task can be accurately modeled. It is used to measure whether the external world knowledge on which the task depends can be formally expressed and computerized. The second score refers to the score obtained after quantitatively evaluating the semantic records of the task on the world modelability dimension. The value ranges from 0 to 1, reflecting the degree to which the domain involved in the task can be modeled.
[0067] Specifically, intelligent agents can quantitatively evaluate the formalizability of a task from the dimension of world modelability to obtain a second score.
[0068] For example, an intelligent agent can query a pre-built domain modelability knowledge base (a mapping table of domain labels to modelability scores, e.g., programming language semantics = 1.0, mathematical axiomatic systems = 0.95, enterprise knowledge base / regulations = 0.6, medical guidelines = 0.5, public opinion / value system = 0.1) based on task semantic records and domain labels (input), and output: world modelability score. .
[0069] S204. Determine the third score of the task semantic record in the strictness dimension of the judgment criteria.
[0070] In this embodiment, the judgment criterion strictness dimension refers to the quantitative dimension of the highest verification strictness achievable by the assessment task, used to measure what level of verification strength the task can support. The third score refers to the score obtained after quantitatively evaluating the semantic records of the task on the judgment criterion strictness dimension, with a value ranging from 0 to 1, reflecting the highest judgment criterion strictness level achievable by the current task.
[0071] Specifically, the agent can quantitatively evaluate the formalizability of the task from the dimension of criterion rigor to obtain a second score.
[0072] For example, the agent can receive Sc and Sw (inputs) and, based on the rule max_rigor = f(Sc, Sw) (specifically: when Sc>0.7 and Sw>0.7, max_rigor reaches the formal proof level; when Sc>0.5 and Sw>0.5, max_rigor reaches the constraint verification level; otherwise, max_rigor drops to the evidence support level or statistical confidence level), calculate the highest achievable judgment criterion strictness. .
[0073] S205. Based on the first score, the second score, and the third score, determine the comprehensive formalizability index of the input task.
[0074] In this embodiment, the comprehensive formalizable index refers to the comprehensive formalizable index obtained by weighting and summing the first score, the second score, and the third score according to preset weights.
[0075] Specifically, the agent can perform a weighted summation of the first, second, and third scores based on preset weights to obtain the three-dimensional Composite Formalizability Index (CFI).
[0076] For example, the Comprehensive Formalizable Index (CFI) can be calculated using the following formula:
[0077]
[0078] Where α, β, and γ are weight parameters (default α=0.4, β=0.35, γ=0.25). The weights can be determined by empirical values or by parameter learning methods based on algorithms. The output is the CFI value and the three-dimensional score tuple.
[0079] S206. Generate a task context record based on the task semantic record, the first score, the second score, and the third score.
[0080] Specifically, the agent can write the first, second, and third scores of the task semantic record into the task context record for storage.
[0081] S207. Based on the task evaluation results, determine the task quadrant category to which the original task request belongs and the initial weight of the dual engines corresponding to the task quadrant category.
[0082] For example, based on the two axes of specification clarity (Sc) and world modelability (Sw), tasks are automatically divided into four quadrants: Q1 (specification clarity + world modelability, strong formal verification), Q2 (specification clarity + world openness, reference constraints + evidence chain verification), Q3 (specification fuzziness + world modelability, local constraint verification + higher-order rule dominance), and Q4 (specification fuzziness + world openness, higher-order rule dominance + value engine), and a differentiated dual-engine weight reference configuration is given for each quadrant. For example, quadrant partitioning is performed based on thresholds θ_spec=0.6 (specification clarity boundary) and θ_world=0.6 (world modelability boundary): Sc≥θ_spec and Sw≥θ_world→Q1; Sc≥θ_spec and Sw<θ_world→Q2; Sc<θ_spec and Sw≥θ_world→Q3; Sc<θ_spec and Sw<θ_world→Q4; Output: quadrant labels ∈{Q1,Q2,Q3,Q4}, receiving quadrant labels and risk sensitivity annotations (input), based on the risk sensitivity annotations (fields extracted from task semantics in step 101), outputting the initial dual-engine weight pair (w_FBE, w_HOE): Q1: Risk sensitivity marked as high risk → (0.95, 0.05), Risk sensitivity marked as medium risk → (0.70, 0.30); Q2: Risk sensitivity marked as high risk → (0.70, 0.30), Risk sensitivity marked as medium risk → (0.60, 0.40); Q3: Risk sensitivity marked as high risk → (0.50, 0.50), Risk sensitivity marked as medium risk → (0.35, 0.65); Q4: Risk sensitivity marked as high risk → (0.50, 0.50), Risk sensitivity marked as medium risk → (0.10, 0.90).
[0083] The technical solution of this invention quantifies and evaluates tasks from three dimensions: specification clarity, world modelability, and the rigor of judgment criteria. Based on the evaluation results, tasks are divided into four quadrant categories, achieving accurate identification and classification of differentiated tasks. This ensures that the subsequent dual-engine verification strength and rule activation strategy are precisely matched with task characteristics. This module effectively solves the technical problems of excessively high verification costs for low-risk tasks and insufficient protection for high-risk tasks caused by the "one-size-fits-all" approach in existing technologies. By incorporating economic constraints as explicit variables into the system design, it achieves refined scheduling of verification resources.
[0084] S208. Construct a goal decomposition tree based on the user value goals recorded in the task context using a goal-oriented engine.
[0085] In this embodiment, the user value goal refers to the ultimate purpose of the task described by the user in natural language or a structured manner, and is an important component of the task semantic record. The goal decomposition tree refers to encoding the user's high-level value goal (fuzzy description in natural language) into a structured hierarchical tree representation, including three levels: top-level value goal, middle-level sub-goals, and bottom-level verifiable sub-constraints.
[0086] Specifically, the intelligent agent can receive the user's value goal description from the task context record through the value goal parser, transforming this vague, high-level value orientation into a structured hierarchical tree representation. The goal decomposition tree contains three levels: Top-level value goal: the most abstract expression of the user's core demands, i.e., "what fundamental problem does the user ultimately hope to solve?" For example, "achieving steady asset appreciation while ensuring principal safety." Middle-level sub-goals: several executable sub-steps that must be completed to achieve the top-level value goal, organized according to logical dependencies. Bottom-level verifiable sub-constraints: quantifiable success criteria corresponding to each middle-level sub-goal, serving as measurable evidence of whether the sub-goal has been achieved. The generated goal decomposition tree is written into the context of the goal-oriented engine.
[0087] S209. Based on the task quadrant category, determine the target high-order mechanism subset from the preset high-order rule set.
[0088] In this embodiment, the target higher-order mechanism subset refers to a specific combination of rules that are activated differently from the nine higher-order order generation mechanisms according to the task quadrant category. Different quadrants correspond to different subsets of activation rules.
[0089] Specifically, the goal-oriented engine receives the goal decomposition tree and task context records (input), and activates and executes nine high-order rule modules in a two-layer structure: Layer 1: Global Rules (always active in any quadrant): The fault tolerance and self-healing module and the feedback loop and continuous adaptation module are forcibly activated in all tasks. The fault tolerance and self-healing module monitors abnormal states during candidate action generation and outputs degradation strategies and alternative path suggestions; the feedback loop and continuous adaptation module collects intermediate execution signals for the current task and outputs goal adjustment suggestions. The underlying algorithms of the above two mechanisms (transaction rollback, online update) are existing technologies and will not be elaborated upon. Layer Two: On-Demand Rules (Activated on demand based on quadrant labels and task semantic labels): The goal-oriented engine reads the quadrant labels and task semantic labels in the task context and activates the corresponding modules according to the following rules. Multiple subsets of higher-order target mechanisms (hereinafter referred to as modules) can be activated simultaneously for the same task: Q1 / Q2 Quadrant Activation: Hierarchical Delegation and Trust Chain Module (outputs delegation path and hierarchical verification suggestions); Q2 / Q3 / Q4 Quadrant Activation: Progressive Specification Emergence Module (outputs specification refinement suggestions); Q3 / Q4 Quadrant Activation: Satisfaction Module (outputs satisfaction threshold and approximate reasoning strategy), Narrative and Meaning Construction Module (outputs deep analysis of user intent), Social Contract and Consensus Module (outputs multi-party coordination suggestions); Q4 Quadrant Activation: Embodied and Contextual Intelligence Module (outputs contextualized experience matching results); Additional activation when the task semantic label contains multi-Agent tags: Multi-Agent Game and Dynamic Equilibrium Module (outputs game equilibrium strategy suggestions). Each module is coded as an independent computing unit and executed in parallel to output candidate strategy suggestions. The outputs of all activated modules are weighted and aggregated (with weights consistent with w_HOE) to form candidate action generation guidelines, which are then passed to the candidate action generator.
[0090] S210. Based on the target decomposition tree and the target higher-order mechanism subset, generate at least one candidate action that satisfies the value target constraint.
[0091] In this embodiment, the value objective constraint refers to the structured constraints from the user's high-level value objectives that the Goal-Oriented Engine (HOE) must satisfy when generating candidate actions.
[0092] Specifically, the goal-oriented engine receives two inputs: a goal decomposition tree and a subset of higher-order goal mechanisms. It then calls a large language model to generate candidate action sequences while satisfying the value objective. The generation process is constrained by the following conditions: Positive constraints: Candidate actions must serve at least one mid-level sub-goal in the goal decomposition tree and ultimately point to the top-level value objective; actions unrelated to or deviating from the user's value objective cannot be generated. Rule constraints: Candidate actions must conform to the strategy suggestions output by each activated rule module in the higher-order goal mechanism subset. For example, if the "satisfaction is sufficient" mechanism is activated, the generated candidate actions do not need to exhaustively compare all possibilities; instead, they can be used as candidate outputs when a preset satisfaction threshold is reached (e.g., goal completion ≥ 80%). Quantitative constraints: The candidate action generator generates N candidate actions by default (N=5), each containing an action description, expected effect, and associated parameters. After generation, the N candidates can be initially ranked according to their value score. The value score is calculated based on the matching degree of the goal decomposition tree—candidate actions with higher alignment to the top-level value objective rank higher.
[0093] S211. Calculate the satisfaction score for each candidate action, sort the candidate actions according to the satisfaction score, and output the set of candidate actions with satisfaction scores.
[0094] In this embodiment, the satisfaction score refers to the comprehensive score obtained after quantitatively evaluating the candidate execution action from three dimensions: functional satisfaction, value creation, and goal alignment.
[0095] Specifically, a satisfaction score (SA) is calculated for each candidate action using a goal-oriented engine. The candidate actions are then sorted according to the satisfaction scores, and a set of candidate actions carrying the satisfaction scores is output.
[0096] For example, the formula for calculating satisfaction rating is as follows:
[0097] SA = ω1 × Functional Satisfaction + ω2 × Value Creation + ω3 × Goal Alignment
[0098] Wherein, ω1, ω2, and ω3 are weight coefficients, satisfying ω1+ω2+ω3=1, with a suggested default value of (0.4, 0.3, 0.3), which can be configured according to the domain (e.g., ω1 can be increased in medical scenarios). Functional satisfaction: Matching the expected output of the candidate action with the bottom-level sub-constraints of the target decomposition tree one by one, calculating the satisfaction ratio, is a structured rule calculation and does not rely on LLM. Value creation degree: LLM scores the degree of fit between the candidate action and the top-level value target, outputting a floating-point number of [0,1]; LLM scoring is existing technology and will not be elaborated. Target alignment degree: Calculating the cosine similarity between the candidate action vector and the original target vector of HOE; vector semantic similarity calculation is existing technology and will not be elaborated.
[0099] The technical solution of this invention embodiment is as follows: (1) Computational coding scheme for nine high-order order generation mechanisms: For the first time, the following mechanisms are incorporated: progressive specification emergence of human society (online learning + specification refinement), social contract and consensus (distributed protocol + implicit rule coding), fault tolerance and self-repair (transaction rollback + degradation strategy + alternative path generation), hierarchical delegation and trust chain (graph structure trust propagation + layered verification delegation), fuzzy reasoning and satisfaction (Satisficing threshold setting + approximate reasoning), feedback loop and continuous adaptation (reinforcement learning framework + online update), narrative and meaning construction (intent understanding + value goal representation), multi-agent game and dynamic equilibrium (game theory equilibrium calculation + emergent coordination), embodiment. The nine mechanisms of contextual intelligence (long-term memory + contextual encoding + experience accumulation) are transformed into executable algorithms or heuristic strategies and integrated into a high-order rule execution engine; (2) Explicit value goal representation and goal decomposition tree method: The user's high-level value goals (fuzzy description in natural language) are encoded into a structured goal decomposition tree, including three levels: top-level value goals, middle-level sub-goals, and bottom-level verifiable sub-constraints. The goal decomposition tree is dynamically updated and drives the generation of candidate actions; (3) Satisfaction assessment and multi-dimensional feedback loop method: A satisfaction assessment function is designed to cover four dimensions: functional satisfaction, user experience, value creation, and constraint guarantee. The assessment results are fed back to HOE to drive the iterative adjustment of the candidate action generation strategy. It realizes the high-order value goal driving of fuzzy open tasks. This module effectively solves the technical problems of existing Agent frameworks lacking high-order coordination mechanisms in human society and relying solely on prompt words for implicit goal alignment, so that the agent can still generate candidate actions that conform to the user's deep values when facing open tasks with missing specifications and fuzzy definitions.
[0100] S212. Based on the risk level recorded in the task context, the bottom-line assurance engine determines the combination of verification layers from the preset verification layer set. The preset verification layer set includes syntax layer verification, constraint layer verification, behavior layer verification, and world layer verification. The verification layer combination includes at least two verification layers.
[0101] In this embodiment, the risk level refers to the risk classification result obtained after risk sensitivity labeling of the original task request in the task context record. The preset verification layer set refers to the multi-level formal verification system pre-built in the bottom-line assurance engine, including four levels: syntax layer verification, constraint layer verification, behavioral layer verification, and world layer verification. Each layer is equipped with an independent toolchain and applicable boundaries. Syntax layer verification, the first verification layer, is used to verify the legality of structured parameters of candidate execution actions, API call parameter schema compliance, and code compilation passability; it is the layer with the lowest verification cost. Constraint layer verification, the second verification layer, is used to verify the field integrity, enumeration value legality, workflow state jump legality, and business rule compliance of candidate execution actions. Behavioral layer verification, the third verification layer, is used to verify the semantic satisfiability of the behavioral behavior of candidate execution actions, calling the SMT / SAT solver and model detector to verify the retention of bottom-line constraints by candidate actions under all possible execution paths. World layer verification, the fourth verification layer, is used to verify the correspondence between factual assertions of candidate execution actions and real-world data, performing evidence chain comparison by querying domain knowledge bases, medical databases, or regulatory texts. A verification layer combination refers to a set of one or more verification layers that are dynamically determined and activated from a preset verification layer set according to the risk level. Different risk levels correspond to different verification layer combinations, presenting a progressive inclusion relationship.
[0102] Specifically, the preparation phase of the bottom-line assurance engine's layered verification is executed by the layered verification orchestrator. The layered verification orchestrator reads the risk level field from the task context record (this field originates from the risk sensitivity annotation in the task context record), and then, based on a preset differentiated activation strategy, determines the combination of verification layers to be activated for the current task from a preset set of verification layers (including four levels: L1 syntax layer, L2 constraint layer, L3 behavior layer, and L4 world layer). The verification layer combinations corresponding to each risk level are as follows: Low-risk tasks: Activate syntax layer verification and constraint layer verification (L1 and L2), verifying only the structural and syntactic correctness and basic constraint satisfaction of candidate actions, resulting in the lowest verification cost and fastest execution speed. Medium-risk tasks: Activate syntax layer verification and constraint layer verification + behavior layer verification (L1, L2, and L3), adding verification of the semantic satisfiability of action behavior on top of L1 and L2, ensuring that actions maintain bottom-line constraints in all possible execution paths. High-risk task: Activate syntax layer validation, constraint layer validation, behavior layer validation, and world layer validation (L1, L2, L3, and L4), activating all four validation layers. Add multi-validator confidence-weighted cross-judgment at layers L3 and L4 to achieve comprehensive formal validation of candidate actions. Very high-risk task: Activate syntax layer validation, constraint layer validation, behavior layer validation, and world layer validation, plus manual review (L1, L2, L3, L4, and manual intervention). In addition to all four layers of validation and multi-validator cross-judgment, add a manual review node as a fallback for the final decision.
[0103] S213. For each candidate execution action in the candidate execution action set, verify the candidate execution action according to the verification layer combination, and determine the verification decision result.
[0104] Specifically, for each candidate action in the candidate action set, a verification process can be independently executed for each candidate action based on the determined combination of verification layers. Each activated verification layer is executed serially in a fixed order L1 to L2 to L3 to L4. If the verification of the previous layer fails, the verification process for that calibration action is immediately terminated and does not proceed to subsequent layers. After each candidate action is verified, its verification decision (including binary judgment of allow / reject, list of violated constraints, minimum counterexample, set of feasible alternative actions, and confidence level) is written to the dual-engine arbitration bus and stored in association with the candidate actions output by the target-oriented engine according to their action IDs.
[0105] For example, L1 syntax layer validation: Performed by the L1 syntax layer validator, it verifies the validity of the structured parameters of the candidate action, the compliance of the API call parameter schema, and the compileability of the code. If the candidate action outputs in JSON format, it verifies whether it conforms to the predefined JSON schema; if code generation is involved, it checks whether the code snippet can be compiled. Output: L1 decision (pass / fail + failure location). This layer of validation has the lowest cost and is performed at any risk level. If the validation passes, it proceeds to L2; if it fails, the calibration action is directly judged as "rejected". L2 constraint layer validation: Performed by the L2 structure / constraint layer validator, it verifies the field completeness of the candidate action (whether required fields are missing), the validity of enumeration values (whether parameters are within the allowed value set), the numerical range constraints (whether numerical parameters are out of bounds), the validity of the workflow (whether there is an illegal state transition or a deadlock path), and the compliance of business rules (such as compliance clause matching). Output: L2 decision. If the validation passes, it proceeds to L3 (if L3 is activated) or is directly judged as "passed" (if L3 is not activated); if it fails, it is judged as "rejected". L3 Behavioral Layer Verification: Performed by the L3 Semantic / Behavioral Layer Verifier, it calls the SMT / SAT solver (such as Z3) to perform constraint satisfiability verification and the model detector to perform state reachability analysis, verifying the preservation of bottom-line constraints for candidate actions across all possible execution paths. For example, if a candidate action contains conditional branching logic, it verifies whether each branch path satisfies the bottom-line constraints, rather than only verifying the main path. When constraints cannot be satisfied, the solver automatically generates a minimum counterexample, precisely indicating the action steps and triggering conditions that violate the constraints. The L3 verifier output format adds a feasible_domain field, which is output synchronously by Z3 during the solution process, providing accurate location of the breach point for subsequent alternative solution generation. The L3 layer is only activated when the risk level is medium risk or higher, and the output includes: L3 ruling, violation path, minimum counterexample, and feasible domain description. L4 World Layer Validation: Performed by the L4 Fact / World Correspondence Layer Validator, this layer queries trusted external data sources such as domain knowledge bases, medical databases, and regulatory texts to verify the correspondence between factual assertions and real-world data, outputting confidence interval labels. The L4 layer is activated only when the risk level is high or higher, outputting the L4 decision with confidence interval labels. Multi-Validator Cross-Decision: When the L3 and L4 layers are activated (i.e., the risk level is high or very high), the hierarchical validation orchestrator calls multiple independent validators in parallel to validate the same candidate action. Each validator uses a different algorithm implementation or a different knowledge base source. Each validator is pre-configured with confidence weights (e.g., validator A has a confidence of 0.9, validator B has a confidence of 0.8, and validator C has a confidence of 0.7), and the decision results of each validator are aggregated according to a weighted voting rule.When conflicts arise between validators (e.g., two validators pass, one rejects), an escalation protocol is triggered. First, the validation strength is increased (e.g., enabling a more precise solver configuration). If the conflict remains unresolved, manual intervention is requested. The theoretical basis of this design is: acknowledging that a single validator may have errors or blind spots, and reducing the risk of misjudgment by a single validator through cross-validation by multiple validators and confidence-weighted aggregation. It achieves mathematically provable hard behavioral guarantees for candidate actions of the agent. This module effectively solves the technical problems of existing mainstream agent frameworks that rely solely on LLM self-judgment and lack formal bottom-line constraints, and that purely formal validation can only reject without fault tolerance when verification fails, achieving a fault-tolerant yet non-blocking engineering implementation.
[0106] The technical solution of this invention adds a four-layer hierarchical verification architecture for AI Agent actions: the formal verification system is systematically divided into L1 syntax layer (JSON validity, API parameter schema, code compilation), L2 structure / constraint layer (field constraints, workflow legal jump, required clauses, type resource interface constraints), L3 semantic / behavioral layer (code specification satisfaction, planning goal attainability, mathematical proof rigor, relying on SMT / SAT and model detection), and L4 fact / world correspondence layer (correspondence between historical facts, medical evidence, financial data and the real world, relying on evidence chain comparison and knowledge base query), each layer is configured with an independent toolchain and applicable boundaries, and the verification results are aggregated by layer; (2) confidence level and differentiated verification scheduling method: according to the task risk level (low / medium / high / extremely high) and Quadrant assignment, dynamic configuration of activated verification layer combination, introduction of economic constraints as explicit design variables to prevent over-verification and under-verification; (3) Multi-verifier confidence weighted cross-judgment method: acknowledging that the verifier itself may have errors, configuring confidence weights for each verifier, executing multi-verifier verification in parallel, aggregating the adjudication results according to the weighted voting rules, triggering upgraded verification or manual intervention when there is a conflict between verifiers; (4) Red line veto and automatic generation method of alternative solutions: when the bottom line verification vetoes candidate actions, the system not only outputs the rejection decision, but also automatically generates a set of alternative actions that satisfy all bottom line constraints. The alternative solution generation strategy is: to perform minimum modification search at the constraint violation point of the veto candidate action, keep the high-order target direction unchanged, and achieve "fault tolerance without blocking".
[0107] Furthermore, the verification decision includes the final decision, a list of violated constraints, and confidence levels. Correspondingly, after obtaining the verification decision by performing hierarchical verification of the candidate action set based on the task context record and using the bottom-line assurance engine, the following is also included:
[0108] When the final decision is rejection, the minimum modification path is determined based on the constraint violation points of the candidate actions in the constraint violation list; and a set of alternative actions is generated based on the minimum modification path.
[0109] In this embodiment, the final decision result refers to the final binary judgment result made by the bottom-line assurance engine on the candidate action, with a value of "allow" or "reject". The constraint violation list refers to the list of all bottom-line constraints violated by the action when it is ultimately decided as "rejected". The confidence level refers to the system's quantitative estimate of the reliability of the verification decision result, ranging from 0 to 1, and is calculated by the multi-validator weighted cross-decision module based on the confidence weights of each validator and the consistency of the decision. The constraint violation point refers to the specific location where the candidate action triggers a bottom-line constraint violation, including three elements: the violation action step, the violation constraint condition, and the triggering state, and is precisely output by the L3 behavior layer validator when generating the minimum counterexample. The minimum modification path refers to the local correction scheme with the smallest modification magnitude executed at the constraint violation point of the rejected candidate action, making it fall back into the feasible region while minimizing the semantic distance from the original HOE target direction. The alternative action set refers to the set of one or more alternative actions that satisfy all bottom-line constraints, automatically generated by the system through the minimum modification path when the original candidate action is rejected.
[0110] Specifically, when a candidate action is ultimately rejected during the hierarchical verification process, the system does not simply output a BLOCKED status and terminate the process. Instead, it first extracts a list of violated constraints from the verification decision result. This list records one or more constraint violation points. Starting from the violation point, it enumerates candidate modification directions within the feasible region boundary to determine the minimum modification path. The modification scheme is then applied to the candidate action sequence, generating one or more corrected complete action sequences. These sequences are sorted according to the modification cost function C = α parameter offset + β semantic distance, and the top K modification schemes with the smallest C are selected as the alternative action set. The parameter offset is the normalized distance of the numerical change, and the semantic distance is the vector cosine distance between the modified action and the original HOE candidate action. The calculation of the vector semantic distance is existing technology and will not be elaborated. The minimum is determined by the minimum value of the cost function C, rather than exhaustively enumerating all feasible solutions. When the feasible region is an empty set (the SMT solver returns UNSAT and there is no relaxation space), the BLOCKED status is output, along with the definition of the cost function C and the default configuration of the K value (K=3 is recommended).
[0111] S214. Based on the verification decision result, filter out the candidate execution actions that are ultimately rejected from the candidate execution action set to obtain the filtered target execution action set.
[0112] In this embodiment, the target execution action set refers to the set of compliant candidate actions that have passed all formal verifications after filtering out all candidate actions marked as "rejected" by the bottom-line guarantee engine from the candidate execution action set.
[0113] Specifically, the agent can receive a set of candidate actions from the goal-oriented engine (each candidate action carries a satisfaction score (SA)) and a verification decision from the bottom-line guarantee engine (each candidate action includes a final decision, a list of violated constraints, and a confidence level). Based on the final decision field in the verification decision, the dual-engine arbitration bus performs an absolute coverage rule filter on the set of candidate actions: retaining all candidate actions with a final decision of "allowed" and removing all candidate actions with a final decision of "rejected," thus obtaining the target action set.
[0114] S215. If the target execution action set is a non-empty set, determine the comprehensive score of each target execution action based on the satisfaction score, confidence level, and initial weight of the dual engines for each target execution action in the target execution action set.
[0115] In this embodiment, the comprehensive score refers to the final score obtained after comprehensively evaluating the candidate action or alternative action during the arbitration stage. It is calculated by combining the dual-engine weights, the bottom-line satisfaction level, and the satisfaction score.
[0116] Specifically, after filtering out violating candidate actions, if the set of target actions is not empty (i.e., at least one candidate action has passed all formal verifications), the dual-engine arbitration bus enters the comprehensive scoring and selection process. The dual-engine arbitration bus reads the satisfaction score (SA in the above representation), the confidence level of the verification decision (from the step two output of the bottom-line assurance engine), and the initial weight pairs of the dual engines (first weight value w_FBE and second weight value w_HOE) in the task context record, and calculates the comprehensive score for each target action according to the comprehensive scoring formula. The comprehensive scoring formula is: Comprehensive score = w_FBE × Confidence level + w_HOE × Satisfaction score.
[0117] S216. The action with the highest overall score will be the final action to be executed.
[0118] Specifically, the dual-engine arbitration bus can compare and sort according to the comprehensive score, and select the target action with the highest comprehensive score as the final action to be executed.
[0119] S217. Otherwise, based on the first weight value and the second weight value, determine the comprehensive score of each alternative action in the alternative action set, and select the alternative action with the highest comprehensive score as the final action to be executed.
[0120] Specifically, when the target action set is empty after filtering out violating candidate actions (i.e., all candidate actions are rejected by the bottom-line guarantee engine), the dual-engine arbitration bus will not directly output a failure status, but will instead proceed to the evaluation process of the alternative action set. At this point, the dual-engine arbitration bus reads the alternative action set generated by the bottom-line guarantee engine in the above steps and calculates the comprehensive score of each alternative action in the set based on the initial weight pair (w_FBE, w_HOE) of the dual engines. This is done using the same formula as the comprehensive score calculation in the above steps. Note that the alternative actions are automatically generated through "minimum modification search," not actively planned by the goal-oriented engine; therefore, their satisfaction scores may require special handling. The agent can calculate the degree of matching between each alternative action and the original user value goal as the basis for its satisfaction score. After calculating the comprehensive score for all alternative actions in the alternative action set, the dual-engine arbitration bus selects the alternative action with the highest comprehensive score as the final execution action. If the set of alternative actions is empty (i.e., there are no alternatives that can satisfy all bottom-line constraints, corresponding to the output BLOCKED state), the dual-engine arbitration bus cannot generate any executable actions, triggering a manual intervention process where a human decides whether to relax the constraints or redefine the task objective. If the set of alternative actions is not empty, the alternative action with the highest overall score is output as the final execution action, and the alternative identifier is recorded for use in the interpretability report. Based on the selected final execution action, a structured interpretability report can be generated, in the format {execution action, FBE decision summary (pass / reject + constraint violation), HOE target alignment summary, alternative description (if any), dual-engine weight pair, confidence level}, which is output to the user interface or the caller.
[0121] Furthermore, embodiments of the present invention implement a cross-Agent constraint propagation and consistency synchronization method: in multi-Agent scenarios, agents share bottom-line constraints through a constraint propagation protocol to prevent local compliance behaviors from causing systemic violations at the overall level; a multi-Agent emergent order design method based on game equilibrium: the multi-Agent collaboration problem is modeled as a game theory framework, and a local incentive mechanism is designed to drive individual agent decisions to converge toward global equilibrium, thereby achieving an emergent coordination order without a central controller.
[0122] The technical solution of this invention achieves a two-stage pipeline arbitration through the FBE absolute coverage rule for hard screening and the weighted comprehensive scoring for soft selection, realizing unified and collaborative decision-making between bottom-line constraints and value objectives. This module effectively solves the technical problem in the prior art where there is a semantic gap between formal verification and neural network decision output, and the two cannot be uniformly adjudicated within the same framework. It ensures that the final execution action satisfies all provable bottom-line constraints and maximizes the achievement of higher-order value objectives within the feasible domain.
[0123] Example 3
[0124] Figure 3 This is a schematic diagram of the structure of a task decision-making device for an intelligent agent provided in Embodiment 3 of the present invention. Figure 3 As shown, the device includes:
[0125] Task evaluation module 31 is used to obtain the original task request, perform multi-dimensional evaluation on the original task request, and generate a task context record.
[0126] The quadrant determination module 32 is used to determine the task quadrant category to which the original task request belongs and the initial weight of the dual engines corresponding to the task quadrant category based on the task evaluation result. The initial weight of the dual engines includes a first weight value of the bottom-line guarantee engine and a second weight value of the goal-oriented engine.
[0127] Action generation module 33 is used to generate a set of candidate execution actions based on the task quadrant category, through the goal-oriented engine and a preset high-order rule set;
[0128] The hierarchical verification module 34 is used to perform hierarchical verification on the candidate execution action set based on the task context record and through the bottom-line guarantee engine to obtain the verification decision result;
[0129] The action determination module 35 is used to determine the final execution action based on the verification decision result, the initial weight of the dual engines, and the set of candidate execution actions.
[0130] The technical solution of this invention achieves differentiated task processing and precise strategy matching through multi-dimensional evaluation and task quadrant classification; it ensures the agent's flexible adaptability to fuzzy open tasks through the high-order rule set of the goal-oriented engine; it provides mathematically provable hard behavioral guarantees through the hierarchical verification of the bottom-line guarantee engine; and it achieves the synergistic unity of bottom-line constraints and value goals through an arbitration mechanism. This solves the technical problems of formal verification and neural network decision-making being mutually exclusive and bottom-line and goals being incompatible in the prior art, and realizes agent decision-making that balances safety and efficiency in high-risk scenarios.
[0131] Furthermore, the task evaluation module 31 is specifically used for:
[0132] Information is extracted from the original task request to determine the task semantic record;
[0133] Determine the first score of the task semantic record in the specification clarity dimension;
[0134] Determine the second score of the task semantic record in the world modelability dimension;
[0135] The task semantic record is determined as the third score in the strictness dimension of the judgment criteria;
[0136] Based on the first score, the second score, and the third score, a comprehensive formalizability index for the input task is determined;
[0137] A task context record is generated based on the task semantic record, the first score, the second score, and the third score.
[0138] Furthermore, the action generation module 33 is specifically used for:
[0139] The goal-oriented engine constructs a goal decomposition tree based on the user value goals recorded in the task context.
[0140] Based on the task quadrant category, determine a subset of target high-order mechanisms from a preset set of high-order rules;
[0141] Based on the target decomposition tree and the target higher-order mechanism subset, at least one candidate action that satisfies the value target constraint is generated.
[0142] Each candidate action is assigned a satisfaction score, and the candidate actions are sorted according to the satisfaction scores. A set of candidate actions carrying the satisfaction scores is then output.
[0143] Furthermore, the layered verification module 34 is specifically used for:
[0144] The bottom-line assurance engine determines a combination of verification layers from a preset set of verification layers based on the risk level recorded in the task context. The preset set of verification layers includes syntax layer verification, constraint layer verification, behavior layer verification, and world layer verification. The combination of verification layers includes at least two verification layers.
[0145] For each candidate execution action in the candidate execution action set, the candidate execution action is verified according to the verification layer combination to determine the verification decision result.
[0146] The verification decision result includes a final decision result, a list of violated constraints, and a confidence level. Correspondingly, the device also includes:
[0147] The alternative generation module is used to determine the minimum modification path based on the constraint violation points of the candidate execution action set in the constraint violation list when the final decision result is rejection, after the layered verification of the candidate execution action set based on the task context record and the bottom-line guarantee engine. The module then generates an alternative action set based on the minimum modification path.
[0148] Furthermore, the replacement generation module is specifically used for:
[0149] Based on the verification decision result, filter out the candidate execution actions whose final decision result is rejection from the candidate execution action set to obtain the filtered target execution action set;
[0150] If the set of target execution actions is a non-empty set, the comprehensive score of each target execution action is determined based on the satisfaction score of each target execution action in the set of target execution actions, the confidence level, and the initial weight of the dual engines;
[0151] The action with the highest overall score will be the final action to be executed.
[0152] Otherwise, based on the first weight value and the second weight value, a comprehensive score for each alternative action in the set of alternative actions is determined, and the alternative action with the highest comprehensive score is selected as the final action to be executed.
[0153] Optionally, the device further includes:
[0154] The calibration module is used to monitor the execution process of the final execution action after the final execution action is determined based on the verification decision result, the initial weights of the dual engines, and the set of candidate execution actions, generate historical execution records, and calibrate the preset dual engine weights and the preset high-order rule set.
[0155] The task decision-making device for an intelligent agent provided in the embodiments of the present invention can execute the task decision-making method for an intelligent agent provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0156] Example 4
[0157] Figure 4 A schematic diagram of an electronic device 40 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0158] like Figure 4As shown, the electronic device 40 includes at least one processor 41 and a memory, such as a read-only memory (ROM) 42 or a random access memory (RAM) 43, communicatively connected to the at least one processor 41. The memory stores computer programs executable by the at least one processor. The processor 41 can perform various appropriate actions and processes based on the computer program stored in the ROM 42 or loaded from storage unit 48 into the RAM 43. The RAM 43 may also store various programs and data required for the operation of the electronic device 40. The processor 41, ROM 42, and RAM 43 are interconnected via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.
[0159] Multiple components in electronic device 40 are connected to I / O interface 45, including: input unit 46, such as keyboard, mouse, etc.; output unit 47, such as various types of monitors, speakers, etc.; storage unit 48, such as disk, optical disk, etc.; and communication unit 49, such as network card, modem, wireless transceiver, etc. Communication unit 49 allows electronic device 40 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0160] Processor 41 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 41 performs the various methods and processes described above, such as the task decision-making methods of intelligent agents.
[0161] In some embodiments, the agent's task decision-making method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 48. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 40 via ROM 42 and / or communication unit 49. When the computer program is loaded into RAM 43 and executed by processor 41, one or more steps of the agent's task decision-making method described above may be performed. Alternatively, in other embodiments, processor 41 may be configured to execute the agent's task decision-making method by any other suitable means (e.g., by means of firmware).
[0162] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0163] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0164] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0165] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0166] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0167] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0168] In one embodiment, the present invention further includes a computer program product, which includes a computer program that, when executed by a processor, implements the task decision-making method of an intelligent agent according to any embodiment of the present invention.
[0169] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0170] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0171] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A task decision-making method for an intelligent agent, characterized in that, include: Obtain the original task request, perform multi-dimensional evaluation on the original task request, and generate a task context record; Based on the task evaluation results, the task quadrant category to which the original task request belongs and the initial weight of the dual engines corresponding to the task quadrant category are determined. The initial weight of the dual engines includes the first weight value of the bottom-line guarantee engine and the second weight value of the goal-oriented engine. Based on the task quadrant categories, the goal-oriented engine generates a set of candidate execution actions in combination with a preset high-order rule set; Based on the task context record, the candidate action set is subjected to hierarchical verification through the bottom-line protection engine to obtain the verification decision result. The final execution action is determined based on the verification decision result, the initial weights of the dual engines, and the set of candidate execution actions.
2. The method according to claim 1, characterized in that, The step of performing a multidimensional evaluation of the original task request and generating a task context record includes: Information is extracted from the original task request to determine the task semantic record; Determine the first score of the task semantic record in the specification clarity dimension; Determine the second score of the task semantic record in the world modelability dimension; The task semantic record is determined as the third score in the strictness dimension of the judgment criteria; Based on the first score, the second score, and the third score, a comprehensive formalizability index for the input task is determined; A task context record is generated based on the task semantic record, the first score, the second score, and the third score.
3. The method according to claim 1, characterized in that, The process of generating a set of candidate execution actions based on the task quadrant categories, using the goal-oriented engine in conjunction with a preset high-order rule set, includes: The goal-oriented engine constructs a goal decomposition tree based on the user value goals recorded in the task context. Based on the task quadrant category, determine a subset of target high-order mechanisms from a preset set of high-order rules; Based on the target decomposition tree and the target higher-order mechanism subset, at least one candidate action that satisfies the value target constraint is generated. Each candidate action is assigned a satisfaction score, and the candidate actions are sorted according to the satisfaction scores. A set of candidate actions carrying the satisfaction scores is then output.
4. The method according to claim 1, characterized in that, The step of performing hierarchical verification on the candidate action set based on the task context record and through the bottom-line guarantee engine to obtain the verification decision result includes: The bottom-line assurance engine determines a combination of verification layers from a preset set of verification layers based on the risk level recorded in the task context. The preset set of verification layers includes syntax layer verification, constraint layer verification, behavior layer verification, and world layer verification. The combination of verification layers includes at least two verification layers. For each candidate execution action in the candidate execution action set, the candidate execution action is verified according to the verification layer combination to determine the verification decision result.
5. The method according to claim 1, characterized in that, The verification decision result includes the final decision result, a list of violated constraints, and a confidence level. Correspondingly, after obtaining the verification decision result by performing hierarchical verification on the candidate action set based on the task context record and the bottom-line assurance engine, the following steps are also included: When the final decision is rejection, the minimum modification path is determined based on the constraint violation points of the candidate execution actions in the constraint violation list; Based on the minimum modification path, generate a set of alternative actions.
6. The method according to claim 5, characterized in that, The step of determining the final execution action based on the verification decision result, the initial weights of the dual engines, and the set of candidate execution actions includes: Based on the verification decision result, filter out the candidate execution actions whose final decision result is rejection from the candidate execution action set to obtain the filtered target execution action set; If the set of target execution actions is a non-empty set, the comprehensive score of each target execution action is determined based on the satisfaction score of each target execution action in the set of target execution actions, the confidence level, and the initial weight of the dual engines; The action with the highest overall score will be the final action to be executed. Otherwise, based on the first weight value and the second weight value, a comprehensive score for each alternative action in the set of alternative actions is determined, and the alternative action with the highest comprehensive score is selected as the final action to be executed.
7. The method according to claim 1, characterized in that, After determining the final execution action based on the verification decision result, the initial weights of the dual engines, and the set of candidate execution actions, the method further includes: The execution process of the final action is monitored, historical execution records are generated, and the preset dual-engine weights and the preset high-order rule set are calibrated.
8. A task decision-making device for an intelligent agent, characterized in that, include: The task evaluation module is used to obtain the original task request, perform multi-dimensional evaluation on the original task request, and generate a task context record. The quadrant determination module is used to determine the task quadrant category to which the original task request belongs and the initial weight of the dual engines corresponding to the task quadrant category based on the task evaluation results. The initial weight of the dual engines includes a first weight value of the bottom-line guarantee engine and a second weight value of the goal-oriented engine. The action generation module is used to generate a set of candidate execution actions based on the task quadrant category, through the goal-oriented engine and a preset high-order rule set; The hierarchical verification module is used to perform hierarchical verification on the candidate execution action set based on the task context record and through the bottom-line guarantee engine to obtain the verification decision result. The action determination module is used to determine the final action to be executed based on the verification decision result, the initial weights of the dual engines, and the set of candidate actions to be executed.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the task decision-making method of the intelligent agent according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the task decision-making method of the intelligent agent as described in any one of claims 1-7.