Method for constructing intelligent agent and storage medium

By defining paradigms and implementing models within enterprise intelligent agent systems, and utilizing natural language to describe data to construct intelligent agents with varying degrees of flexibility, the problems of system complexity and high costs are solved, enabling efficient and low-cost automation and the widespread application of intelligent agents.

CN122113987APending Publication Date: 2026-05-29DIGITAL CHINA CHINA CO LTD +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DIGITAL CHINA CHINA CO LTD
Filing Date
2026-02-25
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Enterprise intelligent agent systems are expensive to operate, consume a lot of computing power, are too complex, are difficult to apply widely, provide answers that are not intelligent or accurate enough, have a low degree of automation, and are difficult to handle complex business processes.

Method used

By determining the paradigm and model of the intelligent agent to be built based on user-input task data, and using natural language description data to replace relevant coding for system orchestration, documents and models with different levels of constraints are matched to build intelligent agents with different flexibility, simplifying the construction process and optimizing resource investment.

Benefits of technology

It reduces the complexity and implementation cost of intelligent agent systems, improves the task execution efficiency and accuracy of intelligent agents, achieves efficient and low-cost automation, and adapts to complex tasks in different business scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122113987A_ABST
    Figure CN122113987A_ABST
Patent Text Reader

Abstract

The application provides a construction method of an intelligent agent and a storage medium. The construction method determines a paradigm and a model to be carried by the intelligent agent based on task data input by a user, matches natural language description data containing the task data and documents with different constraint degrees according to the paradigm, and finally constructs the intelligent agent in combination with the natural language description data and the model. The system arrangement can be performed by replacing relevant coding with the natural language description data, the construction and operation logic of the intelligent agent is simplified, and thus the technical problem that the intelligent agent system is complex and difficult to be widely applied due to coding arrangement in the related art is improved. Meanwhile, the paradigm and the model are selected according to the task demand, unnecessary computing power consumption is avoided, and the resource input for the operation of the intelligent agent is optimized, and thus the technical problem that the intelligent agent system has high operation cost and consumes much computing power in the related art is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data processing technology, specifically to a method for constructing an intelligent agent and a storage medium. Background Technology

[0002] The following problems exist when implementing related enterprise intelligent agent systems: First, the operating cost of intelligent agent systems is high; "expensive" means high cost and high computing power consumption. Second, intelligent agent systems are too complex; "complex" means "elite application," making them difficult to widely apply. The root cause of these problems is that related intelligent agent systems use related coding for system orchestration, resulting in complex systems and high implementation costs, making it difficult to achieve efficient and low-cost automation. Summary of the Invention

[0003] This application provides a method for constructing an intelligent agent and a storage medium, which can reduce the complexity and implementation cost of intelligent agent systems.

[0004] In a first aspect, embodiments of this application provide a method for constructing an intelligent agent, including: The paradigm of the intelligent agent to be built and the model carried by the intelligent agent to be built are determined based on the task data input by the user; the paradigm is used to characterize the flexibility of the intelligent agent to be built. The natural language description data of the agent to be constructed is determined according to the paradigm; the natural language description data includes the task data, a first document, and a second document, wherein the task data is used to provide a task execution basis with a first degree of constraint; the first document is a business process document used to provide a task execution basis with a second degree of constraint, and the second document is a fully structured business guidance document used to provide a task execution basis with a third degree of constraint, wherein the first degree of constraint is less than the second degree of constraint; and the second degree of constraint is less than the third degree of constraint. The target intelligent agent is constructed based on the natural language description data and the model.

[0005] Secondly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, is used to implement the methods described above.

[0006] In the embodiments of this application, the paradigm and model of the intelligent agent to be constructed are determined based on task data input by the user. Then, natural language description data containing task data and documents with different levels of constraints are matched according to the paradigm. Finally, the intelligent agent is constructed by combining the natural language description data and the model. The system can be orchestrated by relying on natural language description data to replace related coding, simplifying the construction and operation logic of the intelligent agent, thereby improving the technical problem of complex intelligent agent systems and difficulty in widespread application caused by coding and orchestration in related technologies. At the same time, by selecting paradigms and models that adapt to task requirements, unnecessary computing power consumption can be avoided, and the resource investment for intelligent agent operation can be optimized, thereby improving the technical problem of high operating cost and high computing power consumption of intelligent agent systems in related technologies. In addition, by using the accurate matching of natural language description data with different levels of constraints and paradigms, lightweight construction of intelligent agents can be achieved, further reducing implementation costs, thereby improving the technical problem of high implementation cost and inability to automate intelligent agent systems efficiently and at low cost in related technologies. Attached Figure Description

[0007] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0008] Figure 1 This is a flowchart illustrating a method for constructing an intelligent agent as provided in an embodiment of this application. Figure 2 This is another structural schematic diagram of the method for constructing an intelligent agent provided in the embodiments of this application; Figure 3 A schematic diagram of the structure of the intelligent agent construction apparatus provided for embodiments of this application; Figure 4 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0009] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0010] The relevant enterprise intelligent agent systems still have the following problems when they are implemented: Third, the answers of the intelligent agents are not intelligent or accurate enough; Fourth, the degree of automation is not high. "Low degree of automation means that small tasks are okay to automate, but large tasks are difficult", making it difficult to handle complex business processes.

[0011] like Figure 1 As shown, Figure 1 This is a flowchart illustrating a method for constructing an intelligent agent provided in this application embodiment. This embodiment provides a method for constructing an intelligent agent, including: Step 101: Determine the paradigm of the intelligent agent to be built and the model carried by the intelligent agent based on the task data input by the user; the paradigm is used to characterize the flexibility of the intelligent agent to be built.

[0012] In this embodiment, the intelligent agent can be a software system or entity with perception, decision-making, and execution capabilities, capable of completing specific tasks based on input information; the task data can be information related to the task to be executed input by the user; the paradigm can be an operational mode that characterizes the flexibility of the intelligent agent. The model can be an algorithmic model that supports the intelligent agent in realizing cognitive and execution functions, such as a large language model, a model with reasoning capabilities, or a model with both reasoning and tool invocation capabilities.

[0013] Specifically, when determining the paradigm and model of the intelligent agent to be built based on user-input task data, the user-input task data can be analyzed first to identify the characteristics of the task, such as whether the task requires highly autonomous decision-making or has a fixed execution process. Then, the corresponding paradigm can be matched according to the task characteristics, and an appropriate model can be selected based on the task requirements. For example, if the task needs to strictly follow a fixed process, a low-flexibility paradigm can be matched. Finally, a model that can accurately parse the natural language description data corresponding to the paradigm can be selected.

[0014] In some embodiments, the paradigm includes a first paradigm representing a first degree of flexibility of the agent to be constructed, a second paradigm representing a second degree of flexibility of the agent to be constructed, and a third paradigm representing a third degree of flexibility of the agent to be constructed; the first degree of flexibility is greater than the second degree of flexibility; the second degree of flexibility is greater than the third degree of flexibility; determining the paradigm of the agent to be constructed based on user-input task data includes: Define the task objectives for the task data; the task objectives include at least one of a first objective for analysis, a second objective for matching, and a third objective for scoring. Determine the first paradigm based on the primary objective; Determine the second paradigm based on the second objective; The third paradigm is determined based on the third objective.

[0015] The first paradigm can be a highly flexible, open-ended pattern; the second paradigm can be a moderately flexible, relatively fixed pattern; and the third paradigm can be a low-flexibility, fixed, and highly precise pattern. The first level of flexibility represents the agent's high degree of autonomous decision-making and adjustment capability; the second level represents the agent's moderate degree of autonomous decision-making and adjustment capability; and the third level represents the agent's low degree of autonomous decision-making and adjustment capability. The task objective can be the specific task direction that the user expects the agent to complete; the first objective can be a task objective for data analysis; the second objective can be a task objective for object matching; and the third objective can be a task objective for result scoring.

[0016] Specifically, when determining the paradigm for building an intelligent agent based on user-input task data, we can first extract the characteristics of the task from the task data and clarify the specific type of the task objective. For example, we can determine whether the task focuses on analysis, matching, or scoring. Then, we can select a paradigm with different levels of flexibility according to the type of task objective. Specifically, if the task data points to a primary objective of analysis, since analysis tasks typically require a high degree of autonomous planning ability, the first paradigm can be adopted. If the task data points to a secondary objective of matching, matching tasks require a certain degree of flexibility and must follow basic procedures, so the second paradigm can be adopted. If the task data points to a tertiary objective of scoring, scoring tasks must strictly follow rules, so the third paradigm can be adopted.

[0017] In the embodiments of this application, by classifying paradigms into types with different levels of flexibility and determining the corresponding paradigm based on the type of task objective in the task data, a precise match between paradigms and task objectives can be achieved. This improves the technical problems in related technologies where paradigm selection lacks clear basis and the mismatch between agent flexibility and task requirements leads to poor task execution results. Simultaneously, by establishing a correspondence between task objectives and paradigms with different levels of flexibility, the paradigm determination process can be simplified, reducing decision-making complexity in agent construction. This improves the technical problems in related technologies where agent construction processes are cumbersome and implementation costs are high. Furthermore, by precisely matching task objectives with paradigms, computational waste caused by agents having excessively high or low flexibility can be avoided, optimizing resource allocation efficiency. This improves the technical problems in related technologies where agents consume excessive computational power and have high operating costs.

[0018] In some embodiments, determining the model to be mounted on the intelligent agent based on user-input task data includes: Determine the task type of the task data; the task type includes inference type or execution type. The agent to be constructed is determined based on the reasoning type, and a first model is selected; the first model is a model with reasoning capabilities. The agent to be built is determined based on the reasoning type, and a second model is constructed. The second model is a model with reasoning and tool invocation functions.

[0019] Among them, the task type can be a category divided according to the functional attributes of the task; the reasoning type can be a task category used for logical deduction, semantic parsing or rule judgment; the execution type can be a task category used for tool invocation, process execution or result implementation; the first model can be an algorithm model with reasoning function, such as an algorithm model focused on business document parsing and task logic decomposition; the second model can be an algorithm model with both reasoning function and tool invocation function, such as a composite model that can complete reasoning planning and drive tool execution.

[0020] Specifically, when determining the model to be used by the intelligent agent based on user-input task data, the functional requirements can be identified from the task data first. This involves determining whether the task type is a reasoning-only type or an execution type requiring both reasoning and execution. If the task data only requires semantic parsing, rule extraction, or logical deduction (reasoning-related tasks), then the task type is a reasoning-only model. If the task data requires further execution by calling tools based on reasoning planning (execution-type tasks), then the task type is a second model that combines reasoning and tool invocation capabilities.

[0021] In some embodiments, when the agent to be built is equipped with a second model, the agent is equipped with a dual-engine intelligent hub, which is the core processing unit of the agent system. This hub is responsible for business semantic parsing, logical deduction, and tool invocation, and is a key technical support for realizing the three agent paradigms. Through the collaborative work of the reasoning function and the tool invocation function, deep parsing and logical deduction of business semantics, as well as accurate tool matching and invocation, are achieved. This dual-engine structure can significantly improve the system's understanding ability and execution efficiency, especially when handling complex tasks, enabling a closed loop of cognition-execution and improving the overall intelligence level of the system. As an example, in the application of contract review tasks: the reasoning function is responsible for understanding the contract content, identifying key clauses and potential risk points, while the tool invocation function is responsible for selecting and invoking appropriate tools, such as Optical Character Recognition (OCR) text extraction tools, clause parsing engines, and risk assessment tools.

[0022] In the embodiments of this application, by determining the model to be carried by the agent based on the task type of the task data, redundant models are avoided, and unnecessary computing power is reduced, thereby improving the technical problem of high computing power consumption and high operating costs of agent systems caused by excessive model functions in related technologies. At the same time, by matching the task type to select the appropriate model, the model architecture of the agent is simplified, avoiding system redundancy caused by complex model combinations, thereby improving the technical problem of agent systems being too complex and difficult to widely apply in related technologies. In addition, by accurately matching the model functions with task requirements, the task execution efficiency of the agent is improved, and resource waste is reduced, thereby improving the technical problem of high implementation costs and inefficiency in automation of agent systems in related technologies.

[0023] Step 102: Determine the natural language description data of the agent to be constructed according to the paradigm; the natural language description data includes task data, a first document and a second document. The task data is used to provide the basis for task execution with a first degree of constraint; the first document is a business process document used to provide the basis for task execution with a second degree of constraint, and the second document is a fully structured business guidance document used to provide the basis for task execution with a third degree of constraint. The first degree of constraint is less than the second degree of constraint; the second degree of constraint is less than the third degree of constraint.

[0024] In this embodiment, the natural language description data can be a set of information presented in natural language form to constrain or guide the construction of an intelligent agent; the first document can be a business process document described in natural language, such as a governance or standard operating procedure (SOP) document; the second document can be a fully structured business guidance document described in natural language, such as a planning paradigm routine task document; the first constraint level can be the constraint strength of the task data on the intelligent agent's execution process; the second constraint level can be the constraint strength of the first document on the intelligent agent's execution process; and the third constraint level can be the constraint strength of the second document on the intelligent agent's execution process.

[0025] In this embodiment, when determining the natural language description data of the agent to be constructed based on the paradigm, the type of natural language description data required can be clarified first based on the selected paradigm. For example, if the paradigm is a high-flexibility paradigm, the natural language description data can mainly include task data; if the paradigm is a medium-flexibility paradigm, the natural language description data can include task data and a first document; if the paradigm is a low-flexibility paradigm, the natural language description data can include task data, a first document, and a second document. Then, the corresponding natural language description data is collected or retrieved to ensure that the data can provide the basis for task execution with the corresponding degree of constraint.

[0026] In this embodiment, the system orchestration is performed using natural language instead of related encoding, which significantly reduces the system's complexity and implementation cost. This allows business experts to directly participate in the system's adjustment and optimization, reducing reliance on a large number of programmers, thereby lowering project implementation costs and solving the problems of "expensive = high cost" and "complex = elite application".

[0027] In some embodiments, determining the natural language description data of the agent to be constructed according to a paradigm includes: Based on the first paradigm, natural language description data is identified as task data; The natural language description data is identified as the first document based on the second paradigm. The natural language description data is identified as the second document based on the third paradigm.

[0028] Among them, the first paradigm can be the operating mode that represents the first degree of flexibility of the intelligent agent to be built; the second paradigm can be the operating mode that represents the second degree of flexibility of the intelligent agent to be built; and the third paradigm can be the operating mode that represents the third degree of flexibility of the intelligent agent to be built.

[0029] Specifically, when determining the natural language description data for the intelligent agent to be built based on the paradigm, the appropriate natural language description data type can be matched according to the flexibility requirements corresponding to different paradigms. Specifically, the first paradigm corresponds to a high degree of flexibility, requiring no additional structured document constraints, therefore the natural language description data can be determined to be only task data; the second paradigm corresponds to a medium degree of flexibility, requiring certain constraints from business process documents, therefore the natural language description data can be determined to be the first document; the third paradigm corresponds to a low degree of flexibility, requiring strong constraints from fully structured business guidance documents, therefore the natural language description data can be determined to be the second document.

[0030] In the embodiments of this application, by determining the corresponding natural language description data type based on different paradigms, the introduction of redundant data is avoided, and the data processing flow for building intelligent agents is simplified, thereby improving the technical problems of complex intelligent agent systems and high implementation difficulty caused by data redundancy in related technologies. At the same time, by accurately matching the paradigm with the natural language description data, unnecessary document parsing and processing steps are reduced, and computing power consumption is reduced, thereby improving the technical problems of high operating costs and wasted computing power in intelligent agent systems in related technologies. In addition, by using natural language description data to replace related coding and arrangement, the construction logic of intelligent agents is further simplified and the construction efficiency is improved, thereby improving the technical problems of high implementation costs and difficulty in achieving efficient automation in intelligent agent systems in related technologies.

[0031] In some embodiments, the method further includes: The enterprise's business systems are split into business domains and transformed into multiple tools to be invoked by the second model; Determine the first number of first tools among multiple tools to be invoked based on the first paradigm; Determine a second number of second tools from among multiple tools to be invoked based on the second paradigm; Determine a third number of third tools from multiple tools to be invoked based on the third paradigm; Among them, the first quantity is greater than or equal to the second quantity; the second quantity is greater than or equal to the third quantity.

[0032] Here, "business system" can refer to the software system or platform used by an enterprise to handle various business operations; for example, a business system can be Customer Relationship Management (CRM), Product Lifecycle Management (PLM), or Supplier Relationship Management (SRM). "Business domain" can refer to the areas where an enterprise's business is divided by function or scenario. "Tools to be invoked" can be standardized tools that transform the functions of the business system and allow the model to invoke them; for example, tools to be invoked can form a specific enterprise toolbox divided by business domain that can be invoked by a large model. The first quantity can be the number of tools to be invoked corresponding to the first paradigm; the first tool can be a tool to be invoked that adapts to the first paradigm; the second quantity can be the number of tools to be invoked corresponding to the second paradigm; the second tool can be a tool to be invoked that adapts to the second paradigm; the third quantity can be the number of tools to be invoked corresponding to the third paradigm; the third tool can be a tool to be invoked that adapts to the third paradigm. As an example, business system transformation can involve turning a company's human resources system, project management system, and time management system into a usable toolkit, including tools for extracting job requirements, resume keyword matching, employee education verification, and project participation record query.

[0033] It's important to note that a "resource foundation" for establishing a ToolOS architecture can be built. This foundation consists of pre-built tool libraries, Natural Language (NL) SOP libraries, and Routines libraries categorized by business domain. This ensures resource coverage of core enterprise business scenarios. Specifically, this is achieved as follows: Based on business system transformation technology, tools from specific internal business systems (such as CRM, PLM, HR systems, project management systems, etc.) that can be called by a large language model are transformed and categorized by business scenario into a pre-built tool library. Business process documents described in natural language for various business scenarios are collected and categorized by business type into a pre-built NL SOP library. Fully structured business guidance documents (routines) described in natural language for core business scenarios are collected and categorized by key task type into a pre-built Routines library. The ToolOS architecture is then constructed. This architecture includes pre-built tool / NL SOP / routine libraries, uses a vector database to store tool / NL SOP / routine features, and supports semantic retrieval.

[0034] In some embodiments, the agent may include a memory module for managing and optimizing the system's memory capabilities, thereby improving the agent's response accuracy and efficiency. By managing long memories (e.g., routines) and short memories (e.g., context / key facts / node outputs) through a hierarchical storage architecture, joint memory of tool / NL SOP / routine call records, knowledge base retrieval trajectories, and external service interaction logs can be achieved. This significantly improves the agent's contextual understanding and task execution efficiency, reduces redundant computations and queries, and thus lowers system response time and resource consumption.

[0035] The memory module employs a hierarchical storage strategy, enabling rapid retrieval and recall of relevant tools and historical records. For example, when processing job matching tasks, the system can quickly retrieve historical matching records and related tool usage, improving matching efficiency and accuracy. Based on the prompt templates set in the model, process prompts or process documentation are optimized to obtain at least one initial routine task (e.g., a Routine). The model collaborates with business experts to create and optimize these initial routine tasks. At least one initial routine task is stored in the process memory, which prevents the set of initial routine tasks from consuming the execution module's context window resources. The length of the result parameters in the output is validated, and a validation result is obtained. If the validation result indicates that the length of the result parameter fails the validation, key-value pairs of the result parameters are generated. The key of the key-value pair is the identifier of the result parameter, and the value is the content of the result parameter. These key-value pairs are stored in a variable memory, which optimizes parameter passing between multi-step tool calls.

[0036] In other embodiments, when the agent to be built is equipped with a second model, the agent is equipped with a collaborative framework, which is a key mechanism for handling complex tasks. This framework facilitates task planning, execution, and reflection through the collaboration of multiple functions. The collaborative framework provides an efficient and robust collaborative approach, effectively breaking down user problems and executing each step. It checks the output of each sub-task step; if the output does not meet preset requirements, it adjusts the sub-task steps or toolbox invocation strategy. This achieves hierarchical decomposition and refined execution of tasks, significantly improving the system's ability to handle complex tasks. Simultaneously, the reflection mechanism continuously optimizes the execution process, enhancing the system's adaptability and execution efficiency. As an example, in the intelligent business opportunity screening scenario, the application of the collaborative framework can include: 1. User submits a request: screening business opportunities related to AI platforms. 2. Model decomposes the task: defining steps such as bid retrieval, information extraction, and deep analysis. 3. Model maps task steps to specific tasks: calling bid retrieval tools, OCR text extraction tools, and the Qichacha open platform, etc. 4. The tool server executes specific operations: such as retrieving 30 relevant leads from the bid pool over the past month. 5. Model review and evaluation: If insufficient information is found, the search scope may be adjusted or other information sources may be added. This collaborative approach can flexibly handle complex opportunity screening tasks, improving the accuracy and efficiency of the screening process.

[0037] Specifically, when splitting a company's business system by business domain, the system can be divided into different business domains based on the functional attributes of the business, such as recruitment, finance, and customer service. Then, the functions within each business domain are extracted and transformed into standardized tools to be invoked, ensuring that these tools have callable interfaces and clear functional outputs. When determining the first number of tools based on the first normal form, since the first normal form is highly flexible and requires more types of tools for selection, a larger number of tools can be chosen as the first tools. When determining the second number of tools based on the second normal form, the second normal form is moderately flexible and requires fewer tools than the first normal form, so a moderate number of tools can be chosen as the second tools. When determining the third number of tools based on the third normal form, the third normal form is less flexible and only requires essential tools, so a smaller number of tools can be chosen as the third tools.

[0038] In the embodiments of this application, by splitting the enterprise business system into tools to be invoked according to business domains, and then determining the number and type of tools to be adapted according to different paradigms, redundant tool configuration and invocation are avoided, and the complexity of tool management during the operation of the intelligent agent is reduced, thereby improving the technical problem of intelligent agent systems being too complex and difficult to be widely applied in related technologies. At the same time, by matching the number of tools according to the flexibility of the paradigm, unnecessary tool invocations are avoided from occupying computing resources, reducing the computing power consumption of the system operation, thereby improving the technical problem of high operating costs and wasted computing power in intelligent agent systems in related technologies. In addition, by using the tool-based transformation of the business system, the docking logic between the model and the business system is simplified, and the execution efficiency of the intelligent agent is improved, thereby improving the technical problem of high implementation costs and inefficiency in automation of intelligent agent systems in related technologies.

[0039] Step 103: Construct the intelligent agent to be built based on the natural language description data and model to obtain the target intelligent agent.

[0040] In this embodiment, the target intelligent agent can be an intelligent agent that, after completion, possesses task data processing capabilities. When constructing the intelligent agent to be built based on natural language description data and a model, the model can be used to parse the constraint rules and execution logic in the natural language description data. For example, the model's reasoning function can be used to parse the constraint relationships between steps and tools in the natural language description data, the model's tool invocation function can be used to match the corresponding tool invocation strategies, and then the model's capabilities and data constraints can be integrated to complete the construction of the intelligent agent's cognitive module, execution module, and collaboration module, ultimately obtaining the target intelligent agent. In some embodiments, the target intelligent agent can be obtained by constructing the intelligent agent to be built based on task data, natural language description data, and a model.

[0041] In some embodiments, the target intelligent agent includes a first intelligent agent with a first degree of flexibility, a second intelligent agent with a second degree of flexibility, and a third intelligent agent with a third degree of flexibility; the target intelligent agent is constructed based on natural language description data and a model, including: Input the task data and the preset first prompt information into the model to obtain the first intelligent agent; The first document and the preset second prompt information are input into the model to obtain the second intelligent agent; The second document and the preset third prompt information are input into the model to obtain the third intelligent agent.

[0042] In this model, the first agent can be an agent with the highest level of flexibility; the second agent can be an agent with the highest level of flexibility; and the third agent can be an agent with the highest level of flexibility. As an example, the first agent can be a highly flexible agent with open-ended results: a process planned and orchestrated entirely autonomously by a large language model and executed by calling the toolbox. The second agent can be an agent with some flexibility and relatively fixed results: a process planned and orchestrated by a large language model based on business-friendly, fully natural language-described business process documents (Governance / SOPs) and executed by calling the toolbox. The third agent can be an agent with low flexibility, fixed results, but absolutely reliable accuracy: a process planned and orchestrated by a large language model based on business-friendly, fully natural language-described, fully structured business guidance documents (Routines) and executed by calling the toolbox. The core orchestration part of all three paradigms is based entirely on natural language rather than traditional coding methods in the implementation state. By transforming enterprise-specific business systems into callable toolboxes, the system can directly utilize the enterprise's internal business knowledge and processes to improve the accuracy and relevance of the agent's responses, solving problems of insufficiently intelligent or accurate answers.

[0043] This application embodiment sets up three agent paradigms: a highly flexible agent can handle complex and open tasks, improving the system's adaptability; a moderately flexible agent utilizes business process documents to ensure the controllability of results while maintaining a certain degree of flexibility; and a low-flexibility but high-precision agent ensures the accurate execution of critical tasks by strictly following structured business guidance documents.

[0044] The preset first prompt information can be a preset instruction used to guide the model in building a first intelligent agent, such as prompts containing autonomous planning logic, wherein the first prompt information can be determined based on task data; the preset second prompt information can be a preset instruction used to guide the model in building a second intelligent agent, such as prompts containing SOP parsing rules, wherein the second prompt information can be determined based on a first document; the preset third prompt information can be a preset instruction used to guide the model in building a third intelligent agent, such as prompts containing rigid rule execution requirements, wherein the third prompt information can be determined based on a second document.

[0045] Specifically, when constructing a target intelligent agent based on natural language description data and a model, the matching natural language description data and preset prompts can be input into the model according to the intelligent agent type corresponding to different paradigms. Specifically, when constructing the first intelligent agent, task data and preset first prompts can be input into the model, allowing the model to generate a highly flexible first intelligent agent based on the autonomous planning logic in the prompts and the task data. When constructing the second intelligent agent, a first document and preset second prompts can be input into the model, allowing the model to generate a moderately flexible second intelligent agent based on the SOP parsing rules in the prompts and the first document. When constructing the third intelligent agent, a second document and preset third prompts can be input into the model, allowing the model to execute requirements according to the rigid rules in the prompts and generate a low-flexibility third intelligent agent based on the second document.

[0046] In the embodiments of this application, by inputting matched natural language description data and preset prompts into the model, agents with varying degrees of flexibility are constructed. This eliminates the need for complex coding development, simplifying the agent construction process and thus addressing the technical problem of overly complex and difficult-to-widely-apply agent systems in related technologies. Simultaneously, the preset prompts guide the model to automatically construct agents, reducing manual intervention costs and computational consumption, thereby addressing the technical problem of high operating and implementation costs of agent systems in related technologies. Furthermore, by combining differentiated input data and prompts, agents adapted to different scenarios are accurately constructed, improving the agent construction efficiency and task adaptability, thus addressing the technical problem of agent systems in related technologies being unable to automate efficiently and at low cost.

[0047] In some embodiments, the method further includes: The task data is processed based on the first or second model mounted on the target intelligent agent to obtain the output results.

[0048] Specifically, when processing task data based on the first or second model mounted on the target agent, the type of model mounted on the target agent can be identified first. If the target agent is mounted on the first model, the reasoning function of the model is used to perform semantic parsing, logical deduction, or rule matching on the task data to generate corresponding reasoning output results. If the target agent is mounted on the second model, the processing flow of the task data is planned first through its reasoning function, then the appropriate tool is called to perform specific operations, and finally the tool execution results and reasoning conclusions are integrated to obtain complete output results.

[0049] In the embodiments of this application, task data is processed according to the first or second model mounted on the target intelligent agent. The automated processing of task data is completed by relying on the adaptation function of the model, without the need for manual intervention in complex process operations. This improves the technical problems of high implementation cost and inefficient automation of intelligent agent systems in related technologies. At the same time, by matching the model type to process task data in a targeted manner, redundant calls to irrelevant functions are avoided, and the ineffective consumption of computing resources is reduced. This improves the technical problems of high operating cost and high computing power consumption of intelligent agent systems in related technologies. In addition, by directly outputting results with the inference or execution capabilities of the model, the intermediate links of task processing are simplified, and the overall processing efficiency is improved. This improves the technical problems of intelligent agent systems being too complex and difficult to widely apply in related technologies.

[0050] In some embodiments, task data is processed according to a second model mounted on the target agent to obtain output results, including: When the target agent is the first agent, the first number of first tools among multiple tools to be invoked are invoked based on the second model to obtain the output result; When the target agent is a second agent, the second number of second tools among multiple tools to be invoked are invoked based on the second model to obtain the output result; When the target agent is a third agent, the third third tool among multiple tools to be invoked is invoked based on the second model to obtain the output result.

[0051] The output results can be conclusions, reports, or execution feedback generated after the tool is invoked; the first quantity can be the number of tools adapted to the first intelligent agent; the first tool can be the tool to be invoked adapted to the first intelligent agent; the second quantity can be the number of tools adapted to the second intelligent agent; the second tool can be the tool to be invoked adapted to the second intelligent agent; the third quantity can be the number of tools adapted to the third intelligent agent; the third tool can be the tool to be invoked adapted to the third intelligent agent.

[0052] Specifically, when processing task data based on the second model mounted on the target agent, the type of the target agent can be identified first, and then the corresponding tool invocation strategy can be matched. Specifically, if the target agent is a first agent, the second model can invoke the first number of first tools from among multiple tools to be invoked, using the collaborative processing of multiple tools to complete the analysis or execution of the task data; if the target agent is a second agent, the second model can invoke the second number of second tools from among multiple tools to be invoked, selecting the appropriate tool according to the constraints of the business process document; if the target agent is a third agent, the second model can invoke the third number of third tools from among multiple tools to be invoked, selecting only the core necessary tools to complete the task processing under rigid rules.

[0053] As an example, a highly flexible agent is used for initial job requirement analysis and candidate competency analysis. A moderately flexible agent is used to execute a standardized resume screening and preliminary matching process. A less flexible but highly accurate agent is used for final job-person matching scoring and recommendations. Human Resources (HR) experts can define the job-person matching process using a natural language-described SOP. Step 1: Use a job hard requirement extraction tool to obtain job requirements. Step 2: Use a resume keyword matching tool to screen resumes that meet the requirements. Step 3: Perform educational verification, project experience lookup, and performance rating lookup for each candidate. Step 4: Use a keyword matching degree calculation tool to evaluate the candidate's match with the job. Step 5: If the match degree is ≥70%, retain the candidate information for the next round of evaluation. In this way, the system can automate the complex job-person matching process while maintaining a high degree of accuracy and efficiency. HR experts can adjust the matching process or standards at any time without relying on the Information Technology (IT) department to modify the system, greatly improving the system's flexibility and usability.

[0054] In the embodiments of this application, by calling corresponding numbers and types of tools according to different types of target agents, redundant tool calling operations are avoided, reducing the complexity of tool management and calling during agent operation, thereby improving the technical problem of overly complex agent systems that are difficult to widely apply in related technologies. At the same time, by matching the flexibility of agent calling tools, the computational power consumption caused by unnecessary tool calls is reduced, and resource utilization efficiency is optimized, thereby improving the technical problem of high operating costs and high computational power consumption of agent systems in related technologies. In addition, by accurately matching tool calls with agent types, the efficiency and accuracy of task data processing are improved, thereby improving the technical problem of high implementation costs and inefficiency in automation of agent systems in related technologies.

[0055] In some embodiments, task data is processed according to a second model mounted on the target agent to obtain output results, including: The second model is defined as having a phase for processing task data. This phase includes a first phase of reasoning through the task data and a second phase of invoking the tool based on the task data. When the second model is in the first stage, determine the first weight of the first stage and the second weight of the second stage; the value of the first weight is greater than the value of the second weight. The first weight is assigned to the inference function and the second weight is assigned to the tool invocation function to generate tool invocation instructions; Based on the tool call command, the second model is controlled to be in the second stage, the value of the second weight is increased and the value of the first weight is decreased, so that the value of the first weight is less than the value of the second weight. The first weight is assigned to the inference function and the second weight is assigned to the tool invocation function to generate the output results.

[0056] Here, "stage" can refer to different work steps in the process of the second model processing task data; "first stage" can refer to the step in which the second model performs inference analysis on the task data; "second stage" can refer to the step in which the second model calls tools based on the inference results; "first weight" can refer to the proportion of resources allocated to the inference function; "second weight" can refer to the proportion of resources allocated to the tool calling function; "tool calling instruction" can refer to the instruction information generated by the second model to control the tool calling; and "output result" can refer to the final result obtained by the second model after processing the task data.

[0057] Specifically, when processing task data using the second model mounted on the target agent, the task processing stages can be divided first, clearly defining the first stage as responsible for inference processing and the second stage as responsible for tool invocation processing. When the second model is in the first stage, the first weight value is determined to be greater than the second weight value, allocating more resources to the inference function to complete the logical analysis of the task data, process planning, and generation of tool invocation instructions. Subsequently, the second model is controlled to enter the second stage, increasing the second weight value and decreasing the first weight value, allowing the tool invocation function to receive more resource support. Based on the tool invocation instructions, tool invocation and result integration are completed, ultimately generating the output result. It should be noted that through a dynamic attention mechanism, the two functions in the second model can work in coordination. For example, when discovering potential risk clauses, the inference function will guide the tool invocation function to select a more specialized legal clause intelligent matcher for in-depth analysis.

[0058] In the embodiments of this application, by dividing the second model into stages for processing task data and dynamically adjusting the weights of inference and tool invocation functions, resources are allocated to the core functions of different stages as needed, avoiding ineffective resource occupation. This improves the technical problems of high computing power consumption and high operating costs caused by unreasonable resource allocation in related technologies. At the same time, by focusing on core functions in stages, the processing logic of the model is simplified, and resource competition between different functions is reduced, thereby improving the technical problems of overly complex intelligent agent systems that are difficult to widely apply in related technologies. In addition, by dynamically adjusting the weights, the processing efficiency of each stage is improved, ensuring the smoothness and accuracy of task data processing. This improves the technical problems of high implementation costs and inefficiency in automation of intelligent agent systems in related technologies.

[0059] In some embodiments, in the agent construction and task processing flow, the first model determined by the reasoning type, the second model determined by the execution type, and the first, second, and third agents built based on these models all rely on the model's task adaptability in specific business scenarios within an enterprise. Targeted fine-tuning training to optimize the model's performance in specific tasks and domains can serve as a key technological supplement to support the implementation of the aforementioned process: on the one hand, fine-tuning training allows the first model to gain a deeper understanding of enterprise business knowledge in reasoning tasks, reducing "illusions" and improving reasoning accuracy; on the other hand, it can optimize the adaptation accuracy of the second model in the tool invocation stage, ensuring a high degree of alignment between tool invocation and task objectives. Therefore, fine-tuning training technology can be an important means to improve the performance and adaptability of agent systems, further ensuring the execution efficiency and result reliability of target agents in complex enterprise business scenarios. Fine-tuning training technology, by combining reinforcement learning (RL) and supervised fine-tuning (SFT), as well as continued pre-training, enables the model to better master the patterns and industry knowledge of specific tasks. This technology can significantly improve the model's performance in specific enterprise scenarios, reduce "illusion" problems, and improve the accuracy and relevance of responses. Meanwhile, by optimizing training methods and hyperparameters, the model can be adapted to specific scenarios with limited computing resources, reducing system deployment and operating costs. As an example, in a job matching scenario, fine-tuning training techniques can include: 1. Building a training dataset based on historical successful matching cases from the company. 2. Using continued pre-training methods to allow the model to learn company-specific job descriptions, skill requirements, and other professional terminology. 3. Improving the model's performance on specific tasks such as resume parsing and skill matching through supervised fine-tuning. 4. Using reinforcement learning methods to optimize the model's matching strategy and improve the matching success rate. 5. Continuously adjusting and optimizing the model based on feedback from actual matching results. In this way, the system can more accurately understand the company's specific job requirements and talent characteristics, improving matching accuracy while reducing unnecessary manual intervention and improving overall efficiency.

[0060] The following describes the method for constructing an intelligent agent provided in the embodiments of this application. Figure 2 As shown, Figure 2This is another structural diagram of the intelligent agent construction method provided in this application embodiment. The method includes: after the user inputs task data, first determining the task objective and task type. The task objective includes at least one of a first objective for analysis, a second objective for matching, and a third objective for scoring. The task type includes a reasoning type or an execution type. Based on the task objective, a corresponding paradigm is matched: the first objective corresponds to a first paradigm, the second objective to a second paradigm, and the third objective to a third paradigm, wherein the first paradigm has a higher degree of flexibility than the second paradigm, and the second paradigm has a higher degree of flexibility than the third paradigm. Based on the task type, the model to be carried by the intelligent agent to be constructed is determined: the reasoning type corresponds to a first model with reasoning functionality, and the execution type corresponds to a second model with both reasoning and tool invocation functionality.

[0061] Simultaneously, specific business systems within the enterprise are split into business domains and transformed into multiple tools that can be invoked by the model. Then, the corresponding number of tools is determined according to the paradigm: the first normal form corresponds to the first number of first tools, the second normal form corresponds to the second number of second tools, and the third normal form corresponds to the third number of third tools, with the first number being greater than or equal to the second number, and the second number being greater than or equal to the third number.

[0062] The model matches the corresponding natural language description data based on the paradigm. The first paradigm corresponds to task data, the second paradigm corresponds to business process documents described in natural language, and the third paradigm corresponds to fully structured business guidance documents described in natural language. The corresponding natural language description data and preset prompts are input into the model. The task data, combined with the first prompt, constructs the first intelligent agent; the business process document, combined with the second prompt, constructs the second intelligent agent; and the fully structured business guidance document, combined with the third prompt, constructs the third intelligent agent.

[0063] Determine the model type equipped by the target agent. If the target agent is equipped with the first model, process the task data directly through the model's inference function and output results by combining the enterprise's specific business knowledge. If the target agent is equipped with the second model, first divide the task processing stage, including a first stage of inference processing of task data and a second stage of tool invocation processing. In the first stage, determine the first weight and the second weight. The first weight is greater than the second weight. Allocate the first weight to the inference function and the second weight to the tool invocation function, generating tool invocation instructions. Then, according to the tool invocation instructions, control the model to enter the second stage, increase the second weight and decrease the first weight, so that the second weight is greater than the first weight, while keeping the weight allocation corresponding to the function unchanged. Invoke the corresponding number and type of tools according to the agent type. The first agent calls the first tool, the second agent calls the second tool, and the third agent calls the third tool. Integrate the tool execution results and inference conclusions to finally obtain the output result.

[0064] like Figure 3 As shown, Figure 3 A schematic diagram of a device for constructing an intelligent agent provided in an embodiment of this application is shown. The device 300 includes: The first determining module 301 is used to determine the paradigm of the intelligent agent to be built and the model carried by the intelligent agent to be built based on the task data input by the user; the paradigm is used to characterize the flexibility of the intelligent agent to be built. The second determining module 302 determines the natural language description data of the intelligent agent to be constructed according to the paradigm; the natural language description data includes task data, a first document and a second document, the task data is used to provide the basis for task execution with a first degree of constraint; the first document is a business process document used to provide the basis for task execution with a second degree of constraint, the second document is a fully structured business guidance document used to provide the basis for task execution with a third degree of constraint, the first degree of constraint is less than the second degree of constraint; the second degree of constraint is less than the third degree of constraint; The construction module 303 is used to construct the intelligent agent to be constructed based on natural language description data and models, so as to obtain the target intelligent agent.

[0065] To implement the method of the embodiments of this application, such as Figure 4 As shown, Figure 4 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this application. This application also provides an electronic device 40 that may include: a memory 401 for storing a computer program; and a processor 402 for implementing the method described above when executing the computer program. The processor 402 can implement the steps of any of the methods described above, which will not be elaborated further here.

[0066] It should be noted that the electronic devices provided in the above embodiments and the above method embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0067] Of course, in practical applications, such as Figure 4 As shown, the electronic device 40 may further include at least one network interface 403. Various components in the electronic device are coupled together via a bus system 404. It is understood that the bus system 404 is used to implement communication between these components. In addition to a data bus, the bus system 404 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 4Various buses are labeled as bus systems 404. The number of processors 402 can be at least one. Network interface 403 is used for wired or wireless communication between electronic devices and other devices. Memory 401 in this embodiment is used to store various types of data to support the operation of the electronic device. The methods disclosed in the above embodiments can be applied to processor 402, or implemented by processor 402. Processor 402 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 402 or by instructions in software form. The processor 402 can be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 402 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this application can be directly reflected as the combined execution of hardware and software modules in a microcontroller. The software module can reside in a storage medium located in memory 401. Processor 402 reads information from memory 401 and, in conjunction with its hardware, completes the steps of the aforementioned method. In an exemplary embodiment, electronic device 40 can be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to execute the aforementioned method.

[0068] Specifically, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, such as a memory 401 storing the computer program, which can be executed by a processor 402 to complete the aforementioned method steps. The computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM.

[0069] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.

[0070] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0071] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, RAM, magnetic disks, or optical disks.

[0072] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed descriptions of other embodiments above, which will not be repeated here.

[0073] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for constructing an intelligent agent, characterized in that, include: The paradigm of the intelligent agent to be built and the model carried by the intelligent agent to be built are determined based on the task data input by the user. The paradigm is used to characterize the flexibility of the intelligent agent to be constructed; The natural language description data of the agent to be constructed is determined according to the paradigm; the natural language description data includes the task data, a first document and a second document, the task data is used to provide the basis for task execution with a first degree of constraint; the first document is a business process document used to provide the basis for task execution with a second degree of constraint, the second document is a fully structured business guidance document used to provide the basis for task execution with a third degree of constraint, and the first degree of constraint is less than the second degree of constraint; The second level of constraint is less than the third level of constraint; The target intelligent agent is constructed based on the natural language description data and the model.

2. The method according to claim 1, characterized in that, The paradigm includes a first paradigm characterizing a first degree of flexibility of the agent to be constructed, a second paradigm characterizing a second degree of flexibility of the agent to be constructed, and a third paradigm characterizing a third degree of flexibility of the agent to be constructed; the first degree of flexibility is greater than the second degree of flexibility; The second level of flexibility is greater than the third level of flexibility; The process of determining the paradigm of the intelligent agent to be constructed based on user-input task data includes: Determine the task objective of the task data; the task objective includes at least one of a first objective for analysis, a second objective for matching, and a third objective for scoring; The first paradigm is determined based on the first objective; The second paradigm is determined based on the second objective; The third paradigm is determined based on the third objective.

3. The method according to claim 2, characterized in that, The step of determining the natural language description data of the agent to be constructed according to the paradigm includes: Based on the first paradigm, the natural language description data is determined to be the task data; The natural language description data is identified as the first document based on the second paradigm. The natural language description data is identified as the second document based on the third paradigm.

4. The method according to claim 2, characterized in that, The process of determining the model to be used by the intelligent agent based on user-input task data includes: The task type of the task data is determined; the task type includes reasoning type or execution type. The agent to be constructed is determined to be equipped with a first model based on the reasoning type; the first model is a model with reasoning capabilities. The agent to be constructed is determined to be equipped with a second model based on the reasoning type; the second model is a model with reasoning function and tool calling function.

5. The method according to claim 4, characterized in that, The method further includes: The enterprise's business systems are split into business domains and transformed into multiple tools to be invoked by the second model; Based on the first paradigm, a first number of first tools are determined from the plurality of tools to be invoked; Based on the second paradigm, a second number of second tools are determined from the plurality of tools to be invoked; Based on the third paradigm, a third number of third tools are determined from the plurality of tools to be invoked; Wherein, the first quantity is greater than or equal to the second quantity; and the second quantity is greater than or equal to the third quantity.

6. The method according to claim 4, characterized in that, The target intelligent agent includes a first intelligent agent with the first degree of flexibility, a second intelligent agent with the second degree of flexibility, and a third intelligent agent with the third degree of flexibility; The process of constructing the target intelligent agent based on the natural language description data and the model includes: The task data and the preset first prompt information are input into the model to obtain the first intelligent agent; The first document and the preset second prompt information are input into the model to obtain the second intelligent agent; The second document and the preset third prompt information are input into the model to obtain the third intelligent agent.

7. The method according to claim 6, characterized in that, The method further includes: The task data is processed according to the first or second model mounted on the target intelligent agent to obtain the output result.

8. The method according to claim 7, characterized in that, The step of processing the task data according to the second model mounted on the target intelligent agent to obtain the output result includes: When the target agent is the first agent, the first number of first tools among multiple tools to be invoked are invoked based on the second model to obtain the output result; When the target agent is the second agent, the second number of second tools among the plurality of tools to be invoked are invoked based on the second model to obtain the output result; When the target agent is the third agent, the third tool among the multiple tools to be invoked is invoked based on the second model to obtain the output result.

9. The method according to claim 8, characterized in that, The step of processing the task data according to the second model mounted on the target intelligent agent to obtain the output result includes: The stage for the second model to process the task data is determined; the stage includes a first stage of reasoning on the task data and a second stage of invoking the tool to be invoked based on the task data; When the second model is in the first stage, a first weight for the first stage and a second weight for the second stage are determined; the value of the first weight is greater than the value of the second weight. The first weight is assigned to the inference function and the second weight is assigned to the tool invocation function to generate a tool invocation instruction; Based on the tool call command, the second model is controlled to be in the second stage, the value of the second weight is increased and the value of the first weight is decreased, so that the value of the first weight is less than the value of the second weight. The first weight is assigned to the inference function and the second weight is assigned to the tool invocation function to generate the output result.

10. A computer-readable storage medium, characterized in that, The computer-readable medium stores a computer program that, when executed by a processor, is used to implement the method according to any one of claims 1 to 9.