Task processing method, task platform, computing device and computer readable storage medium
Through a large language model, a structured execution protocol is generated and dependency analysis is performed, which solves the problem of collaborative skills combinations of multiple agents and realizes efficient autonomous planning and execution of complex tasks.
Patent Information
- Application Number
- CN202510759918.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-09
AI Technical Summary
In the prior art, a single agent is difficult to meet the needs of complex multi-step and multi-task scenarios. How to integrate the skills of multiple agents and independently plan and execute tasks is a key issue that needs to be solved urgently.
Through a large language model, the task description text is semanticly parsed, task elements are extracted, structured execution protocols are generated, dependency analysis and skill combinations are performed, skill tools are scheduled to execute tasks, realize semantic alignment and dynamic adaptation of multimodal skill interfaces, and build an extensible skill combination framework.
It significantly improves the independent planning and execution capabilities of complex tasks, breaks through the limitations of a single agent, and realizes the accuracy and scalability of skill and tool calls to handle target tasks in open scenarios.
Smart Images

Figure CN120297322A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the technical field of artificial intelligence, and particularly to a task processing method, a task platform, a computing device, and a computer-readable storage medium. Background Art
[0002] With the development of deep learning technology, as a new generation of artificial intelligence systems, agents are being widely used in the automated processing of complex tasks and user interaction scenarios.
[0003] Currently, agents can perceive the environment, understand user needs, and autonomously plan and execute tasks, thereby providing users with efficient and accurate services.
[0004] However, a single skill often fails to meet complex project requirements. Especially when facing multi-step and multi-task scenarios, agents need to have the ability to combine multiple skills. How to integrate the skills of multiple agents and the corresponding tools, and autonomously plan and execute tasks is a key problem that needs to be solved urgently. Summary of the Invention
[0005] In view of this, the embodiments of this specification provide a task processing method. One or more embodiments of this specification also relate to a task platform, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.
[0006] In one embodiment of this specification, a task processing method is provided, including: Obtain the task description text of the target task; Use a large language model to extract the task elements of the target task from the task description text. Based on the task elements, determine multiple target agents and the target skills of the multiple target agents from the candidate agent set, and generate a structured execution protocol corresponding to the target skills of the multiple target agents; Based on the structured execution protocols corresponding to the target skills of the multiple target agents, perform a dependency analysis on the target skills of the multiple target agents to obtain a skill combination; Schedule the skill tools corresponding to the skill combination to execute the target task to obtain a task result.
[0007] By using a large language model to perform semantic parsing and element extraction on task description texts, semantic alignment and dynamic adaptation of the multimodal skill interface are achieved, solving the collaboration problem among heterogeneous skill modules; by generating a structured execution protocol and implementing dependency analysis, an extensible skill combination framework is constructed, significantly enhancing the autonomous planning and task execution capabilities for complex tasks. Through the semantic-driven skill combination mechanism, the limitation of a single intelligent agent is broken through. Meanwhile, relying on the generalization and understanding ability of the large language model, the accuracy and scalability of using skills and tools to handle target tasks in an open scenario are realized. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] Figure 1 is a flowchart of a task processing method provided by an embodiment of this specification; Figure 2 is a schematic flowchart of a task processing method provided by an embodiment of this specification; Figure 3 is a schematic diagram of an intelligent agent list page of a task platform applied to a large language model provided by an embodiment of this specification; Figure 4 is a schematic diagram of an intelligent agent multi-version page of a task platform applied to a large language model provided by an embodiment of this specification; Figure 5 is a schematic diagram of an intelligent agent orchestration canvas of a task platform applied to a large language model provided by an embodiment of this specification; Figure 6 is one of the schematic diagrams of adding skills to a combined task on a task platform applied to a large language model provided by an embodiment of this specification; Figure 7 is another schematic diagram of adding skills to a combined task on a task platform applied to a large language model provided by an embodiment of this specification; Figure 8 is a schematic diagram of pre-processing and post-processing on a task platform applied to a large language model provided by an embodiment of this specification; Figure 9 is a schematic diagram of concurrent execution of combined skills on a task platform applied to a large language model provided by an embodiment of this specification; Figure 10 is a schematic diagram of the overall architecture of a task platform applied to a large language model provided by an embodiment of this specification; Figure 11 is a schematic diagram of the combined skill architecture of a task platform applied to a large language model provided by an embodiment of this specification; Figure 12It is a schematic diagram of a generation skill input parameter schema for a task platform applied to a large language model provided by an embodiment of this specification; Figure 13 It is a schematic diagram of parallel execution of multiple tasks by combined skills in a task platform applied to a large language model provided by an embodiment of this specification; Figure 14 It is a schematic diagram of input supplementary information for a task platform applied to a large language model provided by an embodiment of this specification; Figure 15 It is a schematic diagram of the structure of a task platform provided by an embodiment of this specification; Figure 16 It is a structural block diagram of a computing device provided by an embodiment of this specification. Detailed implementation manners
[0009] Many specific details are set forth in the following description in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0010] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.
[0011] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0012] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0013] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, usually including hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than one quadrillion model parameters. A large model can also be referred to as a Foundation Model. Through pre-training of the large model with a large amount of unlabeled corpus, a pre-trained model with more than hundreds of millions of parameters is produced. Such a model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLMs), multi-modal pre-training models, etc.
[0014] When a large model is actually applied, it only needs to be fine-tuned with a small number of samples on the pre-trained model to be applied to different tasks. Large models can be widely applied in fields such as natural language processing (NLP) and computer vision. Specifically, they can be applied to tasks in the field of computer vision such as visual question answering (VQA), image captioning (IC), and image generation, as well as tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0015] First, the noun terms involved in one or more embodiments of this specification are explained.
[0016] Agent: A software system with the capabilities of autonomous perception, decision-making, and execution. It can complete specific tasks based on environmental inputs and user requirements through reasoning, planning, and tool invocation. Its core features include autonomy (without human intervention), adaptability (dynamically adjusting strategies), and interactivity (continuously interacting with users or the environment), and it is widely used in fields such as automated services, intelligent assistants, and complex task orchestration.
[0017] Skill: An ability unit for an agent to perform a specific task, encapsulating specific functional logics (such as text generation, data query, Application Programming Interface (API) call, etc.), and defining input and output specifications through a standardized interface (skill schema). Skills can be deployed independently or used in combination, supporting modular expansion. For example, the translation skill requires input text and returns the result in the target language, and its effectiveness depends on the agent's context understanding and tool scheduling capabilities.
[0018] Tool: External resources or interfaces integrated by an agent to expand the functional boundary, including search engines, databases, third-party Application Programming Interfaces (APIs), or local scripts, etc. Tools are called by the agent through abstract encapsulation (such as API adapters) to solve requirements that cannot be covered by native skills (such as real-time data acquisition), and their collaborative efficiency depends on interface compatibility and the agent's dynamic scheduling strategy.
[0019] Memory: The core module for an agent to store and manage information, divided into short-term memory (maintaining the current conversation state, temporary variables) and long-term memory (persistent historical records, domain knowledge bases). The memory mechanism supports continuous decision-making through context association and retrieval enhancement (such as vector databases). For example, remembering user preferences to optimize subsequent interactions, and its design needs to balance real-time performance and storage overhead.
[0020] Context: A collection of dynamic background information in the current task or conversation, including user intentions, environmental states, historical interaction records, and intermediate results of skill execution. The context is passed to the skill chain through a structured representation (such as key-value pairs or graphs) to solve the problems of anaphora resolution (such as "it" referring to the previous object) and task coherence in multi-turn conversations, and is the key dependence for an agent to achieve accurate responses.
[0021] Skill schema: A standard structured template that defines the behavior of a skill, specifying functional descriptions, input parameters (such as data types, required / optional), output formats (such as JavaScript Object Notation (JSON) schema), and execution constraints (such as permission requirements). The schema enables automatic registration and discovery of skills (such as in a skill pool) through machine-readable metadata, and drives the protocol generation of the agent (such as the Open Application Programming Interface (OPEN-API) specification) to ensure seamless collaboration between heterogeneous skills.
[0022] In this specification, a task processing method is provided. This specification also relates to a task platform, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail one by one in the following embodiments.
[0023] See Figure 1 , Figure 1 which shows a flowchart of a task processing method provided by an embodiment of this specification, including the following specific steps: Step 102: Obtain the task description text of the target task.
[0024] The embodiments of this specification are applied to websites, applications, or system platforms deployed with large language models. For example, the task platform of large language models, or the code development platform of large language models, or the cloud network platform of large language models.
[0025] The target task is a natural language request to be processed. The target task is a composite operation task executed by coordinating the skills of multiple agents, such as cross-system data integration, multi-modal information processing, or dynamic decision support. For example, the target task is: query the weather in a certain place. The task description text is the natural language instruction input by the user. The task description text directly or implicitly includes task elements such as task objectives, execution constraints, and task backgrounds. For example, the task description text of the target task is: "Query the weather in a certain place."
[0026] An optional way to obtain the task description text of the target task is: receive the task description text of the target task sent by the front end, or another optional way is: obtain the task description text of the target task from the database, or another optional way is: parse the task data of the target task to obtain the task description text of the target task, which is not limited here.
[0027] Exemplarily, on the task platform of a certain large language model, the user hopes to automate the execution of the target task by combining multiple agents in the agent pool of the task platform. The user inputs the task description text "Query the weather in a certain place" at the front end.
[0028] In step 102, obtaining the task description text of the target task provides the input data basis for generating the structured execution protocol subsequently.
[0029] Step 104: Use the large language model to extract the task elements of the target task from the task description text, and based on the task elements, determine multiple target agents and the target skills of the multiple target agents from the candidate agent set, and generate the structured execution protocol corresponding to the target skills of the multiple target agents.
[0030] Large language models are pre-trained deep neural networks with natural language understanding and generation capabilities, obtaining semantic reasoning and context modeling capabilities through training with massive amounts of text data. Optionally, large language models use self-attention mechanisms to achieve cross-sequence dependency modeling. The core functions of large language models include intent recognition, entity extraction, logical reasoning, and structured output generation, capable of mapping unstructured task description texts into executable operation instructions. The task elements of the target task are a set of structured task features parsed from the task description text, including the task objective, entities, operation instructions, input and output parameters, execution constraints, and associated information of the target task. For example, in the task description text "Query the weather in a certain place", the task elements include the task objective (obtain weather data), execution constraints (time range is the next 24 hours), input parameters (location: XX, date: XXXX-XX-XX, temperature: 35°C), and associated information (need to call the real-time data interface). An agent is an intelligent agent instance that, based on environmental inputs and user requirements, completes specific tasks through reasoning, planning, and tool invocation. As a carrier for skill invocation, multiple functional modules are registered in its preset skill pool. Each agent declares its skill invocation specifications through a skill schema. For example, agent A is registered with weather query skills and weather map rendering, and agent B is registered with data visualization skills. The candidate agent set is the currently callable agents, containing multiple pre-registered intelligent agent instances with heterogeneous functions. This set is maintained through a dynamic registration mechanism. The preset skills of an agent are callable functional modules pre-registered in the agent's skill pool through a skill schema, and each skill corresponds to problem-solving capabilities in a specific domain. For example, preset skills include but are not limited to natural language processing skills (such as entity recognition), API invocation skills (such as map service interfaces), data processing skills (such as table merging), and decision-making and reasoning skills (such as path planning algorithms). The target agent is the intelligent agent instance among the preset skills of multiple agents that is suitable for the task elements. The screening criteria include the relevance of skill functions to the task objective, input parameter compatibility, and satisfaction of execution constraints. For example, for a weather query task, an agent registered with weather API invocation skills and supporting the specified geocoding format is selected. The target skill of the target agent is the callable functional module in the target agent's skill pool that matches the task elements. For example, the target skill is the "real-time weather data acquisition" skill encapsulated in the weather query agent. The structured execution protocol corresponding to the target skill is a machine-readable invocation specification generated based on the skill schema, defining the input parameter format, output data structure, and invocation constraints required for skill execution. The structured execution protocol is represented using a standardized data description language (such as JSON schema, Protobuf (Protocol Buffers)), containing parameter names, data types, verification rules, and dependencies.For example, under the guidance of a large model Prompt (prompt), a structured execution protocol corresponding to the target skill is generated. The Prompt is as follows: { / / Define the data execution protocol specification for skill call parameters, which is used to standardize the input format and semantic constraints when the XX skill is executed. "description": "The data specification of actionParam when the XX skill is launched", / / Describe the global purpose of this schema: ensure that the skill call parameters conform to the predefined interface standard and solve the semantic alignment problem between heterogeneous skills. "properties": { / / Declare the set of attributes of the skill input parameters, and each attribute corresponds to a specific data field extracted from the task context. "input": { / / The semantic parsing result of the user's original instruction, which needs to be generated after the intention recognition and entity extraction by the large language model. "description": "The problem extracted from the user input", / / Explain the field semantics: the structured problem statement after NLU processing, for example, mapping \"Check the weather in a certain place\" to \"Obtain meteorological data in a certain place\". "type": "string" / / Data type constraint: limited to the string type, and it needs to conform to the UTF-8 encoding and length limit (such as ≤ 1024 characters).} / / Note: Other parameter fields can be extended, such as output format preference (output_format), service quality requirements, etc.} / / Typical application scenario example: When the agent calls the translation skill, actionParam needs to include input = \"HelloWorld\", target_lang = \"zh-CN\". Exemplarily, after the large language model parses the task description text "query the weather in a certain place", it extracts the task elements: the goals are to obtain location data, temperature data, air quality data, and wind speed data, the constraint is the time range of "XXXX-XX-XX", and the input parameter is city=XX. According to the skill schema matching, from the candidate agent set, the "location query", "temperature query", "air quality query", and "wind speed query" skills of the environmental monitoring agent are selected, and a structured execution protocol is generated: input {city: "XX", date: "XXXX-XX-XX"}, output {"location":{"city":"XX","coordinates":{"latitude":"XX.XXXX","longitude":"XXX.XXXX"},"timezone":"XX / XXXXX"},"weather":{"date":"XXXX-XX-XX","temperature":{"value":"XX.X","unit":"°C / °F","description":"[Natural language description]"},"air_quality":{"aqi":"XX","pm2_5":"XX","pollutants":["[Pollutant type and concentration]"]},"wind":{"speed":"X.X","unit":"m / s","direction":"[Wind direction description]","level":"[Wind speed classification]"},"timestamp":"XXXX-XX-XXTXX:XX:XXZ"}}.
[0031] In step 104, by using the large language model to perform semantic parsing and element extraction on the task description text, the semantic alignment and dynamic adaptation of the multimodal skill interface are realized, and the cooperation problem between heterogeneous skill modules is solved; by generating a structured execution protocol and implementing dependency analysis, an extensible skill combination framework is constructed, realizing cross-skill task context transfer and dynamic resource scheduling, providing a combination basis and combined object data for obtaining skill combinations subsequently.
[0032] Step 106: Based on the structured execution protocols corresponding to the target skills of multiple target agents, perform dependency analysis on the target skills of multiple target agents to obtain skill combinations.
[0033] Dependency analysis is a logical reasoning process for determining the execution order and data dependencies between skills. Through dependency analysis, the input and output parameters and execution constraints in the structured execution protocol are parsed to construct a data flow and control flow dependency graph between skills, including predecessor-successor relationships (e.g., Skill B needs to be started after Skill A is completed) and data transfer paths (e.g., the input parameter of Skill C comes from the output result of Skill D). Dependency analysis uses a topological sorting algorithm to identify the execution sequence, avoid circular dependencies, and ensure the feasibility and execution efficiency of the task chain. For example, in a weather query task, dependency analysis identifies that the "query location data" skill needs to be executed before the "weather API call" skill because the input parameter of the latter depends on the geographical coordinate data output by the former. A skill combination is a skill call sequence that is constructed based on the results of dependency analysis and meets the task requirements. By logically arranging multiple target skills according to the execution priority, data dependencies, and resource constraints, a skill combination that can collaboratively complete the target task is formed. The skill combination represents the execution process through a directed acyclic graph (DAG), where nodes represent skill instances and edges represent data transfer or execution order constraints. For example, the skill combination "location query → temperature query → air quality query → wind speed query → data visualization" forms a multi-level skill call chain, and the output parameters of each skill serve as the input parameters of the next skill.
[0034] Based on the structured execution protocols corresponding to the target skills of multiple target agents, perform dependency analysis on the target skills of multiple target agents to obtain a skill combination. An optional method is: based on the structured execution protocols corresponding to the target skills of multiple target agents, perform dependency analysis on the target skills of multiple target agents, construct a skill call topology, and based on the skill call topology, construct a skill combination.
[0035] Exemplarily, for the task description text "Query the weather in a certain place", four skills of "Location Query", "Temperature Query", "Air Quality Query", and "Wind Speed Query" need to be called. Dependency analysis reveals that: 1) The "Temperature Query" needs to receive the standard geographical coordinates output by the "Location Query" (such as latitude XX.XXXX, longitude XXX.XXXX); 2) The "Air Quality Query" depends on the same geographical coordinates and date parameters; 3) The "Wind Speed Query" needs to synchronously obtain the timestamp to ensure data consistency. Based on this, a skill call topology is constructed, and a skill combination execution sequence is generated: Location Query → (Temperature Query, Air Quality Query, Wind Speed Query) are executed in parallel, where: The Location Query outputs {"coordinates": {"latitude": "XX.XXXX", "longitude": "XXX.XXXX"}} as the input parameters for the subsequent three skills; The Temperature Query, Air Quality Query, and Wind Speed Query share the time constraint parameter date: "XXXX-XX-XX"; After all sub-skills are executed, a data aggregation operation is triggered to generate a skill combination for the complete weather report.
[0036] In step 106, by performing dependency analysis on the target skills of multiple target agents, a skill combination is obtained, providing an achievable skill combination for the subsequent execution of the target task.
[0037] Optionally, based on the structured execution protocols corresponding to the target skills of multiple target agents, perform dependency analysis on the target skills of multiple target agents to obtain a skill combination, including the following specific steps: Based on the structured execution protocols corresponding to the target skills of multiple target agents, perform dependency analysis on the target skills of multiple target agents to determine the missing skills, obtain the structured execution protocols corresponding to the missing skills, and based on the structured execution protocols corresponding to the target skills of multiple target agents and the structured execution protocols corresponding to the missing skills, obtain the skill combination.
[0038] Missing skills are functional modules that are identified during the dependency analysis and are not registered in the current skill pool. Missing skills are dynamically determined through the input-output matching gaps of task elements and existing skills, and need to be solved through skill extension (such as temporarily registering new skills) or alternative solutions (such as combining existing skills to simulate functions). Its discovery mechanism depends on the logical reasoning ability of the large language model and the completeness verification of the skill schema. For example, if the task requires "generating a weather report for a certain place" but the skill pool lacks the "Data Visualization" skill, then it is determined as a missing skill, and supplementary registration or calling an external tool (such as the Matplotlib interface) needs to be implemented.
[0039] Among them, the missing skills can be manually set (e.g., an administrator explicitly marks a certain skill as an item to be supplemented), or can be potential required skills predicted through the generalization ability of the large language model (e.g., inferring unregistered but frequently required functional modules based on historical task patterns), or can be interface gaps automatically discovered through the input-output compatibility check of the skill schema (e.g., the output parameters of the existing skill cannot meet the input type constraints of the downstream skill), which is not limited here.
[0040] Exemplarily, based on the skill gaps pre-identified by the developer based on domain knowledge, the developer constructs the missing skill "data visualization" by himself and configures the structured execution protocol corresponding to the missing skill.
[0041] By dynamically identifying and integrating the missing skills and determining the structured execution protocol of the missing skills, the integrity and scalability of the agent system in complex task processing are ensured.
[0042] Step 108: Schedule the skill tools corresponding to the skill combination to execute the target task and obtain the task result.
[0043] The skill tools corresponding to the skill combination are the execution engines or external service interfaces bound to each skill instance in the skill combination. The skill tools are set according to the tool types declared in the skill schema (such as REST API endpoints, database connection configurations, algorithm containers), and are manifested as a set of physically instantiated resources for actually executing the skill function logic. Optionally, the tool scheduling follows the quality of service policy, including response time optimization (such as selecting a data center nearby) and fault tolerance mechanisms (such as failure retry policies). For example, the map service tool may be docked with a certain map API and database query API at the same time for automatic scheduling. The task result is the structured data or operation status output after the execution of the skill combination, reflecting the completion of the target task. The task result includes direct output (such as API response data), indirect products (such as generated files), and execution metadata (such as elapsed time, error code). The format of the task result needs to conform to the output specification implicitly or explicitly defined in the task description text and can be used for subsequent operations through the context passing mechanism. For example, the task result of the weather query task is in JSON format: {"temperature":25,"unit":"°C","humidity":"60%"}, with an execution status code of 200 attached. Scheduling the skill tools corresponding to the skill combination to execute the target task and obtain the task result, an optional way is: using a scheduling engine to schedule the skill tools corresponding to the skill combination to execute the target task and obtain the task result, which is not limited here.
[0044] Exemplarily, the scheduling engine performs the following operations: call the "Location Encoding" skill tool (tool implementation: encapsulation of the geocoding API), input {"city":"XX"}, and output {"coordinates":[121.47,31.23]}; schedule three skill tools in parallel: "Temperature Prediction" (tool implementation: microservice of the meteorological numerical prediction model), input {"coordinates":[121.47,31.23],"hours":24}; "Precipitation Probability" (tool implementation: third-party weather data aggregation interface), with the same input as Temperature Prediction; "Wind Speed Analysis" (tool implementation: local Python wind speed calculation module), with the same input as Temperature Prediction; after aggregating the above results, call the "Report Generation" skill tool (tool implementation: template rendering engine), input structured weather data; finally, schedule the "Data Visualization" skill tool (tool implementation: Matplotlib + Seaborn chart generation), input dashboard data. The task result is {"status":"success","metadata":{"duration":"2.3s"},"data":{Dashboard}}.
[0045] Optionally, scheduling the skill tools corresponding to the skill combination to execute the target task and obtaining the task result includes the following specific steps: scheduling the skill tools corresponding to the skill combination to execute the target task, and in the case of detecting missing information, determining the missing skill, obtaining the structured execution protocol corresponding to the missing skill, and obtaining the skill combination based on the structured execution protocols corresponding to the target skills of multiple target agents and the structured execution protocol corresponding to the missing skill.
[0046] Exemplarily, when the scheduling engine executes to the data aggregation stage, it is found that the meteorological data set {"temperature_series": [25, 26, 24], "precipitation": 0.3} output by the upstream skill does not match the data structure required for the input of the downstream report generation skill in the standardized time series format {"timestamps": ["09:00", "12:00", "15:00"], "metrics": {"temp": [25, 26, 24], "rain_prob": [0.3]}}. After analyzing the context logs by the large language model, it is recognized that the "time series alignment" skill needs to be added, and the function of this skill is to dynamically bind discrete numerical values with timestamps. The structured execution protocol of the missing skill is immediately generated by the large language model as follows: {"description": "A data conversion module that aligns an unordered numerical sequence with a standardized time axis", "input_schema": {... / / Input parameters for time series alignment...}, "output_schema": {... / / Input parameters for time series alignment...}}.
[0047] In the embodiments of this specification, by using the large language model to perform semantic parsing and element extraction on the task description text, semantic alignment and dynamic adaptation of the multi-modal skill interface are achieved, and the cooperation problem between heterogeneous skill modules is solved; by generating a structured execution protocol and implementing dependency analysis, an extensible skill combination framework is constructed, significantly improving the autonomous planning and task execution capabilities for complex tasks. Through the semantic-driven skill combination mechanism, the limitation of a single intelligent agent is broken through. At the same time, relying on the generalization and understanding ability of the large language model, the accuracy and scalability of using skills and tools to call and process target tasks in an open scenario are realized.
[0048] In an optional embodiment of this specification, in step 104, the large language model is used to extract the task elements of the target task from the task description text, including the following specific steps: Using the large language model to perform semantic parsing on the task description text and extract the task elements of the target task from the task description text, where the task elements include at least one of the entity, operation instruction, and associated information of the target task.
[0049] The entities of the target task are the objective objects or sets of data items involved in the task description text, including but not limited to specific instances such as locations, times, numerical indicators, project objects, etc. Optionally, the entities are extracted from natural language through named entity recognition and mapped to standardized semantic identifiers (such as GeoNames codes, timestamps), serving as the basis for input parameters of skill calls. For example, in the task description text "Query the temperature in a certain place in the next 24 hours", the entities include the location entity "XX" (mapped to geoID: 1816670) and the time entity "the next 24 hours" (parsed as the time interval [now, now + 86400s]). The operation instructions of the target task are the functional action requirements implicitly or explicitly defined in the task description text, representing the core operation types that the user expects to execute. The operation instructions are recognized through an intent classification model and mapped to predefined skill call templates (such as Query, Calculate, Compare, etc.), driving the intelligent agent to select the corresponding function modules. For example, in the description text "Compare the PM2.5 values of YY", the operation instruction is "Compare", triggering the combined call of the data comparison skill. The associated information of the target task is the auxiliary constraint conditions or task background related to the task execution context, including user preferences, service quality requirements, data timeliness constraints, and cross-task dependencies, etc. The associated information is parsed through a context reasoning model and used to optimize the skill call strategy (such as preferentially selecting low-latency APIs) and execution parameter configuration (such as setting the cache expiration period). For example, in the description text "Display the air quality trend in the last week in a chart, including the ozone index", the associated information includes the output format preference (chart visualization) and the index filtering condition (ozone concentration).
[0050] Exemplarily, use a large language model to extract entities: {location: {city: "XX", admin_division: true}, time_range: {"XXXX-XX-XX"}, metric: "PM2.5"}, identify the operation instruction sequence: ["aggregate", "compare", "visualize"], parse the associated information: {output_format: "chart", granularity: "monthly"}, and obtain the task elements: the goal is to obtain location data, temperature data, air quality data, and wind speed data, the constraint is the time range of "XXXX-XX-XX", and the input parameter is city = XX.
[0051] In the embodiments of this specification, by using a large language model to perform semantic parsing on the task description text and extract the task elements of the target task from the task description text, the accurate conversion from natural language instructions to structured task parameters is realized, providing accurate data support for determining the target skills of at least one target intelligent agent subsequently.
[0052] In an alternative embodiment of this specification, in step 104, based on the task elements, determining multiple target agents and the target skills of the multiple target agents from the candidate agent set includes the following specific steps: Based on the task elements and the matching degree between the skill descriptions of the preset skills of the agents in the candidate agent set and the execution protocol specifications of the preset skills, determining multiple target agents and the target skills of the multiple target agents from the candidate agent set, where the structured execution protocol specification includes parameter rules and / or operation rules.
[0053] The skill description of the preset skill is the functional description data of the preset skill, which can be defined by natural language or structured metadata. The skill description may include the skill usage (such as "geocoding conversion"), the applicable field (such as "meteorological data processing"), the function limitation (such as "only supports Chinese address input"), and the performance index (such as "response time < 500ms"), and is used for skill discovery and matching decision-making. For example, the description of the weather query skill is "obtain real-time temperature, humidity, and wind data through latitude and longitude coordinates, support global queries but only return Celsius temperature units". The execution protocol specification of the preset skill is the technical standard of the skill call interface defined in a machine-readable form, which restricts the input and output data structures and the execution environment requirements. The execution protocol specification uses a standardized description language (such as JSON schema) to declare the parameter type, required fields, enumeration value range, and error code system to ensure the syntax compatibility of cross-skill calls. For example, the execution protocol specification of the air quality query skill includes the input parameter {"coordinates": {"type": "array", "items": {"type": "number"}, "minItems": 2}}, and the output field {"aqi": {"type": "integer", "minimum": 0}}. The parameter rule is a set of technical constraints for the input and output parameters in the execution protocol specification, which guarantees the semantic consistency of data exchange. The parameter rule may include the data type (such as string / number), the value range (such as 0 ≤ pH ≤ 14), the format regular expression (such as ISO 8601 timestamp), and the dependency relationship (such as when type = "VIP", require_card_id = true). The operation rule is a set of non-functional constraint conditions during skill execution, which defines the resource consumption, concurrency control, and service quality requirements. The operation rule may cover the timeout threshold (such as max_duration = 2s), the permission level (such as auth_level ≥ 3), the resource dependency (such as GPU video memory ≥ 4GB), and the traffic limit (such as 10 times / minute).
[0054] Exemplarily, the task elements include the operation instruction "aggregate" and the input parameters {"city": "XX", "metrics": ["PM2.5","SO2"]}: Agents A (skill score 0.87) and B (skill score 0.79) whose skill descriptions contain "air pollutant aggregation" are matched. It is verified that the execution protocol specification of Agent A requires that "NO2" must be included in metrics, while Agent B supports dynamic metric expansion. Finally, the "multi-pollutant analysis" skill of Agent B is selected.
[0055] In the embodiments of this specification, by establishing a multi-dimensional matching mechanism for task elements, skill descriptions, and execution protocol specifications, the accuracy and adaptability of skill selection are achieved, and the problem of function positioning in heterogeneous skill pools is solved; through the verification of parameter rules and / or operation rules, the compatibility and reliability of skill calls are ensured, laying a technical foundation for the subsequent generation of structured execution protocols.
[0056] In an optional embodiment of this specification, generating the structured execution protocols corresponding to the target skills of multiple target agents in step 104 includes the following specific steps: generating the structured execution protocols corresponding to the target skills of multiple target agents based on the skill descriptions and execution protocol specifications of the target skills of the multiple target agents, where the structured execution protocol specification includes parameter rules and / or operation rules.
[0057] The skill description of the target skill is the functional description data of the target skill, which can be defined by natural language or structured metadata. The skill description may include the skill purpose, applicable field, function limitations, and performance indicators, and is used for skill discovery and matching decisions. The execution protocol specification of the target skill is a technical standard for skill call interfaces defined in a machine-readable form, which restricts the input and output data structures and execution environment requirements. The execution protocol specification uses a standardized description language (such as JSONschema) to declare parameter types, required fields, enumeration value ranges, and error code systems to ensure syntactic compatibility for cross-skill calls.
[0058] Exemplarily, for the target skill "weather report generation", based on its skill description "integrate multi-source meteorological data to generate a comprehensive report" and the execution protocol specification {"input": {"required": ["temperature","humidity"], "properties": {"temperature": {"type": "number"}, "humidity": {"type": "number", "maximum": 100}}}, "output": {"type": "object", "properties": {"report": {"type": "string"}, "severity_level": {"type": "integer", "minimum":1, "maximum": 5}}}}, a structured execution protocol is generated: {"action":"generate_weather_report", / / Define the action type of the skill call, identifying the specific functional module currently being executed."parameters":{ / / Declare all parameter configurations required for the skill call."input":{ / / Input parameter set, corresponding to the input data structure defined in the skill schema."temperature":"humidity":"pressure":},"constraints":{ / / Execution constraints, defining non-functional requirements."output_format":"markdown", / / The output format is forced to be Markdown syntax, corresponding to the format restriction in the skill description"timeout":3000 / / Timeout threshold (unit: milliseconds), and the interruption mechanism will be triggered if it exceeds 3000ms}},"output_schema":{ / / Output data structure definition, declaring the format specification of the skill execution result...}} In the embodiments of this specification, by performing structured conversion on the skill description and the execution protocol specification, a machine-readable standardized interface definition is achieved; through parameter rules, data consistency of input and output is ensured, and operation rules guarantee service quality, forming a verifiable and executable skill call framework, providing an accurate protocol basis for subsequent dependency analysis and skill combination.
[0059] In an alternative embodiment of the present specification, based on the skill description and execution protocol specification of the target skill, a structured execution protocol corresponding to the target skill of multiple target agents is generated, including the following specific steps: Based on the skill description and execution protocol specification of the target skill, context data is queried to generate a structured execution protocol corresponding to the target skill of multiple target agents, where the context data includes at least one of historical interaction data, real-time interaction request data, and external service response data.
[0060] Context data is a collection of dynamic information that can be accessed by an agent during task execution. Context data includes at least one of historical interaction data, real-time interaction request data, and external service response data. Optionally, context data is maintained through key-value storage or a vector database, supporting fast retrieval based on timestamps or semantic similarity, and is used for dynamic parameter filling in skill calls and execution strategy optimization. For example, in a weather query task, context data includes the user's historical location preferences (historical interaction data), the exact time range of the current request (real-time interaction request data), and real-time weather alerts returned by a third-party API (external service response data). Historical interaction data is a persistent record accumulated from past interactions between the agent and the user or the environment, including but not limited to user preference settings, frequent operation patterns, and historical task execution logs. Historical interaction data is stored through a long-term memory module, and its application scenarios include filling default values for personalized parameters (such as automatically filling the user's frequently used cities) and learning skill call priorities (such as preferentially selecting APIs with high user ratings). For example, historical interaction data may record that "in 8 out of the past 10 queries by user A, the user requested to display the temperature in Celsius", driving the temperature query skill to automatically set the parameter unit: "°C". Real-time interaction request data is instantaneous state information generated during the current task execution cycle, including user input data (such as follow-up statements in a multi-turn conversation), temporary variables (such as intermediate calculation results), and system monitoring metrics (such as CPU utilization). Optionally, real-time interaction data is maintained through a short-term memory module, using buffering for efficient reading and writing, and is used to address immediate dependencies (such as downstream skills needing to obtain the real-time output of upstream skills). For example, when the user adds a request "plus precipitation probability", the real-time interaction request data will add a new field {"require_precipitation": true}, triggering the dynamic expansion of skill combinations. External service response data is the structured result returned when the agent calls a third-party service or tool, including the API response body (such as JSON returned by a REST (Representational State Transfer) interface), the result set of a database query, and the status feedback of a hardware device. For example, external service response data may contain {"aqi": 45, "provider": "XWeather", "timestamp": "XXXX-03-20T14:00:00Z"}, which needs to be mapped to an internal standard format for use by downstream skills.
[0061] Exemplarily, when processing the task description text "query the weather in a certain place", query the historical interaction data to obtain the user's default city parameter: {"preferred_city": "XX"}. Parse the real-time interaction request data to extract the time range: {"time_range": ["XXXX-XX-XX"]}. Call the air quality API to obtain the external service response data: {"historical_data": [{"month": "XXXX-XX-XX", "avg_pm25":}, {"month": "XXXX-XX-XX", "avg_pm25":}]}. Integrate these context data to generate a structured execution protocol.
[0062] In the embodiments of the present specification, by establishing a multi-dimensional context data fusion mechanism, the dynamic and accurate generation of skill parameters is realized. The default parameter settings are optimized through the persistent learning of historical interaction data, the streaming processing of real-time interaction data ensures task coherence, and the standardized conversion of external service response data ensures system compatibility, forming a closed-loop parameter generation system, significantly improving the skill call accuracy and system robustness in complex task scenarios.
[0063] In an alternative embodiment of the present specification, step 106 includes the following specific steps: parse the structured execution protocols corresponding to the target skills of multiple target agents to obtain the execution dependencies corresponding to the target skills of multiple target agents; based on the execution dependencies corresponding to the target skills of multiple target agents, construct a skill call topology; based on the skill call topology, choreograph the target skills of multiple target agents to obtain a skill combination.
[0064] The execution dependency corresponding to the target skill is a set of data flow and control flow constraint conditions between skills parsed from the structured execution protocol, including the source of input parameters (e.g., the output field X of skill A is used as the input parameter Y of skill B), the execution order requirement (e.g., skill C must be started after skill D is completed), and the resource competition rule (e.g., skill E and skill F cannot share the GPU instance). The execution dependency can be represented by directed edges to form a precondition graph for skill calls. For example, in a weather query task, the execution dependency of the "temperature prediction" skill is {"required_inputs": ["coordinates"], "provider_skill": "location encoding"}, indicating that it requires the geographical coordinate parameters provided by the location encoding skill. The skill call topology is a graph structure constructed based on the execution dependency, which can be a directed acyclic graph. The nodes represent skill instances, and the edges represent the data transfer or execution order constraints between skills. The topological structure eliminates cyclic dependencies through the topological sorting algorithm and optimizes the parallel execution paths. For example, the skill call topology may include parallel branches: location encoding → (temperature prediction, precipitation analysis, wind speed calculation) → data aggregation → visualization, where the three skills in the parentheses can be executed in parallel.
[0065] Parse the structured execution protocols corresponding to the target skills of multiple target agents to obtain the execution dependencies corresponding to the target skills of multiple target agents. An optional way is: through the graph traversal algorithm, parse the mapping relationship of the input and output parameters of the structured execution protocols corresponding to the target skills of multiple target agents to generate the execution dependencies corresponding to the target skills of multiple target agents. Another optional way is: use a constraint satisfaction solver to calculate the parameter compatibility constraints of the structured execution protocols corresponding to the target skills of multiple target agents and deduce the execution dependencies corresponding to the target skills of multiple target agents. This is not limited here.
[0066] Based on the execution dependencies corresponding to the target skills of multiple target agents, construct a skill call topology. An optional way is: use the critical path analysis algorithm to identify the longest execution path based on the execution dependencies corresponding to the target skills of multiple target agents and construct a skill call topology. Another optional way is: use a constraint satisfaction solver to calculate the execution dependencies corresponding to the target skills of multiple target agents and construct a skill call topology. Another optional way is: use a reinforcement learning algorithm with the execution time as the goal to generate a skill call topology based on the execution dependencies corresponding to the target skills of multiple target agents. This is not limited here.
[0067] Exemplarily, for a weather query task involving four skills, parsing the structured execution protocol yields the following dependencies: Temperature prediction: {"depends_on": ["Location code"], "inputs": ["coordinates"]}; Precipitation analysis: {"depends_on": ["Location code"], "inputs": ["coordinates"]}; Wind speed calculation: {"depends_on": ["Location code"], "inputs": ["coordinates"]}; Data aggregation: {"depends_on": ["Temperature prediction", "Precipitation analysis", "Wind speed calculation"], "inputs": ["temp_data", "precip_data", "wind_data"]}. Build the skill call topology: A[Location code] --> B[Temperature prediction]; A --> C[Precipitation analysis]; A --> D[Wind speed calculation]; B --> E[Data aggregation]; C --> E; D --> E. Orchestrate to generate the skill combination.
[0068] In the embodiments of this specification, precise extraction of dependencies is achieved through structured protocol parsing. A conflict-free call topology is constructed using graph theory algorithms to generate a skill combination that takes into account both efficiency and correctness. This solves the problem of optimizing the execution order during multi-skill collaboration, improving execution feasibility and task execution efficiency.
[0069] In an alternative embodiment of this specification, step 108 includes the following specific steps: Load the skill combination into the asynchronous scheduling engine; Run the asynchronous scheduling engine to schedule the skill tools corresponding to the skill combination to execute the target task and obtain the task result.
[0070] The asynchronous scheduling engine is a distributed scheduling system that supports concurrent task execution and realizes parallel invocation and collaborative management of multi-skill tools through a task queue, a resource pool, and a fault tolerance mechanism. The asynchronous scheduling engine may include non-blocking task dispatching (such as based on an event loop), dynamic load balancing (such as proximity scheduling), and fault isolation (such as a single skill failure not affecting the overall process), and is suitable for the efficient execution of heterogeneous skill combinations. For example, the asynchronous scheduling engine may include a RabbitMQ task queue, a Kubernetes resource scheduler, and a Prometheus monitoring module to achieve parallel execution of the "Temperature prediction", "Precipitation analysis", and "Wind speed calculation" skills in a weather query task.
[0071] Run the asynchronous scheduling engine, schedule the skill tools corresponding to the skill combination to execute the target task, and obtain the task results. An optional method is to use message middleware, run the asynchronous scheduling engine, schedule the skill tools corresponding to the skill combination to execute the target task, and obtain the task results. Asynchronous dispatch and result callback of skill call requests are realized.
[0072] For example, load the skill combination to the task queue of the asynchronous scheduling engine, publish the stage 1 task to the geocoding service cluster through Kafka, and after receiving the completion event, initiate the stage 2 call to the weather model service and the third-party API at the same time. Listen to the response events of all skill tools, trigger the retry / supplement mechanism when timeout or failure occurs, and generate the task result: In the embodiments of this specification, an asynchronous scheduling engine is used to achieve parallel execution and efficient resource utilization of multi-skill tools, which solves the response delay problem caused by serial scheduling and significantly improves the processing throughput and reliability of complex tasks.
[0073] In an optional embodiment of the present specification, the task description text is a plurality of task description texts.
[0074] Multiple task description texts entered by the user in one interaction may involve different but related task objectives. Multiple task description texts can be identified and separated by delimiters (such as line breaks, semicolons) or semantic segmentation algorithms to support multi-task parallel processing. For example, the user input "query the weather in a certain place; then check the air quality in another place tomorrow" contains two task description texts, corresponding to two independent but parallel-processable tasks of weather query and air quality query.
[0075] For example, a user input containing two task description texts is processed: Input text: "Get the temperature data of a certain place for the next three days, and compare it with the data of another place during the same period". After semantic parsing, two subtasks are identified: Task 1: {"task": "Temperature query", "location": "a certain place", "duration": "3 days"}; Task 2: {"task": "Data comparison", "locations": ["a certain place", "another place"], "metric": "temperature"}. Generate reference content for skill combinations for parallel execution. Execute geocoding and temperature prediction of t1 and t2 in parallel.
[0076] In the embodiments of this specification, efficient execution of compound instructions is achieved through intelligent parsing and parallel processing of multi-task description texts. This solution breaks through the limitations of the traditional single-task processing mode and significantly improves the system response efficiency in complex interactive scenarios by dynamically building parallel skill combinations.
[0077] In an alternative embodiment of this specification, step 108 includes the following specific steps: In the case of detecting a missing skill with parameter missing in the structured execution protocol, use a large language model to generate parameter completion guidance information based on the structured execution protocol specification corresponding to the missing skill; feedback the parameter completion guidance information to the front end; receive the supplementary parameters sent by the front end; based on the supplementary parameters, update the structured execution protocol corresponding to the missing skill to obtain an updated skill combination; schedule the skill tools corresponding to the updated skill combination to execute the target task to obtain a task result.
[0078] The missing skill with parameter missing is a skill instance where the required input parameters are not satisfied during the skill execution. The judgment basis includes: 1) The field value marked as "required" in the structured execution protocol is empty; 2) The parameter type verification fails (e.g., a numeric parameter receives a string); 3) The context-passed parameter does not match the skill schema definition (e.g., the weather query skill requires geocoding input but the latitude and longitude coordinates are missing). For example, when executing the "Air quality prediction" skill, it is detected that the "Time range" parameter is missing, and this skill is marked as in the parameter missing state. The parameter completion guidance information is a structured prompt generated by the large language model to guide the user to provide the missing parameters, including the parameter name, data type, value example, and constraint description. Its generation process integrates the protocol specification of the skill schema and context reasoning, and is represented in a mixed form of natural language and structured templates. For example, the generated guidance information is: "Please specify the query time range (format: YYYY-MM-DD to YYYY-MM-DD)". The supplementary parameters are compliant parameter data provided by the user after responding to the guidance information, and need to be integrated into the execution context after passing the protocol verification. The supplementary parameters need to meet: the data type matches (e.g., the date string conforms to ISO 8601); the value range is valid (e.g., the temperature value is between -50~60°C); there is no conflict with the existing parameters (e.g., the time range does not exceed the API service period). For example, the user enters "YYYY-YY-YY" as the time range supplementary parameter. The updated skill combination is a skill call sequence reconstructed after integrating the supplementary parameters. The update mechanism ensures that the dependency relationship is maintained through topological sorting. For example, the single API call in the original skill combination is extended to a batch query by time period.
[0079] Based on the supplementary parameters, to update the structured execution protocol corresponding to the missing skill to obtain an updated skill combination, an alternative way is: adopt a graph reconstruction algorithm, and based on the supplementary parameters, adjust the node parameter configuration and edge dependency relationship of the skill call topology to obtain an updated skill combination.
[0080] Schedule the skill tools corresponding to the updated skill combination to execute the target task and obtain the task result. An optional way is to use a scheduling engine to schedule the skill tools corresponding to the updated skill combination to execute the target task and obtain the task result.
[0081] Exemplarily, when executing the "weather report generation" skill combination, a missing skill is detected. Missing skill identification: It is found that the "data visualization" skill lacks the "chart_type" parameter (the protocol requires the mandatory enumerated values ["line", "bar", "pie"]). Generate guiding information: "Please select the chart type: line chart (line) / bar chart (bar) / pie chart (pie)". The user supplements the parameter: "line". Add "parameters": {"visualization": {"chart_type": "line"}} to the structured execution protocol. Insert a "data transformation" node into the skill call topology to map the temperature data array into a time series format. Generate an interactive report containing a line chart of the temperature change over 24 hours.
[0082] In the embodiments of this specification, by detecting parameter missing in real time and generating accurate guiding information, a closed-loop parameter completion mechanism is achieved; by dynamically updating the protocol and reconstructing the skill combination, the continuous execution ability of the task chain is guaranteed, effectively solving the problem of incomplete parameters in open scenarios, significantly reducing the task interruption rate, and at the same time optimizing the input efficiency of complex parameters through human-computer collaboration.
[0083] In the application scenarios of an agent, skills are the core execution units of the agent, while tools are the specific implementation means of skills. To cope with increasingly complex project requirements, the agent needs to have the following capabilities: 1. Flexible expansion of combined skills: By associating multiple skills, a more powerful composite ability is formed to adapt to diverse task requirements. 2. Context awareness and dynamic parameter generation: During task execution, dynamically generate the parameters required for skill execution according to context data (such as user input, memory, external service return results, etc.) to ensure accurate skill invocation. 3. User interaction optimization: When the parameters required for skill execution are incomplete, the agent can actively interact with the user to clarify and complete the missing information, improving the task completion rate. 4. Multi-task parallel processing: When the user raises multiple questions or tasks, the agent can simultaneously invoke multiple skills to improve the response efficiency and service quality. Figure 2 The flowchart of a task processing method provided by an embodiment of this specification is shown as Figure 2 shown below: Front - end user question: The user inputs a task request through natural language (such as "query the weather in a certain place"), triggering the task - processing flow. Large - language model planning: Semantic parsing: Extract task elements (objectives, entities, constraints, etc.); Skill matching: Screen target agents and target skills according to the skill schema; Protocol generation: Output a structured execution protocol (including input and output specifications). Dependency analysis and skill combination: Topology construction: Analyze the data - flow / control - flow dependency relationships between protocols; Parallel orchestration: Generate a skill - call sequence in DAG structure. Asynchronous scheduling engine: Task distribution: Load the skill combination into the message queue; Dynamic execution: Schedule skill tools to run in parallel; Fault - tolerance processing: Monitor timeouts / failures and trigger a retry mechanism. Large - language model observation: Execution monitoring: Real - time track the execution status of skill tools (such as progress / duration / error code), and analyze the data quality of intermediate results (such as field integrity, numerical rationality); Dynamic adjustment: Protocol optimization: Modify the structured execution protocol according to runtime feedback; Combination reconstruction: Update the skill combination in real time.
[0084] Optionally, parameter - completion mechanism: Missing detection: Identify required but un - provided parameters; Guided generation: Prompt the user to clarify through natural language; Protocol update: Re - construct the skill combination after integrating the supplementary parameters.
[0085] Final - answer feedback: Data aggregation: Merge the outputs of multiple skills; Format conversion: Generate the final answer that meets the user's expectations; Front - end response: Return the structured final answer to the front - end user.
[0086] On the task platform of the large - language model, a set of intelligent - agent execution frameworks based on combined skills is designed.
[0087] The following combines the attached Figures 3 to 14 , taking the application of the task - processing method provided in this specification on the task platform of the large - language model as an example, to further illustrate the task - processing method.
[0088] Among them, Figure 3 shows a schematic diagram of an intelligent - agent list page provided by an embodiment of this specification for the task platform of the large - language model, as Figure 3Shown as follows: On the agent list page of the large language model task platform: The left navigation area includes: a home page button; a practice example button: to view preset task templates (such as the weather query combined skill); an agent center button (highlighted and selected): to enter the agent management module; a toolbox area: categorically display registered skill tools. The right agent management area includes: Agent name: an input box "Please enter content", a query button, a reset button, and a control for obtaining an import key credential. Creation entry: + New agent: to start the custom agent configuration process (corresponding to the determination of the target agent in step 104); Import agent: support for importing third-party agents through password credentials ("Obtain import password credentials"). Agent list: display deployed agents 1 and 2 (expandable). Each agent provides a complete operation chain: Configuration: set the skill schema and execution protocol specifications (corresponding to the generation of the structured execution protocol in step 104); Edit: modify the skill description and parameter rules; API debugging: verify the skill tool interface (corresponding to the scheduling execution in step 108); Evaluation results: view historical task execution metrics (such as the dependency analysis effect in step 106); Export: generate a shareable agent package.
[0089] Among them, Figure 4 shows a schematic diagram of an agent multi-version page of a task platform applied to a large language model provided by an embodiment of this specification, as Figure 4 shown: an input box "Please enter content" for the version name, a selection control "Please select" for the version status, a query button, a reset button, and an effect tracking button. The top includes version query: version name, version status; Creation entry: + New agent version: to start the custom agent version configuration process (corresponding to the determination of the target agent in step 104); Output version configuration; Import agent version: support for importing third-party agent versions through password credentials. The version list records the version name, version code, version status, version creator, version creation time, creation method, and operations of each version. By versioning the skill schema of the agent, ensure that protocol updates do not affect online tasks; Trace the version iteration of the skill through the build path; Use the network protection mechanism to solve resource conflicts during multi-version parallel scheduling.
[0090] Among them, Figure 5 shows a schematic diagram of an agent orchestration canvas of a task platform applied to a large language model provided by an embodiment of this specification, as Figure 5 shown: The agent orchestration canvas details the orchestration process: Start → Script task → Combined task → Result rendering task.
[0091] Among them, Figure 6 shows one of the schematic diagrams of adding skills to a combined task of a task platform applied to a large language model provided by an embodiment of this specification, asFigure 6 As shown: The input parameters of the combined task include name, source, type, required, and description; the execution task includes: selecting a tool or agent, skill name and description, and operations; skills can be added and user clarification skills can be added; the output parameters of the combined task include name, type, and description: test - object - test; Success - boolean - whether successful; errorCode - string text - error code; errorMessage - string text - error message; Content - string text - content; taskId - string text - task code.
[0092] Among them, Figure 7 Figure 2 shows a schematic diagram of adding skills to a combined task in a task platform for large language models provided by an embodiment of this specification. As Figure 7 shown: The input parameters of the script task include name, source, type, required, and description; the execution of the editing tool task includes selecting a tool or agent (testing the api by oneself), skill name (testing), applicable scenario (querying weather forecast), capabilities, limiting conditions (not applicable to scenarios other than querying weather forecast), continuing to execute the task when an exception occurs, pre - processing of input parameters (no processing, script processing, and planning processing), and script. Among them, the script is: def preprocesssTooParam(tool_param)……@param: tool_param, the tool input string @return: the processed tool input string Explanation: Both the input and output parameters are Json strings. For the original {"userId": "111"} input parameter, it is changed to {"userId": "222"}……return {"userId": "222"}.
[0093] Among them, Figure 8 Figure 3 shows a schematic diagram of pre - processing and post - processing in a task platform for large language models provided by an embodiment of this specification. As Figure 8 shown: The input parameters of the script task include name, source, type, required, and description; the execution of the editing tool task includes selecting a tool or agent (testing the api by oneself), skill name (testing), applicable scenario (querying weather forecast), capabilities, limiting conditions (not applicable to scenarios other than querying weather forecast), continuing to execute the task when an exception occurs, pre - processing of input parameters (no processing, script processing, and planning processing), input parameters (parameter name, field source, type, required, and description), post - processing of results (directly returning the result, script conversion, and simulated result output), and script. Among them, the input parameters (parameter name, field source, type, required, and description) and post - processing of results (directly returning the result, script conversion, and simulated result output) can be configured manually.
[0094] Among them, Figure 9Shows a schematic diagram of the concurrent execution of combined skills of a task platform applied to a large language model provided by an embodiment of this specification, as Figure 9 shown: On the left is the dialogue debugging area. Through multiple rounds of dialogue, the user generates multiple batches of test responses for the skill combination: - [{"content": "["errorMessages":],"success":true,"data":"result":"The result information of the user's query is as follows: User name: AAA, User phone: XXXXXXXX, User address: XXXXXXXXXX","errorCode":null,"errorMsg":null,"extraData":null,"originld":null,"env":null,"other":null,"firstErrorMessage":null,"failure":false)", "isStream": false, "listlndex": 1, "success": true, "taskld": "6c111077-9d66-4800-aa3f-331364303eb8","taskName": "Test"}, {"content": ""errorMessages":[],"success":true,"data":"result":"The result information of the user's query is as follows: User name: AAA, User phone: XXXXXXXX, User address: XXXXXXX"},"errorCode":null"erorMsg":null"extraData":null,"originld":null"env":nulL,"firstErrorMessage":null,"failure":false}","isStream":"false,"listindex":0,"success":true,"taskld":"cd38c945-9544-46d4-b347-2bb3a0b9f04d", "taskName":"Test"}]. In the middle is the intelligent agent task rule area where the flow chart is displayed; on the right are the log and model inference process areas, which show the execution status of each skill during the task execution.
[0095] Among them, Figure 10 shows a schematic diagram of the overall architecture of a task platform applied to a large language model provided by an embodiment of this specification, as Figure 10Shown as follows: The scenario layer includes: Scenarios 1 - n, Designer Debugging Module, OPEN - API Interface. The orchestration component includes: Skill Orchestration Component, Skill Combination, Tools, Agents, Document Retrieval, Data Table Retrieval, Graphic and Text Retrieval, Multi - tenant Foot Bone, System Observability, OneRAG Suite, and SQL Executor. The skill capability layer includes: Asynchronous Scheduling Engine, Header Encryption, Header Function Value Retrieval, Input Parameter / Header Default Value, Asynchronous Tools, Asynchronous Callback Expiration Policy, Skill Simulation, Skill Exception Ignoring and Continuing Execution Ability, Debug / Breakpoint Ability, Input Parameter Pre - processing, Skill Invocation, Result Post - processing, Skill Combination Parallel Ability, and Streaming Tools. The skill execution layer includes: API Executor, Knowledge Base Retrieval, OneRAG Suite Executor, Large Language Model Speech Prompt Executor, and SQL Executor. Tool registration includes: Custom API registration: Curl Parsing Registration, Tool Debugging, Request Method, Header Definition, Input and Output Parameter Definition, Whether Asynchronous, Whether Streaming, Tenant & Space Isolation...; Platform Tools: Prompt Large Language Model Tools, Calculator, asr, tts, SQL Executor, OneRAG Suite, Knowledge Base, and Tenant.
[0096] Figure 11 The following shows a schematic diagram of the combined skill architecture of a task platform for large language models provided by an embodiment of this specification, as Figure 11 Shown as follows: From the previous node, multiple API tools are called through skill combination to achieve the streaming variable update of the skill combination, and the subsequent nodes are executed sequentially.
[0097] Figure 12 The following shows a schematic diagram of generating the skill input parameter schema of a task platform for large language models provided by an embodiment of this specification, as Figure 12 Shown as follows: It is possible to select whether to continue executing in case of task exception. The input parameters of the script task include name, source, type, required, and description. The execution of the editing tool task includes input parameter pre - processing (no processing, script processing, and planning processing), input parameter and post - processing after completion (directly return the result, script conversion, simulated result output). In generating the skill input parameter schema, the input parameters include the settings of parameter name, field source, type, required, and description: Input - Prompt: User input extraction question - string - no - open for modification; User - Constant: {"userId": "243243"} - no -...; Username - Constant: XXXX - string - no - not open for modification.
[0098] Figure 13 The following shows a schematic diagram of the parallel execution of multiple tasks in the combined skills of a task platform for large language models provided by an embodiment of this specification, as Figure 13Shown as follows: Multiple skills and their parameters are generated during planning. Multiple skills are generated simultaneously in the combined skills. The asynchronous scheduling engine will load multiple skills and initiate calls to the skill services at the same time. In the left dialogue debugging area, the user generates multiple batches of test responses for skill combinations through multiple rounds of dialogue: - [{"content": "["errorMessages":], "success": true, "data": "result": "The result information queried by the user is as follows: User name: AAA, User phone: XXXXXXXX, User address: XXXXXXXXXX", "errorCode": null, "errorMsg": null, "extraData": null, "originld": null, "env": null, "other": null, "firstErrorMessage": null, "failure": false), "isStream": false, "listlndex": 1, "success": true, "taskld": "6c111077-9d66-4800-aa3f-331364303eb8", "taskName": "Test"}, {"content": ""errorMessages": [], "success": true, "data": "result": "The result information queried by the user is as follows: User name: AAA, User phone: XXXXXXXX, User address: XXXXXXX"},"errorCode": null, "erorMsg": null, "extraData": null, "originld": null, "env": nulL, "firstErrorMessage": null, "failure": false}", "isStream": "false", "listindex": 0, "success": true, "taskld": "cd38c945-9544-46d4-b347-2bb3a0b9f04d", "taskName": "Test"}]. The flowchart is displayed in the middle intelligent agent task rule area; in the right log and model inference process area, the execution status of each skill during the task execution is displayed.
[0099] Figure 14 The schematic diagram of the input supplementary information of a task platform applied to a large language model provided by an embodiment of this specification is shown, as Figure 14As shown: When the necessary parameters for executing a skill are not met, the user inputs supplementary information again. According to the skill schema description, when the large model is planning, it can sense whether the necessary parameters for the execution tool are missing based on the context information, and gives the user a clarification Q&A when missing. In the left dialogue debugging area, the user has a multi-round dialogue: - May I ask what information you need? I will do my best to help you. - Query the weather in a certain place. - To query the weather in a certain place, I need to know which day's weather you want to obtain, and whether you need detailed weather information, such as temperature, humidity, wind speed, etc. The right intelligent agent task rule area shows a flowchart: Start → Task Planning (Prompt full text, reasoning result) → User Intent Clarification Task (output result).
[0100] The above Figures 3 to 14 The intelligent agent execution framework on the task platform of the large language model mentioned above specifically includes the following key features: 1. In the combined skills, by associating multiple skills and adding skills in the intelligent agent by combining the ability description and input parameter schema protocol of the skill, in react, it can accurately launch the skill by combining the planning and thinking ability of the large model; 2. When the intelligent agent plans and executes a skill, through the skill schema defined by the combined skill, the dynamic input parameters for skill execution can be extracted and generated from the context data during the planning to launch the skill; 3. When the user's question matches the required execution skill, but the input parameters of the execution skill are incomplete, the ability to clarify for the user is provided, and the user is asked again to complete the execution skill parameters; 4. When the user asks multiple questions and hits multiple skills in the combined skill, the combined skill supports parallel execution of multiple skills at one time.
[0101] Corresponding to the above method embodiment, this specification also provides a task platform embodiment, Figure 15 showing a schematic structural diagram of a task platform provided by an embodiment of this specification. The task platform 1500 includes a front-end task interface 1510 and a response unit 1520; The front-end task interface 1510 is used to receive the task description text of the target task sent by the front end; the response unit 1520 is used to use the large language model to extract the task elements of the target task from the task description text, and based on the task elements, determine at least one target skill of the target intelligent agent from the preset skills of multiple intelligent agents, and generate a structured execution protocol corresponding to the target skills of multiple target intelligent agents; based on the structured execution protocols corresponding to the target skills of multiple target intelligent agents, perform dependency analysis on the target skills of multiple target intelligent agents to obtain a skill combination; schedule the skill tools corresponding to the skill combination to execute the target task, obtain the task result, and send the task result to the front end.
[0102] Optionally, the response unit 1520 is specifically configured to: use a large language model to semantically parse the task description text, and extract task elements of the target task from the task description text, where the task elements include at least one of an entity of the target task, an operation instruction, and associated information.
[0103] Optionally, the response unit 1520 is specifically configured to: based on the task elements, and the matching degree between the skill descriptions of the preset skills of the agents in the candidate agent set and the line protocol specifications of the preset skills, determine multiple target agents and the target skills of the multiple target agents from the candidate agent set, where the structured execution protocol specification includes parameter rules and / or operation rules.
[0104] Optionally, the response unit 1520 is specifically configured to: based on the skill descriptions and execution protocol specifications of the target skills of the multiple target agents, generate structured execution protocols corresponding to the target skills of the multiple target agents, where the structured execution protocol specification includes parameter rules and / or operation rules.
[0105] Optionally, the response unit 1520 is specifically configured to: based on the skill descriptions and execution protocol specifications of the target skills of the multiple target agents, query context data, and generate structured execution protocols corresponding to the target skills of the multiple target agents, where the context data includes at least one of historical interaction data, real-time interaction request data, and external service response data.
[0106] Optionally, the response unit 1520 is specifically configured to: parse the structured execution protocols corresponding to the target skills of the multiple target agents to obtain the execution dependencies corresponding to the target skills of the multiple target agents; based on the execution dependencies corresponding to the target skills of the multiple target agents, construct a skill call topology; and based on the skill call topology, orchestrate the target skills of the multiple target agents to obtain a skill combination.
[0107] Optionally, the response unit 1520 is specifically configured to: load the skill combination into an asynchronous scheduling engine; run the asynchronous scheduling engine to schedule the skill tools corresponding to the skill combination to execute the target task, and obtain a task result.
[0108] Optionally, the task description text is multiple task description texts.
[0109] Optionally, the response unit 1520 is specifically configured to: when detecting a missing skill with a missing parameter in the structured execution protocol, use a large language model to generate parameter completion guidance information based on the structured execution protocol specification corresponding to the missing skill; feedback the parameter completion guidance information to the front end; receive the supplementary parameters sent by the front end; update the structured execution protocol corresponding to the missing skill based on the supplementary parameters to obtain an updated skill combination; schedule the skill tools corresponding to the updated skill combination to execute the target task to obtain a task result.
[0110] In the embodiments of the present specification, the task platform realizes the semantic alignment and dynamic adaptation of the multi-modal skill interface by using a large language model to perform semantic parsing and element extraction on the task description text, and solves the cooperation problem between heterogeneous skill modules; by generating a structured execution protocol and performing dependency analysis, a scalable skill combination framework is constructed, significantly improving the autonomous planning and task execution capabilities of complex tasks. Through the semantic-driven skill combination mechanism, the limitation of a single intelligent agent is broken through. At the same time, relying on the generalization and understanding ability of the large language model, the accuracy and scalability of processing the target task by skill and tool invocation in an open scenario are realized.
[0111] The above is a schematic solution of a task platform in this embodiment. It should be noted that the technical solution of this task platform and the technical solution of the above task processing method belong to the same concept. For the details not described in detail in the technical solution of the task platform, reference can be made to the description of the technical solution of the above task processing method.
[0112] Figure 16 The structural block diagram of a computing device provided by an embodiment of the present specification is shown. The components of the computing device 1600 include, but are not limited to, a memory 1610 and a processor 1620. The processor 1620 is connected to the memory 1610 through a bus 1630, and the database 1650 is used to store data.
[0113] The computing device 1600 also includes an access device 1640, which enables the computing device 1600 to communicate via one or more networks 1660. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 1640 may include one or more of any type of wired or wireless network interface (e.g., a Network Interface Controller (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC).
[0114] In one embodiment of the present specification, the above components of the computing device 1600 and Figure 16 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 16 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.
[0115] The computing device 1600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 1600 can also be a mobile or stationary server.
[0116] Wherein, the processor 1620 is used to execute the following computer program / instructions, and when the computer program / instructions are executed by the processor, the steps of the above task processing method are implemented.
[0117] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the above task processing method belong to the same concept. For the detailed content not described in the technical solution of the computing device, reference can be made to the description of the technical solution of the above task processing method.
[0118] An embodiment of this specification also provides a computer-readable storage medium, which stores computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above task processing method are implemented.
[0119] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above task processing method belong to the same concept. For the detailed content not described in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above task processing method.
[0120] An embodiment of this specification also provides a computer program product, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above task processing method are implemented.
[0121] The above is a schematic solution of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above task processing method belong to the same concept. For the detailed content not described in the technical solution of the computer program product, reference can be made to the description of the technical solution of the above task processing method.
[0122] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain implementations, multitasking and parallel processing are also possible or may be advantageous.
[0123] The computer instructions include computer program code, which may be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, external hard drives, magnetic disks, optical discs, computer memories, read-only memories (ROM), random access memories (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0124] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the embodiments of this specification are not limited by the described order of actions, because according to the embodiments of this specification, some steps may be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential for the embodiments of this specification.
[0125] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0126] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not elaborate on all the details and do not limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A task processing method, comprising: Obtaining the task description text of the target task; Using a large language model to extract the task elements of the target task from the task description text, and based on the task elements, determining multiple target agents and the target skills of the multiple target agents from a candidate agent set, and generating a structured execution protocol corresponding to the target skills of the multiple target agents; Based on the structured execution protocols corresponding to the target skills of the multiple target agents, performing a dependency analysis on the target skills of the multiple target agents to obtain a skill combination; Scheduling the skill tools corresponding to the skill combination to execute the target task to obtain a task result.
2. The method according to claim 1, wherein the using a large language model to extract the task elements of the target task from the task description text comprises: Using the large language model to perform semantic parsing on the task description text, and extracting the task elements of the target task from the task description text, wherein the task elements include at least one of an entity, an operation instruction, and associated information of the target task.
3. The method according to claim 1, wherein the determining multiple target agents and the target skills of the multiple target agents from a candidate agent set based on the task elements comprises: Based on the task elements, and the matching degree between the skill descriptions of the preset skills of the agents in the candidate agent set and the protocol specifications of the preset skills, determining multiple target agents and the target skills of the multiple target agents from the candidate agent set, wherein the structured execution protocol specifications include parameter rules and / or operation rules.
4. The method according to claim 1, wherein the generating a structured execution protocol corresponding to the target skills of the multiple target agents comprises: Based on the skill descriptions and execution protocol specifications of the target skills of the multiple target agents, generating a structured execution protocol corresponding to the target skills of the multiple target agents, wherein the structured execution protocol specifications include parameter rules and / or operation rules.
5. The method according to claim 4, wherein the generating a structured execution protocol corresponding to the target skills of the multiple target agents based on the skill descriptions and execution protocol specifications of the target skills of the multiple target agents comprises: Based on the skill descriptions and execution protocol specifications of the target skills of the multiple target agents, querying context data to generate a structured execution protocol corresponding to the target skills of the multiple target agents, wherein the context data includes at least one of historical interaction data, real-time interaction request data, and external service response data.
6. The method according to claim 1, wherein the performing a dependency analysis on the target skills of the multiple target agents based on the structured execution protocols corresponding to the target skills of the multiple target agents to obtain a skill combination comprises: Parsing the structured execution protocols corresponding to the target skills of the multiple target agents to obtain the execution dependency relationships corresponding to the target skills of the multiple target agents; Construct a skill invocation topology based on the execution dependencies corresponding to the target skills of the multiple target agents. Based on the skill invocation topology, orchestrate the target skills of the multiple target agents to obtain a skill combination.
7. The method according to claim 1, wherein scheduling the skill tools corresponding to the skill combination to execute the target task to obtain a task result includes: Load the skill combination into an asynchronous scheduling engine. Run the asynchronous scheduling engine to schedule the skill tools corresponding to the skill combination to execute the target task to obtain a task result.
8. The method according to claim 1, wherein the task description text is multiple task description texts.
9. The method according to any one of claims 1-8 further includes: When detecting a missing skill with missing parameters in the structured execution protocol, use the large language model to generate parameter completion guidance information based on the structured execution protocol specification corresponding to the missing skill. Feed back the parameter completion guidance information to the front end. Receive the supplementary parameters sent by the front end. Based on the supplementary parameters, update the structured execution protocol corresponding to the missing skill to obtain an updated skill combination. Schedule the skill tools corresponding to the updated skill combination to execute the target task to obtain a task result.
10. A task platform, comprising a front-end task interface and a response unit; The front-end task interface is used to receive the task description text of the target task sent by the front end. The response unit is used to use the large language model to extract the task elements of the target task from the task description text, determine multiple target agents and the target skills of the multiple target agents from the candidate agent set based on the task elements, and generate the structured execution protocols corresponding to the target skills of the multiple target agents. Based on the structured execution protocols corresponding to the target skills of the multiple target agents, perform dependency analysis on the target skills of the multiple target agents to obtain a skill combination, schedule the skill tools corresponding to the skill combination to execute the target task, and obtain a task result. The front-end task interface is further used to send the task result to the front end.
11. A computing device, comprising: A memory and a processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 9 are implemented.
12. A computer-readable storage medium storing computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
13. A computer program product comprising computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
Citation Information
Patent Citations
Intelligent agent generation method and device, electronic equipment and storage medium
CN117764107A
Function calling method, function packaging method and device
CN117850875A
Problem solving method and device based on intelligent agent, storage medium and product
CN118469023A
Human-computer interaction method and device
CN118551000A
Multi-agent cooperation method and system
CN118917632A
Cited By
Information processing method and device, storage medium and computer equipment
CN120706578A
Information processing method and device, storage medium and computer device
CN120706578B
Task processing method and device based on tool and electronic equipment
CN120892474A
General instrument image processing method and device and electronic equipment
CN120976908A
Intelligent timed task configuration and multi-mode feedback system based on natural language interaction
CN121009303A