Multi-task hybrid execution method and system for an agent

CN122884571APending Publication Date: 2026-10-09NR ELECTRIC CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610926446.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-10-09

AI Technical Summary

Technical Problem

对于无API、无MCP、无CLI、无数据库访问权限的封闭式软件,智能体无法直接调用其功能

Benefits of technology

[0034]第四方面,本申请还提供一种计算机可读存储介质,其中存储有计算机可读指令,当计算机读取并执行计算机可读指令时,实现上述第一方面的方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122884571A_ABST
    Figure CN122884571A_ABST
Patent Text Reader

Abstract

The application provides a multi-task hybrid execution method and system of an intelligent agent, which comprises the following steps: receiving a business task, analyzing the business task to obtain a task context; identifying a target system according to the task context, and querying corresponding capability metadata; judging the execution capability of the target system based on the capability metadata, and determining the execution mode of the business task; when the intelligent agent has a structured calling capability and the structured calling capability can meet the requirements, obtaining a structured calling result based on a structured calling path; when the intelligent agent does not have a structured calling capability or the structured calling capability cannot meet the requirements, obtaining an interface-level execution result based on an interface-level execution path; when multiple target systems or multiple execution modes are involved, obtaining a hybrid execution result based on a hybrid execution path; and verifying the execution process, gathering the execution results, and obtaining a task execution result. In this way, the intelligent agent sets an execution mode selection node according to the target system capability metadata and the task risk level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent power operation and maintenance technology, and in particular to a method and system for multi-task hybrid execution of an intelligent agent. Background Technology

[0002] With the development of large language models and intelligent agent technology, intelligent agent systems can complete task understanding, task decomposition, tool invocation, and result generation based on natural language, and have been gradually applied to scenarios such as knowledge question answering, data analysis, assisted operation and maintenance, and business process automation. Existing intelligent agents typically understand user intent through large language models, autonomously generate execution steps, and select external tools, business interfaces, or knowledge bases to complete tasks.

[0003] In industrial production operations such as power dispatching, equipment maintenance, fault analysis, simulation calculation, inspection management, and report generation, intelligent agents need to call upon different business systems, data services, professional algorithms, and software tools to complete tasks such as data querying, analysis and calculation, result generation, and decision support. The software systems in actual industrial settings are complex, including systems with APIs, Web services, database interfaces, command-line interfaces, or MCP services, as well as a large number of existing industrial software lacking standard interfaces, closed desktop software, professional simulation tools, vendor-specific systems, and legacy business systems.

[0004] For systems with standard interfaces, intelligent agents can complete tasks through structured tool calls, such as calling APIs to query ledger information, obtaining operational data through data services, executing calculation programs through command-line tools, and accessing professional analysis capabilities through MCP services. However, for software systems lacking open interfaces, business personnel typically have to manually operate the interface, such as opening the software, selecting menus, entering parameters, importing files, clicking calculations, waiting for results, and exporting reports. Although such software carries important business capabilities, it is difficult for intelligent agents to directly invoke them.

[0005] Existing agent execution solutions based on API or tool calls are highly dependent on whether the target system has open interfaces. For closed software without APIs, MCPs, CLIs, or database access, agents cannot directly call its functions. Although traditional RPA technology can simulate human operation interfaces, it usually relies on fixed scripts, is sensitive to interface changes, and lacks natural language task understanding, multi-system collaboration, dynamic path planning, and unified governance capabilities. GUI agents based on multimodal large models can understand some interface content, but if the model directly operates the interface autonomously, it is prone to problems such as misidentification, accidental clicks, accidental input, accidental triggering of high-risk actions, uncontrollable states, and difficulty in auditing the process.

[0006] Furthermore, existing solutions typically separate structured interface calls and automated interface operations into different tool systems, lacking a unified description of target system capabilities, execution path selection, task orchestration, result aggregation, and operational governance mechanisms. For complex tasks, it may be necessary to simultaneously call structured interfaces, query knowledge bases, operate desktop software, parse exported files, and generate final reports. If different execution methods cannot be uniformly orchestrated and tracked, it will affect the task's closed-loop capability and the reliability of the final result. Summary of the Invention

[0007] This application provides a method and system for multi-task hybrid execution of intelligent agents, enabling the selection or combination of structured call paths, interface-level execution paths, and hybrid execution paths, thereby achieving low-intrusion, traceable, and controllable execution of interface-enabled systems and interface-less industrial software.

[0008] Firstly, this application provides a method for multi-task hybrid execution of an intelligent agent, which is executed by a computing device. The computing device can be understood as a computer or similar device, and is not limited thereto in this application. The method includes:

[0009] The system receives a business task, parses it to obtain a task context, which includes at least one of the following: task objective, target system, operation object, input parameters, expected output, and execution constraints. Based on the task context, it identifies the target system and queries the corresponding capability metadata. This capability metadata describes at least one of the target system's interface capabilities, interface operation capabilities, permission requirements, risk level, and operation constraints. Based on the capability metadata, it determines the target system's execution capability and, based on the result, determines the execution method of the business task. The execution methods include structured call paths, interface-level execution paths, and hybrid execution paths. When the target system has structured call capabilities and these capabilities meet the requirements, a structured call result is obtained based on the structured call path. When the target system lacks structured call capabilities or its structured call capabilities do not meet the requirements, an interface-level execution result is obtained based on the interface-level execution path. When the business task involves multiple target systems or multiple execution methods, a hybrid execution result is obtained based on the hybrid execution path. The execution processes of the structured call paths, interface-level execution paths, and hybrid execution paths are verified, and the execution results are aggregated to obtain the task execution result.

[0010] By establishing capability metadata for the target system, this application enables the system to identify its structured interface capabilities and interface operation capabilities, and accordingly to split the flow between structured call paths, interface-level execution paths, and hybrid execution paths. This allows the intelligent agent to select the appropriate execution path based on the actual capabilities of the target system, thereby improving its adaptability to both interface-enabled and interface-less industrial software.

[0011] In the aforementioned multi-task hybrid execution method for intelligent agents, the capability metadata includes at least one of the following: target system identifier, target system type, interface type, executable actions, input parameter definition, output parameter definition, structured call address, interface operation entry point, permission level, risk level, operation whitelist, prohibited actions, manual confirmation policy, timeout policy, retry policy, rollback policy, audit policy, interface template, interface element description, status verification rules, and result extraction rules.

[0012] In this way, the present application limits the capability metadata, enabling the intelligent agent to select nodes by setting the execution method according to the target system capability metadata and task risk level, and to judge and distribute the structured interface availability, interface operation availability and multi-system task combination.

[0013] In the aforementioned multi-task hybrid execution method for intelligent agents, the execution capability of the target system is determined based on capability metadata, and the execution method of the business tasks is determined based on the execution capability determination result, including:

[0014] When the capability metadata indicates that the target system supports at least one of the target structured call interfaces in the list, and the structured call interface can meet the current task execution requirements, the subtask corresponding to the target system is determined to use a structured call path; when the capability metadata indicates that the target system does not support structured call interfaces, or the target system supports structured call interfaces but the structured call interface cannot meet the current task execution requirements, and the target system supports graphical interface operation, the subtask corresponding to the target system is determined to use a UI-level execution path; when a business task involves multiple target systems and different target systems correspond to different execution methods, the business task is determined to use a hybrid execution path; when a business task involves high-risk actions, at least one of the following is triggered: manual confirmation, dual confirmation, disabling automatic execution, or generating only operation suggestions.

[0015] Through the above methods, this application integrates different execution methods into a unified system of task orchestration, access control, risk verification, log auditing, anomaly rollback, result traceability, and capability accumulation. By uniformly governing the execution process through access verification, whitelist verification, risk level judgment, manual confirmation, log auditing, screenshot documentation, and anomaly rollback, the risk of mis-calling of large models, misoperation of interfaces, unauthorized calls, and accidental triggering of high-risk actions is reduced.

[0016] In the aforementioned multi-task hybrid execution method for intelligent agents, when a business task involves multiple target systems or multiple execution methods, a hybrid execution result is obtained based on the hybrid execution path, including:

[0017] When a business task involves multiple target systems or multiple execution methods, the business task is broken down into multiple subtasks; the target system and execution method corresponding to each subtask are determined, and the subtasks are executed according to the corresponding execution method; one or more of the following subtasks—structured call subtasks, interface-level execution subtasks, knowledge retrieval subtasks, file parsing subtasks, and report generation subtasks—are arranged into a task chain according to their dependencies; after any subtask is completed, the task context is updated according to the output of that subtask, and the subsequent subtasks are driven to execute based on the updated task context.

[0018] Through the above methods, this application records task input, execution path, calling parameters, interface operation steps, status verification results, exported files, exception information and final results through result aggregation, context recording and audit trail mechanisms, thereby improving the traceability and verifiability of the task execution process.

[0019] In the aforementioned multi-task hybrid execution method for intelligent agents, the structured invocation capability includes:

[0020] The system calls to the target system using at least one of the following methods: application programming interface (API) calls, model context protocol service calls, large model function calls, command-line interface command calls, network service calls, database queries, script calls, file parsing tool calls, message queue calls, plugin calls, or vendor software development kit (SDK) calls. When the target system possesses structured calling capabilities that meet the task execution requirements, the system generates call parameters based on the structured call path and calls the corresponding structured capabilities of the target system to obtain the structured call result. Calling the corresponding structured capabilities of the target system includes: matching callable capabilities according to the task objective; generating call parameters based on the task context; performing format and permission checks on the call parameters; calling the callable capabilities; receiving the returned results; and performing format conversion, result verification, and archiving on the returned results.

[0021] Through the above methods, this application sets up multiple structured calling methods, which increases the application scope of structured calling and improves the reliability of subsequent returned results.

[0022] In the aforementioned multi-task hybrid execution method for intelligent agents, when the target system lacks structured invocation capabilities or its structured invocation capabilities do not meet the requirements, interface-level execution results are obtained based on the interface-level execution path, including:

[0023] When the target system lacks structured calling capabilities or its structured calling capabilities cannot meet the task execution requirements, the interface state of the target system is obtained based on the interface-level execution path. Interface elements are identified, an operation sequence is generated according to the task objective and interface state, interface actions are executed according to the operation sequence, and state verification is performed after the interface actions are executed to obtain the interface-level execution result. Among these methods, obtaining the interface state of the target system includes obtaining desktop software interface elements through the operating system control tree; obtaining web page elements through the browser document object model structure; identifying windows, buttons, input boxes, menus, tables, or pop-ups through interface screenshots; identifying interface text through optical character recognition; identifying interface elements and their executable actions through a multimodal model; and obtaining the current interface state through window handles, coordinate regions, interface templates, or historical operation trajectories.

[0024] In this way, when the target system does not have structured calling capability or the structured calling capability cannot meet the task execution requirements, this application obtains the interface state of the target system through the interface-level execution path, specifies the conditions for using the interface-level execution path, and sets up multiple methods for obtaining the interface state of the target system to improve the feasibility and versatility of the solution.

[0025] In the aforementioned multi-task hybrid execution method for intelligent agents, the sequence of operations for generating task objectives and interface states includes:

[0026] Based on the task objective, determine the interface operation steps to be completed, and based on the interface state, determine the target interface elements corresponding to the interface operation steps; generate an operation description for each interface operation step, which includes at least one of the following: step number, action type, target element, input value, preconditions, expected state, timeout, number of retries, risk level, and whether manual confirmation is required; wherein, the action type includes at least one of the following: click, input, select, drag and drop, upload file, import data, execute menu command, start calculation, wait for result, export file, save file, close pop-up window, and switch tabs.

[0027] Through the above method, this application generates an operation description for each interface operation step, and provides detailed limitations on the content of the operation description, thus expanding the scope of the operation description.

[0028] The aforementioned method for multi-task hybrid execution of intelligent agents also includes:

[0029] When a business task is successfully executed and meets the conditions for capability accumulation, at least one of the following is extracted: task type, target system, input parameter template, execution path, operation steps, status verification rules, exception handling strategy, permission requirements, risk level, and output result format. Based on the extraction results, a reusable capability template is generated and registered as a Skill, tool, plugin, workflow template, interface operation template, automated process, or intelligent entity sub-capability for subsequent similar tasks to call.

[0030] Through the above methods, this application will successfully execute paths as Skills, workflow templates, interface operation templates, or intelligent entity sub-capabilities, transforming existing software capabilities that originally relied on manual operation into reusable, manageable, and governable capability assets.

[0031] In a second aspect, this application provides a multi-task hybrid execution system for an intelligent agent, used to execute the method provided in the first aspect of this application, including: a receiving module, a querying module, a judging module, and a verification module;

[0032] The system comprises the following modules: a receiving module, which receives business tasks, parses them, and obtains the task context; the task context includes at least one of the following: task objective, target system, operation object, input parameters, expected output, and execution constraints; a query module, which identifies the target system based on the task context and queries the corresponding capability metadata of the target system; the capability metadata describes at least one of the following: interface capabilities, interface operation capabilities, permission requirements, risk level, and operation constraints of the target system; and a judgment module, which judges the execution capability of the target system based on the capability metadata and determines the execution method of the business task based on the judgment result; the execution method of the business task includes a structured call path. The system comprises three execution modules: a UI-level execution path and a hybrid execution path; a judgment module, which is also used to obtain a structured call result based on the structured call path when the target system has structured call capability and the structured call capability meets the requirements; a judgment module, which is also used to obtain a UI-level execution result based on the UI-level execution path when the target system does not have structured call capability or the structured call capability does not meet the requirements; a judgment module, which is also used to obtain a hybrid execution result based on the hybrid execution path when the business task involves multiple target systems or multiple execution methods; and a verification module, which is used to verify the execution process of the structured call path, the UI-level execution path, and the hybrid execution path, and to aggregate the execution results to obtain the task execution result.

[0033] Thirdly, this application also provides a computing device, comprising: a memory for storing program instructions; and a processor for calling the program instructions stored in the memory and executing the method described in the first aspect according to the obtained program instructions.

[0034] Fourthly, this application also provides a computer-readable storage medium storing computer-readable instructions, which, when read and executed by a computer, implement the method of the first aspect described above.

[0035] Fifthly, this application provides a computer program product including a computer program executable by a computer device, which, when run on the computer device, causes the computer device to perform the method described in the first aspect.

[0036] Beneficial Effects: By establishing capability metadata for the target system, the system can identify its structured interface capabilities and user interface operation capabilities, and accordingly, distribute traffic among structured call paths, interface-level execution paths, and hybrid execution paths. This allows the agent to select the appropriate execution path based on the actual capabilities of the target system, improving adaptability to both interface-enabled and interface-less industrial software. A closed-loop interface-level execution mechanism is formed through interface perception, interface element recognition, operation planning, action execution, status verification, and exception handling, enabling interface-less desktop software, web systems, remote desktop systems, and professional simulation tools to be incorporated into the agent's task chain without large-scale modifications to the original system. The unified orchestration of capabilities such as structured calls, interface-level execution, knowledge retrieval, file parsing, and report generation into a hybrid task chain supports complex industrial tasks across multiple systems, tools, and execution methods. Unified governance of the execution process through permission verification, whitelist verification, risk level judgment, manual confirmation, log auditing, screenshot documentation, and exception rollback reduces the risks of erroneous large-model calls, interface misoperations, unauthorized calls, and accidental triggering of high-risk actions. By recording task inputs, execution paths, call parameters, interface operation steps, status verification results, exported files, exception information, and final results through result aggregation, context logging, and audit trail mechanisms, the traceability and verifiability of the task execution process are improved. Successful execution paths can be precipitated as Skills, workflow templates, interface operation templates, or intelligent entity sub-capabilities, transforming existing software capabilities that originally relied on manual operation into reusable, manageable, and governable capability assets. Attached Figure Description

[0037] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 A flowchart illustrating a multi-task hybrid execution method for an intelligent agent provided in an embodiment of this application;

[0039] Figure 2is a schematic diagram of the architecture of a multi-task hybrid execution method for an agent provided by the embodiments of the present application;

[0040] Figure 3 is a schematic diagram of the metadata structure of target system capability of a multi-task hybrid execution method for an agent provided by the embodiments of the present application;

[0041] Figure 4 is a schematic diagram of a specific implementation flow of a multi-task hybrid execution method for an agent provided by the embodiments of the present application;

[0042] Figure 5 is an interface-level execution flowchart of an interface-free industrial software of a multi-task hybrid execution method for an agent provided by the embodiments of the present application;

[0043] Figure 6 is a schematic diagram of a hybrid task chain combining structured calling and interface-level execution of a multi-task hybrid execution method for an agent provided by the embodiments of the present application;

[0044] Figure 7 is a structural schematic diagram of a computing device provided by the embodiments of the present application. DETAILED EMBODIMENTS

[0045] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present invention.

[0046] In the following embodiments of the present application, the expression "and / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can indicate: only A exists, both A and B exist, and only B exists, where A and B can be singular or plural. The character " generally indicates that the associated objects before and after are in an "or" relationship. The expression "at least one of the following" or similar expressions refers to any combination of these items, including any combination of a single item or multiple items. For example, at least one of a, b and c can represent: a, b, c, a-b, a-c, b-c, or a-b-c, where a, b and c can be single or multiple. The singular expressions "a", "an", "", "above-mentioned", "the" and "this" are intended to also include expressions such as "one or more" unless explicitly indicated otherwise by the context. Furthermore, unless stated otherwise, the ordinal words such as "first" and "second" mentioned in the embodiments of the present application are used to distinguish multiple objects and are not used to limit the order, timing, priority or importance of multiple objects.

[0047] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0048] Example 1

[0049] Embodiment 1 of this application provides a method for multi-task hybrid execution of an intelligent agent. This method is executed by a computing device, which can be understood as a computer or other similar device, but is not limited thereto in this application. The method flow is as follows: Figure 1 As shown, it includes:

[0050] Step 101: Receive the business task, parse the business task, and obtain the task context.

[0051] The task context includes at least one of the following: task objective, target system, operation object, input parameters, expected output, and execution constraints. The business task includes user natural language instructions, workflow tasks, business events, alarm information, timed tasks, or sub-tasks dispatched by other intelligent agents.

[0052] Step 102: Identify the target system based on the task context and query the capability metadata corresponding to the target system.

[0053] Among them, capability metadata is used to describe at least one of the target system's interface capabilities, interface operation capabilities, permission requirements, risk level, and operational constraints.

[0054] Specifically, capability metadata includes at least one of the following: target system identifier, target system type, interface type, executable actions, input parameter definition, output parameter definition, structured call address, interface operation entry point, permission level, risk level, operation whitelist, prohibited actions, manual confirmation policy, timeout policy, retry policy, rollback policy, audit policy, interface template, interface element description, status verification rules, and result extraction rules.

[0055] Step 103: Determine the execution capability of the target system based on the capability metadata, and determine the execution method of the business task based on the execution capability determination result.

[0056] The execution methods for business tasks include structured call paths, interface-level execution paths, and hybrid execution paths.

[0057] Specifically, based on capability metadata, the execution method selection node is set, and the availability of the structured interface of the target system, the availability of the interface operation, whether the task involves multiple target systems or multiple execution links, and the risk level of the task are judged. The execution path of the business task is determined based on the judgment results.

[0058] In one possible implementation, when the capability metadata indicates that the target system supports at least one of the target structured call interfaces in the target structured call interface list, and the structured call interface can meet the current task execution requirements, the subtask corresponding to the target system is determined to adopt a structured call path.

[0059] When the capability metadata indicates that the target system does not support the structured call interface, or the target system supports the structured call interface but the structured call interface cannot meet the current task execution requirements, and the target system supports graphical interface operation, the subtask corresponding to the target system is determined to adopt the interface-level execution path;

[0060] When a business task involves multiple target systems and different target systems correspond to different execution methods, a hybrid execution path is determined for the business task.

[0061] When a business task involves high-risk actions, trigger at least one of the following: manual confirmation, dual confirmation, disabling automatic execution, or generating only operation suggestions.

[0062] In one possible implementation, when the target system has structured calling capability and the structured calling capability can meet the requirements, the structured calling result is obtained based on the structured calling path.

[0063] Specifically, it includes at least one of the following: Application Programming Interface (API) call, Model Context Protocol (MCP) service call, Function Calling (FC) of large models, Command Line Interface (CLI) call, Web service call, database query, script call, file parsing tool call, message queue call, plugin call, or Software Development Kit (SDK) call; when the target system has structured calling capability and the structured calling capability can meet the task execution requirements, call parameters are generated based on the structured calling path, and the corresponding structured capability of the target system is called to obtain the structured calling result; calling the structured capability corresponding to the target system includes: matching callable capabilities according to the task objective; generating call parameters according to the task context; performing format validation and permission validation on the call parameters; calling the callable capability; receiving the return result, and performing format conversion, result validation, and archiving on the return result.

[0064] In one possible implementation, when the target system does not have structured calling capability or the structured calling capability does not meet the requirements, the interface-level execution result is obtained based on the interface-level execution path.

[0065] Specifically, when the target system lacks structured calling capabilities or its structured calling capabilities cannot meet the task execution requirements, the interface state of the target system is obtained based on the interface-level execution path. Interface elements are identified, an operation sequence is generated according to the task objective and interface state, interface actions are executed according to the operation sequence, and state verification is performed after the interface actions are executed to obtain the interface-level execution result. Among these methods, obtaining the interface state of the target system includes obtaining desktop software interface elements through the operating system control tree; obtaining web page elements through the browser document object model (DOM); identifying windows, buttons, input boxes, menus, tables, or pop-ups through interface screenshots; obtaining interface text through optical character recognition (OCR); identifying interface elements and their executable actions through a multimodal model; and obtaining the current interface state through window handles, coordinate regions, interface templates, or historical operation trajectories.

[0066] For example, generating an operation sequence based on task objectives and interface states includes: determining the interface operation steps to be completed based on the task objectives; determining the target interface elements corresponding to the interface operation steps based on the interface states; generating an operation description for each interface operation step, the operation description including at least one of the following: step number, action type, target element, input value, preconditions, expected state, timeout, number of retries, risk level, and whether manual confirmation is required; wherein, the action type includes at least one of the following: clicking, inputting, selecting, dragging, uploading files, importing data, executing menu commands, starting calculations, waiting for results, exporting files, saving files, closing pop-ups, and switching tabs.

[0067] Furthermore, after executing the interface action, a status check is performed, including at least one of the following: checking whether the window is opened correctly; checking whether the page redirects successfully; checking whether the target control status meets expectations; checking whether the input content is correct; checking whether the file is imported successfully; checking whether the progress bar starts or stops; checking whether the result area is refreshed; checking whether the exported file is generated; checking whether the pop-up window is closed; checking whether the log file generates the specified record; and checking whether the output result meets the preset format or business rules.

[0068] In one possible implementation, when a business task involves multiple target systems or multiple execution methods, a hybrid execution result is obtained based on a hybrid execution path.

[0069] Specifically, when a business task involves multiple target systems or multiple execution methods, the business task is broken down into multiple subtasks; the target system and execution method corresponding to each subtask are determined, and the subtasks are executed according to the corresponding execution method; one or more of the following subtasks—structured call subtasks, interface-level execution subtasks, knowledge retrieval subtasks, file parsing subtasks, and report generation subtasks—are arranged into a task chain according to their dependencies; after any subtask is completed, the task context is updated based on the output of that subtask, and the subsequent subtasks are driven to execute based on the updated task context.

[0070] Step 104: Verify the execution process of the structured call path, the interface-level execution path, and the hybrid execution path, and aggregate the execution results to obtain the task execution result.

[0071] Specifically, the execution process of structured call paths, interface-level execution paths, and hybrid execution paths is subject to permission verification, risk control, exception handling, and log auditing; the results of structured call paths and interface-level execution paths are aggregated to generate and output task execution results.

[0072] For example, this step may include: authenticating the user's identity; verifying the user's access permissions to the target system; performing whitelist verification of structured call capabilities; performing whitelist verification of interface operation actions; determining the risk level based on the action type, target system level, data sensitivity level, and whether it affects production operations; when the risk level meets the preset high-risk conditions, performing manual confirmation, dual confirmation, disabling automatic execution, or only generating operation suggestions; and recording one or more of the following: task input, call parameters, interface operation steps, status verification results, key interface screenshots, exported files, exception information, manual confirmation records, and final output results.

[0073] When the status verification fails, the structured call fails, or the UI-level execution is abnormal, exception handling is performed. Exception handling includes at least one of the following: automatic retry; reacquiring the UI status; re-identifying UI elements; regenerating the operation sequence; rolling back to the previous step; rolling back to the initial state; switching to an alternative execution path; pausing the task; requesting manual confirmation; transferring to manual takeover; terminating the task; and generating an exception report.

[0074] In one possible implementation, when a business task is successfully executed and the capability accumulation conditions are met, at least one of the following is extracted: task type, target system, input parameter template, execution path, operation steps, status verification rules, exception handling strategy, permission requirements, risk level, and output result format of the business task. Based on the extraction results, a reusable capability template is generated, and the reusable capability template is registered as a Skill, tool, plugin, workflow template, interface operation template, automated process, or intelligent entity sub-capability for subsequent similar tasks to call.

[0075] Example 2

[0076] Based on Embodiment 1, Embodiment 2 of this application provides a functional architecture for a multi-task hybrid execution method for an intelligent agent. The architecture is as follows: Figure 2 As shown in the diagram, the execution method selection module is positioned as a decision node after the capability metadata management module, and it routes tasks to structured call paths, interface-level execution paths, or hybrid execution paths based on the capability metadata. This functional architecture includes a task receiving module, a task parsing module, a target system identification module, a capability metadata management module, an execution method selection module, a structured call module, an interface awareness module, an interface element identification module, an operation planning module, an interface execution module, a status verification module, an exception handling module, a result aggregation module, a governance audit module, and a capability accumulation module.

[0077] The task receiving module is used to receive business tasks. Business tasks can come from user natural language commands, agent dialogue requests, workflow node triggers, external system events, alarm events, scheduled tasks, or sub-tasks dispatched by other agents. Tasks can be in the form of text, structured parameters, event objects, or workflow context.

[0078] The task parsing module is used to parse business tasks and obtain the task context. The task context includes at least the task objective, business scenario, target system, operation object, input parameters, expected output, task priority, risk level, execution constraints, whether manual confirmation is required, whether automatic execution is allowed, and whether interface-level operation is allowed. For example, if a user inputs "Import a fault waveform file, run simulation software for short-circuit analysis, and generate an analysis report," the task parsing module can obtain the following: the task objective is short-circuit simulation analysis; the target system includes fault waveform file service, professional simulation software, and report generation service; the input parameters include waveform file, equipment parameters, and analysis model; and the output is a simulation analysis report.

[0079] The target system identification module is used to identify the target system that needs to be invoked or operated based on the task context. Target systems may include scheduling automation systems, equipment ledger systems, defect management systems, reporting systems, simulation analysis software, professional calculation programs, alarm analysis systems, inspection systems, document management systems, desktop utility software, web business systems, and vendor-specific closed software.

[0080] The capability metadata management module is used to maintain the target system's capability metadata. Capability metadata can include the target system identifier, system name, system type, interface type, callable actions, supported execution methods, input parameter definitions, output result definitions, permission levels, risk levels, operation whitelist, prohibited actions, manual confirmation policy, timeout policy, retry policy, rollback policy, audit policy, interface template, and result validation rules. The target system capability metadata structure is as follows: Figure 3 As shown.

[0081] The execution method selection module serves as a decision-making and routing node in the hybrid execution process of the intelligent agent. Based on the task context and target system capability metadata, it assesses the availability of the target system's structured interfaces, the availability of user interface operations, whether the task involves multiple systems, and the task's risk level, and generates corresponding execution path selection results. The execution path selection results include one or more of the following: structured call path, user interface-level execution path, hybrid execution path, manually confirmed path, and path suggestion only. Specifically, when the capability metadata indicates that the target system possesses structured calling capabilities such as API, MCP, CLI, Web services, database interfaces, script tools, or file interfaces, and these structured calling capabilities can meet the current task execution requirements, the execution method selection module will route the task to a structured calling path. When the capability metadata indicates that the target system does not possess structured calling capabilities, or although it possesses structured calling capabilities, it cannot meet the current task execution requirements, but the target system allows operations to be completed through a graphical interface, the execution method selection module will route the task to a graphical interface execution path. When a business task involves multiple target systems, multiple sub-tasks, or multiple execution stages, and different sub-tasks correspond to different execution methods, the execution method selection module will route the task to a mixed execution path and trigger mixed task chain orchestration. When a task involves high-risk production operations, sensitive data access, or insufficient permissions, the execution method selection module will route the task to a manual confirmation path, generate only a suggested path, or reject the execution path. The execution path selection result output by this module is written into the task context and serves as the basis for the subsequent execution of the structured calling module, graphical interface execution module, governance audit module, and result aggregation module.

[0082] The structured call module is used to perform structured calls to systems with standard interfaces. Structured call methods can include API calls, MCP service calls, Function Calling calls, CLI command calls, Web service calls, database queries, script calls, file parsing tool calls, message queue calls, plugin calls, or vendor SDK calls. Before the call, the structured call module performs format validation and permission checks on the call parameters; after the call, it performs format conversion, result validation, and archiving of the returned results.

[0083] The interface perception module, interface element recognition module, operation planning module, interface execution module, and state verification module together constitute the interface-level execution closed loop. Interface perception methods can include operating system control tree recognition, browser DOM structure recognition, screenshot image recognition, OCR text recognition, multimodal model recognition, interface template matching, window handle recognition, coordinate region recognition, and historical operation trajectory matching. The interface element recognition module identifies buttons, menus, input boxes, file upload entries, query buttons, export buttons, calculation buttons, confirmation buttons, table rows and columns, pop-up close buttons, tabs, and drop-down options from the interface state. The operation planning module generates operation sequences such as clicking, inputting, selecting, importing data, executing menu commands, starting calculations, waiting for results, exporting files, saving files, closing pop-ups, and switching tabs. The interface execution module executes the operation sequences. The state verification module checks whether the window state, control state, input content, file import state, progress bar state, result area refresh state, exported file generation state, pop-up state, or log output state meet expectations after each operation.

[0084] The exception handling module is used to handle exceptions during structured calls or UI-level execution. Exception types can include insufficient permissions, API call failure, parameter validation failure, target system unreachable, missing UI elements, UI layout changes, pop-up blocking, operation timeout, file import failure, file export failure, calculation failure, empty return result, abnormal result format, unconfirmed risky actions, and user-interrupted execution. Exception handling methods can include automatic retry, re-identifying the UI, regenerating operation steps, reverting to the previous step, switching to an alternative path, pausing and waiting, requesting manual confirmation, transferring to manual takeover, terminating the task, and generating an exception report.

[0085] The governance audit module is used for unified governance of the entire execution process. Governance audit content includes user authentication, user permission verification, target system permission verification, tool call whitelist verification, interface operation whitelist verification, risk level assessment, manual confirmation, dual-person confirmation, sensitive data anonymization, call log recording, key screenshot documentation, operation trajectory recording, anomaly recording, execution result recording, and execution report generation. For high-risk operations involving remote control, remote adjustment, parameter download, configuration modification, and device start / stop, the system can disable automatic execution, generate only operation suggestions, request manual confirmation before execution, or request dual-person confirmation before execution.

[0086] The results aggregation module aggregates results from different execution paths. Result sources can include API responses, MCP service responses, CLI execution results, database query results, file parsing results, UI result area content, exported files, log files, screenshots, knowledge base retrieval results, and results returned by other agents. The results aggregation module performs unified formatting, deduplication, validation, correlation, and interpretation on these results to form the final task result.

[0087] The capability accumulation module is used to accumulate successfully executed task paths into reusable capabilities. Accumulation formats can include Skills, workflow templates, operation templates, UI automation processes, intelligent entity sub-capabilities, reusable task chains, or scenario-based execution templates. Accumulated content includes task type, target system, input parameters, execution path, operation steps, status verification rules, exception handling rules, permission requirements, risk level, and output result format.

[0088] Example 3

[0089] Embodiment 3 of this application provides a specific implementation of a multi-task hybrid execution method for an intelligent agent, based on Embodiments 1 and 2. For example... Figure 4 As shown, the intelligent agent hybrid execution method for existing industrial software in this embodiment includes the following steps:

[0090] Step 401: Receive task input for business tasks. The system receives business tasks from user natural language, workflows, external system events, alarm events, or other intelligent agents.

[0091] Step 402: Parse the task context. The system parses the business task to obtain the task objective, target system, operation object, input parameters, output requirements, business scenario, execution constraints, and risk level.

[0092] Step 403: Identify the target system. The system identifies the target system that needs to be invoked or operated based on the task context. If the task involves multiple systems, a set of target systems is generated.

[0093] Step 404: Query capability metadata. Query the capability metadata corresponding to each target system.

[0094] Step 405: Execution Method Selection and Path Diversion. The system sets an execution method selection node based on the target system's capability metadata, assessing the target system's interface capabilities, user interface operation capabilities, task complexity, and risk level. If the target system possesses structured call capabilities such as API, MCP, CLI, Web services, database interfaces, script tools, or file interfaces, and these capabilities can satisfy the current task, then the system enters the structured call path. If the target system lacks structured call capabilities, or its structured call capabilities cannot satisfy the current task, but the target system supports graphical interface operation, then the system enters the interface-level execution path. If the business task involves multiple target systems or multiple sub-tasks, and different sub-tasks correspond to different execution methods such as structured calls, interface-level execution, knowledge retrieval, file parsing, and report generation, then the system enters the mixed task chain orchestration. If the task involves high-risk operations, insufficient permissions, unauthorized calls, or prohibited automatic execution actions, then the system enters the manual confirmation, suggestion generation, or execution rejection process.

[0095] Step 406: Execute structured calls. For subtasks that select a structured call path, the system matches the target capability, generates call parameters, verifies the parameter format and user permissions, calls the target interface or tool, receives the return result, verifies the return result, and writes the result to the task context.

[0096] Step 407: Perform UI-level execution. For subtasks that select a UI-level execution path, the system starts or connects to the target software, obtains the current UI state, identifies operable UI elements, generates an operation sequence based on the task objective, performs whitelisting and permission checks on the operation sequence, executes the UI operations step by step, performs status checks after each operation, identifies and handles exceptions, extracts UI results or exports files, and writes the results to the task context.

[0097] Step 408, mixed execution, is carried out through structured calls and interface agent execution.

[0098] Step 409: Unified governance, verification, and auditing. This involves at least one of the following: permission verification, whitelisting, risk assessment, manual confirmation, log recording, and anomaly rollback.

[0099] Step 410: Perform status verification and exception handling. For each UI action or structured call result, the system verifies the status based on window state, control state, file state, log state, interface return status, and business rules. When verification fails, the system performs retry, rollback, replanning, alternative path switching, or manual intervention.

[0100] Step 411, Result Aggregation and Output. The system aggregates results from different sources, including structured data, interface extraction results, file parsing results, log results, and report generation results. It then performs consistency checks, formatting, and natural language interpretation on the results, finally outputting the task results.

[0101] Step 412, Execution path capability consolidation. For successfully executed task chains with reusable value, the system consolidates them as Skills, intelligent agent tools, workflow templates, interface execution templates, or dedicated intelligent agent capabilities for reuse in subsequent similar tasks.

[0102] Example 4

[0103] Based on Embodiment 3, Embodiment 4 of this application provides a non-interface industrial software interface-level execution flow for a multi-task hybrid execution method of an intelligent agent; such as... Figure 5 As shown, the interface-free industrial software interface-level execution flow in this embodiment includes the following steps:

[0104] Step 501: Launch or connect to the target software. The system launches the target desktop software, connects to the web page, enters the remote desktop, or activates the target window based on the target system's capability metadata.

[0105] Step 502: Obtain the current interface state. The system obtains the current interface state through the control tree, DOM, screenshot, OCR, multimodal model, window handle, coordinate region, or interface template.

[0106] Step 503: Identify UI elements. The system identifies UI elements such as buttons, menus, input boxes, drop-down lists, tables, pop-ups, file selection windows, result areas, status bars, and progress bars, and generates UI element description objects.

[0107] Step 504: Generate an operation sequence. The system generates an operation sequence based on the task objective, current interface state, interface element descriptions, operation whitelist, and risk level. Each step in the operation sequence may include a step number, action type, target element, input value, preconditions, desired state, timeout, number of retries, risk level, and manual confirmation requirements.

[0108] Step 505: Perform operation whitelist and permission verification. The system determines whether the current operation belongs to the set of operations that are allowed to be executed automatically, and verifies the permissions of the user, agent, and target system. For high-risk operations, the system either initiates manual confirmation or generates only a suggestion.

[0109] Step 506: Execute interface actions. The system executes actions such as clicking, inputting, selecting, dragging, uploading files, importing data, executing menu commands, starting calculations, waiting for results, exporting files, saving files, closing pop-ups, or switching tabs according to the operation sequence.

[0110] Step 507: Obtain the interface state after the operation. After each action is performed, the system re-obtains the interface state or the status of related files, logs, and interfaces.

[0111] Step 508: Determine if the status verification passes. The system determines whether this step is complete based on preset status verification rules. For example, after clicking "Import File," it verifies whether the filename is displayed on the interface; after clicking "Start Calculation," it verifies whether the progress bar starts; and after clicking "Export Report," it verifies whether a file is generated in the target directory.

[0112] Step 509: Extract execution results and record logs. Once the status verification is passed and all operation steps are completed, the system extracts the execution results from the interface, exported files, log files, or result directory, and records key screenshots, operation steps, input parameters, output files, exception information, and manual confirmation records.

[0113] When the status verification fails, the system can retry, roll back to the previous step, re-identify the interface, regenerate the operation sequence, switch to an alternative path, or transfer the task to manual intervention.

[0114] Example 5

[0115] Embodiment 5 of this application provides a scenario for the automatic execution of a professional simulation software using a multi-task hybrid execution method for intelligent agents. In a power fault analysis scenario, operational personnel need to import fault waveform files into the professional simulation software, select an analysis model, execute a short-circuit simulation, and export the analysis results. This professional simulation software is desktop software and does not provide API or CLI interfaces; tasks can only be completed through a manual interface.

[0116] The user inputs: "Please import the P214 line fault waveform file, run the short-circuit simulation analysis, and generate an analysis report." The system parses the data and finds that the task objective is short-circuit simulation analysis, the operation object is the P214 line fault waveform file, the target system includes waveform file service, professional simulation software, and report generation service, the output is an analysis report, the risk level is analysis-type task, and it is not a direct control operation.

[0117] After querying the target system's capability metadata, the system found that: the waveform recording file service supports API, the professional simulation software only supports GUI, and the report generation service supports MCP. The system generates a hybrid execution path: API retrieves the waveform recording file, GUI operates the simulation software, the file is parsed to obtain the simulation results, and MCP calls the report generation service. This hybrid task chain is illustrated below. Figure 6 As shown.

[0118] The system first obtains the fault recording file of line P214 via API and saves it to a specified directory. Then, it starts the simulation software, senses the interface status, and identifies interface elements such as "File Import," "Model Selection," "Start Calculation," and "Export Results." Subsequently, it generates an operation sequence and automatically executes steps such as importing the file, selecting the model, running the calculation, and exporting the results. After each operation, a status check is performed, such as verifying whether the file was successfully imported, whether the calculation progress has started, and whether the result file has been generated. If the result file does not exist, the export step is re-executed or manual intervention is initiated.

[0119] Finally, the system invokes the report generation service to generate a report from the simulation results, equipment information, fault parameters, and analysis conclusions. It also records the user task, API call parameters, interface operation steps, key screenshots, result files, report files, exception information, and execution time. This embodiment achieves low-intrusion invocation of interface-free simulation software, enabling the intelligent agent to automatically complete software operation processes that were originally dependent on manual intervention. Furthermore, it ensures execution reliability and traceability through status verification and auditing.

[0120] Example 6

[0121] Embodiment 6 of this application provides a hybrid query scenario for a ledger system and a reporting system in a multi-task hybrid execution method for intelligent agents. In an equipment operation and maintenance scenario, a user needs to query the defect records of a certain device over the past year and generate statistical reports. The basic equipment ledger system provides an API, but the reporting system is a historical web system that does not provide a standard interface and can only be filtered and exported through the page.

[0122] After receiving a user's query task, the system parses the device name, time range, and report type; queries basic device information via API; determines that the report system only supports web page operations; initiates interface-level execution, automatically enters the report page, allows the user to fill in the device name and time range, clicks query, verifies whether the table is refreshed, clicks export and verifies whether an Excel or PDF file is generated; then parses the exported file, summarizes the device ledger and defect records, and generates statistical results and natural language descriptions.

[0123] This embodiment enables a hybrid execution of API queries and UI exports, avoiding interface modifications to the historical reporting system and reducing system integration costs.

[0124] Example 7

[0125] Embodiment 7 of this application provides a controlled execution scenario for an operation and maintenance inspection task using a multi-task hybrid execution method for intelligent agents. In this scenario, on-duty personnel require the intelligent agent to perform an inspection of a certain business system, checking the system's operating status, service status, log anomalies, and report generation. Some inspection items can be obtained through an interface, while others require access to the management interface.

[0126] After receiving an inspection task, the system identifies the task as a read-only inspection task; queries the service status via API; parses the log file via file tools; and views the task through interface-level execution for tasks without an interface management interface. Actions involving restarting services or modifying configurations are identified as high-risk and are not executed automatically by default; instead, suggestions are generated or manual confirmation is requested. Finally, the inspection results are summarized and an inspection report is output.

[0127] This embodiment enables automated and controlled execution of operation and maintenance inspection tasks, distinguishes between read-only queries and high-risk operations, and reduces the risk of misoperation.

[0128] Example 8

[0129] Embodiment 8 of this application provides another implementation of the multi-task hybrid execution method for intelligent agents. In other implementations, the structured call path in this invention is not limited to API, MCP, or CLI, but may also include REST API, GraphQL, SOAP WebService, RPC, gRPC, database SQL queries, message queues, file interfaces, FTP / SFTP, command-line tools, local scripts, vendor SDKs, plugin interfaces, and microservice interfaces, etc.

[0130] The interface-level execution in this invention is not limited to desktop software, but can also be applied to web page systems, B / S architecture business systems, C / S architecture desktop software, remote desktop software, virtual machine software, industrial touch screen interfaces, mobile apps, browser plugin interfaces, and operation and maintenance management consoles.

[0131] The interface perception method in this invention can employ control tree recognition, DOM structure recognition, screenshot image recognition, OCR text recognition, multimodal model recognition, coordinate template matching, historical operation template matching, log or process status assisted recognition, etc., or they can be used in combination.

[0132] The operation planning method in this invention can be implemented through preset rules, historical templates, workflow configuration, large model reasoning, multimodal model planning, state machines, behavior trees, task planning algorithms, manual configuration, or a combination of multiple methods.

[0133] The status verification method in this invention can be implemented based on interface changes, control status, file generation, log output, database status changes, interface return status, result field integrity, screenshot similarity, or manual confirmation.

[0134] The unified governance mechanism in this invention can be implemented through user role permissions, target system permissions, tool permissions, data permissions, operation permissions, scenario permissions, time window permissions, multi-person approval, electronic signatures, or temporary authorization.

[0135] The capability accumulation objects in this invention can be represented as Skills, Tools, workflow templates, interface operation templates, RPA processes, intelligent agent plugins, intelligent agent subtasks, service adapters, scenario execution templates, or business process assets.

[0136] Example 9

[0137] Embodiment 9 of this application provides a multi-task hybrid execution system for an intelligent agent, used to execute the content of the multi-task hybrid execution method for an intelligent agent provided in this application. The system includes: a receiving module, a query module, a judgment module, and a verification module.

[0138] The system comprises the following modules: a receiving module, which receives business tasks, parses them, and obtains the task context; the task context includes at least one of the following: task objective, target system, operation object, input parameters, expected output, and execution constraints; a query module, which identifies the target system based on the task context and queries the corresponding capability metadata of the target system; the capability metadata describes at least one of the following: interface capabilities, interface operation capabilities, permission requirements, risk level, and operation constraints of the target system; and a judgment module, which judges the execution capability of the target system based on the capability metadata and determines the execution method of the business task based on the judgment result; the execution method of the business task includes a structured call path. The system comprises three execution modules: a UI-level execution path and a hybrid execution path; a judgment module, which is also used to obtain a structured call result based on the structured call path when the target system has structured call capability and the structured call capability meets the requirements; a judgment module, which is also used to obtain a UI-level execution result based on the UI-level execution path when the target system does not have structured call capability or the structured call capability does not meet the requirements; a judgment module, which is also used to obtain a hybrid execution result based on the hybrid execution path when the business task involves multiple target systems or multiple execution methods; and a verification module, which is used to verify the execution process of the structured call path, the UI-level execution path, and the hybrid execution path, and to aggregate the execution results to obtain the task execution result.

[0139] Having introduced the multi-task hybrid execution system of the intelligent agent in the exemplary embodiments of this application, we will now introduce a computing device in another exemplary embodiment of this application.

[0140] Example 10

[0141] Those skilled in the art will understand that various aspects of this application can be implemented as a system, method, or program product. Therefore, various aspects of this application can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, collectively referred to herein as a "circuit," "module," or "system."

[0142] In some possible implementations, the computing device according to this application may include at least one processor and at least one memory. The memory stores a computer program that, when executed by the processor, causes the processor to perform the steps of the cross-device assisted interaction method for off-site clearing according to various exemplary embodiments of this application described above.

[0143] The following reference Figure 7 To describe a computing device 130 according to this embodiment of the present application. Figure 7 The computing device 130 shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this application. Figure 7 As shown, the computing device 130 is presented in the form of a general-purpose smart terminal (or Bluetooth headset). The components of the computing device 130 may include, but are not limited to: at least one processor 131, at least one memory 132, and a bus 133 connecting different system components (including memory 132 and processor 131).

[0144] Bus 133 represents one or more of several bus architectures, including a memory bus or memory controller, peripheral bus, processor, or local bus using any of the various bus architectures. Memory 132 may include readable media in the form of volatile memory, such as random access memory (RAM) 1321 and / or cache memory 1322, and may further include read-only memory (ROM) 1323. Memory 132 may also include a program / utility 1325 having a set (at least one) of program modules 1324, including but not limited to: an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment.

[0145] The computing device 130 can also communicate with one or more external devices 134 (e.g., keyboard, pointing device, etc.), and / or with any device that enables the computing device 130 to communicate with one or more other smart terminals (e.g., router, modem, etc.). This communication can be performed via input / output (I / O) interface 135. Furthermore, the computing device 130 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 136. As shown, network adapter 136 communicates with other modules used in the computing device 130 via bus 133. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with the computing device 130, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0146] In some possible implementations, various aspects of the multitasking hybrid execution method of the intelligent agent provided in this application can also be implemented in the form of a program product, which includes a computer program. When the program product is run on a computer device, the computer program is used to cause the computer device to perform the steps in the multitasking hybrid execution method of the intelligent agent according to the various exemplary embodiments of this application described above.

[0147] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0148] The program product for multitasking execution of intelligent agents according to the embodiments of this application may employ a portable compact disc read-only memory (CD-ROM) and include a computer program, and may run on a smart terminal. However, the program product of this application is not limited to this. In this document, the readable storage medium may be any tangible medium containing or storing a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.

[0149] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a readable computer program. Such propagated data signals may take various forms, including—but not limited to—electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0150] It should be noted that although several units or sub-units of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0151] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0152] This application is applicable not only to scenarios such as power dispatching, equipment operation and maintenance, fault analysis, and report generation, but also to industrial fields with a large amount of existing software and closed tools, such as oil and petrochemicals, rail transportation, water conservancy and hydropower, manufacturing, aerospace, industrial control, and energy management. This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable access frequency prediction device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable access frequency prediction device, generate instructions for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0153] These computer program instructions may also be stored in a computer-readable storage medium capable of directing a computer or other programmable access predictive device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0154] These computer program instructions can also be loaded onto a computer or other programmable access predictive device to cause a series of operational steps to be performed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable device for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0155] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0156] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A method for multi-task hybrid execution of an intelligent agent, characterized in that, include: The system receives a business task, parses the business task, and obtains a task context; wherein the task context includes at least one of the following: task objective, target system, operation object, input parameters, expected output, and execution constraints. Identify the target system based on the task context, and query the capability metadata corresponding to the target system; wherein, the capability metadata is used to describe at least one of the target system's interface capabilities, interface operation capabilities, permission requirements, risk level, and operation constraints; The execution capability of the target system is determined based on the capability metadata, and the execution method of the business task is determined according to the execution capability determination result; wherein, the execution method of the business task includes structured call path, interface-level execution path and hybrid execution path; When the target system has structured calling capability and the structured calling capability can meet the requirements, the structured calling result is obtained based on the structured calling path; When the target system does not have structured calling capability or the structured calling capability does not meet the requirements, the interface-level execution result is obtained based on the interface-level execution path; When the business task involves multiple target systems or multiple execution methods, a hybrid execution result is obtained based on the hybrid execution path; The execution processes of the structured call path, the interface-level execution path, and the hybrid execution path are verified, and the execution results are aggregated to obtain the task execution result.

2. The method according to claim 1, characterized in that, The capability metadata includes at least one of the following: target system identifier, target system type, interface type, executable action, input parameter definition, output parameter definition, structured call address, interface operation entry point, permission level, risk level, operation whitelist, prohibited actions, manual confirmation policy, timeout policy, retry policy, rollback policy, audit policy, interface template, interface element description, status verification rules, and result extraction rules.

3. The method according to claim 1, characterized in that, The step of determining the execution capability of the target system based on the capability metadata, and determining the execution method of the business task based on the execution capability determination result, includes: When the capability metadata indicates that the target system supports at least one of the target structured call interfaces in the target structured call interface list, and the structured call interface can meet the current task execution requirements, it is determined that the subtask corresponding to the target system adopts a structured call path. When the capability metadata indicates that the target system does not support the structured call interface, or the target system supports the structured call interface but the structured call interface cannot meet the current task execution requirements, and the target system supports graphical interface operation, the subtask corresponding to the target system is determined to adopt the interface-level execution path; When the business task involves multiple target systems and different target systems correspond to different execution methods, it is determined that the business task adopts a hybrid execution path; When the business task involves high-risk actions, at least one of the following can be triggered: manual confirmation, dual confirmation, disabling automatic execution, or generating only operation suggestions.

4. The method according to claim 1, characterized in that, When the business task involves multiple target systems or multiple execution methods, the hybrid execution result is obtained based on the hybrid execution path, including: When the business task involves multiple target systems or multiple execution methods, the business task is broken down into multiple sub-tasks; Determine the target system and execution method for each subtask, and execute the subtasks according to the corresponding execution methods. Arrange one or more of the structured call subtasks, interface-level execution subtasks, knowledge retrieval subtasks, file parsing subtasks, and report generation subtasks into a task chain according to their dependencies; After any subtask is completed, the task context is updated based on the output of that subtask, and subsequent subtasks are executed based on the updated task context.

5. The method according to claim 1, characterized in that, The structured invocation capability includes: At least one of the following: application programming interface (API) call, model context protocol service call, large model function call, command line interface command call, network service call, database query, script call, file parsing tool call, message queue call, plugin call, or vendor software development kit call; When the target system has structured calling capability and the structured calling capability can meet the task execution requirements, calling parameters are generated based on the structured calling path, and the corresponding structured capability of the target system is called to obtain the structured calling result; The process of invoking the structured capabilities corresponding to the target system includes: matching callable capabilities according to the task objective; generating invoking parameters according to the task context; performing format and permission verification on the invoking parameters; invoking the callable capabilities; receiving the returned results; and performing format conversion, result verification, and archiving on the returned results.

6. The method according to claim 1, characterized in that, When the target system lacks structured calling capabilities or the structured calling capabilities do not meet the requirements, obtaining the interface-level execution result based on the interface-level execution path includes: When the target system does not have structured calling capability or the structured calling capability cannot meet the task execution requirements, the interface state of the target system is obtained based on the interface-level execution path, the interface elements are identified, an operation sequence is generated according to the task objective and the interface state, the interface action is executed according to the operation sequence, and the state is verified after the interface action is executed to obtain the interface-level execution result. The step of obtaining the interface state of the target system includes at least one of the following: obtaining desktop software interface elements through the operating system control tree; obtaining web page elements through the browser document object model structure; identifying windows, buttons, input boxes, menus, tables, or pop-ups through interface screenshots; identifying interface text through optical character recognition; identifying interface elements and their executable actions through a multimodal model; and obtaining the current interface state through window handles, coordinate regions, interface templates, or historical operation trajectories.

7. The method according to claim 6, characterized in that, Generate an operation sequence based on the task objective and the interface state, including: Based on the task objective, determine the interface operation steps to be completed, and based on the interface state, determine the target interface element corresponding to the interface operation steps. An operation description is generated for each interface operation step. The operation description includes at least one of the following: step number, action type, target element, input value, preconditions, expected state, timeout, number of retries, risk level, and whether manual confirmation is required. The action type includes at least one of the following: click, input, select, drag and drop, upload file, import data, execute menu command, start calculation, wait for result, export file, save file, close pop-up window, and switch tabs.

8. The method according to claim 1, characterized in that, The method further includes: When the business task is successfully executed and the capability accumulation conditions are met, at least one of the following is extracted: task type, target system, input parameter template, execution path, operation steps, status verification rules, exception handling strategy, permission requirements, risk level, and output result format of the business task. Based on the extraction results, reusable capability templates are generated and registered as Skills, tools, plugins, workflow templates, interface operation templates, automated processes, or intelligent entity sub-capabilities for subsequent similar tasks to call.

9. A multi-task hybrid execution system for an intelligent agent, characterized in that, include: A receiving module is used to receive a business task, parse the business task, and obtain a task context; wherein, the task context includes at least one of a task objective, a target system, an operation object, input parameters, expected output, and execution constraints; The query module is used to identify the target system based on the task context and query the capability metadata corresponding to the target system; wherein, the capability metadata is used to describe at least one of the target system's interface capabilities, interface operation capabilities, permission requirements, risk level, and operation constraints; The judgment module is used to judge the execution capability of the target system based on the capability metadata, and determine the execution method of the business task according to the execution capability judgment result; wherein, the execution method of the business task includes structured call path, interface-level execution path and hybrid execution path; The judgment module is further configured to obtain the structured call result based on the structured call path when the target system has structured call capability and the structured call capability can meet the requirements. The judgment module is also used to obtain the interface-level execution result based on the interface-level execution path when the target system does not have structured calling capability or the structured calling capability does not meet the requirements. The judgment module is also used to obtain a mixed execution result based on the mixed execution path when the business task involves multiple target systems or multiple execution methods; The verification module is used to verify the execution process of the structured call path, the interface-level execution path and the hybrid execution path, and to aggregate the execution results to obtain the task execution result.

10. A computing device, characterized in that, Its features include: Memory, used to store program instructions; A processor is configured to invoke program instructions stored in the memory and execute the method as described in any one of claims 1-8 according to the obtained program instructions.