Task execution method and system

By identifying AI task characteristics and obtaining resource profiles, determining the matching degree, and allocating resources to the target computing domain, the problem of balancing cost and performance in AI tasks is solved, achieving better resource utilization.

CN121858232APending Publication Date: 2026-04-14SHENZHEN COMTOP INFORMATION TECH
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN COMTOP INFORMATION TECH
Filing Date
2025-12-29
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Current technologies cannot achieve an optimal balance between cost and performance when performing AI tasks.

Method used

By identifying the task characteristics of AI tasks, resource profiles of general computing domains and intelligent computing domains in a hybrid cloud environment are obtained, the matching degree between task characteristics and resource profiles is determined, and AI tasks are assigned to target computing domains for execution based on the matching degree relationship.

Benefits of technology

While ensuring the performance of AI tasks, it significantly optimizes the overall cost of computing resources, achieving a better balance between cost and performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121858232A_ABST
    Figure CN121858232A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a task execution method and system. The method comprises the following steps: in response to a to-be-executed artificial intelligence task, identifying task features of the artificial intelligence task, and obtaining a first resource portrait of a general computing domain and a second resource portrait of an intelligent computing domain in a hybrid cloud environment; determining a first matching degree between the task feature and the first resource portrait, and determining a second matching degree between the task feature and the second resource portrait; and according to a numerical relationship between the first matching degree and the second matching degree, determining a target computing domain from the general computing domain and the intelligent computing domain, and allocating the artificial intelligence task to the target computing domain, so that the target computing domain is utilized to execute the artificial intelligence task. According to the technical scheme provided by the embodiment of the invention, the optimal balance between the cost and the performance can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of artificial intelligence technology, and in particular to a task execution method and system. Background Technology

[0002] With the widespread application of artificial intelligence (AI) technology, AI agents have become an important vehicle for performing complex AI tasks. To meet the diverse computing needs of AI tasks, enterprises typically execute AI tasks in a hybrid cloud environment.

[0003] In the process of realizing this invention, the inventors discovered the following technical problems in the prior art: Currently, when performing AI tasks, it is impossible to achieve an optimal balance between cost and performance, which urgently needs to be solved. Summary of the Invention

[0004] This invention provides a task execution method and system that solves the problem of not being able to achieve an optimal balance between cost and performance.

[0005] According to one aspect of the present invention, a task execution method is provided, which may include:

[0006] In response to an AI task to be executed, the task characteristics of the AI ​​task are identified, and a first resource profile of the general computing domain and a second resource profile of the intelligent computing domain are obtained in a hybrid cloud environment.

[0007] Determine the first degree of matching between task features and the first resource profile, and determine the second degree of matching between task features and the second resource profile;

[0008] Based on the numerical relationship between the first and second matching degrees, the target computing domain is determined from the general computing domain and the intelligent computing domain, and the artificial intelligence tasks are assigned to the target computing domain so as to utilize the target computing domain to execute the artificial intelligence tasks.

[0009] According to another aspect of the present invention, a task execution system is provided, which may include:

[0010] The resource profile acquisition module is used to respond to the artificial intelligence task to be executed, identify the task characteristics of the artificial intelligence task, and acquire the first resource profile of the general computing domain and the second resource profile of the intelligent computing domain in the hybrid cloud environment.

[0011] The matching degree determination module is used to determine the first matching degree between the task features and the first resource profile, and to determine the second matching degree between the task features and the second resource profile;

[0012] The task execution module is used to determine the target computing domain from the general computing domain and the intelligent computing domain based on the numerical relationship between the first matching degree and the second matching degree, and to allocate the artificial intelligence task to the target computing domain so as to execute the artificial intelligence task using the target computing domain.

[0013] According to another aspect of the present invention, an electronic device is provided, which may include:

[0014] At least one processor; and

[0015] A memory that is communicatively connected to at least one processor; wherein,

[0016] The memory stores a computer program that can be executed by at least one processor, such that when the at least one processor executes the program, it implements the task execution method provided in any embodiment of the present invention.

[0017] According to another aspect of the present invention, a computer-readable storage medium is provided having computer instructions stored thereon for causing a processor to execute the task execution method provided in any embodiment of the present invention.

[0018] According to another aspect of the present invention, a computer program product is provided, on which a computer program is stored, which, when executed by a processor, implements the task execution method provided in any embodiment of the present invention.

[0019] The technical solution of this invention, in response to an artificial intelligence task to be executed, identifies the task characteristics of the artificial intelligence task and obtains a first resource profile of a general computing domain and a second resource profile of an intelligent computing domain in a hybrid cloud environment; determines a first matching degree between the task characteristics and the first resource profile, and a second matching degree between the task characteristics and the second resource profile; based on the numerical relationship between the first and second matching degrees, determines a target computing domain from the general computing domain and the intelligent computing domain, and allocates the artificial intelligence task to the target computing domain, which can then be used to execute the artificial intelligence task. This technical solution, by first identifying task characteristics and obtaining resource profiles of different computing domains, and then obtaining matching degrees based on these profiles, allows for precise matching of the artificial intelligence task to a more suitable computing domain. This significantly optimizes the overall cost of computing resource usage while ensuring the execution performance of the artificial intelligence task, thus achieving a better balance between cost and performance.

[0020] It should be understood that the description in this section is not intended to identify key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart of a task execution method provided according to an embodiment of the present invention;

[0023] Figure 2 This is a flowchart of another task execution method provided by an embodiment of the present invention;

[0024] Figure 3 This is a flowchart of another task execution method provided by an embodiment of the present invention;

[0025] Figure 4 This is a structural block diagram of a task execution system provided according to an embodiment of the present invention;

[0026] Figure 5 This is a schematic diagram of the structure of an electronic device that implements the task execution method of the present invention. Detailed Implementation

[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0028] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. The same applies to "target," "original," etc., and will not be repeated here. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0029] It should be noted that the collection, gathering, updating, analysis, processing, use, transmission, and storage of user personal information involved in the technical solution of this invention all comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. Necessary measures are taken to prevent unauthorized access to user personal information data and to maintain user personal information security and network security.

[0030] Figure 1 This is a flowchart of a task execution method provided by an embodiment of the present invention. This embodiment is applicable to the execution of AI tasks, especially in the context of hybrid cloud environments. The method can be executed by the task execution system provided by this embodiment, which can be implemented in software and / or hardware. This system can be deployed on an AI Agent, which can be integrated into an electronic device, such as various user terminals or servers.

[0031] See Figure 1 The method of this invention specifically includes the following steps:

[0032] S110. In response to the artificial intelligence task to be executed, identify the task characteristics of the artificial intelligence task, and obtain a first resource profile of the general computing domain and a second resource profile of the intelligent computing domain in a hybrid cloud environment.

[0033] Artificial intelligence tasks (i.e. AI tasks) can be understood as tasks driven by AI agents and executed in a hybrid cloud environment. Optionally, the task can be an image recognition task, an alarm classification task, or a log generation task, which depends on the actual situation and is not specifically limited here. The number of tasks can be one or more, and each task has its own requirements for computing domain (or computing resources).

[0034] For each AI task, task characteristics are identified. These characteristics can be used to describe the inherent attributes and key parameters required by the AI ​​task. Optionally, these task characteristics can be at least one of the following: task type (e.g., inference, generation, or logic tasks), resource requirements (e.g., whether a Graphics Processing Unit (GPU) / Neural Processing Unit (NPU) acceleration is needed), task migration cost (the cost of migrating the AI ​​task between the general computing domain and the intelligent computing domain), model size (the size of the model used to execute the AI ​​task, such as the number of parameters and input dimensions), and predicted execution time (estimated based on historical data and model size). These can be set according to actual needs and are not specifically limited here. In this embodiment of the invention, optionally, the above-mentioned task characteristics can be obtained by analyzing the task metadata of the AI ​​task, which can be obtained from the AI ​​Agent description file or API request headers.

[0035] A hybrid cloud environment can be understood as a unified computing environment based on private cloud, public cloud, and edge cloud. In this embodiment of the invention, it can be further understood as a heterogeneous computing resource pool composed of at least one general-purpose computing domain (hereinafter referred to as the general computing domain) and at least one intelligent computing domain (hereinafter referred to as the intelligent computing domain). The general computing domain is a general-purpose computing resource based on a central processing unit (CPU) architecture, with the CPU as its core hardware, providing flexible and versatile computing resources with relatively limited computing power. The intelligent computing domain, on the other hand, consists of AI-accelerated computing resources represented by GPUs and NPUs, with GPUs and NPUs as their core hardware, offering high-performance AI computing capabilities but at a higher cost.

[0036] Building upon this, the first resource profile can be considered a quantitative description of the current resource status of the general computing domain, and similarly, the second resource profile can be considered a quantitative description of the current resource status of the intelligent computing domain. In this embodiment of the invention, optionally, the types of the first and second resource profiles can be the same or different. For example, the first resource profile may include at least one of CPU core count, memory capacity, network bandwidth, and average response latency, while the second resource profile may include at least one of GPU / NPU count, video memory size, model loading rate, and device temperature. Both the first and second resource profiles may also include the current load status (such as utilization rate, task queue length, and average response time), which can be set according to actual needs and is not specifically limited here. In this embodiment of the invention, optionally, the resource profiles of the two computing domains can be collected in real time by calling the monitoring interface of the hybrid cloud management platform, thereby obtaining real-time resource profiles.

[0037] In this step, by actively identifying task characteristics and acquiring resource profiles in real time, we can accurately understand the task requirements of AI tasks and the real-time resource status of the two types of computing domains in a hybrid cloud environment. This provides a precise and dynamic data foundation for subsequent matching decisions, avoiding blind allocation of AI tasks due to unclear task requirements or outdated resource status, and ensuring the objectivity and timeliness of the decision-making basis.

[0038] S120. Determine the first degree of matching between the task features and the first resource profile, and determine the second degree of matching between the task features and the second resource profile.

[0039] The first matching degree can be considered a quantitative indicator used to measure the extent to which the resource capabilities of the general computing domain (described by the first resource profile) meet the requirements of the current AI task (described by task features). Similarly, the second matching degree can be considered a quantitative indicator used to measure the extent to which the resource capabilities of the intelligent computing domain (described by the second resource profile) meet the requirements of the current AI task (described by task features).

[0040] Two matching degrees are determined. Here's an example: a matching degree evaluation model is established. This model maps and weights various indicators of task characteristics and resource profile indicators to obtain the corresponding matching degree. For example, for a low-precision, high-throughput inference task, its second matching degree with the intelligent computing domain, which excels in low-precision, high-performance computing, might be as high as 0.9 (out of 1.0), while its first matching degree with the general computing domain, which is less efficient at handling such tasks, might only be 0.3.

[0041] In this step, by calculating two independent matching degrees, the suitability of each computing domain for performing the current AI task can be quantitatively evaluated. This allows for a clear and objective comparison of the advantages and disadvantages of the two computing domains in meeting specific task requirements, providing key and comparable quantitative evidence for the final cost and performance trade-off decision, thereby achieving the optimal balance between cost and performance.

[0042] S130. Based on the numerical relationship between the first matching degree and the second matching degree, determine the target computing domain from the general computing domain and the intelligent computing domain, and assign the artificial intelligence task to the target computing domain so as to use the target computing domain to execute the artificial intelligence task.

[0043] The target computing domain can be understood as the computing domain (i.e., general computing domain or intelligent computing domain) that is ultimately selected to perform AI tasks based on the numerical relationship between two matching degrees and the preset decision-making strategy.

[0044] In this embodiment of the invention, optionally, the computing domain with a higher matching degree can be used as the target computing domain; alternatively, a matching degree difference threshold can be preset. When the matching degree advantage of the intelligent computing domain is not obvious (i.e., the second matching degree - the first matching degree ≤ the matching degree difference threshold) but its cost is significantly higher, the general computing domain can be selected first to save costs; otherwise, the intelligent computing domain can be selected first to ensure performance; and so on. This can be set according to actual needs and is not specifically limited here. In addition, for AI tasks that are called multiple times consecutively, computing domains can be bound based on historical performance data, thereby reducing cross-domain migration. The AI ​​task is allocated (i.e., scheduled) to the target computing domain to start the execution of the AI ​​task in the target computing domain.

[0045] In this step, decisions are made based on the numerical relationship of quantified matching degree, so that AI task allocation is no longer a one-size-fits-all approach, but rather an intelligent trade-off between performance and cost, thereby achieving dynamic and adaptive optimization. For example, for some AI tasks with low computational requirements or cost sensitivity, even if they are AI tasks, they can be allocated to the more cost-effective general computing domain, thus avoiding wasting high-cost intelligent computing resources on unnecessary scenarios and saving costs.

[0046] The technical solution of this invention, in response to an artificial intelligence task to be executed, identifies the task characteristics of the artificial intelligence task and obtains a first resource profile of a general computing domain and a second resource profile of an intelligent computing domain in a hybrid cloud environment; determines a first matching degree between the task characteristics and the first resource profile, and a second matching degree between the task characteristics and the second resource profile; based on the numerical relationship between the first and second matching degrees, determines a target computing domain from the general computing domain and the intelligent computing domain, and allocates the artificial intelligence task to the target computing domain, which can then be used to execute the artificial intelligence task. This technical solution, by first identifying task characteristics and obtaining resource profiles of different computing domains, and then obtaining matching degrees based on these profiles, allows for precise matching of the artificial intelligence task to a more suitable computing domain. This significantly optimizes the overall cost of computing resource usage while ensuring the execution performance of the artificial intelligence task, thus achieving a better balance between cost and performance.

[0047] Figure 2 This is a flowchart of another task execution method provided by an embodiment of the present invention. This embodiment is based on and optimized from the above-described technical solutions. In this embodiment, optionally, determining the first matching degree between task features and the first resource profile includes: obtaining at least one of the task type, resource requirements, and task migration cost of the artificial intelligence task based on the task features; and obtaining at least one of the hardware performance, storage availability, and expected response latency of the general computing domain based on the first resource profile; determining the first matching degree between task features and the first resource profile based on target information; wherein, the target information includes at least one of computing power matching degree, task type, task migration cost, storage availability, and expected response latency; the computing power matching degree is determined based on resource requirements and hardware performance. The explanations of terms that are the same as or corresponding to those in the above embodiments are not repeated here.

[0048] See Figure 2 The method in this embodiment may specifically include the following steps:

[0049] S210. In response to an artificial intelligence task to be executed, identify the task characteristics of the artificial intelligence task, and obtain a first resource profile of the general computing domain and a second resource profile of the intelligent computing domain in a hybrid cloud environment.

[0050] S220. Based on the task characteristics, obtain at least one of the following: task type, resource requirements, and task migration cost of the artificial intelligence task; and based on the first resource profile, obtain at least one of the following: hardware performance of the general computing domain, storage availability, and expected response latency.

[0051] Based on the task characteristics, at least one of the following is determined: task type, resource requirements, and task migration cost. The task type describes the type of AI task, such as inference, generation, or logic tasks as illustrated in the example above. Different types of AI tasks have vastly different requirements for computing resources. The resource requirements characterize the specifications of computing resources needed to execute the AI ​​task, particularly whether specific computing resources are needed for acceleration. The task migration cost characterizes the cost of migrating the AI ​​task between the general computing domain and the intelligent computing domain. This cost can be a quantitative or semi-quantitative indicator, such as network transmission time overhead for data migration, progress loss due to task interruption, or operational costs of reconfiguring the environment, etc., without specific limitations here.

[0052] Based on the first resource profile, at least one of the following is obtained for the general computing domain: hardware performance, storage availability, and expected response latency. Hardware performance can be understood as the hardware capability parameters of the computing nodes within the general computing domain, particularly representing CPU performance, such as the single-core / multi-core clock speed of the CPU. Storage availability represents the proportion of currently available storage resources (such as memory or video memory) to the total capacity of the general computing domain, reflecting its ability to handle tasks requiring high data throughput. Expected response latency can be understood as the predicted time required for an AI task to obtain computing resources (or complete) from submission, based on the current load of the general computing domain (load information included in the first resource profile), task queue status, and network conditions. This is crucial for AI tasks with time-sensitive requirements.

[0053] S230. Based on the target information, determine the first matching degree between the task characteristics and the first resource profile, wherein the target information includes at least one of computing power matching degree, task type, task migration cost, storage availability and expected response latency, and the computing power matching degree is determined based on resource requirements and hardware performance.

[0054] Among them, the computing power matching degree is determined based on resource requirements and hardware performance. This computing power matching degree can characterize the degree of matching between hardware performance and resource requirements.

[0055] Then, the first matching degree can be determined based on at least one of the following: computing power matching degree, task type, task migration cost, storage availability, and expected response latency. For example, the first matching degree can be directly determined based on this information (i.e., the target information described in this step). Even more for example, when at least two types of target information exist, the first matching degree can be obtained by comprehensively weighting the various target information according to their respective weight parameters. For instance, for the i-th AI task, its first matching degree Scorei = w1×Ci + w2×Mi + w3×Li + w4×Ti, where Ci represents the computing power matching degree, Mi represents the storage availability, Li represents the expected response latency, Ti represents the task migration cost, and w1, w2, w3, and w4 are weight parameters. Here is an example: for a high I / O demand model training data preprocessing task, the matching degree calculation might assign higher weight parameters to storage availability and computing power matching degree, and lower weight parameters to task migration cost, ultimately obtaining a quantified first matching degree.

[0056] S240. Determine the second degree of matching between task features and the second resource profile.

[0057] The process for determining the second matching degree can be the same as or different from the process for determining the first matching degree. This can be set according to actual needs and is not specifically limited here.

[0058] S250. Based on the numerical relationship between the first matching degree and the second matching degree, determine the target computing domain from the general computing domain and the intelligent computing domain, and assign the artificial intelligence task to the target computing domain so as to utilize the target computing domain to execute the artificial intelligence task.

[0059] The technical solution of this invention significantly improves the accuracy of computing resource matching and the rationality of decision-making by refining the first matching degree calculation process into a comprehensive consideration of one or more dimensions of information such as task type, task migration cost, storage availability, expected response latency and computing capacity matching degree.

[0060] Based on determining the first matching degree using weight parameters, an optional technical solution, after performing the artificial intelligence task using the target computing domain, further includes:

[0061] Acquire execution data for the artificial intelligence task, including execution logs and / or performance data during the execution of the artificial intelligence task; adjust the weight parameters based on the execution data.

[0062] Execution data can be understood as data collected during and after the actual execution of an AI task on the target computing domain, reflecting the execution status and effect of the AI ​​task. In this technical solution, this data can be an execution log, which can be understood as a text record of at least one of the following: key events, status changes, error information, and resource usage details during the execution of the AI ​​task. For example, it can be at least one of the following: task start time, task end time, time spent in each stage, intermediate checkpoint information, and possible anomalies or warnings. This is related to the actual situation and is not specifically limited here. The data can also be performance data, which can be quantifiable indicators collected during the execution of the AI ​​task, reflecting resource utilization efficiency and task execution effect. For example, it can be at least one of the following: task execution latency, energy consumption, GPU / CPU utilization, peak memory usage, storage I / O throughput, and network bandwidth usage. This is related to the actual situation and is not specifically limited here.

[0063] The weight parameters are adjusted based on execution data, using the collected execution data as feedback signals to evaluate the actual effectiveness of the scheduling decision. For example, if an AI task is assigned to the GPU, but the monitored GPU utilization is only 62%, while the execution latency is better than expected, this may indicate that although the AI ​​task is completed quickly, GPU resources are not being fully utilized, resulting in potential resource waste. In this case, a preset optimization algorithm can be applied to fine-tune the weight parameters. For example, the weight parameters related to computing power matching can be appropriately reduced, or the weight parameters related to resource utilization can be increased, thereby optimizing the computing power allocation for subsequent similar AI tasks. This achieves self-learning optimization, representing a closed-loop optimization allocation process for AI tasks.

[0064] Before describing the following embodiments, we will first provide an exemplary description of their application scenarios. For example, based on executing AI tasks in a hybrid cloud environment, serverless architecture, due to its features such as function-level autoscaling and event-driven execution, is increasingly being chosen by enterprises to combine hybrid cloud environments with serverless architecture. Specifically, this involves deploying an AI agent in a hybrid cloud environment and then running the AI ​​agent through a serverless architecture to execute AI tasks in the hybrid cloud environment.

[0065] Next, in this application scenario, the task execution process in the following embodiments will be described in detail.

[0066] Figure 3This is a flowchart of another task execution method provided by an embodiment of the present invention. This embodiment is based on and optimized from the above-described technical solutions. In this embodiment, optionally, executing an artificial intelligence task using a target computing domain includes: creating a serverless runtime on the target computing domain, obtaining the function corresponding to the artificial intelligence task, adding the function to the serverless runtime to obtain a function instance; and executing the function instance using the target computing domain to execute the artificial intelligence task. The explanations of terms that are the same as or corresponding to those in the above embodiments are not repeated here.

[0067] See Figure 3 The method in this embodiment may specifically include the following steps:

[0068] S310. In response to an artificial intelligence task to be executed, identify the task characteristics of the artificial intelligence task, and obtain a first resource profile of the general computing domain and a second resource profile of the intelligent computing domain in a hybrid cloud environment.

[0069] S320. Determine the first degree of matching between the task features and the first resource profile, and determine the second degree of matching between the task features and the second resource profile.

[0070] S330. Based on the numerical relationship between the first matching degree and the second matching degree, determine the target computing domain from the general computing domain and the intelligent computing domain, and assign the artificial intelligence task to the target computing domain.

[0071] S340. Create a serverless runtime on the target computing domain, obtain the function corresponding to the artificial intelligence task, add the function to the serverless runtime, and obtain the function instance.

[0072] In this context, a serverless runtime can be understood as an environment of executable code conforming to serverless architecture standards, created within the target computing domain. This environment is typically containerized, such as an on-demand container instance, which encapsulates the basic runtime libraries and dependencies required for AI task execution. However, its lifecycle (creation, operation, and destruction) is entirely managed by the scheduling system, without requiring users to pre-configure or maintain persistent servers. The serverless runtime is created on the target computing domain.

[0073] A function can be understood as the function corresponding to an AI task, or more specifically, as an executable unit that encapsulates the execution logic of an AI task according to the Serverless functional programming model. This function can be added to the Serverless runtime created above to form a concrete, runnable function instance.

[0074] Building upon this, optionally, functions added to the serverless runtime can be migrated between the general computing domain and the intelligent computing domain via virtualization interfaces corresponding to the general computing domain and the intelligent computing domain. The virtualization interface can be understood as a unified and standardized Application Programming Interface (API) exposed for the two different hardware architectures of the general computing domain and the intelligent computing domain. It provides consistent resource access, memory management, and task scheduling services to the upper-layer serverless runtime, thereby shielding the architectural differences between the general computing domain and the intelligent computing domain. This means that the execution of functions added to the serverless runtime is no longer tightly bound to a specific computing domain. When rescheduling is required due to load balancing, cost optimization, hardware failure, or policy adjustments, the function can be transparently migrated from the current computing domain (i.e., the target computing domain) to another computing domain without code modification by calling the aforementioned virtualization interface. This achieves seamless function migration and, consequently, unified management and automatic elastic migration of general computing resources and intelligent computing resources.

[0075] S350. Perform artificial intelligence tasks by utilizing the target computation domain to execute function instances.

[0076] Specifically, the AI ​​task is actually completed by executing function instances within the target computing domain. Optionally, after the AI ​​task is completed, the Serverless runtime and its function instances can typically be automatically reclaimed to achieve automatic resource release.

[0077] The technical solution of this invention introduces a serverless execution mode, bringing extreme elasticity and resource efficiency to AI tasks in a hybrid cloud environment. Specifically, it can create and destroy function instances on demand in the target computing domain at the millisecond level according to the arrival of AI tasks. This significantly reduces resource idle waste and lowers costs. Secondly, it simplifies the deployment and operation complexity of AI applications, allowing developers to focus more on the business logic itself.

[0078] An optional technical solution, the above task execution method further includes:

[0079] Based on the task characteristics, a model is obtained for performing the artificial intelligence task, and the model parameters are obtained from the cache of the target computing domain.

[0080] Adding a function to a serverless runtime creates a function instance, including:

[0081] Add the function and model parameters to the serverless runtime to obtain a function instance.

[0082] By analyzing task characteristics, the model required to perform the AI ​​task can be determined, such as a visual detection model or a language model. This depends on the specific circumstances and is not specifically limited here. Furthermore, the stored model parameters are directly retrieved from the local cache or high-speed cache of the target computing domain. This cache is persistent and separate from the temporary execution code environment described above.

[0083] Therefore, when creating a function instance, the function and model parameters quickly retrieved from the cache can be injected into the Serverless runtime together. Since the model parameters are already in place, the function instance does not need to undergo the time-consuming model download and loading process, thus achieving hot start of the function instance.

[0084] The above technical solution decouples model parameters from temporary function instances and persists them in the cache through model parameter caching. This enables hot start response in seconds or even milliseconds when the same model is called repeatedly, thereby significantly reducing the average response time of AI tasks and greatly improving the real-time experience and throughput of AI services, especially inference services.

[0085] Another alternative technical solution is that the artificial intelligence task is a task driven by an artificial intelligence agent, and the above task execution method also includes:

[0086] Obtain the context of the AI ​​agent from the object storage of the target computation domain;

[0087] Adding a function to a serverless runtime creates a function instance, including:

[0088] Add the function and context to the serverless runtime to obtain a function instance.

[0089] For AI tasks driven by AI agents, execution relies not only on code and models but also on a context that records historical interactions, decision states, and session memories. In this technical solution, this context is persistently stored in an object storage independent of function instances, achieving state separation. Furthermore, when creating a function instance, in addition to adding the function to the Serverless runtime, the context quickly retrieved from object storage can be injected into the Serverless runtime. This allows newly created function instances to immediately resume to their previous or specified execution state without needing to accumulate context from scratch.

[0090] The above technical solution, through context state separation storage and fast recovery mechanism, ensures the continuity of AI Agent task state, enabling AI Agent to seamlessly warm start between different function instances. This further shortens the cold start latency of complex AI tasks, ensures the coherence of AI Agent interaction and low latency response, and is particularly suitable for multi-turn interaction scenarios that require memory, such as dialogue and decision-making.

[0091] Another alternative technical solution is that the artificial intelligence task is a task driven by an artificial intelligence agent and the function instance is a containerized function instance. The above task execution method also includes:

[0092] Based on the time series of historical tasks driven by the AI ​​agent, predict the number of AI tasks that the AI ​​agent will drive in the future.

[0093] Depending on the quantity, perform a shrink operation or an expand operation and warm up the containerized function instance.

[0094] This involves collecting and analyzing the distribution data of historical tasks driven by the AI ​​Agent over time, forming a time series. This time series is then analyzed, for example using predictive models (such as time series analysis models), to determine the number of AI tasks the AI ​​Agent might initiate in future periods. Based on this, appropriate elastic scaling operations can be performed according to the number of tasks. Specifically,

[0095] Perform scaling operations and warm up containerized function instances. When a significant increase in the number of AI tasks is predicted in the future, the computing resource pool can be proactively scaled up on the corresponding target computing domain. More importantly, a batch of containerized function instances can be warmed up (i.e., created and initialized in advance) to put them in a standby state and prepare the models and basic runtime environment.

[0096] Perform a scaling-down operation. When it is predicted that the number of AI tasks will enter a low period in the future, the system can proactively scale down to release excess computing resources and containerized function instances, thereby avoiding resource idleness.

[0097] The above technical solution proactively responds to AI task load fluctuations through a predictive elastic scaling mechanism, thereby expanding capacity and warming up containerized function instances in advance before the peak period of AI tasks, ensuring that AI tasks are executed without delay, significantly improving the response speed of AI tasks, and automatically shrinking capacity to release idle resources during low load periods. Thus, while ensuring high performance, it maximizes the utilization of computing resources and optimizes the overall cost.

[0098] Building upon this foundation, to better understand the various technical solutions described above, we will use a power inspection AIAgent as an example for illustrative explanation. For instance, in a power scenario, general computing resources are deployed in the headquarters' private cloud for rule judgment and report generation, while intelligent computing resources are located in the provincial intelligent computing center for AI model inference. The computing power abstraction layer masks architectural differences through a unified interface, enabling the serverless runtime to transparently access computing resources of different types, achieving seamless migration of functions between general and intelligent computing.

[0099] When the power inspection AI Agent initiates an AI task, it acquires the task metadata and derives the task characteristics based on this metadata. In this example, the task metadata includes the model type (visual inference), model size (2GB), and real-time requirements (<1s response time). Therefore, the task characteristics are: the task type is inference, and the resource requirement is GPU acceleration. Then, a first matching degree is calculated based on the task characteristics and a first resource profile, and a second matching degree is calculated based on the task characteristics and a second resource profile. Analysis shows that while the general computing domain has sufficient CPUs, it cannot meet the inference latency requirements, while the intelligent computing domain has available GPUs with low load. Therefore, the AI ​​task is assigned to the intelligent computing domain for execution. Based on this, several specific AI tasks are given below: image recognition tasks can be assigned to the intelligent computing domain (GPU nodes), alarm classification tasks can be assigned to the general computing domain (CPU nodes), and log generation tasks can be distributed to edge nodes (lightweight general computing instances).

[0100] Furthermore, a Serverless runtime is created on the target computing domain. At this point, model parameters can be retrieved from the cache and the context can be retrieved from the object storage. Then, the functions, model parameters, and context corresponding to the AI ​​task are added to the Serverless runtime to obtain containerized function instances, thereby achieving warm start response.

[0101] After executing the containerized function instance, the output anomaly detection results can be returned to the business platform through the interface, and logs and performance data can be uploaded to optimize the weight parameters.

[0102] The above example, by integrating serverless architecture with intelligent scheduling algorithms (i.e., algorithms that select target computing domains to schedule AI tasks to target computing domains), enables efficient, intelligent, and elastic execution of AI tasks in a hybrid cloud environment where both general computing domains and intelligent computing domains are online simultaneously.

[0103] Figure 4This is a structural block diagram of a task execution system provided in an embodiment of the present invention. This system is used to execute the task execution methods provided in any of the above embodiments. This system and the task execution methods of the above embodiments belong to the same inventive concept. Details not described in detail in the embodiments of the task execution system can be found in the embodiments of the above task execution methods. See also... Figure 4 The system may specifically include: a resource profile acquisition module 410, a matching degree determination module 420, and a task execution module 430.

[0104] The resource profile acquisition module 410 is used to respond to the artificial intelligence task to be executed, identify the task characteristics of the artificial intelligence task, and acquire the first resource profile of the general computing domain and the second resource profile of the intelligent computing domain in the hybrid cloud environment.

[0105] The matching degree determination module 420 is used to determine the first matching degree between the task features and the first resource profile, and to determine the second matching degree between the task features and the second resource profile.

[0106] The task execution module 430 is used to determine the target computing domain from the general computing domain and the intelligent computing domain based on the numerical relationship between the first matching degree and the second matching degree, and to allocate the artificial intelligence task to the target computing domain so as to execute the artificial intelligence task using the target computing domain.

[0107] Optionally, the matching degree determination module 420 may include:

[0108] The data analysis unit is used to obtain at least one of the following based on the task characteristics: task type, resource requirements, and task migration cost of the artificial intelligence task; and to obtain at least one of the following based on the first resource profile: hardware performance of the general computing domain, storage availability, and expected response latency.

[0109] The first matching degree determination unit is used to determine the first matching degree between the task features and the first resource profile based on the target information;

[0110] The target information includes at least one of the following: computing power matching degree, task type, task migration cost, storage availability, and expected response latency.

[0111] The computing power matching degree is determined based on resource requirements and hardware performance.

[0112] In addition, the target information may optionally include at least two of the following: computing power matching degree, task type, task migration cost, storage availability and expected response latency.

[0113] The first matching degree determination unit is specifically used for:

[0114] Based on at least two types of target information and the weight parameters corresponding to each type of target information, determine the first matching degree between the task features and the first resource profile.

[0115] Based on this, optionally, the above-mentioned task execution system may also include:

[0116] The execution data acquisition module is used to acquire the execution data of the artificial intelligence task after the artificial intelligence task is executed using the target computing domain. The execution data includes execution logs and / or performance data during the execution of the artificial intelligence task.

[0117] The weight parameter adjustment module is used to adjust the weight parameters based on the execution data.

[0118] Optionally, the task execution module 430 may include:

[0119] The function instance is obtained by creating a serverless runtime on the target computing domain, obtaining the function corresponding to the artificial intelligence task, adding the function to the serverless runtime, and obtaining the function instance.

[0120] The task execution unit is used to perform artificial intelligence tasks by utilizing the execution function instance of the target computation domain.

[0121] Based on this, optionally, the above-mentioned task execution system may also include:

[0122] The model parameter acquisition module is used to obtain the model to be used to perform the artificial intelligence task based on the task characteristics, and to obtain the model parameters of the model from the cache of the target computing domain.

[0123] A function instance can obtain a unit, which may include:

[0124] The first sub-unit is used to add functions and model parameters to the serverless runtime to obtain function instances.

[0125] Alternatively, if the artificial intelligence task is a task driven by an artificial intelligence agent, then the above-mentioned task execution system may further include:

[0126] The context acquisition module is used to obtain the context of the AI ​​agent from the object storage of the target computing domain;

[0127] A function instance can obtain a unit, which may include:

[0128] The second subunit is used to add functions and contexts to the serverless runtime to obtain function instances.

[0129] Alternatively, the AI ​​task is a task driven by an AI agent and the function instance is a containerized function instance. The aforementioned task execution system may also include:

[0130] The quantity prediction module can be used to predict the number of AI tasks driven by the AI ​​agent in future periods based on the time series of historical tasks already driven by the AI ​​agent.

[0131] The scaling module is used to perform scaling down or scaling up operations based on the quantity and to preheat containerized function instances.

[0132] Alternatively, for the virtualization interfaces corresponding to the general computing domain and the intelligent computing domain, functions added to the serverless runtime can migrate between the general computing domain and the intelligent computing domain through the virtualization interfaces.

[0133] The task execution system provided in this embodiment of the invention, through a resource profiling module, responds to an AI task to be executed by identifying the task characteristics of the AI ​​task and acquiring a first resource profile of a general computing domain and a second resource profile of an intelligent computing domain in a hybrid cloud environment; through a matching degree determination module, it determines a first matching degree between the task characteristics and the first resource profile, and a second matching degree between the task characteristics and the second resource profile; through a task execution module, based on the numerical relationship between the first and second matching degrees, it determines a target computing domain from the general computing domain and the intelligent computing domain, and allocates the AI ​​task to the target computing domain, which can then be used to execute the AI ​​task. This system, by first identifying task characteristics and acquiring resource profiles of different computing domains, and then obtaining a matching degree based on this, can accurately match the AI ​​task to a more suitable computing domain based on the matching degree. This significantly optimizes the overall cost of computing resource usage while ensuring the execution performance of the AI ​​task, thus achieving a better balance between cost and performance.

[0134] The task execution system provided in the embodiments of the present invention can execute the task execution method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0135] It is worth noting that in the above-described embodiments of the task execution system, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.

[0136] Figure 5A schematic diagram of an electronic device 10, which can be used to implement embodiments of the present invention, is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0137] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0138] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0139] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as task execution methods.

[0140] In some embodiments, the task execution method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the task execution method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to execute the task execution method by any other suitable means (e.g., by means of firmware).

[0141] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chips or system-on-a-chips (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0142] Computer programs used to implement the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs can be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0143] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0144] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0145] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0146] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0147] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication unit 19, or installed from storage unit 18, or installed from ROM 12. When the computer program is executed by processor 11, it performs the functions defined in the methods of the embodiments of the present invention.

[0148] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0149] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A task execution method, characterized in that, include: In response to an artificial intelligence task to be executed, the task characteristics of the artificial intelligence task are identified, and a first resource profile of the general computing domain and a second resource profile of the intelligent computing domain in a hybrid cloud environment are obtained. Determine a first degree of matching between the task features and the first resource profile, and determine a second degree of matching between the task features and the second resource profile; Based on the numerical relationship between the first matching degree and the second matching degree, a target computing domain is determined from the general computing domain and the intelligent computing domain, and the artificial intelligence task is assigned to the target computing domain to execute the artificial intelligence task using the target computing domain.

2. The method according to claim 1, characterized in that, Determining the first matching degree between the task features and the first resource profile includes: Based on the task characteristics, at least one of the following is obtained: task type, resource requirements, and task migration cost of the artificial intelligence task; and based on the first resource profile, at least one of the following is obtained: hardware performance, storage availability, and expected response latency of the general computing domain. Based on the target information, a first matching degree is determined between the task features and the first resource profile; The target information includes at least one of the following: computing power matching degree, task type, task migration cost, storage availability, and expected response latency. The computing power matching degree is determined based on the resource requirements and the hardware performance.

3. The method according to claim 2, characterized in that, The target information includes at least two of the following: computing power matching degree, task type, task migration cost, storage availability, and expected response latency. Determining the first matching degree between the task features and the first resource profile based on the target information includes: A first matching degree between the task features and the first resource profile is determined based on at least two types of target information and weight parameters corresponding to each type of target information.

4. The method according to claim 3, characterized in that, Following the execution of the artificial intelligence task using the target computing domain, the method further includes: Obtain the execution data of the artificial intelligence task, wherein the execution data includes execution logs and / or performance data during the execution of the artificial intelligence task; The weight parameters are adjusted based on the execution data.

5. The method according to claim 1, characterized in that, The execution of the artificial intelligence task using the target computing domain includes: A serverless runtime is created on the target computing domain, and the function corresponding to the artificial intelligence task is obtained. The function is then added to the serverless runtime to obtain a function instance. The artificial intelligence task is performed by executing the function instance using the target computing domain.

6. The method according to claim 5, characterized in that, Also includes: The model to be applied for performing the artificial intelligence task is obtained based on the task characteristics, and the model parameters of the model are obtained from the cache of the target computing domain. Adding the function to the serverless runtime to obtain a function instance includes: The function and the model parameters are added to the serverless runtime to obtain a function instance.

7. The method according to claim 5, characterized in that, The artificial intelligence task is a task driven by an artificial intelligence agent, and the method further includes: Obtain the context of the artificial intelligence agent from the object storage of the target computing domain; Adding the function to the serverless runtime to obtain a function instance includes: The function and the context are added to the serverless runtime to obtain a function instance.

8. The method according to claim 5, characterized in that, The artificial intelligence task is a task driven by an artificial intelligence agent, and the function instance is a containerized function instance. The method further includes: Based on the time series of historical tasks already driven by the AI ​​agent, predict the number of AI tasks that the AI ​​agent will drive in future time periods; Based on the stated quantity, perform a shrinking operation or an expanding operation and warm up the containerized function instance.

9. The method according to claim 5, characterized in that, For the virtualization interfaces corresponding to the general computing domain and the intelligent computing domain, the functions added to the serverless runtime migrate between the general computing domain and the intelligent computing domain through the virtualization interfaces.

10. A task execution system, characterized in that, include: The resource profile acquisition module is used to respond to the artificial intelligence task to be executed, identify the task characteristics of the artificial intelligence task, and acquire a first resource profile of the general computing domain and a second resource profile of the intelligent computing domain in a hybrid cloud environment. The matching degree determination module is used to determine a first matching degree between the task feature and the first resource profile, and to determine a second matching degree between the task feature and the second resource profile; The task execution module is used to determine a target computing domain from the general computing domain and the intelligent computing domain based on the numerical relationship between the first matching degree and the second matching degree, and to allocate the artificial intelligence task to the target computing domain so as to execute the artificial intelligence task using the target computing domain.

Citation Information

Cited By

  • Power business scene-oriented operator scheduling method and device, medium and equipment

    CN122044901A