Task processing method, artificial intelligence chip and storage medium

By receiving task flow in a heterogeneous computing environment, determining task flow characteristics and selecting appropriate prediction models, the problem of low prediction accuracy of task execution time in the task flow is solved, and the balance of scheduling between task flows is improved.

CN120179343APending Publication Date: 2025-06-20CAMBRIAN (KUNSHAN) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311744804.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-18
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In a heterogeneous computing environment, the prior art predicts the accuracy of task execution time in a task flow, which affects the balance of task flow scheduling.

Method used

By receiving pending tasks in at least two task flows, each task flow feature is determined and a suitable target duration prediction model, including a linear fit model or a long short-term memory model, is selected based on these features to predict the duration of the task.

Benefits of technology

Using prediction models that match the task flow characteristics can accurately and quickly predict the task execution time in the task flow, thereby improving the balance of scheduling between task flows.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179343A_ABST
    Figure CN120179343A_ABST
Patent Text Reader

Abstract

The invention provides a task processing method, an artificial intelligence chip and a storage medium, and the method comprises the steps: receiving a to-be-processed task in at least two task flows, and determining task flow features corresponding to the at least two task flows, the task flow features being determined according to historical task data of the task flows; for each task flow, a target duration prediction model corresponding to the task flow is determined according to the task flow characteristics of the task flow, and the target duration prediction model comprises a linear fitting model or a long-short-term memory model; and in response to the fact that the target duration prediction model corresponding to the task flow is the corresponding long-short-term memory model, determining the predicted execution duration of the to-be-processed task in the task flow based on the long-short-term memory model. The task execution duration in the task flow can be accurately and quickly predicted, so that the scheduling balance among the task flows can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a task processing method, an artificial intelligence chip, and a storage medium. Background Art

[0002] In a heterogeneous computing environment using an artificial intelligence (AI) chip, that is, an environment in which a central processing unit (CPU) and an AI chip jointly perform calculations, tasks are sent from the CPU to the AI chip side through a task stream. To support parallel processing of tasks, there can be multiple streams. Tasks within each stream are executed serially, and tasks between streams can be executed in parallel when there is no explicit dependency. A scheduling module can be set in the AI chip, and the scheduling module in the AI chip selects tasks from the stream and sends them to the AI chip hardware for execution.

[0003] To achieve relatively fair use of the hardware resources of the AI chip by different streams, it is necessary to count and continuously update the cumulative execution duration of tasks in each stream over a period of time. When counting the cumulative execution duration of tasks in each stream, it is necessary to predict the execution duration of tasks to be sent to the AI hardware resources for processing in the stream.

[0004] Currently, the accuracy of the prediction results of the execution duration of tasks in the stream is relatively low, which affects the balance of task flow scheduling. Summary of the Invention

[0005] The present disclosure provides a task processing method, an artificial intelligence chip, and a storage medium.

[0006] According to a first aspect of the present disclosure, there is provided a task processing method, the method including: receiving tasks to be processed in at least two task streams, and determining task flow characteristics corresponding to the at least two task streams respectively, where the task flow characteristics are determined according to historical task data of the task streams; for each task stream, determining a target duration prediction model corresponding to the task stream according to the task flow characteristics of the task stream, where the target duration prediction model includes a linear fitting model or a long short-term memory model; in response to the target duration prediction model corresponding to the task stream being the corresponding long short-term memory model, determining a predicted execution duration of the tasks to be processed in the task stream based on the long short-term memory model.

[0007] According to a second aspect of the present disclosure, there is provided an artificial intelligence chip. The artificial intelligence chip is connected to a central processing unit of a host. The artificial intelligence chip includes a processor core and a scheduling module, and the scheduling module is communicatively connected to each processor core. The scheduling module is configured to: receive the to-be-processed tasks corresponding to at least two task flows sent by the central processing unit, and obtain the task flow characteristics corresponding to the at least two task flows respectively, wherein the task flow characteristics of each task flow are determined according to the historical task data of the task flow; for each task flow, determine a target duration prediction model corresponding to the task flow according to the task flow characteristics of the task flow, and the target duration prediction model includes a linear fitting model or a long short-term memory model; in response to the target duration prediction model corresponding to the task flow being the corresponding long short-term memory model, determine the predicted execution duration of the to-be-processed tasks in the task flow based on the target long short-term memory model.

[0008] According to a third aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method provided in the first aspect.

[0009] According to a fourth aspect of the present disclosure, there is provided a computer program product, which includes: a computer program stored in a readable storage medium. At least one processor of an electronic device can read the computer program from the readable storage medium, and the at least one processor executes the computer program to cause the electronic device to execute the method provided in the first aspect.

[0010] According to the task processing method, artificial intelligence chip and storage medium provided by the present disclosure, by receiving the to-be-processed tasks in at least two task flows, the task flow characteristics corresponding to the at least two task flows are determined, wherein the task flow characteristics are determined according to the historical task data of the task flow; for each task flow, a target duration prediction model corresponding to the task flow is determined according to the task flow characteristics of the task flow, and the target duration prediction model includes a linear fitting model or a long short-term memory model; in response to the target duration prediction model corresponding to the task flow being the corresponding long short-term memory model, the predicted execution duration of the to-be-processed tasks in the task flow is determined based on the long short-term memory model. Since the linear fitting model can predict the execution duration of tasks in a task flow with good regularity with less computational effort, and the LSTM model can predict the execution duration of tasks in a complex task flow, by combining the LSTM model and the linear fitting model to predict the execution duration of tasks in different task flows, compared with using the same task execution duration prediction method for different task flows, using the target duration prediction model matching the task flow characteristics to predict the execution duration of the to-be-processed tasks in the task flow can predict the task execution duration in the task flow more accurately and quickly, and thus can improve the balance of scheduling between task flows.

[0011] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0013] Figure 1 is an application scenario diagram of the AI chip of the present disclosure in a heterogeneous computing environment;

[0014] Figure 2 is a schematic structural diagram of the artificial intelligence chip provided by the present disclosure;

[0015] Figure 3 is a schematic flowchart of an embodiment of the task processing method provided by the present disclosure;

[0016] Figure 4 is a schematic flowchart of another embodiment of the task processing method provided by the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0018] In a heterogeneous computing environment, an artificial intelligence chip (AI) can receive different task streams sent by a host (such as a central processing unit (CPU) on a terminal device). Each task stream can include multiple tasks. To make the use of the hardware resources of the AI chip relatively balanced for each stream, the cumulative execution duration of the tasks in each stream can be statistically calculated, and the hardware resources of the AI chip can be configured for the tasks to be processed in each stream according to the cumulative execution duration corresponding to each stream.

[0019] Since the tasks of a stream include executed tasks and tasks to be executed, and there is a real execution time for the executed tasks, in order to statistically calculate the cumulative execution duration of the tasks in the stream, it is necessary to determine the predicted execution duration of the tasks to be executed in the stream.

[0020] When predicting the execution duration of a task, a linear fitting method is commonly used to predict the task execution duration. This method assumes that the execution durations corresponding to the tasks in a stream are stable or regular, that is, the execution durations of the tasks change smoothly. Therefore, the execution durations of each task can be reasonably deduced based on the above stability or regularity. However, this method has poor accuracy in predicting the execution duration of scenarios with sudden changes in execution duration (the execution durations of tasks do not change smoothly), which in turn leads to unreasonable scheduling results among streams.

[0021] Please refer to Figure 1 , Figure 1 which shows an application scenario diagram of the AI chip of the present disclosure in a heterogeneous computing environment. As Figure 1 shown, the AI chip 11 includes a scheduling module 111 and multiple computing cores 112. A host (such as a central processing unit) 12 can send tasks to the AI chip 11 through multiple task streams (streams). The AI chip 11 receives tasks in different streams, and the scheduling module determines and assigns respective computing cores for the tasks in each task stream.

[0022] The AI chip may include multiple scheduling modules and multiple clusters, and each cluster includes multiple computing cores. The scheduling module can schedule tasks on each computing core.

[0023] Please refer to Figure 2 , Figure 2 which is a schematic structural diagram of an AI chip. As Figure 2 shown, the AI chip includes multiple clusters 13 (such as Figure 2 the four clusters in Figure 2 ) and a scheduling module 111. Each cluster 13 may include multiple computing cores (such as

[0024] the four computing cores in

[0025] To solve the problem that in the existing solutions, for task flows with sudden changes in task execution duration, inaccurate prediction of task execution duration may lead to poor rationality in task scheduling among different task flows, the present disclosure predicts the category of a task flow based on historical task data in the task flow, determines a target duration prediction model corresponding to the task according to the category, and the target duration prediction model includes a linear fitting function model and a Long Short Term Memory (LSTM) model. The LSTM model can predict the execution duration of tasks in complex task flows. Compared with using the same task execution duration prediction method for different streams, using the target duration prediction model matching the task flow characteristics to predict the execution duration of the tasks to be processed in the task flow can predict the task execution duration in the task flow more accurately and quickly, thereby improving the balance of scheduling among task flows.

[0026] Please refer to Figure 3 , Figure 3 which is a schematic flowchart of an embodiment of the task processing method provided by the present disclosure. As Figure 3 shown, the task processing method provided in this embodiment includes the following steps:

[0027] S301. Receive the tasks to be processed in at least two task flows, and determine the task flow characteristics corresponding to each of the at least two task flows, where the task flow characteristics of the task flow are determined according to the historical task data of the task flow.

[0028] In this embodiment, the execution subject of the task processing method may be the AI chip itself, specifically, the scheduling module in the AI chip.

[0029] The AI chip can receive at least two task flows sent by the host. Each task flow may include multiple tasks. The tasks in a task flow may include computing tasks or input / output tasks.

[0030] The tasks in each task flow may carry a task flow identifier. After receiving the tasks in each task flow, the AI chip may send each task flow to the scheduling module of the AI chip, and the scheduling module may identify the task flow to which each task belongs according to the task flow identifier of the input task.

[0031] Each task in each task flow may have at least one task characteristic, and the task flow characteristic corresponding to the task flow may be a task characteristic with a relatively high degree of correlation with the target prediction variable (i.e., the predicted execution duration). Among them, the task flow characteristics include, but are not limited to: the number of tasks, the average time interval between task arrivals, the average execution duration of tasks, the task execution duration change characteristics, etc.

[0032] S302. For each task flow, determine the target duration prediction model corresponding to the task flow according to the task flow characteristics of the task flow. The target duration prediction model includes a linear fitting function model or a long short-term memory model.

[0033] The present disclosure can determine the predicted execution durations corresponding to different tasks to be processed by combining a linear fitting function model and a long short-term memory model.

[0034] The long short-term memory model is suitable for predicting the execution duration of more complex tasks, but its time complexity and space complexity are higher than those of the linear fitting function model. Using LSTM to predict the task execution duration takes a longer time and requires more hardware resources. Since the predicted execution duration of the prediction task is executed in the AI chip, in order to ensure that the AI chip has a high processing speed, the scheduling module of the AI chip needs to achieve a relatively accurate prediction of the task execution duration under the premise of a short time and few hardware resources. Therefore, for a task flow with better execution duration stability (that is, the task execution duration of the task flow changes smoothly), the linear fitting function model is preferably used to predict the execution duration of the task to be executed. For a task flow with a sudden change in execution duration, that is, a task flow with a large change in execution duration, the long short-term memory model can be used to predict the execution duration of the tasks in the task flow. Therefore, before selecting the target duration prediction model corresponding to the task flow, it is necessary to identify the task flow characteristics of the task flow.

[0035] The present disclosure can pre-determine the task flow characteristics of the task flows that are respectively suitable for the LSTM model and the linear fitting function model to predict the task execution duration, and record the matching task flow characteristics corresponding to each model. Further, the present disclosure can configure the respective target duration prediction models according to different task flow characteristics.

[0036] For the task flows being processed, the present disclosure can statistically analyze the task flow characteristics of each task flow according to the historical task data of these task flows being processed, and determine the target prediction model corresponding to the task flow according to the matching relationship between the task flow characteristics and the matching task flow characteristics corresponding to each model.

[0037] S303: In response to the target duration prediction model corresponding to the task flow being a long short-term memory model, determine the predicted execution duration of the task to be processed in the task flow based on the long short-term memory model.

[0038] The LSTM model can be a pre-trained model for predicting the task execution duration according to historical data.

[0039] When it is determined that the target duration prediction model corresponding to a task flow is the LSTM model, the above-mentioned scheduling module can input the task to be processed in the task flow into the LSTM model for the LSTM to perform inference operations, so as to predict the execution duration of the above-mentioned task to be processed.

[0040] The above LSTM model is used to predict the execution duration of the Nth task based on the task features of the first N - 1 tasks in the input time series, the actual execution duration of the tasks, and the Nth task feature (task feature excluding the execution duration).

[0041] Schematically, if the determined sequence length is 5, the LSTM model can predict the predicted execution duration of the task corresponding to the x(t + 1) moment based on the task features of the input x(t - 4), x(t - 3), x(t - 2), x(t - 1), x(t). The task features of each task in the task flow are input into the above model in a streaming manner, and the above target LSTM model predicts the execution duration of each task in the task flow.

[0042] In this embodiment, the to-be-processed tasks in at least two task flows are received, and the task flow features corresponding to the at least two task flows are determined, where the task flow features are determined according to the historical task data of the task flow; for each task flow, a target duration prediction model corresponding to the task flow is determined according to the task flow features of the task flow, and the target duration prediction model includes a linear fitting model or a long short-term memory model; in response to the target duration prediction model corresponding to the task flow being a corresponding long short-term memory model, the predicted execution duration of the to-be-processed task in the task flow is determined based on the long short-term memory model. The linear fitting model can predict the execution duration of tasks in a task flow with better regularity with less computational effort, and the LSTM model can predict the execution duration of tasks in a complex task flow. By combining the LSTM model and the linear fitting model to predict the execution duration of tasks in different task flows, compared with using the same task execution duration prediction method for different task flows, using the target duration prediction model matching the task flow characteristics to predict the execution duration of the to-be-processed task in the task flow can predict the task execution duration in the task flow more accurately and quickly, and thus can improve the balance of scheduling between task flows.

[0043] As an implementation, the task flow features can be obtained according to the task data of all tasks issued during the entire life cycle of a task flow from creation to destruction. Specifically, for each task flow, the present disclosure can extract the task features of the historical tasks in the task flow to obtain the task flow features corresponding to the task flow according to the task features of the historical tasks in the task flow. Among them, the above task features can be the features of a single task, including but not limited to: task type, task execution duration, task data parallel scale. The above task types include access storage-intensive and compute-intensive. The above task features can be represented in the form of a task feature vector.

[0044] Among them, the above task types can be preset. Access-intensive can mean that most of the execution time of a single task is used for the input and output of data in the data storage area. Computation-intensive can mean that most of the execution time of a single task is used for computing.

[0045] Among them, the task execution duration of a task refers to the time elapsed from when the computing core receives the task to when the task is executed and completed.

[0046] The task data parallel scale refers to the minimum number of processor cores required for a task to execute. For example, the task data parallel scale can be U1, U2, U3, and U4, where U1 means the task requires 4 computing cores, U2 means the task requires 8 computing cores, U3 means the task requires 12 computing cores; U4 means the task requires 16 computing cores. The number of computing cores here is only for explanation and does not limit the specific number of computing cores.

[0047] For each task flow, the present disclosure can determine the task flow characteristics of the task flow according to the historical task data and the task characteristics of each historical task in the task flow. For example, the number of tasks can be obtained by accumulating the number of historical tasks. The average time interval between task arrivals can be obtained by calculating the time interval between the arrivals of any two adjacent historical tasks based on the time data of each task arrival. The average task execution duration can be obtained by calculating the average task execution duration based on the execution durations of multiple historical tasks. The task execution duration variation characteristic can be calculated by calculating the task execution variance of multiple historical tasks based on the execution durations of multiple historical tasks.

[0048] Furthermore, in the process of determining the task flow characteristics above, in order to eliminate the influence of the dimension of data characteristics, it is necessary to perform feature normalization on the historical task data, that is, perform normalization processing on the task feature vectors corresponding to each historical task, so that different indicators are comparable. It is planned to adopt the zero-mean normalization method (also known as standard deviation standardization), and the mean of the processed data is 0 and the standard deviation is 1.

[0049] Further, embodiments of the present disclosure can, based on the results of normalization processing, respectively evaluate the correlation between each piece of historical task data and the target prediction variable, and determine the task flow characteristics according to the correlation. Specifically, the present disclosure can use the Pearson correlation coefficient to evaluate the linear correlation between a certain feature in the task feature vector and the target prediction variable (i.e., the task execution duration). The value of the correlation coefficient is between [-1, 1]. The closer the correlation coefficient is to 1 or -1, the stronger the correlation degree; the closer the correlation coefficient is to 0, the weaker the correlation degree. Through this method, features with weak correlation can be deleted to reduce storage space and computational overhead.

[0050] Here, the task flow characteristics extracted from the historical task data can be used, and after a new task is executed and becomes historical task data, the new historical task data can be added to the extraction of the task flow characteristics of the task. Therefore, the task flow characteristics extracted according to the historical task data gradually approach the true task flow characteristics.

[0051] In some embodiments, step S302 above may include the following steps:

[0052] First, for each task flow, determine the task flow category of the task flow according to the task flow characteristics.

[0053] Secondly, determine the target duration prediction model corresponding to the task flow according to the task flow category.

[0054] The present disclosure can preset multiple task flow categories, determine the task flow characteristics belonging to each category from multiple historical task flows, and further establish the association relationship between the task flow category and the task flow characteristics.

[0055] For the task flow being processed, the present disclosure can determine the characteristics of the task flow according to the historical data of the task flow, and the method for determining the task flow characteristics is as shown above. Then, the present disclosure can also determine the task flow category to which the task flow belongs according to the association relationship between the task flow category and the task flow characteristics.

[0056] In these embodiments, the duration prediction model corresponding to the task flow category can be set first. By determining the task flow category according to the task flow characteristics and then determining the target duration prediction model corresponding to the task flow according to the task flow category, the efficiency of determining the target duration prediction model corresponding to the task flow can be improved, thereby improving the efficiency of predicting the execution duration of the tasks to be executed in the task flow.

[0057] In some embodiments, determining the task flow category of the task flow according to the task flow characteristics includes:

[0058] Cluster multiple task flow characteristics using a preset clustering method and preset categories, and determine the task flow category of each task flow according to the clustering results.

[0059] In these embodiments, the above clustering method can be various clustering methods, such as the k-means clustering method, the mean shift clustering algorithm, the hierarchical clustering algorithm, etc.

[0060] The present disclosure can preset multiple categories in advance, and then use the above various clustering methods to cluster multiple task flows, and determine the categories corresponding to each task flow according to the clustering results.

[0061] As an implementation solution, the above clustering method can be the k-means clustering method.

[0062] For the recognition of task flow categories, the present disclosure can use the K-means clustering algorithm to cluster the historical data of multiple collected task flows. Specifically, using the k-means clustering method to cluster multiple task flows includes the following steps:

[0063] Step 1: First, preliminarily determine multiple categories of task flows according to the business scenario, suppose it is k categories.

[0064] Step 2: Randomly select k data points from the historical data set of the collected task flows as the cluster center points.

[0065] Step 3: For each point in the data set, calculate its Euclidean distance from each cluster center point, and divide it into the set to which the cluster center point it is closest belongs.

[0066] Step 4: After all the data are grouped, recalculate the cluster center point of each set.

[0067] Step 5: If the distance between the newly calculated cluster center point and the original cluster center point is less than a set threshold, it is considered that the clustering has tended to be stable and the algorithm terminates.

[0068] Step 6: If the distance between the new cluster center point and the original cluster center point changes greatly, then iterate steps 3-5.

[0069] The following is the recognition process of the task flow and the update of the cluster center point:

[0070] Suppose there are a total of k types of scenarios, n i (i ∈ 1,…,k) represents the total number of task flows belonging to the current i-th type of scenario, c i (i ∈ 1,…,k) is the cluster center point of the i-th type of scenario, and f represents the feature vector of the current task flow.

[0071] Calculate the Euclidean distance D between the feature of the current task flow and each cluster center point i = distance(f - c i ), select the value when D is the smallest, when Di <When T (a preset threshold) is met, the task flow is identified as the i-th type of scenario, and the formula c i =(n i ×c i +f) / (n i +1) is used to update the value of the cluster center point. When D≥T, that is, the Euclidean distance between the current task flow and the nearest cluster center point is greater than T, it indicates that the current task flow does not belong to any of the i-th type of scenarios and does not have obvious learnable features. For such task flows, they are uniformly classified into other scenarios, and the exponential smoothing algorithm is default used for prediction.

[0072] Clustering by K-means can quickly obtain multiple converging categories.

[0073] In some embodiments, determining the target duration prediction model corresponding to the task flow according to the task flow category includes:

[0074] For each task flow category, determining the task execution duration change characteristic value corresponding to the task flow category according to multiple task flows belonging to the task flow category;

[0075] Determining the target duration prediction model corresponding to the task flow according to the task execution duration change characteristic value.

[0076] In these embodiments, after obtaining multiple task flow categories, that is, after respectively determining the task flow categories to which multiple task flows belong, for each task flow category, the present disclosure can determine the task execution duration change characteristic value corresponding to the task flow category through multiple task flows belonging to the task flow category.

[0077] As an implementation manner, the above task execution duration change characteristic value may be the mean square deviation of the task execution duration. In this implementation manner, for each task flow category, the present disclosure can use the historical task data respectively corresponding to multiple task flows belonging to the task flow category to calculate the mean square deviation of the task execution duration of the task flow category.

[0078] The present disclosure can better analyze the distribution characteristics and rules of the task execution duration of the task flow category by using the historical task data respectively corresponding to multiple task flows belonging to the same task flow category, so as to better determine the change trend of the task execution duration of the task flow category.

[0079] As an implementation manner, the above determining the target duration prediction model corresponding to the task flow according to the task execution duration change characteristic value includes:

[0080] If the task execution duration change eigenvalue is less than or equal to the first preset threshold, the target duration prediction model corresponding to this task flow category is a linear fitting model;

[0081] If the task execution duration change eigenvalue is greater than the first preset threshold, the target duration prediction model corresponding to this task flow category is a long short-term memory model.

[0082] Taking the task execution duration mean square deviation of the task execution duration change eigenvalue as an example, the first preset threshold can be set in advance. If the task execution duration mean square deviation of a task flow category is less than or equal to the first preset threshold, the target duration prediction model corresponding to this task flow category is a linear fitting model. If the task execution duration mean square deviation of a task flow category is greater than the first preset threshold, the target duration prediction model corresponding to this task flow category is a long short-term memory model.

[0083] When the task execution duration change eigenvalue is less than or equal to the first preset threshold, the execution durations of multiple tasks in the task flow change gently, and a linear fitting model can be used to predict the execution durations of tasks in this task flow. Otherwise, a long short-term memory model is used to predict the execution durations of tasks in this task flow, and the predicted execution durations of tasks to be executed in each task flow can be determined quickly and accurately.

[0084] In some embodiments, the task processing method of the present disclosure further includes the following steps:

[0085] First, for the first task flow category with the task execution duration change eigenvalue greater than the first preset threshold, according to the magnitudes of the task execution duration eigenvalues corresponding to multiple first task flow categories, the multiple first task flow categories are divided into at least one task flow category set, where each task flow category set includes at least one first task flow category, and each task flow category set corresponds to a first long short-term memory model trained using the historical task flow data corresponding to this task flow category set.

[0086] Second, for each task flow, determine the first model parameter of this task flow according to the historical task data of this task flow, and use the first model parameter of this task flow to fine-tune the first long short-term memory model corresponding to the task flow category to which this task flow belongs, to obtain the second long short-term memory model of this task flow.

[0087] For the first task flow category with the task execution duration change eigenvalue greater than the first preset threshold, the range of the task execution duration change eigenvalue may be relatively large. If the same LSTM model is used to predict the execution durations of tasks in each first task flow category, problems such as long calculation time and inaccurate execution duration prediction may occur.

[0088] In these embodiments, in order to reduce the computational complexity of the LSTM model and improve the accuracy of predicting the execution duration of tasks by the LSTM model, the present disclosure may further classify multiple first task flow categories according to the magnitude of the task execution duration change feature value to obtain at least one task flow category set.

[0089] For example, a second preset threshold, a third preset threshold, and a fourth preset threshold may be set.

[0090] If the task execution duration change feature value corresponding to a task flow category is greater than the first preset threshold and less than or equal to the second preset threshold, the scheduling module may determine that this task flow category belongs to the first task flow category set; if the task execution duration change feature value corresponding to a task flow category is greater than the second preset threshold and less than or equal to the third preset threshold, the scheduling module may determine that this task flow category belongs to the second task flow category set; if the task execution duration change feature value corresponding to a task flow category is greater than the third preset threshold and less than or equal to the fourth preset threshold, the scheduling module may determine that this task flow category belongs to the third task flow category set.

[0091] In some application scenarios, the task execution duration change feature value may be the variance of the task execution duration.

[0092] In this implementation manner, the present disclosure may pre-use the historical task flow data corresponding to the task flow category set for different task flow category sets to train the first LSTM model corresponding to the task flow category set. That is to say, for each task flow category set, the corresponding LSTM model of this set may be pre-trained. The LSTM model corresponding to each task flow category set has a higher matching degree with each task flow category in this set.

[0093] In order to further improve the accuracy of the LSTM model in predicting the execution duration of tasks in each task flow, for each task flow in each task flow category set, before using the first LSTM model corresponding to the task flow category set to predict the execution duration of the to-be-executed task in this task flow, the model parameters of this task flow may be determined according to the historical task data of this task flow, and the above first LSTM model may be fine-tuned using the above model parameters to obtain the second LSTM model of this task flow.

[0094] The above model parameters include, but are not limited to, one or more of the following: the length of the periodic task sequence, the preset number of features, the number of model layers, etc.

[0095] The length of the periodic task sequence of a task flow may be obtained by the scheduling module through periodic feature statistics of the task features of multiple historical tasks of this task flow.

[0096] LSTM captures dependencies in sequential data. Therefore, it is necessary to determine the length of the sequence provided to the LSTM model, that is, how far back to trace valuable information. If the input sequence is too long, the amount of information that the LSTM model needs to process will become very large, resulting in longer inference time and increased memory consumption. In addition, if the input sequence contains a large amount of irrelevant information, the LSTM model may learn some useless features, thereby reducing its prediction ability. If the input sequence is too short, the LSTM model may not be able to capture long-term dependencies in the sequence, thus affecting its prediction ability.

[0097] For each task flow, the scheduling module can determine the sequence length based on the task characteristics of multiple historical tasks in the task flow. For example, it can statistically analyze the periodicity based on the task characteristics of multiple historical tasks in the task flow. Correspondingly, the execution duration of the tasks will also exhibit periodicity.

[0098] As an exemplary illustration, for the face-swiping payment task flow, the tasks in this task flow are cycled periodically as follows: (1) image acquisition, (2) image preprocessing, (3) object recognition for the preprocessed image, and (4) payment process handling. Since the tasks in this task flow periodically execute the above four steps, the task characteristics are periodic, and the execution durations of the tasks in the task flow are also periodic. The present disclosure can determine the task sequence length of the LSTM model corresponding to this task flow based on the periodicity of the execution durations of the tasks in the task flow.

[0099] As an exemplary illustration, the execution durations of the tasks in a task flow are as follows: 4us -> 4us -> 4us -> 10ms -> 4us -> 4us -> 4us > 10ms - 4us -> 4us -> 4us -> 10ms, then its period is 4. The present disclosure can select the task sequence length of the LSTM model corresponding to this task flow to be slightly greater than 4, for example, the task sequence length is 5. If the selected task sequence length input to the LSTM model is too long, overfitting will occur.

[0100] In addition, there are two parameters that have a great impact on the complexity and accuracy of the LSTM model, namely the number of LSTM layers and the number of features in each layer. Increasing the number of features has a positive impact on accuracy until a plateau is reached, after which a significant degradation will be observed. Since the LSTM model operates on sequential data, the addition of LSTM layers allows the hidden states at each level to operate at different time scales. However, adding too many LSTM layers will reduce the performance of the network. Once the network starts to converge, adding more layers will cause its accuracy to saturate and then rapidly decline.

[0101] The present disclosure can determine the number of layers and the number of features of the LSTM model according to the required prediction accuracy. For example, the number of layers of the LSTM is less than or equal to 4, and the number of features in each layer is set between 128 and 256. Specifically, the number of layers of the LSTM and the number of features in each layer can be set according to specific application scenarios, and no limitation is imposed here.

[0102] In some embodiments, the above task duration prediction method further includes the following steps:

[0103] First, in response to the target duration prediction model corresponding to the task flow being a linear fitting function model, determine the parameters of the linear fitting function of the task flow according to the task characteristics of multiple historical tasks of the task flow, and obtain the target linear fitting function model.

[0104] Second, based on the target linear fitting function model, determine the predicted execution duration of the task to be processed in the task flow.

[0105] Schematically, when it is determined that the stability and regularity are good according to the task characteristics of the historical tasks of the task flow (for example, the change in task execution duration is gentle), the present disclosure uses a linear fitting function model as the target duration prediction model of the task flow.

[0106] The function involved in the linear fitting function model may include an exponential function, and the exponential function may further include multiple parameters. The scheduling module can determine the parameters of the linear fitting function according to multiple historical task data of the task flow. Furthermore, the target linear fitting function model is determined. The linear fitting function model may be a linear fitting function model with the execution duration of historical tasks as the dependent variable and other task characteristics of historical tasks as the independent variable.

[0107] After determining the target linear fitting function model, the scheduling module can input the task characteristics of the task to be processed except for time (at this time, there is no execution duration yet) into the target linear fitting function model (for example, transfer the task characteristics of the task to be processed through a preset parameter passing interface). The predicted execution duration of the task to be processed can be generated by the linear fitting function model.

[0108] In these embodiments, for a task flow with good task characteristic stability, in order to reduce the occupation of AI chip hardware resources and reduce the prediction duration, the CPU core can perform linear fitting on the historical data in the task flow, and determine the predicted execution duration of the task to be executed according to the obtained linear fitting function model.

[0109] Please refer to Figure 4 , which shows a schematic flowchart of the task scheduling method provided by the present disclosure.

[0110] As Figure 4 shown, the method includes the following steps:

[0111] S401. Receive the to-be-processed tasks in at least two task flows, and determine the task flow characteristics corresponding to each of the at least two task flows, where the task flow characteristics are determined according to the historical task data of the task flow.

[0112] S402. For each task flow, determine the target duration prediction model corresponding to the task flow according to the task flow characteristics of the task flow. The target duration prediction model includes a linear fitting model or a long short-term memory model.

[0113] S403. In response to the target duration prediction model corresponding to the task flow being the corresponding long short-term memory model, determine the predicted execution duration of the to-be-processed tasks in the task flow based on the long short-term memory model.

[0114] For the specific implementation of the above steps S401 to S403, reference can be made to Figure 2 the description of the embodiments shown, which will not be elaborated here.

[0115] S404. For each task flow, determine the cumulative execution duration of the task flow according to the historical cumulative execution duration of the executed tasks in the task flow and the predicted execution duration of the to-be-processed tasks.

[0116] S405. Determine the execution priorities of the to-be-processed tasks according to the cumulative execution durations corresponding to the respective task flows and the balance of the computing cores used by the task flows; allocate the computing cores in the artificial intelligence chip to the to-be-processed tasks according to the execution priorities.

[0117] In this embodiment, the execution entity of step S405 and step S406 can be the scheduling module in the AI chip.

[0118] Taking the task flows including task flow A and B as an example, task flow A includes tasks A1, A2, A3, A4, A5, and A6. Task flow B includes tasks B1, B2, B3, B4, B5, and B6.

[0119] Among them, A1 to A5 in task flow A are completed tasks. B1 to B5 in task flow B are completed tasks. The scheduling module determines the task flow characteristic AP of task flow A according to the completed tasks A1 to A5 in task flow A. The scheduling module determines that the task flow characteristic AP belongs to the first category of task flows, and determines the linear fitting function model according to the task flow characteristic AP. Use the task characteristics of A1 to A5 to adjust the parameters of the linear fitting function model to obtain the target linear fitting function model AL. Input the task characteristic of A6 into the target linear fitting function model AL, and the target linear fitting function model AL outputs the predicted execution duration TA6' of task A6.

[0120] The scheduling module can extract the actual task execution durations TA1, TA2, TA3, TA4, TA5 from the task characteristics of A1, A2, A3, A4, A5. The scheduling module calculates the cumulative execution duration of task flow A: T1 = TA1 + TA2 + TA3 + TA4 + TA5 + TA6’. The scheduling module determines the task flow characteristics BP of task flow B based on the completed tasks B1 to B5 in task flow B. The scheduling module determines that the task flow characteristics BP belong to the second category of task flows, and determines the LSTM model (pre-trained LSTM model) based on the task flow characteristics BP. The parameters of the pre-trained LSTM model are adjusted using the task characteristics of B1 to B5 to obtain the target LSTM model BL. The task characteristics of B6 are input into the target LSTM model BL, and the predicted execution duration TB6’ of task B6 is output by the target LSTM model BL. The scheduling module can extract the actual task execution durations TB1, TB2, TB3, TB4, TB5 from the task characteristics of B1, B2, B3, B4, B5. The scheduling module calculates the cumulative execution duration of task flow B: T2 = TB1 + TB2 + TB3 + TB4 + TB5 + TB6’.

[0121] Assume T1 > T2, the scheduling module determines that the priority of task B6 is higher than that of task A6. The computing cores are allocated to the pending tasks in each task flow according to the priority. The computing core can be preferentially allocated to task B6. Then the computing core is allocated to task A6.

[0122] Compared with Figure 2 the embodiment shown, this embodiment describes the steps of determining the priority for the pending tasks of each task flow according to the predicted execution duration of the pending tasks and the actual execution duration of the historical tasks, and allocating computing cores to the pending tasks of each task flow according to the priority. By determining the priority in the above manner and allocating the corresponding computing cores to the tasks in each task flow according to the above priority, it helps to achieve the balance of the hardware resources occupied by each task flow, and thus helps to achieve the rationality of task scheduling among task flows.

[0123] Please continue to refer to Figure 1 , Figure 1 where the AI chip 11 in Figure 3 or Figure 4Determine the predicted execution duration of the tasks to be executed in each task flow in the manner of the illustrated embodiment. Then, for each task flow, the above scheduling module may calculate the sum of the cumulative execution durations of the historical tasks of the task flow and the predicted execution duration corresponding to the task to be processed in the task flow as the cumulative execution duration of the task flow.

[0124] The above scheduling module may allocate respective corresponding computing cores to the tasks to be processed in different task processes according to the balance based on the cumulative execution durations of multiple task flows.

[0125] According to an embodiment of the present disclosure, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network through a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processing device, the above functions defined in the method of the embodiment of the present disclosure are performed.

[0126] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program (computer-executable instructions) that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0127] The above computer-readable medium can be included in the above electronic device; or it can exist separately without being assembled into the electronic device.

[0128] The above computer-readable medium carries one or more programs, and when the one or more programs are executed by the electronic device, the electronic device is caused to execute the method shown in the above embodiments.

[0129] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0130] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0131] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, by way of non-limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), system on a chip (SOC), complex programmable logic devices (CPLD), and the like.

[0132] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0133] The above specific embodiments do not constitute a limitation on the scope of protection of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present disclosure shall be included within the scope of protection of the present disclosure.

Claims

1. A task processing method, characterized in that, Including: Receiving pending tasks in at least two task flows, and determining task flow characteristics corresponding to each of the at least two task flows, where the task flow characteristics are determined according to historical task data of the task flows; For each task flow, determining a target duration prediction model corresponding to the task flow according to the task flow characteristics of the task flow, where the target duration prediction model includes a linear fitting model or a long short-term memory model; In response to the target duration prediction model corresponding to the task flow being a corresponding long short-term memory model, determining a predicted execution duration of the pending task in the task flow based on the long short-term memory model.

2. The method according to claim 1, characterized in that, The "For each task flow, determining a target duration prediction model corresponding to the task flow according to the task flow characteristics of the task flow" includes: For each task flow, determining the task flow category of the task flow according to the task flow characteristics; Determining a target duration prediction model corresponding to the task flow according to the task flow category.

3. The method according to claim 2, characterized in that, The "Determining the task flow category of the task flow according to the task flow characteristics" includes: Using a preset clustering method and preset categories to cluster each task flow characteristic to obtain a clustering result; Determining the task flow category of each task flow according to the clustering result.

4. The method according to claim 2, characterized in that, The "Determining a target duration prediction model corresponding to the task flow according to the task flow category" includes: For each task flow category, determining a task execution duration change characteristic value corresponding to the task flow category according to multiple task flows belonging to the task flow category; Determining a target duration prediction model corresponding to the task flow according to the task execution duration change characteristic value.

5. The method according to claim 4, characterized in that, The "Determining a target duration prediction model corresponding to the task flow according to the task execution duration change characteristic value" includes: If the task execution duration change characteristic value is less than or equal to a first preset threshold, determining that the target duration prediction model corresponding to the task flow category is a linear fitting model; If the task execution duration change characteristic value is greater than the first preset threshold, determining that the target duration prediction model corresponding to the task flow category is a long short-term memory model.

6. The method according to claim 3, characterized in that, The preset clustering method includes the K-means clustering method.

7. The method according to claim 5, characterized in that, The method further includes: For a first task flow category with a task execution duration change characteristic value greater than the first preset threshold, dividing multiple first task flow categories into at least one task flow category set according to the magnitudes of the task execution duration characteristic values corresponding to the multiple first task flow categories, where each task flow category set includes at least one first task flow category, and each task flow category set corresponds to a first long short-term memory model trained using the historical task flow data corresponding to the task flow category set; For each task flow, determining the model parameters of the task flow according to the historical task data of the task flow, and fine-tuning the first long short-term memory model corresponding to the task flow category to which the task flow belongs using the model parameters of the task flow to obtain a second long short-term memory model of the task flow.

8. The method according to claim 7, characterized in that, The "Determining a predicted execution duration of the pending task in the task flow based on the long short-term memory model" includes: For each task flow, inputting the data of the pending task in the task flow into the second long short-term memory model corresponding to the task flow to obtain the predicted execution duration of the pending task.

9. The method according to claim 7, characterized in that, The model parameters include one or more of the following: The length of the periodic task sequence, the number of preset features, and the number of model layers generate the first model parameters of the task flow.

10. The method according to claim 1, characterized in that, The method further includes; In response to the target duration prediction model corresponding to the task flow being a linear fitting function model, determining the parameters of the linear fitting function of the task flow according to the task characteristics of multiple historical tasks of the task flow, and obtaining the target linear fitting function model; Determining the predicted execution duration of the task to be processed in the task flow based on the target linear fitting function model.

11. The method according to any one of claims 1 to 10, characterized in that, The method further includes: For each task flow, determining the cumulative execution duration of the task flow according to the historical cumulative execution duration of the executed tasks of the task flow and the predicted execution duration of the task to be processed; Determining the execution priorities of the tasks to be processed according to the cumulative execution durations corresponding to the respective task flows and the balance of the computing cores used by the task flows; Allocating the computing cores in the artificial intelligence chip to each task to be processed according to the execution priorities.

12. An artificial intelligence chip, the artificial intelligence chip is connected to the central processing unit of the host; the artificial intelligence chip includes a plurality of processor cores and a scheduling module, and the scheduling module is communicatively connected to each processor core; characterized in that, The scheduling module is used for: Receiving the tasks to be processed corresponding to at least two task flows sent by the central processing unit, and obtaining the task flow characteristics corresponding to the at least two task flows respectively, wherein the task flow characteristics of each task flow are determined according to the historical task data of the task flow; For each task flow, determining the target duration prediction model corresponding to the task flow according to the task flow characteristics of the task flow, and the target duration prediction model includes a linear fitting model or a long short-term memory model; In response to the target duration prediction model corresponding to the task flow being the corresponding long short-term memory model, determining the predicted execution duration of the task to be processed in the task flow based on the target long short-term memory model.

13. The artificial intelligence chip according to claim 12, characterized in that, The scheduling module is further used for: The scheduling module, for each task flow, calculates the cumulative execution duration of the task flow according to the actual task execution duration of the completed tasks of the task flow and the predicted execution duration of the task to be processed; And Based on the balance of the cumulative execution durations of multiple task flows, allocating the corresponding computing cores to each task to be processed.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the task processing method according to any one of claims 1 to 11 is implemented.