Method and device for predicting execution duration of tensor program, equipment and medium

By determining the abstract syntax tree of the tensor program and obtaining device features, predicting the execution time of the tensor program, solving the problems of inaccurate prediction time and low feature extraction efficiency in the prior art, and achieving efficient and accurate prediction across devices and models.

CN119987777APending Publication Date: 2025-05-13DOUYIN GROUP (HK) LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311511143.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict the execution time of tensor programs on different devices and models, and feature extraction methods cannot efficiently process the features of abstract syntax trees, resulting in inefficient training and difficulty in generalizing predictions across devices and models.

Method used

By determining the abstract syntax tree of the tensor program, obtaining the quantitative representation and location of the calculation expression, and combining the hardware characteristics of the target device, predicting the execution time of the tensor program on the target device.

Benefits of technology

The efficiency and accuracy of predicting the execution time of tensor program is improved, the specific structure and internal characteristics of the tensor program are fully considered, and combined with device characteristics, the prediction ability across models and devices is realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987777A_ABST
    Figure CN119987777A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a method and device for predicting the execution duration of a tensor program, equipment and a medium. The method includes determining an abstract syntax tree of the tensor program, the abstract syntax tree including a computational expression of the tensor program. The method further includes determining, based on the abstract syntax tree, a quantized representation of the computational expression and a location of the computational expression in the abstract syntax tree. The method also includes obtaining device features related to hardware of the target device. The method further includes predicting an execution duration of the tensor program on the target device based on the quantized representation, the location, and the device characteristics. Through the method, the features of the tensor program and the features of the computing device can be efficiently processed, and the efficiency and accuracy of predicting the execution duration of the tensor program on the computing device are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field Embodiments of the present disclosure generally relate to the field of machine learning, and more specifically, to methods, devices, apparatuses, and media for predicting the execution time of a tensor program. Background Art With the availability of computing hardware and richer data resources, neural network models have begun to show outstanding performance in various tasks. Today, neural networks have become an important pillar in the field of machine learning and artificial intelligence. It has been widely used in many fields such as image and speech processing, natural language processing, recommendation systems, autonomous driving, and has achieved many remarkable results. The booming development of deep learning has also led to the development of hardware and software tools, enabling people to design, train, and deploy complex neural network models more efficiently. By running these models on computing devices, users can be provided with a variety of useful information. However, there are still many problems that need to be solved when running neural network models on computing devices. Summary of the invention Embodiments of the present disclosure provide a method, apparatus, device, and medium for predicting the execution time of a tensor program. According to a first aspect of the present disclosure, a method for predicting the execution time of a tensor program is provided. The method includes determining an abstract syntax tree of the tensor program, the abstract syntax tree including a computational expression of the tensor program. The method also includes determining, based on the abstract syntax tree, a quantized representation of the computational expression and a position of the computational expression in the abstract syntax tree. The method also includes obtaining device features related to the hardware of the target device. The method also includes predicting the execution time of the tensor program on the target device based on the quantized representation, position, and device features. In the second aspect of the present disclosure, a device for predicting the execution time of a tensor program is provided. The device includes an abstract syntax tree determination module, configured to determine the abstract syntax tree of the tensor program, the abstract syntax tree including the computational expression of the tensor program; a quantization representation and position determination module, configured to determine the quantization representation of the computational expression and the position of the computational expression in the abstract syntax tree based on the abstract syntax tree; a device feature acquisition module, configured to acquire device features related to the hardware of the target device; and a tensor program execution time prediction module, configured to predict the execution time of the tensor program on the target device based on the quantization representation, position and device features. In a third aspect of the present disclosure, an electronic device is provided, comprising at least one processor; and a storage device for storing at least one program, wherein when the at least one program is executed by the at least one processor, the at least one processor implements the method according to the first aspect of the present disclosure. In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method according to the first aspect of the present disclosure is implemented. It should be understood that the contents described in this content section are not intended to limit the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS The above and other objects, features and advantages of the present disclosure will become more apparent through a more detailed description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, wherein like reference numerals generally represent like components throughout the exemplary embodiments of the present disclosure. Figure 1 A schematic diagram illustrating an example environment in which the apparatus and / or method of embodiments of the present disclosure may be implemented; Figure 2 A schematic diagram for determining features of a tensor program according to an embodiment of the present disclosure is illustrated; Figure 3A A schematic diagram illustrating an example of distribution of all nodes of an abstract syntax tree according to an embodiment of the present disclosure; Figure 3B A schematic diagram illustrating an example of leaf node distribution of an abstract syntax tree according to an embodiment of the present disclosure; Figure 4 A schematic diagram illustrating an example method for predicting the execution time of a tensor program according to an embodiment of the present disclosure is illustrated; Figure 5 A schematic diagram illustrating an example framework of a prediction model for predicting the execution time of a tensor program according to an embodiment of the present disclosure is illustrated; Figure 6 A schematic diagram illustrating an example of processing data offset in predicting the execution time of a tensor program according to an embodiment of the present disclosure is illustrated; Figure 7 A schematic block diagram of an apparatus suitable for predicting the execution time of a tensor program is illustrated. Figure 8 A schematic block diagram of an example device suitable for implementing embodiments of the present disclosure is illustrated. In the various drawings, the same or corresponding reference numerals represent the same or corresponding parts. DETAILED DESCRIPTION It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of the relevant laws, regulations and relevant provisions. In response to receiving the user's active request, a prompt message is sent to the user to clearly prompt the user that the operation requested to be performed will require the acquisition and use of the user's personal information. Thus, the user can independently choose whether to provide personal information to the electronic device, application, server or storage medium or other software or hardware that performs the operation of the technical solution of this disclosure according to the prompt message. Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as being limited to the embodiments described herein, which are instead provided for a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not intended to limit the scope of protection of the present disclosure. In the description of the embodiments of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". Other explicit and implicit definitions may also be included below.

[0001] As mentioned above, there are still many problems to be solved when running neural network models on computing devices. For example, in the field of deep learning, in order to improve the performance and efficiency of models, researchers often need to predict the performance of different devices and models. However, due to the differences in devices and models, as well as a large number of computational graph structures, accurately predicting the performance of tensor programs on different devices and models is a challenging task, for example, predicting the execution time of a model on a computing device. In order to more accurately predict the performance of tensor programs, some methods for identifying the performance of tensor programs have been proposed in traditional solutions. For example, in some traditional solutions, the extreme gradient boosting method is used, which uses decision trees and gradient boosting to predict the performance of tensor programs. For example, the extreme gradient boosting method uses a set of predefined rules and features for performance prediction. However, this solution has some defects. First, these predefined rules and features cannot fully capture the differences between different devices and models. Second, the extreme gradient boosting method does not take into account the specific structure and internal characteristics of tensor programs. In other traditional solutions, the Tiramisu method is used, which uses the Abstract Syntax Tree (AST) to capture the internal structure of the tensor program and uses the features in the AST for performance prediction. However, this solution also has some defects. First, due to the complexity and diversity of AST, the Tiramisu method may encounter difficulties in processing different types of tensor programs. Second, the Tiramisu method does not solve the problem of performance differences between different devices. In general, traditional tensor program performance prediction schemes rely on trained deep neural network models, which leads to many limitations. First, the prediction models in traditional schemes often have difficulty accurately predicting the performance of tensor programs on multiple different devices and models; second, the feature extraction methods of traditional schemes cannot efficiently process the features of the abstract syntax tree (AST), resulting in low training efficiency. Third, in traditional schemes, the distribution differences between different devices and models make it difficult to generalize prediction models across devices and models. Fourth, the existing prediction model training objectives often tend to favor a certain metric and cannot balance the errors between different metrics. At least to solve the above and other potential problems, an embodiment of the present disclosure proposes a method for predicting the execution time of a tensor program. In this method, a computing device first determines the abstract syntax tree of the tensor program. The abstract syntax tree of the tensor program includes the calculation expression of the tensor program. The computing device also determines the quantized representation of the calculation expression and the position of the calculation expression in the abstract syntax tree based on the abstract syntax tree. In addition, the computing device also needs to obtain device characteristics related to the hardware of the target device. By utilizing the quantized representation, position and device characteristics of the target device of the calculation expression, the computing device can predict the execution time of the tensor program on the target device. Through this method, since the specific structure and internal characteristics of the tensor program are fully considered, all necessary information is encapsulated in a compact and regular structure, and in combination with the acquired device information, it is possible to efficiently process the characteristics of the tensor program and the characteristics of the computing device, thereby improving the efficiency and accuracy of predicting the execution time of the tensor program on the computing device. The embodiments of the present disclosure will be described in detail below in conjunction with the accompanying drawings. Figure 1An example environment in which the apparatus and / or method of embodiments of the present disclosure may be implemented is shown. In environment 100 , computing device 106 determines execution time 116 of tensor program 102 by processing abstract syntax tree 104 of tensor program 102 . Examples of computing device 106 include, but are not limited to, personal computers, server computers, handheld or laptop devices, mobile devices (such as mobile phones, personal digital assistants (PDAs), media players, etc.), multiprocessor systems, consumer electronics, minicomputers, mainframe computers, distributed computing environments that include any of the above systems or devices, and the like. like Figure 1 As shown, the tensor program 102 is a program to be executed related to the neural network model. In some embodiments, the tensor program is a program from an operator of the neural network model. For example, a tensor program is obtained by compiling an operator of the neural network model. Additionally, different tensor programs can be obtained by compiling the same operator using different compilers. In some embodiments, different operators have different tensor programs. The above examples are only used to describe the present disclosure, and are not specific limitations of the present disclosure. The tensor program 102 can be represented as an abstract syntax tree 104. The abstract syntax tree 104 may include the computation expressions of the tensor program 102, loop information related to the computation expressions, and the locations of the computation expressions. It is understood that Figure 1 The forms of the tensor program 102 and the abstract syntax tree 104 shown are merely examples and are not specific limitations of the present disclosure. The tensor program 102 and the abstract syntax tree 104 may be set to any suitable content. In some embodiments, the abstract syntax tree 104 includes a computation expression that describes a detailed computation type and memory mode. The abstract syntax tree also includes loop information related to the computation expression and the location of the computation expression in the abstract syntax tree. Figure 1 The computing device 106 is shown to receive the abstract syntax tree 104, for example, receiving the abstract syntax tree 104 from other computing devices or remote storage devices in the network, which is only an example and not a specific limitation of the present disclosure. The computing device 106 can also obtain the abstract syntax tree 104 from a local storage device. Additionally, the computing device 106 obtains the tensor program 102 from a local or other external device and then generates an abstract syntax tree locally. In some embodiments, when processing the abstract syntax tree 104, the computing device 106 determines a quantized representation 108 for the computational expression, such as a computational vector of the computational expression, based on the computational expression in the abstract syntax tree 104 and loop information related to the computational expression. The computing device 106 also determines a position 110 of the computational expression as a leaf node based on the abstract syntax tree 104. The position 110 is the position of the computational expression of the tensor program in the abstract syntax tree. For example, the abstract syntax tree of the tensor program is traversed to determine the position of the computational expression as a leaf node in the abstract syntax tree. In addition, when the computing device 106 processes the abstract syntax tree 104, it also obtains the device characteristics 112 of the target device 104 on which the tensor program is to be executed. The device characteristics 112 are the device characteristics of the target device on which the abstract syntax tree is to be used, such as device characteristics related to the hardware specifications of the target device. Therefore, there are quantized representations and locations related to the computational expressions in the abstract syntax tree 104 and device characteristics related to the target device in the computing device. Then, the computing device 106 determines the time required for the tensor program 102 to be executed on the target device 104 based on the above information. Through this method, when predicting the execution time of the tensor program on the target device, the specific structure and internal characteristics of the tensor program are fully considered, and all necessary information is encapsulated in a compact and regular structure, which improves the efficiency of predicting the execution time of the tensor program without losing any useful information. Combined with the above Figure 1 A schematic diagram of an example environment in which the device and / or method of the embodiments of the present disclosure may be implemented is described below. Figure 2 A schematic diagram is described for determining features of a tensor program according to an embodiment of the present disclosure. like Figure 2 As shown, Example 200 refers to Figure 1 The abstract syntax tree 104 in FIG. 104 is a specific example of an abstract syntax tree 104 from Figure 1 The specific example of the tensor program 102 shown in FIG. 1 is obtained. The abstract syntax tree includes four leaf nodes, namely, node 4, node 8, node 9, and node 10, each of which corresponds to a computational expression. The remaining nodes are used to describe loop information. The computing device traverses the abstract syntax tree to obtain an ordered list 202 describing the traversal order of each node. From the abstract syntax tree, vector or quantized representation 204 of the computational expressions A, B, C, and D for the four leaf nodes can be obtained. In one embodiment, the computing device performs a pre-order traversal of the abstract syntax tree to obtain an ordered list corresponding to the nodes in the abstract syntax tree. The computational expressions of the tensor program correspond to the leaf nodes in the abstract syntax tree, which describe the detailed computational type and memory mode. For the leaf nodes to be processed in the tensor program, a special marker (for example, using -1 as a special marker) is added after each leaf node in the ordered list 202 to capture the location of the computational expression (leaf node) and loop nesting information. For example, after leaf node 4 in the ordered list, a list item storing -1 is inserted. In addition, the index of each leaf node in the ordered list is recorded, and a sequential vector 206 for the leaf node or computational expression is generated from the ordered list according to the index. For example, the computational expression A corresponds to index 4 in the ordered list 202. In some embodiments, the ordered list is obtained by performing a post-order traversal on the abstract syntax tree. In some embodiments, the ordered list is obtained by performing an in-order traversal on the abstract syntax tree. In some embodiments, the computing device performs a pre-order traversal on the abstract syntax tree to generate an ordered list corresponding to the nodes in the abstract syntax tree. Then, the computing device forms a sequence vector based on the order of the leaf nodes or the computational expressions in the ordered list. The above examples are only used to describe the present disclosure, and are not specific limitations of the present disclosure. The sequence vector 206 and the quantized representation for the computation expression can form a compact abstract syntax tree. Then, the features of the tensor program can be further obtained based on the sequence vector 206 and the quantized representation for the computation expression. Compared with the original abstract syntax tree, the compact abstract syntax tree reduces the size of the features while retaining all factors that may affect the tensor program and are independent of the target device. Combined with the above Figure 2 A schematic diagram for determining an abstract syntax tree example in an embodiment of the present disclosure is described below. Figure 3A , 3B A schematic diagram of the distribution of all nodes of an abstract syntax tree and a schematic diagram of the distribution of leaf nodes of the abstract syntax tree according to an embodiment of the present disclosure are described. like Figure 3A As shown, in example 300A, the scatter plot represents the distribution of the number of nodes in the abstract syntax tree in the first data set, the horizontal axis represents the number of all node distributions, and the vertical axis represents the frequency of all node distributions. like Figure 3B As shown in example 300B, the scatter plot represents the distribution of the number of leaf nodes of the abstract syntax tree in the first data set, the horizontal axis represents the number of all leaf nodes, and the vertical axis represents the frequency of all leaf nodes. Figure 3A and Figure 3BAs shown, the number of leaf nodes is significantly less than the number of nodes. Since the range of the number of leaf nodes is relatively small in tensor programs, compared with the original abstract syntax tree, a compact abstract syntax tree is used, for example, only information related to leaf nodes is used to represent tensor programs, so that the size of tensor program features is reduced while retaining all device-independent factors that may affect the performance of tensor programs. In some embodiments, after the computing device processes the abstract syntax tree 104, for example, after obtaining a compact abstract syntax tree, the distribution of leaf nodes in the abstract syntax tree becomes relatively scattered compared to the original relatively dense distribution due to the reduction of the size of the features. Combined with the above Figure 2 , Figure 3A , Figure 3B A schematic diagram for determining the characteristics of a tensor program according to an embodiment of the present disclosure is described. Figure 4 A schematic diagram describing an example method for predicting the execution time of a tensor program according to an embodiment of the present disclosure. Figure 4 The process shown can be found in Figure 1 The illustrated computing device 106 or any other suitable computing device may be used. At block 402, an abstract syntax tree of a tensor program is determined. The abstract syntax tree includes computational expressions of the tensor program. For example, Figure 1 The computing device 106 in the embodiment is used to obtain the abstract syntax tree 104, which includes leaf nodes describing the computational expressions and nodes describing loop information related to the computational expressions. The loop information includes the number of loops, the loop length (the iteration range of the loop variable), and the attributes of each loop (whether the loop is vectorized, unrolled, or parallelized). In some embodiments, when determining the abstract syntax tree, the computing device 106 first determines the computational expression of the leaf node. Then, the computing device 106 determines the computational type and memory mode of the computational expression. In some embodiments, the computing device 106 also determines that the loop information related to each computational expression includes the number of loops, the loop length (the iteration range of the loop variable), and the attributes of each loop (whether the loop is vectorized, expanded, or parallelized). The above examples are only used to describe the present disclosure, and are not specific limitations of the present disclosure. Those skilled in the art can also determine by any suitable method. At block 404 , based on the abstract syntax tree, a quantified representation of the computation expression and a position of the computation expression in the abstract syntax tree are determined. For example, computing device 106 determines quantified representation 108 for the computation expression and a position 110 of the computation expression in the abstract syntax tree according to abstract syntax tree 104 . In some embodiments, the computing device first determines the quantized representation of the computation expression using the computation type and memory mode of the computation expression and the loop information of the computation expression. The computing device 106 then determines the position of the computation expression in the abstract syntax tree according to the ordered list generated by traversing the abstract syntax tree. At block 406 , device characteristics related to the hardware of the target device are obtained. For example, the computing device 106 determines the device characteristics based on information related to the hardware specifications of the target device. In some embodiments, the hardware specifications of the target device include at least one of the following: clock frequency, memory bandwidth, number of computing cores, peak number of floating point operations per second FLOPS at different precisions, size of L1 cache, size of L2 cache, or memory size. In some embodiments, the computing device 106 obtains the hardware specifications of the target device by detecting the hardware of the target device. In some embodiments, the computing device 106 obtains the hardware specifications of the target device from the hardware configuration table of the target device. The above examples are only used to describe the present disclosure, not to specifically limit the present disclosure. Those skilled in the art may obtain any suitable hardware information in any suitable manner. At block 408, the execution time of the tensor program on the target device is predicted based on the quantized representation and location of the computation expression and device characteristics of the target device. For example, the computing device 106 predicts the execution time of the tensor program on the target device based on the quantized representation and location of the computation expression and device characteristics related to the hardware specifications of the target device. In some embodiments, the computing device may determine a computation vector or quantized representation of a computation expression based on the computation type and memory mode of the computation expression in the abstract syntax tree and the corresponding loop information. In some embodiments, the computing device determines the location of the computation expression based on an ordered list determined by traversing the abstract syntax tree. In some embodiments, the computing device obtains device features based on hardware specifications of the target device. In some embodiments, the computing device determines a position code corresponding to the computation expression according to the position of the computation expression and the length of the quantized representation of the computation expression. Then, the computing device predicts the execution time of the tensor program on the target device according to the position code, the quantized representation of the computation expression, and the device characteristics of the target device. In some embodiments, when predicting the execution time of a tensor program on a target device, the computing device obtains the execution time by applying the quantized representation, location, and device characteristics to the prediction model. In some embodiments, when predicting the execution time of a tensor program on a target device, the computing device obtains a mapping relationship between the quantized representation, location, and device characteristics and the execution time. The mapping relationship is then used to obtain the execution time corresponding to the quantized representation, location, and device characteristics of the calculation expression and the target device. The above examples are only used to describe the present disclosure and are not specific limitations of the present disclosure. In some embodiments, the computing device may also train the prediction model. When training the prediction model, the computing device first obtains the sample abstract syntax tree of the sample tensor program, and then obtains the sample quantization representation of the sample calculation expression and the sample position of the sample calculation expression in the sample abstract syntax tree from the sample abstract syntax tree. The computing device also obtains the sample device characteristics of the sample device to execute the sample tensor program, and the sample execution time of the sample tensor program on the sample device. Then, the computing device uses the above information to train the prediction model. In some embodiments, when the computing device trains the prediction model, the computing device may also fine-tune the prediction model according to the center distance difference CMD associated with the first group of sample devices and the second group of sample devices. In some embodiments, when a computing device predicts the execution time of a tensor program on a new target device, in order to ensure the accuracy of the prediction, the computing device determines whether the target device has been used to predict the execution time. For example, it determines whether the target device is a new device. If it is determined that the target device has not been used to predict the execution time, indicating that the target device is a new device, multiple tasks are selected from the task set corresponding to the tensor program through a clustering operation. The tensor programs corresponding to the multiple tasks are then executed using the target device to obtain sample data to fine-tune the prediction model. The fine-tuned prediction model is then used to process predictions related to the new target device. Through this method, by fully considering the specific structure and internal characteristics of the tensor program, all necessary information is encapsulated in a compact and regular structure, which improves the efficiency of predicting the execution time of the tensor program without losing any useful information. In further improvements, the scheme also introduces fine-tuning of the prediction model based on center distance differences, so that the method can be quickly fine-tuned on new devices and achieve more accurate performance predictions. And in this scheme, the characteristics of the tensor program and the target device are fully combined, so that the method can predict the execution time of the tensor program across models and devices. Combined with the above Figure 4 A schematic diagram of an example method for determining the trajectory of text in a tensor program according to an embodiment of the present disclosure is described. Figure 5A diagram depicting an example framework for determining the execution time of a tensor program according to an embodiment of the present disclosure. like Figure 5 As shown, example 500 adopts a Transformer-based predictor model framework to predict the execution time of a tensor program. Figure 5 The framework includes an encoder 502 . In one embodiment, when predicting the execution time of a tensor program, the computing device inputs the device-independent compact abstract syntax tree representation χ into the position encoder 504 for position encoding, wherein the representation includes a quantized representation of a computation expression and the position of the computation expression. Then, a fixed-length embedding Z is generated through a Transformer encoder 506 and a linear layer 510-1, 510-2, ... or 510-11 with a specific number of leaf nodes. χ . Figure 5 Eleven linear layers are shown, which are only examples, and any suitable linear layers can be set as needed. The computing device also inputs the device-related feature v into the device encoder 508, such as a multilayer perceptron (MLP) network, to calculate the embedding vector Z v ; Further embedding vector Z χ and the device-dependent embedding vector Z v Aggregation to generate embedding vectors It is then input into the regression layer 512 or the decoder to generate a predicted value of the execution time of the tensor program In some embodiments, when performing position coding, the computing device calculates the position coding of each leaf node (computational expression) using the sequential vector in the compact abstract syntax tree, and encodes the position of the leaf node in the original abstract syntax tree through a unique representation. Specifically, let the length of the calculation vector of each leaf node be N entry , V is a sequential vector, and the position encoding calculation of the εth leaf node is shown in the following formula (1): in, Represents the entry identifier in the output position encoding. θ is a user-defined scalar that affects the speed at which the frequency decreases along the vector dimension. The user can set it to an appropriate number based on actual needs. In some embodiments, during the training process, a cost model, typically based on machine learning, is trained to minimize the mean squared error (MSE). However, when the range of predicted values ​​(i.e., the latency or execution time of the tensor program) varies widely, the MSE often causes the model's predictions to be close to the mean of the performance distribution, underestimating high latencies and overestimating low latencies. The mean absolute percentage error (MAPE) measures the relative error (i.e., the average absolute percentage difference between the predicted values ​​and the actual values), and minimizing this training objective tends to produce large absolute errors. In order to strike a balance between absolute and relative errors, a scale-insensitive hybrid training objective is used that minimizes both MSE and MAPE. Specifically, in the pre-training of the prediction model (on the training set S train The loss function used in the above example is shown in formula (2): Among them, λ is a coefficient used to ensure that the MSE and MAPE terms have the same order of magnitude. Users can select the coefficient that performs better in the experiment, y i represents the sample value, The predicted value obtained from the sample input. In the training of the prediction model, in order to obtain better prediction performance in the target field, we use S train The prediction model is fine-tuned on the input features in the target domain. The goal of fine-tuning is to simultaneously minimize the mixing error and the distribution difference between the latent representation of the source domain (the output of the encoder) and the target domain. The distribution difference is measured using central moment discrepancy (CMD): The goal is to make the fine-tuned predictor perform well on different deep neural network models and devices. One of the most important theoretical results in domain-invariant learning is that the generalization risk of the model can be mitigated by reducing the distance between different domains in the latent space (i.e., the difference between the average error of the cost model evaluated on model M and device D and the average error of the cost model evaluated on model M′ and device D′). Assume h is an encoder that maps the input features to the latent space. With an appropriately chosen distribution difference measure Δ(·), the following bound on the generalization risk can be obtained by equation (3), for D, and M, in is a collection of devices. is a collection of models: in, Represents the predicted loss value when the input of model M is x and the output is y on the target device D The average error, Represents the predicted loss value when the input of model M′ is x and the output is y on the target device D′ This formula indicates that in the latent space, the distance between the corresponding representation distributions can be used to limit the system to any pair of and Based on this distribution difference-based bound, the following objective can be minimized during model fine-tuning to project samples from different domains to close locations in the representation space, as shown in Equation (4): in represents the standard training loss, represents a regularization aimed at minimizing the representation divergence. The actual divergence measure Δ(·) is the Central Moment Discrepancy (CMD). CMD is theoretically well-founded, achievable and computationally efficient, and has excellent empirical success in learning domain-invariant representations. Given two distributions and The CMD distance can be defined as follows (5): Where a, b are distributions and The joint distribution of supports, is the jth order moment. In practice, usually only a limited number of moments are needed (e.g., j ≤ 5). Finally, the training objective for fine-tuning the predictor is obtained as shown in Equation (6): where z s and z t are the latent representations of the source domain (e.g., a set of devices for cross-device performance prediction (CDPP)) and the target domain (e.g., a set of devices that do not overlap with the source devices), respectively. α is the coefficient determined by the auto-tuner. Furthermore, in order to achieve fast fine-tuning on new devices and to achieve accurate performance prediction, the tensor program that best represents the entire dataset should be selected and performance analyzed on the target device. Since tensor programs for different devices may not be exactly the same (for example, a tensor program for a GPU cannot be run directly on a CPU), representative tasks (rather than tensor programs) are selected and the corresponding tensor programs of these tasks are analyzed on the target device. Consider the same set of tensor programs on different devices. tasks, where tasks correspond to operators in the network model. Each task τ in has a set of device-independent features X τ (including the characteristics of its tensor program) and the corresponding latent representation Z of the tensor program τ .remember for The goal is to determine a subset of k tasks It is profiled on the target device such that the distribution difference between the latent representation of the selected task and the latent representation of all tasks is minimized. In some embodiments, X is the feature space of all possible tensor programs, is the latent / embedding space. is the feature set for the fine-tuning task. is the closest to x in the latent space In c, It captures the maximum distance between any tensor program in the input space and the closest tensor program in the fine-tuning sample in the latent space. Therefore, a lower bound on the generalization risk of the fine-tuned model is obtained as shown in Equation (7): In order to minimize ∈ and thus reduce the generalization risk, a clustering-based sampling strategy is proposed. First, K-means clustering is performed on all tensor program features in X and they are divided into κ clusters. and sort them according to the size of the clusters. Then, we compute a table Ψ, whose entries are The feature X that represents the task τ τ The average distance to the center of cluster e. Starting from the cluster with the largest cluster size, select a task for each cluster for performance analysis and remove it from the candidate task set after selection. Specifically, for cluster e, e = 1, 2, ..., κ, select the task with the smallest Ψ[e, τ] value from the candidate set. In some embodiments, the computing device needs to automatically select hyperparameters in the cost model. For example, an automatic adjuster is used to automatically select hyperparameters in the cost model. After testing different configurations, the user can select the best configuration that performs better in all experiments for automatic hyperparameter search. Figure 6A schematic diagram of an example of processing data offset in predicting the execution time of a tensor program according to an embodiment of the present disclosure is illustrated. In example 600, in the process of predicting the execution time of a tensor program through the scheme of the present disclosure, as shown in box 602, there is a long-tail distribution of the execution time of the tensor program in the data set. However, this skewness may seriously affect the accurate prediction model. Figure 604 is the distribution of Y values ​​after BOX-COX conversion, Figure 606 is the distribution of Y values ​​after Yeo-Johnson conversion, and Figure 608 is the distribution of Y values ​​after Quantile conversion. In some embodiments, when the computing device evaluates the tensor program delay distribution, a power transformation is selected to map the non-normal probability distribution to be more similar to a Gaussian distribution to reduce the impact of outliers. Power transformation is a technique for mapping a non-normal probability distribution to be more similar to a Gaussian distribution. An example of a power transformation is the Box-Cox transformation, which fits the optimal parameters of the mapping by maximum likelihood estimation. Another example is the Yeo-Johnson transformation, which can handle negative and zero values. The Quantile transformation transforms variables into standard distributions, including uniform and normal distributions, and is non-parametric. Select one from representative standardization methods to make the data more standardized by evaluating the distribution of tensor program delays after applying each method. As shown in box 604, the Box-Cox transformation generates a more normal and symmetrical distribution with fewer outliers. Therefore, in order to correct the impact of data skew on the prediction model, a ready-made library sklearn can be used to estimate the optimal parameters of the BoxCox transformation based on the training data set, and the inverse Box-Cox transformation is applied to transform the delay back to the original space for error measurement. As shown in box 606, the Y value distribution after Yeo-Johnson transformation is also relatively reasonable. Box 608 is the Y value distribution after Quantile transformation, which can convert the variable into a normal distribution and is non-parametric. Figure 7 FIG. 1 is a schematic block diagram of a device for determining the trajectory of text in a tensor program according to an embodiment of the present disclosure. Figure 7 As shown, the apparatus 700 includes an abstract syntax tree determination module 710, configured to determine an abstract syntax tree of a tensor program, the abstract syntax tree including a computational expression of the tensor program; a quantization representation and position determination module 720, configured to determine, based on the abstract syntax tree, a quantization representation of the computational expression and a position of the computational expression in the abstract syntax tree; a device feature acquisition module 730, configured to acquire device features related to the hardware of a target device; and a tensor program execution duration prediction module 740, configured to predict the execution duration of the tensor program on the target device based on the quantization representation, position and device features. In some embodiments, the quantization representation and position determination module 720 includes: a first loop information determination module, configured to determine a calculation expression and loop information for the calculation expression based on an abstract syntax tree; and a determination module based on loop information, configured to determine the quantization representation based on the calculation expression and the loop information. In some embodiments, a first loop information determination module includes: a leaf node-based expression determination module, configured to determine a calculation expression based on a leaf node of an abstract syntax tree; and a second loop information determination module, configured to determine loop information based on a non-leaf node of the abstract syntax tree. In some embodiments, the quantization representation and position determination module 720 includes: an ordered list determination module, configured to determine an ordered list corresponding to a node in the abstract syntax tree by traversing the abstract syntax tree; and a position determination module, configured to determine the position of a computational expression corresponding to a leaf node in the abstract syntax tree based on the ordered list. In some embodiments, the ordered list determination module includes: a leaf node determination module, configured to determine whether to traverse to a leaf node of an abstract syntax tree; and a tag storage module, configured to add a list item after the list item corresponding to the leaf node in the abstract syntax tree in response to traversing to the leaf node for storing a predetermined tag. In some embodiments, the position determination module includes: an index determination module configured to determine the index of a leaf node in the abstract syntax tree in the ordered list; and an index-based position determination module configured to determine the position of the calculation expression based on the index. In some embodiments, the device feature acquisition module 730 includes a specification determination module configured to determine the hardware specifications of the target device; and a feature determination module configured to determine device features based on the hardware specifications. In some embodiments, the hardware specifications include at least one of the following: clock frequency, memory bandwidth, number of computing cores, peak number of floating-point operations per second (FLOPS) at different precisions, size of the first-level cache L1, size of the second-level cache L2, or memory size. In some embodiments, the tensor program execution duration prediction module 740 includes: a position coding determination module, configured to determine the position coding corresponding to the calculation expression based on the position and the length of the quantized representation; and a duration determination module, configured to predict the execution duration of the tensor program on the target device based on the position coding, the quantized representation and the device characteristics. In some embodiments, the tensor program execution duration prediction module 740 includes: an execution duration determination module configured to obtain the execution duration by applying the quantized representation, location, and device characteristics to the prediction model. In some embodiments, the apparatus 700 further includes: a training module configured to train a prediction model based on sample quantized representations of sample computational expressions in a sample abstract syntax tree of a sample tensor program, sample positions of sample computational expressions in the sample abstract syntax tree, sample device characteristics of a sample device, and sample execution time of a sample tensor program on a sample device. In some embodiments, the training module includes: a first fine-tuning module configured to fine-tune the prediction model based on the center distance differences CMD associated with the first group of sample devices and the second group of sample devices. In some embodiments, the apparatus 700 further includes: a device usage determination module configured to determine whether the target device has been used for the prediction of execution duration; a selection module configured to select a plurality of tasks from a task set corresponding to the tensor program through a clustering operation in response to determining that the target device has not been used for the prediction of execution duration; and a second fine-tuning module configured to fine-tune the prediction model based on the tensor programs and target devices corresponding to the plurality of tasks. Figure 8 A schematic block diagram of an example device 800 that may be used to implement embodiments of the present disclosure is shown. Figure 1 The computing device 106 in the embodiment can be implemented by using the device 800. As shown in the figure, the device 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 902 or computer program instructions loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The CPU 801, the ROM 802, and the RAM 803 are connected to each other via a bus 1304. An input / output (I / O) interface 805 is also connected to the bus 804. A number of components in the device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage page 808, such as a disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks. The various processes and processing described above, such as method 400, may be performed by processing unit 801. For example, in some embodiments, method 400 may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by CPU 801, one or more actions of method 400 described above may be performed. The present disclosure may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present disclosure. Computer readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. Computer readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the above. More specific examples (non-exhaustive list) of computer readable storage medium include: portable computer disk, hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disk read-only memory (CD-ROM), digital versatile disk (DVD), memory stick, floppy disk, mechanical encoding device, such as a punch card or a convex structure in a groove on which instructions are stored, and any suitable combination of the above. The computer readable storage medium used here is not interpreted as a transient signal itself, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated by a waveguide or other transmission medium (e.g., a light pulse by an optical fiber cable), or an electrical signal transmitted by a wire. The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device. The computer program instructions for performing the operation of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed completely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be customized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure. Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions. These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram. Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram. The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of the module, program segment or instruction includes one or more executable instructions for realizing the specified logical function. In some alternative implementations, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of special hardware and computer instructions. The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or technical improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A method for predicting the execution time of a tensor program, comprising: Determine an abstract syntax tree of a tensor program, the abstract syntax tree including a computation expression of the tensor program; Based on the abstract syntax tree, determining a quantized representation of the computational expression and a position of the computational expression in the abstract syntax tree; Obtain device characteristics related to the target device's hardware; as well as An execution time of the tensor program on the target device is predicted based on the quantized representation, the location, and the device characteristics.

2. The method according to claim 1, wherein determining the quantized representation of the computation expression and the position of the computation expression in the abstract syntax tree comprises: Based on the abstract syntax tree, determining the computation expression and loop information for the computation expression; as well as The quantized representation is determined based on the calculation expression and the loop information.

3. The method according to claim 2, wherein determining the computation expression of the tensor program and loop information for the computation expression comprises: Determine the computation expression based on the leaf nodes of the abstract syntax tree; as well as The loop information is determined based on non-leaf nodes of the abstract syntax tree.

4. The method according to claim 1, wherein determining the quantized representation of the computation expression and the position of the computation expression in the abstract syntax tree comprises: Determining an ordered list corresponding to nodes in the abstract syntax tree by traversing the abstract syntax tree; as well as Based on the ordered list, the position of the computation expression corresponding to a leaf node in the abstract syntax tree is determined.

5. The method of claim 4, wherein determining the ordered list corresponding to the nodes in the abstract syntax tree by traversing the abstract syntax tree comprises: Determine whether the leaf node of the abstract syntax tree has been traversed; as well as In response to traversing to the leaf node in the abstract syntax tree, a list item is added after the list item corresponding to the leaf node for storing a predetermined tag.

6. The method of claim 4, wherein determining the position of the computation expression corresponding to a leaf node in the abstract syntax tree comprises: Determine the index of a leaf node in the abstract syntax tree in the ordered list; as well as The location of the calculation expression is determined based on the index.

7. The method according to claim 1, wherein obtaining device characteristics related to hardware of the target device comprises: Determining the hardware specifications of the target device; and Based on the specifications of the hardware, the device characteristics are determined.

8. The method according to claim 7, wherein the specification of the hardware comprises at least one of the following: Clock frequency, memory bandwidth, number of computing cores, peak number of floating-point operations per second (FLOPS) at different precisions, size of L1 cache, size of L2 cache, or memory size.

9. The method according to claim 1, wherein predicting the execution time of the tensor program on the target device comprises: determining a position code corresponding to the computational expression based on the position and the length of the quantized representation; as well as An execution time of the tensor program on the target device is predicted based on the position encoding, the quantized representation, and the device characteristics.

10. The method according to claim 1, wherein predicting the execution time of the tensor program on the target device comprises: The execution time is obtained by applying the quantitative representation, the location, and the device characteristics to a prediction model.

11. The method according to claim 10, further comprising: The prediction model is trained based on sample quantized representations of sample computation expressions in a sample abstract syntax tree of a sample tensor program, sample positions of the sample computation expressions in the sample abstract syntax tree, sample device characteristics of a sample device, and sample execution duration of the sample tensor program on the sample device.

12. The method of claim 11, wherein training the prediction model comprises: The prediction model is fine-tuned based on the center distance differences CMD associated with the first group of sample devices and the second group of sample devices.

13. The method according to claim 10, further comprising: determining whether the target device has been used for the prediction of the execution duration; In response to determining that the target device has not been used for the prediction of the execution duration, selecting a plurality of tasks from a set of tasks corresponding to a tensor program through a clustering operation; as well as The prediction model is fine-tuned based on the tensor programs corresponding to the plurality of tasks and the target device.

14. A device for predicting the execution time of a tensor program, comprising: An abstract syntax tree determination module is configured to determine an abstract syntax tree of a tensor program, wherein the abstract syntax tree includes a computation expression of the tensor program; A quantization representation and position determination module, configured to determine the quantization representation of the computation expression and the position of the computation expression in the abstract syntax tree based on the abstract syntax tree; A device feature acquisition module, configured to acquire device features related to hardware of a target device; as well as The tensor program execution time prediction module is configured to predict the execution time of the tensor program on the target device based on the quantized representation, the location and the device characteristics.

15. An electronic device, comprising: at least one processor; as well as A storage device, used to store at least one program, when the at least one program is executed by the at least one processor, so that the at least one processor implements the method according to any one of claims 1-13.

16. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 13 when executed by a processor.

Citation Information

Cited By

  • Neural network cost model evaluation method and system based on category guidance

    CN121092943A