A method and apparatus for generating a tensor program

By filtering the scheduling primitive to generate tensor programs that match the target hardware platform, the problem of insufficient adaptability of tensor programs in the prior art is solved, and more efficient hardware resource utilization is achieved.

CN115686530BActive Publication Date: 2025-07-08UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211418249.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-14
Publication Date
2025-07-08
Estimated Expiration
2042-11-14

AI Technical Summary

Technical Problem

The prior art is difficult to accurately determine tensor programs suitable for current hardware platforms, resulting in inefficient processing and waste of resources.

Method used

By obtaining the scheduling primitives set of the calculation graph, the target scheduling primitives are filtered based on the target performance indicators and hardware feature information, and a tensor program matching the target hardware platform is generated.

Benefits of technology

It improves the processing efficiency and accuracy of tensor programs on the target hardware platform, reduces resource usage, and improves the accuracy of the generation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115686530B_ABST
    Figure CN115686530B_ABST
Patent Text Reader

Abstract

The present application discloses a method and apparatus for generating a tensor program. The method includes: obtaining a computational graph representing a target algorithm model; determining a set of scheduling primitives for a computational subgraph corresponding to the computational graph; screening for target scheduling primitives from the set of scheduling primitives based on a target performance metric and hardware feature information of a target hardware platform to which the computational graph is to be applied; and generating a tensor program corresponding to the target scheduling primitives so that the tensor program is deployed on the target hardware platform. By screening the scheduling primitives and generating a tensor program for the screened scheduling primitives, the present application improves the processing efficiency and can determine the tensor program based on the hardware feature information of the target hardware platform, making it more compatible with the current target hardware platform and improving the accuracy of generating the tensor program for the current hardware platform.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method and device for generating a tensor program. Background Art

[0002] Currently, the algorithm models developed by developers can be abstracted into a computational graph representation and deployed on a hardware platform for application. In order to enable the computational graph to be better applied on the hardware platform, a tensor compiler is usually used to process the computational graph to obtain a tensor program, and the tensor program is deployed on the corresponding hardware platform to implement the application of the algorithm model developed by the developer.

[0003] During the process of the tensor compiler determining the tensor program finally deployed on the hardware platform based on the computational graph, a large number of tensor programs will be generated, and when selecting the tensor program with the optimal performance, at the same time, different hardware platforms correspond to different deployed tensor programs. Therefore, how to accurately determine the tensor program suitable for the current hardware platform has become an urgent problem to be solved. Summary of the Invention

[0004] In view of the above problems, this application provides a method and device for generating a tensor program, which improves the accuracy of generating the tensor program for the current hardware platform.

[0005] To achieve the above object, this application provides the following technical solutions:

[0006] A method for generating a tensor program, the method includes:

[0007] Obtain a computational graph representing a target algorithm model;

[0008] Determine a set of scheduling primitives for a computational subgraph corresponding to the computational graph;

[0009] Based on a target performance metric and hardware feature information of a target hardware platform to which the computational graph is to be applied, screen out a target scheduling primitive from the set of scheduling primitives;

[0010] Generate a tensor program corresponding to the target scheduling primitive so that the tensor program is deployed on the target hardware platform.

[0011] Optionally, the determining a set of scheduling primitives for a computational subgraph corresponding to the computational graph includes:

[0012] Based on the output connection relationship of operator nodes in the computational graph, determine target operator nodes among the operator nodes;

[0013] Using the target operator node as a splitting point, split the computational graph into several computational subgraphs;

[0014] Obtain a set of scheduling primitive instructions corresponding to the computational subgraph, and combine the instructions in the set of scheduling primitive instructions to obtain multiple scheduling primitives;

[0015] Based on the multiple scheduling primitives, generate a set of scheduling primitives corresponding to the computational subgraph.

[0016] Optionally, the screening of the target scheduling primitive from the scheduling primitives based on the target performance metric and the hardware feature information of the target hardware platform to which the computational graph is to be applied includes:

[0017] Generate a scheduling primitive feature vector corresponding to each scheduling primitive;

[0018] Input each scheduling primitive feature vector into a target processing model to obtain a prediction score output by the target processing model that matches each scheduling primitive feature vector, where the target processing model is a machine learning model trained based on the target performance metric corresponding to each scheduling primitive as the training target value and the model learning structure corresponding to the hardware feature information of each hardware platform;

[0019] Based on the prediction scores, determine a target scheduling primitive feature vector among the scheduling primitive feature vectors;

[0020] Obtain a target scheduling primitive that matches the target scheduling primitive feature vector.

[0021] Optionally, the generating of a scheduling primitive feature vector corresponding to each scheduling primitive includes:

[0022] Extract features from each scheduling primitive to obtain target features, where the target features include primitive characteristics, numerical parameters, and character parameters;

[0023] Based on a feature vector mapping table, determine vector information corresponding to each target feature;

[0024] Generate a scheduling primitive feature vector according to the vector information corresponding to each target feature.

[0025] Optionally, the method further includes:

[0026] Obtain training data, where the training data includes scheduling primitive feature vectors corresponding to scheduling primitives and labeled target performance metric labels for each hardware platform;

[0027] Divide the training data into a training set and a test set;

[0028] Training the initial model based on the training data to obtain a trained initial model, where the model structure of the initial model includes a basic learning layer and a hardware learning layer for learning the hardware characteristics of each processing platform;

[0029] Testing the trained model according to the test set, and adjusting the model parameters of the trained initial model according to the test results to obtain a target processing model.

[0030] Optionally, the testing the trained initial model according to the test set, and adjusting the model parameters of the trained initial model according to the test results to obtain a target processing model includes:

[0031] Inputting the test set into the trained initial model to obtain the prediction results output by the trained initial model;

[0032] Comparing the prediction results with the target performance indicators marked in the test set, if the comparison result does not meet the preset conditions, determining the hardware platform corresponding to the current test set;

[0033] Adjusting the parameters of the hardware learning layer corresponding to the hardware platform based on the comparison results, and adjusting the parameters of the basic learning layer based on the test results of all the test sets;

[0034] Determining the trained initial model with adjusted parameters as the target processing model.

[0035] Optionally, the determining the target scheduling primitive feature vector among each of the scheduling primitive feature vectors based on the prediction score includes:

[0036] Based on the prediction score, obtaining the prediction score corresponding to the target hardware platform;

[0037] According to the prediction score corresponding to the target hardware platform, determining the target scheduling primitive feature vector among each of the scheduling primitive feature vectors.

[0038] A device for generating a tensor program, the device includes:

[0039] An acquisition unit, configured to acquire a computation graph representing a target algorithm model;

[0040] A determination unit, configured to determine a set of scheduling primitives of a computation sub-graph corresponding to the computation graph;

[0041] A screening unit, configured to screen out target scheduling primitives from the set of scheduling primitives based on a target performance indicator and hardware feature information of a target hardware platform to which the computation graph is to be applied;

[0042] A generation unit for generating a tensor program corresponding to the target scheduling primitive, so that the tensor program can be deployed on the target hardware platform.

[0043] Optionally, the determination unit includes:

[0044] A first determination subunit for determining a target operator node among the operator nodes based on the output connection relationship of the operator nodes in the computation graph;

[0045] A splitting subunit for splitting the computation graph into several computation subgraphs with the target operator node as the splitting point;

[0046] A combination subunit for obtaining a set of scheduling primitive instructions corresponding to the computation subgraph, combining the instructions in the set of scheduling primitive instructions, and obtaining multiple scheduling primitives;

[0047] A first generation subunit for generating a set of scheduling primitives corresponding to the computation subgraph based on the multiple scheduling primitives.

[0048] Optionally, the screening unit includes:

[0049] A second generation subunit for generating a scheduling primitive feature vector corresponding to each scheduling primitive;

[0050] A model processing subunit for inputting each scheduling primitive feature vector into a target processing model to obtain a prediction score output by the target processing model that matches each scheduling primitive feature vector, where the target processing model is a machine learning model trained based on the target performance index corresponding to each scheduling primitive as the training target value and the model learning structure corresponding to the hardware feature information of each hardware platform;

[0051] A second determination subunit for determining a target scheduling primitive feature vector among the scheduling primitive feature vectors based on the prediction score;

[0052] A first obtaining subunit for obtaining a target scheduling primitive that matches the target scheduling primitive feature vector.

[0053] Optionally, the second generation subunit is specifically used for:

[0054] Performing feature extraction on each scheduling primitive to obtain target features, where the target features include primitive characteristics, numerical parameters, and character parameters;

[0055] Determining vector information corresponding to each target feature based on a feature vector mapping table;

[0056] Generating a scheduling primitive feature vector according to the vector information corresponding to each target feature.

[0057] Optionally, the apparatus further includes:

[0058] A data acquisition unit, configured to acquire training data, where the training data includes a scheduling primitive feature vector corresponding to a scheduling primitive and a labeled target performance metric label for each hardware platform;

[0059] A partitioning unit, configured to partition the training data into a training set and a test set;

[0060] A training unit, configured to train an initial model based on the training data to obtain a trained initial model, where the model structure of the initial model includes a basic learning layer and a hardware learning layer for learning the hardware features of each processing platform;

[0061] An adjustment unit, configured to test the trained model according to the test set and adjust the model parameters of the trained initial model according to the test results to obtain a target processing model.

[0062] Optionally, the adjustment unit includes:

[0063] An input subunit, configured to input the test set into the trained initial model to obtain a prediction result output by the trained initial model;

[0064] A comparison subunit, configured to compare the prediction result with the target performance metric labeled in the test set, and if the comparison result does not meet a preset condition, determine the hardware platform corresponding to the current test set;

[0065] An adjustment subunit, configured to adjust the parameters of the hardware learning layer corresponding to the hardware platform based on the comparison result and adjust the parameters of the basic learning layer based on the test results of all the test sets;

[0066] A third determination subunit, configured to determine the trained initial model with adjusted parameters as the target processing model.

[0067] Optionally, the second determination subunit is specifically configured to:

[0068] Based on the prediction score, obtain the prediction score corresponding to the target hardware platform;

[0069] According to the prediction score corresponding to the target hardware platform, determine a target scheduling primitive feature vector among all the scheduling primitive feature vectors.

[0070] A storage medium stores executable instructions, and when the instructions are executed by a processor, the method for generating a tensor program as described in any one of the above is implemented.

[0071] An electronic device, comprising:

[0072] A memory for storing programs;

[0073] A processor for executing the program, and the program is specifically used to implement the method for generating a tensor program as described in any one of the above.

[0074] Compared with the prior art, the present application provides a method and device for generating a tensor program. The method includes: obtaining a computational graph representing a target algorithm model; determining a set of scheduling primitives for computational subgraphs corresponding to the computational graph; based on a target performance metric and hardware feature information of a target hardware platform to which the computational graph is to be applied, screening out target scheduling primitives from the set of scheduling primitives; generating a tensor program corresponding to the target scheduling primitives, so that the tensor program is deployed on the target hardware platform. By screening the scheduling primitives, the present application generates a tensor program of the screened scheduling primitives, improving the processing efficiency, and being able to determine the tensor program based on the hardware feature information of the target hardware platform, making it more compatible with the current target hardware platform and improving the accuracy of generating the tensor program for the current hardware platform. Description of the Drawings

[0075] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.

[0076] Figure 1 It is a schematic flowchart of a method for generating a tensor program provided by an embodiment of the present application;

[0077] Figure 2 It is a schematic diagram of an operator node in a computational graph provided by an embodiment of the present application;

[0078] Figure 3 It is a schematic diagram of a feature extraction specification provided by an embodiment of the present application;

[0079] Figure 4 It is a schematic diagram of a feature extraction method provided by an embodiment of the present application;

[0080] Figure 5 It is a schematic diagram of a model structure applied to a target processing model provided by an embodiment of the present application;

[0081] Figure 6 It is a schematic diagram of a model structure for multi-task learning provided by an embodiment of the present application;

[0082] Figure 7 The structural schematic diagram of a tensor program generation device provided by an embodiment of the present application. Detailed implementation manners

[0083] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0084] The terms "first" and "second" etc. in the specification, claims and above-mentioned accompanying drawings of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may include unlisted steps or units.

[0085] In an embodiment of the present application, a method for generating a tensor program is provided. This method can be applied to a tensor compiler. The compilation process of machine learning can be regarded as the transformation between tensor functions. Therefore, it is necessary to generate a tensor program corresponding to the target algorithm model through a tensor compiler so that the target hardware platform can run the tensor program to implement the application of the target algorithm model. Refer to Figure 1 , which is the flowchart of a method for generating a tensor program provided by an embodiment of the present application. This method may include the following steps:

[0086] S101. Obtain a computational graph representing the target algorithm model.

[0087] The target algorithm model is a model to be applied to the target hardware platform. In the embodiments of the present application, the target algorithm model may include deep learning-related algorithm models and image processing-related algorithm models. Among them, deep learning-related algorithm models may include face recognition algorithms, image classification algorithms, speech recognition algorithms, target detection algorithms, etc.; image processing-related algorithm models may include blurring algorithms, edge detection algorithms, etc. Specifically, the computational graph representing the target algorithm model may be the computational graph of a face recognition algorithm or the computational graph of a speech recognition algorithm.

[0088] The computational graph representing the target algorithm model may include multiple operator nodes. Operator nodes are the basic computational units that make up a machine learning network. Operator nodes may be, for example, operations such as convolution and pooling in a neural network.

[0089] S102. Determine a set of scheduling primitives for the computational subgraph corresponding to the computational graph.

[0090] Since the computational graph includes multiple operator nodes, directly processing the computational graph will consume the processing and computing resources of the process and reduce the processing efficiency. Therefore, in the embodiments of the present application, it is necessary to split the computational graph to obtain computational subgraphs. Specifically, the computational graph can be split according to actual splitting requirements, or can be split according to the operator nodes in the computational graph. After obtaining the computational subgraphs corresponding to the computational graph, a set of scheduling primitives corresponding to the computational subgraphs can be determined. Among them, the set of scheduling primitives includes a large number of scheduling primitives, and the scheduling primitives can be used to clarify the specific execution process of the operators in the computational subgraph, so as to obtain the optimal operating efficiency.

[0091] In an implementation manner of the embodiments of the present application, the determining the set of scheduling primitives of the computational subgraph corresponding to the computational graph includes: determining a target operator node among the operator nodes based on the output connection relationship of the operator nodes in the computation; using the target operator node as a splitting point to split the computational graph into several computational subgraphs; obtaining a set of scheduling primitive instructions corresponding to the computational subgraphs, combining the instructions in the set of scheduling primitive instructions to obtain a plurality of scheduling primitives; and generating the set of scheduling primitives corresponding to the computational subgraph based on the plurality of scheduling primitives.

[0092] The output connection relationship of each operator node can be determined according to the node information of the operator node. Among them, the node information of the operator node can include the input connection relationship, output connection relationship, and related attribute information of the operator node, etc. The target node can be determined among the operator nodes according to the actual application requirements of the computational graph or the logical relationship of processing information. The target operator node can be used as a splitting point for splitting. Further, the output end of the target operator node can be used as a splitting point to split the computational graph into a plurality of computationally subgraphs connected in series.

[0093] See Figure 2 , which is a schematic diagram of an operator node in a computational graph provided by the embodiments of the present application. In Figure 2 the shown computational graph, there are operator nodes A, B, C, D, E, and F. The input and output connection relationships of each operator node are as Figure 2 shown. Assuming that D is used as the splitting point, then Figure 2 the dashed box shown is the computational subgraph corresponding to the computational graph. Since each computational subgraph constituting the computational graph has a certain input or output relationship, the processing of the computational graph can be realized by processing the computational subgraphs, thereby reducing the occupation of processing resources and improving the processing efficiency.

[0094] There are various instructions in the scheduling primitive instruction set. For example, instructions such as addition and subtraction representing calculations, and it can also include functions representing processing logic, such as the split() function for splitting strings. Scheduling primitive instructions in the scheduling primitive instruction set can be selected according to actual needs, and the scheduling primitive instructions can be combined according to the processing logic or a specific combination method to obtain multiple scheduling primitives, thereby obtaining a scheduling primitive set.

[0095] S103. Based on the target performance metrics and the hardware feature information of the target hardware platform to which the computational graph is to be applied, screen out the target scheduling primitive from the scheduling primitive set.

[0096] S104. Generate a tensor program corresponding to the target scheduling primitive so that the tensor program can be deployed on the target hardware platform.

[0097] Among them, the target performance metrics refer to parameters such as the execution efficiency and processing rate when applying the computational graph to the corresponding hardware platform. The target hardware platform refers to the hardware platform to which the computational graph is to be applied. Among them, the hardware feature information can include the basic attribute features of the hardware platform, such as the processor parameters and memory information of the hardware platform, and can also include the load parameters of the hardware platform, such as the number of currently executing tasks and the number of tasks to be executed. That is, different hardware platforms may correspond to different tensor programs to be set. Therefore, in the process of generating tensor programs in the embodiments of the present application, the hardware feature information of the target hardware platform also needs to be considered.

[0098] Correspondingly, in the embodiments of the present application, after obtaining the scheduling primitives, tensor programs corresponding to each scheduling primitive are not directly generated. Because the number of obtained scheduling primitives is large, if tensor programs are directly generated, the processing process is relatively complex, and it is necessary to screen out the tensor program with the highest matching degree with the target hardware platform from the generated multiple tensor programs, that is, it is necessary to screen out the tensor program with the highest running efficiency on the target hardware platform. However, tensor programs are nested tree-structured data, and it is difficult to automatically extract features, and often rely on manual feature extraction, and the feature functions are limited by the prior knowledge of professionals. Therefore, the screening of tensor programs is also relatively difficult. Therefore, in the embodiments of the present application, the scheduling primitives are first screened, and after obtaining the target scheduling primitive, a tensor program corresponding to the target scheduling primitive can be generated.

[0099] In order to further improve the processing efficiency, in the embodiments of the present application, the corresponding target scheduling primitive can be determined through a target processing model. That is, the screening out of the target scheduling primitive from the scheduling primitives based on the target performance metrics and the hardware feature information of the target hardware platform to which the computational graph is to be applied includes:

[0100] Generate a scheduling primitive feature vector corresponding to each scheduling primitive;

[0101] Input each scheduling primitive feature vector into the target processing model to obtain the prediction scores output by the target processing model that match each scheduling primitive feature vector;

[0102] Based on the prediction scores, determine the target scheduling primitive feature vector among each scheduling primitive feature vector;

[0103] Obtain the target scheduling primitive that matches the target scheduling primitive feature vector.

[0104] In the embodiments of the present application, feature extraction can be performed on the scheduling primitive to determine the scheduling primitive feature vector corresponding to each scheduling primitive. In one implementation manner, the generation of the scheduling primitive feature vector corresponding to each scheduling primitive includes: performing feature extraction on each scheduling primitive to obtain target features, where the target features include primitive features, numerical parameters, and character parameters; determining the vector information corresponding to each target feature based on the feature vector mapping table; and generating the scheduling primitive feature vector according to the vector information corresponding to each target feature.

[0105] Specifically, preprocess the scheduling primitive to obtain an abstract scheduling primitive sequence. For each abstract scheduling primitive in the sequence, feature extraction generally only retains three basic elements: primitive type, numerical parameter, and character parameter. Secondly, for the preprocessed abstract scheduling primitive sequence, feature extraction can be referred to Figure 3 the schematic diagram of the feature extraction specification shown, which describes the abstract identifier of the scheduling primitive. A scheduling primitive sequence (PrimitiveSequence) consists of several scheduling primitives (Primitive). Each primitive starts with a primitive type (PrimitiveType), followed by several numerical parameters (Number) or character parameters (NameParam). Specifically, it can be referred to Figure 4 the schematic diagram of the feature extraction method shown for feature extraction. The feature vector mapping table is a mapping table between features and vectors. For example, the primitive type can be represented by a specific digital vector, the numerical parameter can be represented by a specific digital vector, and the digital parameter can be directly represented by its own number. Specifically, for the primitive type, it is converted into a one-hot vector. A one-hot vector is a special vector in which only one bit is allowed to be 1 and the other bits must be 0. For the numerical parameter, its value remains unchanged. For the character parameter, they are uniformly converted into tokens encoded as ids. Then, all features are concatenated according to the original position of the elements. Further, post-processing can be performed on the extracted features through processing methods such as clipping, padding, and normalization, so as to generate the scheduling primitive feature vector.

[0106] The scheduling primitive feature vector serves as the input information for the target processing model. Among them, the target processing model is a machine learning model trained based on the target performance metrics corresponding to each scheduling primitive as the training target values, and the model learning structure corresponding to the hardware feature information of each hardware platform. It should be noted that the target performance metric refers to the processing performance metric determined based on historical operation information for applying the tensor program corresponding to the scheduling primitive to the hardware platform, such as the processing time for processing the same amount of data, or it can also be a latency parameter, etc. It should be noted that in order to facilitate the tensor model determined by the embodiments of the present application to meet the setting requirements of various current hardware platforms, when annotating the target performance metrics corresponding to each scheduling primitive, it can be annotated based on different hardware platforms. For example, A represents the scheduling primitive feature vector, and its annotation information B is a set of annotation information, including B1 and B2. Among them, B1 includes B11 representing the first hardware platform and B12 representing the second hardware platform, B2 includes B21 and B22, B21 represents the target performance metric label under the first platform, and B22 represents the target performance metric label under the second platform.

[0107] Therefore, in order to ensure the accurate learning of information by the target processing model, its training target values and model structure are improved based on learning features, that is, including the target performance metrics corresponding to each scheduling primitive as the training target values, and the model learning structure corresponding to the hardware feature information of each hardware platform. The prediction score output by the target processing model represents the prediction score of the target performance metric determined based on the input scheduling primitive feature vector, such as the score predicting its performance as the first type of performance or the score predicting its performance as the second type of performance.

[0108] Furthermore, in the embodiments of the present application, a method for generating a target processing model is also provided. The method may include: obtaining training data, where the training data includes the scheduling primitive feature vector corresponding to the scheduling primitive and the labeled target performance metric labels for each hardware platform; dividing the training data into a training set and a test set; training an initial model based on the training data to obtain the trained initial model, where the model structure of the initial model includes a basic learning layer and a hardware learning layer for learning the hardware features of each processing platform; testing the trained model according to the test set, and adjusting the model parameters of the trained initial model based on the test results to obtain the target processing model.

[0109] Further, testing the trained initial model according to the test set and adjusting the model parameters of the trained initial model according to the test results to obtain a target processing model, including: inputting the test set into the trained initial model to obtain a prediction result output by the trained initial model; comparing the prediction result with the target performance index labeled in the test set, if the comparison result does not meet the preset condition, determining the hardware platform corresponding to the current test set; adjusting the parameters of the hardware learning layer corresponding to the hardware platform based on the comparison result, and adjusting the parameters of the basic learning layer based on the test results of all the test sets; determining the trained initial model with adjusted parameters as the target processing model.

[0110] It should be noted that when dividing the training data in the embodiments of the present application, in addition to dividing the training data into a training set and a test set, the training data can also be divided into a training set, a validation set, and a test set, which can be specifically determined according to actual model training requirements and characteristics such as the number of samples and attributes. As long as it can be ensured that there is training data for training the model, verifying the accuracy of the model, and adjusting the model parameters based on the verification results in the training data, the present application does not limit the division of the training data and the specific application process.

[0111] See Figure 5 , which is a schematic diagram of a model structure applied to a target processing model provided by an embodiment of the present application. The model first upsamples the dimension to 256 or more through multiple linear layers (Linear). Then, an attention mechanism (Self-Attention) module is used to capture context features, followed by two residual blocks (Residual Block). Finally, multiple linear layers and a sum operation are used to obtain a prediction score. If the target performance index label is a latency label, and the target performance index label is a normalized latency, the formula is as follows:

[0112] label = (min_latency) / latency

[0113] Among them, latency can be the latency parameter of the current tensor program, and min_latency is the minimum value among the latency of all tensor programs of the computational subgraph. The reason for performing standardized latency processing on the target performance metric labels is to be able to process all target performance metrics in the same processing dimension, facilitating data processing. Since the training of the model is an iterative update process, that is, a process of repeatedly adjusting the model parameters to minimize the difference between the output of the trained model and the labels actually annotated in the sample data. Therefore, the prediction result is compared with the target performance metrics annotated in the test set, and the model parameters are adjusted according to the comparison result. At the same time, since the prediction result output by the target processing model in the embodiment of the present application is determined based on the hardware feature information of the corresponding hardware platform, therefore, the model structure of the target processing model in the embodiment of the present application at least includes a basic learning layer and a hardware learning layer. Therefore, when adjusting the model parameters, at least the model parameters of the basic learning layer and the model parameters of the hardware learning layer need to be adjusted. Specifically, the prediction result is compared with the target performance metrics annotated in the test set. If the comparison result does not meet the preset condition, the hardware platform corresponding to the current test set is determined, that is, if the deviation between the prediction result and the target performance metrics is large, it is necessary to determine the hardware platform annotated in the test set from which the prediction result is derived, so as to adjust the hardware learning layer corresponding to the hardware platform based on the comparison result corresponding to the test set. At the same time, the parameters of the basic learning layer can also be adjusted according to the comparison result. Further, the parameters of the basic learning layer can also be adjusted based on the test results of all the test sets.

[0114] Specifically, in the embodiment of the present application, the model structure that can use multi-task learning (MTL) technology to solve the problem of unavailable tensor programs across hardware platforms can be seen Figure 6 . For example, tensor programs are common on CPUs or GPUs. Therefore, after the scheduling primitive sequence is converted into features, it is the same on CPUs or GPUs. The difference is that they have different latencies. The present application sets a task for each hardware platform. The loss of the entire network is the sum of the losses of all tasks. Figure 6 Including an example during training, where Figure 6 The rectangular dotted box marked below can ensure the basic learning layer, Figure 6 The 3 rectangular dotted boxes marked side by side above can represent each corresponding learning layer. Further, if the tensor program has no label on the corresponding hardware, we set its label to the default value -100. When calculating the loss, the loss of this task is not calculated. That is to say, tasks without labels do not participate in backpropagation.

[0115] In single-task learning, the backpropagation of gradients often gets stuck in local optima. In multi-task learning, the local optima of different tasks are in different positions. The interaction of different tasks can help the hidden layer get rid of local optima. Multiple tasks share weights in the shallow layer, which can weaken the ability of the network, reduce overfitting of the network, and improve generalization ability.

[0116] Correspondingly, the present application can also solve the problem of insufficient sample quantity when training for a certain hardware platform. During the training process of a neural network model, the larger the sample quantity, the higher the accuracy of model training. If there are only 5 million effective labeled samples for the second hardware platform, while there are 50 million effective labeled samples for the corresponding first hardware platform, then the parameter of the basic learning layer of the model can be adjusted according to the labeled samples of the first hardware platform, and the hardware learning layer of the model can be adjusted according to the labeled samples of the second hardware platform, so that the accuracy of the model is also higher than that of the model that only adjusts the overall parameters of the model using the samples of the second hardware platform.

[0117] Therefore, in the embodiment of the present application, the predicted scores output by the target processing model are multiple groups of predicted scores, each group corresponding to a hardware platform, so that the predicted score corresponding to the target hardware platform can be obtained based on the predicted scores; and the target scheduling primitive feature vector can be determined from each of the scheduling primitive feature vectors according to the predicted score corresponding to the target hardware platform.

[0118] It can be seen that in order to solve the problem that the learning-based cost model depends on complex feature extraction engineering, the present invention patent application proposes to abstract the scheduling primitive sequence into a tensor language and extract features therefrom. In addition, the present invention patent application also proposes to use multi-task learning technology to solve the problem of the inavailability of the cost model across different hardware platforms. Experiments show that compared with the state-of-the-art implementation, the present invention can speed up the average search time of CPU and GPU workloads by 9.1 times and 3.0 times respectively. For the cross-hardware problem, the search time of CPU and GPU workloads can be accelerated by 4.7 times and 2.9 times respectively while only using 7% of the target hardware data. Compared with TenSet, the finally generated tensor program has an acceleration of 1.0 times to 1.3 times in execution time.

[0119] See Figure 7 , in another embodiment of the present application, a tensor program generation device is further provided, and the device includes:

[0120] An obtaining unit 701, configured to obtain a computation graph representing a target algorithm model;

[0121] A determining unit 702, configured to determine a set of scheduling primitives of a computation subgraph corresponding to the computation graph;

[0122] A screening unit 703, configured to screen a target scheduling primitive from the set of scheduling primitives based on the target performance metric and the hardware feature information of the target hardware platform to which the computational graph is to be applied;

[0123] A generating unit 704, configured to generate a tensor program corresponding to the target scheduling primitive, so that the tensor program can be deployed on the target hardware platform.

[0124] Optionally, the determining unit includes:

[0125] A first determining subunit, configured to determine a target operator node among the operator nodes based on the output connection relationship of the operator nodes in the computational graph;

[0126] A splitting subunit, configured to use the target operator node as a splitting point to split the computational graph into a plurality of computational subgraphs;

[0127] A combining subunit, configured to obtain a set of scheduling primitive instructions corresponding to the computational subgraph, combine the instructions in the set of scheduling primitive instructions, and obtain a plurality of scheduling primitives;

[0128] A first generating subunit, configured to generate a set of scheduling primitives corresponding to the computational subgraph based on the plurality of scheduling primitives.

[0129] Optionally, the screening unit includes:

[0130] A second generating subunit, configured to generate a scheduling primitive feature vector corresponding to each scheduling primitive;

[0131] A model processing subunit, configured to input each scheduling primitive feature vector into a target processing model to obtain a prediction score output by the target processing model that matches each scheduling primitive feature vector, where the target processing model is a machine learning model trained based on the target performance metric corresponding to each scheduling primitive as a training target value and the model learning structure corresponding to the hardware feature information of each hardware platform;

[0132] A second determining subunit, configured to determine a target scheduling primitive feature vector among the scheduling primitive feature vectors based on the prediction score;

[0133] A first obtaining subunit, configured to obtain a target scheduling primitive that matches the target scheduling primitive feature vector.

[0134] Optionally, the second generating subunit is specifically configured to:

[0135] Extract features from each scheduling primitive to obtain target features, where the target features include primitive characteristics, numerical parameters, and character parameters;

[0136] Determine the vector information corresponding to each target feature based on the feature vector mapping table;

[0137] Generate a scheduling primitive feature vector according to the vector information corresponding to each target feature.

[0138] Optionally, the device further includes:

[0139] A data acquisition unit, configured to acquire training data, where the training data includes a scheduling primitive feature vector corresponding to a scheduling primitive, and a labeled target performance metric label for each hardware platform;

[0140] A partitioning unit, configured to partition the training data into a training set and a test set;

[0141] A training unit, configured to train an initial model based on the training data to obtain a trained initial model, where the model structure of the initial model includes a basic learning layer and a hardware learning layer for learning the hardware features of each processing platform;

[0142] An adjustment unit, configured to test the trained model according to the test set, and adjust the model parameters of the trained initial model according to the test results to obtain a target processing model.

[0143] Optionally, the adjustment unit includes:

[0144] An input subunit, configured to input the test set into the trained initial model to obtain a prediction result output by the trained initial model;

[0145] A comparison subunit, configured to compare the prediction result with the target performance metric labeled in the test set, and if the comparison result does not meet a preset condition, determine the hardware platform corresponding to the current test set;

[0146] An adjustment subunit, configured to adjust the parameters of the hardware learning layer corresponding to the hardware platform based on the comparison result, and adjust the parameters of the basic learning layer based on the test results of all the test sets;

[0147] A third determination subunit, configured to determine the trained initial model with adjusted parameters as the target processing model.

[0148] Optionally, the second determination subunit is specifically configured to:

[0149] Obtain the prediction score corresponding to the target hardware platform based on the prediction score;

[0150] Determine the target scheduling primitive feature vector among all the scheduling primitive feature vectors according to the prediction score corresponding to the target hardware platform.

[0151] An embodiment of the present application provides a tensor program generation device, which obtains a computational graph representing a target algorithm model; determines a set of scheduling primitives for a computational subgraph corresponding to the computational graph; based on a target performance metric and hardware feature information of a target hardware platform to which the computational graph is to be applied, filters out target scheduling primitives from the set of scheduling primitives; and generates a tensor program corresponding to the target scheduling primitives, so that the tensor program can be deployed on the target hardware platform. By filtering the scheduling primitives and generating a tensor program for the filtered scheduling primitives, the present application improves the processing efficiency and can determine the tensor program based on the hardware feature information of the target hardware platform, making it more compatible with the current target hardware platform and improving the accuracy of generating the tensor program for the current hardware platform.

[0152] It should be noted that the specific implementation of each unit and subunit in this embodiment can refer to the corresponding content in the foregoing, and will not be elaborated herein.

[0153] An embodiment of the present application also provides a storage medium, which stores executable instructions that, when executed by a processor, implement the tensor program generation method described in any one of the above.

[0154] Correspondingly, an embodiment of the present application also provides an electronic device, which may include:

[0155] A memory for storing programs;

[0156] A processor for executing the program, and the program is specifically used to implement the tensor program generation method described in any one of the above.

[0157] It should be noted that the specific implementation of the processor in this embodiment can refer to the corresponding content in the foregoing, and will not be elaborated herein.

[0158] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method part.

[0159] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.

[0160] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well-known in the technical field.

[0161] The above description of the disclosed embodiments enables those skilled in the art to implement or use this application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application will not be limited to the embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for generating a tensor program, characterized in that, The method includes: Obtaining a computation graph representing a target algorithm model; Determining a set of scheduling primitives for a computation subgraph corresponding to the computation graph, where determining the set of scheduling primitives for the computation subgraph corresponding to the computation graph includes: determining a target operator node among the operator nodes based on the output connection relationship of the operator nodes in the computation graph; using the target operator node as a splitting point to split the computation graph into a plurality of computation subgraphs; obtaining a set of scheduling primitive instructions corresponding to the computation subgraphs, combining the instructions in the set of scheduling primitive instructions to obtain a plurality of scheduling primitives; generating a set of scheduling primitives corresponding to the computation subgraphs based on the plurality of scheduling primitives; Filtering out a target scheduling primitive from the set of scheduling primitives based on a target performance metric and hardware feature information of a target hardware platform to which the computation graph is to be applied, where filtering out the target scheduling primitive from the set of scheduling primitives based on the target performance metric and the hardware feature information of the target hardware platform to which the computation graph is to be applied includes: generating a scheduling primitive feature vector corresponding to each scheduling primitive; inputting each scheduling primitive feature vector into a target processing model to obtain a prediction score output by the target processing model that matches each scheduling primitive feature vector, where the target processing model is a machine learning model trained based on the target performance metric corresponding to each scheduling primitive as a training target value and a model learning structure corresponding to the hardware feature information of each hardware platform; determining a target scheduling primitive feature vector among the respective scheduling primitive feature vectors based on the prediction score; obtaining a target scheduling primitive that matches the target scheduling primitive feature vector; Generating a tensor program corresponding to the target scheduling primitive so that the tensor program can be deployed on the target hardware platform.

2. The method according to claim 1, wherein The generating a scheduling primitive feature vector corresponding to each scheduling primitive includes: Performing feature extraction on each scheduling primitive to obtain target features, where the target features include primitive characteristics, numerical parameters, and character parameters; Determining vector information corresponding to each target feature based on a feature vector mapping table; Generating a scheduling primitive feature vector according to the vector information corresponding to each target feature.

3. The method according to claim 1, wherein The method further includes: Obtaining training data, where the training data includes scheduling primitive feature vectors corresponding to scheduling primitives and labeled target performance metric labels for each hardware platform; Dividing the training data into a training set and a test set; Training an initial model based on the training data to obtain a trained initial model, where the model structure of the initial model includes a basic learning layer and a hardware learning layer for learning the hardware features of each processing platform; Testing the trained model according to the test set and adjusting the model parameters of the trained initial model according to the test results to obtain a target processing model.

4. The method according to claim 3, characterized in that, The testing the trained initial model according to the test set and adjusting the model parameters of the trained initial model according to the test results to obtain a target processing model includes: Input the test set into the trained initial model to obtain the prediction result output by the trained initial model; Compare the prediction result with the target performance metrics labeled in the test set. If the comparison result does not meet the preset conditions, determine the hardware platform corresponding to the current test set; Adjust the parameters of the hardware learning layer corresponding to the hardware platform based on the comparison result, and adjust the parameters of the basic learning layer based on the test results of all the test sets; Determine the trained initial model with adjusted parameters as the target processing model.

5. The method according to claim 1, characterized in that, The determining the target scheduling primitive feature vector among the respective scheduling primitive feature vectors based on the prediction score includes: Based on the prediction score, obtain the prediction score corresponding to the target hardware platform; According to the prediction score corresponding to the target hardware platform, determine the target scheduling primitive feature vector among the respective scheduling primitive feature vectors.

6. A generating device for a tensor program, characterized in that, The apparatus includes: An acquisition unit, configured to acquire a computation graph representing a target algorithm model; A determination unit, configured to determine a set of scheduling primitives for a computation subgraph corresponding to the computation graph. The determining the set of scheduling primitives for the computation subgraph corresponding to the computation graph includes: determining a target operator node among the operator nodes based on the output connection relationship of the operator nodes in the computation graph; using the target operator node as a splitting point to split the computation graph into a plurality of computation subgraphs; acquiring a set of scheduling primitive instructions corresponding to the computation subgraph, combining the instructions in the set of scheduling primitive instructions to obtain a plurality of scheduling primitives; generating a set of scheduling primitives corresponding to the computation subgraph based on the plurality of scheduling primitives; A screening unit, configured to screen out a target scheduling primitive from the set of scheduling primitives based on a target performance metric and hardware feature information of a target hardware platform to which the computation graph is to be applied. The screening out the target scheduling primitive from the scheduling primitives based on the target performance metric and the hardware feature information of the target hardware platform to which the computation graph is to be applied includes: generating a scheduling primitive feature vector corresponding to each scheduling primitive; inputting each scheduling primitive feature vector into a target processing model to obtain a prediction score output by the target processing model and matching each scheduling primitive feature vector, where the target processing model is a machine learning model trained based on a target performance metric corresponding to each scheduling primitive as a training target value and a model learning structure corresponding to the hardware feature information of each hardware platform; determining a target scheduling primitive feature vector among the respective scheduling primitive feature vectors based on the prediction score; obtaining a target scheduling primitive matching the target scheduling primitive feature vector; A generating unit, configured to generate a tensor program corresponding to the target scheduling primitive, so that the tensor program can be deployed on the target hardware platform.

7. A storage medium, characterized in that, The storage medium stores executable instructions, and when the instructions are executed by a processor, the method for generating a tensor program as described in any one of claims 1-5 is implemented.

8. An electronic device, characterized in that, It includes: A memory, configured to store a program; A processor for executing the program, which is specifically used to implement the method for generating a tensor program as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Task parallel processing implementation method and device, equipment and medium

    CN111309479A

  • Task scheduling method and device, electronic equipment and storage

    CN114840322A