Training method of cost model, generation method of tensor program and related device
Patent Information
- Application Number
- CN202611168839.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-04
- Publication Date
- 2026-09-01
AI Technical Summary
然而,现有成本模型普遍存在预测精度不足与计算效率偏低的问题
[0011] The beneficial technical effects of the training method for the cost model provided in this disclosure are at least as follows: First, the hardware parameters and initial labeled dataset of the first hardware platform are acquired, and the cost model is initially trained. Then, the following steps are executed iteratively: performance prediction is performed on candidate scheduling primitive sequences in the unlabeled dataset based on the current cost model; a sampling value score is calculated based on the prediction results; and target scheduling primitive sequences with reference value are selected based on this score. The target sequence is then actually measured on the first hardware platform to obtain its actual execution performance parameters, and these parameters are updated to the initial labeled dataset as new labeled samples for the next round of training. Thus, through iterative active sampling, high model accuracy can be achieved with fewer labeled samples, effectively overcoming the problem in existing technologies where a large amount of experimental data is required due to the lack of representativeness in the sampling strategy, and significantly shortening the data acquisition cycle. Simultaneously, the computational complexity of the sequence feature extraction layer in the cost model is linearly related to the length of the input sequence, effectively reducing computational complexity and improving training and inference efficiency, achieving a balance between prediction accuracy and computational efficiency.
Smart Images

Figure CN122674802A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to a method for training a cost model, a method for generating tensor programs, and related equipment. Background Technology
[0002] With the rapid development of deep learning technology, various neural network models have achieved remarkable results in fields such as image recognition and natural language processing. However, due to the significant differences in architectural characteristics among different hardware platforms, how to enable neural network models to execute efficiently on heterogeneous hardware platforms has become a key issue.
[0003] However, deep learning compiler systems typically employ automatic tuning techniques to generate efficient and optimized tensor programs for specific hardware. Existing automatic tuning frameworks usually require the construction of cost models to predict the execution performance of different scheduling primitive sequences on the target hardware, thereby guiding the search algorithm to efficiently locate the optimal solution within the vast scheduling space. However, existing cost models generally suffer from insufficient prediction accuracy and low computational efficiency. Summary of the Invention
[0004] This disclosure provides a method for training a cost model, a method for generating tensor programs, and related equipment to at least solve the above-mentioned technical problems existing in the prior art.
[0005] In a first aspect, embodiments of this disclosure provide a method for training a cost model, the method comprising: Obtain the hardware parameters of the first hardware platform and the first dataset. The first dataset includes multiple scheduling primitive sequences and the actual execution performance parameters of each scheduling primitive sequence on the first hardware platform. The cost model is trained based on hardware parameters and the first dataset. The performance of unlabeled candidate scheduling primitive sequences in the second dataset is predicted based on the cost model, and the sampling value score of the candidate scheduling primitive sequences is calculated based on the performance prediction results. The sampling value score represents the degree of contribution of the candidate scheduling primitive sequences to the cost model. The target scheduling primitive sequence is selected from the second dataset based on the sampling value score; Obtain the actual execution performance parameters of the target scheduling primitive sequence on the first hardware platform, update the target scheduling primitive sequence and the target scheduling primitive sequence to the first dataset; and return to execute the step of training the cost model based on the hardware parameters and the first dataset until the preset stopping condition is met; the computational complexity of the sequence feature extraction layer in the cost model is linearly related to the sequence length of the input scheduling primitive sequence.
[0006] Secondly, embodiments of this disclosure provide a method for generating tensor programs, the method comprising: Obtain the computational subgraph corresponding to the neural network model to be optimized, and generate multiple first scheduling primitive sequences corresponding to the computational subgraph; By inputting multiple first scheduling primitive sequences and the hardware parameters of the target hardware platform into the cost model obtained by the training method of the cost model provided in the first aspect, the predicted execution performance parameters corresponding to each first scheduling primitive sequence are obtained. Based on the predicted execution performance parameters, a second scheduling primitive sequence is determined from multiple first scheduling primitive sequences, and a tensor program that executes on the target hardware platform is generated based on the second scheduling primitive sequence.
[0007] Thirdly, embodiments of this disclosure provide an apparatus for training a cost model, the apparatus comprising: The first acquisition module is used to acquire the hardware parameters of the first hardware platform and the first dataset. The first dataset includes multiple scheduling primitive sequences and the actual execution performance parameters of each scheduling primitive sequence on the first hardware platform. The training module is used to train the cost model based on the hardware parameters and the first dataset; The performance prediction module is used to predict the performance of unlabeled candidate scheduling primitive sequences in the second dataset according to the cost model, and calculate the sampling value score of the candidate scheduling primitive sequences based on the performance prediction results. The sampling value score represents the degree of contribution of the candidate scheduling primitive sequences to the cost model. The filtering module is used to filter the target scheduling primitive sequences obtained from the second dataset based on the sampling value score; The first acquisition module is also used to acquire the actual execution performance parameters of the target scheduling primitive sequence on the first hardware platform, update the target scheduling primitive sequence and the target scheduling primitive sequence to the first dataset; and return to the step of training the cost model according to the hardware parameters and the first dataset until the preset stopping condition is met; the computational complexity of the sequence feature extraction layer in the cost model is linearly related to the sequence length of the input scheduling primitive sequence.
[0008] Fourthly, embodiments of this disclosure provide an apparatus for training a cost model, the apparatus comprising: The second acquisition module is used to acquire the computation subgraph corresponding to the neural network model to be optimized, and generate multiple first scheduling primitive sequences corresponding to the computation subgraph; The input module is used to input multiple first scheduling primitive sequences and the hardware parameters of the target hardware platform into the cost model obtained by the training method of the cost model provided in the first aspect, and to obtain the predicted execution performance parameters corresponding to each first scheduling primitive sequence. The determination module is used to determine a second scheduling primitive sequence from multiple first scheduling primitive sequences based on predicted execution performance parameters, and to generate a tensor program that executes on the target hardware platform based on the second scheduling primitive sequence.
[0009] Fifthly, embodiments of this disclosure provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to execute the training method of the cost model provided in the first aspect or the generation method of the tensor program provided in the second aspect.
[0010] In a sixth aspect, embodiments of this disclosure provide a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the training method of the cost model provided in the first aspect or the generation method of the tensor program provided in the second aspect.
[0011] The beneficial technical effects of the training method for the cost model provided in this disclosure are at least as follows: First, the hardware parameters and initial labeled dataset of the first hardware platform are acquired, and the cost model is initially trained. Then, the following steps are executed iteratively: performance prediction is performed on candidate scheduling primitive sequences in the unlabeled dataset based on the current cost model; a sampling value score is calculated based on the prediction results; and target scheduling primitive sequences with reference value are selected based on this score. The target sequence is then actually measured on the first hardware platform to obtain its actual execution performance parameters, and these parameters are updated to the initial labeled dataset as new labeled samples for the next round of training. Thus, through iterative active sampling, high model accuracy can be achieved with fewer labeled samples, effectively overcoming the problem in existing technologies where a large amount of experimental data is required due to the lack of representativeness in the sampling strategy, and significantly shortening the data acquisition cycle. Simultaneously, the computational complexity of the sequence feature extraction layer in the cost model is linearly related to the length of the input sequence, effectively reducing computational complexity and improving training and inference efficiency, achieving a balance between prediction accuracy and computational efficiency.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0013] Figure 1 This is a schematic diagram of the architecture of a cost model training system provided in an embodiment of this disclosure; Figure 2 This is a flowchart illustrating a training method for a cost model provided in an embodiment of this disclosure. Figure 1 ; Figure 3 This is a schematic diagram of the structure of a cost model provided in an embodiment of this disclosure; Figure 4 This is a flowchart illustrating a training method for a cost model provided in an embodiment of this disclosure. Figure 2 ; Figure 5 This is a schematic diagram illustrating the alternating learning of a knowledge base model and an activity column model provided in an embodiment of this disclosure; Figure 6 This is a flowchart illustrating a method for generating tensor programs according to an embodiment of this disclosure; Figure 7 This is a schematic diagram of the structure of a training device for a cost model provided in an embodiment of this disclosure; Figure 8 This is a schematic diagram of the structure of a tensor program generation device provided in an embodiment of this disclosure; Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0014] To make the objectives, features, and advantages of this disclosure more apparent and understandable, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0015] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0016] If the application documents contain similar descriptions such as "first / second", the following explanation shall be added: In the following description, the terms "first / second / third" are used only to distinguish similar objects and do not represent a specific order of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0018] As mentioned in the background section, the significant differences in architectural characteristics among different hardware platforms are a key factor restricting the efficient deployment of neural network models. Existing technologies predict the execution performance of scheduling primitive sequences on target hardware by constructing cost models to guide search optimization, but in practice, they still have the following shortcomings: First, it is difficult to achieve both accuracy and efficiency in model architecture. LSTM-based models converge slowly and have poor parallelism, while Transformer-based models can capture long-range dependencies, but their self-attention mechanism brings O(n²) computational complexity, resulting in a significant increase in training and inference overhead.
[0019] Secondly, data acquisition is costly. Existing sampling strategies (such as random sampling or uncertain sampling) lack representativeness and often require a large number of real hardware measurements to ensure prediction accuracy, resulting in data acquisition cycles that can last for several days or even tens of days, severely delaying the deployment process of compilers.
[0020] Third, the ability to transfer knowledge across hardware is lacking. Fine-tuning methods can only utilize single-source knowledge and are prone to forgetting, while multi-task learning faces the dilemma of parameter explosion and data synchronization. Therefore, when faced with new hardware, existing models usually need to be trained from scratch, making it difficult to reuse the optimization knowledge accumulated on existing hardware platforms.
[0021] Based on this, the present disclosure provides a method for training a cost model to solve the technical problems existing in the background art. The method for training a cost model provided by the present disclosure will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0022] For ease of understanding, the embodiments disclosed herein are first described in conjunction with the appendix. Figure 1 The architecture diagram of the training system for the cost model of this disclosure is described in detail.
[0023] Figure 1 This is an architecture diagram of a cost model training system provided in an embodiment of this disclosure. like Figure 1 As shown, the method provided in this embodiment can be deployed on a deep learning compiler system to achieve efficient automatic tuning and performance optimization of tensor programs on heterogeneous hardware platforms.
[0024] Specifically, deep learning frameworks such as TensorFlow and PyTorch output deep neural network models, which undergo graph-level optimization, including operator fusion, constant folding, and data layout transformation, to obtain a computational subgraph. This computational subgraph is then fed into a search algorithm to generate corresponding scheduling primitive sequences, supporting the training of the cost model.
[0025] The cost model can be divided into a training phase and an inference phase, as detailed below: Training Phase: The cost models for each other hardware are accumulated through knowledge distillation, with the knowledge being imported into a continuous knowledge distillation framework. The sampler selects a sequence of scheduling primitives and, combined with the cross-hardware knowledge accumulated by the continuous knowledge distillation framework, completes the offline training of the cost model for the target hardware. In addition to using the continuous knowledge distillation framework, an efficient sampler is also needed to sample data when training the cost model for the target hardware platform, in order to save data collection time.
[0026] Inference phase: The search algorithm outputs a sequence of scheduling primitives. The target hardware model performs performance prediction on the tensor program corresponding to the sequence of scheduling primitives and selects the K optimal sequences of scheduling primitives. Code is generated based on the selected K sequences of scheduling primitives and deployed to the target hardware platform.
[0027] Based on this, the training method of the cost model provided in the embodiments of this disclosure will be described in detail below with reference to the accompanying drawings.
[0028] Figure 2 This is a flowchart illustrating a training method for a cost model provided in an embodiment of this disclosure. Figure 1 .
[0029] like Figure 2 As shown, the training method for the cost model provided in this embodiment may include the following steps: S210, obtain the hardware parameters of the first hardware platform and the first dataset.
[0030] The first hardware platform is the target hardware to be optimized (i.e., the aforementioned target hardware platform), such as a central processing unit (CPU), a graphics processing unit (GPU), or other artificial intelligence (AI) accelerators. The hardware parameters of the first hardware platform may include structural parameters that characterize the hardware's computing power, such as the number of CPU cores, cache size, or the number of GPU computing units, memory bandwidth, etc., which are not specifically limited here.
[0031] Furthermore, the aforementioned first dataset is an initially labeled dataset, comprising multiple scheduling primitive sequences and the actual execution performance parameters of each scheduling primitive sequence on the first hardware platform. The scheduling primitive sequence can be several computational subgraphs output by a neural network model after graph-level optimization (such as operator fusion, constant folding, and data layout transformation). Each computational subgraph generates a series of low-level operation instruction sequences through a search algorithm. This scheduling primitive sequence guides the generation of tensor programs. The actual execution performance parameters are used to measure the running efficiency of the corresponding tensor program on the first hardware platform. For example, these performance parameters may include execution latency, i.e., the time consumed by the tensor program to run on the first hardware platform.
[0032] In one example, the first dataset can be obtained through random sampling. Specifically, a small number of samples (e.g., 1% to 3% of the total candidate samples) are randomly selected from the first dataset, and these programs are actually run on the first hardware platform. Their execution latency is collected as a supervision label, thereby constructing the initial labeled dataset.
[0033] S220 trains a cost model based on hardware parameters and the first dataset.
[0034] The aforementioned cost model can be used to predict the performance of an input scheduling primitive sequence on a specific hardware platform (e.g., the first hardware platform). This cost model may include a sequence feature extraction model, the computational complexity of which can be linearly related to the sequence length of the input scheduling primitive sequence.
[0035] In one example, the cost model is based on Linear-Time Sequence Modeling with Selective State Spaces (Mamba architecture) to efficiently capture long-range dependencies in scheduling primitive sequences in a lightweight manner. Figure 3 As shown, the structure of this cost model specifically includes: The encoder consists of three linear layers that map the input features to a high-dimensional space. The output dimensions of each layer are 64, 128, and 128, respectively. The encoder output is then fed into the Mamba block after being normalized by the root mean square. The Mamba block (i.e., the sequence feature extraction layer) contains a linear layer, 1D convolution, activation functions, and a structured state space model (SSM). The results of the two branch operations are fused and then output through the linear layer. This module dynamically adjusts the state transition parameters through a selection mechanism, thereby effectively capturing long-range dependencies in the scheduling primitive sequence. Its computational complexity is O(n), which is much lower than the O(n²) of the Transformer architecture. The output of the Mamba block is then processed by root mean square normalization before being sent to the decoder. Decoder: Consists of three linear layers with output dimensions of 64, 32, and 1 respectively, ultimately outputting performance predictions.
[0036] Furthermore, this cost model employs the LambdaRank loss function during training, focusing on optimizing the ranking accuracy of high-performance tensor programs rather than fitting precise latency values. This model maintains high prediction accuracy while having only about 0.35 MB of parameters.
[0037] Based on the above structure, this cost model takes the scheduling primitive sequence as input and combines it with the structural parameters of the target hardware (such as the number of CPU cores, cache size, or the number of GPU computing units and memory bandwidth) for feature encoding to form a unified input representation. The main body of the model adopts a state-space-based Mamba architecture to efficiently capture long-range dependencies in the scheduling primitive sequence in a lightweight manner, thereby achieving accurate prediction of tensor program performance.
[0038] S230, based on the cost model, perform performance prediction on the unlabeled candidate scheduling primitive sequences in the second dataset, and calculate the sampling value score of the candidate scheduling primitive sequences based on the performance prediction results.
[0039] The second dataset may include a large number of unlabeled (i.e., not run on actual hardware) candidate scheduling primitive sequences. Performance prediction results may refer to the predicted execution performance parameters of the corresponding tensor program output by the cost model after the candidate scheduling primitive sequences are input into the cost model, such as predicted execution latency; no specific limitation is made here.
[0040] The sampling value score is used to characterize the value of candidate scheduling primitive sequences to further training the cost model, that is, the degree to which they contribute to improving the prediction accuracy of the cost model after being added to the first dataset. Specifically, the larger the sampling value score, the higher the contribution of the candidate scheduling primitive sequence to improving the performance of the cost model, and the more it should be prioritized for screening and actual hardware measurement; conversely, the smaller the sampling value score, the lower the contribution of the candidate sequence, and it can be postponed or not measured. This will not be elaborated further here.
[0041] S240, the target scheduling primitive sequence obtained from the second dataset based on the sampling value score.
[0042] The target scheduling primitive sequence can be a candidate scheduling primitive sequence selected from the second dataset based on the sampling value score, which is intended to be measured in actual hardware to obtain a true performance label. There can be one or more target scheduling primitive sequences, and this embodiment of the disclosure does not specifically limit the number of target scheduling primitive sequences.
[0043] S250: Obtain the actual execution performance parameters of the target scheduling primitive sequence on the first hardware platform, update the target scheduling primitive sequence and the target scheduling primitive sequence to the first dataset; and return to execute the step of training the cost model based on the hardware parameters and the first dataset until the preset stopping condition is met.
[0044] The preset stopping condition is a pre-set criterion used to terminate iterative training. For example, the stopping condition can be determined to be met and the iteration can end when the total number of labeled samples reaches the preset sampling budget limit (e.g., 10% of the total number of samples), the model prediction accuracy reaches the expected target, or the model performance no longer improves significantly in multiple consecutive iterations.
[0045] Specifically, the hardware parameters of the first hardware platform and the first dataset are first obtained. The first dataset is an initial labeled dataset, including multiple scheduling primitive sequences and the actual execution performance parameters of each scheduling primitive sequence on the first hardware platform. The hardware parameters and the scheduling primitive sequences in the first dataset are used as input to the cost model, and the cost model is trained by combining the actual execution performance parameters corresponding to each scheduling primitive sequence.
[0046] Next, the iteration phase begins, repeatedly executing the following steps until a preset stopping condition is met: Based on the cost model trained so far, the performance of unlabeled candidate scheduling primitive sequences in the second dataset is predicted to obtain the prediction performance results for each candidate scheduling primitive sequence. Based on the prediction performance results, the sampling value score of each candidate scheduling primitive sequence is calculated. The sampling value score is used to characterize the degree of contribution of the candidate sequence to further training of the cost model. Then, based on the sampling value score, the target scheduling primitive sequence to be used for actual hardware measurement is selected from the second dataset.
[0047] The selected target scheduling primitive sequences were actually run on the first hardware platform, and their actual execution performance parameters were measured and obtained. Subsequently, the target scheduling primitive sequences and their corresponding actual execution performance parameters were added to the first dataset as new labeled samples to expand the labeled dataset.
[0048] After the first dataset is updated, the process returns to the step of training the cost model based on the hardware parameters and the first dataset, i.e., retraining the cost model using the updated first dataset. This process is repeated iteratively until a preset stopping condition is met, resulting in the final trained cost model.
[0049] The training method for the cost model provided in this disclosure first acquires the hardware parameters of a first hardware platform and an initial labeled dataset to perform preliminary training on the cost model. Then, the following steps are executed iteratively: performance prediction is performed on candidate scheduling primitive sequences in the unlabeled dataset based on the current cost model; a sampling value score is calculated based on the prediction results; and target scheduling primitive sequences with reference value are selected based on this score. The target sequence is then actually measured on the first hardware platform to obtain its actual execution performance parameters, which are then updated to the initial labeled dataset as new labeled samples for the next round of training. Thus, through iterative active sampling, high model accuracy can be achieved with fewer labeled samples, effectively overcoming the problem in existing technologies where a large amount of experimental data is required due to the lack of representativeness in the sampling strategy, and significantly shortening the data acquisition cycle. Simultaneously, the computational complexity of the sequence feature extraction layer in the cost model is linearly related to the length of the input sequence, effectively reducing computational complexity and improving training and inference efficiency, achieving a balance between prediction accuracy and computational efficiency.
[0050] In order to provide a comprehensive and detailed description of the training method for the cost model provided in the embodiments of this disclosure, in one embodiment, the above-described S240 may include: Based on the distribution ratio of different operator types in the second dataset, determine the sampling quota corresponding to each type of operator; For each type of operator, a target scheduling sequence that meets the sampling quota is selected from the second dataset according to the preset order of sampling value.
[0051] Here, the operator type refers to the core computational operation type in the tensor program corresponding to the scheduling primitive sequence. For example, the operator type can include convolution operators, matrix multiplication operators, activation function operators, pooling operators, normalization operators, etc., without specific limitations here. The distribution ratio can be the percentage of the number of candidate scheduling primitive sequences corresponding to each type of operator to the total number of candidate scheduling primitive sequences in the second dataset, without specific limitations here.
[0052] In addition, the sampling quota refers to the upper limit of the number of samples that can be selected for each type of operator in this iteration of data collection. It is used to ensure that the selected samples are consistent with the original dataset in terms of operator type distribution.
[0053] The preset order mentioned above can be a screening order pre-set based on actual experience or circumstances. This preset order can be in the order of sampling value scores from high to low, and no specific limitation is made here.
[0054] Specifically, firstly, statistical analysis is performed on all samples in the second dataset to identify the core operator types contained in each scheduling primitive sequence, and the frequency or proportion of each operator type in the entire second dataset is statistically analyzed. Then, based on the statistically obtained distribution proportions and combined with the preset total sampling budget, a corresponding sampling quota is allocated to each type of operator. After determining the sampling quotas for each type of operator, all candidate scheduling primitive sequences belonging to a certain operator type are selected from the second dataset, and the sampling value score calculated for each candidate sequence in the previous step S130 is obtained. Then, all candidate sequences under this type of operator are sorted from high to low (i.e., in descending order) according to the sampling value score. Finally, according to the sampling quota corresponding to this type of operator, a corresponding number of candidate scheduling primitive sequences are selected from the sorted sequences from high to low as the target scheduling sequences selected under this type of operator. The above operation is performed for each type of operator, and finally, the target sequences selected from all categories are merged to form all the target scheduling primitive sequences selected from the second dataset.
[0055] In one example, if the second dataset contains operator type A and operator type B with a distribution ratio of 7:3, and the total sampling quota for this round is 100 samples, then the sampling quota for operator type A is determined to be 70 samples, and the sampling quota for operator type B is determined to be 30 samples. For operator type A, the top 70 sequences are selected from all candidate sequences belonging to this type according to their sampling value scores from high to low; similarly, the top 30 sequences are selected for operator type B. Finally, the two selection results are merged to obtain the complete sequence of target scheduling primitives for this iteration.
[0056] In this embodiment, by combining the distribution ratio of operator types to determine the sampling quota and screening the target scheduling sequence, the problem of insufficient sample representativeness caused by ignoring data distribution in traditional sampling methods can be effectively avoided. This allows the screened target samples to maintain good data distribution consistency while improving model accuracy, further improving the efficiency of iterative sampling and the predictive performance of the final cost model.
[0057] To describe in detail the training method of the cost model provided in the embodiments of this disclosure, in one embodiment, the step of calculating the sampling value score of the candidate scheduling primitive sequence based on the performance prediction result may specifically include the following steps: Obtain first information, which characterizes the degree of difference in prediction performance between unlabeled candidate scheduling primitive sequences and labeled scheduling primitive sequences; Obtain the second information, which characterizes the degree of dispersion of the cost model's prediction performance on unlabeled candidate scheduling primitive sequences; Based on the first and second information, determine the sampling value score of the candidate scheduling primitive sequence.
[0058] The first information measures the amount of new information a candidate sample can bring to the current training set. A greater degree of difference indicates that the region where the sample is located is less covered by labeled samples, meaning that adding it to the training set can bring more new information to the model and help expand the model's ability to recognize unknown regions. For example, for any unlabeled sample, the minimum difference between its prediction performance and the prediction performance of all labeled samples can be calculated as the first information.
[0059] The second piece of information is used to measure the uncertainty of the cost model's prediction for that sample. Higher uncertainty indicates insufficient knowledge of the region where the sample is located, suggesting the sample has greater labeling value and that adding it to the training set can effectively improve the model's prediction accuracy in that region. For example, the variance of the prediction performance of unlabeled samples compared to all labeled samples can be calculated as the second piece of information.
[0060] Specifically, it is possible to obtain first information and second information. Since the first information represents the degree of difference in prediction performance between the unlabeled candidate scheduling primitive sequence and the labeled scheduling primitive sequence, and the second information represents the degree of dispersion of the prediction performance of the current cost model for the unlabeled candidate scheduling primitive sequence, after obtaining the first information and the second information, it is possible to determine the sampling value score of the candidate scheduling primitive sequence based on the first information and the second information.
[0061] In one example, the first piece of information mentioned above is used to measure the difference in prediction performance between the current unlabeled samples and the labeled samples; for each unlabeled sample Calculate its predictive performance With all labeled samples ( ,in Predictive performance (for the total number of labeled samples) The minimum absolute difference between the samples is used as the diversity score for that sample. The calculation formula is as follows: (1) in, This indicates a labeled dataset; This represents the predictive performance value of the cost model for the input sample; The larger the value, the more likely it is to be an unlabeled sample. The greater the difference in prediction performance between the sample and all labeled samples, that is, the more new information the sample can bring to the current training set, the higher its representative value.
[0062] The second piece of information mentioned above is used to measure the cost model's performance on unlabeled samples. Uncertainty in prediction performance. In this example, unlabeled samples are used. Variance of prediction performance with all labeled samples As the uncertainty score of this sample ,Right now The calculation method is as follows: (2) in, Indicates unlabeled samples The mean of the prediction performance of all labeled samples.
[0063] (3) in, Indicates unlabeled samples The variance of the predictive performance against all labeled samples; the larger the variance value, the better the cost model performs against the sample. The more unstable the predictive performance, that is, the less knowledge the model has of the region where the sample is located, the higher the labeling value of the sample.
[0064] In this embodiment, the sampling value score is determined by comprehensively considering the degree of difference between unlabeled and labeled samples and the uncertainty of the model's prediction of unlabeled samples. This allows the selected target samples to effectively fill the gaps in the data distribution of the current training set, while also focusing on areas with low model prediction confidence. This results in a significant improvement in model performance with less labeling cost, further improving the efficiency of iterative sampling and the prediction accuracy of the final cost model.
[0065] Based on this, in one embodiment, the step of determining the sampling value score of the candidate scheduling primitive sequence according to the first information and the second information may include: The sampling value score of the candidate scheduling primitive sequence is determined by multiplying the performance prediction result of the candidate scheduling primitive sequence with the first information and summing it with the second information.
[0066] Specifically, the sampling value score is calculated as follows: the performance prediction result of the candidate scheduling primitive sequence is multiplied by the first information, and then the product is added to the second information. The result is the sampling value score of the candidate scheduling primitive sequence.
[0067] In one example, the product of the performance prediction result based on the candidate scheduling primitive sequence and the first information, plus the sum of the second information, determines the sampling value score of the candidate scheduling primitive sequence according to the following formula: (4) in, Represents any unlabeled candidate scheduling primitive sequence in the second dataset; This represents the candidate scheduling primitive sequence. Sampling value score; This indicates that the current cost model is effective against... The predictive performance value is used as a weighting factor to prioritize samples with higher predictive performance when diversity scores are the same. Indicates the first piece of information; This indicates the second piece of information.
[0068] In this way, by jointly optimizing the three dimensions of representativeness (diversity), uncertainty (model confidence), and prediction performance, the most valuable samples can be efficiently identified from a massive amount of unlabeled samples.
[0069] In this embodiment, by incorporating prediction performance as a weighting factor into the calculation of the sampling value score, samples with superior prediction performance are prioritized when diversity is comparable. This helps guide the sampling direction towards high-performance regions. Simultaneously, combining diversity and uncertainty indicators ensures that the selected samples cover a wider range of data distribution areas and supplement regions lacking model knowledge. This organic combination of three factors allows each iteration to achieve the maximum model performance gain with minimal annotation costs, significantly improving the efficiency of iterative sampling and the prediction accuracy of the final cost model.
[0070] In order to fully and thoroughly describe the training method of the cost model provided in the embodiments of this disclosure, in one embodiment, such as Figure 4 As shown, the steps described above for training the cost model based on hardware parameters and the first dataset can specifically include the following steps: S410, Obtain the knowledge base model.
[0071] The Knowledge Base (KB) model stores common knowledge information from multiple second hardware platforms. This common knowledge information can be shared optimization knowledge learned by each second hardware platform during tensor program performance prediction, such as the general mapping between scheduling primitive sequences and hardware execution performance, or the general prior distribution of the performance impact of different scheduling strategies; no specific limitations are imposed here. Furthermore, the second hardware platform refers to one or more source hardware platforms from which the cost model has been trained. Its platform type can be the same as or different from the first hardware platform; no specific limitations are imposed here.
[0072] It should be noted that the knowledge base model uses the cost models already trained on various second hardware platforms as teacher models. Knowledge distillation is used to transfer the knowledge trained on each second hardware platform to the knowledge base model. Specifically, during the historical optimization process, for each of the multiple source hardware platforms where the cost models have been trained, each platform possesses a pre-trained cost model. These models contain specialized knowledge for tensor program performance prediction on that hardware platform. Through knowledge distillation, this specialized knowledge is extracted and integrated into a unified knowledge base model, enabling this model to store cross-hardware general knowledge information. Furthermore, the structure of the knowledge base model is identical to that of the cost model; both are lightweight sequence prediction models.
[0073] S420 uses the cost model of the first hardware platform as the active column model, trains the active column model based on hardware parameters and the first dataset, and uses the general knowledge information of the knowledge base model to assist in the training of the active column model during the training process.
[0074] The Active Column (AC) model is a neural network with the same structure as the knowledge base model, used to learn knowledge about the current specific new hardware (i.e., the first hardware platform). The parameters of the Active Column model are independent of the knowledge base model; they can be structurally identical but do not share parameters, which is not specifically limited here.
[0075] Specifically, it is possible to acquire a knowledge base model and use the cost model of the first hardware platform as the active column model. The active column model is trained based on hardware parameters and the first dataset. During the training process of the active column model, the general knowledge information of the knowledge base model is used to assist in the training of the active column model. This enables the active column model to reuse existing general knowledge information when learning new hardware data, thereby accelerating the convergence process and improving the prediction accuracy under conditions of a small number of samples.
[0076] like Figure 5As shown, the continuous knowledge distillation framework consists of a fixed-capacity hardware knowledge base model (KB) and an active column model (AC) for learning from the current hardware. Through alternating continuous learning (CL) and knowledge distillation (KD) phases, it achieves continuous accumulation and transfer of multi-source knowledge. KB is a fixed-size neural network (using the Mamba cost model structure described above) used for long-term storage of general optimization knowledge learned from multiple hardware components. AC is a neural network with the same structure as KB, used to learn knowledge about specific new hardware, with independent parameters. In the continuous learning phase, the parameters of the knowledge base model are frozen and used as a source of prior knowledge. Its feature output is introduced into the active column model through layer-by-layer connections, enabling the active column to reuse existing knowledge when learning new hardware data, thereby accelerating the convergence process. In the knowledge distillation phase, the active column model is frozen and used as a teacher model. Its knowledge is transferred to the knowledge base model through a distillation loss function. Simultaneously, regularization constraints based on Fisher information limit drastic changes in the knowledge base parameters, thus avoiding the forgetting of learned historical hardware knowledge. For subsequent new hardware platforms, repeat the alternating process of "continuous learning + knowledge distillation" to enable the knowledge base model to continuously accumulate general optimization knowledge from multiple hardware platforms, while keeping the model parameter size unchanged and without needing to access historical hardware data.
[0077] In this embodiment, when a cost model needs to be trained for a new hardware platform, it is not necessary to start from scratch or access historical hardware data. Only the existing knowledge base model needs to be used as a source of prior knowledge, allowing for rapid training of a high-precision cost model with a small sample size. Furthermore, the parameter size of the knowledge base model is fixed and does not increase with the number of hardware devices, effectively solving the scalability problem in cross-hardware knowledge transfer and enabling the continuous accumulation and efficient reuse of multi-source knowledge.
[0078] To provide a comprehensive and detailed description of the training method for the cost model provided in this disclosure, in one embodiment, the step of using the general knowledge information of the knowledge base model to assist in the training of the active column model during the training process may specifically include: During the training of the active column model, the parameters of the knowledge base model are frozen, and the feature output of the knowledge base model is introduced into the active column model to assist in the training of the active column model by utilizing the general knowledge information of the knowledge base model.
[0079] In this context, freezing the parameters of the knowledge base model can mean that during the training process of the active column model, the network weights of the knowledge base model do not participate in gradient backpropagation and are not updated during the training process, thereby ensuring that the historical hardware knowledge stored in the knowledge base model will not be overwritten by the training process of the current hardware platform.
[0080] Specifically, during the training process of the active column model, the parameters of the knowledge model can be frozen, and the feature output of the knowledge base model can be introduced into the active column model. That is, during the forward propagation of the active column model, the feature information extracted by the intermediate or output layer of the knowledge base model is used as an auxiliary input to provide the active column model with the general knowledge information of the knowledge base model for auxiliary training.
[0081] In this embodiment, the parameters of the knowledge base model can be frozen during the training of the active column model, and the feature output of the knowledge base model can be introduced into the active column model to assist in its training using the general knowledge information of the knowledge base model. In this way, based on the general optimization knowledge provided by the knowledge base model, only a small amount of labeled data from the first hardware platform is needed to quickly adapt to the characteristics of the current hardware platform, effectively shortening the model convergence time, improving prediction accuracy under limited sample conditions, and avoiding the catastrophic forgetting problem of historical knowledge.
[0082] In order to fully and thoroughly describe the training method of the cost model provided in the embodiments of this disclosure, in one embodiment, the method may further include the following steps: After the active column model is trained, its parameters are frozen. The active column model is then used as the teacher model to perform knowledge distillation on the knowledge base model, transferring the knowledge information from the active column model to the knowledge base model.
[0083] The freezing of parameters in the active column model can refer to the fact that the network weights of the active column model no longer change during the knowledge distillation process, which will not be elaborated on here.
[0084] Specifically, after the active column model is trained, it is used as the teacher model to perform knowledge distillation on the knowledge base model. That is, the hardware-specific knowledge contained in the trained active column model is transferred to the knowledge base model through knowledge distillation. This allows the knowledge base model to retain existing historical hardware knowledge and further integrate the knowledge of the current first hardware platform, so that the knowledge base model can learn the hardware-specific knowledge mastered by the active column model.
[0085] In the knowledge distillation stage, to ensure that the knowledge base model can effectively learn the current hardware knowledge contained in the active column model, while maintaining its independent learning ability and preventing catastrophic forgetting of learned historical knowledge, this embodiment designs the following distillation loss function: (5) in, This represents the distillation loss function. This represents the prediction result of the knowledge base model (student model). With real labels The ranking loss between these values is used to maintain the knowledge base model's independent learning ability on the current hardware platform; where, The weighting coefficient for the first term is used to balance the contribution of each loss to the total loss.
[0086] This represents the prediction results of the knowledge base model. Prediction results compared with the activity column model (teacher model) The ranking loss between them is used to transfer the current hardware platform knowledge held by the active column model to the knowledge base model; where, Normalized weights for teacher prediction quality are used to weight the prediction output quality of the active column model. The less accurate the prediction, the less guiding its role, and the lower its corresponding weight value.
[0087] EWC loss is used to consolidate previously learned knowledge from KB; among which, This is the regularization coefficient, used to control the strength of the EWC regularization term; Indicates the index of the current task. Indicates the index of the previous task; This indicates that the knowledge base model is in the current task. The next Each parameter value; This indicates that the knowledge base model was in the previous task. The next Each parameter value; Indicates the first The parameters in the current task The Fisher information value is used to measure the sensitivity of the parameter to the model's prediction results. The larger the Fisher information value, the more important the parameter is, and the greater the penalty for its changes during updates.
[0088] In this embodiment, after knowledge distillation, the knowledge base model is updated to a new version that incorporates general knowledge from more hardware platforms, including the first hardware platform. For subsequent new hardware platforms, the alternating process of "training the active column model + knowledge distillation to update the knowledge base" is repeated, enabling the knowledge base model to continuously accumulate general optimization knowledge from multiple hardware platforms while maintaining the same model parameter size. Furthermore, the entire process does not require access to historical hardware data, achieving scalable, cross-hardware knowledge accumulation and efficient reuse.
[0089] Furthermore, this disclosure also provides a method for generating tensor programs. The training method for the cost model provided by this disclosure will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0090] Figure 6 This is a flowchart illustrating a method for generating tensor programs provided in this embodiment. Figure 1 .
[0091] like Figure 6 As shown, the method for generating this tensor program may include the following steps: S610: Obtain the computational subgraph corresponding to the neural network model to be optimized, and generate multiple first scheduling primitive sequences corresponding to the computational subgraph.
[0092] S620 inputs multiple first scheduling primitive sequences and the hardware parameters of the target hardware platform to the system using, for example... Figure 2 The method steps shown yield a cost model, which provides the predicted execution performance parameters for each first scheduling primitive sequence.
[0093] S630 determines a second scheduling primitive sequence from multiple first scheduling primitive sequences based on predicted execution performance parameters, and generates a tensor program to be executed on the target hardware platform based on the second scheduling primitive sequence.
[0094] For details, please refer to the aforementioned embodiments; further details will not be elaborated here.
[0095] Based on the tensor program generation method provided in this disclosure, the lightweight and high-precision cost model trained in the aforementioned embodiments is used to quickly screen a massive number of candidate scheduling schemes. Only a very small number of candidate schemes need to be tested on actual hardware to find the near-optimal tensor program, thereby significantly shortening the deployment cycle of deep learning models on heterogeneous hardware platforms. Furthermore, since the computational complexity of the cost model is linearly related to the sequence length, it can quickly complete batch prediction of massive candidate sequences, further improving the overall tuning efficiency.
[0096] Based on the method provided in this disclosure, the following technical effects can be achieved: First, data acquisition efficiency is significantly improved: the high-efficiency sampler only needs 10% of the samples to achieve model accuracy comparable to the full dataset, shortening the data acquisition time from tens of days to several days. Second, the cost-effective and lightweight model is efficient: the Mamba-based architecture maintains high prediction accuracy while having only 0.35MB of parameters, with complexity linearly related to sequence length, resulting in faster training and inference speeds. Furthermore, cross-hardware knowledge transfer is scalable: the continuous knowledge distillation framework decouples storage and learning through a knowledge base and active columns, ensuring that model parameters do not increase with the number of hardware components and that data synchronization is unnecessary, avoiding parameter explosion and data dependency problems in multi-task learning. Finally, overall performance is excellent: as shown in Table 1, experiments on multiple CPU and GPU platforms demonstrate that this method accelerates tuning time by 18.5× (CPU) and 23.2× (GPU) compared to Tenset-MLP, and improves inference speed by 1.20× and 1.15×, respectively.
[0097] Table 1
[0098] Furthermore, based on the same inventive concept, embodiments of this disclosure provide a training device for a cost model, which can be specifically combined with the appendix. Figure 7 A training apparatus for a cost model provided in this disclosure will be described in detail.
[0099] Figure 7 This is a schematic diagram of the structure of a cost model training device provided in an embodiment of this disclosure.
[0100] like Figure 7 As shown, the training device 700 for this cost model may include: The first acquisition module 710 is used to acquire the hardware parameters of the first hardware platform and the first dataset. The first dataset includes multiple scheduling primitive sequences and the actual execution performance parameters of each scheduling primitive sequence on the first hardware platform. Training module 720 is used to train the cost model based on hardware parameters and the first dataset; The performance prediction module 730 is used to predict the performance of unlabeled candidate scheduling primitive sequences in the second dataset according to the cost model, and calculate the sampling value score of the candidate scheduling primitive sequences based on the performance prediction results. The sampling value score represents the degree of contribution of the candidate scheduling primitive sequences to the cost model. The filtering module 740 is used to filter the target scheduling primitive sequence obtained from the second dataset based on the sampling value score; The first acquisition module 710 is also used to acquire the actual execution performance parameters of the target scheduling primitive sequence on the first hardware platform, update the target scheduling primitive sequence and the target scheduling primitive sequence to the first dataset; and return to the step of training the cost model according to the hardware parameters and the first dataset until the preset stopping condition is met; the computational complexity of the sequence feature extraction layer in the cost model is linearly related to the sequence length of the input scheduling primitive sequence.
[0101] In one embodiment, the training apparatus for the cost model provided in this disclosure may further include: The determination module is used to determine the sampling quota corresponding to each type of operator based on the distribution ratio of different operator types in the second dataset; The filtering module is specifically used to select target scheduling sequences that meet the sampling quota from the second dataset according to the preset order of sampling value for various operators.
[0102] In one embodiment, the training apparatus for the cost model provided in this disclosure may further include: The first acquisition module is also used to acquire first information, which characterizes the degree of difference in prediction performance between the unlabeled candidate scheduling primitive sequence and the labeled scheduling primitive sequence. The first acquisition module is also used to acquire second information, which characterizes the degree of dispersion of the cost model's prediction performance on unlabeled candidate scheduling primitive sequences. The determination module is also used to determine the sampling value score of the candidate scheduling primitive sequence based on the first information and the second information.
[0103] In one embodiment, the training apparatus for the cost model provided in this disclosure may further include: The determination module is specifically used to determine the sampling value score of the candidate scheduling primitive sequence based on the product of the performance prediction result of the candidate scheduling primitive sequence and the first information, and the sum of the first information and the second information.
[0104] In one embodiment, the training apparatus for the cost model provided in this disclosure may further include: The first acquisition module is also used to acquire a knowledge base model. The knowledge base model is used to store general knowledge information of multiple second hardware platforms. The knowledge base model is obtained by using the cost model that has been trained on each second hardware platform as the teacher model and transferring the knowledge trained on each second hardware platform to the knowledge base model through knowledge distillation.
[0105] The training module is specifically used to train the active column model based on the cost model of the first hardware platform, according to the hardware parameters and the first dataset. During the training process of the active column model, the general knowledge information of the knowledge base model is used to assist in the training of the active column model.
[0106] In one embodiment, the training apparatus for the cost model provided in this disclosure may further include: The freeze module is used to freeze the parameters of the knowledge base model during the training process of the active column model, and introduce the feature output of the knowledge base model into the active column model to assist in the training of the active column model by utilizing the general knowledge information of the knowledge base model.
[0107] In one embodiment, the training apparatus for the cost model provided in this disclosure may further include: The freeze module is used to freeze the parameters of the active column model after training is completed. The active column model is then used as the teacher model to perform knowledge distillation on the knowledge base model, transferring the knowledge from the active column model to the knowledge base model.
[0108] It is understood that the cost model training method apparatus provided in the above embodiments can, as needed, allocate the above processing to different program modules to complete all or part of the processing described above when implementing the corresponding cost model training method. Furthermore, the apparatus and corresponding method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiments, and will not be repeated here.
[0109] Furthermore, based on the same inventive concept, embodiments of this disclosure provide a tensor program generation apparatus, which can be specifically combined with the appendix. Figure 8 A tensor program generation apparatus provided in this disclosure will be described in detail.
[0110] Figure 8 This is a schematic diagram of the structure of a tensor program generation device provided in an embodiment of this disclosure.
[0111] like Figure 8 As shown, the tensor program generation device 800 may include: The second acquisition module 810 is used to acquire the computation subgraph corresponding to the neural network model to be optimized, and generate multiple first scheduling primitive sequences corresponding to the computation subgraph; Input module 820 is used to input multiple first scheduling primitive sequences and hardware parameters of the target hardware platform into a system that utilizes, for example,... Figure 2 The cost model obtained by the method steps shown yields the predicted execution performance parameters corresponding to each first scheduling primitive sequence; The determination module 830 is used to determine a second scheduling primitive sequence from multiple first scheduling primitive sequences based on predicted execution performance parameters, and to generate a tensor program that will be executed on the target hardware platform based on the second scheduling primitive sequence.
[0112] It is understood that the tensor program generation method apparatus provided in the above embodiments can, when implementing the corresponding tensor program generation method, allocate the above processing to different program modules as needed to complete all or part of the processing described above. Furthermore, the apparatus and corresponding method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiments, and will not be repeated here.
[0113] This application provides a computer program product or computer program that includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform a method for training a cost model or a method for generating a tensor program.
[0114] This application provides a computer-readable storage medium storing executable instructions. When the executable instructions are executed by a processor, the processor will execute the cost model training method or the tensor program generation method provided in this application.
[0115] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0116] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0117] As an example, executable instructions may, but do not necessarily, correspond to files in a file system. They may be stored as part of a file that holds other programs or data, for example, in one or more scripts in a Hyper Text Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple collaborating files (e.g., a file that stores one or more modules, subroutines, or code sections).
[0118] As an example, executable instructions can be deployed to execute on a single computing device, or on multiple computing devices located in one location, or on multiple computing devices distributed across multiple locations and interconnected via a communication network.
[0119] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure; as shown below. Figure 9 As shown, the electronic device 90 includes: a processor 901, and a memory 902 communicatively connected to the processor 901; the memory 902 stores instructions executable by the processor 901. The instructions are executed by the processor 901 to enable the processor 901 to perform: Obtain the hardware parameters of the first hardware platform and the first dataset, wherein the first dataset includes multiple scheduling primitive sequences and the actual execution performance parameters of each scheduling primitive sequence on the first hardware platform; A cost model is trained based on the hardware parameters and the first dataset. The performance of unlabeled candidate scheduling primitive sequences in the second dataset is predicted according to the cost model, and the sampling value score of the candidate scheduling primitive sequences is calculated based on the performance prediction results. The sampling value score represents the degree of contribution of the candidate scheduling primitive sequences to the cost model. The target scheduling primitive sequence is obtained by filtering from the second dataset based on the sampling value score; Obtain the actual execution performance parameters of the target scheduling primitive sequence on the first hardware platform, update the target scheduling primitive sequence and the target scheduling primitive sequence to the first dataset; and return to execute the step of training the cost model based on the hardware parameters and the first dataset until a preset stopping condition is met; the computational complexity of the sequence feature extraction layer in the cost model is linearly related to the sequence length of the input scheduling primitive sequence.
[0120] Alternatively, enable processor 901 to execute: Obtain the computational subgraph corresponding to the neural network model to be optimized, and generate multiple first scheduling primitive sequences corresponding to the computational subgraph; The plurality of first scheduling primitive sequences and the hardware parameters of the target hardware platform are input to the application. Figure 1 The cost model obtained by the training method of the cost model described above yields the predicted execution performance parameters corresponding to each first scheduling primitive sequence; Based on the predicted execution performance parameters, a second scheduling primitive sequence is determined from the plurality of first scheduling primitive sequences, and a tensor program to be executed on the target hardware platform is generated based on the second scheduling primitive sequence. The embodiments of the electronic device and the corresponding cost model training method or tensor program generation method provided above belong to the same concept, and their specific implementation processes are detailed in the method embodiments, and will not be repeated here.
[0121] In practical applications, the electronic device 90 may further include at least one network interface 903. The various components of the electronic device 90 are coupled together via a bus system 904. It is understood that the bus system 904 is used to implement communication between these components. In addition to a data bus, the bus system 904 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 9 All buses are labeled as bus system 904. The number of processors 901 and the number of memories 902 can be at least one. The network interface 903 is used for wired or wireless communication between the electronic device 90 and other devices.
[0122] The memory 902 in this embodiment is used to store various types of data to support the operation of the electronic device 90.
[0123] The methods disclosed in the above embodiments of this disclosure can be applied to or implemented by processor 901. Processor 901 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuit of the hardware in processor 901 or by instructions in software form. The processor 901 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 901 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this disclosure. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the methods disclosed in the embodiments of this disclosure can be directly manifested as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software modules may be located in a storage medium, which is located in memory 902. Processor 901 reads the information in memory 902 and, in conjunction with its hardware, completes the steps of the aforementioned cost model training method or tensor program generation method.
[0124] In some embodiments, the electronic device 90 may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the aforementioned methods.
[0125] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0126] In the above description, the term "some embodiments" refers to a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0127] Unless otherwise defined, all technical and scientific terms used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used in this disclosure is for the purpose of describing embodiments of this disclosure only and is not intended to be limiting of this disclosure.
[0128] It should be understood that in the various embodiments of this disclosure, the sequence number of each implementation process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this disclosure.
[0129] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means two or more, unless otherwise explicitly specified.
[0130] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this disclosure is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this disclosure should be determined by the scope of the claims.
Claims
1. A method for training a cost model, characterized in that, The method includes: Obtain the hardware parameters of the first hardware platform and the first dataset, wherein the first dataset includes multiple scheduling primitive sequences and the actual execution performance parameters of each scheduling primitive sequence on the first hardware platform; A cost model is trained based on the hardware parameters and the first dataset. The performance of unlabeled candidate scheduling primitive sequences in the second dataset is predicted according to the cost model, and the sampling value score of the candidate scheduling primitive sequences is calculated based on the performance prediction results. The sampling value score represents the degree of contribution of the candidate scheduling primitive sequences to the cost model. The target scheduling primitive sequence is obtained by filtering from the second dataset based on the sampling value score; Obtain the actual execution performance parameters of the target scheduling primitive sequence on the first hardware platform, update the target scheduling primitive sequence and the target scheduling primitive sequence to the first dataset; and return to execute the step of training the cost model based on the hardware parameters and the first dataset until a preset stopping condition is met; the computational complexity of the sequence feature extraction layer in the cost model is linearly related to the sequence length of the input scheduling primitive sequence.
2. The method according to claim 1, characterized in that, The target scheduling primitive sequence obtained by filtering from the second dataset based on the sampling value score includes: Based on the distribution ratio of different operator types in the second dataset, determine the sampling quota corresponding to each type of operator; For each type of operator, a target scheduling sequence that meets the sampling quota is selected from the second dataset according to the preset order of the sampling value.
3. The method according to claim 1, characterized in that, The step of calculating the sampling value score of the candidate scheduling primitive sequence based on the performance prediction results includes: Obtain first information, which characterizes the degree of difference in prediction performance between unlabeled candidate scheduling primitive sequences and labeled scheduling primitive sequences; Obtain second information, which characterizes the degree of dispersion of the cost model's prediction performance on the unlabeled candidate scheduling primitive sequence; Based on the first information and the second information, the sampling value score of the candidate scheduling primitive sequence is determined.
4. The method according to claim 3, characterized in that, The step of determining the sampling value score of the candidate scheduling primitive sequence based on the first information and the second information includes: The sampling value score of the candidate scheduling primitive sequence is determined by multiplying the performance prediction result of the candidate scheduling primitive sequence with the first information and summing it with the second information.
5. The method according to any one of claims 1 to 4, characterized in that, The step of training the cost model based on the hardware parameters and the first dataset includes: A knowledge base model is obtained, which stores general knowledge information from multiple second hardware platforms. This knowledge base model uses cost models already trained on each second hardware platform as teacher models, and is obtained by transferring the knowledge trained on each second hardware platform to the knowledge base model through knowledge distillation. Using the cost model of the first hardware platform as the active column model, the active column model is trained based on the hardware parameters and the first dataset. During the training process of the active column model, the general knowledge information of the knowledge base model is used to assist in the training of the active column model.
6. The method according to claim 5, characterized in that, During the training process of the active column model, the general knowledge information of the knowledge base model is used to assist in the training of the active column model, including: During the training process of the active column model, the parameters of the knowledge base model are frozen, and the feature output of the knowledge base model is introduced into the active column model to assist in the training of the active column model using the general knowledge information of the knowledge base model.
7. The method according to claim 5, characterized in that, The method includes: After the active column model is trained, its parameters are frozen. Using the active column model as a teacher model, knowledge distillation is performed on the knowledge base model to transfer the knowledge information from the active column model to the knowledge base model.
8. A method for generating tensor programs, characterized in that, The method includes: Obtain the computational subgraph corresponding to the neural network model to be optimized, and generate multiple first scheduling primitive sequences corresponding to the computational subgraph; The plurality of first scheduling primitive sequences and the hardware parameters of the target hardware platform are input into the cost model obtained by the training method of the cost model according to any one of claims 1 to 7, so as to obtain the predicted execution performance parameters corresponding to each first scheduling primitive sequence. Based on the predicted execution performance parameters, a second scheduling primitive sequence is determined from the plurality of first scheduling primitive sequences, and a tensor program is generated to be executed on the target hardware platform based on the second scheduling primitive sequence.
9. A training device for a cost model, characterized in that, The device includes: The first acquisition module is used to acquire the hardware parameters of the first hardware platform and the first dataset, wherein the first dataset includes multiple scheduling primitive sequences and the actual execution performance parameters of each scheduling primitive sequence on the first hardware platform. The training module is used to train a cost model based on the hardware parameters and the first dataset; The performance prediction module is used to predict the performance of unlabeled candidate scheduling primitive sequences in the second dataset according to the cost model, and calculate the sampling value score of the candidate scheduling primitive sequences based on the performance prediction results. The sampling value score represents the degree of contribution of the candidate scheduling primitive sequences to the cost model. The filtering module is used to filter the target scheduling primitive sequence obtained from the second dataset based on the sampling value score; The first acquisition module is further configured to acquire the actual execution performance parameters of the target scheduling primitive sequence on the first hardware platform, update the target scheduling primitive sequence and the target scheduling primitive sequence to the first dataset; and return to execute the step of training the cost model based on the hardware parameters and the first dataset until a preset stopping condition is met; the computational complexity of the sequence feature extraction layer in the cost model is linearly related to the sequence length of the input scheduling primitive sequence.
10. The apparatus according to claim 9, characterized in that, The device includes: The determination module is used to determine the sampling quota corresponding to each type of operator based on the distribution ratio of different operator types in the second dataset; The filtering module is specifically used to filter out target scheduling sequences that meet the sampling quota from the second dataset according to a preset order of the sampling values for various operators.
11. The apparatus according to claim 10, characterized in that, The device includes: The first acquisition module is further configured to acquire first information, wherein the first information characterizes the degree of difference in prediction performance between the unlabeled candidate scheduling primitive sequence and the labeled scheduling primitive sequence; The first acquisition module is further configured to acquire second information, the second information being characterized by the degree of dispersion of the prediction performance of the cost model on the unlabeled candidate scheduling primitive sequence; The determining module is further configured to determine the sampling value score of the candidate scheduling primitive sequence based on the first information and the second information.
12. The apparatus according to claim 11, characterized in that, The device includes: The determination module is specifically used to determine the sampling value score of the candidate scheduling primitive sequence based on the product of the performance prediction result of the candidate scheduling primitive sequence and the first information, and the sum of the product and the second information.
13. A tensor program generation apparatus, characterized in that, The device includes: The second acquisition module is used to acquire the computational subgraph corresponding to the neural network model to be optimized, and generate multiple first scheduling primitive sequences corresponding to the computational subgraph; The input module is used to input the plurality of first scheduling primitive sequences and the hardware parameters of the target hardware platform into the cost model obtained by the training method of the cost model according to any one of claims 1 to 7, so as to obtain the predicted execution performance parameters corresponding to each first scheduling primitive sequence. The determination module is configured to determine a second scheduling primitive sequence from the plurality of first scheduling primitive sequences based on the predicted execution performance parameters, and generate a tensor program to be executed on the target hardware platform based on the second scheduling primitive sequence.
14. An electronic device, characterized in that, include: At least one processor; And, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 8.
15. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 8.