Deep learning compiler automatic tuning method based on reinforcement learning

Through a reinforcement learning-based method, combining multi-level feature similarity matching and variance threshold selection method, the scheduling space of the deep learning compiler is reduced, and the problems of inefficiency and unstable tuning results in the existing technology are solved, and more efficient deep learning model reasoning is achieved.

CN120215900APending Publication Date: 2025-06-27BEIJING UNIV OF CHEM TECH +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510332331.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing deep learning compiler automatic tuning technology is inefficient, the scheduling space is redundant, the tuning results are unstable, and it is difficult to achieve efficient inference under limited hardware resources.

Method used

Using reinforcement learning-based method, the search starting point is determined through a top-down multi-level feature similarity matching strategy, the variance threshold selection method is used to reduce the scheduling space, the DQN model is selected for agent state initialization, and the Monte Carlo sampling strategy and cost model are automatically tuned.

Benefits of technology

Reduce invalid search paths, reduce overall search complexity, improve resource utilization and adaptability, and achieve more efficient deep learning model inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215900A_ABST
    Figure CN120215900A_ABST
Patent Text Reader

Abstract

The invention relates to a deep learning compiler automatic tuning method based on reinforcement learning. The method comprises the steps of comparing a scheduling task with a historical scheduling task in a multi-level offline database through a multi-level feature similarity matching strategy from top to bottom to determine a search starting point scheduling configuration; generating a feature matrix according to the feature vector corresponding to each type of scheduling task, and screening out a significant dimension reduction search space through a variance threshold selection method; based on a reinforcement learning method, selecting a DQN model, generating a random action through a Monte Carlo sampling strategy, and changing scheduling configuration according to the random action; and evaluating the changed scheduling configuration performance by using a cost model, and constructing a reward function based on an evaluation result to complete automatic tuning. By constructing a multi-level off-line database and utilizing feature similarity matching from top to bottom, starting point scheduling configuration can be quickly retrieved and determined; and a variance threshold selection method is adopted to carry out dimension reduction on feature dimension vectors, so that the retrieval efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning compiler optimization, and particularly to an automatic tuning method for deep learning compilers based on reinforcement learning. Background Art

[0002] With the rapid development of deep learning technology, it has been increasingly widely applied in fields such as image recognition and natural language processing, especially in scenarios with high response requirements, such as autonomous driving, mobile devices, and on-orbit processing. However, these application scenarios are often limited by hardware resources. How to achieve efficient inference of deep learning models under limited resources has become an urgent technical challenge. To address this challenge, various model lightweight methods have emerged, such as model pruning, knowledge distillation, and quantization, which effectively reduce the requirements of models for hardware resources by reducing the number of model parameters and computational volume, but often at the cost of sacrificing model accuracy. In contrast, deep learning compiler technology improves the inference efficiency by optimizing the execution instructions of the model, achieving efficient inference without sacrificing accuracy.

[0003] Deep learning compiler technology can be divided into two categories: general compilers and special compilers. Among them, special deep learning compilers such as OpenVINO, TensorRT, and Roller are optimized for specific hardware architectures (such as Intel CPUs and NVIDIA GPUs); general compilers such as TVM (Tensor Virtual Machine) have become important tools for deploying deep learning models across various hardware platforms due to their wide applicability and open source nature. Deep learning compilers achieve efficient inference of deep learning models on different hardware architectures by converting tensor operations in deep learning models into instructions executable on specific hardware platforms. During the compilation process of deep learning models by deep learning compilers, a key link is automatic tuning (Auto-tuning), and its core idea is to discover the best scheduling strategies for various operators suitable for the target hardware through automated search algorithms, thereby maximizing the utilization rate of hardware resources and improving computational efficiency.

[0004] The existing automatic tuning process usually relies on random sampling to determine the search starting point. This method is not only inefficient, consuming a large amount of time and hardware resources, but also introduces a large degree of uncertainty, making it difficult to ensure the stability of the tuning results; the automatic tuning process mostly uses heuristic algorithms for search. These algorithms are inefficient in high-dimensional scheduling spaces and are prone to falling into local optimal solutions, making it difficult to discover the global optimal scheduling configuration; it fails to effectively reduce the scheduling search space, resulting in a large number of invalid or low-optimization-potential scheduling configurations being generated during the search process, increasing the computational burden of the search and reducing the overall efficiency.

[0005] Therefore, traditional automatic tuning techniques have problems such as low efficiency, redundant scheduling space, and unstable tuning results. Summary of the Invention

[0006] Based on this, to solve the above technical problems, a method for automatically tuning a deep learning compiler based on reinforcement learning is provided, which can reduce invalid search paths, reduce the complexity of overall search, and improve resource utilization and adaptability.

[0007] A method for automatically tuning a deep learning compiler based on reinforcement learning, the method comprising:

[0008] Obtain a scheduling task, and compare the scheduling task with historical scheduling tasks in a multi-level offline database through a top-down multi-level feature similarity matching strategy, and determine a starting scheduling configuration for search according to the comparison result;

[0009] Classify the scheduling tasks according to task types, generate a feature matrix according to the feature vectors corresponding to each type of scheduling task, screen each feature dimension vector in the feature matrix by a variance threshold selection method, extract significant dimensions, and obtain a reduced scheduling space;

[0010] Based on the reinforcement learning method, select a DQN model according to the feature vector and significant dimensions corresponding to the scheduling task, determine the initial state of the agent according to the starting scheduling configuration and the reduced scheduling space, generate random actions through a Monte Carlo sampling strategy, and change the scheduling configuration according to the random actions;

[0011] Perform performance evaluation on the changed scheduling configuration through a cost model, and construct a reward function according to the evaluation result to complete automatic tuning.

[0012] In one embodiment, the method further comprises:

[0013] Collect various deep learning frameworks, and extract various scheduling tasks from each of the deep learning frameworks;

[0014] Perform tuning compilation on each of the scheduling tasks, and collect the target data of the scheduling tasks after tuning compilation;

[0015] Construct a multi-level offline database based on the target data of the scheduling tasks.

[0016] In one embodiment, comparing the scheduling task with historical scheduling tasks in a multi-level offline database through a top-down multi-level feature similarity matching strategy, and determining a starting scheduling configuration for search according to the comparison result, includes:

[0017] Extract the target upper-level features in the scheduling task and traverse the reference upper-level features in the multi-level offline database;

[0018] If the similarity between the target upper-level feature and the reference upper-level feature is greater than the first threshold, extract the target middle-level feature in the scheduling task and compare it with the reference middle-level feature in the multi-level offline database; otherwise, select the default scheduling configuration as the starting scheduling configuration;

[0019] If the similarity between the target middle-level feature and the reference middle-level feature is greater than the second threshold, extract the target lower-level feature in the scheduling task and compare it with the reference lower-level feature in the multi-level offline database; otherwise, select the default scheduling configuration as the starting scheduling configuration;

[0020] If the similarity between the target lower-level feature and the reference lower-level feature is greater than the third threshold, select the starting scheduling configuration from the multi-level offline database; otherwise, select the default scheduling configuration as the starting scheduling configuration.

[0021] In one embodiment, classify the scheduling tasks according to the task type, and generate a feature matrix according to the feature vectors corresponding to each type of scheduling task, including:

[0022] Collect all scheduling tasks during the compilation and tuning of the deep learning model, and classify the scheduling tasks according to the task type;

[0023] Obtain the feature vectors corresponding to each type of scheduling task, and determine the indexes of each feature vector and the indexes of the dimensions within each feature vector to generate a feature matrix;

[0024] Among them, the feature vectors of the same type of scheduling task have the same feature vector dimension.

[0025] In one embodiment, screen each feature dimension vector in the feature matrix by the variance threshold selection method to extract significant dimensions and obtain a reduced scheduling space, including:

[0026] Calculate the variance of each feature dimension vector in the feature matrix using the variance threshold selection method;

[0027] Set a variance threshold, compare the variance with the variance threshold, and screen and extract significant dimensions from each feature dimension vector according to the comparison result;

[0028] Retain the corresponding values of the significant dimensions in each feature dimension vector to form a reduced scheduling space.

[0029] In one embodiment, select a DQN model based on the feature vectors corresponding to the scheduling tasks and significant dimensions, including:

[0030] Combine each feature vector corresponding to the scheduling task with the significant dimensions and encode them into respective target feature vectors;

[0031] Calculate the similarity between each of the target feature vectors and the scheduling task using Spearman similarity, and compare the similarity with a reference threshold;

[0032] Determine the DQN model according to the comparison result.

[0033] In one embodiment, generate random actions through a Monte Carlo sampling strategy and change the scheduling configuration state according to the random actions, including:

[0034] Generate random actions through a Monte Carlo sampling strategy and represent the random actions as action feature vectors corresponding to the scheduling configuration;

[0035] Based on the action feature vectors, indicate the adjustment direction and amplitude of the scheduling configuration parameters to complete the change of the scheduling configuration.

[0036] In one embodiment, perform performance evaluation on the changed scheduling configuration through a cost model and construct a reward function according to the evaluation result, including:

[0037] Perform performance evaluation on the scheduling configuration corresponding to the scheduling configuration before the change through a cost model to obtain performance evaluation metrics;

[0038] Compare the performance evaluation metrics with the evaluation result and construct a reward function according to the comparison result.

[0039] In one embodiment, the method further includes:

[0040] Guide the agent to converge to the optimal scheduling configuration based on the reward function and record the reward increment;

[0041] When the cumulative reward increment is lower than a preset threshold or the search reaches the maximum number of iterations, the search process terminates and the result is output.

[0042] In one embodiment, the method further includes:

[0043] Store the scheduling configuration state, random actions, and rewards generated by the reward function during each interaction process as historical data in a buffer;

[0044] When performing model training, use a random sampling mechanism to extract historical data from the buffer;

[0045] Complete model training using the historical data.

[0046] The above-mentioned automatic tuning method for deep learning compilers based on reinforcement learning can quickly retrieve and determine the starting scheduling configuration by constructing a multi-level offline database and using top-down feature similarity matching, which can reduce the invalid search paths; the variance threshold selection method is used to screen out the significant dimensions, and by ignoring the unimportant dimensions, the spatial complexity is greatly reduced while effectively retaining the representativeness of the subspace, thereby improving the search efficiency; the reinforcement learning algorithm is used to iteratively optimize the scheduling configuration, so as to achieve automatic tuning optimization, which can improve the data utilization rate and training stability, and improve the resource utilization rate and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 FIG. is an application environment diagram of the automatic tuning method for deep learning compilers based on reinforcement learning in an embodiment;

[0048] Figure 2 FIG. is a schematic flowchart of the automatic tuning method for deep learning compilers based on reinforcement learning in an embodiment;

[0049] Figure 3 FIG. is a schematic flowchart of the top-down multi-level feature similarity matching strategy in an embodiment;

[0050] Figure 4 FIG. is a structural block diagram of a reinforcement learning architecture in an embodiment;

[0051] Figure 5 FIG. is a schematic diagram of an agent changing the current state according to an action in an embodiment;

[0052] Figure 6 FIG. is a schematic flowchart of the automatic tuning method for deep learning compilers based on reinforcement learning in another embodiment;

[0053] Figure 7 FIG. is an internal structure diagram of an agent in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0055] It can be understood that the terms "first", "second", etc. used in the present application may be used herein to describe thresholds, but these thresholds are not limited by these terms. These terms are only used to distinguish the first threshold from another threshold. For example, without departing from the scope of the present application, the first threshold may be referred to as the second threshold, and similarly, the second threshold may be referred to as the first threshold. Both the first threshold and the second threshold are thresholds, but they are not the same threshold.

[0056] The automatic tuning method of the deep learning compiler based on reinforcement learning provided by the embodiments of the present application can be applied to an application environment as Figure 1 shown. As Figure 1 shown, the application environment includes an agent 110. The agent 110 can obtain scheduling tasks, compare the scheduling tasks with historical scheduling tasks in a multi-level offline database through a top-down multi-level feature similarity matching strategy, and determine the starting scheduling configuration for searching according to the comparison result; the agent 110 can classify the scheduling tasks according to the task type, generate a feature matrix according to the feature vectors corresponding to each type of scheduling task, screen each feature dimension vector in the feature matrix by the variance threshold selection method, extract significant dimensions, and obtain a reduced scheduling space; the agent 110 can select a DQN model based on the reinforcement learning method according to the feature vectors and significant dimensions corresponding to the scheduling tasks, determine the initial state of the agent according to the starting scheduling configuration and the reduced scheduling space, generate random actions through the Monte Carlo sampling strategy, and change the scheduling configuration according to the random actions; the agent 110 can perform performance evaluation on the changed scheduling configuration through a cost model, construct a reward function according to the evaluation result, and complete automatic tuning. Among them, the agent 110 can be, but is not limited to, various devices such as personal computers, laptop computers, smartphones, robots, unmanned aerial vehicles, and tablet computers.

[0057] In one embodiment, as Figure 2 shown, an automatic tuning method of the deep learning compiler based on reinforcement learning is provided, including the following steps:

[0058] Step 202, obtain scheduling tasks, compare the scheduling tasks with historical scheduling tasks in a multi-level offline database through a top-down multi-level feature similarity matching strategy, and determine the starting scheduling configuration for searching according to the comparison result.

[0059] After generating a complete scheduling space according to the scheduling tasks and manual templates at the backend of the deep learning compiler, a pre-constructed multi-level offline database can be used to quickly compare the features of the current scheduling tasks with the historical scheduling tasks in the database through a top-down multi-level feature similarity matching strategy, and screen out the result most similar to the current task as the starting point for searching.

[0060] Among them, carefully planning the starting point for searching in the scheduling space allows the deep learning compiler to use past experience to provide a starting point close to the optimal region for the compilation and tuning process of the deep learning model, thereby improving the tuning efficiency and result quality.

[0061] In this embodiment, the deep learning model used can be the classic network AlexNet. AlexNet is a pioneering deep learning model composed of an 8-layer network, including 5 convolutional layers, 3 fully connected layers, as well as technologies such as ReLU activation and dropout. The network has a relatively low complexity and high representativeness, making it very suitable for data analysis and result display.

[0062] Specifically, scheduling tasks can be extracted from the AlexNet deep learning model, and the extracted task data is shown in the following table:

[0063]

[0064]

[0065] In this embodiment, the operators to be optimized mainly include two categories: convolution and fully connected. Among them, the convolution operator is further subdivided into two calculation methods. According to data such as the target hardware and parameters, scheduling task name, scheduling task parameters, and optimal scheduling configuration, a multi-level offline database can be constructed to facilitate the subsequent rapid retrieval of the optimal scheduling configuration.

[0066] In one embodiment, a method for automatically tuning a deep learning compiler based on reinforcement learning may further include the process of constructing a multi-level offline database. The specific process includes: collecting various deep learning frameworks, extracting various scheduling tasks from each deep learning framework; performing tuning compilation on each scheduling task, and collecting the target data of the scheduling tasks after tuning compilation; constructing a multi-level offline database based on the target data of the scheduling tasks.

[0067] The construction of the multi-level offline database first requires extracting different scheduling tasks from various popular deep learning architectures, performing tuning compilation, and collecting and organizing key data related to the scheduling tasks, including the target hardware and parameters, scheduling task name, scheduling task parameters, and optimal scheduling configuration, etc. Then, based on the collected data, a multi-level feature structure is constructed to facilitate subsequent rapid retrieval at each level. In this embodiment, taking a three-level example, the multi-level feature structure of the three-level offline database is shown in the following table:

[0068]

[0069] After the construction of the multi-level offline database is completed, a top-down multi-level feature similarity matching strategy can be adopted to layer by layer screen out the optimal starting scheduling configuration.

[0070] In one embodiment, a method for automatically tuning a deep learning compiler based on reinforcement learning may further include a process of screening out a starting point scheduling configuration. The specific process includes: extracting the target upper-level features in the scheduling task and traversing the reference upper-level features in the multi-level offline database; if the similarity between the target upper-level features and the reference upper-level features is greater than the first threshold, then extracting the target middle-level features in the scheduling task and comparing them with the reference middle-level features in the multi-level offline database; otherwise, selecting the default scheduling configuration as the starting point scheduling configuration; if the similarity between the target middle-level features and the reference middle-level features is greater than the second threshold, then extracting the target lower-level features in the scheduling task and comparing them with the reference lower-level features in the multi-level offline database; otherwise, selecting the default scheduling configuration as the starting point scheduling configuration; if the similarity between the target lower-level features and the reference lower-level features is greater than the third threshold, then selecting the starting point scheduling configuration from the multi-level offline database; otherwise, selecting the default scheduling configuration as the starting point scheduling configuration.

[0071] Specifically, in this embodiment, different algorithms can be designed according to the characteristics of each layer of features to calculate the similarity between the input data and the corresponding layer of features in the database. If the maximum similarity calculated for the current layer does not exceed the preset threshold, or there is no matching similar scheduling task in the offline database, the default scheduling configuration is used as the search starting point. Otherwise, the value corresponding to the feature with the highest similarity is selected as the input for the next layer of screening, and the process progresses layer by layer until the final optimal scheduling configuration is determined, or the default scheduling configuration is reverted according to the situation.

[0072] In this embodiment, taking a three-level offline database as an example, the similarities of the features in the upper level, middle level, and lower level can be compared in sequence to determine the optimal scheduling configuration.

[0073] Specifically, after constructing the multi-level offline database, when there is a new scheduling task that needs to be tuned, a top-down multi-level feature similarity matching strategy is adopted, as Figure 3 shown. First, extract the target hardware and parameter features F1 of the upper level of the new scheduling task and traverse all the upper-level features F i in the offline database, where i is the serial number. Calculate the similarity between F1 and F i in sequence using cosine similarity, and select the maximum value as the result of the retrieval for this layer and compare it with the threshold T1. If it is smaller than the threshold, it means that there is no task in the offline database that matches the new scheduling task, and the retrieval process ends, and the default scheduling configuration is selected as the starting point of the search; otherwise, then extract the task name feature F2 of the middle level of the new scheduling task, and use the normalized edit distance to compare it with all the middle-level features F jCalculate the similarity, select the maximum value and compare it with the threshold T2. If it is greater than the threshold, finally extract the task parameter feature F3 at the middle level in the new scheduling task, and use the normalized Euclidean distance to sequentially compare with all the lower-level features F in the maximum similarity at the middle level k Calculate the similarity, select the maximum value and compare it with the threshold T3. If it is greater than the threshold, obtain the optimal scheduling configuration of the most similar task as the search starting point of the new scheduling task.

[0074] Among them, when calculating the similarity between the n-dimensional input feature vector A and a certain n-dimensional feature vector B in the three-level offline database, the algorithms for different levels are as follows:

[0075] For the features in the upper level, considering the strong representation ability of the target hardware and parameter features, the cosine similarity is used to calculate the similarity metric, and the calculation formula can be expressed as:

[0076]

[0077] For the features in the middle level, considering that the task name feature has high requirements for accuracy, the normalized edit distance is used as the similarity metric, and the calculation formula can be expressed as: Among them, EditDistance(A,B) represents the minimum number of edits to convert A to B.

[0078] For the features in the lower level, considering that the task parameter features are numerical features of different dimensions, the normalized Euclidean distance is used to provide a consistent similarity metric, and the calculation formula can be expressed as:

[0079]

[0080] In this embodiment, by constructing a pre-built multi-level offline database for storing the optimal scheduling configurations of various historical scheduling tasks, based on the top-down feature similarity matching strategy at multiple levels, through feature similarity matching, quickly retrieve the historical task results most similar to the current task, and use this as the starting point for reinforcement learning search, thereby effectively avoiding the inefficient process of starting from scratch in traditional methods and significantly improving the accuracy and tuning efficiency of starting point planning.

[0081] Step 204: Classify the scheduling tasks according to the task type, generate a feature matrix based on the feature vectors corresponding to each type of scheduling task, screen each feature dimension vector in the feature matrix by the variance threshold selection method, extract the significant dimensions, and obtain the reduced scheduling space.

[0082] After reasonably planning the search starting point, it is possible to further perform effective screening of feature dimensions for a scheduling space with high complexity. Each scheduling configuration in the scheduling space can be represented as a feature vector, where the value of each dimension corresponds to the adjustable parameter in the manual template. By performing an analysis of variance on the feature vectors of the scheduling configurations, significant dimensions with high information content are screened out, reducing the scale of the scheduling space, removing redundant scheduling configurations, ensuring the representativeness of the subspace while reducing the space complexity, improving the search efficiency, and significantly enhancing the efficiency of the compilation optimization process.

[0083] In one embodiment, a method for automatic tuning of a deep learning compiler based on reinforcement learning may further include a process of organizing and generating a feature matrix. The specific process includes: collecting all scheduling tasks during the compilation optimization process of the deep learning model, and classifying the scheduling tasks according to the task type; obtaining the feature vectors corresponding to each type of scheduling task, and determining the indexes of each feature vector and the indexes of the dimensions within each feature vector to generate a feature matrix; wherein, the feature vectors of the same type of scheduling task have the same feature vector dimensions.

[0084] Among them, based on the fact that the feature vectors of the same type of scheduling task have the same number of feature vector dimensions, first, all scheduling tasks during the compilation optimization process of the deep learning model can be classified according to the task type. The scheduling tasks can be mainly divided into three categories. For each type of scheduling task, organize its corresponding feature vectors into a feature matrix X ij , where i represents the index of the feature vector and j represents the index of the dimension within the feature vector.

[0085] Then, the variance threshold selection method can be used to screen out significant dimensions with higher information content from the feature vectors, ignoring unimportant dimensions, thereby greatly reducing the complexity of the scheduling space while effectively retaining the representativeness of the subspace, laying a foundation for the efficient training of subsequent reinforcement learning.

[0086] Specifically, in one embodiment, a method for automatic tuning of a deep learning compiler based on reinforcement learning may further include a process of reducing the scheduling space using the variance threshold selection method. The specific process includes: calculating the variance of each feature dimension vector in the feature matrix using the variance threshold selection method; setting a variance threshold, comparing the variance with the variance threshold, and screening and extracting significant dimensions from each feature dimension vector according to the comparison result; retaining the corresponding values of the significant dimensions in each feature dimension vector to form a reduced scheduling space.

[0087] Specifically, the variance threshold selection method can be used. For the obtained feature matrix X ij in each feature dimension vector X j, calculate its variance to evaluate the contribution of this dimension to the feature vector in the current category. The calculation formula can be expressed as: Among them, Among them, j is the index of the dimension in the feature vector, i is the index representing the category of the feature vector, N is the number of feature vectors in this category, and u j is the average value of the j-th dimension of all feature vectors in this category. Among them, variance reflects the difference and information content of features among samples: the larger the variance, the stronger the discrimination ability and higher information content of the feature, and the higher the optimization potential; the smaller the variance, the weaker the discrimination effect of the feature, less information content, and it may be a redundant feature.

[0088] Next, in this embodiment, a variance threshold θ can be set to select the top k dimensions with variances exceeding the threshold from all feature dimension sets S as significant dimensions, retain these significant dimensions, and ignore other dimensions to form a new representative subspace. The formula for this process can be expressed as: S = {i|Var(X j ) > θ, i ≤ k)}. Among them, the variance threshold reduction scheduling space process data can be as shown in the following table:

[0089] Category Dimensional variance Parameter θ Parameter k Significant dimension Id Conv2d_nchw.cuda [0,1735,152,13,0,0] 10 4 [5,3,2,1] Conv2d_nchw_winograd.cuda [20444,49,49,25,0,0,0,0] 10 4 [3,1,2,0]

[0090] Step 206, based on the reinforcement learning method, select a DQN model according to the feature vector and significant dimensions corresponding to the scheduling task, determine the initial state of the agent according to the starting scheduling configuration and the reduced scheduling space, generate random actions through the Monte Carlo sampling strategy, and change the scheduling configuration according to the random actions.

[0091] In the reduced scheduling space, a reinforcement learning framework can be used to implement the search for the optimal scheduling configuration. The architecture of reinforcement learning is as Figure 4As shown below. Specifically: First, the agent uses the feature vector after dimensionality reduction as the state of reinforcement learning, and each state corresponds to a potential scheduling configuration. Then, through the Monte Carlo sampling strategy, actions are randomly generated. The actions indicate the adjustment direction and amplitude of the scheduling configuration parameters, driving the transition between states, that is, realizing the dynamic conversion between different scheduling configurations. After executing the actions, the state changes, that is, it transfers from one scheduling configuration to another. At the same time, the agent will receive a reward, which is based on the performance of the new scheduling configuration on the target hardware. Through continuous interaction with the environment, the agent learns and optimizes the strategy to maximize the long-term cumulative reward, so as to find the scheduling configuration with the optimal performance. Subsequently, a cost model is used to predict the performance of the scheduling configuration, and a reward function is constructed based on this to guide the reinforcement learning agent to learn the optimal strategy. In the case of high task similarity, similar tasks are allowed to share and reuse the parameters of the policy network, thereby reducing repeated training and accelerating the convergence speed of the model using the existing learning results. The experience replay mechanism is introduced in the reinforcement learning process to improve data utilization and training stability. At the same time, the agent gradually approaches the optimal scheduling configuration through continuous exploration and feedback, thus significantly improving the model operation efficiency.

[0092] The reinforcement learning method is used to search the scheduling space. The scheduling configuration search problem is modeled as a deterministic Markov decision process (MDP), and the optimal scheduling configuration is quickly searched through the Deep Q-Learning (DQN) algorithm. Through the interaction between the agent and the environment, the optimal strategy is learned to maximize the reward, and the similarity between scheduling tasks is used to achieve model sharing, accelerating training and convergence.

[0093] In one embodiment, a method for automatically tuning a deep learning compiler based on reinforcement learning may further include the process of selecting a DQN model. The specific process includes: combining each feature vector corresponding to the scheduling task with the significant dimensions and encoding them into each target feature vector; calculating the similarity between each target feature vector and the scheduling task using the Spearman similarity, and comparing the similarity with a reference threshold; determining the DQN model according to the comparison result.

[0094] Specifically, the features of the scheduling task and the extracted significant dimensions can be combined first and encoded into a feature vector. To measure the correlation between scheduling tasks, the Spearman similarity formula is used as follows: where d iIt represents the rank difference of the i-th feature, and n is the total number of eigenvalues. According to the calculation result of the similarity between scheduling tasks, a threshold standard is set, and tasks with similarity exceeding this threshold are classified into the same category. The scheduling tasks are mainly classified into three categories: Conv2d, Conv2d_winograd, and Dense. That is, only three DQN models are needed to complete the compilation and optimization of AlexNet, avoiding repeated training between independent tasks, thus effectively improving the training efficiency of the model and significantly accelerating the learning and convergence process of the agent.

[0095] Next, in an embodiment, a method for automatic tuning of a deep learning compiler based on reinforcement learning may further include a process of changing the scheduling configuration. The specific process includes: generating a random action through a Monte Carlo sampling strategy, and representing the random action as an action feature vector corresponding to the scheduling configuration; based on the action feature vector, indicating the adjustment direction and amplitude of the scheduling configuration parameters, and completing the change of the scheduling configuration.

[0096] After selecting the DQN model, the agent first enters the initial state s0. Among them, the initial state can be represented as a feature vector X of the scheduling configuration, which is jointly determined by the planned search starting point and the reduced scheduling space. Subsequently, the agent generates a random action a according to the Monte Carlo sampling strategy, which can also be represented as a feature vector, corresponding to the state, indicating the adjustment direction and amplitude of the scheduling configuration parameters. Then the agent changes the current state according to the action, so as to realize the transition from state s i to state s j as shown in Figure 5 Subsequently, the agent generates a random action according to the Monte Carlo sampling strategy. The action can be used to indicate the adjustment direction and amplitude of the scheduling configuration parameters; the agent changes the current state according to the action, so as to realize the transition from one scheduling configuration state to another scheduling configuration state.

[0097] Step 208, perform performance evaluation on the changed scheduling configuration through a cost model, and construct a reward function according to the evaluation result to complete automatic tuning.

[0098] In an embodiment, a method for automatic tuning of a deep learning compiler based on reinforcement learning may further include a process of constructing a reward function. The specific process includes: performing performance evaluation on the scheduling configuration before the change through a cost model to obtain performance evaluation indicators; comparing the performance evaluation indicators with the evaluation results, and constructing a reward function according to the comparison results.

[0099] After the state transfer is completed, the cost model XGBoost can be further called to evaluate the performance of the scheduling configuration corresponding to the new state, and the evaluation result is compared with the performance metrics of the current configuration, so as to construct the reward function r as the core basis for measuring the quality of actions. The higher its value, the closer the scheduling configuration in the new state is to the optimal solution. With the reward function r, the agent can preferentially explore the configuration paths that can significantly improve performance, thereby continuously optimizing the scheduling configuration and gradually approaching the optimal performance solution.

[0100] In one embodiment, a method for automatically tuning a deep learning compiler based on reinforcement learning may further include a process of outputting the optimal scheduling configuration. The specific process includes: guiding the agent to converge to the optimal scheduling configuration based on the reward function and recording the reward increment; when the cumulative reward increment is lower than a preset threshold, or the search reaches the maximum number of iterations, the search process terminates and the result is output.

[0101] As the iteration process progresses, the reward function r gradually guides the agent to converge to the optimal scheduling configuration. When the cumulative reward increment of the agent is lower than the preset threshold T, or the search reaches the maximum number of iterations trials, the search process terminates, and the system outputs the scheduling configuration with the highest performance as the final result.

[0102] In one embodiment, a method for automatically tuning a deep learning compiler based on reinforcement learning may further include a process of model training. The specific process includes: storing the scheduling configuration state, random actions, and rewards generated by the reward function during each interaction process as historical data in a buffer; when performing model training, using a random sampling mechanism to extract historical data from the buffer; and completing model training using the historical data.

[0103] To further improve the training efficiency of reinforcement learning, the data of the state s, action a, reward r, and new state s' generated during each interaction process are stored in the experience replay buffer D. When the agent updates the policy network φ later, a random sampling mechanism is used to extract historical experience data (s i ,a i ,r i ,s i+1 ) from the buffer. The random sampling mechanism improves the stability of model training by reducing data correlation and significantly increases the reuse rate of data. The policy network updates the expected return of actions in each iteration, thereby continuously enhancing the agent's evaluation ability and decision-making efficiency for various scheduling configurations.

[0104] In one embodiment, a method for automatically tuning a deep learning compiler based on reinforcement learning is provided as Figure 6As shown in the figure, it mainly includes a search starting point planning part S1, a scheduling space search part S2, and a reinforcement learning search part S3, where:

[0105] Search starting point planning part S1: First, a pre-built multi-level offline database is constructed to store the optimal scheduling configurations of various historical scheduling tasks. Through feature similarity matching, the historical task results most similar to the current task are quickly retrieved and used as the starting point for reinforcement learning search, thus avoiding inefficient search from scratch;

[0106] Scheduling space search part S2: By performing variance analysis on the feature vectors of the scheduling configurations, significant dimensions with high information content are screened out, greatly reducing the space complexity while effectively retaining the representativeness of the subspace;

[0107] Reinforcement learning search part S3: A deep reinforcement learning algorithm is used to search for scheduling configurations. The agent regards each feature vector representing a scheduling configuration as a state and generates random actions according to the Monte Carlo sampling strategy. This action indicates the adjustment direction and amplitude of the scheduling configuration parameters. After the action is executed, the state changes, that is, it transfers from one scheduling configuration to another. At the same time, the agent will receive a reward based on the performance of the new scheduling configuration on the target hardware. Through continuous interaction with the environment, the agent learns and optimizes the strategy to maximize the long-term cumulative reward, thereby finding the scheduling configuration with the optimal performance. To improve the learning efficiency, the present invention introduces a model sharing mechanism. In the case of high task similarity, similar tasks are allowed to share and reuse the parameters of the policy network, thereby reducing repeated training and using the existing learning results to accelerate the convergence speed of the model.

[0108] It should be understood that although the steps in the above flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0109] In one embodiment, a computer device is provided. This computer device can be an agent, and its internal structure diagram can be as Figure 7As shown in the figure. The agent includes a processor, a memory, a network interface, a display screen, and an input device connected by a system bus. Among them, the processor of the agent is used to provide computing and control capabilities. The memory of the agent includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the agent is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it implements a method for automatically tuning a deep learning compiler based on reinforcement learning. The display screen of the agent can be a liquid crystal display screen or an electronic ink display screen. The input device of the agent can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the shell of the agent, or an external keyboard, touchpad, or mouse, etc.

[0110] Those skilled in the art can understand that Figure 7 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the agent to which the solution of this application is applied. The specific agent may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0111] In one embodiment, an agent is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, it implements the steps of a method for automatically tuning a deep learning compiler based on reinforcement learning.

[0112] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by the processor, it implements the steps of a method for automatically tuning a deep learning compiler based on reinforcement learning.

[0113] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0114] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0115] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A deep learning compiler automatic tuning method based on reinforcement learning, characterized in that: The method comprises: Obtain the scheduling task, and compare the scheduling task with the historical scheduling tasks in the multi-level offline database through a top-down multi-level feature similarity matching strategy, and determine the starting scheduling configuration for the search based on the comparison results; The scheduling tasks are classified according to the task type, and a feature matrix is ​​generated according to the feature vector corresponding to each type of scheduling task, and each feature dimension vector in the feature matrix is ​​screened by a variance threshold selection method to extract a significant dimension to obtain a reduced scheduling space; Based on the reinforcement learning method, a DQN model is selected according to the feature vector and significant dimension corresponding to the scheduling task, the initial state of the agent is determined according to the starting scheduling configuration and the reduced scheduling space, random actions are generated through the Monte Carlo sampling strategy, and the scheduling configuration is changed according to the random actions; The performance of the changed scheduling configuration is evaluated through the cost model, and a reward function is constructed based on the evaluation results to complete automatic tuning.

2. The automatic tuning method for deep learning compiler based on reinforcement learning according to claim 1 is characterized in that: The method further comprises: Collect various deep learning frameworks, and extract various scheduling tasks from each of the deep learning frameworks; Perform tuning and compilation on each of the scheduling tasks, and collect the scheduling task target data after tuning and compilation; Based on the scheduling task target data, a multi-level offline database is constructed.

3. The automatic tuning method of deep learning compiler based on reinforcement learning according to claim 1, characterized in that: Through a top-down multi-level feature similarity matching strategy, the scheduling task is compared with the historical scheduling tasks in the multi-level offline database, and the starting scheduling configuration of the search is determined based on the comparison results, including: Extracting target upper-level features in the scheduling task, and traversing reference upper-level features in the multi-level offline database; If the similarity between the target upper-level feature and the reference upper-level feature is greater than a first threshold, extracting the target middle-level feature in the scheduling task and comparing it with the reference middle-level feature in the multi-level offline database; otherwise, selecting the default scheduling configuration as the starting scheduling configuration; If the similarity between the target mid-level feature and the reference mid-level feature is greater than a second threshold, extracting the target lower-level feature in the scheduling task and comparing it with the reference lower-level feature in the multi-level offline database; otherwise, selecting the default scheduling configuration as the starting scheduling configuration; If the similarity between the target lower-level feature and the reference lower-level feature is greater than a third threshold, a starting scheduling configuration is selected from the multi-level offline database; otherwise, a default scheduling configuration is selected as the starting scheduling configuration.

4. The automatic tuning method of deep learning compiler based on reinforcement learning according to claim 1 is characterized in that: The scheduling tasks are classified according to the task type, and a feature matrix is ​​generated according to the feature vector corresponding to each type of scheduling task, including: Collect all scheduling tasks during the deep learning model compilation and tuning process, and classify the scheduling tasks according to task types; Obtain the feature vector corresponding to each type of scheduling task, determine the index of each feature vector and the index of the dimension within each feature vector, and generate a feature matrix; Among them, the feature vectors of the same scheduling task type have the same feature vector dimension.

5. The automatic tuning method for deep learning compiler based on reinforcement learning according to claim 4 is characterized in that: Each feature dimension vector in the feature matrix is ​​screened by using the variance threshold selection method to extract significant dimensions, thereby obtaining a reduced scheduling space, including: Calculate the variance of each feature dimension vector in the feature matrix using a variance threshold selection method; Setting a variance threshold, comparing the variance with the variance threshold, and screening and extracting significant dimensions from each feature dimension vector according to the comparison result; The corresponding values ​​of the significant dimensions in each feature dimension vector are retained to form the reduced scheduling space.

6. The automatic tuning method for deep learning compiler based on reinforcement learning according to claim 1, characterized in that: Select the DQN model based on the feature vector and significant dimension corresponding to the scheduling task, including: Combine the feature vectors and significant dimensions corresponding to the scheduling tasks and encode them into target feature vectors; Calculate the similarity between each of the target feature vectors and the scheduling task using Spearman similarity, and compare the similarity with a reference threshold; According to the comparison results, the DQN model is determined.

7. The automatic tuning method for deep learning compiler based on reinforcement learning according to claim 1, characterized in that: Generate random actions through Monte Carlo sampling strategy, and change the scheduling configuration according to the random actions, including: Generate random actions through a Monte Carlo sampling strategy, and characterize the random actions as action feature vectors corresponding to the scheduling configuration; The scheduling configuration parameters are adjusted in a direction and magnitude based on the action feature vector, thereby completing the change of the scheduling configuration.

8. The automatic tuning method of deep learning compiler based on reinforcement learning according to claim 1, characterized in that: The performance of the changed scheduling configuration is evaluated through the cost model, and a reward function is constructed based on the evaluation results, including: The performance of the scheduling configuration before the change is evaluated through the cost model to obtain the performance evaluation index; The performance evaluation index is compared with the evaluation result, and a reward function is constructed according to the comparison result.

9. The automatic tuning method of deep learning compiler based on reinforcement learning according to claim 8, characterized in that: The method further comprises: Guiding the agent to converge to the optimal scheduling configuration based on the reward function and recording the reward increment; When the cumulative reward increment is lower than a preset threshold, or the search reaches a maximum number of iterations, the search process is terminated and the result is output.

10. The automatic tuning method of deep learning compiler based on reinforcement learning according to claim 1, characterized in that: The method further comprises: The scheduling configuration state, random actions, and rewards generated by the reward function generated during each interaction are stored in the buffer as historical data; When training the model, a random sampling mechanism is used to extract historical data from the buffer; Model training is completed using the historical data.

Citation Information

Cited By

  • Multi-agent based controller policy generation method and generation system

    CN122546613A