A method for evaluating execution time of a deep learning model
Patent Information
- Application Number
- CN202311203914.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-18
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-09-18
AI Technical Summary
但是获取Cost(*)需要在实际硬件上运行,由此也带来了两个问题:第一是深度学习模型的执行时间一般较长,而搜索空间一般较大,这会导致寻找该问题最优解的时间过长,而且大部分时间都消耗在深度学习模型执行而不是搜索最优解的过程上;第二是深度学习模型实际执行时间可能会随着深度学习模型外的其他临时环境因素影响,导致时间记录不准确,从而干扰找到最优解的过程,甚至可能无法找到最优解
[0055]本发明上述实施例基于深度学习处理器设计的深度学习模型执行时间的评估方法对于深度学习模型执行时间的评估粒度高、效率高以及不需要依赖工程师经验。因此,本发明的深度学习模型执行时间的评估方法可以有效降低最优化搜索开销,为搜索提供指导,进一步地,经过优化后的深度学习模型能够在保证准确率不下降的同时还缩短了执行时间。
Smart Images

Figure CN117422957B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of object recognition in machine learning, specifically to the field of deep learning model optimization, and more specifically to a method for evaluating the execution time of a deep learning model. Background Technology
[0002] With the continuous development and popularization of deep learning methods, higher demands have been placed on processor computing power and power consumption. Researchers in the field of system architecture have designed a series of dedicated deep learning processors to address the characteristics of deep learning methods and have put them into practical applications. Deep learning processors, with their high computing power and low power consumption, enable deep learning methods to be deployed more efficiently on platforms such as servers, terminal devices, and embedded devices, promoting the application of deep learning methods. However, its ecosystem is still incomplete.
[0003] Cambricon's MLU series chips are representative of deep learning processors, with main products including the MLU270 for inference and the MLU220 for edge computing devices. The software toolchain is largely complete, including the Cambricon Neuware Machine Learning Library (CNML), which supports basic neural network operators such as convolution, pooling, matrix-vector multiplication (MLP), and batch normalization. Users can develop and build deep learning programming frameworks based on this toolchain and deploy them on the MLU.
[0004] Deep learning programming frameworks provide flexible programming interfaces for the development of deep learning applications, allowing deep learning model developers to implement deep learning methods simply and efficiently. However, during model inference, there is currently a lack of mature, general-purpose deep learning compilation frameworks to support the optimization and transformation of deep learning models.
[0005] Deep learning compilers can accept deep learning models written in various deep learning programming frameworks and optimize them into executable files that can be deployed on target hardware platforms. A key method in this optimization process is computation graph optimization, which mainly involves rearranging, fusing, and resolving operators. The goal is to improve the inference speed of deep learning models (i.e., reduce execution time) after a series of operations on the computation graph. Therefore, this problem can be abstracted into the following optimization problem:
[0006]
[0007] Where X represents all possible inter-operator operations performed on specified operator nodes of the computation graph corresponding to the deep learning model, x1(n, mp) represents the inter-operator operation mp performed on n operator nodes in the computation graph, st is a general symbol for describing optimization problems, N represents all operator nodes in the computation graph, MP represents the set of inter-operator operations on all operator nodes in the computation graph, and mp t Let $T$ represent the $t$-th inter-operator operation, and $Cost(G(X))$ represent the execution time of the deep learning model corresponding to all inter-operator operations in the computation graph.
[0008] As shown in the formula above, the goal of this optimization problem is to minimize Cost(*), which is to minimize the execution time of the deep learning model corresponding to the computation graph. The aim is to shorten the execution time (i.e., inference time) of the deep learning model as much as possible while ensuring accuracy. However, obtaining Cost(*) requires running on actual hardware, which brings two problems: First, the execution time of deep learning models is generally long, while the search space is generally large. This will lead to an excessively long time to find the optimal solution to the problem, and most of the time will be consumed in the execution of the deep learning model rather than in the search for the optimal solution. Second, the actual execution time of the deep learning model may be affected by other temporary environmental factors outside the deep learning model, resulting in inaccurate time recording, which will interfere with the process of finding the optimal solution, or even prevent the optimal solution from being found.
[0009] To address the above issues, using a deep learning model inference cost model (i.e., a method for evaluating the execution time of a deep learning model) to approximate the execution time overhead of a deep learning model can effectively reduce the time overhead of optimization search and provide guidance for the search. Compared to executing on actual hardware, the method of using a cost model to evaluate the execution time of a deep learning model has the following advantages: (1) Reduced computational overhead: The cost model can quickly estimate the execution time under different decisions or parameter choices without actually running the model, which avoids the overhead of executing the model multiple times, especially when the search space is large or the search process requires iteration; (2) Improved search efficiency: By introducing a cost model into the optimization search, less likely optimization schemes can be filtered out more quickly, thus focusing on options that are expected to achieve better results.
[0010] Current methods for evaluating the execution time of deep learning models on processors are based on the cost assessment of traditional parallel computing tasks using traditional multi-core processors or CPU-GPU heterogeneous platforms. They typically predict the execution time of hierarchical convolutional neural network models on traditional platforms like CPUs and GPUs, evaluating the model's execution time at the layer or loop level, without assessing the execution time at the operator level. However, this approach lacks granularity in evaluating deep learning model execution time, leading to problems such as low accuracy, low efficiency, high cost, and over-reliance on engineer experience. Furthermore, it does not address the emerging general-purpose deep learning model execution time modeling methods targeting deep learning processors.
[0011] In summary, existing methods for evaluating the execution time of deep learning models suffer from several problems: low granularity (previous methods were built at the loop level), inefficiency, and over-reliance on engineer experience. Furthermore, these methods are not designed for deep learning processors (previous methods were based on general-purpose GPUs). Therefore, existing methods fail to effectively reduce optimization search overhead, provide guidance for the search, and, while not guaranteeing accuracy, may even reduce the efficiency of the optimization search process or increase the execution time of the deep learning model due to inappropriate evaluation methods. Summary of the Invention
[0012] Therefore, the purpose of this invention is to overcome the shortcomings of the prior art and provide a method for evaluating the execution time of deep learning models.
[0013] The objective of this invention is achieved through the following technical solution:
[0014] According to a first aspect of the present invention, a method for evaluating the execution time of a deep learning model is provided, the method comprising:
[0015] S1. Obtain a deep learning model trained using a deep learning processor;
[0016] S2. Quantize the deep learning model and convert the quantized deep learning model into a computation graph, wherein the computation graph includes multiple operators and the connection relationships between each operator;
[0017] S3. Obtain the parameter information of each operator in the computation graph;
[0018] S4. Based on the parameter information of each operator in the computation graph, the execution time of each operator in the computation graph is obtained by using the cost evaluation function corresponding to each operator;
[0019] S5. Calculate the migration time between the input and output data of each operator, and calculate the blocking waiting time of each operator;
[0020] S6. Evaluate the execution time of the deep learning model based on the execution time of each operator in the computation graph, the blocking waiting time of each operator, and the migration time between the input and output data of each operator.
[0021] In some embodiments of the present invention, in step S4, the cost evaluation function corresponding to each operator is obtained in the following manner:
[0022] Construct an operator cost dataset, wherein the operator cost dataset includes: operator type, computational precision, parameter information, input data size, and execution time;
[0023] Statistical regression analysis was performed on the operator cost dataset to obtain the statistical distribution of the execution time for each operator;
[0024] Based on the statistical distribution of execution time for each operator and the computational and storage architectures of deep learning processors, a cost evaluation function for each operator is obtained through polynomial fitting regression.
[0025] In some embodiments of the present invention, in step S4, the parameter information of each operator in the computation graph includes: preset hyperparameters and the size of preset input data, and the execution time of each operator in the computation graph is obtained in the following manner:
[0026] Cost op (op) = Cost_op_F(input) size [hp(op)])
[0027] Where Cost_op_F(*) is the cost evaluation function for the operator op, and Cost_op_F(input) size [hp(op)] represents the preset hyperparameter hp and the preset input data size input. size The following calculates the execution time of the operator op in the graph.
[0028] In some embodiments of the present invention, in step S5, the migration time of the input data and output data of each operator in the computation graph is obtained in the following manner:
[0029]
[0030] Where sizeof(input) represents the memory space occupied by the input data of the operator op, sizeof(output) represents the memory space occupied by the output data of the operator op, and bandwidth(global) represents the memory bandwidth of the deep learning processor.
[0031] When traversing the computation graph using the graph traversal method and finding connections between operators that contain a preset fusion operator, the migration time between the input and output data of the preset fusion operator in the computation graph is obtained as follows:
[0032] Cost_edge_data(op i _op j )
[0033] =Cost_edge_data(op i-1 →op i )
[0034] +Cost_edge_data(op j →op j+1 )
[0035] Among them, op i Op represents the first operator of the preset fusion operators. j This represents the last operator in the preset fusion operators, Cost_edge_data(op i-1 →op i ) represents the first operator op of the preset fusion operators in the computation graph. i A previous operator op i-1 The migration time between input and output data, Cost_edge_data(op j →op j+1 ) represents the last operator op of the preset fusion operators in the computation graph. j An operator op following j j+1 The migration time between input and output data.
[0036] In some embodiments of the present invention, in step S5, the computation graph is traversed based on a graph traversal method to obtain the blocking waiting time of each operator in the computation graph. Specifically, when the computation graph is traversed based on a graph traversal method and the connection relationship between operators contains a preset fusion operator, the blocking time of the preset fusion operator in the computation graph is:
[0037] Cost_edge_block(op i _op j ) = 0
[0038] In some embodiments of the present invention, in step S6, the execution time of the deep learning model is evaluated in the following manner:
[0039]
[0040] According to a second aspect of the present invention, a method for optimizing a deep learning model is provided, the method comprising:
[0041] T1. Obtain the deep learning model trained on the deep learning processor and convert it into a computation graph;
[0042] T2. The optimization space of the computation graph is searched using the optimal search function to obtain multiple optimization methods for the computation graph;
[0043] T3. Using the deep learning model execution time evaluation method described in the above embodiments, the execution time of the deep learning model corresponding to each optimization method is evaluated in turn, and an optimization method is selected to optimize the computation graph according to the preset rules based on the calculated execution time.
[0044] In some embodiments of the present invention, in step T3, after evaluating the execution time of the deep learning model corresponding to each optimization method, the cost evaluation function corresponding to each operator is updated in the following manner:
[0045] Cost_op_F_update()=Cost_op_F()*α+Cost_op_F_new()*β
[0046] Where Cost_op_F_update(*) is the cost evaluation function after the operator op is updated, α and β are preset hyperparameters, Cost_op_F(*) is the cost evaluation function at the time of the previous evaluation of the operator op, and Cost_op_F_new(*) is the incremental cost evaluation function of the operator op. The incremental cost evaluation function corresponding to each operator is obtained in the following way:
[0047] The execution time of each operator in the computation graph is added to the operator cost dataset to obtain a new operator cost dataset;
[0048] Statistical regression analysis was performed on the new operator cost dataset to obtain the statistical distribution of the execution time for each operator;
[0049] Based on the statistical distribution of execution time for each operator and the computational and storage architectures of deep learning processors, the incremental cost evaluation function for each operator is obtained through polynomial fitting regression.
[0050] According to a third aspect of the present invention, an image recognition method is provided, the method comprising:
[0051] Acquire the image to be identified;
[0052] The target recognition model is optimized using the deep learning model optimization method described in the above embodiments, and then the image to be recognized is recognized.
[0053] According to a fourth aspect of the present invention, an electronic device is provided, comprising: one or more processors; and a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the electronic device performs the steps of the methods described in the first, second, and third aspects.
[0054] Compared with the prior art, the advantages of the present invention are as follows:
[0055] The deep learning model execution time evaluation method based on a deep learning processor design described in the above embodiments of the present invention offers high granularity and efficiency in evaluating deep learning model execution time, and does not rely on engineer experience. Therefore, the deep learning model execution time evaluation method of the present invention can effectively reduce optimization search overhead, provide guidance for the search, and further, the optimized deep learning model can shorten the execution time while ensuring that the accuracy does not decrease. Attached Figure Description
[0056] The embodiments of the present invention will be further described below with reference to the accompanying drawings, wherein:
[0057] Figure 1 This is a schematic diagram of a method for evaluating the execution time of a deep learning model according to an embodiment of the present invention.
[0058] Figure 2 This is a schematic diagram illustrating the experimental results of the execution time of the convolution operator evaluated by the cost evaluation function corresponding to the convolution operator according to an embodiment of the present invention;
[0059] Figure 3 This is a schematic diagram illustrating the experimental results of the execution time of the pooling operator evaluated by the cost evaluation function corresponding to the pooling operator according to an embodiment of the present invention. Detailed Implementation
[0060] To make the objectives, technical solutions, and advantages of this invention clearer, the invention is further described in detail below through specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0061] As described in the background section, existing methods for evaluating the execution time of deep learning models suffer from problems such as low granularity (previous methods were built at the loop level), low efficiency, and excessive reliance on engineer experience. Furthermore, these methods are not designed for deep learning processors (previous methods were based on general-purpose GPUs).
[0062] To address the aforementioned problems, this invention proposes a method for evaluating the execution time of deep learning models, based on a dedicated processor (i.e., a deep learning processor) and starting from the operator level (oplevel). The invention is described in detail below with reference to the accompanying drawings and embodiments.
[0063] According to one embodiment of the present invention, such as Figure 1 As shown, this diagram illustrates the flowchart of the method for evaluating the execution time of the deep learning model according to the present invention, which includes steps S1-S6. Each step is described in detail below.
[0064] In step S1, a deep learning model trained on a deep learning processor is obtained. This involves searching for and downloading publicly available deep learning models trained on a deep learning processor. For example, typical deep learning object recognition models include MobileNet, DenseNet, and ResNet. The deep learning compilation framework PyTorch is used to load the deep learning object recognition model. It should be noted that besides PyTorch, other deep learning programming frameworks such as TensorFlow and Keras can also be used; no specific limitation is made here.
[0065] Due to the accuracy requirements of deep learning models, the general architecture and components of deep learning processors are designed for quantized deep learning models. Therefore, deep learning models need to be quantized to fully utilize the processor components and improve the inference speed of deep learning models (i.e., reduce the execution time of deep learning models).
[0066] In step S2, the deep learning model is quantized and the quantized deep learning model is converted into a computation graph, wherein the computation graph includes multiple operators and the connection relationships between each operator. Specifically, the present invention can quantize a deep learning model with a precision of fp32 to an int8 or int16 precision. In this embodiment of the present invention, the precision of the deep learning model is quantized from fp32 to int8. The deep learning model is parsed using PyTorch to generate a computation graph, and the computation graph is presented in the form of an Abstract Syntax Tree (AST).
[0067] In step S3, parameter information for each operator in the computation graph is obtained. This parameter information includes preset hyperparameters and the preset input data size. It should be noted that the computational components of general deep learning processors suffer from memory access alignment issues. This means that when the input data size is a multiple of the processor's memory transaction length (typically a multiple of 8 or 32), the operator's time overhead is generally better, significantly less than for similar sizes. However, when the size is slightly larger than a multiple of the memory transaction length, memory access conflicts may occur, leading to higher time overhead. Therefore, in designing the operator cost evaluation function, this invention discusses several cases based on whether the input data size is a multiple of the memory transaction length to more accurately evaluate the operator's execution time.
[0068] In step S4, based on the parameter information of each operator in the computation graph, the execution time of each operator in the computation graph is obtained using the cost evaluation function corresponding to each operator. According to an embodiment of the present invention, in step S4, the parameter information of each operator in the computation graph includes: preset hyperparameters and preset input data dimensions, and the execution time of each operator in the computation graph is obtained in the following manner:
[0069] Cost op (op) = Cost_op_F(input) size [hp(op)])
[0070] Where Cost_op_F(*) is the cost evaluation function for the operator op, and Cost_op_F(input) size [hp(op)] represents the preset hyperparameter hp and the preset input data size input. size The following calculates the execution time of the operator op in the graph.
[0071] To better understand the evaluation process of the execution time of each operator in the computation graph corresponding to the deep learning model in the above embodiments, this process is represented as:
[0072] hp(op)=[hp0,hp1,...,hp m ] T
[0073] Cost_op_F = f(hp(op))
[0074] input size =[input0, input1,..., input k ] T
[0075] Cost op (op) = Cost_op_F(input)size [hp(op)]
[0076]
[0077] Where hp(op) represents all the preset hyperparameters of operator op in the computation graph, m represents the number of preset hyperparameters, op can be a convolution operator (conv) or a pooling operator (pooling), k represents the number of preset input data, and input size The value represents the size of the preset input data for the operator `op`, and `Cost_op_F(*)` is the cost evaluation function for the operator `op`. As can be seen from the above formula, this function is jointly defined by the operator type and the preset hyperparameters of the operator. That is, the input is the size of the input data corresponding to the operator and the preset hyperparameters, and the output is the operator execution time. The specific form of this function is usually determined through theoretical analysis and actual statistics. This invention uses the method of constructing an operator cost dataset and combining it with theoretical analysis to determine `Cost_op_F(*)`. According to an embodiment of the present invention, in step S4, the cost evaluation function corresponding to each operator is obtained as follows: Constructing an operator cost dataset, wherein the operator cost dataset includes: operator type, computational precision, parameter information, input data size, and execution time; performing statistical regression analysis on the operator cost dataset to obtain the statistical distribution of the execution time of each operator; based on the statistical distribution of the execution time of each operator and the computational and storage architectures of the deep learning processor (such as GPR, NRAM, WRAM, SRAM, and GDRAMs, etc.), obtaining the cost evaluation function corresponding to each operator through polynomial fitting regression. It should be noted that the operator cost dataset was collected by calling the deep learning processor's dedicated operator library (including convolution operators, pooling operators, and activation operators, etc., under different parameters and computational precision, and their execution time on the MLU), setting different data items for traversal testing, and statistically analyzing the execution time; at the same time, when considering the computational architecture and storage architecture of the MLU, the following three factors need to be included: (1) the memory alignment problem of the computational components; (2) the migration time between global memory and on-chip local memory; (3) the time for operators that do not support computation on the MLU to migrate back to the host (HOST) for execution.
[0078] After evaluating the execution time of each operator in the computation graph of the deep learning model in step S4 above, the next step is to evaluate the migration time between the input and output data of each operator in the computation graph and the blocking wait time of each operator to calculate the complete execution time of the deep learning model. It should be noted that the migration time between the input and output data of an operator refers to the time overhead incurred by migrating data between the host and device's global memory. Generally, the input and output data of a deep learning model can be entirely placed in global memory after quantization; therefore, this overhead only occurs at the beginning and end of the deep learning model's execution. Since the computation graph of a deep learning model may be executed in parallel or with multiple branches, an operator may need to wait for the completion of multiple preceding operators to obtain all the data before it can begin execution. Therefore, this invention abstracts the cost of the computation graph edge as the blocking wait time (Cost_edge_block) of the entire computation graph running in parallel. Furthermore, general deep learning compilation frameworks (such as PyTorch and TensorFlow) typically employ a graph optimization strategy that merges memory-intensive operators like BatchNorm and Activate with computationally intensive operators like Conv, into a single operator, aiming to reduce memory access time overhead. Therefore, this invention takes this into account and incorporates a dedicated deep learning processor operator library for targeted processing. The following is a schematic diagram illustrating the calculation of the input and output data migration time for each operator and the blocking latency for each operator.
[0079] In step S5, the migration time between the input and output data of each operator is calculated, as well as the blocking waiting time of each operator is calculated.
[0080] According to an embodiment of the present invention, in step S5, the migration time of the input data and output data of each operator in the computation graph is obtained in the following manner:
[0081]
[0082] Where `sizeof(input)` represents the memory space occupied by the input data of operator `op`, `sizeof(output)` represents the memory space occupied by the output data of operator `op`, and `bandwidth(global)` represents the memory bandwidth of the deep learning processor. Specifically, when traversing the computation graph using the graph traversal method and traversing to the point where the connection relationships between operators contain a preset fusion operator, the migration time between the input and output data of the preset fusion operator in the computation graph is obtained as follows:
[0083] Cost_edge_data(opi _op j )
[0084] =Cost_edge_data(op i-1 →op i )
[0085] +Cost_edge_data(op j →op j+1 )
[0086] Among them, op i Op represents the first operator of the preset fusion operators. j This represents the last operator in the preset fusion operators, Cost_edge_data(op i-1 →op i ) represents the first operator op of the preset fusion operators in the computation graph. i A previous operator op i-1 The migration time between input and output data, Cost_edge_data(op j →op j+1 ) represents the last operator op of the preset fusion operators in the computation graph. j The next operator op j+1 The migration time between input and output data is considered. Here, we take the GDRAM deep learning processor as an example to introduce the range of memory bandwidth values. GDRAM refers to the MLU270 global memory, with a capacity of 16GB and a bandwidth of 102.4GB / s. Based on actual measurements, the bandwidth in this embodiment is taken as 80GB / s. In practical applications, due to data scale or other possible limitations imposed by processes and hardware technology, memory bandwidth often cannot reach theoretical performance. Based on actual calculations and empirical estimations, in MLU series graphics cards, the memory bandwidth is generally taken as 80% of the peak bandwidth. When the data to be transmitted is large enough, the memory bandwidth can be taken as 85% of the peak bandwidth.
[0087] According to another embodiment of the present invention, in step S5, the computation graph is traversed based on the graph traversal method to obtain the blocking waiting time of each operator in the computation graph. Specifically, when the computation graph is traversed based on the graph traversal method and the connection relationship between operators contains a preset fusion operator, the blocking time of the preset fusion operator in the computation graph is:
[0088] Cost_edge_block(op i _op j ) = 0
[0089] To better understand the concept of preset fusion operators in the above embodiments, taking the connection relationship between conv-batchnorm-relu operators as an example, when an operator of the form conv-batchnorm-relu (i.e., a preset fusion operator) is identified during the traversal of the computation graph, the operator blocking wait time between preset fusion operators will not be calculated again; instead, only the transition time between the input and output of this series of operators will be calculated. It should be noted that, depending on the deep learning model and specific application scenarios, the form of the preset fusion operator can also be other forms, which are not specifically limited here.
[0090] In step S6, the execution time of the deep learning model is evaluated based on the execution time of each operator in the computation graph, the blocking wait time of each operator, and the migration time between the input and output data of each operator. According to an embodiment of the present invention, in step S6, the execution time of the deep learning model is evaluated in the following manner:
[0091]
[0092] The deep learning model execution time evaluation method designed in the above embodiments performs cost evaluation at the operator (op level). Therefore, the granularity of operator execution time evaluation is higher, the integrity of operator nodes is stronger, and the adaptability to dedicated processors with higher encapsulation granularity is stronger. The calculated deep learning model execution time has high accuracy (i.e., high evaluation granularity), high efficiency (using a simple fitting regression method for cost estimation, without the need for complex algorithms such as deep learning, resulting in faster results), and does not rely on engineer experience (constructing formulas and datasets for automatic cost evaluation and updating, without relying on the experience of domain experts).
[0093] Considering that the optimization search process of deep learning models is mainly focused on optimizing the execution time of the model, by applying the deep learning model execution time evaluation method designed in the embodiments of the present invention to the field of deep learning model optimization, the execution time of deep learning models can be shortened as much as possible while ensuring accuracy.
[0094] According to an embodiment of the present invention, a method for optimizing a deep learning model is proposed. The method includes: T1, obtaining a deep learning model trained on a deep learning processor and converting it into a computation graph; T2, using an optimal search function to search the optimization space of the computation graph to obtain multiple optimization methods for the computation graph; T3, using the deep learning model execution time evaluation method described in the above embodiment to sequentially evaluate the execution time of the deep learning model corresponding to each optimization method, and selecting an optimization method to optimize the computation graph according to a preset rule based on the execution time of the computation.
[0095] Furthermore, due to the inaccuracy of hardware time measurements, when using the deep learning model execution time evaluation method of this invention, it is necessary to update the cost evaluation function corresponding to each operator based on the actual execution time of the final deep learning model. Therefore, this invention uses an update strategy based on the smooth averaging principle to update the cost evaluation functions corresponding to the various operators. According to an embodiment of this invention, in step T3, after evaluating the deep learning model execution time corresponding to each optimization method, the cost evaluation function corresponding to each operator is updated in the following manner:
[0096] Cost_op_F_update()=Cost_op_F()*α+Cost_op_F_new()*β
[0097] Wherein, Cost_op_F_update(*) is the cost evaluation function after the operator op is updated, α and β are preset hyperparameters, Cost_op_F(*) is the cost evaluation function at the time of the previous evaluation of the operator op, and Cost_op_F_new(*) is the incremental cost evaluation function of the operator op. The incremental cost evaluation function for each operator is obtained as follows: the execution time of each operator in the computation graph is added to the operator cost dataset to obtain a new operator cost dataset; statistical regression analysis is performed on the new operator cost dataset to obtain the statistical distribution of the execution time of each operator; based on the statistical distribution of the execution time of each operator and the computational and storage architecture of the deep learning processor, the incremental cost evaluation function for each operator is obtained through polynomial fitting regression. It should be noted that the use of an optimal search function to search the optimization space in this optimization method is a widely used technique that can help optimize the model and improve its performance. An optimal search function is a method that can find the variable values that minimize or maximize the objective given a target and constraints. Since the search optimization method for optimizing deep learning models is an existing technology, it will not be elaborated on here.
[0098] By using the deep learning model execution time evaluation method of the present invention to optimize the deep learning model, the execution time of the deep learning model corresponding to various optimization methods searched can be effectively reduced while ensuring that the accuracy does not decrease.
[0099] In practical applications, the models optimized by the deep learning model optimization methods described in the above embodiments effectively shorten the execution time of deep learning models while maintaining accuracy. The application of the optimized deep learning model from the above embodiments will be illustrated using a deep learning object recognition model as an example.
[0100] According to an embodiment of the present invention, the method includes: acquiring an image to be identified; optimizing the target recognition model using the optimization method of the deep learning model described in the above embodiment, and then recognizing the image to be identified.
[0101] As can be seen from the above description, the evaluation method for the execution time of deep learning models designed in this invention can shorten the execution time of deep learning models.
[0102] To verify that the cost evaluation functions based on various operators in the embodiments of the present invention can more accurately evaluate the execution time of deep learning models, the inventors conducted the following experiments:
[0103] The preset hyperparameters were initialized to: α = 0.95, β = 0.05. Taking the Cambricon MLU270 second-generation chip as an example, the update strategy of the above embodiment (based on the smooth averaging principle) was adopted, and a total of 1000 rounds of iterative acquisition were performed. The expressions for Cost_op_F(x) of the convolution operator at kernel size = 1, 3 and the pooling operator at kernel size = 2, stride = 2 (max pooling and average pooling) were obtained. size () refers to the size of the input data, x = 32i means the size of the input data is a multiple of 32, with no padding). Specifically:
[0104] Convolution operator (Conv):
[0105] kernel size = 1:
[0106]
[0107] kernel size = 3:
[0108]
[0109] Experimental results are as follows Figure 2 As shown, from Figure 2 As shown in the figures ((a) shows the experimental results for kernel size = 1; (b) shows the experimental results for kernel size = 3), the horizontal axis represents the size of the input data for the convolution operator, the vertical axis represents the execution time of the convolution operator, the scatter plot represents the execution time of the convolution operator on actual hardware, and the smooth curve represents the execution time of the operator cost evaluation function. The results show that the convolution operator cost evaluation function can effectively classify and evaluate the time cost of the fitting operator according to the memory alignment principle (i.e., classify according to whether the size of the input data is a multiple of the memory transaction length), and the evaluation results are basically close to the actual execution results. This shows that the deep learning model execution time evaluation method proposed in this invention is accurate and effective.
[0110] Pooling operator:
[0111] Max pooling:
[0112] kernel_size=2, stride=2
[0113]
[0114] Average pooling:
[0115] kernel size=2, stride=2
[0116]
[0117] Experimental results are as follows Figure 3 As shown, from Figure 3 As shown in the figures ((a) shows the experimental results of the max pooling operator; (b) shows the experimental results of the average pooling operator), the horizontal axis represents the size of the input data for the pooling operator, the vertical axis represents the execution time of the pooling operator, the scatter plot represents the execution time of the pooling operator on actual hardware, and the smooth curve represents the execution time of the operator cost evaluation function. The results show that the pooling operator cost evaluation function can effectively classify and evaluate the time overhead of the fitting operator according to the memory alignment principle (i.e., classify it according to whether the size of the input data is a multiple of the memory transaction length), and the evaluation results are basically close to the actual execution results. This shows that the deep learning model execution time evaluation method proposed in this invention is accurate and effective.
[0118] It should be noted that the preset hyperparameters α and β are generally set to α = 0.95 and β = 0.05, which will achieve a relatively accurate fitting effect. When the sample data size is limited, resulting in a small number of iterations, α can be adjusted to 0.9 and β = 0.1. For other embodiments, appropriate adjustments can be made according to the data size. The basic principle is that for data samples with high uncertainty, the value of β should be larger. The specific value of the preset hyperparameters is not limited here. In addition, this embodiment also obtains the expressions for Cost_op_F(x) of other commonly used operators, which are not listed here due to space limitations.
[0119] In summary, the deep learning model execution time evaluation method based on the deep learning processor design in the above embodiments offers high granularity and efficiency in evaluating deep learning model execution time, and does not rely on engineer experience. Therefore, the deep learning model execution time evaluation method of the present invention can effectively reduce the optimization search overhead, provide guidance for the search, and further, the optimized deep learning model can shorten the execution time while ensuring that the accuracy does not decrease.
[0120] It should be noted that although the steps are described in a specific order above, it does not mean that the steps must be executed in the above specific order. In fact, some of these steps can be executed concurrently, or even in a different order, as long as the required function can be achieved.
[0121] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for causing a processor to implement various aspects of the invention.
[0122] Computer-readable storage media can be tangible devices that hold and store instructions for use by an instruction execution device. Computer-readable storage media can be, for example, including but not limited to, electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination thereof.
[0123] The various embodiments of the present invention have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical application, or technical improvements to the embodiments in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A method for evaluating the execution time of a deep learning model, characterized in that, The method includes: S1. Obtain a deep learning model trained using a deep learning processor; S2. Quantize the deep learning model and convert the quantized deep learning model into a computation graph, wherein the computation graph includes multiple operators and the connection relationships between each operator; S3. Obtain the parameter information of each operator in the computation graph; S4. Based on the parameter information of each operator in the computation graph, the execution time of each operator in the computation graph is obtained using the cost evaluation function corresponding to each operator; wherein, the cost evaluation function corresponding to each operator is obtained in the following way: Construct an operator cost dataset, wherein the operator cost dataset includes: operator type, computational precision, parameter information, input data size, and execution time; Statistical regression analysis was performed on the operator cost dataset to obtain the statistical distribution of the execution time for each operator; Based on the statistical distribution of execution time for each operator and the computational and storage architectures of deep learning processors, the cost evaluation function for each operator is obtained through multinomial fitting regression. S5. Calculate the migration time between the input and output data of each operator, and calculate the blocking waiting time of each operator; S6. Evaluate the execution time of the deep learning model based on the execution time of each operator in the computation graph, the blocking waiting time of each operator, and the migration time between the input and output data of each operator.
2. The method according to claim 1, characterized in that, In step S4, the parameter information for each operator in the computation graph includes: preset hyperparameters and preset input data dimensions, and the execution time of each operator in the computation graph is obtained in the following manner: in, For operators The cost evaluation function, Indicates the preset hyperparameters and the size of the preset input data Operators in the computation graph below Execution time.
3. The method according to claim 2, characterized in that, In step S5, the migration time of the input and output data of each operator in the computation graph is obtained in the following manner: in, Operator Memory space occupied by input data Operator The memory footprint of the output data This represents the memory bandwidth of the deep learning processor, where, When traversing the computation graph using the graph traversal method and finding connections between operators that contain a preset fusion operator, the migration time between the input and output data of the preset fusion operator in the computation graph is obtained as follows: in, This represents the first operator of the preset fusion operators. This indicates the last operator in the preset fusion operators. This represents the first operator of the preset fusion operators in the computation graph. A previous operator The migration time between input and output data. This represents the last operator among the preset fusion operators in the computation graph. The next operator The migration time between input and output data.
4. The method according to claim 3, characterized in that, In step S5, the computation graph is traversed using a graph traversal method to obtain the blocking waiting time of each operator in the computation graph. Specifically, when the computation graph is traversed using the graph traversal method and the connection relationship between operators contains a preset fusion operator, the blocking time of the preset fusion operator in the computation graph is: 。 5. The method according to claim 4, characterized in that, In step S6, the execution time of the deep learning model is evaluated in the following manner: 。 6. An optimization method for a deep learning model, characterized in that, The method includes: T1. Obtain the deep learning model trained on the deep learning processor and convert it into a computation graph; T2. The optimization space of the computation graph is searched using the optimal search function to obtain multiple optimization methods for the computation graph; T3. The execution time of the deep learning model corresponding to each optimization method is evaluated sequentially using the method described in any one of claims 1-5, and an optimization method is selected to optimize the computation graph based on the execution time of the computation according to a preset rule.
7. The method according to claim 6, characterized in that, In step T3, after evaluating the execution time of the deep learning model corresponding to each optimization method, the cost evaluation function corresponding to each operator is updated in the following manner: in, For operators Updated cost evaluation function, All of these are preset hyperparameters. For operators The cost evaluation function from the previous evaluation. For operators The incremental cost evaluation function is obtained as follows: The execution time of each operator in the computation graph is added to the operator cost dataset to obtain a new operator cost dataset; Statistical regression analysis was performed on the new operator cost dataset to obtain the statistical distribution of the execution time for each operator; Based on the statistical distribution of execution time for each operator and the computational and storage architectures of deep learning processors, the incremental cost evaluation function for each operator is obtained through polynomial fitting regression.
8. An image recognition method, characterized in that, The method includes: Acquire the image to be identified; The target recognition model is optimized using any one of the optimization methods described in claims 6-7, and then the image to be recognized is recognized.
9. A computer-readable storage medium, characterized in that, It stores a computer program that can be executed by a processor to implement the steps of the method according to any one of claims 1-5, 6-7, and 8.
10. An electronic device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the electronic device to perform the steps of the method as described in any one of claims 1-5, 6-7, and 8.
Citation Information
Patent Citations
Parallel implementation method of particle mesh method on ARMv8 processor
CN110275732A
Method for acquiring training cost of distributed deep learning model based on multiple GPUs (Graphics Processing Unit)
CN114862656A