A method and device for predicting the hierarchical inference time of a convolutional neural network
Through layered type division and machine learning algorithms, feature engineering and model training are carried out for different types of deep learning models, which solves the problems of poor feature engineering and time-consuming measurement in the existing technology, and realizes efficient and accurate inference time prediction, which improves the model segmentation and inference speed in cloud-edge-end collaborative scenarios.
Patent Information
- Application Number
- CN202210133010.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-14
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-02-14
AI Technical Summary
In the inference time prediction of deep learning models, feature engineering is poor, resulting in too many zero elements in the feature matrix, and the prediction accuracy of using a unified model is low, and the actual measurement method is time-consuming and labor-intensive, especially in the cloud-edge-end collaborative scenario.
Hierarchical type division and machine learning algorithms are used to perform feature engineering and model training on different hierarchical types, and multiple inference time prediction models are built, combined with operator fusion strategies to reduce feature screening and actual measurement needs.
It improves the accuracy and efficiency of inference time prediction of deep learning models, reduces the time-consuming measurement, and improves the model segmentation accuracy and inference speed in cloud-edge-end collaborative scenarios.
Smart Images

Figure CN114648123B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly relates to a method and device for predicting the hierarchical inference time of a convolutional neural network. Background Art
[0002] In recent years, with the explosive growth of data, the substantial improvement of hardware computing power, and the increasing maturity of deep learning algorithms, artificial intelligence technology has witnessed a blowout development, and has made breakthrough progress in a series of fields such as image recognition, object detection, speech recognition, and knowledge graphs, and has great application value in many aspects such as robot applications, industrial manufacturing, and the Internet of Things. Currently, most deep learning models use GPUs for large-scale concurrent computing during training to reduce the model training time; while in the application of the model, the inference ability of the deep neural network is utilized to complete the forward calculation of the network, and this calculation is deployed as an application service to generate commercial value at the same time. When a deep learning model is inferring, it only includes the forward calculation process, and the deployment environment is diverse. And the inference application scenarios of users are diverse. For example, inference on the mobile phone side requires lightweight processing of the model and adaptation to different hardware platforms (Arm CPUs, Mali GPUs, etc.), and applications on robots require combining edge computing and cloud computing to accelerate the model inference process to solve the problem of limited computing power of the robot itself. Sometimes, inference also needs to be jointly completed by multiple models, and these models may be separately and independently trained during the training phase. When inferring and running, for the joint deployment of specific models, it is necessary to comprehensively consider how to maximize the efficiency of joint model inference.
[0003] Research on the inference acceleration problem of deep learning models mainly focuses on model compression and cloud-edge-end collaborative inference acceleration of models. The former is applied to scenarios with extremely limited computing power such as mobile phones, which will cause a certain loss of inference accuracy. The latter is commonly used in scenarios where the edge-side computing power is limited and there are many inference tasks, such as robot applications, the Internet of Things, and autonomous driving. It can use edge computing nodes and cloud servers to achieve inference acceleration of the model without loss of accuracy. In the research on cloud-edge-end inference acceleration, it is usually necessary to obtain the inference time information of each layer of the model to complete the subsequent splitting process of the model. Due to the large number of layers in deep learning models and the changing deployment environment, relying on actual measurement of the inference time of model layers is extremely laborious. Currently, domestic and foreign research mainly focuses on the prediction of model inference time. By extracting feature parameters that may affect the model inference time, machine learning algorithms are used to predict the model layers and the overall inference time, including the nn-Meter inference time prediction system described in the best paper of MobiSys, "nn-Meter: Towards Accurate Latency Prediction of Deep-Learning Model Inference on Diverse Edge Devices", the inference time prediction method for CNN networks described in the CF conference paper, "Performance Prediction for Convolutional Neural Networks on Edge GPUs", and the prediction of the training time of the CNN model layer by layer described in the IEEE BigData conference paper, "Predicting the Computational Cost of Deep Learning Models".
[0004] The key idea of nn-Meter is to divide the entire model into kernels and then perform kernel-level prediction. nn-Meter is based on two key technologies, which can accurately predict the inference time of different models during deployment. One is kernel detection, which can automatically identify the optimization strategies of the deployment platform, and then decompose the model into actual running kernels based on these strategies. The other is adaptive data sampling, which effectively samples the most beneficial configurations from the entire design space to efficiently build an accurate kernel-level latency predictor.
[0005] Five machine learning algorithms are used in Performance Prediction to predict the inference time of the CNN model, and an importance assessment of the feature parameters is made. Eleven features are selected in the paper and screened by the XGBoost algorithm. Based on the importance of the selected features, subsequent inference time prediction is carried out.
[0006] In the paper "Predicting the Computational Cost", the training time of the CNN model is predicted layer by layer in batches, and finally the sum is obtained to get the model training time in each epoch. The paper lists the feature classifications that may affect the training time, including layer types, parameters of specific layers, and hardware feature parameters, and uses traditional linear regression algorithms for prediction.
[0007] The above method has the following problems:
[0008] (1) For the problem of predicting the inference time of deep learning models, the most crucial and challenging step is to screen feature parameters and perform feature engineering. Existing technologies consider all influencing factors and then perform unified screening, which will result in a large number of zero elements in the feature matrix, thus making the result of feature engineering poor;
[0009] (2) Existing methods all establish a prediction model to predict the inference time of each layer of the model. However, the types of deep learning model layers vary greatly, resulting in different influencing factors for the inference time of each layer. Relying on a single prediction model to predict all layers will lead to low prediction accuracy for some layers.
[0010] In actual application scenarios, the latency (inference time) of a trained convolutional neural network model in actual deployment is an important indicator to determine whether the image classification model is usable. In existing technologies, NAS (Neural Architecture Search, automatic network structure search) is usually used to search for convolutional neural network models. During the NAS search process, the inference time of the model is required as feedback to evaluate the quality of the model. Directly relying on the measured method is time-consuming and laborious, and a lot of useless test code needs to be implanted in the code. On the other hand, in cloud-edge collaboration scenarios, such as robotics applications, the Internet of Things, and autonomous driving, where the computing power at the edge is limited and there are many inference tasks, it is usually necessary to obtain the inference time information of each layer of the model to complete subsequent splitting processing of the model. Due to the large number of layers in deep learning models and the ever-changing deployment environment, relying on measuring the inference time of each layer of the model manually is extremely laborious. Summary of the Invention
[0011] To solve the deficiencies of the existing technology, achieve efficient search for multiple models under different software and hardware platforms, predict the combined inference time, avoid the time-consuming and laborious nature of actual measurement, and improve the efficiency and accuracy of inference time prediction, the present invention adopts the following technical solutions:
[0012] A method for predicting the inference time of convolutional neural network layers includes the following steps:
[0013] S101, Data collection: Collect the hierarchical operator information of various convolutional neural network models, determine multiple hierarchical types according to the characteristics of the operator information, divide the operator information into each hierarchical type, and collect the platform framework information;
[0014] The convolutional neural network model is used for image classification. By the inference time when performing image processing on various convolutional neural network models, the convolutional neural network model is searched, thereby improving the efficiency and accuracy of the search, further improving the efficiency of the model, and ultimately improving the efficiency of image processing, avoiding the method that relies on actual measurement, which is time-consuming and laborious, and a lot of useless test codes need to be implanted in the actual measurement;
[0015] The convolutional neural network model is used for cloud-edge-end collaborative inference acceleration. Since the computing power at the edge side is limited and there are many inference tasks, obtain the inference time information of each layer of the model to complete the subsequent splitting process of the model. Due to the large number of layers of the deep learning model and the variable deployment environment, relying on the actual measurement of the inference time of the model layers is a huge workload. By predicting the inference time of the model, the inference acceleration of the model can be achieved without loss of accuracy.
[0016] Collect the operator layer information of the deep convolutional neural network model. The convolutional neural network models include: AlexNet, VGG16, ResNet, DenseNet, MobileNetv1, MobileNetv2, MobileNetv3, GoogleNet, ShuffleNet. Divide the operator layers with the same function into one category. The convolutional layer includes convolutional operators, depthwise separable convolutional operators, and dilated convolutional operators; the pooling layer includes max pooling operators, min pooling operators, and average pooling operators; the batch normalization layer includes batch normalization operators; the fully connected layer includes fully connected operators; the activation function layer includes Relu operators, Relu6 operators, Sigmoid operators, and Swish operators; the element-wise calculation layer includes Add operators, Concat operators, and Multiply operators.
[0017] S102, Feature engineering construction: For multiple hierarchical types, extract the layer feature parameters of the corresponding convolutional neural network models under each hierarchical type, and extract the platform framework feature parameters closely related to the inference time in the platform framework information, and fuse the model layer feature parameters and the platform framework feature parameters to form the feature parameters of multiple hierarchical types;
[0018] S103. Inference time prediction: Classify multiple hierarchical types according to the data characteristics of the feature parameters, group the hierarchical types with the same feature parameters into one group, and divide them into p groups in total. Separate the hierarchical types with different feature parameters into one group, and divide them into q groups in total. Use a machine learning algorithm to build an inference time prediction model for each group, that is, establish p inference time prediction models for the p groups of hierarchical types respectively, and establish q prediction models for the remaining q groups of hierarchical types, obtaining a total of p + q prediction models for predicting the inference time of the convolutional neural network model.
[0019] Further, the platform framework information includes hardware platform information and / or inference software framework information.
[0020] Further, the data acquisition in S101 includes the following steps:
[0021] Step 201: Collect various convolutional neural network model structures, including deep learning models such as AlexNet, VGG16, ResNet, DenseNet, MobileNetv1, MobileNetv2, MobileNetv3, etc.
[0022] Step 202: Collect platform framework information, including the name, type (such as CPU, GPU, TPU, VPU), hardware processing capacity, memory bandwidth of the hardware platform, as well as the model number of the hardware platform and the corresponding hardware configuration information, and / or the name of the inference software framework (such as TensorRT, TFLite, Openvino, MNN, NCNN), and the design features and relevant parameters under each inference software framework.
[0023] Step 203: Obtain the hierarchical operator information of the model from the model structure, including: convolution operator, depthwise separable convolution operator, dilated convolution operator, max pooling operator, min pooling operator, average pooling operator, batch normalization operator, fully connected operator, Relu operator, Relu6 operator, Sigmoid operator, Swish operator, Add operator, Concat operator, Multiply operator.
[0024] Step 204: Determine multiple hierarchical types according to the hierarchical operator information, including: convolutional layer, pooling layer, batch normalization layer, fully connected layer, activation function layer, element-wise calculation layer.
[0025] Step 205, divide the hierarchical operator information into one of multiple hierarchical types, and the corresponding relationships include: the convolutional layer contains convolutional operators, depthwise separable convolutional operators, and dilated convolutional operators; the pooling layer contains max pooling operators, min pooling operators, and average pooling operators; the batch normalization layer contains batch normalization operators, the fully connected layer contains fully connected operators, the activation function layer contains Relu operators, Relu6 operators, Sigmoid operators, and Swish operators; the element-wise calculation layer contains Add operators, Concat operators, and Multiply operators.
[0026] Further, the construction of the feature engineering in S102 includes the following steps:
[0027] S301, extract the model layer feature parameters for 6 hierarchical types, where:
[0028] The model layer feature parameters corresponding to the convolutional layer include: the size of the input feature map of this layer, the number of channels of the input feature map, the size of the convolutional kernel, the stride size, the padding size, the dilation size, the group size, and the number of channels of the output feature map of this layer;
[0029] The model layer feature parameters corresponding to the pooling layer include: the size of the input feature map of this layer, the number of channels of the input feature map, the size of the convolutional kernel, the stride size, the padding size; the size of the input feature map refers to the height and width of the feature map;
[0030] The model layer feature parameters corresponding to the normalization layer include: the size of the input feature map of this layer and the number of channels of the input feature map;
[0031] The model layer feature parameters corresponding to the fully connected layer include: the size of the input feature map of this layer, the number of channels of the input feature map, and the number of neurons;
[0032] The model layer feature parameters corresponding to the activation function layer include: the size of the input image of this layer and the number of channels of the input feature map;
[0033] The model layer feature parameters corresponding to the element-wise calculation layer include: the size of the input image of this layer and the number of channels of the input feature map;
[0034] S302, extract the feature parameters of the platform framework, including the number of floating-point operations per second and the memory access bandwidth of the hardware, and the adopted strategies include operator fusion strategies and / or parallel computing strategies;
[0035] S303, fuse the model layer feature parameters and the platform framework feature parameters as the feature parameters of each hierarchical type, and perform the determination of the feature parameters to decide the hierarchical types that need to perform feature screening in each hierarchical type;
[0036] S304, traverse each stratification type and perform feature screening on the feature parameters.
[0037] Further, the fusion of the feature parameters in S303 is to convert the model layer feature parameters and the platform framework feature parameters into the same dimension through feature engineering.
[0038] Further, the fusion of the feature parameters in S303 is to perform multi-dimensional fusion of the model layer feature parameters and the platform framework feature parameters, that is, the feature parameters of the platform framework are respectively in a one-to-one and / or one-to-many relationship with the model layer feature parameters pairwise.
[0039] Further, in S304, count the number of feature parameters of each stratification type. For the n stratification types with the number of feature parameters greater than the first threshold, construct data sets for various convolutional neural network models respectively, form n sets of data sample collections, use machine learning to train the feature screening model, and obtain the weight value coefficients of each feature under the n stratification types respectively. Sort according to the coefficients from large to small, set a weight threshold, filter out the feature parameters smaller than the weight threshold, and finally obtain the feature parameters after feature screening for the n different stratification types as the input for inference time prediction, and jointly form the feature parameters of m stratification types with the feature parameters in the remaining m - n stratification types. The purpose of selection is to perform different processing for different stratification types. For those with a small number of feature parameters, retaining all information can improve the final prediction accuracy. For those with a large number of feature parameters, some less important feature parameters can be filtered through feature screening. The machine learning algorithms used for feature screening include: XGBoost, LightGBM, random forest, and logistic regression.
[0040] Further, the inference time prediction in S103 includes the following steps:
[0041] S401, obtain the feature parameters of multiple stratification types after feature screening, and further classify the feature parameters of multiple stratification types;
[0042] S402, traverse the feature parameters of multiple stratification types, divide the stratification types with the same feature parameters into the same group, form p groups of stratification types, and the feature parameters in each group of the p groups contain the same model layer feature parameters and platform framework feature parameters;
[0043] S403, divide the stratification types with different feature parameters into q groups of stratification types;
[0044] S404. For the p+q group stratification types, the characteristic parameters between groups are all different. Machine learning algorithms are used to model the p+q group stratification types respectively for training the inference time prediction model. The machine learning algorithms may include: linear regression, non-linear regression, random forest, XGBoost, neural network.
[0045] S405. Based on the training of the prediction model, p+q inference time prediction models are obtained to complete the inference time prediction of various stratification types of the convolutional neural network.
[0046] Furthermore, after merging some operators using the operator fusion strategy and running them as a fused operator, the memory transfer can be reduced and the operator inference running speed can be improved. The stratification operator type is increased with the fused operator, and the characteristic parameters of the fused operator are a combination of the characteristic parameters of each operator before fusion. The detection mechanism of operator fusion is to introduce a time difference in the code of the model forward inference to determine whether operator fusion has occurred. When fusion occurs, the running time of two connected operators is less than the sum of the running times of the operators running alone. The stratified inference time prediction of operator fusion includes the following steps:
[0047] S501. Obtain the convolutional neural network model.
[0048] S502. Operator fusion detection. Search for fused operators for the operators of the convolutional neural network model using the fusion detection mechanism.
[0049] S503. When the model is running for inference, after operator fusion detection, in the actual inference process of the model, it runs with the fused operators under the platform framework.
[0050] S504. Model stratified inference time prediction. The extraction of characteristic parameters is to extract the characteristic parameters of the model layer after operator fusion, combine the information of the platform framework to form the characteristic parameters of various stratification types, and then predict the inference time according to S102 and S103.
[0051] A device for predicting the stratified inference time of a convolutional neural network includes: a data acquisition module, a feature engineering module, and an inference time prediction module connected in sequence.
[0052] The data acquisition module collects the stratification operator information of various convolutional neural network models, determines various stratification types according to the characteristics of the operator information, divides the operator information into each stratification type, and collects the platform framework information.
[0053] For the construction of the feature engineering, for multiple hierarchical types, extract the layer feature parameters of the corresponding convolutional neural network model under each hierarchical type, and extract the platform framework feature parameters closely related to the inference time in the platform framework information, and fuse the model layer feature parameters and the platform framework feature parameters to form the feature parameters of multiple hierarchical types;
[0054] For the inference time prediction, classify multiple hierarchical types according to the data characteristics of the feature parameters, divide the hierarchical types with the same feature parameters into one group, and a total of p groups are divided. The hierarchical types with different feature parameters are separately divided into one group, and a total of q groups are divided. Use machine learning algorithms to construct an inference time prediction model for each group, that is, establish p inference time prediction models for p groups of hierarchical types respectively, establish q prediction models for the remaining q groups of hierarchical types, and a total of p + q prediction models are obtained for predicting the inference time of the convolutional neural network model.
[0055] The advantages and beneficial effects of the present invention are as follows:
[0056] In order to solve the problem in the prior art that when predicting the inference time of a deep learning model, all influencing factors are considered comprehensively for feature engineering, resulting in too many zero elements in the feature matrix, and using a single model for prediction leads to low prediction accuracy. The present invention performs feature engineering on different hierarchical types respectively, and at the same time uses machine learning algorithms to train models for hierarchical types with different feature parameters respectively to complete the prediction of the hierarchical inference time, thereby greatly improving the prediction accuracy of the model hierarchical inference time. Description of the Drawings
[0057] Figure 1 It is a schematic diagram of the automatic network structure principle search.
[0058] Figure 2 It is a flowchart of the method of the present invention.
[0059] Figure 3 It is a flowchart of the operation of the data acquisition module in the embodiment of the present invention.
[0060] Figure 4 It is a flowchart of the operation of the feature engineering module in the embodiment of the present invention.
[0061] Figure 5 It is a flowchart of the operation of the inference time prediction module in the embodiment of the present invention.
[0062] Figure 6a It is a schematic diagram of the operator fusion process in the embodiment of the present invention.
[0063] Figure 6b It is a schematic diagram of the operator fusion process in the embodiment of the present invention.
[0064] Figure 6cIt is a schematic diagram of the operator fusion process in an embodiment of the present invention.
[0065] Figure 7 It is a schematic diagram of the operator fusion detection structure in an embodiment of the present invention.
[0066] Figure 8 It is a schematic diagram of the inference time prediction structure using operator fusion in an embodiment of the present invention.
[0067] Figure 9 It is a schematic diagram of the device structure of the present invention. Detailed implementation manners
[0068] The following will describe in detail the specific implementation manners of the present invention with reference to the accompanying drawings. It should be understood that the specific implementation manners described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0069] A method for predicting the hierarchical inference time of a convolutional neural network. The convolutional neural network is used for image classification. The latency (inference time) of a trained convolutional neural network model in actual deployment is an important indicator for determining whether the image classification model is available. In the prior art, NAS (Neural Architecture Search, automatic network structure search) is usually used to search for the convolutional neural network model to improve the model accuracy and achieve a better effect than the manually designed network. During the NAS search process, the inference time of the model is required as feedback for evaluating the quality of the model. Directly relying on the measured method is time-consuming and laborious, and a lot of useless test codes need to be implanted in the code. Therefore, it is necessary to predict the inference time of the model. As Figure 1 shown, the model inference time prediction is used in the performance evaluation strategy link in NAS. After using this method for prediction, the efficiency and accuracy of the NAS search network can be improved, thereby improving the accuracy of the searched convolutional neural network model.
[0070] In another example, the method for predicting the hierarchical inference time of a convolutional neural network is applied to the co-inference acceleration of the model at the cloud, edge, and terminal. For scenarios where the computing power at the terminal side, such as robot applications, the Internet of Things, and autonomous driving, is limited and there are many inference tasks, the edge computing node and the cloud server can be used to achieve the inference acceleration of the model without loss of accuracy. In the research on the co-inference acceleration problem at the cloud, edge, and terminal, it is usually necessary to obtain the inference time information of each layer of the model to complete the subsequent splitting process of the model. Due to the large number of layers of the deep learning model and the changing deployment environment, relying on the measured inference time of the model layer by layer is a huge workload, and it is also necessary to predict the inference time of the model. After using this method for inference time prediction, the accuracy of the cloud-edge-terminal model segmentation can be improved, so that the final segmentation effect can achieve an improvement in the model inference speed.
[0071] Specifically, asFigure 2 As shown in the figure, it includes the following steps:
[0072] S101, Data collection. Collect the hierarchical operator information of various convolutional neural network models, determine m = 6 hierarchical types according to the characteristics of the operator information, and divide the operator information into one of the 6 hierarchical types; collect the inference software framework information; collect the hardware platform information; specifically, collect the layer information of typical deep convolutional neural network models, such as the operator layer information of models like AlexNet, VGG16, ResNet, DenseNet, MobileNetv1, MobileNetv2, MobileNetv3, GoogleNet, ShuffleNet, etc., which can be printed out by a computer.
[0073] The specific division method is to divide the operator layers with the same function into one category. For example: the convolutional layer includes convolutional operators, depthwise separable convolutional operators, and dilated convolutional operators; the pooling layer includes the max pooling operator Max pool, min pooling operator, and average pooling operator; the batch normalization layer includes batch normalization operators; the fully connected layer includes fully connected operators; the activation function layer includes Relu operators, Relu6 operators, Sigmoid operators, Swish operators; the element-wise calculation layer includes Add operators, Concat operators, Multiply operators, etc.
[0074] When operator fusion occurs under a specific hardware platform and inference software framework, the fused operators are also divided into one hierarchical category.
[0075] In this embodiment, as Figure 3 shown, data collection includes the following steps:
[0076] Step 201, collect various convolutional neural network model structures, and the models include AlexNet, VGG16, ResNet
[0077] , DenseNet, MobileNetv1, MobileNetv2, MobileNetv3 and other deep learning models;
[0078] Step 202, collect the hardware platform information, including the name of the hardware platform, such as CPU, GPU, TPU, VPU, etc., the model of the hardware platform, and the hardware configuration information corresponding to the model;
[0079] Hardware parameters Hardware type CPU, GPU, TPU, VPU... Hardware processing capacity TeraFLOPS (TFLOPs), number of hardware cores, model hierarchical computational volume... Memory bandwidth Memory access bandwidth (GB / s), model hierarchical memory access volume...
[0080] Collect the inference software framework information, including the name of the inference software framework, such as TensorRT, TFLite, Openvino, MNN, NCNN, etc., and the design characteristics and related parameters under each inference software framework;
[0081] Software parameters Software framework TensorRT, TFLite, OpenVino, MNN, NCNN... Others Operator fusion strategy, operator parallel strategy, model hierarchical parameter quantity, memory occupancy...
[0082] Step 203: Obtain the hierarchical operator information of the model from the model structure, including: convolution operator, depthwise separable convolution operator, dilated convolution operator, max pooling operator, min pooling operator, average pooling operator, batch normalization operator, fully connected operator, Relu operator, Relu6 operator, Sigmoid operator, Swish operator, Add operator, Concat operator, Multiply operator;
[0083] Step 204: Determine 6 hierarchical types according to the hierarchical operator information, including: convolutional layer, pooling layer, batch normalization layer, fully connected layer, activation function layer, element-wise calculation layer;
[0084] Step 205: Divide the model hierarchical operator information into one of the 6 hierarchical types. The corresponding relationships include: the convolutional layer contains the convolution operator, depthwise separable convolution operator, and dilated convolution operator; the pooling layer contains the max pooling operator, min pooling operator, and average pooling operator; the batch normalization layer contains the batch normalization operator, the fully connected layer contains the fully connected operator, the activation function layer contains the Relu operator, Relu6 operator, Sigmoid operator, and Swish operator; the element-wise calculation layer contains the Add operator, Concat operator, and Multiply operator.
[0085] S102: Construct feature engineering. For the 6 hierarchical types, extract the layer feature parameters of the corresponding neural network model under each hierarchical type, extract the inference software framework feature parameters closely related to the inference time in the inference software framework information, and extract the hardware platform feature parameters closely related to the inference time in the hardware platform information. Combine the model layer feature parameters, inference software framework feature parameters, and hardware platform feature parameters to form the feature parameters of the 6 hierarchical types. Count the number of feature parameters of each hierarchical type. For the hierarchical types with the number of feature parameters greater than the feature parameter threshold, form n hierarchical types. For these n hierarchical types, construct the data sets of various convolutional neural network models respectively to form n sets of data sample collections. Use machine learning algorithms to train the feature screening model, and obtain the weight value coefficients of each feature under the n hierarchical types respectively. Sort according to the coefficients from large to small, and set a weight threshold to filter out the feature parameters smaller than the weight threshold. Finally, obtain the feature information after feature screening of the n hierarchical types, and jointly form the feature information of the 6 hierarchical types with the feature parameters in the remaining m - n hierarchical types; The specific hierarchical types can be seen in the following table:
[0086] Layer type Convolutional layer Convolution, depthwise separable convolution, dilated convolution, grouped convolution... Pooling layer Max pooling, min pooling, average pooling, global pooling... Batch normalization layer BatchNorm Fully connected layer FullyConnect Activation function layer Relu, Relu6, Sigmoid, Tanh, Swish... Element-wise calculation layer Add, Concat, Multiply... Operator fusion layer Conv + BN + Relu, Conv + BN...
[0087] For example, it is selected according to the number of characteristic parameters. If the number of characteristic parameters of the i-th stratification type is more than 6, feature screening is performed. If it is less than 6, all features are retained without screening. The purpose of the selection is to perform different treatments for different stratification types. For those with a small number of characteristic parameters, retaining all the information can improve the final prediction accuracy. For those with a large number of characteristic parameters, some less important feature information can be filtered out through feature screening.
[0088] Feature screening generally uses traditional machine learning algorithms, such as XGBoost, LightGBM, etc., to filter out some less important characteristic parameters, which belongs to data preprocessing, reduces the data dimension, and can speed up the subsequent training of the prediction model.
[0089] In this embodiment, as Figure 4 shown, constructing the feature engineering includes the following steps:
[0090] S301, extracting the model layer characteristic parameters for 6 stratification types, where:
[0091] The model layer characteristic parameters corresponding to the convolutional layer include: the size of the input feature picture of this layer, the number of channels of the input feature picture, the size of the convolutional kernel, the stride size, the padding size, the dilation size, the group size, and the number of channels of the output feature picture of this layer;
[0092] The model layer characteristic parameters corresponding to the pooling layer include: the size of the input feature picture of this layer, the number of channels of the input feature picture, the size of the convolutional kernel, the stride size, and the padding size;
[0093] The model layer characteristic parameters corresponding to the normalization layer include: the size of the input feature picture of this layer and the number of channels of the input feature picture;
[0094] The model layer characteristic parameters corresponding to the fully connected layer include: the size of the input feature picture of this layer, the number of channels of the input feature picture, and the number of neurons;
[0095] The model layer characteristic parameters corresponding to the activation function layer include: the size of the input picture of this layer and the number of channels of the input feature picture;
[0096] The model layer characteristic parameters corresponding to the element-wise calculation layer include: the size of the input picture of this layer and the number of channels of the input feature picture.
[0097] The size of the input feature picture refers to the height, width, etc. of the feature picture.
[0098] S302, extracting the characteristic parameters of the hardware platform, including the number of floating-point operations per second and the memory access bandwidth of the hardware; extracting the characteristic parameters of the inference software framework, including the operator fusion strategy and the parallel computing strategy;
[0099] S303, integrating the model layer feature parameters, the reasoning software framework feature parameters and the hardware platform feature parameters as feature parameters of multiple stratification types, and determining the feature parameters to determine the stratification type that needs to be feature screened among the multiple stratification types;
[0100] The fusion of feature parameters is to put the three parameters together as the overall feature parameters. If the data types of the three parameters are very different, some feature engineering processing is needed to convert the data into one dimension for subsequent processing.
[0101] For example, the first type of hierarchical feature parameters are convolutional layer feature parameters, software parameters, and hardware parameters, which include: input feature image size, number of input feature image channels, convolution kernel size, stride size, padding size, dilation size, group size, number of output feature image channels of this layer, hardware type, convolutional layer calculation amount, convolutional layer memory access amount, convolutional layer parameter amount, whether operator fusion is used, etc. Among these feature parameters, there may be numerical and categorical features. For example, the feature parameters of the convolutional layer are basically numerical, while whether operator fusion is used and the hardware type are categorical. Therefore, the categorical features need to be converted into numerical features.
[0102] In another embodiment, the fusion of feature parameters is to fuse the three parameters in multiple dimensions, that is, the feature parameters of the hardware platform correspond to the feature parameters of the inference software framework and the feature parameters of the model layer respectively, and the relationship between them can be one-to-one or one-to-many, and finally they are still fused into 6 hierarchical types.
[0103] S304, traverse each stratification type, count the number of feature parameters of each stratification type, and for stratification types whose number of feature parameters is greater than the feature parameter number threshold, use a machine learning algorithm to perform feature screening, assign weights to features, and remove feature parameters that are less than the weight threshold; for stratification types whose number of feature parameters is less than the feature parameter number threshold, retain all feature parameters, and finally obtain feature information of 6 stratification types as input for inference time prediction. The machine learning algorithm may include: XGBoost, LightGBM, random forest, and logistic regression.
[0104] S103, inference time prediction, classify the six stratification types according to the data characteristics of the feature information, divide the stratification types with the same feature information into one group (for example, if the feature parameters of the batch normalization layer and the activation function layer are the same, they can be divided into one group), divided into p groups in total, and divide the stratification types with different feature information into q groups. Use machine learning algorithms to establish p inference time prediction models for the p groups of stratification types, and establish q prediction models for the remaining q groups of stratification types, and obtain a total of p+q prediction models for predicting the inference time.
[0105] In this embodiment, as Figure 5 shown, the inference time prediction includes the following steps:
[0106] S401, obtain the 6 types of hierarchical type feature information after feature screening, and further classify the feature information of the 6 types of hierarchical types;
[0107] S402, traverse the feature information of the 6 types of hierarchical types, divide the hierarchical types with the same feature information into the same group to form p groups of hierarchical types, and the feature information in each group of the p groups contains the same model layer feature parameters, inference software framework feature parameters, and hardware platform feature parameters;
[0108] S403, divide the hierarchical types with different feature information into q groups of hierarchical types;
[0109] S404, for the p+q groups of hierarchical types, where the feature information between groups is different, use machine learning algorithms to model the p+q groups of hierarchical types respectively to train the inference time prediction model; the machine learning algorithms may include: linear regression, non-linear regression, random forest, XGBoost, neural network;
[0110] S405, based on the training of the prediction model, obtain p+q inference time prediction models to complete the inference time prediction of the 6 types of hierarchical types of the convolutional neural network.
[0111] In this embodiment, as Figure 6a shown in -c, during the actual operation of a convolutional neural network, some operator layers can be run as a fused operator after operator fusion. For example, the convolutional operator Conv1, batch normalization operator Batch Norm, and Relu operator can be combined and run as a fused operator, which can reduce memory transfer and improve the inference operation speed of the operator. Whether operator fusion can be performed is determined by the characteristics of the hardware platform and software inference framework. If it is run as a fused operator, then the aforementioned hierarchical operator types need to add a fused operator. The feature parameters of the fused operator are a combination of the feature parameters of each operator before fusion, and feature screening also needs to be performed. The processing steps after feature screening are the same as the previous steps. The detection mechanism of operator fusion is mainly for different hardware platforms and software inference frameworks. By introducing a time difference in the code of the model forward inference to determine whether operator fusion has occurred (if fusion occurs, the running time of two consecutive operators should be less than the sum of the running times of the operators running alone), it can be represented by Figure 7 and the corresponding formula:
[0112] T op1 + T op2 - T (op1,op2)> µ * min(T op1, T op2 )
[0113] where T op1 represents the individual inference time of the first operator, T op2 represents the individual inference time of the second operator, T (op1,op2) represents the joint inference time of the first operator and the second operator, and µ is a set empirical value. If there are multiple operators to be fused, the above formula is repeated. First, it is judged whether op1 and op2 can be fused. If op1 and op2 can be fused, then the fused op1 and op2 are regarded as one operator and fused with op3 for judgment.
[0114] A process for predicting the inference time of operator fusion is as Figure 8 shown:
[0115] S501, Obtain the convolutional neural network model;
[0116] S502, Operator fusion detection, search for fused operators for the operators of the convolutional neural network model by using a fusion detection mechanism (see the above formula);
[0117] S503, During the model inference operation, after operator fusion detection, during the actual inference process of the model, it will run with the fused operators under a specific inference software framework and hardware platform;
[0118] S504, Model hierarchical inference time prediction. Feature extraction is to extract the feature information of the model layer after operator fusion, and combine the information of the hardware platform and the inference software framework to form 6 types of hierarchical feature parameters. The remaining steps are the same as S102 and S103.
[0119] As Figure 9 shown, a device for predicting the hierarchical inference time of a convolutional neural network includes a data acquisition module, a feature engineering module, and an inference time prediction module that are connected in sequence.
[0120] The data acquisition module is responsible for collecting the hierarchical operator information, inference software framework information, and hardware platform information of various convolutional neural network models, determining m types of hierarchical types according to the characteristics of the collected hierarchical operator information, where m is a positive integer greater than 1, and all operator information will be classified into one of the m types of hierarchical types. The m types of hierarchical types, inference software framework information, and hardware platform information are all inputs to the feature engineering module.
[0121] The feature engineering module is responsible for extracting the model layer feature parameters corresponding to each of the m hierarchical types, extracting the inference software framework feature parameters closely related to the inference time in the inference software framework information, and the hardware platform feature parameters closely related to the inference time in the hardware platform information. It forms the feature parameters of the m hierarchical types by combining the model layer feature parameters, the inference software framework feature parameters, and the hardware platform feature parameters. Then, it uses machine learning algorithms to perform feature screening on the feature parameters of n hierarchical types among the m hierarchical types, obtaining the feature information of the n hierarchical types after feature screening, and jointly forms the feature information of the m hierarchical types with the feature parameters of the remaining m - n hierarchical types. The feature information is used as the input of the inference time prediction module.
[0122] The inference time prediction module is responsible for classifying the m hierarchical types according to the data characteristics of the feature information. It divides the hierarchical types with the same feature information into one group, a total of p groups are divided, and the remaining hierarchical types with different feature information are divided into q groups. It uses machine learning algorithms to establish p inference time prediction models for the p groups of hierarchical types respectively, and establish q prediction models for the remaining q groups of hierarchical types, obtaining a total of p + q prediction models.
[0123] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for predicting the hierarchical inference time of a convolutional neural network, characterized in that It includes the following steps: S101, data collection. Collect the hierarchical operator information of various convolutional neural network models, determine multiple hierarchical types according to the characteristics of the operator information, divide the operator information into each hierarchical type, and collect the platform framework information. The data collection includes the following steps: Step 201, collect various convolutional neural network model structures; Step 202, collect the platform framework information; Step 203, obtain the hierarchical operator information of the model from the model structure, including: convolutional operator, depthwise separable convolutional operator, dilated convolutional operator, max pooling operator, min pooling operator, average pooling operator, batch normalization operator, fully connected operator, Relu operator, Relu6 operator, Sigmoid operator, Swish operator, Add operator, Concat operator, Multiply operator; Step 204, determine multiple hierarchical types according to the hierarchical operator information, including: convolutional layer, pooling layer, batch normalization layer, fully connected layer, activation function layer, element-wise calculation layer; Step 205, divide the hierarchical operator information into one of multiple hierarchical types, and the corresponding relationships include: the convolutional layer contains the convolutional operator, depthwise separable convolutional operator, dilated convolutional operator; the pooling layer contains the max pooling operator, min pooling operator, average pooling operator; the batch normalization layer contains the batch normalization operator, the fully connected layer contains the fully connected operator, the activation function layer contains the Relu operator, Relu6 operator, Sigmoid operator, Swish operator; the element-wise calculation layer contains the Add operator, Concat operator, Multiply operator; S102, construct feature engineering. Extract the layer feature parameters of the convolutional neural network model corresponding to each hierarchical type, and extract the platform framework feature parameters related to the inference time in the platform framework information, and fuse the model layer feature parameters and the platform framework feature parameters to form the feature parameters of multiple hierarchical types. The construction of feature engineering includes the following steps: S301, extract the model layer feature parameters for 6 hierarchical types, where: The model layer feature parameters corresponding to the convolutional layer include: the size of the input feature image of this layer, the number of channels of the input feature image, the size of the convolutional kernel, the stride size, the padding size, the dilation size, the group size, the number of channels of the output feature image of this layer; The model layer feature parameters corresponding to the pooling layer include: the size of the input feature image of this layer, the number of channels of the input feature image, the size of the convolutional kernel, the stride size, the padding size; The model layer feature parameters corresponding to the normalization layer include: the size of the input feature image of this layer and the number of channels of the input feature image; The model layer feature parameters corresponding to the fully connected layer include: the size of the input feature image of this layer, the number of channels of the input feature image, the number of neurons; The model layer feature parameters corresponding to the activation function layer include: the size of the input image of this layer and the number of channels of the input feature image; The model layer feature parameters corresponding to the element-wise calculation layer include: the size of the input image of this layer and the number of channels of the input feature image; S302, extract the feature parameters of the platform framework; S303. Integrate the model layer feature parameters and the platform framework feature parameters as the feature parameters of each stratification type, and perform the determination of the feature parameters to decide the stratification types that need feature screening in each stratification type; S304. Traverse each stratification type and perform feature screening on the feature parameters; S103. Inference time prediction. Classify multiple stratification types according to the data characteristics of the feature parameters. Group the stratification types with the same feature parameters into one group, and separately group the stratification types with different feature parameters into one group. Build an inference time prediction model for each group to predict the inference time of the convolutional neural network model. The inference time prediction includes the following steps: S401. Obtain the feature parameters of multiple stratification types after feature screening, and further classify the feature parameters of multiple stratification types; S402. Traverse the feature parameters of multiple stratification types, group the stratification types with the same feature parameters into the same group to form p groups of stratification types. The feature parameters in each group of the p groups include the same model layer feature parameters and platform framework feature parameters; S403. Group the stratification types with different feature parameters into q groups of stratification types; S404. For the p+q groups of stratification types, where the feature parameters between groups are all different, use machine learning to build models for the p+q groups of stratification types respectively to train the inference time prediction model; S405. According to the training of the prediction model, obtain p+q inference time prediction models to complete the inference time prediction of multiple stratification types of the convolutional neural network.
2. The method for predicting the hierarchical inference time of a convolutional neural network according to claim 1, wherein The platform framework information includes hardware platform information and / or inference software framework information.
3. A method for predicting the hierarchical inference time of a convolutional neural network according to claim 1, characterized in that The integration of the feature parameters in S303 is to convert the model layer feature parameters and the platform framework feature parameters into the same dimension through feature engineering.
4. A method for predicting the hierarchical inference time of a convolutional neural network according to claim 1, characterized in that The integration of the feature parameters in S303 is to perform multi-dimensional integration of the model layer feature parameters and the platform framework feature parameters, that is, the feature parameters of the platform framework are respectively in a one-to-one and / or one-to-many relationship with the model layer feature parameters.
5. A method for predicting the hierarchical inference time of a convolutional neural network according to claim 1, characterized in that In S304, count the number of feature parameters of each stratification type. For n stratification types with the number of feature parameters greater than the first threshold, build datasets for various convolutional neural network models respectively to form n sets of data sample collections. Use machine learning to train the feature screening model, and obtain the weight value coefficients of each feature under the n stratification types respectively. Set a weight threshold to filter the feature parameters smaller than the weight threshold. Finally, obtain the feature parameters after feature screening of different stratification types as the input for inference time prediction.
6. A method for predicting the hierarchical inference time of a convolutional neural network according to claim 1, characterized in that Adopt the operator fusion strategy. After merging some operators, run them as a fused operator. The stratification operator type adds the fused operator. The feature parameters of the fused operator are a combination of the feature parameters of each operator before fusion. The detection mechanism of operator fusion is to judge whether operator fusion occurs by introducing a time difference in the forward inference of the model. When fusion occurs, the running time of two connected operators is less than the sum of the running times of the operators running separately. The stratification inference time prediction of operator fusion includes the following steps: S501. Obtain the convolutional neural network model; S502, Operator fusion detection, where a fusion detection mechanism is used to search for fused operators for the operators of the convolutional neural network model; S503, During model inference operation, after operator fusion detection, during the actual inference process of the model, the fused operators are used to run under the platform framework; S504, Model hierarchical inference time prediction, where the extraction of feature parameters is to extract the feature parameters of the model layers after operator fusion, combine the information of the platform framework to form feature parameters of multiple hierarchical types, and then predict the inference time according to S102 and S103.
7. A convolutional neural network hierarchical inference time prediction device adopting the convolutional neural network hierarchical inference time prediction method described in claim 1, comprising: A data acquisition module, a feature engineering module, and an inference time prediction module that are connected in sequence, characterized in that: The data acquisition module collects the hierarchical operator information of various convolutional neural network models, determines multiple hierarchical types according to the characteristics of the operator information, divides the operator information into each hierarchical type, and collects the platform framework information; Construct feature engineering, extract the layer feature parameters of the convolutional neural network model corresponding to each hierarchical type, and extract the platform framework feature parameters related to the inference time in the platform framework information, and fuse the model layer feature parameters and the platform framework feature parameters to form feature parameters of multiple hierarchical types; The inference time prediction classifies multiple hierarchical types according to the data characteristics of the feature parameters, divides the hierarchical types with the same feature parameters into one group, separately divides the hierarchical types with different feature parameters into one group, and constructs an inference time prediction model for each group to predict the inference time of the convolutional neural network model.