A Method for Constructing a Runtime Deep Neural Network for Edge Inference
By dividing DNN into parallel training and optimization block combinations of sparse blocks, the efficiency and accuracy problems of derived models generated on edge devices are solved, and efficient model scaling and resource utilization are achieved.
Patent Information
- Application Number
- CN202110586606.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-27
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-05-27
AI Technical Summary
It is difficult for the prior art to efficiently generate derivative models with different fine-grained sizes on edge devices, and there are problems such as large loss of inference accuracy, long training time and high memory switching costs during model scaling.
By dividing the deep neural network into multiple sparse blocks, retraining them separately, and using integer linear programming optimization to select appropriate derivative block combinations, a new DNN model is built to adapt to resource changes.
It realizes the rapid generation of derived blocks with different sparseness, reduces memory switching costs, improves model scaling efficiency, adapts to multiple DNN models, balances accuracy and delay, and has good universality and generalization.
Smart Images

Figure CN115409178B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a method for constructing a runtime deep neural network for edge inference. Background Art
[0002] With the continuous development of artificial intelligence technology, the application of deep learning technology (DL: Deep Learning, which uses deep neural networks to learn the mapping relationship between input data and output results) in the field of computer vision has become very common, including image classification, object detection, and object tracking. Specifically, network models such as VGG16, YOLOv3, and ResNet18 are constructed, and samples with annotation values are processed. The samples are fed into the network in batches for feedforward calculation to obtain the calculation results of the network. The gap between the results and the true annotation values of the samples is measured, expressed as an error value (Loss), and the error value is used to update the weight parameters in the network according to the backpropagation algorithm to achieve the gradient descent process. In this way, through continuous optimization and adjustment of the deep neural network model, the knowledge contained in the samples can be "learned". Using the constructed DNN (Deep Neural Network), it is deployed in the operating environment to generate accurate inferences for samples with unknown true values.
[0003] In recent years, in order to improve the performance of DNNs, especially the generalization performance when inferring new samples, researchers have adopted methods such as expanding the scale of weight parameters and increasing the number of network layers. The inference performance (such as accuracy) of DNN models is greatly affected by their size. Larger models have higher inference performance, and reducing the size will inevitably reduce the model's inference performance. However, nowadays, the trend of intelligent applications serving edge devices (terminal devices on the edge of the network, such as smartphones) is becoming increasingly obvious. Such applications require extremely low processing latency. Therefore, it is necessary to directly deploy DNNs on edge devices (such as smartphones, tablets, drones, and other embedded systems) that are closer to the data generation source. On the one hand, edge devices are restricted by factors such as size and mobility and cannot provide sufficient energy (such as battery power) and computing resources (such as powerful processors and sufficient memory space) for large-sized DNNs to perform inference tasks. On the other hand, edge devices do not run a single application but multiple intelligent applications simultaneously. That is, multiple DNNs need to run simultaneously, which undoubtedly intensifies the competition for resources between different tasks and makes the operating environment on the device constantly change. This poses a huge challenge to the deployment of traditional DNNs on the edge side. Therefore, it is necessary to scale the deep neural network according to the operating conditions of edge devices. When resources are sufficient, a larger-sized DNN is constructed for inference to improve accuracy. When resources are insufficient, a small-sized DNN is constructed to meet the processing latency while achieving the highest possible inference accuracy or achieving the shortest possible processing latency while meeting the accuracy requirements.
[0004] To address the contradiction between the DNN resource requirements and the environmental conditions of edge devices, the deep neural network scaling technology constructs a volume-compressed DNN model when the operating environment resources decrease to reduce the computational amount required for DNN edge inference (i.e., edge inference, using the deep neural network to perform inference and inference tasks on edge devices) and reduce the resource requirements of DNNs. Structured pruning is a commonly used technique that can derive multiple smaller-sized models from the original DNN. When a DNN needs to be run on an edge device, a derived model with a suitable size (a deep neural network model composed of derived blocks) is selected from the scaling results to support the needs of user applications. However, the existing structured pruning techniques have the following problems:
[0005] (1) There are a limited number of derived models. Since traditional methods compress the model as a whole, a series of derived models can only provide a coarse granularity difference, that is, the volume gap between derived models is large. When facing different resource conditions, it is difficult to select a very suitable derived model, which results in a lot of unused resources, unable to achieve high resource utilization, and thus faces a greater loss of inference performance (such as result accuracy). According to experiments, when the available resources gradually decrease, traditional methods will switch the current model to a neighboring smaller-volume derived model, and this switching process leads to a significant drop in the inference accuracy of the deep neural network. When the available memory is less than 50MB, using the filter pruning and NestDNN methods will replace the model from 37MB to 28MB, and using the low-rank matrix factorization method will replace the model from 38MB to 24MB. This causes a large leap in the model volume. On the one hand, it inevitably leads to a large loss in the inference accuracy of the replaced model. On the other hand, due to the sharp reduction in the volume of the replaced model, although the memory condition is met, it also results in low memory utilization. In contrast, the ideal model scaling (adjusting the volume of the model, using techniques of model compression and magnification to construct a new deep neural network model) method will enable the scaled-down model to meet higher resource utilization as much as possible and thus incur a very low loss of inference accuracy.
[0006] (2) The training of derived models is slow. To generate derived models that can be used for scaling, the compressed derived models need to be retrained, which requires the use of a powerful GPU server for several hours. For the filter pruning and low-rank matrix factorization methods, training a derived model using a Quadro RTX8000 graphics card with 48GB of memory takes 7 to 8 hours. To achieve high resource utilization and obtain more derived models with a smaller granularity difference, it takes several months of retraining, which further leads to the ineffective improvement of the number of derived models in traditional methods.
[0007] (3) The cost of model scaling is high. In the running environment, a large number of derived models will be scaled according to resource changes, that is, switched to a larger-volume or smaller-volume model that is more suitable for the resource conditions. The overall replacement of the model will cause a non-negligible memory switching cost on edge devices, further exacerbating the energy tension situation. Although the NestDNN method can partially solve this problem by designing a nested structure of derived models to reduce the number of memory page faults, it leads to two additional cost increases: the generation of nested models extends the training time by 50% to 200% compared to the standard filter pruning method, and the enlargement operation of the model requires additional calculations to generate larger network layers.
[0008] For example, the patent application with the patent number CN202011517057.0 needs to significantly modify the original neural network model. It splits the convolutional layer into two convolutional layers, namely a depth convolutional layer and a pointwise convolutional layer, which reduces the generality of the method. Similarly, taking the entire model as a granularity, a fixed - volume compressed model needs to be generated before running the DNN, facing the problems described above.
[0009] The patent application with the patent number CN201911348985.6 uses sparse learning to train the model to be compressed. And it needs to add the regularization of the scaling factor in each channel of the neural network as a penalty term to the training loss function, which modifies the training process and the loss function of the original model. At the same time, it uses a genetic algorithm to search for a large number of generated sub - networks to obtain the optimal sub - network, which further increases the generation time cost of the sub - network (derived model).
[0010] The patent application with the patent number CN201711097406.6 reconstructs the underlying features and high - level features of the deep neural network model with "layers" as the granularity, but still uses a unified compression ratio and cannot efficiently generate a large number of derived models with small differences in sparsity. Summary of the Invention
[0011] The purpose of the present invention is to provide a method for constructing a runtime deep neural network for edge inference that can overcome the above - mentioned technical problems. The technical problems solved by the method of the present invention are how to efficiently generate derived models with fine - grained differences for scaling within an acceptable time and how to effectively manage the derived models according to dynamic environmental changes during runtime to achieve lower model switching costs and higher performance, such as inference accuracy and processing latency, in the resource - constrained environment of edge devices.
[0012] The method of the present invention includes the following steps:
[0013] Step 1: Generate blocks with different sparsities (i.e., Block: a block composed of one or more DNN layers) from the original DNN:
[0014] Step 1.1: Analyze the given original DNN model, and form each layer or multiple layers in the neural network into a block b i (1 ≤ i ≤ m), until the entire DNN is divided, obtaining m blocks;
[0015] Step 1.2: Set s1, s2,... i sparsities for b and apply the standard pruning and filtering technique to each block b i , generating n iA derived block b i,1 ,b i,2 ,..., Each derived block corresponds to a sparsity level. The weight parameters in the neural network model are represented as matrices. The derived block is a derivative generated by pruning, i.e., compressing the calculations, on the blocks in the original deep neural network model;
[0016] Step 1.3, construct the training environment for each block b i For each block in the original DNN, construct the training environment, i.e., the input data and output data required for each block;
[0017] Convert each Mini - Batch in the original training samples through the calculations of the original DNN into the input data required for each block b i as the input for training the derived block of each block b i ; convert each Mini - Batch through the calculations of the original DNN into the output data required for each block b i as the target output value for training the derived block of each block b i , and obtain the input and output required for each block, i.e., the training environment;
[0018] Step 2, retrain the blocks with different sparsity levels:
[0019] Step 2.1, take out the training environment of each block b i obtained in Step 1.3 as the training environment for the corresponding derived block b i,1 ,b i,2 ,…, ;
[0020] Step 2.2, perform a feed - forward calculation on the input data in the training environment using each derived block to obtain the output of the derived block;
[0021] Step 2.3, obtain the error value of this calculation using the output value of the derived block and the target output value in the corresponding training environment;
[0022] Step 2.4, back - propagate the error value and update the parameters of the derived block according to the gradient descent algorithm. When the training end condition is reached, i.e., the maximum number of iterations is reached or the error value is less than the preset threshold, end the training of the derived block; otherwise, return to Step 2.2;
[0023] Step 3, perform a multi - dimensional performance evaluation on the trained derived blocks to prepare for selecting appropriate derived blocks in the running environment:
[0024] Step 3.1, obtain the number of parameters of the derived block b i,j ;
[0025] Step 3.2: Generate a memory occupancy scale metric from the parameter scale;
[0026] Step 3.3: Accuracy loss evaluation module: The volume of the derived block is reduced to varying degrees compared to the block in the original DNN. The inference accuracy of the derived block for the input data decreases. The accuracy loss evaluation module evaluates the degree of decline in accuracy for blocks with different sparsities. For the same derived block, the final effect, i.e., the accuracy loss, produced when it is placed in models with different configurations is also different. That is, the degree of accuracy decline caused by the derived block is related to its environment (e.g., the sparsity of other derived blocks in the final model). Use the accuracy loss evaluation module to perform accuracy evaluation and select k representative sparsities from the n sparsities of the derived block; i
[0027] Step 3.4: For each block b i , select the derived blocks corresponding to the k sparsities respectively and form a complete DNN model with the derived blocks of the same sparsity;
[0028] Step 3.5: Evaluate the accuracy loss of each derived block b i,j . Replace b i,j successively into the k models with different sparsities obtained in Step 3.4 to form the models to be tested Replace the original block b of b i,j successively into the k models with different sparsities to form the models to be tested i
[0029] Step 3.6: Test the models to be tested on the test set to obtain the inference accuracy of each model to be tested;
[0030] Step 3.7: Combine the results of b i and b i,j (j = 1, 2,..., n i ) on the models to be tested with the same sparsity to form a tuple. That is, each derived block will obtain k tuples corresponding to the k sparsities. Each original block b i will obtain n i × k tuples. Regard these tuples as the accuracy evaluation matrix of b i . Each row corresponding to each derived block b i,j has k element values, and each element value is a binary tuple. Take the difference between the two values in the tuple. The difference represents the difference in accuracy loss between the original block b i and the derived block b i,j under the model to be tested with a certain sparsity. Take the average of the k accuracy difference results of each derived block as the accuracy loss evaluation result of the derived block;
[0031] Step 3.8, Delay evaluation module, DNN model compression to reduce the computational load and accelerate the execution of the DNN inference process. On actual devices, the processing delay of each block is affected by aspects of the operating environment (such as resource contention, performance interference) and random system events (such as system maintenance, garbage collection). Using the ImageNet dataset and ResNet18 model to test on edge devices (Raspberry Pi 4B), the absolute value of the reduced value of the derived block delay fluctuates greatly. For example, a derived block with a sparsity of 0.8 reduces the delay by 4 seconds when there is only 1 running application, while when there are 5 running applications, it can reduce the delay by 17 seconds. The delay evaluation module uses "percentage" rather than "absolute value" for evaluation. For derived block b i,j , measure the delay of the currently running corresponding derived block b i,v ;
[0032] Step 3.9, According to the attribute relationship (such as parameter scale, structural difference) between derived block b i,j and b i,v , calculate the delay of b i,j ;
[0033] Step 4, Monitor the status of the operating environment. When the status of the operating environment changes, re-seek the optimal combination scheme of derived blocks according to the requirements. The steps are as follows:
[0034] Step 4.1, Model an integer linear programming problem to maximize the inference accuracy of the scaled DNN. Then, at the k-th scaling, the objective function is as follows in formula (1):
[0035]
[0036] In formula (1), the decision variable takes a value of 1 or 0. 1 means selecting derived block B i,j , 0 means not selecting. The first constraint condition is as follows in formula (2):
[0037]
[0038] The overall processing delay of the DNN obtained by accumulating the processing delays of the selected derived blocks does not exceed the set delay upper limit t max , as follows in formula (3):
[0039]
[0040] The volume s of the model before scaling base and the reduced volumes (S i,jThe difference between the sum and the scaled volume does not exceed the available memory s in the current operating environment max , as shown in the following formula (4):
[0041]
[0042] For each original block b in the DNN i , at the same time, select a corresponding derivative, as shown in the following formula (5):
[0043]
[0044] Based on the above formula for modeling, an optimization problem with 4 constraints is obtained, and it is modeled for maximizing the inference accuracy;
[0045] Step 4.2, Solve the integer linear programming problem, relax the integer constraints and convert the integer linear programming problem into a linear programming problem;
[0046] Step 4.3, Use the branch and bound method to search for the optimal solution in polynomial time complexity, that is, for each original block b i of value;
[0047] Step 5, Scale the DNN model in the operating environment according to the derivative block combination scheme obtained in Step 4, replace the derivative blocks, and construct a new DNN:
[0048] Step 5.1, When obtaining the derivative block combination scheme P of the DNN running in the system current , then go to Step 5.2, otherwise go to Step 5.3;
[0049] Step 5.2, Compare the new derivative block combination scheme P new with the m blocks in P current one by one: When the derivative blocks used for the same position b in the DNN i are different, replace the derivative blocks at the corresponding positions in the memory with the derivative blocks in the new scheme; when the derivative blocks used for the same position b in the DNN i are the same, do not operate, repeat Step 5.2 until the entire DNN has been replaced with the new scheme, and go to Step 5.4;
[0050] Step 5.3, Put each derivative block in the new scheme into the memory to construct a complete DNN model;
[0051] Step 5.4, Use the newly constructed DNN to perform the inference task.
[0052] The superior effects of the present invention are:
[0053] 1. The method of the present invention can quickly derive blocks with different sparsities from the original DNN and efficiently retrain each sparse block;
[0054] 2. The method of the present invention takes "blocks" rather than "the whole model" as the granularity, and can train a certain number of blocks in a shorter time and combine them into a large number of finally optional derived models with different performances to finely meet the resource changes of the model running environment without retraining the whole model with different sparsities; the model scaler provided by the method of the present invention can solve the derived block combination optimization problem with polynomial time complexity when the resource conditions change in the running environment of the system, calculate the optimal solution according to the performance indicators of each derived block. Since the method of the present invention only needs to put the derived blocks different from those in the running DNN in the new derived block combination scheme into the memory and replace the original derived blocks at the corresponding positions in the memory, the construction of the new DNN can be completed. Therefore, there is no need to perform memory read / write page swapping on the whole model, which can greatly reduce the memory switching cost and improve the model scaling efficiency, while fully meeting the resource condition constraints and optimization objectives of the system;
[0055] 3. The framework provided by the method of the present invention can be widely applied to various DNN models, can adapt to multiple optimization objectives of maximizing model accuracy, minimizing processing delay, and balancing the two, and has good generality and generalization;
[0056] 4. The derivation scheme with block granularity of the method of the present invention provides up to 1000 - 10000 times of model scaling selection without increasing additional costs. The derived model generation scheme proposed by the method of the present invention takes "blocks" as the granularity, and generates derived blocks with different sparsities for each part of a DNN model. At the same time, the generation processes of these derived blocks can be carried out simultaneously, and different derived blocks can be trained in parallel without increasing additional training costs;
[0057] 5. The training processes of different derived blocks of the method of the present invention do not interfere with each other and can be executed in parallel, while the traditional retraining method with the whole model as the granularity needs to retrain the whole model, with a large amount of calculation and low parallelism, and cannot reach the execution efficiency of the method of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 is a schematic diagram of a deep neural network scaling framework based on block granularity of the method of the present invention;
[0059] Figure 2 is a flowchart of step 1 of the method of the present invention;
[0060] Figure 3 is a flowchart of step 2 of the method of the present invention;
[0061] Figure 4 It is the flowchart of step 3 of the method described in the present invention;
[0062] Figure 5 It is the flowchart of step 4 of the method described in the present invention;
[0063] Figure 6 It is the flowchart of step 5 of the method described in the present invention. Detailed implementation manners
[0064] The implementation manners of the present invention will be described in detail below with reference to the accompanying drawings. As Figure 1-6 shown, the method described in the present invention includes the following steps:
[0065] Step 1, generating blocks with different sparsities composed of one or more DNN layers from the original DNN (i.e., Block: a block composed of one or more layers in a deep neural network):
[0066] Step 1.1, analyzing the given original DNN model, and forming each layer or multiple layers in the neural network into a block b i (1 ≤ i ≤ m), until the entire DNN is divided to obtain m blocks;
[0067] Step 1.2, setting s1, s2,... i sparsities for b , and applying the standard pruning filtering technique to each block b i to generate n i derived blocks b i,1 , b i,2 ,... Each derived block corresponds to a sparsity, the weight parameters in the neural network model are represented as matrices, and the derived blocks are generated by pruning, i.e., compressing the calculations of the blocks in the original deep neural network model;
[0068] Step 1.3, constructing the training environment for each block b i , and constructing the training environment for each block in the original DNN, i.e., the input data and output data required for each block;
[0069] Converting each Mini - Batch in the original training samples through the calculation of the original DNN into the input data required for each block b i as the input for training the derived blocks of each block b i ; converting each Mini - Batch through the calculation of the original DNN into the output data required for each block b i as the target output value for training the derived blocks of each block b i , and obtaining the input and output required for each block, i.e., the training environment;
[0070] Step 2, retrain the blocks with different sparsities:
[0071] Step 2.1, take out each block b obtained in Step 1.3 i 's training environment as the corresponding derived block b i,1 , b i,2 ,... 's training environment;
[0072] Step 2.2, perform feed - forward calculation on the input data in its training environment using each derived block to obtain the output of the derived block;
[0073] Step 2.3, use the output value of the derived block and the target output value in the corresponding training environment to obtain the error value of this calculation;
[0074] Step 2.4, back - propagate the error value, update the parameters of the derived block according to the gradient descent algorithm. When the training end condition is reached, that is, the maximum number of iterations is reached or the error value is less than the preset threshold, end the training of the derived block; otherwise, return to Step 2.2;
[0075] Step 3, perform multi - dimensional performance evaluation on the trained derived blocks to prepare for selecting suitable derived blocks in the running environment:
[0076] Step 3.1, obtain the number of parameters of the derived block b i,j ;
[0077] Step 3.2, generate a memory occupancy scale index from the parameter scale;
[0078] Step 3.3, accuracy loss evaluation module: The volume of the derived block is reduced to varying degrees compared to the block in the original DNN, and the inference accuracy of the derived block for the input data decreases. The accuracy loss evaluation module evaluates the degree of decline in accuracy of blocks with different sparsities. For the same derived block, the final effect, that is, the accuracy loss, generated when it is placed in models with different configurations is also different. That is, the degree of accuracy decline caused by the derived block is related to its environment (for example, the sparsity of other derived blocks in the final model). Use the accuracy loss evaluation module to select k representative sparsities from the n i sparsities of the derived block;
[0079] Step 3.4, for each block b i , respectively select the derived blocks corresponding to the k sparsities and form a complete DNN model with the derived blocks of the same sparsity;
[0080] Step 3.5, evaluate the accuracy loss of each derived block b i,j and take bi,j Replace them into the k models with different sparsities obtained in step 3.4 in sequence to form the models to be tested. Take b i,j 's original block b i Replace them into the k models with different sparsities in sequence to form the models to be tested.
[0081] Step 3.6: Test the models to be tested on the test set to obtain the inference accuracy of each model to be tested.
[0082] Step 3.7: Take b i and b i,j (j = 1, 2,..., n i ) The results on the models to be tested with the same sparsity are combined into a tuple. That is, each derived block will obtain k tuples corresponding to k sparsities. Each original block b i Obtain n i ×k tuples. Regard these tuples as the accuracy evaluation matrix of b i . In the matrix, each row corresponding to each derived block b i,j has k element values. Each element value is a binary tuple. Calculate the difference between the two values in the tuple. The difference represents the difference in the accuracy loss between the original block b i and the derived block b i,j under a certain sparsity of the model to be tested. Take the average of the k accuracy difference results of each derived block as the evaluation result of the accuracy loss of the derived block.
[0083] Step 3.8: Delay evaluation module. DNN model compression is used to reduce the computational amount and accelerate the execution of the DNN inference process. On actual devices, the processing delay of each block is affected by aspects such as the running environment (such as resource competition, performance interference) and random system events (such as system maintenance, garbage collection). Test using the ImageNet dataset and the ResNet18 model on edge devices (Raspberry Pi 4B). The absolute value of the reduction in the delay of the derived block fluctuates greatly. For example, a derived block with a sparsity of 0.8 reduces the delay by 4 seconds when there is only 1 running application, while when there are 5 running applications, it can reduce the delay by 17 seconds. The delay evaluation module uses the "percentage" rather than the "absolute value" method for evaluation. For the derived block b i,j , measure the delay of the corresponding derived block b i,v that is currently running.
[0084] Step 3.9: According to the attribute relationship (such as parameter scale, structural difference) between the derived block b i,j and b i,v , calculate the delay of b i,j .
[0085] Step 4: Monitor the status of the running environment. When the status of the running environment changes, re-seek the optimal combination scheme of derived blocks according to the requirements. The steps are as follows:
[0086] Step 4.1: Model an integer linear programming problem to maximize the inference accuracy of the scaled DNN. When scaling for the k-th time, the objective function is as shown in the following formula (1):
[0087]
[0088] In formula (1), the decision variable takes a value of 1 or 0. 1 means selecting the derived block B i,j , and 0 means not selecting. The first constraint condition is the following formula (2):
[0089]
[0090] The overall processing delay of the DNN obtained by accumulating the processing delays of the selected derived blocks does not exceed the set delay upper limit t max , as shown in the following formula (3):
[0091]
[0092] The volume s before model scaling base minus the sum of the reduced volumes (S i,j ) of the selected derived blocks is the scaled volume, which does not exceed the available memory s in the current running environment max , as shown in the following formula (4):
[0093]
[0094] For each original block b in the DNN i , at the same time, select a corresponding derivative, as shown in the following formula (5):
[0095]
[0096] Based on the above formulas for modeling, an optimization problem with 4 constraint conditions is obtained, and it is modeled for maximizing the inference accuracy;
[0097] Step 4.2: Solve the integer linear programming problem, relax the integer constraints, and convert the integer linear programming problem into a linear programming problem;
[0098] Step 4.3: Use the branch and bound method to search for the optimal solution in polynomial time complexity, that is, for each original block b i of value;
[0099] Step 5. Scale the DNN model in the running environment according to the derived block combination scheme obtained in Step 4, replace the derived blocks, and construct a new DNN:
[0100] Step 5.1. When the derived block combination scheme P of the DNN running in the system is obtained current , go to Step 5.2; otherwise, go to Step 5.3.
[0101] Step 5.2. Compare the new derived block combination scheme P new with P current one by one for the m blocks in it: When the derived blocks used for the same position b i in the DNN are different, replace the derived blocks at the corresponding positions in the memory with the derived blocks in the new scheme; when the derived blocks used for the same position b in the DNN are the same, do not operate, repeat Step 5.2 until the entire DNN has been replaced with the new scheme, and then go to Step 5.4.
[0102] Step 5.3. Put each derived block in the new scheme into the memory and construct a complete DNN model.
[0103] Step 5.4. Use the newly constructed DNN to perform inference tasks.
[0104] On small edge devices, such as Huawei smartphones with the ARM Cortex architecture, Xiaomi smartphones with the Qualcomm Kryo architecture, and Raspberry Pi 4B with the ARM architecture, use the representative ResNet18 model to implement the image classification task widely used in the mobile vision system. In a specific embodiment, using the method of the present invention in Step 1, divide the neural network layers of the ResNet18 model, and delimit every two adjacent convolutional layers except the first convolutional layer as a block, so that the granularity of subsequent compression is "block" rather than the whole model, and at the same time, the residual unit structure of ResNet18 itself can be maintained, giving full play to the design advantages of the DNN itself. Set 5 derived blocks with sparsities starting from 0.1, ending at 0.5, and with an interval of 0.1 for each original block. For the 8 blocks of the model, a total of 40 different sparsity-derived blocks can be generated, that is, 5 8 kinds of derived models available for scaling. Compared with the traditional method, the method of the present invention can propose up to 1000 - 10000 times more candidate derived models without increasing the training time.
[0105] Preprocess the ImageNet dataset and convert it into the sample input format required by the ResNet18 model. Input the samples into the original ResNet18 model. Use the calculation result of the neural network part before the first block as the input data of the first block. Use the output result calculated by the first block as the target output value of the derived block of the first block. Use the above input and output as the training environment of the derived block of the first block. And so on. Obtain the training environment of each derived block in turn.
[0106] In a specific embodiment, in step 2, each derived block is retrained: the input data and target output value in the corresponding training environment are obtained; the input data is input into the derived block in the form of a mini-batch, and an inference result is obtained through feedforward calculation; the inference result is compared with the target value to calculate the error value; the error value is back-propagated, and the parameters of the derived block are adjusted using a gradient descent algorithm; the above steps are repeated until a preset number of iterations is reached, or the inference result of the derived block is close enough to the target value, and the retraining is completed. The above retraining of different derived blocks can be performed in parallel and speeds up the retraining process of the DNN.
[0107] In a specific embodiment, the following evaluation is performed in step 3: a memory evaluation module is used to obtain a memory usage scale for each derived block according to its parameter scale; an accuracy evaluation module is used to construct a test model with k=3 sparsities for each original block and the corresponding derived block, and each test model is tested on the ImageNet test set to obtain three groups of test results, i.e., the inference accuracy of each model, and the results of the derived block and its original block on the test model with each sparsity are subtracted, and the average of the three groups of differences is calculated as the accuracy evaluation index of the derived block; and a delay evaluation module is used to obtain the processing delay of the derived block actually running.
[0108] In a specific embodiment, in step 4, the model scaler implemented by the method of the present invention will continuously monitor the operating status of the system. When the resource conditions (such as the amount of available memory) change, the results of the three dimensions (memory usage, accuracy loss, and processing delay) of evaluating each derived block in step 3 are used as inputs of the integer linear programming model, and the optimal solution is solved based on linear relaxation and branch-and-bound algorithm, as shown in the following formula (6):
[0109]
[0110] Among them, the first model scaling matrix B (1) The value of the i-th row and j-th column in indicates whether the j-th derived block is selected to assemble the complete DNN for the i-th original block.
[0111] In a specific embodiment, according to the optimization result, in step 5, the new derived block combination scheme is compared with the derived blocks (b 1,1 ,b 2,2 ,b 3,2 ,b 4,2 ,b 5,4 ,b 6,4 ,b 7,3 ,b 8,4 ) that are running in the system. The new derived blocks (b 2,1 ,b 3,3 ,b 8,5 ) are put into memory and respectively replaced at the corresponding positions (the positions of the 2nd, 3rd, and 8th blocks) in the running DNN. The model scaling method provided by the method of the present invention is granular with "blocks", and can replace only the derived blocks that are different from those in the existing memory in the new scaling scheme, without replacing the entire model. Compared with the traditional method, it can achieve an average energy consumption saving of 78.42% and 56.38% (overall 71.07%) in model switching on two types of smartphones and embedded systems on the edge side respectively.
[0112] After completing the DNN scaling, the newly constructed DNN will continue to perform the inference task. Since the method of the present invention can efficiently generate a large number of derived models with fine-grained volume differences, the new DNN can achieve a very high resource utilization rate, such as the memory occupancy ratio. When setting the processing delay constraint, the method of the present invention can improve the average inference accuracy by 31.74% compared with the traditional method, and even improve the accuracy by 68.77% in the case of severe memory constraints; when setting the inference accuracy constraint, it can reduce the processing delay by 27.88%; when no constraint is set, it can reduce the average accuracy loss by 36.47% and at the same time reduce the average processing delay by 16.31%.
[0113] The above is only the specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the scope disclosed by the present invention should be covered by the protection scope of the claims of the present invention.
Claims
1. A method for constructing a runtime deep neural network for edge inference, characterized in that It includes the following steps: Step 1, generate blocks with different sparsity levels consisting of one or more DNN layers from the original DNN: Step 1.1, analyze the given original DNN model, and form each single layer or multiple layers in the neural network into a block b i , where 1 ≤ i ≤ m, until the entire DNN is partitioned to obtain m blocks; Step 1.2, for b i Set the sparsity, and apply the standard pruning filtering technique to each block b i to generate n i derived blocks Each derived block corresponds to a sparsity. The weight parameters in the neural network model are represented as matrices. The derived blocks are generated by pruning, i.e., compressing the calculations, of the blocks in the original deep neural network model; Step 1.3, construct each block b i 's training environment, and construct a training environment for each block in the original DNN, that is, the input data and output data required for each block; Convert each Mini - Batch in the original training samples through the calculation of the original DNN into each block b i The required input data is used as the input for training each derived block of block b i Convert each Mini - Batch through the calculation of the original DNN into each block b i The required output data is used as the target output value for training each derived block of block b i Obtain the input and output required for each block, that is, the training environment Step 2, retrain the blocks with different sparsity levels; Step 2.1, take out each block b obtained in Step 1.3 i 's training environment as the corresponding derived block 's training environment; Step 2.2, perform a feed-forward calculation on the input data in its training environment using each derived block to obtain the output of the derived block; Step 2.3, obtain the error value of this calculation using the output value of the derived block and the target output value in the corresponding training environment; Step 2.4, backpropagate the error value, update the parameters of the derived block according to the gradient descent algorithm. When the training end condition is reached, i.e., the maximum number of iterations is reached or the error value is less than the preset threshold, the training of the derived block ends. Otherwise, return to Step 2.2; Step 3, perform multi-dimensional performance evaluation on the trained derived blocks to prepare for selecting appropriate derived blocks in the running environment; Step 4, monitor the status of the running environment. When the status of the running environment changes, re-seek the optimal derived block combination scheme according to the requirements; Step 5, scale the DNN model in the running environment according to the derived block combination scheme obtained in Step 4, replace the derived blocks, and construct a new DNN; Test using the ImageNet dataset and the ResNet18 model on edge devices.
2. The method for constructing a runtime deep neural network for edge inference according to claim 1, wherein The said Step 3 includes the following steps: Step 3.1, obtain the number of parameters of the derived block b i,j ; Step 3.2, generate a memory occupancy scale metric from the parameter scale; Step 3.3, accuracy loss evaluation module: The volume of the derived block is reduced to varying degrees compared with the block in the original DNN. The inference accuracy of the derived block for the input data decreases. The accuracy loss evaluation module evaluates the degree of decline in accuracy of blocks with different sparsities. For the same derived block, the final effect (i.e., accuracy loss) produced when it is placed in models with different configurations is also different. That is, the degree of accuracy decline caused by the derived block is related to the sparsity of other derived blocks in the final model. Use the accuracy loss evaluation module to perform accuracy evaluation and select k representative sparsities from the n i sparsities of the derived block; Step 3.4, for each block b i , respectively select the derived blocks corresponding to k sparsities Form the complete DNN model by combining the derived blocks with the same sparsity; Step 3.5, evaluate the accuracy loss of each derived block b i,j and successively replace b i,j into the k models with different sparsities obtained in Step 3.4 to form the models to be tested Successively replace the original block b i,j of b i into the k models with different sparsities to form the models to be tested Step 3.6, test the model to be measured on the test set to obtain the inference accuracy of each model to be measured; Step 3.7, take b i and b i,j , where j = 1, 2, …, n i , and the results on the model under test with the same sparsity form a tuple. That is, each derived block will obtain k tuples corresponding to k sparsities. Each original block b i obtains n i ×k tuples. These tuples are regarded as the accuracy evaluation matrix of b i . Each row in the matrix corresponding to each derived block b i,j has k element values. Each element value is a binary tuple. Take the difference between the two values in the tuple. The difference represents the difference in accuracy loss between the original block b i and the derived block b i,j under the model under test with a certain sparsity. Take the average of the k accuracy difference results of each derived block as the accuracy loss evaluation result of the derived block; Step 3.8, Delay Evaluation Module, DNN model compression to reduce the computational load and accelerate the execution of the DNN inference process. On actual devices, the processing delay of each block is affected by aspects such as the operating environment and random system events. Using the ImageNet dataset and ResNet18 model to test on edge devices, the absolute value of the derived block delay reduction value fluctuates greatly. The delay evaluation module uses "percentage" rather than "absolute value" for evaluation. For the derived block b i,j , measure the delay of its corresponding derived block b i,v that is currently running; Step 3.9, according to the attribute relationship of the derived blocks b i,j and b i,v , calculate the latency of b i,j .
3. A method for constructing a runtime deep neural network for edge inference according to claim 1, characterized in that, The said Step 4 includes the following steps: Step 4.1, model an integer linear programming problem to maximize the inference accuracy of the scaled DNN. Then, at the k-th scaling, the objective function is as follows in formula (1): The decision variable in formula (1) takes a value of 1 or 0, where 1 indicates the selection of the derived block B i,j , 0 indicates non-selection, and the first constraint is the following formula (2): The processing delay of the overall DNN obtained by cumulatively adding the processing delays of the selected derived blocks does not exceed the set delay upper limit t max , as shown in the following formula (3): The volume s before model scaling base The reduced volume S of each selected derived block i,j , the difference between their sum and the scaled volume does not exceed the available memory s in the current operating environment max , as shown in the following formula (4): For each original block b in the DNN i , at the same time, select a corresponding derivative as shown in the following formula (5): Based on the above formula for modeling, an optimization problem with 4 constraint conditions is obtained, and it is modeled for maximizing the inference accuracy; Step 4.2, solve the integer linear programming problem, relax the integer constraints and convert the integer linear programming problem into a linear programming problem; Step 4.3, using the branch and bound method, search for the optimal solution under polynomial time complexity, that is, for each original block b i of value 4. A method for constructing a runtime deep neural network for edge inference according to claim 1, characterized in that The said Step 5 includes the following steps: Step 5.1, when obtaining the derived block combination scheme p of the DNN running in the system current , go to Step 5.2, otherwise go to Step 5.3; Step 5.2, combine the new derived block combination scheme p new with p current and compare them with the m blocks in p one by one: When the derived blocks used for the same position b in the DNN i are different, replace the derived block at the corresponding position in the memory with the derived block in the new scheme; when the derived blocks used for the same position b in the DNN i are the same, do not operate, repeat Step 5.2 until the entire DNN has been replaced with the new scheme, and go to Step 5.4; Step 5.3, put each derived block in the new scheme into memory to construct a complete DNN model; Step 5.4, use the newly constructed DNN to perform inference tasks.
Citation Information
Patent Citations
General miniaturization method of deep neural network
CN107748913A
Neural network pruning method based on combination of sparse learning and genetic algorithm
CN111105035A
Neural network compression method for remote sensing image target detection
CN112488070A