Neural Network Compression via Pruning Quantization and Distillation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Deep neural networks are too large for deployment on edge devices like smartphones, causing computational and memory resource constraints, and existing model compression techniques like pruning and knowledge distillation face challenges with smaller models having fewer than one million parameters, resulting in poor performance.
Innovation Solution
The method involves a processor-implemented approach that prunes an initial neural network model based on a threshold, applies quantization, generates a teacher model using the pruned weights, and trains a student model using knowledge distillation to produce a compressed neural network, incorporating iterative pruning and quantization-aware training with a learnable step size.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If model compression techniques like pruning and knowledge distillation are applied to smaller models (fewer than one million parameters), then model size is reduced, but performance deteriorates
Solution Approach 1:
The patent applies preliminary pruning to the teacher model before knowledge distillation to identify and remove redundant weights. By performing pruning in advance and using the pruned weights to guide the distillation process, the method prepares the model structure optimally before compression, preventing performance deterioration that would otherwise occur with direct compression of small models
Solution Approach 2:
The patent incorporates feedback mechanisms where the pruned weights from the teacher model are used to inform and guide the training of the student model. This feedback loop ensures that the student model learns from the optimized structure of the teacher, maintaining performance while achieving compression. The iterative refinement process allows continuous improvement based on performance metrics
2Adaptability or versatility
If deep neural networks are deployed on edge devices, then application capability is enhanced, but computational and memory resource constraints are exceeded
Solution Approach 1:
The patent extracts and removes redundant weights and connections from the neural network through systematic pruning. By identifying and eliminating unnecessary parameters, the method reduces the model size and computational complexity while retaining the essential functionality needed for edge device deployment, thus resolving the contradiction between capability and resource constraints
Solution Approach 2:
The patent changes key parameters of the neural network including weight magnitudes, activation thresholds, and network architecture parameters. By optimizing these parameters through pruning and quantization, the model achieves better efficiency-characteristic tradeoffs suitable for edge devices while maintaining application capability
Data Source
AI summary
A processor-implemented method for compressing a deep neural network model includes receiving an initial neural network model. The initial neural network is pruned based on a first threshold to generate a pruned network and a set of pruned weights. A quantization process is applied to the pruned network to produce a pruned and quantized network. A teacher model is generated by incorporating the pruned set of weights with the pruned network. In addition, an initial student model is generated from the quantized and pruned network. The initial student model is trained using the teacher model to output a trained student model.


