Deep Neural Network Compression via Joint Pruning and Quantization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods of pruning and quantizing neural networks as separate operations complicate the optimization process, limiting the selection of optimal models and increasing computational burden.
Innovation Solution
A method that jointly prunes and quantizes the weights and output feature maps of a neural network using an analytic threshold function, optimizing the parameters through a cost function to balance accuracy and complexity, allowing a wider range of solutions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If pruning and quantization are performed as separate independent operations, then the neural network can be compressed, but the optimization process becomes more complex and the selection of optimal models is limited
Solution Approach 1:
The patent combines pruning and quantization into a single joint operation that processes weights and output feature maps simultaneously. This merging of operations allows the system to optimize both compression aspects together, reducing the complexity of the optimization process while achieving better model selection and compression results.
2Quantity of substance
If separate pruning and quantization operations are applied, then some compression is achieved, but the computational burden increases
Solution Approach 1:
By merging pruning and quantization into a joint operation, the patent reduces the total computational burden. The simultaneous processing of weights and output feature maps through unified threshold functions and cost functions eliminates redundant computations that would occur if the operations were performed separately, thereby improving computational efficiency while achieving the same compression ratio.
3Ease of manufacture
If conventional separate operations are used for pruning and quantization, then the process is simpler to implement, but the allowable state-space of the network is shrunk considering only one pruned model
Solution Approach 1:
The joint operation merges pruning and quantization into a unified process that considers multiple models simultaneously through the cost function. This approach maintains implementation feasibility while significantly improving model selection flexibility, as the system can evaluate and select from a broader state-space of pruned and quantized models rather than being limited to a single sequentially processed model.
Data Source
AI summary
A system and a method generate a neural network that includes at least one layer having weights and output feature maps that have been jointly pruned and quantized. The weights of the layer are pruned using an analytic threshold function. Each weight remaining after pruning is quantized based on a weighted average of a quantization and dequantization of the weight for all quantization levels to form quantized weights for the layer. Output feature maps of the layer are generated based on the quantized weights of the layer. Each output feature map of the layer is quantized based on a weighted average of a quantization and dequantization of the output feature map for all quantization levels. Parameters of the analytic threshold function, the weighted average of all quantization levels of the weights and the weighted average of each output feature map of the layer are updated using a cost function.


