Neural Network Layer Sensitivity for Mixed-Precision Quantization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Selecting a reasonable configuration for mixed precision quantization in neural network models to balance precision and processing speed is time-consuming.
Innovation Solution
A method that involves performing quantization on neural network layers, calculating layer sensitivity, and iteratively updating and selecting neural network models based on layer sensitivity and noise change to determine an optimal configuration.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If mixed precision quantization is applied to neural network layers, then processing speed is improved and model size is compressed, but selecting a reasonable configuration becomes time-consuming
Solution Approach 1:
The patent applies preliminary action by pre-calculating layer sensitivity values for all layers before performing quantization. This sensitivity information is computed in advance using a sensitivity calculation module that analyzes the impact of quantization on each layer's output. By having this sensitivity data ready beforehand, the system can quickly determine the optimal quantization configuration without time-consuming trial and error during the actual quantization process, thus resolving the contradiction between achieving fast processing speed and avoiding time-consuming configuration selection.
2Quantity of substance
If mixed precision quantization is applied to neural network layers, then model size is compressed, but configuration selection becomes time-consuming
Solution Approach 1:
The patent uses preliminary action by pre-computing sensitivity values for all layers before quantization. The sensitivity calculation module analyzes each layer's contribution to the final output and determines how much quantization error each layer can tolerate. This pre-computed sensitivity information enables rapid configuration selection that achieves model compression without requiring extensive trial and error, thus resolving the contradiction between reducing model size and minimizing configuration selection time.
Solution Approach 2:
The patent applies parameter changes by dynamically adjusting the quantization precision parameter for each layer based on its sensitivity value. Layers with low sensitivity are assigned lower precision (fewer bits) to maximize compression, while layers with high sensitivity maintain higher precision to preserve accuracy. This parameter adaptation allows the system to achieve significant model size reduction while the pre-computed sensitivity values ensure the configuration is determined quickly without time-consuming searches.
3Speed
If quantization is performed on multiple layers, then processing speed is improved, but model accuracy may deteriorate
Solution Approach 1:
The patent applies local quality by assigning different quantization precision levels to different layers based on their individual sensitivity characteristics. Instead of uniformly quantizing all layers to the same precision, the system identifies which layers are more sensitive to quantization errors and preserves higher precision for those layers while applying aggressive quantization to less sensitive layers. This localized differentiation maintains model accuracy for critical layers while achieving speed improvements through quantization of non-critical layers.
Solution Approach 2:
The patent uses parameter changes by dynamically adjusting the precision parameter for each layer according to its sensitivity value. The quantization configuration is adapted layer-by-layer, with each layer receiving a precision level optimized for its specific role in the network. This parameter adaptation ensures that layers contributing most to accuracy maintain high precision while other layers are quantized to achieve faster processing, thus resolving the contradiction between speed improvement and accuracy preservation.
Data Source
AI summary
Disclosed are a quantization method of a neural network model and an electronic device. The quantization method comprises: performing quantization on layers in a first neural network model, to obtain a second neural network model; calculating, using the second neural network model as a current neural network model, a layer sensitivity of each layer, on which the quantization has been performed, in the current neural network model; updating the current neural network model by canceling the quantization of a layer, having a highest layer sensitivity, to obtain a third neural network model; calculating, using the third neural network model as the current neural network model, a layer sensitivity of each layer, in the current neural network model; repeating the updating and the calculating until a predetermined number of third neural network models are obtained; and selecting an optimal neural network model.


