Neural Network Quantization with Layer-Wise Accuracy Preservation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Modifications to neural network models, such as making them sparse and/or quantized, often result in reduced accuracy, necessitating improvements in performance.
Innovation Solution
A method involving processors that modify neural networks to become sparse and/or quantized by differently weighting indications of similarity between full-precision and quantized layers, focusing on layers most influential to accuracy through iterative processes using loss functions and feature distillation techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If neural network models are made sparse and/or quantized, then computational efficiency and resource usage are improved, but model accuracy deteriorates
Solution Approach 1:
The patent applies local quality by differentiating treatment across neural network layers. Instead of uniformly quantizing all layers, the system identifies and prioritizes influential layers (such as early convolutional layers in vision models) for quantization while preserving full precision in less critical layers. This selective approach maintains computational efficiency benefits while preserving model accuracy where it matters most.
Solution Approach 2:
The patent utilizes parameter changes by dynamically adjusting quantization parameters including bit-width, rounding modes, and quantization axes based on layer importance analysis. The system modifies these parameters iteratively during training to optimize the balance between computational efficiency and accuracy preservation for each specific layer.
2Productivity
If neural network models are made sparse and/or quantized, then processing resources are reduced, but model performance deteriorates
Solution Approach 1:
The patent implements preliminary action through iterative training processes where the model is trained multiple times with progressively applied quantization and sparsification. During these iterations, the system identifies which layers contribute most to performance and prioritizes their preservation, effectively preparing the model architecture in advance to withstand resource constraints while maintaining reliability.
Solution Approach 2:
The system employs feedback mechanisms where performance metrics from each training iteration are used to adjust subsequent quantization strategies. The feedback loop continuously monitors model performance and refines which layers should be quantized versus preserved, ensuring that resource reduction does not compromise critical performance aspects.
Data Source
AI summary
Apparatuses, systems, and techniques are to modify a precision and/or sparsity of one or more neural network layers. In at least one embodiment, a precision and/or sparsity of one or more neural network layers are based, at least in part on, a comparison of activations of a sparse version and a dense version of a neural network.


