Neural Network Quantization with Layer-Wise Accuracy Preservation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modifications to neural network models, such as making them sparse and/or quantized, often result in reduced accuracy, necessitating improvements in performance.

Innovation Solution

A method involving processors that modify neural networks to become sparse and/or quantized by differently weighting indications of similarity between full-precision and quantized layers, focusing on layers most influential to accuracy through iterative processes using loss functions and feature distillation techniques.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Use of energy by moving object

If neural network models are made sparse and/or quantized, then computational efficiency and resource usage are improved, but model accuracy deteriorates

Engineering Contradiction:
Improvecomputational efficiencyVSAvoidmodel accuracy
Core Design Contradiction:
Use of energy by moving objectVSMeasurement precision

Solution Approach 1:

The patent applies local quality by differentiating treatment across neural network layers. Instead of uniformly quantizing all layers, the system identifies and prioritizes influential layers (such as early convolutional layers in vision models) for quantization while preserving full precision in less critical layers. This selective approach maintains computational efficiency benefits while preserving model accuracy where it matters most.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent utilizes parameter changes by dynamically adjusting quantization parameters including bit-width, rounding modes, and quantization axes based on layer importance analysis. The system modifies these parameters iteratively during training to optimize the balance between computational efficiency and accuracy preservation for each specific layer.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If neural network models are made sparse and/or quantized, then processing resources are reduced, but model performance deteriorates

Engineering Contradiction:
Improveprocessing resourcesVSAvoidmodel performance
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent implements preliminary action through iterative training processes where the model is trained multiple times with progressively applied quantization and sparsification. During these iterations, the system identifies which layers contribute most to performance and prioritizes their preservation, effectively preparing the model architecture in advance to withstand resource constraints while maintaining reliability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system employs feedback mechanisms where performance metrics from each training iteration are used to adjust subsequent quantization strategies. The feedback loop continuously monitors model performance and refines which layers should be quantized versus preserved, ensuring that resource reduction does not compromise critical performance aspects.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250225395A1Neural network modification
Publication Date: 2025.07.10 NVIDIA CORP
  • US20250225395A1 patent drawing
  • US20250225395A1 patent drawing
  • US20250225395A1 patent drawing

AI summary

Apparatuses, systems, and techniques are to modify a precision and/or sparsity of one or more neural network layers. In at least one embodiment, a precision and/or sparsity of one or more neural network layers are based, at least in part on, a comparison of activations of a sparse version and a dense version of a neural network.