Hierarchical Model Quantization for Low-Bit Performance Loss

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing model quantization methods face challenges in reducing performance loss when converting floating-point parameters to fixed-point parameters, especially at lower bits, due to varying sensitivity across neural network layers and device compatibility issues.

Innovation Solution

A hierarchical quantization method is employed, where model layers are sorted by performance deviations and quantized in descending order, with parameter adjustments using gradient derivation to minimize performance loss.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If model quantization is performed to reduce storage consumption and computation amount, then storage efficiency and device compatibility are improved, but performance loss increases especially at lower bits

Engineering Contradiction:
Improvestorage consumptionVSAvoidperformance loss
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent applies different quantization bit depths to different model layers based on their sensitivity characteristics. Important layers use higher bit depths (e.g., 16-bit or 32-bit) to maintain performance, while less sensitive layers use lower bit depths (e.g., 8-bit or 4-bit) to reduce storage. This local differentiation resolves the contradiction by optimizing storage efficiency without uniformly sacrificing performance across all layers.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent dynamically adjusts quantization parameters (bit depth, precision levels) based on layer-specific sensitivity analysis. By changing the quantization parameter settings according to each layer's importance and sensitivity to quantization error, the system achieves better performance preservation at lower overall storage consumption compared to uniform quantization approaches.

Inventive Principle:
Principle #35Parameter changes

2Adaptability or versatility

If uniform quantization is applied to all model layers, then device compatibility is improved, but performance loss varies significantly across layers with some layers being more sensitive than others

Engineering Contradiction:
Improvedevice compatibilityVSAvoidquantization performance consistency
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent implements layer-specific quantization strategies where each model layer is analyzed for its sensitivity to quantization and assigned appropriate quantization parameters. This local quality approach ensures that critical layers maintain high precision while less sensitive layers can use lower precision, achieving consistent performance across devices without uniform quantization penalties.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the model into different layers and applies different quantization schemes to each segment. By dividing the model and treating each layer individually based on its characteristics, the system achieves both device compatibility (through quantization) and performance consistency (through tailored quantization strategies for each segment).

Inventive Principle:
Principle #1Segmentation

3Use of energy by moving object

If low-bit quantization is used to reduce computation amount, then resource consumption is reduced, but decoding performance deteriorates due to sensitivity variations across layers

Engineering Contradiction:
Improvecomputation amountVSAvoiddecoding performance
Core Design Contradiction:
Use of energy by moving objectVSReliability

Solution Approach 1:

The patent applies different quantization bit depths to different model layers based on their sensitivity characteristics. Important layers use higher bit depths (e.g., 16-bit or 32-bit) to maintain performance, while less sensitive layers use lower bit depths (e.g., 8-bit or 4-bit) to reduce storage. This local differentiation resolves the contradiction by optimizing storage efficiency without uniformly sacrificing performance across all layers.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent dynamically adjusts quantization parameters (bit depth, precision levels) based on layer-specific sensitivity analysis. By changing the quantization parameter settings according to each layer's importance and sensitivity to quantization error, the system achieves better performance preservation at lower overall storage consumption compared to uniform quantization approaches.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4730205A1Model quantization method and apparatus, and device and storage medium
Publication Date: 2026.04.22 HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
  • EP4730205A1 patent drawingFigure 1~2
  • EP4730205A1 patent drawingFigure 3
  • EP4730205A1 patent drawingFigure 4~5

AI summary

A model quantization method and apparatus, and a device and a storage medium, which belong to the technical field of model quantization. The method comprises: acquiring a performance deviation corresponding to each model layer in a model to be quantized; constructing a hierarchical quantization sequence according to the performance deviations corresponding to the respective model layers; and performing hierarchical quantization processing on the model layers in said model on the basis of the hierarchical quantization sequence. A hierarchical quantization sequence is constructed according to performance deviations corresponding to respective model layers in a model to be quantized, hierarchical quantization processing is respectively performed on the model layers in said model according to the sequence order of the hierarchical quantization sequence, and after quantization during the processing parameter adjustment and correction are used to reduce a performance loss of the model to the greatest possible extent, which performance loss is generated by low-bit quantization.