Neural Network Quantization Parameter Control for Low-Bit Inference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Neural networks require large storage space and processing bandwidth due to high-precision data representation, leading to increased costs and resource consumption.

Innovation Solution

A method to adjust data bit width during quantization, converting high-precision data to low-precision fixed-point data using quantization parameters, reducing storage space and improving calculation performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If high-precision data representation is used in neural networks, then measurement precision is improved, but storage space and processing bandwidth increase

Engineering Contradiction:
Improvedata precisionVSAvoidstorage space
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies parameter changes by dynamically adjusting the precision (bit width) of data representation during different stages of neural network operation. During training, higher precision is maintained for accuracy, while during inference, lower precision is used to reduce storage and computational overhead. This dynamic parameter adjustment resolves the contradiction between maintaining measurement precision and reducing storage space requirements.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements dynamics by making the data precision adaptive rather than static. The system automatically adjusts the bit width of data based on the operational context (training vs. inference) and performance requirements. This dynamic adaptation allows the neural network to optimize the trade-off between precision and resource consumption in real-time, resolving the technical contradiction.

Inventive Principle:
Principle #15Dynamics

2Quantity of substance

If data bit width is reduced during quantization, then storage space is reduced, but measurement precision deteriorates

Engineering Contradiction:
Improvestorage spaceVSAvoiddata precision
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent uses dynamic precision scaling where the bit width is adjusted based on the operational phase and performance requirements. During training, sufficient precision is maintained to ensure accurate gradient computation, while during inference, precision is reduced to optimize storage and speed. This dynamic approach resolves the contradiction by adapting precision to actual needs rather than using a fixed low precision throughout.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies local quality by allowing different parts of the neural network or different operational phases to use different precision levels. Critical operations during training maintain higher precision, while less critical inference operations use lower precision. This localized differentiation of precision quality resolves the contradiction between storage reduction and precision maintenance.

Inventive Principle:
Principle #3Local quality

3Quantity of substance

If dynamic precision scaling is applied during training, then storage space is reduced, but training stability may be affected

Engineering Contradiction:
Improvestorage spaceVSAvoidtraining stability
Core Design Contradiction:
Quantity of substanceVSStability of the object's composition

Solution Approach 1:

The patent implements dynamic precision adjustment during training where the bit width is adaptively changed based on the training progress and performance metrics. The system monitors training stability and adjusts precision levels accordingly, maintaining higher precision when stability is compromised and reducing precision when it is sufficient. This dynamic control resolves the contradiction between storage reduction and training stability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent uses feedback mechanisms to monitor training stability and adjust precision levels in response. The system continuously evaluates the impact of precision reduction on training convergence and stability, and dynamically adjusts the precision settings to maintain optimal performance. This feedback-driven approach resolves the contradiction by using performance information to guide precision adjustments.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP3998554B1Method for determining quantization parameter of neural network, and related product
Publication Date: 2025.11.26 SHANGHAI CAMBRICON INFORMATION TECH CO LTD
  • EP3998554B1 patent drawingFigure 1~2
  • EP3998554B1 patent drawingFigure 3~4
  • EP3998554B1 patent drawingFigure 5A~5B

AI summary

The technical solution involves a board card including a storage component, an interface apparatus, a control component, and an artificial intelligence chip. The artificial intelligence chip is connected to the storage component, the control component, and the interface apparatus, respectively; the storage component is used to store data; the interface apparatus is used to implement data transfer between the artificial intelligence chip and an external device; and the control component is used to monitor a state of the artificial intelligence chip. The board card is used to perform an artificial intelligence operation.