Neural Network Quantization via Layer-Specific Parameter Adjustment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing complexity of neural networks leads to challenges in data processing efficiency, storage capacity, and access efficiency due to the large amount of data and increasing data dimensions, with traditional quantization methods resulting in low precision and operation inefficiency.

Innovation Solution

A method for neural network quantization that determines specific data to be quantized based on storage capacity and uses a loop processing method with corresponding quantization parameters to achieve efficient data compression while maintaining precision, involving the quantization of input neurons and gradients in batches.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If the same quantization scheme is adopted for the entire neural network, then data processing efficiency is improved, but measurement precision deteriorates

Engineering Contradiction:
Improvedata processing efficiencyVSAvoidquantization precision
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies different quantization schemes to different parts of the neural network based on their specific characteristics. Each layer or operation is assigned a quantization scheme tailored to its data distribution and computational requirements, rather than using a uniform quantization approach across the entire network. This local optimization resolves the contradiction by maintaining high precision where needed while achieving overall efficiency improvement.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The neural network is segmented into multiple layers or operational units, each processed with its own quantization parameters. The patent divides the quantization process into discrete steps where different quantization schemes can be applied to different segments (layers) of the network, allowing simultaneous optimization of precision and efficiency across the entire system.

Inventive Principle:
Principle #1Segmentation

2Quantity of substance

If floating point data is converted to fixed point data for quantization, then storage capacity is reduced, but manufacturing precision deteriorates

Engineering Contradiction:
Improvestorage capacityVSAvoiddata precision
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent dynamically adjusts quantization parameters (such as bit width, scaling factors, and offset values) based on the specific requirements of each layer and operation. By changing these parameters adaptively rather than using fixed conversion rules, the system achieves efficient storage reduction while maintaining the necessary precision for accurate computations.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The quantization scheme transitions from a static fixed-point conversion to a dynamic process where quantization parameters are adjusted based on runtime conditions, data characteristics, and layer-specific requirements. This dynamic approach allows the system to optimize the balance between storage efficiency and computational precision adaptively.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12165039B2Neural network quantization data processing method, device, computer equipment and storage medium
Publication Date: 2024.12.10 ANHUI CAMBRICON INFORMATION TECH CO LTD
  • US12165039B2 patent drawing
  • US12165039B2 patent drawing
  • US12165039B2 patent drawing

AI summary

The present disclosure provides a data processing method, a board card device, a computer equipment, and a storage medium for data quantization. The board card provided in the present disclosure includes a storage component, an interface device, a control component, and an artificial intelligence chip of a data processing device, where the artificial intelligence chip is connected to the storage device, the control device, and the interface apparatus, respectively. The storage component is configured to store data; the interface device is configured to implement data transmission between the artificial intelligence chip and an external equipment; and the control component is configured to monitor a state of the artificial intelligence chip. According to the data processing method, the device, the computer equipment, and the storage medium provided in the embodiments of the present disclosure, data to be quantized is quantized according to a corresponding quantization parameter, which may reduce the storage space of data while ensuring the precision, as well as ensure the accuracy and reliability of the operation result and improve the operation efficiency.