Neural Network Quantization Using Adaptive Bit Widths

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional neural network quantization methods use a fixed bit width for all data, leading to low precision and inefficiencies in data processing, storage, and operation due to the differences in data characteristics within the network.

Innovation Solution

A neural network quantization method that determines and applies specific quantization parameters to each piece of data based on its characteristics, allowing for local quantization and reducing storage space while maintaining precision and improving operation efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a same quantization scheme is adopted for the entire neural network, then the data processing complexity is reduced, but the precision of data operation deteriorates

Engineering Contradiction:
Improvedata processing complexityVSAvoidprecision of data operation
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent segments the neural network data into different types (first type and second type) based on their characteristics, and applies different quantization schemes to each type. This segmentation allows the system to maintain low complexity while achieving high precision by treating different data types appropriately.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements local quality by applying different quantization parameters to different data types within the neural network. First type data uses one quantization scheme while second type data uses another, allowing each data type to be processed with the optimal precision level for its characteristics.

Inventive Principle:
Principle #3Local quality

2Ease of manufacture

If a same quantization scheme is adopted for the entire neural network, then the implementation simplicity is improved, but the operation precision deteriorates

Engineering Contradiction:
Improveimplementation simplicityVSAvoidoperation precision
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent divides the neural network data into distinct segments (first type and second type) that can be processed with different quantization schemes. This segmentation maintains implementation simplicity through clear classification while achieving high operation precision through type-specific processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes quantization parameters based on data type characteristics. By adjusting quantization parameters according to whether data belongs to the first or second type, the system achieves both implementation simplicity and high operation precision.

Inventive Principle:
Principle #35Parameter changes

3Quantity of substance

If fixed bit width quantization is used, then the storage capacity is reduced, but the precision of operation data deteriorates

Engineering Contradiction:
Improvestorage capacityVSAvoidprecision of operation data
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent applies local quality by using different bit widths for different data types. First type data may use a smaller bit width while second type data uses a larger bit width, allowing the system to reduce overall storage capacity while maintaining the precision required for each specific data type's operations.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12112257B2Data processing method, device, computer equipment and storage medium
Publication Date: 2024.10.08 ANHUI CAMBRICON INFORMATION TECH CO LTD
  • US12112257B2 patent drawing
  • US12112257B2 patent drawing
  • US12112257B2 patent drawing

AI summary

The present disclosure provides a data processing method, a data processing device, a computer equipment, and a storage medium. The data processing device includes a board card and the board card provided in the present disclosure includes a storage component, an interface device, a control component, and an artificial intelligence chip of a data processing device. According to the data processing method, the data processing device, the computer equipment, and the storage medium provided in the embodiments of the present disclosure, data to be quantized is quantized according to a corresponding quantization parameter, which may reduce the storage space of data while ensuring the precision, as well as ensure the accuracy and reliability of the operation result and improve the operation efficiency.