Neural Network Weight Quantization for Analog Computing-In-Memory
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Neural network models implemented on digital computing systems face challenges in being applied to edge devices due to high complexity and limited computing speed and power, and existing quantization methods for crossbar-enabled analog computing-in-memory systems do not fully consider the distribution characteristic of weights, leading to increased quantization errors.
Innovation Solution
A quantization method and apparatus that acquires the distribution characteristic of neural network weights to determine an initial quantization parameter, reducing quantization errors by dynamically adjusting the precision of weights based on their distribution, and optionally adding noise to improve robustness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If quantization is applied to reduce mapping overhead, then computing efficiency on edge devices is improved, but quantization error increases
Solution Approach 1:
The patent applies different quantization parameters to different weight elements based on their distribution characteristics. Instead of uniform quantization, the system analyzes the distribution of each weight and selects appropriate quantization parameters, thereby reducing quantization error in critical regions while maintaining efficiency gains elsewhere.
Solution Approach 2:
The patent dynamically adjusts quantization parameters based on the distribution characteristics of weight elements. By changing quantization parameters (such as quantization level and step size) according to the specific distribution properties of each weight, the system optimizes the balance between mapping overhead reduction and quantization error minimization.
2Loss of energy
If existing quantization methods are used, then mapping overhead is reduced, but quantization error increases due to not considering weight distribution characteristics
Solution Approach 1:
The system analyzes the distribution characteristics of each weight element and applies localized quantization strategies. This allows different parts of the weight matrix to be quantized differently according to their specific distribution properties, thereby reducing overall mapping overhead while minimizing quantization error in critical regions.
Solution Approach 2:
The patent incorporates feedback mechanisms where the quantization process considers the distribution characteristics of weights and adjusts parameters accordingly. This feedback loop ensures that quantization parameters are optimized based on actual weight distribution patterns, reducing both mapping overhead and quantization error simultaneously.
Data Source
AI summary
Disclosed are a quantization method and quantization apparatus for a weight of a neural network, and a storage medium. The neural network is implemented on the basis of a crossbar-enabled analog computing-in-memory (CACIM) system, and the quantization method includes: acquiring a distribution characteristic of a weight; and determining, according to the distribution characteristic of the weight, an initial quantization parameter for quantizing the weight to reduce a quantization error in quantizing the weight. The quantization method provided by the embodiments of the present disclosure does not pre-define the quantization method used, but determines the quantization parameter used for quantizing the weight according to the distribution characteristic of the weight to reduce the quantization error, so that the effect of the neural network model is better under the same mapping overhead, and the mapping overhead is smaller under the same effect of the neural network model.


