Near Data Processing for Deep Neural Network Quantization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current deep neural network accelerators lack hardware support for efficient quantization, leading to inefficient processing and high resource consumption due to offline data quantization, which requires general processors for assistance.
Innovation Solution
A processing system and integrated circuit with a near data processing apparatus and acceleration apparatus that enables online dynamic quantization, reducing unnecessary data access and allowing high-precision parameter updates, thereby enhancing the accuracy and efficiency of neural network models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If offline quantization is used with general processors, then quantization can be performed, but processing efficiency is poor and resource consumption is high
Solution Approach 1:
The system separates quantization operations from general processing by dedicating specific hardware components (quantization units) within the memory controller to handle quantization tasks, while general processors focus on other operations. This segmentation resolves the contradiction by improving quantization efficiency without increasing overall resource consumption.
Solution Approach 2:
The patent introduces a memory controller with integrated quantization units as an intermediary between memory and general processors. This intermediary handles quantization operations locally, eliminating the need for general processors to assist with quantization, thereby improving efficiency and reducing computing resource occupation.
2Productivity
If data is quantized offline, then quantization can be achieved, but unnecessary data access occurs and processing efficiency decreases
Solution Approach 1:
The system performs quantization operations preliminarily within the memory controller before data is transferred to processing units. By preparing quantized data in advance during memory access operations, the system eliminates subsequent quantization delays and unnecessary data access, directly improving processing speed and reducing time loss.
Solution Approach 2:
The memory controller performs quantization operations on its own without requiring external processor assistance. This self-service capability allows the system to quantize data during normal memory operations, eliminating additional data access cycles and improving overall processing efficiency.
3Measurement precision
If high-precision floating-point numbers are used, then model accuracy is maintained, but memory bandwidth and storage requirements increase
Solution Approach 1:
The system dynamically changes the precision parameter of data representations by performing quantization operations that convert high-precision floating-point numbers to lower-precision formats. This parameter change reduces memory bandwidth and storage requirements while maintaining sufficient model accuracy through controlled quantization processes.
Solution Approach 2:
The patent applies different precision levels to different parts of the neural network processing pipeline. High-precision operations are performed only where necessary for model accuracy, while quantized lower-precision operations are used for other computations. This local quality approach optimizes the balance between accuracy and resource usage.
Data Source
AI summary
A device for optimizing parameters of a deep neural network is included in an integrated circuit apparatus. The integrated circuit apparatus includes a general interconnection interface and other processing apparatus. A computing apparatus interacts with other processing apparatus to jointly complete a computing operation specified by a user. The integrated circuit apparatus further includes a storage apparatus. The storage apparatus is connected to the computing apparatus and other processing apparatus, respectively. The storage apparatus is used for data storage of the computing apparatus and other processing apparatus.


