Data processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The high computational resource demand and inefficiency in matrix calculations due to the time-consuming dequantization process from low-precision to high-precision formats in large-scale model training and fine-tuning processes for neural networks.
Innovation Solution
A data processing device and method that converts quantized first matrix elements from a low-precision format to a high-precision format using a data transmission unit within a systolic array, performing element multiplication and accumulated sum calculations without dequantization, utilizing a floating-point multiplication unit in a k-dimensional accumulator to calculate the product of the accumulated sum multiplied by the quantization parameter.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If dequantization is performed to convert low-precision quantized data to high-precision format before matrix calculation, then calculation precision is improved, but computational resource consumption and processing time increase significantly
Solution Approach 1:
The patent applies preliminary action by performing format conversion from low-precision to high-precision format during the data transmission phase before the matrix calculation begins. This preliminary conversion ensures that the data is ready in the correct format for efficient calculation without requiring additional dequantization steps during the computation process, thereby resolving the contradiction between precision and efficiency
Solution Approach 2:
The patent introduces an intermediary format conversion mechanism that acts as a mediator between the low-precision quantized data storage and the high-precision calculation requirements. This intermediary conversion layer allows the system to maintain storage efficiency with low-precision formats while enabling high-precision calculations through the conversion interface, eliminating the need for resource-intensive dequantization operations
2Measurement precision
If dequantization operations are performed to process quantized data, then matrix calculation can be performed with high precision, but hardware resource consumption increases
Solution Approach 1:
The patent performs the format conversion action preliminarily during data transmission from external memory, before the data reaches the calculation unit. This preliminary conversion ensures that only the necessary conversion resources are consumed once, and the converted high-precision data can then be processed efficiently without requiring repeated dequantization operations, thereby reducing overall hardware resource consumption while maintaining calculation precision
3Measurement precision
If traditional dequantization process is used to convert low-precision format to high-precision format, then data can be processed with higher precision, but the process is time-consuming
Solution Approach 1:
The patent implements preliminary action by completing the format conversion during the data transmission phase from external memory to the processing unit. This timing optimization ensures that the conversion is performed in parallel with data fetch operations rather than sequentially, significantly reducing the total processing time while maintaining the ability to perform high-precision matrix calculations
Data Source
AI summary
A device, a chip, and a method for data processing is provided. The device includes: a data transmission unit for receiving a plurality of groups of quantized data, each of the groups of quantized data including a quantization parameter and k first matrix elements in a first low-precision format; and converting the format of the k first matrix elements into a first high-precision format; and a systolic array including a processing element, which further includes: an arithmetic logic unit for calculating an accumulated sum of the element products of k first matrix elements in the first high-precision format multiplied by corresponding k second matrix elements; and a k-dimensional accumulator configured to calculate, by a floating-point multiplication unit arranged therein, a plurality of products of the quantization parameter included in each group of quantized data and a corresponding accumulated sum, and perform accumulation to the products.


