数据处理的方法、计算单元、电子设备、存储介质和程序产品
By parsing the weight matrix into multiple data segments and determining the base value and scaling factor, efficient multiplication calculation between the weight matrix and the input matrix is achieved, solving the bottleneck problems of computational cost and memory bandwidth in deep learning models and improving data processing efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- VASTAI TECH (SHANGHAI) INC
- Filing Date
- 2025-12-16
- Publication Date
- 2026-07-17
AI Technical Summary
As the scale of deep learning model parameters expands, the computational cost and memory bandwidth of model training and inference become the main bottlenecks. Existing mixed-precision computing and quantization techniques suffer from precision loss, gradient instability, and separation of storage and computation in large-scale data scenarios, leading to increased latency and energy consumption.
The weight matrix is obtained through the storage module of the computing unit, parsed into multiple data segments by the parsing module, the base value and scaling factor are determined by the mapping module, and the multiplier and accumulator perform multiplication calculations to achieve efficient multiplication of the weight matrix and the input matrix.
It effectively improves the efficiency of data decoding and computation, reduces the structural complexity of computing units, reduces latency and energy consumption, and balances data storage and computation, making it suitable for large-scale model training and inference scenarios.
Smart Images

Figure CN121350397B_ABST