一种基于混合量化精度键值缓存的自注意力机制计算装置
By designing a self-attention mechanism computing device based on hybrid quantization precision and optimizing the storage and computation process of key-value cache, the problems of high computational complexity and low storage resource utilization in the self-attention mechanism are solved, and efficient computation and optimized utilization of storage resources are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SOUTHEAST UNIV
- Filing Date
- 2024-08-07
- Publication Date
- 2026-07-17
AI Technical Summary
The self-attention mechanism suffers from high computational complexity and low storage resource utilization, especially in large-scale data processing where computational efficiency is low. Furthermore, existing hybrid quantization precision techniques have failed to effectively address the time issues associated with key-value caching and nonlinear unit computation.
Design a self-attention mechanism computing device based on hybrid quantization precision, including an input data quantization module, a self-attention mechanism computing module, an nm dequantization operation module, a computation difference-load difference matching module, and a key-value cache module. By dynamically adjusting the quantization precision and optimizing the storage and computation process of the key-value cache, a balance between precision and computational performance is achieved.
It significantly reduces computational complexity, improves computational efficiency, reduces storage overhead, and lowers energy consumption. It is suitable for a variety of complex computing scenarios and achieves the best balance between inference accuracy and computational efficiency.
Smart Images

Figure CN119047527B_ABST