Computing chip and computer device
By integrating storage and computation modules into the computing chip, the problem of frequent memory access in the attention mechanism is solved, enabling efficient computation of Q, K, and V matrices and improving the processing efficiency of AI models.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- HUAWEI TECH CO LTD
- Filing Date
- 2025-10-13
- Publication Date
- 2026-07-23
AI Technical Summary
In the attention mechanism, the calculation of the Q matrix, K matrix, and V matrix involves frequent memory read and write operations, resulting in low processing efficiency of the AI model.
The system adopts a Partition In-Memory (PIM) architecture, which integrates the storage module and the computing module. The computing chip performs the multiplication of the Q matrix and the K matrix and the softmax operation, reducing memory access operations and improving computing efficiency.
By reducing memory access latency and lowering memory footprint, the computational efficiency of the Q, K, and V matrices in the attention mechanism is significantly improved, thereby enhancing the processing efficiency of the AI model.
Smart Images

Figure CN2025127402_23072026_PF_FP_ABST