Computing chip and computer device

By integrating storage and computation modules into the computing chip, the problem of frequent memory access in the attention mechanism is solved, enabling efficient computation of Q, K, and V matrices and improving the processing efficiency of AI models.

WO2026152793A1PCT designated stage Publication Date: 2026-07-23HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-10-13
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

In the attention mechanism, the calculation of the Q matrix, K matrix, and V matrix involves frequent memory read and write operations, resulting in low processing efficiency of the AI ​​model.

Method used

The system adopts a Partition In-Memory (PIM) architecture, which integrates the storage module and the computing module. The computing chip performs the multiplication of the Q matrix and the K matrix and the softmax operation, reducing memory access operations and improving computing efficiency.

Benefits of technology

By reducing memory access latency and lowering memory footprint, the computational efficiency of the Q, K, and V matrices in the attention mechanism is significantly improved, thereby enhancing the processing efficiency of the AI ​​model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025127402_23072026_PF_FP_ABST
    Figure CN2025127402_23072026_PF_FP_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of artificial intelligence (AI). Disclosed are a computing chip and a computer device. The computing chip comprises a memory module and an operation module, wherein the memory module is used for storing a K matrix and a V matrix, which are obtained from an inference request of an AI model, and the operation module comprises a first operation unit and a second operation unit; the first operation unit is used for performing a matrix vector multiplication operation on the K matrix and Q vectors included in a Q matrix that is obtained from the inference request, so as to obtain a first operation result; the second operation unit is used for performing a softmax operation on the first operation result, so as to obtain a second operation result; and the first operation unit is also used for performing a matrix vector multiplication operation on the second operation result and the V matrix, so as to obtain a third operation result, and providing the third operation result to a processor. By means of the present application, the operation efficiency of a Q matrix, a K matrix and a V matrix in an attention mechanism can be improved, thereby improving the processing efficiency of an AI model.
Need to check novelty before this filing date? Find Prior Art