Attention head configuration method for large model, electronic device, and storage medium

By identifying and classifying attention heads in large models and dynamically scheduling computation patterns, the problems of high computational complexity and insufficient flexibility of traditional large models and sparse attention mechanisms are solved, thus achieving efficient and flexible attention computation.

CN122433798APending Publication Date: 2026-07-21UC MOBILE CHINA CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610249085.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-02
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Traditional full attention mechanisms for large models have high computational complexity, while sparse attention mechanisms lack flexibility after training and have high engineering implementation costs, making them difficult to apply efficiently in already trained models.

Method used

By identifying the attention heads of large models, classifying them into different types based on performance representation parameters, and adding type identifiers, we can dynamically call full or sparse computations. Combined with memory rearrangement and processor load balancing optimizations, we can achieve on-demand computation.

Benefits of technology

Without retraining the model, it improves computational efficiency and resource utilization, balances inference accuracy and efficiency, and reduces engineering costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122433798A_ABST
    Figure CN122433798A_ABST
Patent Text Reader

Abstract

The application provides an attention head configuration method for a large model, an electronic device and a storage medium. According to an embodiment of the application, the attention heads of each network layer in the large model are identified, and the plurality of attention heads are divided into a plurality of different types of attention heads according to the performance representation parameters of the attention heads, including a first type of attention head configured to perform full-amount attention calculation and a second type of attention head configured to perform sparse attention calculation. By adding corresponding type identifiers to the different types of attention heads, when the large model performs attention calculation based on an input text sequence, the different types of attention heads are distinguished and called according to the type identifiers, so that the attention heads perform different attention mechanisms, and on-demand calculation is realized at the granularity of the attention heads, taking into account the inference effect and inference efficiency of the large model. The full-amount attention calculation can retain complete context information and maintain the inference accuracy of the model, and the sparse attention calculation can reduce the calculation overhead and improve the inference efficiency.
Need to check novelty before this filing date? Find Prior Art