Attention head configuration method for large model, electronic device, and storage medium
By identifying and classifying attention heads in large models and dynamically scheduling computation patterns, the problems of high computational complexity and insufficient flexibility of traditional large models and sparse attention mechanisms are solved, thus achieving efficient and flexible attention computation.
Patent Information
- Application Number
- CN202610249085.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-02
- Publication Date
- 2026-07-21
AI Technical Summary
Traditional full attention mechanisms for large models have high computational complexity, while sparse attention mechanisms lack flexibility after training and have high engineering implementation costs, making them difficult to apply efficiently in already trained models.
By identifying the attention heads of large models, classifying them into different types based on performance representation parameters, and adding type identifiers, we can dynamically call full or sparse computations. Combined with memory rearrangement and processor load balancing optimizations, we can achieve on-demand computation.
Without retraining the model, it improves computational efficiency and resource utilization, balances inference accuracy and efficiency, and reduces engineering costs.
Smart Images

Figure CN122433798A_ABST