模型推理任务处理系统及模型推理任务处理方法
By employing a cascaded collaborative approach of a systolic array and a multiply-accumulate tree hardware accelerator in the model inference task processing system, the inefficiency caused by a single optimized design of the hardware accelerator is resolved, achieving a balance between the first token generation time and the single token generation time, thereby improving the model inference efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INSPUR SUZHOU INTELLIGENT TECH CO LTD
- Filing Date
- 2026-04-30
- Publication Date
- 2026-07-17
AI Technical Summary
Existing hardware accelerators are only optimized for one type of computational operation in GEMM or GEMV, which cannot balance the time for generating the first token and the time for generating a single token, resulting in low model inference efficiency.
Design a model inference task processing system, which includes a pre-filling stage hardware accelerator and a decoding stage hardware accelerator. The pre-filling stage hardware accelerator adopts a systolic array, and the decoding stage hardware accelerator adopts a multiply-accumulate tree. The two are of the same size and have compatible computation instructions. They work together in a cascaded manner to adapt to the computational characteristics of different stages.
It achieves a balance between the first token generation time and the single token generation time, improves the overall efficiency of model inference, and supports distributed cluster deployment with different hardware accelerators.
Smart Images

Figure CN122114197B_ABST