模型推理任务处理系统及模型推理任务处理方法

By employing a cascaded collaborative approach of a systolic array and a multiply-accumulate tree hardware accelerator in the model inference task processing system, the inefficiency caused by a single optimized design of the hardware accelerator is resolved, achieving a balance between the first token generation time and the single token generation time, thereby improving the model inference efficiency.

CN122114197BActive Publication Date: 2026-07-17INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INSPUR SUZHOU INTELLIGENT TECH CO LTD
Filing Date
2026-04-30
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing hardware accelerators are only optimized for one type of computational operation in GEMM or GEMV, which cannot balance the time for generating the first token and the time for generating a single token, resulting in low model inference efficiency.

Method used

Design a model inference task processing system, which includes a pre-filling stage hardware accelerator and a decoding stage hardware accelerator. The pre-filling stage hardware accelerator adopts a systolic array, and the decoding stage hardware accelerator adopts a multiply-accumulate tree. The two are of the same size and have compatible computation instructions. They work together in a cascaded manner to adapt to the computational characteristics of different stages.

Benefits of technology

It achieves a balance between the first token generation time and the single token generation time, improves the overall efficiency of model inference, and supports distributed cluster deployment with different hardware accelerators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122114197B_ABST
    Figure CN122114197B_ABST
Patent Text Reader

Abstract

本申请公开了一种模型推理任务处理系统及模型推理任务处理方法,涉及模型推理技术领域,计算节点包括多个预填充阶段硬件加速器和多个解码阶段硬件加速器;预填充阶段硬件加速器包括脉动阵列,脉动阵列对模型推理过程中的矩阵乘矩阵计算任务进行处理;解码阶段硬件加速器包括多个乘加树,乘加树对模型推理过程中的矩阵乘向量计算任务进行处理;脉动阵列的规模与多个乘加树的总规模相同;预填充阶段硬件加速器和解码阶段硬件加速器的计算指令兼容;计算节点用于基于模型推理任务确定目标预填充阶段硬件加速器和目标解码阶段硬件加速器,以处理模型推理任务。达到了均衡TTFT和TPOT,提高模型推理效率的技术效果。
Need to check novelty before this filing date? Find Prior Art