基于多核加速卡的AI加速系统

By employing NOC connectivity and unidirectional pipelining communication in the multi-core hardware architecture of the large language model, the data transmission and computational division between cores are optimized, solving the problems of low utilization of computing resources and low efficiency of inter-core communication, and improving system performance and support for ultra-long text sequences.

CN121365040BActive Publication Date: 2026-07-17STORAGEX TECH INC +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
STORAGEX TECH INC
Filing Date
2025-11-03
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing large language models suffer from low utilization of multi-core hardware architectures, inefficient inter-core communication, difficulty in supporting computation of ultra-long text sequences, and limited overall system performance.

Method used

It adopts an architecture based on a network on-chip (NOC) to connect scalar computing cores, matrix computing cores and high-bandwidth memory, and optimizes data transmission and computing division of labor between cores through unidirectional pipelined cascaded communication and dynamic task scheduling of the computing management core.

Benefits of technology

It improves the utilization of computing resources, optimizes the efficiency of inter-core communication, enhances the support for ultra-long text sequences, reduces the consumption of wiring resources and the wiring difficulty of EDA tools, and improves the overall system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121365040B_ABST
    Figure CN121365040B_ABST
Patent Text Reader

Abstract

本申请公开基于多核加速卡的AI加速系统,涉及模型加速领域。包括挂载在NOC上的标量计算核、N个矩阵计算核、计算管理核、HBM,以及Host CPU;标量计算核通过NOC分别与N个矩阵计算核路由连接;N个矩阵计算核通过NOC首尾连接;计算管理核接收Host CPU的任务,将指令与数据发送到NOC上,通过路由链路送入目标计算核;矩阵计算核基于接收到的指令数据进行矩阵运算,以及向标量计算核输出中间数据;标量计算核基于接收的指令和中间数据进行标量计算和数据汇聚,返回加速结果。本方案通过NOC路由链路和流水级联通信机制,可提升计算资源利用率、优化核间通信效率、增强超长文本序列支持能力。
Need to check novelty before this filing date? Find Prior Art