一种基于Systolic阵列的计算结构可配置卷积神经网络加速系统

By designing a computational structure based on a Systolic array and configurable units, and combining deep parallel convolution and Winograd convolution modes, the flexibility and efficiency issues of existing hardware accelerators when dealing with different CNN models are solved, achieving low-power, high-performance acceleration of convolutional neural networks.

CN121189398BActive Publication Date: 2026-07-17HUAZHONG UNIV OF SCI & TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAZHONG UNIV OF SCI & TECH
Filing Date
2025-08-29
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing convolutional neural network hardware accelerators lack flexibility when dealing with CNN models with different layer structures, and insufficient algorithm and hardware optimization leads to limitations in power consumption and real-time performance of the acceleration system.

Method used

It adopts a computational structure design based on a Systolic array, combined with configurable units and data flow scheduling, to support flexible configuration of different convolutional layers. It is accelerated by deep parallel convolution and Winograd convolution mode, and the computation process is optimized by using a double buffer mechanism and a control signal pulsation mechanism.

Benefits of technology

It achieves high computational performance acceleration of convolutional neural networks with low power consumption, supports the applicability of various CNN models, improves computational efficiency and resource utilization, simplifies the design process, and reduces instruction processing latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121189398B_ABST
    Figure CN121189398B_ABST
Patent Text Reader

Abstract

本发明属于半导体集成电路技术以及深度学习硬件加速技术领域,涉及一种基于Systolic阵列的计算结构可配置卷积神经网络加速系统,该系统包括控制处理器和协处理器,其中,控制处理器与协处理器之间通过高速扩展接口与低速信号接口连接;控制处理器用于通过高速扩展接口向协处理器发送自定义计算结构配置数据与图像数据;控制处理器还用于通过低速信号接口向协处理器发送层指示信号、卷积模式信号和计算模式信号。本发明可以同时完成多个基本计算路径的并行计算,经过计算结构配置后,可以完成卷积计算过程中不同输入通道维度与不同输出通道维度下的计算任务,有着运算吞吐量达、能耗比低的特点,十分适合卷积神经网络的高算力需求。
Need to check novelty before this filing date? Find Prior Art