Systems and methods for accelerating neural network convolution and training
By minimizing the connection distance between processing elements and memory in an ASIC and integrating processing blocks and memory using a 3D-IC architecture, efficient training and inference of neural networks are achieved, solving the problem of low memory access efficiency and improving computing performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- RAMBUS INC
- Filing Date
- 2020-11-24
- Publication Date
- 2026-07-07
AI Technical Summary
In existing technologies, the low memory access efficiency during the training and inference processes of neural networks leads to limitations in processing speed and performance.
It adopts an application-specific integrated circuit (ASIC) design, supports concurrent forward and backward propagation by minimizing the connection distance between processing elements and memory, utilizes hierarchical buffers and ring buses for data communication, and integrates processing blocks and memory through a 3D-IC architecture to achieve efficient data transmission and computation.
It improves the training and inference efficiency of neural networks, reduces idle time, and enhances hardware performance and computing speed.
Smart Images

Figure CN114902242B_ABST
Abstract
Citation Information
Patent Citations
Neural network, encoder, decoder, learning method, control method, and program
JP2019016166A
Deep neural network processing on hardware accelerators with stacked memory
US20160379115A1
System and method for speeding up general matrix-vector multiplication on GPU
US20170372447A1
System and method for an optimized winograd convolution accelerator
US20190042923A1
System and method for implementing neural networks in integrated circuits
US20190080223A1