Triple-Mode DRAM Cell for Dynamic AI Accelerator Core
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional PIM-based processors face challenges in reducing power consumption of the input/output feature map memory, due to their static core architecture, which leads to decreased energy efficiency when computing layers of varying sizes, and are limited by the characteristics of static RAM, including high transistor count and leakage current issues.
Innovation Solution
A DRAM using a triple-mode memory cell and an AI accelerator with a dynamic core structure that allows for varying sizes of internal memory and calculator according to the structure and layer of a deep neural network, enabling improved integration, area efficiency, and energy efficiency by supporting computation, memory, and data conversion modes within a single cell, and reconfiguring dataflow based on neural network requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If static core architecture is used in PIM-based processors, then the structure is simple and easy to manufacture, but the energy efficiency decreases when computing layers of varying sizes due to fixed memory and calculator sizes
Solution Approach 1:
The patent implements a dynamic core architecture where the calculator size and memory configuration can be flexibly adjusted according to the specific layer being computed. The core includes reconfigurable components that allow the system to adapt its structure dynamically, enabling optimal energy efficiency for different computation scenarios while maintaining manufacturing feasibility through standardized cell designs.
Solution Approach 2:
The system changes operational parameters by allowing the calculator to have variable sizes (e.g., different numbers of MAC units) and memory to have configurable capacities. This parameter flexibility enables the processor to match its resources to the actual computation requirements of each layer, avoiding the energy waste associated with fixed-size architectures that must accommodate the largest possible layer.
2Speed
If SRAM-based PIM is used, then computation speed is fast, but the transistor count per cell is high (6-18 transistors) reducing the degree of integration
Solution Approach 1:
The patent merges the memory storage function with the computation function into a single integrated cell structure. By combining the SRAM cell with MAC (multiply-accumulate) computation units and analog-to-digital converters within the same memory cell footprint, the design achieves high computation speed while reducing the overall transistor count per functional unit compared to separate memory and processor architectures.
Solution Approach 2:
The memory cell is designed to perform multiple functions: data storage, analog computation (multiplication and accumulation), and analog-to-digital conversion. This multi-functionality eliminates the need for separate dedicated computation units and converters, thereby reducing the total transistor count while maintaining fast computation capabilities.
3Reliability
If DRAM-based PIM with large capacitor is used, then leakage current effect is reduced, but the degree of integration in memory is lowered
Solution Approach 1:
The system dynamically adjusts the capacitor size parameter based on the operational requirements. For computations where leakage current is a concern, larger capacitors are used to maintain reliability. For applications where integration density is prioritized and computation speed is sufficient, smaller capacitors are employed. This parameter flexibility allows the system to optimize between reliability and integration density depending on the specific use case.
Solution Approach 2:
The memory architecture employs dynamic capacitor sizing where the capacitor value can be changed or selected based on the computational task at hand. This dynamic adaptation allows the system to use larger capacitors only when necessary for leakage-resistant operation, while maintaining high integration density for tasks where smaller capacitors suffice, thereby resolving the contradiction between reliability and integration density.
4Measurement precision
If analog operation parallelism is increased to compensate for leakage current, then high-accuracy operation is maintained, but computational efficiency and area efficiency decrease
Solution Approach 1:
The patent introduces an intermediate analog-to-digital conversion stage within the memory cell that allows analog computation to proceed with high parallelism while converting results to digital form for accurate accumulation. This intermediary conversion mechanism enables the system to maintain computational accuracy without being constrained by leakage current, thereby achieving both high accuracy and high computational efficiency simultaneously.
Solution Approach 2:
The system changes the operational mode parameter, allowing flexible switching between fully analog operation and hybrid analog-digital operation. By adjusting this parameter, the system can optimize for either accuracy or efficiency depending on the computational requirements, resolving the contradiction between computation accuracy and computational efficiency.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The proposed solution improves the degree of integration and area efficiency in memory, enhances the utilization rate of the calculator, and reduces energy consumption by up to 31% compared to fixed core architectures, while also reducing area consumption by the analog-to-digital converter and minimizing operation delays due to refresh synchronization.
Implementation Method 1
a dynamic random access memory (DRAM) using a triple-mode memory cell
Data Source
AI summary
A DRAM is configured using a triple-mode memory cell that supports a computation mode, a memory mode, and a data conversion mode by one cell and converts modes as necessary, and an AI accelerator using the same is provided, so that a dataflow may be reconfigured according to a structure and a size of an AI neural network (so-called deep neural network) to be trained.


