In-Memory Bit Partitioning for Low-Latency AI Computation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current digital accelerators for artificial neural networks face high power consumption, high latency, and low transmission rates due to frequent data communication with memory, and CIM-based accelerators have high hardware complexity and limited design flexibility due to the need for high-dimensional analog-to-digital converters.
Innovation Solution
A memory system with a plurality of first memory units, read word lines, and read bit lines, where each memory unit includes transistors for controlling currents and voltages, allowing linear combinations of currents and voltages according to weightings, and using multiple low-dimensional analog-to-digital converters for bit partitioning and internal computation, reducing power consumption and latency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If digital accelerators frequently communicate with memory during data accessing, then data can be accessed, but power consumption increases, latency increases, and transmission rate decreases
Solution Approach 1:
The patent merges memory storage and computation functions into a single integrated structure. Memory cells store data while transistors perform computation operations directly on the stored data, eliminating the need for separate memory access and computation stages. This integration allows data to be processed in-place, simultaneously achieving fast data access and low power consumption by removing the energy-intensive data movement between memory and processor.
Solution Approach 2:
The memory cells are designed to serve multiple functions: storing data during computation operations and performing arithmetic operations on the stored data. The same memory structure that holds weights and inputs also executes the multiplication and accumulation operations, making the system universally capable of both memory access and computation without requiring separate dedicated components for each function.
2Measurement precision
If CIM-based accelerators use high-dimensional analog-to-digital converters, then computation accuracy improves, but hardware complexity increases and design flexibility decreases
Solution Approach 1:
The patent segments the computation process into distinct phases: analog computation phase where multiply-accumulate operations are performed using memory cells and transistors, followed by a conversion phase where only the final results are converted to digital format using simple analog-to-digital converters. This segmentation allows high-precision computation to be achieved through analog operations while avoiding the need for complex high-dimensional ADCs, as only low-dimensional conversion is required for the output results.
Solution Approach 2:
The patent replaces complex mechanical/digital conversion systems with analog computation mechanisms. Instead of using complex high-dimensional ADCs to achieve computation accuracy, the system uses analog electrical signals and transistor characteristics to perform computation operations directly, substituting the need for complex digital conversion hardware with simpler analog processing that achieves the same accuracy goal.
Data Source
AI summary
A memory system includes a plurality of first memory units, a plurality of read word lines, and a plurality of read bit lines. Each first memory unit of the plurality of first memory units includes a second memory unit, a first transistor coupled to the second memory unit, and a second transistor coupled to the second memory unit and the first transistor. Each read word line of the plurality of read word lines is coupled to a plurality of first transistors disposed along a corresponding row. Each read bit line of the plurality of read bit lines is coupled to a plurality of second transistors disposed along a corresponding column.


