In-Memory Bit Partitioning for Low-Latency AI Computation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current digital accelerators for artificial neural networks face high power consumption, high latency, and low transmission rates due to frequent data communication with memory, and CIM-based accelerators have high hardware complexity and limited design flexibility due to the need for high-dimensional analog-to-digital converters.

Innovation Solution

A memory system with a plurality of first memory units, read word lines, and read bit lines, where each memory unit includes transistors for controlling currents and voltages, allowing linear combinations of currents and voltages according to weightings, and using multiple low-dimensional analog-to-digital converters for bit partitioning and internal computation, reducing power consumption and latency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If digital accelerators frequently communicate with memory during data accessing, then data can be accessed, but power consumption increases, latency increases, and transmission rate decreases

Engineering Contradiction:
Improvedata access speedVSAvoidpower consumption
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The patent merges memory storage and computation functions into a single integrated structure. Memory cells store data while transistors perform computation operations directly on the stored data, eliminating the need for separate memory access and computation stages. This integration allows data to be processed in-place, simultaneously achieving fast data access and low power consumption by removing the energy-intensive data movement between memory and processor.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The memory cells are designed to serve multiple functions: storing data during computation operations and performing arithmetic operations on the stored data. The same memory structure that holds weights and inputs also executes the multiplication and accumulation operations, making the system universally capable of both memory access and computation without requiring separate dedicated components for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If CIM-based accelerators use high-dimensional analog-to-digital converters, then computation accuracy improves, but hardware complexity increases and design flexibility decreases

Engineering Contradiction:
Improvecomputation accuracyVSAvoidhardware complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the computation process into distinct phases: analog computation phase where multiply-accumulate operations are performed using memory cells and transistors, followed by a conversion phase where only the final results are converted to digital format using simple analog-to-digital converters. This segmentation allows high-precision computation to be achieved through analog operations while avoiding the need for complex high-dimensional ADCs, as only low-dimensional conversion is required for the output results.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces complex mechanical/digital conversion systems with analog computation mechanisms. Instead of using complex high-dimensional ADCs to achieve computation accuracy, the system uses analog electrical signals and transistor characteristics to perform computation operations directly, substituting the need for complex digital conversion hardware with simpler analog processing that achieves the same accuracy goal.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20220262426A1Memory System Capable of Performing a Bit Partitioning Process and an Internal Computation Process
Publication Date: 2022.08.18 NAT CHENG KUNG UNIV
  • US20220262426A1 patent drawing
  • US20220262426A1 patent drawing
  • US20220262426A1 patent drawing

AI summary

A memory system includes a plurality of first memory units, a plurality of read word lines, and a plurality of read bit lines. Each first memory unit of the plurality of first memory units includes a second memory unit, a first transistor coupled to the second memory unit, and a second transistor coupled to the second memory unit and the first transistor. Each read word line of the plurality of read word lines is coupled to a plurality of first transistors disposed along a corresponding row. Each read bit line of the plurality of read bit lines is coupled to a plurality of second transistors disposed along a corresponding column.