Current-Mode SRAM Compute-in-Memory for Low-Power Edge AI
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current digital AI & ML solutions are power-hungry and costly, making them unsuitable for edge devices and sensors, where low-power, low-cost, and asynchronous signal processing is required for applications like robotics, medical, and surveillance, and they rely on expensive deep sub-micron manufacturing.
Innovation Solution
Implementing Compute-In-Memory (CIM) or Compute-Near-Memory (CNM) architectures with current-mode data-converters, multipliers, and multiply-accumulate circuits that integrate with digital systems, utilizing analog and mixed-signal processing to reduce power consumption and cost, and operate with minimal digital circuitry, enabling low-power, low-cost, and asynchronous AI & ML signal processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If standard digital solutions are used for AI & ML applications, then ease of interface compatibility and programming flexibility are achieved, but power consumption and manufacturing cost increase significantly
Solution Approach 1:
The patent combines memory and computation functions into a single integrated structure where crossbar arrays perform multiply-accumulate operations directly on stored data. This merging eliminates the need for separate digital processors, reducing power consumption while maintaining computational capability for AI & ML applications.
Solution Approach 2:
The patent replaces traditional digital mechanical switching and processing with analog current-mode computation. Current signals naturally perform multiplication and accumulation operations as they flow through the crossbar array, substituting complex digital logic with simpler analog physics-based computation that consumes less power.
2Productivity
If bleeding edge deep sub-micron manufacturing is used for digital AI & ML chips, then computational performance is improved, but manufacturing cost and power consumption increase
Solution Approach 1:
The patent changes the operational parameters from digital voltage switching to analog current flow. This parameter change allows computation to be performed using standard CMOS fabrication processes at larger feature sizes, eliminating the need for expensive deep sub-micron manufacturing while maintaining computational performance through the inherent parallelism of the crossbar architecture.
3Adaptability or versatility
If digital processors are used for AI & ML tasks, then programming flexibility is maintained, but latency increases due to memory read/write cycles
Solution Approach 1:
By merging memory storage and computation operations into the same physical location within the crossbar array, the patent eliminates the memory wall problem. Data remains stored in the crossbar memory during computation, allowing instant access and eliminating latency associated with traditional memory read/write cycles while maintaining programming flexibility through configurable synaptic weights.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach reduces dynamic power consumption, eliminates latency, and enhances privacy by performing AI & ML tasks at the edge or on sensors, using mainstream CMOS fabrication and low-voltage power supplies, while maintaining high accuracy and small size, thus overcoming the limitations of conventional digital processors.
Implementation Method 1
utilizing analog and mixed-signal processing to reduce power consumption and cost
Implementation Method 2
Current-mode mixed-signal SRAM based compute-in-memory
Data Source
AI summary
Multipliers and Multiply-Accumulate (MAC) circuits are fundamental building blocks in signal processing, including in emerging applications such as machine learning (ML) and artificial intelligence (AI) that predominantly utilize digital-mode multipliers, and MACs. Typically, digital multipliers and MACs can operate at high speed with high resolution, and synchronously. As the resolution and speed of digital multipliers, and MACs increase, usually the dynamic power consumption and chip size of digital implementations increases substantially that makes them impractical for some ML and AI segments, including in portable, mobile, near edge, or near sensor applications. The multipliers and MACs utilizing the disclosed current mode data-converters are manufacturable in main-stream digital CMOS process, and they can have medium to high resolutions, capable of low power consumptions, having low sensitivity to power supply and temperature variations, as well as operating asynchronously, which makes them suitable for high-volume, low cost, and low power ML and AI applications. Moreover, the multipliers and MACs disclosed in this invention can be placed near conventional CMOS memory cells, such as Static-Random-Access-Memory (SRAM) or Electrically Programmable Read-Only Memory (EPROM) or Electrically Erasable Programmable Read-Only Memory (E2PROM), which facilitates In-Memory-Compute (IMC) and or near-memory-compute (NMC), that can further reduce dynamic power consumption.


