Digital CIM Accumulator Using Column Counters for Accurate ML
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional computation-in-memory (CIM) processes using analog signals result in inaccurate neural network computations, and there is a need for more efficient and accurate processing of machine learning model data, particularly in resource-constrained devices like mobile and IoT devices.
Innovation Solution
A digital computation-in-memory (DCIM) architecture utilizing digital counters and accumulators to perform in-memory computations, with energy-efficient and high-speed operation, reducing electromagnetic interference and energy consumption by using self-timed operations and phase-shifting of local clocks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional analog CIM processes are used, then computation can be performed in memory, but computation accuracy deteriorates
Solution Approach 1:
The patent replaces analog signal processing with digital signal processing in the CIM architecture. Digital counters and digital logic circuits substitute for analog accumulators, eliminating the inherent inaccuracies of analog computation while maintaining in-memory processing capabilities. This substitution enables precise digital counting and addition operations directly within the memory structure.
Solution Approach 2:
The patent changes the fundamental operating parameter from analog voltage levels to digital binary states (0 and 1). By encoding weights and computations in digital form rather than analog signal magnitudes, the system achieves exact representation and computation without the noise and precision limitations of analog signals.
2Productivity
If dedicated machine learning accelerators are used, then processing capacity is enhanced, but space and power consumption increase
Solution Approach 1:
The patent merges the memory function and computation function into a single integrated structure. The same memory cells that store data are used to perform computations through digital counting and addition operations, eliminating the need for separate accelerator hardware and reducing overall power consumption while maintaining high processing capacity.
Solution Approach 2:
The memory structure is designed to serve multiple functions: data storage, digital counting, and accumulation operations. This multi-functionality allows the system to perform machine learning computations without requiring dedicated accelerator hardware, thereby reducing power consumption and space requirements.
3Productivity
If data is moved across common data busses, then processing can be performed, but power usage increases and latency is introduced
Solution Approach 1:
The patent segments the computation process into column-level operations that can be performed independently within the memory structure. Each column has its own digital counter, allowing parallel processing without requiring data to be moved across shared data busses, thereby reducing power consumption and latency.
4Productivity
If more processing hardware is added to edge devices, then machine learning capability is improved, but device area and power constraints are violated
Solution Approach 1:
The patent combines storage and computation functions within the same memory structure, eliminating the need for separate processing hardware. This integration enables edge devices to perform machine learning computations without adding extra hardware components, thereby respecting area and power constraints while improving ML capability.
Data Source
AI summary
Certain aspects provide an apparatus for performing machine learning tasks, and in particular, to computation-in-memory architectures. One aspect provides a method for in-memory computation. The method generally includes: accumulating, via each digital counter of a plurality of digital counters, output signals on a respective column of multiple columns of a memory, wherein a plurality of memory cells are on each of the multiple columns, the plurality of memory cells storing multiple bits representing weights of a neural network, wherein the plurality of memory cells of each of the multiple columns correspond to different word-lines of the memory; adding, via an adder circuit, output signals of the plurality of digital counters; and accumulating, via an accumulator, output signals of the adder circuit.


