Digital Compute in Memory Architecture for Neural Network Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional computation-in-memory (CIM) processes for machine learning tasks face challenges in accuracy due to the use of analog signals, which can lead to inaccuracy in neural network computations, and require significant power and space, making them unsuitable for edge devices with power and packaging constraints.
Innovation Solution
A digital CIM architecture that includes a circuit with multiple memory cells on bit-lines configured to store neural network weights and accumulators to perform computations and accumulate output signals after sequential activation of word-lines, reducing data transfer bottlenecks and enhancing accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional CIM processes use analog signals for computation, then computation can be performed in memory, but accuracy deteriorates due to susceptibility to noise and computational errors
Solution Approach 1:
The patent replaces analog signal-based computation with digital computation in memory. Instead of using continuous analog voltages that are susceptible to noise, the invention uses discrete digital signals (0 and 1) to represent data and perform computations. This substitution of analog with digital fundamentally resolves the accuracy problem while maintaining in-memory computation capability.
Solution Approach 2:
The patent changes the parameter of signal representation from continuous analog values to discrete digital values. By transforming the computation domain from analog to digital, the system achieves both in-memory computation and high accuracy, as digital signals are inherently more robust against noise and degradation.
2Productivity
If dedicated hardware accelerators are used for machine learning processing, then processing capacity is enhanced, but power consumption and device space increase
Solution Approach 1:
The patent merges the computation function with the memory function into a single integrated structure. Instead of having separate memory and processing units that require data transfer between them, the invention combines both functions in the same physical location, eliminating the need for dedicated hardware accelerators and reducing overall power consumption.
Solution Approach 2:
The patent makes the memory device multi-functional by enabling it to perform both storage and computation operations. This universal approach allows the same hardware structure to serve multiple purposes, eliminating the need for specialized accelerators and reducing the overall hardware footprint and power requirements.
3Productivity
If data is moved across common data busses for processing, then processing can be performed, but power usage increases and latency is introduced
Solution Approach 1:
The patent extracts the computation function from external processing units and moves it directly into the memory array. By taking out the computation capability from separate processors and embedding it within the memory structure itself, the system eliminates the need for data movement across data busses, thereby reducing power consumption and latency.
4Productivity
If more processing capabilities are added to edge devices, then machine learning tasks can be performed, but space and packaging constraints are violated
Solution Approach 1:
The patent combines memory and processing functions into a single integrated structure, effectively doubling the functionality of the same physical space. This merging allows edge devices to gain machine learning processing capabilities without increasing the overall device footprint, as the computation is performed within the existing memory array.
Data Source
AI summary
Certain aspects generally relate to performing machine learning tasks, and in particular, to computation-in-memory architectures and operations. One aspect provides a circuit for in-memory computation. The circuit generally includes multiple bit-lines, multiple word-lines, an array of compute-in-memory cells, and a plurality of accumulators, each accumulator being coupled to a respective one of the multiple bit-lines. Each compute-in-memory cell is coupled to one of the bit-lines and to one of the word-lines and is configured to store a weight bit of a neural network.


