Digital Compute-in-Memory Multicast Weights for Lower Area and Latency
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing digital compute-in-memory (DCIM) systems are constrained by a unicasting architecture that replicates analog computer-in-memory (ACIM) systems, leading to redundant weight-vectors and increased area consumption, signal latencies, and reduced reliability due to quantization challenges as semiconductor components shrink.
Innovation Solution
Implementing a digital multicasting architecture that eliminates redundant weight-vectors and reduces area consumption by multicast weight-vectors to multiple multipliers, reducing signal line lengths and propagation losses, and enhancing operational speed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a unicasting architecture is used to replicate analog computer-in-memory systems, then the system can perform compute-in-memory operations, but redundant weight-vectors are created leading to increased area consumption
Solution Approach 1:
The patent merges multiple weight-vector storage locations into a single shared weight-vector storage unit. Instead of replicating weight-vectors across multiple storage locations as in traditional unicasting architectures, the invention allows multiple multipliers to access and share the same weight-vectors from a common storage location, thereby eliminating redundant storage and reducing area consumption while maintaining compute-in-memory functionality
2Ease of operation
If weight-vectors are replicated across multiple storage locations, then each multiplier can access its required weights, but signal line lengths increase leading to propagation losses and increased latency
Solution Approach 1:
The patent extracts the redundancy of weight-vector replication and removes it from the system. By taking out the unnecessary copies of weight-vectors from multiple storage locations and consolidating them into a single shared storage location, the system eliminates the long signal lines required to distribute replicated weights, thereby reducing signal propagation time and latency while maintaining full weight access capability for all multipliers
3Area of moving object
If semiconductor components are shrunk to increase transistor density, then device size is reduced, but quantization challenges increase reducing reliability
Solution Approach 1:
The patent substitutes the mechanical/analog approach of replicating weight-vectors in analog computer-in-memory systems with a digital sharing approach. By using digital weight-vector storage and sharing mechanisms, the system achieves higher precision and reliability in quantized environments, overcoming the limitations imposed by component shrinking while maintaining compact device size
Data Source
AI summary
A digital compute-in-memory (DCIM) system includes in a first region of a semiconductor die, memory cells, multipliers and adder trees. The memory cells and the multipliers are arranged in corresponding two-dimensional weighting-arrays (two-dimensional matrices) and multiplying-arrays which are organized into pairs. Each of the multiplying-arrays is coupled to each of input-rows (input-channels) of the input-matrix For each of the pairs, and for a selected weight-row (one-dimensional weight-vector) of the corresponding weighting-array, the selected weight-row is multicast to each of the multipliers in the multiplying-array of the pair. The weighting-arrays together represent a two-dimensional weight-matrix. Each of the multiplying-arrays is configured to perform input-matrix-by-weight-vector multiplication resulting in products corresponding to the input-channels for a combined effect of the CIM system overall being configured to perform matrix-by-matrix multiplication. The adder trees are configured to operate on an input-channel-specific basis including adding the products resulting in sums corresponding to the input-channels.


