VLSI Data Path Elements Using Microcell Parallel Topologies
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing ASIC designs face challenges in efficiently implementing computational logic due to data-flow dependencies, which hinder the exploitation of parallelism and result in reduced processing throughput.
Innovation Solution
The implementation of data path elements using a plurality of microcells connected in various topologies (serial, parallel, cascade) to perform atomic operations on atomic input data, allowing for the split of operations into sub-atomic operations and data fragments to enhance parallelism and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If standard-cell methodology is used to design ASICs with pre-laid out digital-logic gates, then manufacturing scalability is improved, but processing throughput deteriorates due to data-flow dependencies that hinder parallelism exploitation
Solution Approach 1:
The patent segments the computational logic into multiple independent processing elements (PEs) that operate in parallel. Each PE is further divided into sub-PEs that can process different data fragments simultaneously. This segmentation breaks the data-flow dependencies that limit throughput in standard-cell designs, allowing multiple operations to execute concurrently while maintaining manufacturing scalability through modular architecture.
Solution Approach 2:
The patent introduces a new dimension of parallelism by processing data at the bit-level rather than at the word-level. By organizing processing elements to handle individual bits or small bit-groups in parallel across multiple clock cycles, the design achieves throughput improvement without increasing the overall data path width, effectively adding a temporal dimension to the computational process.
2Productivity
If parallelism is exploited at data-word boundaries to enhance processing throughput, then productivity is improved, but device complexity increases due to the need to manage data-flow dependencies
Solution Approach 1:
The patent segments both the data and the processing units into matching granularities. Data is divided into fragments that correspond to the capabilities of individual processing elements, and each PE is segmented into sub-PEs that handle specific bit-operations. This matching segmentation eliminates the need for complex dependency management because each segment operates independently with well-defined interfaces, reducing control logic complexity while maintaining high parallelism.
Solution Approach 2:
The patent changes the operational parameters of the processing elements by allowing them to operate at different clock cycles and with different data fragment sizes. This parameter flexibility enables simple control logic to manage complex parallel operations, as the same basic PE structure can be configured for different operations without requiring complex interdependence management.
3Area of stationary object
If data width is reduced to consume less silicon resources, then area is reduced, but manufacturing precision deteriorates as operations are performed with less precision
Solution Approach 1:
The patent segments wide data operations into multiple narrower operations performed by smaller processing elements. Instead of using a single wide data path that consumes large silicon area, the design uses multiple narrow data paths processed in parallel by smaller PEs. Each PE operates with reduced data width (e.g., processing 8-bit fragments instead of 64-bit words), reducing individual PE area, while the parallel combination achieves the same overall precision and throughput.
Solution Approach 2:
The patent employs periodic action by cycling through multiple clock cycles to complete wide data operations. Each clock cycle processes a fragment of the overall computation using a narrow data path, and the results are accumulated over successive cycles. This periodic processing allows the system to achieve high-precision results through repeated narrow operations rather than requiring a single wide operation, thereby reducing silicon area while maintaining precision.
Data Source
AI summary
Data path elements implemented using a plurality of logic primitives are provided. The plurality of logic primitives are connected in one or more topologies to perform an atomic operation on at least two atomic input data of a pre-defined bit-size. The atomic operation is split into a plurality of sub-atomic operations and the at least two atomic input data is split into a plurality of sub-atomic data fragments. Each of the plurality of logic primitives perform a sub-atomic operation from the plurality of sub-atomic operations on at least two sub-atomic data fragments from the plurality of sub-atomic data fragments to generate at least one partial sub-atomic output data. The at least one partial sub-atomic output data is generated based on a partial arithmetic operation, a shift operation, or a logical operation performed by each of the plurality of logic primitives in each clock-cycle. The architectures of a plurality of data-path elements comprised of a plurality of logic primitives in a plurality of connection topologies are provided.


