Hierarchical Clock Tree Distribution for IC Skew Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current clock distribution methods for integrated circuit chips face challenges in achieving low skew with flexible planning and efficient resource use across N-level hierarchical entities, often requiring extensive tuning and large design resources, and are limited by availability of thick metal wires.
Innovation Solution
The method involves independently constructing local and top clock tree distributions using high-performance buffers, with a structured clock buffer floor planner to create a hybrid tree, allowing for balanced clock block regions and eliminating the need for random logic macro basining, thereby enabling parallel timing closure and consistent routing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If clock grids or meshes are used to achieve very low skew, then skew is reduced, but wiring and design resources are costly
Solution Approach 1:
The patent divides the clock distribution system into hierarchical levels (top-level clock tree and local clock trees within RLMs). This segmentation allows skew control at each level independently, achieving low overall skew while reducing the total wiring resources needed compared to a full mesh structure.
Solution Approach 2:
The patent applies different clock distribution strategies to different hierarchical levels. The top-level uses a structured clock buffer tree, while local RLMs use their own optimized buffer trees. This local optimization allows each region to use resources efficiently while maintaining overall skew performance.
2Adaptability or versatility
If buffer or inverter trees are used, then flexibility is improved, but skew control is limited and extensive tuning is needed
Solution Approach 1:
The patent pre-balances the local clock trees within each RLM during the design phase, creating pre-balanced buffer trees that compensate for latency differences. This preliminary action eliminates the need for extensive post-placement tuning, as the trees are already optimized for their specific hierarchical contexts.
3Reliability
If large buffers and thick wires are used, then performance is improved, but thick metal is not available inside child hierarchical blocks
Solution Approach 1:
The patent segments the clock distribution into hierarchical levels where thick metal is available at the parent/chip level for the top clock tree, while local RLMs use thinner metal layers appropriate for their scale. This segmentation allows each level to use the metal resources available to it, achieving overall performance without requiring thick metal throughout the entire chip.
Solution Approach 2:
The patent moves the heavy buffering and thick wire requirements to the top-level clock distribution, while local RLMs use lighter resources. This dimensional separation in the hierarchical structure allows performance-critical paths to use thick metal where available, while less critical paths use thinner metal, optimizing resource utilization across the chip.
4Productivity
If random logic macros are used, then data management is improved, but clocking complexity increases
Solution Approach 1:
The patent isolates clocking complexity within individual RLMs by creating independent local clock trees for each macro. This segmentation confines the complexity to manageable units, while the top-level clock distribution remains simple and systematic, reducing overall clocking complexity despite using multiple RLMs.
Solution Approach 2:
The patent pre-balances the clock trees within each RLM during design, creating standardized clocking structures that can be replicated across multiple macros. This preliminary optimization eliminates the need for complex per-macro tuning, allowing data management benefits of RLMs to be realized without proportionally increasing clocking complexity.
Data Source
AI summary
A method, system and computer program product for implementing enhanced clock tree distributions to decouple across N-level hierarchical entities of an integrated circuit chip. Local clock tree distributions are constructed. Top clock tree distributions are constructed. Then constructing and routing a top clock tree is provided. The local clock tree distributions and the top clock tree distributions are independently constructed, each using an equivalent local clock distribution of high performance buffers to balance the clock block regions.


