Multi-Clocked Netlist Retiming via Domain Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional retiming methods for integrated circuit designs face challenges in reducing the number of state-holding elements and combinational-path latency, while also suffering from significant runtime issues, especially on large netlists, which hinders the optimization of area, power, and delay.
Innovation Solution
The method involves translating the netlist into a retiming graph with implicitly clocked register primitives, partitioning it into regions with identical clock domains, and combining compatible domains to generate new partitions, which are then retimed using a solver to produce a behaviorally equivalent netlist, thereby achieving hybrid domain-based and free-running retiming.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional retiming methods are used to reduce the number of state-holding elements and combinational-path latency, then area and delay are improved, but runtime increases significantly especially on large netlists
Solution Approach 1:
The patent segments the netlist into multiple partitions based on clock domains and register density. Each partition is retimed independently using retiming solvers, allowing the large-scale retiming problem to be divided into smaller, more manageable sub-problems that can be solved more efficiently while maintaining overall optimization quality
Solution Approach 2:
The patent applies different retiming strategies to different regions of the netlist based on their characteristics. Clock-domain-based retiming is applied to synchronous regions, while free-running retiming is applied to asynchronous regions. This localized approach optimizes each region according to its specific requirements rather than applying a uniform retiming method across the entire netlist
2Ease of operation
If the entire netlist is converted to low-level design representation for monolithic retiming, then retiming can be performed uniformly, but runtime consumption increases significantly
Solution Approach 1:
Instead of performing monolithic retiming on the entire converted netlist, the patent divides the netlist into multiple partitions that can be retimed independently. This segmentation maintains the uniformity of retiming operations within each partition while dramatically reducing the overall runtime by avoiding the computational complexity of processing the entire netlist as a single unit
3Reliability
If domain-based retiming is performed within each identical-clock-domain partition, then clock domain integrity is maintained, but additional runtime is consumed compared to unified retiming
Solution Approach 1:
The patent merges the benefits of domain-based retiming with free-running retiming by identifying compatible clock domains and combining them into unified retiming regions. This allows retiming to span across traditional domain boundaries where appropriate, reducing the number of separate retiming operations needed while maintaining clock domain integrity through compatibility checking
4Area of stationary object
If register count is reduced through retiming, then area and power are reduced, but verification algorithms suffer run-time degradation
Solution Approach 1:
The patent performs retiming as a preliminary optimization step before verification, transforming the netlist into a more efficient form with reduced register count and optimized timing. This preliminary retiming action prepares the netlist for verification in a state that reduces verification complexity and runtime, rather than dealing with the original unoptimized structure
Data Source
AI summary
Embodiments of the present disclosure provide enhanced systems and methods for implementing enhanced retiming of multiple clock netlists to improve integrated circuit (IC) design quality and provide enhanced retiming with reduced retiming runtime. Disclosed embodiments provide effective and efficient retiming without sacrificing netlist quality, and yield significant speedup of retiming runtime over traditional retiming.


