Method, tool, computer equipment, medium and program for concurrent clock optimization and data optimization

Through concurrent clock optimization and data optimization methods, the clock tree is balanced by utilizing useful skew, and the clock optimization function is integrated into the data optimizer. This solves the problem of low clock tree optimization efficiency and achieves efficient timing convergence and resource utilization.

CN120633585AActive Publication Date: 2025-09-12HUAXIN GIANTS (HANGZHOU) MICROELECTRONICS CO LTD

Patent Information

Application Number
CN202511118440.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-09-12
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

Existing clock tree optimization methods perform clock path optimization and data path optimization separately, resulting in low optimization efficiency, inability to effectively utilize useful skew, resulting in resource waste and poor timing convergence.

Method used

Adopting concurrent clock optimization and data optimization methods, the concurrent optimizer balances the clock tree so that some paths obtain useful skew. The useful skew is used for clock and data optimization. The clock optimization function is integrated into the data optimizer to adjust the clock path in real time to improve timing convergence.

Benefits of technology

Significantly save optimization time, reduce computing resources, effectively utilize design resources, reduce power consumption, and improve timing convergence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633585A_ABST
    Figure CN120633585A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of electronic design automation, in particular to a concurrent clock optimization and data optimization method and tool, computer equipment, a medium and a program. The concurrent clock optimization and data optimization method comprises the following steps: providing a clock optimizer, and constructing a clock tree; initial skew on the clock tree is obtained, and the clock optimizer carries out balance processing on the clock tree based on the initial skew, so that the clock tree obtains useful skew; integrating the clock optimizer and the data optimizer to obtain a concurrent optimizer; and simultaneously performing data optimization on the balanced clock tree by using a concurrent optimizer, and performing clock optimization on the balanced clock tree based on the useful skew. The problem that an existing optimization tool is poor in clock tree optimization efficiency is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electronic design automation technology, and in particular to a method, tool, computer equipment, medium and program for concurrent clock optimization and data optimization. Background Art

[0002] In the integrated circuit field, clock tree optimization is performed after layout optimization based on the clock tree results. Existing clock tree optimization methods generally include two methods: clock path optimization and data path optimization. Clock tree synthesis is a key component connecting the layout and routing stages. The quality of the clock tree significantly impacts the timing closure and power consumption of the entire design.

[0003] Typically, after layout optimization is complete, a zero-skew clock tree is built. The expected quality of the results is consistent with that after layout optimization. However, after the clock tree is built, the skew is generally not completely zero, and there are also on-chip variations. Therefore, the timing results after clock tree construction are often worse than those after layout optimization. Therefore, the tool needs to optimize the data path based on the current clock skew.

[0004] For a register, its capture clock path also serves as the trigger clock path for the next-level register. Therefore, a zero-skew clock tree is not always optimal for timing closure. Therefore, clock skew that contributes to timing closure is called useful skew. Using useful skew can reduce the complexity of datapath optimization, effectively utilize design resources, and avoid the need for numerous buffers introduced by optimizations in the datapath, thereby reducing power consumption.

[0005] Existing methods optimize clock paths and data paths separately. Clock optimization focuses on clock delay and clock skew, but fails to account for the direct impact of clock tree changes on timing closure. Furthermore, clock tree optimization fails to account for on-chip variations, clock gating, and clock pessimism elimination, resulting in a significant gap between actual optimization goals and intended optimization objectives. This separation of clock and data optimization often leaves the impact on timing closure unpredictable, leading to poor optimization efficiency. Summary of the Invention

[0006] In order to solve the problem of poor efficiency of existing optimization tools for clock tree optimization, the present invention provides a method, tool, computer device, medium and program for concurrent clock optimization and data optimization.

[0007] In order to solve the above technical problems, the present invention provides the following technical solutions: a method for concurrent clock optimization and data optimization, the method comprising the following steps: providing a clock optimizer and constructing a clock tree; providing a concurrent optimizer; obtaining the initial skew of each path on the clock tree, balancing the clock tree based on the initial skew so that some clock tree paths obtain useful skew and some clock tree paths obtain zero skew; using the concurrent optimizer to perform data optimization on the balanced clock tree, and simultaneously performing clock optimization on the balanced clock tree based on the useful skew; or, using the concurrent optimizer to perform data optimization on the balanced clock tree, and simultaneously performing clock optimization on the balanced clock tree.

[0008] Preferably, balancing the clock tree based on the initial skew includes: providing preset design conditions to determine whether the constructed clock tree meets the preset design conditions; if not, the clock optimizer directly balances the clock tree based on the initial skew; if it meets, using a concurrent optimizer to perform data optimization on the clock tree, providing a timing updater, and after data optimization, the timing updater combines with the concurrent optimizer to calculate the offset based on the initial skew, optimizes the clock tree based on the offset, and balances the clock tree after data optimization and clock optimization.

[0009] Preferably, before balancing the clock tree after data optimization and clock optimization, the process also includes: obtaining the initial delay change of each node in the constructed clock tree; after calculating the offset, obtaining the updated delay change of each node in the clock tree based on the offset; and re-clustering each node in the constructed clock tree based on the updated delay change.

[0010] Preferably, the data-optimized timing updater combines with the concurrent optimizer to calculate the offset based on the initial skew, including: after data optimization, obtaining the worst timing margin of the fan-in path and the worst timing margin of the fan-out path based on the timing updater; obtaining the offset based on the worst timing margin of the fan-in path and the worst timing margin of the fan-out path; wherein the offset includes the effective skew of the clock advance and the effective skew of the clock delay; the effective skew of the clock advance is calculated as: Formula 1: When d_slack<0, if q_slack>0, the effective skew is the smaller value of q_slack and negative d_slack. If q_slack<0, the effective skew is (q_slack-d_slack) / 2; The effective skew of the clock delay is calculated as follows: Formula 2: Used when q_slack < 0. If d_slack > 0, the effective skew is the smaller value of d_slack and negative q_slack. If d_slack < 0, the effective skew is (d_slack – q_slack) / 2. In Formula 1 and Formula 2, d_slack is the worst timing margin of the fan-in path, and q_slack is the worst timing margin of the fan-out path.

[0011] Preferably, providing a concurrent optimizer includes: providing a buffer, a register and a data optimizer, configuring an interface for connecting the register and the buffer in the data optimizer, the register is the starting point and end point of the timing path, and the buffer is a resource that can optimize the timing path; and integrating the clock optimization function into the data optimizer.

[0012] Preferably, the clock optimization function is integrated into the data optimizer, including: reverse traversal of the clock ports on the starting point and end point registers to find the clock path; traversing the elements on the clock path, obtaining the load driven by the element through the clock path and whether there is an end point with poor timing in the driven load; judging whether the end point can be optimized; if so, adjusting the clock path; updating the timing information of the clock path and the associated timing path through the timing updater to perceive the result of timing convergence after the optimization of the end point; judging whether to cancel the operation of adjusting the clock path based on the result of timing convergence; if so, looking for the next end point that can be adjusted until the timing converges or the clock path traversal is completed.

[0013] In order to solve the above technical problems, the present invention provides another technical solution as follows: the tool includes: Clock optimizer: used to build and balance the clock tree; Concurrent optimizer: includes concurrent clock data optimizer and optimization processor; Concurrent clock data optimizer: used to optimize data of the balanced clock tree and schedule clock optimization of the balanced clock tree based on useful skew; Optimization processor: used to perform data optimization or clock optimization operations.

[0014] Preferably, the tool further comprises: Library component processor: connected to the optimization processor, used to process the units on the time path and data path respectively, and directly act on the optimization processor; Timing Updater: Used to update the optimization results after each optimization is completed, to determine whether the timing path has converged and to obtain the timing margin.

[0015] Offset converter: used to transmit the offset calculated by the concurrent clock data optimizer.

[0016] In order to solve the above technical problems, the present invention provides another technical solution as follows: a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the method according to the above concurrent clock optimization and data optimization.

[0017] In order to solve the above technical problems, the present invention provides another technical solution as follows: a computer device, applied to the above-mentioned concurrent clock optimization and data optimization method, including a memory, a processor and a computer program stored in the memory, and the processor executes the above-mentioned computer program to implement the concurrent clock optimization and data optimization method.

[0018] In order to solve the above technical problems, the present invention provides another technical solution as follows: a computer program product, including a computer program or instructions, which implements the above-mentioned concurrent clock optimization and data optimization method when executed by a processor.

[0019] Compared with the prior art, the concurrent clock optimization and data optimization methods, tools, computer equipment, media, and programs provided by the present invention have the following beneficial effects: 1. An embodiment of the present invention provides a method for concurrent clock optimization and data optimization, the method comprising the following steps: providing a clock optimizer and constructing a clock tree; A concurrent optimizer is provided; the initial skew of each path on the clock tree is obtained, and the clock tree is balanced based on the initial skew, so that some clock tree paths have useful skew and other clock tree paths have zero skew; the concurrent optimizer is used to perform data optimization on the balanced clock tree, and the clock optimization of the balanced clock tree is performed based on the useful skew; or, the concurrent optimizer is used to perform data optimization on the balanced clock tree, and the clock optimization of the balanced clock tree is performed simultaneously. In this embodiment, by providing a concurrent optimizer to perform clock optimization and data optimization simultaneously, the optimization time of the clock tree is greatly saved, and the computing resources required for the optimization process are also reduced. In addition, the impact of useful skew on the clock tree is taken into account, and different paths of the clock tree are selectively optimized, further reducing the waste of resources in the optimization process.

[0020] 2. The embodiment of the present invention balances the clock tree based on the initial skew, including: providing preset design conditions to determine whether the constructed clock tree meets the preset design conditions; if not, the clock optimizer directly balances the clock tree based on the initial skew; if it meets, the concurrent optimizer is used to optimize the clock tree data, and a timing updater is provided. After data optimization, the timing updater combines with the concurrent optimizer to calculate the offset based on the initial skew, optimizes the clock tree based on the offset, and balances the clock tree after data optimization and clock optimization. In this embodiment, two operating modes are set based on the preset design conditions to reasonably deal with different scenarios. In addition, in extreme mode, data optimization is performed by the concurrent optimizer; the timing updater calculates the offset based on the initial skew; and after clock optimization is performed based on the offset, balancing is performed. This solves the inefficiency problem caused by repeated adjustments of the clock tree on a zero-skew clock tree.

[0021] 3. An embodiment of the present invention also provides a concurrent clock optimization and data optimization tool, which has the same beneficial effects as the above-mentioned concurrent clock optimization and data optimization method, and will not be described in detail here.

[0022] The embodiment of the present invention further provides a computer-readable storage medium, which has the same beneficial effects as the above-mentioned concurrent clock optimization and data optimization method, and will not be described in detail here.

[0023] 5. An embodiment of the present invention also provides a computer device having the same beneficial effects as the above-mentioned concurrent clock optimization and data optimization methods, which will not be described in detail here.

[0024] 6. A computer program product provided by an embodiment of the present invention includes a computer program or instructions, which, when executed by a processor, implements the above-mentioned concurrent clock optimization and data optimization method. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is a flowchart of a method for concurrent clock optimization and data optimization provided by the first embodiment of the present invention.

[0026] Figure 2 This is a flowchart of a method for concurrent clock optimization and data optimization provided by the first embodiment of the present invention, in which a clock optimization function is integrated into a data optimizer.

[0027] Figure 3 2 is a schematic diagram of the structure of a concurrent clock optimization and data optimization tool provided by the second embodiment of the present invention.

[0028] Figure 4 It is a schematic structural diagram of a computer-readable storage medium provided by the third embodiment of the present invention.

[0029] Figure 5 It is a structural diagram of a computer device provided by the fourth embodiment of the present invention.

[0030] Figure 6 It is a schematic diagram of the structure of a computer program product provided by the fifth embodiment of the present invention. DETAILED DESCRIPTION

[0031] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and implementation examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0032] In the embodiments provided herein, it should be understood that "B corresponding to A" means that B is associated with A and B can be determined based on A. However, it should also be understood that determining B based on A does not mean determining B based solely on A; B can also be determined based on A and / or other information.

[0033] It should be understood that references to "one embodiment" or "an embodiment" throughout this specification mean that specific features, structures, or characteristics associated with the embodiment are included in at least one embodiment of the present invention. Therefore, the appearance of "in one embodiment" or "in an embodiment" throughout this specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics may be combined in any suitable manner in one or more embodiments. Those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required for the present invention.

[0034] In various embodiments of the present invention, it should be understood that the size of the serial numbers of the above-mentioned processes does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0035] The flow charts and block diagrams in the accompanying drawings of the present invention illustrate the possible implementation architecture, functions and operations of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementation schemes, the functions marked in the box can also occur in a different order than those marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which is determined based on the functions involved. It should be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.

[0036] In the integrated circuit field, clock tree optimization is performed after layout optimization based on the clock tree results. Existing clock tree optimization methods generally include two methods: clock path optimization and data path optimization. Clock tree synthesis is a key component connecting the layout and routing stages. The quality of the clock tree significantly impacts the timing closure and power consumption of the entire design.

[0037] Typically, after layout optimization is complete, a zero-skew clock tree is built. The expected quality of the results is consistent with that after layout optimization. However, after the clock tree is built, the skew is generally not completely zero, and there are also on-chip variations. Therefore, the timing results after clock tree construction are often worse than those after layout optimization. Therefore, the tool needs to optimize the data path based on the current clock skew.

[0038] For a register, its capture clock path also serves as the trigger clock path for the next-level register. Therefore, a zero-skew clock tree is not always optimal for timing closure. Therefore, clock skew that contributes to timing closure is called useful skew. Using useful skew can reduce the complexity of datapath optimization, effectively utilize design resources, and avoid the need for numerous buffers introduced by optimizations in the datapath, thereby reducing power consumption.

[0039] Existing methods optimize clock paths and data paths separately. Clock optimization focuses on clock delay and clock skew, but fails to account for the direct impact of clock tree changes on timing closure. Furthermore, clock tree optimization fails to account for on-chip variations, clock gating, and clock pessimism elimination, resulting in a significant gap between actual optimization goals and intended optimization objectives. This separation of clock and data optimization often leaves the impact on timing closure unpredictable, leading to poor optimization efficiency.

[0040] To facilitate understanding, some definitions of clock tree optimization are first explained: Clock path optimization: Clock optimization, also known as clock optimization, refers to optimizing the path of the clock signal from the clock source to the clock input of each register. The purpose is to reduce clock skew and delay and ensure that the clock signal reaches all registers synchronously.

[0041] Data path optimization: Data optimization for short, refers to optimizing the path of data signals from register output to the input of the next-level register in order to meet timing constraints (such as setup time and hold time).

[0042] Trigger clock path: This refers to the path of the clock signal from the clock source to the clock input of the current register. The delay of this path affects when the register starts sampling data.

[0043] The capture clock path refers to the path of the clock signal from the clock source to the clock input of the next-level register. The difference between the delay of this path and the delay of the trigger clock path is called clock skew, which directly affects the results of timing checks.

[0044] Data path: This refers to the path along which a data signal is transmitted from the output of one register to the input of the next register. Its delay must meet the setup and hold time requirements.

[0045] Timing setup time: This refers to the minimum time that a data signal must remain stable before the valid edge of the clock signal arrives. If the data path delay is too long, it may cause a setup time violation.

[0046] Clock cycle: This refers to the duration of one complete cycle of a clock signal, which determines the maximum operating frequency of a circuit. Timing constraints are typically based on the clock cycle.

[0047] On-chip variation: This refers to the uncertainty in the actual delay of the same path on the same chip due to local differences in manufacturing process, temperature, or voltage. This variation affects the accuracy of timing analysis.

[0048] Gated clock: A low-power technology that controls the switching of clock signals through logic gates, providing clock signals to registers only when needed to reduce dynamic power consumption.

[0049] Clock Pessimism Removal (CPR) refers to removing overly pessimistic estimates caused by on-chip variations during timing analysis. For example, delay variations on a common clock path may be double-counted in the trigger and capture paths. CPR eliminates this double pessimism.

[0050] Setup Check: Verifies that the data signal can reach the next-level register stably before the clock capture edge. The formula is: Trigger Clock Path Delay + Data Path Delay ≤ Capture Clock Path Delay + Setup Time + Clock Period.

[0051] Hold Check: Verifies that the data signal does not change prematurely after the clock capture edge. The formula is: Trigger Clock Path Delay + Data Path Delay ≥ Capture Clock Path Delay + Hold Time.

[0052] Skew: The difference between the timing delay of the capture clock path and the timing delay of the trigger clock path is called clock skew. Its unit is nanoseconds or picoseconds. Skew directly affects the setup time and hold time margin.

[0053] The impact of the clock tree on timing closure is primarily manifested as follows: A timing path is typically divided into three parts: the trigger clock path, the capture clock path, and the data path. The difference between the timing delay of the capture clock path and the timing delay of the trigger clock path is called clock skew. For a path's setup timing check, positive clock skew is conducive to setup timing closure; for hold timing checks, negative clock skew is conducive to timing closure. For a register, its capture clock path is also the trigger clock path for the next-level register. Therefore, a zero-skew clock tree is not always optimal for timing closure. We refer to clock skew that contributes to timing closure as useful skew. Using useful skew can reduce the complexity of datapath optimization, effectively utilize design resources, and avoid the need for numerous buffers introduced by optimizations along the datapath, thereby reducing power consumption.

[0054] Existing clock tree optimization methods perform clock path optimization and data path optimization separately, which clearly fails to consider the impact of useful skew on optimization. Furthermore, performing clock path optimization and data path optimization separately not only results in structural redundancy and poor optimization results, but also significantly wastes computing resources.

[0055] To solve the problem of poor clock tree optimization efficiency of existing optimization tools, please refer to Figure 1 A first embodiment of the present invention provides a method for concurrent clock optimization and data optimization, the method comprising the following steps: S1, provides clock optimizer and builds clock tree; S2, provides a concurrent optimizer; S3, obtaining the initial skew of each path on the clock tree, and balancing the clock tree based on the initial skew, so that some clock tree paths have useful skew and other clock tree paths have zero skew; S4, uses the concurrent optimizer to optimize the data of the balanced clock tree, and optimizes the clock of the balanced clock tree based on the useful skew; S5, or, using a concurrent optimizer to perform data optimization on the balanced clock tree and clock optimization on the balanced clock tree at the same time.

[0056] As can be understood, first, this embodiment uses a conventional clock optimizer to construct a clock tree. It should be noted that the constructed clock tree contains multiple paths, also known as timing paths. The length and skew of each path in the constructed clock tree are random and uncertain. Skew is a natural consequence of clock signal transmission and cannot be directly adjusted, but can only be changed indirectly through clock tree optimization. Furthermore, the clock tree is balanced, and during this balancing process, some paths are given useful skew, while others are given zero skew. It should be understood that balancing the clock tree is also a process of optimizing the clock tree, which can be performed using a conventional clock optimizer or a concurrent optimizer. In this embodiment, by ensuring that some paths in the clock tree have useful skew during the balancing process, the impact of skew on some paths in the clock tree is effectively utilized during subsequent clock tree optimization. It should be noted that the concurrent optimizer in this embodiment differs from both conventional clock optimizers and data optimizers in that its specific structure integrates clock optimization functionality into the data optimizer, thereby enabling the data optimizer to perform some of the functions of the clock optimizer. This eliminates the need for the concurrent optimizer to perform clock optimization and data optimization separately when optimizing the clock tree, as is common in the prior art. The concurrent optimizer in this embodiment can perform both clock and data optimization simultaneously, significantly saving optimization time and reducing the computing resources required for the optimization process.

[0057] Specifically, after balancing, for the paths that achieve useful skew, the concurrent optimizer will perform data optimization on the balanced clock tree, while simultaneously optimizing the clocks based on the useful skew. It should be understood that using useful skew during simultaneous clock optimization can reduce the complexity of data path optimization, effectively utilize design resources, and avoid the need for large amounts of buffers introduced by optimization on the data path, thereby reducing power consumption. For the remaining paths that achieve zero skew, which are timing paths that do not require useful skew to aid convergence, the concurrent optimizer can directly perform both data and clock optimization on these clock tree paths.

[0058] It should be understood that in this embodiment, by providing a concurrent optimizer to simultaneously perform clock optimization and data optimization, the optimization time for the clock tree is greatly saved, and the computing resources required for the optimization process are reduced. In addition, the impact of useful skew on the clock tree is taken into account. By selectively optimizing different paths of the clock tree, the waste of resources in the optimization process is further reduced.

[0059] Furthermore, in the above step S3, balancing the clock tree based on the initial skew includes: Provide preset design conditions to determine whether the constructed clock tree meets the preset design conditions; If not, the clock optimizer directly balances the clock tree based on the initial skew; If the conditions are met, the concurrent optimizer is used to optimize the clock tree data and provide a timing updater. After data optimization, the timing updater combines with the concurrent optimizer to calculate the offset based on the initial skew. The clock tree is optimized based on the offset and the clock tree is balanced after data and clock optimization.

[0060] It should be noted that in step S3, the preset design conditions can be freely defined according to user needs. For example, the preset design conditions can be determined based on the design complexity of the clock tree. For example, if the constructed clock tree design is relatively complex, the constructed clock tree is automatically determined to meet the preset design conditions. If the constructed clock tree design is relatively simple, the constructed clock tree is automatically determined to not meet the preset design conditions. The design complexity of the clock tree can be determined by the design process node, design scale, and register scale, and the details are not detailed here.

[0061] It should be understood that whether the constructed clock tree meets the preset design requirements can be viewed as one of two modes of clock tree optimization. The first is the simple mode. For example, if the clock tree is relatively simple and does not meet the preset design requirements, the existing clock optimizer can be used to balance the clock tree. This is simple, but the clock optimizer in this mode requires a two-step operation. The first step is to level the clock tree. Because the paths in the initially constructed clock tree may be long or short, direct clock tree optimization without balancing will complicate the subsequent optimization workflow. Therefore, the simplest approach is to first make each path in the clock tree of similar length. This step is called leveling. However, leveling makes all paths in the clock tree of similar length. If subsequent optimization is performed directly, all paths in the clock tree will be in a zero-skew state. This means that the clock tree will be optimized with zero skew, which wastes computing resources. Therefore, the clock optimizer in the simple mode requires a second step. This step aims to readjust the partially leveled clock tree paths to obtain useful skew. This involves lengthening some paths and shortening others to restore useful skew, facilitating subsequent optimization. However, this method of flattening the clock tree in order to balance the clock tree, while shortening this part of the clock tree based on the useful skew, is a very inefficient way of processing. If dealing with a slightly more complex clock tree, this approach will be extremely time-consuming. Therefore, the other of the two modes is the extreme mode: in the extreme mode, the existing clock optimizer is not used for balancing processing, but a timing updater is provided. The role of the timing updater is to update the optimization results after each optimization is completed. It can also be used to determine whether the timing path has converged and obtain the timing margin. Specifically, the concurrent optimizer can be used to optimize the clock tree data first. After the data is optimized, the timing margin can be obtained based on the timing updater. The offset can then be calculated using the timing margin and the initial skew of the clock tree path. It should be noted that the offset is a design parameter used to achieve useful skew. This means that the clock tree is optimized using the offset to balance the clock tree after data and clock optimization. It should be understood that in extreme mode, calculating the offset value is a computational process that does not involve any clock tree operations. Clock tree operations are only performed when optimizing the clock tree based on the offset. That is, unlike simple mode, extreme mode only performs a single operation. After balancing, some clock tree paths directly achieve useful skew, eliminating the need to first level the paths and then lengthen or shorten them, as in simple mode.

[0062] Understandably, this embodiment sets two operating modes based on preset design conditions to appropriately address different scenarios. Furthermore, in extreme mode, data optimization is performed using the concurrent optimizer; the timing updater calculates the offset based on the initial skew; clock optimization is performed based on the offset, followed by balancing. This resolves the inefficiency issue caused by repeated adjustments to the clock tree on a zero-skew clock tree. Furthermore, users can design a button to enable either simple mode or extreme mode, making it easier for them to switch between these two modes.

[0063] Furthermore, before balancing the clock tree after data optimization and clock optimization, the following steps are also performed: Get the initial delay variation of each node in the constructed clock tree; After calculating the offset, the update delay change of each node in the clock tree is obtained based on the offset; The nodes in the constructed clock tree are re-clustered based on the update delay changes.

[0064] It should be understood that after adjusting the clock path based on the offset, the delay distribution of each node in the clock tree changes unevenly: for example, some nodes are shortened due to positive offset requirements; some nodes are lengthened due to negative offset requirements. If the nodes are not adjusted to maintain the original delay, directly balancing the clock tree will result in Nodes with large latency differences will be forced to cluster into groups, which will increase power consumption and computing resources.

[0065] Therefore, by obtaining the update delay changes of each node in the clock tree, the clock tree nodes are re-sampled. Specifically, nodes with similar updated delay values ​​are grouped together, for example, nodes with delays between 80-85ps are grouped together, and nodes with delays between 95-105ps are grouped together. This replaces the traditional grouping based on physical location or logical hierarchy. This reduces the difficulty of clock tree balancing and, consequently, reduces power consumption during the balancing process.

[0066] Furthermore, in step S4, after data optimization, the timing updater combines with the concurrent optimizer to calculate the offset based on the initial skew, including: After data optimization, the worst timing margin of the fan-in path and the worst timing margin of the fan-out path are obtained based on the timing updater; Obtaining an offset based on the worst timing margin of the fan-in path and the worst timing margin of the fan-out path; The offset includes the effective skew of the clock advance and the effective skew of the clock delay; The effective skew of the clock advance is calculated as: Formula 1: When d_slack<0, if q_slack>0, the effective skew is the smaller value of q_slack and negative d_slack. If q_slack<0, the effective skew is (q_slack-d_slack) / 2; The effective skew of the clock delay is calculated as follows: Formula 2: Used when q_slack < 0. If d_slack > 0, the effective skew is the smaller value of d_slack and negative q_slack. If d_slack < 0, the effective skew is (d_slack – q_slack) / 2. In Formula 1 and Formula 2, d_slack is the worst timing margin of the fan-in path, and q_slack is the worst timing margin of the fan-out path.

[0067] It should be understood that this embodiment transforms pessimism elimination into a skew redistribution problem under mathematical constraints, and dynamically calculates the offset using the worst timing margin of the real-time fan-in path and the worst timing margin of the fan-out path of the fan-in / fan-out path.

[0068] Further, see Figure 1 , providing concurrent optimizers including: Provide buffers, registers, and data optimizers. Configure interfaces for connecting registers and buffers in the data optimizer. Registers are the starting and ending points of the timing path, and buffers are resources that can optimize the timing path. The clock optimization function is integrated into the data optimizer.

[0069] It should be understood that traditional effective skew-based clock optimization and data optimization have separate data optimizers and clock optimizers. Clock optimization selects clock paths and focuses on optimizing clock delay and clock skew, failing to directly assess the impact of clock tree changes on timing closure. Furthermore, clock tree optimization fails to consider timing closure-related information such as on-chip variations, clock gating, and clock pessimism elimination. This leads to a significant gap between the actual optimization goal and the intended optimization goal. Data optimization, on the other hand, considers timing margins and optimizes the worst timing path, directly assessing the impact of optimization on timing closure. The concurrent clock and data optimizer in this implementation is a program module responsible for scheduling the data optimizer and the timing updater. It achieves design timing closure by selecting the worst timing path, finding optimized components, executing optimization operations, and performing timing updates. Registers and buffers are components in electronic designs processed by the program. Registers are the starting and ending points of timing paths, while buffers are components that the program can process to add optimized timing paths to the design. Registers are usually fixed when designing the front end. Their data paths implement complex functions through various combinational logic elements. Buffers do not change the state of the signal and can be inserted by the program.

[0070] Specifically, please further combine Figure 2 , integrating clock optimization functions into the data optimizer includes: s221, reversely traverse the clock ports on the start and end registers to find the clock path; s222, traverse the elements on the clock path, obtain the load driven by the element through the clock path, and whether there is an endpoint with poor timing among the driven loads; s223, determine whether the endpoint can be optimized; s224, if yes, adjust the clock path; s225, updating the timing information of the clock path and the associated timing paths through the timing updater to perceive the timing closure result after the optimization of the endpoint; s226, judging whether to cancel the operation of adjusting the clock path based on the result of timing closure; s227, if yes, search for the next endpoint of the adjustable process until the timing converges or the clock path traversal is completed.

[0071] It should be understood that an ordinary data optimizer will select the end point of the timing path with a negative timing margin. This end point is generally the data port of the register, and the worst data path is obtained from this port, and then the data path is optimized. However, the concurrent clock data optimizer of this embodiment integrates the clock optimization part into the data optimizer. The clock optimization optimizes the path with poor timing after path optimization. The clock path is found by reverse traversing the clock ports on the start and end registers, traversing the elements of the clock path, and judging whether the end point can be optimized by looking at the load driven by the current element of the clock path and whether there is an end point with poor timing in the driven load. If it can be optimized, the clock optimization will call the optimization processing module to adjust the clock path by inserting buffers and replacing elements on the clock path. After the clock path is adjusted, the timing information of the clock path and the associated timing path can be updated through the timing update module, so that the impact of the clock optimization on timing convergence can be directly perceived and the timing convergence result can be obtained. For example, if the modification of this endpoint causes timing degradation, this step will be undone and the next executable endpoint will be searched for until timing convergence or the clock path traversal is completed. It should be understood that the concurrent clock path optimization provided in this embodiment can more accurately optimize the clock path, and the next round of optimization can directly perceive the changes in the timing convergence results, which can achieve better optimization results than the separate data optimization and clock optimization in the prior art.

[0072] To demonstrate the effects of the concurrent clock optimization and data optimization methods in this embodiment, the following is experimental data from the first embodiment: Experiment 1: The test case set in Table 1 was tested using the method without concurrent clock data optimization, the method with concurrent clock data optimization, and the method with offset-based concurrent clock data optimization.

[0073] Table 1: Test case set:

[0074] Table 2: Test results

[0075] Table 1 shows the test case set, and Table 2 shows the test results for the test case set in Table 1 using the three methods in Experimental Example 1. In Table 2, the Worst Negative Slack (WNS) is the most severe timing violation value among all timing paths. The Total Negative Slack (TNS) is the sum of the negative slacks of all paths that violate the timing constraints (paths with negative slack). It should be noted that concurrent clock data optimization corresponds to the simple mode in the first embodiment, while offset-based concurrent clock data optimization corresponds to the extreme mode in the first embodiment. It can be seen that both the simple mode and the extreme mode significantly improve timing closure and runtime compared to traditional data optimization using effective skew. Specifically, the simple mode achieves better optimization results and very short runtimes. While the extreme mode increases runtime, it is still faster than traditional methods and is primarily used for complex design scenarios.

[0076] Please combine Figure 1 and Figure 3 The second embodiment of the present invention further provides a concurrent clock optimization and data optimization tool, which applies the above-mentioned concurrent clock optimization and data optimization method, and the tool includes: Clock optimizer: used to build and balance the clock tree; Concurrent optimizer: includes concurrent clock data optimizer and optimization processor; Concurrent clock data optimizer: used to optimize data of the balanced clock tree and schedule clock optimization of the balanced clock tree based on useful skew; Optimization processor: used to perform data optimization or clock optimization operations.

[0077] It can be understood that the tool provided in this embodiment greatly saves the optimization time of the clock tree and reduces the computing resources required for the optimization process. In addition, the impact of the useful skew on the clock tree is taken into account. By selectively optimizing different paths of the clock tree, the waste of resources in the optimization process is further reduced. Among them, the concurrent clock data optimizer and the clock optimizer both act as schedulers. Specifically, the concurrent clock data optimizer can coordinate data path optimization and clock path optimization in real time, such as concurrent optimization based on skew, or concurrent optimization based on offset. It should be noted that both clock optimization using offset and clock optimization with effective skew, both optimization processes can complete the joint optimization of clock path and data path, and there will be a big difference mainly in the clock path optimization part. The optimization processor plays the role of executor, which performs the optimization operation.

[0078] Furthermore, the tool also includes: Library element processor: connected to the optimization processor, used to process the units on the time path and data path respectively, and directly act on the optimization processor. It should be understood that the optimization processor can selectively call the clock / data path unit in the library element processor. For example, if the clock path only uses a low-power buffer, the clock path unit can be called first.

[0079] Offset converter: used to transmit the offset calculated by the concurrent clock data optimizer. Specifically, after the offset calculation is completed, it is transmitted to other modules through the offset converter.

[0080] Timing Updater: Used to update optimization results after each optimization is completed, determine whether the timing path has converged, and determine the timing margin. It should be understood that the Timing Updater can update both datapath timing and clockpath timing. Datapath timing updates have a wide update range but are slow, while clockpath updates have a narrow update range and are fast.

[0081] Please combine Figure 1 and Figure 4 The third embodiment of this embodiment also provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the method according to the above-mentioned concurrent clock optimization and data optimization.

[0082] The computer-readable storage medium provided in the embodiment of the present invention has the same beneficial effects as the above-mentioned concurrent clock optimization and data optimization methods, and will not be described in detail here.

[0083] Please combine Figure 1 and Figure 5The fourth embodiment of this embodiment also provides a computer device, which is applied to the above-mentioned concurrent clock optimization and data optimization method, including a memory, a processor and a computer program stored in the memory, and the processor executes the above-mentioned computer program to implement the concurrent clock optimization and data optimization method.

[0084] The computer device provided by the embodiment of the present invention has the same beneficial effects as the above-mentioned concurrent clock optimization and data optimization methods, which will not be described in detail here.

[0085] Please combine Figure 1 and Figure 6 The fifth embodiment of this embodiment also provides a computer program product, including a computer program or instructions, which implements the above-mentioned concurrent clock optimization and data optimization method when executed by a processor.

[0086] The computer program product provided by the embodiment of the present invention has the same beneficial effects as the above-mentioned concurrent clock optimization and data optimization methods, and will not be described in detail here.

[0087] The above is a detailed introduction to the method, tool, computer equipment, medium and program for concurrent clock optimization and data optimization disclosed in the embodiment of the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention. Any modifications, equivalent replacements and improvements made within the principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for concurrent clock optimization and data optimization, characterized by: The method comprises the following steps: Provides clock optimizer and builds clock tree; Provide concurrent optimizer; Obtain the initial skew of each path on the clock tree and balance the clock tree based on the initial skew so that some clock tree paths have useful skew and other clock tree paths have zero skew. Use the concurrent optimizer to perform data optimization on the balanced clock tree, and also perform clock optimization on the balanced clock tree based on the useful skew. Alternatively, use a concurrent optimizer to perform data optimization on the balanced clock tree and clock optimization on the balanced clock tree at the same time.

2. The method for concurrent clock optimization and data optimization according to claim 1, wherein: Balancing the clock tree based on initial skew involves: Provide preset design conditions to determine whether the constructed clock tree meets the preset design conditions; If not, the clock optimizer directly balances the clock tree based on the initial skew; If the conditions are met, the concurrent optimizer is used to optimize the clock tree data and provide a timing updater. After data optimization, the timing updater combines with the concurrent optimizer to calculate the offset based on the initial skew. The clock tree is optimized based on the offset and the clock tree is balanced after data and clock optimization.

3. The method for concurrent clock optimization and data optimization according to claim 2, wherein: Before balancing the clock tree after data and clock optimization, the following steps are also performed: Get the initial delay variation of each node in the constructed clock tree; After calculating the offset, the update delay change of each node in the clock tree is obtained based on the offset; The nodes in the constructed clock tree are re-clustered based on the update delay changes.

4. The method for concurrent clock optimization and data optimization according to claim 2, wherein: The data optimized timing updater combines with the concurrent optimizer to calculate the offset based on the initial skew, including: After data optimization, the worst timing margin of the fan-in path and the worst timing margin of the fan-out path are obtained based on the timing updater; Obtaining an offset based on the worst timing margin of the fan-in path and the worst timing margin of the fan-out path; The offset includes the effective skew of the clock advance and the effective skew of the clock delay; The effective skew of the clock advance is calculated as: Formula 1: When d_slack<0, if q_slack>0, the effective skew is the smaller value of q_slack and negative d_slack. If q_slack<0, the effective skew is (q_slack-d_slack) / 2; The effective skew of the clock delay is calculated as follows: Formula 2: Used when q_slack < 0. If d_slack > 0, the effective skew is the smaller value of d_slack and negative q_slack. If d_slack < 0, the effective skew is (d_slack – q_slack) / 2. In Formula 1 and Formula 2, d_slack is the worst timing margin of the fan-in path, and q_slack is the worst timing margin of the fan-out path.

5. The method for concurrent clock optimization and data optimization according to claim 2, wherein: Concurrent optimizers include: Provide buffers, registers, and data optimizers. Configure interfaces for connecting registers and buffers in the data optimizer. Registers are the starting and ending points of the timing path, and buffers are resources that can optimize the timing path. The clock optimization function is integrated into the data optimizer.

6. The method for concurrent clock optimization and data optimization according to claim 5, wherein: Clock optimization features integrated into the data optimizer include: Perform reverse traversal on the clock ports on the start and end registers to find the clock path; Traverse the components on the clock path, obtain the loads driven by the components through the clock path, and check whether there are endpoints with poor timing in the driven loads; Determine whether the endpoint can be optimized; If yes, adjust the clock path; Updating the timing information of the clock path and the associated timing paths through the timing updater to perceive the timing convergence result after the optimization of the endpoint; Determine whether to cancel the clock path adjustment operation based on the timing closure result; If so, look for the next endpoint of the adjustable process until the timing converges or the clock path traverses.

7. A concurrent clock optimization and data optimization tool, applying the concurrent clock optimization and data optimization method according to any one of claims 1 to 6, characterized in that: The tools include: Clock optimizer: used to build and balance the clock tree; Concurrent optimizer: includes concurrent clock data optimizer and optimization processor; Concurrent clock data optimizer: used to optimize data of the balanced clock tree and schedule clock optimization of the balanced clock tree based on useful skew; Optimization processor: used to perform data optimization or clock optimization operations.

8. The concurrent clock optimization and data optimization tool according to claim 7, Its characteristics are: The tools also include: Library component processor: connected to the optimization processor, used to process the units on the time path and data path respectively, and directly act on the optimization processor; Timing Updater: used to update the optimization results after each optimization is completed, to determine whether the timing path has converged and to obtain the timing margin; Offset converter: used to transmit the offset calculated by the concurrent clock data optimizer.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the method for concurrent clock optimization and data optimization according to any one of claims 1 to 6.

10. A computer device, applied to the method for concurrent clock optimization and data optimization according to any one of claims 1 to 6, characterized in that: The system comprises a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the method of concurrent clock optimization and data optimization.

11. A computer program product, characterized in that: The method comprises a computer program or an instruction, which, when executed by a processor, implements the method for concurrent clock optimization and data optimization described in any one of items 1-6.

Citation Information

Patent Citations

  • Clock offset locality optimizing analysis method

    CN101504680A

  • Processor performance optimization method based on clock planning deviation algorithm

    CN103324774A

  • Layout and wiring method suitable for improving CPU core frequency

    CN109783984A

  • Data path optimization method and device, electronic equipment and storage medium

    CN118627463A

  • Clock tree adaptive optimization method using useful deviation based on RISC-V architecture

    CN119358505A

Cited By

  • Time sequence convergence method and device based on clock tree

    CN121503414A