Clock network optimization method, computer equipment and storage medium
By optimizing the connection adjustments of the signal nodes and receiver endpoints of the clock network, the number of cloned components is reduced, solving the problems of area and power consumption overhead, load balancing, and wiring congestion in the prior art, and improving the chip design effect.
Patent Information
- Application Number
- CN202511687260.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2045-11-18
AI Technical Summary
Existing clock network optimization methods introduce too many cloned components during the optimization process, which increases the area and power consumption in chip design, and also causes load balancing and wiring congestion problems, affecting the overall design performance.
By obtaining the chip design layout, the set of signal nodes and receiving endpoints is determined, the initial connection results are established, and the connection results are adjusted according to the set goals to reduce the number of cloned components and optimize the clock network.
It reduces the area and power consumption of the clock network, solves load balancing and wiring congestion problems, and improves chip design performance.
Smart Images

Figure CN121145784A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the chip layout planning technical field, and particularly relates to a clock network optimization method, computer equipment and a storage medium. BACKGROUND
[0002] Electronic Design Automation (EDA) refers to a design method of using computer-aided design (CAD) software to complete the functional design, synthesis, verification, physical design (including layout, wiring, layout, design rule checking, etc.) and other processes of a very large scale integrated circuit (VLSI) chip.
[0003] In the design process, the clock network has become a key determinant of power, timing and area (PPA) in integrated circuits. Through the clock network, the latency and skew of signals of each component in the chip can be determined, and optimization is to minimize clock latency and clock skew. However, in the existing optimization method, although the corresponding purpose is achieved to a certain extent, due to the introduction of too many cloned elements, it causes large area and power consumption overhead, and leads to load balancing, wiring congestion and other problems, which seriously affects the overall chip design effect.
[0004] It should be noted that the information disclosed in the above background section is only used to strengthen the understanding of the background of the present disclosure, and therefore can include information that does not constitute prior art known to those skilled in the art. SUMMARY
[0005] Therefore, the present disclosure provides a clock network optimization method, computer equipment and a storage medium to solve or partially solve the above problems.
[0006] In order to achieve the above purpose, in a first aspect, the present disclosure provides a clock network optimization method, comprising: obtaining a chip design layout, determining a signal node set and a receiving endpoint set according to the chip design layout; establishing a connection between the signal node set and the receiving endpoint set, and generating an initial connection result; in response to the initial connection result not meeting the set target, adjusting the connection of the initial connection result, determining whether to accept the adjustment through a set strategy, and generating an intermediate connection result; in response to the intermediate connection result meeting the set target, determining the intermediate connection result as an optimization result.
[0007] In a second aspect, the present disclosure provides a computer device, comprising one or more processors, a memory; and one or more programs, wherein the one or more programs are stored in the memory and executed by the one or more processors, and the programs comprise instructions for performing the method according to the first aspect.
[0008] In a third aspect, the present disclosure provides a non-volatile computer readable storage medium containing a computer program, which, when executed by one or more processors, causes the processors to perform the method of the first aspect.
[0009] From the above, it can be seen that the present disclosure provides a clock network optimization method, a computer device and a storage medium. The method comprises: obtaining a chip design layout, determining a signal node set and a receiving endpoint set according to the chip design layout; establishing a connection between the signal node set and the receiving endpoint set, generating an initial connection result; in response to the initial connection result not meeting a set target, adjusting the initial connection result, determining whether to accept the adjustment through a set strategy, generating an intermediate connection result; in response to the intermediate connection result meeting the set target, determining the intermediate connection result as an optimization result. When the clock network is set, after the signal node and the receiving endpoint to be connected are determined, the initial connection is performed first, and then the connection effect is detected according to the set requirement target. If the target is not met, the adjustment is performed on the basis of the currently established connection. In the adjustment process, with the switching of the signal node to which the receiving endpoint is connected, the corresponding parameters change at the same time, and the cloned devices in the connection also increase or decrease, so as to adjust the number of elements in the clock network. When the number of elements is reduced to meet the requirement, the number of cloned elements in the original optimization process is reduced, thereby reducing the area and power consumption overhead, and the fewer cloned element quantity further solves the problems of load balancing and wiring congestion, and improves the chip design effect. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the related art, the drawings needed to be used in the embodiments or related art description will be briefly introduced. Obviously, the drawings in the following description are only embodiments of the present disclosure, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.
[0011] Figure 1 A hardware structure schematic diagram of an exemplary computer device provided by the embodiments of the present disclosure is shown.
[0012] Figure 2 A basic structure schematic diagram of an EDA tool provided by the embodiments of the present disclosure is shown.
[0013] Figure 3 A schematic diagram showing a basic execution flow of a computing command of an EDA tool provided by an embodiment of the present disclosure is shown.
[0014] Figure 4a A structural schematic diagram of a clock network without optimization provided by an embodiment of the present disclosure is shown.
[0015] Figure 4b A structural schematic diagram of an optimized clock network provided by an embodiment of the present disclosure is shown.
[0016] Figure 5 A flow schematic diagram of an exemplary method provided by an embodiment of the present disclosure is shown.
[0017] Figure 6a A schematic diagram showing determination of a signal node set and a receiving endpoint set before optimization provided by an embodiment of the present disclosure is shown.
[0018] Figure 6b A schematic diagram showing connection based on the signal node set and the receiving endpoint set for optimization provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0019] For the purpose of making the purpose, technical scheme and advantages of the present disclosure more clear, the present disclosure is further described in detail below with reference to the embodiments and the accompanying drawings.
[0020] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should be understood as the common meanings understood by those skilled in the art to which the present disclosure belongs. The terms “first”, “second” and similar terms used in the embodiments of the present disclosure do not represent any order, number or importance, but are only used to distinguish different components. The terms “include”, “contain” and similar terms mean that the elements, objects or method steps listed before the terms cover the elements, objects or method steps listed after the terms and their equivalents, and do not exclude other elements, objects or method steps. The terms “connect” or “connected” and similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms “upper”, “lower”, “left”, “right” and the like only represent relative positional relationships, and when the absolute positions of the described objects are changed, the relative positional relationships can also be changed accordingly.
[0021] Figure 1A structural diagram of a computer device 100 is shown. The computer device 100 can include a processor 102, a memory 104, a network interface 106, a peripheral interface 108, and a bus 110. The processor 102, the memory 104, the network interface 106, and the peripheral interface 108 are communicatively connected to each other via the bus 110.
[0022] The processor 102 can be a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), a microcontroller (MCU), a programmable logic device, a digital signal processor (DSP), an application specific integrated circuit (ASIC), or one or more integrated circuits. The processor 102 can be configured to perform functions related to the techniques described in the present disclosure. In some embodiments, the processor 102 can also include multiple sub-processors integrated as a single logical component, such as Figure 1 As shown, the processor 102 can include a sub-processor a 102a, a sub-processor b 102b, and a sub-processor c 102c, etc.
[0023] The memory 104 can be configured to store data (e.g., instruction sets, computer code, intermediate data, etc.). For example, as shown, the stored data can include program instructions (e.g., program instructions for implementing the technical solutions of the present disclosure) and data to be processed. The processor 102 can also access the stored program instructions and data, and execute the program instructions to operate on the data to be processed. The memory 104 can include volatile storage or non-volatile storage. In some embodiments, the memory 104 can include random access memory (RAM), read only memory (ROM), optical disk, magnetic disk, hard disk, solid state disk (SSD), flash memory, memory stick, etc. Figure 1 The network interface 106 can be configured to provide communication with other external devices to the computer device 100 via a network. The network can be any wired or wireless network capable of transmitting and receiving data. For example, the network can be a wired network, a local wireless network (e.g., Bluetooth, WiFi, near field communication (NFC), etc.), a cellular network, the Internet, or a combination thereof. It can be understood that the type of network is not limited to the specific examples described above. In some embodiments, the network interface 106 can include any combination of any number of network interface controllers (NICs), radio frequency modules, transceivers, modems, routers, gateways, adapters, cellular network chips, etc.
[0024]
[0025] The peripheral interface 108 can be configured to connect the computer device 100 with one or more peripheral devices to realize information input and output. For example, the peripheral devices can include input devices such as a keyboard, a mouse, a touchpad, a touch screen, a microphone, various sensors, and the like, and output devices such as a display, a speaker, a vibrator, an indicator light, and the like.
[0026] The bus 110 can be configured to transmit information between various components (e.g., the processor 102, the memory 104, the network interface 106, and the peripheral interface 108) of the computer device 100, such as an internal bus (e.g., a processor-memory bus), an external bus (a USB port, a PCI-E bus), and the like.
[0027] It should be noted that although the above device only shows the processor 102, the memory 104, the network interface 106, the peripheral interface 108, and the bus 110, in the specific implementation process, the device can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device can also only contain components necessary for implementing the embodiments of the present disclosure, and does not necessarily contain all the components shown in the figure.
[0028] Figure 2 A basic structure diagram of the EDA tool 200 according to an embodiment of the present disclosure is shown.
[0029] As shown in Figure 2 , the part above the dashed line is the user part; the part below the dashed line is the EDA tool 200, which can be implemented by the device 100 shown in Figure 1 . In some embodiments, the EDA tool 200 can be implemented as EDA software. More specifically, the EDA tool 200 can be software for layout (Placement) and routing (Routing) based on chip design. The EDA tool 200 can include a Tcl command module 204 (or can be a graphical / window interface module), various computing modules (e.g., a Place computing module 206, a Route computing module 208, an Optimization computing module 210, and the like), and a database system 212. The user 202 can operate the EDA tool 200 by inputting relevant commands in the Tcl command module 204 (or can be a graphical / window interface module).
[0030] The Tcl command module 204 mainly functions as message passing or command passing. The Tcl command module 204 can read the instructions input by the user 202 to the simulation tool 200, and can distribute and pass to the corresponding computing module according to the specific content of the instructions to perform specific tasks.
[0031] According to the different computing tasks, the computing modules can be divided into, for example, a Place computing module 206, a Route computing module 208, an Optimization computing module 210, and the like. The Place computing module 206 can be used to calculate a reasonable placement position for all components, the Route computing module 208 can be used to calculate a reasonable wire connection mode between components, and the Optimization computing module 210 can be used to optimize the placement position and the wire connection mode between components. The computing processes of these computing modules can be performed in, for example, the processor 102. Figure 1
[0032] The database system 212 can be used to record and store all information (such as position, direction, size, structure, wire connection mode, and the like) of the simulated or designed chip completely and comprehensively. These information can be stored in, for example, the memory 104. Figure 1
[0033] Figure 3 A basic execution flow 300 of one computing command of the EDA tool 200 according to an embodiment of the present disclosure is shown. As shown in Figure 3 step 302, the user 202 can issue a command (for example, a do_place command) to the EDA tool 200 through a command interface or a graphical user interface (GUI) provided by the Tcl command module 204. Then, in step 304, the Tcl command module 204 parses the command and distributes it to the corresponding computing module (for example, the Place computing module 206). In step 306, the computing modules perform the specific calculations required by each of them. During this period, as shown in step 308, the computing modules need to (frequently and repeatedly) call the data in the database system 212 for calculation. After the calculation is completed, as shown in step 310, the computing modules can write the calculation results to the database system 212 and return the calculation results to the Tcl command module 204. In step 312, the Tcl command module 204 returns the calculation results to the user 202 through the command interface or the graphical user interface (GUI), and the processing process of one computing command of the EDA tool 200 ends. In step 314, the user can evaluate according to the calculation results, and then determine the next plan.
[0034] As described in the background section, clock networks typically need to drive large-scale register flip-flops, with fan-out numbers often reaching tens of thousands or even millions. In some EDA tools, multi-level buffering, hierarchical topologies (such as H-trees, balanced trees, etc.) are commonly used to construct clock distribution networks to minimize clock latency and clock skew. Clock skew here refers to the difference in delay between clock signals arriving at different sink points. In high-performance chip design, except for specific requirements (useful skew), skew must be strictly controlled to ensure flip-flop synchronization and timing correctness.
[0035] In some embodiments, clock network optimization can also be performed using the Multi-Tap (multi-branch drive) method. The basic idea of the Multi-Tap method is to introduce local drive through multiple "Tap points" in the early stages of clock tree synthesis, allowing clock signals to enter different regions of the chip in parallel from multiple locations. This shortens local wiring length, reduces latency, and balances the clock load across different regions. This method can improve the significant latency of long-distance transmission from a single clock source and shorten the construction depth of local subtrees, reducing latency differences between different regions.
[0036] like Figure 4a , Figure 4b As shown, where Figure 4a This represents a traditional or basic network structure without clock optimization. The timing relationships of each component are clearly defined in this structure. However, in this structure, the clock signal drives all registers and macrocells through a single tap driver, such as... Figure 4a In this network, clock signals for all components are transmitted through the "Tap point" in the upper left corner. Due to significant differences in the distance between different components and the Tap point, the clock signal arrival delay varies considerably between components, resulting in a large clock skew. Then, as... Figure 4b This is a schematic diagram illustrating the optimization of the Multi-Tap method. The Multi-Tap method introduces multiple parallel tap points and constructs an upper-level clock tree using methods such as H-trees, ensuring that clock signals arrive at the tap points synchronously. Then, by driving local cells through different tap points, it effectively reduces the difference in clock signal transmission path length and significantly reduces clock skew. However, from... Figure 4b As can be seen from this, although the Multi-Tap method can effectively optimize clock performance, its application also comes with significant costs. Figure 4a and Figure 4b Because of timing, the two components at the bottom of the diagram need to pass through an intermediate "clock gate" before receiving a signal. Initially, there is only one clock gate, and thus...Figure 4b In the prior art, a "clock gate" needs to be cloned to meet the timing and element requirements. Therefore, when the Multi-Tap method is applied to an actual chip, a large number of cloned elements are formed. For example, after the Multi-Tap method is applied to an area originally having only 100 elements, the number of elements may reach 150 or even 200. The main problems caused by this include: (1) area and power consumption overhead: Multi-Tap needs to copy the original clock tree trunk and distribute it to different Tap points to share the load. This process introduces a large number of additional buffers and clock gate units, resulting in significant increases in network power consumption and area, which is particularly problematic in high-frequency and high-performance designs. (2) Load balancing problem: Due to differences in chip floor planning and standard cell placement results, unreasonable Tap location planning may cause local area drive capacity redundancy, while other areas still have problems of insufficient drive and large delay bias, thereby forming a timing bottleneck. (3) Risk of wiring congestion: Each Tap point needs to obtain a signal from the global clock tree and construct a local clock sub-tree. Additional global wiring resources and more local fan-out networks are required, which may cause congestion on the critical wiring level.
[0037] In view of this, the present disclosure provides a clock network optimization method. When setting a clock network, the present disclosure determines a signal node and a receiving endpoint to be connected, performs initial connection, detects the connection effect according to the set requirement target, adjusts the connection based on the current established connection if the connection effect does not meet the target, and adjusts the number of elements in the clock network by switching the signal node connected to the receiving endpoint and changing the corresponding parameters and the cloned devices in the connection during the adjustment process. When the number of elements is reduced to meet the requirements, the number of cloned elements in the optimization process is reduced, thereby reducing the area and power consumption overhead, and the small number of cloned elements further solves the problems of load balancing and wiring congestion, thereby improving the chip design effect.
[0038] Figure 5 A flowchart of an exemplary method 500 provided by an embodiment of the present disclosure is shown. The method 500 can be implemented by a computer device 100 and can be implemented as part of a function of an EDA tool 200. As shown in Figure 1 , the method 500 can further include the following steps. Figure 2 Figure 5
[0039] Step 502: Obtain a chip design layout, and determine a signal node set and a receiving endpoint set according to the chip design layout.
[0040] Generally, the chip design layout contains each layer structure of the chip processing, specific transistor layout, wiring, routing, channel via layer connection position, etc. According to the chip design layout, the chip processing service provider can directly run and carry out the batch processing production of the chip. Further, the chip design layout itself is drawn step by step, accompanied by various optimizations, and finally the entire chip design layout is completed. At the initial stage of the chip design layout, it can only indicate the functional information of the level, such as a layer is a routing layer, a layer is an insulating layer, etc.; or only sets up various functional elements, indicating the position, size, etc. of each functional element. Then use EDA tools and the like to design and optimize step by step, and finally form a complete version of the chip design layout. Here, the element (inst) can be a standard processing unit, module or hard core, input / output terminal, register, etc. in the chip.
[0041] In this step, since the embodiment is the optimization adjustment of the clock network, the positions of various elements, the initial timing connection relationship between elements, etc. can be indicated in the chip design layout. For the network optimized by the Multi-Tap method, a clock tree is generally constructed to represent the timing relationship, which can include: Tap point: a local driving signal element of the EDA tool branched from the main clock source or intermediate element. The Tap point can be an input port of the chip, or a buffer driver, or other standard elements (such as integrated clock gating ICG). Sink (signal terminal): the final receiving end of the clock signal, which is usually a register, a latch or other clock triggered functional module. Trunk Node (trunk node): devices in the clock tree from Tap point to Sink before, which are cloned in the Tap distribution process. Leaf Net (leaf node network): network in the clock tree finally connected to the Sink clock load. Trunk Net (trunk network): network in the clock tree other than Leaf Net. In this embodiment, the Tap point is referred to as a signal node, the Sink point is referred to as a receiving end point, and a set containing multiple Tap points is referred to as a signal node set, and a set containing multiple Sink points is referred to as a receiving end point set.
[0042] In a specific embodiment, in a Leaf Net of the clock tree, at least one Sink point is generally included, and the Sink points in the set of receiving end points can be further divided by the Leaf Net to form at least one group of receiving end points. When optimizing the clock network, the above information is generally set and can be obtained by reading the chip design layout. In a specific application scenario, the initial data information such as the connection relationship of elements and the position of elements can be obtained through the DEF (Design Exchange File, physical information of design library) file and / or LEF (Library Exchange File, physical information of process library) file of the chip design layout.
[0043] Step 504, establishing the connection between the set of signal nodes and the set of receiving end points to generate an initial connection result.
[0044] In this step, after the set of signal nodes and the set of receiving end points are determined, the initial connection between the signal nodes and the receiving end points can be established according to the two sets. When connecting, generally, one receiving end point is only connected to one signal node, but one signal node can be connected to more than one receiving end point. In this way, the initial connection can be selected according to the position of each receiving end point to connect the nearest signal node, or connected according to the set connection relationship according to the relevant setting, or the receiving end points of the same group (Leaf Net) are connected to a signal node, and so on, to complete the initial connection and form an initial connection result.
[0045] In some embodiments, in order to minimize the factors such as delay, line length in layout, the initial connection can be performed by selecting the nearest distance. The distance here can be the Manhattan distance between the signal nodes and the receiving end points. The Manhattan distance represents the sum of the horizontal and vertical distances between two points in a two-dimensional (or multi-dimensional) coordinate system. For example, the Manhattan distance between point A and point B is . That is, in some embodiments, the establishing the connection between the set of signal nodes and the set of receiving end points includes: for any receiving end point in the set of receiving end points, determining the nearest signal node in the set of signal nodes, and establishing the connection between the any receiving end point and the signal node.
[0046] In different initial connection modes, for a group (Leaf Net) of receiving endpoints, it is inevitable to be connected to different signal nodes, so it is necessary to clone and copy other elements on the connection line between the two points, thereby causing the number of optimized elements to be more than the original number of elements after completing the initial connection. As shown in Figure 6a and Figure 6b , in a specific application scenario, the chip design layout can read design netlist, physical information, timing constraints and other information, and then construct the data structure of the clock tree. Then further, the signal node set and the receiving endpoint set can be constructed, as shown in Figure 6a , wherein other elements between the two are omitted. For the receiving endpoint set, it can also be grouped according to the Leaf Net. Then, as shown in Figure 6b , the initial allocation is performed, and all receiving endpoints are allocated to the signal nodes closest to their Manhattan distance. As shown in Figure 6b , it can be seen that Leaf Net 1 is divided into two parts, and Leaf Net 2 is divided into three parts, and the elements upstream of them need to be copied accordingly.
[0047] Step 506, in response to the initial connection result not meeting the set target, adjusting the connection of the initial connection result, determining whether to accept the adjustment through a set strategy, and generating an intermediate connection result.
[0048] In this step, first, the target is set for evaluating the specific effect of the current clock network, which may be different in different application scenarios, and can be set according to the specific application scenario.
[0049] In some application scenarios, the clock offset, power consumption, and the number of elements contained are set to the corresponding target. Specifically, assuming that the current signal node set Tap point set: , represents all pre-arranged M Tap elements available for selection; the receiving endpoint set Sink point set: , represents N Sink points downstream of the Tap points that need to be allocated. Then, a set of binary decision variables can be used to formalize the allocation operation, wherein represents that the Sink point is allocated to the Tap point , represents no connection. Usually, a Sink point can only be allocated to one Tap point, that is .
[0050] Then, for the clock offset, it is expected to minimize the clock offset in this scenario, assuming that the clock offset , wherein represents clock arrival time at the Sink point). Generally, the Tap points are evenly distributed in the upstream clock tree, and the design goal is to have the same delay from the clock source to all Tap points, so all Tap points generally have the same clock arrival time. The clock arrival time from a Tap point to a Sink point is proportional to the wire length. Therefore, minimizing clock skew is equivalent to minimizing the distance difference between all Sink points and their assigned Tap points. The distance difference can be measured by the bounding box of the coordinates of the Sink points driven by a Tap point, which can be simplified to the clock skew distance where represents the Tap point assigned to all Sink points.
[0051] After that, for power consumption, in this scenario, it is also desired to minimize the power consumption, which mainly includes dynamic switching power and static power, where the dynamic switching power is proportional to the total capacitance, i.e. For , it mainly includes the load capacitance and the wire capacitance , where the input capacitance of all Tap points and the input capacitance of all Sink points are fixed, i.e. the load capacitance is fixed, so minimizing is equivalent to minimizing , and the optimization goal directly translates to minimizing the topology distance of the clock tree (i.e. the sum of the Manhattan distances from all Sink points to their assigned Tap points, denoted as ). On the other hand, for the static power, it mainly comes from the leakage current of transistors, which is proportional to the number of devices, so reducing the number of buffers and gated cells can directly reduce the overall static power, i.e. minimizing the number of elements can also achieve the purpose of minimizing the static power.
[0052] Finally, for the limitation of the number of elements, when the Multi-Tap method is used for clock performance optimization in the foregoing embodiments, the allocation of Tap points means that multiple cloning of the original clock tree is needed. For example, if the receiving end points Sink points of a certain group (Leaf Net) are allocated to 4 different Tap points, then the clock trunk device upstream of the Net needs to be copied at least 3 times, which significantly increases the power consumption and area of the clock tree. Obviously, the number of Leaf Nets is positively correlated with the number of device cloning, so by controlling the number of receiving end points Sink in each group of receiving end points (Leaf Net) allocated, the scale of Tap allocation and the number of device replication such as clock gating can be controlled. Thus, under the premise of achieving the target Skew, the degree of replication of the clock tree is reasonably limited to achieve a controllable balance between power consumption, area and performance. Specifically, assuming that the number of elements before optimization is , the number of elements obtained by each optimization iteration is , and the target number of elements is set to , wherein . In this way, the target can be set.
[0053] Finally, by comprehensively considering the above targets, it can be determined that the set target is a constrained mixed integer nonlinear programming problem (MINLP), and the objective function can be expressed as follows, wherein represents the wire score, and are weight coefficients, which can be adjusted as needed.
[0054]
[0055] That is, in some embodiments, the set target at least includes target requirements for the wire score and the number of elements; wherein the wire score corresponds to the topology distance converted from the clock skew distance and the wire capacitance, and the number of elements is used to count the number of elements on the wire when the signal node set and the receiving end point set are connected.
[0056] After obtaining the initial connection result in step 504, it can be determined whether it meets the set target by a target function similar to the foregoing. If the set target is directly met, it can be considered that the initial connection result directly meets the optimization requirement, and then it can be directly considered that the optimization is completed, and the initial connection result is output as the final optimization result. If the set target is not met, the connection mode can be adjusted on the basis of the initial connection result to form a new connection result, and it is determined again whether the set target is met. If it is not met, the adjustment is continued, and the iteration is performed in this way. In this process, whether to accept the adjusted connection scheme can be determined by using various set strategies, for example, whether to accept the adjustment can be determined by using a heuristic optimization algorithm such as a simulated annealing strategy, a greedy algorithm, etc. After that, whether to accept or not to accept, the connection result output by the set strategy is the intermediate connection result, that is, if the adjustment is accepted, the adjusted connection scheme is the intermediate connection result of this time, and if the adjustment is not accepted, the connection scheme before the adjustment is the intermediate connection result of this time.
[0057] Various adjustment strategies can be employed to modify the connection method. For example, at least one receiving endpoint (Sink point) can be arbitrarily selected from the current connection result and reassigned from its original connected signal node (Tap point) to another Tap point; or at least two Sink points can be arbitrarily selected and their connected Tap points can be swapped; or for at least one group (Leaf Net) of receiving endpoints, the Sink points originally connected to the same Tap point can be adjusted to connect to another Tap point, and so on. In other embodiments, adjustments can also be made to Tap points, such as replacing the Sink point connected to a certain Tap point, or swapping two Tap points, etc. It can be seen that any scheme that provides a new connection result can be considered as a connection adjustment. Of course, when making adjustments, only one of the above methods can be executed, multiple methods can be used alternately during iteration, and multiple adjustments can be performed in a single iteration. Therefore, when making adjustments, the specific adjustment method for the connection scheme can be determined according to the specific scenario. Then, considering factors such as convergence speed, a scheme with fewer changes each time can be selected for adjustment. That is, in some embodiments, the connection adjustment of the initial connection result includes: randomly selecting at least one receiving endpoint and replacing the signal node connected to the at least one receiving endpoint; or randomly selecting at least two receiving endpoints and exchanging the signal nodes connected to the at least two receiving endpoints; or determining at least one group of receiving endpoints in the set of receiving endpoints and adjusting the receiving endpoints connected to the same signal node in the at least one group of receiving endpoints to connect to another signal node; wherein the grouping of the at least one group of receiving endpoints is determined according to the chip design layout. Here, "replacement" means replacing the signal node connected to the receiving endpoint with another one; "exchange" means swapping the signal nodes connected to two receiving endpoints; in a specific embodiment, a set of receiving endpoints refers to all receiving endpoints under a Leaf Net. Among these receiving endpoints, at least one receiving endpoint connected to the same signal node is selected, and its connected signal node is adjusted to another one. Of course, if there are multiple sets of receiving endpoints connected to the same signal node under a Leaf Net, one set can be selected for adjustment, or multiple sets can be selected.
[0058] In specific scenarios, through the above adjustments, each group of receiving endpoints within the original Leaf Net will be re-divided based on the existing Leaf Net, combined with... Figure 6a and Figure 6b, the original Leaf Net 1 corresponds to four receiving endpoints A, B, C and D, after initial connection, receiving endpoints A, B and C are connected to the signal node Tap0 (which can be called Root0), and receiving endpoint D is connected to the road signal node Tap1. Since the original group of receiving endpoints corresponding to Leaf Net 1 is divided into two groups, the intermediate elements required by the original receiving endpoint D need to be cloned. In the subsequent adjustment, if receiving endpoint C is connected to Tap2, the number of divided groups becomes three, and further cloning of the intermediate elements required by the original receiving endpoint C is needed. Compared with the two groups after the initial connection adjustment, the number of elements increases due to the increase in the number of groups after the adjustment. Further, in the iteration process, the number of elements will increase or decrease as the adjustment is adjusted in a similar manner. In combination with the foregoing, the initial leaf node network (Leaf Net) division is set by the chip design layout, and the leaf node network can be read from the chip design layout. That is, in some embodiments, the connection adjustment of the initial connection result includes: determining at least one group of receiving endpoints based on the chip design layout, the at least one group of receiving endpoints being divided based on the leaf node network of the chip design layout; and adjusting the number of elements based on the adjustment of the number of groups of the at least one group of receiving endpoints in the adjustment process.
[0059] After the adjustment is completed, it is necessary to determine whether to accept the adjustment. Here, an adjustment acceptance condition can be set for judgment. If the condition is met, the adjustment is accepted. If the condition is not met, the adjustment is not accepted, and the adjustment is continued or restarted. Here, the set condition can be based on the foregoing, and the wire score and the number of elements are used as the judgment standard. For the wire score, the current adjusted wire score can be set to be less than the wire score before adjustment or less than the target wire score. For the number of elements, the current adjusted number of elements can be set to be less than the target number of elements or less than the number of elements before adjustment. When the current adjustment meets the requirements of the wire score and the number of elements, it is considered that the adjusted connection scheme is better, and the adjustment is directly accepted. That is, in some embodiments, the determination of whether to accept the adjustment by setting the strategy includes: determining whether to accept the adjustment based on the comparison result of the wire score and the number of elements before and after the adjustment.
[0060] On this basis, the present embodiment further selects an annealing strategy from a plurality of set strategies to determine whether to accept the adjustment. It also uses the wire score and the number of elements as the judgment standard. As in the foregoing embodiment, when the wire score and the number of elements both meet the conditions, it is considered that the adjusted scheme is better, and the adjustment is directly accepted. However, unlike the foregoing embodiment, if the current adjustment does not meet the foregoing judgment conditions, i.e., the wire score does not meet the condition or the number of elements does not meet the condition, the probability is determined according to the annealing strategy to determine whether to accept the adjustment scheme, wherein For the current temperature, the annealing strategy sets a starting temperature at the beginning, and then gradually decreases the temperature with the iteration of the loop, and is the adjusted connection score after this adjustment and the connection score before the adjustment. It can be seen that with the increase of the number of iterations, the temperature will be lower and lower, and the probability of accepting a solution that does not meet the conditions will be lower and lower. That is, in some embodiments, the determination of whether to accept the adjustment includes: in response to the adjusted connection score being less than the connection score before the adjustment, and the number of elements after the adjustment being less than the number of elements before the adjustment or the number of elements after the adjustment being less than a set target number, accepting the adjustment; otherwise, determining the probability of accepting the adjustment based on the adjusted connection score, the connection score before the adjustment, and the current temperature set by the set strategy, and determining whether to accept the adjustment according to the probability.
[0061] Step 508, in response to the intermediate connection result meeting the set target, determining the intermediate connection result as the optimization result.
[0062] In this step, after obtaining the intermediate connection result, it is re-judged whether the intermediate connection result meets the set target. If the set target is met, the current intermediate connection result is output as the final optimization result.
[0063] If the set target is not met, the aforementioned connection adjustment can be repeated, the intermediate connection result of this time is taken as input, adjusted, and it is judged by the set strategy whether to accept the adjustment, and then a new intermediate connection result is output and formed. After that, it is continuously judged whether the intermediate connection result meets the set target, and the iteration is repeated in this way, so that the intermediate connection result can be continuously updated. That is, in some embodiments, after the intermediate connection result is generated, the method further includes: in response to the intermediate connection result not meeting the set target, performing the connection adjustment on the intermediate connection result, determining whether to accept the adjustment by the set strategy, and iterating the intermediate connection result.
[0064] Afterwards, in the process of cyclic iteration of the intermediate connection result, if the set target cannot be met all the time, in order to save resources, a stop condition can be set for iteration. Then after obtaining the intermediate connection result each time, before judging whether the set target is met according to the intermediate connection result, a judgment can be made whether the stop condition is met, if the stop condition is met, the judgment of the set target can be skipped, and the current intermediate connection result can be directly output as the optimization result. In some embodiments, the stop condition that can be set includes: (1) in continuous multiple iterations, the connection score and the element quantity have no obvious change (i.e. in continuous set number of iterations, the change range of the connection score and the element quantity is within the set threshold range); (2) the current temperature set in the set strategy (such as annealing strategy, etc.) is reduced to the set threshold temperature; (3) the number of iterations reaches the set threshold iteration number, etc. That is, in some embodiments, the method further includes: in response to the change of the connection score and the element quantity within the set range in the continuous set number of iterations, or in response to the current temperature set in the set strategy being reduced to the set temperature, or in response to the set iteration number being reached, determining the current intermediate connection result as the optimization result.
[0065] Finally, the optimization result formed can be output, and the optimization result can be output to the downstream EDA process for further processing, or the chip design layout can be adjusted directly according to the optimization result. Of course, the output mode of the optimization result can also include multiple modes, for example, the output can be displayed on the corresponding device to give the operator corresponding feedback. Of course, in other embodiments, the output mode of the optimization result can not be limited to output display, and the optimization result can also be used to store, display, use or further process the optimization result. According to different application scenarios and implementation needs, the output mode of the optimization result can be flexibly selected.
[0066] Specifically, for example, for the application scenario of the method of the present embodiment executed on a single device, the optimization result can be directly output in the form of display on the display component (display, projector, etc.) of the current device, so that the operator of the current device can directly see the content of the optimization result from the display component.
[0067] For example, for the application scenario that the method of the embodiment is executed on a system composed of multiple devices, the optimization result can be sent to other preset devices in the system as a receiving end, i.e., a synchronization terminal, through any data communication mode (wired connection, NFC, Bluetooth, wifi, cellular mobile network, etc.), so that the synchronization terminal can perform subsequent processing. Optionally, the synchronization terminal can be a preset server, which is generally set in the cloud as a data processing and storage center, and can store and distribute the optimization result; the receiving end of the distribution is a terminal device, and the holder or operator of the terminal device can be a management personnel, a designer, a maintenance personnel, etc. of each link of chip design.
[0068] For example, for the application scenario that the method of the embodiment is executed on a system composed of multiple devices, the optimization result can be sent to other preset devices in the system as a receiving end, i.e., a synchronization terminal, through any data communication mode (wired connection, NFC, Bluetooth, wifi, cellular mobile network, etc.), so that the synchronization terminal can perform subsequent processing. Optionally, the synchronization terminal can be a preset server, which is generally set in the cloud as a data processing and storage center, and can store and distribute the optimization result; the receiving end of the distribution is a terminal device, and the holder or operator of the terminal device can be a management personnel, a designer, a maintenance personnel, etc. of each link of chip design.
[0069] As can be seen from the above, the embodiment of the disclosure provides a clock network optimization method. The method comprises: obtaining a chip design layout, determining a signal node set and a receiving endpoint set according to the chip design layout; establishing a connection between the signal node set and the receiving endpoint set, generating an initial connection result; in response to the initial connection result not meeting a set target, adjusting the connection of the initial connection result, determining whether to accept the adjustment through an annealing strategy, generating an intermediate connection result; and in response to the intermediate connection result meeting the set target, determining the intermediate connection result as an optimization result. When the clock network is set, after the signal node and the receiving endpoint to be connected are determined, the initial connection is performed first, and then the connection effect is detected according to the set requirement target. If the target is not met, the adjustment is performed on the basis of the currently established connection. In the adjustment process, with the switching of the signal node to which the receiving endpoint is connected, the corresponding parameters change, and at the same time, the cloned devices in the connection also increase or decrease, so as to adjust the number of elements in the clock network. When the number of elements is reduced to meet the requirement, the number of cloned elements in the original optimization process is reduced, thereby reducing the area and power consumption overhead, and the fewer cloned element quantity further solves the problems of load balancing and wiring congestion, and improves the chip design effect.
[0070] It should be noted that the method of the embodiment of the disclosure can be executed by a single device, such as a computer or a server. The method of the embodiment of the disclosure can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this distributed scenario, one of the multiple devices can only execute one or more steps in the method of the embodiment of the disclosure, and the multiple devices can interact with each other to complete the method.
[0071] It is to be understood that the above description is based on certain specific embodiments of the disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order and still achieve desirable results. Additionally, the processes depicted in the figures do not necessarily require the particular order shown or sequential order to achieve desirable results. In certain implementations, multitasking and parallel processing can be advantageous.
[0072] Based on the same inventive concept, the disclosure also provides a non-transitory computer readable storage medium containing a computer program, which stores computer instructions for causing the computer to perform the method 500 according to any of the above embodiments.
[0073] The computer readable storage medium of the embodiments can include permanent and non-permanent, removable and non-removable media, which can be implemented by any method or technology to store information. The information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette, magnetic tape / disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible to a computing device.
[0074] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to perform the method 500 according to any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which are not repeated here.
[0075] Based on the same inventive concept, the disclosure also provides a computer program product, which includes a computer program. In some embodiments, the computer program is executable by one or more processors to cause the processors to perform the method 500. Corresponding to the execution subject of each step in each embodiment of the method 500, the processor performing the corresponding step can belong to the corresponding execution subject.
[0076] The computer program product of the above embodiments is configured to enable a processor to perform the method 500 as described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which are not repeated here.
[0077] It should be understood by those of ordinary skill in the art that the above discussion of any of the embodiments is merely exemplary and is not intended to suggest that the scope of the present disclosure is in any way limited to these examples; the embodiments or technical features among different embodiments can be combined, the steps can be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of brevity. It should be understood that the above discussion is merely exemplary and is not intended to suggest that the scope of the present disclosure is in any way limited to these examples.
[0078] In addition, in order to simplify the description and discussion, and so as not to make the embodiments of the present disclosure difficult to understand, the well-known power / ground connections of integrated circuit (IC) chips and other components can or can not be shown in the provided drawings. Furthermore, the apparatuses can be shown in the form of block diagrams in order to avoid making the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform to be implemented, i.e., these details should be fully within the understanding of those skilled in the art. Where specific details (e.g., circuitry) are set forth in order to describe an illustrative embodiment of the present disclosure, it should be apparent to those skilled in the art that the present disclosure can be practiced without such specific details or with an alternative and / or equivalent technique.
[0079] Although the present disclosure has been described in conjunction with the specific embodiments thereof, it is evident that many alternatives, modifications and variations will be apparent to those skilled in the art in light of the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) can use the embodiments discussed.
[0080] The embodiments of the present disclosure are intended to cover all such alternatives, modifications and variations as falling within the broad scope of the above described embodiments. Accordingly, any and all such modifications, variations or equivalents that fall within the spirit and scope of the embodiments of the present disclosure are intended to be included within the scope of the present disclosure.
Claims
1. A clock network optimization method, characterized in that, include: Obtain the chip design layout, and determine the set of signal nodes and the set of receiving endpoints based on the chip design layout; Establish a connection between the set of signal nodes and the set of receiving endpoints, and generate an initial connection result; In response to the initial connection result not meeting the set target, the initial connection result is adjusted, and whether to accept the adjustment is determined by a set strategy to generate an intermediate connection result. In response to the intermediate connection result meeting the set target, the intermediate connection result is determined as the optimized result.
2. The method according to claim 1, characterized in that, Establishing the connection between the signal node set and the receiving endpoint set includes: For any receiving endpoint in the set of receiving endpoints, determine the nearest signal node in the set of signal nodes, and establish a connection between the receiving endpoint and the signal node.
3. The method according to claim 1, characterized in that, The set target includes at least target requirements for connection score and number of components; wherein, the connection score corresponds to the topology distance obtained by converting the clock offset distance and the connection capacitance, and the number of components is used to count the number of components on the connection when the signal node set is connected to the receiving endpoint set.
4. The method according to claim 3, characterized in that, The process of determining whether to accept adjustments by setting a strategy includes: Based on the comparison results of the connection scores and the number of components before and after the adjustment, it is determined whether to accept the adjustment.
5. The method according to claim 4, characterized in that, The process of determining whether to accept the adjustment includes: If the adjusted connection score is less than the original connection score, and the adjusted number of components is less than the original number of components, or the adjusted number of components is less than the set target number, then the adjustment is accepted. Conversely, based on the adjusted connection score, the connection score before adjustment, and the current temperature set by the set strategy, the probability of accepting the adjustment is determined, and whether to accept the adjustment is determined according to the probability.
6. The method according to claim 3, characterized in that, The connection adjustment of the initial connection result includes: At least one set of receiving endpoints is determined based on the chip design layout, and the at least one set of receiving endpoints is partitioned based on the leaf node network of the chip design layout; The number of components is adjusted based on the adjustment of the number of groups of the at least one set of receiving endpoints during the adjustment process.
7. The method according to claim 1, characterized in that, The connection adjustment of the initial connection result includes: Randomly select at least one receiving endpoint and replace the signal node connected to the at least one receiving endpoint; or Randomly select at least two receiving endpoints and exchange the signal nodes connected to the at least two receiving endpoints; or At least one set of receiving endpoints is determined in the set of receiving endpoints, and receiving endpoints in the at least one set of receiving endpoints that are connected to the same signal node are adjusted to connect to another signal node; wherein the grouping of the at least one set of receiving endpoints is determined according to the chip design layout.
8. The method according to claim 1, characterized in that, After generating the intermediate connection results, the method further includes: In response to the intermediate connection result not meeting the set target, the connection result is adjusted, and the set strategy determines whether to accept the adjustment, so as to iterate the intermediate connection result.
9. The method according to claim 8, characterized in that, Before the intermediate connection result meets the set target, the method further includes: In response to the fact that the changes in the connection score and the number of components are within a set range during a set number of iterations, or in response to the current temperature set by the set strategy decreasing to a set temperature, or in response to reaching a set number of iterations, the current intermediate connection result is determined as the optimized result.
10. A computer device, characterized in that, It includes one or more processors, memory; and one or more programs, wherein the one or more programs are stored in the memory and executed by the one or more processors, the programs including instructions for performing the method according to any one of claims 1 to 9.
11. A non-volatile computer-readable storage medium containing a computer program, characterized in that, When the computer program is executed by one or more processors, the processors perform the method of any one of claims 1 to 9.
Citation Information
Patent Citations
Clock node clustering method and clock network structure
CN103150435A
A backbone network of chip multi-source clock tree
CN109976503A
Method and system for optimizing unbalanced clock network of integrated circuit
CN111859836A
Clock tree synthesis and layout hybrid optimization method and device, storage medium and terminal
CN113807043A
Clock tree optimization method, optimization device and related equipment
CN114997087A