Clock timing in replicated arrays
By employing a block group timing model and worst-case analysis, the method addresses the challenge of simulating clock timing in large node arrays, enhancing simulation efficiency and accuracy.
Patent Information
- Application Number
- JP2025509090
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-08-19
- Filing Date
- 2023-08-16
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2043-08-16
AI Technical Summary
Accurately simulating and optimizing clock signal timing in large, complex node arrays is challenging due to the exponential increase in netlist size and clock phase variations, which can lead to inaccuracies and prolonged simulation times using existing electronic design automation (EDA) tools.
A method and system for simulating clock signal timing in node arrays by modeling a block group timing model, focusing on a subset of nodes with consistent clock distribution circuitry, and adjusting the clock distribution network based on worst-case timing analysis, reducing the need for full array simulation.
This approach significantly reduces simulation time and improves accuracy by modeling a majority of the node array using block group timing, allowing for efficient optimization of clock distribution networks in large node arrays.
Smart Images

Figure 2025529819000001_ABST
Abstract
Description
[Technical Field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 371,993, entitled "CLOCK TIMING IN REPLICATED ARRAYS," filed August 19, 2022, the disclosure of which is incorporated herein by reference in its entirety for all purposes.
[0002] FIELD OF THE DISCLOSURE This disclosure relates generally to distributed clocking, and more particularly to techniques for modeling clocking in an array. [Background technology]
[0003] A high-density processing system can be built using an array of processing nodes. The nodes can communicate with neighboring nodes to perform processing tasks. Communication between the nodes can be synchronous and / or asynchronous. A clock signal can be provided to each node to synchronize the nodes, which can enable communication between the nodes. Summary of the Invention [Means for solving the problem]
[0004] Each claimed innovation has several aspects, no single one of which is solely responsible for its desirable attributes. Without limiting the scope of the claims, some prominent features of this disclosure will now be discussed briefly.
[0005] One aspect of the present disclosure is a method of simulating a node array. The method includes accessing a timing model of computational nodes of the node array stored in non-transitory computer-readable memory and simulating, using one or more computers, clock signal timing of a majority of the nodes of the node array using the timing model of the computational nodes. The timing model of the computational node represents timing data related to clock signal propagation between the computational node and four neighboring nodes of the node array, each abutting the computational node.
[0006] The method may also include determining worst-case timing of clock signals in the node array based on the simulating.
[0007] The method may also include adjusting a clock distribution network of the node array based on the worst-case timing. Additionally, adjusting the clock distribution network may include updating one or more files representing circuit configurations of the node array. Furthermore, the method may further include accessing a global nodal timing model of the node array stored in a non-transitory computer-readable memory. The determinization may be based on a simulation of clock signal timing using the global nodal timing model.
[0008] In the method, simulating may include simulating mesochronous clocking in the node array.
[0009] In this method, the timing model of the node array can model a computational node that receives a clock signal from a first pair of four neighboring nodes and a computational node that provides a clock signal to a second pair of four neighboring nodes. In addition, the clock signal can be delayed by one delay unit at the computational node relative to the first pair of neighboring nodes. Furthermore, the clock signal can be delayed by two delay units at the second pair of neighboring nodes relative to the first pair of neighboring nodes.
[0010] In this method, the timing model of the node array is a block group timing model. Furthermore, the node array can be essentially of instances of compute nodes and instances of global nodes.
[0011] In this method, the majority of the nodes of the node array may include at least 90% of the nodes of the node array.
[0012] The method may further include generating a timing model of the compute node by at least simulating clock signal propagation between the compute node and four neighboring nodes.
[0013] Another aspect of the present disclosure is non-transitory computer-readable storage including instructions that, when executed by one or more processors, cause the storage to perform a method for simulating a node array. The method includes accessing a timing model of computational nodes of the node array stored in non-transitory computer-readable memory and simulating, using one or more computers, clock signal timing of a majority of the nodes of the node array using the timing model of the computational nodes. Further, the timing model of the computational node represents timing data related to clock signal propagation between the computational node and four neighboring nodes of the node array, each abutting the computational node.
[0014] Another aspect of the present disclosure is a computer system for simulating a node array. The system includes a non-transitory computer-readable memory that stores a computational node timing model of a node, and one or more processors configured to execute instructions to access at least the computational node timing model and to simulate clock signal timing for a majority of the nodes of the node array using the computational node timing model. Further, the computational node timing model represents timing data related to clock signal propagation between the computational node and four neighboring nodes of the node array, each abutting the computational node.
[0015] Another aspect of the present disclosure is a system for simulating clock timing distribution across a node array including a plurality of computing nodes. The system may include one or more computing devices configured to store timing models corresponding to the computing node and four neighboring computing nodes to which each neighboring computing node abuts. Further, each computing device of the one or more computing devices is configured to access the timing model of the computing node and, using the computing device, simulate clock signal timing distribution for a majority of the nodes of the node array using the timing model of the computing node.
[0016] Another aspect of the present disclosure is a non-transitory computer-readable storage medium storing instructions for simulating clock timing distribution across a node array. When executed by a processor, the instructions cause the processor to perform operations including accessing a timing model of a computational node and, using a computing device, simulating clock signal timing distribution for a majority of the nodes of the node array using the timing model of the computational node. Further, the timing model of the computational node represents timing data related to clock signal propagation between the computational node and four neighboring nodes of the node array, each abutting the computational node.
[0017] In the non-transitory computer-readable storage medium, the operations may include determining worst-case timing of clock signals in the node array based on the simulating. Further, the non-transitory computer-readable storage medium may include adjusting a clock distribution network of the node array based on the worst-case timing. Adjusting the clock distribution network may include updating one or more files representing a circuit configuration of the node array.
[0018] In a non-transitory computer-readable storage medium, the operations may include accessing a global nodal timing model of the node array stored in non-transitory computer-readable memory, the determination being based on a simulation of clock signal timing using the global nodal timing model.
[0019] In the non-transitory computer-readable storage medium, simulating includes simulating mesochronous clocking in the node array.
[0020] In the non-transitory computer-readable storage medium, the timing model of the node array may model a computational node that receives a clock signal from a first pair of four neighboring nodes and a computational node that provides a clock signal to a second pair of four neighboring nodes. Additionally, the clock signal may be delayed by one delay unit at the computational node relative to the first pair of neighboring nodes. Furthermore, the clock signal may be delayed by two delay units at the second pair of neighboring nodes relative to the first pair of neighboring nodes.
[0021] In the non-transitory computer-readable storage medium, the timing model of the node array may be a block group timing model. Further, the node array may consist essentially of instances of computational nodes and instances of global nodes.
[0022] In the non-transitory computer-readable storage medium, the majority of the nodes of the node array may include at least 90% of the nodes of the node array.
[0023] In the non-transitory computer-readable storage medium, the operations may also include generating a timing model of the compute node by at least simulating clock signal propagation between the compute node and four neighboring nodes.
[0024] For purposes of summarizing the disclosure, certain aspects, advantages, and novel features of the innovations have been described herein. It should be understood that not all such advantages may necessarily be achieved in accordance with any particular embodiment. Thus, the innovations may be embodied or implemented to achieve or optimize one advantage or group of advantages as taught herein without necessarily achieving other advantages as may be taught or suggested herein. [Brief explanation of the drawings]
[0025] Embodiments of the present disclosure will now be described, by way of non-limiting examples, with reference to the accompanying drawings, in which:
[0026] [Figure 1] FIG. 1 is a schematic block diagram of an exemplary chip according to aspects of the present disclosure.
[0027] [Figure 2A] FIG. 1 is a schematic diagram of a clock distribution network according to one embodiment.
[0028] [Figure 2B] 2B illustrates an example implementation of clock distribution circuitry within an example node of the node array of FIG. 2A.
[0029] [Figure 2C] 2B illustrates another exemplary implementation of clock distribution circuitry within an exemplary node of the node array of FIG. 2A.
[0030] [Figure 3A] 2B is a node clock level map associated with an exemplary node array, such as the node array of FIG. 2A.
[0031] [Figure 3B] 3B is a node clock level topology corresponding to the node clock level map of FIG. 3A.
[0032] [Figure 4] FIG. 2 illustrates an example implementation of the node array of FIG. 1.
[0033] [Figure 5] FIG. 2 illustrates a characterization of the node array of FIG. 1 to model clock timing distribution across the node array.
[0034] [Figure 6] FIG. 6 is a diagram illustrating an example of a block of the node array of FIG. 5.
[0035] [Figure 7] FIG. 1 illustrates an exemplary embodiment of a computing device. DETAILED DESCRIPTION OF THE INVENTION
[0036] In the following detailed description of several embodiments, various descriptions of specific embodiments are presented. However, the innovations described herein can be embodied in many different ways, for example, as defined and encompassed by the claims. This description refers to the drawings, in which like reference numbers and / or terminology may indicate identical or functionally similar elements. It will be understood that the elements depicted in the figures are not necessarily drawn to scale. It will also be understood that some embodiments may include more elements and / or a subset of the elements depicted in the figures. Furthermore, some embodiments may incorporate any suitable combination of features from two or more figures. The headings provided herein are for convenience only and do not necessarily affect the scope or meaning of the claims. Overview of distributed clocking for node arrays
[0037] The present disclosure relates to clock distribution networks with clock signals that arrive at different times at various nodes in a node array. Generating clock signals with a fixed offset may be referred to as mesochronous clocking. Embodiments disclosed herein relate to mesochronous clock networks that are modularly built with common circuitry. Clock signals in such networks can be locally low-skew and mesochronous at coarser levels.
[0038] The principles and advantages disclosed herein can be applied to any suitable circuit chip. In particular applications, the clock signal distribution disclosed herein can be applied to a chip that includes an array of smaller computational nodes, each of which may be referred to as a processor or core. In this manner, the clock signal can form an arrival time wave across the array of computational nodes. In one embodiment, each computational node can receive a low-skew clock. The computational nodes of the array can be designed with only interfaces to neighboring computational nodes, taking into account the arrival time difference (skew) of the mesochronous clock phase. The techniques described herein can be applied to square (equal rows and columns) node arrays or rectangular node arrays with a different number of rows than columns.
[0039] FIG. 1 is a schematic block diagram of an exemplary chip 100 according to aspects of the present disclosure. The chip 100 may be an integrated circuit die. The chip 100 may include a node array 102 (also referred to as a compute node array) with distributed clocking, one or more serializer / deserializer (SerDes) clock blocks 104, a clock generator 106, and a clock controller 108. The SerDes clock block 104 may interface with other chips 100 to form an array of chips 100. In certain example applications, the node array 102 may be included in a chip 100 in a system-on-wafer system, an array of chips 100 on a printed circuit board, or the like. In certain applications, the node array 102 of FIG. 1 may be implemented in a system-on-wafer packaged in a wafer-level packaging structure. As shown in the embodiment of FIG. 1, the clock generator 106 may be implemented external to the node array 102. In some embodiments, the clock generator 106 may include a phase-locked loop (PLL). The clock generator 106 can be positioned to provide clock signals to the computational nodes at the corners of the node array 102. The clock controller 108 can also be implemented outside the node array 102. Nodes within the node array 102 can include inter-node interfaces that can be configured to communicate synchronously. The core-to-serializer / deserializer (SerDes) interface can be asynchronous. In some embodiments, the PLL operates at a 100 MHz reference frequency band, although other frequency bands are contemplated within the scope of this disclosure. The PLL can be configured to generate source clocks for various operating modes, including functional mode, bypass mode, and test mode. The PLL can be configured to operate without glitch source clock selection and to manage thermal issues and maximum current through clock throttling. In some examples, clock throttling, on-chip clock control (OGG) for scan capture, and clock ramp-up / down are implemented by cycle skipping to modulate the effective frequency, for example, using a first-in, first-out (FIFO) 32 x 32 pattern.
[0040] In the node array 102 with distributed clocking of FIG. 1, each node may be an instance of computing circuitry (also referred to as a processing core or computational node). In certain applications, most of the nodes may be implemented as instances of computing circuitry, and one or more of the nodes may be implemented as instances of different circuitry. Each node of the node array 102 may include an instance of substantially the same clock distribution circuitry, even if other circuitry of at least some of the nodes differs from that of the other nodes. For example, most of the nodes may be implemented as instances of computing circuitry, and one or more of the nodes may be implemented as instances of global nodes. The global nodes may include process, voltage, and temperature (PVT) sensors (not shown in FIG. 1). In the node array 102, the nodes may be tiled and abutted. For example, each node of the node array 102 may be self-contained and interconnected to adjacent nodes (e.g., abutting nodes). At the same time, the node array 102 may be implemented without using top-level wiring, gates, or channels. Thus, nodes can be configured to communicate with neighboring nodes using low-level wiring over relatively short connections. In some embodiments, the nodes of the node array 102 can be tiered without mirroring or rotation. In particular implementations, nodes can be configured to communicate with neighboring nodes using power lines (V DD / V SS ) grid pitch. For example, the height and width of each node may be a multiple of the power grid pitch. The power grid pitch may be further aligned to the bump pitch.
[0041] Each node in the node array 102 may include substantially identical instances of clock distribution circuitry. Nodes may be designed so that the node's output clock wiring is aligned with its neighboring node's input clock wiring. Nodes may be tiered and tiled within the node array so that the clock output wiring aligns with and electrically connects to the clock input wiring of a neighboring node located downstream to receive the clock signal. Such electrical connections allow the node array to be implemented without channels or top-level wiring for clock distribution. In certain embodiments, the fanout of the clock distribution circuitry may be balanced relative to inverters.
[0042] As described herein, a clock signal received at a root node may propagate from the root node to two neighboring nodes with one delay unit. The root node may be located at a corner of the node array 102. The delay unit may be a fixed offset for a given node array. The delay unit may correspond to the delay from buffering the clock signal (e.g., using an inverter) and the wiring delay associated with the clock signal propagating to that neighboring node. For example, in one embodiment, one of the two neighboring nodes is in the same row as the root node, and the other of the two neighboring nodes is in the same column as the root node. As an example, the neighboring nodes are positioned south and east of the root node. In this configuration, the clock signal continues to propagate from the two neighboring nodes in the node array to the south and east neighboring nodes, in this example, with one more delay unit. Such clock signal propagation continues through the clock distribution network until the clock signal reaches a node in the node array at the opposite corner from the root node. In some examples, signals routed from an originating node (e.g., a node located in the southeast corner of node array 102) to a node located north or west may travel upstream and lose one unit delay in the node array, while signals routed from an originating node (e.g., a node located in the southeast corner of node array 102) to a node located south or east may travel downstream and gain one unit delay in the node array. In some applications, signals traveling upstream may be routed faster than signals traveling downstream to account for the unit delay and meet setup and hold time specifications.
[0043] One of the two neighboring nodes may be located in the same row as the root node, and the other of the two neighboring nodes may be located in the same column as the root node. In some embodiments, the neighboring nodes abut the root node. As an example, the neighboring nodes are south and east of the root node as shown in FIG. 2A. For example, the neighboring nodes of node 206A may be nodes 206B and 206C. The clock signal continues to propagate from the two neighboring nodes of the root node in the node array to the neighboring nodes south and east with one more unit of delay in this example. Such clock signal propagation continues through the clock distribution network in node array 102 until the clock signal reaches a node in node array 102 at the opposite corner from the root node. In some examples, a signal routed from an originating node (e.g., node 206D) generating the signal to a neighboring node north or west of the originating node may travel upstream to the destination node (e.g., node 206A) and lose one unit of delay in node array 102. This signal routing is referred to as upstream signal travel. Signals routed from an originating node (e.g., node 206A) to a neighboring node to the south or east may travel downstream to a destination node (e.g., node 206D) and experience a unit delay in the node array 102. This may be referred to as downstream signal travel. Signals traveling upstream may be routed faster than signals traveling downstream to account for the unit delay and meet setup and hold time specifications.
[0044] FIG. 2A is a schematic diagram of a clock distribution network 200 according to one embodiment. Clock distribution network 200 includes a clock management unit (CMU) 202 and clock distribution circuitry in a node array 204 of nodes 206 (also referred to as a clock distribution node array). Each node 206 includes an instance of clock distribution circuitry for distributing a clock signal within node array 204. In the embodiment of FIG. 2A, clock distribution network 200 has a 2D distributed, strapped H-tree topology. CMU 202 is configured to output a clock signal, which is received at a root node 206A of node array 204.
[0045] 2A, the root may be located at the input to a node 206 at a corner of the node array 204. For example, the root may be located at the input to a node 206 at the northwest or upper left corner of the node array 204 (e.g., node 206A) shown in FIG. 2A. In other embodiments, the root may be the input to a node 206 at another corner of the node array 204 (e.g., node 206D) when the clock signal propagates in different directions along the rows and / or columns of the nodes. A node 206 that receives a clock signal from outside the node array 204 may be referred to as a root node 206.
[0046] 2A , the clock distribution network 200 can be implemented using a node array 204. The node array 204 shown in FIG. 2A is an example of the node array 102 with distributed clocking of FIG. 1. In particular embodiments, each node 206 can be an instance of a computing circuit. In particular applications, a majority of the nodes 206 include instances of computing circuitry, and one or more of the remaining nodes 206 include instances of different circuits, such as global nodes. A global node may refer to a node 206 that does not include circuitry for performing processing tasks. In some examples, a global node may include process, voltage, and temperature (PVT) sensors. In some implementations, both the computational nodes and the global nodes may include communication interfaces that enable communication with neighboring nodes 206. In some implementations, the communication interface for the computational nodes may be the same as the communication interface for the global nodes.
[0047] In particular embodiments, each node 206 of a node array 204 may include an instance of the same clock distribution circuitry, even if one or more other circuitry of the node 206 differs from that of the other nodes 206. In one embodiment of a node array 204, the nodes 206 may be tiled and abutted. At the same time, the node array 204 may be implemented without top-level wiring or gates. Thus, the nodes 206 may communicate with neighboring nodes 206 using lower-level wiring via short connections. The nodes 206 of a node array 204 may be stepped without mirroring or rotation. The nodes 206 may also be aligned to the grid pitch of power (VDD / VSS) lines. For example, the height and width of each node 206 may be a multiple of the power grid pitch. In some embodiments, the power grid pitch may be further aligned to the bump pitch.
[0048] As shown in Figure 2A, each node 206 may include substantially the same instance of clock distribution circuitry. Figure 2B illustrates an example implementation of clock distribution circuitry within an example node 206 of the node array 204 of Figure 2A. With reference to Figures 2A and 2B, the clock distribution circuitry includes a first input clock wire 222, a second input clock wire 224, a first inverter 226, a second inverter 228, a third inverter 230, a fourth inverter 232, a clock tap point 234, a first output clock wire 236, and a second output clock wire 238.
[0049] The clock distribution circuitry of each of the nodes 206 is designed so that the output clock wires 236 and 238 of the node 206 are aligned with the input clock wires 222 and 224 of a neighboring node 206. The nodes 206 can be stepped and tiled within the node array 204 so that the output clock wires 236 and 238 align with and electrically connect to the two input clock wires 222 and 224 of a neighboring node 206. Using these electrical connections, the node array 204 can be implemented without using channels or top-level wiring for clock distribution.
[0050] 2B , input wires 222 and 224 may receive input clock signals from two of the neighboring nodes 206. For example, first input clock wire 222 receives the input clock signal from the neighboring node 206 above the current node 206, while second input clock wire 224 receives the input clock signal from the neighboring node 206 to the left of the current node 206. First input clock wire 222 and second input clock wire 224 provide clock signals to first inverter 226 and second inverter 228. First inverter 226 inverts the clock signal and provides the inverted clock signal to clock tap point 234, which is then provided to the primary circuitry (e.g., computing circuitry or global circuitry in certain embodiments) of the corresponding node of computational node array 102.
[0051] The second inverter 228 inverts the clock signal and provides the inverted clock signal to a third inverter 230 and a fourth inverter 232. The third inverter 230 and the fourth inverter 232 each invert the inverted clock signal and output the resulting clock signals on a first output clock wire 236 and a second output clock wire 238. The first output clock wire 236 and the second output clock wire 238 output the clock signals to neighboring nodes 206 to the right and below the current node 206.
[0052] Referring back to FIG. 2A , a clock signal received at a root node 206 (e.g., node 206A) propagates to its two neighboring nodes below and to the right (e.g., nodes 206B, 206C) with one delay unit. The delay unit may be a fixed offset across the node array 204. In some implementations, the delay unit may correspond to the delay from buffering the clock signal (e.g., via inverters 228-232) combined with the wiring delay associated with the clock signal propagating to the downstream neighboring node 206. In FIG. 2A , one of the downstream neighboring nodes 206 is to the right in the same row as the root node 206, and the other downstream neighboring node 206 is below in the same column as the root node 206. In other words, the neighboring nodes 206 may be located south and east of the root node 206.
[0053] The clock signal continues to propagate to neighboring nodes 206 to the south and east with one more unit of delay as the clock signal traverses the entire node array 204 of Figure 2A. Such clock signal propagation continues through the clock distribution network until the clock signal reaches a node 206 of the node array 204 at the opposite corner from the root node 206 (e.g., node 206D).
[0054] As the clock signal propagates through the node array 204, a node 206 within the node array 204 may receive the clock signal from two other neighboring nodes 206 with substantially the same delay. A recombined mesh topology may combine two clock signals received from two neighboring nodes 206 at a given node 206 of the node array 204. For example, in FIG. 2B , the clock signals received via the first input clock wire 222 and the second input clock wire 224 may be combined and received at each of the first inverter 226 and the second inverter 228. In some embodiments, the clock signals are combined by connecting the first input clock wire 222 and the second input clock wire 224 directly together. Other implementations for providing a recombined mesh topology are also possible.
[0055] The clock distribution circuitry disclosed herein enables flexible array structures that support a wide range of array designs. For example, the node array 204 can be square, with a substantially equal number of rows and columns. Alternatively, the node array 204 can be rectangular, with a substantially different number of rows than columns. The clock distribution circuitry disclosed herein also provides a relatively simple reconfiguration of the array relative to the clock, which can also enable relatively late design decisions regarding node array shape. In contrast, array size and shape with other clock distribution networks are typically expensive decisions to postpone due to the amount of clock design time involved. However, in certain cases, such late decisions can result in overall chip design optimization and therefore may be desirable.
[0056] FIG. 2C shows another exemplary implementation of the clock distribution circuit configuration within an exemplary node 206 of the node array 204 of FIG. 2A. Referring to FIG. 2C, the clock distribution circuit configuration includes a first input clock wiring 222, a second input clock wiring 224, a second inverter 228, a third inverter 230, a fourth inverter 232, a clock tap point 234, a first output clock wiring 236, and a second output clock wiring 238. In FIG. 2C, each clock distribution circuit configuration of node 206 is designed such that the output clock wirings 236 and 238 of node 206 are aligned with the input clock wirings 222 and 224 of adjacent nodes 206. Node 206 can be made to be in a stepped and tiled pattern within the node array 204 such that the output clock wirings 236 and 238 are aligned and electrically connected to the two input clock wirings 222 and 224 of adjacent nodes 206. Using these electrical connections, the node array 204 can be implemented without using channels or top-level wiring for clock distribution.
[0057] In an example of the clock distribution circuit configuration, as shown in FIG. 2C, the input wirings 222 and 224 can receive input clock signals from two of the adjacent nodes 206. For example, the first input clock wiring 222 receives an input clock signal from the adjacent node 206 above the current node 206, while the second input clock wiring 224 receives an input clock signal from the adjacent node 206 to the left of the current node 206. The first input clock wiring 222 and the second input clock wiring 224 provide the clock signals to the second inverter 228. The second inverter 228 inverts the clock signal and provides the inverted clock signal to the third inverter 230 and the fourth inverter 232. Each of the third inverter 230 and the fourth inverter 232 inverts the inverted clock signal and outputs the resulting clock signals to the first output clock wiring 236 and the second output clock wiring 238. The first output clock wiring 236 and the second output clock wiring 238 output the clock signals to the adjacent nodes 206 to the right and below the current node 206.
[0058] FIG. 3A is a node clock level map associated with an exemplary node array, such as node array 204 of FIG. 2A. The exemplary node array 204 has 18 rows and 18 columns. With 18 rows and 18 columns, there can be 324 nodes. As another example, node array 204 can include 360 nodes arranged in rows and columns. Nodes 206 of node array 204 can have clock distribution circuitry corresponding to the clock distribution circuitry of FIG. 2B, for example. Nodes 206 of node array 204 can also have clock distribution circuitry corresponding to the clock distribution circuitry of FIG. 2C. This clock map indicates the number of unit delays of the clock signal output for nodes 206 of node array 204. For example, root node 206 has one unit delay. Two nodes 206 neighboring root node 206 have two unit delays. Nodes 206 on a diagonal line from southwest to northeast can have the same unit delay. Using the clock distribution circuitry described herein, the unit delays can be fixed offsets. Nodes 206 along these diagonals may receive clock signals with substantially the same timing delay. These diagonals may be referred to as phases or waves. The phases correspond to different clock signal arrival times at the nodes 206. Clock signal distribution corresponding to the map of FIG. 3A may implement a 35-phase mesochronous clock. The number of phases of the mesochronous clock signal for a node array having the clock distribution circuitry described herein may be the number of rows + the number of columns minus one.
[0059] In particular embodiments, rather than a clock signal traversing node array 204 in a wave formed along the diagonal of node array 204, clock distribution network 200 may be configured to generate a wave that traverses node array 204 in a row or column direction. For example, rather than outputting a clock signal south and east, each node 206 may output a clock signal either south or east. In this manner, a clock signal may propagate in a wave traveling south or east. However, aspects of the present disclosure are not limited to a particular direction of travel for the clock signal; the clock signal may propagate along other diagonals and / or north or west.
[0060] The offsets in Figure 3A may be taken into account when routing signals between nodes 206. Signals routed to nodes north or west of the originating node generating the signal may travel upstream and lose one unit delay in the node array 204 corresponding to Figure 3A. Signals routed to nodes south or east of the originating node may travel downstream and gain one unit delay in the node array 204 corresponding to Figure 3A. Signals traveling upstream may be routed faster than signals traveling downstream to account for the unit delay and meet setup and hold time specifications. The timing clock distribution shown in Figure 3A may also be referred to as wave clock distribution.
[0061] Figure 3B is a node clock level topology 310 corresponding to the node clock level map of Figure 3A. As shown in Figure 3B, nodes 206 included in groups 310, 320, 330, and 340 may have unit delays of 1, 2, 3, and 4, respectively. Nodes 206 included in the same group may have the same number of unit delays from the clock signal received at the root node. Timing Modeling
[0062] Timing in node arrays can be important for computational accuracy, performance, and so on, but accurately simulating timing in highly replicated node arrays can be difficult. For example, the size of network-on-chip (NoC) data buses can exponentially increase the netlist size of node arrays, and clock phase variations can cause inaccuracies in simulating timing distribution in node arrays. Furthermore, available electronic design automation (EDA) software can struggle to calculate timing for node arrays. In some cases, especially when the node array is large and / or complex, the EDA software may be unable to calculate timing for the node array. Alternatively, the EDA software may require a significant amount of time to calculate timing for the node array.
[0063] To address at least a portion of the above-described technical challenges, one or more aspects of the present disclosure correspond to systems and methods for modeling the timing of clock signal distribution in a node array. According to aspects, timing distribution of a node array can be simulated by modeling the timing delay of the node array based on timing delay analysis on a block of nodes (e.g., a portion of the node array). As described above, each node 206 of a node array 204 may include an instance of substantially the same clock distribution circuitry, even if one or more other circuitry of the node 206 differs from that of the other nodes 206. For example, the compute node 406 and the global node 408 shown in FIG. 4 may have the same clock distribution circuitry. In this manner, timing distribution can be modeled by performing timing analysis on one block (e.g., a node). For example, because a signal is traveling in one direction (e.g., upstream or downstream), the time delay depending on the signal direction can be determined based on the interface delay between the node and its neighboring (e.g., abutting or adjacent) nodes. This method may be advantageous because it avoids analyzing the time delays of all nodes in the node array and calculating the analyzed time delays to model the timing distribution of the node array.
[0064] As shown in FIG. 4 , in some embodiments, a node array may include compute nodes 406, global nodes 408, SerDes components 410, general-purpose input / output (GPIO) / security processing 412, etc. The compute nodes 406 may include circuitry for performing processing tasks. The global nodes 408 may not include circuitry for performing processing tasks. For example, the global nodes 408 may include PVT sensors for monitoring the operational status of the node array. In some implementations, both the compute nodes 406 and the global nodes 408 may include communication interfaces that enable communication with neighboring nodes. In some implementations, the communication interface for the compute nodes 406 may be the same as the communication interface for the global nodes 408.
[0065] In some embodiments, the techniques disclosed herein can be used to perform static timing analysis of a heavily replicated array having multiple compute nodes 406 and multiple global nodes 408. In some embodiments, dummy blocks (which may be similar to or the same as the global blocks described above) can have most of their internal circuitry removed, but can maintain interfaces and / or communications at their edges so that they can interface directly with functional blocks (e.g., compute nodes) in the array. In some embodiments, the array can distribute clock signals to the functional blocks and dummy blocks using mesh clock distribution.
[0066] In some implementations of a node array, nodes within the node array may communicate only with adjacent nodes. In such implementations, there may be no "flyover" signals, bypass signals, or other signals that cross the nodes. Routes connecting adjacent nodes may be horizontal or vertical. Nodes may be connected to neighboring nodes by horizontal and vertical routes. A computational node 406 that is not on the edge of the node array may interface with four neighboring nodes in the node array: the two nodes adjacent to the computational node 406 in the same row of the node array and the two nodes adjacent to the computational node 406 in the same column of the node array. A global node 408 may interface with four neighboring computational nodes 406 in the node array: the two computational nodes 406 adjacent to the global node 408 in the same row of the node array and the two computational nodes 406 adjacent to the global node 408 in the same column of the node array.
[0067] As mentioned above, modeling clock distribution across an array of nodes can be difficult, time-consuming, or even impossible using available EDA tools. Therefore, a simplified approach that can accurately model clock timing is desirable.
[0068] In some embodiments, the modeling techniques may include creating different global static timing model replicas of functional blocks and dummy blocks. The static timing model of a functional block may include models of the functional block with different surrounding environments (e.g., completely surrounded by functional blocks, with a dummy block on one side and a functional block on the other side, etc.). In some embodiments, there may be only a limited number of arrangements of functional blocks and / or dummy blocks around a functional block in a node array. The static timing model of a dummy block may take into account surrounding functional blocks. In some embodiments, a dummy block may be surrounded by functional blocks.
[0069] Large node arrays with mesochronous clocking present technical challenges for static timing analysis using traditional timing tools. Wide two-dimensional buses, even in hierarchical designs, can significantly increase netlist size and, consequently, simulation run time. Timing analysis of clock signals in node arrays described herein relies on the directionality of data propagation relative to the clock propagation direction. Because interfaces are only between neighboring nodes abutting each other within the node array, timing can be performed using a block group timing model of a node and its neighboring nodes that have communication interfaces with that node. Block group timing can involve simulations involving five nodes, which may be a small subset of the node array. The block group timing approach can avoid the need for arrival time annotation and the effort of correlating the simulation with the actual design, which may exist with other approaches. By creating a block group for each scenario within the node array, a complete node array can be accurately simulated based on several models. The block group timing approach is described with reference to Figure 5. Timing of node arrays in this disclosure may utilize one or more of the following simplifications: one block occupies the majority of the array (e.g., 98%), there is only an interface between abutting neighboring nodes, clock phases are systematic and matched, and most cross-delay variation is common node.
[0070] FIG. 5 illustrates an example of selecting nodes within a node array for modeling. In some embodiments, the timing of a computational node 406A, a neighboring node 406B around the computational node 406A, a global node 408, and a neighboring node 406C around the global node 408 may be modeled while other nodes 416 may be depopulated. Additionally, in some embodiments, a lane aggregator 414A and neighboring circuitry 414B of a node array may be modeled by modeling corners, while other nodes 416 within the node array may be depopulated by removing internal circuitry from the timing model. Such an approach can significantly reduce the number of nets, logic gates, wires, parasitic capacitances, etc. to be modeled. For example, the number of logic gates to be modeled may be reduced by about 80%, about 90%, etc. The reduction may depend, for example, on the number of nodes in the node array, the types of nodes within the node array, etc. Such an approach can provide significant speed improvements in modeling and may be particularly beneficial for replicated designs without global signals. Modeling computation time can be reduced, for example, from days or weeks to hours.
[0071] Referring to FIG. 5, the timing of clock distribution in a node array can be modeled using a computational node 406A, a neighboring node 406B around the computational node 406A, a global node 408, a neighboring node 406C around the global node, a lane aggregator 414A, and neighboring circuitry 414B. Clock timing within a node array can be simulated using a timing model of a small subset of the node array. The computational node 406A and the neighboring computational node 406B can be simulated to create a timing model of the computational node. The timing model of the computational node can be used for each computational node in the array. The global node 408 and the neighboring computational node 406C can be simulated to create a timing model of the global node. The timing model of the global node can be used for each global node in the array.
[0072] In some embodiments, a time delay of the clock signal between the computational node 406A and each neighboring node 406B can be created based on the determined delay. Each other computational node in the node array can use the same timing mode as the computational node 406A. For example, because each node in the node array includes the same instance of clock distribution circuitry and interface circuitry and abuts the same neighboring node instances, the same timing model can be used for each of the computational nodes in FIG. 3A.
[0073] A time delay of the clock signal between the global node 408 and each of the neighboring computational nodes 406C can be created based on the determined delay. Each other global node in the node array can use the same timing mode as the global node 408.
[0074] A time delay of the clock signal between the lane aggregator 414A and the neighboring circuitry 414B can be determined and used for each similar instance of such circuitry. A model can be determined and used for each similarly positioned lane aggregator. The lane aggregator model can simulate a device under test (DUT) block and associated interface paths.
[0075] The node array model can cover functional blocks (e.g., compute nodes) and depopulated blocks. The node array model can cover global communication node timing at a fraction of the size of the design without a gray-box model.
[0076] A method for generating a model of a circuit design having replicated instances of circuit blocks (e.g., node arrays) may include parsing a hardware description language (e.g., Verilog) model of the circuit design and generating models of the circuit blocks (e.g., functional blocks). The method may also include removing redundant similar instances of the circuit blocks to reduce the size of the model. The same or similar method may be performed for all other types of block instances (e.g., global blocks) of the design, in which case the model may include only unique scenarios. Static timing analysis may be performed on the model. Worst case timing
[0077] In some implementations of node arrays, block group timing models based on blocks of nodes in the node array are modeled to include various timing delay scenarios. For example, the delay between any two node clock arrivals within the node array may be different. Block distribution may be variable. For example, the delay between two nodes within the node array (e.g., near the center of the node array) may be different from the delay between two nodes near the edge of the node array. This may occur because, for example, capacitance may be different in different regions of the node array (e.g., higher in the center) due to the manufacturing process.
[0078] In some embodiments, timing can be simulated to determine worst-case scenarios for early and late arrivals. This information can be used to ensure that any node in the array can meet its arrival and hold time specifications. Such an approach can be used, for example, to ensure that modeling results can be applied to any node in a node array, even if the modeling was done using a particular node or set of nodes in the array.
[0079] In some embodiments, the mesh clock distribution can move in a wave-like distribution across the node array. The EDA tool can allow for annotating clock delays for early and late arrivals of the clock signal. In some embodiments, static timing analysis can be performed using only worst-case scenarios. Clock arrival differences can be generated between different compute nodes and global nodes to identify the worst-case scenario for any node with respect to arrival across the node array. Static timing analysis can be performed for the worst-case possible combinations of arrivals of all nodes (e.g., all compute nodes and global nodes) in the node array. Applying worst-case scenarios can reduce the run time for static timing analysis of the node array. Worst-case scenarios for both setup and hold times can be applied. Together, this can ensure that the static timing of the design converges for all combinations of communication traffic flowing between any two adjacent nodes in the node array.
[0080] In contrast, in a typical approach, a static timing analysis may be performed for every node in the array, which may require significant computational resources to complete.
[0081] Timing can be generated between any two nodes in a node array. Figure 6 shows determining a portion of a node array (e.g., a block group of a node array) timing between a node (e.g., compute node 406A) and an adjacent node (e.g., neighboring compute node 406B) in both the vertical and horizontal directions. Worst case early and late arrivals can be annotated for these adjacent nodes. This can ensure that setup and hold time specifications are met.
[0082] Worst-case timing data can be generated by parsing the results of a grid-wide node array circuit simulation (e.g., SPICE simulation). The fastest and slowest arrival times through the buffer can be selected. The buffer delays can be annotated in a timing model. This timing model can be optimized.
[0083] The worst-case timing data can be used in the model of the node array described in the previous section. Because only unique scenarios are used in the model, worst-case setup times and worst-case hold times can be used for each unique scenario in the model. The model and worst-case timing data can be used together to efficiently simulate the static timing of the node array and ensure that setup and hold time specifications are met.
[0084] The worst-case timing data can be used to modify the design of the clock distribution circuitry. If the worst-case timing is outside the timing specifications, the clock distribution circuitry can be adjusted until the timing specifications are met. Adjusting the clock distribution circuitry can include adjusting the size of one or more clock drivers (e.g., increasing or decreasing the size of a driver depending on whether a setup or hold time specification is not met) and / or adjusting the width of one or more wires carrying the clock signal (e.g., widening or narrowing a wire regardless of whether a setup or hold time specification is not met). Such adjustments to the clock distribution circuitry can be applied to each node of the array. Design automation tools can automate the process of adjusting the clock distribution circuitry until the worst-case timing meets the timing specifications. In some other applications, circuit designers can use the worst-case timing data to update the clock distribution circuitry. Exemplary Embodiment of a Node Array Timing Distribution Model
[0085] In some embodiments, a node array timing distribution model can be used to simulate clock timing in a node array. For example, simulation results of the node array timing distribution model can be used to determine worst-case scenarios for clock timing in a node array. These simulation results can be utilized to design a clock distribution network, including interfaces and / or inverters and / or wiring that carry signals between nodes in the node array.
[0086] 7 illustrates an example of a computing device 710 capable of simulating clock timing distribution across a node array. As shown in FIG. 7, the computing device 710 implements a timing distribution simulation component 720, a non-volatile storage device 714, and a main processor 712.
[0087] In some examples, the main processor 712 can provide dedicated computing resources to be used by the timing distribution simulation component 720. Furthermore, according to examples disclosed herein, the main processor 712 can utilize the designated computing resources to process data generated from the timing distribution simulation component 720.
[0088] As shown in FIG. 7 , the timing distribution simulation component 720 may include a timing model generator 722 and a timing model simulator 724. The timing model generator 722 may be configured to model the time delays of a node array by analyzing the timing delays on a block of nodes (e.g., a portion of the node array). With reference to FIGS. 4 and 5 , the timing model generator 722 may utilize a modeling technique that includes creating different global static timing model replicas for functional blocks and dummy blocks. The static timing model of a functional block may include a model of the functional block having different surrounding environments (e.g., completely surrounded by functional blocks, having a dummy block on one side and a functional block on the other side, etc.). In some embodiments, there may be only a limited number of arrangements of functional blocks and / or dummy blocks around a functional block in a node array. The static timing model of a dummy block may consider surrounding functional blocks. In some embodiments, the dummy block may be surrounded by functional blocks. The timing analysis of clock signals in a node array described herein depends on the directionality of data propagation relative to the clock propagation direction. Because the only interfaces are between neighboring nodes abutting each other within the node array, timing can be performed using a block group timing model of a node and its neighboring nodes that have communication interfaces with that node. Block group timing can involve simulations involving five nodes (e.g., a central node abutted by neighboring nodes), which can be a small subset of the node array. The block group timing approach can avoid the need to annotate arrival times and the effort of correlating the simulation with the actual design, which may exist with other approaches. By creating a block group for each scenario within the node array, a complete node array can be accurately simulated based on several models.
[0089] In some embodiments, the timing model generator 722 may model and select a computational node, neighboring nodes around the computational node (e.g., four abutting neighboring nodes), a global node, and neighboring nodes around the global node, while other nodes in the node array may be depopulated. Additionally, in some embodiments, the lane aggregator and neighboring circuitry of the node array may be modeled by modeling corners, while other nodes in the node array may be depopulated by removing internal circuitry from the timing model. Such an approach may significantly reduce the number of nets, logic gates, wires, parasitic capacitances, etc. to be modeled. For example, the number of logic gates to be modeled may be reduced by about 80%, about 90%, etc. The reduction may depend, for example, on the number of nodes in the node array, the types of nodes in the node array, etc. Such an approach may provide significant speed improvements in modeling and may be particularly beneficial for replicated designs without global signals. The computational time for modeling may be reduced, for example, from days or weeks to a few hours. In some embodiments, the models generated from the timing model generator 722 may be stored in the non-volatile storage device 714 .
[0090] The timing model simulator 724 can be configured to simulate the models stored in the non-volatile storage device 714. In some embodiments, the timing model simulator 724 can access the models by accessing the non-volatile storage device 714. In some embodiments, a model including a computational node and neighboring computational nodes can be simulated to create a timing model for the computational node. The timing model for the computational node can also be used for each computational node in the array. The global node and neighboring computational nodes can also be simulated to create a timing model for the global node. The timing model for the global node can also be used for each global node in the array.
[0091] In some embodiments, the timing model simulator 724 creates time delays for clock signals between the computational node and each neighboring node based on the determined delays. Each other computational node in the node array can use the same timing mode as the computational node. For example, the same timing model can be used for each of the computational nodes in FIG. 3A because each node in the node array includes the same instance of clock distribution circuitry and interface circuitry and abuts the same instances of neighboring nodes. Time delays for clock signals between the global node and each of the neighboring computational nodes can also be created based on the determined delays. Each other global node in the node array can use the same timing mode as the global node.
[0092] In some embodiments, the timing model simulator 724 can simulate node arrays by modeling node arrays that cover functional blocks (e.g., compute nodes) and depopulated blocks. A method for generating a simulation model of a circuit design having replicated instances of circuit blocks (e.g., node arrays) may include parsing a hardware description language (e.g., Verilog) model of the circuit design and generating models of the circuit blocks (e.g., functional blocks). The method may also include removing redundant similar instances of the circuit blocks to reduce the size of the model. The same or similar method can be performed for all other types of block instances (e.g., global blocks) of the design. In that case, the model may include only unique scenarios.
[0093] In some embodiments, the timing model simulator 724 performs timing analysis to determine worst-case timing data and includes the determined worst-case timing data in the simulation. For example, timing can be generated between any two nodes in a node array (e.g., FIG. 6 illustrates determining a portion of a node array (e.g., a block group of a node array) timing between a node (e.g., compute node 406A) and an adjacent node (e.g., neighboring compute node 406B) in both vertical and horizontal directions). Worst-case early and late arrivals can be annotated for these adjacent nodes. This can ensure that setup and hold time specifications are met.
[0094] Worst-case timing data can be generated by parsing the results of a grid-wide node array circuit simulation (e.g., SPICE simulation). The fastest and slowest arrival times through the buffer can be selected. The buffer delays can be annotated in a timing model. This timing model can be optimized.
[0095] The worst-case timing data can be used in the model of the node array generated by timing model generator 722. Because only unique scenarios are used in the model, a worst-case setup time and a worst-case hold time can be used for each unique scenario in the model. The model and worst-case timing data can be used together to efficiently simulate the static timing of the node array and ensure that setup and hold time specifications are met.
[0096] In some embodiments, the worst-case timing data can be used to modify the design of the clock distribution circuitry. If the worst-case timing is outside the timing specifications, the clock distribution circuitry can be adjusted until the timing specifications are met. Adjusting the clock distribution circuitry can include adjusting the size of one or more clock drivers (e.g., increasing or decreasing the size of a driver depending on whether a setup or hold time specification is not met) and / or adjusting the width of one or more wires carrying the clock signal (e.g., widening or narrowing a wire regardless of whether a setup or hold time specification is not met). Such adjustments to the clock distribution circuitry can be applied to each node of the array. Design automation tools can automate the process of adjusting the clock distribution circuitry until the worst-case timing meets the timing specifications. In some other applications, circuit designers can use the worst-case timing data to update the clock distribution circuitry.
[0097] For ease of discussion and not to limit the present disclosure, FIG. 7 shows only a timing distribution simulation component 720, a non-volatile storage device, and a main processor, although multiple sub-components or systems may be used. Uses, Terminology, and Conclusions
[0098] The node arrays disclosed herein can be implemented in a variety of processing systems. Such processing systems may be used in and / or specifically configured for high-performance computing and / or computationally intensive applications, such as neural network training, neural network inference, machine learning, artificial intelligence, complex simulations, etc. In some applications, the processing systems may be used to perform neural network training. For example, such neural network training may generate data for a vehicle's (e.g., automobile) autopilot system, other autonomous vehicle functions, or advanced driver assistance system (ADAS) functions.
[0099] Unless the context clearly dictates otherwise, throughout the specification and claims, words such as "comprise," "comprising," "include," "including," and the like, are to be construed in an inclusive sense, i.e., "including, but not limited to," as opposed to an exclusive or exhaustive sense. The term "coupled," as generally used herein, refers to two or more elements, which may be directly connected or connected by one or more intermediate elements. Similarly, the term "connected," as generally used herein, refers to two or more elements, which may be directly connected or connected by one or more intermediate elements. Furthermore, the words "herein," "above," "below," and words of similar import, when used in this application, shall refer to this application as a whole and not to particular portions of this application. Where the context permits, words in the above detailed description using the singular or plural may also include the plural or singular, respectively. The word "or" in connection with a list of two or more items covers all of the following interpretations of that word: any of the items in the list, all of the items in the list, and any combination of the items in the list.
[0100] Additionally, conditional language used herein, particularly "can," "could," "might," "may," "eg," "for example," "such as," and the like, unless expressly stated otherwise or understood otherwise within the context in which it is used, is generally intended to convey that certain embodiments include certain features, elements, and / or conditions, while other embodiments do not. Thus, such conditional language is generally not intended to imply that features, elements, and / or conditions are in any way required for one or more embodiments.
[0101] The above description has been given with reference to specific embodiments. However, the above exemplary discussion is not intended to be exhaustive or to limit the invention to the precise form described. Many modifications and variations are possible in light of the above teachings. This will enable those skilled in the art to best utilize the techniques and various embodiments with various modifications suitable for various applications.
[0102] Although the present disclosure and examples have been described with reference to the accompanying drawings, various changes and modifications will become apparent to those skilled in the art, and such changes and modifications should be understood to be included within the scope of the present disclosure.
Claims
1. 1. A method for simulating a node array, comprising: accessing a timing model of a computational node of the node array, the timing model representing timing data related to clock signal propagation between the computational node and four neighboring nodes of the node array, each neighboring node abutting the computational node; using one or more computers to simulate clock signal timing for a majority of the nodes in the node array using the timing model of the computational nodes; A method comprising:
2. The method of claim 1 , further comprising determining worst-case timing of the clock signals in the node array based on the simulating step.
3. 3. The method of claim 2, further comprising adjusting a clock distribution network of the node array based on the worst-case timing.
4. 4. The method of claim 3, wherein adjusting the clock distribution network comprises updating one or more files representing circuit configurations of the node array.
5. 3. The method of claim 2, further comprising accessing a global nodal timing model of the node array stored in the non-transitory computer-readable memory, and wherein the determining is based on a simulation of clock signal timing using the global nodal timing model.
6. 2. The method of claim 1, wherein the simulating step includes simulating mesochronous clocking in the node array.
7. 2. The method of claim 1, wherein the timing model of the node array models the computational node receiving the clock signal from a first pair of the four neighboring nodes and the computational node providing the clock signal to a second pair of the four neighboring nodes, the clock signal being delayed by one delay unit at the computational node relative to the first pair of neighboring nodes, and the clock signal being delayed by two delay units at the second pair of neighboring nodes relative to the first pair of neighboring nodes.
8. The method of claim 1 , wherein the node array timing model is a block group timing model.
9. The method of claim 1 , wherein the majority of nodes in the node array comprises at least 90% of the nodes in the node array.
10. The method of claim 9 , wherein the node array consists essentially of instances of the compute nodes and instances of global nodes.
11. The method of claim 1 , further comprising generating a timing model of the compute node by at least simulating clock signal propagation between the compute node and the four neighboring nodes.
12. 1. A non-transitory computer-readable storage device comprising instructions that, when executed by one or more processors, cause a method for simulating a node array to be performed, the method comprising: accessing a timing model of a computational node of the node array, the timing model representing timing data related to clock signal propagation between the computational node and four neighboring nodes of the node array, each neighboring node abutting the computational node; using one or more computers to simulate clock signal timing for a majority of the nodes in the node array using the timing model of the computational nodes; 4. Non-transitory computer readable storage, including:
13. 1. A computer system for simulating a node array, comprising: a non-transitory computer readable memory that stores a timing model of a computational node of a node, the timing model representing timing data related to clock signal propagation between the computational node and four neighboring nodes of the node array, each neighboring node abutting the computational node; one or more processors configured to access at least a timing model of the computational nodes and execute instructions to simulate clock signal timing for a majority of nodes in the node array using the timing model of the computational nodes; A computer system comprising:
14. 1. A system for simulating clock timing distribution across a node array including a plurality of compute nodes, comprising: one or more computing devices configured to store timing models corresponding to a computing node and four neighboring computing nodes to which each neighboring computing code is attached, wherein each of the one or more computing devices comprises: accessing a timing model of the compute node; using a computing device to simulate clock signal timing distribution for a majority of the nodes in the node array using a timing model of the compute node; one or more computing devices configured to: A system comprising:
15. 1. A non-transitory computer-readable storage medium storing instructions for simulating clock timing distribution across a node array, the instructions, when executed by a processor, causing the processor to: accessing a timing model of a computational node, the timing model representing timing data related to clock signal propagation between the computational node and four neighboring nodes of the node array, each neighboring node abutting the computational node; using a computing device to simulate clock signal timing distribution for a majority of the nodes in the node array using a timing model of the compute node; A non-transitory computer-readable storage medium for causing a computer to perform operations including:
16. 16. The non-transitory computer-readable storage medium of claim 15, further comprising determining worst-case timing of the clock signals in the node array based on the simulating.
17. 20. The non-transitory computer-readable storage medium of claim 16, further comprising adjusting a clock distribution network of the node array based on the worst-case timing.
18. 20. The non-transitory computer-readable storage medium of claim 17, wherein adjusting the clock distribution network comprises updating one or more files representing circuit configurations of the node array.
19. 17. The non-transitory computer-readable storage medium of claim 16, further comprising accessing a global nodal timing model of the node array stored in the non-transitory computer-readable memory, wherein the determining is based on a simulation of clock signal timing using the global nodal timing model.
20. 16. The non-transitory computer-readable storage medium of claim 15, wherein the simulating comprises simulating mesochronous clocking in the node array.
21. 16. The non-transitory computer-readable storage medium of claim 15, wherein the timing model of the node array models the computational node receiving the clock signal from a first pair of the four neighboring nodes and the computational node providing the clock signal to a second pair of the four neighboring nodes, wherein the clock signal is delayed by one delay unit at the computational node relative to at the first pair of neighboring nodes and the clock signal is delayed by two delay units at the second pair of neighboring nodes relative to at the first pair of neighboring nodes.
22. 16. The non-transitory computer-readable storage medium of claim 15, wherein the node array timing model is a block group timing model.
23. 16. The non-transitory computer-readable storage medium of claim 15, wherein the majority of nodes in the node array comprises at least 90% of the nodes in the node array.
24. 23. The non-transitory computer-readable storage medium of claim 22, wherein the node array consists essentially of instances of the compute nodes and instances of global nodes.
25. 16. The non-transitory computer-readable storage medium of claim 15, further comprising generating a timing model of the compute node by at least simulating clock signal propagation between the compute node and the four neighboring nodes.
Citation Information
Patent Citations
Clock distribution network rapid design method
CN110688723A
Method and apparatus for determining timing uncertainty in mesh circuit
JP2007242015A
Method, device and system for analyzing clock mesh
JP2007328788A
Method for balanced-delay clock tree insertion
US6698006B1