Clock timing in replicating arrays

By accessing and simulating the timing model of node arrays in a high-density processing system, the problem of difficult to accurately model clock signal propagation and synchronization between nodes is solved, and the accurate simulation and optimization of clock signal propagation is achieved, and the system performance is improved.

CN119998813APending Publication Date: 2025-05-13TESLA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380070829.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-08-19
Filing Date
2023-08-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In high-density processing systems, clock signal propagation and synchronization between nodes is difficult to accurately model and simulate, resulting in communication instability and performance degradation.

Method used

By accessing the compute node timing model of the node array in non-transitory computer-readable memory, the computer simulates the clock signal timing of most nodes in the node array, including determining the worst-case timing and adjusting the clock allocation network.

Benefits of technology

Accurate simulation and optimization of the propagation of the clock signal in the node array is realized, and synchronization stability and system performance are improved between nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998813A_ABST
    Figure CN119998813A_ABST
Patent Text Reader

Abstract

The present disclosure relates to systems and methods for simulating clock timing allocation across an array of nodes (204). An example method includes accessing a timing model of a compute node (206), where the timing model of the compute node (206) represents timing data associated with clock signal propagation between the compute node (206) of an array of nodes (204) and four neighboring nodes each contiguous with the compute node (206), and using a computing device, transmitting the timing model of the compute node (206) to the array of nodes (204), where the timing model of the compute node (206) represents timing data associated with clock signal propagation between the compute node (206). A timing model of the compute nodes (206) is used to simulate a clock signal timing allocation for a majority of the nodes in the array of nodes (204).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to priority application

[0002] This application claims priority to U.S. Provisional Application No. 63 / 371,993, filed on August 19, 2022, entitled “CLOCK TIMING IN REPLICATE DARRAYS,” the disclosure of which is hereby incorporated by reference in its entirety for all purposes. Technical Field

[0003] The present disclosure relates generally to distributed clocks, and in particular to techniques for modeling clocks in an array. Background Art

[0004] An array of processing nodes can be used to build a high-density processing system. A node can communicate with an adjacent node to perform a processing task. Communication between nodes can use synchronous and / or asynchronous methods. A clock signal can be provided to each node so that the nodes can be synchronized, thereby enabling communication between nodes. Summary of the invention

[0005] The innovations described in the claims each have several aspects, no single one of which is solely responsible for its desirable attributes. Without limiting the scope of the claims, some of the prominent features of the disclosure will now be briefly described.

[0006] One aspect of the present disclosure is a method for simulating a node array. The method includes accessing a timing model of a computing node of the node array stored in a non-transitory computer readable memory, and using one or more computers, using the timing model of the computing node to simulate clock signal timing for most nodes in the node array. The timing model of the computing node represents timing data associated with clock signal propagation between the computing node and four adjacent nodes of the node array, each of the four adjacent nodes being adjacent to the computing node.

[0007] The method may also include determining, based on simulations, worst case timing for clock signals in the array of nodes.

[0008] The method may also include adjusting a clock distribution network of the node array based on the worst case timing. Further, adjusting the clock distribution network may include updating one or more files representing the circuit system of the node array. Further, the method may also include accessing a timing model of a global node of the node array stored in a non-transitory computer readable memory. The determination may be based on a simulation of the clock signal timing, the simulation of the clock signal timing using the timing model of the global node.

[0009] In the method, the simulation may include simulating mesochronous clocking in the array of nodes.

[0010] In the method, the timing model of the node array can model a computing node that receives a clock signal from a first pair of adjacent nodes among four adjacent nodes and a computing node that provides a clock signal to a second pair of adjacent nodes among the four adjacent nodes. In addition, the clock signal can be delayed by one unit delay in the computing node relative to the first pair of adjacent nodes. In addition, the clock signal can be delayed by two unit delays in the second pair of adjacent nodes relative to the first pair of adjacent nodes.

[0011] In the method, the timing model of the node array is a block group timing model. In addition, the node array may be substantially composed of instances of compute nodes and instances of global nodes.

[0012] In the method, the majority of nodes of the node array may include at least 90% of the nodes in the node array.

[0013] The method may also include generating a timing model of the compute node by simulating at least clock signal propagation between the compute node and four neighboring nodes.

[0014] Another aspect of the present disclosure is a non-transitory computer-readable storage including instructions that, when executed by one or more processors, cause execution of a method for simulating a node array. The method includes accessing a timing model of a computing node of the node array stored in a non-transitory computer-readable memory, and using one or more computers, using the timing model of the computing node to simulate the clock signal timing for most of the nodes in the node array. In addition, the timing model of the computing node represents timing data associated with clock signal propagation between the computing node and four adjacent nodes of the node array, and each of the four adjacent nodes is adjacent to the computing node.

[0015] Another aspect of the present disclosure is a computer system for simulating a node array. The system includes: a non-transitory computer-readable memory storing a timing model of a computing node of the node; and one or more processors configured to execute instructions to at least access the timing model of the computing node and use the timing model of the computing node to simulate the clock signal timing for most nodes in the node array. In addition, the timing model of the computing node represents the timing data associated with the clock signal propagation between the computing node and four adjacent nodes of the node array, and each of the four adjacent nodes is adjacent to the computing node.

[0016] Another aspect of the present disclosure is a system for simulating clock timing distribution on a node array including multiple computing nodes. The system may include one or more computing devices configured to store a timing model, the timing model corresponding to a computing node and four adjacent computing nodes, and the four adjacent computing codes are each adjacent to the computing node. In addition, each computing device in the one or more computing devices is configured to access the timing model of the computing node, and use the computing device to simulate the clock signal timing distribution for most nodes in the node array using the timing model of the computing node.

[0017] Another aspect of the present disclosure is a non-transitory computer-readable storage medium that stores instructions to simulate clock timing distribution on a node array. The instructions, when executed by a processor, cause the processor to perform operations, the operations including accessing a timing model of a computing node, and using a computing device, using the timing model of the computing node to simulate the clock signal timing distribution for most nodes in the node array. In addition, the timing model of the computing node represents timing data associated with clock signal propagation between a computing node and four adjacent nodes of the node array, and each of the four adjacent nodes is adjacent to the computing node.

[0018] In the non-transitory computer readable storage medium, the operation may include determining a worst-case timing for a clock signal in the node array based on the simulation. Additionally, the non-transitory computer readable storage medium may include adjusting a clock distribution network of the node array based on the worst-case timing. Adjusting the clock distribution network may include updating one or more files representing the circuit system of the node array.

[0019] In the non-transitory computer readable storage medium, the operation may include accessing a timing model of a global node of the node array stored in the non-transitory computer readable memory. The determination is based on a simulation of the clock signal timing, the simulation of the clock signal timing using the timing model of the global node.

[0020] In the non-transitory computer readable storage medium, the simulation includes simulating an average synchronized clock in an array of nodes

[0021] In the non-transitory computer readable storage medium, the timing model of the node array can model a computing node that receives a clock signal from a first pair of adjacent nodes among four adjacent nodes and a computing node that provides a clock signal to a second pair of adjacent nodes among the four adjacent nodes. In addition, the clock signal can be delayed by one unit delay in the computing node relative to the first pair of adjacent nodes. In addition, the clock signal can be delayed by two unit delays in the second pair of adjacent nodes relative to the first pair of adjacent nodes.

[0022] In the non-transitory computer readable storage medium, the timing model of the node array may be a block group timing model. In addition, the node array may be substantially composed of instances of compute nodes and instances of global nodes.

[0023] In the non-transitory computer readable storage medium, the majority of the nodes of the node array may include at least 90% of the nodes in the node array.

[0024] In the non-transitory computer-readable storage medium, the operations may also include generating a timing model of the computing node by simulating at least clock signal propagation between the computing node and four neighboring nodes.

[0025] To summarize the present disclosure, certain aspects, advantages and novel features of the innovations are described herein. It should be understood that not all such advantages may be realized according to any particular embodiment. Therefore, the innovation may be embodied or performed in a manner that realizes or optimizes one advantage or group of advantages taught herein, without necessarily realizing other advantages taught or implied herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Embodiments of the present disclosure will be described by way of non-limiting examples with reference to the accompanying drawings.

[0027] Figure 1 is a schematic block diagram of an example chip according to aspects of the present disclosure.

[0028] Figure 2A is a schematic diagram of a clock distribution network according to an embodiment.

[0029] Figure 2B Pictured Figure 2A An example implementation of clock distribution circuitry within an example node of an array of nodes.

[0030] Figure 2C Pictured Figure 2A Another example implementation of clock distribution circuitry within an example node of an array of nodes.

[0031] Figure 3A is with Figure 2A Node clock-level diagram associated with an example node array such as a node array of .

[0032] Figure 3B is with Figure 3A The node clock level diagram corresponds to the node clock level topology.

[0033] Figure 4 Pictured Figure 1 An example implementation of a node array.

[0034] Figure 5 Pictured Figure 1 Characterization of the node array to model the clock timing distribution on the node array.

[0035] Figure 6 Pictured Figure 5 An example of a block of node arrays.

[0036] Figure 7 An example embodiment of a computing device is illustrated. DETAILED DESCRIPTION

[0037] The following detailed description of certain embodiments presents various descriptions of specific embodiments. However, the innovations described herein may be embodied in a variety of different ways, for example, as defined and covered by the claims. In this description, reference is made to the accompanying drawings in which the same reference numerals and / or terminology may indicate identical or functionally similar elements. It will be understood that the elements illustrated in the figures are not necessarily drawn to scale. Furthermore, it will be understood that certain embodiments may include more elements than illustrated in the figures and / or a subset of the elements illustrated in the figures. Furthermore, some embodiments may combine any suitable combination of features from two or more of the figures. The headings provided herein are for convenience only and do not necessarily affect the scope or meaning of the claims. Introduction to Distributed Clocks for Arrays of Nodes

[0038] The present disclosure relates to a clock distribution network having clock signals that arrive at various nodes of a node array at different times. Generating a clock signal with a fixed offset can be referred to as an average synchronous clock. Embodiments disclosed herein relate to an average synchronous clock network constructed by modularization of a common circuit system. The clock signal of such a network can be locally low-skew and averagely synchronized at a coarser level.

[0039] The principles and advantages disclosed herein can be applied to any suitable circuit chip. In some applications, the distribution of clock signals disclosed herein can be applied to several chips, each of which includes a smaller computing node array. The computing node can be referred to as a processor or core. In this way, the clock signal can form an arrival time wave on the computing node array. In one embodiment, each computing node can receive a low-skew clock. The computing nodes of the array can be designed to use only interfaces with adjacent computing nodes to consider the arrival time difference (offset) of the average synchronous clock phase. The technology described herein can be applied to square (rows and columns are equal) node arrays or rectangular node arrays with different numbers of rows and columns.

[0040] Figure 11 is a schematic block diagram of an example chip 100 according to aspects of the present disclosure. The chip 100 may be an integrated circuit die. The chip 100 may include a node array 102 (also referred to as a compute node array) with a distributed clock, one or more serializer / deserializer (SerDes) clock blocks 104, a clock generator 106, and a clock controller 108. The SerDes clock block 104 may be joined with other chips 100 to form the chip array 100. In some application examples, the node array 102 may be included on a chip 100 in a system on a wafer, a chip array 100 on a printed circuit board, etc. In some applications, Figure 1 The node array 102 can be implemented on a system on wafer packaged with a wafer level packaging structure. Figure 1 As shown in the embodiment of the present invention, the clock generator 106 can be implemented outside the node array 102. In some embodiments, the clock generator 106 may include a phase-locked loop (PLL). The clock generator 106 can be arranged to provide clock signals to the computing nodes at the corners of the node array 102. The clock controller 108 can also be implemented outside the node array 102. The nodes within the node array 102 can include node-to-node interfaces, which can be configured to communicate synchronously. The core to serializer / deserializer (SerDes) interface can be asynchronous. In some embodiments, the PLL operates on a 100MHz reference frequency band, however other frequency bands are also considered to be within the scope of the present disclosure. The PLL can be configured to generate source clocks for various operating modes, including functional, bypass, and test modes. The PLL can be configured to operate without glitch source clock selection and is configured to manage any thermal issues and the maximum current through clock throttling. In some examples, clock throttling, OGG (on-chip clock control) for scan capture, and clock up / down are implemented by cycle skipping to modulate the effective frequency using, for example, a 32x32 first-in-first-out (FIFO) mode.

[0041] exist Figure 1 In the node array 102 with distributed clocks, each node can be an instance of a computing circuit (also referred to as a processing core or computing node). In some applications, most of the nodes can be implemented as instances of computing circuits, and one or more of the nodes can be implemented as instances of different circuits. Each node of the node array 102 can include an instance of substantially the same clock distribution circuit system even if the other circuit systems of at least some of the nodes are different from those of other nodes. For example, most of the nodes can be implemented as instances of computing circuits, and one or more of the nodes can be implemented as instances of global nodes. Global nodes can include process, voltage, and temperature (PVT) sensors ( Figure 1102). In the node array 102, the nodes can be tiled and adjacent. For example, each node in the node array 102 can be independent and interconnected to (multiple) adjacent nodes (e.g., adjacent (multiple) nodes). At the same time, the node array 102 can be implemented without using top-level wires, gates, or channels. Therefore, a node can be configured to communicate with an adjacent node with a lower-level wire through a relatively short connection. In some embodiments, the nodes of the node array 102 can be arranged in a stepped manner without mirroring or rotation. In some implementations, the nodes can be connected to the power line (V DD / V SS ) grid pitch. For example, the height and width of each node can be a multiple of the power grid pitch. The power grid pitch can be further aligned with the bump pitch.

[0042] Each node in the node array 102 may include an instance of substantially identical clock distribution circuitry. The nodes may be designed so that the output clock line of the node is aligned with the input clock line of its adjacent node. The nodes may be arranged in a stepped manner and tiled in the node array so that the clock output line is aligned and electrically connected with the clock input line of the adjacent node arranged downstream to receive the clock signal. With this electrical connection, a node array may be implemented without channels or top-level wiring for clock distribution. In some embodiments, the fan-out of the clock distribution circuitry may be balanced for use with inverters.

[0043] As described herein, the clock signal received at the root node can be propagated from the root node to two adjacent nodes with a unit delay. The root node can be located at the corner of the node array 102. The unit delay can be a fixed offset for a given node array. The unit delay can correspond to the delay from buffering the clock signal (e.g., using an inverter) and the line delay associated with the clock signal propagating to its (multiple) adjacent nodes. For example, in one embodiment, one of the two adjacent nodes is located in the same row as the root node, and the other of the two adjacent nodes is located in the same column as the root node. As an example, the adjacent nodes are positioned to the south and east of the root node. In this configuration, the clock signal continues to propagate from the two adjacent nodes in the node array in this example to the adjacent nodes in the south and east with a unit delay. Such clock signal propagation continues through the clock distribution network until the clock signal arrives at the node in the node array that is in the diagonal position of the root node. In some examples, a signal routed from an originating node (e.g., a node at the southeast corner of node array 102) to a node to the north or west may travel upstream and lose one unit delay in the node array, while a signal routed from an originating node (e.g., a node at the southeast corner of node array 102) to a node to the south or east may propagate downstream and gain one unit delay in the node array. In some applications, signals traveling upstream may be routed faster than signals traveling downstream to account for unit delay and meet setup and hold time specifications.

[0044] One of the two adjacent nodes may be located in the same row as the root node, and the other of the two adjacent nodes may be located in the same column as the root node. In some embodiments, the adjacent nodes are adjacent to the root node. As an example, Figure 2A, the adjacent nodes are located to the south and east of the root node. For example, the adjacent nodes of node 206A may be nodes 206B and 206C. The clock signal continues to propagate from the two adjacent nodes of the root node in the node array in this example to the adjacent nodes to the south and east with one or more unit delays. This clock signal propagation continues through the clock distribution network in the node array 102 until the clock signal arrives at the node in the node array 102 that is in the diagonal position of the root node. In some examples, the signal that is routed from the originating node (e.g., node 206D) that generates the signal to the adjacent nodes to the north or west of the originating node can travel upstream to the target node (e.g., node 206A), and loses a unit delay in the node array 102. This signal routing is referred to as upstream signal travel. The signal that is routed from the originating node (e.g., node 206A) to the adjacent nodes to the south or east can travel downstream to the target node (e.g., node 206D), and obtains a unit delay in the node array 102. This can be referred to as downstream signal travel. Signals traveling upstream can be routed faster than signals traveling downstream to account for unit delays and meet setup and hold time specifications.

[0045] Figure 2A is a schematic diagram of a clock distribution network 200 according to an embodiment. Clock distribution network 200 includes a clock management unit (CMU) 202, and clock distribution circuitry of a node array 204 (also referred to as a clock distribution node array) of nodes 206. Each node 206 includes an instance of clock distribution circuitry for distributing clock signals within node array 204. Figure 2A In the embodiment of the present invention, the clock distribution network 200 has a 2D distributed striped H-tree topology. The CMU 202 is configured to output a clock signal, which is received at the root node 206A of the node array 204.

[0046] refer to Figure 2A , the root may be located at the input of a node 206 at a corner of the node array 204. For example, the root may be located at Figure 2A 204 (e.g., node 206A). In other embodiments, when the clock signal propagates in different directions along the rows and / or columns of nodes, the root may be the input of a node 206 at another corner (e.g., node 206D) of the node array 204. A node 206 that receives a clock signal from outside the node array 204 may be referred to as a root node 206.

[0047] Further references Figure 2A , the clock distribution network 200 can be implemented using a node array 204 . Figure 2A The node array 204 shown in FIG. Figure 1An example of a node array 102 with distributed clocks. In some embodiments, each node 206 can be an instance of a computing circuit. In some applications, a majority of the nodes 206 include instances of computing circuits, and one or more of the remaining nodes 206 (such as global nodes) include instances of different circuits. A global node may refer to a node 206 that does not include a circuit system for performing a processing task. In some examples, a global node may include a process, voltage, and temperature (PVT) sensor. In some implementations, both the computing nodes and the global nodes may include a communication interface to enable communication with adjacent nodes 206. In some implementations, the communication interface of the computing node may be the same as the communication interface for the global node.

[0048] In some embodiments, each node 206 in the node array 204 may include an instance of the same clock distribution circuitry, even if other circuitry of one or more nodes 206 is different from that of other nodes 206. In one embodiment of the node array 204, the nodes 206 may be tiled and contiguous. Also, the node array 204 may be implemented without any top-level wires or gates. Thus, a node 206 may communicate with an adjacent node 206 having lower-level wires via short connections. The nodes 206 of the node array 204 may be arranged in a stepped pattern without mirroring or rotation. The nodes 206 may also be connected to a power supply (V DD / V SS ) lines. For example, the height and width of each node 206 can be a multiple of the power grid pitch. In some embodiments, the power grid pitch can be further aligned with the bump pitch.

[0049] like Figure 2A As shown in , each node 206 may include substantially identical instances of clock distribution circuitry. Figure 2B The diagram shows Figure 2A An example implementation of clock distribution circuitry within an example node 206 of the node array 204 is shown. Figure 2A and Figure 2B The clock distribution circuit system includes a first input clock line 222, a second input clock line 224, a first inverter 226, a second inverter 228, a third inverter 230, a fourth inverter 232, a clock tap point 234, a first output clock line 236, and a second output clock line 238.

[0050] The clock distribution circuitry of each of the nodes 206 is designed so that the output clock lines 236 and 238 of the node 206 are aligned with the input clock lines 222 and 224 of the adjacent nodes 206. The nodes 206 can be arranged in a stepped manner and tiled in the node array 204 so that the output clock lines 236 and 238 are aligned and electrically connected with the input clock lines 222 and 224 of two adjacent nodes 206. Using these electrical connections, the node array 204 can be implemented without using channels or top-level wiring to distribute clocks.

[0051] Return to Figure 2B , the input lines 222 and 224 may receive input clock signals from two of the neighboring nodes 206. For example, the first input clock line 222 receives the input clock signal from the neighboring node 206 above the current node 206, and the second input clock line 224 receives the input clock signal from the neighboring node 206 to the left of the current node 206. The first and second input clock lines 222 and 224 provide clock signals to the first and second inverters 226 and 228. The first inverter 226 inverts the clock signal and provides the inverted clock signal to the clock tap point 234, which is then provided to the main circuit system (e.g., in some embodiments, a computing circuit or a global circuit) of the corresponding node of the computing node array 102.

[0052] The second inverter 228 inverts the clock signal and provides the inverted clock signal to the third and fourth inverters 230 and 232. Each of the third and fourth inverters 230 and 232 inverts the inverted clock signal and outputs the resulting clock signal to the first and second output clock lines 236 and 238. The first and second output clock lines 236 and 238 output the clock signal to the adjacent node 206 below and to the right of the current node 206.

[0053] Return to reference Figure 2A , a clock signal received at a root node 206 (e.g., node 206A) propagates to its two neighboring nodes to the right and below (e.g., nodes 206B, 206C) with a unit delay. The unit delay may be a fixed offset for the entire node array 204. In some implementations, the unit delay may correspond to a combination of a line delay associated with propagating the clock signal to the downstream neighboring node 206 and a delay of the buffered clock signal (e.g., via inverters 228-232). Figure 2A , one of the downstream neighbor nodes 206 is in the same row as the root node 206 and is located to the right of the root node 206, and another of the downstream neighbor nodes 216 is in the same column as the root node 206 and is located below the root node 206. In other words, the neighbor nodes 206 may be located to the south and east of the root node 206.

[0054] When the clock signal passes through Figure 2A When the clock signal reaches the entire node array 204, the clock signal will continue to propagate to the adjacent nodes 206 to the south and east with a unit delay. This clock signal propagation continues through the clock distribution network until the clock signal reaches the node 206 diagonally opposite the root node 206 in the node array 204 (e.g., node 206D).

[0055] As the clock signal propagates through the node array 204, a node 206 in the node array 204 may receive a clock signal with substantially the same delay from two other adjacent nodes 206. The recombined mesh topology may combine two clock signals received from two adjacent nodes 206 at a given node 206 of the node array 204. Figure 2B , the clock signals received via the first input clock line 222 and the second input clock line 224 may be combined and may be received at each of the first inverter 226 and the second inverter 228. In some embodiments, the clock signals are combined by directly connecting the first input clock line 222 and the second input clock line 224 together. Other implementations for providing a recombined mesh topology are also possible.

[0056] The clock distribution circuit system disclosed herein allows for a flexible array structure that supports a wide range of array designs. For example, the node array 204 can be substantially square, having the same number of rows and columns. Alternatively, the node array 204 can be substantially rectangular, having a different number of rows and columns. The clock distribution circuit system disclosed herein also provides for relatively simple array reconstruction relative to the clock, which can also allow for relatively late schedule design decisions regarding the shape of the node array. In contrast, due to the amount of clock design time involved, the array size and shape of other clock distribution networks are typically expensive decisions that should not be postponed. However, in some cases, such delayed decisions may result in overall chip design optimization and are therefore desirable.

[0057] Figure 2C Pictured Figure 2A Another example implementation of clock distribution circuitry within an example node 206 of the node array 204 is shown. Figure 2C The clock distribution circuit system includes a first input clock line 222, a second input clock line 224, a second inverter 228, a third inverter 230, a fourth inverter 232, a clock tap point 234, a first output clock line 236 and a second output clock line 238. For each node 206, Figure 2CThe clock distribution circuit system 2C in FIG. 2 is designed so that the output clock lines 236 and 238 of the node 206 are aligned with the input clock lines 222 and 224 of the adjacent node 206. The nodes 206 can be arranged in a stepped manner and tiled in the node array 204 so that the output clock lines 236 and 238 are aligned and electrically connected with the input clock lines 222 and 224 of two adjacent nodes 206. Using these electrical connections, the node array 204 can be implemented without using channels or top-level wiring to distribute clocks.

[0058] In the example of clock distribution circuitry, such as Figure 2C , the input lines 222 and 224 may receive input clock signals from two of the neighboring nodes 206. For example, the first input clock line 222 receives the input clock signal from the neighboring node 206 above the current node 206, and the second input clock line 224 receives the input clock signal from the neighboring node 206 to the left of the current node 206. The first and second input clock lines 222 and 224 provide clock signals to the second inverter 228. The second inverter 228 inverts the clock signal and provides the inverted clock signal to the third and fourth inverters 230 and 232. Each of the third and fourth inverters 230 and 232 inverts the inverted clock signal and outputs the resulting clock signal to the first and second output clock lines 236 and 238. The first and second output clock lines 236 and 238 output the clock signal to the neighboring nodes 206 to the right and below the current node 206.

[0059] Figure 3A is with an example node array such as Figure 2A 204). The example node array 204 has 18 rows and 18 columns. Since there are 18 rows and 18 columns, there can be 324 nodes. As another example, the node array 204 can include 360 ​​nodes arranged in rows and columns. For example, the node 206 of the node array 204 can have the same Figure 2B The node 206 of the node array 204 may also have a clock distribution circuit system corresponding to the clock distribution circuit system of the node array 204. Figure 2CA clock distribution circuit system corresponding to the clock distribution circuit system of the node array 204. This clock diagram illustrates the number of unit delays of the clock signal output of the node 206 of the node array 204. For example, the root node 206 has 1 unit delay. The two nodes 206 adjacent to the root node 206 have 2 unit delays. The nodes 206 on the diagonal from southwest to northeast can have the same unit delay. Using the clock distribution circuit system described in this article, the unit delay can be a fixed offset. The nodes 206 along these diagonals can receive clock signals with substantially the same timing delay. These diagonals can be referred to as phases or waves. The phase corresponds to the different arrival times of the clock signals in the nodes 206. Figure 3A The clock signal distribution corresponding to the figure can achieve a 35-phase average synchronous clock. The number of phases of the average synchronous clock signal of the node array with the clock distribution circuit system described herein can be the number of rows plus the number of columns minus 1.

[0060] In some embodiments, the clock distribution network 200 can be configured to generate a wave that passes through the node array 204 in a row or column direction, rather than a clock signal that passes through the node array 204 in a wave formed along a diagonal of the node array 204. For example, each node 206 can output a clock signal to the south or east, rather than outputting a clock signal to the south and the east. In this way, the clock signal can propagate in the form of a wave that travels to the south or the east. However, aspects of the present disclosure are not limited to a particular direction of travel of the clock signal, and the clock signal can propagate along other diagonals and / or to the north or the west.

[0061] When routing signals between nodes 206, it may be considered Figure 3A A signal routed to a north or west node from the originating node where the signal is generated can travel upstream and Figure 3A A signal routed from an originating node to a south or east node can travel downstream and at the node corresponding to Figure 3A A unit delay is obtained in the node array 204. In order to solve the unit delay problem and meet the setup and hold time specifications, the signal traveling upstream can be routed faster than the signal traveling downstream. Figure 3A The sequential clock distribution shown in may also be referred to as wave clock distribution.

[0062] Figure 3B is with Figure 3A The node clock level diagram corresponds to the node clock level topology 310. Figure 3BAs shown in FIG. 1 , the node(s) 206 included in groups 310, 320, 330, and 340 may have unit delays of 1, 2, 3, and 4, respectively. The nodes 206 included in the same group may have the same number of unit delays as the clock signal received at the root node.

[0063] Timing Modeling

[0064] Although the timing in the node array may be important for computational accuracy, performance, etc., it may be difficult to accurately simulate the timing in a high-replication node array. For example, the size of the network-on-chip (NoC) data bus may cause the netlist size of the node array to surge, and clock phase changes may cause inaccurate simulation of the timing distribution in the node array. In addition, available electronic design automation (EDA) software may have difficulty calculating the timing of the node array. In some cases, the EDA software may not be able to calculate the timing of the node array, especially when the node array is large and / or complex. Alternatively, the EDA software may need to spend a lot of time to calculate the timing of the node array.

[0065] To address at least a portion of the above technical challenges, one or more aspects of the present disclosure correspond to systems and methods for modeling the timing of clock signal distribution in a node array. According to these aspects, the timing distribution of the node array can be modeled by modeling the timing delay of the node array based on the timing delay analysis of the node block (e.g., a portion of the node array). As described above, each node 206 of the node array 204 can include an instance of substantially the same clock distribution circuit system, even if one or more of the other circuit systems in the node 206 are different from other nodes 206. For example, Figure 4 The computing nodes 406 and global nodes 408 shown in the figure can have the same clock distribution circuit system. Therefore, the timing distribution can be modeled by performing timing analysis on a block (e.g., a node). For example, since the signal travels in one direction (e.g., upstream or downstream), the time delay according to the signal direction can be determined based on the interface delay between the node and its adjacent (e.g., adjacent or neighboring) node. This method is advantageous because it can avoid analyzing the time delay of each node in the node array and calculating the analyzed time delay to model the timing distribution of the node array.

[0066] like Figure 4As shown in , in some embodiments, the node array may include a computing node 406, a global node 408, a SerDes component 410, a general purpose input / output (GPIO) / security processing 412, and the like. The computing node 406 may include a circuit system for performing a processing task. The global node 408 may not include a circuit system for performing a processing task. For example, the global node 408 may include a PVT sensor to monitor the operating conditions of the node array. In some implementations, both the computing node 406 and the global node 408 may include a communication interface to enable communication with adjacent nodes. In some implementations, the communication interface of the computing node 406 may be the same as the communication interface of the global node 408.

[0067] In some embodiments, the techniques disclosed herein can be used to perform static timing analysis on a large number of replicated arrays having multiple compute nodes 406 and multiple global nodes 408. In some embodiments, a virtual block (which can be similar or identical to a global block as described above) can have most of the internal circuitry removed, but can retain interfaces and / or communications on its edges so that it can interface directly with functional blocks (e.g., compute nodes) in the array. In some embodiments, the array can use grid clock distribution to distribute clock signals to functional blocks and virtual blocks.

[0068] In some implementations of node arrays, nodes in the node array may communicate only with adjacent nodes. In such implementations, there may be no "flyover" signal, bypass signal, or other cross-node signal. The route connecting adjacent nodes may be horizontal or vertical. Nodes may be connected to adjacent nodes by horizontal routes and vertical routes. The computing node 406 that is not on the edge of the node array may be engaged with four adjacent nodes of the node array, two nodes adjacent to the computing node 406 in the same row of the node array, and two nodes adjacent to the computing node 406 in the same column of the node array. The global node 408 may be engaged with four adjacent computer nodes 406 of the node array, two computing nodes 406 adjacent to the global node 408 in the same row of the node array, and two computing nodes 406 adjacent to the global node 408 in the same column of the node array.

[0069] As mentioned above, modeling the clock distribution for an entire array of nodes can be difficult, time consuming, or even impossible using available EDA tools. Therefore, a simplified approach that can accurately model clock timing is needed.

[0070] In some embodiments, a modeling technique may include creating different copies of global static timing models of functional blocks and virtual blocks. The static timing model of a functional block may include models of the functional block with different surroundings (e.g., completely surrounded by functional blocks, with virtual blocks on one side, functional blocks on the other side, etc.). In some embodiments, within a node array, only a limited number of functional blocks and / or virtual blocks are arranged around the functional block. The static timing model of the virtual block may take into account the surrounding functional blocks. In some embodiments, the virtual block may be surrounded by functional blocks.

[0071] Large node arrays with average synchronous clocks present technical challenges for static timing analysis with traditional timing tools. Wide two-dimensional buses can significantly increase the size of the netlist and, therefore, the simulation run time, even for hierarchical designs. The timing analysis of clock signals in a node array as described herein depends on the directionality of data propagation relative to the clock propagation direction. Since interfaces are only between adjacent nodes that are adjacent to each other in the node array, timing can be performed with a block group timing model of a node and adjacent nodes that have a communication interface with the node. Block group timing can involve a simulation involving five nodes, which can be a small subset of the node array. The block group timing approach can avoid the annotation of arrival times and the effort of correlating simulations to the actual design that may exist in other approaches. By creating a block group for each scenario in the node array, the entire node array can be accurately simulated based on a few models. The reference Figure 5 The timing of the node array in the present disclosure can take advantage of one or more of the following simplifications: one block fills most (e.g., 98%) of the array, there are interfaces only between adjacent neighbor nodes, clock phases are systematic and matched, and most Clos delay variations are common nodes.

[0072] Figure 5An example of selecting nodes in a node array for modeling is shown. In some embodiments, the timing of the adjacent nodes 406B around the computing node 406A, the global node 408, and the adjacent nodes 406C around the global node 408 can be modeled, and other nodes 416 can be sparse. In addition, in some embodiments, the channel aggregator 414A and the adjacent circuit system 414B of the node array can be modeled by modeling the corners, and the other nodes 416 in the node array can be sparse by removing the internal circuit system from the timing model. This method can significantly reduce the number of networks, logic gates, lines, parasitic capacitances, etc. to be modeled. For example, the number of logic gates to be modeled can be reduced by about 80%, about 90%, etc. For example, the reduction can depend on the number of nodes in the node array, the node type in the node array, etc. This method can significantly improve the modeling speed, and is particularly beneficial for replicating designs without global signals. For example, the modeling calculation time can be reduced from several days or even weeks to a few hours.

[0073] refer to Figure 5 , computing node 406A, adjacent nodes 406B around computing node 406A, global node 408, adjacent nodes 406C around global node, channel aggregator 414A, and adjacent circuit system 414B can be used to model the clock distribution timing in node array. Timing models of small subsets of node array can be used to simulate the clock timing in node array. Computing node 406A and adjacent computing node 406B can be simulated to create the timing model of computing node. The timing model of computing node can be used for each computing node of array. Global node 408 and adjacent computing node 406C can be simulated to create the timing model of global node. The timing model of global node can be used for each global node of array.

[0074] In some embodiments, a time delay of the clock signal between the computing node 406A and each adjacent node 406B can be created based on the determined delay. Each other computing node of the node array can use the same timing mode as the computing node 406A. For example, because each node of the node array includes an instance of the same clock distribution circuit system and interface circuit system and is adjacent to an instance of the same adjacent node, Figure 3A Each compute node in the cluster uses the same timing model.

[0075] A time delay of the clock signal between the global node 408 and each adjacent compute node 406C may be created based on the determined delay. Each other global node of the node array may use the same timing mode as the global node 408 .

[0076] The time delay of the clock signal between the channel aggregator 414A and the adjacent circuit 414B can be determined and used for each similar instance of such circuit system. A model can be determined and used for each similarly located channel aggregator. The channel aggregator model can simulate the device under test (DUT) block and the associated interface path.

[0077] The model of the node array can cover functional blocks (e.g., compute nodes) and sparsification blocks. The model of the node array can cover global communication node timing whose size is only a small fraction of the design size without any gray box model.

[0078] A method of generating a circuit design model having a replicated instance (e.g., a node array) of a circuit block may include parsing a hardware description language (e.g., Verilog) model of the circuit design and generating a model of the circuit block (e.g., a functional block). The method may also include removing redundant similar instances of the circuit block to reduce the model size. The same or similar method may be performed for all other types of block instances (e.g., global blocks) of the design. The model may then include only unique scenarios. Static timing analysis may be performed on the model.

[0079] Worst Case Timing

[0080] In some implementations of node arrays, a block group timing model based on a node block of a node array is modeled to include various timing delay scenarios. For example, the delay between any two node clock arrival points in the node array can be different. Block allocation can be variable. For example, the delay between two nodes in the node array (e.g., near the middle of the node array) can come from the delay between two nodes near the edge of the node array. This situation may occur because the capacitance in different areas of the node array may be different (e.g., higher in the middle) due to, for example, the manufacturing process.

[0081] In some embodiments, the timing can be simulated to determine the worst case of early arrival and late arrival. This information can be used to ensure that any node in the array can meet the arrival time specification and the hold time specification. For example, this method can be used to ensure that even if a specific node or node group in the array is used for modeling, the modeling results can be applied to any node in the node array.

[0082] In some embodiments, the grid clock distribution can be distributed in waves to travel through the node array. EDA tools can allow clock delays to be annotated for early and late arrivals of clock signals. In some embodiments, static timing analysis can be run using only the worst case. Clock arrival differences can be generated between different computing nodes and global nodes, and the worst case for any node to arrive through the node array can be confirmed. Static timing analysis can be run for the worst possible arrival combination of all nodes of the node array (e.g., all computing nodes and global nodes). By applying the worst case, the running time of the static timing analysis of the node array can be reduced. The worst case of setup time and hold time can be applied. At the same time, this can ensure that the static timing of the design converges for all communication traffic combinations carried out between any two adjacent nodes of the node array.

[0083] In contrast, in a typical approach, static timing analysis can be performed on all nodes in the array, which can require significant computing resources to complete.

[0084] Timing can be generated between any two nodes in the node array. Figure 6 The timing of determining a portion of a node array (e.g., a block group of the node array) to and from a node (e.g., compute node 406A) and neighboring nodes (e.g., adjacent compute node 406B) in both vertical and horizontal directions is illustrated. The worst early and late arrivals can be annotated for these neighboring nodes. This can ensure that setup and hold time specifications are met.

[0085] Worst case timing data can be generated by parsing circuit simulation (e.g., SPICE simulation) results of an array of nodes throughout the grid. The fastest and slowest arrival times through the buffer can be selected. The delays of the buffer can be annotated in the timing model. This timing model can be optimized.

[0086] The worst-case timing data can be used with the node array model described in the previous section. Since only unique scenarios are used in the model, the worst-case scenario for setup time and the worst-case scenario for hold time can be used with each unique scenario in the model. Together, the model and the worst-case timing data can be used to effectively simulate the static timing of the node array and ensure that the setup and hold time specifications are met.

[0087] Worst-case timing data can be used to modify the design of the clock distribution circuit system. If the worst-case timing exceeds the timing specifications, the clock distribution circuit system can be adjusted until the timing specifications are met. Adjusting the clock distribution circuit system can involve adjusting the size of one or more clock drivers (e.g., increasing or decreasing the driver size depending on whether the setup time or hold time specifications are met) and / or adjusting the width of one or more lines carrying the clock signal (e.g., widening or narrowing the line regardless of whether the setup time or hold time specifications are met). This adjustment of the clock distribution circuit system can be applied to each node of the array. Design automation tools can automate the process of adjusting the clock distribution circuit system until the worst-case timing meets the timing specifications. In some other applications, circuit designers can use worst-case timing data to update the clock distribution circuit system.

[0088] Example embodiment of a node array timing allocation model

[0089] In some embodiments, the node array timing distribution model can be used to simulate the clock timing in the node array. For example, the simulation results of the node array timing distribution model can be used to determine the worst case of the clock timing in the node array. These simulation results can be used to set the clock distribution network, including interfaces and / or inverters and / or lines that carry signals between nodes of the node array.

[0090] Figure 7 An example of a computing device 710 that can simulate clock timing distribution across an array of nodes is shown. Figure 7 As shown in FIG. 7 , computing device 710 implements a timing distribution simulation component 720 , a non-volatile storage device 714 , and a main processor 712 .

[0091] In some examples, the main processor 712 can provide dedicated computing resources for use by the timing distribution simulation component 720. In addition, the main processor 712 can utilize the designated computing resources to process data generated from the timing distribution simulation component 720 according to examples disclosed herein.

[0092] like Figure 7 As shown in , the timing distribution simulation component 720 may include a timing model generator 722 and a timing model simulator 724. The timing model generator 722 may be configured to model the time delay of the node array by analyzing the timing delay on the node block (e.g., a portion of the node array). Figure 4 and Figure 5, the timing model generator 722 can utilize a modeling technique that includes creating different global static timing model copies of functional blocks and virtual blocks. The static timing model of the functional block can include models of functional blocks with different surrounding environments (e.g., completely surrounded by functional blocks, with virtual blocks on one side, with functional blocks on the other side, etc.). In some embodiments, only a limited number of functional blocks and / or virtual blocks may be arranged around the functional blocks in the node array. The static timing model of the virtual block can take into account the surrounding functional blocks. In some embodiments, the virtual block can be surrounded by the functional block. The timing analysis of the clock signal in the node array described herein depends on the directionality of data propagation relative to the clock propagation direction. Since the interface is only between adjacent nodes adjacent to each other in the node array, the timing can be performed with a block group timing model of the node and the adjacent node with a communication interface to the node. The block group timing can involve a simulation involving 5 nodes (e.g., an intermediate node with adjacent adjacent nodes), which can be a small subset of the node array. The block group timing method can avoid the annotation of arrival time and the effort of associating the simulation with the actual design that may exist in other methods. By creating a block group for each scene in the node array, the entire node array can be accurately simulated based on several models.

[0093] In some embodiments, the timing model generator 722 can model the computational nodes, the adjacent nodes around the computational nodes (e.g., 4 adjacent adjacent nodes), the global nodes, and the adjacent nodes around the global nodes, while sparsely ...

[0094] The timing model simulator 724 can be configured to simulate the model stored in the non-volatile storage device 714. In some embodiments, the timing model simulator 724 can access the model by accessing the non-volatile storage device 714. In some embodiments, a model including a computing node and an adjacent computing node can be simulated to create a timing model of the computing node. The timing model of the computing node can also be used for each computing node of the array. The global node and the adjacent computing nodes can also be simulated to create a timing model of the global node. The timing model of the global node can be used for each global node of the array.

[0095] In some embodiments, the timing model simulator 724 creates a time delay for the clock signal between the computing node and each adjacent node based on the determined delay. Each other computing node of the node array can use the same timing model as the computing node. For example, because each node of the node array includes an instance of the same clock distribution circuit system and interface circuit system and is adjacent to an instance of the same adjacent node, Figure 3A Each computing node in the node array uses the same timing model. A time delay of a clock signal between the global node and each computing node in the adjacent computing nodes can also be created based on the determined delay. Each other global node of the node array can use the same timing model as the global node.

[0096] In some embodiments, the timing model simulator 724 can simulate node arrays by modeling node arrays covering functional blocks (e.g., computing nodes) and sparse blocks. The method of generating a simulation model of a circuit design with a circuit block replication instance (e.g., a node array) may include parsing a hardware description language (e.g., Verilog) model of the circuit design and generating a model of the circuit block (e.g., a functional block). The method may also include removing redundant similar instances of the circuit block to reduce the size of the model. The same or similar method can be performed on all other types of block instances (e.g., global blocks) of the design. Then, the model can only include unique scenarios.

[0097] In some embodiments, the timing model simulator 724 performs timing analysis to determine worst-case timing data and includes the determined worst-case timing data in the simulation. For example, timing can be generated between any two nodes of the node array (e.g., Figure 6 The diagram illustrates determining timing to and from a portion of a node array (e.g., a block group of a node array) to and from a node (e.g., compute node 406A) and neighboring nodes (e.g., neighboring compute node 406B) in both vertical and horizontal directions. The worst early and late arrivals can be annotated for these neighboring nodes. This can ensure that setup and hold time specifications are met.

[0098] Worst case timing data can be generated by parsing circuit simulation (e.g., SPICE simulation) results of an array of nodes throughout the grid. The fastest and slowest arrival times through the buffer can be selected. The delays of the buffer can be annotated in the timing model. This timing model can be optimized.

[0099] The worst case timing data can be used with the node array model generated by the timing model generator 722. Since only unique scenarios are used in the model, the worst case of setup time and the worst case of hold time can be used with each unique scenario in the model. The model and the worst case timing data together can be used to effectively simulate the static timing of the node array and ensure that the setup and hold time specifications are met.

[0100] In some embodiments, worst-case timing data can be used to modify the design of the clock distribution circuit system. If the worst-case timing exceeds the timing specification, the clock distribution circuit system can be adjusted until the timing specification is met. Adjusting the clock distribution circuit system may include adjusting the size of one or more clock drivers (e.g., increasing or decreasing the driver size depending on whether the setup time or hold time specification is met) and / or adjusting the width of one or more lines carrying the clock signal (e.g., widening or narrowing the line regardless of whether the setup time or hold time specification is met). This adjustment of the clock distribution circuit system can be applied to each node of the array. Design automation tools can automate the process of adjusting the clock distribution circuit system until the worst-case timing meets the timing specification. In some other applications, circuit designers can use worst-case timing data to update the clock distribution circuit system.

[0101] To simplify the discussion and not to limit the present disclosure, Figure 7 Only the timing distribution analog component 720, non-volatile storage device, and host processor are illustrated, although multiple subcomponents or systems may be used.

[0102] Applications, terminology, and conclusions

[0103] The node arrays disclosed herein can be implemented in various processing systems. Such processing systems can be used and / or specifically configured for high performance computing and / or computationally intensive applications, such as neural network training, neural network reasoning, machine learning, artificial intelligence, complex simulations, and the like. In some applications, the processing system can be used to perform neural network training. For example, such neural network training can generate data for an autonomous driving system of a vehicle (e.g., a car), other autonomous vehicle functionality, or advanced driver assistance system (ADAS) functionality.

[0104] Unless the context clearly requires otherwise, throughout the specification and claims, the words "including", "comprising", etc. should be interpreted in an inclusive sense, rather than an exclusive or exhaustive sense; that is, interpreted in the sense of "including but not limited to". The word "coupled" as generally used in this article refers to two or more elements that can be directly connected or connected through one or more intermediate elements. Similarly, the word "connected" as generally used in this article refers to two or more elements that can be directly connected or connected through one or more intermediate elements. In addition, when the words "herein", "above", "below" and words with similar meanings are used in this application, they should refer to the entire application, not to any specific part of the application. Where the context permits, the words used in the singular or plural in the above "specific embodiments" may also include the plural or singular, respectively. The word "or" refers to a list of two or more items, which covers all the following interpretations of the word: any item in the list, all items in the list, and any combination of items in the list.

[0105] Furthermore, unless expressly stated otherwise or understood otherwise in context, conditional language used herein, such as "can," "may," "might," "for example," "such as," and the like, is generally intended to convey that certain embodiments include certain features, elements, and / or states, while other embodiments do not. Thus, such conditional language is generally not intended to imply that features, elements, and / or states are in any way essential to one or more embodiments.

[0106] The foregoing description has been described with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the invention to the precise form described. In view of the above teachings, many modifications and variations are possible. This will enable others skilled in the art to best utilize these techniques and various embodiments and make various modifications for various uses.

[0107] Although the present disclosure and examples have been described with reference to the accompanying drawings, various changes and modifications will be apparent to those skilled in the art. It should be understood that these changes and modifications are included within the scope of the present disclosure.

Claims

1. A method for simulating a node array, the method comprising: accessing a timing model of a compute node of the node array stored in a non-transitory computer readable memory, wherein the timing model of the compute node represents timing data associated with clock signal propagation between the compute node and four neighboring nodes of the node array, the four neighboring nodes each being adjacent to the compute node; as well as Using one or more computers, the timing models of the computing nodes are used to simulate clock signal timing for a majority of the nodes in the node array. 2 . The method of claim 1 , further comprising determining a worst-case timing for the clock signal in the array of nodes based on the simulation. 3 . The method of claim 2 , further comprising adjusting a clock distribution network of the node array based on the worst-case timing. 4 . The method of claim 3 , wherein adjusting the clock distribution network comprises updating one or more files representing circuitry of the array of nodes.

5. The method of claim 2 further comprising accessing a timing model of a global node of the node array stored in the non-transitory computer readable memory, wherein the determining is based on a simulation of the clock signal timing, the simulation of the clock signal timing using the timing model of the global node. The method of claim 1 , wherein the simulating comprises simulating an average synchronized clock in the array of nodes.

7. The method of claim 1 , wherein the timing model of the node array models a computing node that receives the clock signal from a first pair of adjacent nodes among the four adjacent nodes and a computing node that provides the clock signal to a second pair of adjacent nodes among the four adjacent nodes, wherein the clock signal is delayed by one unit delay in the computing node relative to that in the first pair of adjacent nodes, and wherein the clock signal is delayed by two unit delays in the second pair of adjacent nodes relative to that in the first pair of adjacent nodes. The method of claim 1 , wherein the timing model of the node array is a block group timing model.

9. The method of claim 1, wherein the majority of nodes of the node array comprises at least 90% of the nodes in the node array.

10. The method of claim 9, wherein the node array consists essentially of instances of the compute nodes and instances of global nodes.

11. The method of claim 1, further comprising generating the timing model of the computing node by at least simulating clock signal propagation between the computing node and the four neighboring nodes.

12. A non-transitory computer readable memory comprising instructions that, when executed by one or more processors, cause execution of a method of simulating a node array, wherein the method comprises: accessing a timing model of a compute node of the node array stored in a non-transitory computer readable memory, wherein the timing model of the compute node represents timing data associated with clock signal propagation between the compute node and four neighboring nodes of the node array, the four neighboring nodes each being adjacent to the compute node; as well as Using one or more computers, the timing models of the computing nodes are used to simulate clock signal timing for a majority of the nodes in the node array.

13. A computer system for simulating a node array, the computer system comprising: a non-transitory computer readable memory storing a timing model of a compute node of the node, wherein the timing model of the compute node represents timing data associated with clock signal propagation between the compute node and four neighboring nodes of the node array, the four neighboring nodes each being adjacent to the compute node; as well as One or more processors configured to execute instructions to at least access the timing model of the computing node and simulate clock signal timing for a majority of nodes in the node array using the timing model of the computing node.

14. A system for simulating clock timing distribution on a node array including a plurality of computing nodes, the system comprising: One or more computing devices are configured to store a timing model, the timing model corresponding to a computing node and four adjacent computing nodes, the four adjacent computing nodes are each adjacent to the computing node, wherein each of the one or more computing devices is configured to: Accessing the timing model of the computing node; as well as Using a computing device, clock signal timing distribution for a majority of nodes in the node array is simulated using the timing model of the computing node.

15. A non-transitory computer readable storage medium storing instructions to simulate clock timing distribution on a node array, the instructions when executed by a processor causing the processor to perform operations comprising: accessing a timing model of a compute node, wherein the timing model of the compute node represents timing data associated with clock signal propagation between the compute node and four neighboring nodes of the node array, the four neighboring nodes each being adjacent to the compute node; as well as Using a computing device, clock signal timing distribution for a majority of nodes in the node array is simulated using the timing model of the computing node. 16 . The non-transitory computer readable storage medium of claim 15 , further comprising determining worst case timing for the clock signals in the array of nodes based on the simulation.

17. The non-transitory computer readable storage medium of claim 16, further comprising adjusting a clock distribution network of the node array based on the worst case timing. 18 . The non-transitory computer-readable storage medium of claim 17 , wherein adjusting the clock distribution network comprises updating one or more files representing circuitry of the array of nodes.

19. The non-transitory computer-readable storage medium of claim 16, further comprising accessing a timing model of a global node of the node array stored in the non-transitory computer-readable memory, wherein the determining is based on a simulation of the clock signal timing, the simulation of the clock signal timing using the timing model of the global node.

20. The non-transitory computer readable storage medium of claim 15, wherein the simulating comprises simulating an average synchronized clock in the array of nodes.

21. The non-transitory computer-readable storage medium of claim 15, wherein the timing model of the node array models a computing node that receives the clock signal from a first pair of adjacent nodes among the four adjacent nodes and a computing node that provides the clock signal to a second pair of adjacent nodes among the four adjacent nodes, wherein the clock signal is delayed by one unit delay in the computing node relative to the first pair of adjacent nodes, and wherein the clock signal is delayed by two unit delays in the second pair of adjacent nodes relative to the first pair of adjacent nodes.

22. The non-transitory computer-readable storage medium of claim 15, wherein the timing model of the node array is a block group timing model.

23. The non-transitory computer-readable storage medium of claim 15, wherein the majority of nodes in the node array comprises at least 90% of the nodes in the node array.

24. The non-transitory computer-readable storage medium of claim 22, wherein the node array consists essentially of instances of the compute nodes and instances of global nodes.

25. The non-transitory computer-readable storage medium of claim 15, further comprising generating the timing model of the computing node by at least simulating clock signal propagation between the computing node and the four neighboring nodes.