Clock distribution with clock offset

A modular clock distribution network with timed offsets addresses inefficiencies in high-density processing systems, simplifying design and reducing power consumption while enhancing electrical robustness and flexibility in chip layout.

JP2025528226AActive Publication Date: 2025-08-26TESLA INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2025509088
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-19
Filing Date
2023-08-16
Publication Date
2025-08-26
Estimated Expiration
2043-08-16

AI Technical Summary

Technical Problem

Existing clock distribution methods in high-density processing systems result in inefficient design, increased area and power consumption, and lack of flexibility in chip layout due to custom top-level routing, leading to sub-optimal performance and power dissipation.

Method used

A modular clock distribution network with mesochronous clocking, where clock signals are distributed through a 2D array of nodes with timed offsets, allowing for local lockstep operation and reduced power consumption by aligning clock arrival times across the chip.

Benefits of technology

The proposed method simplifies chip design, reduces power dissipation, and enhances electrical robustness by minimizing supply rail noise, enabling flexible array configurations and late design decisions for optimal chip performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025528226000001_ABST
    Figure 2025528226000001_ABST
Patent Text Reader

Abstract

Techniques for clock distribution with a fixed clock offset are disclosed. In one aspect, a clock distribution network for a node array includes clock distribution circuitry in a plurality of nodes. At least one of the nodes is configured to receive a clock signal, provide the clock signal to computing circuitry, and provide the clock signal to neighboring nodes. A delay unit may exist between the clock signal at the node and the neighboring node. In a particular embodiment, a node may provide the clock signal to a first neighboring node in the same column and a second neighboring node in the same row, where the first and second neighboring nodes receive the clock signal with substantially the same delay.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 373,024, filed August 19, 2022, the disclosure of which is incorporated herein by reference in its entirety for all purposes.

[0002] FIELD OF THE DISCLOSURE The present disclosure relates generally to clock distribution for electronic circuits and related systems and methods. [Background technology]

[0003] A high-density processing system can be built using an array of processing nodes. The nodes can communicate with neighboring nodes to perform processing tasks. Communication between the nodes can be synchronous and / or asynchronous. A clock signal can be provided to each node to synchronize the nodes, which can enable communication between the nodes. Summary of the Invention [Means for solving the problem]

[0004] Each claimed innovation has several aspects, no single one of which is solely responsible for its desirable attributes. Without limiting the scope of the claims, some prominent features of this disclosure will now be discussed briefly.

[0005] In one aspect, an integrated circuit having a clock distribution network for a computational node array is provided, the integrated circuit comprising: a node array including a plurality of nodes, the plurality of nodes including a first node and a second node abutting the first node, the first node including clock distribution circuitry configured to receive a clock signal; provide the clock signal to computing circuitry of the first node; and provide the clock signal to the second node, where the clock signal is delayed at the second node relative to the first node by a delay unit.

[0006] In a particular embodiment, a first node is configured to receive clock signals from two upstream nodes and to provide clock signals with delay units to two downstream nodes.

[0007] In a particular embodiment, the nodes are arranged in rows and columns, and the node array is configured to propagate a clock signal through the node array such that nodes along a diagonal of the node array have substantially the same timing delay with respect to the clock signal.

[0008] In a particular embodiment, the first node includes a first input clock wiring configured to receive a clock signal from a first upstream node, a second input clock wiring configured to receive a clock signal from a second upstream node, a first output clock wiring configured to provide a clock signal with a delay unit to a first downstream node, and a second output clock wiring configured to provide a clock signal with a delay unit to a second downstream node.

[0009] In a particular embodiment, the first node further comprises a first inverter coupled between the first input clock wiring and the computing circuitry, the first inverter being also coupled between the second input clock wiring and the computing circuitry, a second inverter, and a third inverter, the second and third inventors being coupled between the first input clock wiring and the first output clock wiring, the second inverter and the third inverter also being coupled between the second input clock wiring and the first output clock wiring.

[0010] In certain embodiments, the first upstream node is located north of the first node, the second upstream node is located west of the first node, the first downstream node is located east of the first node, and the second downstream node is located south of the first node.

[0011] In a particular embodiment, the node array includes a plurality of computational nodes and a plurality of global nodes.

[0012] In a particular embodiment, the integrated circuit further comprises a clock management circuit including: a clock generation circuit configured to receive a system clock signal and generate a functional clock signal; a first multiplexer configured to receive the functional clock signal and an alternate clock signal and to selectively output one of the functional clock signal and the alternate clock signal; and a second multiplexer configured to receive an output from the first multiplexer and a test clock signal and to output one of the output from the first multiplexer and the test clock signal to a root node of the node array.

[0013] In a particular embodiment, the integrated circuit further includes a multiplexer configured to receive the functional clock signal and the test clock signal from the clock generation circuit and to output one of the functional clock signal and the test clock signal as a clock signal to a root node of the node array.

[0014] In a particular embodiment, the node array has a strapped H-tree clock distribution topology.

[0015] In another aspect, a node array having mesochronous clock distribution is provided, the node array including a plurality of nodes arranged in rows and columns, the node array including a root node at a corner of the node array, the root node configured to receive a clock signal from outside the node array and provide the clock signal with a delay unit to a first neighboring node in the same column of the node array and provide the clock signal with the delay unit to a second neighboring node in the same row of the node array, wherein nodes along a diagonal of the node array receive the clock signal with the same number of unit clock delays.

[0016] In particular embodiments, the root node includes computing circuitry, and the root node is further configured to provide a clock signal to the computing circuitry.

[0017] In a particular embodiment, the plurality of nodes includes a first node configured to receive clock signals from two upstream nodes and to provide clock signals with one unit clock delay to two downstream nodes.

[0018] In a particular embodiment, the plurality of nodes includes a first node including a first input clock wiring configured to receive a clock signal from a first upstream node, a second input clock wiring configured to receive a clock signal from a second upstream node, a first output clock wiring configured to provide a clock signal to a first downstream node, and a second output clock wiring configured to provide a clock signal to a second downstream node.

[0019] In a particular embodiment, the first node further comprises a first inverter coupled between the first input wiring and the computing circuitry of the first node, and the first inventor is also coupled between the second input wiring and the computing circuitry.

[0020] In certain embodiments, the first upstream node is located north of the first node, the second upstream node is located west of the first node, the first downstream node is located east of the first node, and the second downstream node is located south of the first node.

[0021] In a particular embodiment, the node array further includes a multiplexer configured to receive the functional clock signal and the test clock from the clock generation circuit and to output one of the functional clock signal and the test clock as a clock signal to the root node.

[0022] In a particular embodiment, the node array has a strapped H-tree clock distribution topology.

[0023] In yet another aspect, a method of clock distribution in a node array is provided, the method including receiving a clock signal at a first node of the node array, providing the clock signal to computing circuitry of the first node, and providing the clock signal to a neighboring node of the node array, the neighboring node abutting the first node, the clock signal having a delay unit at the neighboring node relative to that at the first node.

[0024] In a particular embodiment, the method further includes receiving, at the first node, a clock signal from two upstream nodes with a delay unit relative to the two upstream nodes, one of the two upstream nodes being in the same row of the node array as the first node and the other of the two upstream nodes being in the same column of the node array as the first node; and providing the clock signal to two downstream nodes with the delay unit relative to the first node.

[0025] For purposes of summarizing the disclosure, certain aspects, advantages, and novel features of the innovations have been described herein. It should be understood that not all such advantages may necessarily be achieved in accordance with any particular embodiment. Thus, the innovations may be embodied or implemented to achieve or optimize one advantage or group of advantages as taught herein without necessarily achieving other advantages as may be taught or suggested herein. [Brief explanation of the drawings]

[0026] [Figure 1] FIG. 1 is a schematic block diagram of an exemplary chip according to aspects of the present disclosure.

[0027] [Figure 2A] FIG. 1 is a schematic diagram of a clock distribution network according to one embodiment.

[0028] [Figure 2B] FIG. 1 is a schematic diagram of a clock management unit (CMU) according to an aspect of the present disclosure.

[0029] [Figure 2C] 2B illustrates an example implementation of clock distribution circuitry within an example node of the node array of FIG. 2A.

[0030] [Figure 2D] 2B illustrates an alternative exemplary implementation of clock distribution circuitry within an exemplary node of the node array of FIG. 2A.

[0031] [Figure 3] 2B is a node clock level map associated with an exemplary node array, such as the node array of FIG. 2A.

[0032] [Figure 4A]FIG. 1 is a schematic diagram of a clock distribution network having a node array with a 2D distributed, strapped H-tree clock distribution topology, according to one embodiment of the present disclosure.

[0033] [Figure 4B] 2B illustrates an example implementation of clock distribution circuitry within an example node of the node array of FIG. 2A.

[0034] [Figure 4C] FIG. 4B illustrates the node array of FIG. 4A rearranged to show the strapped H-tree topology of the node array. DETAILED DESCRIPTION OF THE INVENTION

[0035] The following description of some embodiments presents various descriptions of specific embodiments. However, the innovations described herein may be embodied in many different ways, for example, as defined and encompassed by the claims. This description refers to the drawings, in which like reference numbers may indicate identical or functionally similar elements. It will be understood that the elements depicted in the figures are not necessarily drawn to scale. Furthermore, it will be understood that some embodiments may include more elements than shown in the drawings and / or a subset of the elements depicted in the drawings. Furthermore, some embodiments may incorporate any suitable combination of features from two or more drawings.

[0036] The present disclosure provides a new method for distributing clock signals across a chip, allowing the clock circuitry to be modularly configured by assembling identical sub-pieces of the overall clock distribution circuitry. The clock distribution circuitry disclosed herein can save area, simplify design, and reduce power. Noise can be reduced compared to clock distribution of synchronous clock signals. The embodiments disclosed herein can also significantly reduce supply rail noise in certain frequency ranges, which can improve the electrical robustness of the chip and help further reduce power dissipation.

[0037] Traditionally, clock signals are built and routed at the top level of a chip, which adds effort, area, and power costs to the design. In such cases, clock distribution is custom designed at the top level of the chip. One way to do this is to route the clock signal with channels between sub-blocks, which can fragment the design and consume area. Another way is to push the top-level clock down to the sub-blocks, which can slow the design process and cause identical parts of the design to diverge, where unique copies are made. Traditional techniques can result in a clock signal that arrives at all receivers at approximately the same time. The circuit can then operate in lockstep.

[0038] In the clock distribution network disclosed herein, the clock arrives at various receivers at different times. The clock signal can be distributed through a two-dimensional (2D) array of nodes such that the clock signal arrives at different nodes with different timing offsets. Due to the clock distribution structure, arrival times can be grouped into a curve or wave across the die. At a local level, the circuitry of a node can operate in lockstep. More globally, the circuitry within different nodes of the node array can operate with timing offsets relative to each other. Peak currents from the power grid can be reduced by having different nodes perform computing with timing offsets relative to each other. Such computing can also improve the quality of the power signal. The computing circuitry can be designed to handle differences in the arrival times of the clock signal.

[0039] The clock distribution network disclosed herein can simplify the top-level design and configuration of clock circuitry on a chip. Clocking with a fixed offset can be referred to as mesochronous clocking. The embodiments disclosed herein enable mesochronous clock networks to be built modularly of instances of common subsection designs. Clock signals in such networks can be locally low-skew and mesochronous at coarser levels.

[0040] The clock distribution disclosed herein can be applied to any suitable chip. In particular applications, the clock distribution disclosed herein can be applied to a chip that includes an array of smaller computational nodes. The computational nodes may be referred to as processors or cores. In this manner, a clock signal can form an arrival time wave across the array. Each computational node can receive a low-skew clock signal. The computational nodes of the array can be designed with interfaces only to neighboring computational nodes, taking into account the arrival time difference (skew) of the mesochronous clock phases. A chip with the clock distribution network disclosed herein can have, for example, a 35-phase mesochronous clock or a 41-phase mesochronous clock. The clock distribution described herein can be used in a square (equal rows and columns) node array or in a rectangular node array with a different number of rows than columns.

[0041] FIG. 1 is a schematic block diagram of an exemplary chip 100 according to aspects of the present disclosure. The chip 100 may be an integrated circuit die. The chip 100 may include a node array 102 (also referred to as a compute node array) with distributed clocking, one or more serializer / deserializer (SerDes) clock blocks 104, a clock generator 106, and a clock controller 108. The SerDes clock block 104 may interface with other chips 100 to form an array of chips 100. In certain example applications, the node array 102 may be included in a chip 100 in a system-on-wafer system, an array of chips 100 on a printed circuit board, or the like. In certain applications, the node array 102 of FIG. 1 may be implemented in a system-on-wafer packaged in a wafer-level packaging structure. As shown in the embodiment of FIG. 1, the clock generator 106 may be implemented external to the node array 102. In some embodiments, the clock generator 106 may include a phase-locked loop (PLL). A clock generator 106 may be positioned to provide clock signals to computational nodes at the corners of the node array 102. A clock controller 108 may also be implemented outside the node array 102. The nodes within the node array 102 may include inter-node interfaces that may be configured to communicate synchronously. The core-to-serializer / deserializer (SerDes) interface may be asynchronous.

[0042] In the node array 102 with distributed clocking of FIG. 1, each node may be an instance of computing circuitry (also referred to as a processing core or compute node). In certain applications, most of the nodes may be implemented as instances of computing circuitry, and one or more of the nodes may be implemented as instances of different circuitry. Each node of the node array 102 may include an instance of substantially the same clock distribution circuitry, even if other circuitry of at least some of the nodes differs from that of the other nodes. In the node array 102, the nodes may be tiled and abutted. For example, each node of the node array 102 may be self-contained and interconnected to adjacent nodes. At the same time, the node array 102 may be implemented without using top-level wiring or gates. Thus, nodes may be configured to communicate with neighboring nodes using lower-level wiring via short connections. In some embodiments, the nodes of the node array 102 may be tiered without mirroring or rotation. In certain implementations, the nodes may be connected to a power supply line (V DD / V SS ) grid pitch. For example, the height and width of each node can be a multiple of the power grid pitch. The power grid pitch can be further aligned to the bump pitch.

[0043] Each node in the node array 102 may include substantially identical instances of clock distribution circuitry. Nodes may be designed so that the node's output clock wiring is aligned with its neighboring node's input clock wiring. Nodes may be tiered and tiled within the node array so that the clock output wiring aligns with and electrically connects to the clock input wiring of a neighboring node located downstream to receive the clock signal. Such electrical connections allow the node array to be implemented without channels or top-level wiring for clock distribution. In certain embodiments, the fanout of the clock distribution circuitry may be balanced relative to inverters.

[0044] As described herein, a clock signal received at a root node may propagate from the root node to two neighboring nodes with one delay unit. The root node may be located at a corner of node array 102. A delay unit may be a fixed offset for a given node array. A delay unit may correspond to the delay from buffering the clock signal (e.g., using an inverter) and the wiring delay associated with the clock signal propagating to that neighboring node.

[0045] One of the two neighboring nodes may be located in the same row as the root node, and the other of the two neighboring nodes may be located in the same column as the root node. The neighboring nodes abut the root node. As an example, the neighboring nodes are located south and east of the root node in FIG. 2A. In this example, the clock signal continues to propagate from the two neighboring nodes of the root node in the node array to the neighboring nodes south and east with one additional unit of delay. Such clock signal propagation continues through the clock distribution network in the node array 102 until the clock signal reaches a node in the node array 102 located at the opposite corner from the root node. In this example, a signal routed from the originating node to a neighboring node north or west of the originating node may travel upstream and lose one unit of delay in the node array 102, while a signal routed from the originating node to a neighboring node south or east may travel downstream and gain one unit of delay in the node array 102. Signals traveling upstream may be routed faster than signals traveling downstream to account for the unit delay and meet setup and hold time specifications.

[0046] FIG. 2A is a schematic diagram of a clock distribution network 200 according to one embodiment. The clock distribution network 200 includes a clock management unit (CMU) 202 and clock distribution circuitry in a node array 204 of nodes 206 (also referred to as a clock distribution node array). Each node 206 includes an instance of the clock distribution circuitry for clock distribution within the node array 204. In the embodiment of FIG. 2A, the clock distribution network 200 has a 2D distributed, strapped H-tree topology. The CMU 202 is configured to output a clock signal, which is received at a root node 206 of the node array 204.

[0047] 2B is a schematic diagram of a CMU 202 according to an aspect of the present disclosure. The CMU 202 includes a PLL 212, a first multiplexer 214, and a second multiplexer 216. The PLL 212 is configured to receive a system clock signal, sysclk, and generate a functional clock signal, funcclk. The first multiplexer 214 is configured to receive the functional clock signal, funcclk, at a first input and an alternate clock signal at a second input, and to selectively output one of the functional clock signal, funcclk, and the alternate clock signal at an output of the first multiplexer 214. Depending on the embodiment, the alternate clock signal may include one or more of a bypass clock signal, a reference clock signal generated on-chip or off-chip 100, a divided clock signal, or any other suitable clock signal. The second multiplexer 216 is configured to receive the clock output signal from the first multiplexer 214 at a first input and the test clock signal testclk at a second input, and to selectively output one of the clock output signal from the first multiplexer 214 or the test clock signal testclk at the output of the second multiplexer 216. Thus, the CMU 202 can be configured to selectively output one of the functional clock signal funcclk, the test clock signal, or the alternate clock signal to the root node of the node array 204. The CMU 202 can provide clock signals to the clock distribution network 200 for operating and / or testing the chip 100. For example, the CMU 202 can provide the test clock signal testclk for clock distribution to test the chip 100. As another example, the CMU 202 can provide the functional clock signal funcclk for typical operation of the chip 100.

[0048] 2A , the root may be located at the input to a node 206 at a corner of the node array 204. For example, the root may be located at the input to the node 206 at the northwest or upper left corner of the node array 204 shown in FIG. 2A . In other embodiments, the root may be the input to another corner node 206 of the node array 204 when the clock signal propagates in different directions along the rows and / or columns of the nodes. The node 206 that receives the clock signal from outside the node array 204 may be referred to as the root node 206.

[0049] Referring again to FIG. 2A , the clock distribution network 200 can be implemented using a node array 204. The node array 204 shown in FIG. 2A is an example of the node array 102 with distributed clocking of FIG. 1. In particular embodiments, each node 206 can be an instance of a computing circuit. In particular applications, a majority of the nodes 206 include instances of computing circuitry, and one or more of the remaining nodes 206 include instances of different circuitry, such as global nodes. A global node may refer to a node 206 that does not include circuitry for performing processing tasks. In some implementations, both the computational nodes and the global nodes can include communication interfaces that enable communication with neighboring nodes 206. In some implementations, the communication interface for the computational nodes may be the same as the communication interface for the global nodes.

[0050] In particular embodiments, each node 206 of a node array 204 may include an instance of the same clock distribution circuitry, even if one or more other circuitry of the node 206 differs from that of the other nodes 206. In a node array 204, the nodes 206 may be tiled and abutted. At the same time, the node array 204 may be implemented without top-level wiring or gates. Thus, the nodes 206 may communicate with neighboring nodes 206 using lower-level wiring via short connections. The nodes 206 of a node array 204 may be stepped without mirroring or rotation. The nodes 206 may also be aligned to the grid pitch of power (VDD / VSS) lines. For example, the height and width of each node 206 may be a multiple of the power grid pitch. In some embodiments, the power grid pitch may be further aligned to the bump pitch.

[0051] As shown in Figure 2A, each node 206 may include substantially the same instance of clock distribution circuitry. Figure 2C illustrates an example implementation of clock distribution circuitry within an example node 206 of the node array 204 of Figure 2A. With reference to Figures 2A and 2C, the clock distribution circuitry includes a first input clock wire 222, a second input clock wire 224, a first inverter 226, a second inverter 228, a third inverter 230, a fourth inverter 232, a clock tap point 234, a first output clock wire 236, and a second output clock wire 238.

[0052] The clock distribution circuitry of each of the nodes 206 is designed so that the output clock wires 236 and 238 of the node 206 are aligned with the input clock wires 222 and 224 of a neighboring node 206. The nodes 206 can be stepped and tiled within the node array 204 so that the output clock wires 236 and 238 align with and electrically connect to the two input clock wires 222 and 224 of a neighboring node 206. Using these electrical connections, the node array 204 can be implemented without using channels or top-level wiring for clock distribution.

[0053] 2C , input wires 222 and 224 may receive input clock signals from two of neighboring nodes 206. For example, first input clock wire 222 receives the input clock signal from the neighboring node 206 above the current node 206, while second input clock wire 224 receives the input clock signal from the neighboring node 206 to the left of the current node 206. First input clock wire 222 and second input clock wire 224 provide clock signals to first inverter 226 and second inverter 228. First inverter 226 inverts the clock signal and provides the inverted clock signal to clock tap point 234, which is then provided to the primary circuitry (e.g., computing circuitry or global circuitry in certain embodiments) of the corresponding node of computational node array 102.

[0054] The second inverter 228 inverts the clock signal and provides the inverted clock signal to a third inverter 230 and a fourth inverter 232. The third inverter 230 and the fourth inverter 232 each invert the inverted clock signal and output the resulting clock signals on a first output clock wire 236 and a second output clock wire 238. The first output clock wire 236 and the second output clock wire 238 output the clock signals to neighboring nodes 206 to the right and below the current node 206.

[0055] Referring back to FIG. 2A , a clock signal received at the root node 206 propagates from the root node 206 to its two neighboring nodes below and to the right with one delay unit. The delay unit may be a fixed offset throughout the node array 204. In some implementations, the delay unit may correspond to the delay from buffering the clock signal (e.g., via inverters 228-232) combined with the wiring delay associated with the clock signal propagating to the downstream neighboring node 206. In FIG. 2A , one of the downstream neighboring nodes 206 is to the right in the same row as the root node 206, and the other downstream neighboring node 206 is below in the same column as the root node 206. In other words, the neighboring nodes 206 may be located south and east of the root node 206.

[0056] The clock signal continues to propagate to neighboring node 206 to the south with one more unit of delay as the clock signal traverses the entire node array 204 of Figure 2A. Such clock signal propagation continues through the clock distribution network until the clock signal reaches node 206 in the node array 204 at the opposite corner from the root node 206 (e.g., at the bottom right of the figure).

[0057] As the clock signal propagates through the node array 204, a node 206 within the node array 204 may receive the clock signal from two other neighboring nodes 206 with substantially the same delay. The recombined mesh topology may combine two clock signals received from two neighboring nodes 206 at a given node 206 of the node array 204. For example, in FIG. 2C , the clock signals received via the first input clock wire 222 and the second input clock wire 224 may be combined and received at each of the first inverter 226 and the second inverter 228. In some embodiments, the clock signals are combined by connecting the first input clock wire 222 and the second input clock wire 224 directly together. Other implementations for providing a recombined mesh topology are also possible.

[0058] The clock distribution circuitry disclosed herein enables flexible array structures that support a wide range of array designs. For example, the node array 204 can be square, with a substantially equal number of rows and columns. Alternatively, the node array 204 can be rectangular, with a substantially different number of rows than columns. The clock distribution circuitry disclosed herein also provides a relatively simple reconfiguration of the array relative to the clock, which can also enable relatively late design decisions regarding node array shape. In contrast, array size and shape with other clock distribution networks are typically expensive decisions to postpone due to the amount of clock design time involved. However, in certain cases, such late decisions can result in overall chip design optimization and therefore may be desirable.

[0059] 2A. Node 206 of FIG. 2D is similar to node 206 shown in FIG. 2C, except that the outputs of third inverter 230 and fourth inverter 232, respectively, are not coupled to each other. Thus, third inverter 230 independently provides an output clock signal to first output clock wire 238, while fourth inverter 232 independently provides an output clock signal to second output clock wire 236.

[0060] In summary, clock distribution network 200 may be implemented such that each of nodes 206 is configured to receive a clock signal from at least one neighboring node (or CMU 202, in the case of root node 206), provide the clock signal to a corresponding node of the array of compute nodes (e.g., via clock tap point 234), and provide the clock signal to a neighboring clock distribution node 206 when located adjacent to the downstream clock distribution node 206. For example, for a node 206 located adjacent to four neighboring nodes 206, node 206 may receive a clock signal from two upstream clock distribution nodes and provide the clock signal with a unit delay to two downstream clock distribution nodes.

[0061] FIG. 3 is a node clock level map associated with an exemplary node array, such as the node array 204 of FIG. 2A. The exemplary node array 204 has 18 rows and 18 columns. With 18 rows and 18 columns, there can be 324 nodes. As another example, the node array 204 can include 360 ​​nodes arranged in rows and columns. The nodes 206 of the node array 204 can have clock distribution circuitry corresponding to the clock distribution circuitry of FIG. 2C or 2D, for example. This clock map shows the number of unit delays of the clock signal output for the nodes 206 of the node array 204. For example, the root node 206 has one unit delay. Two nodes 206 neighboring the root node 206 have two unit delays. The nodes 206 on the diagonal from southwest to northeast can have the same unit delay. Using the clock distribution circuitry described herein, the unit delays can be a fixed offset. The nodes 206 along these diagonals can receive clock signals with substantially the same timing delay. These diagonal lines may be referred to as phases or waves. The phases correspond to different clock signal arrival times at nodes 206. The clock signal distribution corresponding to the map of FIG. 3 may implement a 35-phase mesochronous clock. The number of phases of the mesochronous clock signal for a node array having the clock distribution circuitry described herein may be the number of rows + the number of columns minus one.

[0062] In particular embodiments, rather than a clock signal traversing node array 204 in a wave formed along the diagonal of node array 204, clock distribution network 200 may be configured to generate a wave that traverses node array 204 in a row or column direction. For example, rather than outputting a clock signal south and east, each node 206 may output a clock signal either south or east. In this manner, a clock signal may propagate in a wave traveling south or east. However, aspects of the present disclosure are not limited to a particular direction of travel for the clock signal; the clock signal may propagate along other diagonals and / or north or west.

[0063] The offsets in Figure 3 may be taken into account when routing signals between nodes 206. A signal routed to a node north or west of the originating node that generates the signal may travel upstream and lose one unit delay in the node array 204 corresponding to Figure 3. A signal routed to a node south or east of the originating node may travel downstream and gain one unit delay in the node array 204 corresponding to Figure 3. Signals traveling upstream may be routed faster than signals traveling downstream to account for the unit delay and meet setup and hold time specifications.

[0064] Figure 4A is a schematic diagram of a clock distribution network 400 having a node array 404 with a 2D distributed, strapped H-tree clock distribution topology, in accordance with one embodiment of the present disclosure. Figure 4B shows an example implementation of the clock distribution circuitry within an example node 406 of the node array 404 of Figure 2A. Figure 4C shows the clock distribution circuitry of the node array 404 of Figure 4A rearranged to illustrate the strapped H-tree topology of the node array 404.

[0065] The clock distribution network 400 includes a CMU 402 and a node array 404. The CMU 402 includes a PLL 412 and a multiplexer 416. The PLL 412 is configured to receive a system clock signal and generate a functional clock signal. The multiplexer 416 is configured to receive a functional clock signal and a scan clock signal and to selectively provide one of the functional clock signal and the scan clock signal to a root node 406 of the node array 404. The node array 404 includes a plurality of nodes 406. Each of the nodes 406 includes a first input clock wire 422, a second input clock wire 424, a first inverter 426, a second inverter 428, a clock tap point 434, a first output clock wire 438, and a second output clock wire 436, as shown in FIG. 4B .

[0066] The input wires 422 and 424 can receive input clock signals from two of the neighboring nodes 406. For example, the first input clock wire 422 receives the input clock signal from the neighboring node 406 above the current node 406, while the second input clock wire 424 receives the input clock signal from the neighboring node 406 to the left of the current node 406. If the node 406 is the root node, the clock signal is received from the CMU 402. The first input clock wire 422 and the second input clock wire 424 provide the clock signal to the first inverter 426 and the second inverter 428. The first inverter 426 inverts the clock signal and provides the inverted clock signal to the clock tap point 434, which is then provided to the primary circuitry of the node 406 (e.g., computing circuitry or global circuitry in certain embodiments).

[0067] The second inverter 428 inverts the clock signal and provides the inverted clock signal to a first output clock wire 436 and a second output clock wire 438. The first output clock wire 436 and the second output clock wire 438 output the clock signal to neighboring nodes 406 to the right and below the current node 406.

[0068] 4A-4C, each node 406 along a diagonal of a node array 404 can receive a clock signal with the same number of unit delays in a 2D distributed, strapped H-tree clock distribution network. For example, there are four nodes 406 of the node array 404 along a diagonal that receive a clock signal with three unit delays from the clock root. As another example, there are three nodes 406 along another diagonal of the node array 404 that receive a clock signal with four unit delays from the root node 406. These diagonal nodes 406 can receive clock signals with the same number of unit delays from two neighboring nodes 406 and combine the two received clock signals 406.

[0069] The node arrays disclosed herein can be implemented in a variety of processing systems. Such processing systems may be used in and / or specifically configured for high-performance computing and / or computationally intensive applications, such as neural network training, neural network inference, machine learning, artificial intelligence, complex simulations, etc. In some applications, the processing systems may be used to perform neural network training. For example, such neural network training may generate data for a vehicle's (e.g., automobile) autopilot system, other autonomous vehicle functions, or advanced driver assistance system (ADAS) functions. conclusion

[0070] The foregoing disclosure is not intended to limit the disclosure to the precise form or particular field of use disclosed. Accordingly, various alternative embodiments and / or modifications to the disclosure, whether expressly described or implied herein, are contemplated in light of the present disclosure. While embodiments of the present disclosure have been described in this manner, those skilled in the art will recognize that changes can be made in form and detail without departing from the scope of the present disclosure. Accordingly, the present disclosure is limited only by the claims.

[0071] In the foregoing specification, the present disclosure has been described with reference to specific embodiments. However, as those skilled in the art will understand, the various embodiments disclosed herein can be modified or otherwise implemented in various other ways without departing from the spirit and scope of the present disclosure. Accordingly, this description should be considered illustrative and is for the purpose of teaching those skilled in the art how to make and use various embodiments of the disclosed vent assembly. It should be understood that the forms of the disclosure shown and described herein should be construed as representative embodiments. Equivalent elements, materials, processes, or steps may be substituted for those typically shown and described herein. Furthermore, certain features of the present disclosure can be utilized independently of the use of other features, all of which will become apparent to those skilled in the art after having the benefit of this description of the present disclosure. The terms "including," "comprising," "incorporating," "consisting of," "having," "being," and the like, used to describe and claim the present disclosure, are intended to be construed in a non-exclusive manner, i.e., allowing for the presence of items, components, or elements not expressly described. References to the singular should also be construed to relate to the plural.

[0072] Furthermore, the various embodiments disclosed herein should be construed in an illustrative and explanatory sense and in no way as limiting the present disclosure. All coupling references (e.g., attached, secured, coupled, connected, etc.) are used solely to aid the reader's understanding of the present disclosure and do not create any limitations with respect to the position, orientation, or use of the systems and / or methods disclosed herein, among other things. Accordingly, coupling references, if present, should be interpreted broadly. Furthermore, such coupling references do not necessarily imply that two elements are directly connected to one another. Furthermore, without limitation, all numerical terms such as "first," "second," "third," "primary," "secondary," "main," or any other conventional and / or numerical terminology should also be construed solely as identifiers to aid the reader's understanding of the various elements, embodiments, variations, and / or modifications of the present disclosure, and in particular do not create any limitations with respect to the order or preference of any element, embodiment, variation, and / or modification relative to or over another element, embodiment, variation, and / or modification.

[0073] It will also be understood that one or more of the elements shown in the drawings / figures may also be implemented in a more separated or integrated manner, or may be removed or rendered inoperable in certain cases, as may be useful depending on the particular application.

Claims

1. 1. An integrated circuit having a clock distribution network for an array of compute nodes, comprising: a node array including a plurality of nodes, the plurality of nodes including a first node and a second node abutting the first node; The first node: receiving a clock signal; providing the clock signal to computing circuitry of the first node; and providing the clock signal to the second node, the clock signal being delayed by a delay unit at the second node relative to the first node.

2. The first node: receiving the clock signals from two upstream nodes; and providing the clock signal with the delay unit to two downstream nodes.

3. the nodes are arranged in rows and columns; 2. The integrated circuit of claim 1, wherein the node array is configured to propagate the clock signal through the node array such that nodes along a diagonal of the node array have substantially the same timing delay with respect to the clock signal.

4. The first node: a first input clock line configured to receive the clock signal from a first upstream node; a second input clock line configured to receive the clock signal from a second upstream node; a first output clock wiring configured to provide the clock signal with the delay unit to a first downstream node; a second output clock wiring configured to provide the clock signal with the delay unit to a second downstream node.

5. The first node: a first inverter coupled between the first input clock wiring and the computing circuitry, the first inventor also coupled between the second input clock wiring and the computing circuitry; a second inverter; 5. The integrated circuit of claim 4, further comprising: a third inverter, the second and third inventors coupled between the first input clock wiring and the first output clock wiring, the second inverter and the third inverter also coupled between the second input clock wiring and the first output clock wiring.

6. the first upstream node is located north of the first node; the second upstream node is located west of the first node; the first downstream node is east of the first node; 5. The integrated circuit of claim 4, wherein the second downstream node is located south of the first node.

7. The integrated circuit of claim 1 , wherein the node array comprises a plurality of computational nodes and a plurality of global nodes.

8. 1. A clock management circuit, comprising: a clock generation circuit configured to receive a system clock signal and generate a functional clock signal; a first multiplexer configured to receive the functional clock signal and the alternate clock signal and to selectively output one of the functional clock signal and the alternate clock signal; a second multiplexer configured to receive the output from the first multiplexer and a test clock signal and to output one of the output from the first multiplexer and the test clock signal to a root node of the node array.

9. 2. The integrated circuit of claim 1, further comprising a multiplexer configured to receive a functional clock signal and a test clock signal from a clock generation circuit and to output one of the functional clock signal and the test clock signal as the clock signal to a root node of the node array.

10. The integrated circuit of claim 1 , wherein the node array has a strapped H-tree clock distribution topology.

11. 1. A node array with mesochronous clock distribution, comprising: a node array including a plurality of nodes arranged in rows and columns; the node array includes a root node at a corner of the node array; the root node is configured to receive a clock signal from outside the node array, provide the clock signal with a delay unit to a first neighboring node in the same column of the node array, and provide the clock signal with the delay unit to a second neighboring node in the same row of the node array; A node array, wherein nodes along a diagonal of the node array receive the clock signal with the same number of unit clock delays.

12. The node array of claim 11 , wherein the root node includes computing circuitry, the root node further configured to provide the clock signal to the computing circuitry.

13. the plurality of nodes receiving the clock signals from two upstream nodes; 12. The node array of claim 11, comprising: a first node configured to: provide the clock signal with one unit clock delay to two downstream nodes.

14. the plurality of nodes a first input clock line configured to receive the clock signal from a first upstream node; a second input clock line configured to receive the clock signal from a second upstream node; a first output clock wiring configured to provide the clock signal to a first downstream node; a second output clock wiring configured to provide the clock signal to a second downstream node.

15. 15. The node array of claim 14, wherein the first node further comprises a first inverter coupled between the first input wiring and computing circuitry of the first node, the first inverter also coupled between the second input wiring and the computing circuitry.

16. the first upstream node is located north of the first node; the second upstream node is located west of the first node; the first downstream node is east of the first node; The node array of claim 14 , wherein the second downstream node is located south of the first node.

17. the node array:

12. The node array of claim 11, further comprising: a multiplexer configured to receive a functional clock signal and a test clock from a clock generation circuit and to output one of the functional clock signal and the test clock as the clock signal to the root node.

18. The node array of claim 11 , wherein the node array has a strapped H-tree clock distribution topology.

19. 1. A method of clock distribution in a node array, comprising: receiving a clock signal at a first node of the node array; providing the clock signal to computing circuitry of the first node; providing the clock signal to a neighboring node of the node array, the neighboring node abutting the first node, the clock signal having a delay unit at the neighboring node relative to that at the first node; A method comprising:

20. receiving, at the first node, the clock signal from two upstream nodes with the delay units for the two upstream nodes, one of the two upstream nodes being in the same row of the node array as the first node and the other of the two upstream nodes being in the same column of the node array as the first node; 20. The method of claim 19, further comprising: providing the clock signal to two downstream nodes with the delay unit relative to the first node.

Citation Information

Patent Citations

  • Method and apparatus, integrated circuit and node for providing timing signals to multiple circuits

    JP2008543188A

  • System and Method for Calibrating a Frequency Doubler

    US20210194605A1

  • Scan test device and scan test method

    US20210373074A1