Clock allocation with clock skew

By designing a clock distribution network in the node array of a high-density processing system, uniform clock delay between nodes is achieved, and the complex design and high resource consumption caused by clock signal synchronization and asynchronous communication in the prior art is solved, thereby achieving a simpler and more energy-saving clock distribution solution.

CN119998756APending Publication Date: 2025-05-13TESLA INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380071273.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-08-19
Filing Date
2023-08-16
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, when building a high-density processing system, the synchronization and asynchronous communication of clock signals lead to complex design, large area occupancy, high power consumption, and difficulty in achieving uniform clock delay between nodes.

Method used

An integrated circuit is designed to include an array of nodes with a clock distribution network, which receives and delays the clock signals through the clock distribution circuit system and provides them to the computing circuit system and adjacent nodes, ensuring that the nodes along the diagonal of the node array have substantially the same timing delay.

Benefits of technology

A uniform clock delay between nodes is achieved, chip design is simplified, area and power is saved, noise is reduced, electrical robustness of the chip and power dissipation is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119998756A_ABST
    Figure CN119998756A_ABST
Patent Text Reader

Abstract

Techniques for clock allocation with a fixed clock offset are disclosed. In one aspect, a clock distribution network for an array of nodes includes clock distribution circuitry of a plurality of nodes. At least one of the nodes is configured to receive a clock signal, provide the clock signal to the computing circuitry, and provide the clock signal to an adjacent node. There may be a delay unit between a clock signal at a node and an adjacent node. In some embodiments, a node may provide a clock signal to a first adjacent node in the same column and a second adjacent node in the same row, where the first and second adjacent nodes receive the clock signal at substantially the same delay.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of U.S. Provisional Patent Application No. 63 / 373,024, filed on August 19, 2022, the disclosure of which is hereby incorporated by reference in its entirety for all purposes. Technical Field

[0003] The present disclosure generally relates to clock distribution for electronic circuits and related systems and methods. Background Art

[0004] An array of processing nodes can be used to build a high-density processing system. A node can communicate with an adjacent node to perform a processing task. Communication between nodes can use synchronous and / or asynchronous methods. A clock signal can be provided to each node so that the nodes can be synchronized, thereby enabling communication between nodes. Summary of the invention

[0005] The innovations described in the claims each have several aspects, no single one of which is solely responsible for its desirable attributes. Without limiting the scope of the claims, some of the prominent features of the disclosure will now be briefly described.

[0006] In one aspect, an integrated circuit with a clock distribution network for a computing node array is provided, comprising: a node array including a plurality of nodes, the plurality of nodes including a first node and a second node adjacent to the first node, wherein the first node includes a clock distribution circuit system, the clock distribution circuit system being configured to: receive a clock signal, provide the clock signal to the computing circuit system of the first node, and provide the clock signal to the second node, wherein the clock signal is delayed by a delay unit in the second node relative to the first node.

[0007] In some embodiments, the first node is configured to receive a clock signal from two upstream nodes and provide the clock signal to two downstream nodes in a delay unit.

[0008] In certain embodiments, the nodes are arranged in rows and columns, and the node array is configured to propagate a clock signal through the node array such that nodes along a diagonal of the node array have substantially the same timing delay to the clock signal.

[0009] In some embodiments, the first node includes: a first input clock line configured to receive a clock signal from a first upstream node; a second input clock line configured to receive a clock signal from a second upstream node; a first output clock line configured to provide a clock signal to a first downstream node in a delay unit; and a second output clock line configured to provide a clock signal to a second downstream node in a delay unit.

[0010] In some embodiments, the first node also includes: a first inverter, coupled between the first input clock line and the computing circuit system, and the first inverter is also coupled between the second input clock line and the computing circuit system; a second inverter; and a third inverter, the second inverter and the third inverter are coupled between the first input clock line and the first output clock line, and the second inverter and the third inverter are also coupled between the second input clock line and the first output clock line.

[0011] In some embodiments, the first upstream node is located north of the first node, the second upstream node is located west of the first node, the first downstream node is located east of the first node, and the second downstream node is located south of the first node.

[0012] In some embodiments, the node array includes a plurality of compute nodes and a plurality of global nodes.

[0013] In some embodiments, the integrated circuit also includes: a clock management circuit, including: a clock generation circuit, configured to receive a system clock signal and generate a functional clock signal; a first multiplexer, configured to receive the functional clock signal and an alternative clock signal, and selectively output one of the functional clock signal and the alternative clock signal; and a second multiplexer, configured to receive a test clock signal and an output from the first multiplexer, and output one of the test clock signal and the output from the first multiplexer to a root node of the node array.

[0014] In some embodiments, the integrated circuit further includes: a multiplexer configured to receive the test clock signal and the functional clock signal from the clock generation circuit, and output one of the functional clock signal and the test clock signal as a clock signal to a root node of the node array.

[0015] In some embodiments, the node array has a bounded H-tree clock distribution topology.

[0016] In another aspect, a node array with mesochronous clock distribution is provided, comprising: a node array including a plurality of nodes arranged in rows and columns, wherein the node array includes a root node at a corner of the node array, wherein the root node is configured to receive a clock signal from outside the node array, provide the clock signal to a first adjacent node in the same column of the node array in a delay unit, and provide the clock signal to a second adjacent node in the same row of the node array in a delay unit, and wherein nodes along a diagonal of the node array receive clock signals with the same number of unit clock delays.

[0017] In certain embodiments, the root node includes computing circuitry, and the root node is further configured to provide a clock signal to the computing circuitry.

[0018] In some embodiments, the plurality of nodes includes a first node configured to receive clock signals from two upstream nodes and provide clock signals to two downstream nodes with a unit clock delay.

[0019] In some embodiments, the plurality of nodes include a first node, the first node including: a first input clock line configured to receive a clock signal from a first upstream node; a second input clock line configured to receive a clock signal from a second upstream node; a first output clock line configured to provide a clock signal to a first downstream node; and a second output clock line configured to provide a clock signal to a second downstream node.

[0020] In certain embodiments, the first node further includes a first inverter coupled between the first input line and the computing circuitry of the first node, the first inverter further coupled between the second input line and the computing circuitry.

[0021] In some embodiments, the first upstream node is located north of the first node, the second upstream node is located west of the first node, the first downstream node is located east of the first node, and the second downstream node is located south of the first node.

[0022] In some embodiments, the node array further includes: a multiplexer configured to receive the test clock and the functional clock signal from the clock generation circuit, and output one of the functional clock signal and the test clock as the clock signal to the root node.

[0023] In some embodiments, the node array has a bounded H-tree clock distribution topology.

[0024] On the other hand, a clock distribution method in a node array is provided, comprising: receiving a clock signal at a first node of the node array; providing the clock signal to a computing circuit system of the first node; and providing the clock signal to an adjacent node of the node array, wherein the adjacent node is adjacent to the first node, and wherein the clock signal has a unit delay in the adjacent node relative to the first node.

[0025] In some embodiments, the method further includes: at a first node, receiving a clock signal from two upstream nodes in a delay unit relative to the two upstream nodes, wherein one of the two upstream nodes is in the same row of the node array as the first node, and wherein the other of the two upstream nodes is in the same column of the node array as the first node; and providing a clock signal to two downstream nodes in a delay unit relative to the first node.

[0026] To summarize the present disclosure, certain aspects, advantages and novel features of the innovations are described herein. It should be understood that not all such advantages may be realized according to any particular embodiment. Therefore, the innovation may be embodied or performed in a manner that realizes or optimizes one advantage or group of advantages taught herein, without necessarily realizing other advantages taught or implied herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a schematic block diagram of an example chip according to aspects of the present disclosure.

[0028] Figure 2A is a schematic diagram of a clock distribution network according to an embodiment.

[0029] Figure 2B is a schematic diagram of a clock management unit (CMU) according to aspects of the present invention.

[0030] Figure 2C Picture shows Figure 2A An example implementation of clock distribution circuitry within an example node of an array of nodes.

[0031] Figure 2D Shows Figure 2A An alternative example implementation of a clock distribution circuit system within an example node of an array of nodes.

[0032] Figure 3 is with Figure 2A Node clock-level diagram associated with an example node array such as a node array of .

[0033] Figure 4A is a schematic diagram of a clock distribution network with a node array having a 2D distributed bounded H-tree clock distribution topology according to an embodiment of the present disclosure.

[0034] Figure 4B Shows Figure 2A An example implementation of clock distribution circuitry within an example node of an array of nodes.

[0035] Figure 4C Shows Figure 4A A node array is rearranged to show a bound H-tree topology of the node array. DETAILED DESCRIPTION

[0036] The following description of certain embodiments presents various descriptions of specific embodiments. However, the innovations described herein may be embodied in a variety of different ways, for example, as defined and covered by the claims. In this description, reference is made to the accompanying drawings, in which the same reference numerals may indicate the same or functionally similar elements. It will be understood that the elements illustrated in the figures are not necessarily drawn to scale. In addition, it will be understood that certain embodiments may include more elements than illustrated in the figures and / or a subset of the elements illustrated in the figures. In addition, some embodiments may combine any suitable combination of features from two or more figures.

[0037] The present disclosure provides a new method for distributing clock signals across chips, so that clock circuits can be modularly constructed by assembling the same sub-chip of the entire clock distribution circuit system. The clock distribution circuit system disclosed herein can save area, simplify design, and reduce power. Noise can be reduced relative to clock distribution for synchronous clock signals. The embodiments disclosed herein can also significantly reduce supply rail noise in certain frequency ranges, which can help improve chip electrical robustness and further reduce power dissipation.

[0038] Traditionally, clock signals are constructed and routed at the top level of the chip, which results in design effort, area, and power costs. In this case, clock distribution is a custom design at the top level of the chip. One way to do this is to route the clock signal in the channels between the sub-blocks. This can disrupt the design and consume area. Another way is to push the top-level clock down into the sub-blocks. This can slow down the design process and cause the same part of the design to be forked, where unique copies are created. Traditional methods can generate clock signals that arrive at all receivers almost simultaneously. The circuit can then operate in lock-step.

[0039] In the clock distribution network disclosed in this article, the clock arrives at each receiver at different times. The clock signal can be distributed through a two-dimensional (2D) array of nodes so that the clock signal arrives at different nodes with different timing offsets. Due to the clock distribution structure, the arrival time can be grouped into profiles or waves across the tube core. At the local level, the circuit system of the node can operate in lock step. More comprehensively, the circuit systems in different nodes of the node array can operate with timing offsets relative to each other. By making different nodes perform calculations with timing offsets relative to each other, the peak current from the power grid can be reduced. The quality of the power supply signal can also be improved by such calculations. The computing circuit system can be designed to process the arrival time difference of the clock signal.

[0040] The clock distribution network disclosed herein can simplify the top-level design of chips and clock circuit system structures. A clock with a fixed offset can be called an average synchronous clock. The embodiments disclosed herein relate to an average synchronous clock network built by modularizing common circuit systems. The clock signals of such a network can be locally low-skew and averagely synchronized at a coarser level.

[0041] The clock distribution disclosed herein can be applied to any suitable chip. In some applications, the clock distribution disclosed herein can be applied to chips that each include a smaller array of computing nodes. The computing node can be referred to as a processor or a core. In this way, the clock signal can form an arrival time wave across the array. Each computing node can receive a low skew clock signal. The computing nodes of the array can be designed to have only interfaces to adjacent computing nodes to consider the arrival time difference (skew) of the average synchronous clock phase. For example, the chip with a clock distribution network disclosed herein can have a 35-phase average synchronous clock or a 41-phase average synchronous clock. The clock distribution described herein can be used in a square (row and column equal) node array, or in a rectangular node array with a different number of rows and columns.

[0042] Figure 1 1 is a schematic block diagram of an example chip 100 according to aspects of the present disclosure. The chip 100 may be an integrated circuit die. The chip 100 may include a node array 102 (also referred to as a compute node array) with a distributed clock, one or more serializer / deserializer (SerDes) clock blocks 104, a clock generator 106, and a clock controller 108. The SerDes clock block 104 may be joined with other chips 100 to form the chip array 100. In some application examples, the node array 102 may be included on a chip 100 in a system on a wafer, a chip array 100 on a printed circuit board, etc. In some applications, Figure 1 The node array 102 can be implemented on a system on wafer packaged with a wafer level packaging structure. Figure 1 In the embodiment shown in FIG. 1 , the clock generator 106 can be implemented outside the node array 102. In some embodiments, the clock generator 106 can include a phase-locked loop (PLL). The clock generator 106 can be arranged to provide clock signals to the computing nodes at the corners of the node array 102. The clock controller 108 can also be implemented outside the node array 102. The nodes within the node array 102 can include node-to-node interfaces, which can be configured for synchronous communication. The core to serializer / deserializer (SerDes) interface can be asynchronous.

[0043] exist Figure 1In the node array 102 with distributed clocks, each node can be an instance of a computing circuit (also referred to as a processing core or computing node). In some applications, most of the nodes can be implemented as instances of computing circuits, and one or more of the nodes can be implemented as instances of different circuits. Each node of the node array 102 can include an instance of substantially the same clock distribution circuit system, even if the other circuit systems of at least some of the nodes are different from those of other nodes. In the node array 102, the nodes can be tiled and adjacent. For example, each node in the node array 102 can be independent and interconnected to (multiple) adjacent nodes. At the same time, the node array 102 can be implemented without using top-level wires, gates. Therefore, the node can be configured to communicate with adjacent nodes with lower-level wires through shorter connections. In some embodiments, the nodes of the node array 102 can be arranged in a stepped manner without mirroring or rotation. In some implementations, the nodes can be connected to the power line (V DD / V SS ). For example, the height and width of each node can be a multiple of the power grid pitch. The power grid pitch can be further aligned with the bump pitch.

[0044] Each node in the node array 102 may include an instance of substantially identical clock distribution circuitry. The nodes may be designed so that the output clock line of the node is aligned with the input clock line of its adjacent node. The nodes may be arranged in a stepped manner and tiled in the node array so that the clock output line is aligned and electrically connected with the clock input line of the adjacent node arranged downstream to receive the clock signal. With this electrical connection, a node array may be implemented without channels or top-level wiring for clock distribution. In some embodiments, the fan-out of the clock distribution circuitry may be balanced for use with inverters.

[0045] As described herein, a clock signal received at a root node may propagate from the root node to two neighboring nodes with a unit delay. The root node may be located at a corner of the node array 102. The unit delay may be a fixed offset for a given node array. The unit delay may correspond to the delay from buffering the clock signal (e.g., using an inverter) and the wire delay associated with propagating the clock signal to its neighboring node(s).

[0046] One of the two adjacent nodes may be located in the same row as the root node, and the other of the two adjacent nodes may be located in the same column as the root node. The adjacent nodes are adjacent to the root node. As an example, Figure 2AThe adjacent nodes in the node array 102 are located on the south and east sides of the root node. The clock signal continues to propagate from the two adjacent nodes of the root node in the node array in this example to the adjacent nodes on the south and east sides with one or more unit delays. This clock signal propagation continues through the clock distribution network in the node array 102 until the clock signal reaches the node that is diagonally opposite the root node in the node array 102. In this example, a signal that is routed from the originating node that generates the signal to an adjacent node located on the north or west side of the originating node can travel upstream and lose a unit delay in the node array 102, and a signal that is routed from the originating node to an adjacent node on the south or east side can travel downstream and gain a unit delay in the node array 102. Signals traveling upstream can be routed faster than signals traveling downstream to resolve unit delays and meet setup and hold time specifications.

[0047] Figure 2A is a schematic diagram of a clock distribution network 200 according to an embodiment. Clock distribution network 200 includes a clock management unit (CMU) 202, and clock distribution circuitry of a node array 204 (also referred to as a clock distribution node array) of nodes 206. Each node 206 includes an instance of clock distribution circuitry for clock signal distribution within node array 204. Figure 2A In the embodiment of the present invention, the clock distribution network 200 has a 2D distributed striped H-tree topology. The CMU 202 is configured to output a clock signal, which is received at a root node 206 of the node array 204.

[0048] Figure 2B2 is a schematic diagram of a CMU 202 according to various aspects of the present disclosure. The CMU 202 includes a PLL 212, a first multiplexer 214, and a second multiplexer 216. The PLL 212 is configured to receive a system clock signal sysclk and generate a functional clock signal funcclk. The first multiplexer 214 is configured to receive the functional clock signal funcclk at a first input and receive an alternative clock signal at a second input, and selectively output one of the functional clock signal funcclk and the alternative clock signal at an output of the first multiplexer 214. Depending on the embodiment, the alternative clock signal may include one or more of the following: a bypass clock signal, a reference clock signal generated on or outside the chip 100, a divided clock signal, or any other suitable clock signal. The second multiplexer 216 is configured to receive the clock output signal from the first multiplexer 214 at a first input and receive the test clock signal testclk at a second input, and selectively output one of the clock output signal from the first multiplexer 214 or the test clock signal testclk at an output of the second multiplexer 216. Therefore, the CMU 202 can be configured to selectively output one of the following to the root node of the node array 204: the functional clock signal funcclk, the test clock signal, or the alternative clock signal. The CMU 202 can provide a clock signal to the clock distribution network 200 for operating and / or testing the chip 100. For example, the CMU 202 can provide the test clock signal testclk to the clock distribution for testing the chip 100. As another example, the CMU 202 can provide the functional clock signal funcclk for typical operation of the chip 100.

[0049] refer to Figure 2A , the root may be located at the input of a node 206 at a corner of the node array 204. For example, the root may be located at Figure 2A 204. The root node 206 is located at the input of the node 206 at the northwest corner or the upper left corner of the node array 204 illustrated in FIG. In other embodiments, when the clock signal propagates in different directions along the rows and / or columns of the nodes, the root may be the input of another corner node 206 of the node array 204. The node 206 that receives the clock signal from outside the node array 204 may be referred to as the root node 206.

[0050] Refer to Figure 2A , the clock distribution network 200 can be implemented using a node array 204 . Figure 2A The node array 204 shown in FIG. Figure 1An example of a node array 102 with distributed clocks. In some embodiments, each node 206 can be an instance of a computing circuit. In some applications, most of the nodes 206 include instances of computing circuits, and one or more of the remaining nodes 206 (such as global nodes) include instances of different circuits. A global node can refer to a node 206 that does not include circuitry for performing processing tasks. In some implementations, both the computing nodes and the global nodes can include a communication interface to enable communication with neighboring nodes 206. In some implementations, the communication interface of the computing node can be the same as the communication interface for the global node.

[0051] In some embodiments, each node 206 in the node array 204 may include an instance of the same clock distribution circuit system, even if the other circuit systems of one or more nodes 206 are different from those of other nodes 206. In the node array 204, the nodes 206 may be tiled and adjacent. At the same time, the node array 204 may be implemented without any top-level wiring or gates. Therefore, the node 206 may communicate with the adjacent node 206 having the lower layer wiring through a short connection. The nodes 206 of the node array 204 may be arranged in a stepped manner without mirroring or rotation. The nodes 206 may also be aligned with the grid spacing of the power supply (VDD / VSS) line. For example, the height and width of each node 206 may be a multiple of the power grid spacing. In some embodiments, the power grid spacing may be further aligned with the bump spacing.

[0052] like Figure 2A As shown in , each node 206 may include substantially identical instances of clock distribution circuitry. Figure 2C The diagram shows Figure 2A An example implementation of clock distribution circuitry within an example node 206 of the node array 204 is shown. Figure 2A and Figure 2C The clock distribution circuit system includes a first input clock line 222, a second input clock line 224, a first inverter 226, a second inverter 228, a third inverter 230, a fourth inverter 232, a clock tap point 234, a first output clock line 236, and a second output clock line 238.

[0053] The clock distribution circuitry of each of the nodes 206 is designed so that the output clock lines 236 and 238 of the node 206 are aligned with the input clock lines 222 and 224 of the adjacent nodes 206. The nodes 206 can be arranged in a stepped manner and tiled in the node array 204 so that the output clock lines 236 and 238 are aligned and electrically connected with the input clock lines 222 and 224 of two adjacent nodes 206. Using these electrical connections, the node array 204 can be implemented without using channels or top-level wiring to distribute clocks.

[0054] Back to Figure 2C , the input lines 222 and 224 may receive input clock signals from two of the neighboring nodes 206. For example, the first input clock line 222 receives the input clock signal from the neighboring node 206 above the current node 206, and the second input clock line 224 receives the input clock signal from the neighboring node 206 to the left of the current node 206. The first and second input clock lines 222 and 224 provide clock signals to the first and second inverters 226 and 228. The first inverter 226 inverts the clock signal and provides the inverted clock signal to the clock tap point 234, which is then provided to the main circuit system (e.g., in some embodiments, a computing circuit or a global circuit) of the corresponding node of the computing node array 102.

[0055] The second inverter 228 inverts the clock signal and provides the inverted clock signal to the third and fourth inverters 230 and 232. Each of the third and fourth inverters 230 and 232 inverts the inverted clock signal and outputs the resulting clock signal to the first and second output clock lines 236 and 238. The first and second output clock lines 236 and 238 output the clock signal to the adjacent node 206 below and to the right of the current node 206.

[0056] Return to reference Figure 2A , the clock signal received at the root node 206 propagates to its two neighboring nodes to the right and below with a unit delay. The unit delay may be a fixed offset for the entire node array 204. In some implementations, the unit delay may correspond to a combination of the line delay associated with the clock signal propagating to the downstream neighboring node 206 and the delay of the buffered clock signal (e.g., via inverters 228-232). Figure 2A , one of the downstream neighboring nodes 206 is in the same row as the root node 206 and is located to the right of the root node 206, and another of the downstream neighboring nodes 216 is in the same column as the root node 206 and is located below the root node 206. In other words, the neighboring nodes 206 may be located to the south and east of the root node 206.

[0057] When the clock signal passes through Figure 2A When the clock signal reaches the entire node array 204 of the root node 206, the clock signal will continue to propagate to the adjacent node 206 on the south side with a unit delay. This clock signal propagation continues through the clock distribution network until the clock signal reaches the node 206 in the node array 204 that is diagonally opposite to the root node 206 (e.g., at the bottom right of the figure).

[0058] As the clock signal propagates through the node array 204, a node 206 in the node array 204 may receive a clock signal with substantially the same delay from two other adjacent nodes 206. The recombined mesh topology may combine two clock signals received from two adjacent nodes 206 at a given node 206 of the node array 204. Figure 2C , the clock signals received via the first input clock line 222 and the second input clock line 224 may be combined and may be received at each of the first inverter 226 and the second inverter 228. In some embodiments, the clock signals are combined by directly connecting the first input clock line 222 and the second input clock line 224 together. Other implementations for providing a recombined mesh topology are also possible.

[0059] The clock distribution circuitry disclosed herein allows for a flexible array structure that supports a wide range of array designs. For example, the node array 204 can be substantially square, having the same number of rows and columns. Alternatively, the node array 204 can be substantially rectangular, having a different number of rows and columns. The clock distribution circuitry disclosed herein also provides for relatively simple array reconfiguration relative to the clock, which can also allow for relatively late schedule design decisions regarding the shape of the node array. In contrast, due to the amount of clock design time involved, the array size and shape of other clock distribution networks are typically expensive decisions that should not be postponed. However, in some cases, such a delayed decision may result in overall chip design optimization and may therefore be desirable.

[0060] Figure 2D Picture shows Figure 2A An alternative example implementation of clock distribution circuitry within an example node 206 of the node array 204 of FIG. Figure 2D Node 206 is similar to Figure 2C Node 206 is shown, except that the outputs of the third and fourth inverters 230 and 232, respectively, are not coupled to each other. Thus, the third inverter 230 independently provides an output clock signal to the first output clock line 238, while the fourth inverter 232 independently provides an output clock signal to the second output clock line 236.

[0061] In summary, the clock distribution network 200 may be implemented such that each of the nodes 206 is configured to receive a clock signal from at least one neighboring node (or CMU 202, in the case of a root node 206), provide the clock signal to a corresponding node of the compute node array (e.g., via a clock tap 234), and provide the clock signal to an adjacent clock distribution node 206 when arranged adjacent to a downstream clock distribution node 206. For example, for a node 206 arranged adjacent to four neighboring nodes 206, the node 206 may receive a clock signal from two upstream clock distribution nodes and provide the clock signal to two downstream clock distribution nodes with a unit delay.

[0062] Figure 3 is with an example node array such as Figure 2A 204). The example node array 204 has 18 rows and 18 columns. Since there are 18 rows and 18 columns, there can be 324 nodes. As another example, the node array 204 can include 360 ​​nodes arranged in rows and columns. For example, the node 206 of the node array 204 can have the same Figure 2C or Figure 2D A clock distribution circuit system corresponding to the clock distribution circuit system of the node array 204. This clock diagram illustrates the number of unit delays of the clock signal output of the node 206 of the node array 204. For example, the root node 206 has 1 unit delay. The two nodes 206 adjacent to the root node 206 have 2 unit delays. The nodes 206 on the diagonal from southwest to northeast can have the same unit delay. Using the clock distribution circuit system described in this article, the unit delay can be a fixed offset. The nodes 206 along these diagonals can receive clock signals with substantially the same timing delay. These diagonals can be referred to as phases or waves. The phase corresponds to the different arrival times of the clock signals in the nodes 206. Figure 3 The clock signal distribution corresponding to the figure can achieve a 35-phase average synchronous clock. The number of phases of the average synchronous clock signal of the node array with the clock distribution circuit system described herein can be the number of rows plus the number of columns minus 1.

[0063] In some embodiments, the clock distribution network 200 can be configured to generate a wave that passes through the node array 204 in a row or column direction, rather than a clock signal that passes through the node array 204 in a wave formed along a diagonal of the node array 204. For example, each node 206 can output a clock signal to the south or east, rather than outputting a clock signal to the south and the east. In this way, the clock signal can propagate in the form of a wave that travels to the south or the east. However, aspects of the present disclosure are not limited to a particular direction of travel of the clock signal, and the clock signal can propagate along other diagonals and / or to the north or the west.

[0064] When routing signals between nodes 206, it may be considered Figure 3 A signal routed to a north or west node from the originating node where the signal is generated can travel upstream and Figure 3 A signal routed from an originating node to a south or east node can travel downstream and at the node corresponding to Figure 3 A unit delay is obtained in the node array 204. In order to solve the unit delay problem and meet the setup and hold time specifications, the signal traveling upstream can be routed faster than the signal traveling downstream.

[0065] Figure 4A 4 is a schematic diagram of a clock distribution network 400 having a node array 404 having a 2D distributed bounded H-tree clock distribution topology according to an embodiment of the present invention. Figure 4B Shows Figure 2A An example implementation of clock distribution circuitry within an example node 406 of node array 404 is shown. Figure 4C Shows Figure 4A 4. The clock distribution circuitry of the node array 404 is rearranged to illustrate the bound H-tree topology of the node array 404.

[0066] The clock distribution network 400 includes a CMU 402 and a node array 404. The CMU 402 includes a PLL 412 and a multiplexer 416. The PLL 412 is configured to receive a system clock signal and generate a functional clock signal. The multiplexer 416 is configured to receive a functional clock signal and a scan clock signal, and selectively provide one of the functional clock signal and the scan clock signal to a root node 406 of the node array 404. The node array 404 includes a plurality of nodes 406. Each of the nodes 406 includes a first input clock line 422, a second input clock line 424, a first inverter 426, a second inverter 428, a clock tap point 434, a first output clock line 438, and a second output clock line 436, as shown in FIG. Figure 4B shown.

[0067] Input lines 422 and 424 may receive input clock signals from two of the adjacent nodes 406. For example, the first input clock line 422 receives an input clock signal from an adjacent node 406 above the current node 406, and the second input clock line 424 receives an input clock signal from an adjacent node 406 to the left of the current node 406. For the case where the node 406 is a root node, the clock signal is received from the CMU 402. The first and second input clock lines 422 and 424 provide clock signals to first and second inverters 426 and 428. The first inverter 426 inverts the clock signal and provides the inverted clock signal to the clock tap point 434, which is then provided to the primary circuit of the node 406 (e.g., a computing circuit system or a global circuit in some embodiments).

[0068] The second inverter 428 inverts the clock signal and provides the inverted clock signal to the first and second output clock lines 436 and 438. The first and second output clock lines 436 and 438 output the clock signal to the adjacent nodes 406 to the right and below the current node 406.

[0069] like Figures 4A-4C As shown, each node 406 along the diagonal of the node array 404 can receive a clock signal with the same number of unit delays in the 2D distributed bound H-tree clock distribution network. For example, there are four nodes 406 of the node array 404 along the diagonal, which receive clock signals with 3 unit delays from the clock root. As another example, there are three nodes 406 along another diagonal of the node array 404, which receive clock signals with 4 unit delays from the root node 406. The nodes 406 along these diagonals can receive clock signals with the same number of unit delays from two adjacent nodes 406, and combine the two received clock signals 406.

[0070] The node arrays disclosed herein can be implemented in various processing systems. Such processing systems can be used and / or specifically configured for high performance computing and / or computationally intensive applications, such as neural network training, neural network reasoning, machine learning, artificial intelligence, complex simulations, and the like. In some applications, the processing system can be used to perform neural network training. For example, such neural network training can generate data for an autonomous driving system of a vehicle (e.g., a car), other autonomous vehicle functionality, or advanced driver assistance system (ADAS) functionality.

[0071] in conclusion

[0072] The foregoing disclosure is not intended to limit the disclosure to the precise form or specific field of use disclosed. Therefore, it is contemplated that various alternative embodiments and / or modifications of the disclosure are possible in light of the disclosure, whether explicitly described or implied herein. Having thus described the embodiments of the disclosure, it will be appreciated by those of ordinary skill in the art that changes may be made in form and detail without departing from the scope of the disclosure. Therefore, the disclosure is limited only by the claims.

[0073] In the foregoing description, the present disclosure has been described with reference to specific embodiments. However, as will be appreciated by those skilled in the art, the various embodiments disclosed herein may be modified or otherwise implemented in various other ways without departing from the spirit and scope of the present invention. Therefore, this description is considered to be illustrative and is intended to teach those skilled in the art to make and use the various embodiments of the disclosed vent assembly. It should be understood that the forms of the disclosure shown and described herein should be considered as representative embodiments. Equivalent elements, materials, processes or steps may replace those representatively shown and described herein. In addition, certain features of the present disclosure may be utilized independently of the use of other features, all of which are obvious to those skilled in the art after benefiting from this specification of the present disclosure. Expressions such as "including", "comprising", "incorporated", "composed of", "having", and "is" are used to describe and claim the present disclosure, and are intended to be interpreted in a non-exclusive manner, i.e., to allow for the presence of items, components or elements that are not explicitly described. References to the singular are also interpreted as involving the plural.

[0074] In addition, various embodiments disclosed herein should be understood in an illustrative and explanatory sense, and should never be interpreted as limiting the present invention.All connection references (e.g., attachment, attachment, coupling, connection, etc.) are only used to help readers understand the present disclosure, and may not produce restrictions, particularly about the position, orientation or use of the system and / or method disclosed herein.Therefore, connection references (if any) will be interpreted broadly.In addition, this connection reference does not necessarily mean that two elements are directly connected to each other.In addition, all numerical terms such as but not limited to "first", "second", "third", "primary", "minor", "primary" or any other common and / or numerical terms should also be regarded as identifiers only, to help readers understand the various elements, embodiments, variations and / or modifications of the present disclosure, and may not produce any restrictions, particularly about any element, embodiment, variation and / or modification relative to or exceeding the order or preference of another element, embodiment, variation and / or modification.

[0075] It should also be understood that one or more of the elements depicted in the figures may also be implemented in a more separate or integrated manner, or even removed or rendered inoperable in some cases, as may be useful depending on the particular application.

Claims

1. An integrated circuit having a clock distribution network for an array of computing nodes, comprising: a node array including a plurality of nodes, the plurality of nodes including a first node and a second node adjacent to the first node, Wherein the first node comprises a clock distribution circuit system, and the clock distribution circuit system is configured to: Receive clock signal, providing the clock signal to computing circuitry of the first node, and The clock signal is provided to the second node, wherein the clock signal is delayed by a delay unit in the second node relative to the clock signal in the first node.

2. The integrated circuit of claim 1 , wherein the first node is configured as: receiving the clock signal from two upstream nodes, and The clock signal is provided to two downstream nodes at the delay unit.

3. The integrated circuit of claim 1 , wherein: The nodes are arranged in rows and columns, and The node array is configured to propagate the clock signal through the node array such that nodes along a diagonal of the node array have substantially the same timing delay to the clock signal.

4. The integrated circuit of claim 1 , wherein the first node comprises: a first input clock line configured to receive the clock signal from a first upstream node; a second input clock line configured to receive the clock signal from a second upstream node; a first output clock line configured to provide the clock signal to a first downstream node in the delay unit; as well as The second output clock line is configured to provide the clock signal to a second downstream node with the delay unit.

5. The integrated circuit of claim 4, wherein the first node further comprises: a first inverter coupled between the first input clock line and the computing circuitry, the first inverter also coupled between the second input clock line and the computing circuitry; A second inverter; as well as A third inverter, the second inverter and the third inverter are coupled between the first input clock line and the first output clock line, and the second inverter and the third inverter are also coupled between the second input clock line and the first output clock line.

6. The integrated circuit of claim 4, wherein: The first upstream node is located on the north side of the first node, The second upstream node is located west of the first node, The first downstream node is located to the east of the first node, and The second downstream node is located on the south side of the first node.

7. The integrated circuit of claim 1, wherein the node array comprises a plurality of compute nodes and a plurality of global nodes.

8. The integrated circuit of claim 1 , further comprising: Clock management circuit, including: A clock generation circuit configured to receive a system clock signal and generate a functional clock signal; a first multiplexer configured to receive the functional clock signal and the alternative clock signal and selectively output one of the functional clock signal and the alternative clock signal; and A second multiplexer is configured to receive a test clock signal and the output from the first multiplexer, and output one of the test clock signal and the output from the first multiplexer to a root node of the node array.

9. The integrated circuit of claim 1 , further comprising: The multiplexer is configured to receive a test clock signal and a functional clock signal from a clock generation circuit, and output one of the functional clock signal and the test clock signal as the clock signal to a root node of the node array.

10. The integrated circuit of claim 1, wherein the node array has a tied H-tree clock distribution topology.

11. A node array with average synchronous clock distribution, comprising: a node array, comprising a plurality of nodes arranged in rows and columns, wherein the node array includes a root node at a corner of the node array, wherein the root node is configured to receive a clock signal from outside the node array, provide the clock signal to a first adjacent node in the same column of the node array in a delay unit, and provide the clock signal to a second adjacent node in the same row of the node array in a delay unit, and The nodes along a diagonal of the node array receive the clock signal with the same number of unit clock delays.

12. The node array of claim 11, wherein the root node includes computing circuitry, and the root node is further configured to provide the clock signal to the computing circuitry.

13. The node array of claim 11, wherein the plurality of nodes includes a first node, the first node being configured to: receiving the clock signal from two upstream nodes, and The clock signal is provided to two downstream nodes with a unit clock delay.

14. The node array of claim 11, wherein the plurality of nodes includes a first node, the first node comprising: a first input clock line configured to receive the clock signal from a first upstream node; a second input clock line configured to receive the clock signal from a second upstream node; a first output clock line configured to provide the clock signal to a first downstream node; as well as The second output clock line is configured to provide the clock signal to a second downstream node.

15. The node array of claim 14, wherein the first node further comprises a first inverter coupled between the first input line and the computational circuitry of the first node, the first inverter also coupled between the second input line and the computational circuitry.

16. The node array of claim 14, wherein: The first upstream node is located on the north side of the first node, The second upstream node is located west of the first node, The first downstream node is located to the east of the first node, and The second downstream node is located on the south side of the first node.

17. The node array according to claim 11, wherein the node array further comprises: The multiplexer is configured to receive a test clock and a functional clock signal from a clock generation circuit, and output one of the functional clock signal and the test clock to the root node as the clock signal.

18. The node array of claim 11, wherein the node array has a bounded H-tree clock distribution topology.

19. A clock distribution method in a node array, comprising: receiving a clock signal at a first node of the node array; providing the clock signal to computing circuitry of the first node; as well as The clock signal is provided to an adjacent node of the node array, wherein the adjacent node is adjacent to the first node, and wherein the clock signal has a unit delay in the adjacent node relative to the first node.

20. The method according to claim 19, further comprising: receiving, at the first node, the clock signal from two upstream nodes at the delay unit relative to the two upstream nodes, wherein one of the two upstream nodes is in the same row of the node array as the first node, and wherein the other of the two upstream nodes is in the same column of the node array as the first node; as well as The clock signal is provided to two downstream nodes with the delay unit relative to the first node.